← Files Plugin CreatorARCHIVED FILE

skills/prepare-plugin-submission/references/review-materials.md

9.3 KB · Oct 4, 2026 · 12:13 UTC

↓ Download file

# MCP review examples and evidence

For each app needing initial public review, draft five positive and three
negative cases from its actual tools, schemas, and intended user workflows.
Skills-only plugins need no MCP cases, demo, or reviewer credentials; existing
reviewed dependencies need an eligible version rather than a new review.

Before handoff, verify each MCP tool's annotations include boolean `readOnlyHint`,
`openWorldHint`, and `destructiveHint` values matching its behavior. Correct
missing or inaccurate annotations in the server source. Annotation justifications
are not required.

Guide the creator through a useful test set:

1. Identify the main tasks users should accomplish and the limits of the app.
   Propose coverage from the implementation, then ask the creator only about
   unclear intended behavior, realistic examples, and required test data.
   Draft the cases yourself; do not ask the creator to author them from scratch.
2. Make each positive case test a distinct supported behavior: a core workflow,
   meaningful input variation, valid boundary, or multi-step task where supported.
   Use natural user prompts instead of telling the model which tool to call.
   Do not pad the set with paraphrases or invent capabilities to fill the count.
3. Give each case a purpose, reproducible setup, and observable pass condition.
   Name exact expected tools and important argument values in the expectations.
   State what the user should receive and which returned data must support it;
   avoid vague outcomes such as "works correctly". For random, live, or generated
   results, assert valid ranges, required properties, and consistency with the
   tool response rather than an invented exact output.
4. Choose negative prompts that plausibly resemble the app's purpose but are
   outside its capabilities, so they test when its tools should not be invoked.
   State the expected explanation or clarification and prohibited tool use in
   the case description. Keep these distinct from supported workflows that
   return no results or encounter authentication/validation errors; those need
   explicit recovery expectations when tested. Never treat a fabricated success
   as acceptable error handling.
5. Use reviewer-accessible fixtures and include any necessary conversation
   context or setup in the description. Keep secrets in secure reviewer-access
   fields. Present the draft set with a short explanation of its coverage, invite
   corrections, and resolve missing setup or ambiguous pass conditions before
   calling the cases ready to run.
6. Run cases where access and authorization allow, comparing actual tool calls,
   arguments, and user-visible results with the expectations. Record **Passed**,
   **Failed**, **Blocked**, or **Not run** with evidence in preparation notes;
   do not add unsupported result fields to the manifest. Fix the app or the
   expectation based on intended behavior, not merely to make a failure pass.
   Preserve unverified status when execution is unavailable.

For example, "Search works" is too vague. For a supported catalog search,
"Find blue mugs under $20" can expect the actual search tool with the color and
price filters, and a response containing only matching returned items. Use a
known test catalog to make this reproducible. A request to purchase a found mug
is a useful negative case only if purchasing is unsupported: explain that limit
without attempting a purchase or claiming an order was placed. Adapt examples
to the plugin's real contract.

Declare cases in one of these locations; never combine them:

- **Exactly one MCP server:** root `plugin.json` under
  `extensions.com.openai.review.test_cases`, with `positive` and `negative`
  arrays. Each case has `description` and `prompt`; positive cases also have
  `tools_triggered` and `expected_behavior` as strings, not arrays. Optional
  evidence uses `file_attachment_urls` and `expected_output_url`.
- **Multiple servers:** each exact `mcp.json` server's
  `extensions.com.openai.review`, with `testCases` and `negativeTestCases`
  arrays. Each case has `description` and `userPrompt`; positive cases also
  have `toolsTriggered` and `expectedOutput`. Optional evidence uses
  `fileAttachmentUrls` and `expectedOutputUrl`.

Other package fields under `extensions.com.openai`:

- `review.demo_recording_url`: a reviewer-accessible recording of the submitted
  behavior. Verify the link.
- `review.commerce` and `review.commerce_description`: truthful declarations
  about commerce behavior, when applicable.
- `publication.release_notes`: what changed in this release.
- `publication.countries`: intended supported country codes. Omission preserves
  targeting; an empty list deliberately removes restrictions.
- `publication.translations`: localized listing descriptions and subtitles keyed
  by locale. Base listing fields supply English; do not synthesize an `en-US` entry.

In review information, countries and translations occupy the first page;
reviewer credentials follow on a separate page. When the dashboard supports
field-level ZIP ownership, ZIP-supplied countries and translations are prefilled
and locked, including explicit empty lists/maps. Update the ZIP to change them.
Fields omitted from the ZIP remain editable, including after saving and
reopening. Check the saved country targeting and translation text before review.
Older dashboard/backend versions may keep the entire publication section locked;
use the ZIP for those versions rather than assuming inline edits are available.

### Record the demo video

Every app entering initial public review needs a verified demo recording. Guide recording and hosting during preparation, before the final submission-portal upload, connection, and testing step. Help the creator produce it rather than stopping at a script or a request for a URL.

Use an existing development installation or preview in the intended host to demonstrate the version being packaged. A real demo may need a working development connection and sign-in; that is distinct from connecting the saved submission in the portal. If no runnable setup is available, explain that specific recording dependency, finish the independent listing and publication details, and guide the minimum supported development setup. Do not prescribe a submission draft upload merely to begin preparation, create a duplicate private plugin, or claim a recording is complete without real interactions.

1. Draft a short walkthrough from the review cases: show the version being packaged
   in the intended host, any necessary development connection steps, realistic prompts,
   actual tool results, and the main supported interactions. Include a relevant
   boundary or error-handling example. Use a dedicated test account and sample
   data, keeping passwords, tokens, and unrelated private content off screen.
2. Offer to use the **Computer** plugin to perform the visible walkthrough.
   Follow its available tool instructions and the creator's preferences. Check
   recording capabilities before starting: computer control and screenshots do
   not by themselves provide video capture. Use an available recording tool;
   otherwise guide the creator to start and stop their screen recorder while
   you demonstrate the app. If Computer is unavailable or declined, give the
   creator the same walkthrough to record manually. Do not make Computer a
   submission requirement.
3. Rehearse the flow and resolve failures before recording. Capture real
   interactions and results, leaving enough time to read them. Stay within
   authorized actions; use test or sandbox flows for actions with side effects.
4. Play back the saved video to verify readable prompts/results, successful
   playback, coverage of the walkthrough, and absence of exposed secrets. Do
   not substitute mockups, a script, or fabricated results for a working demo.
5. Help the creator host the recording at their chosen reviewer-accessible
   destination within the authorized scope. Verify playback and permissions,
   write the actual URL to `extensions.com.openai.review.demo_recording_url`,
   and rebuild the ZIP. Until a working recording link exists, keep this item
   marked incomplete; a local video alone does not complete the URL field.

For example, a dice plugin walkthrough can show “Roll a d20,” “Roll 2d6+3 and show the calculation,” and “Roll a 2.5-sided die.” Explain what to capture: the actual d20 result, both d6 results and their sum plus 3, and the invalid-input explanation without a tool call. Adapt the prompts to the actual plugin. Give the creator the recording steps and hosting guidance now; ask for the link after helping them produce the recording, then write the verified URL into the package.

### Reviewer access

Supply reviewer credentials and instructions only through secure dashboard
fields, never the package. When authentication is required, guide the creator
to provide a dedicated test account with sample data and the permissions needed
for every case. Include its exact login URL, tenant/workspace, and sign-in steps.
Verify the reviewer can sign in without access to the creator's mailbox, phone,
or private network, run the cases, and keep that access working during review.
If access cannot be verified, report the blocker rather than marking review
information complete. Verify imported values in the saved draft. If ZIP
import is unsupported, use editable dashboard fields where available and report
remaining gaps.

SHA-256: bd361a69856bb59d1bb4797e9ec6806559b224f54c7464be7db1a1b472cbd415