
The demonstration is usually the easy part. A customer asks a natural-language question, the app calls a tool, and an attractive result appears in ChatGPT.
The real product begins when the inventory has changed, the customer repeats an instruction, the account token expires or an upstream system accepts only half of a request. If the app can reserve, amend, order, quote or submit, those are product states rather than technical footnotes.
A customer-facing ChatGPT app should therefore be built against release evidence, not a list of features. The decisive question is whether the business can prove that one useful customer action is accurate, authorised, recoverable and supportable.
Answer in 60 seconds
Start with one customer job and the system that has authority to fulfil it. Define the action as a controlled state change: understand the request, fetch current state, draft the action, revalidate the exact terms, show the consequence, obtain approval, commit once, issue a receipt and provide a route to undo or escalate where possible.
Expose only the narrow tools that job needs. Separate read tools from write tools. Add a visual interface only when it improves comparison, input or confirmation. Keep permissions minimal and enforce them in the backend.
Before public release, assemble evidence for the ordinary path, negative prompts and the ugly cases: duplicate submission, stale price, expired identity, partial write, timeout and upstream outage. Assign an owner to monitoring, support, policy changes and rollback.
Submission and publication are delivery steps. They do not guarantee prominent discovery or customer adoption.
In this article
- The product is the capability behind the conversation
- Define a state machine before defining screens
- Build the smallest useful tool surface
- Add interface only where it carries decision weight
- Treat identity and permissions as product behaviour
- The Release Evidence Pack
- The Ugly-Case Control Matrix
- Public review is part of the release plan
- A release sequence tied to proof
- What to put in the supplier brief
1. The product is the capability behind the conversation
Current OpenAI documentation describes a plugin as a package that can include instructions or skills, a remote MCP server, an interface and lifecycle hooks. A customer-facing integration typically uses an MCP server to expose controlled tools. An optional interface runs inside the host when cards, comparisons, forms or confirmations are more useful than prose. Plugin architecture, MCP server guide, UI guide.
The labels are moving: buyers still search for “ChatGPT app” and “Apps SDK”, while current developer documentation increasingly uses “plugins”. The business architecture matters more than the current label:
ChatGPT conversation → selected tool → integration service → authoritative business system → validated result
For a customer action, the same path returns in the other direction with a result, reference, status and support route. The model may interpret the request and select a tool. It should not become the source of truth for a price, account permission or completed transaction.
This has an immediate scoping consequence. The work is not “connect our website to ChatGPT”. It is to design and operate a bounded product across an external conversational surface and the company’s own systems.
2. Define a state machine before defining screens
A consequential action needs a visible sequence:
understand intent → fetch authoritative state → draft action → revalidate terms → show scope and consequences → obtain explicit approval → commit once → issue receipt → support undo or escalation
Each transition needs an owner and evidence.
Suppose a customer asks to move an appointment. The assistant can identify possible slots without changing anything. Once the customer chooses, the app should fetch the latest slot state, show the date, location, price difference and cancellation consequence, then ask for confirmation. The backend checks identity and permission, submits an idempotent change, returns the booking reference and records the result.
If the confirmation is repeated because the response was slow, the app should not create a second appointment. If the chosen slot disappeared, it should not claim success from an earlier read. If only one of several downstream steps completed, support needs to know what happened.
The same logic applies to an order, a service quotation, a return, a booking or a B2B request. The financial and operational consequences determine how much control is required.
3. Build the smallest useful tool surface
An MCP tool should represent a clear business capability, not a vague instruction to “handle the customer”. Useful boundaries might be:
- search available locations;
- retrieve eligible options;
- calculate a provisional quotation;
- draft a rescheduling request;
- confirm a selected change;
- retrieve the resulting receipt or status.
Separate reads from writes. A search tool can be called while the customer is exploring. A tool that commits an order needs stronger identity, confirmation, retry and audit controls.
Tool inputs and outputs should use stable identifiers and explicit schemas. Customer-facing labels can change; a booking ID or product ID should not. Return structured error states that the app can explain, rather than hiding every failure behind an invented conversational answer.
OpenAI’s metadata guidance recommends a versioned “golden prompt” set with direct, indirect and negative examples. The goal is to measure when the correct tool is called, when a different tool is better and when no tool should be called. Change one metadata variable at a time and replay the set after changes. Metadata evaluation guidance.
That is invocation evidence. It does not prove that customers will discover the app in the directory or use it at commercial scale.
4. Add interface only where it carries decision weight
Conversation is good at gathering an incomplete request and clarifying trade-offs. It is less reliable as the only representation of a dense comparison or a consequential confirmation.
An in-chat interface earns its scope when it helps the customer:
- compare several options without losing important attributes;
- edit dates, quantities or selections precisely;
- understand a total, restriction or cancellation condition;
- confirm the exact action that will be submitted;
- see a stable receipt or status.
Keep a non-visual fallback for accessibility and host differences. Test keyboard use, focus order, labels, error messaging, zoom, contrast and screen-reader meaning. Automated checks are useful, but they do not replace human evaluation of whether the decision remains understandable. WCAG 2.2 is the current W3C Recommendation to use as the accessibility baseline. WCAG 2.2.
The interface should make uncertainty and consequences clearer. Decoration alone does not justify another component to build, review and maintain.
5. Treat identity and permissions as product behaviour
Public information may need no customer identity. Private records or account actions do.
For authenticated tools, define:
- which identity provider is authoritative;
- which user and tenant the token represents;
- the minimum permission scope for each tool;
- how issuer, audience, signature and expiry are validated;
- what the customer sees when access is missing or expires;
- which authorisation rule is enforced server-side before every action;
- what personal data is logged, retained and deleted.
OpenAI’s current authentication guidance uses OAuth for user-specific access. Its security guidance emphasises least privilege, explicit consent, server-side validation, constrained content-security policies and careful handling of external content. Authentication guidance, security and privacy guidance.
The model can help interpret intent. It should not decide that a user is entitled to view another account, approve a refund or bypass a commercial rule. We set out the wider permission model in managing AI agent permissions and the product-level controls in what changes when AI can take actions.
6. The Release Evidence Pack
Ask the delivery team to produce this pack with the working integration. It turns “the demo works” into inspectable acceptance evidence.
| Evidence item | What it contains | The question it closes |
|---|---|---|
| Outcome contract | One customer job, eligible users, start and finish state, excluded actions and measurable acceptance criteria | What exactly are we releasing? |
| Architecture and truth map | Systems, data flows, source of truth for each field, acceptable age and validation point | Which system is allowed to make each claim? |
| Tool-contract register | Tool name, purpose, inputs, outputs, side effects, error states, owner and version | What can the model request, and what can the tool actually do? |
| Permissions and data inventory | Identity, scope, tenant rules, personal data, retention, deletion and log access | Who can do what, and what data crosses the boundary? |
| Golden prompt corpus | Direct, indirect, negative, ambiguous and adversarial prompts with expected behaviour | Does invocation match the intended customer language? |
| Action test evidence | Revalidation, confirmation, idempotency, authorisation, receipt and audit checks | Can a consequential action be committed once and traced? |
| Ugly-Case Control Matrix | Timeouts, duplicates, stale state, partial writes, outages and recovery behaviour | What happens outside the happy path? |
| Interface and accessibility evidence | Display modes, fallback behaviour, keyboard/screen-reader checks and human review | Can customers understand and operate the decision surface? |
| Release and review map | Domain and identity ownership, policy/listing assets, reviewer access, deployment and re-review triggers | Can the team submit, publish and change the product responsibly? |
| Operations and rollback plan | Monitoring, alerts, support, incident owner, kill switch, rollback and customer remediation | Who keeps the capability safe after launch? |
The pack should be versioned with the release. A diagram or test result from an earlier architecture is not evidence for the system currently in production.
7. The Ugly-Case Control Matrix
Do not accept “handled gracefully” as a requirement. Define the expected result.
| Ugly case | Expected customer behaviour | Required system evidence |
|---|---|---|
| Missing or ambiguous identifier | Ask for the minimum clarification; do not guess the record | No read or write against an uncertain account object |
| Expired authentication | Explain that reconnection is required and preserve safe context where permitted | Denied action, no data leakage, traceable auth failure |
| Price, stock or slot changed | Requote the authoritative current terms before confirmation | Old and new version recorded; no commitment on stale state |
| Customer repeats submit | Return the existing result or safe status | Idempotency key and one downstream commitment |
| Upstream timeout | Report pending/unknown accurately; avoid claiming failure or success without evidence | Correlation ID, retry rule and reconciliation path |
| Partial write | Stop further action, show an accurate status and route to remediation | Completed steps, compensation/rollback attempt and owner |
| Permission denied | Explain the permitted next step without exposing protected detail | Server-side denial and audit event |
| External service outage | Offer a truthful fallback or later route | Health signal, circuit breaker/limit and support message |
| Customer rejects confirmation | Make no change | No write and a recorded cancelled draft only where appropriate |
Add sector-specific cases. A retailer needs variant and fulfilment errors. An appointment service needs double-booking and timezone cases. A B2B supplier needs account-price, quantity and approval-rule conflicts. Regulated services need a narrower action scope and stronger human escalation.
8. Public review is part of the release plan
For a public remote MCP integration, current OpenAI guidance requires a stable public HTTPS endpoint, verified identity, accurate tool annotations, policy and listing information, declared network domains and reviewer access where authentication is involved. The current submission guidance also asks for at least five positive and three negative test cases. Approval is followed by a separate publication action, and material metadata changes can trigger another review cycle. Submission guidance, review requirements.
Treat these as current platform requirements, not permanent constants. Assign ownership for developer verification, domain control, policy pages, credentials, submission responses and emergency unpublishing.
Public availability and discovery are separate. Current review guidance says enhanced distribution is limited and cannot be requested. A delivery contract can include successful submission support, exact-name search and direct-link checks. It should not promise recommendation, featured placement, traffic or sales.
9. A release sequence tied to proof
1. Prove the customer job without a write action
Use realistic requests and authoritative data. Confirm that the result improves the customer’s decision enough to justify the integration.
2. Stabilise the read path
Measure data accuracy, latency, permission handling, empty results and upstream failures. Repair the source system or adapter where the truth is unreliable.
3. Add a draft action
Let the app prepare the change without committing it. Make the intended scope and consequences visible.
4. Add a controlled commit
Implement revalidation, explicit approval, backend authorisation, idempotency, receipt and audit evidence. Start with the narrowest eligible action.
5. Run the ugly cases and operational rehearsal
Test retries, outages, revoked access, stale objects, partial completion, monitoring, support and rollback. Include the people who will handle a real failure.
6. Complete platform review and a bounded release
Submit the evidence required by the host, publish after approval and expose the capability to a defined group or journey where possible. Monitor invocation quality and business outcomes separately.
10. What to put in the supplier brief
Give prospective teams the customer job, authoritative systems, read/write boundary, identity model, target markets, expected interface and consequence level. Ask every supplier to return the same evidence structure:
- included tools, systems and UI states;
- assumptions about existing APIs and data quality;
- authentication, security and privacy boundaries;
- ordinary, negative, edge and adversarial test coverage;
- release/review work and exclusions;
- monitoring, support and incident ownership;
- source code, cloud account, domain, secrets, telemetry and documentation ownership;
- change, re-review, rollback and exit terms.
This makes proposals comparable and exposes a common risk: an attractive front-end estimate that assumes the difficult integration and operating work already exists.
Build for a real promise
The customer does not experience an MCP server or an iframe. They experience whether the business gives the correct option, respects their permission, commits the intended action once and helps when something goes wrong.
That is the release standard. Build the smallest useful capability, then require evidence that the ordinary path and the ugly paths can be operated.
Bring one customer action and its source system. We can turn it into an outcome contract, define the Release Evidence Pack and identify what must be repaired before a public build is worth funding. Designing and operating that bounded capability is the scope of our AI Agents & LLM Products service.
Sources and scope
Platform documentation was checked on 9 September 2026. Product naming, review requirements, distribution, commerce eligibility and host capabilities can change and should be rechecked before submission. The state machine, Release Evidence Pack and Ugly-Case Control Matrix are IZZY’s delivery framework. They reduce ambiguity and expose risk; they do not guarantee platform approval, discovery, error-free operation or commercial return.