What should a ChatGPT app cost? Compare the quote before the price

“How much does a ChatGPT app cost?” sounds like a pricing question. In practice, it is a scope-comparison problem.

One proposal may price a demonstration connected to a clean test API. Another may include production authentication, three difficult legacy systems, confirmation for write actions, accessibility, public review, monitoring and incident support. Both can be called a “ChatGPT app”, while describing very different products.

A useful estimate begins by normalising what each supplier has included, assumed and left for the client. Only then can you compare the price.

Answer in 60 seconds

There is no responsible market price for “an app inside ChatGPT” without a defined customer job and system boundary. The main cost drivers are the number of customer journeys, the readiness of authoritative APIs, private-data access, the consequences of actions, interface states, exception handling, markets, evaluation evidence and ongoing operations.

Separate build cost from run cost. Separate fixed work from volume-dependent meters. Record who owns each account, component and operational responsibility. Treat optional model/API usage as its own line: an MCP service does not automatically incur model-token charges unless the implementation itself calls a metered model or tool.

Ask every supplier to complete the same Quote Normaliser. A lower number is meaningful only when the scope, evidence, ownership and exclusions are comparable.

In this article

  1. Why published price bands mislead
  2. Scope from the customer job, not the platform label
  3. The Quote Normaliser
  4. Split build cost into inspectable work packages
  5. Separate run cost from build cost
  6. The Cost Ownership and Metering Map
  7. Model three budgets instead of one false forecast
  8. Red flags in a ChatGPT app quote
  9. What to bring to a pricing conversation

1. Why published price bands mislead

The visible conversation is a poor unit of estimation. A search result, a customer-specific answer and a confirmed transaction might each occupy one compact card, yet their operating requirements are very different.

Consider four scopes:

Apparent requestWork that may sit behind itCost consequence
“Show our public locations”One reliable public source, a read-only tool and a simple resultBounded integration if the source is clean and current
“Show my eligible account options”Identity, tenant boundaries, permissions, private data and expired-access statesAuthentication and authorisation become core product work
“Move my booking”Fresh availability, business rules, exact confirmation, idempotent write, receipt and supportConsequential action and exception handling expand build and run scope
“Recommend and submit a regulated application”Eligibility, explanation, sensitive data, auditability and human decision ownershipThe action may need to remain narrow or outside the app altogether

The price cannot be inferred from the number of screens or prompts. It depends on the sources of truth, failure consequences and evidence required to release the promise safely.

A vendor’s public starting price is evidence of that vendor’s offer under its own assumptions. It is not a market benchmark. Timelines have the same problem: “four weeks” can mean a prototype, a read-only pilot, a submission-ready product or a production service with support. Ask which.

2. Scope from the customer job, not the platform label

Current OpenAI documentation allows a plugin to combine instructions, remote MCP tools, an optional interface and lifecycle hooks. A basic capability may need only a small read-only tool surface. A customer product can also require OAuth, several backend systems, interactive UI, write actions, public review and ongoing measurement. Plugin architecture.

Before estimating, write an outcome contract:

  • one customer job in the customer’s language;
  • the eligible user and market;
  • the start and finish state;
  • the authoritative system for every commercial fact;
  • the actions the app may read, draft and commit;
  • the consequences of an incorrect answer or duplicate action;
  • the evidence required for acceptance;
  • the first operational owner after launch.

If a supplier cannot tell which of these assumptions drives the estimate, the number is not yet decision-grade.

3. The Quote Normaliser

Give the same table to every prospective supplier. Require explicit quantities, assumptions and exclusions rather than “included” where the boundary matters.

Scope dimensionSupplier must specifyWhy it changes cost
Customer journeysNamed journeys, eligible users, finish state and excluded variantsEach journey adds tools, states, tests and support cases
Authoritative systemsAPIs, owners, environments, data quality and documented limitsMissing or unreliable APIs create adapter and remediation work
Tool surfaceRead, draft and write tools; schemas; side effects; rate limitsWrite consequences require stronger control and evidence
Identity and tenancyAnonymous, account login, roles, organisations and permission scopesPrivate or multi-tenant data adds authentication and backend authorisation
InterfaceCards, comparisons, forms, confirmations, receipts, responsive states and fallbacksEvery state needs implementation, accessibility and host testing
ExceptionsStale state, duplicates, timeouts, partial writes, outages and denied accessThe “ugly cases” often determine production reliability
MarketsCountries, languages, currencies, taxes, policy and data locationLocalisation is more than translating interface text
EvaluationPositive, negative, ambiguous, adversarial, load, security and accessibility evidenceAcceptance needs repeatable tests, not a recorded demo
Platform releaseVerification, policies, listing assets, reviewer access, submission and re-review supportPublic release is a governed delivery workstream
OperationsMonitoring, logs, alerts, support hours, incidents, patching and change regressionA live integration has continuing service obligations
Ownership and exitSource, cloud, domain, credentials, telemetry, documentation and handoverThe apparent build saving may become dependency or migration cost

Add a column for the client’s responsibility. If the quotation assumes “API provided by client”, name the endpoints, expected quality, authentication method, test environment, rate limits and delivery date. Otherwise a major dependency is hidden inside four words.

4. Split build cost into inspectable work packages

Product and journey definition

This covers customer research, the outcome contract, conversation and UI states, action boundaries, accessibility intent, markets and acceptance criteria. A vague discovery phase should produce concrete decisions that reduce the estimate’s uncertainty.

Data and integration readiness

This covers source-system access, adapters, schemas, stable identifiers, caching, rate limits and commercial validation. If price or availability is inconsistent across the business’s own systems, the app team must either repair the source, reconcile it or narrow the promise.

The “AI” layer cannot decide which conflicting record is contractually correct.

Identity, security and privacy

Authenticated customer data adds identity-provider work, tenant boundaries, minimum scopes, server-side permissions, consent, retention, deletion and incident considerations. OpenAI’s current guidance uses OAuth for user-specific access and requires authorisation checks in the backend. Authentication guidance, security and privacy guidance.

Tool and action engineering

Read-only search is different from an action that creates, pays, changes or cancels something. Consequential tools need fresh-state validation, explicit confirmation, idempotency, receipts, audit records, retry rules and rollback or remediation paths.

Interface and accessibility

Optional in-chat UI can improve comparisons, form input and confirmation. Scope every meaningful state: loading, no result, partial result, validation error, denied permission, changed price, pending action and success receipt. Include non-visual fallback and human accessibility evaluation.

Evaluation and release evidence

Current OpenAI metadata guidance recommends direct, indirect and negative prompt sets, with versioned precision/recall measurement and regression replay. Current public review guidance also requires test cases, accurate tool metadata, policy/listing assets, production endpoints and reviewer access where needed. Metadata evaluation, review requirements.

A line called “QA” is not sufficient. The quote should name the test corpus, environments, action evidence, security checks, accessibility coverage and acceptance owner.

Handover and launch

Include documentation, operator training, secrets transfer, deployment ownership, alert routes, support procedures, known limits and rollback rehearsal. Submission support should be described as work performed; platform approval, featured placement and adoption cannot be sold as guaranteed deliverables.

5. Separate run cost from build cost

A proposal can appear cheaper by moving work into unpriced client operations. Build a second map for recurring cost.

Run componentFixed or variable meterQuestions to resolve
MCP hosting and networkBase capacity plus traffic, compute and egressWho scales it, what availability is needed and what is the budget alert?
Data stores, queues and backupsCapacity, retention and transaction volumeWhich state must persist, for how long and in which region?
Identity providerActive users, authentications or enterprise contractWho owns the tenant and what happens if pricing or provider changes?
External business APIsRequest volume, licence tier or partner feeAre production rights, quotas and support already included?
Optional model/API callsModel, input/output tokens and tool usage where the backend invokes themIs a model call actually required, and which team controls the budget?
Monitoring and logsData volume, retention, seats and alertingCan the team diagnose a customer action without retaining unnecessary personal data?
Support and incident responseHours, service level and event volumeWho answers the customer and who fixes the integration?
Security and privacy operationsReviews, patches, requests and incidentsWhich obligations sit with supplier and client?
Evaluation and re-reviewChange frequency, corpus replay and submission workWhich changes trigger regression work or platform review?
Payment and commerce operationsProvider fees, refunds, fulfilment and disputes where eligibleWhich party remains merchant of record and owns customer remediation?

Current OpenAI API pricing applies when an implementation calls a metered OpenAI model or tool. A remote MCP backend that returns data from the business’s own systems does not create a universal “ChatGPT token fee” by definition. Ask the supplier to identify the exact model-dependent component before accepting that cost line. OpenAI model documentation.

Do not freeze variable run costs into one annual figure without the assumptions. Record the expected volume, unit, source, base period, budget threshold and owner.

6. The Cost Ownership and Metering Map

Create one entry for every component in the proposed service and complete the same seven fields for each. Three example components:

FieldIdentity providerMCP hostingRegression evaluation
Setup costSupplier estimateSupplier estimateInitial corpus and automation
Recurring meterMonthly active usersCompute/requests/egressPer release or monthly replay
Volume assumptionClient-provided rangePilot and scale casesPlanned release frequency
Account ownerClientTo decideClient
Operational ownerClient IAM teamSupplier or client platform teamNamed product/QA owner
Evidence sourceCurrent provider quoteCloud calculator and architectureDelivery plan
Exit costMigration and user reconnectionDeployment transfer and DNS/change workCorpus, results and tooling handover

The map is deliberately operational. It reveals whether the client will own a reusable product or rent an opaque service whose costs and evidence cannot move with it.

7. Model three budgets instead of one false forecast

Use scenarios based on your own volume and consequence assumptions:

Bounded pilot

One read-oriented customer job, one or two clean systems, limited users or route, minimal UI and a defined evidence question. The purpose is to learn whether the capability improves the journey, not to simulate a complete public product.

Production service

Real identity, operational systems, defined service levels, ugly-case handling, accessibility, public review where relevant, monitoring, support and handover. Estimate the build and twelve-month operation separately.

Expansion case

Additional journeys, markets, languages, hosts or write actions. Price it only after the first service’s data shows where complexity and value actually sit.

For each scenario, set a stop condition. A pilot should not become an indefinite programme because the team keeps expecting adoption after one more feature.

8. Red flags in a ChatGPT app quote

  • The proposal uses total ChatGPT audience as the acquisition forecast.
  • Submission, approval, publication, discovery and sales are treated as the same deliverable.
  • The estimate assumes a production-ready API without naming or inspecting it.
  • Read tools and consequential write actions receive the same security and test scope.
  • Authentication is included, but tenant roles, token failures and server-side permissions are absent.
  • The interface is priced by screen while loading, error, confirmation and accessibility states are omitted.
  • “Testing” means a demonstration prompt rather than a versioned positive, negative and ugly-case corpus.
  • Hosting is included, but monitoring, incident response, support, patching and re-review are not owned.
  • Model-token cost appears as a mandatory platform fee without an identified metered model call.
  • The supplier retains the production accounts, source or telemetry without a clear exit path.
  • The price promises a complete public channel while eligibility and platform rules remain unchecked.

A red flag is a question to resolve, not automatic evidence of a bad supplier. The response should become a written assumption, scope item or exclusion.

9. What to bring to a pricing conversation

Bring enough evidence to replace guesswork:

  • one customer job and the current journey;
  • systems and named owners for price, availability, account and transaction state;
  • API documentation or an explicit statement that it does not yet exist;
  • read, draft and commit actions;
  • roles, markets, languages and sensitive-data boundaries;
  • consequence of error and required human approval;
  • expected user and transaction volumes as ranges;
  • support and availability expectations;
  • required ownership, handover and procurement conditions;
  • the business decision the first release must inform.

The supplier can then show which unknowns require discovery, which work is optional and which assumptions would materially change the price.

Compare the promise before the number

The cheapest credible route may be better public information, a product feed or an existing provider integration. When a proprietary app is justified, its cost should correspond to the customer promise and the evidence needed to operate it.

Normalise the quotes. Separate build and run. Assign ownership. Price the ugly cases. Then compare the number.

Bring one customer journey, the relevant systems and any proposal you already have. We can turn the request into a Quote Normaliser and Cost Ownership and Metering Map before a fixed scope is proposed. When the app is justified, the build itself is the scope of our AI Agents & LLM Products service.

Sources and scope

Platform documentation was checked on 9 September 2026. Terminology, review requirements, commerce eligibility, model pricing and host capabilities are volatile and should be rechecked before procurement or publication. No market price band is stated because comparable current delivery and operating scopes could not be verified. The Quote Normaliser and Cost Ownership and Metering Map are IZZY’s procurement tools; they do not constitute a supplier quote or guarantee a commercial result.

izzy.agency teamEngineering & product insights from the izzy.agency team.We use AI in our research and preparation. The analysis, the sourcing and the writing are ours. How we work