
If you are the Head of Digital Commerce, the pressure is no longer abstract. Everyone is calling shopping through AI chats “the future”, and leadership wants to know whether customers will be able to find, compare and buy your products there. Yet the present is less tidy: ask an AI assistant to find a product with several constraints, compare two variants or confirm a current price, and the answer can still be incomplete or wrong.
Does that mean ecommerce sites are badly built? Sometimes the merchant’s product data is part of the problem. But a poor AI result can also come from missing platform access, catalogue normalisation, retrieval or the agent’s reasoning. The answer is not visible from the chat alone.
There is another source of confusion. “The AI can buy” may describe a useful handoff to the checkout you already operate. It may also describe an agent completing the transaction itself - a capability that some platforms document but restrict by access or eligibility. Catalogue or website work cannot unlock that access on its own.
This guide is for established European retailers and manufacturers whose product truth moves through several systems, variants or markets. Before funding a feed, integration or “agentic commerce” build, test one real category from source data to purchase path. The result should tell you what is broken, who controls it and whether anything new needs to be built. If your catalogue is simple and maintained in one commerce platform, start with that platform’s current feed and checkout guidance; a cross-system review may be unnecessary.
Answer in 60 seconds
For most merchants today, the realistic purchase path is a handoff to the checkout they already operate. Some platforms document agent-completed checkout for eligible integrations, but a merchant cannot activate that capability through catalogue or website work alone. Treat it as unavailable until the chosen platform confirms access.
Begin with one commercially important product category and a risk-based set of real products: a normal bestseller, a complex variant, a current price or promotion change, an availability change, a delivery or market restriction and a product whose specifications matter in comparison.
For each case, compare the authoritative source, product page, structured data, current merchant feed and target AI channel. Check the exact product and variant, title, attributes, price, currency, availability, images, delivery, returns and eligibility. Then run five buyer tasks: exact lookup, constrained discovery, comparison, price or stock verification, and the available purchase handoff.
Classify every result as a pass, a merchant-controlled defect, a channel-side issue observed, or unknown. If the platform is partner-gated or the product was never ingested, do not call the result a website failure.
Keep the existing merchant checkout when a reliable handoff solves the journey. Additional integration work begins when the handoff must preserve a selected variant or quantity or create a prefilled cart. Only scope agent-completed checkout after the platform confirms that the merchant and integration are eligible.
The decision should be specific: repair product truth, repair public access, operate a reliable feed, improve the checkout handoff, scope a checkout integration after eligibility is confirmed - or monitor the channel without building yet.
In this article
- Two practical paths and one gated capability
- Why AI product discovery and comparison still fail
- How to test one real product category
- How to turn the findings into a decision
- What to fix now and what not to build yet
- When an external review earns its place
1. Two practical paths and one gated capability
Use the customer’s experience to define the project. Do not group discovery, checkout handoff and agent-completed purchase under one “agentic commerce” label.
| Capability | What the customer experiences | What the merchant needs |
|---|---|---|
| Find and compare | The AI identifies suitable products and explains relevant differences | Reachable or ingested product data that is current and comparison-ready |
| Merchant checkout handoff | The AI opens the correct product, basket or checkout on the merchant’s site | A stable destination that preserves the selected product, variant and, where supported, quantity |
| Gated agent-completed checkout | An eligible agent confirms totals and completes an authorised order | Confirmed platform access plus authentication, cart state, live pricing, payment, confirmation, order and exception controls |
For most merchants, the first two paths are the practical ones today. The checkout handoff can reuse infrastructure the merchant already operates, although preserving variant, quantity or cart context may require an integration. The third path exists in documented platform flows, but access is gated; it is not a capability a merchant can assume or switch on through catalogue work.
This article covers external AI shopping channels. If the assistant sits inside your own store, the decision is different; read our guide to on-site AI shopping assistants.
What is documented today?
Platform access changes quickly, so the following status is dated 20 August 2026.
| Platform path | Current documented position | What it does not prove |
|---|---|---|
| ChatGPT product feeds | OpenAI product-feed onboarding is available to approved partners. Its product schema separates eligibility for product search from eligibility for checkout. | Approval, surfacing, ranking or checkout for a particular merchant or product |
| ChatGPT merchant checkout | OpenAI currently describes merchant-hosted external checkout as the recommended generally available route for most plugin developers. | That an AI can create a cart or complete payment for every merchant |
| ChatGPT embedded payment | OpenAI currently describes its embedded payment sheet as a private beta for select marketplace partners. | Near-term access for an unapproved merchant |
| Shopify agentic commerce | Shopify documents merchant checkout handoff for general access and direct completion for eligible agents. | Universal availability across agents, merchants, markets or products |
Published infrastructure is real. Universal merchant access is not. A crawlable page or accepted feed can support discovery without enabling agent-completed checkout.
2. Why AI product discovery and comparison still fail
The difficulty is not imaginary. Recent benchmarks show that shopping agents still struggle with grounded product tasks:
- ShoppingBench, published in the AAAI 2026 proceedings, tested a controlled environment with more than 2.5 million products. GPT-4.1 achieved an absolute success rate below 50% on its tasks.
- EComAgentBench, a preprint with 662 tasks, distributed requirements across the shopper’s query, profile and a clarification step. The strongest of seven evaluated models reached 57.1% overall accuracy.
- ShoppingComp, also a preprint, evaluated 145 instances and 558 scenarios built around real products. The best model it reports, GPT-5.2, scored 17.76%. Its authors attribute the failures to grounding in open-world product data, verifying multi-constraint requirements, reasoning over noisy or conflicting evidence and risk-aware decisions.
The percentages are not directly comparable because the benchmarks use different tasks and metrics. They are not estimates of how often an ordinary shopper receives a bad answer. They establish a narrower point: finding, filtering and verifying products remains difficult even when an agent sounds confident.
A failed answer can enter the chain in several places:
| Where the problem sits | What can go wrong | Main control |
|---|---|---|
| Merchant product truth | Variant IDs change; attributes are missing; page, feed and checkout disagree on price or stock | Merchant |
| Public access | Product pages or images are blocked, unstable, duplicated or dependent on fragile rendering | Merchant, hosting and security controls |
| Channel data | The feed is missing, stale, rejected, out of scope for the market or not available to that merchant | Merchant and platform |
| Cross-merchant normalisation | Different catalogues model the same product, bundle or variant differently | Mostly platform |
| Shopper intent | Important constraints remain implicit or change during the conversation | Shopper and agent |
| Retrieval and reasoning | The agent finds the wrong candidate, loses a constraint or makes an unsupported comparison | Mostly agent and platform |
Two merchants can also structure the same legitimate product differently without either storefront being “bad”. Shopify describes this catalogue heterogeneity as a product-identity and clustering problem in its Catalog API engineering work.
The practical question is therefore not “Is our ecommerce good?” It is: where does one tested product journey stop matching the evidence?
3. How to test one real product category
This is a diagnostic test, not a statistical survey of the whole catalogue. Its purpose is to cover the product rules most likely to expose a broken identity, mapping, update or handoff.
Step 1: fix the test boundary
Choose:
- one commercially important category;
- one market, currency and delivery destination;
- one target AI channel, including its model or mode where visible;
- one test date and time;
- the merchant systems that should hold the authoritative values.
Do not begin with the whole catalogue or several AI platforms. A smaller fixed boundary makes it possible to distinguish a product-data problem from ordinary variation between channels.
Step 2: build a risk-based product set
Do not select only clean, simple products. Include at least one case from every applicable group below.
| Test case | Selection rule | Failure it can expose |
|---|---|---|
| Normal product | A popular or commercially important product with an ordinary buying path | Baseline identity, page, feed and discovery failure |
| Variant product | A product where size, colour, material, capacity or another option changes the offer | Parent/variant confusion, wrong image, price or stock |
| Price event | A live promotion or a product whose price changed recently | Stale feed, expired sale or currency mismatch |
| Availability event | A low-stock, out-of-stock, back-order or recently restocked item | Update delay and unsupported availability claims |
| Market or policy case | A product affected by delivery limits, returns rules or market eligibility | A recommendation the customer cannot actually buy |
| Comparison case | A product whose technical or category-specific attributes determine suitability | Missing fields and unsupported comparisons |
| Non-standard model | A bundle, subscription, configurable or made-to-order product, if the category contains one | Incorrect product boundaries, totals or fulfilment assumptions |
The first-pass sample is sufficiently covered when every relevant rule in the category appears at least once, including one current price or stock event. That does not make it statistically representative.
If one case fails, add further products governed by the same rule. The purpose is to learn whether the defect belongs to one record or to a repeated mapping, ownership or update process. Record that distinction; do not turn a small diagnostic sample into a catalogue-wide percentage.
Step 3: follow each case across the truth chain
Create one row per product or variant and record the same fields at each available layer.
| Layer | Evidence to capture |
|---|---|
| Authoritative source | Product and variant ID, title, attributes, price, currency, promotion, availability, eligibility and named owner |
| Live product page | Customer-visible values, selected variant, canonical URL, image, delivery and returns information |
| Structured page data | Product, offer and variant fields actually rendered in the page source |
| Current merchant feed | Submitted values, update timestamp, accepted, rejected and warning states |
| Target AI result | Product selected, variant, claims made, visible source or destination, test conditions and timestamp |
| Purchase path | Product or cart destination, retained variant and quantity, recalculated price, availability and customer confirmation |
For each field, write down the acceptable update delay before testing. A made-to-order product and fast-moving stock do not need the same tolerance. Without a declared tolerance, “fresh” has no operational meaning.
If a layer does not exist, mark it not applicable. If access or eligibility cannot be confirmed, mark it unknown. Do not fill a missing platform result with an assumption.
Google’s current product structured-data guidance recommends product-page markup, a Merchant Center feed, or both for Google’s own experiences. Google says the combination can improve eligibility and help it understand and verify product data. That is a Google-specific benefit, not a guarantee for other AI systems.
Step 4: test five buyer tasks
Use the sampled products in a clean conversation and preserve the prompt, response, sources, date, locale and platform conditions.
- Exact lookup: find a named product and exact variant from the merchant.
- Constrained discovery: find a product in the category that meets the buyer’s budget and two or more relevant constraints.
- Comparison: compare two sampled products using the attributes that actually determine the decision.
- Freshness check: confirm the current price, promotion, availability and delivery position for a volatile case.
- Purchase path: select the correct variant and use whatever next step the channel genuinely supports - product-page referral, basket or checkout.
Repeat surprising or inconsistent results under the same recorded conditions before treating them as a pattern. AI outputs can vary, and one answer does not prove a stable platform behaviour.
Step 5: give every check one honest status
| Status | Use it when | Do not claim |
|---|---|---|
| Pass | The tested stage produced the expected product, variant and values within the declared tolerance | That the whole catalogue or every prompt will pass |
| Merchant-controlled defect | An authoritative source, page, image, structured value, feed mapping or checkout destination is contradictory, unavailable or rejected for a reason the merchant controls | That repairing it guarantees an AI recommendation |
| Channel-side issue observed | The merchant-controlled evidence is correct and available, but the tested AI result is missing, wrong or loses a constraint | That the model or platform is the proven root cause after one run |
| Unknown | Platform access, ingestion, eligibility or a necessary result cannot be inspected | That the merchant passed or failed |
The useful output is an issue register with the product, field, affected layer, evidence, owner and next action. A single “AI-ready” score hides the information needed to repair anything.
4. How to turn the findings into a decision
Read the first broken layer, not the most fashionable possible solution.
| Observed finding | Appropriate next action | What not to assume |
|---|---|---|
| Product and variant IDs or customer-visible values disagree before the AI channel | Repair product ownership, mappings and update rules | That another feed or chatbot will reconcile the conflict |
| Pages or images cannot be reached as intended | Repair URLs, rendering, canonicalisation, crawler policy or security configuration | That adding structured data alone will make the product discoverable |
| Product pages are correct but the feed is missing, stale or rejected | Repair the channel mapping, validation, update and rejection process | That producing one valid export solves ongoing freshness |
| Source, page and accepted feed are correct but the AI result is wrong or absent | Repeat the controlled test, preserve evidence and treat it as a channel-side limitation or unknown | That the ecommerce site must be rebuilt |
| The correct product is found and a reliable merchant handoff is available | Keep and improve the existing checkout path | That embedded payment is necessary |
| The handoff loses variant, quantity or context | Scope the smallest cart or checkout handoff integration | That an agent-completed transaction is required |
| The business wants the agent to complete orders, but platform access is unconfirmed | Mark the capability unavailable for planning purposes and monitor the target platform | That catalogue, structured-data or website work can unlock access |
| The platform confirms the merchant and integration are eligible for agent-completed checkout | Scope authentication, state, confirmation, payment, order and exception handling | That eligibility removes transaction, customer-experience or operational risk |
This is also where the distinction between discovery and checkout becomes practical. A product can be eligible for search but not for checkout. A customer can still complete a useful AI-assisted journey when the final purchase happens on the merchant’s site.
If the failure continues through payment, fulfilment, returns or analytics, move beyond this channel test and inspect the wider ecommerce conversion chain.
5. What to fix now and what not to build yet
Fix now when the evidence supports it
- Stabilise product and variant identifiers.
- Name the authoritative system and owner for price, availability, attributes, delivery and eligibility.
- Remove contradictions between the customer-facing page, structured data, feed and checkout.
- Make intended product pages and images reliably reachable.
- Validate feeds before submission and record accepted, rejected and warning states afterwards.
- Define update tolerances and an incident owner for commercially dangerous mismatches.
- Preserve the selected product and variant when the buyer moves to the merchant’s site.
This work supports ordinary ecommerce operations as well as emerging AI channels.
Do not build yet without a verified path
- a platform-specific feed when the merchant has no confirmed access;
- a multi-platform translation layer before one end-to-end route is understood;
- agent-completed payment without confirmed platform eligibility and a controlled customer journey;
- an MCP server whose only objective is generic “AI visibility”;
- an
llms.txtfile presented as a substitute for product data, feeds or transaction tools; - a campaign promising ChatGPT inclusion, AI rankings or automatic purchasing.
A crawler, product feed, platform catalogue and MCP tool solve different problems. Our guide to llms.txt and WebMCP explains the difference between describing a site and exposing a controlled action.
Different AI engines may also retrieve and cite different sources under different conditions. If the product evidence is correct but one platform behaves differently, use a controlled multi-engine method rather than a generic visibility diagnosis; see why AI search engines cite different sources.
6. When an external review earns its place
As Head of Digital Commerce, you do not need an external supplier merely to tell you that product data should be accurate. Outside help earns its place when catalogue, ecommerce, merchandising, engineering and platform teams cannot agree on the first broken layer - or when the proposed repair crosses several systems, markets and owners.
IZZY’s Advisory model can be scoped around one category and one target channel before anyone commits to a larger build. The useful evidence to bring is:
- the category and market being tested;
- the systems holding product, price and inventory data;
- the target AI channel and current access status;
- the risk-based product set;
- one mismatch, rejection or uncertain result already observed.
If the product is found correctly and the failure begins after the buyer reaches your site - for example at JavaScript rendering, consent, anti-bot controls, cart or checkout - the more relevant route is IZZY’s Agent Readiness Audit. It tests revenue-critical site journeys with real agents. It is not a substitute for diagnosing catalogue truth, feed ingestion or platform eligibility.
The Advisory should end in a bounded investment decision: repair internally, bring in an Embedded Expert for a defined gap, scope a controlled implementation through a Full Partnership, or do not build yet. It should not promise inclusion in ChatGPT, a higher AI ranking, automatic purchases or support from a specific platform.
Conclusion
Current OpenAI and Shopify documentation covers product discovery, merchant checkout handoff and gated agent-completed checkout. That documented capability does not make today’s product results consistently reliable or every merchant automatically eligible.
The sensible response is neither to dismiss the channel nor to rebuild ecommerce around the hype. Test one meaningful category. Cover the product rules most likely to fail. Follow each item from authoritative data to the AI result and purchase path. Label merchant defects, channel-side observations and unknowns separately.
Then fund the first repair the evidence reveals. Sometimes that will be catalogue work. Sometimes it will be a feed or checkout handoff. Sometimes the merchant will have clean foundations and no platform access - and the correct decision will be to monitor rather than build.
Turn one category into a defensible investment decision
As Head of Digital Commerce, bring one product category, its source systems, the AI shopping channel you are considering and one mismatch or uncertainty you can already see. We will scope whether the next step is an internal repair, a bounded Advisory, an Embedded Expert, a controlled implementation - or no build yet.
Frequently asked questions
Not by itself. Contradictory product data, inaccessible pages or stale feeds can contribute. But platform coverage, catalogue normalisation, shopper intent and the agent’s retrieval or reasoning can also cause a poor result. Test the chain before assigning blame.
They can make product facts public and easier for supported systems to interpret. They do not guarantee that ChatGPT or another assistant ingests, retrieves, recommends or ranks the product.
It depends on the target channel. A feed can provide a normalised product record, update control and validation feedback. Follow the channel’s current specification and confirm merchant access before building a new export.
OpenAI documents a Google-compatible core feed path after it confirms that the registered feed supports that representation. Compatibility does not mean every existing export is complete or accepted unchanged. Validate the current OpenAI product specification.
Not necessarily. MCP can expose controlled tools or actions when a target journey requires them. It is not a generic replacement for crawlable pages, accurate product data or a documented merchant feed.
No. An assistant can often send a shopper to a merchant page or checkout. Agent-completed purchase requires a supported integration, confirmed eligibility, transaction controls and customer confirmation. Until the target platform confirms that access, plan for a merchant-checkout handoff.
Sources and evidence note
This article uses current official platform documentation for access, feed, crawler and checkout claims, plus methodological research on shopping-agent performance. Material sources include:
- OpenAI Agentic Commerce onboarding
- OpenAI stable product-feed specification
- OpenAI checkout documentation
- OpenAI crawler documentation
- Shopify agentic commerce documentation
- Shopify Checkout MCP documentation
- IZZY Agent Readiness Audit
- Google Product structured-data guidance
- Google guidance on sharing product data
- ShoppingBench, AAAI 2026
- EComAgentBench preprint
- ShoppingComp preprint
- Shopify Engineering on catalogue clustering
ShoppingBench is peer-reviewed. EComAgentBench and ShoppingComp are preprints. Their metrics are not directly comparable and do not estimate ordinary consumer failure rates. The risk-based category test and four result statuses are IZZY diagnostic methods, not an industry standard, statistical audit or readiness certification.
Platform access, specifications and beta status can change. Recheck availability, eligibility and checkout statements immediately before publication. This article does not promise platform inclusion, ranking, recommendation, traffic, sales or automated purchasing, and it is not legal, tax, payments or compliance advice for a specific implementation.