# izzy.agency - Full Content Dump (generated at build) > We design, build and ship AI products, SaaS platforms and Web3 apps. Fixed-quote engagements. Operator-led team. Trusted by Pernod Ricard, L'Oréal, NTT Docomo. > Website: https://izzy.agency > Email: hello@izzy.agency --- ## Services (English) ### 001 - Corporate AI Training Subtitle: Everyone has a license. Nobody works differently. Corporate AI training only works if it changes how your teams work - not if it explains what a token is. Six months after rolling out ChatGPT, most teams use it to rewrite emails, while one developer quietly automates half their job and nobody asks how. Our hands-on workshops close that gap, role by role: built on your real workflows, your real tools - ChatGPT, Claude, Copilot, Gemini - and your real data rules, run by the engineers who put LLM systems into production. From a half-day AI-literacy kickoff to a multi-week program, on-site or remote. Fixed scope, fixed quote. Deliverables: - Workshops built on your team's real tasks - people leave with prompts they reuse Monday morning - Role-based tracks: what sales needs from AI isn't what finance, support, or engineering needs - A usage policy your team actually understands - what to paste, what never to paste, and why - Working reflexes on the tools you already pay for - ChatGPT, Claude, Copilot, Gemini - A shortlist of automation candidates surfaced in the room, priced in hours saved per month - A playbook that outlives the workshop - prompts, patterns, and guardrails your team keeps Approach: Most corporate AI training fails the same way: a day of slides about how LLMs work, a round of applause, and two weeks later nobody's workflow has changed. The training wasn't wrong - it was generic. Your finance team doesn't need a lecture on transformers; they need to watch AI take apart the export they fight with every Friday. The gap between "we bought licenses" and "this is how we work now" is where the ROI evaporates - quietly, one unused seat at a time. We train on your workflows, or we don't train at all. Before the first session, we collect the tasks your teams actually do; the workshops are hands-on, on live tools, and everyone leaves with prompts that already work on their own job. The trainers are the engineers who ship LLM systems into production - training is how we hand over what production taught us, not a course we resell. And since February 2025, Article 4 of the EU AI Act expects companies using AI to ensure their staff's AI literacy: we document the program so compliance gets its answer - but the goal is a team that works differently, not a certificate that says it attended. Tech stack: AI Training, ChatGPT, Claude, Copilot, Gemini, Prompt Engineering, AI Act, AI Literacy, Workshops --- ### 002 - Agent Readiness Audit Subtitle: AI can cite you. Can it buy from you? The next visitor who abandons your checkout might not be a person. ChatGPT's agent mode, Perplexity shopping, and Google's agentic checkout now navigate, click, fill forms, and buy on your customers' behalf - and they get stuck on things human visitors never notice: content that only renders in JavaScript, consent walls, anti-bot filters, buttons that aren't really buttons. Our Agent Readiness Audit answers one blunt question: can an AI agent actually complete the journeys that make you money on your site - or does it fail halfway? Deliverables: - Your revenue-critical journeys tested end to end - cart to checkout, landing to signup, content to contact form - Real agents driven through every journey: ChatGPT agent mode, Perplexity, and a scripted Playwright + LLM agent - Screenshots and traces of the exact point where each agent gets stuck - Per-journey pass/fail and an overall agent completion score - A machine-readability and gating review - DOM vs JS rendering, schema.org, semantic HTML, CAPTCHA, WAF, consent, robots.txt / llms.txt - A prioritized fix list, split into quick config wins and dev work Approach: We don't crawl your site: we pick, with you, the journeys your revenue actually runs on - search → product → cart → checkout, landing → signup → first activation, content → quote form → submit - and drive real agents through them. ChatGPT agent mode, Perplexity, and a headless Playwright + LLM agent: we record exactly where each one breaks, with screenshots and traces. Then we check the layer underneath: JS-only rendering, schema markup, semantic HTML, CAPTCHA and WAF rules, robots.txt and llms.txt. You get a per-journey pass/fail with the exact stuck point, a completion score, and a prioritized fix list split into quick config wins and dev work. Let's be honest about the timing: agentic traffic is still a small share of most pipelines. That's why we give you scenario ranges rather than an invented "€X lost", and why most clients run this alongside SEO Auto Review - same buyer, adjacent question: one measures whether AI engines find and cite you, the other whether AI agents can transact with you once they arrive. The point is being ready before your competitors, not panicking. And because the deliverable is a prioritized fix list, it feeds straight into the rest of the studio: the quick config wins your team ships in a day, and the dev and CRO work we can implement for you. Tech stack: Agentic AI, ChatGPT Agent, Perplexity, Playwright, Schema.org, Semantic HTML, llms.txt, Anti-bot / WAF, CRO --- ### 003 - n8n AI Automation Subtitle: The POC worked. Production is a different job. An n8n workflow that breaks on Friday night tells no one - you find out on Monday, three days of leads gone. Our n8n AI automation work takes your processes (or the workflow spaghetti you already have) and turns them into a system the business can run on: error handling and retries, alerting, secured secrets, documentation, monitoring. Self-hosted with a trusted European hosting provider, GDPR-compliant - your data never passes through a US SaaS or an offshore contractor. Fixed scope, fixed quote. Deliverables: - Your processes mapped, every automation priced in hours saved per month - Production-grade automations - error handling, retries, alerting on every workflow - Self-hosted n8n with a European hosting provider - your data stays in Europe, GDPR-compliant - Secrets and access managed properly - no more credentials pasted into nodes - Your existing workflows taken over: we keep what holds, rebuild what breaks - Monitoring that catches the broken workflow before your team does Approach: n8n made automation accessible - that's its strength and its trap. The first workflow takes an afternoon. Six months later, workflows trigger each other, credentials are pasted into nodes, nobody knows what breaks when an API changes, and the person who built it all has left. That's not a failure of the tool: it's the difference between a POC and production. A low-cost contractor sells you the first; we sell the second. We work fixed-scope, in stages: an audit that maps your processes and prices every automation in hours saved per month, a build sprint that ships the production-grade automations - errors, retries, alerting, secrets, documentation, handover training - then a care plan that absorbs API changes. All of it on n8n self-hosted in Europe: that's why you chose n8n over Zapier in the first place - we just take the logic all the way. Tech stack: n8n, AI Automation, Self-hosted, GDPR, EU Hosting, LLM, Monitoring, API, Webhooks --- ### 004 - AI Visibility Audit Subtitle: First on Google. Invisible in ChatGPT. Your buyers ask AI before they ask Google, and the answer is three names: a shortlist you're either on, or not. There is no page two. Ranking #1 on Google doesn't put you there - AI engines don't read the top 10 and repeat it, they synthesize from sources they trust and can parse. Our AI Visibility Audit measures how often ChatGPT, Perplexity, Google AI Overview and Gemini actually cite you when it matters, diagnoses why they skip you, and hands you a prioritized plan to fix it. Measurement first; opinions second. Deliverables: - An AI share-of-voice score across four engines - ChatGPT, Perplexity, Google AI Overview, Gemini - A prompt set built from real buyer questions in your category, every answer scored - A benchmark against your main competitors: who gets cited instead of you, and how often - Full answer transcripts - you read what the AI says about you, word for word - Root-cause diagnosis for every missed citation, ranked by impact - Re-measurement after implementation: a before/after delta, not a claim Approach: Most agencies sell "GEO optimization" without ever measuring your baseline. If nobody counted how often you're cited today, nobody can prove you improved - and you're paying for a conviction. We run this audit as an engineering process: real buyer prompts, every major engine, competitor benchmarks, repeatable scoring. We measure your AI share of voice first, then trace each miss to its cause - content the engines can't parse, nothing quotable to cite, no third-party sources corroborating you, or competitors simply answering buyer questions better - and hand you a prioritized roadmap split into config-level quick wins and deeper work. Then we re-measure. You found this page because we do for ourselves exactly what we sell. We're not an SEO agency that bolted "AI" onto its pricing page - we're a product & AI studio that builds LLM systems in production, so we know how the engines retrieve, parse and cite. The audit runs on our own tooling - SEO Auto Review - and the people who measure are the engineers who fix. The audit is the fixed-scope front door; SEO Auto Review takes over continuously when you want to track and optimise week after week. Tech stack: GEO, ChatGPT, Perplexity, Google AI Overview, Gemini, Share of Voice, Structured Data, LLM, Schema.org --- ### 005 - AI Agents & LLM Products Subtitle: Ship intelligence, not demos. Agents that earn their seat. "What are we doing on AI?" shouldn't be the hardest question in your board pack. Our AI development work builds agents that earn their seat - pipeline qualified while you sleep, support handled overnight, ops running without a person in the loop. Marketing and sales first, because that's where AI pays back fastest. Every AI project starts with a business metric, not a tech demo. If it doesn't move revenue, retention, or efficiency, we don't ship it. Deliverables: - LLM apps that do the job, not just demo it - Agents handling work your team would otherwise hire for - Retrieval systems that know your business, not just the internet - Monitoring that catches drift before your customers do - AI features embedded in the products you already ship - Prompts tuned to your cost and quality, not generic templates Approach: Most AI projects fail the same way: they demo well, get approval, then fall apart the first time a real customer uses them. The wrong answer from an AI in front of your customers isn't a bug - it's a brand event. The cost isn't the token spend. It's the board meeting where you explain why the thing you announced last quarter embarrassed a client. We ship AI that's dull to look at and reliable to run. Evals before production so you know how often it's wrong. Guardrails before launch so the wrong answer doesn't reach a customer. Monitoring after launch so drift shows up in a dashboard, not a support ticket. Senior engineers who built this before the word "agent" became a marketing term. Tech stack: LLM Apps, Autonomous Agents, RAG, Evals, Guardrails, Observability, OpenAI, Anthropic, LangChain --- ### 006 - Full-Stack Engineering Subtitle: Your roadmap already slipped. We build the code that unblocks it. Your roadmap already slipped once this year. Full-stack engineering is how we unblock it - web, mobile, SaaS, coded to be owned by the engineer who joins next quarter, not rewritten when they do. Typed, tested, observable. Deliverables: - SaaS and web apps your roadmap actually needs - Mobile apps shipped alongside the same codebase your web runs on - APIs your integrations can live with - Data models that don't bottleneck your next feature - Performance tuned to what your users actually feel - Handoff so complete your next CTO won't need us Approach: Every shortcut in code compounds later. The rushed API becomes the reason integrations take six weeks. The skipped tests become the reason features break every sprint. The missing documentation becomes the reason your next hire rewrites the whole module. We build to the opposite standard - the one where the shortcut doesn't exist because the right thing was roughly the same effort. Senior engineers write less code, and the code they write is the code that stays. You pay for speed and quality in the same invoice. Weekly staging deploys, PR review on every merge, documentation that doesn't rot - and a handoff your team actually uses. Tech stack: Vue, React, Next.js, Node, NestJS, TypeScript, React Native, Drupal, WordPress --- ### 007 - Rescue Missions & AI-Code Hardening Subtitle: Make it survive production. Make it survive the handoff. Inherited an AI-built MVP that can't pass review? A migration stuck halfway? A codebase nobody wants to own? Our codebase rescue work takes it over, fixes the foundations, and hands back something your board, your oncall, and your acquirer can live with. Deliverables: - A map of where the risks actually live - AI-generated code cleaned, typed, tested, traceable - Incremental fixes shipped while the product keeps running - Security and performance signed off by someone accountable - A tech-debt plan your team can run after we're gone - Documentation your next hire won't have to reverse-engineer Approach: The costly mistake with a broken codebase is the rewrite. Six months of feature freeze, a team learning each other's code on the wrong project, and a "new" product that ships late and doesn't solve the original problem. We've rescued codebases that almost went that way. The pattern is always the same: the code looks worse than it is, and the expensive fix isn't the one you need. Our rescue engagements ship fixes in the same rhythm your team ships features - incremental refactoring, strangler-fig migrations, progressive modernization - so the business doesn't pause for a foundation repair. When we leave, your team owns a codebase they can reason about, and a plan they can execute without us. Tech stack: Refactor, AI-Code Remediation, Security, Migration, Tech Debt, Architecture, Testing --- ### 008 - Product Design Subtitle: Interfaces people understand on first try. And come back to on the hundredth. Leads that arrive and don't convert cost more than no leads at all - and most of the time that's a design problem, not a marketing one. Our product design work builds interfaces people understand on first try, flows built on how your customers actually think, and a design system your engineers can build from without a translator. We set the conversion target before the first wireframe, and we test against it. Deliverables: - Flows your customers navigate without help docs - UI that survives the A/B test your team runs in month three - accessible from day one, not bolted on - A design system your engineers can build from without a translator - Research and framing workshops that tell you what your customers actually think - and get your team aligned on what to build - Motion and micro-interactions that guide, not decorate - Prototypes validated before engineering spends a sprint on them Approach: Most design work fails at one of two checkpoints: the internal review, where stakeholders argue about colours, or the live launch, where users bounce because the flow didn't match how they think. Bad design isn't a taste problem - it's a conversion problem. Both failure modes come from skipping research, and both cost you weeks plus the goodwill of the team that shipped it. We start earlier, on the real users in the real product - which is where the answer already is. Senior designers come in with patterns from products your customers already use. The conventions they expect. The friction that signals untrustworthy. The micro-moments that close a sale. Busy doesn't mean smart - great design concerns itself with the details so users don't have to. You need the interface that makes the business case you signed up for, not an award-winning one - and a team that won't argue about buttons when the quarter's on the line. Tech stack: UX Strategy, UI Systems, Figma, Prototyping, User Research, Design Tokens, Storybook --- ### 009 - Cloud, DevOps & Platform Subtitle: Ship daily. Sleep nightly. The infrastructure either slows releases or wakes you at 3am - usually both. Our cloud, DevOps and platform work builds the layer that lets your team ship daily and your bill stop climbing. Quiet reliability, tuned cost, monitoring you don't have to babysit. Set up, documented, and handed over to engineers who can run it without us. Deliverables: - Deployment pipelines your team pushes through without a pager - Infrastructure your next audit won't flag as a surprise - Cost trends that don't creep up every quarter - Monitoring that pages humans only when it should - Secrets, rollbacks, and disaster recovery your board would want in writing - A runbook your team can actually follow at 3am Approach: The real cost of bad infrastructure isn't the cloud bill. It's the engineer who spends Friday debugging a deploy instead of shipping the feature that was supposed to ship. It's the board meeting where you explain the outage. It's the acquirer who reads your architecture doc and marks down the valuation. Every broken deploy has a second invoice attached, and nobody tracks it until the year's over. We build the layer where none of that shows up. Right-sized for where you are - not where a vendor's pitch says you should be. Kubernetes when you need it, managed services when you don't, and honesty about which is which. Senior platform engineers who've run infrastructure at telecom scale and at pre-seed scale, and know why the two shouldn't look the same. Tech stack: AWS, GCP, Kubernetes, Terraform, GitHub Actions, Docker, Datadog, Grafana --- ### 010 - Blockchain & Web3 Subtitle: The market won't wait. The code can't break. Both are our job. Crypto projects live or die at the audit. Our blockchain and Web3 work builds token systems, DeFi protocols, and dApps that reach mainnet, pass scrutiny, and keep running when the market turns. In the ecosystem since before the hype - still here after it. Deliverables: - Smart contracts your audit firm won't flag - DeFi protocols composable with the ecosystem you're entering - dApp frontends that actually match the wallet flow - Token systems and governance designed for the long game - Cross-chain bridges that don't make headlines for the wrong reasons - Gas optimisation baked in, not retrofitted Approach: Every Web3 exit story has the same first line: "The contract had a bug nobody caught." In crypto, the bug is the business. There's no hotfix, no redeploy, no quiet rollback - only the treasury that's gone and the Twitter thread that explains why. The real cost of cutting audit corners isn't what you save on engineering. It's the runway the exploit ends. We build to the opposite standard. Formal verification, fuzz testing, independent third-party audit, and the kind of boring code that reads the same the tenth time as the first. Senior engineers who were in the ecosystem when Solidity was the hard part, and who know which patterns hold when the market conditions they were designed for go away. Tech stack: Solidity, Rust, DeFi, dApps, Audits, EVM, Hardhat, Foundry, ethers.js --- ### 011 - Fractional CTO Subtitle: Experienced CTO leadership. None of the guesswork. Every vendor quote lands differently, and none feel honest. Fractional CTO services give you someone who sits on your side of the table - architecture calls, vendor picks, hiring, board updates - run by someone who's shipped the thing they're advising on. Accountable from day one. Deliverables: - Architecture calls that survive your next hire - Vendor picks you can defend in a board meeting - Engineering hires who stay past month three - Technical strategy that matches your business plan - Board and investor updates in the language they read - Process that ships - not process that reports Approach: Most technical mistakes compound silently. The wrong stack, the wrong hire, the wrong vendor - each is a six-figure problem your business ships with for a year before anyone notices. Pattern recognition is what closes the gap: someone who's seen the same decisions at many other companies and knows which ones don't undo later. A full-time CTO is the wrong tool for most stages of a business. You pay C-suite salary for years of work that happens in bursts - hiring sprints, board prep, platform picks, due diligence windows. We're here for those weeks, out of your way the rest, with the experience to tell the difference. Tech stack: Strategy, Architecture, Team Building, Process Design, Vendor Eval, Board Reporting --- ### 012 - Technical Due Diligence Subtitle: What you're buying. Not what you were shown. What you're acquiring rarely matches what you were shown. Our technical due diligence audits architecture, code quality, AI-code origin, security, and team capability before you sign - so the price matches the asset, and surprises don't arrive in month two. Decision-grade report, delivered in the window your deal team actually needs it. Deliverables: - An honest read on what you're actually buying - Architecture assessed against the business plan, not a generic checklist - AI-generated code provenance traced where it matters for IP and compliance - Security and compliance gaps priced into your negotiation - Team capability mapped to the roadmap you plan to ship - A remediation plan you can point at in the term sheet Approach: Most due diligence reports are written for the file, not for the decision. Two hundred pages, hedged language, every risk flagged equally, no position on what actually matters. You sign anyway, and the issues show up in month two - after the valuation is locked and the team is halfway out the door. The cost of a soft report isn't the invoice. It's the purchase price you couldn't renegotiate. We write reports that close deals or protect against them. Red flags, yellow flags, green lights - with a ranking, a dollar estimate on remediation, and a clear position on what breaks the deal versus what's a price-negotiation lever. Senior engineers who've built what they're evaluating, not just read about it. Delivered in the window your deal team needs it, not on vendor time. Tech stack: Code Audit, Architecture, Security, Scalability, Risk Assessment, Compliance, AI-Code Provenance --- ### 013 - Brand & Identity Subtitle: Said once. Remembered everywhere. If buyers can't explain what you do in one sentence, they won't remember you in six months. Brand identity is the personality of your company - and everything downstream gets easier when that personality is sharp. Our brand and identity work builds your positioning, your visual system, and the foundations that make the rest of your marketing pull its weight. Deliverables: - A one-sentence answer buyers actually repeat - Positioning that makes category choice obvious - A visual system that scales from favicon to booth - Naming, voice, and messaging that hold up outside the pitch deck - Motion and interaction guidelines your product team can apply - Brand governance your next marketing hire can run without calling us Approach: A weak brand taxes every other dollar you spend. Ad cost-per-click goes up. Demo-to-close rate goes down. Hiring takes longer because candidates can't explain where they work. The lost revenue never shows up on a line item - it shows up as a worse quarter than the numbers said you'd have. The work of a strong brand isn't aesthetic. It's the reason the same ad budget converts differently. We start with the one-sentence answer - the thing a buyer can repeat after one pitch. Then we build the visual, verbal, and interaction system that reinforces that sentence everywhere it shows up. That sentence now has a second audience: the AI engines your buyers ask. ChatGPT and Perplexity cite companies whose positioning reads the same everywhere they look - one clear sentence, consistent naming, coherent messaging across your site, your socials, your directories. Brand work decides whether machines can repeat what you do, and machines increasingly give the first pitch. Senior designers and strategists who've positioned products against commodity competitors and come out of the category the clear pick. You leave with a system your team can run - and a positioning nobody on the inside has to explain twice. "Can I just do this with AI?" You can generate a logo in thirty seconds. What you can't generate is the one-sentence answer a buyer repeats after one pitch - the positioning that makes the category choice obvious. AI gives you a thousand acceptable options. Craft is knowing which one is the right one. The difference shows up in conversion, not on a mood board. Tech stack: Positioning, Visual Identity, Naming, Motion, Guidelines, Figma, After Effects --- ### 014 - Content Strategy Subtitle: Content that reads like a human wrote it. Because one did. Content that reads like a bot wrote it won't sell to humans. Our content strategy builds the messaging that converts, the SEO that earns traffic, and the content systems that keep publishing - including the social feed that went quiet - without losing the voice that makes it yours. Deliverables: - A messaging framework your whole team can write against - Website copy that closes the loop a visitor opened - Product marketing content organised around buyer decisions, not topic clusters - SEO architecture that compounds - pages that rank, and keep ranking - GEO coverage so AI search engines cite you, not your competitors - An editorial system your team runs without a monthly agency invoice Approach: Most content marketing fails quietly. The team keeps shipping, metrics keep stalling, and the cost-per-lead stays higher than paid. The real cost isn't the content team's salary - it's the compounding advantage you lose to the competitor whose content actually ranks. By the time you audit the funnel, they own the search results, the AI citations, and the first pitch your buyer sees. Catching up takes twice the work you'd have done on time. We write to land, not to fill a calendar. Buyer interviews, real decision paths, and the one thing your content has to do at each step - then the copy, the structure, and the editorial system that makes the whole thing repeatable. Senior writers who've built content engines for technical products people actually read. You leave with a system your team runs, a voice nobody can clone, and a distribution plan the next hire doesn't have to reinvent. Tech stack: Copywriting, SEO, Messaging, Content Ops, Analytics, Editorial, GEO --- ### 015 - Long-term Support & Maintenance Subtitle: Browsers change. Hackers try. Your product keeps running. Technology doesn't stand still, and neither does what's trying to break it. Browsers update, dependencies change, security threats evolve, and the platform you shipped last year needs the attention this year. Our long-term support and maintenance work keeps your product live, current, and monitored - by the same team that shipped it, on a predictable retainer that scales with what you need. Deliverables: - Uptime monitoring with alerts that reach humans before your customers do - Security patches applied before vulnerabilities reach production - Performance tuning tied to what your users actually feel - Bug fixes ranked by business impact, not ticket volume - Dependency updates and platform upgrades managed without disruption - Monthly reporting your team actually reads Approach: The hidden cost of neglected maintenance isn't the outage itself - it's the day the outage happens during a launch, a board meeting, or a customer demo. The work of long-term support is making sure the call never comes. Our retainers aren't about logging tickets and running down a checklist. They're about knowing your stack well enough to catch the pattern that precedes the incident, and having the senior engineer on call who built the thing in the first place. Most agencies treat support as a post-sale service line. We treat it as the continuation of a partnership that started with a build - and charge accordingly. Fixed monthly retainer scoped to what your product actually needs, not a bundle of hours you hope to use. When your business changes, we reprice against the new scope. When it doesn't, the retainer doesn't either. Tech stack: Monitoring, Security Patches, Performance, Dependency Updates, Drupal, WordPress, Magento, Node, React --- --- ## Products (English) ### News Platform - AI CONTENT ENGINE Tagline: Your accounts went quiet. Your expertise didn't. Buyers look at your feed before they look at your website. If the last post is from four months ago, that's the first impression. News Platform takes the publishing job off your team's desk: it reads your industry's news, writes branded content in your voice, and publishes to your site, LinkedIn and newsletter on a schedule you set. You approve; it ships. Features: - Automated industry news monitoring and aggregation - AI-powered content generation in your brand voice - Multi-channel publishing - website, social, newsletter - Editorial dashboard with approval workflows - Content calendar with automated scheduling - Performance analytics per article and channel - SEO-optimised output from day one - Position your brand as the industry authority Tech stack: NLP, GPT-4, RSS / API Ingestion, Multi-Channel, Editorial AI, SEO Optimisation, Analytics, Headless CMS Pricing: Fixed quote per scope, set during scoping. What we scope is what we quote, and what we quote is what you're invoiced. URL: https://izzy.agency/en/products/news-platform/ --- ### Smart Lead Conversion - AI SALES AUTOMATION Tagline: Your prospects followed up, 24/7. Most leads don't pick the best vendor. They pick the one who answered first, properly. Smart Lead Conversion answers first: personalised email sequences, voice follow-ups, and AI-analysed responses from the moment a lead comes in. Your sales team only talks to warm leads. Features: - Automated personalised sequences - AI-powered response analysis - Voice & email follow-ups - Works nights, weekends, holidays - Lead scoring & qualification - CRM integration - Real-time engagement analytics - Sales team handoff with full context Tech stack: AI, Email Automation, Voice AI, CRM Integration, NLP, Analytics Pricing: Fixed quote per scope, set during scoping. What we scope is what we quote, and what we quote is what you're invoiced. URL: https://izzy.agency/en/products/smart-lead-conversion/ --- ### SEO Auto Review - AI SEARCH VISIBILITY Tagline: Visible on Google. And on every AI your buyers ask. Your buyers ask ChatGPT before they ask you. When the answer names a competitor, you never find out - the deal simply doesn't call. SEO Auto Review watches those answers for you: automated positioning analysis, keyword recommendations, and optimisation for ChatGPT, Claude, Perplexity, Google AI Overview, and whatever your buyers ask next. Features: - Automated SEO positioning analysis - AI search engine optimisation (GEO) - Content & keyword recommendations - Competitor visibility tracking - Schema markup generation - Conversational AI citation tracking (ChatGPT, Claude, Perplexity, Google AI Overview, and beyond) - Weekly automated reports Tech stack: SEO, GEO, AI Search, Analytics, Structured Data, ChatGPT, Claude, Perplexity, Google AI Overview Pricing: Fixed quote per scope, set during scoping. What we scope is what we quote, and what we quote is what you're invoiced. URL: https://izzy.agency/en/products/seo-auto-review/ --- --- ## Personalised Solutions URL: https://izzy.agency/en/personalised-solutions/ Three ways to work with izzy - one senior team, fixed quote per scope, no account managers between you and the work. - Advisory - when you need the thinking, not the building. Architecture review, vendor evaluation, technical due diligence, AI & search-visibility audits, fractional CTO time. Fixed scope per engagement, or monthly retainer. - Embedded Expert - when you need specialist capability, not a full team. An experienced engineer, designer, AI specialist, or platform expert embedded with your team. Fixed monthly or per-sprint. - Full Partnership - when you need us to own the outcome. End-to-end delivery - discovery, design, build, ship - run by one team, with monitoring and advice after launch. Fixed quote per scope. Constant across all three: experienced operators, fixed quote per scope, one point of contact, honest read. --- ## Guides (English) - The company brain: search and automate your company knowledge (9 parts): https://izzy.agency/en/blog/guides/company-brain/ A nine-part guide to building a searchable operating memory for your company with n8n: from meeting notes to permission-aware RAG, approvals and evaluation. --- ## Blog articles (English) ### Campaign strategy and execution: resolve decision debt before AI scales it URL: https://izzy.agency/en/blog/campaign-decision-debt/ Published: 2026-09-04 Summary: Find the campaign decisions your team only thinks it has made - before production, media and AI make the hidden assumptions expensive. Campaigns rarely arrive as a blank page. They arrive with a launch date, a deck marked “approved”, assets already moving and several teams carrying different versions of what has supposedly been decided. That gap is **campaign decision debt**. AI, production and media spend do not create it. They make unresolved assumptions expensive faster. ## Answer in 60 seconds An approved campaign idea is not necessarily an executable campaign. The objective may be broad, the audience action implicit, the mechanic untested, the dependencies unowned and the measurement disconnected from any later decision. Before scaling the work, label every consequential part as **fixed, assumed, missing, conflicting or owned**. These are not maturity levels: a date can be fixed and still create debt if nobody owns its consequences. IZZY first establishes what has actually been decided. Every choice must connect to an audience action and credible brand benefit. If complexity threatens real use or the fixed date, simplify the execution while preserving the objective. This does not guarantee campaign performance. It prevents the team from spending, producing and automating as if unresolved questions were settled facts. ## In this article 1. [Approved is not the same as decided](#1-approved-is-not-the-same-as-decided) 2. [What campaign decision debt includes](#2-what-campaign-decision-debt-includes) 3. [The IZZY Campaign Decision Debt Map](#3-the-izzy-campaign-decision-debt-map) 4. [Use the map from A to Z or after idea approval](#4-use-the-map-from-a-to-z-or-after-idea-approval) 5. [Protect the campaign promise from impression to action](#5-protect-the-campaign-promise-from-impression-to-action) 6. [A non-standard mechanic carries the highest interest](#6-a-non-standard-mechanic-carries-the-highest-interest) 7. [Let AI scale owned decisions, not decide through the gaps](#7-let-ai-scale-owned-decisions-not-decide-through-the-gaps) ## 1. Approved is not the same as decided Campaign work tends to begin in one of two states. Sometimes the brand needs the journey from A to Z. There is a product, audience, budget or launch date, but the campaign itself still needs a purpose, proposition, idea, experience, production system and measurement logic. Sometimes the creative territory is already approved. The work appears further ahead, but approval may cover only the presentation: the visual world, headline or high-level mechanic. It may not cover what a person will actually do, which content and data the experience needs, how markets will adapt it, what a media partner receives or how success will change the next decision. Both situations can contain the same chaos. Marketing, creative, product, content, development, markets and media partners each hold part of the campaign. A slide becomes “final” while its consequences remain distributed across people who believe somebody else decided them. This is why we do not begin by adding deliverables to the list. We begin by separating constraints from assumptions and approvals from executable decisions. Campaigns often lose value through small gaps rather than one dramatic error: a changing promise, a mechanic with no job, mobile considered after desktop, late tracking or a metric nobody uses. ## 2. What campaign decision debt includes Campaign decision debt is the cost created when a team proceeds as though an important decision has been made when it has not. We use five labels to expose it: - **Fixed** - a real constraint that will not move, such as a launch date, budget ceiling, market, approved territory or media commitment. - **Assumed** - something treated as true without enough evidence or explicit agreement: people will understand the mechanic, assets will arrive, consent will cover the data use, or the platform metric represents value. - **Missing** - a decision that has not been made at all, even though production depends on it. - **Conflicting** - two teams are making valid decisions against different priorities, definitions or versions of the campaign. - **Owned** - one accountable person can state the choice, its deadline, its dependencies, the trade-off accepted and the brand value it is meant to protect. In compact form: **decision debt = assumptions + missing decisions + unresolved conflicts + fixed constraints with unowned consequences**. An **owned decision = choice + owner + deadline + dependencies + accepted trade-off + value protected**. The labels are not sequential. “Fixed” does not mean “resolved”: a fixed date increases debt when the mechanic, content or approval path remains assumed. An owned decision may remain provisional when its owner defines what evidence will confirm or change it. Not every question needs an immediate answer. If uncertainty is visible, has an owner and resolution date, and work cannot outrun it, deferral is planning. Hidden deferral is debt. Decision debt gains “interest” when production makes change costly, localisation multiplies it, media or AI scales it, and the fixed date removes recovery time. ## 3. The IZZY Campaign Decision Debt Map The map is not a renamed campaign checklist. Its job is to reveal where the campaign is acting on an assumption and what value the brand risks by leaving it unresolved. | Campaign layer | What may look decided | Debt to expose | The decision is owned when… | Brand value protected | |---|---|---|---|---| | **Objective and benefit** | “Awareness, engagement and conversion” are listed | There is no priority or later decision | One change, baseline, trade-off and next decision are named | Budget serves a defined purpose | | **Audience exchange** | A segment and message are approved | Why people should care or act remains implicit | Tension, promise, action and value exchange are explicit | Attention can become relevant response | | **Idea and mechanic** | The concept works in the deck | Novelty replaces function; failure paths stay assumed | The mechanic has a job, acceptance criteria and fallback | Creative ambition survives real use | | **Production reality** | Assets and suppliers are listed | Content, markets, technology, data and approvals follow different plans | Owners, dependencies, critical path and release criteria are visible | Quality and the fixed date survive rework | | **Distribution handoff** | Channels and budget are selected | Assets, destination, signals and events are not one brief | Media receives valid inputs, boundaries and optimisation objective | Spend scales a coherent campaign | | **Measurement and learning** | A dashboard is planned | Metrics have no interpretation owner or consequence | Each signal has meaning, owner and possible next action | Evidence improves the next investment decision | A decision does not become owned because a meeting ended. It becomes owned when somebody accepts its consequences and can explain the value being protected. The map needs three passes: label the current state without solving it; resolve the debt connected to an immovable date, irreversible build, committed spend, consent or several markets; then define the simplest objective-preserving fallback. ## 4. Use the map from A to Z or after idea approval When IZZY works from A to Z, the map stops formats being chosen too early. A product launch does not automatically need a microsite, game or complex personalisation. The format earns its place after the intended brand change and audience exchange are clear. When an idea already exists, the map does something different. It does not reopen taste by default. It identifies what the approval did and did not cover. Across engagements, we repeatedly encounter a fixed date, approved direction, incomplete content, several markets, an external media partner and a mechanic whose mobile behaviour or data dependency is still assumed. This is a recurring condition, not one reconstructed client story. The point is to expose contradictions without flattening the idea. Marketing may maximise participation, brand protect premium expression, development reduce states to meet the date, and media need a faster first action. The debt sits in pretending they are aligned. Structuring the chaos means returning every conflict to an owner and the campaign objective: what stays constant, what can adapt, which dependency can invalidate the route and what a simpler version must preserve. Every action and production choice should be able to answer one question: what benefit is the campaign paying this cost to create for the brand? If the answer is only “because the idea contains it”, the decision is not yet owned. ## 5. Protect the campaign promise from impression to action A customer does not experience the media plan, design file and website as separate workstreams. They experience one promise. If an advertisement leads with one product, benefit or cultural idea and the destination opens with something else, the campaign spends attention and then asks the customer to reconstruct the connection. The page can be attractive and technically correct while still breaking the campaign. Before multiplying formats, reduce the campaign to a small message spine: - **Audience:** who is this for in this context? - **Tension:** what matters to them now? - **Promise:** what useful change or experience is being offered? - **Proof:** what makes the promise credible? - **Action:** what should happen next? The spine is an acceptance test, not a fixed set of headlines. Language, visual execution and cultural references can adapt by market. A variation becomes a different strategic hypothesis when it changes the audience, promise, proof or action. Treating that change as routine production hides another decision. Continuity also needs to survive the click. Preserve the relevant product or offer, market and language, price or eligibility conditions, chosen state and expected next action. Do not advertise a specific bottle, collaboration, event or configuration and send the visitor to a generic destination where they must find it again. The owned experience should then be designed as a short product journey, not a poster with a button. Define what the visitor already expects, what must be recognised immediately, what evidence they need, the primary action, the confirmation state and the recovery path. Recovery is part of conversion design. Products become unavailable, promotions expire, forms fail, integrations slow down, consent choices change measurement and campaign links continue circulating after the planned end date. An owned decision explains what happens in those states and who can correct them while traffic is arriving. This is where IZZY keeps marketing inside design and engineering. A cleaner interface is not automatically a better campaign experience. The change must preserve the promise, support the intended action and protect the evidence the business needs next. For a lead-generation website rather than a campaign, the same continuity test is covered in [Your website gets visitors but no enquiries? Check these five things before spending more on ads](https://izzy.agency/en/blog/website-visitors-no-enquiries/). ## 6. A non-standard mechanic carries the highest interest A beautiful mechanic can fail as soon as a person uses it. We often see concepts become slow or broken on mobile; translated content no longer fits; the front end expects data the back end cannot supply; or analytics cannot distinguish meaningful participation. Late technical failure consumes time reserved for content, QA and approval. The campaign may then miss a date tied to a launch, media booking, partnership or event. This is decision debt in a concentrated form. The hidden assumption is not simply “the technology will work”. It is that the exact technical expression is essential to the campaign’s value. Challenge that assumption while the mechanic can change. Product design, front-end and back-end perspectives should test the behaviour most likely to invalidate it using realistic screens, content, loading, errors and data. Measurement belongs in the prototype, not after it. At the same time, define a simpler route. The fallback is not a cheaper imitation of the original. It should preserve the audience action, brand promise and launch date while removing the dependency most likely to fail. IZZY would rather simplify spectacle than protect complexity that makes the experience unusable, unmeasurable or late. Technical feasibility is therefore not separate from marketing value. It decides whether that value can reach a person at all. The build boundary matters too. One hard-coded campaign page may be the right answer for a small, genuinely singular activation. A continuous campaign calendar may need reusable components, content fields, market permissions, analytics identifiers and expiry states. The decision should follow the programme, not an automatic preference for either a disposable page or a universal platform. A component is not reusable if markets cannot operate its content model. A flexible CMS is not useful if its permissions ignore the approval path. The campaign system is owned when the team can state what repeats, what remains idea-specific, who can publish or pause it and how the experience retires without becoming a dead end. ## 7. Let AI scale owned decisions, not decide through the gaps Advertising products such as Google Ads Performance Max use AI across bidding, budget optimisation, audiences, creatives and attribution. Generative tools can accelerate synthesis, variants, resizing, draft localisation, tagging and reporting. They act on goals, assets, events, signals and rules supplied by people. They do not know whether a blank is intentional or whether the team simply avoided a decision. Unless the distinction is encoded, automation treats both as input. The debt labels create a practical boundary: - automate within **fixed and owned** constraints; - use **assumed** items as hypotheses to test, not truths to scale; - stop automation from filling a **missing** strategic decision by convenience; - escalate **conflicting** objectives to the accountable owner instead of optimising whichever signal is easiest to measure. Distribution still belongs to the system when a specialist manages media buying. Brand, destination, assets, events, exclusions and measurement need one handoff; otherwise the platform optimises a contradictory campaign. How a small business should size and read a first paid test is covered in [Google Ads vs Meta Ads for small businesses: what to check before you spend](https://izzy.agency/en/blog/google-ads-vs-meta-ads-small-business/). Measurement closes the loop only when it changes something. Reach, participation, qualified visits, registrations, product exploration, retailer clicks or sales can all matter in the right campaign. None is valuable merely because a dashboard can display it. Before launch, write an evidence contract for the actions that matter. It should name: - the event and exact trigger; - the campaign, market, content and product context attached to it; - the systems through which the action passes; - consent-dependent, duplicate, retry, cancellation and unavailable states; - the source that confirms the downstream outcome; - the person responsible for interpretation; - the next decision the evidence is allowed to change. A submitted form may be the experience conversion while qualification and opportunity progression live in the CRM. A purchase event may still need reconciliation with cancellations, returns, tax, fulfilment and margin. A game completion can show participation without proving memory, preference or incremental sales. The same boundary applies to ROI. Agree which media, creative, production, technology, localisation, incentive, support and internal costs count; what kind of value is being evaluated; and what the attribution or comparison design can support. Sometimes the result is incremental return. Sometimes it is attributed revenue, qualified pipeline, conversion cost or directional evidence. Do not promote the easiest dashboard ratio into profit. Decide who interprets each signal, what evidence is sufficient and which result changes the next allocation. AI can surface patterns; it cannot turn attribution into causation or choose the next investment without context. Once the campaign is live, this discipline continues. Operational defects may need immediate repair. Performance hypotheses need an owner, one declared change, an observation window and a rollback. If the team changes media, message, offer, experience and tracking together, the campaign may improve while its learning value collapses. If the campaign is already live, continue with [Your campaign is live. How do you decide what to fix, scale or stop?](https://izzy.agency/en/blog/post-launch-campaign-decisions/). The Campaign Decision Debt Map is not a performance guarantee. Its value is more disciplined: fewer hidden assumptions, clearer trade-offs and a campaign that can explain why it is doing what it is doing before speed and spend magnify the answer. ## What IZZY's public work supports and what it does not IZZY's current public cases support the delivery territory behind this approach. The [Havana Club case](https://izzy.agency/en/case-studies/#havana-club) describes UX/UI optimisation, front-end and back-end work, interactive mini-games, promotional campaigns produced end to end and technical consulting for a multi-country brand with a continuous campaign calendar. The [Lillet case](https://izzy.agency/en/case-studies/#lillet) describes UX/UI, front-end and back-end delivery, technical consulting and promotional pages for the Emily in Paris × Lillet collaboration and a new-bottle launch. These pages verify that IZZY designs and builds owned campaign experiences where brand, interaction, content and technology meet. They contain no campaign-level data sufficient to establish a particular conversion lift, return on investment or causal performance result. This article therefore uses the cases as bounded capability evidence, not as proof that the method guarantees an outcome. ## Bring the unresolved decisions, not only the polished brief Bring the campaign as it is: launch challenge, contradictory brief or approved idea. IZZY can expose its decision debt, test the highest-risk mechanic and define the production system and handoffs. If media buying belongs with an internal team or specialist partner, we make that boundary explicit. If the work needs strategy, product design, brand/content, AI or full-stack delivery in one partnership, we scope the responsibilities rather than assuming them. The three ways to work with IZZY are described on [Personalised Solutions](https://izzy.agency/en/personalised-solutions/): advisory, embedded expert or full partnership. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### What is campaign decision debt? It is the later cost of treating an assumed, missing, conflicting or unowned decision as settled: rework, weak handoffs, unusable experiences, wasted spend or lost time. ### Can the map be used when the creative idea is already approved? Yes. The map identifies what approval covered and which audience, content, technical, distribution or measurement decisions remain open. ### Where should AI be used? Use it inside clear goals, evidence and boundaries. Treat assumptions as hypotheses, escalate conflicts and do not let convenience fill a missing strategic decision. ### Does IZZY manage paid media? This article does not present IZZY as a paid-media agency. Media may remain internal or with a specialist; responsibilities and the creative, experience, data and measurement handoff must be explicit. ### Is this an anonymous case study? No. It reflects IZZY’s combined experience without naming clients or merging engagements into a fictional chronology or result. ## Evidence boundaries Campaign Decision Debt is IZZY practice guidance, not an external standard or guarantee. Platform AI capabilities change and should be checked when a campaign is scoped. No client-specific outcome is claimed. Results still depend on proposition, audience, offer, media, product, implementation, market conditions and follow-up. The evidence contract, message spine and decision-debt labels are IZZY practice guidance. They do not replace an appropriate analytics, research or incrementality design. ## Sources and evidence note Checked on 4 September 2026: - [IZZY case studies](https://izzy.agency/en/case-studies/) for the bounded Havana Club and Lillet scope above. - [IZZY Personalised Solutions](https://izzy.agency/en/personalised-solutions/) for the current Advisory, Embedded Expert and Full Partnership routes. - [How IZZY uses AI](https://izzy.agency/en/how-we-use-ai/) for the current human-accountability boundary. - [Google Ads guidance on Performance Max](https://support.google.com/google-ads/answer/10724817?hl=en) for the current example of AI used in bidding, budget optimisation, audiences, creatives and attribution. - [Google Ads guidance on advertisements and landing pages](https://support.google.com/google-ads/answer/6238826/optimising-your-ad-and-landing-page?hl=en-GB) for message and call-to-action continuity. - [Google Analytics guidance on key events](https://support.google.com/analytics/answer/9267568?hl=en) for the current definition of a business-important event and the distinction between event reporting and attribution. --- ### Your campaign is live. How do you decide what to fix, scale or stop? URL: https://izzy.agency/en/blog/post-launch-campaign-decisions/ Published: 2026-09-04 Summary: Use live campaign evidence to separate defects from hypotheses and decide what to fix, hold, test, scale or stop without losing the learning. One platform reports an efficient conversion. The campaign page shows more completed actions. The CRM says the leads are weak. Finance cannot see the expected value yet. Each view may be accurate within its own boundary. The dangerous response is to change the media, message, offer, experience and tracking at once. The numbers may improve, but the team will no longer know what changed the result. At IZZY, marketing stays inside the design and engineering work after launch. We separate an operational defect from a performance hypothesis, connect platform activity to the owned experience and downstream outcome, and record each live change before the campaign scales it. ## Answer in 60 seconds Do not ask only whether the dashboard is up or down. Ask whether the campaign is observable, which level of evidence changed and what decision that evidence is strong enough to support. Then choose one of five routes: - **Fix now** when the experience is not behaving as approved: a broken form, wrong product, failed route, missing content or invalid event. - **Hold** when the system works but the evidence is immature, delayed or too sparse for a responsible change. - **Test** when a plausible explanation exists but several causes could produce the same pattern. - **Scale** when the result appears in the evidence layer that matters to the business and the conditions that produced it can be preserved. - **Stop** when the route is invalid, unsafe, commercially indefensible or consuming more value than the evidence can justify. Record the observation, source, evidence level, decision, owner, exact change, what remains fixed, review window and rollback. That record is the **Campaign Change Register**. This method does not guarantee conversion or ROI. It helps a team act on live evidence without letting the loudest dashboard or fastest stakeholder silently redefine the campaign. If the campaign is not live and important approvals remain unresolved, start with [Campaign strategy and execution: resolve decision debt before AI scales it](https://izzy.agency/en/blog/campaign-decision-debt/). ## In this article 1. [Confirm that the campaign is observable](#1-confirm-that-the-campaign-is-observable) 2. [Reconstruct the live outcome chain](#2-reconstruct-the-live-outcome-chain) 3. [Separate defects from performance hypotheses](#3-separate-defects-from-performance-hypotheses) 4. [Choose fix now, hold, test, scale or stop](#4-choose-fix-now-hold-test-scale-or-stop) 5. [Use one Campaign Change Register](#5-use-one-campaign-change-register) 6. [Report conversion and ROI at the level the evidence supports](#6-report-conversion-and-roi-at-the-level-the-evidence-supports) 7. [Turn live decisions into the next campaign brief](#7-turn-live-decisions-into-the-next-campaign-brief) ## 1. Confirm that the campaign is observable A live campaign can produce numbers before it produces reliable evidence. The advertisement may be serving while the destination is unavailable in one market. Analytics may count a button click while the form behind it fails. A key event may fire twice after a retry. Consent behaviour may change what is visible by device or geography. The CRM may receive the lead without the campaign identifier needed to reconnect it to the experience. These are not performance findings. They are observability or operational defects. Before interpreting the result, check the complete path a person actually takes: 1. Can the intended audience receive the correct campaign promise? 2. Does each entry route reach the expected product, offer, language and state? 3. Does the primary action work on the relevant devices and browsers? 4. Does the confirmation state reflect what the downstream system received? 5. Are event names, triggers and identifiers stable across variants and markets? 6. Can the CRM, commerce or operational system identify the outcome that matters? 7. Did a release, content change, consent update or tracking edit alter the comparison? Use real records where possible. A successful test submission does not prove that production leads reached the correct queue. An analytics purchase event does not prove that the order remained valid after payment, cancellation or return. A retailer click does not reveal whether stock existed at the destination. For an ecommerce destination, the full path from product data to fulfilled order is audited in [Why is your ecommerce website not converting? Audit the path from product data to fulfilled order](https://izzy.agency/en/blog/ecommerce-website-not-converting/). The first live decision may therefore be to repair the evidence path rather than optimise the campaign. That is not delaying marketing. It prevents the team from rewarding or punishing work based on a broken observation. ## 2. Reconstruct the live outcome chain Once the campaign is observable, put the evidence in the order the customer and business experience it. | Evidence layer | What it can show | What it cannot establish alone | Useful source of truth | |---|---|---|---| | **Distribution** | Delivery, reach, clicks, spend and platform-reported actions | Whether the destination worked or the action created value | Media platform and partner records | | **Promise** | Which audience, offer, message and creative led to the visit | Whether recognition, comprehension or trust caused the next action | Campaign taxonomy, creative record and research where available | | **Owned experience** | Entry behaviour, meaningful interaction, completion and failure states | Lead quality, fulfilled revenue, margin or incremental effect | Product analytics, logs, form and commerce records | | **Downstream handoff** | Qualification, opportunity progression, order state, attendance or follow-up | Whether the campaign caused the outcome | CRM, commerce, operations or research system | | **Commercial outcome** | Revenue, contribution, qualified pipeline or another agreed value | Incrementality unless the comparison design supports it | Finance, revenue operations or an agreed business record | The layers are not a scoreboard where every campaign must reach the bottom immediately. A brand activation, event campaign, retailer-driving campaign and direct ecommerce campaign have different jobs. The purpose of the chain is to stop an early signal from quietly becoming a later claim. If click-through improves while qualified outcomes remain flat, the campaign may have become better at attracting a cheap action rather than a useful one. If the owned experience converts but CRM quality falls, the proposition, qualification rule, audience or downstream response may be misaligned. If one market appears weak only in the advertising platform while fulfilled orders remain healthy, the attribution boundary deserves investigation before the experience is rebuilt. This is why one campaign needs a shared interpretation, not merely access to several dashboards. Each team can report accurately and still optimise a different definition of success. ## 3. Separate defects from performance hypotheses IZZY draws a hard line between a **defect** and a **hypothesis**. A defect means the live campaign is not doing what the approved system is supposed to do. Examples include: - the advertisement and destination show different offers; - a market route returns the wrong language or product; - the primary action fails or loses the visitor's input; - an event fires on the wrong trigger; - a promotion continues after its eligibility period; - a third-party service fails without a recovery path; - the downstream team does not receive the information it needs to respond. The default response is to contain or fix the defect, verify the repair and record the period affected. Do not preserve a broken experience to protect an experiment. A performance hypothesis is different. The system behaves as designed, but the result raises a question: - the audience may not recognise the promise; - the proof may be insufficient for the action; - the offer may be attractive but poorly qualified; - the interaction may add friction without adding value; - the media route may bring the wrong context; - the reported conversion may be a weak proxy for the commercial outcome. These explanations are not facts because a chart moved. They compete with other explanations. The appropriate response is a bounded test or a decision to hold, not an unrecorded redesign presented as a repair. The distinction protects both speed and learning. Teams can repair an invalid experience immediately while resisting the pressure to change every plausible cause of underperformance at once. ## 4. Choose fix now, hold, test, scale or stop The same observed pattern can support different decisions depending on evidence quality, business risk and reversibility. There is no universal number of days, clicks or conversions that makes a campaign ready to scale. | Decision | Use it when | Required action | Boundary | |---|---|---|---| | **Fix now** | Approved behaviour, factual truth, availability, accessibility, tracking or handoff is broken | Contain the impact, repair, verify and mark the affected period | Do not call the repair an optimisation win | | **Hold** | The system works but volume is sparse, outcomes lag, seasonality or another change clouds the comparison | Preserve the current conditions and set the next review point | Waiting without an owner or date is not a decision | | **Test** | A changeable cause is plausible and the evidence can distinguish it from alternatives | State the hypothesis, change one meaningful layer where practical, define guardrails and rollback | Cosmetic variation is not automatically a useful test | | **Scale** | The relevant downstream result is credible, capacity can absorb more demand and the winning conditions are understood well enough to protect | Increase exposure in bounded steps, monitor quality and keep the message, experience and outcome definition coherent | More platform volume can expose operational or commercial limits | | **Stop** | The proposition, route, timing, safety, legality, availability or economics are invalid - or further spend cannot answer a useful question | Stop the affected route, preserve evidence and record what should not be repeated | Stopping one route does not automatically condemn the whole campaign idea | The decision owner should be able to explain why the selected route protects more value than the alternatives. For example, low volume is not always an invitation to broaden an audience. The destination may be failing, the offer may be unavailable, the conversion window may still be open or the campaign may serve a deliberately narrow high-value group. Conversely, a large number of inexpensive platform conversions is not a scale signal when downstream qualification, margin or fulfilment deteriorates. The strongest decision is not the most aggressive one. It is the one the available evidence can support and the operating system can absorb. ## 5. Use one Campaign Change Register A media-platform change log can show that a bid, asset or configuration changed. A deployment log can show when the owned experience changed. Neither usually explains the complete business reason across the advertisement, destination, CRM, commerce flow and follow-up. The Campaign Change Register is the shared layer above those specialist logs. | Field | What to record | Why it matters | |---|---|---| | **Timestamp and affected period** | When the observation began, when the decision was made and when the change reached people | Creates a valid before-and-after boundary | | **Observation** | What changed in the evidence, without adding an explanation | Separates fact from interpretation | | **Evidence level and source** | Distribution, promise, owned experience, downstream handoff or commercial outcome; exact system or report | Shows what the signal can and cannot prove | | **Alternative explanations** | Other credible causes that could create the same pattern | Prevents the first convenient story becoming fact | | **Decision** | Fix now, hold, test, scale or stop | Makes the action explicit | | **Exact change and owner** | The layer, configuration, content, component or process that will change; accountable person | Prevents distributed, invisible optimisation | | **What remains fixed** | Audience, promise, offer, destination, market, measurement or another protected condition | Preserves some interpretability | | **Expected effect and guardrail** | What should move and what must not deteriorate | Connects the action to campaign value | | **Observation window and next review** | When the evidence should be reconsidered and why that timing is appropriate | Prevents premature or forgotten conclusions | | **Rollback or fallback** | How to restore the previous route or move to the simpler safe version | Makes change reversible where possible | The register should be small enough to use while the campaign is live. It is not another report deck and it does not replace platform, product or engineering logs. Its purpose is decision memory. When results move, the team can see whether a new creative, budget change, landing-page release, product-availability update, consent change or sales-response issue happened in the same period. When the next campaign begins, the team can recover the reasoning rather than inherit only the final asset. If several urgent changes must ship together, record the bundle and its reason. Do not later attribute the result to one element as though the others stayed fixed. ## 6. Report conversion and ROI at the level the evidence supports Google Analytics currently defines a key event as an action particularly important to the success of the business. That is a useful product definition. Marking an event as important does not by itself establish the event's downstream value or incremental effect. Use precise performance language: - **Platform-reported action:** the advertising platform assigned an action under its settings and attribution logic. - **Experience conversion:** the person completed the action the owned campaign experience was designed to influence. - **Qualified or fulfilled outcome:** the downstream system confirmed that the action met the agreed business condition. - **Attributed value:** a model assigned credit to the campaign or touchpoint. - **Incremental effect:** an appropriate comparison supports the claim that the campaign changed what would otherwise have happened. - **ROI:** the included incremental value and included costs are both defensible enough for the stated calculation. Attribution remains useful for reporting and tactical decisions. It is not the same question as causation. Different models can distribute credit differently across touchpoints, and a campaign can contribute to a result without being solely responsible for it. Before presenting ROI, agree which costs count. Media, strategy, creative, production, technology, licences, incentives, localisation, support and internal time may all matter. Agree whether the value is revenue, contribution, qualified pipeline or another defensible outcome. Record cancellations, returns, fulfilment and sales capacity where they materially change that value. Sometimes the available result is attributed revenue or qualified pipeline. Sometimes it is conversion cost or a directional comparison. Reporting the supported level is more useful than forcing every campaign into a confident return figure. The reporting language also controls optimisation. If the business cares about qualified opportunities but the platform receives only form submissions, automated delivery may become excellent at finding people who submit and poor at finding people the business can serve. The remedy is not a more persuasive label on the dashboard. It is a better evidence and feedback path. The CRM side of that path, including what each automated follow-up is allowed to send and why, is covered in [CRM outreach and consent: does your automation know why it is sending each message?](https://izzy.agency/en/blog/crm-outreach-consent-automation/). ## 7. Turn live decisions into the next campaign brief Post-launch work should improve the current campaign where useful and reduce uncertainty in the next one. At each review, preserve: - the audience, promise and action actually exposed to people; - which markets, routes and versions genuinely differed; - the operational incidents and periods they affected; - the evidence at each layer, including missing or conflicting records; - every fix, hold, test, scale or stop decision and its owner; - which explanation became stronger, weaker or remained unresolved; - which component, content rule, tracking pattern or fallback should be reused; - which assumption must return to the next brief as an owned decision. This closes the loop with campaign decision debt. A live campaign can retire an assumption by producing useful evidence. It can also reveal a new conflict: the platform optimises one action, the destination encourages another and the business values a third. Do not preserve only the winning visual or final dashboard. Preserve the conditions under which the result appeared and the change history needed to interpret it. Otherwise the next team receives files without the campaign knowledge that made them valuable. If the next brief includes a first paid test, [Google Ads vs Meta Ads for small businesses: what to check before you spend](https://izzy.agency/en/blog/google-ads-vs-meta-ads-small-business/) sets out what to check before you spend. ## Bring one live campaign and the disagreement IZZY's public [Havana Club and Lillet work](https://izzy.agency/en/case-studies/) verifies experience across campaign UX/UI, interactive work, campaign pages, front-end, back-end and technical consulting. The public cases do not disclose campaign-level ROI or a post-launch performance chronology, so this article does not claim one. If your live campaign has several correct dashboards and no shared decision, bring: - the objective and customer action; - the live entry routes and owned destination; - the main platform, product and downstream evidence; - the last meaningful changes and their dates; - the result people disagree about; - the decision that cannot wait. A bounded IZZY Advisory can reconcile the evidence, identify whether the next move is a repair, test or stop decision, and define the owner, observation window and rollback. If implementation work is needed, scope it after the diagnosis rather than assuming it. Media buying can remain with your internal team or specialist partner; the handoff and evidence boundary still need to be explicit. Advisory is the smallest of the three engagement models described on [Personalised Solutions](https://izzy.agency/en/personalised-solutions/). [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### How soon after launch should a campaign be reviewed? Operational and measurement checks begin immediately. A performance decision waits for the evidence window appropriate to the campaign's volume, buying cycle, attribution settings and downstream lag. There is no universal day that makes every campaign mature. ### What should be changed immediately? Repair invalid approved behaviour: broken journeys, wrong or expired content, unavailable offers, inaccessible primary actions, incorrect events and failed downstream handoffs. Record the affected period and verify the repair. ### When is a campaign ready to scale? When the result is credible at the evidence layer the business values, the conditions that produced it can be protected, and fulfilment, sales or operations can absorb more demand. A low platform cost alone may not be enough. ### What if Google Ads, Meta Ads, analytics and the CRM disagree? First compare definitions, attribution windows, identifiers, consent boundaries, time zones, duplicates and downstream status. The systems may be counting different valid things. Choose the source of truth for the specific decision rather than forcing one dashboard to answer every question. ### Should we change one thing at a time? When practical and when learning matters, keep one meaningful layer stable so the result remains interpretable. Urgent defects and bundled operational repairs may require several changes; record the bundle and do not claim one element caused the outcome. ### Does IZZY manage paid media? This article does not present IZZY as a paid-media agency. IZZY can align campaign purpose, the owned experience, full-stack implementation, events, downstream handoffs and the evidence used with an internal or specialist media team. ### Can a live campaign prove ROI? Only when the value, included costs, attribution or comparison method and data quality support that claim. Otherwise report the strongest level available, such as attributed revenue, qualified pipeline, fulfilled conversion or directional evidence, without upgrading it to causal return. ## Evidence boundaries The Campaign Change Register and five decision routes are IZZY practice guidance, not an external standard, universal optimisation cadence or performance guarantee. No threshold, conversion lift, ROAS, revenue result or client-specific post-launch outcome is claimed. Appropriate timing and evidence depend on the campaign objective, audience, market, volume, product, media, implementation, operational capacity and measurement design. ## Sources and evidence note Checked on 4 September 2026: - [IZZY case studies](https://izzy.agency/en/case-studies/) for the bounded Havana Club and Lillet campaign-experience scope. - [IZZY Personalised Solutions](https://izzy.agency/en/personalised-solutions/) for the current Advisory route. - [Google Analytics guidance on key events](https://support.google.com/analytics/answer/9267568?hl=en) for the current definition of a key event and the use of attribution reports. - [Adobe Advertising campaign change-log documentation](https://experienceleague.adobe.com/en/docs/advertising/dsp/campaign-management/campaigns/campaign-change-log) as a current example of a platform-level record of campaign changes. The cross-system Campaign Change Register described here is IZZY's practice method, not an Adobe feature. --- ### How to index company documents without duplicates, stale answers or a data swamp URL: https://izzy.agency/en/blog/document-ingestion-freshness/ Published: 2026-08-30 Summary: Design a RAG ingestion pipeline that tracks document identity, updates, permissions, deletion, replay and freshness without hiding failures. The same brief appears several times in search while an older, deleted or newly restricted copy can influence an answer. Nobody can see where the update stopped. A reliable **RAG ingestion pipeline** uses a controlled lifecycle: discover, identify, fetch, parse, chunk, upsert, reconcile and expose state. Keep its four identities linked but separate. A webhook starts work; it does not prove current search. This article owns ingestion and freshness. Source authority ([part 3 of this guide](https://izzy.agency/en/blog/systems-of-record-drive-notion-slack-github-linear/)), solution selection ([part 4](https://izzy.agency/en/blog/notion-enterprise-search-vs-custom-rag/)) and full evaluation and value ([part 9](https://izzy.agency/en/blog/guides/company-brain/)) remain separate. No live tenant, corpus, parser, workflow, permission model, vector index or client outcome was tested. ## The answer in 60 seconds **RAG ingestion pipeline:** 1. Discover the change and retain a checkpoint. 2. Resolve stable identity. 3. Fetch the permitted current version. 4. Parse and classify under a versioned contract. 5. Chunk with provenance and access metadata. 6. Upsert with retry-safe controlled keys. 7. Reconcile access, deletion and supersession. 8. Expose freshness, failure and recovery. The aim is retry-safe, observable reconciliation - not exactly-once delivery, immediate freshness, universal deduplication or automatic permission correctness. ## In this article 1. [Discover the source change](#1-discover-the-source-change) 2. [Resolve stable identity](#2-resolve-stable-identity) 3. [Fetch the permitted current version](#3-fetch-the-permitted-current-version) 4. [Parse and classify](#4-parse-and-classify) 5. [Chunk with source metadata](#5-chunk-with-source-metadata) 6. [Upsert idempotently](#6-upsert-idempotently) 7. [Reconcile deletion, permission and supersession](#7-reconcile-deletion-permission-and-supersession) 8. [Expose freshness and failure state](#8-expose-freshness-and-failure-state) ## 1. Discover the source change **Decision:** which signal starts work, and what proves its scope is still covered? Use an initial scan, then a checkpoint. Drive search, [user/shared-drive change logs](https://developers.google.com/workspace/drive/api/guides/about-changes), [page tokens](https://developers.google.com/workspace/drive/api/guides/manage-changes) and [shared-drive parameters](https://developers.google.com/workspace/drive/api/guides/enable-shareddrives) have distinct scopes. Record scope, owner and rescan route. [Drive notifications](https://developers.google.com/workspace/drive/api/guides/push), [Notion events](https://developers.notion.com/reference/webhooks-events-delivery), [Slack events](https://docs.slack.dev/apis/events-api/), [GitHub deliveries](https://docs.github.com/en/webhooks/using-webhooks/best-practices-for-using-webhooks) and [Linear webhooks](https://linear.app/developers/webhooks) have different renewal, ordering, retry, payload and access limits. Each still needs current state. Backfill differs: [Notion search is not exhaustive or immediate](https://developers.notion.com/reference/search-optimizations-and-limitations), [Slack history is scoped and paginated](https://docs.slack.dev/reference/methods/conversations.history/), and [Linear pagination](https://linear.app/developers/pagination) is not an immutable event log. **Illustrative shared-Drive scenario - not an IZZY client case.** Rename, move, revision, permission and deletion enter here. An expired channel without a checkpoint sets `replay_required`, not current. ## 2. Resolve stable identity **Decision:** what key survives presentation changes and addresses every derived record? A [Drive file ID](https://developers.google.com/workspace/drive/api/guides/about-files) survives a rename and [parent change](https://developers.google.com/workspace/drive/api/guides/folder). A copy has its own source identity even when content matches. Provider identities differ: Notion [pages](https://developers.notion.com/reference/page), [blocks](https://developers.notion.com/reference/block) and data sources; [Slack channel and message timestamp](https://docs.slack.dev/reference/events/message/message_changed/); GitHub [delivery ID](https://docs.github.com/en/webhooks/webhook-events-and-payloads) versus [path plus ref/SHA](https://docs.github.com/en/rest/repos/contents); and [Linear delivery UUID](https://linear.app/developers/webhooks) for a payload, not its object. | Identity | Controlled meaning | Must not become | |---|---|---| | Source record | Native authority: `source_system + source_id` | Name, URL or delivery | | Extracted document | Parsed observed source version | Source authority | | Chunk | Unit linked to source, version and transform | Independent truth | | Vector/index entry | Store record mapped to `chunk_id` | Universal provider ID | | Exact field | Minimum meaning | Boundary | |---|---|---| | `source_system` | Provider plus connection/tenant | Do not infer wider coverage | | `source_id` | Native ID within that scope | Not a display name | | `source_version` | Provider version observed | `unknown` when unavailable | | `source_updated_at` | Provider-reported source update time | Not fetch or index time | | `indexed_at` | Observed index-write time | Not query-visibility proof | | `permission_class` | ACL reference for reconciliation | Not store enforcement | | `content_hash` | Hash under recorded normalisation | Not semantic dedupe | | `chunk_id` | Addressable chunk key | Links to source/version | | `sync_state` | Processing/failure state | Not boolean `synced` | | `superseded_by` | Replacement pointer | Not automatic | | `deleted_at` | Deletion/tombstone time | Not downstream proof | ## 3. Fetch the permitted current version **Decision:** can the ingestion identity still read the intended current bytes or object graph? Drive uses [media download for blobs and export for Workspace documents](https://developers.google.com/workspace/drive/api/guides/manage-downloads); format and size constrain parser input. Notion needs current sharing and [block traversal](https://developers.notion.com/reference/block); Slack needs [token and scopes](https://docs.slack.dev/authentication/installing-with-oauth/); GitHub needs [path, ref, size and permission](https://docs.github.com/en/rest/repos/contents); Linear can return [HTTP 200 with partial data and errors](https://linear.app/developers/graphql). Separate fetch from parse: `denied` is not `partial`. A Drive move keeps `source_id` but may change inherited access. Recheck before serving. ## 4. Parse and classify **Decision:** did the intended document class transform completely under the current contract? [Notion child blocks and unsupported types](https://developers.notion.com/reference/block) show why a response is not completeness evidence. n8n's [Default Data Loader](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.documentdefaultdataloader) loads binary or JSON, attaches metadata and connects a splitter; it does not supply identity, permission reconciliation or parser assurance. Version the parser/classifier, class, transform and required units. Quarantine unsupported input as `schema_parser_failure` for replay. A revised brief missing a required section is `partial`, not current. ## 5. Chunk with source metadata **Decision:** can each chunk be traced, filtered, superseded and removed? n8n exposes [chunk size and overlap](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.textsplitterrecursivecharactertextsplitter); settings do not create identity or prove completeness. Its [component map](https://docs.n8n.io/build/integrate-ai/understand-ai-components/store-and-search-data-with-vectors) separates loaders, splitters, embeddings, stores and retrievers without recovery. Pinecone recommends [structured IDs and source metadata](https://docs.pinecone.io/guides/index-data/data-modeling); Qdrant stores application-supplied [JSON payload](https://qdrant.tech/documentation/manage-data/payload/). Metadata supports filters, not proof of current source permissions. For a revision, derive addressable chunks, retain the source/version link and mark old chunks for supersession. ## 6. Upsert idempotently **Decision:** can a retry run without accumulating another uncontrolled current representation? Workflow surfaces differ: n8n's [Pinecone node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.vectorstorepinecone) documents ID-based update; its [Qdrant node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.vectorstoreqdrant) exposes another surface. Neither defines the provider's complete API. n8n [Remove Duplicates](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.removeduplicates) compares configured workflow fields within input or prior executions; it is not semantic document deduplication. Pinecone [upsert overwrites the same record ID](https://docs.pinecone.io/guides/index-data/upsert-data) but does not remove sibling chunks. Qdrant documents [point-specific writes](https://qdrant.tech/documentation/manage-data/points/) and [several synchronisation patterns](https://qdrant.tech/documentation/data-synchronization/), not one automatic source guarantee. Own event, source-version, chunk and store keys. A same-hash Drive copy is a duplicate candidate, not the same authority. Write current entries, link supersession and reconcile siblings. ## 7. Reconcile deletion, permission and supersession **Decision:** what must leave search, become inaccessible or be marked replaced? Drive's [file resource](https://developers.google.com/workspace/drive/api/reference/rest/v3/files) exposes trash, parent, version, permission and caller-relative capability state; moves can change inherited access. Notion's [trash endpoint changes `in_trash`](https://developers.notion.com/reference/trash-page), not permanent deletion. [Slack deletion events expose channel and deleted timestamp](https://docs.slack.dev/reference/events/message/message_deleted/); access still depends on [OAuth scope](https://docs.slack.dev/authentication/installing-with-oauth/). Store removal is separate. Pinecone supports [ID-, metadata- and namespace deletion](https://docs.pinecone.io/guides/manage-data/delete-data) with eventual consistency; Qdrant has [point deletion](https://qdrant.tech/documentation/manage-data/points/). Acceptance is not immediate query absence. For the Drive brief: rename preserves identity; move requires inherited-access re-evaluation; revision supersedes chunks; a same-content copy is a candidate duplicate; permission loss sets `denied`; trash/delete sets `deleted_at`, records a tombstone, removes or quarantines entries and verifies downstream state. ## 8. Expose freshness and failure state **Decision:** can an operator see what is current, what failed and where replay starts? GitHub's [delivery view](https://docs.github.com/en/webhooks/testing-and-troubleshooting-webhooks/viewing-webhook-deliveries) is bounded; failures need [manual or API redelivery](https://docs.github.com/en/webhooks/testing-and-troubleshooting-webhooks/redelivering-webhooks). [n8n error workflows](https://docs.n8n.io/build/flow-logic/handle-errors-gracefully) need configuration, visibility is access-scoped and [failed runs can be retried](https://docs.n8n.io/build/understand-workflows/understand-executions/view-all-executions). Neither is a durable source checkpoint. | `sync_state` | Observed trigger | Owner action | Clear condition | |---|---|---|---| | `duplicate` | Event/key/hash repeat | Classify layer; reconcile | One evidenced current representation | | `stale` | Source ahead of index evidence | Find lagging stage; replay | Class freshness evidence is current | | `denied` | Fetch/use not permitted | Stop serving; reconcile | Reprocessing or removal confirmed | | `deleted` | Trash/delete observed | Tombstone; remove/quarantine | Removal confirmed; no resurrection | | `partial` | Required units failed | Preserve and isolate | Success or recorded exclusion | | `replay_required` | Gap, expiry or failure | Replay checkpoint/backfill | New checkpoint; no gap | | `schema_parser_failure` | Parser cannot produce form | Quarantine; record versions | Compatible reprocessing succeeds | Measure source observation, `source_updated_at`, checkpoint, processing and `indexed_at` separately; set class freshness, not a universal SLA. Only gap-free replay plus reconciled Drive version, access, chunks and removal observably clears `replay_required`. ## Conclusion: make current state an evidenced result A RAG ingestion pipeline is a controlled lifecycle, not a connector plus `synced`. Stable identities, reconciliation and observable freshness/failure are the controls. This design has not been live-tested. ## Scope one controlled ingestion-pipeline audit > Bring one source, document class, change-volume pattern, access model, freshness expectation and known duplicate or stale-answer failure. IZZY will define what can be observed, reconciled and tested. The decision may be to improve source governance first. This is exactly the scope of our [n8n AI Automation service](https://izzy.agency/en/services/automatisation-ia-n8n/). [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9). ## Frequently asked questions ### What is a RAG ingestion pipeline? It discovers, identifies, fetches, parses, chunks, upserts, reconciles and exposes freshness or failure. ### Can content hashes or n8n remove every duplicate? No. Hashes flag equal content; [n8n compares configured fields](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.removeduplicates). Event, source, document, chunk and index duplicates need separate rules. ### Does a webhook mean the knowledge base is current? No. A [webhook](https://docs.github.com/en/webhooks/using-webhooks/best-practices-for-using-webhooks) signals work. Fetch, permission, parse, index, reconciliation and evidence remain separate. ### What should happen when source permissions change? Set `denied`, stop serving or quarantine, reconcile access, then reprocess or confirm removal. [Drive sharing](https://developers.google.com/workspace/drive/api/guides/manage-sharing) remains an input, not proof. ### How should deleted documents leave a vector index? Record a tombstone and `deleted_at`, use the [provider removal route](https://docs.pinecone.io/guides/manage-data/delete-data), prevent replay restoration and confirm absence. ## Sources and method Research was checked on 2026-07-27 using current official Google Drive, Notion, Slack, GitHub, Linear, n8n, Pinecone and Qdrant documentation. Official links sit beside product claims. This is IZZY architecture guidance. No live tenant, webhook, parser, workflow, index, permission propagation, deletion, replay, retrieval result, SLA or client outcome was tested. --- ### Notion Enterprise Search or custom RAG: when should you buy, connect or build? URL: https://izzy.agency/en/blog/notion-enterprise-search-vs-custom-rag/ Published: 2026-08-30 Summary: A practical SME decision framework for testing Notion Enterprise Search, a bounded connection and custom RAG without assuming a universal winner. Staff need current answers across Notion, Drive, Slack and GitHub. Yet the business has not established coverage, access, freshness or weak-answer handling. For **Notion AI vs custom RAG**, test **buy** first, **connect** one gap, and **build** only where required coverage, controls, evaluation or actions remain unavailable with a named owner. This reference architecture is not a comparison test: no tenant, corpus, access model, workflow, answer set, price or client outcome was tested. ## The answer in 60 seconds - **Buy** when native search and connectors pass the question, access, freshness and evidence tests. - **Connect** when one bounded query, monitor or handoff closes a named gap. - **Build** when gaps remain and the organisation accepts retrieval, access, evaluation, monitoring and exit. If the evidence is unknown or nobody owns the failure, reduce the requirement or defer automation. ## In this article 1. [Define the decision: buy, connect or build](#1-define-the-decision-buy-connect-or-build) 2. [Establish the questions and acceptance conditions](#2-establish-the-questions-and-acceptance-conditions) 3. [Assess buy: where native Notion search may be enough](#3-assess-buy-where-native-notion-search-may-be-enough) 4. [Assess connect: close one defined gap](#4-assess-connect-close-one-defined-gap) 5. [Assess build: own the retrieval product](#5-assess-build-own-the-retrieval-product) 6. [Compare the same scenario](#6-compare-the-same-scenario) 7. [Calculate operating burden without invented prices](#7-calculate-operating-burden-without-invented-prices) 8. [Run the decision gate](#8-run-the-decision-gate) ## 1. Define the decision: buy, connect or build This is not a contest over a universal winner or a whole-suite replacement. It is a choice of ownership based on required questions. **Buy** uses native Enterprise Search and configured connectors. **Connect** adds a bounded query, transformation, monitor or handoff without creating a general retrieval product. **Build** owns custom ingestion and retrieval plus access, evidence, evaluation, monitoring and exit. Native can be correct; "do not build" is valid. A small connection avoids a forced binary. ## 2. Establish the questions and acceptance conditions **Illustrative SME scenario - not an IZZY client case.** One SME needs authorised staff to answer the same five questions across Notion, Drive, Slack and GitHub: 1. What client scope is approved, and what changed? 2. Which delivery milestone is current, and what is blocking it? 3. Which decision changed the work, and where is its evidence? 4. Which GitHub pull request, issue or file implements that decision? 5. What may this user see, how current is the answer, and what supports it? Freeze the criteria: coverage, access, freshness, citations, structured retrieval, actions, evaluation, operating burden, lock-in and exit. Set freshness per question and name evidence, failures and owners. Test allowed and denied users, a changed role, departed user and disconnected source. Vendor documentation is an input, not tenant assurance. Freshness is a selection condition here; ingestion, deletion, reconciliation and replay belong to [part 5 of this guide](https://izzy.agency/en/blog/document-ingestion-freshness/). ## 3. Assess buy: where native Notion search may be enough Notion documents [Enterprise Search](https://www.notion.com/en-gb/help/enterprise-search) for Business and Enterprise plans: workspace and connected-app search, source narrowing and citations. Model choice can affect connected information. This does not show that the scenario passes. Connector presence is not coverage. The [overview](https://www.notion.com/en-gb/help/notion-ai-connectors) favours finding and summarising over complex calculations or broad aggregation; app-specific boundaries differ: - [Slack](https://www.notion.com/en-gb/help/notion-ai-connectors-for-slack) covers configured public and some user-added private content, with Slack Connect, Canvas and List exclusions plus bounded history and indexing. - [Drive](https://www.notion.com/en-gb/help/notion-ai-connectors-for-google-drive) supports named file types under ownership, group and shared-drive rules, with target-audience, spreadsheet-analysis and update limits. - [GitHub](https://www.notion.com/en-gb/help/notion-ai-connector-for-github) covers code, pull requests, issues, files and READMEs, with wiki/fork exclusions, history differences, authentication and indexing constraints. [SharePoint and OneDrive](https://www.notion.com/en-gb/help/notion-ai-connector-for-microsoft-sharepoint-and-onedrive) differ again in permissions, files, exclusions, lookback and updates. Limits cannot be flattened into one promise. Notion's [security statements](https://www.notion.com/en-gb/help/notion-ai-security-practices) also do not prove tenant enforcement or compliance. Choose buy only if the five questions pass for representative people, permissions, freshness and evidence. ## 4. Assess connect: close one defined gap If the milestone question needs exact status and client filters, a bounded connection may close that gap. Notion's [Search endpoint](https://developers.notion.com/reference/post-search) searches shared page and data-source titles. Its [limitations](https://developers.notion.com/reference/search-optimizations-and-limitations) say results are neither exhaustive nor immediate and are not optimised for data-source filtering. A narrower pattern can retrieve a shared [schema](https://developers.notion.com/reference/retrieve-a-data-source), apply [property or compound filters](https://developers.notion.com/reference/filter-data-source-entries), paginate and return source identity. It still needs least-required [capability](https://developers.notion.com/reference/capabilities), explicit sharing, [pagination](https://developers.notion.com/reference/pagination), [limit](https://developers.notion.com/reference/request-limits) handling, monitoring and an owner. Keep this category narrow. If it starts to require general multi-source ingestion, semantic retrieval, identity propagation and evaluation, reassess it as build. ## 5. Assess build: own the retrieval product Build is not a vector-store node. n8n documents [load, split, embed, store and retrieve](https://docs.n8n.io/build/integrate-ai/understand-ai-components/retrieve-relevant-context), with metadata and agent or direct retrieval. Loader, retriever, answer components and an [example path](https://docs.n8n.io/build/integrate-ai/ai-examples/use-website-content) remain building blocks, not a production system. The system must own permission-aware retrieval, source validation, citations, grounded review, expected outcomes, metrics, monitoring and regression response. n8n documents [test datasets](https://docs.n8n.io/build/integrate-ai/test-and-improve-ai-workflows/understand-why-to-test) and evaluation, but no universal threshold or accuracy result. Documented patterns can send an [unanswered query to a person](https://docs.n8n.io/build/integrate-ai/ai-examples/set-a-human-fallback-for-ai-workflows), pause for [action approval](https://docs.n8n.io/build/integrate-ai/ai-examples/human-in-the-loop-for-tools), and trigger [error workflows](https://docs.n8n.io/build/flow-logic/handle-errors-gracefully). None proves recovery or continuity. The [NIST GenAI Profile](https://doi.org/10.6028/NIST.AI.600-1) covers purpose, evaluation, sources, suppliers, fallbacks and monitoring. [OWASP](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) says RAG does not fully mitigate prompt injection; its [RAG guidance](https://cheatsheetseries.owasp.org/cheatsheets/RAG_Security_Cheat_Sheet.html) covers access checks, attribution, deletion, logging and fail-closed behaviour. Choose build only when buy and connect leave required gaps and an owner accepts these obligations. Ingestion engineering remains outside this article. Adjacent IZZY routes cover [AI-agent security](https://izzy.agency/en/blog/ai-agent-security-product-controls/) and [permission lifecycle](https://izzy.agency/en/blog/ai-agent-permissions-access-control/). ## 6. Compare the same scenario This is a transparent decision model, not a client case. Keep the same five questions, four sources, representative users and failure conditions for every route. | Requirement | Buy | Connect | Build | Evidence needed | Owner | |---|---|---|---|---|---| | coverage | unknown - test connector exclusions | gap - intentionally bounded | unknown - design and test | five-question results | knowledge lead | | access | unknown - test mapped access | unknown - test shared scope | unknown - test retrieval filters | allowed/denied lifecycle tests | security owner | | freshness | unknown - measure each source | unknown - expose query delay | unknown - define acceptance | timestamps and delay cases | source owner | | citations | unknown - inspect answer links | unknown - retain source identity | unknown - design attribution | reachable evidence | content owner | | structured retrieval | unknown - test exact fields | unknown - validate API filter | unknown - design if required | expected records | operations lead | | actions | gap - search is the scope | gap - controlled handoff only | unknown - approve selected actions | action boundary and audit | process owner | | evaluation | unknown - run question set | unknown - test handoff | unknown - own dataset and metrics | expected outcomes | evaluation owner | | operating burden | unknown - record administration | unknown - include integration | unknown - include full lifecycle | burden ledger | service owner | | lock-in | unknown - map suite dependency | unknown - map API/schema dependency | unknown - map model/store suppliers | dependency inventory | technical owner | | exit | unknown - test export/replace route | unknown - remove handoff safely | unknown - preserve source authority | stop and recovery plan | executive owner | Record no answer, weak or ungrounded answer, denied access, connector or index delay, workflow failure and action awaiting approval as different failures. Do not hide a blocking failure inside a weighted total. ## 7. Calculate operating burden without invented prices Use a qualitative ledger: `operating burden = setup + recurring ownership + incident/recovery + supplier/change response + re-evaluation + exit` Record task, owner, trigger, evidence, failure consequence and exit dependency. Buy includes connector administration, tests and exit; connect adds integration monitoring; build adds retrieval, access, evaluation, incidents and supplier choices. Do not invent prices, days, headcount, ROI or savings. The [NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework) is voluntary, under revision and not certification. NCSC guidance on [design](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-design), [development](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-development) and [deployment](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-deployment) informs threat, supplier, failover and deployment questions without selecting a product. [Production governance](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/) and [technical due diligence](https://izzy.agency/en/services/due-diligence/) are adjacent reviews, not evidence that a route passes this gate. ## 8. Run the decision gate 1. **Buy** if native capability meets all five questions and every blocking access, freshness, evidence and failure condition. 2. **Connect** if one bounded, owned integration closes the remaining defined gap without becoming a general retrieval product. 3. **Build** if documented gaps remain in required coverage, controls, evaluation or actions and the organisation accepts operating and exit burden. 4. **Reduce or defer** if the requirement is unnecessary, ownership is missing or evidence remains unknown. Record the evidence, owner, unresolved unknowns, reassessment trigger and exit condition. Category selection may depend on freshness, but [part 5](https://izzy.agency/en/blog/document-ingestion-freshness/) owns stable identity, change capture, chunking, upsert, deletion, permission reconciliation and replay implementation. ## Conclusion: choose the least-owned route that passes Assess Notion Enterprise Search first, connect a defined residual gap, and justify custom RAG only through unmet requirements and accepted ownership. No tenant or comparative performance test was performed; use the organisation's question, access, freshness and failure evidence. ## Test your shortlist before assuming you must build > Bring five recurring questions, your current sources, your access model, one unacceptable failure and your current tool shortlist. IZZY will determine whether native search covers the requirement, a bounded connection closes a defined gap, or a custom build is justified. The assessment may recommend the native option, or less automation. This is exactly the scope of our [n8n AI Automation service](https://izzy.agency/en/services/automatisation-ia-n8n/). [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9). ## Frequently asked questions ### Is Notion Enterprise Search enough for an SME? It may be. Test the five questions, app-specific coverage, representative access, freshness and usable citations. Documentation is not the result. ### What does "connect" mean between native search and custom RAG? A bounded query, transformation, monitor or handoff. It does not own general multi-source ingestion, semantic retrieval and evaluation. ### When is custom RAG justified? When buy and connect leave required coverage, control, evaluation or action gaps, and an owner accepts operating and exit burden. ### How should we compare cost without reliable price data? Use the burden ledger and comparable supplier data. Do not turn missing prices into assumed totals. ### Does connecting a source make every answer current? No. Connector coverage and indexing differ. Test freshness per question. ## Sources and method Research was checked on 2026-07-27 using current official Notion and n8n documentation for product claims, and primary NIST, OWASP and NCSC guidance for controls. Links sit beside material claims. This is IZZY architecture guidance. No tenant, corpus, access model, workflow, answer set, comparative performance, price or client outcome was tested. --- ### MCP, skills or CLI? Put the safety boundary where the model cannot ignore it URL: https://izzy.agency/en/blog/mcp-skills-cli-ai-agent-access/ Published: 2026-08-28 Summary: Choose MCP, agent skills or CLI by what the agent can actually execute. A production database example shows where instructions end and controls begin. An AI agent needed to inspect a production database on an IZZY project. We had two practical routes. One was a skill-guided CLI workflow. It was convenient and loaded its detailed instructions only when needed. But in that environment, the shell, database program and available credentials created a broad execution path that included mutation-capable operations. The other was an MCP server exposing a small set of read methods. It added tool definitions and infrastructure, but it gave the model fewer actions to call. We chose the constrained MCP route and added agent-harness hooks that rejected mutating actions on that path. The lesson was not that MCP is secure and CLI is dangerous. It was more useful: > **For production or hard-to-recover systems, pay for a smaller mechanically available capability surface. For disposable or easily restored state, optimise for simplicity.** The safety boundary should sit where the model cannot decide to ignore it. ## Answer in 60 seconds MCP, skills and CLI are not equivalent alternatives. - **MCP** is a protocol through which a server can expose named tools with described inputs and outputs. A deliberately narrow server can make only a small set of operations available through that route. - **An agent skill** is an on-demand package of instructions and resources. It can teach an agent when and how to use a tool, but the prose in the skill is not, by itself, an enforced permission boundary. - **A CLI or direct connection** is an execution route. Its real authority depends on the programs, network, filesystem, credentials and downstream permissions reachable from the agent environment. - **Hooks and policy checks in the agent harness** can reject operations deterministically before execution. They strengthen a route, but only the route they actually intercept. Use a skill plus CLI when the state is local, disposable or reliably recoverable and the simpler workflow is worth the broader surface. Use a narrow MCP server or another constrained API when the agent touches production data, customer systems or actions whose failure is expensive. Often the best design is hybrid: **the skill explains the procedure; the constrained tool performs the action**. ## In this article 1. [MCP, skills and CLI solve different layers](#1-mcp-skills-and-cli-solve-different-layers) 2. [The production database decision](#2-the-production-database-decision) 3. [The same choice appears in every operational tool](#3-the-same-choice-appears-in-every-operational-tool) 4. [What each route is genuinely good at](#4-what-each-route-is-genuinely-good-at) 5. [Do not choose on context cost alone](#5-do-not-choose-on-context-cost-alone) 6. [Choose by consequence and recoverability](#6-choose-by-consequence-and-recoverability) 7. [Map the complete capability path](#7-map-the-complete-capability-path) 8. [Build the boundary so it fails closed](#8-build-the-boundary-so-it-fails-closed) ## 1. MCP, skills and CLI solve different layers The comparison becomes confusing when all three are treated as interchangeable tool formats. They sit at different points in the system: | Layer | What it decides | What it does not prove | |---|---|---| | **Skill** | Which procedure the agent should follow; what to inspect; which tool or script to use | That the agent cannot choose another available route or prohibited command | | **MCP tool surface** | Which named operations a particular server exposes and which inputs those operations accept | That the handlers, credentials and downstream service enforce the intended policy correctly | | **CLI or direct access** | How a program is invoked or a service is reached | That the program, shell, network path or credential is narrowly scoped | | **Agent harness hooks** | Which attempted calls or commands the configured harness allows, rejects or escalates | That another client, shell, service or credential cannot bypass that harness | | **Downstream identity and policy** | What the database, cloud account, CMS, CRM or other target ultimately authorises | That sensitive results will be minimised before they return to the model | The current [MCP tools specification](https://modelcontextprotocol.io/specification/2025-11-25/server/tools) defines discoverable tools with names, descriptions and JSON schemas. It also requires servers to validate inputs and implement access controls. The protocol gives a server a structured way to expose capabilities; it does not make every MCP server least-privileged by default. The [Agent Skills implementation guide](https://agentskills.io/client-implementation/adding-skills-support) uses progressive disclosure. The model sees skill names and descriptions at session start, loads the full `SKILL.md` when a skill is activated and loads referenced resources as needed. That can make skills economical for specialised procedures. It does not convert an instruction such as “never run `DELETE`” into a permission the runtime must obey. A skill can invoke a carefully sandboxed script. An MCP server can expose an unrestricted shell. A CLI can run with a read-only identity. The label does not settle the risk; the complete path does. ## 2. The production database decision The IZZY project did not begin as an abstract protocol evaluation. The task was to let an agent retrieve information from a production database without giving that agent a convenient way to change the data. ### Route A: skill plus CLI The skill could have described the approved workflow: 1. inspect the schema; 2. generate a read query; 3. run it with the database CLI; 4. format the result; 5. never insert, update, delete or alter anything. That was simple and context-efficient. The detailed procedure would load only when relevant. But the actual environment still gave the agent a shell path to a database program with broader capabilities. “Never mutate data” would have remained an instruction inside the same decision-making system that generated the command. If the model misunderstood the task, followed malicious retrieved content or simply chose an unexpected command, the instruction was not the final authority. ### Route B: constrained MCP tools plus harness hooks The selected route exposed concrete read methods. No generic SQL executor or shell escape hatch was presented through that MCP surface. Hooks in the agent harness inspected the configured path and rejected mutating actions before execution. Credentials were kept in `.env` rather than placed in chat, prompts or normal tool arguments and results. This reduced unnecessary model exposure, but it was not treated as magic isolation. If an agent can read arbitrary files, inspect process environments, invoke an unrestricted shell or reach verbose logs, an environment variable may still be accessible. OWASP notes that [environment variables can be visible to processes and may appear in logs or system dumps](https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html). The useful distinction looked like this: | Boundary | Skill + broad CLI route considered | Constrained MCP route selected | |---|---|---| | Procedure | Loaded on demand from the skill | Could still be supplied by a skill or system instructions | | Visible operations | Broad shell and database-program grammar | Small set of named read operations | | Mutation rule | Primarily expressed as an instruction | No mutation method on that MCP surface; harness hooks rejected mutation attempts | | Credential handling | Available to the execution environment | Kept outside prompt and normal tool payloads in `.env` | | Generic escape hatch | Present in the considered environment | Excluded from the configured MCP route | | Trade-off | Less standing tool context; simpler setup | More tool/infrastructure overhead; smaller exposed capability surface | This is one first-hand project decision, not a benchmark or proof that every MCP implementation is safer. The public account does not specify the database-native role configuration. A robust production design should still minimise the downstream identity, because MCP methods and harness hooks are defence-in-depth around that authority, not replacements for it. The choice reduced the model’s mutation capability **through the configured agent path**. It did not make the database safe against every other client, credential, bug or administrator. ## 3. The same choice appears in every operational tool The database makes the contrast easy to see, but the architecture question is generic: > Is the agent receiving a small set of purpose-built capabilities, or a general execution surface accompanied by instructions? | Utility | Narrow capability surface | Broad execution surface | |---|---|---| | **Database** | Inspect approved schema; run bounded read query; retrieve named report | Execute arbitrary SQL through a native client | | **Filesystem** | List or read files in an approved directory | Shell access able to write, move or delete across reachable paths | | **Git** | Inspect status, diff or selected history | Run arbitrary Git commands, including commit, push, reset or credentialed remote operations | | **Cloud** | Describe approved resources; read selected logs or metrics | Use a cloud CLI with create, change, delete and identity-management authority | | **CRM or CMS** | View records or prepare a draft | Update contacts, publish content, delete records or change permissions | | **Email** | Search a bounded mailbox or create a draft | Send, forward, delete or change mailbox rules | “Read” is not automatically low-risk. A read capability may expose confidential data, cross tenant boundaries, return more rows than needed or run an expensive query. But separating observation from mutation removes one class of consequence and gives the remaining risks a clearer shape. This is why [OWASP’s AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) recommends minimum tools, per-tool permission scoping and explicit authorisation for sensitive operations. It also warns against unrestricted wildcard access and relying solely on model output for authorisation decisions. ## 4. What each route is genuinely good at There is no useful universal winner. Each route buys something different. ### MCP: a declared and reusable tool contract A well-designed MCP server can provide: - a small list of named operations; - structured input schemas and, where useful, structured outputs; - one place to validate parameters and filter results; - consistent logging and error handling; - portability across clients that support the protocol; - remote authorisation flows when the deployment needs them. Those benefits make MCP attractive when several agents or clients need the same controlled capability. The server becomes a maintained product boundary rather than a prompt convention. The costs are real: - another component to build, deploy, patch and observe; - authentication, network and lifecycle complexity; - tool metadata that may consume model context, depending on the client and how it presents tools; - a new high-value service if it aggregates access; - false confidence if a “read” tool accepts arbitrary expressions, the handler is flawed or another broad tool bypasses it. MCP’s [HTTP authorisation specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) is optional and transport-specific; stdio implementations normally retrieve credentials from their environment. OAuth support is not proof that the downstream resource, action and data scope is correct. ### Skills plus CLI: low-friction procedure over existing tools A skill-guided CLI route is often the fastest way to reuse a mature toolchain. It can provide: - on-demand instructions rather than a large always-active procedure; - familiar commands engineers can reproduce outside the agent; - easy composition of local utilities; - minimal integration code; - a strong fit for temporary, local and developer-owned work. The central limitation is not that “a CLI can do everything”. Many CLIs are narrow, and every CLI only does what its program and identity allow. The problem appears when the agent also has a general shell, mutation-capable programs, broad credentials and a network path to valuable systems. In that situation, the skill describes the safe route while the environment still contains unsafe alternatives. The runtime needs another boundary: sandboxing, command allowlists, a proxy, restricted credentials, harness hooks or a purpose-built wrapper. ### Direct access: a path, not a control model “Direct” can mean a native database driver, a cloud SDK, an HTTP API or a CLI talking to the target without an MCP broker. It can still be carefully constrained. A direct database connection using a purpose-specific read identity and an isolated worker may be safer than a badly designed MCP server exposing arbitrary SQL. The architecture review should therefore ask what direct access bypasses. Does it bypass central validation, result filtering, attribution or revocation? Or does a downstream policy already enforce those controls adequately? ### Hooks: useful enforcement with a precise scope Hooks in an agent harness can reject a tool call, command or parameter pattern before execution. Unlike a prose instruction, a deterministic hook does not need the model to agree. But a hook is only as strong as its interception point. If the agent can reach the same target through another shell, plugin, network client or credential, the hook is not the boundary for that alternative path. Hooks should fail closed, produce reviewable denials and be tested against bypasses - not merely the expected command syntax. ## 5. Do not choose on context cost alone Context overhead matters. Tool names, descriptions and schemas have a cost, especially when a client exposes many of them to the model. Skills use progressive disclosure precisely to avoid loading every full procedure at startup. But neither “MCP always bloats context” nor “CLI is always cheaper” is a reliable architecture rule. The MCP protocol does not dictate one model-context strategy; implementations are free to present tools through different interface patterns. A client can filter tools, load them dynamically or route them through a smaller discovery layer. Conversely, a CLI workflow can spend substantial context and tokens reading help, recovering from errors and interpreting unstructured output. A current [controlled MCP-versus-CLI preprint](https://arxiv.org/abs/2608.08654) found that cost comparisons were unstable across seven agent scaffoldings, five models and one software task. The scaffolding had the dominant effect. Because the study covers one task and is a preprint, it should not be generalised into a universal performance ranking. Measure the route in your own harness on representative tasks: - context added before the task begins; - tool calls and retries; - task completion verified in the target system; - latency and compute cost; - failed actions and their cost; - engineering and operational overhead; - the consequence of one boundary failure. For a production database, we accepted context and infrastructure overhead because it bought a smaller exposed capability surface. For a local database that could be dropped and recreated, we would be much less interested in paying that cost. ## 6. Choose by consequence and recoverability Start with the worst action the configured path can complete, not with the interface label. | Situation | Sensible starting route | Why | |---|---|---| | Local test database with reliable fixtures | Skill plus CLI/direct access | Fast, transparent and cheap to restore if the accepted scope is actually local | | Local repository work on a recoverable branch | Skill plus CLI in an isolated workspace | Existing developer tools are useful and version control provides a recovery path | | Production investigation requiring bounded reads | Skill plus narrow MCP tools or another constrained service | Procedure stays on demand while execution exposes only necessary operations | | Production cloud diagnosis | Read-only identity plus bounded tools, network controls and output limits | “Read” still reaches sensitive configuration and logs; broad cloud authority is unnecessary | | CRM/CMS drafting | Read and draft capabilities, separate publish/update path | Preparation and external commitment should not share one undifferentiated permission | | Payment, deletion, access-right change or production write | Separate purpose-built action with downstream authorisation and independent approval | High-consequence operations need a decision outside the model and a tested recovery or stop path | The UK NCSC’s current [agentic-AI cyber-risk guidance](https://www.ncsc.gov.uk/blogs/managing-the-cyber-risk-of-agentic-ai) makes the same proportionality point at system level: more autonomy and greater potential impact require stronger controls. It explicitly advises against relying on prompting alone and asks operators to consider sandbox, network, credentials, data, monitoring and emergency shutdown together. The choice can change as the workflow matures. A CLI may be the right route while proving a task on disposable state. Once the same workflow reaches production, multiple users or customer data, the smallest viable execution surface may justify a server or wrapper. ## 7. Map the complete capability path Before choosing MCP, a skill or CLI, draw one path from user intent to business consequence. We use an **Agent Capability Path** with seven fields: | Field | Question to answer | Production-database example | |---|---|---| | **Purpose** | What bounded job should the agent complete? | Retrieve information needed for an investigation | | **Instruction layer** | Which skill, prompt or operating procedure guides the model? | Read workflow and query guidance | | **Execution surface** | Which exact methods, commands, scripts or network routes can it invoke? | Named read tools; no generic SQL or shell route through the configured MCP surface | | **Credential boundary** | Where does the secret live, and which processes or services can use it? | Outside prompt and normal tool payloads; available to the configured execution component | | **Enforcement points** | Which independent checks can reject the action? | Tool handler validation plus agent-harness hooks; downstream permissions should add another layer | | **Result boundary** | Which rows, fields, files or logs can return to the model and user? | Bounded read result with sensitive fields and volume considered explicitly | | **Stop and recovery** | How is the agent halted, access revoked and state checked? | Disable the route or credential, inspect logs and verify database state | The artefact is small enough for an architecture review and concrete enough to test. It also exposes false boundaries. If “do not delete” appears only under Instruction layer, the deletion rule is advisory. If the credential can be read by the same unrestricted shell the hook is meant to control, the credential boundary is weak. If a read tool can return an entire customer table, the result boundary remains broad even though integrity is protected. For continuing access, connect this decision to the wider [AI-agent permission lifecycle](https://izzy.agency/en/blog/ai-agent-permissions-access-control/). For consequential actions, use the [capability-contract and product-control model](https://izzy.agency/en/blog/ai-agent-security-product-controls/). This article owns the path-selection decision; those guides own the ongoing access and autonomy controls around it. If the agent needs production-shaped data rather than live production access, [snapshots, replicas and controlled retrieval](https://izzy.agency/en/blog/ai-coding-agents-production-context/) can remove the question altogether. ## 8. Build the boundary so it fails closed Once the path is visible, implement the smallest route that can complete the task. ### 1. Separate allowed, prohibited and approval-gated actions Do not write “database access” or “manage the CMS”. Name operations: inspect these schemas, read these records, prepare this draft, publish this exact revision, or delete nothing. ### 2. Remove generic escape hatches from high-consequence paths A server with three read methods plus `execute_any_command` is not a narrow surface. Look for shell tools, arbitrary code execution, generic SQL, unrestricted HTTP clients, filesystem traversal and inherited credentials that recreate the broader path. ### 3. Restrict the downstream identity too The target system should reject authority the agent does not need. A narrow interface around a broad administrative credential creates avoidable single-layer dependence. ### 4. Keep secrets out of prompts, arguments, results and ordinary logs Use the secret mechanism appropriate to the deployment, minimise which process can retrieve it, rotate it and test revocation. `.env` may be a practical local configuration method; it is not a substitute for process isolation or a managed secret boundary when consequences justify one. ### 5. Add deterministic validation outside the model Validate tool names, parameters, resource scope, operation class, row or object limits and approval state. Harness hooks can reject disallowed calls; server handlers, proxies and downstream systems should enforce the rules closest to the consequence. ### 6. Constrain results as well as actions Limit fields, records, time range, response size and sensitive values. Treat tool results and retrieved content as untrusted input before they re-enter model context. ### 7. Preserve attributable evidence Record the requesting identity, agent version, attempted action, target, policy decision and result without recording credentials or unnecessary sensitive data. A model transcript alone is not a complete audit trail. ### 8. Test denials and bypasses Attempt prohibited operations through alternate syntax and alternate available tools. Test malformed parameters, encoded commands, indirect requests, large outputs, retries, timeouts and a missing policy service. The safe failure state is a refusal, not silent fallback to the broad route. ### 9. Test the stop path Know which server, token, process, network route, queued job and session must stop. Revocation is incomplete if the agent can continue through an inherited credential or another client. ## Conclusion: the interface is not the safety boundary MCP can expose a narrow capability surface. It can also expose an unrestricted one. A skill can make a complex CLI workflow efficient and repeatable. It cannot, through prose alone, remove commands and credentials from the agent’s environment. Choose the architecture from the consequence backwards. If state is disposable and recovery is trivial, skill plus CLI may be exactly right. If the agent reaches production or hard-to-recover systems, reduce what is mechanically callable, remove alternate routes and enforce the same decision at more than one layer. The practical hybrid is often the strongest: **use a skill to load the right procedure when needed, and a narrow MCP tool or service to execute only the capabilities that procedure is allowed to request**. ## Map one agent access path before you expose it Bring one workflow, the utility it needs and the highest-impact action currently reachable. In a 30-minute scoping call, IZZY can map the instruction layer, execution surface, credential boundary, enforcement points, returned data and stop path, then compare the smallest viable implementation routes. This is the kind of decision IZZY makes inside its [AI integration engagements](https://izzy.agency/en/services/ai-integration/), before any agent is connected to a real system. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) ## Frequently asked questions ### Is MCP safer than a CLI? Not by default. A narrow MCP server can expose fewer operations than a broad shell route, but an MCP server can also expose arbitrary commands. Compare the actual tools, handlers, credentials, alternate paths and downstream permissions. ### Can a skill enforce read-only behaviour? A skill can instruct the agent and invoke a restricted script, but its prose is not an enforced permission. Read-only behaviour needs a runtime, proxy, handler, hook, sandbox or downstream identity that rejects mutation independently of the model. ### Does keeping credentials in `.env` hide them from the model? It can keep credentials out of prompts and ordinary tool payloads when configured carefully. It does not guarantee that an agent with filesystem, process, shell or logging access cannot reach them. ### If an MCP server exposes no write tool, is mutation impossible? Only through that surface, and only if no generic escape hatch or alternate route recreates write access. Handler validation, harness policy and downstream permissions should preserve the same restriction. ### Does MCP always use more context than skills or CLI? No universal rule has been verified. MCP tool metadata can add context, while skills progressively load full instructions on demand. Actual cost depends on the client, tool count, task, output format, retries and agent scaffolding. Measure the configured system. ### What should we use for a local database? If it is genuinely local, disposable and easily restored, a skill plus CLI or direct driver is often the simpler route. Make sure the credentials and network path cannot silently reach production. ### Is read-only production access safe? It reduces mutation risk, not confidentiality, tenant-isolation, expensive-query or result-exposure risk. Scope resources and outputs, set operational limits and log access proportionately. ## Sources, method and limitations Primary technical and official sources checked on 28 August 2026: - [Model Context Protocol: tools](https://modelcontextprotocol.io/specification/2025-11-25/server/tools) - [Model Context Protocol: authorisation](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) - [Agent Skills: adding skills support](https://agentskills.io/client-implementation/adding-skills-support) - [Agent Skills specification](https://agentskills.io/specification) - [OWASP AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) - [OWASP Secrets Management Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html) - [UK NCSC: managing the cyber risk of agentic AI](https://www.ncsc.gov.uk/blogs/managing-the-cyber-risk-of-agentic-ai) - [Alier Forment et al.: controlled MCP/CLI comparison, preprint](https://arxiv.org/abs/2608.08654) The production-database example is a de-identified first-hand IZZY project account. No client outcome, incident or comparative performance result is claimed. The public account does not specify the database-native role configuration. Current practitioner comparisons were used only to test crowding and terminology, not as neutral proof of security or demand. No search-volume estimate was available, and no universal token, latency or success-rate conclusion is made. This article provides general product and engineering guidance. It is not a penetration test, database-security review, compliance assessment or legal opinion. Production credentials, personal or confidential data, write access and high-impact actions require review against the actual architecture by appropriately qualified owners. --- ### One brain does not mean one database: what should stay in Drive, Notion, Slack, GitHub and Linear? URL: https://izzy.agency/en/blog/systems-of-record-drive-notion-slack-github-linear/ Published: 2026-08-22 Summary: Decide what should stay, be indexed, linked, copied or excluded across Drive, Notion, Slack, GitHub and Linear - without duplicate truths. The same launch date exists in a Drive brief, a Notion checklist, a Slack thread and a Linear milestone. Search can find all four. It cannot tell your team which one wins unless you define that rule. Do not centralise company data by copying everything into one app. Assign one authoritative system per record class and business process. Keep the original there; index or link it elsewhere. Copy only for a declared reason, with provenance and an expiry or reconciliation rule. This article provides a conditional decision method, not a universal tool hierarchy or a client migration result. Your processes, permissions, product plans and record obligations need organisation-specific review. ## The answer in 60 seconds - Authority is a governance role, not a software feature. - Classify each record: keep, index, link, copy or exclude. - Allow one maintained authority per record class or process. - Give every integration stable identity, access and freshness rules. - Write changes through the authority, then record the result. - Stop new duplicates before migrating legacy content. ## In this article 1. [Authority is a role, not an app feature](#1-authority-is-a-role-not-an-app-feature) 2. [Choose one of five placement actions](#2-choose-one-of-five-placement-actions) 3. [Use a conditional tool-ownership matrix](#3-use-a-conditional-tool-ownership-matrix) 4. [Give authority an operational contract](#4-give-authority-an-operational-contract) 5. [Place one product launch across five systems](#5-place-one-product-launch-across-five-systems) 6. [Repair an ownership conflict](#6-repair-an-ownership-conflict) 7. [Migrate without a big-bang move](#7-migrate-without-a-big-bang-move) 8. [Run the five-question governance test](#8-run-the-five-question-governance-test) ## 1. Authority is a role, not an app feature [IBM describes](https://www.ibm.com/think/topics/system-of-record-vs-source-of-truth) a system of record as the authoritative source for a domain or process, while a source of truth may harmonise data from several systems. The distinction is functional. Buying a search tool or creating a master Notion space does not assign operational authority. For each record class, answer: where is it created, who maintains it and which version resolves a conflict? “We usually look in Slack first” is a habit, not a rule. “The release manager maintains the approved launch date in the launch record” is a rule. [NIST's enterprise-architecture definition](https://csrc.nist.gov/glossary/term/enterprise_architecture) includes how systems are configured, integrated, operated and related to security and mission. For an SME, that can begin with a one-page record map. You do not need an enterprise programme to stop two tools owning the same field. ## 2. Choose one of five placement actions Give every in-scope record one primary action: | Action | Meaning | Example | |---|---|---| | **Keep** | Leave the maintained authority in its native system | Signed brief in Drive | | **Index** | Make permitted content searchable without independent authority | Brief text in a retrieval index | | **Link** | Point to the current native record | Notion launch page links the Linear project | | **Copy** | Store a justified snapshot | Approved terms attached to a delivery record | | **Exclude** | Deliberately keep it outside the memory | Secrets or unsupported private content | Link when people can reach the current source. Index when cross-system retrieval needs content. A copy is the expensive option because it can drift. Require `source_id`, source URL/version, copied time, purpose, owner and expiry or reconciliation rule. “Copy for convenience” is not enough. If nobody can say which copy may be edited or how it expires, you have created another authority contest. ## 3. Use a conditional tool-ownership matrix | Tool | Often owns when… | Prefer index/link when… | Do not assume | |---|---|---|---| | **Drive** | A governed file is reviewed and maintained there | Another system needs discovery or workflow context | Every file is approved | | **Notion** | A structured operating record is actively maintained there | It is aggregating records owned elsewhere | Search makes Notion authoritative | | **Slack** | The organisation deliberately governs a conversation record | A thread is evidence or context for a maintained decision | Chat is always temporary - or durable | | **GitHub** | Work is repository-scoped: code, PR, issue, release evidence | Product/delivery approval is cross-functional | All engineering work belongs there | | **Linear** | The team runs delivery ownership and state there | The record is code-specific or documentary | Every organisation needs Linear as delivery truth | Google says [shared-drive files belong to the team](https://support.google.com/drive/answer/7286514?hl=en), subject to edition and policy. Notion can [query structured data sources](https://developers.notion.com/reference/query-a-data-source), but API search is [not exhaustive](https://developers.notion.com/reference/search-optimizations-and-limitations). GitHub supports [issues](https://docs.github.com/en/issues/tracking-your-work-with-issues/using-issues); Linear issues [belong to one team](https://linear.app/docs/creating-issues) and require title and status. Capabilities make a role possible; your process assigns it. Slack deserves precision. It is searchable, retention is configurable and messages can be edited or deleted. So “Slack is temporary” is not a reliable architecture rule. Decide whether a thread is governed evidence or whether a decision must be promoted to a maintained record, then document that promotion. ## 4. Give authority an operational contract For each authoritative record class, capture: - native and cross-system identity; - creator, maintainer and conflict owner; - permitted readers and writers; - fields it governs - and fields it does not; - source update time and required freshness; - downstream indexes, links and copies; - write-back and retirement rules. Use stable native IDs, not editable titles. Google Drive exposes file IDs, metadata and permissions. In Linear, [moving an issue between teams](https://linear.app/docs/editing-issues) can change its human-facing identifier and URL while old links redirect; integrations need redirect-aware or immutable mappings. Freshness is a contract, not a webhook checkbox. Drive notifications tell a consumer to read the change feed. GitHub and Linear publish webhooks for supported events. Delivery, signature validation, renewal, retry and downstream failure still need handling. Store `source_updated_at`, `event_received_at`, `downstream_updated_at` and `sync_state`. If a copy or index is behind, show that state rather than serving it as current. Write through the authority whenever possible. After an approved write, record the native ID, URL, result and timestamp in the coordinating record - the same discipline that keeps [meeting notes from becoming duplicate tickets](https://izzy.agency/en/blog/meeting-notes-to-actions-n8n/). If a second system needs an editable field, specify direction: `Linear → Notion`, not “two-way sync”. Two-way editing without field ownership is conflict automation. ## 5. Place one product launch across five systems Consider the Atlas launch: | Record | Governing system | Other placement | |---|---|---| | Approved brief and acceptance criteria | Drive | Indexed; linked from Notion | | Launch control page and decision log | Notion | Links to all native records | | Pricing discussion | Slack | Indexed as context; approved outcome promoted | | Release branch, PRs and security fix | GitHub | Status linked into launch control | | Delivery work, owners and milestone | Linear | Summary linked into launch control | There is no universal “launch record”. There are accountable records for different questions. “What was approved?” resolves to Drive/Notion according to the decision policy. “Is the fix merged?” resolves to GitHub. “Who owns the remaining work?” resolves to Linear. The [searchable operating memory](https://izzy.agency/en/blog/searchable-operating-memory-company/) can retrieve across all five. It must preserve these roles. The Notion launch page may coordinate the view, but it should not silently copy five editable launch dates. ## 6. Repair an ownership conflict Suppose Drive says launch on 14 October, Notion says 21 October and Linear says 18 October. 1. Freeze automatic propagation for that field. 2. Gather each value, source ID, editor and update time. 3. Apply the documented conflict rule - or escalate to the named owner. 4. Update the authoritative record through its normal approval path. 5. Replace downstream editable copies with links or controlled projections. 6. Write back the resolution and test the next update. Do not resolve the conflict by choosing the newest timestamp unless that is the agreed rule. A recent Slack message can still be an unapproved proposal. Make `conflict`, `partial_write`, `denied` and `replay_required` visible workflow states. Error routing and duplicate comparison can support the repair, but they do not decide which business value wins. ## 7. Migrate without a big-bang move Start with record classes, not folders. Inventory five to ten high-consequence examples: launch date, customer commitment, price, contract version, release approval, incident owner or delivery deadline. Record all current locations and editors. Then stop creating new duplicates. Change templates, automation and team instructions so new records use the chosen authority. A migration that moves history while workflows keep producing parallel records cannot converge. Move or redirect one class at a time: 1. Declare the target authority and transition owner. 2. Add native identity and backlinks. 3. Redirect integrations and search. 4. Reconcile active records. 5. Mark old copies read-only, archived or superseded. 6. Test permission, freshness, conflict and rollback. Do not delete historical sources merely to make the architecture diagram tidy. Retention and evidence needs require separate review. ## 8. Run the five-question governance test Before a record class enters the operating memory, ask: 1. Where is the record created? 2. Who is accountable for maintaining it? 3. Which version wins a conflict? 4. How do permitted changes reach indexes, links or copies? 5. How is the record retired or superseded? If the team cannot answer one question, mark the class `unresolved` and keep automation read-only. Search may still expose evidence, but the answer should say that authority could not be verified. Review the map when a tool, team, workflow, permission model or record obligation changes. Governance that exists only in a launch workshop will drift as quickly as the copies it was meant to control. ## Conclusion: one operating memory, many accountable systems One company brain does not require one database. It requires clear record roles. Keep authority close to the people and workflow that maintain it. Index and link for discovery. Copy with provenance and expiry. Exclude deliberately. Then make identity, access, freshness, conflict and write-back observable. ## Map five representative records > Bring five records that currently appear in more than one tool, their locations and one unresolved conflict. IZZY can map authority, placement and workflow boundaries before recommending migration or automation. The useful result may be fewer integrations, not more. This is exactly the scope of our [n8n AI Automation service](https://izzy.agency/en/services/automatisation-ia-n8n/). [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9). ## Frequently asked questions ### What is the difference between a system of record and a source of truth? A system of record maintains authoritative data for a domain or process. A source-of-truth view may reconcile data across several systems. ### Should Notion be the source of truth for a company? Only for record classes the team deliberately creates and maintains there. It can index or coordinate records owned elsewhere. ### Is Slack too temporary to contain business records? Not inherently. Retention and editing are configurable. Decide whether Slack is governed evidence or whether decisions must be promoted. ### Is copying data between systems always wrong? No. A copy can be justified, but needs provenance, purpose, owner, access, freshness and expiry or reconciliation. ### How do we choose the first system-of-record decision? Choose a high-consequence record that already conflicts, name its maintainer and governing workflow, then stop new duplicates. ## Sources and method Research was checked on 2026-07-27 using current official Google, Notion, Slack, GitHub, Linear and n8n documentation plus IBM and NIST architecture definitions. Links sit beside material claims. This is IZZY architecture guidance. No client data model, tenant, migration, sync or permissions were tested. Completeness, latency, security, outcomes, ROI and legal compliance remain unverified. --- ### When generic design stops being useful URL: https://izzy.agency/en/blog/when-generic-design-stops-being-useful/ Published: 2026-08-22 Summary: Generic design is useful for proving a concept. Learn when to keep the scaffold, personalise the system or give a validated product a direction of its own. A generic interface can be exactly right at the beginning. Fast design has one job: make the idea testable. It lets a team check whether a feature makes sense, whether the main interaction works and whether the idea deserves more investment. At that stage, a mature visual identity can wait. The problem begins when that temporary design outlives the question it was built to answer. Once the idea, feature or business has been validated, the interface has a different job. It must help unfamiliar people understand what to do. It must behave coherently on a small screen, recover from mistakes, support different access needs and express a direction appropriate to this product - not simply resemble a plausible product in the same category. AI makes the first stage faster. That is useful. It also makes the transition between a functional scaffold and a considered product easier to miss. ## Answer in 60 seconds Generic design stops being useful when it begins to hide product-specific decisions rather than helping you test them. - Keep the generic scaffold while you are still proving the problem, feature or core interaction. - A fast prototype still needs clear logic, understandable information architecture, accessibility, clean execution and small-screen behaviour. Otherwise you may test defects in the prototype rather than the idea. - When the concept works, decide whether to personalise the existing system or establish a new direction. A full redesign is not the automatic answer. - Put usability before visual distinctiveness. Fix broken logic, confusing structure, errors and access barriers first. - Keep AI in the workflow. Use it for speed and breadth, but keep the design decision accountable to people who understand the user, product and context. This is not an argument against AI or reusable patterns. It is a method for knowing when their defaults are no longer enough. ## In this guide 1. [Generic design is a stage, not a failure](#1-generic-design-is-a-stage-not-a-failure) 2. [Start with what is wrong before asking what is distinctive](#2-start-with-what-is-wrong-before-asking-what-is-distinctive) 3. [Make the generic-to-specific design decision](#3-make-the-generic-to-specific-design-decision) 4. [Keep AI in the workflow, and keep judgement accountable](#4-keep-ai-in-the-workflow---and-keep-judgement-accountable) 5. [Use one page to decide what the product needs next](#5-use-one-page-to-decide-what-the-product-needs-next) ## 1. Generic design is a stage, not a failure Fast design is valuable because it makes an assumption testable. Use the simplest prototype that can answer the current question: a sketch for hierarchy, a clickable flow for sequence, or working code when the real interaction matters. The point is to learn before committing to a full build. An experiment and a live product need different evidence. If the immediate question is “Can somebody understand and use this feature?”, a restrained component library and a fast prototype may be exactly right. They reduce the temptation to debate brand expression before the team knows whether the feature should exist. Generic does not have to mean careless. A useful prototype should still be contextualised enough to produce meaningful feedback. Labels should make sense to the intended user. The main action should follow a coherent sequence. Obvious errors, broken states and strange interaction logic should not distort the test. People should be able to try it on the screens they are likely to use. At IZZY, we start with the small screen. It forces the hierarchy and action priorities into the open: what must remain, what can move and what the user needs first. That does not mean every prototype needs a finished responsive system. It means a wide desktop preview is not proof that the interaction works in context. Accessibility belongs in this baseline too. [WCAG 2.2](https://www.w3.org/TR/WCAG22/) applies across desktop and mobile web content and treats responsive variations as part of the page for conformance. A prototype review is not an accessibility certification. It can, however, catch choices that would otherwise invalidate learning or become expensive to unwind later. The aim is not to make the experiment look finished. It is to make the experiment honest. If the open question is whether the business can sell, deliver and support the idea - not how its design should develop - use our separate guide to [turning a prototype into a commercial product](https://izzy.agency/en/blog/prototype-to-commercial-product/). ## 2. Start with what is wrong before asking what is distinctive When we review a fast-built interface, we do not begin by asking whether it has enough personality. We look for mistakes. That includes bugs, incoherent behaviour, unclear logic, weak information architecture, inaccessible controls, missing states and visual choices that interfere with the task. These are not secondary details. They determine whether the interface can be understood and used. A polished screen is not evidence of a usable product. The result has to be judged with particular users, goals and conditions in mind. Our review order is deliberately practical: | Layer | First question | Typical evidence | |---|---|---| | Task and logic | Can the intended user complete the important action, and does the result make sense? | Observed attempts, task outcomes, bugs and contradictions | | Information architecture | Can people find and understand what they need without a guided tour? | Navigation paths, labels, hierarchy and points of hesitation | | Accessibility and screen behaviour | Can people perceive, operate and understand the interface across relevant access needs and screen sizes? | Keyboard and assistive-technology checks, responsive states and user testing | | Errors and edge states | What happens when information is missing, invalid, delayed or unavailable? | Empty, loading, error, partial, success and return states | | Visual direction | Does the experience feel appropriate to the product, audience and brand? | Direction criteria, brand-system decisions, component use and comprehension testing | The order is not a claim that visual design is unimportant. It protects visual work from being used to disguise a more fundamental problem. If the task model is wrong, a new colour palette will not fix it. If the information architecture is unclear, a more expressive typeface may make the confusion more attractive. If a key control cannot be used with a keyboard, distinctiveness is not the next priority. Usability comes first because it is the condition under which the rest of the design can matter. ## 3. Make the generic-to-specific design decision Once the idea has produced useful evidence, teams often jump to one of two extremes: keep the generated or template interface untouched, or commission a total redesign. We use three routes. ### Route 1: keep the scaffold Keep it when the team is still testing the problem, the audience is controlled and the current interface is not distorting the result. The design still needs a baseline of clarity, accessibility and responsive behaviour appropriate to the test. It does not need a mature brand system merely to demonstrate that the core action is plausible. The next investment should buy evidence: another prototype, a corrected flow or research with the intended user. ### Route 2: personalise the system Choose this route when the core interaction works and the structure is fundamentally sound, but the experience still feels borrowed, incomplete or inconsistent. Personalisation can include: - defining the brand identity, or translating an existing one into the product - from typography, colour, iconography and imagery to tone of voice, layout principles and motion behaviour; - rewriting labels and guidance in the customer’s real language; - clarifying hierarchy and information density; - designing purposeful microinteractions and animations that provide feedback, clarify state changes, show progress and help people stay oriented; - making components, states and interaction patterns consistent; - improving mobile behaviour and accessibility; - documenting the resulting decisions in reusable design tokens, components and guidance. This is often focused product-design work rather than a rebuild. Keep the parts that have earned their place. Change the parts that still behave like placeholders. ### Route 3: establish a new direction Choose this route when the current design encodes the wrong assumptions. That may be the case when the navigation follows the tool rather than the user’s mental model, the visual tone signals the wrong category or level of trust, the interface cannot accommodate real product states, or the product has moved to a different audience and use case since the prototype was made. Here, styling the existing scaffold can preserve the wrong structure. The team needs to return to the problem, define a direction and decide which patterns should survive. This is the important boundary: a validated concept can justify more design investment, but validation does not automatically justify more decoration. The investment should address the next unproven decision. ## 4. Keep AI in the workflow - and keep judgement accountable AI is too useful to ban. It can be valuable across many parts of design and product work. It can make an idea tangible sooner, generate alternatives, scaffold a flow, explore content variations, enumerate edge cases and help a team compare directions. Used well, it creates more material to evaluate and leaves more time for decisions that require context. But speed does not remove accountability. | AI can accelerate | The product team still owns | |---|---| | Prototype screens and interaction alternatives | The user problem and the hypothesis being tested | | Variations in copy, layout or visual treatment | The criteria for accepting or rejecting a direction | | Lists of possible states, errors and edge cases | Verification with real product rules, users and conditions | | Initial accessibility and consistency checks | Accessibility scope, human evaluation and remediation decisions | | Production handoff drafts and component documentation | The final system, trade-offs and release decision | The human contribution is not a layer of visual polish added after the machine has done the real work. It is the framing, selection, correction and validation that make the output appropriate. AI can produce many competent versions. Product design decides which question those versions are answering - and whether any of them should ship. ## 5. Use one page to decide what the product needs next Before commissioning more screens, write down six things: 1. **The current stage:** idea, functional proof, tested pilot or customer-facing product. 2. **What has actually been proved:** technical possibility, task comprehension, repeated use, willingness to pay or something else. 3. **The important user task:** the one path the design must make clearer or more reliable. 4. **The first observed break:** bug, logic, information architecture, accessibility, screen behaviour, error handling or visual fit. 5. **The direction criteria:** what the experience should communicate and which constraints it must respect. 6. **The next decision:** keep the scaffold, personalise the system or establish a new direction. This short record prevents three common mistakes: treating a polished demo as finished, spending on brand expression before the idea is credible, and preserving a generic interface after the product has earned the need for something more specific. It also creates a better brief. Instead of asking a designer to “make it less AI” or “add personality”, the team can explain what has been validated, what users are trying to do, what currently fails and what the new direction must accomplish. ## Conclusion: the right design investment changes with the evidence The purpose of a prototype is not to look original. It is to help a team learn. Use fast design to test the functional idea. Keep that early experience clean, coherent, accessible enough for the research and designed for the screens people will actually use. Then, if the idea proves useful, change the design question. First fix the mistakes, logic, information architecture, accessibility and missing states. After that, decide whether the current system needs focused personalisation or a direction of its own. Judge the result by usefulness and applicability - not visual novelty alone. AI remains part of the process. Usability remains the priority. Personalisation becomes valuable when it helps a product fit its task, its audience and its identity more precisely. ## Choose the design investment the evidence supports You do not need to arrive with a finished prototype. If you want to test an idea without investing in a custom design system too early, IZZY can help define and design the smallest useful proof of concept: the critical flow, necessary states, clear content, accessible interaction and reliable small-screen behaviour. The aim is to create a credible test - not to make an unproven product look finished. If you already have a fast-built product and the idea has been proved, we can help determine whether the next step is a focused UX correction, system personalisation or a new product direction. This is exactly the scope of our [Product Design service](https://izzy.agency/en/services/product-design/). Bring the idea or prototype, the important user task and any evidence you already have. We will scope the right level of design investment from the question you need to answer - not assume a custom redesign from the start. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9). ## Frequently asked questions ### Does a proof of concept need custom branding? Not necessarily. If the immediate goal is to test whether the feature or interaction works, a restrained generic system may be the right choice. It should still be clear, coherent and appropriate enough that poor execution does not distort the test. ### Is generic UI bad design? No. Familiar patterns can reduce learning and help a team move quickly. Generic UI becomes a problem when it no longer reflects the user’s task, the product’s real states or the organisation’s direction. ### Should usability or visual identity come first? Usability comes first in the review order. Fix broken logic, confusing structure, access barriers and missing states before using visual identity to differentiate the product. The two should eventually work together. ### Can AI produce a distinctive design? It can contribute useful and original options. Distinctiveness is not guaranteed by the tool or by avoiding the tool. It comes from specific inputs, clear constraints, informed selection, refinement and validation in the product’s context. ## Sources and method This article draws on IZZY’s experience working with real client products and on external material checked on 19 August 2026. Client identities, project details and case material are not disclosed because of confidentiality and NDAs. The generic-to-specific decision and review order are IZZY practice guidance, not external standards or guarantees. No representative market-demand claim is made. - [W3C - Web Content Accessibility Guidelines (WCAG) 2.2](https://www.w3.org/TR/WCAG22/) --- ### Can AI chats find, compare and buy your products? A practical ecommerce test URL: https://izzy.agency/en/blog/ai-shopping-agent-catalog-readiness/ Published: 2026-08-20 Summary: For ecommerce leaders: test whether product data, feeds, AI retrieval or checkout handoff is blocking AI shopping before funding the wrong build. If you are the Head of Digital Commerce, the pressure is no longer abstract. Everyone is calling shopping through AI chats “the future”, and leadership wants to know whether customers will be able to find, compare and buy your products there. Yet the present is less tidy: ask an AI assistant to find a product with several constraints, compare two variants or confirm a current price, and the answer can still be incomplete or wrong. Does that mean ecommerce sites are badly built? Sometimes the merchant's product data is part of the problem. But a poor AI result can also come from missing platform access, catalogue normalisation, retrieval or the agent's reasoning. The answer is not visible from the chat alone. There is another source of confusion. “The AI can buy” may describe a useful handoff to the checkout you already operate. It may also describe an agent completing the transaction itself - a capability that some platforms document but restrict by access or eligibility. Catalogue or website work cannot unlock that access on its own. This guide is for established European retailers and manufacturers whose product truth moves through several systems, variants or markets. Before funding a feed, integration or “agentic commerce” build, test one real category from source data to purchase path. The result should tell you what is broken, who controls it and whether anything new needs to be built. If your catalogue is simple and maintained in one commerce platform, start with that platform's current feed and checkout guidance; a cross-system review may be unnecessary. ## Answer in 60 seconds For most merchants today, the realistic purchase path is a handoff to the checkout they already operate. Some platforms document agent-completed checkout for eligible integrations, but a merchant cannot activate that capability through catalogue or website work alone. Treat it as unavailable until the chosen platform confirms access. Begin with one commercially important product category and a risk-based set of real products: a normal bestseller, a complex variant, a current price or promotion change, an availability change, a delivery or market restriction and a product whose specifications matter in comparison. For each case, compare the authoritative source, product page, structured data, current merchant feed and target AI channel. Check the exact product and variant, title, attributes, price, currency, availability, images, delivery, returns and eligibility. Then run five buyer tasks: exact lookup, constrained discovery, comparison, price or stock verification, and the available purchase handoff. Classify every result as a pass, a merchant-controlled defect, a channel-side issue observed, or unknown. If the platform is partner-gated or the product was never ingested, do not call the result a website failure. Keep the existing merchant checkout when a reliable handoff solves the journey. Additional integration work begins when the handoff must preserve a selected variant or quantity or create a prefilled cart. Only scope agent-completed checkout after the platform confirms that the merchant and integration are eligible. The decision should be specific: repair product truth, repair public access, operate a reliable feed, improve the checkout handoff, scope a checkout integration after eligibility is confirmed - or monitor the channel without building yet. ## In this article 1. [Two practical paths and one gated capability](#1-two-practical-paths-and-one-gated-capability) 2. [Why AI product discovery and comparison still fail](#2-why-ai-product-discovery-and-comparison-still-fail) 3. [How to test one real product category](#3-how-to-test-one-real-product-category) 4. [How to turn the findings into a decision](#4-how-to-turn-the-findings-into-a-decision) 5. [What to fix now and what not to build yet](#5-what-to-fix-now-and-what-not-to-build-yet) 6. [When an external review earns its place](#6-when-an-external-review-earns-its-place) ## 1. Two practical paths and one gated capability Use the customer's experience to define the project. Do not group discovery, checkout handoff and agent-completed purchase under one “agentic commerce” label. | Capability | What the customer experiences | What the merchant needs | |---|---|---| | Find and compare | The AI identifies suitable products and explains relevant differences | Reachable or ingested product data that is current and comparison-ready | | Merchant checkout handoff | The AI opens the correct product, basket or checkout on the merchant's site | A stable destination that preserves the selected product, variant and, where supported, quantity | | Gated agent-completed checkout | An eligible agent confirms totals and completes an authorised order | Confirmed platform access plus authentication, cart state, live pricing, payment, confirmation, order and exception controls | For most merchants, the first two paths are the practical ones today. The checkout handoff can reuse infrastructure the merchant already operates, although preserving variant, quantity or cart context may require an integration. The third path exists in documented platform flows, but access is gated; it is not a capability a merchant can assume or switch on through catalogue work. This article covers external AI shopping channels. If the assistant sits inside your own store, the decision is different; read our guide to [on-site AI shopping assistants](https://izzy.agency/en/blog/ai-shopping-assistant-ecommerce-site/). ### What is documented today? Platform access changes quickly, so the following status is dated **20 August 2026**. | Platform path | Current documented position | What it does not prove | |---|---|---| | ChatGPT product feeds | [OpenAI product-feed onboarding](https://developers.openai.com/commerce/guides/get-started) is available to approved partners. Its [product schema](https://developers.openai.com/commerce/specs/file-upload/products) separates eligibility for product search from eligibility for checkout. | Approval, surfacing, ranking or checkout for a particular merchant or product | | ChatGPT merchant checkout | OpenAI currently describes [merchant-hosted external checkout](https://developers.openai.com/plugins/build/monetization) as the recommended generally available route for most plugin developers. | That an AI can create a cart or complete payment for every merchant | | ChatGPT embedded payment | OpenAI currently describes its embedded payment sheet as a private beta for select marketplace partners. | Near-term access for an unapproved merchant | | Shopify agentic commerce | Shopify documents [merchant checkout handoff for general access and direct completion for eligible agents](https://shopify.dev/docs/agents/carts-and-checkout/checkout-mcp). | Universal availability across agents, merchants, markets or products | Published infrastructure is real. Universal merchant access is not. A crawlable page or accepted feed can support discovery without enabling agent-completed checkout. ## 2. Why AI product discovery and comparison still fail The difficulty is not imaginary. Recent benchmarks show that shopping agents still struggle with grounded product tasks: - [ShoppingBench](https://ojs.aaai.org/index.php/AAAI/article/view/40640), published in the AAAI 2026 proceedings, tested a controlled environment with more than 2.5 million products. GPT-4.1 achieved an absolute success rate below 50% on its tasks. - [EComAgentBench](https://arxiv.org/abs/2606.17698), a preprint with 662 tasks, distributed requirements across the shopper's query, profile and a clarification step. The strongest of seven evaluated models reached 57.1% overall accuracy. - [ShoppingComp](https://arxiv.org/abs/2511.22978), also a preprint, evaluated 145 instances and 558 scenarios built around real products. The best model it reports, GPT-5.2, scored 17.76%. Its authors attribute the failures to grounding in open-world product data, verifying multi-constraint requirements, reasoning over noisy or conflicting evidence and risk-aware decisions. The percentages are not directly comparable because the benchmarks use different tasks and metrics. They are not estimates of how often an ordinary shopper receives a bad answer. They establish a narrower point: finding, filtering and verifying products remains difficult even when an agent sounds confident. A failed answer can enter the chain in several places: | Where the problem sits | What can go wrong | Main control | |---|---|---| | Merchant product truth | Variant IDs change; attributes are missing; page, feed and checkout disagree on price or stock | Merchant | | Public access | Product pages or images are blocked, unstable, duplicated or dependent on fragile rendering | Merchant, hosting and security controls | | Channel data | The feed is missing, stale, rejected, out of scope for the market or not available to that merchant | Merchant and platform | | Cross-merchant normalisation | Different catalogues model the same product, bundle or variant differently | Mostly platform | | Shopper intent | Important constraints remain implicit or change during the conversation | Shopper and agent | | Retrieval and reasoning | The agent finds the wrong candidate, loses a constraint or makes an unsupported comparison | Mostly agent and platform | Two merchants can also structure the same legitimate product differently without either storefront being “bad”. Shopify describes this catalogue heterogeneity as a product-identity and clustering problem in its [Catalog API engineering work](https://shopify.engineering/catalog-clustering). The practical question is therefore not “Is our ecommerce good?” It is: **where does one tested product journey stop matching the evidence?** ## 3. How to test one real product category This is a diagnostic test, not a statistical survey of the whole catalogue. Its purpose is to cover the product rules most likely to expose a broken identity, mapping, update or handoff. ### Step 1: fix the test boundary Choose: - one commercially important category; - one market, currency and delivery destination; - one target AI channel, including its model or mode where visible; - one test date and time; - the merchant systems that should hold the authoritative values. Do not begin with the whole catalogue or several AI platforms. A smaller fixed boundary makes it possible to distinguish a product-data problem from ordinary variation between channels. ### Step 2: build a risk-based product set Do not select only clean, simple products. Include at least one case from every applicable group below. | Test case | Selection rule | Failure it can expose | |---|---|---| | Normal product | A popular or commercially important product with an ordinary buying path | Baseline identity, page, feed and discovery failure | | Variant product | A product where size, colour, material, capacity or another option changes the offer | Parent/variant confusion, wrong image, price or stock | | Price event | A live promotion or a product whose price changed recently | Stale feed, expired sale or currency mismatch | | Availability event | A low-stock, out-of-stock, back-order or recently restocked item | Update delay and unsupported availability claims | | Market or policy case | A product affected by delivery limits, returns rules or market eligibility | A recommendation the customer cannot actually buy | | Comparison case | A product whose technical or category-specific attributes determine suitability | Missing fields and unsupported comparisons | | Non-standard model | A bundle, subscription, configurable or made-to-order product, if the category contains one | Incorrect product boundaries, totals or fulfilment assumptions | The first-pass sample is sufficiently covered when every relevant rule in the category appears at least once, including one current price or stock event. That does not make it statistically representative. If one case fails, add further products governed by the same rule. The purpose is to learn whether the defect belongs to one record or to a repeated mapping, ownership or update process. Record that distinction; do not turn a small diagnostic sample into a catalogue-wide percentage. ### Step 3: follow each case across the truth chain Create one row per product or variant and record the same fields at each available layer. | Layer | Evidence to capture | |---|---| | Authoritative source | Product and variant ID, title, attributes, price, currency, promotion, availability, eligibility and named owner | | Live product page | Customer-visible values, selected variant, canonical URL, image, delivery and returns information | | Structured page data | Product, offer and variant fields actually rendered in the page source | | Current merchant feed | Submitted values, update timestamp, accepted, rejected and warning states | | Target AI result | Product selected, variant, claims made, visible source or destination, test conditions and timestamp | | Purchase path | Product or cart destination, retained variant and quantity, recalculated price, availability and customer confirmation | For each field, write down the acceptable update delay before testing. A made-to-order product and fast-moving stock do not need the same tolerance. Without a declared tolerance, “fresh” has no operational meaning. If a layer does not exist, mark it `not applicable`. If access or eligibility cannot be confirmed, mark it `unknown`. Do not fill a missing platform result with an assumption. Google's current [product structured-data guidance](https://developers.google.com/search/docs/appearance/structured-data/product) recommends product-page markup, a Merchant Center feed, or both for Google's own experiences. Google says the combination can improve eligibility and help it understand and verify product data. That is a Google-specific benefit, not a guarantee for other AI systems. ### Step 4: test five buyer tasks Use the sampled products in a clean conversation and preserve the prompt, response, sources, date, locale and platform conditions. 1. **Exact lookup:** find a named product and exact variant from the merchant. 2. **Constrained discovery:** find a product in the category that meets the buyer's budget and two or more relevant constraints. 3. **Comparison:** compare two sampled products using the attributes that actually determine the decision. 4. **Freshness check:** confirm the current price, promotion, availability and delivery position for a volatile case. 5. **Purchase path:** select the correct variant and use whatever next step the channel genuinely supports - product-page referral, basket or checkout. Repeat surprising or inconsistent results under the same recorded conditions before treating them as a pattern. AI outputs can vary, and one answer does not prove a stable platform behaviour. ### Step 5: give every check one honest status | Status | Use it when | Do not claim | |---|---|---| | **Pass** | The tested stage produced the expected product, variant and values within the declared tolerance | That the whole catalogue or every prompt will pass | | **Merchant-controlled defect** | An authoritative source, page, image, structured value, feed mapping or checkout destination is contradictory, unavailable or rejected for a reason the merchant controls | That repairing it guarantees an AI recommendation | | **Channel-side issue observed** | The merchant-controlled evidence is correct and available, but the tested AI result is missing, wrong or loses a constraint | That the model or platform is the proven root cause after one run | | **Unknown** | Platform access, ingestion, eligibility or a necessary result cannot be inspected | That the merchant passed or failed | The useful output is an issue register with the product, field, affected layer, evidence, owner and next action. A single “AI-ready” score hides the information needed to repair anything. ## 4. How to turn the findings into a decision Read the first broken layer, not the most fashionable possible solution. | Observed finding | Appropriate next action | What not to assume | |---|---|---| | Product and variant IDs or customer-visible values disagree before the AI channel | Repair product ownership, mappings and update rules | That another feed or chatbot will reconcile the conflict | | Pages or images cannot be reached as intended | Repair URLs, rendering, canonicalisation, crawler policy or security configuration | That adding structured data alone will make the product discoverable | | Product pages are correct but the feed is missing, stale or rejected | Repair the channel mapping, validation, update and rejection process | That producing one valid export solves ongoing freshness | | Source, page and accepted feed are correct but the AI result is wrong or absent | Repeat the controlled test, preserve evidence and treat it as a channel-side limitation or unknown | That the ecommerce site must be rebuilt | | The correct product is found and a reliable merchant handoff is available | Keep and improve the existing checkout path | That embedded payment is necessary | | The handoff loses variant, quantity or context | Scope the smallest cart or checkout handoff integration | That an agent-completed transaction is required | | The business wants the agent to complete orders, but platform access is unconfirmed | Mark the capability unavailable for planning purposes and monitor the target platform | That catalogue, structured-data or website work can unlock access | | The platform confirms the merchant and integration are eligible for agent-completed checkout | Scope authentication, state, confirmation, payment, order and exception handling | That eligibility removes transaction, customer-experience or operational risk | This is also where the distinction between discovery and checkout becomes practical. A product can be eligible for search but not for checkout. A customer can still complete a useful AI-assisted journey when the final purchase happens on the merchant's site. If the failure continues through payment, fulfilment, returns or analytics, move beyond this channel test and inspect the wider [ecommerce conversion chain](https://izzy.agency/en/blog/ecommerce-website-not-converting/). ## 5. What to fix now and what not to build yet ### Fix now when the evidence supports it - Stabilise product and variant identifiers. - Name the authoritative system and owner for price, availability, attributes, delivery and eligibility. - Remove contradictions between the customer-facing page, structured data, feed and checkout. - Make intended product pages and images reliably reachable. - Validate feeds before submission and record accepted, rejected and warning states afterwards. - Define update tolerances and an incident owner for commercially dangerous mismatches. - Preserve the selected product and variant when the buyer moves to the merchant's site. This work supports ordinary ecommerce operations as well as emerging AI channels. ### Do not build yet without a verified path - a platform-specific feed when the merchant has no confirmed access; - a multi-platform translation layer before one end-to-end route is understood; - agent-completed payment without confirmed platform eligibility and a controlled customer journey; - an MCP server whose only objective is generic “AI visibility”; - an `llms.txt` file presented as a substitute for product data, feeds or transaction tools; - a campaign promising ChatGPT inclusion, AI rankings or automatic purchasing. A crawler, product feed, platform catalogue and MCP tool solve different problems. Our guide to [llms.txt and WebMCP](https://izzy.agency/en/blog/llms-txt-webmcp-website-ai-agents/) explains the difference between describing a site and exposing a controlled action. Different AI engines may also retrieve and cite different sources under different conditions. If the product evidence is correct but one platform behaves differently, use a controlled multi-engine method rather than a generic visibility diagnosis; see [why AI search engines cite different sources](https://izzy.agency/en/blog/why-ai-search-engines-cite-different-sources/). ## 6. When an external review earns its place As Head of Digital Commerce, you do not need an external supplier merely to tell you that product data should be accurate. Outside help earns its place when catalogue, ecommerce, merchandising, engineering and platform teams cannot agree on the first broken layer - or when the proposed repair crosses several systems, markets and owners. [IZZY's Advisory model](https://izzy.agency/en/personalised-solutions/) can be scoped around one category and one target channel before anyone commits to a larger build. The useful evidence to bring is: - the category and market being tested; - the systems holding product, price and inventory data; - the target AI channel and current access status; - the risk-based product set; - one mismatch, rejection or uncertain result already observed. If the product is found correctly and the failure begins after the buyer reaches your site - for example at JavaScript rendering, consent, anti-bot controls, cart or checkout - the more relevant route is IZZY's [Agent Readiness Audit](https://izzy.agency/en/services/agent-readiness-audit/). It tests revenue-critical site journeys with real agents. It is not a substitute for diagnosing catalogue truth, feed ingestion or platform eligibility. The Advisory should end in a bounded investment decision: repair internally, bring in an Embedded Expert for a defined gap, scope a controlled implementation through a Full Partnership, or do not build yet. It should not promise inclusion in ChatGPT, a higher AI ranking, automatic purchases or support from a specific platform. ## Conclusion Current OpenAI and Shopify documentation covers product discovery, merchant checkout handoff and gated agent-completed checkout. That documented capability does not make today's product results consistently reliable or every merchant automatically eligible. The sensible response is neither to dismiss the channel nor to rebuild ecommerce around the hype. Test one meaningful category. Cover the product rules most likely to fail. Follow each item from authoritative data to the AI result and purchase path. Label merchant defects, channel-side observations and unknowns separately. Then fund the first repair the evidence reveals. Sometimes that will be catalogue work. Sometimes it will be a feed or checkout handoff. Sometimes the merchant will have clean foundations and no platform access - and the correct decision will be to monitor rather than build. ## Turn one category into a defensible investment decision As Head of Digital Commerce, bring one product category, its source systems, the AI shopping channel you are considering and one mismatch or uncertainty you can already see. We will scope whether the next step is an internal repair, a bounded Advisory, an Embedded Expert, a controlled implementation - or no build yet. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) or [send the team a short brief](https://izzy.agency/en/contact/). ## Frequently asked questions ### Does poor AI product discovery mean our ecommerce site is badly built? Not by itself. Contradictory product data, inaccessible pages or stale feeds can contribute. But platform coverage, catalogue normalisation, shopper intent and the agent's retrieval or reasoning can also cause a poor result. Test the chain before assigning blame. ### Do product pages and structured data make products visible in ChatGPT? They can make product facts public and easier for supported systems to interpret. They do not guarantee that ChatGPT or another assistant ingests, retrieves, recommends or ranks the product. ### Do we need a product feed? It depends on the target channel. A feed can provide a normalised product record, update control and validation feedback. Follow the channel's current specification and confirm merchant access before building a new export. ### Can we reuse a Google Merchant Center feed for ChatGPT? OpenAI documents a Google-compatible core feed path after it confirms that the registered feed supports that representation. Compatibility does not mean every existing export is complete or accepted unchanged. Validate the current [OpenAI product specification](https://developers.openai.com/commerce/specs/file-upload/products). ### Do we need an MCP server for AI shopping visibility? Not necessarily. MCP can expose controlled tools or actions when a target journey requires them. It is not a generic replacement for crawlable pages, accurate product data or a documented merchant feed. ### Can an AI assistant buy from any ecommerce website? No. An assistant can often send a shopper to a merchant page or checkout. Agent-completed purchase requires a supported integration, confirmed eligibility, transaction controls and customer confirmation. Until the target platform confirms that access, plan for a merchant-checkout handoff. ## Sources and evidence note This article uses current official platform documentation for access, feed, crawler and checkout claims, plus methodological research on shopping-agent performance. Material sources include: - [OpenAI Agentic Commerce onboarding](https://developers.openai.com/commerce/guides/get-started) - [OpenAI stable product-feed specification](https://developers.openai.com/commerce/specs/file-upload/products) - [OpenAI checkout documentation](https://developers.openai.com/plugins/build/monetization) - [OpenAI crawler documentation](https://developers.openai.com/api/docs/bots) - [Shopify agentic commerce documentation](https://shopify.dev/docs/agents) - [Shopify Checkout MCP documentation](https://shopify.dev/docs/agents/carts-and-checkout/checkout-mcp) - [IZZY Agent Readiness Audit](https://izzy.agency/en/services/agent-readiness-audit/) - [Google Product structured-data guidance](https://developers.google.com/search/docs/appearance/structured-data/product) - [Google guidance on sharing product data](https://developers.google.com/search/docs/specialty/ecommerce/share-your-product-data-with-google) - [ShoppingBench, AAAI 2026](https://ojs.aaai.org/index.php/AAAI/article/view/40640) - [EComAgentBench preprint](https://arxiv.org/abs/2606.17698) - [ShoppingComp preprint](https://arxiv.org/abs/2511.22978) - [Shopify Engineering on catalogue clustering](https://shopify.engineering/catalog-clustering) ShoppingBench is peer-reviewed. EComAgentBench and ShoppingComp are preprints. Their metrics are not directly comparable and do not estimate ordinary consumer failure rates. The risk-based category test and four result statuses are IZZY diagnostic methods, not an industry standard, statistical audit or readiness certification. Platform access, specifications and beta status can change. Recheck availability, eligibility and checkout statements immediately before publication. This article does not promise platform inclusion, ranking, recommendation, traffic, sales or automated purchasing, and it is not legal, tax, payments or compliance advice for a specific implementation. --- ### Do You Need a Rebrand - or Is the Real Problem Somewhere Else? URL: https://izzy.agency/en/blog/brand-identity-roi-technical-founders/ Published: 2026-08-20 Summary: A diagnostic guide for technical founders: separate positioning and brand problems from UX, accessibility, conversion, acquisition and engineering before commissioning a redesign. Something on the website feels wrong. The homepage is crowded. The sales deck tells a clearer story than the site. The product has moved on, but the company still describes what it used to be. Visitors browse, hesitate or leave before the action the business cares about. The brief becomes: “We need a new design.” That instinct may be pointing at a real problem. It has not identified the problem yet. A team without specialist language may use *design* as one word for the whole visible experience. The underlying issue may be positioning, messaging, user experience, accessibility, the offer, acquisition or a journey that is technically broken. The expensive mistake is translating a reasonable feeling directly into a rebrand. ## The short answer Do not begin by asking how the website should look. Begin by asking where understanding or action first breaks. - If different people describe the company differently, investigate positioning and messaging. - If the intended user understands the offer but cannot find or complete the next action, investigate UX and product design. - If people are blocked by keyboard access, labels, contrast, focus or assistive technology, investigate accessibility directly. - If forms, authentication, checkout, payments or integrations fail, investigate engineering and operations. - If the journey works but attracts the wrong audience or presents an unconvincing offer, investigate acquisition and conversion. - If the product, market or company has changed while the story and identity have not, a rebrand may be the right move. A redesign is not a diagnosis. A rebrand should be the result of diagnosis, not the starting assumption. ## In this guide - [Why “design” becomes the catch-all](#why-design-becomes-the-catch-all) - [The IZZY Rebrand Triage](#the-izzy-rebrand-triage) - [The decision changes with the company](#the-decision-changes-with-the-company) - [What brand work can and cannot change](#what-brand-work-can-and-cannot-change) - [Scope the smallest useful intervention](#scope-the-smallest-useful-intervention) - [Evaluate progress without manufacturing an ROI promise](#evaluate-progress-without-manufacturing-an-roi-promise) ## Why “design” becomes the catch-all Design is the part everyone can see. That makes it the natural name for dissatisfaction with the whole experience. A homepage can look incoherent because the visual system is inconsistent. It can also look incoherent because nobody has decided which buyer it serves, which problem matters or how the product should be explained. A polished interface can still contain an inaccessible form. A clear message can still lead to a failed payment. A working funnel can still receive the wrong traffic. These problems can appear together, but they are not interchangeable. **Positioning** defines the market context: who the company is for, which problem it owns and why a buyer should consider it. **Messaging** turns that position into language people can understand and repeat across the website, sales material and product. **Brand identity** creates a coherent verbal and visual system for expressing the company: name, voice, typography, colour, imagery, motion and rules for using them. **Product design and UX** organise information, choices and interactions so a user can understand the product and complete a task. **Accessibility** addresses barriers that affect people with disabilities. It overlaps with usability, design, content and code, but it requires dedicated evaluation rather than a visual opinion. **Conversion work** examines whether the audience, offer, evidence, journey and follow-up support the intended business action. **Engineering** makes the journey function under real conditions: data, forms, authentication, payments, integrations, errors, monitoring and recovery. Changing the visual layer can help when the visual layer is the problem. It cannot substitute for a decision the business has not made or repair a system that does not work. ## The IZZY Rebrand Triage Use the table to choose what to investigate. It is a routing tool, not a score and not an industry standard. What you observe | Evidence to inspect | Possible problem lane | Sensible next move --- | --- | --- | --- The homepage, pitch deck and sales team describe the company differently | Compare the audience, category, problem, promise and proof used on each surface | Positioning and messaging | Resolve the core position and message hierarchy before changing the identity Relevant buyers need a founder to explain what the company actually does | Observe how buyers interpret the homepage and compare that with the intended meaning | Positioning, messaging or information architecture | Test a clearer proposition and page structure before commissioning a full rebrand People understand the offer but cannot find or complete the next action | Run representative tasks and inspect hesitation, backtracking, errors and abandonment points | UX or product design | Repair the affected journey and test it again The form, login, checkout or payment route fails or produces inconsistent states | Compare the browser experience with application, payment, integration and operational evidence | Engineering or operations | Fix the first repeated technical break before redesigning the surface People encounter keyboard, focus, label, contrast, zoom or assistive-technology barriers | Combine relevant standards-based checks, knowledgeable review and user evaluation | Accessibility | Define and remediate the accessibility scope; do not treat a visual refresh as conformance Traffic arrives but enquiries are consistently irrelevant | Compare acquisition promise, landing message, intended buyer, search terms and submitted needs | Audience, acquisition or offer | Correct targeting or the offer before interpreting the result as a brand failure The message is stable, but every team recreates visual decisions | Compare current assets, templates, components and governance | Visual identity or design system | Build or repair the reusable system rather than reopening positioning without evidence The product, audience or category has changed, but the company still tells its old story | Compare the current business with the position and associations carried by the existing brand | Repositioning and brand identity | Decide the new position first, then update the verbal and visual system Visitors spend time across several pages but do not take the intended action | Combine behavioural data with task observation, enquiries and journey evidence | Could be discovery, positioning, offer, UX or measurement | Treat the pattern as a question; find the first repeated break before prescribing work The last row matters. Time on site and pages viewed show behaviour, not motive. A visitor may be interested, comparing, lost, unable to find information or simply leaving a tab open. The metric cannot tell the team which explanation is true. Start with one relevant buyer and one representative journey. Ask what they understood, what they expected, what they tried and where the observed experience diverged from the intended one. ## The decision changes with the company The right level of brand work depends on what the business already knows. Company stage is context, not a price card or an automatic trigger. ### Before launch An early product needs enough clarity to be understood and enough consistency to be credible. It may need a working name, a clear audience, a concise promise, a basic verbal and visual system and rules the team can apply. It may not need an elaborate identity built around assumptions that customer conversations will soon change. Before launch, ask: - Is the intended buyer specific enough to recognise themselves? - Can the team explain the problem and useful outcome without listing features? - Does the basic identity work across the product, website and sales material? - Which parts are decisions, and which remain hypotheses to revisit? The useful outcome is a coherent foundation that can learn, not a finished monument. ### After early customer proof This is where the brand question can become more concrete. The product has users. Customer language exists. The team knows which explanations create interest and which create confusion. At the same time, the website, deck, product and sales process may have evolved separately. Brand work becomes credible when it uses that evidence to resolve a real inconsistency: - customers value something the homepage barely mentions; - the product now solves a different problem from the original pitch; - sales repeatedly translates internal language into buyer language; - every new page or campaign recreates the message and visual rules; - the company looks and sounds different across important touchpoints. The case is not “the company is old enough for a rebrand”. The case is that the current system no longer represents or supports the business the company can now describe with evidence. ### During repositioning Repositioning is a business decision before it is a design exercise. The company may be entering a new category, addressing a different buyer, moving upmarket or changing the scope of its product. If the position changes, the existing name, message and identity may no longer carry the right meaning. But changing visuals before resolving the position simply gives the uncertainty a new appearance. Sequence the work: - establish the intended buyer, category, problem and difference; - identify which existing associations should be kept, changed or retired; - test whether the new language is understood as intended; - build the verbal and visual system around the resolved position; - plan the rollout across the website, product, sales material, directories and operational touchpoints. Growth alone is not a reason to rebrand. A meaningful mismatch between the business and the way it is understood is. ## What brand work can and cannot change Brand work can improve the quality and consistency of decisions the company controls. It can: - clarify who the company is trying to reach and what it wants to be known for; - align the language used across marketing, sales and product; - create a verbal and visual system teams can apply without reinventing it; - make the company easier to describe consistently; - give future pages, campaigns and product surfaces clearer constraints; - expose disagreements the organisation has been hiding inside copy or design reviews. It cannot, by itself: - create demand for an unwanted product; - make an unconvincing offer valuable; - repair a confusing product journey; - fix a broken form, checkout, integration or analytics implementation; - establish accessibility conformance without the required design, implementation and evaluation work; - guarantee conversion growth, investment, pricing power, faster sales or any other commercial result. The responsible claim is narrower: > Brand work can make a company easier to understand, describe and represent consistently. Commercial impact depends on the product, audience, offer, distribution, experience and implementation, so it should be measured rather than promised. ## Scope the smallest useful intervention Not every diagnosis should end in a full rebrand. Evidence | Appropriate scope --- | --- The position is clear, but the homepage buries it | Messaging and information-architecture refresh The verbal position works, but visual execution is inconsistent | Visual identity refresh or design-system repair The business has changed, and the old story no longer fits | Repositioning followed by verbal and visual identity work The buyer understands the offer, but the journey is difficult | Product or website UX work Accessibility barriers block use | Accessibility requirements, remediation and evaluation across design and implementation The journey fails technically | Engineering or rescue work The audience or offer is wrong | Acquisition, offer or product work before redesign The evidence is inconclusive | A bounded framing and research step; no redesign decision yet This protects the team from two opposite mistakes: commissioning a large rebrand to solve a contained problem, or polishing one page when the company has outgrown its position. A useful brief should name: - the business change or observed symptom; - the intended buyer and journey; - the evidence already available; - the decision the work must enable; - what is explicitly outside the scope; - how the team will check whether the chosen problem was addressed. “Make it modern” is not enough. Neither is “increase conversion”. The first describes taste. The second names an outcome without identifying the mechanism. ## Evaluate progress without manufacturing an ROI promise Define evidence before the work begins. The evidence should match the diagnosed problem. For positioning and messaging, examine whether relevant buyers interpret the company as intended, which words they repeat and where they still need explanation. For consistency, compare the core promise, terminology and identity across the homepage, product, sales material and the channels that matter to the business. For usability, observe whether intended users can complete representative tasks and explain where they hesitate or fail. For accessibility, use an agreed evaluation scope that combines appropriate checks and human expertise. A visual approval and an automated scan are not enough to establish accessibility. For implementation, verify that the system shipped correctly across the surfaces and journeys included in the scope. Commercial measures such as qualified enquiries or sales conversations can be monitored, but they should not be attributed to brand work automatically. Campaigns, traffic quality, product changes, pricing, seasonality and sales follow-up may change at the same time. The objective is not to prove that brand causes every positive movement. It is to make the intended change explicit, observe whether it happened and avoid claiming what the evidence cannot isolate. ## Conclusion: diagnose before you redesign “We need a new design” can be a useful starting signal. It is not yet a useful scope. Find the first place where the real experience diverges from the intended one. If buyers cannot understand or repeat the company’s position, investigate brand and messaging. If they understand but cannot act, investigate UX, accessibility, conversion or engineering. If the business has changed and the old story no longer fits, reposition before redesigning the expression of it. Sometimes the right answer is a rebrand. Sometimes it is a smaller messaging repair, a product-design intervention, an accessibility programme, an engineering fix or no design project yet. Choosing correctly is more valuable than making the wrong answer look polished. ## Bring the symptom, not a preselected package Bring the homepage, current sales explanation, important buyer journey and any evidence showing where understanding or action breaks. IZZY can help frame whether the next move belongs to positioning, brand identity, product design or another part of the system. If the evidence points to brand, [IZZY Brand & Identity](/en/services/branding/) starts with the position and message before building the visual system. If it points elsewhere, not recommending a rebrand is still a useful result. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) ## Frequently asked questions ### Is brand identity the same as a logo or visual identity? No. A logo and visual identity are parts of the system. Brand identity also includes the verbal signals and rules that keep the company’s expression coherent. Positioning sits underneath that expression: it defines the market context the identity needs to support. ### Does a technical startup need a complete brand before launch? Not automatically. It needs enough clarity and consistency for the intended buyer to understand the product and for the team to communicate coherently. The appropriate depth depends on which decisions are stable and which are still being tested. ### Can a rebrand improve conversion? It may remove confusion or inconsistency that affects a decision, but it cannot guarantee a conversion change. Conversion also depends on audience, offer, product, UX, accessibility, technical reliability, distribution and follow-up. Diagnose the mechanism and define evidence before making the claim. ### Should we rebrand before redesigning the website? Resolve positioning first when the company’s audience, category or promise is unclear or changing. If the position is sound and the problem is confined to site structure, interaction or implementation, a website intervention may be enough. ### How do we know whether the problem is UX or brand? Ask whether the intended user understands the offer before the difficulty appears. If they cannot explain what the company does or why it matters, investigate positioning and messaging. If they understand but cannot navigate or complete the task, investigate UX and implementation. Both can be present, so test the sequence rather than choosing by opinion. ### Does a redesign solve accessibility problems? Only if accessibility requirements are included in the design and implementation and the result is evaluated appropriately. A new visual style or automated scan alone does not establish that people with disabilities can use the experience. ## Sources and evidence note The Rebrand Triage is an IZZY diagnostic framework, informed by founder input and the current division between IZZY Brand & Identity, Product Design and engineering capabilities. It is not an external standard or a measured prevalence model. Only primary official guidance was used for the bounded factual claims in this article: - [W3C Web Accessibility Initiative: Evaluating Web Accessibility](https://www.w3.org/WAI/test-evaluate/) - accessibility should be evaluated during design and development; tools alone cannot determine whether a site meets accessibility standards. - [Google Analytics Help: Engagement overview report](https://support.google.com/analytics/answer/13391283?hl=en) - average engagement time records time with a website in focus or an app in the foreground; it does not establish why a person stayed or what they understood. No external conversion benchmark, price range, market prevalence claim, fundraising claim, client outcome or promised ROI is used. Checked on 20 August 2026. --- ### AI text watermarks: what does a detected mark actually prove? URL: https://izzy.agency/en/blog/ai-text-watermark-authorship-proof/ Published: 2026-08-18 Summary: What an AI text watermark means when an LLM drafts, rewrites, translates or proofreads - and why it cannot calculate human versus AI authorship. An AI provider announces text watermarking. The headline is simple. The conclusion many people may draw from it is not: *watermarked means written by AI*. For IZZY, the reason to examine this news is to make a more useful distinction. A text can be **processed with AI** without being **autonomously authored by AI**. If people collapse those two events, a technical signal can become a false account of who contributed what. Imagine a colleague writes a press release. They ask an AI assistant to improve the structure, translate one paragraph and fix the punctuation. Later, a detector reports a watermark. Did the AI write the release? That question asks the signal to do more than it can. A text watermark can support a narrower conclusion: a compatible model was probably involved in generating enough of the tested wording for its statistical pattern to be detectable. It is not a percentage meter for human versus AI authorship. The distinction matters because AI now sits at many points in a content workflow. “Used AI” can mean generating a first draft, rewriting a human draft, translating it, changing five sentences or correcting three commas. Those are different contributions, even when the final document carries the same binary label in a dashboard. This article explains what generation-time text watermarking is, what a positive or negative result can establish, and which records a publishing team still needs. ## Answer in 60 seconds An AI text watermark is usually not a visible stamp or a hidden character. In the statistical schemes covered here, the model embeds a pattern while choosing the next tokens in its response. A detector with the relevant key or scoring method tests whether the word sequence contains enough evidence of that pattern. ([Anthropic](https://www.anthropic.com/news/claude-text-watermark), [Dathathri et al., Nature](https://www.nature.com/articles/s41586-024-08025-4)) A detected watermark can support that the watermarked model was likely involved in producing or processing the tested text. It does **not** tell you: - who originated the ideas or supplied the source draft; - what percentage of the document is “human” or “AI”; - whether the model drafted the text or heavily edited it; - whether the claims are true, sourced, original or legally usable; - who owns the text or is responsible for publishing it. The reverse is also important. No detected mark does not prove human authorship. The text may be too short or factual, lightly proofread, generated by an unmarked or older model, heavily edited after generation, or tested with an incompatible detector. Anthropic's current explanation makes the mixed-authorship problem explicit. Its announced Claude watermark can only estimate the likelihood that Claude was partly involved; it cannot distinguish “Claude wrote this” from “Claude heavily edited this”. Anthropic also says a grammar-and-punctuation-only proofread may leave too few changed words for the mark to register, while a translation produced by Claude can carry a watermark because Claude chooses the translated words. ([Anthropic](https://www.anthropic.com/news/claude-text-watermark)) Treat a watermark as one provenance signal. Keep the draft history, the model-use record, the human review and the publication decision separate. ## In this article 1. [What “AI watermark” can mean](#1-what-ai-watermark-can-mean) 2. [How a statistical text watermark works](#2-how-a-statistical-text-watermark-works) 3. [How IZZY distinguishes AI authorship from AI assistance](#3-how-izzy-distinguishes-ai-authorship-from-ai-assistance) 4. [What a positive or negative result can establish](#4-what-a-positive-or-negative-result-can-establish) 5. [What the EU AI Act changes, and what it does not](#5-what-the-eu-ai-act-changes-and-what-it-does-not) 6. [How the IZZY workflow records AI contribution](#6-how-the-izzy-workflow-records-ai-contribution) 7. [What to do when a detector flags a document](#7-what-to-do-when-a-detector-flags-a-document) 8. [Conclusion](#the-operational-conclusion) 9. [Frequently asked questions](#frequently-asked-questions) 10. [Sources](#sources-method-and-limitations) ## 1. What “AI watermark” can mean The word *watermark* is being used for several different mechanisms. They should not be treated as interchangeable. | Mechanism | What it records or tests | What it does not establish | | --- | --- | --- | | **Generation-time text watermark** | A statistical pattern embedded as the model chooses tokens | The human/AI authorship percentage, factual quality or ownership | | **Content Credential or provenance metadata** | Signed statements about an asset's origin, tools or editing history | That every statement in the credential is true in a wider real-world sense, or that missing metadata means no AI was used | | **Visible AI label** | A disclosure presented to a reader | The detailed contribution of each person or tool | | **Post-hoc AI-writing classifier** | Stylistic or learned features associated with AI-written text | A provider-specific watermark unless it has the relevant watermark method or key | | **Internal process record** | Who used which tool, for what task, with which review and approval | A technical signal embedded in the published text | Anthropic illustrates the first two mechanisms in the same announcement. It describes a statistical watermark for Claude-generated text, but a C2PA Content Credential for supported image and file formats. The text pattern is embedded in generation; the file credential is signed provenance metadata. ([Anthropic](https://www.anthropic.com/news/claude-text-watermark), [C2PA specification](https://spec.c2pa.org/specifications/)) C2PA itself is careful about the boundary. Its specification provides a way to attach cryptographically verifiable provenance assertions to an asset. It does not turn provenance into a judgement that the content is “good”, “bad” or factually trustworthy. This is why “the detector found an AI watermark” is not enough information for a governance decision. First ask: **which mechanism, from which provider or standard, tested by which tool?** ## 2. How a statistical text watermark works An autoregressive language model produces a response one token at a time. At each step, it estimates a probability distribution over possible next tokens and samples from the available choices. Some choices are tightly constrained. In “2 + 2 =”, replacing “4” with another number changes the answer. Other choices have more freedom. A model may be able to select “grey” or “overcast” without materially changing the sentence. A generation-time watermark uses those flexible choices to introduce a keyed statistical pattern. The detector later scores the relationship between the observed tokens and the pattern expected under that key. The result is evidence with a confidence threshold, not an invisible sentence saying “this document was written by AI”. The general approach is established in peer-reviewed work. A 2023 ICML paper described a scheme that promotes a randomised set of candidate tokens and detects the resulting pattern with a statistical test. The 2024 SynthID-Text paper described a different sampling method, keyed scoring and production-scale deployment. Anthropic says its Claude implementation is a version of the SynthID-Text approach. ([Kirchenbauer et al., ICML](https://proceedings.mlr.press/v202/kirchenbauer23a.html), [Dathathri et al., Nature](https://www.nature.com/articles/s41586-024-08025-4), [Anthropic](https://www.anthropic.com/news/claude-text-watermark)) Three practical limits follow from the mechanism. ### Length matters Longer passages usually offer more token choices and therefore more evidence. Short samples can leave a detector uncertain. No universal minimum applies across providers, models, languages, prompts and detector thresholds. ### The type of text matters Creative or varied writing offers more alternative phrasings than exact quotations, code, fixed facts or tightly constrained answers. Google DeepMind says its text watermark works best on longer, diverse responses and is less effective where little variation is possible. Anthropic makes the same point for factual passages, proofreading and much code. ([Google DeepMind](https://deepmind.google/blog/watermarking-ai-generated-text-and-video-with-synthid/), [Anthropic](https://www.anthropic.com/news/claude-text-watermark)) ### Later editing matters Light editing may preserve enough of the pattern to detect. A substantial rewrite, paraphrase or translation by another system can weaken it. The SynthID-Text paper identifies editing, spoofing and scrubbing as continuing limitations; Anthropic says a complete rewrite can remove its mark. ([Dathathri et al., Nature](https://www.nature.com/articles/s41586-024-08025-4), [Anthropic](https://www.anthropic.com/news/claude-text-watermark)) These are scheme-level observations, not a guaranteed threshold for a particular document. Until a provider publishes its detector, operating parameters and validation results, do not turn “probably robust to light edits” into a numeric rule. ## 3. How IZZY distinguishes AI authorship from AI assistance The following distinction is an IZZY editorial definition. It is not a legal test, an industry standard or a claim that authorship can always be reduced to one rule. ### AI authorship: the model is given the topic and creates the content autonomously In the clearest AI-authorship scenario, a person supplies a topic or short instruction and the model produces the material without a human-designed research framework, source process, decision structure or specialist editorial direction. The person may still decide whether to use the output. But the substance and wording of the draft were produced by the model from a minimal brief. ### AI assistance: people design and own the content system In an AI-assisted workflow, the human contribution is not limited to correcting a finished model response. People define the problem, build the research and editorial process, set the evidence boundaries, direct the model's task, review the output and remain responsible for the publication. The model proposes content *inside that system*. A specialist then checks, changes, accepts or rejects it. This means “AI was used” and “AI was the author” are not equivalent statements. ### A public version of the IZZY content workflow The simplified workflow used for an evidence-led article is: 1. **Select the signal and reader decision.** Decide why the topic matters now, who needs the answer and which decision the article should support. 2. **Check existing coverage.** Review current IZZY articles and the live service context so the new piece has a distinct job and does not repeat an existing article. 3. **Research current questions.** Examine recent discussions for language, confusion and objections, while treating community evidence as directional rather than representative. 4. **Build the evidence base.** Open primary, official, peer-reviewed and relevant practitioner sources; separate verified facts from observations and inference. 5. **Create the claim boundaries.** Record what each source supports, what it does not support and which volatile claims will need rechecking. 6. **Define the editorial position.** Set the angle, structure, exclusions, first-party IZZY contribution and the practical reader outcome. 7. **Use AI for synthesis and structure.** Give the model the research and evidence boundaries to organise, compare and stress-test, not to invent an article from a one-line topic. 8. **Apply specialist editing.** Check the reasoning, facts, sources, technical or regulatory nuance, IZZY voice and whether the proposed framework is genuinely useful. 9. **Run publication QA.** Recheck links, metadata, unsupported claims, limitations, conversion path and any need for legal or subject-matter review. 10. **Keep human responsibility.** A named person approves, revises, delays or rejects the article. The model does not make the publication decision. This is deliberately a public, simplified description. It explains the control points without disclosing IZZY's internal prompts, qualification logic, research operations or other proprietary methods. This article began with a founder-selected news signal. The technical and regulatory claims were checked against primary and peer-reviewed sources; the interpretation of authorship versus assistance was then reviewed with the founder; and the article remains subject to human editorial and legal review. It is not presented as a client case or a fully autonomous AI article. ### What changes across drafting, editing, translation and proofreading The original idea behind this article was that even a typo check could watermark human text. The evidence supports a more precise answer: **a light proofread can produce too few model-chosen words to create a detectable signal.** The probability of detection grows as the model makes more wording decisions. | Workflow event | How much wording the model chooses | What a watermark result may show | What it cannot tell you | | --- | --- | --- | --- | | Human supplies a draft; model corrects punctuation and a few errors | Very little | A mark may be too sparse to detect | Whether the proofread happened if no mark appears | | Human supplies a draft; model line-edits several sentences | Some | A mark may show likely involvement if enough marked choices remain | Which ideas or passages originated with the human | | Human supplies a draft; model substantially rewrites it | A great deal | A positive result may be consistent with heavy model editing | Whether the model or human should be called the “author” | | Model creates a full first draft | Most or all output wording | A sufficiently long result may carry strong watermark evidence | Whether the facts are correct, sources were checked or a human later approved it | | Model translates human-written text | All wording in the target language | Anthropic says its translated output carries its watermark | Who authored the underlying ideas or source-language text | | Human lightly edits model-generated text | The model's earlier choices mostly remain | The watermark may survive | The scale or quality of the human review | | Human or another model completely rewrites the output | Few original choices remain | The original watermark may weaken or disappear | That no AI was involved earlier in the workflow | | Human text is pasted into an AI chat only for analysis, with no returned text published | None of the published wording | A generated response may be marked, but the untouched source is not rewritten by that act | Whether AI was consulted, unless a separate process record exists | The table explains why authorship percentages are the wrong output. A detector looks for evidence in the final token sequence. It does not replay the drafting history, compare every version, assign intellectual contribution or decide whether “author”, “editor”, “translator” or “proofreader” is the right description. Two documents could therefore produce similar detector results while having very different histories: - one started as a human draft and received a substantial model rewrite; - another was model-drafted and then carefully reviewed and edited by a person. The watermark alone cannot reconstruct which path occurred. ## 4. What a positive or negative result can establish Use calibrated language. The safest interpretation depends on whether the detector is authentic, compatible with the scheme and applied to enough unaltered text. For IZZY, the most immediate risk is not the existence of the watermark but its interpretation. A binary result can create false precision about the model's contribution and become the basis for accusations against an employee, writer, supplier or student. The detector cannot supply the missing drafting history, so the organisation must not pretend that it can. ### If a compatible watermark is detected You may be able to say: > “This test found statistical evidence consistent with the named model or watermarking scheme having generated or processed at least part of this text.” You should not upgrade that to: - “AI wrote 100% of this document.” - “The named person did not write it.” - “The text was not human-reviewed.” - “The content is false, plagiarised or low quality.” - “The publication complies with AI-transparency law.” - “The provider owns the output.” Anthropic states that its watermark says nothing about ownership or authorship and cannot identify a specific user, organisation or chat. It is designed to test likely Claude involvement. Those are different claims. ([Anthropic](https://www.anthropic.com/news/claude-text-watermark)) ### If no watermark is detected You may be able to say: > “This detector did not find enough evidence of this particular watermark in the tested sample.” You should not upgrade that to “a human wrote it”. Possible explanations include: - the provider or model did not apply that watermark; - the text predates the model's marking support; - the wrong detector or key was used; - the passage is too short, exact or factual; - the model only made a few proofreading changes; - later editing weakened the pattern; - another model with a different scheme generated the text. In other words, a negative result has a defined technical scope. It is not a certificate of human authorship. ### A detector score is not an authorship score Detection performance is normally discussed through thresholds, true-positive rates and false-positive rates. The SynthID-Text research, for example, evaluates detection as a function of text length at a fixed false-positive rate. That is useful for measuring a detector. It does not convert a 99% watermark probability into “99% of the words were written by AI”. ([Dathathri et al., Nature](https://www.nature.com/articles/s41586-024-08025-4)) The percentage belongs to the detector's hypothesis test, not to the division of creative or editorial labour. ## 5. What the EU AI Act changes, and what it does not The current attention to text marking is partly regulatory. Article 50(2) of the EU AI Act requires providers of systems that generate synthetic audio, image, video or text to make outputs machine-readable and detectable as artificially generated or manipulated, as far as technically feasible. It also tells providers to account for effectiveness, interoperability, robustness, reliability, implementation cost and the state of the art. ([EU AI Act, Article 50](https://eur-lex.europa.eu/eli/reg/2024/1689/oj?locale=en)) The provider marking duty is not absolute. Article 50(2) excludes systems to the extent they perform an assistive function for standard editing or do not substantially alter the input data or its semantics. The European Commission's current guidelines identify standard editing as an example of what can fall outside scope. ([European Commission guidelines](https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-transparency-obligations)) Two distinctions are operationally important. ### Provider marking and publisher disclosure are separate A machine-readable mark is a technical signal added or enabled by the AI provider. A visible label is information presented by a deployer or publisher to an audience. Article 50 gives them different scopes. For text, the deployer disclosure rule in Article 50(4) concerns AI-generated or manipulated text published to inform the public on matters of public interest. It also contains an exception where the content has undergone human review or editorial control and a person or organisation holds editorial responsibility. The facts of a specific publication still matter. ### A watermark is not a compliance decision A positive mark does not prove that the right visible disclosure was made. A missing mark does not automatically show a provider breached the law. The system, model date, use case, editing function, technical feasibility, provider/deployer role and applicable exceptions all matter. The Commission published a voluntary Code of Practice to help providers and deployers implement these duties, which have applied since 2 August 2026. The current guidelines and Code should be checked when designing a real workflow. This article explains the evidence boundary; it is not legal advice. ([European Commission Code of Practice](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content), [European Commission guidelines](https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-transparency-obligations)) ## 6. How the IZZY workflow records AI contribution Watermark detection becomes useful when it sits inside the human-owned workflow described above. The following five records make that process reviewable. They are IZZY operating guidance, not a technical standard or legal-compliance certificate. For each material public asset, keep five records. ### 1. Signal: what exactly was detected? Record: - provider, model and version where known; - detector name, version and access route; - date of the test; - text sample tested; - raw score, category or threshold result; - whether the detector is designed for that watermark. Avoid screenshots without context. A future reviewer needs to understand what the tool actually tested. ### 2. Event: what did the model do? Classify the use: generate, expand, rewrite, translate, summarise, line-edit, proofread or analyse. Preserve the prompt or task instruction where policy and confidentiality permit. This event record supplies the context the watermark lacks. “Proofread punctuation only” and “rewrite for a new audience” should not collapse into the same checkbox. ### 3. Contribution: what came from people and prior sources? Keep the original draft, meaningful versions and tracked changes. Name the subject-matter contributor and editor. Record sources used for material factual claims. This does not produce a perfect authorship percentage either. It creates reviewable provenance for the decision that matters: can the organisation stand behind the final asset? ### 4. Assurance: what was independently checked? Record the checks relevant to the content: - factual and source verification; - rights, confidentiality and personal-data review; - brand and accessibility review; - subject-matter or legal review where risk requires it; - final human approval and unresolved limitations. A watermark is not a substitute for any of these controls. ### 5. Decision: what will be published, labelled or escalated? Name the accountable approver. Record whether the asset can be published, needs a visible disclosure, must be revised, or requires specialist review. State the reason and the policy or rule applied. The result is not “AI: yes/no”. It is an evidence-backed publishing decision. If your organisation already uses an AI pilot policy, add this record to it rather than creating a parallel bureaucracy. Our [AI pilot-to-production governance checklist](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/) covers the wider allow, review, disclose, restrict and block decisions; the signal-to-decision record handles the narrower watermark-evidence problem. ## 7. What to do when a detector flags a document Do not begin by rewriting the text. Preserve the artefact and the result first. 1. **Confirm the detector.** Is it the provider's detector or a tool validated for that exact scheme? A generic AI-writing classifier is not equivalent. 2. **Preserve the tested version.** Keep the complete text, the sample boundaries and the raw result. 3. **Check model coverage.** Was the named model capable of applying that watermark on the generation date? Anthropic says older Claude model support is being rolled out over time, so the date and version matter. 4. **Retrieve workflow evidence.** Look for the original draft, version history, tracked changes, model task and named reviewers. 5. **Classify the model's role.** Separate proofreading, line editing, substantial rewriting, translation and full generation. 6. **Run the checks the mark cannot perform.** Verify claims and sources, assess rights and confidentiality, and confirm the responsible editor. 7. **Apply the relevant policy.** Decide on approval, visible disclosure, revision or escalation based on use and risk - not on the detector alone. 8. **Escalate high-consequence decisions.** Do not use a watermark result by itself to make an employment, academic-misconduct, legal or regulatory determination. Those decisions need the applicable process and specialist review. For a low-risk blog draft, this may take minutes. For investor communications, regulated advice, employment material or public-interest reporting, it should be more formal. If the team cannot reconstruct the workflow, label that evidence gap honestly. “Could not verify the drafting history” is safer than inventing a precise authorship story from a statistical signal. ## The operational conclusion Text watermarking can make model involvement more detectable. That is useful. It is also a narrower achievement than “solving AI authorship”. The mark lives in model-generated choices. Authorship lives across ideas, drafts, edits, sources, approvals and responsibility. A detector sees the first layer; your content process must preserve the rest. At IZZY, we therefore do not define authorship by whether AI touched a document. We ask who designed the content process, selected and verified the evidence, directed the work, edited the result and accepted responsibility for publication. That does not make every AI-assisted document human-authored by default. It makes the contribution question answerable with workflow evidence instead of a binary guess. For publishing teams, the practical rule is simple: > **Use the watermark to open an evidence review, not to close the authorship question.** If you need to map where AI generates, rewrites, translates or approves public content - and define the evidence each step must keep - IZZY's [Advisory work](https://izzy.agency/en/personalised-solutions/) can help turn the workflow into owned controls and decision records. ## Map where AI touches your public content In a 30-minute scoping call we will map where AI generates, rewrites, translates or approves your public content, and define the evidence each step should keep. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) ## Frequently asked questions ### Are all LLM outputs watermarked? No. Providers can use different marking methods, keys and rollout schedules, and some models or deployments may not apply a generation-time watermark. A detector for one scheme cannot automatically detect every other model. ### Does Claude watermark text? Anthropic announced on 14 August 2026 that future Claude models will generate watermarked text using a version of SynthID-Text, with support for older models to be added over the following months. It also said a detection API would be offered, but implementation details were still being worked out on the research date. Check the model and current provider documentation before assuming a particular response is marked. ([Anthropic](https://www.anthropic.com/news/claude-text-watermark)) ### Will a typo check watermark my human-written text? Not necessarily in a detectable way. Anthropic says a grammar-and-punctuation-only proofread may change too few words for its watermark to register. Heavier editing creates more model choices and therefore more room for a signal. ### If a watermark is detected, did AI write the whole document? No. A compatible positive result can support likely model involvement in the tested wording. It does not calculate the model's share of the ideas, source draft or editorial contribution, and it cannot distinguish full drafting from heavy editing on its own. ### If no watermark is detected, is the text human-written? No. The sample may be short, factual, lightly edited, generated before marking support, produced by another model, extensively rewritten or tested with the wrong detector. ### Can copying and pasting remove a statistical text watermark? Ordinary copying does not change the token sequence, so it should not by itself remove a pattern embedded in word choices. Editing can weaken the pattern, and a complete rewrite may remove it. This is different from deleting metadata or invisible characters. ### Can a text watermark identify the user or chat? Anthropic says its announced watermark contains no identifying information and cannot be traced to a specific person, organisation or chat. Do not generalise that privacy property to every possible marking system without checking its design. ### Is a watermark the same as a Content Credential? No. In the mechanisms discussed here, a text watermark is embedded statistically during token generation. A C2PA Content Credential is a cryptographically signed provenance record associated with an asset. They can complement one another, but they carry different evidence. ### Does the EU AI Act require every AI-assisted text to have a visible label? No such universal conclusion follows from Article 50. It separates provider machine-readable marking from deployer disclosure and includes scope conditions and exceptions, including standard editing in the provider duty and human review/editorial responsibility in the public-interest text disclosure rule. Apply the current official guidance to the specific workflow and seek legal advice where the consequence warrants it. ## Sources, method and limitations Researched on 18 August 2026. The core evidence was checked against: - [Anthropic's primary explanation of Claude's announced text watermark](https://www.anthropic.com/news/claude-text-watermark), published 14 August 2026; - [the peer-reviewed SynthID-Text paper in Nature](https://www.nature.com/articles/s41586-024-08025-4); - [the foundational ICML/PMLR statistical-watermark paper](https://proceedings.mlr.press/v202/kirchenbauer23a.html); - [Google DeepMind's SynthID text explainer](https://deepmind.google/blog/watermarking-ai-generated-text-and-video-with-synthid/); - [Article 50 of the EU AI Act](https://eur-lex.europa.eu/eli/reg/2024/1689/oj?locale=en), the [European Commission's current transparency guidelines](https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-transparency-obligations) and the [final Code of Practice](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content); - [the C2PA technical specifications](https://spec.c2pa.org/specifications/). Recent community discussions were also reviewed to identify live questions about proofreading, editing, translation and false authorship inferences. They informed the questions addressed, not the technical or legal conclusions. Automated community retrieval returned no ranked candidates, and X, TikTok, Instagram and usable YouTube coverage were unavailable in that pass, so the demand signal is directional rather than representative. First-party editorial input was collected from the IZZY founder on 18 August 2026. It informed the distinction between autonomous AI authorship and a human-designed AI-assisted content process, the public workflow description and the warning about false contribution claims and accusations. It did not supply or replace the technical and legal evidence. No client case is claimed. Important limitations: - Anthropic had announced a forthcoming detector API but had not published its operating details in the sources checked. - No universal minimum sample length or edit threshold was verified; performance depends on the scheme and text. - Provider rollout and EU implementation guidance can change. - Watermark research demonstrates capabilities and limitations under defined tests; it does not make every real-world detector result conclusive. - The IZZY authorship/assistance distinction and workflow are editorial and operating definitions, not a universal authorship test. - This article is operational guidance, not legal advice or a determination of authorship, copyright, employment, academic integrity or regulatory compliance. --- ### Open-source code security: can your team absorb the next vulnerability report? URL: https://izzy.agency/en/blog/open-source-code-security-maintainer-checklist/ Published: 2026-08-18 Summary: A practical security operating model for founders maintaining public code: threat modelling, repository controls, scanning, triage, fixes and releases. Your repository is public. It moves money, processes customer data, controls infrastructure or sits inside products you do not operate. You know that defenders can inspect it. You also have to assume that curious researchers, opportunistic actors and well-resourced attackers can inspect it too. The instinctive answer is: *scan the code before they do*. That is necessary, but incomplete. A scanner can create a finding. It cannot, by itself, decide whether the finding is reachable in your deployment model, reproduce it safely, develop and test the right fix, coordinate disclosure, ship a release or help downstream users install it. The founder-level question is therefore not “Do we have an AI security scanner?” It is: > **Can our team turn the next credible security finding into a validated, released and adopted fix - without losing control of the product?** This article explains the minimum operating model behind that capability. It applies to public codebases in general, with extra urgency where a defect could expose funds, credentials, customer data, signing keys, critical availability or downstream users. ## Answer in 60 seconds AI-assisted security research is now capable of finding real vulnerabilities in mature open-source software. Google Project Zero reported an experimental AI-assisted SQLite finding in 2024; Mozilla reported that a 2026 collaboration produced reproducible findings that Firefox engineers validated and fixed. These are selected defensive programmes, not proof that every model or scan works equally well. ([Google Project Zero](https://projectzero.google/2024/10/from-naptime-to-big-sleep.html), [Mozilla](https://blog.mozilla.org/en/firefox/hardening-firefox-anthropic-red-team/)) At the same time, more candidate findings mean more noise. A 2026 project-scale preprint evaluating five LLM-based methods and two traditional tools against 222 known vulnerabilities and 24 active projects found low recall and high false-discovery rates. The result is limited to the studied tools, languages and projects, but it is a useful warning: a scan result is neither proof of exploitability nor proof of safety. ([Li et al., 2026 preprint](https://arxiv.org/abs/2601.19239)) For a founder or CTO, the practical response is a continuous loop: 1. Model what can go wrong and which code paths matter most. 2. Prevent avoidable repository, CI/CD, secret, dependency and release failures. 3. Detect continuously with several methods, using AI as one layer. 4. Validate findings against the real threat model and require reproducible evidence. 5. Fix, test, release, communicate, measure adoption and feed the lesson back into the system. If any gate has no owner, evidence or response path, you do not yet have a security programme. You have a collection of tools. ## In this article 1. [What the original idea gets right, and what it overstates](#1-what-the-original-idea-gets-right-and-what-it-overstates) 2. [Why capable teams still delay proactive security work](#2-why-capable-teams-still-delay-proactive-security-work) 3. [The IZZY security absorption loop](#3-the-izzy-security-absorption-loop) 4. [A 30-minute readiness test](#4-a-30-minute-readiness-test) 5. [Where AI-assisted review belongs in the tool mix](#5-where-ai-assisted-review-belongs-in-the-tool-mix) 6. [What to implement first](#6-what-to-implement-first) 7. [When to bring in specialist security help](#7-when-to-bring-in-specialist-security-help) 8. [Conclusion](#the-real-advantage-is-not-scanning-first) 9. [Frequently asked questions](#frequently-asked-questions) 10. [Sources](#evidence-and-limits) ## 1. What the original idea gets right, and what it overstates The idea behind this article is straightforward: widely available language models make source-code analysis easier, public repositories are inspectable, and projects with valuable assets may be attractive targets. Why do teams not use the same capability proactively? Three parts hold up. First, AI-assisted vulnerability research has crossed from benchmark demonstrations into real defensive work. Mozilla’s March 2026 account is particularly useful because it describes the quality that made the reports actionable: minimal test cases, engineer validation and a collaborative remediation process - not merely model-generated prose. OpenAI’s more recent Patch the Planet programme describes the same pattern at a wider operational level: security engineers manually reproduce, deduplicate, reassess and prioritise candidate findings before they reach maintainers. This is vendor-published programme evidence, not an independent performance comparison, but the workflow is instructive. ([Mozilla](https://blog.mozilla.org/en/firefox/hardening-firefox-anthropic-red-team/), [OpenAI](https://openai.com/index/patch-the-planet/)) Second, public code can be reviewed at scale by defenders and attackers. That makes “security through obscurity” a poor operating assumption. It does **not** make open source inherently insecure. Public review can expose weaknesses; it can also enable independent scrutiny, faster fixes and reusable defensive tooling. The relevant risk is the combination of an exploitable weakness, a reachable attack path and insufficient prevention or response - not repository visibility by itself. Third, discovery is only the first part of the pipeline. CNCF describes the operational sequence as scanning, triage, fixing and releasing, followed by downstream consumption of the fix. Its warning is that attention is concentrating on discovery while the later stages remain bottlenecks. ([CNCF](https://www.cncf.io/blog/2026/04/16/the-ai-driven-shift-in-vulnerability-discovery-what-maintainers-and-bug-finders-need-to-know/)) One part could not be verified: there is no representative evidence in this research showing that attackers systematically begin with wallets or rank repositories strictly by direct monetary value. Wallets, payment infrastructure, identity systems and signing services are reasonable **high-consequence examples**, but not a verified universal order of attack. The useful thesis is therefore narrower: > AI is increasing the amount of code-analysis capability and the possible flow of findings. Teams responsible for public, high-consequence software should improve their capacity to prevent, validate and remediate vulnerabilities before that flow overwhelms them. ## 2. Why capable teams still delay proactive security work Most founders do not consciously choose insecurity. They postpone a poorly defined body of work. ### “Security review” has no natural boundary Does it mean dependency scanning, static analysis, secret scanning, fuzzing, manual code review, penetration testing, infrastructure review or a formal smart-contract audit? Without a threat model, every answer can sound both necessary and insufficient. ### Findings create work before they create value A new scanner can generate hundreds of alerts within minutes. Each one still needs context: Is the code reachable? Is the configuration realistic? Does an attacker already need privileged access? Is the affected version shipped? Is the severity overstated? A possible issue that nobody can validate is operational debt, not yet a security outcome. This burden is already visible among maintainers. Directus describes a growing gap between the capacity to find possible issues and the capacity to verify, remediate and publish fixes: its CTO reports 230 vulnerability reports in early 2026 against a historical average of 30 to 40 a year, of which roughly 5% were validated as genuine. That is one vendor’s inbox rather than an industry measurement, but it shows the shape of the problem. An ongoing OpenSSF working-group issue is gathering practices for handling high volumes of low-quality AI-generated reports; because that work remains open, it should be treated as a current community discussion rather than a finished standard. ([Directus](https://directus.com/resources/ai-is-straining-vulnerability-disclosure-for-maintainers), [OpenSSF working-group issue](https://github.com/ossf/wg-vulnerability-disclosures/issues/178)) ### Product pressure rewards visible features Security work competes with revenue, hiring, reliability and customer commitments. Its value is often expressed as an avoided event, while the cost of a release delay is immediate. Unless leadership defines a minimum baseline and gives someone authority to stop a release, the short-term incentive usually wins. ### The team confuses a tool with an accountable process Buying a scanner feels bounded. Creating owners, response times, release gates and disclosure rules feels organisational. Yet it is the organisational layer that determines whether a real finding is fixed or left in a queue. This is one reason security belongs inside the delivery model rather than in a quarterly clean-up. NIST’s Secure Software Development Framework is explicitly designed to integrate secure practices into each software-development lifecycle, including preparation, protection, production and vulnerability response. ([NIST SSDF 1.1](https://csrc.nist.gov/pubs/sp/800/218/final)) ### What IZZY has encountered in client codebases This is not an abstract concern for our team. In codebases we have reviewed, inherited or helped to harden, we have encountered: - keys or other credentials stored in the repository; - basic login credentials available in repository files; - API routes left without appropriate protection; - dependencies with known vulnerabilities; - little or no automated test coverage; - weaknesses in the application code itself; and - reliance on obscurity - the assumption that an endpoint, convention or implementation detail would remain undiscovered. Our response is not to add one scanner and declare the code secure. Depending on the codebase, the work has included repository hygiene, key rotation, protection of API routes, stronger authentication, dependency updates, attention to software-supply-chain risk, automated tests, structured releases and pull-request review workflows. Why had these controls not been applied earlier? In our experience, the main constraints are time and security experience - especially the experience needed to turn a changing body of guidance, dependency information and tooling into one maintained process. The information exists, but the team has not systematised how it stays current and reaches day-to-day delivery. There is an important evidence limit. Client NDAs mean we cannot identify the organisations, expose their repositories or present these observations as independently verifiable case studies. This is qualitative IZZY experience, not a prevalence estimate, benchmark or claim of a measured outcome. ## 3. The IZZY security absorption loop The following is an IZZY operating framework, not a certification or external standard. It converts established control categories into five questions a small or growing product team can own. ### Gate 1. Model: do we know what must not fail? Start with consequence, not tools. Map the actors, assets, trust boundaries, privileged operations, data flows and external interfaces. Identify the code paths that can move funds, authorise users, access secrets, sign or publish artefacts, mutate customer data, execute untrusted input or disable recovery. For each critical path, record: - who or what can reach it; - the preconditions an attacker would need; - the unacceptable outcome; - the existing prevention and detection controls; - the owner who can accept, reduce or escalate the risk. The [OpenSSF OSPS Baseline](https://baseline.openssf.org/versions/2026-02-19) places a security assessment at Level 2 and formal threat modelling and attack-surface analysis at Level 3. You do not need to wait until you are a large project to use the logic. A small wallet or identity service may warrant deeper modelling than a much larger low-consequence library. **Evidence to keep:** a versioned threat model, a critical-path inventory and named risk owners. ### Gate 2. Prevent: have we removed avoidable exposure? Prevention is broader than application code. A sound algorithm can still be undermined by a compromised maintainer account, an over-permissioned workflow, a leaked token or a substituted release artefact. At minimum, assess: - multi-factor authentication and least-privilege access for maintainers; - protected primary branches and independent review for sensitive changes; - CI/CD isolation, especially when workflows process untrusted pull requests or metadata; - secret detection, rotation and a documented secret-management policy; - explicit dependency inventories and rules for vulnerable or malicious packages; - signed or otherwise verifiable releases, unique versions and security change logs; - supported-version and end-of-life statements. These are not arbitrary checklist items. The current OSPS Baseline covers repository access, branch protection, CI/CD credentials, secrets, dependency policy, security assessments, vulnerability disclosure and release integrity across its maturity levels. CISA’s Secure by Design guidance similarly places responsibility on software producers to prioritise customer security throughout the product lifecycle. ([OpenSSF OSPS Baseline](https://baseline.openssf.org/versions/2026-02-19), [CISA and FBI guidance](https://www.cisa.gov/news-events/alerts/2025/01/17/cisa-and-fbi-release-updated-guidance-product-security-bad-practices)) **Evidence to keep:** exported settings, review rules, workflow permissions, release records and documented exceptions - not a screenshot that nobody revisits. ### Gate 3. Detect: are several methods watching the right surfaces? No single technique covers every weakness class. Choose a portfolio that matches the threat model: - software-composition analysis for known dependency vulnerabilities; - secret scanning and push protection; - static analysis for code patterns and data flows; - dynamic testing against a deployed service; - fuzzing for unexpected inputs and state transitions; - property, invariant or differential tests for security-critical behaviour; - manual design and code review; - AI-assisted variant analysis or test generation where it adds coverage. If you use GitHub, its current repository-security guide covers the dependency graph, Dependabot alerts and updates, dependency review, code-security features, secret protection and security policies. Feature availability varies by repository type and plan, so verify what is enabled rather than assuming the platform default is sufficient. ([GitHub repository-security quickstart](https://docs.github.com/en/code-security/getting-started/quickstart-for-securing-your-repository)) Schedule checks on meaningful events: every pull request for cheap deterministic controls, every release for release-critical checks, and periodically for deeper analysis. A one-off scan becomes stale as soon as code, dependencies, configuration or attack knowledge changes. **Evidence to keep:** tool coverage, last successful run, versioned configuration, accepted suppressions and the tests attached to critical paths. ### Gate 4. Validate: can we separate a vulnerability from a plausible paragraph? Every incoming finding needs a triage contract. Require enough evidence to reproduce it without forcing a reporter to publish exploit details. A useful report normally includes: - affected component and version or commit; - prerequisites and assumed attacker access; - minimal reproduction or test case; - observed result and expected security property; - impact tied to the project’s threat model; - duplicate checks and relevant recent fixes; - a private, responsive communication path. Assign one owner to acknowledge reports and another technically qualified person to validate high-consequence findings. Define how to handle suspected duplicates, false positives, contested severity and reports that reveal a design concern rather than an exploitable vulnerability. Publish a `SECURITY.md` file with supported versions, a private reporting route, expected response stages and disclosure expectations. GitHub also supports private vulnerability reporting for eligible public repositories. ([GitHub vulnerability-reporting guidance](https://docs.github.com/en/code-security/how-tos/report-and-fix-vulnerabilities/configure-vulnerability-reporting)) Do not auto-close a report because the wording looks AI-generated. Do not accept it because the wording sounds technical. Judge the evidence and the threat model. **Evidence to keep:** acknowledgement time, validation decision, reproducer, severity rationale and decision owner. ### Gate 5. Remediate and learn: did the protection reach users? A validated finding is not the finish line. The fix must be designed, reviewed, tested, released and adopted. For each confirmed issue: 1. contain exposure if immediate mitigation is possible; 2. fix the root cause, not only the demonstrated input; 3. search for variants in similar code paths; 4. add a regression test or security invariant; 5. review whether the patch introduces a new failure mode; 6. publish the appropriate advisory and supported fixed versions; 7. notify affected operators or downstream maintainers through the agreed channel; 8. monitor installation or deployment where you have visibility; 9. update the threat model, development rule or test suite that should prevent recurrence. Measure outcomes that the team can act on: time to acknowledge, time to validate, time from confirmation to fixed release, age of unresolved high-consequence findings, percentage of critical paths with current tests, and - where observable - fixed-version adoption. Raw finding count is not a useful success metric on its own. **Evidence to keep:** patch, tests, release and advisory links, affected/fixed versions, notification record and post-incident learning. ## 4. A 30-minute readiness test Use this test with the founder, engineering lead and release owner. Do not prepare a presentation; open the repository and show evidence. | Question | Evidence available now | Warning sign | | --- | --- | --- | | What are the three highest-consequence attack paths? | Current threat model linked to code and architecture | The answer is “the whole codebase” or depends on one person’s memory | | Who can change code, CI/CD settings, secrets and releases? | Current access list, MFA and least-privilege controls | Former contributors, shared accounts or unclear workflow permissions | | Which checks run before a sensitive change can merge? | Enforced branch/ruleset settings and passing checks | Checks are advisory, regularly bypassed or absent from critical repos | | How are dependencies and secrets controlled? | Dependency inventory, alert policy, secret scanning and rotation path | Alerts exist but have no owner, threshold or deadline | | How can a researcher report privately? | `SECURITY.md`, security contact and tested private route | Public issue is the only route or the address is unmonitored | | Can the team validate a report safely? | Triage owner, isolated environment and reproducibility standard | Production is the test environment or nobody can make a severity decision | | Can the team ship and communicate a security fix? | Release owner, advisory process and supported-version policy | A patch can merge but there is no emergency release or notification route | | Do downstream users actually receive protection? | Deployment visibility, release adoption signal or explicit limitation | “Fixed on main” is treated as equivalent to user protection | Score the test conservatively: - **0–2 evidence-backed answers:** exposure is poorly bounded; prioritise a security-posture review before adding more scanners. - **3–5:** basic controls exist, but the hand-offs are fragile; run a report-to-release exercise and close the broken gate. - **6–7:** the operating loop exists; test it against a high-consequence scenario and inspect exceptions. - **8:** the loop is evidenced today, not guaranteed tomorrow; keep measuring it as the system changes. This is an IZZY prioritisation aid, not a risk score, security certification or substitute for specialist testing. ## 5. Where AI-assisted review belongs in the tool mix Use AI where it creates evidence or expands a well-defined search - not where it merely creates confidence. Promising uses include: - finding variants of a known vulnerability pattern; - identifying candidate critical paths for human review; - generating fuzzing harnesses, test scaffolding or attack-taxonomy drafts; - comparing implementations of the same protocol; - checking code behaviour against a written specification; - explaining a complex alert to speed human triage; - proposing a patch that is then reviewed and tested under the normal release controls. The Project Zero SQLite work illustrates a bounded, target-specific approach: the agent was given a previously fixed pattern and asked to search for related issues. The researchers also stressed that the work was experimental and that a target-specific fuzzer might, at that point, be at least as effective. ([Google Project Zero](https://projectzero.google/2024/10/from-naptime-to-big-sleep.html)) Weak uses include: - asking one model to “audit the repository” without a threat model; - accepting severity or exploitability without reproduction; - sending unreviewed model output directly to maintainers; - allowing an agent to patch and release without human approval; - treating a clean run as evidence that no vulnerability exists; - uploading sensitive private code or secrets without an approved data-handling path. The final point matters even when the repository is public. Build configuration, incident context, private branches, tokens and customer data may not be. If your team uses coding agents more broadly, define their access and approval boundaries separately; our guides to [AI-agent security controls](https://izzy.agency/en/blog/ai-agent-security-product-controls/) and [AI-agent permissions](https://izzy.agency/en/blog/ai-agent-permissions-access-control/) cover that adjacent decision. ## 6. What to implement first The order should follow consequence and evidence gaps, not whichever tool has the loudest dashboard. ### If you have no documented baseline Start with the authoritative repository and release path: 1. Inventory the repositories, packages, artefacts and supported versions that form the product. 2. Name an executive risk owner, a technical triage owner and a release owner. 3. Map the highest-consequence assets and attack paths. 4. Protect maintainer accounts, the primary branch, CI/CD credentials and release permissions. 5. Publish and test a private reporting route. ### If controls exist but findings accumulate Fix the validation system: 1. Define evidence required for triage. 2. Deduplicate and group findings by root cause and critical path. 3. Separate known-vulnerability alerts from novel code or design findings. 4. Create severity and remediation thresholds tied to the threat model. 5. Reserve engineering capacity for confirmed work. ### If fixes merge but users remain exposed Fix the release and adoption system: 1. Document supported versions and security-update expectations. 2. Make emergency releases repeatable and independently reviewed. 3. Produce clear affected/fixed-version advisories. 4. Improve update mechanisms and operator notification. 5. Track adoption where possible and state clearly where it is not observable. ### If the product is changing faster than the model Move security into change review. Update the threat model when trust boundaries, authentication, signing, fund movement, data access, plugin execution or infrastructure ownership changes. This often intersects with [technical-debt decisions](https://izzy.agency/en/blog/true-cost-of-technical-debt/): a component that nobody can safely modify is also hard to secure quickly. ## 7. When to bring in specialist security help Internal ownership does not mean doing every security task internally. IZZY’s boundary is straightforward: when a client needs a certified or formally independent service - such as a qualified penetration test, a security certificate or another formal assurance conclusion - the work should involve an appropriately qualified external security provider. IZZY can help expose the codebase, architecture and delivery gaps, prepare the remediation work and implement fixes, but that is not the same as issuing an independent certification. Escalate when: - a plausible issue could expose funds, signing keys, authentication, sensitive data or widespread downstream systems; - the team cannot safely reproduce the finding; - exploitability or severity remains contested; - cryptography, protocol design, memory safety, sandbox escape or cross-system trust is involved; - launch, acquisition, regulation, insurance or a customer requires a defined independent assessment; - the fix may reveal the vulnerability before users can update; - the team needs penetration testing, exploit validation, malware or incident forensics, or a formal smart-contract security conclusion. Define the external scope precisely: architecture or threat-model review, targeted code review, penetration test, smart-contract audit, release-pipeline assessment or incident response are different engagements. Ask what artefacts you will receive, what is excluded, how retesting works and who owns disclosure. An external report is still an input to your operating loop. If nobody can implement, release and monitor the recommendations, the engagement has not yet reduced the full product risk. ## The real advantage is not scanning first The best defensive outcome is not the largest alert queue. It is a team that knows what matters, prevents cheap failures, searches continuously, validates quickly and gets safe fixes into users’ hands. AI-assisted review can expand what a small team can inspect. It can also expand noise and create a false sense of coverage. Treat it as one component in an evidence-led security system. If your repository is already important enough that a defect could harm customers, funds or critical operations, do not begin with “Which scanner should we buy?” Begin with the 30-minute test above. The first missing piece of evidence usually tells you where the next security investment belongs. If that test exposes unclear ownership, fragile release controls or an unmanageable findings queue, IZZY can help scope the codebase and operating risks through a bounded [Rescue Mission & AI-Code Hardening review](https://izzy.agency/en/services/rescue-scale/). We will define the scope and evidence before proposing implementation; specialist penetration testing, exploit validation or formal smart-contract conclusions should be commissioned separately where required. ## Run the 30-minute test with us Bring the repository. We will work through the eight questions together and tell you which gate is missing evidence, before you buy another scanner. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) ## Frequently asked questions ### Does making a repository private solve this problem? No. Private visibility can reduce casual access, but it does not replace access control, dependency hygiene, secret management, secure design, testing, release integrity or incident response. It can also reduce community scrutiny. Choose visibility for the product and collaboration model, then secure the resulting system. ### Is open-source software less secure than closed-source software? Repository visibility alone does not answer that question. Security depends on design, implementation, maintainer capacity, dependency and release controls, deployment and response. Both open and closed software can contain exploitable weaknesses; public code can be examined by both defenders and attackers. ### How often should an open-source project run security scans? There is no universal interval. Cheap deterministic checks should normally run on relevant changes; release-critical checks should run before release; deeper reviews should follow the threat model, material architecture changes and risk. Record the trigger and owner rather than relying on an undocumented calendar reminder. ### Can an LLM perform a complete code-security audit? Current evidence does not support that claim. AI-assisted methods can find valuable issues and help create tests or patches, but they can also miss vulnerabilities and produce false positives. Use them with threat-model context, reproducible evidence, human validation and normal release controls. ### What should a `SECURITY.md` file contain? At minimum: supported versions, a private reporting route, the information needed to validate a report, expected response stages, disclosure expectations and any relevant safe-harbour language reviewed for your jurisdiction. This article is not legal advice. ### Do we need a penetration test if we already scan the source code? Possibly. Source scanning and penetration testing answer different questions. A penetration test can examine deployed behaviour, configuration, authentication, integration and attack chains that a source-code tool may not cover. The need and scope should follow the threat model and any customer, regulatory or insurance obligations. ## Evidence and limits Research checked on 18 August 2026. The control model draws on [NIST SSDF 1.1](https://csrc.nist.gov/pubs/sp/800/218/final), the [OpenSSF OSPS Baseline version 2026.02.19](https://baseline.openssf.org/versions/2026-02-19), [CISA Secure by Design guidance](https://www.cisa.gov/news-events/alerts/2025/01/17/cisa-and-fbi-release-updated-guidance-product-security-bad-practices) and current [GitHub repository-security documentation](https://docs.github.com/en/code-security/getting-started/quickstart-for-securing-your-repository). Evidence on AI-assisted discovery comes from selected project reports and one January 2026 preprint; it should not be generalised into a universal detection rate or guarantee. Recent community research across Reddit, YouTube, Hacker News and GitHub was used only to understand current questions and maintainer concerns. Coverage was partial, several clusters were noisy, and no community popularity metric or quote is used as factual support here. This article provides general product-engineering guidance. It is not a penetration test, formal security audit, smart-contract certification, compliance assessment or legal advice. Product-specific conclusions require access to the actual code, architecture, configuration, releases and operating context. --- ### One brain, many systems: how to build a searchable operating memory for your company URL: https://izzy.agency/en/blog/searchable-operating-memory-company/ Published: 2026-08-16 Summary: A practical architecture for searchable company knowledge across Drive, Notion, Slack, GitHub and Linear - with authority, permissions and freshness. Your company already has knowledge. The problem is that a product decision is in Slack, the signed brief is in Drive, implementation lives in GitHub and delivery status is in Linear. Search returns fragments; people still ask which one is current. An AI company knowledge base should not begin with “move everything into one tool”. Build a searchable operating memory: keep authoritative records in their working systems, index permitted context, refresh it through observable workflows, answer with sources and limits, then route approved actions back to the right owner. This is a reference architecture, not a deployed IZZY client case or performance result. Product plans, permissions, indexing latency, retrieval quality and operational controls need tenant-specific testing. ## The answer in 60 seconds - Keep records where teams create and maintain them. - Add an index; do not create a second uncontrolled truth. - Give each indexed item identity, authority, access and freshness metadata. - Require answers to cite evidence and expose uncertainty. - Treat stale, denied and contradictory results as designed states. - Start read-only. Add actions only after named owners accept the controls. ## In this article 1. [Searchability is not a migration project](#1-searchability-is-not-a-migration-project) 2. [Build five layers, not one brain-shaped database](#2-build-five-layers-not-one-brain-shaped-database) 3. [Define the minimum viable operating record](#3-define-the-minimum-viable-operating-record) 4. [Ingest selectively and preserve the source](#4-ingest-selectively-and-preserve-the-source) 5. [Make every answer prove itself](#5-make-every-answer-prove-itself) 6. [Follow one question across four systems](#6-follow-one-question-across-four-systems) 7. [Design the failure state first](#7-design-the-failure-state-first) 8. [Roll out from questions to bounded actions](#8-roll-out-from-questions-to-bounded-actions) ## 1. Searchability is not a migration project [IBM distinguishes](https://www.ibm.com/think/topics/system-of-record-vs-source-of-truth) systems that create and maintain domain records from systems that harmonise information across domains. The practical point is simple: one search surface does not need to replace the applications that run the work. A search layer answers “where is the evidence?” An authority rule answers “which record wins?” Confusing them creates a polished new silo: copied documents with no owner, old decisions that still look valid and summaries that cannot update the source. [Notion Enterprise Search](https://www.notion.com/en-gb/help/enterprise-search) and [AI Connectors](https://www.notion.com/en-gb/help/notion-ai-connectors) show one native approach: search a workspace and connected applications, with citations and permission mapping. Their own documentation also describes plan, setup, coverage, lookback, indexing and aggregation limits. Connected is not synonymous with complete, immediate or authoritative. ## 2. Build five layers, not one brain-shaped database The “one brain” metaphor is useful only when it produces clear responsibilities. Use five layers: | Layer | Job | Must not decide | |---|---|---| | **Source systems** | Create and maintain records | Cross-system interpretation | | **Knowledge/index** | Make permitted content discoverable | Which record is authoritative | | **Orchestration** | Detect, transform, refresh and report state | Business meaning | | **Interpretation** | Retrieve, compare, cite and express uncertainty | Permission to act | | **Action** | Write through an approved system and owner | Its own authority | [n8n's RAG documentation](https://docs.n8n.io/build/integrate-ai/understand-ai-components/retrieve-relevant-context) separates ingestion from querying: fetch content, split it into chunks, create embeddings, add metadata and retrieve a limited set of relevant chunks. That is useful plumbing. Your operating design still has to define authority, access, freshness, evaluation and failure. ## 3. Define the minimum viable operating record Before choosing a vector store or connector, define what the index knows about each item: | Field | Question it answers | |---|---| | `record_id` | Can we update the same item rather than duplicate it? | | `source_url` and `source_system` | Where is the original? | | `record_type` and `authority` | What is it, and when can it win? | | `owner` and `status` | Who maintains it, and is it active? | | `permission_class` | Who may retrieve it? | | `source_updated_at` / `indexed_at` | How fresh is source and index? | | `effective_from` / `review_at` | Is the rule current? | | `relations` | Which project, customer or decision does it concern? | Use durable organisational locations for durable records. For example, Google says [shared-drive files belong to the team](https://support.google.com/drive/answer/7286514?hl=en), not an individual, subject to edition and policy. But “in a shared drive” still does not mean “authoritative”: the record needs an owner, status and rule for what it governs. If you cannot fill these fields for a content class, do not quietly index it as trusted knowledge. Classify it as reference-only, exclude it or give it a review queue. ## 4. Ingest selectively and preserve the source Select sources by recurring business question, not by connector availability. Slack has search and configurable retention; messages may be edited or deleted. GitHub has repository search, roles and webhooks. Linear [searches issues, projects and documents](https://linear.app/docs/search), with documented mode and result limits. Each system exposes useful evidence under different access, history and update rules. Choose ingestion behaviour per record type: - **Link** when users need the current native record. - **Index** text needed for cross-system retrieval. - **Copy** only a justified immutable snapshot with provenance and expiry. - **Exclude** secrets, noise or content whose access cannot be mapped. [n8n's data loader](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.documentdefaultdataloader) can attach metadata for later filtering. In Notion, API search is [not exhaustive](https://developers.notion.com/reference/search-optimizations-and-limitations); use structured data-source queries when complete database retrieval is required. Freshness also needs state. Google Drive notifications say that something changed; the consumer must [read the change feed](https://developers.google.com/workspace/drive/api/guides/manage-changes). Store the last successful source event, retrieval and index update. A green connector icon is not evidence that every record is current. ## 5. Make every answer prove itself Semantic retrieval returns plausible matches, not a business verdict. n8n's vector tools expose a result limit; Notion says its connectors are better suited to finding and summarising than complex calculation or broad aggregation. Require an answer contract: 1. **Direct answer** - concise and scoped to the question. 2. **Evidence** - links to the exact records or passages. 3. **Authority** - which source governs each part. 4. **As-of time** - source update and index time. 5. **Conflict/uncertainty** - missing, stale or contradictory evidence. 6. **Next action** - permitted read-only step or named approval route. The answer “Launch is 14 October” is unsafe if it hides that Slack says 14 October, the approved Drive brief says 21 October and Linear has no committed milestone. Preserve access boundaries. If the system cannot map a searcher's rights to a source, deny or escalate rather than copy the content into a broader index. Connector vendors document permission mapping; treat those statements as capabilities to test, not independent assurance. See IZZY's guide to [AI-agent permissions](https://izzy.agency/en/blog/ai-agent-permissions-access-control/). ## 6. Follow one question across four systems Suppose a founder asks: “Are we ready to launch Atlas on 14 October?” - Drive contains the approved launch brief and acceptance criteria. - Slack contains a later pricing discussion, but not an approved change. - GitHub shows the release branch and an open security fix. - Linear shows delivery items, owners and current states. The operating memory retrieves all four, but does not flatten them. It reports: the approved brief names 14 October; one required security fix remains open; delivery state comes from Linear; the Slack pricing discussion is context, not approval. It links every statement, shows when each source changed and routes the readiness decision to the named launch owner. That is more useful than a confident paragraph assembled from the most semantically similar chunks. It helps a person decide without pretending the retrieval layer made the decision. ## 7. Design the failure state first | State | What the user should see | Operating response | |---|---|---| | **Stale** | Last source/index times | Refresh or label answer time-bounded | | **Denied** | Source exists but is inaccessible | Stop; request access through owner | | **Contradictory** | Both records and authority gap | Escalate; do not choose silently | | **Missing** | No adequate evidence | Say “could not verify” | | **Duplicate** | Same source/version appears twice | Reconcile identity and re-index | | **Partial** | Some sources processed, others failed | Return bounded result and failure state | n8n can [route execution failures](https://docs.n8n.io/build/flow-logic/handle-errors-gracefully), but business recovery remains yours. Record the failed source, stage, last successful checkpoint, owner and replay eligibility. Duplicate comparison helps only when identity fields, history and scope are designed; it is not end-to-end idempotency. Do not hide a failure behind a fluent fallback answer. “No evidence found”, “access denied” and “index older than source” are useful outputs. ## 8. Roll out from questions to bounded actions Start with ten to twenty recurring questions - not as a universal sample, but as a working inventory. For each, name expected sources, authority, access class, acceptable freshness and what a good “could not verify” response looks like. Run a read-only pilot. Review citations, missing evidence, permission behaviour, contradictions and update lag. Add sources only when they improve a defined question. This is also where you decide whether native enterprise search is sufficient or a custom retrieval layer deserves its operating cost; that later build-versus-connect decision needs its own assessment. Enable actions only for a narrow pattern with a named owner, exact write parameters, approval rules, write-back, monitoring and fallback. n8n can [pause selected AI tool calls for human review](https://docs.n8n.io/build/integrate-ai/ai-examples/human-in-the-loop-for-tools), but only tools wired to that review step are gated. The rollout gate is evidence from your corpus and workflows, not a generic accuracy claim. If the team cannot maintain authority and source metadata manually, automation will not repair the ownership problem. ## Conclusion: build a company memory that knows its limits A useful company brain is not one database and not one chatbot. It is a governed path from authoritative records to permitted retrieval, evidence-led interpretation and controlled action. Keep working systems accountable. Make the index observable. Make answers cite, date and qualify themselves. Then expand only where the business question, owner and failure path are clear. ## Map one recurring company question > Bring one question that currently requires searching several tools, the systems involved, your current authority rule and one recent wrong or contradictory answer. IZZY can map the smallest knowledge-and-action architecture worth testing. The right recommendation may be better record ownership or native search - not a custom AI build. This is exactly the scope of our [n8n AI Automation service](https://izzy.agency/en/services/automatisation-ia-n8n/). [Book a 30-minute scoping call](https://izzy.agency/en/contact/). ## Frequently asked questions ### What is an AI company knowledge base? It is a searchable layer over permitted company knowledge. A production design also needs identity, authority, freshness, citations, failure state and ownership. ### Should all company knowledge be moved into one database? No. Keep authoritative records where teams maintain them; index or link only what defined questions need. ### How should permissions work in company knowledge search? Retrieval should preserve source access. If permissions cannot be mapped reliably, deny or escalate instead of widening access through the index. ### Do we need custom RAG to build a company brain? Not necessarily. Native search/connectors may cover the questions. Custom retrieval is justified only when the required coverage, controls or workflow cannot be achieved acceptably without it. ### Where should we start? Start with recurring questions, expected evidence, authority and freshness. Pilot read-only before adding sources or actions. ## Sources and method Research was checked on 2026-07-27 using current official Notion, Google, Slack, GitHub, Linear, n8n, IBM and NIST pages. Links sit beside material claims. This article is IZZY architecture guidance. No live tenant, retrieval corpus, permission model or end-to-end workflow was tested. Search quality, completeness, latency, security, outcomes, ROI and legal compliance remain unverified. --- ### Why AI search engines cite different sources URL: https://izzy.agency/en/blog/why-ai-search-engines-cite-different-sources/ Published: 2026-08-13 Summary: Why do AI search engines cite different sources? Compare ChatGPT, Perplexity and Google, map each evidence gap and measure brand visibility responsibly. ChatGPT, Perplexity, Gemini and Google can answer the same buyer question with different evidence. Do not chase every citation: find where the evidence chain changes and which change deserves action. ## The 60-second answer AI search engines can cite different sources because they do not all: - decide to search the web in the same circumstances; - turn the original question into the same searches; - access the same source pool; - use the same model, mode, interface or conversation context; - retrieve and select the same documents; or - use retrieved material in the answer in the same way. Results can change between runs. Treat one answer as an observation. Test important buyer journeys, preserve the evidence, and keep retrieval, citation, support, recommendation, referral and qualified enquiries separate. ## In this article 1. [Why AI search engines can cite different sources](#1-why-ai-search-engines-can-cite-different-sources) 2. [Why one platform and one opening prompt give an incomplete view](#2-why-one-platform-and-one-opening-prompt-give-an-incomplete-view) 3. [A source can be found, cited and still fail to support the answer](#3-a-source-can-be-found-cited-and-still-fail-to-support-the-answer) 4. [Measure the seven-stage evidence chain](#4-measure-the-seven-stage-evidence-chain) 5. [Build a controlled multi-engine test](#5-build-a-controlled-multi-engine-test) 6. [Turn source differences into an action map](#6-turn-source-differences-into-an-action-map) 7. [Combine native and sampled data without inventing certainty](#7-combine-native-and-sampled-data-without-inventing-certainty) ## 1. Why AI search engines can cite different sources A generative system may transform a question, run several searches and assemble an answer from selected documents. A 2026 comparison found differences in external-source use, diversity and stability across Google organic search and five generative systems ([Findings of ACL 2026](https://aclanthology.org/2026.findings-acl.526/)). Six layers can change the result: | Layer | Why the source set can change | |---|---| | Search activation | A system may search automatically, only in a selected mode or not at all | | Query transformation | One prompt may become several targeted searches or subtopics | | Available source pool | Indexes, crawl access, search partners and selected source modes differ | | Surface and context | Interface, model, mode, locale, location, memory and conversation history can vary | | Retrieval and selection | Systems can retrieve, rerank and expose different documents | | Answer construction | A page may be displayed, omitted, weakly used or materially shape the answer | OpenAI says ChatGPT Search can rewrite a prompt into targeted searches and may use location or relevant memory ([OpenAI: ChatGPT Search](https://help.openai.com/en/articles/9237897-chatgpt-search)). Google says AI Overviews and AI Mode can fan a question out into related searches and may use different models and techniques ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)). Perplexity users can select different models and source modes ([Perplexity Pro Search](https://www.perplexity.ai/help-center/en/articles/10352903-what-is-pro-search)). These disclosures do not reveal complete ranking logic. They show why reports must record the surface and conditions, not rely on permanent platform stereotypes. ## 2. Why one platform and one opening prompt give an incomplete view A single-platform check can be valid for that condition but incomplete across several environments. One opening question is also limited: buyers add location, budget, integration, risk and proof requirements. Google treats each AI Mode follow-up as a new query for reporting ([Search Console measurement guidance](https://support.google.com/webmasters/answer/7042828?hl=en)). A conversational journey also adds new context at every step. A brand can therefore lead the first response and disappear after qualification. Possible explanations include: - mismatch with the new constraint; - unclear positioning or missing comparison facts; - an inaccurate third-party description; or - ordinary response variation. Only the first is necessarily correct exclusion. Test a stable journey instead: **discovery → qualification → comparison → objection → shortlist** Preserve the conversation. A screenshot of the first answer cannot show where the brand was lost. ## 3. A source can be found, cited and still fail to support the answer “Source” is often used for three different events: 1. **Retrieved:** the system found or consulted the page. 2. **Displayed:** the page appeared as a visible citation or source link. 3. **Supporting:** the page actually supports the statement that matters. These events are not interchangeable. Peer-reviewed research evaluates citation completeness and claim support separately ([Findings of EMNLP 2023](https://aclanthology.org/2023.findings-emnlp.467/)). More recent work distinguishes source credibility from answer groundedness ([EACL 2026](https://aclanthology.org/2026.eacl-long.115/)). For brand analysis, ask: > Does the cited page actually support the product, price, market, capability or suitability claim being made? A page may support only part of a sentence or another product version. Counting the URL without checking the claim can turn reporting success into reputation risk. ## 4. Measure the seven-stage evidence chain IZZY uses the following as an operating model, not as an external standard: | Stage | Business question | |---|---| | Search activation and eligibility | Was web retrieval used, and could the relevant page be reached? | | Retrieval | Was the page or domain present in the available source pool? | | Displayed citation | Was it shown as a source? | | Support or absorption | Did it support or materially shape the relevant claim? | | Brand outcome | Was the brand named, compared, recommended, retained or excluded? | | Referral and on-site response | Did an observable visit occur, and what happened on the site? | | Commercial outcome | Did the journey produce a qualified enquiry, opportunity or sale? | One stage does not prove the next. Citation does not prove recommendation; recommendation does not prove a visit; a visit does not prove a qualified lead. A visibility dashboard cannot prove revenue it does not observe. ## 5. Build a controlled multi-engine test Start with buyer decisions, using questions from sales, lost opportunities, support, site search and search demand. Select platforms for the market and record the exact product and mode. For each observation, retain: - prompt and follow-up path; - platform, interface and mode; - model, when visible; - language, country, location and session condition; - collection date; - raw answer and exposed source URLs. There is no verified universal number of prompts or repetitions. A 2026 preprint examining three platforms and three consumer topics found that single-run estimates could look more precise than the underlying results. Its scope cannot prescribe one number for every company ([Quantifying Uncertainty in AI Visibility](https://arxiv.org/abs/2603.08924)). Repeat decision-critical conditions until the range is useful. Label a one-run finding **observed**, a limited repeated pattern **directional**, and reserve stronger language for results that persist under comparable conditions. ## 6. Turn source differences into an action map The useful output is not “our score is 42.” It is a map connecting each gap to its first investigation. | Observed pattern | First investigation | |---|---| | A relevant owned page is not retrieved | Crawl access, indexation, internal linking, topical match and accessible evidence | | A third-party page is cited but describes the brand incorrectly | Directory, review, partner, press or reputation correction | | The brand is named without supporting evidence | Accessible product proof and independent corroboration | | The brand appears during discovery but leaves the comparison | Comparison-ready facts, constraints, pricing principles, positioning and proof | | A page is cited but does not support the answer’s claim | Factual accuracy, source relevance and manual support review | | AI referrals arrive but do not convert | Landing-page message match, proof, form, CRM hand-off and follow-up | Use “investigate,” not “cause.” An absent citation does not prove a crawl problem, and a later exclusion does not prove weak positioning. In a fictional example, a consultancy disappears when a buyer asks for French delivery and a regulated-sector reference. Its site supports both, but a cited directory is old. Confirm retrieval, correct the profile, expose the proof and retest. ## 7. Combine native and sampled data without inventing certainty Native reporting is platform-specific. Bing reports citations, cited pages, sampled grounding queries and trends, while warning that totals do not indicate answer placement, authority or role ([Bing AI Performance](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview)). Google’s reports cover impressions and pages on Google surfaces ([Google Search Central](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports)). Third-party metrics are not interchangeable. Ahrefs separates mentions, citations and found pages ([Ahrefs](https://help.ahrefs.com/en/articles/15501968-ai-visibility-metrics)), while Semrush uses several prompt databases and collection methods across its products ([Semrush](https://www.semrush.com/kb/1607-semrush-ai-visibility-data)). Before comparing dashboards, ask: - Where did the prompts come from, and which conditions were tested? - What is the denominator? - Are raw answers, sources and per-platform results retained? - Are methodology changes distinguished from performance changes? Use native reports for their own surfaces, controlled monitoring for comparable buyer journeys, and analytics and CRM data for qualified demand. ChatGPT referral links include `utm_source=chatgpt.com` where the referral survives ([OpenAI publisher guidance](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq)). Combine it with on-site behaviour, conversion events, CRM qualification and self-reported discovery. ## Conclusion: find the broken stage before funding the fix Different sources reflect different search decisions, source pools, surfaces, contexts and answer construction. The objective is not to “win every AI.” It is to find where reliable evidence stops and which intervention deserves budget. Begin with a controlled baseline. Preserve the answers. Separate retrieval from citation, citation from support and visibility from qualified demand. Improve the first broken stage, then measure again under comparable conditions. ## Map the AI sources shaping your buyer journey Bring the buyer questions, markets and competitors that matter. IZZY can establish a multi-engine baseline, show where your brand is retrieved, cited, retained or misrepresented, and turn the gaps into a prioritised SEO, content, reputation and conversion plan. When the need is repeatable, the method can become recurring monitoring. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Why do ChatGPT, Perplexity, Gemini and Google cite different sources? They can activate search differently, transform questions, access different source pools and use different models, modes and answer methods. Location, context and time can also affect results. ### Is a citation the same as a recommendation? No. A citation means a source was displayed. The brand may be recommended, criticised, excluded or absent. Check its framing and whether the source supports the statement. ### Can one AI visibility score compare different platforms? Only within a method showing the prompts, platforms, denominator, conditions and per-platform results. Otherwise a blended score can hide an absence or error. ### How often should AI visibility be measured? No universal cadence was verified. Measure often enough to distinguish patterns from variation. Keep conditions stable and annotate platform or methodology changes. ### How do you track AI citations and brand mentions? Use controlled buyer journeys, preserve answers and source URLs, and record mentions, recommendations, citations, support and accuracy separately. Add webmaster, analytics and CRM evidence. ### Can AI visibility be connected to leads or revenue? It can be connected to downstream evidence, but not reduced to a simple causal claim. Track referrals, on-site behaviour, conversions, self-reported discovery and CRM qualification, while stating where attribution remains incomplete. ## Sources and evidence note This IZZY synthesis uses research, platform documentation and public methodologies checked on 29 July 2026. Recheck product interfaces before implementation. The seven-stage evidence chain, buyer-journey test and source-to-action map are IZZY operating frameworks. They are not external standards, platform ranking disclosures or guarantees of visibility, traffic or commercial outcomes. --- ### Should you add an AI shopping assistant to your ecommerce site? URL: https://izzy.agency/en/blog/ai-shopping-assistant-ecommerce-site/ Published: 2026-08-12 Summary: Learn when an on-site AI shopping assistant can improve product discovery, what it needs and how to test one without replacing search. A customer wants a sofa that fits a narrow room, survives life with a dog and can arrive before a move. Your filters know dimensions, materials and delivery dates. They do not necessarily know how to turn the whole situation into a useful shortlist. This is the problem an on-site AI shopping assistant is meant to solve. It lets a customer describe a need in ordinary language, refine it through follow-up questions and move from an uncertain brief to products they can actually buy. That can be useful. It is not automatically a conversion strategy. An assistant built on incomplete product data, unreliable stock information or weak measurement may produce confident recommendations without improving the buying journey. Before adding another interface to the site, a retailer needs to decide where conversation genuinely helps, what the system is allowed to say and how success will be proven. ## Answer in 60 seconds An AI shopping assistant is worth testing when customers face a real product-finding problem that search, filters and category pages do not handle well. The strongest candidates usually involve several constraints, uncertain terminology, comparison or follow-up questions. Do not start with a site-wide assistant. Start with one task, one part of the catalogue and a defined group of eligible visitors. Keep conventional search and navigation available. Before the pilot, confirm that the assistant can retrieve authoritative product, price, stock, delivery and policy data. Decide what it must not infer, how personal data will be handled and who owns incorrect answers. Measure the whole path from exposure to fulfilled order, including activation, useful recommendations, product views, basket additions, cancellations, returns and complaints. A shorter session or a high chat count is not, by itself, evidence of a better shopping outcome. The decision should end in one of three places: run a bounded pilot, repair the foundations first, or do not build. ## In this article 1. [What is an AI shopping assistant?](#1-what-is-an-ai-shopping-assistant) 2. [When does conversational product discovery help?](#2-when-does-conversational-product-discovery-help) 3. [Does customer interest justify the investment?](#3-does-customer-interest-justify-the-investment) 4. [Five signs your ecommerce site is not ready](#4-five-signs-your-ecommerce-site-is-not-ready) 5. [What the assistant needs behind the interface](#5-what-the-assistant-needs-behind-the-interface) 6. [How to run a bounded pilot](#6-how-to-run-a-bounded-pilot) 7. [How to measure the result](#7-how-to-measure-the-result) ## 1. What is an AI shopping assistant? An on-site AI shopping assistant is a conversational product-discovery interface operated by, or embedded into, a retailer's website. It interprets a customer's request, asks for missing information and recommends products using the retailer's catalogue and commercial rules. It is useful to separate four experiences that are often grouped under the same label: | Experience | Main job | Typical risk | |---|---|---| | Customer-service chatbot | Answers questions about orders, delivery or policies | Gives generic or outdated support answers | | On-site shopping assistant | Helps a visitor discover and compare products | Recommends unsuitable or unavailable products | | External AI referral | Sends a shopper from an AI search or answer engine to the retailer | Attribution is incomplete or lost | | Purchasing agent | Selects, orders or pays on the customer's behalf | Permission, payment and accountability become more sensitive | This article is about the second category. If the system can change a basket, place an order or trigger another action, it moves into a higher-risk product and security problem. Our guide to [AI agent security and product controls](https://izzy.agency/en/blog/ai-agent-security-product-controls/) covers that wider boundary. The distinction matters because a pleasant chat interface does not prove that the underlying product-finding system works. The commercial value comes from the quality of the shortlist and the next step, not from the conversation alone. ## 2. When does conversational product discovery help? Conversation is most useful when a customer's need is difficult to express as a short keyword or a sequence of independent filters. Good candidate tasks include: - choosing among products with several interacting constraints; - translating an intended use into attributes when the customer does not know the category language; - refining a broad brief and explaining the final trade-offs. Research on conversational recommender systems supports the basic mechanism: dialogue can help elicit preferences, resolve uncertainty and collect feedback. It does not establish a universal commercial uplift. Evaluation methods vary, and a recommendation that sounds helpful can still be wrong. Controlled evidence points to a trade-off. A Microsoft experiment found faster, more satisfying consumer-choice tasks when the LLM result was correct, but also a risk of over-reliance when it was wrong. A smaller hybrid-store experiment found that conversation complemented normal browsing and did not improve every measured outcome. Do not make a customer chat when they already know the product name, want to apply a simple size filter or need to see the whole range. Search, filters, comparison tools, category pages and merchandising still have jobs to do. Baymard's product-finding research continues to document avoidable failures in these conventional interfaces. Adding AI does not repair them. ## 3. Does customer interest justify the investment? Recent French surveys suggest that AI-assisted shopping is no longer a fringe behaviour. They do not show that every retailer needs an assistant. The results need careful reading because the studies ask different questions: | Study | What it suggests | What it does not prove | |---|---|---| | FEVAD/Odoxa, January 2026 | In a survey of 1,500 French online shoppers, 31% reported using generative AI during shopping. Trust was higher before purchase than at the point of purchase. | That an on-site assistant improves conversion or that shoppers will delegate a transaction | | Adyen/Censuswide, May 2026 | In a survey of 2,000 French consumers, 42% were open to an AI-supported journey extending to payment. | Actual use, or unconditional trust in autonomous purchasing | | Checkout.com/Censuswide, March 2026 | In a survey of 2,002 French respondents, substantial groups expressed reluctance to delegate or uncertainty about who should manage an agent. | That customers reject all forms of AI-assisted discovery | These figures should not be averaged: using AI during research, accepting AI-supported payment and delegating a purchase are different behaviours. Together, the studies suggest both interest and hesitation. That favours assistance first - useful, transparent and easy to leave - while the business case comes from a specific journey problem, not adoption headlines. ## 4. Five signs your ecommerce site is not ready ### 1. Search and product data are already unreliable If customers cannot trust availability, variants, delivery dates or product attributes, the assistant will inherit the same weaknesses. Repair the source data and ordinary discovery path first. ### 2. There is no defined high-friction task “We need AI on the site” is not a use case. “Help customers choose a compatible replacement part without knowing the model number” might be. Start with a customer decision that is valuable, repeated and currently difficult. ### 3. The system cannot reach an authoritative answer Recommendations may depend on catalogue attributes, live stock, regional delivery, promotion rules, returns policies or compatibility data. If the assistant cannot retrieve the right source at the right time, it must not improvise. ### 4. Nobody owns multi-turn quality The first answer is only the start. Recommendations can forget constraints or deteriorate later in the comparison. A recent preprint shopping benchmark also found weaker performance on optional criteria and later turns. Test complete conversations against expected outcomes and assign an evaluation owner. ### 5. There is no measurement or stop decision Without a defined eligible audience, control group, event model and exit criteria, the team may celebrate usage without knowing whether the assistant improved the journey. Decide in advance what would justify expansion, correction or shutdown. ## 5. What the assistant needs behind the interface The visible conversation is only one layer. A dependable implementation needs five foundations: - **Structured, current product data.** Use stable product and variant identifiers, accurate attributes, price, availability and any delivery or returns information needed for the decision. The assistant should retrieve product facts, not reconstruct them from marketing copy. - **Approved sources.** Define which systems are authoritative for catalogue, stock, price and policy answers. The assistant must be able to say that it cannot confirm something. - **A hybrid experience.** Let customers open product pages, edit filters, compare alternatives and return to browsing. Preserve their constraints and explain recommendations with checkable facts. - **Privacy, transparency and control.** Make the AI interaction clear, explain relevant data use and do not request personal information that the task does not need. - **Operational ownership.** Assign owners for catalogue quality, model behaviour, analytics, incidents and commercial results. For France and the EU, the exact obligations depend on the implementation and data flow. [European Commission guidance](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act) states that relevant Article 50 AI transparency obligations apply from 2 August 2026. The CNIL and CIANum have also highlighted the additional privacy, cybersecurity and responsibility risks created by agentic systems with memory, external connections or complex action chains. This is not a substitute for a legal assessment. Map the provider, deployer, processor and data-controller roles for the actual product before launch. If the pilot may later become a permanent capability, use the same governance discipline described in our [AI pilot-to-production checklist](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/). ## 6. How to run a bounded pilot A useful pilot is small enough to diagnose and realistic enough to affect a genuine decision. 1. **Choose one shopping job.** Use service transcripts, search logs, zero-result queries, exits and customer research to find a repeated difficulty. Write it as an outcome: “Help a first-time buyer shortlist three compatible products within budget and explain the trade-offs.” 2. **Limit the catalogue and audience.** Select a category with usable data and meaningful choice. Define who is eligible and who remains in the comparison group. 3. **Define the answer contract.** Record approved sources, facts requiring live verification, prohibited inferences, uncertainty language, hand-off rules and the route out of the conversation. 4. **Build an evaluation set.** Include ambiguous briefs, conflicting constraints, unavailable products and follow-ups that change the request. Score constraint retention, accuracy, shortlist quality, explanation and final hand-off. 5. **Instrument the journey.** Connect assistant events to product views, baskets, orders, fulfilment, returns and support outcomes. Keep the unassisted route measurable. 6. **Launch gradually.** Use controlled exposure, review real failures and maintain a kill switch. ## 7. How to measure the result The measurement chain should follow the customer, not the interface: **Eligible visit → assistant exposure → activation → first useful result → constraint refinement → product view → basket → order → fulfilled order → return, cancellation or complaint** Use that chain to avoid five mistakes: - **Exposure is not activation.** Report who saw the assistant and who chose to use it. - **Conversation volume is not success.** More messages may mean engagement, confusion or recovery. - **A placed order is not a retained order.** Watch cancellations, returns and support contacts. - **Speed is not always efficiency.** A short session can be a fast decision or an early exit. - **Assisted and unassisted customers differ.** A controlled rollout or randomised experiment is stronger than comparing self-selected users after the fact. Before launch, set: - a primary customer or commercial outcome; - guardrail metrics for accuracy, returns, complaints and latency; - a minimum evidence window appropriate to your traffic and buying cycle; - explicit expand, repair and stop criteria. If the team cannot agree on those four items, the pilot is not ready. ## Conclusion An AI shopping assistant can improve a difficult product-finding journey. It can also add an impressive interface to unresolved catalogue, UX and measurement problems. The right starting question is not “Which model should we add?” It is “Which customer decision is currently hard, and can conversation improve it without reducing accuracy, control or trust?” If the answer is specific, the product data is dependable and the result can be measured through to the retained order, run a bounded pilot. If those foundations are missing, repair them first. If conversation adds no clear advantage over search, filters or comparison, do not build it. ## Turn one product-finding problem into a testable decision Bring us one journey where customers struggle to choose. In a 30-minute scoping call, we will examine the task, catalogue readiness, customer path, measurement requirements and risk boundaries. You will get an honest first view: pilot, repair the foundations, or keep the existing experience. [Book a scoping call with izzy.agency](https://izzy.agency/en/contact/) ## Frequently asked questions ### What is an AI shopping assistant for ecommerce? It is a conversational interface that helps customers discover, refine and compare products using the retailer's catalogue and rules. Unlike a support chatbot, its main job is product selection rather than order service. ### Should an AI shopping assistant replace site search? Usually not. Conversation is better suited to exploratory or constraint-heavy tasks. Search, filters, navigation and product pages remain more efficient for exact or simple requests. A hybrid experience lets the customer use the right tool at each point. ### Does an AI shopping assistant improve conversion? It may improve a specific discovery journey, but there is no universal conversion guarantee. The result depends on the use case, product data, recommendation quality, UX and measurement design. Test it against an appropriate comparison path and track downstream outcomes. ### What product data does the assistant need? The exact fields depend on the category. Common requirements include product and variant identifiers, attributes, compatibility, price, stock, delivery and returns information. Each fact should have an identified authoritative source. ### Is an AI shopping assistant subject to GDPR? GDPR can apply when personal data is processed. The obligations depend on the data collected, purpose, legal basis, vendors, retention, profiling and international transfers. The assistant may also fall within AI transparency requirements. Complete a data-flow and legal review for the actual implementation. ### What should an ecommerce AI pilot test first? Start with one repeated product-finding problem in a limited category. Test factual accuracy, constraint retention across follow-ups, shortlist usefulness, hand-off quality and the full journey to fulfilment and returns. ## Sources and evidence note This article was developed from independent research rather than from a single market article. Material sources included: - [FEVAD/Odoxa research on French consumers and AI-assisted shopping](https://www.fevad.com/les-francais-sont-de-plus-en-plus-nombreux-a-adopter-lia-pour-acheter-sur-internet/) - [Adyen France consumer research on AI-supported shopping](https://www.adyen.com/fr_FR/presse-et-medias/shopping-algorithmique-plus-de-2-francais-sur-5-feraient-confiance-a-l-ia-pour-faire-leurs-achats) - [Checkout.com France research on delegation and trust](https://www.checkout.com/fr-fr/newsroom/demande-achats-ia-consommateurs-francais) - [Conversational recommender systems: survey and research directions](https://arxiv.org/abs/2004.00646) - [Shopping Reasoning Bench: multi-turn product-recommendation evaluation](https://arxiv.org/abs/2606.12608) - [Microsoft Research randomised comparison of traditional and LLM-based consumer search](https://www.microsoft.com/en-us/research/publication/comparing-traditional-and-llm-based-search-for-consumer-choice-a-randomized-experiment/) - [NIM hybrid-store experiment on conversational shopping](https://www.nim.org/en/publications/detail/from-clicks-to-conversations) - [Baymard product-finding research](https://baymard.com/research/eCommerce-search) - [Google product structured-data guidance](https://developers.google.com/search/docs/appearance/structured-data/product) - [Shopify catalogue search documentation for agents](https://shopify.dev/docs/agents/get-started/search-catalog) - [European Commission guidance on AI Act transparency obligations](https://digital-strategy.ec.europa.eu/en/library/guidelines-transparency-obligations-providers-and-deployers-ai-systems) - [European Commission Q&A on Article 50 transparency obligations](https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act) - [CNIL and CIANum note on agentic AI](https://www.cnil.fr/fr/ia-agentique-cnil-cianum-note) The cited surveys use different samples and question wording, so their percentages are not directly comparable. The experimental and benchmark literature tests particular tasks and settings; it does not establish a universal ecommerce revenue effect. Legal and regulatory status should be rechecked immediately before publication. --- ### Why is your ecommerce website not converting? Audit the path from product data to fulfilled order URL: https://izzy.agency/en/blog/ecommerce-website-not-converting/ Published: 2026-08-11 Summary: Trace ecommerce conversion leaks from product data and discovery through checkout, fulfilment and analytics before adding more traffic or tools. Your ecommerce website may not have one conversion problem. It may have six connected handoffs: the catalogue describes a product, discovery surfaces it, the product page helps someone decide, checkout takes payment, operations fulfil the order and measurement records what happened. When one handoff fails, adding traffic, a marketplace, another plugin or an AI shopping feature can distribute the same fault more widely. A recent French retail commentary made the AI version of this point: automation exposes incomplete attributes, weak variant relationships and inconsistent categories rather than repairing them ([Journal du Net](https://www.journaldunet.com/retail/1552813-l-ia-ne-remplacera-pas-les-catalogues-produits-elle-revelera-surtout-leurs-faiblesses/)). That is a useful signal, not proof that AI is the cause. The underlying business problem is older and more practical: **can a customer move from a reliable product promise to a correctly fulfilled order?** ## Answer in 60 seconds If your ecommerce website gets visits, product views or cart additions but too few verified orders, do not begin with a redesign or a benchmark conversion rate. Trace one representative order through six handoffs: 1. **Catalogue truth:** are identifiers, variants, attributes, price and stock correct? 2. **Discovery:** can search, filters, feeds and external channels find the right product? 3. **Product decision:** can the buyer confirm fit, cost, delivery, returns and the correct variant? 4. **Cart and checkout:** can they complete the purchase on their device, in their market and with a valid payment method? 5. **Order and fulfilment:** do payment, status, email, stock, delivery or download access complete correctly? 6. **Measurement:** do analytics events reconcile with payment and order records? Find the first repeated break, fix the smallest affected part and run the same scenario again. More traffic makes sense only after the path can keep the promise. ## In this guide - [1. Define conversion as a verified business outcome](#1-define-conversion-as-a-verified-business-outcome) - [2. Check whether the catalogue tells one product truth](#2-check-whether-the-catalogue-tells-one-product-truth) - [3. Test discovery and the product decision together](#3-test-discovery-and-the-product-decision-together) - [4. Test checkout as a system, not a screen](#4-test-checkout-as-a-system-not-a-screen) - [5. Follow the order after payment](#5-follow-the-order-after-payment) - [6. Reconcile the funnel with real orders](#6-reconcile-the-funnel-with-real-orders) - [7. Turn observations into a prioritised repair](#7-turn-observations-into-a-prioritised-repair) ## 1. Define conversion as a verified business outcome An `add_to_cart` event is not an order. A checkout success page is not proof that payment settled. A paid order is not complete if the stock, confirmation or promised download never follows. Start with the outcome the business can verify: a valid order linked to the correct payment and fulfilment state. For subscriptions, bookings or digital products, include the activation, renewal or access state that matters. Then choose three representative scenarios: - a high-volume product and common payment method; - a difficult variant, promotion, shipping zone or device; - a high-consequence case such as a refund or digital delivery. Record the expected result at every handoff. This gives the team a testable path instead of a general feeling that “conversion is low”. ## 2. Check whether the catalogue tells one product truth A product exists in the catalogue or PIM, storefront, structured data, merchant feeds, marketplaces, inventory services and checkout. Those surfaces should not disagree about what the customer can buy. Check: - stable product and variant identifiers; - complete selection attributes such as size, material, dimensions and compatibility; - the same price, currency, promotion conditions and tax treatment; - current availability and delivery expectations; - clear returns, warranty and restriction information; - a named source of truth and update owner for each critical field. Google recommends using product structured data, a Merchant Center feed or both. It says using both can help it understand and verify product information across surfaces ([Google Search Central](https://developers.google.com/search/docs/appearance/structured-data/product)). Its Merchant Center specification also requires price and availability to match the landing page, structured data and checkout ([Google Merchant Center](https://support.google.com/merchants/answer/7052112?hl=en)). The wider commercial principle is simple: a channel cannot reliably sell a product record the business cannot keep consistent. Sample one product family and find where the first contradiction is created before rewriting every description. ## 3. Test discovery and the product decision together A visitor cannot buy a product they cannot find. They also cannot confidently choose a product whose important constraints are hidden. Run task-based searches instead of checking whether a category page “looks good”: - “Show me the version compatible with X.” - “Find the item available in size Y and deliverable to Z.” - “Compare the two variants under a fixed budget.” Follow each task through search, filters, product page and selected variant. Check whether the URL, image, price, stock and delivery message update together. Zero results for a valid attribute suggest taxonomy or indexing. A correct result without enough information to verify fit suggests the product record or decision design. A correct page with a different external-feed price points upstream. A prettier page will not reconnect a variant, and a new search tool will not create missing attributes. ## 4. Test checkout as a system, not a screen Cart abandonment is not one diagnosis. Some people are comparing or saving items and were never ready to purchase. Baymard's 2025 checkout research explicitly separates that natural behaviour from avoidable checkout usability problems; its findings come from moderated usability sessions, eye tracking, checkout benchmarking and quantitative studies ([Baymard Institute](https://baymard.com/blog/ecommerce-checkout-usability-report-and-benchmark)). Its market average cannot tell you what is broken on your store. Test the conditions that can change the outcome: - mobile and desktop, guest and account checkout; - main markets, currencies and delivery zones; - standard, discounted and invalid promotion codes; - stock changes and each material payment method; - success, failure, cancellation and retry paths. At each step, inspect total cost, delivery timing, required fields, validation, loading, payment response and recovery. An error should tell the buyer what happened and what to do next without deleting valid information. Accessibility belongs in this test, not in a separate design review. W3C's form guidance recommends explicit labels and instructions and asking only for information required to complete the transaction or process ([W3C Web Accessibility Initiative](https://www.w3.org/WAI/tutorials/forms/)). The objective is a checkout that asks for necessary information, explains the commitment and survives real customer conditions. ## 5. Follow the order after payment Payment gateways, tax services, stock systems, email providers, warehouses and download permissions can all change what the customer receives. For each test order, confirm: 1. gateway and commerce platform agree on payment state; 2. one correct order is created and stock updates; 3. customer and operational notifications arrive; 4. fulfilment, booking, subscription or download access starts at the right status; 5. cancellation, refund and retry produce the expected reversal; 6. support can see enough evidence to resolve a failure. WooCommerce's own documentation recommends test orders for evaluating payment methods, checkout and related integrations. It warns that test orders can trigger emails and appear in analytics, and recommends controlled testing on staging ([WooCommerce: testing orders](https://woocommerce.com/document/managing-orders/testing-orders/)). Digital fulfilment adds a status dependency. WooCommerce documents that download access can begin after payment or wait until an order is complete, depending on product and store settings ([WooCommerce: downloadable products](https://woocommerce.com/document/digital-downloadable-product-handling/)). The customer must receive the promised asset under the intended access rules. ## 6. Reconcile the funnel with real orders Analytics should locate a break, not replace the transaction record. Google Analytics provides ecommerce events for distinct behaviours such as viewing a product, adding it to a cart and purchasing ([Google Analytics](https://support.google.com/analytics/answer/14434488?hl=en)). Build a funnel for the events your implementation genuinely sends, then reconcile it with: - commerce-platform orders; - gateway payments, failures and refunds; - fulfilment or access records; - test and duplicate exclusions; - consent, tag and release changes. If `purchase` events exceed valid orders, inspect duplicate event firing and failed-payment handling. If orders exceed analytics purchases, inspect consent, redirects, cross-domain payment and collection. If analytics and orders agree but customers still complain, inspect fulfilment and support rather than the acquisition dashboard. Do not prescribe a redesign from one site-wide rate and a public average. Where data is sufficient, segment by product family, device, market, source and release date. Ask: **where does the first verified mismatch appear?** ## 7. Turn observations into a prioritised repair A useful ecommerce conversion audit should finish with an evidence-backed backlog, not 80 generic best practices. For every material finding, record: | Finding | Business consequence | Evidence | Smallest useful repair | Retest | |---|---|---|---|---| | Variant is missing from filters | Valid products are not discoverable | Attribute and search comparison | Correct mapping and re-index affected family | Repeat selection task | | Feed and checkout prices disagree | Promise changes before payment | Feed, page and test order | Repair source and synchronisation rule | Recheck three surfaces | | Payment succeeds but download is absent | Paid customer cannot use the product | Gateway, order state and access log | Correct state/access configuration | Place controlled test order | | `purchase` fires twice | Funnel and campaign reporting are overstated | Debug trace and order record | Deduplicate implementation | Reconcile events with orders | Prioritise by consequence, recurrence, reach and reversibility. Repair a repeated payment or fulfilment failure before polishing a low-traffic page. Fix the source of bad product data before manually correcting every destination. ## Conclusion: repair the first broken handoff An ecommerce conversion problem is not automatically a traffic, design or checkout problem. Trace one product promise from source data to fulfilled order. Compare what the customer sees with what the systems record. Find the first repeated break, make one bounded repair and run the same scenario again. Then decide whether the next investment should be product data, discovery, UX, engineering, integration, measurement - or no project yet. ## Bring us one product family. We'll trace the conversion chain. Bring a priority product family, three buying scenarios, recent funnel data and access to the relevant order evidence. IZZY can scope an ecommerce UX and product-data readiness audit, locate the first material breaks and turn them into a prioritised implementation sprint. The result is a decision-ready repair plan - not a promised conversion rate or a list of plugins. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Why is my ecommerce website getting traffic but no sales? Traffic may not match the product; discovery, checkout, payment, fulfilment or analytics may fail. Trace one order from source to fulfilment before buying more traffic. ### What is an ecommerce conversion audit? It reviews how a customer discovers, evaluates, buys and receives a product - and how those steps are measured. It should identify specific breaks, evidence, consequences and the smallest repair to test. ### Can poor product data reduce ecommerce conversion? Missing or inconsistent attributes, variants, price, stock or delivery information can obstruct discovery and decisions. Verify the effect through catalogue, behaviour and order evidence. ### How do I know whether my checkout is broken? Run controlled scenarios across relevant devices, markets, payment methods and outcomes. Compare the browser, gateway, order, notifications, fulfilment and analytics - not only the success page. ### Should we redesign the store to improve conversion? Not before locating the first material break. Redesign can address repeated navigation or interaction failures; it will not fix inventory, variant, payment-state or analytics errors. ### How should downloadable ecommerce products be tested? Place a controlled order and verify payment, order state, email, access, file permission, refund behaviour and support visibility. Test integrations on staging where possible. ## Sources and evidence note Primary documentation was checked on 22 July 2026. JDN and the WooCommerce support item were treated as editorial signals, not demand proof. Baymard's aggregate research gives context; it does not diagnose one store. Recheck changing platform details before implementation. --- ### An AI share link is not private: what should your company control? URL: https://izzy.agency/en/blog/ai-share-link-privacy/ Published: 2026-08-10 Summary: Claude, ChatGPT or another AI assistant: learn what a shared link can expose and how to control access, publication, revocation and incidents. You prepare a summary in Claude, ChatGPT or another AI assistant. To ask a colleague for feedback, you click Share. You think you are sending a document to one person. Depending on the product and account, you may have created a page that anyone with the link can open. The problem is not that every AI conversation is public. The problem is sharing without knowing **which object leaves the workspace, who can open it and how access will end**. ## Answer in 60 seconds - A private conversation, an internal share and an anyone-with-the-link page are not the same space. - “Anyone with the link” does not mean “approved recipients only”. - What becomes visible depends on the object and account: conversation, artifact, file, export or connected-tool output. - Revoking a link, unpublishing an artifact and removing a search result are separate actions. - A useful AI policy gives staff a workable route: what they may share, where, who owns it and what to do when something goes wrong. ## In this article 1. [What the Claude incident shows](#1-what-the-claude-incident-shows) 2. [Private, internal, public or indexed](#2-private-internal-public-or-indexed) 3. [What becomes visible depends on the object](#3-what-becomes-visible-depends-on-the-object) 4. [Use the Share button test](#4-use-the-share-button-test) 5. [What if the link has already circulated?](#5-what-if-the-link-has-already-circulated) 6. [What your AI use policy should add](#6-what-your-ai-use-policy-should-add) 7. [Conclusion](#conclusion-control-the-exit-not-only-the-input) 8. [Frequently asked questions](#frequently-asked-questions) ## 1. What the Claude incident shows Anthropic says Claude chats are private by default. On Free, Pro and Max accounts, a user can then create a snapshot that anyone with the link can view ([Anthropic](https://support.claude.com/en/articles/10593882-share-and-unshare-chats)). Some of those public pages appeared in Google or Bing. Anthropic said it did not provide a directory or sitemap. A link could still be discovered after somebody placed it somewhere a search engine could crawl ([WIRED](https://www.wired.com/story/private-claude-chats-exposed-in-google-and-bing-search-results/)). The exact technical cause remains uncertain. Checks performed at different times found different configurations. This should not be reduced to a simple story about one missing tag ([Search Engine Journal](https://www.searchenginejournal.com/indexed-claude-chats-show-why-disallow-is-not-noindex/583852/)). The more useful lesson is stable: **a public link can travel beyond its first recipient**. This is not unique to Claude. OpenAI warns that a copy imported by another user may remain in that user’s history after the original ChatGPT shared link is deleted ([OpenAI](https://help.openai.com/en/articles/7925741-chatgpt-shared-links-faq)). Researchers have also studied public links from several assistant platforms. That establishes a cross-platform publication surface; it does not prove that every share was accidental ([ShareChat](https://arxiv.org/abs/2512.17843)). ## 2. Private, internal, public or indexed Inside a team, the word “shared” sounds reassuring. In a product, it can describe four different states. | State | Who can open it? | Useful control | | --- | --- | --- | | **Private** | Authorised users | Identity, permissions and revocation | | **Internal** | Signed-in organisation members | Membership, project access and attachments | | **Anyone with the link** | Anyone who receives or finds the URL | Clean content, owner and inventory | | **Indexed or copied** | Search users, an archive or somebody holding a copy | Source revocation, search removal and copy handling | A difficult-to-guess URL is not access control. It can be forwarded, published or retained elsewhere. Google also explains that a `noindex` instruction works only when its crawler can read it. A `robots.txt` rule controls crawling; it does not make a public page confidential ([Google Search Central](https://developers.google.com/search/docs/crawling-indexing/block-indexing)). For sensitive information, the relevant control remains access: authentication, limited permissions or no public publication. ## 3. What becomes visible depends on the object Do not rely on the word Share. Check the object and account. | Object | Audience currently described by Anthropic | Point to check | | --- | --- | --- | | **Free, Pro or Max chat** | Anyone with the link | Earlier messages and displayed artifacts are visible | | **Consumer-published artifact** | Anyone with the link | It can be used, copied or embedded elsewhere | | **Team or Enterprise chat or artifact** | Authenticated organisation members | Project rights and attachments still matter | For a consumer chat, Anthropic says the original attached file is not included in the snapshot. The conversation and Claude’s answers remain visible. Information quoted or summarised from the file can therefore appear on the shared page. For an artifact shared inside Team or Enterprise, Anthropic says authorised viewers may gain access to attachments from the source conversation ([Anthropic - artifacts](https://support.claude.com/en/articles/9547008-publish-and-share-artifacts)). “Files are never shared” would therefore be the wrong rule. Apply the same check to screenshots, exports, copied text and connected-tool output. ## 4. Use the Share button test Before creating a link, answer five questions. 1. **What does the recipient need?** The complete conversation, one answer or only the conclusion? 2. **What is actually on the page?** Customer or employee data, contracts, non-public pricing, code, credentials, internal documents or strategy? 3. **Who can really open it?** One person, a project, the organisation or anyone with the link? 4. **Who remains responsible?** Who records, reviews and revokes the link? 5. **What private alternative should staff use?** An internal project, permissioned folder, business workspace or cleaned document? OWASP identifies sensitive-information disclosure through LLM inputs and outputs as a risk that requires data handling and access controls ([OWASP](https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/)). NIST’s voluntary AI Risk Management Framework also places governance, context, measurement and response across the AI lifecycle; using it does not prove a control works in your environment ([NIST](https://www.nist.gov/itl/ai-risk-management-framework)). The inventory does not need to duplicate the content. Product, object, audience, owner and status may be enough. Prohibition without a usable route encourages shadow tools. The approved path should be the easiest path. ## 5. What if the link has already circulated? Start by reducing the exposure. 1. Revoke the chat share and unpublish related artifacts separately. 2. Preserve the useful facts without redistributing the content. 3. Identify what was visible and who could access it. 4. Rotate exposed passwords, keys or tokens immediately. 5. Look for public posts, search results, previews and copies. 6. Involve the privacy lead, security lead, legal counsel, HR or contract owner according to the content. 7. Fix the route that allowed the mistake. Removing a Google result does not remove the source page. Disabling the source does not guarantee the immediate disappearance of every copy. Track the actions separately ([Google](https://support.google.com/websearch/answer/11080680?hl=en)). If personal data was exposed without authorisation, the organisation may need a breach assessment under the laws that apply to it. Under the GDPR, documentation, authority notification and communication to affected people depend on the facts and level of risk; there is no universal response for every public link ([European Data Protection Board](https://www.edpb.europa.eu/topics/security-data-breaches/personal-data-breaches_en)). ## 6. What your AI use policy should add Many policies name approved tools and prohibited inputs. They often miss one question: **how does a conversation leave the tool?** Add: - approved accounts and sharing functions; - prohibited data and audiences; - covered objects: conversations, artifacts, files, exports and screenshots; - an owner and review or revocation trigger; - the approved private alternative; - an incident route and the specialists to involve. For organisations operating in the EU, [Article 4 of the AI Act](https://eur-lex.europa.eu/eli/reg/2024/1689/2024-07-12/eng/) requires AI literacy measures suited to people and context. Showing staff what the real Share control does is more useful than generic awareness. A policy or training session does not, by itself, establish compliance or security. For the wider operating model, see [our AI governance checklist for SMEs](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/). ## Conclusion: control the exit, not only the input A useful policy does more than tell staff what they may ask an AI assistant. It explains how to share, with which audience, who remains responsible and how access ends. Take one real sharing route. Check it from end to end. Fix the first missing control. ## Bring us one sharing path. We’ll give you an honest read. Bring the approved-tool list, current policy and one sharing route without sensitive content. IZZY can map the move from private conversation to internal collaboration or public publication, then identify the first missing operating control. The answer may be a clearer rule, an account setting, a private workspace, training, an [Agent Readiness Audit](https://izzy.agency/en/services/agent-readiness-audit/), a more controlled [AI integration](https://izzy.agency/en/services/ai-integration/) or a specialist decision. If a larger engagement is not justified, the call should make that clear too. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is a Claude conversation public by default? No. Anthropic describes it as private by default. A user must create a public snapshot or publish an artifact. Team and Enterprise currently use organisation-only sharing. ### Is an anyone-with-the-link page private? No. It may require neither identity nor individual authorisation. The link can be forwarded, published, archived or copied. ### Does disabling the link remove the conversation from Google? Not necessarily. Revoking the source, removing a search result and dealing with external copies are separate actions. ### Does every public AI link create a reportable data breach? No. That depends on the content, audience, applicable law and risk. Preserve the facts and involve the responsible privacy or legal specialist. ### Is an AI use policy enough? No. Connect it to the accounts people actually use, product settings, a private route, an inventory and a workable incident process. ## Sources - [Anthropic - Share and unshare chats](https://support.claude.com/en/articles/10593882-share-and-unshare-chats) - [Anthropic - Publish and share artifacts](https://support.claude.com/en/articles/9547008-publish-and-share-artifacts) - [WIRED - Private Claude Chats Exposed in Google and Bing Search Results](https://www.wired.com/story/private-claude-chats-exposed-in-google-and-bing-search-results/) - [Search Engine Journal - Indexed Claude Chats Show Why Disallow Is Not Noindex](https://www.searchenginejournal.com/indexed-claude-chats-show-why-disallow-is-not-noindex/583852/) - [Google Search Central - Block Search indexing with noindex](https://developers.google.com/search/docs/crawling-indexing/block-indexing) - [Google - Remove web results from Google Search](https://support.google.com/websearch/answer/11080680?hl=en) - [OpenAI - ChatGPT Shared Links FAQ](https://help.openai.com/en/articles/7925741-chatgpt-shared-links-faq) - [OWASP - Sensitive Information Disclosure](https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/) - [NIST - AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) - [European Data Protection Board - Personal data breaches](https://www.edpb.europa.eu/topics/security-data-breaches/personal-data-breaches_en) - [European Union - Artificial Intelligence Act](https://eur-lex.europa.eu/eli/reg/2024/1689/2024-07-12/eng/) - [ShareChat - A Dataset of Chatbot Conversations in the Wild](https://arxiv.org/abs/2512.17843) *Sources checked 29 July 2026. This article provides general operational guidance. It is not legal advice, a compliance assessment, a security audit or a forecast of commercial results.* --- ### How to automate meeting notes to tasks with n8n without duplicate tickets URL: https://izzy.agency/en/blog/meeting-notes-to-actions-n8n/ Published: 2026-08-09 Summary: Turn Google Meet notes into approved, owned tasks with n8n, Notion and Linear - without duplicate tickets, invented owners or hidden workflow failures. Your meeting notes exist, but follow-up still disappears: someone copies actions by hand, two systems create the same ticket, or a task arrives without an owner or context. Yes, you can automate meeting notes to tasks with n8n - but not by turning every AI-detected action into a ticket. A controlled handoff retrieves authorised evidence, proposes actions, requires approval for consequential writes, creates one execution record and writes back the result. Original Drive or Docs records remain authoritative. Notion indexes meetings and control state; it does not copy every artefact. In this reference design, Linear owns execution, with bounded Jira or GitHub variants. This is an evidence-led reference architecture, not a deployed IZZY client case, benchmark or ROI result. Tenant permissions, schemas, latency and failure behaviour need implementation testing. ## The answer in 60 seconds - The stack can support this handoff, subject to access, configuration and testing. - Meet AI notes are fallible summaries, not verbatim. - Notion holds review state; original records stay authoritative. - Claude proposes. A person approves consequential creation. - Create each approved action once, then write it back. - Give retries, failures and notifications explicit controls. ## In this article 1. [Feasible does not mean fully automatic](#1-yes-the-workflow-is-feasible---but-automatic-is-the-wrong-design-goal) 2. [Give every system one job](#2-give-every-system-one-job) 3. [The reference workflow](#3-the-reference-workflow-end-to-end) 4. [Define the meeting record](#4-define-the-meeting-record-before-automating-it) 5. [Let Claude propose](#5-let-claude-extract-proposals-not-silently-create-truth) 6. [Route to one execution system](#6-route-each-action-to-one-execution-system) 7. [Make retries safe](#7-make-retries-safe-and-failures-visible) 8. [Roll out in three stages](#8-roll-out-in-three-stages) ## 1. Yes, the workflow is feasible - but “automatic” is the wrong design goal [Google documents](https://support.google.com/meet/answer/14754931?hl=en) that generated meeting notes are saved in Docs, attached to the Calendar event and placed in the organiser's Drive. Google warns that notes may be incomplete, inaccurate or absent. Access varies by edition, host settings and sharing. AI notes are generated summaries, not verbatim; accuracy is not assured. A [Meet transcript](https://support.google.com/meet/answer/12849897?hl=en-GB) is a separate spoken-word artefact; a [recording](https://support.google.com/meet/answer/9308681/record-a-video-meeting?hl=en-GB) is a separate video artefact. None is an approved commitment. Google [states](https://workspace.google.com/intl/en_id/security/ai-privacy/) that Gemini for Workspace respects access controls and does not use customer data without permission to train or improve models outside Workspace. This vendor statement is not independent assurance. ## 2. Give every system one job | System | One job | Record it owns | It must not become | |---|---|---|---| | Drive/Docs | Preserve evidence | Source | Task tracker | | Notion | Index meetings/review | Control record | Source copy or second task truth | | Linear | Govern work | Primary issue | Meeting archive | | Jira/GitHub | Serve bounded workflows | Jira/repo issue | Parallel defaults | | n8n/Claude | Orchestrate/propose | Workflow/proposal | Business authority | | Slack/Mattermost | Notify after approval | Message/link | The work record | An activated [n8n Drive Trigger](https://docs.n8n.io/integrations/builtin/trigger-nodes/n8n-nodes-base.googledrivetrigger) regularly checks; it does not promise instant delivery. A Drive notification announces change, not details or completed processing; [retrieve details](https://developers.google.com/workspace/drive/api/guides/manage-changes). Channels [expire and need renewal](https://developers.google.com/workspace/drive/api/guides/push). Request the smallest OAuth scopes needed and add scope in context, as [Google recommends](https://developers.google.com/identity/protocols/oauth2/resources/best-practices). Consent depends on the integration. ## 3. The reference workflow, end to end 1. Detect a Drive artefact through a tested trigger/feed. 2. Validate type, identity, eligibility, access and completeness. 3. Deduplicate the event with a stable source identity. 4. Create or update the Notion meeting record. 5. Retrieve evidence through documented [Drive](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.googledrive) or [Docs](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.googledocs) operations. 6. Ask Claude for proposals, not commitments. 7. Show evidence/ambiguity for review. 8. Route an approved action to one execution system. 9. Write its ID, destination and result back. 10. Notify after approval; route failures visibly. n8n can [pause selected AI tool calls](https://docs.n8n.io/build/integrate-ai/ai-examples/human-in-the-loop-for-tools), [resume workflows](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.wait) on time or an event and [route failures](https://docs.n8n.io/build/flow-logic/handle-errors-gracefully). The workflow needs explicit approval rules, wait timeouts, retry limits and a named failure owner. ## 4. Define the meeting record before automating it | Field | Purpose/rule | |---|---| | `meeting_id` | Stable source identity | | `date` | Time/timezone | | `source_url` | Authorised source link | | `participants` | In-scope people | | `confidentiality_class` | Handling class | | `decisions` | Reviewed proposals | | `actions` | Evidence-linked actions | | `evidence_locator` | Passage or timestamp | | `owner_state` | `proposed`, `confirmed`, `unknown`, `not-applicable` | | `destination_system` | Approved destination | | `external_task_id` | Returned ID | | `sync_state` | `pending`, `created`, `updated`, `failed`, `conflict`, `replay-required` | | `review_status` | `unreviewed`, `changes-requested`, `approved`, `rejected`, `expired` | [n8n and Notion](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.notion) can create/query records with access and matching properties. Use `meeting_id` and a [data-source query](https://developers.notion.com/reference/query-a-data-source), not fuzzy titles: API [search is title-oriented](https://developers.notion.com/reference/post-search), can lag and is [not exhaustive](https://developers.notion.com/reference/search-optimizations-and-limitations). Optional [Enterprise Search](https://www.notion.com/en-gb/help/enterprise-search) and [AI Connectors](https://www.notion.com/help/notion-ai-connectors) depend on plan, permissions, configured connectors, indexing and model choice. Notion remains a control/index layer, not a universal retrieval promise. ## 5. Let Claude extract proposals, not silently create truth ```json { "proposals": [ { "action_text": "Keep shadow mode", "decision_or_action": "decision", "owner_name": null, "owner_state": "not-applicable", "deadline": null, "deadline_state": "absent", "destination_hint": "Notion", "evidence_locator": "Decisions, item 1", "ambiguity_flags": [], "confidence_band": "high", "review_required": true }, { "action_text": "Confirm launch owner", "decision_or_action": "action", "owner_name": null, "owner_state": "unknown", "deadline": null, "deadline_state": "absent", "destination_hint": "Linear", "evidence_locator": "Actions, item 2", "ambiguity_flags": ["owner_missing"], "confidence_band": "medium", "review_required": true } ] } ``` [Anthropic structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) can constrain supported models to a schema; refusal and schema failures need handling. Unknown owners stay `null`; “soon” is not a date; ambiguity needs clarification; confidence grants no permission. For client tools, Claude requests; the application validates, authorises and executes, as [Anthropic explains](https://platform.claude.com/docs/en/agents-and-tools/tool-use/how-tool-use-works). Human approval precedes task creation and person-level notification. See IZZY's [security controls](https://izzy.agency/en/blog/ai-agent-security-product-controls/). ## 6. Route each action to one execution system | Destination | Use when | Do not use when | Required write-back | |---|---|---|---| | Linear (primary) | Delivery runs there | Authority is elsewhere | ID/URL, team, owner, result | | Jira (variant) | Jira governs delivery | It creates a second tracker | Key/URL, project, state, result | | GitHub (variant) | Work is repo-scoped | Work lacks repo context | Number/URL, repository, result | | Notion | Decision/evidence/review | Execution is expected | Review state and task link | n8n documents issue creation for [Linear](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.linear), [Jira](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.jira) and [GitHub](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.github), subject to target rules. One action gets one authoritative item. Notion indexes it; notifications link it. [Slack](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.slack) or [Mattermost](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.mattermost) can notify; chat never owns completion. ## 7. Make retries safe and failures visible | Failure | Consequence | Response | |---|---|---| | Duplicate | Parallel tickets | Key, fingerprint, destination lookup | | Partial write | Conflicting records | Write back, notify, reconcile | | Denied access | Stalled work | Error queue, owner, replay decision | | Schema change | Wrong fields | Stop, remap, reapprove | The [Remove Duplicates node](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.removeduplicates) compares items or history; it is not end-to-end idempotency. An [Error Trigger](https://docs.n8n.io/build/flow-logic/handle-errors-gracefully) can start failure handling. The minimal error queue stores `meeting_id`, `action_fingerprint`, `stage`, `error_class`, `last_attempt`, `next_owner` and `replay_eligibility` - never secrets. Alerts go to the operating owner. Before replay, test denied access, expired credentials, unavailable APIs, schema changes, duplicate events and partial success; then re-check approval, destination and identity. Use least-necessary access to read, propose, approve, write and notify. See IZZY's [permissions guide](https://izzy.agency/en/blog/ai-agent-permissions-access-control/). n8n [redaction](https://docs.n8n.io/deploy/host-n8n/configure-n8n/security/redact-execution-data) has edition/version and coverage limits; [external secrets](https://docs.n8n.io/administer/manage-credentials/use-external-secret-stores) are plan-limited and credential-only. For personal data, minimise collection and justify retention. CNIL guides [minimisation](https://www.cnil.fr/fr/minimiser-les-donnees-collectees), [retention](https://www.cnil.fr/fr/passer-laction/les-durees-de-conservation-des-donnees) and [access profiles](https://www.cnil.fr/fr/securite-gerer-les-habilitations). This does not establish compliance. ## 8. Roll out in three stages | Stage | Permitted reads/writes | Reviewer responsibility | Evidence to collect | Rollback condition | |---|---|---|---|---| | **Shadow extraction** | Read artefacts; write isolated proposals. No task/person notification. | Compare with human follow-up; mark unsupported/ambiguous items. | Links, corrections, unknowns, access/schema failures. | Return to manual work if traceability, review or handling fails. | | **Approval-required creation** | After named approval, write one task, write back, then notify. | Approve destination, body, owner and deadline; confirm result. | Decision, parameters, response, conflicts, order, errors, replay. | Disable writes if approval, identity, mapping, write-back, order or recovery fails. | | **Bounded automation** | Pre-approved patterns; review exceptions/changes. | Own patterns, monitor drift, suspend automation. | Matches, overrides, conflicts, failures, permissions, fallback. | Revert when pattern, provider, permissions, consequence, monitoring or fallback changes. | This is not deployment evidence. Approval-required creation begins only after shadow evidence demonstrates traceability, usable review and safe handling. Bounded automation begins only after a named owner accepts evidence for approval binding, identity, write-back, notification order, failure recovery and fallback for the selected pattern. n8n offers Git-backed [source control and environments](https://docs.n8n.io/administer/use-source-control-and-environments) with plan/role limits; rollback remains separate. Use workflow evidence - not a universal sample, accuracy or time threshold - and apply the [pilot-to-production checklist](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/) as scope expands. ## Conclusion: automate the handoff, not the judgment Use this workflow when handoff failure is visible and systems have clear jobs. Without decisions or owners, integration moves ambiguity faster. Automation preserves evidence, routes one approved action and exposes failure. People decide whether a commitment exists, who accepts it and where it belongs. ## Map one meeting-to-action workflow > Bring one recurring meeting, the current notes destination, the task system, your ownership rules and two recent examples of missed or duplicated follow-up. IZZY can map the workflow and scope the smallest useful pilot. If the real problem is meeting discipline rather than integration, the right result may be no implementation. This is exactly the scope of our [n8n AI Automation service](https://izzy.agency/en/services/automatisation-ia-n8n/). [Book a 30-minute scoping call](https://izzy.agency/en/contact/). ## Frequently asked questions ### Can Google Meet notes be turned into tasks with n8n? Yes, subject to access, credentials, schemas and testing. Notes may be incomplete, inaccurate, absent or non-verbatim. ### Are Google Meet notes the same as a meeting transcript? No. AI notes are generated summaries; a transcript is a separate spoken-word artefact. ### Should meeting actions live in Notion or Linear? Notion indexes the meeting and review; the authoritative tracker owns execution. Here, that tracker is Linear. ### How do you stop n8n from creating duplicate tickets? Use stable identities, destination lookup, write-back and replay checks. Duplicate removal alone is not exactly-once delivery. ### When can the workflow create tasks without human approval? Not initially. Consider bounded automation only for a stable, low-risk pattern with named ownership, monitoring and fallback. ## Sources and method Research was checked on 2026-07-22. Official pages support bounded capabilities, not outcomes; links are adjacent. This is an IZZY reference architecture. No tenant or end-to-end workflow was tested. Latency, duplicates, accuracy, writes, notification order, outcomes, ROI and legal compliance remain unverified. --- ### Google Ads vs Meta Ads for small businesses: what to check before you spend URL: https://izzy.agency/en/blog/google-ads-vs-meta-ads-small-business/ Published: 2026-08-03 Summary: Google Ads or Meta Ads? Check your offer, margins, landing page, measurement and follow-up before buying more traffic for your small business. ## The answer in 60 seconds Do not choose between Google Ads and Meta Ads first. Check whether your business is ready to turn paid attention into a customer. You are probably ready for a limited test when: 1. you know which offer you are promoting, to whom and in what buying situation; 2. you know how much contribution a new customer can leave after variable costs; 3. the landing page continues the promise made in the advert; 4. you can measure a sale or useful enquiry - not only a click or form submission; 5. every enquiry or order has a definition, an owner and a recorded outcome; 6. the business can handle the volume without weakening service; 7. your tracking and audience use follow the rules that apply in the markets you serve. Google Search is usually strongest when a potential customer can already describe the need and search for it. Meta can introduce a relevant offer before that search exists. Neither platform repairs a vague offer, weak economics, an unconvincing page or neglected enquiries. The practical first move is a bounded test: one commercial hypothesis, one main platform and stop, repair or continue criteria written before the spend. ## In this guide - [1. Is visibility really the constraint?](#1-is-visibility-really-the-constraint) - [2. Seven checks before buying traffic](#2-seven-checks-before-buying-traffic) - [3. Google Ads or Meta Ads: how should a small business choose?](#3-google-ads-or-meta-ads-how-should-a-small-business-choose) - [4. How to set a test budget without inventing a magic number](#4-how-to-set-a-test-budget-without-inventing-a-magic-number) - [5. A practical 30-day paid-advertising test](#5-a-practical-30-day-paid-advertising-test) - [6. Is your business ready?](#6-is-your-business-ready) ## 1. Is visibility really the constraint? Paid advertising is attractive because the spend is visible and the platform produces immediate numbers. That does not mean reach is the first problem to solve. International evidence supports the need for better marketing and measurement capability, but not the performance of any particular campaign. In the OECD's 2025 survey of 1,009 platform-using SMEs across ten countries, 42% selected digital marketing and SEO as a training need and 33% selected data analytics. The sample is not representative of all SMEs, so treat it as a directional skills signal - not an advertising benchmark ([OECD](https://www.oecd.org/content/dam/oecd/en/networks/oecd-digital-for-smes-global-initiative/D4SME-2025-Policy-Highlights.pdf)). Before buying reach, separate four situations: | What you observe | First question | Possible first move | |---|---|---| | Few people know the offer, but relevant prospects already buy | Can we reach more people in the same buying situation? | Test one paid route | | Prospects do not quickly understand the offer | Does the message name a clear problem, buyer and next step? | Clarify the offer | | Enquiries arrive, but few are suitable | Is the advert or form attracting the wrong need, location or budget? | Repair targeting, message or qualification | | Suitable enquiries arrive, then disappear | Who responds, when and with what information? | Repair ownership and follow-up | Advertising buys an opportunity to be seen or considered. It does not create commercial clarity, trust, operational capacity or sales discipline. ## 2. Seven checks before buying traffic ### 1. One offer, one buyer and one situation “Promote the company” is not a testable brief. Choose one offer, one buyer, one recognisable situation and one next step. Complete this sentence: > We help **[specific customer]** facing **[observable problem]** move towards **[valuable outcome or next step]** through **[specific offer]**. If sales, delivery and marketing complete the sentence differently, clarify it before choosing a platform. Otherwise, the campaign will generate activity without telling you which promise worked. ### 2. Acquisition economics the business can support Start with what a customer can contribute, not a public “average cost per lead”. For a product business, estimate: > **First-order contribution before acquisition = net revenue − product cost − payment fees − fulfilment − merchant-funded delivery − expected returns and service cost** For a service business, use expected gross contribution from the initial engagement and include the cost of delivery. Adjust the logic for your accounting model, cash timing and credible repeat value. Then calculate the observed result of the test: > **Observed customer acquisition cost = attributable test costs ÷ new customers attributed to the test** Media is not the only test cost. Include creative production, landing-page work, tools and external management where they materially change the decision. Do not count uncertain lifetime value as cash already earned. ### 3. A landing page that keeps the advert's promise The advert, audience or search term creates an expectation. The landing page should continue the same offer, customer context and next step. Check the page on a phone: - Can a suitable visitor understand the offer without reading the entire page? - Are price, geography, eligibility or engagement conditions clear enough to prevent obviously unsuitable enquiries? - Is the main action visible, usable and connected to a real operational process? - Does the page supply proof and answer the concern most likely to block this decision? Google describes landing-page experience in terms including useful and relevant information, ease of navigation and whether the page meets the expectations created by the advert ([Google Ads Help](https://support.google.com/google-ads/answer/14086?hl=en)). These are platform quality inputs - not a guarantee of position, cost or conversion. ### 4. Proof that matches the decision An advert can make a claim in seconds. The page must help the buyer evaluate risk. Use proof that is close to the decision: a relevant example, an authorised client reference, named expertise, a transparent process, product detail, delivery conditions or an answer to a material objection. Explain what was done, for whom and under what conditions. Avoid decorating the page with generic testimonials or unsupported numbers. Useful proof narrows uncertainty. ### 5. A conversion that has business meaning Decide what the platform should count - and what the business must confirm. Keep three levels separate: 1. **Interaction:** impression, click, visit or video view. 2. **Response:** form, call, message, booking request or checkout. 3. **Commercial outcome:** accepted enquiry, qualified conversation, opportunity, paid order or retained customer. Google defines conversion actions as meaningful customer actions that the advertiser chooses, such as purchases or phone calls. It also distinguishes primary actions used for bidding from secondary actions retained for observation ([Google Ads Help](https://support.google.com/google-ads/answer/10995103?hl=en)). If every form submission is treated as equally valuable, the campaign cannot distinguish a suitable buyer from spam, a job application or a request outside your service area. ### 6. A qualified-outcome definition and real follow-up Define a qualified enquiry using observable conditions: the need is covered, the market or area is served, minimum commercial fit is present, timing is plausible and a next step is agreed. Then name: - who receives the response; - who qualifies it; - the expected response process; - the next commercial step; - where acceptance, rejection and eventual outcome are recorded. Google Ads can receive qualified- and converted-lead outcomes from a CRM, file or other internal source ([Google Ads Help](https://support.google.com/google-ads/answer/11459091?hl=en-AU)). The capability is useful only after your business has stable definitions and reliable records. A spreadsheet can be sufficient for an early test. The requirement is not expensive CRM software; it is a trace that another person can understand. ### 7. Measurement and targeting that respect the applicable rules There is no single global tracking rule. Requirements depend on where your organisation and users are located, what data is collected, which technologies are used and how platforms or other parties receive it. For example, the UK Information Commissioner's Office says storage and access technologies generally require prior consent unless an exception applies. Its current online-advertising guidance says consent is required where these technologies are used for ad selection, delivery, tracking, profiling and related measurement ([ICO: PECR rules](https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/guidance-on-the-use-of-storage-and-access-technologies/what-are-the-pecr-rules/), [ICO: online advertising](https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/guidance-on-the-use-of-storage-and-access-technologies/how-do-the-rules-apply-to-online-advertising/)). EU guidance likewise requires consent, when used as the legal basis, to be a genuine, informed choice rather than passive scrolling or a forced acceptance with no real alternative ([EDPB consent summary](https://www.edpb.europa.eu/system/files/2026-04/edpb-summary-consent_en.pdf)). These are EU/UK examples, not worldwide legal advice. Before launch, inventory the tags, pixels, SDKs, events, audience data, recipients, purposes and retention. Check the rules for the markets you actually serve and test what happens when a user declines non-essential tracking. ## 3. Google Ads or Meta Ads: how should a small business choose? Choose according to how demand appears - not according to which platform is currently fashionable. Google positions Search campaigns around reaching people while they are actively searching for products and services. Meta describes lead generation as happening in a discovery environment where businesses can create demand and nurture intent beyond people already searching ([Google Ads Help](https://support.google.com/google-ads/answer/9510373?hl=en), [Meta for Business](https://www.facebook.com/business/ads/ad-objectives/lead-generation?locale=en_GB)). | Buying situation | Google Search | Meta Ads | |---|---|---| | The buyer knows the problem and searches for a provider or product | Often coherent for capturing expressed demand | Can provide repetition or reassurance | | The problem exists, but the buyer does not know the solution category | Search demand may be small or ambiguous | Visual and explanatory creative can introduce the problem and offer | | The offer depends on demonstration, identity or repeated exposure | Text-led Search space may be restrictive | Image, video and sequence can make the offer tangible | | Search language is specific but the market is narrow | Useful if enough relevant demand exists | Can broaden discovery, but targeting does not create product-market fit | | The business cannot produce credible creative regularly | Search may be operationally simpler | Meta may be difficult to learn because the creative is part of the targeting and message test | This is a starting hypothesis, not a platform verdict. Google also offers non-Search inventory, and Meta can capture existing brand demand. Begin with the buying situation you can describe and measure. Do not launch both platforms merely to “be everywhere”. With a limited budget and team, one readable test often produces a better decision than two incomplete ones. ## 4. How to set a test budget without inventing a magic number There is no responsible universal starting budget for every country, sector, sales cycle and margin structure. Set three boundaries instead: 1. **Financial boundary:** the amount the business can treat as learning expenditure without weakening payroll, fulfilment or customer service. 2. **Operational boundary:** the number of enquiries or orders the team can respond to, qualify and fulfil properly. 3. **Decision boundary:** the date and evidence that will trigger stop, repair, continue or cautious expansion. Work backwards from the decision. If the available budget cannot produce enough relevant exposure, responses or completed sales cycles to answer the question, reduce the scope: one location, offer, product family, buying situation or creative proposition. “More visibility” is not a useful test objective. “Can this offer produce qualified consultations from UK operations leaders searching for this problem?” is closer to a decision. Do not change the offer, audience, advert, page and conversion definition at the same time. You may improve the campaign, but you will not know what the test taught you. ## 5. A practical 30-day paid-advertising test This is an IZZY operating framework, not a platform standard. Thirty days may be too short to judge a long sales cycle, seasonal market or low-volume offer. Preserve each response until its later commercial outcome is known. ### Before day 1: write the commercial hypothesis Record: - the offer and intended buyer; - the buying situation; - why the selected platform fits that situation; - the landing page and primary action; - the definition of a valid response and customer; - the contribution or cost boundary; - stop, repair and continue conditions. Save the initial advert, page, tracking events and qualification rules. Without a baseline, every later explanation becomes possible. ### Week 1: prove the circuit Test the advert destination, page, form or checkout, consent choices, notifications and customer record. Confirm that a recognisable test response reaches the named owner and can be linked to its source. Repair broken handoffs before interpreting demand. ### Weeks 2 and 3: observe without rebuilding everything Review: - search terms, audiences or placements; - page behaviour and response quality; - accepted and rejected enquiries; - response time and next actions; - orders, cancellations, refunds or lost opportunities; - mismatches between platform, analytics and business records. Change one material variable at a time. Record what you changed and what decision it was meant to improve. ### Week 4: make the decision | Decision | Evidence that can support it | |---|---| | Stop | The intended demand does not appear, economics exceed the boundary or the business cannot process responses safely | | Repair | The platform reaches relevant people, but the page, tracking, qualification or follow-up breaks the path | | Continue the test | Early responses are relevant, but the sales cycle or order volume is not mature enough for a verdict | | Expand cautiously | Accepted responses become reconciled customers within the defined economics and operations can absorb more volume | Do not scale because clicks became cheaper or the platform reported more conversions. Scale only when the deeper business outcome remains visible. ## 6. Is your business ready? This scorecard is a decision aid, not a scientific rating. | Check | Ready to test | Repair first | Too early | |---|---|---|---| | Offer and buyer | One offer, buyer and situation | Competing messages | Team cannot explain the same offer | | Economics | Contribution and internal boundary known | Some variable costs missing | Decision uses revenue alone | | Landing path | Promise continues and action is tested | Localised friction | Generic or broken destination | | Proof | Evidence matches the buyer's risk | Proof is broad or distant | Unsupported promise | | Measurement | Business-valued actions are tested | Events are incomplete | Only traffic and clicks are counted | | Follow-up | Definition, owner and next step exist | Handling is inconsistent | Responses have no owner | | Capacity | Sales and delivery can absorb the test | Capacity needs a limit | New demand would weaken service | | Data use | Technologies, purposes and choices are documented | Configuration needs review | Tags and audience flows are unknown | One “too early” condition can invalidate the whole test. Correct it, run the customer path again and then decide whether paid reach is the next useful investment. ## Conclusion: buy a useful answer before buying volume The first paid campaign should answer a commercial question: > Can this offer, presented to this buyer in this situation, produce customers that the business can identify, serve and acquire within an acceptable economic boundary? When the offer, landing path, measurement and follow-up agree, advertising can accelerate learning. When they contradict one another, it mostly makes the uncertainty more expensive. Build the acquisition system first. Choose the channel second. ## Bring us the complete path before buying more traffic Bring one offer, the proposed or current adverts, the landing page, tracked events and a sample of recent enquiries or orders. In a 30-minute scoping call, IZZY can frame whether the next useful move is a bounded paid test, a conversion repair, measurement and CRM work - or no project yet. We do not promise a lead volume, acquisition cost or return from a call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is Google Ads or Meta Ads better for a small business? Neither is universally better. Google Search often fits an already expressed need; Meta can introduce an offer before someone searches. Choose from the buying situation, evidence and operating capacity. ### How do I know whether my business is ready for paid advertising? You need a specific offer and buyer, supportable economics, a working landing path, business-valued measurement, reliable follow-up, delivery capacity and governed data use. Repair any “too early” condition before launch. ### How much should a small business spend testing ads? There is no universal amount. Set a financially safe boundary, narrow the test enough to generate a useful observation and define the decision date and stop conditions before spending. ### Should I create a dedicated landing page for paid ads? Not automatically. The destination must continue the advert's promise and make the next action clear. A focused existing service or product page can work if it does that under real mobile conditions. ### Can a small business start paid advertising without a CRM? Yes. A structured spreadsheet can record source, qualification, owner, next step and outcome for an early test. The essential requirement is a reliable shared trace. ### How should a small business measure advertising results? Keep interactions, responses and commercial outcomes separate. Reconcile platform and analytics events with accepted enquiries, customers, payments, refunds and delivery records. ### Is 30 days long enough to test Google Ads or Meta Ads? It can be a useful operating window, but it may be too short for low-volume or long-cycle sales. Preserve every response and judge it when the relevant commercial outcome becomes visible. ## Sources and evidence note Sources were checked on 27 July 2026. Platform documentation describes product capabilities, not expected results. The OECD survey is directional and not representative of all SMEs. EU and UK regulator guidance illustrates jurisdiction-specific requirements and is not worldwide legal advice. IZZY formulas, scorecards and the 30-day structure are operating methods, not market standards or forecasts. --- ### CRM outreach and consent: does your automation know why it is sending each message? URL: https://izzy.agency/en/blog/crm-outreach-consent-automation/ Published: 2026-07-31 Summary: Review CRM permissions, message purpose, suppression, vendors and email tracking before automating sales outreach across EU markets. Your CRM has a name, an email address and an active sequence. It knows what to send tomorrow. Does it know where the contact came from, why the message may be sent in this market and what must stop the next one? Automation becomes fragile when it repeats a decision no one has formalised. ## In 60 seconds Reliable outreach starts with a rule: **for every message, your system needs to know its purpose, recipient, channel, permission or legal condition, evidence and stop mechanism.** In practice: - do not treat an available email address as universal permission to communicate; - classify the message by its real purpose before building the trigger; - preserve source, notice, permission state, evidence and objections; - apply an objection across every CRM, sales and messaging tool; - assess email tracking separately from permission to send; - block or review records whose market rule or evidence is unknown. This makes a workflow inspectable. It does not replace local privacy and ePrivacy advice. ## In this guide 1. [One contact is not one universal permission](#1-one-contact-is-not-one-universal-permission) 2. [Classify the message before you automate it](#2-classify-the-message-before-you-automate-it) 3. [What an audit-ready CRM record should show](#3-what-an-audit-ready-crm-record-should-show) 4. [Audit triggers, vendors, suppression and tracking](#4-audit-triggers-vendors-suppression-and-tracking) 5. [Use this ten-question trusted-outreach gate](#5-use-this-ten-question-trusted-outreach-gate) 6. [Know when to pause and involve a specialist](#6-know-when-to-pause-and-involve-a-specialist) ## 1. One contact is not one universal permission An email address can enter through an enquiry, purchase, webinar, event list, import, enrichment product or public page. Those events do not create the same permission. Someone who requests a demo expects an answer. That does not automatically include a newsletter, a dormant-lead sequence and individual tracking of every opening. A business address may still identify a person and remain personal data. During migration, `marketing_allowed = yes` may survive while the notice, purpose, route, market and timestamp disappear. The value remains; the decision cannot be reconstructed. Public or enriched data need the same discipline. European Commission guidance says transparency for indirectly collected data should include its source and intended purpose. A public profile is not permission for any campaign. Before you automate, ask a more useful question than “do we have the email?” Ask: **what could this person reasonably expect us to do with it, in this channel and market?** ## 2. Classify the message before you automate it Commercial, service and relationship messages are not interchangeable. Whatever the label, ask what the message is trying to achieve. | Message purpose | Primary job | Example | Common classification failure | |---|---|---|---| | Direct marketing | Promote an offer, service or the organisation | Sales sequence, reactivation offer, promotional newsletter | Presenting promotion as neutral “information” | | Transactional or service | Deliver a requested service, account action or transaction | Confirmation, invoice, security alert, password reset | Adding significant promotion to a necessary message | | Customer relationship | Help a customer use or manage an existing service without a direct promotional purpose | Product-use guidance, operational notice, account support | Allowing support content to become an upsell sequence | Electronic marketing sits across the GDPR, the ePrivacy Directive and national law. Article 13 of the Directive sets a prior-consent frame for natural-person subscribers and a bounded existing-customer exception for the sender's own similar offers, with easy objection. Parts of implementation remain with Member States. Do not copy a French B2B rule silently into Germany, Ireland or another market. Legitimate interests may support some processing, but do not automatically settle a channel-specific electronic-marketing rule. Where consent is used, it must be freely given, specific, informed and unambiguous, and withdrawal must be as easy as giving it. Your CRM does not need to contain a legal essay. It does need a market-aware decision instead of allowing vague labels such as `lead`, `customer` or `newsletter` to control the send on their own. ## 3. What an audit-ready CRM record should show No regulator prescribes one universal CRM schema. This is IZZY's operational model for making a decision testable: 1. **Source and date.** Form, purchase, event, partner, import, enrichment or public source. 2. **Relationship.** Consumer, professional contact, customer, prospect or user. 3. **Market and channel.** Relevant country plus email, SMS, phone, messaging, push or post. 4. **Purpose.** Direct marketing, transactional, relationship or another documented purpose. 5. **Condition and state.** Consent, contract, existing-customer condition, legitimate-interest assessment, objection or review needed. 6. **Evidence.** Notice, action, timestamp, collection route, provider and proof location. 7. **Tracking choice.** Separate state for pixels or comparable tracking. 8. **Objection or withdrawal.** Date, scope and durable block. 9. **Retention.** Deletion, archive or re-evaluation rule. 10. **Owner.** Person responsible for the campaign, exceptions and stop decision. These elements may live in several tools, but must remain connected. If the email platform stores proof and the CRM a Boolean, document their synchronisation. Start with one source and one sequence. Trace the decision from collection to final message; the first missing control will appear faster than in a broad “GDPR CRM project”. ## 4. Audit triggers, vendors, suppression and tracking Map the route the data really takes: `source -> CRM -> segment -> trigger -> message -> vendor -> tracking -> response or objection -> CRM update` At each hand-off, ask what changes and which system is authoritative. Common failures live between tools: unsubscribe in platform A, old audience in platform B, an import that restores the record or a sales extension that ignores the stop state. Under the GDPR, a person may object to direct-marketing processing at any time. The objection must survive later imports and enrichment - not disappear with the campaign that received it. Inspect vendor defaults: who inserts the pixel, receives the data and decides its use? A settings change can alter the decision you reviewed. Do not infer permission to track from permission to send. France's [2026 CNIL recommendation](https://www.cnil.fr/fr/recommandation-pixel-suivi-courriels) distinguishes performance, personalisation and profiling uses requiring consent from narrow exemptions. It is a national example, not a universal EU ruling, but shows why “open tracking: on” is not neutral. Ask what decision opening data changes. Aggregated delivery information or a business outcome may be enough. Prefer qualified replies, confirmed meetings and accepted opportunities to activity that is merely easy to collect. ## 5. Use this ten-question trusted-outreach gate Before activating a sequence, ask: - [ ] Do we know where every contact came from and when it was collected? - [ ] Have we identified the market, recipient type and channel? - [ ] Is the real purpose of every message classified? - [ ] Is the chosen permission or legal condition documented and still applicable? - [ ] Does the information given to the person match the current use? - [ ] Can the relevant proof be retrieved without a forensic search through backups? - [ ] Does an objection or withdrawal block every tool before the next send? - [ ] Do pixels and comparable tracking have their own purpose, configuration and evidence? - [ ] Do vendors, imports and enrichment follow the same rules? - [ ] Can a named owner stop the campaign and explain the decision? A “no” does not require a new CRM. The fix may be one field, a synchronisation rule, a default exclusion, clearer information or human approval. The wrong response is to compensate for an unknown rule with more automation. ## 6. Know when to pause and involve a specialist Pause activation and obtain qualified advice when: - the legal basis or electronic-marketing rule is disputed; - the workflow spans markets with different national implementations; - sensitive data, children or vulnerable people are involved; - scraping or enrichment makes provenance unclear; - a vendor reuses signals for its own purposes; - scoring, profiling or automated decisions produce a significant effect; - individual tracking is considered necessary but its conditions are unresolved; - no one can demonstrate the notice, choice or objection handling. IZZY can map collection, CRM fields, triggers, vendors, evidence and stop paths. A DPO, lawyer or sector specialist should own the final legal conclusion when the facts require one. ## Conclusion: automate a decision, not an ambiguity Your CRM needs to know why a message is sent, what evidence supports it and which event must stop it. Start with one source, one audience and one sequence. Classify messages, connect evidence, test the objection path and inspect vendors. If the rule remains unknown, do not send automatically. Automation is valuable when it accelerates an explicit decision. Otherwise it accelerates invisible debt. ## Bring one campaign. We’ll trace the complete path. Bring one lead source, CRM fields, templates, automation map, suppression behaviour and tracking settings. In 30 minutes, we can give a bounded read: pilot, focused fix, specialist review or automation not justified yet. See [Smart Lead Conversion](https://izzy.agency/en/products/smart-lead-conversion/) for the productised version of this work. We will not promise compliance or a conversion increase from a call. If visitors never become usable enquiries, start with [the acquisition-path diagnosis](https://izzy.agency/en/blog/website-visitors-no-enquiries/). If AI makes decisions across the workflow, use the [AI production-governance checklist](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/). [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Does GDPR always require consent for B2B sales email? No answer covers every EU market and channel. GDPR governs personal data; ePrivacy and national laws add marketing rules. Professional relevance or legitimate interests may matter, but neither creates automatic permission. ### Is legitimate interest enough for an automated sequence? Not by itself. Document purpose, necessity and balancing, provide transparency and make objection effective. Then check the channel rule in the relevant market. ### Is an unsubscribe link enough? It is important, but incomplete. The stop state must reach every sending tool and survive imports, enrichment and future campaigns. Source, purpose and original decision still matter. ### Are email tracking pixels illegal in the EU? Not universally. Treatment depends on purpose, configuration, national ePrivacy implementation and subsequent processing. Permission to send is not automatic permission to track an opening. ### Can we use contact details found on LinkedIn or a public website? Public availability does not remove GDPR duties. Document source, expected use, transparency, market and channel rules, and the applicable objection or consent condition. ## Official sources - [European Data Protection Board - Consent under GDPR, 2026 summary](https://www.edpb.europa.eu/system/files/2026-04/edpb-summary-consent_en.pdf) - [European Data Protection Board - Guidelines 05/2020 on consent](https://www.edpb.europa.eu/documents/guideline/guidelines-052020-on-consent-under-regulation-2016679_en) - [EUR-Lex - ePrivacy Directive, Article 13](https://eur-lex.europa.eu/eli/dir/2002/58/oj?locale=en) - [European Commission - GDPR principles for organisations](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/principles-gdpr_en) - [European Commission - Objections to direct-marketing processing](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/dealing-requests-individuals_en) - [EUR-Lex - General Data Protection Regulation](https://eur-lex.europa.eu/eli/reg/2016/679/oj) - [CNIL - Electronic communications to prospects and customers, 10 June 2026](https://www.cnil.fr/fr/communication-electronique-quelles-regles) - [CNIL - Email tracking-pixel recommendation, adopted 12 March 2026](https://www.cnil.fr/sites/default/files/2026-04/recommandation-pixels_de_suivi.pdf) Sources and links checked 22 July 2026. This article provides general operational information. It is not legal advice, an assessment of your organisation's compliance or a forecast of commercial results. --- ### How to turn a prototype into a product customers will pay for URL: https://izzy.agency/en/blog/prototype-to-commercial-product/ Published: 2026-07-31 Summary: A practical route from working prototype to paid proof, usable onboarding and a product your business can operate - without building every feature first. Your prototype works. People understand the demo. Then the next invoice, support request or integration exposes the real problem: you have proved that the idea can work, not that a business can sell, deliver and support it repeatedly. The expensive mistake is to respond with more features. Commercialisation is a different job. It connects a specific buyer, a paid promise, a usable path to value, reliable delivery and workable economics. ## Answer in 60 seconds To turn a software prototype into a commercial product, prove six things around one customer journey: 1. **Buyer proof:** a specific buyer is already trying to solve the problem. 2. **Paid-value proof:** they will commit to a bounded outcome. 3. **Usage proof:** the user can reach that outcome without a founder-led performance. 4. **Delivery proof:** the service handles real data, errors, support and recovery. 5. **Ownership proof:** the business controls its code, accounts and operating knowledge. 6. **Economic proof:** price, delivery cost and manual effort support another sale - or expose what must change. You need enough evidence to choose: controlled pilot, focused productisation or pause before uncertainty becomes expensive code. ## In this guide - [1. Identify what the prototype has and has not proved](#1-identify-what-the-prototype-has-and-has-not-proved) - [2. Establish buyer proof before expanding the roadmap](#2-establish-buyer-proof-before-expanding-the-roadmap) - [3. Turn interest into a controlled commercial test](#3-turn-interest-into-a-controlled-commercial-test) - [4. Productise the shortest credible path to value](#4-productise-the-shortest-credible-path-to-value) - [5. Make the promise operable after the demo](#5-make-the-promise-operable-after-the-demo) - [6. Test the economics before you scale](#6-test-the-economics-before-you-scale) - [7. Decide: pilot, focused productisation or pause](#7-decide-pilot-focused-productisation-or-pause) ## 1. Identify what the prototype has and has not proved A prototype is an experiment made tangible. It may test comprehension, usability or technical possibility. It does not yet show that customers will buy or that the company can deliver reliably. [GOV.UK's prototyping guidance](https://www.gov.uk/service-manual/design/making-prototypes) cautions that code built for realistic testing may omit the security and load qualities expected of a live service. Moving it into production therefore needs a separate engineering decision. This is public-service guidance rather than a universal software rule, but the boundary travels well. Use these working definitions for the commercial decision: | Stage | Question it should answer | Evidence it does not provide by itself | |---|---|---| | Prototype | Can we make the idea understandable or technically plausible? | Purchase intent, reliable delivery or sustainable economics | | MVP | Can a chosen user complete the smallest useful journey in real conditions? | Repeatable acquisition, retention or a scalable operating model | | Commercial product | Can we sell, deliver, support and improve a defined promise under accountable ownership? | Permanent product-market fit or guaranteed growth | The same codebase can move through all three stages, or it may need selective replacement. Current evidence decides - not the label. Qualitative startup studies connect prototypes with customer learning, technical debt and testing, and report vague planning and evolving throw-away prototypes among observed challenges ([study of 20 startups](https://arxiv.org/abs/1712.00674), [study of 40 startups](https://arxiv.org/abs/2103.07999), [related 20-startup study](https://arxiv.org/abs/1711.07045)). They do not provide a success formula. ## 2. Establish buyer proof before expanding the roadmap “Small businesses” is not a useful first market. Neither is “marketing teams” or “people who need productivity.” Choose one buyer in one situation with one problem expensive enough - in money, time, risk or missed opportunity - to justify change. Write the commercial hypothesis in one sentence: > When **this situation occurs**, **this buyer** needs to achieve **this outcome**, because the current approach causes **this observable cost or consequence**. Investigate behaviour, not compliments. Ask what triggered the problem, how the buyer handles it now, what the workaround costs and what would make a purchase impossible. A feature request is a clue until you understand the job behind it. [GOV.UK's user-needs guidance](https://www.gov.uk/service-manual/user-research/start-by-learning-user-needs) recommends combining interviews and observation with existing evidence such as analytics and support data. Opinions not grounded in actual or likely users remain assumptions to investigate. Buyer proof is stronger when you know who experiences the problem, who approves spending, what makes action urgent, what the current workaround costs and which buying condition can block progress. A recurring pattern you can reach again is more useful than scattered enthusiasm. [Y Combinator](https://www.ycombinator.com/blog/ycs-essential-startup-advice) similarly advises staying small and focused before product-market fit; this is accelerator guidance, not outcome evidence. ## 3. Turn interest into a controlled commercial test A commercial test asks for a real commitment: payment, operational data, staff time, a named decision-maker or an agreed procurement step. For early B2B software, a controlled pilot can bridge demo and product while exposing what the product does and what the team still performs manually. A useful pilot brief fits on one page: | Decision | What to write down | |---|---| | Buyer, user and problem | Who buys, who uses, who approves and what current consequence changes | | Outcome | The observable result the pilot is intended to produce | | Scope | Included journey, users, data, integrations and exclusions | | Responsibilities and terms | What each party provides; price, payment timing and separately priced setup or support | | Evidence and end decision | What to record, then whether to continue, revise, stop or widen the rollout | Do not use a pilot to hide indefinite custom development. If every prospect receives a different promise, integration and success definition, you may be building a service business. That can be a good business; it is different evidence from a repeatable product. [Stripe Atlas's guide to initial customers](https://stripe.com/guides/atlas/starting-sales) advises founders to recruit early customers actively and feed those conversations into positioning and product decisions. This is practitioner guidance, not a guaranteed route. Free research can still answer usability questions; it simply does not provide purchase proof. ## 4. Productise the shortest credible path to value Once a real buyer and outcome are visible, stop asking which features make the product look complete. Ask which steps must work for the customer to receive the promised value. Map one path: 1. **Trigger:** what makes the buyer or user start now? 2. **Entry:** how do they understand the offer and obtain access? 3. **Setup:** what data, permission, configuration or help is genuinely required? 4. **Core action:** what must the user do inside the product? 5. **Result:** what useful output or changed state do they receive? 6. **Confirmation:** how do the user and the business know it worked? 7. **Return:** what gives them a reason and a route to use it again? Test that path with the intended user. Record where they hesitate, stop or need manual intervention. Manual onboarding is not automatically a failure at this stage; invisible manual work is. It changes support capacity, pricing and the promise you can make. Choose one product event tied to the promised outcome. Google Analytics calls an action important to business success a “key event” ([Google Analytics Help](https://support.google.com/analytics/answer/13965727?hl=en-GB)). Reconcile it with the customer's result and your operational record before using it for a roadmap decision. If people sign up but do not reach value, more acquisition feeds the same break. Diagnose it separately - as you would [a website that attracts interest but no qualified conversations](https://izzy.agency/en/blog/website-visitors-no-enquiries/). ## 5. Make the promise operable after the demo The founder can rescue a demo in real time. A customer expects the product to handle ordinary mistakes, slow services and unavailable support without theatre. For IZZY, delivery proof is a bounded operating record, not a “production-ready” badge. At minimum, inspect: | Customer promise | Evidence to request | |---|---| | The right people can use it | Roles, account setup, access removal and support route have been exercised | | Customer data is handled deliberately | Data flow, suppliers, retention, permissions and sensitive fields are known | | The core journey survives failure | Error, retry, duplicate, timeout and partial-completion cases have been tested | | Releases and failures can be controlled | Source, deployment, rollback, useful signals, alert owner and recovery route are exercised for the critical scope | | Another person can operate it | Accounts, documentation, decisions and practical knowledge are under company control | Google SRE treats production-readiness review as context-specific evidence for accepting operational responsibility ([Google SRE](https://sre.google/sre-book/evolving-sre-engagement-model/)). Keep the principle, not Google's organisation. Security and accessibility belong in the product, not an appendix. The [NIST framework](https://csrc.nist.gov/pubs/sp/800/218/final) provides lifecycle practices, not certification. [W3C](https://www.w3.org/WAI/test-evaluate/) recommends evaluating accessibility throughout development and says tools alone cannot determine conformance. Review depth should follow consequence. Payments, health, finance, employment decisions, sensitive data or privileged actions may require qualified specialists. This article does not classify obligations or approve a release. If the prototype uses AI in a consequential workflow, use a separate [AI pilot-to-production governance review](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/). Model behaviour, provider change and human authority add distinct questions. ## 6. Test the economics before you scale A sale can validate urgency while still hiding an unworkable delivery model. After each pilot or early customer, build a simple evidence ledger: - price and payment actually collected; - setup, migration and support effort; - direct infrastructure, API and supplier costs; - custom work that cannot be reused; - steps required to reach the decision-maker; - usage and outcome after onboarding; - what the next similar customer would require. Do not force mature SaaS metrics onto a handful of customers. Use the ledger to expose the mechanism. If price excludes implementation every customer needs, separate or redesign that work. If one integration serves only one buyer, decide whether it is strategic, chargeable or out of scope. Stripe's low-touch SaaS pricing guide links packaging to customer segment and value ([Stripe Atlas](https://stripe.com/guides/atlas/saas-pricing)). It is practitioner guidance, not a universal price card. Before scaling, record why the customer paid, what was repeatable or bespoke, whether the user repeated the promised value and what must improve before the next sale. Metrics should support a decision. GOV.UK's service guidance connects measures to intended benefits and uses beta evidence to continue, change or stop ([GOV.UK Service Manual](https://www.gov.uk/service-manual/measuring-success/measuring-service-benefits)). Learning that the current offer should not scale is still a useful result. ## 7. Decide: pilot, focused productisation or pause Do not average the six proofs into a score. One material gap can change the route. | What the evidence shows | Decision | Next move | |---|---|---| | A reachable buyer, urgent problem and bounded outcome are credible; delivery still needs close supervision. | Sell a controlled pilot | Agree scope, commercial terms, evidence and an end decision before expanding the build | | Paid value and usage are visible; contained UX, reliability, ownership or cost gaps block repetition. | Focus productisation | Fix the first gap on the critical journey, retest it and update the operating model | | No specific buyer commits, or material data, ownership, security or delivery unknowns prevent a responsible transaction. | Pause and resolve | Stop feature expansion; investigate the commercial assumption or run a bounded product/code rescue | A rewrite is not the default. Repair the journey when the problem is contained. Replace a component when evidence shows it blocks the target outcome or cannot be owned responsibly. If the whole architecture is under question, assess it separately before turning a commercial uncertainty into a rebuild. ## Conclusion: the next milestone is proof, not more product A convincing prototype has already earned its place: it made the idea testable. The next job is to make the business testable. Choose one buyer, one paid promise and one route to value. Record what the customer does, what your team must do and what the transaction costs to deliver. Then strengthen only the part that blocks the next responsible sale. The result may be a controlled pilot, a focused productisation sprint or a pause. All three are better than an expanding roadmap built on applause. ## Bring the prototype. We’ll identify the next proof it needs. Bring one working prototype, the buyer you believe it serves and any customer or usage evidence you already have. In a 30-minute scoping call, IZZY will give you an initial, bounded read on whether the next move appears to be a commercial pilot, focused productisation, deeper technical review or more customer discovery before an engagement. This is not a free audit, product-market-fit verdict, legal or security assessment, launch approval or promise of commercial results. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### What is the difference between a prototype and an MVP? A prototype makes an idea testable. An MVP lets a chosen user complete the smallest useful journey in real conditions. Define the evidence instead of relying on an inconsistent label. ### Do I need a paying customer before building the product? Not in every case. Some products need technical, safety or regulatory work first. Still seek evidence from likely buyers and users, and record which commercial assumptions remain untested. ### How many customer interviews are enough? There is no universal number. Continue until you can distinguish a recurring buyer problem from isolated opinions and make a specific decision. A few conversations do not establish prevalence. ### Should prototype code be rewritten for production? Not automatically. Review the critical journey, architecture, security, tests, operations and ownership. Keep what is fit; harden or replace what evidence shows is fragile or uneconomic. ### What should a paid pilot include? Define buyer, users, outcome, scope, exclusions, responsibilities, commercial terms, evidence, data boundaries and the end decision. A pilot is not an open-ended feature queue. ## Sources and method This article is an original IZZY synthesis of current founder-language signals, official guidance, practitioner material and direct software-startup research checked on 22 July 2026. Community discussions were used to identify the question, not to establish market size, success rates or willingness to pay. The six-proof framework and decision model are IZZY practice guidance; they are not an external standard or guarantee. - [GOV.UK Service Manual - Making prototypes](https://www.gov.uk/service-manual/design/making-prototypes) - [GOV.UK Service Manual - Learning about users and their needs](https://www.gov.uk/service-manual/user-research/start-by-learning-user-needs) - [GOV.UK Service Manual - Measuring the benefits of your service](https://www.gov.uk/service-manual/measuring-success/measuring-service-benefits) - [Stripe Atlas - Your first 10 customers](https://stripe.com/guides/atlas/starting-sales) - [Stripe Atlas - Pricing low-touch SaaS](https://stripe.com/guides/atlas/saas-pricing) - [Y Combinator - Essential Startup Advice](https://www.ycombinator.com/blog/ycs-essential-startup-advice) - [Google Analytics Help - Events and key events](https://support.google.com/analytics/answer/13965727?hl=en-GB) - [Google SRE - Production Readiness Review](https://sre.google/sre-book/evolving-sre-engagement-model/) - [NIST - Secure Software Development Framework 1.1](https://csrc.nist.gov/pubs/sp/800/218/final) - [W3C WAI - Evaluating Web Accessibility](https://www.w3.org/WAI/test-evaluate/) - [Nguyen Duc, Wang and Abrahamsson - What influences the speed of prototyping?](https://arxiv.org/abs/1712.00674) - [Nguyen-Duc, Kemell and Abrahamsson - The entrepreneurial logic of startup software development](https://arxiv.org/abs/2103.07999) - [Nguyen Duc et al. - Towards understanding startup product development as effectual entrepreneurial behaviours](https://arxiv.org/abs/1711.07045) --- ### Your AI prototype looks ready. Is it ready to win trust and convert? URL: https://izzy.agency/en/blog/ai-prototype-conversion-readiness/ Published: 2026-07-30 Summary: Before launching an AI-generated website or app, check whether customers can understand it, trust it, complete key actions and receive a reliable response. The prototype looks polished. Stakeholders can click through it and the team can see the idea. Launching - or sending paid traffic to it - feels like the natural next move. That speed is valuable. It is also the moment to change the question. The demo has shown that the idea can be presented. Before customers depend on it, the business needs evidence that they can understand it, trust it, complete the important action and receive a reliable outcome. ## Answer in 60 seconds Use these seven readiness tests before you launch or increase acquisition spend: 1. A new visitor understands the offer and the next step. 2. The experience carries your company’s trust signals and visual rules consistently. 3. The most important customer journey works from beginning to confirmed outcome. 4. Mobile and accessibility have been checked with people and devices, not only previews or automated scores. 5. Real-world speed, failure states and integrations are acceptable. 6. The business can measure meaningful outcomes, not only visits and clicks. 7. A named person owns changes, incidents, customer response and recovery. These tests support a decision; as IZZY practice guidance, they do not guarantee conversion, compliance or overall product quality. ## 1. The prototype has done its job. Now ask a different question Today’s tools can produce something convincing quickly. Figma documents that Make can create functional prototypes and web apps through chat. Canva reports that its Code product can create interactive websites, apps and experiences. Those capabilities do not establish that a particular result is ready for customers. A realistic prototype is not a failed product. It is a successful way to make an idea tangible, collect reactions and decide what deserves further investment. [GOV.UK’s guidance on making prototypes](https://www.gov.uk/service-manual/design/making-prototypes) makes the distinction clearly: realistic code prototypes can support user research, while prototype code may not meet production security or high-traffic performance needs and should not simply be copied into production. That is transferable guidance, not a universal launch method. The question is now broader than “Does the screen work in the demo?” It is “Can we depend on the experience when the user is unfamiliar, the phone is small or something goes wrong?” You do not answer that by adding more polish. You answer it by gathering evidence along one real customer path. ## 2. Can a customer understand and trust it? For this readiness review, treat comprehension and trust as prerequisites: someone outside the project should understand the offer, proof and next step without a tour. Our guide to [diagnosing visitors who do not become enquiries](https://izzy.agency/en/blog/website-visitors-no-enquiries/) covers that path through contact and follow-up. Fix it first rather than treating AI-prototype readiness as a substitute. Then inspect the generated experience beyond its ideal screen. As IZZY practice guidance, map the empty, loading, partial, error, success and returning-user states. Each should follow the same rules, explain what is happening and give a sensible next action. One polished screen does not show whether the experience is coherent. Then check consistency. In ordinary terms, a design system means colours, type, buttons, form behaviour and language follow recognisable rules. A button should not change meaning between pages, and checkout should not feel like another company. Figma says a Make kit can carry shared product styles and guidance into generated work, and recommends checking where output departs from those rules. That helps; it does not create an automatic match. Compare each state with the product rules the business approved. ## 3. Can the important journey survive real use? Choose one journey that matters commercially: an enquiry, booking, checkout, onboarding or account change. Connect its states from entry to confirmed outcome, including when the customer leaves and returns. The decision concerns the journey, not the screen that made the demo look finished. Treat the preview as one environment, not the result. Repeat the journey on expected devices and browsers. Open the keyboard, rotate the screen, switch away, return and slow the connection. Check whether controls remain visible, progress survives and the next action makes sense. Draw a boundary around each outside service: payment, booking, sign-in, CRM or email. Decide what both sides see when it is slow, rejects a request, sends the same confirmation twice or does not respond. Record which failures can be retried, need a person or must stop the journey. Accessibility belongs inside the journey. The [W3C accessibility evaluation guidance](https://www.w3.org/WAI/test-evaluate/) says no tool alone can determine whether a site meets accessibility standards; knowledgeable human evaluation is required. Under WCAG 2.2, conformance covers the whole page, including versions shown at different screen sizes such as mobile, and all the pages needed for every step of a task. This readiness review uses that whole-journey scope; it does not certify WCAG conformance. A score can point to problems, not certify this prototype or replace testing with relevant people and devices. These checks turn a vague concern into something specific to fix and show whether it reaches the core journey. ## 4. Can the business see and operate what happens next? Customers experience performance as waiting, jumping content and delayed response - not as a technical score. Core Web Vitals provide signals for loading, interaction responsiveness and visual stability. They help teams inspect and monitor the experience, but they do not prove overall usability or predict a commercial result by themselves. In one company’s month-long 50/50 landing-page test, Rakuten 24 reported improvements in conversion rate and revenue per visitor for its Core-Web-Vitals-optimised version. Those results belong to that case, not yours; use the example as a reason to measure your own essential journey, not as a forecast. Measurement should follow the outcome. Google Analytics calls an action important to business success a “key event”. That could be a confirmed booking, qualified enquiry or purchase - not merely a click. Configuration alone does not make the choice meaningful or the data accurate. Run a known test and verify that customer confirmation, the business system and reporting agree. Finally, move the knowledge from the person who ran the demo to the people who will operate the outcome. As IZZY practice guidance, hand over the state map, approved product rules, outside-service dependencies, known limitations, test evidence, reporting and recovery steps. Name one business owner who knows who approves changes, watches the release, responds to customers and starts recovery. If the experience connects to broader AI workflows, our [AI pilot-to-production governance guide](https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/) provides a deeper boundary. ## 5. Decide: launch, focused fix or rescue The seven tests support a practical decision, not a score. Record what you observed, what remains uncertain and which limitation could affect the essential journey. | Route | Evidence | Decision | |---|---|---| | Launch and monitor | Essential journey works; known limitations are minor; outcome and owner are visible | Release to a bounded audience and monitor | | Focused fix | Offer, trust, accessibility, tracking or one journey has a contained problem | Fix the first meaningful break, retest, then launch | | Pause and harden | Customer/data risk, repeated failure, unclear ownership or fragile architecture affects the core journey | Stop public exposure or acquisition spend and run a deeper review | A rebuild is not the default. If the offer is unclear, rewrite it. If one form fails, repair and retest that path. If visual rules drift, align the components customers actually encounter. Pause when the problem cannot be isolated or the consequence is serious. A journey handling payment, account or sensitive data may need product-specific review. NIST’s Secure Software Development Framework supports integrating secure-development practices across the lifecycle; it is not a launch approval. Where fragile architecture blocks a contained fix, our guide to [the true cost of technical debt](https://izzy.agency/en/blog/true-cost-of-technical-debt/) explains repair versus replacement. ## Conclusion: speed to prototype changes the next question AI tools reduce the distance between an idea and something people can try. That is progress. But visual plausibility is only the first kind of evidence. Before launch, prove that a new customer can understand the offer, trust the experience, complete one important journey and receive confirmation. Then prove that the business can see the result and knows who acts when reality differs from the demo. The answer may be launch and monitor, a focused fix or a deeper review - not automatically a rebuild. ## Bring the prototype. We’ll test the path to a real customer outcome. Bring one prototype and one essential customer journey to a 30-minute scoping call. We will give you a bounded initial read on what is visible, what evidence is missing and whether the next step appears to be launch and monitor, a focused fix, a deeper review or no engagement yet. The call is not a conversion forecast, accessibility certification, security audit or launch approval. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Does an AI-generated website need a full rebuild before launch? Not by default. Clearer wording, consistent controls, a repaired form or reliable confirmation may solve a contained problem. Consider rescue work when the core journey remains fragile or consequential risk cannot be isolated. ### Is a high accessibility or performance score enough? No. Automated checks provide useful signals, but accessibility requires knowledgeable human evaluation and performance metrics do not prove the whole experience. Test the complete journey with relevant people and devices. ### Which customer journey should we test first? Choose the action closest to a real business outcome: a qualified enquiry, booking, purchase, onboarding completion or important account action. Follow it through confirmation and the business response. ### Can we launch while some issues remain? Possibly, if the essential journey works, limitations are minor and an owner can monitor a bounded release. Do not use it to hide customer, data or repeated core-journey risk. ### What must the demo owner hand over before launch? Hand over states, product rules, dependencies, known limits, test results, measurement and recovery contacts. The business owner should not need the builder beside them when the journey changes or fails. ## Sources - [Canva Code 2.0 announcement](https://www.canva.com/newsroom/news/Canva-Code/) - [Figma Help: Create and edit a Figma Make file](https://help.figma.com/hc/en-us/articles/31304485164695-Create-and-edit-a-Figma-Make-file) - [Figma Help: Get started with Make kits](https://help.figma.com/hc/en-us/articles/39241689698839-Get-started-with-Make-kits) - [GOV.UK Service Manual: Making prototypes](https://www.gov.uk/service-manual/design/making-prototypes) - [W3C: Web Content Accessibility Guidelines (WCAG) 2.2](https://www.w3.org/TR/WCAG22/) - [W3C WAI: Evaluating Web Accessibility Overview](https://www.w3.org/WAI/test-evaluate/) - [web.dev: Web Vitals](https://web.dev/articles/vitals?hl=en) - [web.dev: Rakuten 24 Core Web Vitals case study](https://web.dev/case-studies/rakuten?hl=en) - [Google Analytics Help: About key events](https://support.google.com/analytics/answer/9267568?hl=en) - [NIST: Secure Software Development Framework (SSDF) Version 1.1](https://csrc.nist.gov/pubs/sp/800/218/final) --- ### Can another team own your AI-built product? URL: https://izzy.agency/en/blog/ai-built-product-codebase-handoff/ Published: 2026-07-29 Summary: Assess whether another team can operate an AI-built product, identify bounded hardening needs and make a defensible handoff decision. The deal room is open, but the buyer’s technical team cannot tell which account controls production or how the deployed build relates to the repository. That is a handoff problem, whether the software was written by people, generated with AI or both. Code origin is not the verdict. Verifiable ownership is. ## The answer in 60 seconds Look for six signals: - a current map of the system and its dependencies; - a repeatable build, release and rollback path; - controlled identities, data access and secrets; - critical journeys tested beyond the happy path; - observable operations, recovery and a workable support model; - named human owners, recorded decisions and known debt. Those signals support three decisions: **accept and operate**, **accept with bounded hardening**, or **pause and rescue**. ## In this article - [The real test is not “who wrote the code?”](#real-test) - [Three business moments that expose a weak handoff](#business-moments) - [The six evidence packs a receiving team needs](#evidence-packs) - [The ten-question takeover test](#takeover-test) - [Decide: take over, harden or pause](#decide) - [What a professional assessment should produce](#professional-assessment) - [Conclusion: ownership has to be demonstrated](#conclusion) ## The real test is not “who wrote the code?” AI-assisted development ranges from suggested lines to agent-led implementation. Research reports mixed, context-dependent results; it does not justify treating AI-built software as inherently better or worse ([NCSC](https://www.ncsc.gov.uk/blogs/the-vibe-coding-spectrum-approach-to-ai-assisted-software-development), [Geruslu, Aliyeva and Tüzün](https://arxiv.org/abs/2603.25146)). The useful question is narrower: can an accountable team explain important flows, create a release, investigate failure and make a responsible change? Review depth should rise with the consequences and sensitivity of failure. A named owner does not prove correctness, but production decisions should not be ownerless. This distinction helps avoid two poor decisions: rejecting a workable product because AI touched it, or accepting an impressive demo without evidence of production readiness. ## Three business moments that expose a weak handoff ### An enterprise customer starts technical review The buyer may ask who can access customer data, what validates a release and what follows a failed deployment. Written policy is useful; current artefacts and demonstrated procedures are stronger evidence. AI-agent [permissions and access control](https://izzy.agency/en/blog/ai-agent-permissions-access-control/) can be examined separately without replacing the handoff question. ### An investor or acquirer begins technical due diligence A functioning product does not reveal whether critical dependencies, release or the support model will transfer. The decision needs boundaries: what is understood, what requires hardening, and what remains too uncertain to accept. ### The builder leaves A repository export is not operational ownership. Cloud accounts, deployment knowledge and incident history may still sit with a founder, freelancer or outgoing agency. The gap becomes visible when the new team cannot ship or respond to a fault independently. ## The six evidence packs a receiving team needs An evidence pack combines current artefacts with actions the receiving team can perform. Documentation shows the intended path; execution tests whether it still works. ### 1. System and dependency map **Buyer question:** What must we understand and control to run the critical journey? **Concrete evidence:** A current architecture view, request and data flows, services, suppliers, versions, owners, maintenance status and exit constraints for critical dependencies. **Failure exposed when missing:** An unknown single point of failure, abandoned component or supplier dependency with no owner. Google’s launch guidance puts architecture, dependencies and failure modes within readiness review, while warning that checklists must fit context ([Google SRE](https://sre.google/sre-book/reliable-product-launches/)). ### 2. Reproducible build and release path **Buyer question:** Can the receiving team release the product from its source? **Concrete evidence:** Declared environment, build command, pipeline, source revision, linked artefact, deployment procedure and exercised rollback. **Failure exposed when missing:** A release dependent on the outgoing builder’s laptop, memory or credentials. Build provenance can link an artefact to source and build steps; it does not prove quality, security or maintainability ([SLSA](https://slsa.dev/spec/v1.2/provenance)). ### 3. Security, identity, data and secrets **Buyer question:** Who actually controls sensitive access and assets after the handoff? **Concrete evidence:** An identity and account inventory, effective permissions, organisational ownership, demonstrated revocation or rotation, secret locations without exposed values, dependency status and inspected CI/CD changes. **Failure exposed when missing:** An irreplaceable personal account, shared secret, excessive permission or pipeline that nobody owns. OWASP treats these as a connected control surface and calls for human accountability and independent review ([OWASP](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Coding_with_AI_Cheat_Sheet.html)). ### 4. Critical-journey tests and failure evidence **Buyer question:** Does important behaviour hold under invalid, edge and plausible abuse conditions as well as the happy path? **Concrete evidence:** Requirements linked to normal, boundary, negative and adversarial cases; recorded results; known defects; and competent review of security-critical behaviour and test changes. **Failure exposed when missing:** A green test suite that checks only assumptions made by the same agent that changed the code. Passing tests remain useful evidence, but not independent security assurance. ### 5. Operations, recovery and support **Buyer question:** Can the team detect, decide, communicate and restore when customers experience failure? **Concrete evidence:** Logs, metrics or traces tied to critical journeys, owned alerts, operating procedures, incident records, an observed rollback or restoration exercise and clear support responsibilities. **Failure exposed when missing:** A backup never restored, an alert with no responder or a support promise with no operating owner. Recovery guidance emphasises scoped, monitored exercises and documented results, not a backup checkbox ([Google Cloud](https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-failures)). ### 6. Human ownership, decisions, documentation and known debt **Buyer question:** Who will make decisions after transfer, and what knowledge will those decisions rely on? **Concrete evidence:** Named owners for product, code, data, operations and support; decision records; prioritised debt; walkthroughs; and hands-on exercises performed by the receiving team. **Failure exposed when missing:** Orphaned documentation, implicit responsibility or recurring dependence on the outgoing builder. Debt should be connected to business consequences; it is not an automatic argument for a rewrite. See IZZY’s separate [technical-debt decision framework](https://izzy.agency/en/blog/true-cost-of-technical-debt/). ## The ten-question takeover test This is a decision aid, not a score. Each answer should point to a current artefact, a demonstrated action or a named owner. 1. Which map connects the critical customer journey to the services, data stores and suppliers behind it? 2. Which inventory records the version, owner, maintenance state and exit constraint of each critical dependency? 3. Which source revision produced the deployed artefact, through which verifiable build steps? 4. Can someone on the receiving team run the build, deployment and intended rollback now? 5. Which record connects accounts, identities, roles, secrets and revocation rights to organisational owners? 6. Which results cover the critical journey under normal, boundary, invalid and abuse conditions? 7. Which usable signal lets a named responder diagnose a customer-visible failure? 8. Which rollback or restoration of the critical scope has been exercised, observed and recorded? 9. Who detects, decides, responds, communicates and restores service during an incident? 10. Which known debt, decision records and practical transfer exercises has the receiving team examined? ## Decide: take over, harden or pause The result depends on the product, transaction and consequences - not a universal threshold. | Observed situation | Decision | Next move | |---|---|---| | Critical journeys are understood; build, release, observation and recovery have been demonstrated; responsibilities are assigned. | Accept and operate | Transfer access progressively and watch the first changes under the new owners. | | The core is operable, but bounded gaps remain in security, testing, dependency exit or operations. | Accept with bounded hardening | Define risks, owners, exit criteria and the order of remediation. | | The team cannot build, deploy or restore the critical scope, or material unknowns prevent accountable ownership. | Pause and rescue | Stabilise access and service, reconstruct missing evidence, then reassess the handoff. | ## What a professional assessment should produce A useful assessment does more than list defects. It should map architecture, dependencies and ownership; record which builds, releases, controls, signals and recovery exercises were actually observed; and separate verified facts from unresolved questions. The [NIST Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) can provide shared language for buyer-supplier discussions, but it is not a takeover certification. Findings should be prioritised by business consequence: an unavailable journey, exposed data, blocked release, fragile recovery or dependency on one person. The output should define owners, sequencing and exit criteria for an executable remediation path. It should also identify where specialist work - such as penetration testing, legal review or intellectual-property analysis - is outside scope. Automated scanners can surface candidates for investigation. Material findings still need architectural context, exploitability review and human validation before they become decision evidence. A clean scan cannot establish safety or production readiness. ## Conclusion: ownership has to be demonstrated Possessing the files is not the same as owning the product operationally. A credible handoff exists when another team can understand, change, release, observe and restore the service under explicit human responsibility. Where the evidence is already present, the decision may be accept and operate. Other products need a bounded hardening plan. A rewrite is not the default verdict. ## Prepare the handoff before it becomes an emergency For a 30-minute scoping conversation, bring repository read access or an export; any architecture and deployment notes available, even if incomplete; one critical customer journey; known incidents or support pain; and the date and context of the handoff, enterprise review or deal. The conversation establishes an initial decision path: accept and operate, define bounded hardening, or prepare a rescue. This is not a free audit, certification or promise of buyer acceptance. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is AI-generated code inherently bad? No. The available findings vary by task, method and context ([Molison et al.](https://arxiv.org/abs/2508.00700)). Assess the product’s critical behaviour, security, operability and accountable ownership rather than its origin label. ### Are automated scanners sufficient before a handoff? No. They can identify items for investigation, but relevant findings need validation and context. Scanners do not show by themselves whether the team can release, recover or support the service. ### When may a rewrite be justified? When the architecture persistently blocks the target journey, critical dependencies can no longer be maintained, or restoring an operable core is less viable than rebuilding the useful scope. Compare explicit scenarios and consequences; “AI-built” is not a sufficient reason. ### Does the original builder need to remain during the transfer? Their availability can help with walkthroughs and bounded support, but it should not become a permanent dependency. An important evidence point is that the receiving team performs critical procedures itself. Google SRE describes training and progressive transfer as parts of accepting production responsibility, without establishing a universal process ([Google SRE](https://sre.google/sre-book/evolving-sre-engagement-model/)). ### Does a handoff assessment replace a security or legal review? No. It can identify escalation needs, but the scope must be explicit. Penetration testing, regulatory analysis, contract review and intellectual-property checks require their own evidence and expertise. For transactions, see IZZY’s separate [technical due diligence service](https://izzy.agency/en/services/due-diligence/). ## Sources and method This framework synthesises official guidance, technical specifications and bounded open research checked on 22 July 2026. The sources inform what evidence to request; they do not establish a universal handoff standard or measure every AI-built product. Research findings remain limited to their stated methods and samples. - [NCSC - The “vibe coding spectrum” approach to AI-assisted software development](https://www.ncsc.gov.uk/blogs/the-vibe-coding-spectrum-approach-to-ai-assisted-software-development) - [OWASP - Secure Coding with AI Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Secure_Coding_with_AI_Cheat_Sheet.html) - [NIST - Secure Software Development Framework, SP 800-218](https://csrc.nist.gov/pubs/sp/800/218/final) - [Google SRE - The Evolving SRE Engagement Model](https://sre.google/sre-book/evolving-sre-engagement-model/) - [Google SRE - Reliable Product Launches at Scale](https://sre.google/sre-book/reliable-product-launches/) - [Google Cloud - Perform testing for recovery from failures](https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-failures) - [SLSA - Provenance, specification v1.2](https://slsa.dev/spec/v1.2/provenance) - [NIST - Incident Response Recommendations, SP 800-61 Rev. 3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) - [OpenTelemetry - Observability primer](https://opentelemetry.io/docs/concepts/observability-primer/) - [Molison et al. - Is LLM-Generated Code More Maintainable & Reliable than Human-Written Code?](https://arxiv.org/abs/2508.00700) - [Geruslu, Aliyeva and Tüzün - Factors Influencing the Quality of AI-Generated Code](https://arxiv.org/abs/2603.25146) --- ### Do you need an AEO tool? A practical guide for small businesses URL: https://izzy.agency/en/blog/do-you-need-an-aeo-tool/ Published: 2026-07-28 Summary: Learn what you can measure for free, what paid AI-visibility tools add and which evidence to request before your business subscribes. ## The 60-second answer Probably not on day one. First establish a repeatable baseline around the questions that precede a purchase. Record whether your company is eligible to be found, mentioned, recommended, cited and - when somebody clicks - responsible for a qualified visit. Search Console, Bing Webmaster Tools, analytics and controlled manual tests can reveal parts of that picture without a specialist subscription. A paid tool becomes useful when repeated collection, multiple markets, competitor tracking and reporting cost more to run manually than the subscription saves. It should also expose the questions and answers behind its score. Otherwise, you have bought a dashboard before defining the decision. An AEO tool can help you observe a market. It cannot guarantee that an assistant will recommend you. ## In this article 1. [The five layers hidden inside AI visibility](#1-the-five-layers-hidden-inside-ai-visibility) 2. [How to establish a useful baseline without an AEO tool](#2-how-to-establish-a-useful-baseline-without-an-aeo-tool) 3. [What paid AI-visibility tools actually add](#3-what-paid-ai-visibility-tools-actually-add) 4. [When to stay manual, subscribe or commission an audit](#4-when-to-stay-manual-subscribe-or-commission-an-audit) 5. [Seven questions to ask before buying](#5-seven-questions-to-ask-before-buying) 6. [What to fix after the measurement](#6-what-to-fix-after-the-measurement) ## 1. The five layers hidden inside “AI visibility” Here, an AEO tool means software that monitors how a brand or its pages appear in AI-generated answers. “Are we visible?” sounds like one question. It is at least five. | Layer | What you are checking | What it does not prove | |---|---|---| | **Eligibility** | Can the relevant system access and use the page or business information? | That it retrieved, cited or recommended the company | | **Mention** | Does the company appear anywhere in the answer? | That it is presented positively or as a suitable option | | **Recommendation** | Is the company proposed for the buyer's stated need? | That the supporting facts are correct or that your site is cited | | **Citation** | Is your page used as a visible source? | That the company itself is recommended or that every linked claim is supported | | **Qualified referral** | Did a person arrive and take a useful next step? | That the AI interaction caused the entire buying decision | These layers can move independently. An assistant may cite your research while recommending a competitor. It may name your business without linking to it. A visitor may arrive from ChatGPT after several earlier interactions that analytics cannot see. Tool vendors also use different definitions. [Ahrefs](https://help.ahrefs.com/en/articles/15501968-ai-visibility-metrics) separates mentions, citations, estimated impressions and AI Share of Voice. [Semrush](https://www.semrush.com/kb/1596-visibility-overview-report) presents its own visibility score alongside mentions, citations and cited pages. Both can be useful within their declared methods; the headline numbers are not interchangeable. Start with the layer your decision requires. Inspect answers to correct product facts, citations to understand source use and attributable visits to assess website actions - while keeping the attribution limit visible. ## 2. How to establish a useful baseline without an AEO tool A manual baseline is not a free version of an enterprise platform. It is a smaller experiment designed around a real decision. ### Choose questions from the buying process Collect questions from sales calls, proposals, support conversations and Search Console. Include situations in which a buyer: - defines the category; - compares approaches or suppliers; - introduces a constraint such as market, budget, integration or regulation; - checks risk or credibility; - asks for a shortlist. Avoid variations invented only to make the company appear. [OpenAI's evaluation guidance](https://developers.openai.com/api/docs/guides/evaluation-best-practices) recommends task-specific tests that reflect real-world use, logged evidence and human judgement. It addresses AI applications, but the discipline transfers: define the decision, then choose tests that represent it. ### Record the conditions For every test, retain: - the exact question; - the assistant or search product; - the market and language; - the date; - whether the session was new or already personalised; - the full answer and visible sources; - the rule used to label mention, recommendation and citation. Generated answers can vary. Repeat decision-critical questions and report the observed range. There is no universal responsible number of prompts or repetitions; the method must justify its coverage. ### Use first-party signals where they exist Google says the same SEO foundations apply to AI Overviews and AI Mode. Eligible pages must be indexed and able to appear with a snippet, but there is no special AI schema or extra technical requirement - and eligibility does not guarantee inclusion. AI-feature traffic is included in Search Console's Web performance data rather than a complete cross-platform AI report ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)). Bing's AI Performance preview can show citations, cited pages, sampled grounding queries, URL-level activity and trends across supported Microsoft AI experiences. Bing says these measures do not indicate answer placement, page importance, authority or ranking ([Bing Webmaster Tools](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview)). OpenAI says publishers can allow `OAI-SearchBot` to access content intended for ChatGPT search. ChatGPT referral links include `utm_source=chatgpt.com`, allowing attributable visits to be analysed ([OpenAI publisher guidance](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq)). That records clicks, not unclicked mentions or private conversations. Together, these sources provide pieces of the baseline. None is a universal scoreboard. ## 3. What paid AI-visibility tools actually add The strongest reason to buy a tool is operational reliability, not access to a secret optimisation method. A suitable platform may add: - scheduled collection instead of manual checks; - consistent storage of questions, answers and sources; - coverage across more assistants, countries or languages; - competitor comparison under the same protocol; - history that makes movement easier to investigate; - exports and access controls for several stakeholders; - alerts when an important answer or source changes. These capabilities do not remove the need to choose representative questions, define counted events or inspect answers. Two tools may use different question databases, engines, locations, dates, entity definitions and formulas. Their scores are compressed views of different methods - not fixed ranks across the AI market. Automation also scales mistakes. A biased question set remains biased; an ambiguous brand can distort comparison; a tool that treats every mention as success can reward an inaccurate recommendation. ## 4. When to stay manual, subscribe or commission an audit Use the lightest method that can support the decision. | Situation | Sensible starting point | Why | |---|---|---| | One brand, one principal market and a small number of important buying questions | **Controlled manual baseline** | You need to learn what should be measured before automating it | | Repeated monitoring across several products, competitors, markets or languages | **Paid tracker** | Collection, history and comparison are becoming an ongoing operational burden | | The company does not agree on the buyer questions, offer, evidence or meaning of success | **Scoped independent audit** | A subscription will automate an unresolved measurement and ownership problem | | A dashboard already exists but nobody trusts or acts on it | **Method and evidence review** | The problem may be definitions, sampling or missing raw evidence rather than data volume | Do not choose only by subscription price. Compare it with the time needed to run the protocol, check evidence, explain changes and coordinate corrections. A managed service should likewise show what was tested, observed and left uncertain - not hide the method behind a score. ## 5. Seven questions to ask before buying Use this evidence test during a demo or proposal review. | Question | Evidence worth requesting | Warning sign | |---|---|---| | **Where do the questions come from?** | Buyer, search, sales or research provenance; branded balance; market scope and exclusions | A large prompt count with no connection to your buying process | | **Which systems and conditions are tested?** | Named products, interfaces, markets, languages, session conditions and dates | Results blended across undisclosed environments | | **How is answer variation handled?** | Repeated observations for important questions, retained dates and visible ranges | One answer presented as a stable rank | | **What does the score count?** | Definitions, formula, denominator and weighting | Mentions, recommendations and citations treated as equivalent | | **Can we inspect full answers and sources?** | Raw responses, source URLs, exports and a change log | Only a score, chart or selected screenshot | | **Are accuracy and buyer fit reviewed?** | Factual checks, recommendation criteria and documented valid exclusions | Every appearance counted as a win | | **What decision will this report change?** | Named owner, correction path and review cadence | More monitoring with no operational next step | A historical study of four generative search systems treated citation coverage and citation support as separate properties ([Liu, Zhang and Liang, EMNLP 2023](https://aclanthology.org/2023.findings-emnlp.467/)). Its results do not describe today's market, but the distinction remains useful: a link is not automatic proof of accuracy. ## 6. What to fix after the measurement Measurement earns its cost when it narrows the next decision. If a page is not eligible or retrievable, inspect crawl access and indexability. If the company is described inaccurately, compare product pages, documentation, profiles and external sources. If competitors are recommended, inspect the criteria and sources supporting them before commissioning more generic content. Some gaps cannot be solved by SEO alone. Product owns offer clarity, experts own evidence and communications influences external recognition. Someone must own the combined outcome. See [GEO is not just an SEO problem](https://izzy.agency/en/blog/geo-ai-visibility-company-ownership/). If credible third parties used in answers do not know you, the work may involve research, partnerships, specialist communities or earned coverage - not schema. See [what to build before AI visibility becomes more paid](https://izzy.agency/en/blog/ai-visibility-citations-earned-media/). Do not confuse publishing an AI-facing file with proving retrieval or a useful outcome. [Our llms.txt and WebMCP analysis](https://izzy.agency/en/blog/llms-txt-webmcp-website-ai-agents/) explains the difference. Correct the first explainable gap, keep the protocol stable and measure again. No movement is also information: the source may not have been retrieved, the evidence may remain weak or the correction may not affect the tested answer. ## Conclusion: buy measurement after you define the decision A small business does not need an enterprise dashboard to begin observing AI visibility. It needs real buying questions, clear definitions, stable conditions and evidence it can inspect. Start manually. Learn which answers and sources matter. Buy software when repetition, scale and coordination justify automation. Commission an audit when the harder problem is deciding what to measure, why the gap exists and who can correct it. The objective is not to own a better score. It is to make a better decision. ## Establish an AI-visibility baseline your team can inspect Bring us your principal market, buyer questions, current dashboard - if you have one - and the decisions the report needs to support. IZZY can scope a fixed [AI Visibility Audit](https://izzy.agency/en/services/ai-visibility-audit/) that separates eligibility, mentions, recommendations, citations and attributable visits, then prioritises the first explainable corrections. No guaranteed citations and no invented universal rank: a bounded method, retained evidence and a clear next decision. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Are AEO tools worth it for a small business? Yes, once repeated monitoring across important questions, competitors or markets becomes operational work. Establish a manual baseline first so you know what is worth paying for. ### Can I measure AI visibility without a paid tool? Yes, within a limited scope. Test selected buyer questions under recorded conditions, inspect relevant first-party tools and measure attributable referrals. Call the result a bounded baseline, not total market visibility. ### Can an AEO tool make my business rank in ChatGPT? No. A platform can collect observations, identify cited sources and organise corrections, but it cannot guarantee a stable ChatGPT rank or recommendation. ### Are AI visibility scores reliable? A documented score can benchmark the questions, systems and conditions tested. It is not a universal measure of every AI conversation. Request the formula, question provenance, repeated results and raw answers. ### How many prompts should an AI visibility audit test? There is no universal number. The set should represent the relevant buying decisions, markets, languages and constraints. The provider should explain its coverage, exclusions and fitness for the decision. ### Does Google Search Console show ChatGPT traffic? No. Search Console reports Google Search performance. ChatGPT visits should be inspected in analytics using available referral and UTM source information. OpenAI says its referral links include `utm_source=chatgpt.com`, but analytics still cannot observe unclicked mentions or private conversations. *Sources and product documentation checked on 20 July 2026. Product features can change. This article describes a measurement and purchase-decision framework; it does not guarantee a citation, recommendation, ranking or commercial outcome.* --- ### Google AI Overviews in France: How to Measure the Impact on SEO Traffic and Enquiries URL: https://izzy.agency/en/blog/ai-overviews-france-seo-traffic-enquiries/ Published: 2026-07-27 Summary: AI Overviews is rolling out in France. Learn how to measure its impact on Google visibility, SEO traffic and qualified enquiries. ## The 60-second answer Google began rolling out **AI Overviews** - labelled *Aperçus IA* in the French interface - and **AI Mode** in France on 22 July 2026. AI Overviews can display a generated answer above the traditional search results. AI Mode lets people continue their research as a conversation, including comparisons and follow-up questions ([Google France announcement](https://blog.google/intl/fr-fr/nouveautes-produits/explorez-obtenez-des-reponses/recherche-ia-apercus-mode/)). For an SME operating in France, the useful question is not "Do we need to redo all our SEO?" It is: > Can prospects still find us when they describe their problem, compare their options and decide which company to contact? During the first few weeks: 1. preserve a baseline in Search Console, GA4 and your CRM; 2. select the questions that genuinely precede a sale; 3. observe the AI Overview, then continue the journey with follow-up questions in AI Mode; 4. record the pages and competitors cited, as well as the point where your company disappears; 5. check whether visible pages lead to a clear offer, credible evidence and an appropriate next step; 6. fix the observed break before producing more content. The French rollout is too recent to conclude that your traffic will fall. First establish what is changing for your market, queries and enquiries. ## In this guide - [1. What is actually changing in Google Search](#1-what-is-actually-changing-in-google-search) - [2. Is your company exposed?](#2-is-your-company-exposed) - [3. What Search Console shows (and what it does not)](#3-what-search-console-shows-and-what-it-does-not) - [4. The diagnostic to run during the first 30 days](#4-the-diagnostic-to-run-during-the-first-30-days) - [5. What to fix first](#5-what-to-fix-first) - [6. False urgencies to avoid](#6-false-urgencies-to-avoid) - [7. Should you exclude your site from generative Search?](#7-should-you-exclude-your-site-from-generative-search) ## 1. What is actually changing in Google Search An AI Overview attempts to answer a question directly using several sources. It does not appear for every search: Google says it is triggered when its systems determine that a generated answer adds something beyond the traditional results. AI Mode goes further. It can break a request into several subtopics and run multiple related searches. Google calls this **query fan-out**. A prospect can ask for a solution, add a budget, compare providers, specify a location and raise an objection without restarting the journey ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)). This creates three changes for a business. ### Search becomes a journey A position for one query no longer represents the whole visibility picture. Your brand may appear in the initial answer and disappear when the user asks for a comparison table or adds a constraint. ### Visibility and clicks become more separate A page may be used as a source without receiving a visit. Conversely, someone may discover your brand in a generated answer and return later through a branded search, direct visit or another channel. ### A visible page must do more A click made after a detailed generated answer may come from someone looking for evidence, a precise detail, pricing, a methodology or a person to contact. The landing page must continue that intent instead of repeating a generic introduction. This does not make SEO irrelevant. Google says its generative features still rely on its index and core Search systems. Traditional rankings, AI citations and commercial outcomes should nevertheless be measured separately. ## 2. Is your company exposed? Not every search presents the same risk or opportunity. Start by classifying the questions your prospects ask. | Search type | Example | What to observe | |---|---|---| | Brand | "izzy agency reviews" | Are the brand, evidence and official information still represented accurately? | | Problem | "website gets traffic but no enquiries" | Is your expertise visible before the prospect knows your offer? | | Solution | "website conversion audit for an SME" | Are your services understood and connected to the right problem? | | Comparison | "agency or freelancer for rebuilding a SaaS product" | Are you still proposed when criteria and alternatives enter the answer? | | Local | "web development agency Brittany" | Are the service area, Business Profile, reviews and services consistent? | | Decision | "how much does an SEO and GEO audit cost" | Does the answer lead to a page that genuinely helps someone decide? | Focus on searches capable of changing a commercial decision. Fewer clicks on a general definition do not have the same consequence as disappearing from a provider comparison. A Pew Research Center study observed the browsing behaviour of 900 US adults in March 2025. Participants clicked external results less often when the page contained an AI-generated summary. The study does not measure France in 2026 or your business; it is a reason to separate impressions, clicks and commercial outcomes - not a forecast for a French SME ([Pew Research Center](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/)). A separate longitudinal study published in 2026 analysed more than 55,000 queries and observed that some domains cited in AI Overviews did not appear among the traditional results displayed on the first page. Again, this is not French market data or a promise that smaller companies will gain access. It shows why an organic position alone cannot represent AI visibility ([study on arXiv](https://arxiv.org/abs/2605.14021)). ## 3. What Search Console shows (and what it does not) Google has started rolling out a dedicated performance report for its generative AI features. It includes impressions from experiences such as AI Overviews and AI Mode. The report can show: - how often a URL from your site appeared in a generative feature; - which pages received those impressions; - the countries, devices and dates involved. Access is still rolling out. If the report is not available, your property may not yet be included or may not have enough impressions ([Search Console documentation](https://support.google.com/webmasters/answer/16984139?hl=en)). | Question | How Search Console can help | Current limitation | |---|---|---| | Does the site appear in Google's generative features? | Impressions in the dedicated report | Not every property has access yet | | Which pages appear? | Page dimension | The page alone does not explain why it was selected | | In which country and on which device? | Country and device dimensions | These do not describe the buyer's profile or intent | | Which exact question triggered the citation? | The general Web report has a Queries dimension | It does not cleanly isolate generative questions or reconstruct the entire journey | | Did the person click and then submit an enquiry? | Search Console counts outbound clicks; GA4 and the CRM can track subsequent actions | The tools do not automatically connect the full conversation to a prospect | Google treats every follow-up question in AI Mode as a new query. Impressions, positions and clicks in the new response are attributed to that new question ([Search Console methodology](https://support.google.com/webmasters/answer/7042828?hl=en)). The consequence is simple: one screenshot or average visibility score does not prove that your brand remains present throughout the journey. To measure what happens after a click, use [GA4 with a controlled classification for AI referrals](https://izzy.agency/en/blog/track-ai-traffic-ga4/). To investigate a possible decline, keep demand, visibility, clicks and measurement failures separate, as explained in [Your website traffic is down: find the break before restarting marketing](https://izzy.agency/en/blog/website-traffic-drop/). ## 4. The diagnostic to run during the first 30 days The following is an IZZY framework. It is not a Google standard, and it does not mean that one month is enough to judge a long sales cycle. ### Before day 1: preserve a comparable baseline Export at least: - Search Console performance for priority pages and queries; - sessions, landing pages and key events from GA4; - enquiries received, accepted and converted in the CRM; - recent website, campaign and measurement changes. Record the period, country, device and metric definitions. Without a comparable baseline, a change after the rollout will remain difficult to attribute. ### Week 1: build the buying-question map Ask sales, support and leadership which questions repeatedly arise before a contact or sale. For each offer, select a small set that covers: 1. the initial problem; 2. the search for a solution; 3. comparison criteria; 4. objections; 5. the desired next step. Do not turn this into one hundred artificial variants. Google says its systems understand synonyms and that pages do not need to be rewritten for every possible AI query. ### Week 2: observe the journey, not only the first answer For each scenario: - check whether an AI Overview appears; - record the brands, pages and sources shown; - continue with a realistic comparison or objection in AI Mode; - note where your brand enters, remains or leaves the response; - preserve the date, country, device and wording. One run remains directional. Responses can vary with context and change over time. The objective is not to manufacture an absolute ranking but to identify patterns that can be investigated. ### Week 3: examine the pages and route to enquiry For every visible page, ask: - does it answer the question that caused it to appear? - does it distinguish fact, experience, opinion and commercial promise? - does it present evidence the company can defend? - are the offer, service area, conditions and intended client understandable? - does the next step work on mobile? - does the enquiry reach the right person and receive a response? Being cited on a page that does not explain the offer or help someone contact the company does not solve the acquisition problem. If visitors already arrive but do not become enquiries, begin with [the five conversion checks](https://izzy.agency/en/blog/website-visitors-no-enquiries/). ### Week 4: decide what to repair | Observed signal | First hypothesis to test | Possible work | |---|---|---| | No generative impressions, including for relevant questions | Eligibility, indexing, access or insufficient data | Technical and Search Console diagnostic | | The site appears for general questions and disappears during comparisons | Insufficient coverage of decision criteria | Comparison content, evidence and positioning | | A page is cited but receives almost no clicks | The answer satisfies the intent or the link offers little additional value | Improve post-click usefulness; monitor without overreacting | | Clicks exist but enquiries remain weak | Offer, evidence, friction or commercial follow-up | Conversion diagnostic | | Enquiries increase but are poorly qualified | Wrong promise, context or missing criteria | Clarify targeting and qualification | The useful deliverable is not a list of trends. It is an order of work connected to an observable break. ## 5. What to fix first ### 1. Technical access and indexing To be eligible as a supporting link in Google's generative features, a page must be indexable and eligible to appear in Search with a snippet. Check crawling, `robots.txt` rules, indexing directives, JavaScript rendering, canonical URLs and internal links. ### 2. Information your company can genuinely prove Add what generic summaries cannot provide: field experience, methodology, limits, publishable internal data, precise examples, decision criteria and answers to objections. A commercial claim without context remains difficult to evaluate for both users and search systems. AI visibility is not owned only by an SEO writer. Product, sales, brand, data and PR teams often hold the information that makes content verifiable. This is why [GEO is also a company ownership problem](https://izzy.agency/en/blog/geo-ai-visibility-company-ownership/). ### 3. Product and local information For a local company, keep Google Business Profile information current. For a catalogue, check Merchant Center data such as prices, availability, identifiers, variants and consistency with the website. These foundations also support traditional Search; they do not guarantee a citation. ### 4. The journey after the answer A page intended for someone who already has a summary should help them move forward: understand the difference, verify evidence, assess risk or contact the right person. Connect explanatory content to the relevant services, cases, methods and next steps. ### 5. Measurement and commercial follow-up Bring Search Console, GA4 and CRM evidence together without pretending they see the same thing. Search Console observes Google visibility and clicks. GA4 observes part of the resulting visits. The CRM shows what the business did with an enquiry. ## 6. False urgencies to avoid Google has published explicit guidance on practices that are not required to appear in its generative features ([official guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide)). Do not urgently commission: - an `llms.txt` file presented as a Google visibility factor; - special "AI" markup or a new schema presented as mandatory; - mechanically breaking every page into small content chunks; - rewriting every page "for LLMs"; - a campaign of manufactured mentions; - a large volume of articles repeating information already available elsewhere; - one score presented as your brand's real position across every generated answer. `llms.txt` may have other experimental uses, but Google says it does not use the file for visibility in Search. We explain the distinction in [Does llms.txt improve your visibility in ChatGPT?](https://izzy.agency/en/blog/llms-txt-webmcp-website-ai-agents/). Good preparation looks less like a hunt for hacks and more like sound operational work: an accessible website, accurate information, non-commodity content, evidence, useful pages and honest measurement. ## 7. Should you exclude your site from generative Search? Google says a Search Console control is available in France that lets website owners decide whether a site may appear in generative Search and help ground its answers ([French announcement](https://blog.google/intl/fr-fr/nouveautes-produits/explorez-obtenez-des-reponses/recherche-ia-apercus-mode/)). A site that opts out no longer receives the impressions or traffic associated with these features. According to Google, the choice is not used as a ranking signal for results outside generative Search ([how the control works](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/)). Do not treat this toggle as a routine SEO setting. | Situation | Question to answer before changing anything | |---|---| | SME looking for prospects | Are you willing to give up this discovery surface and its potential visits? | | Publisher funded by audience traffic | What value are you comparing across visibility, clicks, revenue and content use? | | Sensitive or contract-governed content | Which obligations, permissions and restrictions actually apply? | | Decision remains uncertain | Can you preserve the current setting, establish a baseline and decide using evidence? | For questions involving rights, contracts or compliance, have the decision validated by the appropriate adviser. This article is not legal advice. ## Conclusion: measure useful demand, not only presence in an answer AI Overviews and AI Mode add a new layer between a prospect's question and your website. They do not turn every French SME into an SEO emergency, and they do not prove a decline in French traffic on the day of launch. Start with the questions that precede a sale. Observe the initial answer and the follow-up questions. Connect visible pages to clicks, enquiries and their quality. Then correct the first demonstrated break: technical access, content, evidence, offer, conversion or follow-up. The right question is not "Are we cited?" It is "Do we remain a credible option when the prospect decides what to do next?" ## Bring us the questions where your company must remain visible Bring your Search Console and GA4 exports, commercial pages, a sample of recent enquiries and the questions prospects ask before contacting you. In 30 minutes, IZZY can frame whether the break is in Google visibility, the generated answer, the content, the page or commercial follow-up - and define the first useful test. We do not promise citations, rankings or recovered traffic from one call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is Google AI Overviews available in France? Yes. Google announced the rollout of AI Overviews and AI Mode in France on 22 July 2026 across desktop, mobile and the Google app. An AI Overview does not appear for every query. ### How can a company appear in Google AI Overviews? There is no special markup that guarantees inclusion. A page must be accessible, indexable and eligible for a Search snippet. Google recommends established SEO fundamentals, useful original content, clear technical structure and current business information. ### Is traditional SEO still useful? Yes. Google says its generative features still use its Search index and ranking systems. A traditional ranking does not guarantee a citation, and a citation does not guarantee a click. ### How can a company measure visibility in AI Mode? Check the generative Search Console report if it is available for your property. Combine it with a stable set of scenarios, observation of follow-up questions, GA4 and CRM data. Report visibility, clicks and commercial outcomes separately. ### Will AI Overviews reduce SEO traffic to my French website? This cannot be verified without the site's own post-rollout data. Research from other markets has observed fewer clicks when generated summaries appear, but those results cannot automatically be applied to your market, pages or queries. ### Do we need an llms.txt file or special schema? Not for Google Search. Google says it does not use `llms.txt` and does not require a special file or schema for its generative features. Continue using structured data that supports established rich results and matches the visible content. ### How long should impact be monitored? Start with a baseline and an initial 30-day cycle, then adapt the period to search volume and the sales cycle. A seasonal business or low-volume B2B company will require more time. *Features and documentation checked on 22 July 2026. Search Console access, generated-answer presentation and feature availability may change. This article provides a diagnostic method, not a guarantee of citations, traffic or commercial outcomes.* --- ### Users sign up but do not activate. Where is your onboarding breaking? URL: https://izzy.agency/en/blog/onboarding-activation/ Published: 2026-07-23 Summary: Users sign up but never activate? Define activation as real value, validate events and cohorts, find the first break in the journey, then fix the smallest part. ## Answer in 60 seconds If users create an account but do not use the product, do not begin by adding tooltips or extending welcome sequences. First: - define the user, job and context you are analysing; - name the observable event that demonstrates first meaningful value; - verify that the event and cohort denominator are measured correctly; - find where users stop between expectation, access, setup, first task and useful result; - combine behavioural data with interviews, observation and support evidence; - test the smallest justified change on a bounded cohort. Signup, login, onboarding completion and activation are different events. A user can complete every onboarding screen without achieving the outcome they came for. There is no universal “good activation rate”. The useful rate depends on the product, eligible population, value event, time window and acquisition context. ## In this guide - [1. Define activation as demonstrated value](#1-define-activation-as-demonstrated-value) - [2. Validate the event and the denominator](#2-validate-the-event-and-the-denominator) - [3. Locate the first break in the journey](#3-locate-the-first-break-in-the-journey) - [4. Redesign the shortest credible path to value](#4-redesign-the-shortest-credible-path-to-value) - [5. Choose measurement, onboarding, product or human support](#5-choose-measurement-onboarding-product-or-human-support) ## 1. Define activation as demonstrated value Start with a sentence: > For this user in this context, activation means completing **this meaningful task** and receiving **this useful result** within **this relevant period**. The event should represent evidence of value, not evidence that your interface was displayed. “Viewed the tour” may describe exposure. “Invited a colleague” or “published the first page” may be closer to value when that action matters to the user. Learn the need before choosing the milestone. GOV.UK’s service guidance recommends using analytics, search logs and support data alongside interviews and observation to understand who users are and what they are trying to do. It also warns that suggestions from people who are not users are assumptions to investigate, not user needs ([GOV.UK](https://www.gov.uk/service-manual/user-research/start-by-learning-user-needs)). Treat activation as an operational definition, not a universal product metric. ## 2. Validate the event and the denominator Before redesigning onboarding, check: 1. Who entered the cohort: every account, eligible users, invited users or verified humans? 2. When does the clock start: registration, invitation acceptance or first eligible session? 3. Does the event fire once, repeatedly or before the task succeeds? 4. Are internal users, tests, bots, duplicates and failed imports excluded? 5. Can you reconcile product events with account and support records? 6. Are web, mobile and regional journeys measured consistently? Benchmarks can help you ask questions, but they cannot define your target. Amplitude’s 2025 Product Benchmark Report analysed customer-product data from September 2023 to September 2024 and reported that 69% of products in the top quartile for seven-day activation were also in the top quartile for three-month retention ([Amplitude](https://info.amplitude.com/rs/138-CDN-550/images/the-product-benchmark-report.pdf)). That is a correlation in a vendor dataset, not proof that pushing one activation event will cause retention. Instrumentation, product mix, customer selection and opt-outs limit transferability. ## 3. Locate the first break in the journey Use one cohort and trace the path: | Journey point | Evidence to inspect | First question | |---|---|---| | Expectation | landing page, sales promise, campaign and signup reason | Did the acquired user expect the job the product actually supports? | | Access | verification, authentication, invitation and permissions | Could the user enter without avoidable cognitive or technical burden? | | Setup | data, integration, configuration and team dependencies | Is setup necessary for value, and can users understand its consequence? | | First task | event sequence, screen recording and observation | Can the user identify and complete the next meaningful action? | | Result | output quality, error state and time to result | Does the product return a useful and credible outcome? | | Continuation | later sessions, collaboration, support and cancellation | Is there a reason and route to return? | Do not force one explanation onto every segment. A founder arriving from a referral, an employee invited into a team and a buyer evaluating a trial may have different tasks and stopping points. A systematic review of 147 articles on information-system continuance grouped antecedents into psychological, technological, social and behavioural factors and highlighted an intention–behaviour gap ([International Journal of Information Management](https://doi.org/10.1016/j.ijinfomgt.2021.102315)). It supports a multi-factor investigation, not a SaaS activation formula. Prediction is not explanation either. A 2023 digital-health study found that first-seven-day login patterns could help predict early dropout in four smoking-related interventions, with model performance varying across interventions ([JMIR](https://www.jmir.org/2023/1/e43629/PDF)). The population and products are not a general SaaS benchmark, and predictive association does not prove why a user left. ## 4. Redesign the shortest credible path to value Once the break is observed, remove work that does not protect the user, service or result. - Ask only for information needed at that point. - Explain why a permission, connection or configuration matters. - Provide realistic sample data where empty states block understanding. - Preserve progress when users leave or encounter an error. - Put contextual help beside the decision, not in a detached feature tour. - Offer a human route when setup carries commercial, technical or organisational risk. Authentication can itself block activation. W3C’s guidance for accessible authentication explains that users should not be required to complete a cognitive-function test unless an alternative or assisting mechanism is available; support for password managers and copy-and-paste can reduce burden ([W3C](https://www.w3.org/WAI/WCAG22/Understanding/accessible-authentication-minimum.html)). That guidance addresses a specific WCAG success criterion. It does not establish full accessibility conformity, and client-specific assessment still requires qualified review. ## 5. Choose measurement, onboarding, product or human support | Decision | Use it when | Guardrail | |---|---|---| | Repair measurement | events, identity or denominators do not reconcile | do not optimise an untrusted funnel | | Repair expectation | acquisition promise and actual first job differ | align offer, copy and product reality | | Simplify onboarding | users can reach value but encounter unnecessary work | protect necessary security and consent | | Repair the product | users complete onboarding but the result is weak, late or unreliable | do not hide product defects with education | | Add human onboarding | setup is high-value, infrequent or organisationally complex | define ownership and escalation | | Run a bounded experiment | one observed barrier and success measure are clear | compare eligible cohorts and watch harms | | Stop | no meaningful value event or affected segment is defined | investigate before producing more screens | A professional activation review should leave a segment definition, event dictionary, reconciled cohort, observed journey, evidence register, prioritised hypotheses and measurement plan. ## Conclusion: the next sensible move Onboarding is not the number of screens between signup and the dashboard. Define the first meaningful result. Verify who had a fair opportunity to reach it. Observe the point where the journey stops, then repair the smallest defensible part of the system. ## Bring us the cohort that signed up and disappeared. Bring the acquisition promise, onboarding journey, event definitions, cohort export, setup dependencies, support themes and examples of users who did and did not reach value. In a 30-minute call, IZZY can scope whether the first step is measurement repair, journey research, onboarding redesign, product work or no engagement yet. We will not promise an activation uplift, retention result or accessibility conformity from a call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is activation the same as completing onboarding? No. Onboarding completion shows that a user passed through the defined flow. Activation should represent an observable result that matters to that user and product. ### What is a good user activation rate? There is no universal rate. Define the eligible cohort, value event, window and segment, then establish a trustworthy baseline and compare like with like. ### Should every SaaS product use a product tour? No. Use a tour only when it resolves an observed comprehension problem. A contextual instruction, realistic example, simpler default or human setup route may be more appropriate. ### How should we measure time to value? Measure from a clearly defined eligible start event to a verified value event. Report the distribution by meaningful segment; one average can hide users who never succeed or wait much longer. ### What should we bring to an activation review? Bring acquisition sources, user and job definitions, event schemas, cohort data, onboarding screens, setup dependencies, support records, research notes and examples of successful and unsuccessful first sessions. *Research and source status checked 16 July 2026. This article provides general product and onboarding decision guidance. It is not a product benchmark, accessibility audit, privacy assessment or performance guarantee.* --- ### Is your slow website costing you sales - or hiding a bigger reliability problem? URL: https://izzy.agency/en/blog/slow-fragile-site/ Published: 2026-07-22 Summary: Separate speed, reliability and measurement before you rebuild a slow website: field vs lab evidence, the first bottleneck, releases and tested recovery. ## Answer in 60 seconds A slow website may be a front-end problem, a server bottleneck, an unreliable integration, a bad release or a recovery process nobody has tested. Do not approve a rebuild from one PageSpeed score. First separate three questions: - **Performance:** are real users waiting too long on a critical journey? - **Reliability:** are pages, forms, payments or logins failing intermittently? - **Recovery:** can your team detect the failure, return to a known-good state and prove that data is intact? Use field evidence where it exists, repeated lab tests for diagnosis and telemetry for errors. Stabilise outages and recovery before cosmetic optimisation. Then fix the smallest measured bottleneck. No public benchmark can tell you how much revenue your site is losing. Your journeys, traffic, errors and commercial data must establish the consequence. ## In this guide - [1. Decide whether the problem is speed, reliability or measurement](#1-decide-whether-the-problem-is-speed-reliability-or-measurement) - [2. Read PageSpeed and Core Web Vitals without chasing one score](#2-read-pagespeed-and-core-web-vitals-without-chasing-one-score) - [3. Trace the first material bottleneck](#3-trace-the-first-material-bottleneck) - [4. Stabilise recovery and releases before polishing](#4-stabilise-recovery-and-releases-before-polishing) - [5. Choose tune, stabilise, investigate, rebuild or stop](#5-choose-tune-stabilise-investigate-rebuild-or-stop) ## 1. Decide whether the problem is speed, reliability or measurement “The site is slow” is a symptom, not yet a technical brief. A performance problem is usually repeatable: a meaningful element appears late, the interface responds slowly or content moves while the user is trying to act. A reliability problem is intermittent. A form disappears, checkout times out or an update breaks a template. The team is afraid to deploy because recovery is uncertain. A measurement problem begins with an unqualified complaint or a single lab test being treated as proof of what every visitor experiences. Start with one business-critical journey: enquiry, checkout, login, booking or publishing. Record the device, page, delay or failure, time and completion state. If you cannot name the journey, do not start with site-wide optimisation. ## 2. Read PageSpeed and Core Web Vitals without chasing one score Google's PageSpeed Insights presents two evidence types. **Field data** comes from eligible real Chrome users over a trailing 28-day period. **Lab data** runs Lighthouse in a simulated environment and is useful for diagnosis. They can disagree without either being “wrong” ([Google PageSpeed Insights](https://developers.google.com/speed/docs/insights/v5/about)). A page may lack enough eligible field samples, so PageSpeed may show origin-level data - or none. A green lab score does not prove good real-user experience. Google's current Core Web Vitals cover loading, interactivity and visual stability: - Largest Contentful Paint: 2.5 seconds or less; - Interaction to Next Paint: 200 milliseconds or less; - Cumulative Layout Shift: 0.1 or less; Assessment uses the 75th percentile and should be segmented across mobile and desktop ([web.dev](https://web.dev/articles/vitals)). These are technical thresholds, not revenue targets. In May 2026, the Chrome UX Report covered 18,445,974 eligible origins; 55.9% passed all three Core Web Vitals. This global, origin-level dataset cannot estimate one site's loss or IZZY-market prevalence ([Chrome UX Report](https://developer.chrome.com/docs/crux/release-notes/)). ## 3. Trace the first material bottleneck Do not optimise the homepage because it is the easiest page to test. Follow the journey that matters. | Observed signal | First question | Likely workstream | |---|---|---| | Slow field LCP across one template | Is the delay server-side, rendering-related or the main asset? | measured performance repair | | Good lab result, poor field experience | Which devices, networks, regions or third parties differ? | real-user monitoring and segmentation | | Fast pages with timeouts or errors | Which service, database or integration fails? | reliability investigation | | Failures start after releases | What changed, what was tested and how is rollback authorised? | release-control repair | | Only one inconsistent test looks poor | Can the symptom be reproduced across runs and conditions? | measure before changing | | Conversion fell but the journey is technically stable | Is the first break actually offer, trust, form or follow-up? | conversion diagnosis | Inspect browser requests, JavaScript, server response, database queries, APIs and third parties. Segment by device, template, geography and release where relevant. The 2024 Web Almanac found materially different mobile and desktop results. It remains a dated global sample, not a client diagnosis ([HTTP Archive](https://almanac.httparchive.org/en/2024/performance)). ## 4. Stabilise recovery and releases before polishing If the site is intermittently failing, the first job is not image compression. Establish: 1. critical-journey monitoring and error evidence; 2. the last known-good release; 3. a restoration test and recovery route; 4. a dependency and version inventory; 5. release checks and rollback authority. In Uptime Institute's 2025 report, 80% of 97 data-centre operators believed better management, processes or configuration could have prevented their most recent impactful downtime incident. This small, self-reported, data-centre-heavy signal is not a website outage rate, but it supports looking beyond hosting hardware ([Uptime Institute](https://uptimeinstitute.com/uptime_assets/d7c049ef5b02a6e0a15540a3e5cb8fbf742c7fa54a1af6caeaaab32b7c15d443-GA-2025-05-annual-outage-analysis.pdf)). The UK NCSC recommends inventorying dependencies and controlling how versions enter the system ([NCSC](https://www.ncsc.gov.uk/blogs/software-supply-chain-attacks-check-your-dependencies)). An older package does not prove compromise; version ownership must still be visible. NIST places secure release and vulnerability response across the software lifecycle ([SSDF](https://csrc.nist.gov/pubs/sp/800/218/final)); its configuration guidance adds controlled baselines and monitoring ([SP 800-128 update](https://csrc.nist.gov/pubs/sp/800/128/upd1/final)). Both require tailoring. ## 5. Choose tune, stabilise, investigate, rebuild or stop Use the evidence to choose the smallest justified action: | Decision | Use it when | Do not proceed when | |---|---|---| | Tune performance | Field or repeatable lab evidence isolates a bottleneck and the journey is otherwise stable | The work is based only on a generic checklist | | Stabilise reliability | Errors, outages, release regressions or untested recovery create the larger risk | An active incident requires specialist containment | | Investigate | Field, lab and telemetry disagree or coverage is insufficient | A stakeholder wants a predetermined fix | | Rebuild or replatform | The platform is a demonstrated constraint and parity, migration and recovery can be tested | “The score is low” is the whole business case | | Stop | No material journey, consequence or repeatable symptom has been established | More activity is being used as a substitute for evidence | A performance review should name the bottleneck, controlled change, test, guardrail and recovery route. ## Conclusion: the next sensible move A slow website does not automatically need a new platform. A fragile website does not become reliable because Lighthouse reaches 100. Start with the critical journey. Separate field from lab evidence, and speed from errors, releases and recovery. Stabilise first; then make one controlled change and verify it against the same evidence. ## Bring us the symptom. We'll give you an honest read. Bring one critical journey and the evidence you already have. In 30 minutes, IZZY can scope whether the first step is measurement, a bounded performance repair, reliability stabilisation or no engagement yet. We will not promise a speed, uptime, ranking or revenue result from a call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### What is a good PageSpeed score? Google describes 90 or above as a good Lighthouse lab score, but warns that good lab data does not guarantee good real-user experience. ### Does failing Core Web Vitals mean the website is broken? No. A failed assessment means eligible field data did not meet all three thresholds. If the data is insufficient, no field verdict should be inferred. Neither result identifies the cause or proves lost revenue. ### What is the difference between website performance and reliability? Performance is how quickly a journey responds. Reliability is whether it works consistently and recovers from failure. A site can be fast but unreliable. ### Should we rebuild a slow website? Not before isolating the cause. Consider rebuilding only when the platform is a demonstrated constraint and migration continuity can be tested. ### What should we bring to a website performance review? Bring the affected journey, recent complaints, PageSpeed or real-user data, uptime and error history, recent releases, dependency inventory, hosting or CDN information, and evidence that backup, restore or rollback works. *Research and source status checked 16 July 2026. This article provides general performance and reliability decision guidance. It is not an incident-response service, penetration test, security certification or client-specific architecture recommendation.* --- ### Production context for AI coding agents URL: https://izzy.agency/en/blog/ai-coding-agents-production-context/ Published: 2026-07-21 Summary: Learn how to give AI coding agents useful production context through snapshots, replicas and controlled retrieval - without broad production access. A coding agent can understand a repository and still miss why a change is risky. Code and tests may not show which deployment introduced a failure, which path affects real traffic or what the on-call team already learned. **Production context for an AI coding agent is the task-specific operational evidence that may change an engineering decision.** It can include deployment, incident, error, trace, configuration and ownership information. Production context is evidence. It is not permission to change production. ## Answer in 60 seconds - Start with one engineering task, not a general connection to production. - Select only the evidence that could change the decision. - Define how current and sensitive that evidence is. - Separate permission to retrieve, propose and execute. - Show the reviewer the material evidence behind the recommendation. - Set an access lifetime, revocation path and recovery owner. ## In this article 1. [What production context does an AI coding agent need?](#1-what-production-context-does-an-ai-coding-agent-need) 2. [What can go wrong with production context?](#2-what-can-go-wrong-with-production-context) 3. [Which context boundary should you choose?](#3-which-context-boundary-should-you-choose) 4. [How do you compare the options?](#4-how-do-you-compare-the-options) 5. [How do you implement one workflow?](#5-how-do-you-implement-one-workflow) 6. [When should you pause and escalate?](#6-when-should-you-pause-and-escalate) ## 1. What production context does an AI coding agent need? The useful evidence depends on the task. For incident diagnosis, it may include the affected deployment, a bounded error window, trace examples and the current runbook. For a regression review, it may be a before-and-after comparison tied to two releases. A repository already provides types, dependencies, tests and change history. Operational context adds evidence such as: - deployment identifiers and configuration changes; - incident and alert history for the affected service; - recurring error signatures or bounded metric changes; - trace examples for the failing path; and - current ownership and recovery instructions. Use identifiers to connect sources where the instrumentation supports it. For example, [OpenTelemetry specifies trace and span identifiers for correlating logs with traces](https://opentelemetry.io/docs/specs/otel/compatibility/logging_trace_context/). Correlation does not explain every incident or recover context the team never recorded. Before adding a connector, ask: **what evidence could change this engineering decision?** If the answer is “all production data”, the task is not bounded enough. ## 2. What can go wrong with production context? Design the boundary around failure, not around the longest list of available integrations. | Failure | Operating problem | Design response | |---|---|---| | Thin context | The change ignores deployment or incident history | Add task-specific operational evidence | | Stale context | Evidence belongs to an older release or configuration | Bind it to a deployment, incident or time window | | Overexposed context | The agent receives data the task does not need | Minimise, filter and control retrieval | | Authority creep | Diagnostic access acquires write or execution capability | Separate retrieve, propose and execute permissions | | Unreviewable change | The reviewer cannot reconstruct or reverse the decision | Preserve evidence, approval and recovery steps | [NIST defines least privilege](https://csrc.nist.gov/glossary/term/least_privilege) as the minimum authorisations and resources required for an assigned task. Apply that distinction to evidence, tools and actions separately. [OWASP's AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) treats scoped tools, human approval, monitoring, interruption and rollback as separate controls. No one control establishes assurance. Telemetry also needs an environment-specific data decision. [OpenTelemetry notes that its tooling cannot identify what is sensitive for an operator](https://opentelemetry.io/docs/security/handling-sensitive-data/). Filtering and redaction require configuration and do not remove every disclosure risk. ## 3. Which context boundary should you choose? These five patterns are options, not a maturity ladder. Use the least-exposing pattern that still supports the task. ### Approved static bundle Provide versioned architecture decisions, schemas, tests, ownership and runbooks. This is useful for design work but does not describe current production behaviour. ### Redacted production snapshot Package a bounded set of error signatures, trace examples, deployment metadata and incident notes. It avoids standing access but inherits gaps in instrumentation and redaction. ### Bounded replica or replay Reproduce the behaviour in an isolated environment with synthetic or separately approved data. The replica may differ from the live system, so record that limitation. ### Mediated retrieval Put an identity-aware service between the agent and the operational source. Expose named queries, resources and time windows; filter responses, log calls and expire access with the task. [Google Cloud documents dedicated service identities, temporary privilege and short-lived tokens](https://docs.cloud.google.com/iam/docs/best-practices-service-accounts). These are implementation examples, not a universal agent architecture. ### Exceptional live read Some incident workflows may justify narrow, time-bounded access to current evidence. Give it a named owner, explicit scope, logging and tested revocation. Read-only access reduces change authority but may still expose confidential data or sensitive architecture. Production write authority remains a separate decision and should normally use the team's existing review, deployment and rollback path. [Visual Studio Code provides one product-specific example](https://code.visualstudio.com/docs/agents/approvals) of separating tools, terminal commands, file-system access, network access and approvals. ## 4. How do you compare the options? Do not approve “production access” as one setting. Record the task and boundary together. | Engineering task | Evidence and freshness | Agent authority and access | Reviewer evidence and recovery | |---|---|---|---| | Local refactor | Versioned code, tests and architecture notes | Read and propose a diff | Full diff, tests and revert path | | Incident diagnosis | Redacted logs, traces and current incident window | Retrieve, compare and propose; short-lived identity | Queries, evidence links and no production change | | Regression comparison | Deployment IDs and bounded before/after telemetry | Read and compare until review closes | Selected signals and recorded conclusion | | Sensitive data-fix preparation | Approved sample or bounded replica | Prepare a script; no live execution | Exact script, test evidence and rollback procedure | | Urgent remediation | Minimum current evidence for the named incident | Diagnose and propose through incident authority | Approved change, deployment log and tested rollback | One field is easy to miss: what does the human reviewer see? Private agent evidence turns review into approval of an unexplained conclusion. ## 5. How do you implement one workflow? Start with a workflow that has a clear owner and bounded consequence. 1. **Name the task.** Define the decision, affected service and business consequence. 2. **Inventory the evidence.** List the facts that could change the decision. 3. **Mark freshness and sensitivity.** Decide what may be static, current or excluded. 4. **Choose the boundary.** Use a bundle, snapshot, replica, mediated retrieval or exceptional live read. 5. **Split authority.** Define retrieve, compare, propose, prepare and execute separately. 6. **Design the review.** Show the evidence, proposed change, tests and uncertainty. 7. **Test failure.** Exercise denied access, expiry, revocation, unavailable tools and rollback. 8. **Reopen after change.** Reassess when data, tools, permissions, models or consequences change. Keep the workflow inside the normal development lifecycle. [NIST's Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) provides high-level practices for an SDLC; it does not certify a coding-agent workflow. ## 6. When should you pause and escalate? Pause before expanding the boundary when the workflow may expose personal or regulated data, credentials, confidential client information or security-sensitive architecture; crosses tenants or legal entities; uses privileged write access; affects an active incident; or lacks tested revocation and recovery. IZZY can map the workflow, architecture and decision points. Qualified security, privacy, legal and sector specialists should own conclusions in their domains. ## Conclusion: give the agent evidence, not a control plane The goal is not maximum context. It is enough current evidence for one engineering decision, delivered through a boundary the team can explain and operate. Name the task. Choose the freshness. Minimise exposure. Separate retrieval from execution. Then make the evidence, approval and recovery path visible to the people accountable for production. ## Map one coding-agent context boundary Bring one coding-agent workflow, the context it currently sees, the production evidence it lacks and the review or recovery path you already use. In a 30-minute scoping call, IZZY can map the missing evidence, compare boundary options and identify the decisions that need engineering or specialist ownership. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### 1. What is production context for an AI coding agent? It is task-specific operational evidence - such as deployment, incident, error and trace information - that may change an engineering decision. ### 2. Does a coding agent need direct production access? It should not be the starting assumption. Many tasks can use a static bundle, snapshot, replica or mediated retrieval. ### 3. Is read-only access enough? No. It reduces change authority but may still expose personal data, secrets or sensitive architecture. Scope, filtering, logging and revocation still matter. ### 4. Should the agent and reviewer see the same evidence? The reviewer should be able to inspect the material evidence behind the recommendation. The interface may differ, but the decision should remain reconstructable. ### 5. When should a coding agent change production? That is a separate release decision. Consider impact, reversibility, identity, tests, approval, monitoring and recovery through the team's existing change process. ## Technical evidence note - [OWASP: AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) - [NIST: least privilege](https://csrc.nist.gov/glossary/term/least_privilege) - [Google Cloud: service-account security guidance](https://docs.cloud.google.com/iam/docs/best-practices-service-accounts) - [OpenTelemetry: handling sensitive data](https://opentelemetry.io/docs/security/handling-sensitive-data/) - [OpenTelemetry: trace context in logs](https://opentelemetry.io/docs/specs/otel/compatibility/logging_trace_context/) - [Visual Studio Code: approvals and permissions](https://code.visualstudio.com/docs/agents/approvals) - [NIST: Secure Software Development Framework 1.1](https://csrc.nist.gov/pubs/sp/800/218/final) *This is evidence-led IZZY synthesis from current primary and first-party technical guidance checked 17 July 2026. The context-boundary patterns are operating options, not a universal standard. This article is general product and engineering guidance, not a security assessment, privacy review, legal opinion, compliance determination or forecast of outcomes.* --- ### Why did your website traffic drop? Diagnose the loss before you buy more marketing URL: https://izzy.agency/en/blog/website-traffic-drop/ Published: 2026-07-21 Summary: Confirm the loss, locate the break and separate measurement, visibility and demand causes before you buy more marketing to fix a traffic drop. ## Answer in 60 seconds A falling traffic graph does not yet tell you whether demand disappeared, search visibility changed, tracking broke or a release made pages harder to find. Before you buy more media, rewrite the site or publish another batch of articles: - confirm the loss in more than one evidence source; - locate the affected channel, page group, device, country and date; - compare it with releases, migrations, indexing problems and changes in demand; - fix the first evidenced break, then measure again. If visits are stable but enquiries fell, you have moved into a conversion problem. Keep that diagnosis separate. There is no public benchmark that can calculate the traffic or revenue your website “should” recover. That requires your own Search Console, analytics and commercial evidence. ## In this guide - [1. Confirm that traffic actually dropped](#1-confirm-that-traffic-actually-dropped) - [2. Locate where and when the loss began](#2-locate-where-and-when-the-loss-began) - [3. Separate measurement, visibility, demand and technical causes](#3-separate-measurement-visibility-demand-and-technical-causes) - [4. Fix the smallest evidenced break](#4-fix-the-smallest-evidenced-break) - [5. Choose monitor, repair, investigate, migrate or stop](#5-choose-monitor-repair-investigate-migrate-or-stop) ## 1. Confirm that traffic actually dropped “Traffic is down” may refer to users in analytics, clicks from Google Search, impressions, sessions, page views or qualified visits. Those are not interchangeable. Start with the business consequence and the date it appeared. Then compare: - Google Search clicks and impressions; - organic sessions in your analytics platform; - leads, calls or orders from the same period; - tag, consent and analytics-configuration changes; - releases, redirects and migrations. Google Analytics collection depends on the implementation and consent state. For example, Google states that Analytics does not store its client identifier when analytics storage is disabled through Consent Mode ([Google Analytics](https://support.google.com/analytics/answer/11593727?hl=en)). That does not mean consent should be weakened. It means a measurement change must not be mistaken for a demand collapse. If Search Console clicks are stable while analytics sessions fall, investigate collection before changing search pages. If both decline at the same time, continue into visibility and demand. ## 2. Locate where and when the loss began Site-wide totals hide the diagnosis. Google recommends using the Search Console Performance report to inspect pages, queries, countries, devices and search types. It also recommends long enough comparisons to reveal seasonality and the shape of the decline ([Google Search Central](https://developers.google.com/search/docs/monitor-debug/debugging-search-traffic-drops)). Compare like with like. Weekly or monthly aggregation can reduce day-of-week distortion ([Search Console comparison guidance](https://support.google.com/webmasters/answer/17011165?hl=en)). Record the exact filters and property used. Search Console is not a complete ledger. Google omits anonymised queries and states that tables contain only the most important rows because of data truncation ([Search Console data groupings](https://support.google.com/webmasters/answer/17011259?hl=en)). A filtered query table may therefore disagree with its chart total. The objective is not to produce one perfect number. It is to find the first coherent pattern: one directory, one device, one market, one search type or the whole property. ## 3. Separate measurement, visibility, demand and technical causes Use the pattern to choose the next branch: | Observed pattern | Plausible branch | First check | |---|---|---| | Search Console stable, analytics down | tracking, consent or property change | tag and consent release history | | Impressions and clicks both down | demand, indexing, ranking or security | affected queries/pages and indexing reports | | Impressions stable, clicks down | result presentation, intent or search-page change | query-level CTR, title and search appearance | | One template or directory falls | release, canonical, noindex, rendering or internal-link issue | compare affected URLs with recent deployments | | Loss begins after URL or domain changes | migration continuity | redirects, canonicals, sitemap and old-to-new parity | | Visits remain stable but leads fall | conversion or handoff | move to the conversion diagnosis | Google's own diagnostic guidance lists technical, security, spam, algorithmic, migration, seasonality and changing-interest causes. Its crawling guidance also shows that robots directives, canonicals, status codes and JavaScript rendering can affect discovery and indexing ([Google crawling and indexing](https://developers.google.com/search/docs/crawling-indexing)). Stable rankings do not guarantee stable clicks. In a 2025 Pew Research Center study of 68,879 Google searches from 900 consenting US adults, traditional-result clicks occurred on 8% of visits with an AI summary and 15% without one. The study covers one country, one platform and a tracked panel; it cannot estimate the effect on your site ([Pew Research Center](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/)). Treat AI-result changes as one hypothesis, not a universal explanation. ## 4. Fix the smallest evidenced break Match the repair to the evidence: 1. **Measurement:** restore a documented, consent-respecting implementation and annotate the change. 2. **Indexing:** remove accidental blockers, correct status or canonical errors and verify representative URLs. 3. **Migration:** reconcile old and new URL inventories, redirects, canonicals, sitemaps and critical-page parity. 4. **Search fit:** improve a page only when the affected query and page no longer answer the same decision. 5. **Demand:** if the whole market moved, change the forecast or channel mix rather than “fixing SEO”. Do not change titles, templates, internal links and content simultaneously. You will replace one unexplained graph with another. Set a baseline, release a bounded change, record the date and compare the same segments again. Google explicitly notes that changes do not guarantee a visible search impact. ## 5. Choose monitor, repair, investigate, migrate or stop | Decision | Use it when | Guardrail | |---|---|---| | Monitor | the movement is small, seasonal or not yet commercially material | preserve the baseline and review date | | Repair | one measured implementation or indexing break is isolated | change only the affected layer | | Investigate | Search Console, analytics and business outcomes disagree | keep competing hypotheses visible | | Migrate or rescue | the loss follows a material platform or URL change | require parity, rollback and recovery evidence | | Stop | no repeatable loss or business consequence is established | do not manufacture activity | A useful traffic review ends with an affected-page inventory, evidence timeline, cause register and prioritised recovery backlog - not a generic list of SEO tasks. ## Conclusion: the next sensible move Prove the loss before you prescribe the recovery. Confirm the measurement, locate the affected segment and compare the date with changes in demand, search results and your own releases. Repair the first material break, then measure the same evidence again. ## Bring us the dates. We'll help you find the break. Bring the affected period, Search Console comparison, analytics or consent changes, recent releases and any migration record. In a 30-minute call, IZZY can scope whether the first step is measurement repair, traffic-loss diagnosis, migration rescue or no engagement yet. We will not promise restored traffic, rankings, leads or revenue from a call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Why did my website traffic suddenly drop? Common branches include broken measurement, technical or indexing changes, a migration, ranking changes, seasonality, reduced demand, security issues or a changed search-results page. The date and affected segment should determine which branch you test first. ### How can I tell whether tracking broke or SEO traffic fell? Compare Search Console clicks with analytics sessions for the same dates and landing pages. If Search clicks remain stable while recorded sessions fall, inspect tags, consent and analytics configuration before changing search content. ### Can rankings stay stable while traffic falls? Yes. Impressions, click-through behaviour, result features, demand and query mix can change while a tracked position appears stable. Search Console page and query evidence is more useful than a small rank-tracker list alone. ### How long does traffic recovery take? There is no universal recovery period. It depends on the cause, crawl and indexing behaviour, demand and the scope of the repair. Define the next review window before release and avoid promising a date without client evidence. ### What should we bring to a traffic-loss review? Bring Search Console exports, analytics definitions, the affected dates, release and CMS history, redirect or migration records, key landing pages, lead or order evidence and the commercial definition of a qualified visit. *Research and source status checked 16 July 2026. This article provides general traffic-diagnosis guidance. It is not a ranking forecast, analytics implementation approval or client-specific recovery plan.* --- ### Managing AI agent permissions URL: https://izzy.agency/en/blog/ai-agent-permissions-access-control/ Published: 2026-07-20 Summary: Learn how to inventory, scope, review and revoke AI agent permissions with an access register built for real business systems. An AI integration can reach a business system. The connection alone does not establish its continuing purpose, owner, allowed actions or end condition. **An AI-agent permission is the complete operating decision that links an identity and authority source to a resource, action, data class, duration, approval rule, evidence record and revocation path.** This is an IZZY working definition, not an external standard. The practical question is how to manage AI agent permissions as a lifecycle, not as a list of app connections. ## Answer in 60 seconds - Inventory active and dormant AI integrations, automations and agent workflows. - Assign a business owner and a technical owner to each one. - Record whether it uses a named user's authority, a workload identity, temporary access or a shared credential. - Scope the grant by resource, action and duration. - Attach approval to the exact consequential action and its parameters. - Record attempts, decisions and results without recording credentials or secrets. - Set review triggers, expiry conditions, a stop procedure and a tested revocation path. Keep those decisions in an **agent access register**: a maintained IZZY operating record, not a technical agent directory, compliance certificate or legal register. ## In this article 1. [What counts as an AI agent permission?](#1-what-counts-as-an-ai-agent-permission) 2. [Where does permission debt appear?](#2-where-does-permission-debt-appear) 3. [What belongs in an agent access register?](#3-what-belongs-in-an-agent-access-register) 4. [Which identity and access model should you use?](#4-which-identity-and-access-model-should-you-use) 5. [How should approvals, logs and delegation work?](#5-how-should-approvals-logs-and-delegation-work) 6. [How do you review and revoke access?](#6-how-do-you-review-and-revoke-access) ## 1. What counts as an AI agent permission? For operating purposes, treat each permission as nine connected decisions: **identity + authority source + resource + action + data class + duration + approval + logging + revocation** - **Identity:** the user, application or workload presenting the request. - **Authority source:** the person or business mandate behind it. - **Resource:** the bounded tenant, mailbox, folder, repository, project, table or records. - **Action:** search, read, draft, create, update, send, publish, delete, approve or administer. - **Data class:** personal, confidential, client, financial or other restricted information. - **Duration:** one request, task, pilot or continuing access with a review trigger. - **Approval:** the exact action needing another decision. - **Logging:** the attempt, decision and result available for review. - **Revocation:** the identity, grant, session, job or delegated route that must stop. This broader definition matters because a consent screen or role assignment records only part of the operating decision. Agent identity and authorisation are developing areas. A [2026 NIST NCCoE concept paper](https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd) explores identification, authorisation and auditing for software and AI agents. It is an initial public draft, not a final standard. NIST's [AI Agent Standards Initiative](https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative) describes related work on authentication and identity infrastructure. Do not wait for one universal model. Make permissions visible, apply established practices and document what the architecture cannot express. ## 2. Where does permission debt appear? Permission debt is the gap between access that exists and access the organisation can still explain, review and end. The following five patterns are an IZZY diagnostic model, not measured prevalence. | Pattern | Operating problem | Decision to reopen | |---|---|---| | No accountable owner | An administrator can maintain the connection, but nobody owns its continuing business purpose. | Name who can renew, change or retire the access. | | Ambiguous authority | A human session or shared credential hides which workflow acted and under whose mandate. | Separate the identity or record the authority path. | | Accumulating scope | New resources and actions are added while older grants remain. | Compare the current grant with the current job. | | No end condition | Access has no expiry, renewal decision or material-change trigger. | Define when the decision must be reopened. | | Hidden delegation | One service invokes another without preserving the initiating authority and intermediate route. | Make the delegation path reviewable or redesign it. | ## 3. What belongs in an agent access register? An agent access register gives product, operations, platform and security stakeholders one maintained view of the permission lifecycle. It should record: - integration or agent name, purpose and owners; - identity, authority source and credential type; - connected systems, resources, actions and data classes; - approval rule, duration, expiry and review trigger; - attempted and completed action records; - delegation path; and - revocation and stop procedure. These IZZY-defined fields can live in an existing service catalogue, identity system, risk register or operating document. A dedicated tool is not the first requirement. Consider a support-triage workflow that classifies cases and prepares draft replies. An application identity reads one queue and selected knowledge articles. A support operator approves sending. Changes to the queue, data class, provider or available actions trigger review. Revocation disables the grant and checks pending jobs before manual operation resumes. `Access to support` would not tell a reviewer which queue, actions, authority or revocation behaviour applies. Microsoft Entra Agent ID provides one product example separating [technical owners from business sponsors](https://learn.microsoft.com/en-us/entra/agent-id/agent-owners-sponsors-managers). It is not a universal model, but the distinction makes administration and continuing business accountability explicit. ## 4. Which identity and access model should you use? AI agent identity management begins with choosing whose authority the workflow uses. - **User-delegated authority** preserves a named user's context and should remain within the approved purpose. [Microsoft documents delegated and app-only sign-in activity](https://learn.microsoft.com/en-us/entra/agent-id/sign-in-audit-logs-agents) in one product. - **Application or workload authority** provides a separate technical identity. It can improve visibility and independent revocation, but does not prove human intent. - **Temporary task authority** exists for a bounded operation. Its lifetime depends on the task, system and consequence. - **Shared authority** covers several workflows with one credential, limiting distinction and independent revocation. Least privilege for AI agents requires explicit resource, action and duration decisions: | Dimension | Decision the owner must record | |---|---| | Resource | Which tenant, mailbox, folder, project, repository, table or record set? | | Action | Read, search, draft, create, update, send, publish, delete, approve or administer? | | Duration | One request, session, task, pilot or standing access with a named review trigger? | | Authority | Which user, application, workload or temporary identity? | | Consequence | Which exact action requires independent approval or another execution path? | | Evidence | Which attempt, permission decision and result are recorded without secrets? | This table is an operating framework, not a universal access-control standard. [RFC 9700](https://www.rfc-editor.org/info/rfc9700) recommends restricting OAuth token privileges and binding them to intended resources and actions. For MCP implementations, the [authorisation specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) supports requesting additional scopes when needed. Neither protocol represents the whole permission lifecycle. The UK NCSC recommends [minimum access for the shortest necessary period, temporary credentials where possible and revocation when work ends](https://www.ncsc.gov.uk/blogs/thinking-carefully-before-adopting-agentic-ai). This does not create one expiry period for every integration. ## 5. How should approvals, logs and delegation work? Installation and action approval are different decisions. For sending, publishing, deletion, record changes or access changes, bind approval to the exact target and displayed parameters. `CRM access` cannot replace a decision about a particular update. [OWASP's AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) recommends limiting tools and permissions and explicitly authorising sensitive operations. Enforcement belongs in the application or policy layer, not the model's judgement. Logs should help a reviewer answer three questions: 1. What was attempted - the tool, target, action and relevant parameters? 2. What was decided - the acting identity, authority source, applicable rule and approval result? 3. What happened - completed, denied, failed, reversed or stopped? Do not record authorisation headers, tokens, credentials or secrets. MCP's [authorisation guidance](https://modelcontextprotocol.io/docs/tutorials/security/authorization) states this for MCP implementations. When one integration calls another service, record the initiating authority, intermediate client, target and outcome where supported. MCP's [security guidance](https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices) explains that token passthrough can obscure client identity and downstream audit trails. Other delegated architectures need the same operating answer, though their controls differ. Logs support review; they do not prove intent or complete attribution. ## 6. How do you review and revoke access? Start with one bounded portfolio: a team, product, process or group of integrations that reach the same systems. 1. **Inventory.** Compare app dashboards, identity records, automation platforms, secret stores and team knowledge. Record active, dormant and unresolved routes. 2. **Classify.** Name the owners, purpose, authority, resources, actions, data classes, duration and highest-consequence action. 3. **Narrow.** Remove unused systems, actions and broad scopes where more precise grants are supported. Separate workflows needing different owners or revocation paths. 4. **Add review triggers.** Reopen the decision when purpose, owner, system, data, action, provider, credential or consequence changes. 5. **Test stop and revocation.** Disable the identity or grant; inspect sessions, tokens, delegated services, queued work and scheduled jobs. Confirm the fallback. 6. **Assign operation.** Name who reviews renewal, monitors exceptions and can suspend the workflow. The result may be a narrower grant, separate identity, architectural change or retirement. The register matters only when it changes a grant, renewal or revocation decision. Pause before expansion when a workflow involves personal or regulated data, crosses organisational boundaries, enables high-consequence actions, relies on an inseparable privileged identity, hides delegation or fails a stop test. For personal data, the European Commission describes [purpose limitation, data minimisation and storage limitation](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/principles-gdpr_en). Applying them to legal basis, roles, notices, transfers or impact assessment requires qualified review. IZZY can map product, identity, cloud and operating choices. Legal, privacy, security and sector specialists should own conclusions in their domains. ## Conclusion: treat access as a lifecycle A connection is a technical fact. A permission is a continuing business and operating decision. Keep the purpose, owner, authority, resource, action and duration visible. Connect approval to consequences, preserve the delegation path and test how access stops. When the workflow changes, reopen the grant instead of inheriting yesterday's decision. ## Map one AI-access portfolio Bring one bounded list of AI tools, automations and agent workflows, the systems they reach, known owners and any available scope, log, expiry or revocation records. In a 30-minute scoping call, IZZY can map the first agent access register, identify unresolved decisions and compare implementation paths across product, cloud and identity. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### 1. Is an OAuth consent screen enough to manage AI agent permissions? No. It records part of an authorisation flow. The operating decision also needs purpose, ownership, resource, action, duration, delegation, evidence and revocation. ### 2. Does every AI integration need its own identity? No universal rule fits every system. Separate workload identity can improve visibility and independent revocation; delegated access can preserve user context. Avoid shared authority the team cannot operate. ### 3. How often should AI agent permissions be reviewed? Use a suitable scheduled review plus event-based triggers. Reopen access when its purpose, owner, resource, data, action, provider, credential or consequence changes. ### 4. Is an agent access register the same as an AI agent permission audit? No. The register is an operating record. A formal security, legal or compliance audit has a separate scope, method, evidence standard and qualified owner. ### 5. What should happen to permissions when an AI pilot ends? Revoke the identity, token, role or consent; end supported sessions; stop delegated and queued work; address retained data; and record retirement or a newly approved scope. ## Sources and evidence note - [NIST NCCoE: Software and AI Agent Identity and Authorization concept paper](https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd) - [NIST: AI Agent Standards Initiative](https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative) - [IETF: RFC 9700, OAuth 2.0 Security Best Current Practice](https://www.rfc-editor.org/info/rfc9700) - [Model Context Protocol: understanding authorisation](https://modelcontextprotocol.io/docs/tutorials/security/authorization) - [Model Context Protocol: authorisation specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization) - [Model Context Protocol: security best practices](https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices) - [OWASP: AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) - [Microsoft Learn: owners, sponsors and managers in Entra Agent ID](https://learn.microsoft.com/en-us/entra/agent-id/agent-owners-sponsors-managers) - [Microsoft Learn: Entra Agent ID logs](https://learn.microsoft.com/en-us/entra/agent-id/sign-in-audit-logs-agents) - [UK NCSC: thinking carefully before adopting agentic AI](https://www.ncsc.gov.uk/blogs/thinking-carefully-before-adopting-agentic-ai) - [European Commission: data-protection principles for organisations](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/principles-gdpr_en) *How this was prepared: evidence-led IZZY synthesis from current primary, official and first-party guidance. Sources checked 17 July 2026. The agent access register, permission definition, permission-debt model and review path are IZZY operating frameworks, not external standards. This article is general product and operational guidance, not a security assessment, legal opinion, privacy review, compliance determination or forecast of outcomes.* --- ### How to track AI traffic in GA4: setup, validation and limits URL: https://izzy.agency/en/blog/track-ai-traffic-ga4/ Published: 2026-07-20 Summary: Track ChatGPT, Perplexity and other AI referrals in GA4 with a custom channel group - and report clearly what your analytics still cannot see. ## The 60-second answer To track AI traffic in GA4, start with the new **AI Assistant** channel. When GA4 recognises an AI-assistant referrer, it assigns the medium `ai-assistant` and places the visit in that channel. Google announced this change on 13 May 2026 ([Google Analytics release notes](https://support.google.com/analytics/answer/9164320?hl=en)). That is useful, but the channel is not a complete measure of AI influence. Some identifiable visits may still carry a different medium. Visits without usable source information can appear as Direct. Google’s AI Overviews and AI Mode remain within Organic Search in GA4. Use this measurement process: 1. inspect **Session source / medium**, not only the channel total; 2. create a custom channel based on verified source values; 3. validate the rules against Referral, Unassigned and Organic Search; 4. measure engagement, key events and qualified leads; 5. report the result as **identifiable AI referral traffic**, not total AI traffic. A channel group improves classification. It cannot recover data that was never collected or measure a recommendation that produced no website visit. ## 1. Decide what “AI traffic” means before opening GA4 “AI traffic” can describe several different things: | Question | Evidence source | What it can show | |---|---|---| | Did an identifiable AI assistant send a website visit? | GA4 session source and medium | Referred sessions from a recognised source | | Did our pages appear in Google’s generative search features? | Search Console generative-AI report, where available | Impressions in AI Overviews and AI Mode | | Did an assistant cite or recommend the brand without a click? | Repeatable AI-visibility monitoring | Mentions, citations and answer position | | Did AI influence a buyer who returned later? | CRM, attribution and customer research | Evidence beyond the final website session | | Did a visit arrive without source information? | GA4 Direct traffic | The visit, but not a reliable original source | These measurements are related, not interchangeable. Generative engine optimisation, or GEO, concerns whether AI systems can find, understand and cite a company. Referral traffic covers only users who click. ## 2. Audit your existing ChatGPT and AI referral traffic Before building a GA4 custom channel group, inspect what the property already records. Open: **Reports → Acquisition → Traffic acquisition** Use **Session source / medium** as the primary or secondary dimension. The Traffic acquisition report uses session-scoped dimensions, which describe how a particular session began. User acquisition uses first-user dimensions and answers a different question: how the person was initially acquired ([Google Analytics: traffic-source scopes](https://support.google.com/analytics/answer/11080067?hl=en)). Search the source rows for assistant names and domains already present in your data. Review at least: - the default AI Assistant channel; - Referral; - Unassigned; - Organic Search; - Direct, as a limitation rather than a source you can safely reclassify. Google’s current default definition places a visit in AI Assistant when its medium is exactly `ai-assistant`. The current description names sources such as ChatGPT, Gemini, DeepSeek, Copilot and Grok, while explicitly excluding Google AI Overviews and AI Mode ([Google’s default channel definitions](https://support.google.com/analytics/answer/9756891?hl=en)). Do not assume every row from one assistant has the same source-and-medium combination. Source and medium are separate fields, so inspect the combinations in your property rather than relying on the channel name. Create a small audit table: | Session source | Session medium | Current channel | Sessions | Intended treatment | |---|---|---|---:|---| | `chatgpt.com` | `ai-assistant` | AI Assistant | Property data | Include | | `chatgpt.com` | `referral` | Referral | Property data | Review and include if genuine | | Known AI domain | `(not set)` | Unassigned | Property data | Review source and tagging | | `(direct)` | `(none)` | Direct | Property data | Do not claim as AI | The rows in your account are the source of truth for the rule you build. ## 3. Build a GA4 custom channel group for identifiable AI referrals Google provides an AI-assistant example in its custom-channel documentation. Rules can use Source, Medium and Source platform, and the first matching channel wins ([Google Analytics: custom channel groups](https://support.google.com/analytics/answer/13051316?hl=en)). To create the group: 1. Open **Admin → Data display → Channel groups**. 2. Select **Create new channel group**. 3. Add a channel named **Identifiable AI referrals**. 4. Set the condition to **Source matches regex**. 5. Enter a pattern based on source values you have verified. 6. Move the new channel above Referral. 7. Save the group and test it as a secondary dimension. A cautious starter pattern could look like this: ```text ^(www\.)?(chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com|deepseek\.com|grok\.com|you\.com)$ ``` This is a template, not a universal list. Use only verified source values. Avoid broad fragments such as `ai`, `gpt` or `google`, which can match unrelated traffic. Placing the custom channel above Referral allows it to claim matching AI referrals first. Do not move it above Organic Search until you have tested that the rule cannot absorb ordinary search traffic. ## 4. Understand the difference between retroactive reports and the primary channel group This GA4 detail is easy to miss. Google states that custom channel groups can be applied retroactively to reports. However, if you make one the property’s **primary channel group**, that primary definition populates reporting from that point forward. The safer sequence is: 1. keep the Google default group unchanged as a reference; 2. use the custom group as a secondary dimension; 3. compare both definitions across the same date range; 4. document the rule and approval date; 5. change the primary group only if the organisation needs one shared custom definition. Custom channel groups are not available in the GA4 BigQuery export schema. Teams using BigQuery must reproduce the rules in their queries. ## 5. Validate the custom AI channel before reporting it Run five checks before reporting: ### Coverage Confirm that every verified AI source row reaches the custom channel. ### False positives Check that ordinary search, referral and campaign traffic has not been pulled in. ### Priority Confirm that the AI channel sits above Referral. GA4 assigns traffic to the first matching channel in the ordered group. ### Missing information Review Unassigned and `(not set)` rows. Unassigned appears when no channel rule matches, while Direct represents traffic without a clear referral source ([Google Analytics: Direct traffic](https://support.google.com/analytics/answer/15258820?hl=en)). If the source itself is absent, a regex cannot recover it. Do not estimate an “AI share of Direct” and present it as observed traffic. ### Change control Save the domain list, rule, test result, owner and review date. A channel without an owner will become stale. ## 6. Report business value, not only AI session volume Once classification is stable, report buyer quality: - identifiable AI referral sessions; - engaged sessions and engagement rate; - key events, such as a qualified form or booked call; - session key-event rate; - landing pages receiving AI referrals; - source-level performance; - qualified opportunities and revenue where CRM matching is possible. Compare landing pages and business outcomes, not only channel volume. A few service-page visits may matter more than many low-intent article visits. Use consistent labels in reports: - **Default AI Assistant** for Google’s maintained definition; - **Identifiable AI referrals** for your verified source-based group; - **AI visibility** for mentions and citations measured outside GA4; - **AI-influenced pipeline** only when CRM or customer evidence supports the claim. This naming prevents a classification improvement from being presented as new demand. ## 7. What even a well-built GA4 AI channel cannot measure These are AI traffic attribution limits, not problems a larger regex can solve. ### AI visits with no identifiable source Google classifies `(direct) / (none)` when traffic has no clear referral source. Some of those visitors may have used an assistant, but GA4 does not provide evidence to identify which ones. ### Google AI Overviews and AI Mode as a separate GA4 channel Google includes these experiences in Organic Search rather than AI Assistant. Do not create a source rule that captures all `google / organic`; it would combine ordinary search and generative-search visits. Search Console now has a separate generative-AI report for impressions in AI Overviews and AI Mode. It is rolling out to a subset of properties and adds visibility evidence, not GA4 referral attribution ([Google Search Console documentation](https://support.google.com/webmasters/answer/16984139?hl=en)). ### Recommendations without clicks An assistant may mention the brand, answer the question and end the journey without sending a visit. Website analytics cannot measure that exposure. ### Influence before a later return A buyer may discover the company in an assistant, return through branded search and convert later. Use CRM source questions, sales notes and attribution analysis to investigate. ## 8. Use a four-layer AI measurement model | Layer | Business question | Recommended evidence | |---|---|---| | Identifiable referral | Which assistants sent visits? | GA4 custom channel and Session source / medium | | Google generative visibility | Where did pages appear in AI Overviews or AI Mode? | Search Console generative-AI report, where available | | Citation visibility | Which brands and sources appear in assistant answers? | Controlled prompt and citation monitoring | | Commercial influence | Did AI contribute to qualified demand or revenue? | GA4 key events, CRM, sales notes and customer research | Review the layers together, but do not merge them into one unsupported score. The system should separate observed visits, observed visibility and inferred commercial influence. ## Conclusion: improve classification without overstating attribution The GA4 AI Assistant channel is a useful default. A verified custom channel can provide a more consistent view of AI referrals across historical reports. Neither one measures total AI influence. Start with Session source / medium, build the rule from observed sources, test for false positives and report the result as identifiable AI referral traffic. Then combine it with Search Console, citation monitoring and CRM evidence. The strongest report is not the one with the largest AI number. It is the one whose definition, limits and business consequences can be explained. ## Can you connect AI visibility to qualified demand? Bring us your GA4 acquisition export, key-event definitions and lead journey. IZZY can audit classification, build a defensible measurement model and connect referral traffic with visibility and CRM evidence - without presenting one channel as total AI performance. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Does GA4 track AI traffic automatically now? GA4 classifies recognised referrals into AI Assistant. It cannot identify visits without source information, and Google AI Overviews and AI Mode remain in Organic Search. ### How do I find ChatGPT traffic in GA4? Open Traffic acquisition and use Session source / medium. Search for ChatGPT source values across AI Assistant, Referral and Unassigned. A different or missing medium may not satisfy the default `ai-assistant` rule. ### Are GA4 custom channel groups retroactive? They can be applied retroactively in reports. Making one the primary channel group affects the primary definition going forward, so test it as a secondary dimension first. ### Can GA4 identify clicks from Google AI Overviews and AI Mode? Not as a separate AI Assistant channel. Google classifies those visits as Organic Search. Search Console’s generative-AI report can provide impression data for eligible properties. ### Should I copy a large AI-referral regex from another website? Use it as a starting point. Build the production rule from sources observed in your property, test false positives and assign an owner. ### Is AI referral traffic the same as GEO performance? No. Referral traffic measures identifiable clicks. GEO performance also includes whether the brand is understood, mentioned, compared and cited - even when no visitor reaches the website. *Primary platform documentation checked on 16 July 2026. GA4 and Search Console features, channel definitions and referral behaviour can change. This article provides measurement guidance, not a guarantee of complete attribution.* --- ### GEO is not just an SEO problem. It is a company problem URL: https://izzy.agency/en/blog/geo-ai-visibility-company-ownership/ Published: 2026-07-19 Summary: SEO, brand, product or PR: learn how to assign responsibility for AI visibility without creating another organisational silo. **What does GEO mean?** GEO stands for *Generative Engine Optimisation*. In plain English, it is the work that helps tools such as ChatGPT, Gemini and AI-powered search engines find, understand and potentially use reliable information about your company in their answers. It cannot guarantee that an assistant will recommend you; it improves the clarity and availability of the signals those answers may use. ## The 60-second answer Making the SEO team solely responsible for AI visibility creates an ownership problem. Technical SEO remains essential: systems are less likely to find and understand pages that are inaccessible, poorly structured or unclear. But a generated recommendation also draws on what you sell, how clearly it is positioned, what evidence exists and what credible third parties say. Those inputs belong to product, brand, PR, content and sometimes legal teams. The answer is not a separate GEO silo. Name one owner for the outcome, bring together the functions that produce its signals and use SEO as the technical and measurement foundation. A cross-functional result without an owner quickly becomes nobody’s priority. ## In this article 1. [Why SEO cannot own the result alone](#1-why-seo-cannot-own-the-result-alone) 2. [The four systems that shape AI visibility](#2-the-four-systems-that-shape-ai-visibility) 3. [A workable governance model](#3-a-workable-governance-model) 4. [The first 90 days](#4-the-first-90-days) ## 1. Why SEO cannot own the result alone “Who owns GEO?” often receives a default answer: the SEO team. It already understands search, crawlers, structured data and content performance, so it should certainly be central. The problem begins when that team becomes accountable for inputs it cannot control: - an offer that is too broad or difficult to categorise; - customer evidence that is unavailable or not approved; - weak external reputation; - inconsistent product information; - subject experts who do not contribute to content; - legal policies that prevent precise answers. A technical audit can show that important pages are retrievable. It cannot invent a proposition, an authoritative study or a reputation. SEO is therefore neither dead nor secondary. It is necessary but insufficient. That distinction matters for budgets: strengthening technical delivery when product evidence is missing produces more work, not more trust. ## 2. The four systems that shape AI visibility ### 1. Technical foundations Important pages should be accessible, indexable, stable and properly connected. Structured data can clarify certain entities. Server logs and measurement show what systems retrieve. This is the natural territory of SEO and engineering. ### 2. Product clarity An assistant needs to understand what you offer, who it serves, where it is available and where its limits sit. If commercial pages, documentation and sales teams give different answers, the problem exists before GEO begins. ### 3. Editorial evidence Content should answer buying questions with facts, method and honest boundaries. Marketing can structure publication, but product, sector and compliance specialists must supply the substance that makes an answer defensible. ### 4. External reputation In its study of more than 25 million links in ChatGPT, Claude and Gemini responses, Muck Rack categorised **84% of citations** as earned media ([Muck Rack](https://muckrack.com/blog/what-is-ai-reading-may-2026)). The figure belongs to that study’s corpus, but the principle is useful: your website is not the sole narrator of your company. Trade press, partners, professional bodies, industry publications and relevant communities all take part. Earning accurate coverage is not a technical SEO task, although keeping it precise requires product and editorial coordination. ## 3. A workable governance model Begin by appointing an **outcome owner**, not an owner of every tactic. Depending on the company, this may sit in brand, product, digital marketing or growth. The title matters less than the authority to convene teams and make trade-offs. | Responsibility | Likely lead | Expected contribution | |---|---|---| | Visibility measurement | SEO / data | Stable prompts, captures, sources and limitations | | Technical access | SEO / engineering | Indexability, structure, retrieval and fixes | | Offer clarity | Product / marketing | Category, audience, use cases and boundaries | | Evidence and answers | Content / experts | Data, method, approval and updates | | External presence | Brand / PR | Relevant sources, relationships and accurate mentions | | Risk | Legal / compliance | Review of sensitive claims and usage rules | A smaller business does not need six departments. One person may hold several roles. The important point is to make decisions explicit: who observes, who corrects, who approves and who decides that the evidence is sufficient? Avoid two familiar mistakes. The first is launching an “AI task force” without stable measurement. The second is expecting SEO to create evidence and third-party recognition that other functions must produce. ## 4. The first 90 days ### Days 1–30: establish the baseline Select important buying questions, record the brands and sources present, then review the pages associated with those decisions. Separate mentions, recommendations and citations. A single run remains directional; retain the protocol so it can be repeated. ### Days 31–60: correct the first disagreement Find the first place the company contradicts itself. Is the category unclear? Is an important answer missing? Does evidence exist only in sales decks? Does an external source repeat an obsolete offer? Correct a small number of traceable problems. Do not commission fifty generic articles to compensate for a proposition nobody can explain. ### Days 61–90: publish, distribute and measure again Turn internal expertise into answers that can be cited. Place them in front of sources your buyers genuinely use. Repeat the baseline with the same question set and document what changed - or did not. No movement is also information. The source may not have been retrieved, the evidence may remain weak or the tactic may not influence that model. The outcome owner prevents each team from inventing its own explanation. ## Conclusion: GEO reveals the organisation you already have AI visibility does not create the gaps between technology, product, content and reputation. It makes them visible. Keep SEO at the centre of technical foundations and measurement. Give the outcome to someone who can make the other functions work together. Then run limited corrections with before-and-after evidence. Coordination, not another box on the organisation chart. ## Who owns this outcome in your company? If the answer is unclear, bring us the real organisation, buying questions and observed results. We can frame a cross-functional workshop, define the protocol and assign the first decisions. The purpose is not to sell a new label. It is to make the work manageable. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Does GEO replace SEO? No. GEO expands the observed outcome to generated answers, recommendations and citations. SEO remains essential to make information accessible and understandable. ### Who should own AI visibility? The person able to make trade-offs between technical, product, editorial and reputation work. The role may sit in brand, product or digital marketing, but it needs a clear mandate and shared measurement. ### Do we need a GEO team? Not necessarily. Start with one owner, a small working group and a repeatable protocol. Create a dedicated team only when the volume and value of decisions justify it. ### How should each team’s work be measured? Connect each correction to an observation: retrieved pages, offer consistency, published evidence, external mentions and movement across a stable question set. Do not promise that one tactic will produce a citation. *Sources checked on 14 July 2026. This article proposes an operating model and does not guarantee a change in generated answers.* --- ### AI visibility is becoming paid. What should you build before the window narrows? URL: https://izzy.agency/en/blog/ai-visibility-citations-earned-media/ Published: 2026-07-18 Summary: Paid placements are growing in AI search. Learn how to build and measure earned AI visibility before increasing your media budget. ## The 60-second answer You do not need to believe that AI will replace search to act. The decision is simpler: buyers already use assistants to compare options, understand categories and prepare shortlists, while platforms are beginning to monetise those environments. The right response is neither to wait nor to chase every model. Build three assets now: 1. useful answers to real buying questions; 2. credible presence in the external sources assistants cite; 3. repeatable measurement of your visibility and your competitors’ visibility. The objective is not to “hack ChatGPT”. It is to understand where your brand already contributes to a buyer’s decision - and where it is absent. ## In this article 1. [Why the organic window may narrow](#1-why-the-organic-window-may-narrow) 2. [What citation data actually tells us](#2-what-citation-data-actually-tells-us) 3. [What to build now](#3-what-to-build-now) 4. [How to measure without a black-box score](#4-how-to-measure-without-a-black-box-score) ## 1. Why the organic window may narrow Organic visibility has unusual value while a new discovery surface is taking shape. Rules are unsettled, competitors have not all structured their content and advertising does not occupy every useful position. That situation rarely lasts unchanged. An eMarketer forecast reported by Axios projects that US AI search advertising spend could reach **$25.9 billion in 2029**, or 13.6% of forecast search advertising spend that year ([Axios](https://www.axios.com/newsletters/axios-ai-plus-eafb70c7-71fb-4dab-ba52-5006df10529e)). It is a forecast, not an established outcome, but it indicates the commercial direction. “The free window is closing” is therefore a strategic interpretation, not a known deadline. Organic answers are not due to disappear tomorrow. Paid placements, data partnerships and platform rules can nevertheless make attention more expensive and competition more organised. The cost of waiting is not merely a larger media bill. It is discovering late that your company is absent from the sources assistants use to describe the market. ## 2. What citation data actually tells us Muck Rack analysed more than **25 million links** in responses from ChatGPT, Claude and Gemini across 17 industries. In its study, earned media accounted for **84% of citations**, while paid or advertorial material accounted for **0.3%** ([Muck Rack](https://muckrack.com/blog/what-is-ai-reading-may-2026)). Two qualifications matter. First, this describes a particular corpus, period and method. It does not mean that 84% of every AI answer comes from journalism or that identical sources dominate every prompt. Second, it does not prove that a press release or mention will cause a citation. Earned media is a broader external signal: trade publications, independent analysis, professional bodies, relevant comparison sites and trusted communities. Accuracy and relevance to the question remain essential. The useful conclusion is less dramatic: your website does not control the entire account of your brand. What credible third parties say about you forms part of your visibility infrastructure. ## 3. What to build now ### Answer questions that precede a purchase Start where buyers compare options, assess risk or form a shortlist. “Which provider…?”, “What is the difference between…?”, “How do we avoid…?” and “What should we budget for…?” are more commercially useful than topics chosen only for traffic. Each page should give a direct answer, decision criteria, limitations and a next step. A specific, supportable statement can be cited. A page of interchangeable promises gives a system little reason to choose it. ### Create evidence other people can use Publish what you genuinely know: a method, testing protocol, documented comparison, verifiable proprietary dataset or authorised case study. Give journalists, partners and communities a reason to mention the business beyond “we are excellent”. For an English-speaking B2B market, the useful mix may include trade press, professional bodies, industry newsletters, review or comparison platforms and specialist communities. Choose outlets for relevance to the buying decision, not the volume of links they can provide. ### Correct disagreement between brand and product If your website, profiles, partners and product pages describe the offer differently, an assistant encounters several versions of the company. Clarify the name, category, served markets, evidence and boundaries of the offer. AI visibility is not merely an editorial programme. It exposes disagreements that already existed between product, marketing, sales and communications. ## 4. How to measure without a black-box score Generated answers vary. One test is an observation, not a trend. Use a simple protocol: 1. select 10 to 20 stable buying questions; 2. record language, market, model and date; 3. capture named brands, your position and cited sources; 4. repeat the measurement using the same method; 5. retain screenshots or exports so the result can be reviewed. Do not combine three different outcomes: **being mentioned**, **being recommended** and **being cited as a source**. A brand can appear without a link, be cited without being recommended or be recommended using incorrect information. Then find the first explainable gap. Do your pages fail to answer the question? Do cited sources not know the brand? Is the proposition ambiguous? Useful measurement should lead to a decision, not merely produce a score. ## Conclusion: build the asset before buying the placement Advertising is likely to occupy more space inside AI interfaces. That does not make organic work irrelevant; it increases the value of visibility that is earned, credible and measurable. Start with buying questions, evidence and external sources. Measure them through a stable protocol. Then decide where media spend belongs. Work first, budget second. ## On which buying decisions does AI cite your competitors and not you? Bring us your market, principal competitors and the questions that precede a purchase. We can establish a traceable baseline, separate observations from assumptions and prioritise corrections. No magic score: a method your team can inspect. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Must a business pay to be cited by an AI assistant? Not necessarily. Muck Rack observed very few citations from paid material in its corpus. That does not guarantee an organic citation or predict every future commercial format. ### How long does AI visibility take to improve? There is no dependable timeline for every brand. Retrieval frequency, source strength, the prompt and the model all affect the result. Measure across several dates before drawing a conclusion. ### Does GEO replace SEO? No. An accessible, indexable and clear website remains a foundation. GEO adds observation of generated answers, cited sources and external reputation. ### Which prompts should we track? Track the questions buyers ask before selection: category, comparison, risk, cost, compliance, alternatives and decision criteria. Keep the set stable enough to observe change. *Sources checked on 14 July 2026. The cited studies describe their own corpora and do not guarantee a citation, recommendation or commercial result.* --- ### Your llms.txt is a sign an AI can read outside a door it cannot open URL: https://izzy.agency/en/blog/llms-txt-webmcp-website-ai-agents/ Published: 2026-07-16 Summary: An llms.txt file does not make a website usable by an AI agent. Learn what to check and where transactional sites should invest. ## The 60-second answer An `llms.txt` file can give machines a concise guide to the content you consider important. It does not prove that ChatGPT, Gemini or another assistant will retrieve that file, and it does not enable an agent to do anything on your website. There are two different jobs: - **identity**: describing your business and useful resources; - **capability**: exposing controlled functions an agent can use to search, filter, book or buy. For an informational website, a correct `llms.txt` may be a low-cost experiment. For ecommerce, booking platforms and SaaS products, agent capability is the more important engineering decision. A brochure can explain the shop. It cannot operate the till. ## In this article 1. [What an llms.txt file actually does](#1-what-an-llmstxt-file-actually-does) 2. [Why publishing it does not prove value](#2-why-publishing-it-does-not-prove-value) 3. [What WebMCP changes](#3-what-webmcp-changes) 4. [What to inspect now](#4-what-to-inspect-now) ## 1. What an llms.txt file actually does The `llms.txt` proposal gives a website a root-level file containing a concise, structured guide to selected resources. The intention is sensible: reduce navigation noise and point language models towards useful information. Publishing a file, however, is not the same as having it used. No mechanism requires an assistant to fetch it, trust it or prefer it to ordinary pages. Treat it as a descriptive layer, not a guaranteed shortcut into an AI answer. Start with the obvious check: open `yourdomain.com/llms.txt`. If a plugin generated it, read it as if it were a page for a prospective customer. Are the products, services, markets and links still correct? Does it update when the main site changes? An automated file can be current at launch and quietly wrong six months later. ## 2. Why publishing it does not prove value An Ahrefs analysis reported by Search Engine Journal found roughly 38,000 valid `llms.txt` files in a sample of 137,000 domains. In May 2026, **97% of those files received no requests**. AI retrieval bots accounted for only 1.1% of observed requests ([Search Engine Journal](https://www.searchenginejournal.com/97-of-llms-txt-files-got-no-requests-ahrefs-data-shows/579478/)). That finding does not prove the format will never be adopted. The sample skewed towards technically active websites and covered one month. It does show that, in 2026, the presence of a file is not evidence of meaningful use. Separate three levels when you review it: | Level | Useful question | Evidence required | |---|---|---| | Presence | Does the file exist? | Accessible URL and reviewed content | | Retrieval | Does any relevant system request it? | Server logs and identified bots | | Outcome | Does it change a citation or action? | Repeatable before-and-after testing with limitations | Most discussions stop at presence. An investment decision should depend on retrieval and outcome. ## 3. What WebMCP changes WebMCP tackles a different problem: how can a website expose structured tools that an agent is permitted to use? Chrome is testing it as an experimental feature in a **Chrome 149 origin trial**. A site can describe an action, its parameters and its rules, allowing an agent to invoke it more reliably than by guessing how the visual interface works ([Chrome for Developers](https://developer.chrome.com/blog/ai-webmcp-origin-trial?hl=en)). The status matters. WebMCP is not a universal production standard, and Chrome describes its Gemini integration as a separate later step ([Chrome AI developer preview](https://groups.google.com/a/chromium.org/g/chrome-ai-dev-preview-discuss/c/2rVDhuBOUAY)). Rebuilding a product around one trial would be premature. Ignoring the operating model it reveals would also be short-sighted. Transactional businesses can prepare sound foundations now: - clearly defined functions such as search, stock checks, quotes and booking; - stable inputs and outputs; - explicit permissions for sensitive actions; - human confirmation before payment, submission or an irreversible change; - logs showing what an agent requested and what the system executed. This is product and engineering work. The descriptive file can support it; the file cannot replace it. ## 4. What to inspect now Begin with a limited review, not a broad transformation programme. ### If you already have llms.txt Check its source, update date, links and owner. Generate it from maintained content where possible rather than creating a second manual version of your offer. Then inspect server logs. If no relevant system requests the file, that is useful evidence: keep its maintenance cost proportionate to its observed use. ### If your website supports transactions Choose one valuable customer task. It might be “find an item available in this size” or “check an appointment in this postcode”. Define the input, output, business rules, possible errors and point at which a person must confirm the action. This work remains useful if WebMCP changes. Clear, controlled and observable functions strengthen APIs, automation and the product experience as well. ### If AI visibility is the priority Do not reduce the diagnosis to `llms.txt`. Review crawlability, page clarity, factual consistency and the external sources that describe your business. Test a stable set of buyer questions afterwards. Visibility and agent capability are related, but they are not the same outcome. ## Conclusion: measure the door, not only the sign An accurate, automatically maintained `llms.txt` can be reasonable experimental cover. The available evidence does not justify treating the file alone as an AI visibility strategy. For a transactional website, the more durable question is what an agent can safely do, under which permissions and with which evidence. Prepare one useful capability before multiplying files. Capability, not compliance theatre. ## Is your website merely readable, or genuinely usable? Bring us one buyer task and the current journey. We can separate the visibility, product architecture and agent-capability questions, then scope the smallest defensible test. No promise of universal compatibility and no unnecessary transformation programme. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Does llms.txt improve visibility in ChatGPT? That outcome has not been demonstrated. The file can present selected resources, but its presence does not prove retrieval or a change in an assistant’s response. ### Is llms.txt useless? Not necessarily. If it is accurate, generated from maintained content and inexpensive to operate, it can be a sensible experiment. It should not displace higher-priority work. ### Is WebMCP ready for production? Not as a universal standard. In July 2026, Chrome presents it through an experimental origin trial. Transactional sites can still prepare structured functions, permissions and logging. ### What should an ecommerce team do first? Choose one high-value task, define its rules precisely and test whether it can be completed without ambiguity. Keep human confirmation for sensitive actions. *Sources and experimental status checked on 14 July 2026. This article is operational guidance, not a guarantee of visibility or future compatibility.* --- ### AI agent security: what changes when AI can take actions? URL: https://izzy.agency/en/blog/ai-agent-security-product-controls/ Published: 2026-07-15 Summary: When AI agents can call tools, change data or trigger transactions, security becomes a product-design decision. Use a capability contract to control risk. ## The 60-second answer An AI assistant produces an answer. An **AI agent** can go further: it can use software tools, consult external data, remember context and take actions. That may mean retrieving a delivery status, changing a customer record, issuing a refund or controlling a connected device. The security question is no longer limited to “Can the model say something wrong?” It becomes “What can the system do when the model is wrong, manipulated or misunderstood?” Prompts are not permissions. Model confidence is not authorisation. Before giving an agent more autonomy, define its purpose, tools, identity, trusted evidence, damage limits and recovery path. We call this a **capability contract**. ## 1. The risk boundary moves from content to consequence A conventional chatbot can provide an inaccurate answer. That matters, but the answer normally remains on screen until a person decides what to do with it. An action-taking agent shortens that distance. It can translate a model output into an API call, database change or physical operation. A mistake can become a business event before anyone notices it. The same model can therefore create very different levels of risk: - with a read-only catalogue, it can return the wrong item; - with CRM or payment access, it can change a record or create a financial loss; - with access-management or equipment controls, it can create security, operational or safety consequences. Impact depends on the tools available, the identity and permissions they use, whether actions are reversible and how quickly abnormal behaviour can be detected. OWASP describes the underlying failure as **excessive agency**: too much functionality, too much permission or too much autonomy for the task ([OWASP LLM06](https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM06_ExcessiveAgency.md)). That is a product-design problem before it is a model problem. ## 2. Write a capability contract before connecting a tool A capability contract is a concise agreement between product, engineering, security and operations about what the agent may do and how that boundary will be enforced. | Contract clause | Product question | Evidence before release | |---|---|---| | Purpose | What job may the agent perform, and what is explicitly outside it? | Defined use case, non-goals and misuse cases | | Tools | What is the smallest set of functions required? | Named, granular tools; unused and open-ended functions removed | | Authority | On whose behalf does each action run? | User or service identity, narrow scopes and downstream authorisation | | Evidence | Which inputs may influence an action? | Trusted-source rules and separation of instructions from external content | | Damage | What is the maximum acceptable effect of one error or attack? | Limits on spend, recipients, records, calls, time and retries | | Recovery | How will a person stop, inspect or reverse the action? | Preview, confirmation, audit trail, kill switch and rollback procedure | This turns “make the agent safe” into requirements a team can test. It also exposes weak concepts early. If the intended value requires a generic administrator account, unrestricted tools and no recovery path, the proposed capability is too broad. ## 3. Control actions by consequence, not model confidence Not every action needs the same friction. Classify it by consequence: A workable action model has four levels: ### Observe The agent reads without changing state. Access should still be scoped and logged because reading can disclose confidential data. ### Prepare The agent drafts an email, refund or order, but a person reviews the exact result before it is committed. This is often the right starting point. ### Reversible action The agent changes state, but the change can be reliably undone. Use narrow permissions, clear feedback and a tested reversal path. ### High-impact or irreversible action Payments, deletion, publication, access-right changes and physical operations need an independent control: human approval, a deterministic policy check, a second authorised person or a separate execution service. Approval must be attached to the exact action and parameters. It should not authorise a later request with a different amount, recipient or resource. The model may recommend an action. It should not be the final authority on whether that action is permitted. OWASP recommends enforcing authorisation in downstream systems rather than asking the model to police itself ([OWASP excessive-agency guidance](https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM06_ExcessiveAgency.md)). ## 4. Design for four attack paths - not only a bad user prompt ### Instructions hidden in external content An agent may read a webpage, document or email containing instructions designed to change its behaviour. The operator may never see them. Treat retrieved content and tool responses as untrusted data. Separate them from system instructions, restrict what they may influence and validate actions before execution. The [OWASP AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) includes external documents, APIs and emails in this boundary. ### Borrowed authority An agent helping one user may call a downstream system through a shared privileged account, reaching records that user could not access directly. Execute in the user’s context where possible. Recheck identity, role, resource and operation at the point of action. ### Poisoned memory Unverified content stored as memory can affect later tasks or users. Define what may be remembered, for how long and how it can be inspected or deleted. Isolate memory between users and trust levels. ### Cascades and runaway loops Tool chains can multiply actions and costs. Set budgets for steps, calls, elapsed time, records and spend. Stop automatically when the workflow exceeds its declared purpose or normal pattern. MITRE ATLAS maps agent-specific behaviours including tool invocation, context poisoning and tool poisoning. Use them as abuse cases for the actual workflow, not as a list for a slide ([MITRE ATLAS](https://atlas.mitre.org/)). ## 5. Put agent security requirements in the product backlog Security cannot remain a review performed after the agent appears to work. Each capability needs acceptance criteria alongside its functional requirements. For an action-taking agent, those criteria may include: - tools expose only allowed operations and resources; - downstream services verify identity and authorisation on every action; - external content cannot silently become a control instruction; - parameters are validated and approval is bound to the displayed action; - memory has retention, isolation and deletion rules; - calls, retries, cost and transaction size have enforced limits; - logs connect the user, model, policy, tool request and result without exposing secrets; - suspension, rollback and incident procedures are tested; - release tests cover manipulation, privilege escalation, data leakage and unavailable tools. This belongs in the normal software lifecycle. NIST’s Secure Software Development Framework covers organisational preparation, software protection, well-secured releases and response to residual vulnerabilities. It is intended to be adapted to business risk, not treated as a mechanical checklist ([NIST SSDF](https://csrc.nist.gov/projects/ssdf)). For broader governance, the NIST AI Risk Management Framework addresses trustworthiness across AI design, development, use and evaluation. NIST notes that AI RMF 1.0 is being revised, so check the current version before fixing it into policy ([NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework)). ## 6. Assign ownership across the product, not to “the AI team” | Owner | Decision to own | |---|---| | Product | Intended outcome, prohibited uses, user friction and acceptable consequence | | Engineering | Tool design, identity, permissions, validation and failure handling | | Security | Threat model, control requirements, adversarial testing and incident criteria | | Operations | Monitoring, suspension, escalation, recovery and post-incident learning | | Legal, privacy and domain specialists | Regulatory, contractual, data and sector-specific constraints | The structure varies by company, but responsibility should follow the business consequence. CISA’s secure-by-design guidance asks software manufacturers to take greater ownership of customer security outcomes and integrate security into product decisions from the start ([CISA and international partners](https://www.cisa.gov/sites/default/files/2023-10/Shifting-the-Balance-of-Cybersecurity-Risk-Principles-and-Approaches-for-Secure-by-Design-Software.pdf)). An agent should not transfer that burden to the customer through obscure settings and broad defaults. ## 7. Ten questions for an agent launch review Before releasing a new capability, ask: 1. What is the highest-impact action this agent can take? 2. Which tool or permission is not strictly necessary? 3. Does each action run as the correct user or a generic privileged identity? 4. Can external content influence an action without being treated as untrusted? 5. What exact parameters does a person see before approving? 6. Which actions can be reversed, and has reversal been tested? 7. What limits stop repeated calls, large transactions or runaway cost? 8. Can an operator reconstruct why an action occurred without exposing sensitive data in logs? 9. Who can suspend the agent, and how quickly? 10. Which evidence would block the release rather than become a post-launch task? A demo proves that the happy path can work. A launch review must cover malicious data, wrong permissions, tool failure, model change and reasonable human mistakes. ## Conclusion: autonomy should be earned capability by capability AI agent security is not achieved by one confirmation screen or a list of prohibited words. Review the complete action path: input, model decision, tool, identity, permission, consequence and recovery. Start with a narrow capability contract and minimum authority. Increase autonomy only when limits, approvals, monitoring and recovery work under realistic failure conditions. The objective is not zero autonomy. It is autonomy whose consequences the organisation can explain, contain and own. ## Before you add more autonomy, review the action path Bring us one agent workflow, its tools and the highest-impact action it can take. We can map the capability contract, permission model, approval points and release evidence - without turning a prototype into a security-theatre exercise. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Is prompt injection the only important AI agent risk? No. It is one route to failure. Excessive permissions, shared privileged identities, unsafe memory, compromised tools, weak validation, cascading actions and poor recovery can create serious consequences even without a malicious user prompt. ### Does every agent action need human approval? No. Approval should follow consequence. Read-only and low-impact reversible actions may be automated within enforced limits. Irreversible, financial, externally visible or safety-relevant actions need independent control proportionate to the risk. ### Is a read-only agent safe? It usually has a lower integrity risk because it cannot change records, but it may still expose confidential data, retrieve malicious instructions or access information outside the user’s rights. Scope and authorise read access carefully. ### Can the model decide whether a user is authorised? The model can help interpret intent, but the downstream system should make the final authorisation decision using deterministic identity and policy controls. ### When is an AI agent ready for production? When each capability has defined boundaries, test evidence, monitoring, an incident owner and a workable recovery path. A successful demonstration or benchmark score is not enough. *Primary guidance checked on 15 July 2026. This article provides general product and operational guidance; it is not a security certification, penetration test or legal opinion.* --- ### Website redesign or migration: how to protect your Google visibility URL: https://izzy.agency/en/blog/website-migration-protect-google-visibility/ Published: 2026-07-15 Summary: Protect search visibility during a website migration with an URL inventory, redirect proof, launch controls and evidence-based incident response. ## The 60-second answer A redesign becomes a search migration when it changes how engines discover, identify or interpret important pages. The risk is not the new visual design itself. It is the broken continuity between the old evidence and the new website. Protect that continuity through five controls: 1. record what is changing; 2. map every valuable old URL to an appropriate outcome; 3. test search and commercial journeys before launch; 4. monitor by failure type rather than one traffic total; 5. keep ownership until the new system reaches a stable baseline. Google may need time to process correct changes. Waiting is appropriate only after you have proved that the implementation is sound. ## 1. Start with the migration surface “We are changing the website” is too vague for risk planning. List the systems that will move. ### Address changes Domain, subdomain, protocol, folders, product handles, language paths and parameters. ### Content changes Removed pages, consolidated services, rewritten categories, altered headings and changed internal links. ### Rendering changes New JavaScript framework, server-side rendering, client-side navigation, image delivery or consent controls. ### Commercial changes Checkout, forms, account journeys, product feeds, stock, currencies and analytics events. ### Search instructions Canonicals, robots rules, sitemaps, hreflang, structured data and status codes. This creates the migration surface: the complete set of things that can break or require reprocessing. A redesign can look modest and still have a large migration surface. ## 2. Build the evidence pack before building redirects The redirect map is important, but it should not be the first artefact. Begin with evidence of what the existing site does. For every priority URL, capture: - current address and intended destination; - page purpose and target market; - organic clicks and impressions; - conversions or assisted commercial actions; - significant external and internal links; - indexability, canonical and status code; - replacement, consolidation or retirement decision. Do not preserve a page merely because it exists. Do not delete it merely because traffic is low. A low-traffic page may support a valuable journey, earn links or explain a product constraint. The evidence pack also creates a pre-launch baseline. Without it, the team cannot tell whether a post-launch change is a fault, expected processing or an existing weakness. ## 3. Give every old URL an explicit outcome Each discovered URL should have one of four outcomes: | Outcome | Use when | Required proof | |---|---|---| | Retain | Address and purpose remain valid | Same URL returns the intended page | | Redirect | A clear replacement exists | Permanent redirect to the closest relevant destination | | Consolidate | Several pages now serve one intent | Destination covers the material need of the old pages | | Retire | No replacement or continuing value exists | Deliberate not-found or gone response, with links removed | Avoid universal redirects to the homepage. They hide planning gaps and give users a destination unrelated to their request. Google’s own site-move guidance recommends mapping old and new URLs, using permanent server-side redirects and avoiding redirect chains where possible ([Google Search Central: site moves](https://developers.google.com/search/docs/crawling-indexing/site-move-with-url-changes)). Test the map as data. A spreadsheet reviewed by eye is not proof that the server implements it. ## 4. Define launch acceptance tests The new site should not launch because it looks finished. It should launch because agreed tests pass. ### Search controls - Priority destinations return HTTP 200. - Old priority URLs return the planned permanent redirect. - Redirects do not loop or pass through avoidable chains. - Canonicals point to indexable final URLs. - Production pages are not blocked by `robots.txt` or `noindex`. - Sitemaps contain the final canonical addresses. - Internal links use final URLs rather than relying on redirects. - Structured data matches visible page content. ### Commercial controls - Forms reach the correct owner. - Checkout, booking and account journeys work on mobile. - Prices, stock and product feeds agree. - Consent and analytics events fire as designed. - Error messages and recovery paths are usable. ### Operational controls - Monitoring starts before DNS or routing changes. - A named person can roll back or patch critical faults. - The team knows which errors require immediate intervention. Record the result. “The agency checked it” is not an acceptance test. ## 5. Triage post-launch movement by failure type One traffic chart cannot diagnose a migration. Split the evidence. | Observation | Likely investigation | |---|---| | Widespread crawl errors | Hosting, routing, robots or server response | | Old URLs remain indexed | Redirect discovery, chains or mapping quality | | New pages classed as duplicates | Canonicals, similarity and content differentiation | | One section loses visibility | Mapping, navigation, content or template issue | | Traffic holds but conversion falls | UX, forms, checkout or tracking | | One market disappears | Hreflang, localisation, routing or regional content | Some correct changes take time to settle. Google notes that after certain content corrections, pages can remain in a duplicate cluster while canonical selection is re-evaluated ([Google canonicalisation troubleshooting](https://developers.google.com/search/docs/crawling-indexing/canonicalization-troubleshooting)). That is a reason to inspect before reacting, not a reason to ignore errors for a fixed period. Server failures, accidental blocking and broken redirects require immediate correction. A canonical still being reconsidered may require patience and evidence. ## 6. Use recovery criteria, not a launch anniversary A migration is not stable because a particular number of days has passed. It is stable when the new system behaves coherently. Look for: - priority redirects repeatedly crawled and respected; - destination pages indexed as intended; - old addresses declining without valuable pages becoming orphaned; - errors returning to an agreed operating range; - search visibility forming a consistent new baseline; - commercial journeys and reporting reconciled; - unresolved exceptions assigned to owners. Review by page type and market. A stable homepage can hide a failed product category. Total traffic can hide a conversion loss. Keep the migration register open until the evidence supports closure. Handover is an operational decision, not a project-management ceremony. ## Conclusion: protect continuity, not just rankings A website migration moves more than pages. It moves addresses, evidence, commercial journeys and the instructions search engines use to connect old and new. Map the change surface, prove the implementation and diagnose movement by type. The goal is not a promise of zero fluctuation. It is a controlled transition in which errors are visible and recoverable. ## Can your team prove the migration is ready to launch? Bring us the URL inventory, proposed architecture and critical customer journeys. We can turn them into a migration register, executable checks and a post-launch control plan - without pretending all movement can be eliminated. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Will a redesign automatically damage search visibility? No. Risk depends on what changes and how continuity is managed. URL, rendering, content and technical-instruction changes deserve explicit controls. ### Should every old page be redirected? No. Redirect when an appropriate replacement exists. Retire pages deliberately when no useful equivalent remains, and remove links pointing to them. ### When should the team intervene after launch? Immediately for server errors, blocking, broken journeys and incorrect redirects. For processing or canonical reassessment, inspect the configuration and evidence before changing it again. ### How long should redirects remain? Google recommends keeping them for at least a year, and users or external links may justify retaining them longer. Treat removal as a separate risk decision. ### What is the most important migration document? The live migration register connecting each important URL, intended outcome, test result, post-launch observation and owner. *Primary Google documentation checked on 15 July 2026. Migration effects vary by site, architecture and scale; this article is operational guidance, not a performance guarantee.* --- ### From AI pilot to production: the governance checklist SMEs actually need URL: https://izzy.agency/en/blog/ai-pilot-to-production-governance-checklist/ Published: 2026-07-14 Summary: Move one AI use from demo to controlled production with six decisions, a ten-question gate and clear controls for operations, content and code. Your AI demo works. Then it meets real data, deadlines, more users, a changing supplier - and the owner is absent when something goes wrong. A demo proves possibility. Production requires ownership of the outcome, evidence, operating limit and stop decision. ## Answer in 60 seconds - Start with **one business workflow**, not a company-wide AI programme. - Choose an operating default: **allow, assist, review, disclose, restrict or block**. - Do not release until ten questions about ownership, data, testing, review, security, monitoring and fallback have usable answers. - Apply different controls to operations, public content and production code. - Pause for legal, privacy, employment, IP, security or sector review when the facts trigger it. For an SME, governance can be one current operating record with owners and workflow evidence - not a compliance badge or safety promise. ## In this article 1. [A working AI demo is not a production system](#1-a-working-ai-demo-is-not-a-production-system) 2. [Choose: allow, assist, review, disclose, restrict or block](#2-choose-allow-assist-review-disclose-restrict-or-block) 3. [Use this ten-question production gate](#3-use-this-ten-question-production-gate) 4. [Operations, content and code need different controls](#4-operations-content-and-code-need-different-controls) 5. [What to do in the first ten working days](#5-what-to-do-in-the-first-ten-working-days) 6. [When to pause and bring in a specialist](#6-when-to-pause-and-bring-in-a-specialist) 7. [Conclusion](#conclusion-production-requires-ownership) 8. [Frequently asked questions](#frequently-asked-questions) 9. [Sources](#sources) ## 1. A working AI demo is not a production system AI use is already material: [Eurostat reported that 20.0% of covered EU enterprises with at least ten people employed used AI in 2025](https://ec.europa.eu/eurostat/web/products-statistical-reports/w/ks-01-26-009). The frame excludes smaller micro-enterprises and some sectors; it measures use, not maturity, demand or return. A demo has favourable inputs and a builder nearby. Production adds workload, unusual cases, customer consequences, provider changes, incidents and maintenance. Govern the complete workflow: people, process, data, model, retrieval, tools, suppliers, destinations and recovery. The voluntary, non-prescriptive [NIST AI Risk Management Framework](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10) connects governance, context, measurement and response; using it does not prove a control works in your system. Could a named owner explain the release, supporting evidence, stop condition and who can execute it? ## 2. Choose: allow, assist, review, disclose, restrict or block Approve a use, not a tool in the abstract. | Default | Use it when | Minimum control | |---|---|---| | **Allow** | Reversible, low-consequence internal work | Approved tool, owner, input rules, user check | | **Assist** | A person remains the primary creator or analyst | Task boundary, approved sources, competent owner | | **Review** | Output may affect customers, operations or brand | Tested cases, approver, traceability, correction | | **Disclose** | The context creates a transparency decision | Scope decision, relevant notice or marking, record | | **Restrict** | Sensitive data, code, privileges or material actions | Least privilege, stronger tests, staged release, specialist input | | **Block** | Ownership, lawful use, critical tests or recovery are missing | Redesign or qualified decision before re-evaluation | Categories can combine: allow ideation, review publication, restrict confidential inputs and block automatic publishing. Set control depth from consequence, reversibility, sensitivity and agency. Small-company status does not make a consequential use harmless. Binding rules override the default. The [EU AI Act](https://op.europa.eu/en/publication-detail/-/publication/dc8116a1-3fe6-11ef-865a-01aa75ed71a1/language-en) and [GDPR](https://op.europa.eu/en/publication-detail/-/publication/3e485e15-11bd-11e6-ba9a-01aa75ed71a1/language-en) can create role- and use-specific duties; this table neither classifies your system nor establishes compliance. ## 3. Use this ten-question production gate Before a workflow receives production data, reaches an external audience, changes code or acts through tools, answer these questions: 1. **Purpose and owner:** what task and outcome are in scope, and who owns the decision? 2. **Baseline:** what happens without AI, including rework, review burden and failure consequence? 3. **System boundary:** which service, version, instructions, sources, plugins and destinations are involved? 4. **Data and rights:** what personal, confidential or protected material enters or leaves the flow? 5. **Consequence and authority:** what may the system propose or do, and what remains human-owned? 6. **Acceptance evidence:** which normal, edge, refusal and abuse cases were tested; which failures remain? 7. **Human responsibility:** can the reviewer recognise and reject an error under real workload? 8. **Security and supplier controls:** which permissions, components, changes, incidents and exit conditions are controlled? 9. **Operations:** which signals have an owner, threshold, permitted response and appropriate retention rule? 10. **Fallback and change:** how will the team restore a known-safe or manual route, and which change reopens approval? Keep one production record linked to tests and decisions. The cited guidance does not provide a universal score, test-set size or risk-reduction percentage; evidence must fit the use and consequence. The [NIST AI RMF Playbook](https://airc.nist.gov/airmf-resources/playbook/) is a voluntary resource, not a checklist that produces a release verdict. ## 4. Operations, content and code need different controls One policy cannot replace workflow-specific control. ### Operations Limit what the system may see, recommend and execute. Start in an isolated route; name abort conditions, alert owners and a manual fallback. Measure review, correction, incidents, maintenance and operating cost - not generation speed alone. The [NCSC secure-AI guidelines](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines) are principle-level guidance, not assurance. ### Content Separate factual checks, rights, brand review, disclosure and correction. Fluency passes none by default. The final [Article 50 transparency code](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content) is voluntary and scope-specific; not every AI-assisted item is covered. The [Commission and AI Board assessed it as adequate for relevant obligations](https://digital-strategy.ec.europa.eu/en/library/commission-opinion-assessment-code-practice-transparency-ai-generated-content), but adherence is not conclusive evidence of compliance. Scope remains a specialist decision. ### Code Treat AI-assisted code as an untrusted contribution under human ownership. Protect repositories and credentials, review the full diff, derive tests from requirements and keep secure-development gates. The [NIST Secure Software Development Framework](https://csrc.nist.gov/pubs/sp/800/218/final) is a high-level structure, not certification. ## 5. What to do in the first ten working days Use the first ten working days to make one workflow visible and testable: 1. Name the business owner, technical owner and specialist contacts. 2. Select one bounded use with a visible baseline and reversible pilot path. 3. Inventory users, data, versions, retrieval, plugins, permissions and destinations. 4. Record affected people and screen legal, privacy, employment, IP, security and sector triggers. 5. Set the allow/assist/review/disclose/restrict/block defaults. 6. Define accepted outcomes, severe failures, refusal cases and review. 7. Confirm approved data, rights, retention, provider reuse and transfer questions. 8. Create an isolated pilot with least privilege and a manual fallback. 9. Agree signals, owners, permitted actions, incident authority and change triggers. 10. Route unresolved consequential decisions before exposure. This is not a ten-day production promise. It exposes the first real control gap; the outcome may be a pilot, redesign, specialist decision or stop. ## 6. When to pause and bring in a specialist Pause before exposure when the workflow involves: - prohibited- or high-risk-use questions, provider/deployer roles, conformity or Article 50 scope; - personal or special-category data, automated decisions, transfers, monitoring or a possible data protection impact assessment (DPIA); - hiring, worker evaluation, allocation, surveillance or material changes to work; - training or source-data rights, licences, reservations, similarity or important output ownership; - privileged actions, internet exposure, sensitive architecture, vulnerabilities or an incident; or - health, finance, insurance, education, critical infrastructure, public services or regulated products. IZZY can map the workflow, evidence and decision points. Qualified specialists own the legal, privacy, employment, IP, security, conformity and sector conclusions. ## Conclusion: production requires ownership A policy matters only when it changes a workflow. Set one operating default, close the ten-question gate and preserve a working fallback. Sometimes the answer is a smaller AI role - or none. That is still a useful decision. ## Bring one AI workflow. We'll give you an honest read. Bring your policy or tool list and available evidence. In 30 minutes, we'll give you an initial, bounded read on whether the workflow is ready for a bounded pilot, is missing a production control, needs a specialist decision or should stay out of production. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) or explore [IZZY's AI integration approach](https://izzy.agency/en/services/ai-integration/). ## Frequently asked questions ### What is the difference between an AI demo and a production-ready AI system? A demo shows capability on selected inputs. A production-ready workflow has an owner, defined boundaries, acceptance evidence, competent review, monitoring, an incident path and an executable fallback. ### Does a small business need an AI governance policy? A small business needs operating rules when AI affects staff, customers, data, content, code or systems. Keep them proportionate and connect approved uses to owners, evidence, escalation and stop authority. ### What should an AI usage policy include? Include approved tools, permitted and prohibited uses, data restrictions, review responsibility, content and code rules, incident reporting, supplier changes and stop authority. ### How do we manage shadow AI without banning every tool? Ask where unregistered use occurs and why. Offer a usable approved path, protect sensitive inputs and escalate consequential uses; prohibition alone may not remove the underlying need. ### Does the EU AI Act apply to SMEs? SME status is not a universal exemption. Applicability and duties depend on the system, intended purpose, role, modifications and facts. Obtain qualified advice for client-specific classification or legal conclusions. ## Sources - [Eurostat: The use of artificial intelligence technologies in the European Union - key results, 2026 edition](https://ec.europa.eu/eurostat/web/products-statistical-reports/w/ks-01-26-009) - [NIST: Artificial Intelligence Risk Management Framework 1.0](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10) - [NIST: AI RMF Playbook](https://airc.nist.gov/airmf-resources/playbook/) - [UK NCSC and international partners: Guidelines for Secure AI System Development](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines) - [NIST: Secure Software Development Framework 1.1](https://csrc.nist.gov/pubs/sp/800/218/final) - [European Union: Regulation (EU) 2024/1689 - Artificial Intelligence Act](https://op.europa.eu/en/publication-detail/-/publication/dc8116a1-3fe6-11ef-865a-01aa75ed71a1/language-en) - [European Union: Regulation (EU) 2016/679 - General Data Protection Regulation](https://op.europa.eu/en/publication-detail/-/publication/3e485e15-11bd-11e6-ba9a-01aa75ed71a1/language-en) - [European Commission: Code of Practice on Transparency of AI-Generated Content](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content) - [European Commission and European AI Board: assessment of the Transparency Code](https://digital-strategy.ec.europa.eu/en/library/commission-opinion-assessment-code-practice-transparency-ai-generated-content) *How this was prepared: evidence-led synthesis from official statistics, public frameworks, regulator and legal sources, and IZZY's signed internal evidence dossier. Sources and legal/guidance status checked 13 July 2026. This is general operational information, not legal advice, a security or conformity assessment, or a forecast of commercial results.* --- ### Refactor, replatform or rescue? How to decide before a website migration URL: https://izzy.agency/en/blog/migration-rescue/ Published: 2026-07-14 Summary: Rescue, refactor, replatform or selective rebuild? How to decide before a website migration - using evidence, essential journeys, tests and a safe way back. Your website feels fragile, or the project meant to improve it has stalled. Suppliers recommend different solutions, but you need to decide without becoming a software specialist. Start by asking what is known, what must work and how change will be checked. A migration may be justified, but it should not carry uncertainty into a larger project. Choose the smallest change that solves the real problem. ## 1. Regain control before choosing a solution When releases fail or access is unclear, the whole website can look beyond repair. That appearance is not yet a diagnosis. A broken connection, missing account or uncertain release may be the immediate problem. Establish who can access the live website, hosting, domain, code, content and connected services. Identify the last stable version and recent release or incident notes. Name one decision-maker and a person responsible for each important system. This is a control exercise, not a hidden rebuild. If a cyber incident may be active, stop treating the work as a normal migration. ANSSI places remediation alongside investigation and crisis management, so appropriate incident specialists should guide the response before ordinary transformation work resumes ([ANSSI](https://cyber.gouv.fr/securisation/gestion-de-crise/piloter-la-remediation-dun-incident-cyber/)). Replace competing opinions with a shared picture. If a supplier cannot explain the current state, the problem and the evidence behind their recommendation, you do not yet have enough to approve a route. ## 2. Identify what must keep working Before discussing platforms, identify the business journeys that cannot stop. These may include an enquiry reaching the right person, checkout, login, publishing or data reaching another business system. For each journey, record what happens now, who owns it and what a correct result looks like. This creates a baseline: a documented starting point for judging change. NIST's security-focused guidance is not a website migration manual, but its principles of documenting a baseline, controlling change and monitoring the result are useful here ([NIST SP 800-128](https://csrc.nist.gov/pubs/sp/800/128/upd1/final)). You may hear the word **parity**. In ordinary language, parity means the new or changed website can do the agreed things the business needs at least as well as the accepted starting point. It does not mean copying every old page or preserving every defect. Decide what must remain, what should improve and what can be retired. Before approval, ask for the access record, last stable release, essential journeys and owners, content and URL inventory, connected services, agreed checks showing that each essential journey still works, and known gaps. If URLs change, Google recommends mapping old addresses to new ones, testing redirects and monitoring the move. It also recommends separating major changes, such as a domain, content-management system and layout change, where possible ([Google Search Central](https://developers.google.com/search/docs/crawling-indexing/site-move-with-url-changes)). This is guidance, not a promise about visibility. ## 3. Choose rescue, refactor, replatform or selective rebuild An owner may need four choices, with different treatments for parts of one website. - **Rescue** means restoring control and stability before making a larger commitment. Choose it when access, ownership, recent changes or the current condition are unclear. - **Refactor** means improving a limited part of the existing website's code or set-up while keeping its intended behaviour. Choose it when the platform is still suitable but a defined area is hard to change safely. - **Replatform** means moving the required behaviour to a different platform. Choose it when the platform itself blocks a needed change - for example, it cannot connect to a payment or booking service the business requires - and that limit cannot reasonably be removed in place. - **Selective rebuild** means replacing only the parts that cannot be made safe or workable, while preserving the rest. Choose it when the boundary and expected behaviour can be stated clearly. This is IZZY's decision framework, not a set of universal definitions. A proposal should connect its route to a specific constraint, explain what stays unchanged and show how the result will be tested. If you searched for "**refactor vs replatform website**", remember that those labels describe only two choices. The evidence may point to rescue or selective rebuilding instead. If you are wondering whether to rebuild or migrate, ask why repair or selective replacement is insufficient. "Starting clean" is not evidence, nor is keeping the platform from habit. Follow the business journeys, connected services, recovery options and quality of what is known. A failed project can be rescued without pretending earlier work never happened. Preserve what exists, separate usable parts from assumptions and agree one controlled next decision: migrate, make a smaller repair or say "not yet". ## 4. Approve change only with tests, responsibility and a safe way back A credible proposal names the person responsible for each essential journey, the test that shows whether it works and the signal watched after release. It also states who may proceed or stop. Ask to see representative tests before the main change. Pages loading is not enough. Forms, payments, logins, publishing, search, analytics and connected services need appropriate checks. Compare content, URLs and transferred data with the baseline too. The plan needs a **safe way back**. Teams may call this a **rollback**, meaning a controlled return to the last stable version if a release causes an unacceptable problem. The plan should say who can trigger it, for how long, what happens to new information and how recovery will be checked. Sometimes switching visitors, restoring data or correcting the new version is safer than reversing every change; make the route explicit. Do not approve a launch date while access, responsibility, testing or recovery is "to be confirmed". Proceed when evidence gaps are visible, accepted by named people and matched with stop conditions. ## Conclusion You need not choose on instinct or referee suppliers. Regain control, define what must continue, choose the smallest justified route, and approve it only when tests, responsibility and a safe way back are clear. Migration is one answer, not the starting assumption. ## Bring the evidence, not a predetermined answer Bring the current website, the access situation, recent release history, essential customer journeys and any proposed migration plan. IZZY can assess which evidence is present, which gaps prevent a responsible decision and whether the next step is rescue, a focused review or a migration assessment. That is an evidence-gap assessment, not a promised migration outcome. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) or [send the team a short brief](https://izzy.agency/en/contact/). ## Frequently asked questions ### Should I rebuild or migrate my website? Neither choice is automatic. Rebuild where required behaviour cannot change safely in place. Replatform when the platform itself blocks a change the business needs - for example, it cannot support the required checkout - and essential journeys can be tested elsewhere. Rescue or refactor may be enough when the problem is control, stability or one bounded area. ### Can a failed website project be rescued? Preserve current work, restore access, identify the last stable release and separate verified facts from assumptions. The evidence may support repair, a revised migration or a stop. It cannot guarantee that every previous investment is recoverable. ### What evidence should I request before a website migration? Request a list of who can access each system, the release history, baseline, essential journeys, content and URL inventory, the person responsible for each connected service, agreed checks, known gaps, the person who decides whether the launch goes ahead, and the recovery plan. ### Does every migration need a rollback plan? For a material release, agree a recovery route, though not always a literal reversal. Name the decision-maker, timing, data consequences and checks that confirm an acceptable state. ### Can we change the platform, domain and design together? Google recommends separating major changes where possible. If they must happen together, first trial an essential customer journey on the new platform and domain with the proposed design. Name the person who reviews the results and decides whether to proceed, pause or use the safe way back. Combined changes make problems harder to trace. --- ### How can a small team create content consistently - without making it a full-time operation? URL: https://izzy.agency/en/blog/small-team-content-workflow/ Published: 2026-07-13 Summary: A practical weekly content workflow for small teams - from buyer question and evidence to distribution, follow-up and measurable learning. Your team probably does not lack ideas. It lacks a dependable way to turn one useful idea into a finished business asset. When the same person must find the question, verify the answer, write, chase approval, publish, adapt, respond and report, the feed does not go quiet because somebody forgot the calendar. The operating chain broke. A small team can publish consistently by limiting work in progress, starting with real buyer questions, assigning evidence and response owners, and treating distribution and learning as part of delivery. One finished cycle beats a full queue. ## Answer in 60 seconds - **What is probably happening:** work is stalling between evidence, review, distribution and follow-up - not at the idea stage. - **What to check first:** map the last five content pieces from question to business follow-up. Mark where each one waited or disappeared. - **What not to change yet:** do not add a channel, tool or higher posting target before you find the first material break. - **When outside help is useful:** when the expertise exists but nobody can protect the workflow, proof standard or conversation handoff. ## In this article 1. [Why small-team content breaks](#1-why-small-team-content-breaks) 2. [Why a calendar is not the operating system](#2-why-a-calendar-is-not-the-operating-system) 3. [The weekly workflow](#3-a-weekly-content-workflow-a-small-team-can-actually-run) 4. [Roles without a newsroom](#4-name-the-roles-without-building-a-newsroom) 5. [Where AI helps - and where it does not](#5-use-ai-after-the-thinking-starts) 6. [What to measure](#6-measure-decisions-not-content-noise) 7. [When to keep it in-house or get help](#7-keep-it-in-house-get-help-or-stop) ## 1. Why small-team content breaks The problem is wider than production. In a 2026 CMI/MarketingProfs study of 1,015 B2B marketers, respondents selected content that prompts action (40%), resource constraints (39%) and measurement (33%) among their leading challenges. The sample is broad, mostly North American and not a census of small teams. It still shows that output, capacity and business usefulness are connected problems - not three separate to-do lists. [See the study and methodology.](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research) Look for the first break: - topics are chosen without a buyer question or evidence owner; - approval arrives from several people in several places; - publication is treated as the finish line, so distribution and replies disappear; - reporting counts posts and clicks but changes no decision. More volume makes each of these failures more expensive. ## 2. Why a calendar is not the operating system Keep the editorial calendar. It is useful for dates, topics, formats and channels. But a calendar schedules work. An operating system finishes it. It also records: | Before production | After publication | |---|---| | buyer question and reader decision | planned distribution completed | | evidence owner and material claims | replies triaged and routed | | original point of view | sales or CRM follow-up recorded | | reviewer and release criteria | keep, improve, repurpose or stop decision | If those decisions live only in one person's head, you do not have a lightweight process. You have a single point of failure. ## 3. A weekly content workflow a small team can actually run Use one visible cycle: 1. **Capture a real question.** Pull it from sales, support, search data, workshops or recurring objections. Record its origin. 2. **Select one priority.** Choose one question and one backup. A backlog is storage, not a commitment. 3. **Build an evidence-first brief.** Name the reader decision, sources, expert, claims, caveats, format, distribution and next step. 4. **Extract the point of view.** Ask the person who did the work for examples, distinctions, objections and what they would *not* recommend. 5. **Produce and review.** Create the smallest complete asset. Consolidate factual, editorial and risk feedback. 6. **Publish and distribute.** Adapt the useful parts only for channels where the audience already pays attention. 7. **Converse and learn.** Answer, route qualified problems, capture new questions and decide what changes next. The cycle is complete only when the asset is live, material claims are checked, planned distribution is closed or deliberately stopped, replies have an owner and the next decision is recorded. That is consistency: dependable completion, not compulsory frequency. ## 4. Name the roles without building a newsroom You do not need seven new hires. You need seven decisions with names beside them. | Decision | Minimum owner | |---|---| | What is worth this week's capacity? | business owner | | What do we know, and what proves it? | subject expert | | How does the asset move to done? | content operator | | Is it accurate, useful and safe to release? | named reviewer | | Where will it be distributed? | channel owner | | Who answers and escalates? | conversation owner | | Where does a qualified signal go? | sales or CRM owner | One person may cover several roles. The point is not headcount; it is explicit authority. The [GOV.UK service-team model](https://www.gov.uk/service-manual/the-team/what-each-role-does-in-service-team) is designed for a different context, but its useful principle transfers: skills can be inside the team or available to it, while decision responsibility remains clear. ## 5. Use AI after the thinking starts AI can expand query ideas, organise interview notes, propose outlines, transform an approved asset into channel variants and run consistency checks. It cannot own the buyer question, the commercial judgment, the proof, the expert point of view or the release decision. Google says generative AI can help with research and structure, while scaled pages without added value may breach its spam policies. Its guidance asks for original information, analysis and clear sourcing; it also says there is no preferred word count. [Read the people-first guidance](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) and [the generative-AI guidance](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content?hl=en). NIST identifies confident false output - confabulation - as a generative-AI risk and recommends checking sources and citations. For a small team, the practical rule is simple: AI may assist the work; a named human still signs off the facts. [See the NIST Generative AI Profile.](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf) ## 6. Measure decisions, not content noise A useful scorecard separates activity from outcome. The [AMEC Integrated Evaluation Framework](https://amecorg.com/amecframework/) makes the same distinction: outputs record what was produced; outcomes ask what changed. | Layer | Useful question | Examples | |---|---|---| | Capacity | Can the system finish what it starts? | cycle time, queue age, review delay | | Evidence | Is the asset trustworthy and distinct? | traced claims, corrections, expert acceptance | | Delivery | Did the complete plan run? | publish and distribution completion | | Conversation | Did it create a useful exchange? | questions, objections, qualified conversations | | Commercial learning | Did it assist a real business decision? | sales use, accepted handoff, opportunity with content touch | Do not collapse these into one vanity number. A click is not a lead. A lead is not revenue. Record the evidence separately, then decide: keep, change, expand, refresh or stop. ## 7. Keep it in-house, get help or stop Keep the system internal when a decision owner can protect one weekly priority, experts are available, review is prompt and somebody owns follow-up. Consider bounded outside help when the expertise exists but the workflow repeatedly stalls, evidence is scattered, distribution is improvised or useful enquiries have nowhere reliable to go. Stop and fix a different problem first when there is no credible buyer question, no proof owner, no response capacity, or the real constraint is offer clarity, product quality, sales capacity, analytics integrity or a broken website. Content cannot compensate for a business system that is not ready to receive attention. ## Conclusion: install the loop, not more pressure Small teams do not need to behave like publishers. They need a content system sized to their actual capacity. Start with one question, one evidence owner, one useful asset and one complete route to a conversation. Finish it. Learn from it. Then decide whether the next cycle deserves to exist. ## Bring us the symptom. We'll give you an honest read. If your content cycle keeps disappearing into urgent work, IZZY can map the first material break and scope the smallest useful fix. Bring your last five pieces, current queue, available analytics and a few replies or enquiries. If a project is not justified, the call should make that clear too. [Book a 30-minute scoping call](https://calendar.app.google/Eq7USk7KKoTwzGiA9) or [send the team a short brief](https://izzy.agency/en/contact/). ## Frequently asked questions ### How often should a small team publish? There is no universal rate. Choose the slowest cadence that your team can finish with evidence, review, distribution and follow-up. Increase frequency only after the full cycle is stable. ### What is the difference between a content calendar and a content workflow? A calendar records what and when. A workflow records how a buyer question becomes an evidenced asset, who approves it, how it is distributed, who responds and what the team learns. ### Can one person run the whole content system? One person can operate the flow, but should not implicitly own every fact, approval, public response and commercial handoff. Keep the process small while making authority explicit. ### Should a small team use AI to write content? Yes, for bounded assistance such as research leads, structure, transcription and approved variants. A human should verify sources, protect confidential data, supply the point of view and approve the final asset. ### What should we fix first if content is inconsistent? Map the last five pieces and locate the earliest repeated delay: question selection, evidence, drafting, review, publication, distribution or response. Fix that break before adding tools or output targets. ## Sources - [CMI/MarketingProfs: B2B Content and Marketing Trends - Insights for 2026](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research) - [Google Search Central: Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) - [Google Search Central: Guidance on using generative AI content](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content?hl=en) - [NIST: Artificial Intelligence Risk Management Framework - Generative AI Profile](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf) - [GOV.UK Service Manual: What each role does in a service team](https://www.gov.uk/service-manual/the-team/what-each-role-does-in-service-team) - [AMEC Integrated Evaluation Framework](https://amecorg.com/amecframework/) --- ### Your website gets visitors but no enquiries? Check these five things before spending more on ads URL: https://izzy.agency/en/blog/website-visitors-no-enquiries/ Published: 2026-07-13 Summary: Your website gets visitors but few calls or messages? Check the audience, offer, contact path and follow-up before spending more on advertising. ## Answer in 60 seconds If people visit but do not call, write or book, more advertising may not be the first answer. Potential customers can disappear at five points: 1. The wrong people arrive. 2. The right people do not quickly understand what you offer. 3. They do not see enough reason to trust you. 4. Calling or sending a message is harder than it should be. 5. A genuine enquiry arrives, but nobody receives or follows it up. Check these points in order. A **useful enquiry** is a real person or business you can help asking for a relevant next step - and someone knows they must reply. ## In this guide - [1. Are the right people visiting your website?](#1-are-the-right-people-visiting-your-website) - [2. Can they understand your offer?](#2-can-they-understand-your-offer) - [3. Is it easy and reassuring to contact you?](#3-is-it-easy-and-reassuring-to-contact-you) - [4. Do genuine enquiries reach a person?](#4-do-genuine-enquiries-reach-a-person) - [5. When should you spend more on advertising?](#5-when-should-you-spend-more-on-advertising) ## 1. Are the right people visiting your website? A real visitor may still be wrong for your business: another service, location, job, free advice or price range. Place the advertisement, Google result or social post beside the page it opens. Do both describe the same service, customer, location and next step? A mismatch can attract attention that was never likely to become an enquiry. Do not jump straight to “the clicks are fake”. Google’s own guidance says a low response rate does not automatically mean invalid traffic; targeting, wording and website usability can also be responsible ([Google Ads Help](https://support.google.com/google-ads/answer/11182074?hl=en)). | What you notice | What it might mean | What to check first | |---|---|---| | Many visits, almost no contact | Wrong visitors or message | Compare the advert and page | | Calls for the wrong service | Wording is too broad | Service names and location | | Forms started, not finished | Form is difficult or intrusive | Length, questions and errors | | Messages, but slow replies | No clear owner | Inbox, alerts and responsibility | | Reported “calls”, few remembered | Clicks may count as calls | Phone records and conversations | These are places to investigate, not automatic answers. Find the first point where expectation and reality stop matching. ## 2. Can they understand your offer? Your homepage may look polished and still leave a customer unsure. It should answer: - What does this company do? - Is it for a business like mine? - Does it work in my area or market? - Why should I trust it? - What should I do next? Ask someone outside the business to view the page without your explanation, then answer those questions. Their hesitation is useful evidence. Show useful proof: a relevant example, a client name you have permission to show, a clear process, named expertise or an answer to a common concern. Give the visitor one obvious next step instead of competing buttons. This does not automatically require a redesign. The useful fix may be clearer wording, stronger proof or a more obvious contact route. ## 3. Is it easy and reassuring to contact you? Test the contact journey on a phone with details your team will recognise. Check that: - the phone number can be tapped; - the form works without zooming or guessing; - every error explains what to correct; - the form asks only for information needed for the next conversation; - a clear confirmation appears after sending; - the message reaches the place your team actually checks. Government design guidance recommends knowing why every question is present and asking only for information you need ([question-page guidance](https://design-system.service.gov.uk/patterns/question-pages/)). Removing a field does not promise a particular result. A form requesting a phone number, budget and company details without explaining what happens next can feel risky. Say who will respond, what the conversation covers and whether it creates any commitment. ## 4. Do genuine enquiries reach a person? A successful click is not the same as a conversation. Google Business Profile, for example, defines its “calls” figure as clicks on the call button - not calls answered or suitable customers found ([Google Business Profile Help](https://support.google.com/business/answer/9918094?hl=en-GB)). Forms can fail quietly. A message may reach an old employee, spam folder or unowned inbox. Run one controlled test and record: 1. when the form was sent; 2. where it arrived; 3. who received the alert; 4. who was responsible for replying; 5. whether the reply happened. If the form uses an anti-spam box, ask whether it is also checked securely after submission. Cloudflare says its browser widget alone does not protect a form ([Cloudflare Turnstile documentation](https://developers.cloudflare.com/turnstile/get-started/server-side-validation/)). Passing the check does not make someone a suitable customer. ## 5. When should you spend more on advertising? Spend more when you can answer “yes” to four questions: - Are the visitors relevant to the service and location? - Can they understand the offer and see credible proof? - Can they contact you without unnecessary effort? - Does every genuine enquiry reach a named person who responds? Then run a limited test. Count useful enquiries and real conversations - not only visits, clicks or forms. Use the same definition before and after the change. Google Ads can receive later information about which enquiries became useful outcomes ([Google Ads Help](https://support.google.com/google-ads/answer/11459091?hl=en)). Ask whether campaigns are judged on useful outcomes rather than every form. The feature does not guarantee results. ## Conclusion: check the path before buying more attention More advertising is an amplifier, not a repair. It can help when the right people understand your offer, can contact you and receive a response. It can also enlarge the same broken path. You do not need to become a specialist. Find where a potential customer disappears, fix the first meaningful break, test again, then decide whether more reach makes sense. ## Bring us the symptom. We'll give you an honest read. Bring your website, advertisement, contact route and a few recent enquiries if available. In 30 minutes, we can identify what is known, what is missing and whether the next step is a focused review, specific fix or no engagement yet. We will not promise more leads from a call. [Book a 30-minute scoping call](https://izzy.agency/en/contact/) ## Frequently asked questions ### Why do people visit my website but not contact me? They may be wrong for the business, misunderstand the offer, lack trust, struggle to contact you or receive no reply. Check in that order before considering a rebuild. ### How do I know whether my advertising or website is the problem? Compare the advertisement with its page. Irrelevant visitors point towards advertising; relevant visitors who do not contact you point towards the page; sent messages without conversations point towards delivery or follow-up. ### Could my contact form be losing enquiries? Yes. Test it from a phone and confirm the message reaches the right person. Check errors, confirmation, spam, destination inbox and responsibility. ### Should I rebuild my website to get more enquiries? Not before identifying the problem. A clearer offer, working form or reliable follow-up may be enough. Rebuild when the existing website prevents the required fix. ### When is it sensible to spend more on advertising? Spend more when relevant visitors see a clear offer, contact works, genuine enquiries reach an owner and real conversations can be counted consistently. *Research and official source links checked 13 July 2026. This article is general operational guidance, not legal advice, a diagnosis of your business or a forecast of commercial results.* --- ### OpenAI's GPT-5.6 Redefines the Frontier: Three Models, One Paradigm Shift URL: https://izzy.agency/en/blog/openai-gpt-5-6-flagship-model-launch/ Published: 2026-07-11 Summary: OpenAI launches GPT-5.6 - a family of models that trades raw scale for precision, delivering better results at lower cost across coding, knowledge work, and reasoning. On July 9, OpenAI announced the general availability of GPT-5.6 - a family of three models that signals a maturation in AI development. The headline is not raw capability; it's efficiency. GPT-5.6 Sol (the flagship), Terra (balanced), and Luna (cost-optimized) represent a shift in how frontier AI works: same intelligence, fewer tokens, lower bill. This matters beyond the OpenAI announcement. It's a signal that the age of "bigger is better" is softening. The real competition now is outcome-per-dollar. ## The Family Structure: Why Three Models Matter GPT-5.6 Sol performs at the frontier. On Agents' Last Exam - a benchmark measuring long-running professional workflows across 55 fields - Sol scores 53.6, beating Claude Fable 5 (adaptive reasoning) by 13.1 points. That's the headline metric. But the subheading is sharper: Sol achieves this while using fewer tokens and at lower estimated cost than competing models. Terra and Luna are not second-class citizens. Luna outperforms Claude Opus 4.8 at approximately one-quarter the estimated cost. For teams building tools, integrations, or internal systems, that's a constraint-removal conversation with budget owners. The three-model strategy also reflects a market maturity. Not every task needs Sol. A customer-support chatbot, a code autocomplete tool, a data-processing pipeline - these have different cost-quality tradeoffs. OpenAI is making it rational to choose the right tier rather than defaulting to maximum capability. ## What Changed: Token Efficiency and Multi-Agent Reasoning The technical shift is real. GPT-5.6 was trained to extract more useful work from every token. The result: on Artificial Analysis Intelligence Index - which spans agentic work, coding, scientific reasoning, and general capabilities - Sol with max reasoning comes within one point of Claude Fable 5 while completing tasks in 61% less time at roughly half the estimated cost. For coding work, the gains are sharper. On the Artificial Analysis Coding Agent Index, GPT-5.6 Sol sets a new state of the art (80 points), 2.8 points above Fable 5, while using less than half the output tokens and taking less than half the time. But the innovation that will shape adoption is simpler: Programmatic Tool Calling in the Responses API. Rather than every tool response passing back through the model, GPT-5.6 can now write and run lightweight programs that coordinate tools, filter intermediate data, and decide the next step. That's fewer round trips, less guidance needed, and materially lower latency for tool-heavy tasks. For parallel, complex work, GPT-5.6 introduces "ultra" - a mode that coordinates four agents in parallel by default. On demand-heavy tasks (research synthesis, complex browsing, security testing), this shift of compute for speed creates a new tradeoff: accept higher token use to finish faster. On Terminal-Bench 2.1, ultra adds parallel agents to shift the score-latency frontier upward and left. ## Design, Security, Cybersecurity: Raising the Floor GPT-5.6 Sol now reliably handles design tasks - creating interfaces from high-level direction, inspecting rendered results, catching visual and functional issues, and refining before delivery. This is the closing of a gap. For teams building internal tools or product prototypes, this is meaningful. On cybersecurity, OpenAI has added a tier: Trusted Access for Cyber. Qualified individuals and organizations in the Daybreak program get more of GPT-5.6's defensive capability (secure code review, threat modeling, blue teaming) through more precise safeguards in authorized environments. This is a pattern: frontier capability with access controls, not capability removal. The safeguards themselves are layered. OpenAI reports that GPT-5.6 Sol cyber safeguards block roughly ten times more potentially harmful activity than prior models. They also added a reasoning monitor that reviews conversation context to determine potential harm - not just classifier flags. ## Pricing and Availability: The Competitive Move Pricing is: - Sol: $5 input / $30 output per 1M tokens - Terra: $2.50 input / $15 output - Luna: $1 input / $6 output For comparison, OpenAI's own prior models and competing frontier models carry different prices. Luna's $1 input price flattens the cost curve for teams running high-volume, lower-complexity tasks. Terra's $2.50 / $15 pricing sits between cost and capability. GPT-5.6 is available now across ChatGPT (Plus, Pro, Business, Enterprise), ChatGPT Work, Codex, and the OpenAI API. The rollout began globally on July 9 and will reach full availability within 24 hours. ## Why This Matters for Technical Founders and CTOs Frontier model capability is no longer just a chatbot question. This is about infrastructure decisions. Teams now have a rational choice: which model tier suits which workload? That choice was murkier when fewer models existed or when capability gains tracked tightly with cost. GPT-5.6 also signals that the parameter-count arms race has relaxed. Better training, better inference, better reasoning architecture - these now outpace "more parameters." For teams building on top of these models, this is good news. Efficiency games are more solvable than capability games. The multi-agent default in ultra is also a subtle signal. Parallel reasoning is becoming a primitive. For teams building autonomous agents or complex reasoning pipelines, this is a platform shift. Finally, the three-model strategy with clear pricing tiers is a market maturation signal. Providers are competing on cost-quality tradeoffs, not just capability claims. That's when procurement conversations shift from "which is better" to "which fits our constraints." ## Sources - [OpenAI: GPT-5.6: Frontier intelligence that scales with your ambition](https://openai.com/index/gpt-5-6/) - [OpenAI API Documentation: Programmatic Tool Calling](https://openai.com/docs) - [Artificial Analysis Intelligence Index](https://artificialanalysis.ai) --- ### Platform Engineering in 2026: Why Your Developers Should Never Touch Infrastructure URL: https://izzy.agency/en/blog/platform-engineering-2026/ Published: 2026-06-22 Summary: Platform engineering is the fastest-growing discipline in software. Here's why abstracting infrastructure away from developers makes your team ship 3x faster. The most productive engineering teams in 2026 have something in common: their application developers never touch infrastructure. They don't write Terraform. They don't configure Kubernetes. They don't debug networking issues or manage secrets. They push code, and a platform handles the rest. This isn't magic. It's platform engineering - and it's the single biggest force multiplier we've seen in the past two years. ## What platform engineering actually is Platform engineering is the practice of building internal developer platforms - self-service systems that abstract away cloud complexity so application developers can focus on writing business logic. Instead of filing a Jira ticket to get a new environment, the developer clicks a button. Instead of learning Terraform syntax to deploy a service, they push to a branch and the platform handles provisioning, deployment, monitoring, and rollback. The result: developers spend 80% of their time writing application code instead of 40%. ## Why this matters now Three things converged in 2025–2026: **Cloud complexity exploded.** AWS alone has 200+ services. The number of decisions required to deploy a simple application has grown exponentially. **Developer experience became a retention issue.** Teams with poor developer experience lose engineers at 2x the rate of teams with strong internal platforms. **AI made application code faster to write.** With AI copilots accelerating application development, infrastructure is now the bottleneck. ## What a good internal platform looks like **Self-service environment provisioning.** Developers create environments through a simple interface - no Ops ticket required. **Automated deployment pipelines.** Push to main, deploy to staging. Merge a PR, deploy to production. Rollback with a single command. **Observability by default.** Every service gets logging, metrics, and tracing automatically. **Secret management.** Secrets are injected at runtime, never stored in code. Rotation is automated. **Cost visibility.** Each team sees their infrastructure cost in real-time. ## The build vs. buy decision For teams under 30 engineers: a thin platform layer on top of existing tools (GitHub Actions, Terraform modules, simple CLI). For teams of 30–100: a more sophisticated platform (Backstage, Humanitec, or custom portal). For teams over 100: platform engineering is a dedicated team with its own roadmap. ## What to start with **Week 1:** Automated deployment pipeline. This single change eliminates 60% of infrastructure-related developer frustration. **Week 2:** Environment provisioning. One command to create a full staging environment. **Week 3:** Observability setup. Logging and metrics for every service, injected automatically. **Week 4:** Secret management. Vault or AWS Secrets Manager, automated injection. Four weeks. That's all it takes to give your developers a platform that makes them measurably faster. ## The numbers Teams with internal developer platforms report: - 60–70% reduction in time spent on infrastructure tasks - 3x faster deployment frequency - 50% reduction in production incidents caused by configuration errors - 25% improvement in developer satisfaction scores ## The infrastructure tax you're already paying If your developers spend more than 20% of their time on infrastructure, you're paying an invisible tax on every feature they build. A $150K/year engineer spending 40% of their time on Ops work is a $60K/year Ops cost disguised as an engineering salary. Multiply by your team size. Platform engineering turns that tax into a one-time investment that pays dividends on every engineer you hire afterward. Four weeks to set up. Immediate impact on shipping speed. The ROI is measurable from week one. --- ### The True Cost of Technical Debt (And a Framework for When to Refactor, Rewrite, or Leave It) URL: https://izzy.agency/en/blog/true-cost-of-technical-debt/ Published: 2026-06-15 Summary: Technical debt compounds like interest. Here's the decision framework we use after rescuing 15+ codebases - when to refactor, when to rewrite, and when to leave it alone. Technical debt is the most expensive line item that never appears on your balance sheet. We've rescued 15+ codebases over the past seven years. Some were salvageable. Some needed a controlled demolition. The expensive lesson most companies learn too late: the decision of what to do about technical debt matters more than how you do it. Here's the framework we actually use. ## First, quantify the cost Before you decide anything, measure the damage. Technical debt has four costs that most teams only partially track: **Velocity cost.** How much slower is your team shipping compared to 12 months ago? If a feature that used to take a week now takes three, that delta is your velocity tax. We typically see 30–60% velocity degradation in codebases with significant debt. **Incident cost.** Count production incidents in the last 90 days. Multiply by the average hours spent responding. Multiply by fully loaded engineer cost. That number is usually larger than anyone expects. **Hiring cost.** Engineers talk. If your codebase has a reputation, your hiring funnel is narrower and your offers need to be higher. We have seen companies pay 15–20% salary premiums because their tech reputation preceded them. **Opportunity cost.** The features you did not build because your team was fighting fires. The market you did not enter because shipping was too slow. This is the hardest to measure and usually the largest. ## The decision matrix Once you've got the numbers, you face three options. Here's how to choose: **REFACTOR when:** - Core architecture is sound but execution is messy - Less than 40% of the codebase needs significant changes - Your team can maintain velocity while refactoring incrementally - The domain model still fits your business Refactoring is surgery, not demolition. You fix specific modules, extract concerns, add tests, and clean up interfaces - while the product keeps shipping features. This is the right call 60% of the time. **REWRITE when:** - The architecture cannot support your next 2 years of growth - The tech stack is fundamentally wrong for the problem - More than 60% of the codebase would need rewriting anyway - You have budget for 3–6 months of parallel development Rewrites are expensive and risky. The classic mistake is assuming a rewrite will take half the time it actually takes. Budget double what your estimate says. **LEAVE IT when:** - The system is stable and meeting business needs - The debt is cosmetic, not structural - The product is nearing end-of-life - You do not have the budget or bandwidth to do it properly Sometimes the right answer is "not now." ## The strangler fig pattern Our preferred approach for most rescues is the strangler fig: wrap the old system in new interfaces, migrate functionality piece by piece, and eventually turn off the legacy code. No big bang. No six-month silence followed by a scary launch day. Here's how it works in practice: 1. Identify the highest-pain module in the legacy system 2. Build a replacement behind a clean API boundary 3. Route traffic to the new module 4. Validate, monitor, confirm 5. Repeat with the next module Each cycle takes 2–4 weeks. Each cycle delivers measurable value. ## What investors actually look at If you're raising or planning an exit, your technical debt is on trial. Here's what they flag: No tests: immediate red flag. No documentation: key-person risk. Outdated dependencies: security liability. No CI/CD: means deploys are manual and terrifying. The good news: all of these are fixable. The bad news: fixing them takes time. ## Start with the audit Don't start refactoring randomly. Start with a structured audit. We look at dependency graphs, test coverage, performance bottlenecks, security vulnerabilities, deployment pipeline, and documentation. The output is a prioritized plan. Every rescue we've done started with 40 hours of audit work. It's the highest-ROI investment you can make before writing a single line of new code. Technical debt is a business problem, not an engineering problem. Quantify the cost, pick the right strategy, and start with an audit. Every week you wait, the compound interest grows. --- ### What a $40K–$200K Software Engagement Actually Looks Like URL: https://izzy.agency/en/blog/what-software-engagement-costs/ Published: 2026-06-08 Summary: A transparent breakdown of what you actually get when you hire a software studio. Phases, deliverables, timelines, and why fixed quotes beat hourly billing. Every founder asks the same question: "How much does it cost to build this?" The honest answer is always "it depends," but that's a useless answer when you're trying to budget, pitch investors, or decide whether to hire internally or hire a studio. So here's the transparent breakdown we give every prospective client - the same framework we've used across 53 products. ## The range: $40K–$200K. Here's why. A $40K engagement is typically a focused MVP: 4–6 weeks, single product, one platform, no legacy complexity. A $200K engagement is a full-scale platform build: 12–16 weeks, multiple integrations, complex architecture, production-grade infrastructure. Most of our projects land between $60K–$120K. That covers a solid product with design, engineering, deployment, and documentation. The variance comes from four factors: scope complexity, integration depth, team size, and timeline pressure. ## What you're actually paying for Here's what a typical $80K engagement includes - a SaaS product from zero to production in 8 weeks: **Week 0: Discovery & scoping ($5K–$8K value)** We tear apart your brief. User interviews if you have users. Competitive audit. Technical feasibility check. Architecture proposal. The output is a blueprint both sides sign off on before a single line of code gets written. This phase kills more bad ideas than any other. If we discover that your scope doesn't match your budget, we tell you. If we discover the problem is different from what you described, we tell you that too. **Weeks 1–2: Architecture & design ($15K–$20K value)** System design, data models, deployment strategy, CI/CD pipeline setup. In parallel, UX research, wireframes, and UI design for the core flows. You see the design and architecture before we build. Changes here cost hours. Changes in week 6 cost weeks. **Weeks 3–6: Build ($35K–$45K value)** This is where the money lives. A dedicated team - typically 2 engineers, 1 designer, 1 project lead - building in focused sprints. You see working software every Friday on a staging environment. Not a slide deck. Not a Figma prototype. Real, running software. Every pull request is reviewed. Every deploy is automated. Every decision is documented. **Weeks 7–8: Polish, test, launch ($10K–$15K value)** Performance optimization, security hardening, QA, user acceptance testing, production deployment, monitoring setup, documentation, and team handoff. When we leave, you don't inherit a codebase you can't read. You inherit a system with documentation, test coverage, and a deployment pipeline your team can run. ## What isn't included in these numbers Custom AI model training (adds $15K–$40K depending on complexity). Smart contract audits by third parties (we cover internal audits, but independent audits for DeFi protocols are separate). Ongoing hosting costs (AWS/GCP - we set it up, you pay the cloud bill). Content creation (copywriting, marketing assets). ## Fixed quotes vs. hourly billing We don't do hourly billing. Ever. Here's why: hourly billing creates a perverse incentive. The slower the agency works, the more they earn. The more meetings they schedule, the higher the invoice. You end up paying for inefficiency. Fixed quotes force us to be efficient. We scope aggressively, cut anything that doesn't serve the core metric, and deliver on a timeline we committed to. If we underestimated the effort, that's our problem, not yours. You know the total cost before we start. No surprise invoices. No "we discovered additional complexity" emails in week 4. ## How to evaluate whether a studio is worth the price Ask these questions: Do they show working software weekly, or just status updates? Do they give you a fixed quote or hourly estimate? Do they document what they build? Who owns the IP after payment? Can your team actually maintain the code they leave behind? If the answer to any of these is unsatisfying, keep looking. The cheapest studio is never the cheapest option when you factor in the rewrite six months later. A studio engagement isn't cheap - but it's dramatically cheaper than building the wrong thing with the wrong team and rewriting it six months later. Fixed quotes, weekly staging deploys, and 53 products of experience. That's what $40K–$200K buys. --- ### Smart Contracts in Production: What We Learned Shipping DeFi Protocols URL: https://izzy.agency/en/blog/smart-contracts-in-production/ Published: 2026-06-01 Summary: What separates a testnet deploy from a mainnet-proven protocol. Gas optimization, security audits, upgradability patterns, and lessons from real deployments. Deploying a smart contract to testnet takes an afternoon. Getting a protocol to mainnet - and keeping it there - takes months of engineering discipline that most teams underestimate. We've shipped DeFi protocols, token systems, and dApps to mainnet. Here's what we learned about the gap between "it works on Goerli" and "it handles $10M in TVL on mainnet without breaking." ## Lesson 1: Security isn't a phase. It's the architecture. The most expensive smart contract bugs are not in the code. They are in the design. Reentrancy, oracle manipulation, flash loan attacks - these aren't exotic edge cases. They're the standard attack surface of any DeFi protocol. If your architecture doesn't account for them from day one, no amount of auditing will save you. Our approach: every contract starts with a threat model. Before we write a single line of Solidity, we map the attack vectors. What happens if an oracle reports a stale price? What happens if a user calls this function recursively? What happens if gas prices spike 10x during execution? The threat model shapes the architecture. Not the other way around. ## Lesson 2: Gas optimization is a design decision, not a post-build task Storage layout matters. A single misplaced storage variable can cost users thousands in gas over the lifetime of a contract. We've seen protocols where a simple storage restructuring - packing related variables into single slots - cut gas costs by 30%. Here's what we optimize at the architecture level: **Storage layout.** Pack variables that are read together into the same slot. Use mappings instead of arrays when possible. Avoid string storage on-chain. **Function design.** Minimize external calls. Batch operations where possible. Use `calldata` instead of `memory` for read-only parameters. **Event patterns.** Emit events instead of storing data you only need for off-chain indexing. Events cost a fraction of storage writes. **Proxy patterns.** If the contract needs upgradeability, choose the right proxy pattern upfront. UUPS vs. Transparent Proxy vs. Diamond - each has different gas profiles and security tradeoffs. ## Lesson 3: Formal verification catches what testing misses Fuzz testing is essential. We run thousands of randomized inputs against every function. But fuzz testing is probabilistic - it finds bugs by chance, not by proof. Formal verification is deterministic. It mathematically proves that certain properties always hold. We use a combination of both. Together, they provide a confidence level that no amount of unit testing alone can match. ## Lesson 4: Independent audits are non-negotiable We audit our own contracts internally. Then we send them for independent audit. Always. Budget $30K–$80K for a thorough independent audit of a medium-complexity protocol. Yes, it's expensive. It's also significantly cheaper than the exploit it prevents. ## Lesson 5: Mainnet monitoring is not optional Deploying to mainnet isn't the finish line. It's the starting line. We set up monitoring for every protocol we deploy: transaction monitoring for unusual patterns, price oracle health checks, liquidity depth tracking, gas cost anomalies. The first 72 hours after mainnet deployment are the highest-risk period. We monitor them around the clock. ## Lesson 6: Upgradability is a tradeoff, not a feature "Make it upgradeable" is the most common request we push back on. Upgradability adds complexity, increases gas costs, and introduces a centralization risk. Our rule: if the contract manages user funds, immutability is the default. Upgradability is earned by a clear governance model, time-locked upgrades, and multi-sig controls. ## The real cost of cutting corners We've been hired to rescue protocols that skipped audits, ignored gas optimization, and deployed without monitoring. The cost of fixing the problems afterward was 3–5x what it would've cost to do it right the first time. In smart contracts, you can't push a hotfix. The code is immutable. The stakes are real money. And the users don't forgive. ---