# Recomma > Recomma measures how ChatGPT, Gemini, Google AI Overviews and AI Mode describe your brand — captured from the answers readers see, not from the APIs — and turns every gap into work you can ship and verify. Recomma (recomma.ai) captures the answers ChatGPT, Gemini, Google AI Overviews and AI Mode give to the questions a brand's buyers ask — from the interfaces readers use, not the APIs — records which brands and sources each answer names, prices the market behind the questions, audits the brand's own site for the faults an assistant trips on, and reports the visits assistants send. ## Pages - [Home](https://recomma.ai/): What Recomma measures and how. - [Pricing](https://recomma.ai/pricing): Starter, Growth and Scale plans, and the managed services. - [Blog](https://recomma.ai/blog): Notes on measuring what AI says. - [Fix what AI search gets wrong about you, from Claude Code (MCP)](https://recomma.ai/blog/ai-visibility-mcp-claude-code): Connect Claude Code to Recomma over MCP, read the site audit's findings, fix the template, close the ticket — and let the re-measure say if it worked. - [The GEO audit checklist: 15 markup rules AI crawlers trip on](https://recomma.ai/blog/geo-audit-checklist): The fifteen markup rules our site audit runs on every page, in queue order: what each checks, why an engine cares, and the fix that fits in a diff. - [What is AI search visibility? The metrics, defined](https://recomma.ai/blog/what-is-ai-search-visibility): The share of AI answers that name your brand, plus share of voice, position and sentiment — each defined the way it is counted, and over what. - [AI visibility vs SEO: the same keywords, a different question](https://recomma.ai/blog/ai-visibility-vs-seo): Same keyword data, different result: an answer that names a few brands, not a list of pages. What carries over from SEO, and what does not. - [What sampling a model can and cannot see](https://recomma.ai/blog/what-sampling-can-see): How an AI answer becomes a number: captured from the interface a reader sees, sources kept as shown, and 'named' kept apart from 'cited'. - [Every gap becomes a ticket, and every ticket gets a verdict](https://recomma.ai/blog/gaps-become-tickets): Closing a ticket freezes a baseline; fourteen days later the same prompts are re-measured on fresh answers. The verdict is kept even when nothing moved. - [Why every number here comes with the room it could be wrong by](https://recomma.ai/blog/confidence-intervals): A visibility percentage without a sample size is a coin flip with a decimal point. Every metric here carries the answers behind it and a 95% Wilson interval. - [Compare](https://recomma.ai/compare): Recomma beside the other AI search visibility tools. - [Recomma vs Otterly.AI (2026): which AI visibility tool fits you](https://recomma.ai/vs/otterly): Recomma and Otterly.AI side by side: engines, prices with add-ons, cadence, API and MCP, and how each turns answers into numbers. Facts linked. - [Otterly.AI alternatives (2026): 8 AI visibility tools compared](https://recomma.ai/alternatives/otterly): Eight AI visibility tools to weigh against Otterly.AI — prices, engines, cadence and who each suits, from their own pricing pages. - [Recomma vs Peec AI (2026): plans, engines and method compared](https://recomma.ai/vs/peec-ai): Recomma and Peec AI side by side: plans, prompt allowances, engines, API and MCP access, and what each does after the measurement. - [Profound alternatives (2026): self-serve AI visibility tools](https://recomma.ai/alternatives/profound): Profound sells brand tracking by contract. Self-serve AI visibility tools with published prices, compared on engines, cadence and access. - [Privacy policy](https://recomma.ai/privacy) - [Terms of service](https://recomma.ai/terms) - [Everything above, in full](https://recomma.ai/llms-full.txt): this file with every post's text inline. ## Documentation - [Docs](https://docs.recomma.ai/): every concept and metric, defined. - [Docs for language models](https://docs.recomma.ai/llms.txt) ## Posts ## Fix what AI search gets wrong about you, from Claude Code (MCP) https://recomma.ai/blog/ai-visibility-mcp-claude-code · 2026-10-10 A dashboard can tell you that three hundred of your pages point their canonical somewhere else. It cannot open the template and fix it. An assistant running in your repository can do both, if it can read what the dashboard knows. That is what our MCP server is for, and this is the loop we use it for: read the findings, change the code, mark the work done, and let the re-measure say whether it moved anything. On this page- [What the server is](#what-the-server-is) - [Connect your client](#connect) - [The tools, read and write](#the-tools) - [The loop: findings to verdict](#the-loop) - [What a write cannot do](#guard-rails) - [Ready-made workflows](#ready-made-workflows) ### What the server is MCP — the Model Context Protocol — is the standard way for an assistant such as Claude, ChatGPT or Codex to call tools outside itself. Ours lives at `https://api.recomma.ai/mcp`. Every tool on it calls the same procedures the product's own screens call, so a connection sees exactly what the person who made it sees, in the workspace they have open, and nothing from any other. Reading is what a connection gets by default. Changing anything is a second permission, granted separately, and every change previews before it applies. Two pieces of reference ride along with the tools — a glossary of every term the figures use, and the method behind the market-value numbers — so an assistant can look up what "visibility" or "sample" means here before it explains a figure to you, rather than guessing. It is not the only MCP server in this category, and it does not need to be. What it is built around is the part of the product that ends in a change you can make: the site audit, whose findings come with the markup as it is and as it should be, and the action queue, whose tickets are re-measured after they close. ### Connect your client The shortest path is in the app, under Settings → Connect AI: pick a client and it gives you that client's one-step install. For reference, here is what each one is. How each client connects. All but n8n sign in through the browser; n8n and custom agents use a key made on the same page.ClientHow it connectsClaudeSettings → Connectors on claude.aiChatGPTSettings → Connectors on chatgpt.comClaude CodeOne command in the terminal (below)Codex CLI`codex mcp add recomma --url … && codex mcp login recomma`Gemini CLIAn httpUrl entry in ~/.gemini/settings.jsonVS CodeA one-click install link, for Copilot's agent modeWindsurfA serverUrl entry in ~/.codeium/windsurf/mcp_config.jsonn8nAn MCP Client Tool node, HTTP Streamable, Bearer auth with a keyFor Claude Code, which is where the rest of this post happens: `claude mcp add --transport http --scope user recomma https://api.recomma.ai/mcp`The first call opens a browser. You sign in as usual, and a consent screen says what is being connected and what it may read. Approve it and the client has a token; the connection then appears under Connected apps on the same settings page, where disconnecting it stops its very next call. If you belong to several workspaces, the connection follows whichever one you last had open in the app. ### The tools, read and write Nineteen tools in all. Fifteen read; four write. The reads cover the same ground as the product's pages: QuestionToolsWhich brands can I see?`list_brands`How often am I named, and which way is it going?`get_visibility`, `get_trend`Which questions am I losing, and what did the model say?`list_prompts`, `get_prompt`, `read_answers`Which sites do the answers lean on?`list_sources`, `get_source`What should I work on, and did it work?`list_actions`, `get_action`, `get_impact`What is wrong with my own pages?`get_site_findings`What is the category worth, and who is it sending me?`get_market`, `get_keywords`, `get_referrals`The four writes are `create_prompts` and `suggest_prompts`, which propose new questions to track; `set_prompt_status`, which starts or pauses sampling them; and `set_action_status`, which moves a ticket on the queue. Every figure a read returns carries the number of answers behind it, and a question nothing has sampled yet comes back as null rather than zero — "we have not looked" and "they never name you" are opposite findings, and an assistant should not be able to confuse them. ### The loop: findings to verdict Open Claude Code in the repository your site is built from, and work through it in this order. - Find the brand. `list_brands` returns the id every other tool takes, and whether this connection may write. - Read what fails. `get_site_findings` with no rule lists every audit rule the site fails, with how many pages each touches and how much of the site has been read so far. A partial crawl is a true statement about the pages read, and the tool says how many. - Pick one rule and get the diff. Called again with `rule`, it returns the failing pages with `before` and `after`: the markup as it is and as it should be, plus the steps to get there. Where `after` is null, the fix needs a person to write something, and the assistant should say so rather than invent it. - Fix the template, not the page. A rule failing on a large share of pages is nearly always one line in a layout file. Ask the assistant to find it, change it, and show you the diff before you ship. - Close the ticket. `list_actions` with `kinds: ["owned_audit"]` finds the ticket for that rule. `set_action_status` moves it to `in_progress` while the work happens and to `done` once it is live — previewing first, then applying. - Wait for the verdict. Marking done freezes the current reading on the questions the ticket targets and schedules a re-measure two weeks later. `get_impact` reports it as pending until then, and as a verdict after — kept whichever way it went. The fourth step is the one a dashboard could never take, and the sixth is the one that keeps the rest honest. A fix that changed nothing is worth knowing about, and a queue that only ever reports progress is a queue nobody should believe. We wrote about why the verdict is kept in [every gap becomes a ticket](/blog/gaps-become-tickets). ### What a write cannot do An assistant in a loop can call a tool twenty times in thirty seconds, so the writes are built for that caller: - Writes need the recomma:write scope, which a client asks for separately. A connection that may only read and calls a write gets a challenge naming the missing scope — which a client that supports it turns into a second approval, not a failure. - Every write previews by default and changes nothing until it is called again with apply: true. - New questions land as suggested. Nothing is sampled, and nothing spent, until a person accepts them on the Prompts page. - Activating questions spends the plan's allowance, so the preview says how many runs a week it would start — and a call that would commit more than half of what the plan has left is refused and left to a person. - A ticket already done is never re-marked or reopened: that would overwrite or discard the reading frozen when it closed. That call belongs to a person, on the Impact page. - Nothing deletes. ### Ready-made workflows Clients that support MCP prompts will also list five workflows, each a sequence of the calls above with the rules for reading what comes back: WorkflowWhat it produces`weekly_brief`What moved this week, where the brand is losing, which sites the answers lean on, what is on the queue`visibility_drop`Which engine and which questions a fall came from, and what the answers say instead`competitor_wins`Who leads, on which questions, and which sites carry them without mentioning you`prompt_set_audit`Self-flattering, duplicate or untopiced questions, and priced demand nothing asks about`site_fix_plan`Every failing audit rule, template faults first, with the markup to change on each page`site_fix_plan` is the loop above written down. It reads and reports and edits nothing unless you ask — and when you do, it starts with the template faults. The rules it works from are the fifteen in our [GEO audit checklist](/blog/geo-audit-checklist). **Where this lives.** The full reference — both ways to authenticate, keys for agents without a browser, and every tool's inputs — is in the [MCP server docs](https://docs.recomma.ai/reference/mcp). ## The GEO audit checklist: 15 markup rules AI crawlers trip on https://recomma.ai/blog/geo-audit-checklist · 2026-10-10 Before an engine can name a page in an answer, it has to fetch the page, decide what it is about, cut it into passages and pick one to quote. Every step can go wrong in the markup, and most of the ways it goes wrong can be seen by pressing Ctrl-U. These are the fifteen rules our site audit runs on every page it reads, in the order its queue puts them. On this page- [How the list was drawn](#how-the-list-was-drawn) - [Rule 1: the content has to be in the HTML](#content-in-the-html) - [Rules 2–7: one address, one subject, one shape](#one-address-one-subject) - [Rules 8–10: dates an engine can read](#dates-an-engine-can-read) - [Rules 11–13: say what kind of page it is](#say-what-the-page-is) - [Rules 14–15: what a crawler cannot see](#what-a-crawler-cannot-see) - [The checklist at a glance](#the-checklist) - [Held to the same rules](#held-to-the-same-rules) ### How the list was drawn A rule is on this list only if two things are true of it: a person can check it from the page source, and its fix fits in a diff. That rules out a lot of real advice. "Improve the introduction" is sound and cannot be decided from a shape; "move the author bio into an aside" is a matter of taste. A finding nobody can verify costs more trust than it is worth, so neither is here. The rules are also written to be fixed once. A finding is grouped by rule, not by page — "add alt text to the images that have none" is one job whether it touches one page or forty — because a fault that shows up on forty pages almost always lives in one template. Each rule carries a severity from one to five, and the order below is the order a queue should be worked in: a page whose content is not in its HTML has no headings worth fixing yet. ### Rule 1: the content has to be in the HTML Serve the content in the HTML. Severity 5. The check fires when a page answers with at least 20,000 bytes of markup and fewer than 120 words of text: the content exists, it is just being assembled in the browser. Most of what feeds these engines reads the HTML directly and does not run scripts, so to them the page is empty — and nothing else on this list matters for a page in that state. Check it with JavaScript off, or with `curl`; move the body copy into the server-rendered document. The interactive parts can stay client-rendered. The prose cannot. ### Rules 2–7: one address, one subject, one shape 2. Point each canonical at the page itself. Severity 5, the joint highest. A canonical tag names the real address of a page, and everything the page earns is credited there. When it points somewhere else — often a template hard-coding the home page or a staging host — the engines quote the page under another name, sometimes on a site you do not control. The fix is to build the tag from the page's own address, then check a few sibling pages, because a wrong canonical is almost always a template. ` `3. Declare a canonical on every page. Severity 3. Without one, every address that reaches the page — with a tracking parameter, with and without a trailing slash, through a staging host — is a separate page to a crawler, and the citations split across the copies. Make it absolute, not a path, and keep your analytics parameters out of it. 4. Mark where the content begins and ends. Severity 3. A `main` or `article` element tells an extractor where the page's own content stops. Without either, the boundary is guessed, and the guess usually takes the navigation, the cookie banner and the footer along — so the passage that gets quoted is padded with your menu. Either element passes; keep the navigation, sidebar and footer outside it. 5. Give every page an H1. Severity 3. With no H1, the subject is inferred from whatever the chunker reaches first — the title tag, the URL, or, on templates where the visible title is a styled `div`, the navigation. Mark the text a reader would call the title as an H1. 6. Cut each page down to one H1. Severity 2. Two H1s are two answers to "what is this page", and the one that wins is usually the first — on most templates, the site name. Keep the heading that names the subject and demote the rest to H2; the styling can stay identical. It is the cheapest fix on the list. 7. Step heading levels down one at a time. Severity 2, checked on pages with three or more headings. Levels are how a document says which section contains which. An H2 followed by an H4 breaks that, and a chunker will attach the orphaned section to the wrong parent or to nothing — so the quoted passage arrives without the context that made it correct. If a level was skipped to get a smaller font, change the font. `

Pricing

What the trial includes

Pricing

What the trial includes

` ### Rules 8–10: dates an engine can read 8. Publish a machine-readable publish date. Severity 3, on written pages only — articles, comparisons, alternatives pages, how-tos and lists, and not reference documentation, where an undated page is not a stale one. "September 2026" in a paragraph is a date only a reader can use. When an answer has to choose between two pages that disagree, recency is one of the few tie-breakers it has, and an undated page does not get to compete on it. Wrap the date in a `time` element with an ISO `datetime`, or put `datePublished` in the structured data. ``9. Add a last-updated date where there is a publish date. Severity 2. A page that declares when it was published and never says whether it has been touched since looks the same as one abandoned three years ago. Add `dateModified` beside `datePublished` and wire it to the content's real last edit — not the build date. A nightly build that stamps today on everything is worse than no date at all. 10. Keep the year in a time-sensitive title current. Severity 4, on comparisons, alternatives pages, how-tos and lists. A title that says "best tools of" a year that has passed is now a statement against you: a model choosing between two lists has been handed the reason to pick the other one. Bring the content up to date first — changing the year alone makes the title a false claim — then the title, the H1 and the description. Keep the URL; a new slug throws away the links the page has earned. ### Rules 11–13: say what kind of page it is 11. Add the Open Graph type. Severity 2. `og:type` is the one Open Graph field that says what kind of thing a page is rather than what it looks like. Missing, it defaults to "website" everywhere the page is unfurled, so a substantial post is indexed with the same shape as your home page. Use `article` for posts and `website` for the marketing pages. 12. Mark up the FAQs you already have. Severity 3, checked when a page has three or more headings phrased as questions and no `FAQPage` schema. The page is already an FAQ; the schema pairs each question with its answer explicitly instead of leaving them to be paired by proximity. Mark up only questions the page genuinely answers. 13. Add a contents list to long pages. Severity 2, checked on pages over 1,200 words with four or more H2 sections and fewer than three in-page links. A long page with no contents has no addressable parts: an answer that wants to point at one section can only point at the whole page, and the reader it sends lands at the top. Give each H2 an `id`, link them at the top, and keep the list in the HTML rather than building it in a script — the one at the top of this post is there because this rule asked for it. ### Rules 14–15: what a crawler cannot see 14. Add alt text to the images that have none. Severity 2. Nothing that reads the page can see its images, so alt text is the only description of them in the HTML. A diagram carrying the point of a section is, downstream, a blank. Write one sentence for each image that carries meaning, describing what it shows rather than naming the file, and give purely decorative images an empty `alt`, which is the correct answer for them. 15. Say who each quote came from. Severity 1, checked when two or more quotes on a page name nobody. A blockquote without attribution in the markup is a sentence that appears to be yours, so when it is quoted onward it is quoted as your claim — losing the credibility of a third party saying it, which is why the quote is on the page. Put a `cite` inside each blockquote; for a customer, name the company as well as the person. ### The checklist at a glance The fifteen page rules in queue order, with the severity the audit gives each (5 is the most severe).RuleFails whenSeverity1. Content in the HTML≥20 KB of markup, <120 words52. Canonical points at itselfIt names another address53. Canonical declaredNo canonical tag34. Content boundaryNeither main nor article35. An H1No H1 at all36. One H1Two or more H1s27. Heading orderA level is skipped28. Publish dateNone machine-readable39. Updated datedatePublished, no dateModified210. Current year in titleTitle names a past year411. og:typeTag missing212. FAQ schema≥3 question headings, no FAQPage313. Contents list>1,200 words, ≥4 H2s, <3 anchors214. Image alt textAny image without alt215. Quote attribution≥2 quotes naming nobody1 ### Held to the same rules It would be an odd thing to sell an audit from a site that failed it. Every page on recomma.ai that a crawler can reach is rendered in our test suite exactly as the build renders it and run through these same fifteen rules; a page that breaks one fails the suite. This post is one of those pages. In the product, the rules run across every page in your sitemap, not only the ones somebody has already quoted — the pages with problems are usually the ones nothing has cited yet. [The Site page](https://docs.recomma.ai/product/site) shows what was read and which rules fail, with each failing page's markup as it is and as it should be, and a rule that touches enough of the site becomes a ticket on the same queue as everything else, re-measured fourteen days after it closes. The [plans](/pricing) say what each tier tracks. **What this list leaves out.** Whether AI crawlers are allowed in at all is a site-wide question, not a page one, and the Site page reads it separately from your robots.txt. A page that passes all fifteen rules behind a robots.txt that shuts the retrieval crawlers out is still a page no answer will quote. ## What is AI search visibility? The metrics, defined https://recomma.ai/blog/what-is-ai-search-visibility · 2026-09-27 AI search visibility is the share of AI answers, to the questions your buyers ask, that name your brand. It is measured over answers, not pages, and it comes with four other figures that say how you were named. Here is each one, defined the way we count it. The phrase gets used loosely. Some tools mean "the brand appears somewhere in ChatGPT", others mean "a page of yours is cited", and a few mean a score they have invented. The definitions below are the ones this product prints beside every number, and they are the same ones the other serious tools in the category use — so a figure from here can be compared with a figure from elsewhere. ### Visibility The share of chats that name the brand at least once. A chat naming it three times still counts once; this measures reach, not repetition. A "chat" is one question, asked of one engine, on one day, in one country — so visibility is always stated over a set of chats, and the set is named. ### Share of voice Your share of every brand mention in those chats. Visibility asks how often you appear; share of voice asks how much of the conversation is yours once everyone is counted. A brand can be visible in most chats and still hold a small share of voice, because the answer names six rivals beside it. ### Position Where you are named among the brands in a chat, averaged over the chats that name you. First is #1. Being named late in a list is weaker than being named first, and position is the only figure that says so. ### Sentiment How favourably a chat speaks about the brand, on a scale from 0 to 100: 0 hostile, 50 a neutral mention, 100 a strong endorsement. Averaged over the chats that named it. A chat that could not be scored is shown as a dash, never as 50. ### Named, cited, retrieved, used These are four different things, and a definition of visibility that blurs them is not one. A brand is _named_ when the answer says its name. A page is _cited_ when the answer links to it. A source is _retrieved_ when the engine opened it while answering, whether or not it went on to quote it, and _used_ when the chat cited it at least once. Retrieved is broader than used: a model can read a page and quote nothing from it, and that gap — pages the engines read and did not cite — is where most of the work on your own site lives. ### What every figure is measured over Sample size is the number of chats behind a figure: prompts times engines times sampling days in the window. Below twenty, a percentage moves too much on a single answer to be reported as a reading, and it is greyed out. Every metric ships with a 95% Wilson interval over its sample, so a reading from thirty answers cannot be mistaken for one from three hundred. ### Where the answers come from From the interface a reader uses, not the model's API. The two share a model name and little else: the API runs no retrieval, applies no system prompt, and does not know where the reader is. Asked the same twelve questions both ways, the two paths named brands at similar rates and cited almost nothing in common — 26 of the 35 sources shown to a reader never appeared in the API's answer. Visibility measured over API answers is visibility in a product nobody uses. Each figure has its full definition, with the edge cases, in the docs: [visibility](https://docs.recomma.ai/metrics/visibility), [share of voice](https://docs.recomma.ai/metrics/share-of-voice), [position](https://docs.recomma.ai/metrics/position) and [sentiment](https://docs.recomma.ai/metrics/sentiment). **What it is not.** A rank. There is no position on a results page to track, because there is no results page. An answer names a handful of brands and cites a handful of pages, and the questions are whether you are among them, how you are described, and which pages the engine leaned on to decide. ## AI visibility vs SEO: the same keywords, a different question https://recomma.ai/blog/ai-visibility-vs-seo · 2026-09-24 Search engine optimisation is the work of getting a page to rank for a query. AI visibility is the work of getting a brand named in the answer to a question. They start from the same keyword data and end in different places, and most of what a team knows from one carries over to the other — but not all of it, and the parts that do not are the parts that cost money. ### The result is an answer, not a list A search result is ten links and a reader who picks one. An AI answer is a paragraph that names three to five brands, cites a few pages, and is read as a recommendation. There is no position two; there is named or not named, and, if named, how. That is why the figures are visibility, share of voice, position within the answer and sentiment, rather than a rank. ### What carries over: the keyword data The same keywords, volumes, bids and intents that rank tracking runs on still describe the market — what people are asking about, and what a buyer of each intent is worth. We use exactly that data, but to price the questions rather than to track positions: a keyword is worth its monthly searches, times the share of clicks a top result takes, times what advertisers bid for it. Summed over a category, that is what the market pays to own the intents a brand competes on, and it is the denominator every visibility figure needs. A brand can be highly visible on a corner of its category worth very little. ### What carries over: the technical half An engine has to be allowed in, has to read the page correctly, and has to be able to quote it. Each of those is an old SEO discipline with a new reader: - robots.txt, checked against the AI crawlers by name — GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest — rather than against Googlebot alone. A site that shuts a retrieval crawler out of one path has made a decision; a site that shuts it out everywhere has opted out of the answer. - Markup a chunker can follow: one H1, heading levels that step down one at a time, an article element that says where the content begins and ends, a canonical that points at the page itself, a machine-readable publish date. Fifteen such rules, run across every page in the sitemap. - Content served in the HTML, not assembled by a script after the page loads. Most AI crawlers do not run scripts; a page that is empty until they do is empty to them. ### What does not carry over Rank tracking. Nothing here is derived from a search position, because the thing being measured is an answer. Backlink counting, too: what matters is which sources the engine actually leaned on for this question, and those are read off the answer, not off a link graph. And the sources are not the ones a search result would show. Over the answers we hold, the pages engines cite split into a few kinds — editorial, reference, community threads, other companies' sites, and the brand's own — and the mix is different for every category. Being absent from the roundup an engine keeps citing is a gap; being absent from a directory it never reads is not. ### How the work is different Ranking work is measured by movement on a results page. Visibility work is measured by re-asking the question. A gap — a source that feeds the answer and skips you — becomes a ticket; when the ticket is closed, the reading on the prompts it targets is frozen, and fourteen days later the same prompts are asked again on fresh answers only. The verdict is recorded whichever way it goes. Without that, the report is a percentage nobody can act on, and the meeting is an argument about whether the number is real. **In one line.** SEO asks where your page ranks for a query. AI visibility asks whether the answer names you, how, and what the engine read to decide — and prices the question with the same keyword data SEO already has. ## What sampling a model can and cannot see https://recomma.ai/blog/what-sampling-can-see · updated 2026-10-05 **Rewritten October 2026.** The first version of this piece described a sampler that asked the models' APIs. It now captures each answer from the interface a reader uses, and keeps the API as a fallback; the sections below describe it as it runs today. Every measurement product has a boundary, and most of them keep it in a footnote. Ours is worth reading before you trust a single number on the dashboard, so here it is in full. An AI answer is not a document sitting on a server that we fetch. It is generated, once, for whoever asked. To measure it at all you have to ask the question yourself — and every decision about how you ask changes what you can honestly claim afterwards. ### We capture the interface, not the API Each question is put to the product a reader uses — chatgpt.com, Gemini, Google's AI Overviews and AI Mode — in the reader's country, and the answer is stored as it was shown: the text, and the sources the interface attached to it. The model's API is used only when the interface cannot be reached, and every stored answer records which of the two it came from. The two share a model name and little else. The API runs no retrieval, applies no system prompt, and does not know where the reader is. Asked the same twelve questions both ways, the two named brands at similar rates and cited almost nothing in common: of the 35 sources shown to a reader, 26 never appeared in the API's answer. A capture is still not your screen. It carries no account history and no memory of earlier conversations, so an answer shaped by one reader's past can differ from ours. What it closes is the larger gap — the retrieval, and the sources that come with it. The [docs on fidelity](https://docs.recomma.ai/reference/models#fidelity-interface-or-api) say how each answer was obtained, and [how sampling works](https://docs.recomma.ai/reference/sampling) walks through one cycle. ### Sources are the ones the interface showed A captured answer arrives with its own list of sources — the pages the interface linked beside or beneath the text — and that list is what counts as cited. Each address is reduced to a canonical form first, because one answer will happily cite the same page as both `https://x.com` and `https://www.x.com/`, and counting that as two citations of two sources is wrong twice over. The first version of this sampler parsed links out of the answer text instead, and missed every source an engine shows in its own panel rather than writing into the prose. Reading the interface's list is what removed that blind spot. ### Named and cited are two different lists An answer that links your domain without ever writing your name has cited you; it has not mentioned you. We keep those apart, down to the detail of ignoring brand names that appear only inside a URL — the characters around "acme" in `https://www.acme.com/` are dots, so a naive match reads the domain as prose and credits you with a mention nobody made. > A brand the answer only linked to is on the citation list. A brand the answer recommended by name is on the other one. The difference is the whole point of the measurement. ### The engines we sample, and the ones we don't ChatGPT, Gemini, AI Overviews and Google AI Mode are sampled on every pass, and the product records which engine and which market produced each answer it stores. Perplexity is not part of the standard pass. Claude is not sampled at all: it has no consumer search interface to capture, and an answer from its API is a measurement of something no buyer reads. ### Nothing before the day you start There is no archive of what ChatGPT said about your category last March. Nobody has one — not us, not the tools that imply otherwise. Tracking begins when a prompt is added, which is the least convenient honest answer and the only one available. **Where this lives in the product.** Every page that shows a number also shows which model, which market and how many answers it came from. The boundary is not a disclaimer page; it travels with the figure. ## Every gap becomes a ticket, and every ticket gets a verdict https://recomma.ai/blog/gaps-become-tickets · 2026-07-29 Two quarters of reporting AI visibility, and nothing in the company changed because of it. That is the normal outcome for a dashboard, and it is what the action queue exists to break. A number tells you where you stand. It does not tell you what to do on Monday, and it will not tell you in six weeks whether the thing you did on Monday worked. Both halves of that are fixable, and they are the same fix: make the gap a ticket, and make the ticket carry its own proof. ### A gap is a source, not a feeling Answers cite things. Some of the sites feeding an answer about your category already mention you; some feed it and skip you entirely. The second list is the work. It is concrete — a roundup you are absent from, a directory listing nobody claimed, a thread where the question gets answered without you — and each entry gets scored by how much of the answer traffic actually runs through it. Advice that cannot be acted on is filtered out before you see it. A competitor's own domain is never a ticket: you cannot publish there, and "cover what they cover, better" is not a task anyone can pick up. ### Marking it done freezes a baseline When a ticket is closed, the current reading of the prompts it targets is frozen at that moment. Not the project average — those prompts, that day. Without a snapshot taken at the moment of the change there is nothing honest to compare against later, and the comparison quietly becomes "did anything at all happen this quarter". Fourteen days later the same prompts are measured again, using only answers recorded since the freeze. Reusing the baseline's own runs would compare a period with itself and report progress that is arithmetic, not movement. **The mechanism.** Done → baseline captured → re-measure after 14 days on fresh answers only → verdict written to the ticket. The window is fixed so it cannot be chosen after the fact to flatter the result. ### Including the tickets that did nothing The verdict is recorded whichever way it goes. A ticket that moved nothing is the most useful thing this product can tell a team — it is the difference between a playbook and a superstition — and hiding it would make every other number in the account worth less. > A tool that only ever shows you wins is not measuring your work. It is measuring its own marketing. ### What it changes about the meeting - The report is a list of shipped tickets with verdicts, not a percentage nobody can act on. - Arguments move from whether the number is real to which ticket to pick up next. - Tactics that keep failing get dropped, because there is a record of them failing. None of this makes the work easier. It makes it accountable, which is the part that was missing. ## Why every number here comes with the room it could be wrong by https://recomma.ai/blog/confidence-intervals · 2026-07-08 Ask a model the same question twice and you can get two different lists of brands. A product that reports a single percentage from that is reporting a coin flip with a decimal point on it. This is the first thing anyone notices about AI visibility and the last thing most tools admit. Answers are sampled from a distribution, not looked up in an index. "Are we in the answer" is not a question with an answer; "how often are we in the answer, out of how many asks" is. ### The number is the range Every metric here ships with the number of answers behind it and a 95% confidence interval. Not as a hover tooltip for the curious — beside the figure, in the same row, because the figure is not interpretable without it. Visibility of 62% on 300 answers and visibility of 62% on nine are different claims about the world, and only one of them survives a follow-up question. **Why Wilson, not the textbook interval.** Visibility clusters at the ends of the scale, and the normal approximation falls apart there — a brand named in all eight answers gets an upper bound above 100% and a lower bound no data supports. The Wilson score interval stays inside the range that exists, which is exactly the case this metric lives in. ### A figure has to be narrower than its own error bar A thin window does not get a confident-looking number. Under 20 answers the dashboard dims the figure rather than drawing a trend line through it, because a trend drawn through six answers is a drawing. Counting answers is not enough on its own, though, and for a while it was all we did. One answer in twenty-eight clears any floor you like and still reports 4% when the interval runs from 0.6% to 17.7% — the same evidence saying the brand might be absent, and saying it might be in nearly one answer in five. So the interval has to be narrower than the reading it brackets before the number is drawn as a reading. Zero is held to the same standard from the other end: nought out of twenty-eight is not “nobody names you”, it is “we have not asked enough to tell”. ### What this costs, honestly Intervals make the product look less certain than its competitors, and that is a real cost in a demo. A rival dashboard says 41%. Ours says 41% ± 6 on 300 answers. The first one is easier to put on a slide; the second one is the one that holds up when someone asks what it was last month and whether the difference means anything. > Often the honest reading of a week-over-week move is "we cannot tell yet". A dashboard without intervals is structurally incapable of saying it. ### Defensible beats impressive The people who eventually decide whether this budget survives are not impressed by a big number. They are looking for the seam where it falls apart. Handing them the sample size and the interval before they go looking is not modesty — it is the only version of the number that is still standing at the end of the meeting.