All posts
Method

Why every number here comes with the room it could be wrong by

Ask a model the same question twice and you can get two different lists of brands. A product that reports a single percentage from that is reporting a coin flip with a decimal point on it.

This is the first thing anyone notices about AI visibility and the last thing most tools admit. Answers are sampled from a distribution, not looked up in an index. "Are we in the answer" is not a question with an answer; "how often are we in the answer, out of how many asks" is.

The number is the range

Every metric here ships with the number of answers behind it and a 95% confidence interval. Not as a hover tooltip for the curious — beside the figure, in the same row, because the figure is not interpretable without it. Visibility of 62% on 300 answers and visibility of 62% on nine are different claims about the world, and only one of them survives a follow-up question.

A figure has to be narrower than its own error bar

A thin window does not get a confident-looking number. Under 20 answers the dashboard dims the figure rather than drawing a trend line through it, because a trend drawn through six answers is a drawing.

Counting answers is not enough on its own, though, and for a while it was all we did. One answer in twenty-eight clears any floor you like and still reports 4% when the interval runs from 0.6% to 17.7% — the same evidence saying the brand might be absent, and saying it might be in nearly one answer in five. So the interval has to be narrower than the reading it brackets before the number is drawn as a reading. Zero is held to the same standard from the other end: nought out of twenty-eight is not “nobody names you”, it is “we have not asked enough to tell”.

What this costs, honestly

Intervals make the product look less certain than its competitors, and that is a real cost in a demo. A rival dashboard says 41%. Ours says 41% ± 6 on 300 answers. The first one is easier to put on a slide; the second one is the one that holds up when someone asks what it was last month and whether the difference means anything.

Often the honest reading of a week-over-week move is "we cannot tell yet". A dashboard without intervals is structurally incapable of saying it.

Defensible beats impressive

The people who eventually decide whether this budget survives are not impressed by a big number. They are looking for the seam where it falls apart. Handing them the sample size and the interval before they go looking is not modesty — it is the only version of the number that is still standing at the end of the meeting.

Find out what AI says about your brand

Add a domain and the first pass runs while you read the next one.