iSeer
Back to field notes
Field notes · methodology · July 11, 2026 · 7 min read

Your AI visibility didn't drop. Or did it?

Noise, runs, and receipts — why one changed answer is an observation, not a market event, and what a confirmed drop actually requires.

IT
iSeer Team
AI visibility intelligence

A model mentioned your brand on Monday and omitted it on Tuesday.

Did visibility drop?

Possibly. It may also be ordinary run-to-run variation.

Language models are stochastic systems. The same prompt can produce a different ordering, a different shortlist or a different answer structure on another run. Any monitoring system that treats each change as a market event will manufacture drama from noise.

iSeer does not alert on every changed answer. It samples weekly, calculates statistical confidence and sends a presence alert only when a drop is confirmed by the measurement rule.

That restraint is a feature.

One run means very little.

Here is a real sequence from the iSeer ledger — the same brand and question, sampled weekly on Claude: notion.so on the question “best note taking apps”:

RunBrand present?
Run 1Present
Run 2Present
Run 3Present
Run 4Present
Run 5Present

Five for five. That is also a finding. A stable sequence like this is what earned confidence looks like: every run in the set points the same way, so it supports a presence claim with narrow uncertainty. The claim is not “the brand appeared once” — it is “the brand keeps appearing, run after run.”

But stability on one question is not the whole picture. Across the ledger, ChatGPT and Claude disagreed on 18% of matched buyer-intent questions — 143 of 795 pairs, latest report per domain, as of 11 July 2026. A brand that is rock-solid on one platform and one question can still be contested elsewhere.

If a brand appears in four runs of five and disappears in one, reporting the final run alone creates a false alarm. If it appears once in five runs, reporting the favourable run alone creates a false sense of security.

The unit of evidence is the run set, not the screenshot.

A percentage without uncertainty is incomplete.

Suppose a brand appears in 6 of 10 runs. Its observed presence rate is 60 per cent.

That is not the same as knowing that its true underlying probability of appearance is exactly 60 per cent. Ten runs are a sample. Another ten may differ.

Confidence intervals express that uncertainty. iSeer uses Wilson confidence intervals because simple normal approximations behave poorly with small samples and proportions near zero or one.

The glossary defines confidence in the product's reporting language. The full sampling and calculation rules are documented in the methodology.

The purpose is not to decorate a dashboard with statistics. It is to stop a weak sample from pretending to be a precise result.

Why Wilson intervals.

A common shortcut estimates uncertainty around a proportion using a symmetric interval. That approach can produce implausible bounds below zero or above one and tends to perform badly when the sample is small.

The Wilson interval behaves better in the conditions that matter here: limited runs, binary presence observations and rates that may be close to zero.

It also makes an uncomfortable fact visible. With a small sample, uncertainty can remain wide. The honest answer may be that the available evidence does not yet support a strong claim.

That is preferable to a confident but unstable score.

What counts as a confirmed drop.

A changed answer is an observation. A confirmed drop is a conclusion produced by a defined rule.

In iSeer, an alert is tied to the weekly sampling cycle and the confidence calculation. In the current product the presence alert is sent on Mondays at 08:00 UTC, as one email, and only when the confirmation rule is met — an answer the brand held in at least 60% of runs going to zero across a full week of runs, on both platforms.

The result should be a quiet inbox. If alerts arrive constantly, either the brand is undergoing an extraordinary sequence of changes or the system is confusing randomness with signal.

“Rare and earned” is a better standard for an alert than “technically different from last time”.

Receipts matter more than the alert.

An alert without evidence asks the recipient to trust the monitoring vendor.

A useful alert should allow the team to inspect:

This is what turns “your AI visibility fell” from a marketing message into an auditable claim.

The distinction is central to iSeer. We measured the runs. We retained them. We can show why the system did or did not call the change significant.

Noise can look persuasive.

Stochastic output is especially misleading because each individual answer is well written.

A random omission does not look random. The model still provides reasons, comparisons and a coherent shortlist. Humans are inclined to interpret that coherence as evidence of a stable ranking process.

The answer can be internally polished and externally unstable at the same time.

This is why manual spot checks produce arguments. One person has a screenshot showing the brand. Another has a screenshot showing its absence. Each screenshot is genuine. Neither establishes the underlying rate.

The companion article, ChatGPT and Claude disagree about you more than you think, adds another source of variation: the platform itself.

What teams should do.

Use a fixed set of buyer-intent questions. Run them repeatedly. Store the answers. Separate platforms. Compare periods using the same measurement rule.

Do not rewrite the strategy after one omission. Do not celebrate one inclusion. Treat both as observations until the run set supports a conclusion.

When a confirmed change appears, inspect the answer evidence before changing content. The cause may be category framing, a platform-specific gap, a crawl problem or an ordinary shift in which alternatives the model chose to mention.

The next step should follow the measured failure mode.

Measurement without theatre.

AI visibility is likely to remain variable. A good monitoring product should not conceal that fact. It should quantify it.

The aim is not to remove uncertainty from the system. It is to distinguish uncertainty from change well enough that a team knows when action is justified.

One run is a story. Repeated runs are evidence. Confidence tells you how much weight the evidence can carry.

Try it on your brand

Measure your current AI visibility.

No signup. Buyer-intent prompts across ChatGPT and Claude, with confidence bands and run-level evidence you can inspect.