iSeer
Back to field notes
Field notes · teardown · July 13, 2026 · 7 min read

The 67-point gap.

LastPass scores 73 on ChatGPT and 6 on Claude. Same brand, same questions, same week. If you only check one model, you have measured half the channel — and possibly the wrong half.

IT
iSeer Team
AI visibility intelligence

Most AI-visibility advice treats “AI” as one thing. Ask ChatGPT what it says about you, screenshot the answer, call it a finding. We seeded twenty buyer-intent categories across ChatGPT and Claude this month, and the single most consistent result is that the two models do not agree with each other — sometimes to a degree that changes the entire story a buyer hears.

The most extreme case we measured is LastPass. On ChatGPT it scores 73 out of 100 for password-manager buyer questions. On Claude, 6.

The widest gaps we measured

Each row is one brand, measured the same week with the same questions. The score is 0–100: how often and how strongly a model named the brand when asked a buyer question in its category. A positive gap means ChatGPT rates it higher; negative means Claude does.

BrandCategoryChatGPTClaudeGap
lastpass.compassword manager736+67
klaviyo.comemail marketing054-54
mangools.comseo tools042-42
keepersecurity.compassword manager4612+34
nordpass.compassword manager397+32
brevo.comemail marketing030-30
linear.appproject management027-27
basecamp.comproject management3713+24

Why a gap is not a bug

It is tempting to read a gap as one model being “wrong.” That is the wrong frame. The two models learned from different corpora, weighted differently, cut off at different times. A gap tells you the public record about you is inconsistent — and that different buyers are therefore hearing different things.

LastPass is the clearest illustration. It is a widely-known brand with a widely-known security history. One plausible reading of a 73/6 split is that one model surfaces the brand on name recognition while the other weights the reputational record more heavily. We want to be careful here: we measured the gap, we did not measure the reason. Anyone claiming to know exactly why a model down-ranks a brand is guessing.

We can prove the disagreement exists and how large it is. We cannot prove what is going on inside the model. Those are different claims, and conflating them is how AI-visibility advice turns into astrology.

The zero problem

Four brands in the table score exactly 0 on one platform and something real on the other. Klaviyo: 0 on ChatGPT, 54 on Claude. Linear: 0 on ChatGPT, 27 on Claude. Mangools: 0 and 42.

A zero is the result that most deserves suspicion, because a broken API key produces exactly the same number as genuine absence. That is why every batch we run starts with a control scan: we measure a brand that is unambiguously well-known in its category first, and if that comes back at zero we throw the entire run away rather than publish it. For this dataset the control returned 78 out of 100 — so the zeros above are real absence, not a dead key.

That distinction is the whole ballgame. Without a control, “we found your brand is invisible” and “our scanner was broken” are indistinguishable, and only one of them is worth paying for.

What to do if you have a gap

Check both. That is the boring, correct answer, and it is the reason this product measures two platforms rather than one. If your visibility differs by 30, 40, 60 points across models, an optimisation aimed at the model you happened to check may do nothing for the one your buyer actually opened.

Then ask what the models are reading. Neither of them is crawling your homepage and grading your copy. They are synthesising what everyone else has written about you — reviews, comparisons, forum threads, tutorials. A gap usually means that body of writing is thin, contradictory, or weighted toward a story you would not choose.

Method, so you can argue with it

Twenty categories, eight brands each, measured 2026-07-13 against ChatGPT and Claude with buyer-intent questions. Every question is asked repeatedly rather than once, because a single answer from a stochastic model is an anecdote, not a measurement. Scores are 0–100 and reflect how often and how prominently a brand is named. A control scan gates every batch. Perplexity is not currently measured.

These numbers are a snapshot of one week, not a permanent ranking. They will move. That is the point of measuring them on a schedule rather than screenshotting them once.

See your own gap

Run a free report and see how ChatGPT and Claude each answer for your brand. No signup, both platforms, about thirty seconds.