The 67-point gap.
LastPass scores 73 on ChatGPT and 6 on Claude. Same brand, same questions, same week. If you only check one model, you have measured half the channel — and possibly the wrong half.
Most AI-visibility advice treats “AI” as one thing. Ask ChatGPT what it says about you, screenshot the answer, call it a finding. We seeded twenty buyer-intent categories across ChatGPT and Claude this month, and the single most consistent result is that the two models do not agree with each other — sometimes to a degree that changes the entire story a buyer hears.
The most extreme case we measured is LastPass. On ChatGPT it scores 73 out of 100 for password-manager buyer questions. On Claude, 6.
The widest gaps we measured
Each row is one brand, measured the same week with the same questions. The score is 0–100: how often and how strongly a model named the brand when asked a buyer question in its category. A positive gap means ChatGPT rates it higher; negative means Claude does.
| Brand | Category | ChatGPT | Claude | Gap |
|---|---|---|---|---|
| lastpass.com | password manager | 73 | 6 | +67 |
| klaviyo.com | email marketing | 0 | 54 | -54 |
| mangools.com | seo tools | 0 | 42 | -42 |
| keepersecurity.com | password manager | 46 | 12 | +34 |
| nordpass.com | password manager | 39 | 7 | +32 |
| brevo.com | email marketing | 0 | 30 | -30 |
| linear.app | project management | 0 | 27 | -27 |
| basecamp.com | project management | 37 | 13 | +24 |
Why a gap is not a bug
It is tempting to read a gap as one model being “wrong.” That is the wrong frame. The two models learned from different corpora, weighted differently, cut off at different times. A gap tells you the public record about you is inconsistent — and that different buyers are therefore hearing different things.
LastPass is the clearest illustration. It is a widely-known brand with a widely-known security history. One plausible reading of a 73/6 split is that one model surfaces the brand on name recognition while the other weights the reputational record more heavily. We want to be careful here: we measured the gap, we did not measure the reason. Anyone claiming to know exactly why a model down-ranks a brand is guessing.
We can prove the disagreement exists and how large it is. We cannot prove what is going on inside the model. Those are different claims, and conflating them is how AI-visibility advice turns into astrology.
The zero problem
Four brands in the table score exactly 0 on one platform and something real on the other. Klaviyo: 0 on ChatGPT, 54 on Claude. Linear: 0 on ChatGPT, 27 on Claude. Mangools: 0 and 42.
A zero is the result that most deserves suspicion, because a broken API key produces exactly the same number as genuine absence. That is why every batch we run starts with a control scan: we measure a brand that is unambiguously well-known in its category first, and if that comes back at zero we throw the entire run away rather than publish it. For this dataset the control returned 78 out of 100 — so the zeros above are real absence, not a dead key.
That distinction is the whole ballgame. Without a control, “we found your brand is invisible” and “our scanner was broken” are indistinguishable, and only one of them is worth paying for.
What to do if you have a gap
Check both. That is the boring, correct answer, and it is the reason this product measures two platforms rather than one. If your visibility differs by 30, 40, 60 points across models, an optimisation aimed at the model you happened to check may do nothing for the one your buyer actually opened.
Then ask what the models are reading. Neither of them is crawling your homepage and grading your copy. They are synthesising what everyone else has written about you — reviews, comparisons, forum threads, tutorials. A gap usually means that body of writing is thin, contradictory, or weighted toward a story you would not choose.
Method, so you can argue with it
Twenty categories, eight brands each, measured 2026-07-13 against ChatGPT and Claude with buyer-intent questions. Every question is asked repeatedly rather than once, because a single answer from a stochastic model is an anecdote, not a measurement. Scores are 0–100 and reflect how often and how prominently a brand is named. A control scan gates every batch. Perplexity is not currently measured.
These numbers are a snapshot of one week, not a permanent ranking. They will move. That is the point of measuring them on a schedule rather than screenshotting them once.