iSeer
Back to field notes
Field notes · research · July 11, 2026 · 6 min read

ChatGPT and Claude disagree about you more than you think.

The same buying question, two platforms, different brands in the answer — and why a single-platform score cannot describe that market.

IT
iSeer Team
AI visibility intelligence

Ask ChatGPT and Claude the same buying question and you may receive two competent answers with different brands in them.

That is not an edge case. Across the tracked questions currently available in the iSeer ledger, the two platforms disagreed on brand presence in 18% of cases — 143 of 795 matched question-pairs.

The percentage above must be read with its sample definition: the latest complete report per scanned domain (163 reports, measured as of 11 July 2026); a matched pair is the same buyer-intent question answered without error on both platforms; a disagreement means the brand was mentioned on exactly one of the two.

The point is not that one model is right and the other is wrong. The point is that a single-platform score cannot describe a market in which buyers use more than one system.

The same question does not create the same answer.

Language models do not consult a common, fixed index and then apply identical ranking rules.

ChatGPT and Claude differ in model behaviour, retrieval, product design and the sources available to them. They may interpret the category differently. They may favour different evidence. They may produce a longer or shorter shortlist. They may attach your brand to different use cases.

A company can therefore be:

Only the first two cases are visible in a blended yes-or-no score. The other three contain much of the useful information.

A real divergence.

The example below comes from iSeer's own public category scan — the same data behind the public note-taking leaderboard — not from a customer account.

In the note-taking category scan of 10 July 2026, obsidian.md was mentioned by both platforms on the broad question “What are the best note taking app options available right now?” But on three buyer-intent questions ChatGPT omitted it while Claude included it: “Can you compare the top 5 note taking app options for me?”, “I'm new to note taking app. Where should I start?”, and “What's the best note taking app for someone on a budget?” Its per-platform quality scores in that report: ChatGPT 30, Claude 68.

No conclusion should be drawn from this example alone. It is included because it makes the measurement problem concrete. The same company, question and measurement period produced materially different visibility depending on the platform.

A single “AI visibility” percentage would have hidden the distinction.

Why blended scores mislead.

A blended score can be useful as a summary, but it becomes misleading when it removes the underlying platform evidence.

Suppose a brand is present in every ChatGPT run and absent in every Claude run. An average may place it somewhere in the middle. That number sounds moderate. The actual finding is not moderate at all: one channel is working and one is not.

The remediation is also different.

If the brand is absent everywhere, the likely problem is broad category recognition or positioning. If it is absent on one platform only, the first task is to inspect the sources, framing and answer patterns that differ between platforms. The diagnosis should follow the evidence rather than a generic checklist.

This is why iSeer retains the run-level records and shows platform-level results.

The summary is not permitted to replace the receipts.

Presence is only the first layer.

Two platforms can both mention a brand while still disagreeing in ways that matter.

One may name it first and describe it as a category leader. The other may place it in a secondary list and add a caveat. One may associate it with the intended audience. The other may frame it around a legacy use case.

A binary presence metric cannot capture all of this. It remains useful, but it is only the first layer.

The earlier field note, Why ChatGPT doesn't recommend you, describes the failure modes iSeer uses to separate absence, weak positioning, outframing and platform-specific gaps.

Why manual checking fails.

Teams often test AI visibility by opening a chat window and asking a question once.

That method has four problems.

First, the prompt is usually invented by someone who already knows the brand. Buyers ask different questions.

Second, the result is stochastic. Another run may change the shortlist.

Third, the tester often checks only the platform they personally use.

Fourth, there is no retained baseline. When the answer changes a month later, nobody can establish whether visibility fell or the original observation was simply noise.

A repeatable measurement needs defined prompts, both platforms, stored runs and a rule for interpreting change.

iSeer's methodology explains how prompts, runs and confidence are handled.

What a useful comparison looks like.

A practical platform comparison should show, for each buyer-intent question:

QuestionChatGPTClaudeInterpretation
What are the best note taking app options available right now?MentionedMentionedPresent on both
Can you compare the top 5 note taking app options for me?AbsentMentionedPlatform-specific gap
I'm new to note taking app. Where should I start?AbsentMentionedPlatform-specific gap
What's the best note taking app for someone on a budget?AbsentMentionedPlatform-specific gap

The important column is the last one. A table of mentions is descriptive. A diagnosis identifies what to investigate.

Category context matters.

Platform disagreement is easier to understand when the surrounding category is visible.

A brand may be absent because neither platform recognises it as part of the category. It may also be absent because the category itself is unstable and each model constructs a different comparison set.

The relevant note-taking leaderboard shows which brands appeared across the measured prompts and platforms.

Leaderboards are not permanent league tables. They are dated measurements of a defined prompt set. Their value lies in making the comparison inspectable.

Measure both, then decide.

The sensible response to platform divergence is not to produce separate marketing strategies for every model after one scan.

Measure the disagreement first. Check whether it persists. Inspect how the brand is framed. Then choose the smallest change that addresses the observed gap and re-measure it.

That process is slower than declaring victory from a favourable screenshot. It is also more likely to tell you something true.

Try it on your brand

Run the free report across ChatGPT and Claude.

30 seconds. No signup. The same buyer-intent questions asked on both platforms, with platform-level results retained — the summary never replaces the receipts.