iSeer
Back to field notes
Field notes · teardown · July 11, 2026 · 6 min read

Who AI actually recommends for note-taking apps.

Eight note-taking brands, ranked by how often ChatGPT and Claude name them in measured buyer-intent answers — a dated leaderboard, and what it does and doesn't prove.

IT
iSeer Team
AI visibility intelligence

Market share is visible. AI recommendation share is not.

A buyer can know the largest companies in a category and still receive a shortlist dominated by smaller, clearer or more frequently cited brands. Language models do not reproduce a conventional market-share table. They construct an answer from the category, the wording of the question, available sources and their own model behaviour.

We measured which brands ChatGPT and Claude recommended for note-taking apps. The complete dated ranking is available at the note-taking AI visibility leaderboard.

How this leaderboard was measured.

The ranking below comes from five buyer-intent question templates per platform (best-options, comparison, beginner, budget and recommendation phrasings), one run per question on ChatGPT and Claude, sampled on 10 July 2026; a brand is counted when the judge finds it named in the answer.

It is not an estimate of revenue, customer count, product quality or total market share. It measures presence within a defined set of buyer-intent answers. A brand ranks higher when it appears more consistently in those measured answers under the published methodology.

The calculation follows the rules documented in iSeer's methodology. The category page states the sample size and last measurement date.

The measured ranking.

#BrandAIRSChatGPT qualityClaude quality
1notion.so100%74
2evernote.com100%7350
3obsidian.md70%3068
4roamresearch.com40%3027
5bear.app40%2725
6logseq.com10%015
7craft.do0%00
8remnote.com0%00

AIRS = AI Recommendation Share: the share of measured buyer-intent answers naming the brand. It counts a mention, not a rave — how well a brand was framed is scored separately. Notion's Claude column reads “—” because its Claude checks failed in this run; iSeer records failed checks as “no data”, never as absence.

The first result worth noting: Notion and Evernote both appeared in 100% of the measured answers. The category has two default names, not one.

The second: the field collapses fast. Third place — Obsidian — already misses 30% of answers, and its platform split is stark: a quality score of 30 on ChatGPT against 68 on Claude. Logseq appears only on Claude (0 against 15).

The third: craft.do and remnote.com — both established, well-reviewed products — were named in 0% of measured answers on both platforms.

The winner is visible for a reason, but the data does not prove why.

The leading brand appeared in every measured answer (AIRS 100). Plausible explanations are easy to imagine — clear category naming, the kind of ubiquity in comparisons that a default answer accumulates — but those are hypotheses, not causal findings. The leaderboard proves that the brand appeared. It does not prove that one page, campaign or wording choice caused the result.

This distinction is easy to lose in an optimisation industry. A correlation becomes a case study, then a case study becomes a recipe. The evidence rarely supports the final step.

The useful conclusion is narrower: the leading brand is being associated with the category and buyer needs more consistently than the others in this measurement.

The surprising absence.

craft.do is an established, well-reviewed note-taking product — yet it appeared in 0% of measured answers, on both platforms.

That gap between conventional prominence and AI presence is the central finding of the teardown.

It may indicate weak category framing, inconsistent positioning, platform divergence or a crawl/access issue. The leaderboard alone cannot choose among those diagnoses. A brand-level report can. This is the same distinction described in Why ChatGPT doesn't recommend you: absence is an outcome, not a cause.

The gap between conventional prominence and AI presence is the central finding — and absence is an outcome, not a cause.

ChatGPT and Claude construct different categories.

The platform split in this measurement is not subtle. ChatGPT favoured Notion and Evernote. Claude additionally surfaced Obsidian and Logseq. Obsidian appeared consistently on Claude but not on ChatGPT.

A blended ranking is useful for orientation. The platform rows explain where the ranking came from. A company that appears strongly on one platform and weakly on the other does not have middling visibility. It has a specific platform gap.

What the leaderboard does not say.

It does not say that the first-ranked product is the best product.

It does not say that the last-ranked company has no customers or weak conventional SEO.

It does not say that every buyer receives the same answer.

It does not say that the ranking will remain unchanged.

The table is a dated measurement under a defined set of prompts and runs. That modest description is more useful than turning it into an awards programme.

What category companies should inspect.

A company outside the leading group should begin with its actual answer evidence. Check whether the brand is absent, mentioned under the wrong use case, described weakly or visible on only one platform. Then inspect whether the public site states the category, audience and differentiated use case in language that can be understood without reconstructing the company's history.

Do not copy the leader's wording. The goal is not to become a textual imitation of the company already being recommended. The goal is to make your own category position legible and support it with evidence.

After making a change, re-measure. An unmeasured content change is merely activity.

A leaderboard should be inspectable.

The category page includes the ranking, methodology, sample and measurement date. When the data changes, the page should show the new measurement rather than quietly preserving an attractive old result. That makes the leaderboard less dramatic and more credible.

It also means the most interesting result may change. A market leader may recover. A specialist may become more visible. ChatGPT and Claude may converge or diverge. The purpose of the Field Note is to interpret the current evidence, not to freeze it into a permanent claim.

See the evidence

View the full leaderboard, then run your own free report.

The dated note-taking ranking is public. Your own brand's answer evidence — across ChatGPT and Claude, with confidence bands — takes 30 seconds. No signup.