Answer engines often disagree on who belongs in a category shortlist and how those brands are framed. Brand teams should measure that disagreement as a first-class metric—not noise to average away.

What we measured
BrandAI ran a methodological sample (not a census) designed to mirror how operating teams should instrument BrandAEO:
- 48 locked unbranded category prompts (discovery + best-of phrasing)
- 3 assistants sampled on the same wording the same week
- Peer set of 6 brands pre-registered before scoring
- Scores: inclusion (mention), first-mention on best-of, and trust-framed language among mentions
How much do engines agree?
On the same 48 prompts, pairwise shortlist overlap (Jaccard on mentioned peer-set brands) looked like this:
| Pair | Overlap (Jaccard) | Read |
|---|---|---|
| ChatGPT ↔ Perplexity | 0.61 | Shared core, different edges |
| ChatGPT ↔ Gemini | 0.54 | Larger framing / source drift |
| Perplexity ↔ Gemini | 0.58 | Citation habits diverge |
What this means: If you only monitor one engine, you will misread category weather. A "win" on one assistant can coexist with absence or harsh framing on another.
Where disagreement concentrates
| Prompt family | Highest disagreement driver | Typical team mistake |
|---|---|---|
| Category discovery | Category nouns / synonyms | Measuring only branded vanity prompts |
| Best-of | First-mention + trust adjectives | Treating SOV as preference |
| Comparison | Citation to docs vs reviews | Fixing homepage copy only |
A worked pattern (illustrative)
On implementation-heavy category prompts in the sample, one assistant repeatedly cited vendor docs; another leaned on roundup reviews with stale feature matrices. Shortlist overlap looked "fine" at the brand-name layer, while framing diverged: "powerful but complex" vs "best fit for mid-market."
Operating translation: do not celebrate name inclusion while the citation graph teaches two different stories. Fix the stale matrix and the docs conflict as separate tickets—then re-measure the band, not a single engine screenshot.
Operating rules for brand teams
- Report a disagreement band, not a single SOV number—min/max mention rate across engines for the same prompt family.
- Separate inclusion from preference on every scoreboard (Share of Voice Is Not Preference).
- Assign repairs by evidence class—owned specs, encyclopedic spine, third-party corroboration—not by which engine embarrassed you in a screenshot.
- Re-run the same 48 next week. One dramatic day is not a strategy.
- Version the instrument—prompt-set ID + peer set + sample week printed on every leadership slide.
Pair the readout with Brand Hub identity hygiene and BrandSight structure when gaps look architectural, not merely editorial.
What disagreement is not
| Anti-pattern | Why it fails |
|---|---|
| Abandon measurement | You lose the instrument that shows which evidence graphs diverge |
| Call AI "random" | Disagreement is structured—retrieval and safety priors differ by design |
| Rewrite prompts until one engine flatters you | You break comparability and hide real buyer-journey gaps |
| Skip honest peers and locked wording | You cannot explain gaps you never pre-registered |
Disagreement is weather. Your job is instruments and repairs—not mood.
FAQ
Who should own the disagreement readout each week?
The same AEO DRI who runs the locked prompt scoreboard—usually brand ops or an insights lead. They report min/max mention rate by family, flag one framing split worth a ticket, and refuse to close the meeting without an evidence-class owner.
How do we attach engine bands to repairs without thrashing?
Map each split to an evidence class (owned spec, encyclopedic spine, third-party corroboration), not to "fix ChatGPT." One ticket per fact conflict; re-measure the full band on the same 48 prompts—never swap wording after a bad week.
What is the most common anti-pattern?
Celebrating name inclusion on one engine while citation graphs teach two different stories on another. Shortlist overlap can look "fine" at the brand-name layer while framing diverges on price, complexity, or fit.
Should we optimize for the worst engine?
Optimize for buyer journeys that matter, then watch the band. Chasing a single hostile screenshot without a locked set creates thrash and breaks your Tuesday rhythm.
Where BrandAI fits
- BrandAEO — run the locked panel and print disagreement bands on every leadership slide
- Brand Hub — identity hygiene when category nouns or naming drive splits
- BrandSight — structural context when gaps look architectural, not editorial
- BrandWiki — encyclopedic spine repairs when reference-class facts diverge across engines
Bottom line
AI search is not one leaderboard. It is a set of partially overlapping shortlists.
Measure the overlap. Explain the gaps. Repair the evidence. Then the Tuesday meeting has something real to own.
