Skip to content

Chapter 10. Getting quoted at the comparison gate

By the end of this chapter you will know what the second gate tests, why it is the gate that most visible companies are stuck at, and what properties a comparative claim has to have before a model can use it. This chapter sets standards for the evidence, not tactics for producing it.

The claim this chapter defends

Measured, Volume II. Companies named in seven or more of ten answers lost head-to-head comparisons 68.8% of the time, across 8 companies and 16 absence records. Measured, Volume II. Across all 616 absence records, visibility correlates with absence-being-a-comparison at r = 0.234, p = 4.1 × 10⁻⁹, and the effect survives at r = 0.184, p = 5.9 × 10⁻⁶ over 600 records when the small top tier is excluded entirely. Reasoned. Once a company is a recognised member of its category, nearly everything it still loses is a question that names a specific rival, and that is a different problem from membership.

Operator disclosure. Broadcastwell ran this measurement and sells services in the category it measures. Broadcastwell is excluded from the measured sample and from every ranking. The mitigation is not that the conflict is absent, it is that the raw data and the code are public and the result can be recomputed by anyone who disagrees.

What counts as a comparison question

Measured, Volume II. A question was classified as a head-to-head comparison if it contained "vs", "versus" or "comparison", and that rule took precedence over every other shape, so a question containing both "best" and "vs" was counted as a comparison. Reasoned. The precedence matters for reading the tier table. Comparison is the greediest category in the scheme, which makes the 68.8% at the top tier a conservative reading of the alternatives-and-shortlist share rather than an inflated reading of the comparison share.

The gate is where visible companies live

Measured, Volume II. Comparison absence rises across the tiers from 20.0% at zero visibility, to 24.6%, to 41.6%, to 68.8%. Reasoned. The ordering is monotonic across all four tiers, which is what a sequential process looks like when cut into slices. It also means that success at the first gate delivers you to the second rather than to the end. A company that does category work well should expect its absence mix to shift toward comparisons, and should read that shift as progress rather than as a new failure.

Why the top-tier figure needs its base attached

Measured, Volume II. Nine companies scored 7 or above, but one scored 10 of 10 and produced no absence records at all, so the top tier rests on 8 companies and 16 records. Reasoned. A percentage quoted to one decimal place on a base of 16 records is a small base carrying the most quotable number in the ladder, which is a bad combination. The reason to trust the direction anyway is the correlation that survives dropping the tier entirely, at n = 600. Quote the correlation when you need the finding to hold, and quote 68.8% only with its base beside it.

The comparison surface is the one that always exists

Measured, Volume III. On a separate four-engine collection, Google returned an AI Overview for 47 of 47 comparison questions, a rate of 100%, against 88.8% for best-of questions. Reasoned. Comparison questions reliably produce a generated answer. So unlike the category door, where some questions produce no surface at all, the comparison gate is a contest that always takes place. Losing one is losing a contest that happened rather than missing one that did not.

Standard one: a specific named rival appears

Reasoned. A model answering "A versus B" has to find material that addresses A against B. Content that compares a company against unnamed alternatives, or against a generic category, does not supply that. Volume II's own stated practical consequence for companies at this gate is comparison content that a model can quote against a specific named competitor, and the specificity is the operative part. A comparison page that avoids naming anybody is a page that cannot be retrieved for the question being asked.

Standard two: the claim is liftable as a unit

Reasoned. A generated answer is assembled from fragments. A claim that only makes sense after three paragraphs of setup cannot survive extraction, whereas a sentence carrying its own subject, its own object and its own qualifier can. This is the same property that makes a claim checkable by a human reader skimming, and it is why the chapters in this manual are built from short self-contained sections. It is inference about retrieval behaviour rather than a measured result, and it is marked as such.

Standard three: the claim is falsifiable

Reasoned. "Faster" is not liftable in any useful sense, because it has no referent, no measure and no way to be wrong. A claim with a number, a condition and a scope is a claim a model can attribute and a reader can check. The discipline is the same one this manual applies to itself: every figure carries the sample it was computed on. Content that cannot be wrong is content that carries no information, and there is no reason to expect a retrieval system to prefer it.

Standard four: the comparison exists somewhere you do not own

Measured, Volume I. Among the 100 most-cited domains, 77.4% of citations resolved to vendor-authored content, with review platforms at 10.2%. Reasoned. Vendor-authored material clearly does get cited, so a vendor's own comparison page is not disqualified. But a comparison published by the party that wins it carries an obvious discount, and the measured presence of review platforms in the evidence layer indicates the engine reaches for third-party comparison material too. Breadth across both is the standard, not a preference for either.

Standard five: being quoted is not the same as winning

Measured, Volume II. 29.4% of companies were cited more often than they were named, and nine were cited as a source while never being named at all. Reasoned. A comparison page can be retrieved, quoted and used to support an answer that recommends the other company. Supplying the evidence layer for your own defeat is a real and measured position. The standard is therefore not "be quotable in comparisons" but "be quotable in a way that positions you as the answer", and the two come apart often enough to be worth separating. Chapter 3 at /cited-not-recommended/ has the numbers.

Standard six: the material is reachable without a gate

Reasoned. A comparison behind a form, a login or a download is not available to the process that assembles the answer. This is the least arguable standard in the chapter because it follows from how retrieval works rather than from any inference about ranking. It is also the one most often violated, because comparison material is exactly the kind of asset marketing teams instinctively put behind a lead capture.

The engines disagree most about the tail

Measured, Volume III. Treating engines as raters, agreement across three engines was 0.015 over all vendors and 0.609 when restricted to the study universe of established names. Reasoned. Engines converge on the well-known vendors in a category and diverge on everybody else. A challenger fighting a comparison against a named incumbent is fighting on the incumbent's side of that split, where the engines agree, which means the contest is more consistent across engines than category membership is. The practical reading is that comparison outcomes should be more stable engine to engine than category presence, and that a loss on one engine is more likely to be a loss on all of them.

The rival you are compared against is not your choice

Measured, Volume II. For every answer the study recorded which competitor was named most often. Reasoned. The pairing in a comparison question comes from the buyer and from whatever the engine considers the natural counterpart, not from the vendor's competitive positioning deck. A company may be routinely compared with a rival it does not consider a competitor, and comparison material aimed at the rival it wishes it were compared with will not be retrieved for the question actually asked. The absence list tells you which pairings are real, which is the only reliable source for that.

Superlatives are the common failure

Reasoned. Most vendor comparison material is built from claims that cannot be false: better, faster, more flexible, more scalable. Those survive legal review precisely because they assert nothing. A model assembling an answer that has to distinguish two vendors gets no discriminating information from them, and neither does a buyer. The standard that follows is uncomfortable for marketing teams and simple to state: if the claim could be made word for word by the competitor, it is not doing the work.

What the middle tier tells you about timing

Measured, Volume II. The 4 to 6 tier sits at 20.8% category-level absence against 41.6% comparison absence, over 21 companies and 101 records. Reasoned. By the middle of the range the mix has already flipped. That suggests comparison work becomes the binding constraint well before a company reaches high visibility, rather than only at the top. A company at 4 or 5 of 10 that is still spending exclusively on category presence is probably working on the gate it has largely cleared.

What the ladder does not say about causation

Measured, Volume II. The ladder describes a cross-section of 85 companies at one moment, and Volume II states that nothing in the dataset is an intervention study. Reasoned. So the observation that high-visibility companies lose comparisons does not establish that producing comparison content raises visibility. It establishes what the remaining losses look like at each level. Any supplier converting that into a promised mechanism is adding a claim the data does not carry, and this chapter is not making it either.

The variance caveat applies here more than anywhere

Measured, Volume III. Asked the identical question again in the same window, engines agreed with roughly half of their own previous shortlist, at 0.499 and 0.442 mean agreement over 62 and 74 repeat pairs. Reasoned. A comparison outcome measured once is a single draw. Losing one comparison answer is very weak evidence about anything, and a supplier reporting movement on a handful of comparison questions between two single-run measurements is reporting noise. Chapter 6 at /measurement-noise/ is the reason to require repeats before believing any of this moved.

How to tell you are at this gate

Reasoned. Classify your absence list using the published rules in Appendix B at /absence-rules/, applying comparison first because it takes precedence. If comparison-shaped questions dominate what you lose, you are at the second gate. If the two buckets are close to even you are in transition, and the sequencing argument in Chapter 9 at /category-door/ says to clear the membership constraint first because it bounds the return on everything else.

One company in this sample cleared both gates

Measured, Volume II. One company scored 10 of 10 and therefore produced no absence records at all. Reasoned. That is a single case and it proves nothing about method, but it does establish that the ceiling in this dataset is reachable rather than theoretical. It also means the ladder has an end: there is a state in which there is no absence list left to classify, and a company approaching it should expect its remaining losses to be almost entirely comparisons before they disappear.

What would falsify the standards in this chapter

Reasoned. A published before-and-after on a fixed question set, with repeat runs, in which a company at the comparison gate added checkable, third-party, named-rival comparison material and did not move. Nobody has published that test either way. Until somebody does, these six standards are inferences from a cross-sectional pattern and one stated practical consequence, and they are labelled as inference throughout rather than presented as a mechanism.

What this means for your buying decision

Reasoned. If comparison questions dominate your absence list, ask a supplier which named rivals the work will address and how many of the resulting claims will be checkable rather than adjectival. Ask what will exist off your own domain. Then ask how movement will be measured, and refuse an answer that involves a single run of a handful of questions. Chapter 12 at /how-to-buy-geo/ has the full list of questions to put to a vendor before signing.

Where to go next

Reasoned. Chapter 9 at /category-door/ is the first gate and the one to clear before this one. Chapter 11 at /how-long-it-takes/ is honest about cadence and about how little the evidence says on elapsed time. Chapter 2 at /absence-ladder/ is the measured basis for both gates.

Sources

The measured figures in this chapter come from The 2026 State of GEO, Volume II: the absence ladder across all four tiers, the comparison shape rule and its precedence, the correlation between visibility and comparison-shaped absence with and without the top tier, the composition of the top tier, and Volume II's stated practical consequence for companies at this gate. The trigger rate on comparison questions and the repeat agreement figures are Volume III. The citation composition figure is Volume I. Everything is at github.com/Broadcastwell/state-of-geo-2026.

About this manual

Author. Sairam Sivakumar, Broadcastwell.

Operator disclosure. Broadcastwell ran this measurement and sells services in the category it measures. Broadcastwell is excluded from the measured sample and from every ranking. The mitigation is not that the conflict is absent, it is that the raw data and the code are public and the result can be recomputed by anyone who disagrees.

Self-audit, July 2026. In the four-engine, five-run self-audit published alongside Volume II in July 2026, Broadcastwell was named in 0 of 200 answers and cited 0 times among 663 citations. That published figure stands with its date and is never replaced.

Self-audit, 18 August 2026. Re-measured on the same ten published questions across the same four engines, at one run per question rather than five, between 03:51 and 04:03 UTC on 18 August 2026: named in 1 of 40 answers, and cited once among 674 citations. The single naming and the single citation are the same answer, in which the engine quoted Broadcastwell's own published visibility page as a source. The two lines are not directly comparable, because one rests on five runs per question and the other on one.

Not peer reviewed. This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. It has not been through academic peer review. Read it as measurement, and check the measurement. If you disagree with a number here, recompute it from the public data and publish what you get.

Licence. Prose and figures CC BY 4.0. Site code MIT.

The three volumes.

Volume What it covers DOI
Volume I 85 companies, 61 categories, 860 scored answers and 5,160 citations, one engine held constant. The dataset README additionally records 1,753 unique domains cited 10.5281/zenodo.21537014
Volume II The Absence Ladder. All 616 absence records classified by question shape 10.5281/zenodo.21586091
Volume III Cross-engine divergence. 280 questions, 40 categories, four engines, 853 answers 10.5281/zenodo.21789120

Data and analysis code for all three volumes: github.com/Broadcastwell/state-of-geo-2026.

Version 1.0, August 2026.