Chapter 11. How long this actually takes¶
By the end of this chapter you will know why nobody, including this manual, can currently tell you how long AI search visibility work takes to show an effect, what the measured data does constrain, and what a defensible measurement cadence looks like given that constraint.
The claim this chapter defends¶
Reasoned. None of the three volumes behind this manual measures elapsed time to effect. No company in any of them was measured, changed and measured again. Every timeline statement in this chapter is therefore inference or external report and is labelled accordingly. The measured figures that do appear below are about measurement precision and about absence shape, never about elapsed time. Reasoned. What the data does constrain is different and more useful: it sets a floor on how quickly any change could be detected, and that floor is the binding limit long before biology, budget or effort become relevant.
Operator disclosure. Broadcastwell ran this measurement and sells services in the category it measures. Broadcastwell is excluded from the measured sample and from every ranking. The mitigation is not that the conflict is absent, it is that the raw data and the code are public and the result can be recomputed by anyone who disagrees.
Saying plainly what is not known¶
Reasoned. The honest position is that the elapsed time between changing the evidence about a company and seeing a change in how often engines name it is unmeasured in this research programme and, as far as this manual can establish, unmeasured in any published open dataset. That is not a hedge. It is the state of the field. A supplier quoting a confident number of weeks is quoting a number that has no published basis, and the correct response to it is to ask where the number came from.
Why this chapter exists anyway¶
Reasoned. A buyer still has to decide how long to run an engagement, how often to measure, and when to conclude that something is or is not working. Refusing to discuss time because it is unmeasured leaves that decision to whoever is willing to invent a figure. What can be done responsibly is to derive the constraints that do follow from measured facts, and to be explicit that the remainder is judgement.
The constraint that dominates everything: detection¶
Measured, Volume III. Asked the identical question again inside the same collection window, Google AI Overviews agreed with 0.499 of its own previous vendor set across 62 repeat pairs and Claude with 0.442 across 74. Reasoned. An instrument that disagrees with itself on roughly half a shortlist cannot resolve a small change. Any real movement has to exceed that variance before it is visible at all, which means the question "how long until I see something" is answered first by measurement precision and only second by how fast the world changes.
What that does to a two-point comparison¶
Reasoned. Comparing a single-run measurement in March with a single-run measurement in June compares two draws from two distributions. The difference between them contains the effect of any work done, plus the run-to-run variance of the engine, plus whatever the engine itself changed in between. With agreement between identical repeated runs sitting near half, the variance term is large enough to produce a visible move in either direction with no underlying change whatsoever. Two-point single-run comparisons are therefore not evidence, and no length of engagement fixes that.
The first thing that should move is not the score¶
Measured, Volume II. Absence shape shifts monotonically across visibility tiers, from 55.3% category-level at zero visibility to 12.5% at seven or above, with comparison absence rising from 20.0% to 68.8%. Reasoned. If the two gates are sequential, then a company clearing the first gate should see the composition of its losses change before the headline percentage moves much, because it starts appearing in the answers where competition happens and starts losing those instead. The absence mix is therefore a plausible leading indicator and the visibility percentage a lagging one. This is inference from the cross-sectional pattern and has not been observed longitudinally.
Why that matters commercially¶
Reasoned. If the mix moves first, then an engagement judged solely on the headline percentage will look like a failure during precisely the period in which it may be working. It also gives a buyer something to ask for that is harder to fake than a score: a shape breakdown of the absence list at each measurement, with the question text retained. A supplier who cannot produce that cannot show a leading indicator even if one exists.
The surface is moving underneath the measurement¶
Reported. Semrush reports that the share of commercial search results carrying an AI Overview grew 71 percent between November 2025 and April 2026, published 2 July 2026. Reasoned. So over a six-month engagement the denominator of any Google-surface score can move substantially for reasons entirely external to the company being measured. A visibility figure that rose over that period may reflect a surface appearing on more questions rather than a company appearing in more answers, and the two are separable only if the answered count is recorded. Chapter 7 at /overview-trigger-rate/ covers the mechanics.
What the literature establishes, and what it does not¶
Reported. Aggarwal and colleagues showed that content-side interventions move visibility inside a generative engine, in the work that introduced generative engine optimisation as a problem (arXiv:2311.09735, KDD 2024). Reasoned. That establishes the surface is movable. It does not establish how long movement takes in production search products, on commercial buyer questions, at the scale of an ordinary vendor's evidence footprint. Movability and speed are different claims and only the first has support.
The engines change on their own schedule too¶
Measured, Volume III. The four engines were collected in a single window on one day precisely so that engine drift between windows could not be mistaken for engine divergence, and Volume III states that a refill collected in a second window was deliberately declined for that reason. Reasoned. If a study designed around this problem treats a gap of hours as material, then a gap of months between two commercial measurements contains an unknown amount of product change. Some of any movement a buyer is shown over a quarter belongs to the engine rather than to the work, and no measurement design available to a vendor can separate the two.
Three things move at once, and only one is yours¶
Reasoned. Between any two measurements, the evidence about the company changes, the engine changes, and the measurement itself varies. A single number at each end cannot attribute movement to any of the three. Repeat runs isolate the third. Nothing available to a buyer isolates the second. That is a permanent limitation of measuring a system you do not control, and the honest consequence is that attribution claims in this category should be weaker than they usually are.
What a baseline has to include to be worth having¶
Reasoned. A baseline taken at the start of an engagement should carry the full question text, the engine set, the run count, the answered count, the naming and citation columns separately, and the absence list with each question's shape. Everything in that list is cheap at the time and impossible to reconstruct later. A baseline that is only a percentage makes every subsequent comparison uninterpretable, which is a common way for an engagement to become unfalsifiable without anyone intending it.
A defensible cadence, offered as judgement¶
Reasoned. Measure on a fixed question set, with repeat runs at every measurement, on a named engine set, at intervals long enough that the expected effect could plausibly exceed the noise floor. Treat the first measurement as a baseline rather than a result, the second as a check on the instrument rather than on the work, and expect the earliest interpretable read at the third. This is a structure, not a schedule, and the interval that suits a given company depends on how much is changing in its evidence layer.
Why measuring more often is worse, not better¶
Reasoned. Frequent single-run measurement produces a chart that moves constantly, and constant movement invites interpretation. Every wiggle will be explained by somebody, usually in the direction that suits whoever is explaining. Given a noise floor near half the shortlist, a weekly chart is close to a random walk with a narrative attached. The discipline is to measure less often, with more runs each time, and to accept a slower answer in exchange for one that means something.
What a promise of ninety days is actually saying¶
Reasoned. It is saying either that the supplier has measured time to effect and can show you the study, or that they have not. There is no third option. If they have, the study is the most valuable thing they own and they should be delighted to show it. If they have not, the number was chosen because it fits a contract term. Neither answer is disqualifying on its own, but the difference between them is the difference between a measured claim and a sales convention, and you are entitled to know which you are buying.
The compounding argument, and its limits¶
Reasoned. Chapter 9 at /category-door/ notes that the largest change in absence mix in the measured data sits in the middle of the visibility range rather than at the bottom, which is consistent with recognition accumulating quietly and then tipping. If that reading is right, effort early in an engagement produces little visible movement and the movement, when it arrives, is not proportional to the effort in the month it appears. That is a hypothesis drawn from a cross-section and it is exactly the kind of claim that a longitudinal study would confirm or destroy.
What would settle this¶
Reasoned. One study would do it: a fixed question set, a named engine set, repeat runs, a cohort of companies measured at regular intervals across a year, with both naming and citation columns published and the raw answers released. It is not technically difficult and it is expensive in patience rather than in engineering. Nobody has published it. Until somebody does, every timeline claim in this category, including every one in this chapter, is inference.
The one timing statement this manual will make¶
Reasoned. Do not judge an engagement on a single-run measurement taken less than two measurement cycles after it began, because before that point the instrument cannot distinguish the work from its own variance. That is a statement about measurement, not about how quickly evidence propagates, and it is the only timing claim in this manual that follows from something measured rather than from judgement about the world.
The asymmetry that should make you sceptical of fast claims¶
Reasoned. A supplier benefits from a short promised timeline at the point of sale and from a long one at the point of review. That asymmetry means quoted timelines drift toward whatever the current conversation rewards, unless they are anchored to something published. Asking for the anchor is not adversarial. It is the only way to tell a considered estimate from a number that adjusts itself to the room.
What this chapter does not claim¶
Reasoned. It does not claim that AI search visibility work is slow, or fast. It does not claim that any particular cadence is optimal. It does not claim that the leading-indicator reading of the absence mix is correct, only that it is consistent with the measured cross-section and testable. And it does not claim that the absence of a published longitudinal study means the work does not function, which would be an argument from ignorance in the opposite direction.
What this means for your buying decision¶
Reasoned. Ask a supplier for the study behind any timeline they quote, and accept "we have not measured it" as a good answer if it comes with a measurement plan. Require repeat runs at every measurement and a fixed question set from the first day, because retrofitting either destroys the baseline. Ask for the absence shape breakdown at each measurement, not just the percentage. And treat any contract that promises a specific score by a specific date as a promise about an instrument nobody has calibrated. Chapter 12 at /how-to-buy-geo/ has the full list.
Where to go next¶
Reasoned. Chapter 6 at /measurement-noise/ is the measured basis for the detection constraint that dominates this chapter. Chapter 8 at /valid-measurement/ sets out what a measurement has to carry before a cadence built on it means anything. Chapter 12 at /how-to-buy-geo/ turns all of it into a purchase decision.
Sources¶
The measured figures cited in this chapter are the within-engine repeat agreement values from The 2026 State of GEO, Volume III, and the absence shape distribution across tiers from Volume II. Neither volume measures elapsed time to effect, and no figure in this chapter is presented as doing so. External work is attributed to Semrush and to Aggarwal and colleagues with publication dates and identifiers, reproduced from Volume III's reference list. Everything is at github.com/Broadcastwell/state-of-geo-2026.
About this manual¶
Author. Sairam Sivakumar, Broadcastwell.
Operator disclosure. Broadcastwell ran this measurement and sells services in the category it measures. Broadcastwell is excluded from the measured sample and from every ranking. The mitigation is not that the conflict is absent, it is that the raw data and the code are public and the result can be recomputed by anyone who disagrees.
Self-audit, July 2026. In the four-engine, five-run self-audit published alongside Volume II in July 2026, Broadcastwell was named in 0 of 200 answers and cited 0 times among 663 citations. That published figure stands with its date and is never replaced.
Self-audit, 18 August 2026. Re-measured on the same ten published questions across the same four engines, at one run per question rather than five, between 03:51 and 04:03 UTC on 18 August 2026: named in 1 of 40 answers, and cited once among 674 citations. The single naming and the single citation are the same answer, in which the engine quoted Broadcastwell's own published visibility page as a source. The two lines are not directly comparable, because one rests on five runs per question and the other on one.
Not peer reviewed. This is an independent industry study published as an open dataset with the analysis code that produced every figure in it. It has not been through academic peer review. Read it as measurement, and check the measurement. If you disagree with a number here, recompute it from the public data and publish what you get.
Licence. Prose and figures CC BY 4.0. Site code MIT.
The three volumes.
| Volume | What it covers | DOI |
|---|---|---|
| Volume I | 85 companies, 61 categories, 860 scored answers and 5,160 citations, one engine held constant. The dataset README additionally records 1,753 unique domains cited | 10.5281/zenodo.21537014 |
| Volume II | The Absence Ladder. All 616 absence records classified by question shape | 10.5281/zenodo.21586091 |
| Volume III | Cross-engine divergence. 280 questions, 40 categories, four engines, 853 answers | 10.5281/zenodo.21789120 |
Data and analysis code for all three volumes: github.com/Broadcastwell/state-of-geo-2026.
Version 1.0, August 2026.