deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Most AI Overview and ChatGPT citations come from the first 30% of a page

CXL's 100-citation study and Kevin Indig's 18,012-citation ChatGPT dataset both show AI search citations clustering in the top 30% of source pages, reshaping how technical docs should be structured.

Most AI Overview and ChatGPT citations come from the first 30% of a page

Two datasets, one pattern

Two independent citation analyses, pulled together in a dev.to write-up, arrive at the same conclusion: when AI search engines cite a source, the quoted passage almost always sits near the top of the page. CXL mapped 100 Google AI Overview citations to where each cited passage appeared relative to total page length, and found that 55% fell within the first 30% of the document. The 10–20% band alone generated more citations than any other segment, the top 20% of pages accounted for 48 of the 100 citations, and the bottom 40% produced just 21.

Kevin Indig ran a far larger exercise across 1.2 million search results and 18,012 verified ChatGPT citations. The shape was the same: 44.2% of citations came from the first 30% of the document, and likelihood dropped sharply after the first third. Indig described the distribution as a "ski ramp" — a steep fall-off followed by a long, slow tail.

The retrieval mechanics behind the bias

The write-up points to Google's own documentation for the mechanism. AI Overviews use query fan-out: the model breaks the original query into multiple related sub-queries, retrieves pages for each, and assembles a response from what comes back. That retrieval step works with a limited amount of context and reads from the top of each fetched page. An answer sitting at the 60% mark of a long article may simply never make it into the context window before the budget is exhausted.

One structural pattern survives deep positions: FAQ sections. According to the CXL data, FAQ entries still get cited from far down the page because each question-and-answer pair is self-contained — the model can lift a complete answer out without needing the surrounding narrative. Ordinary prose does not get that treatment; for it, position is decisive.

Ranking for the keyword matters less than expected

Ahrefs analysed 863,000 keywords and 4 million AI Overview URLs and found that only 38% of cited pages also appeared in the organic top 10 for the same query, down from 76% in July 2025. A separate BrightEdge analysis put the overlap at just 17%. The two figures disagree, but the direction is consistent: most cited pages do not rank for the primary keyword. They are being retrieved for related sub-queries generated by fan-out, which means a page does not need the top organic position to be quoted — it needs a clear answer to a related question, placed near the top.

What technical writers should change

The practical implications for documentation, tutorials and blog posts are direct:

  • Put the answer first and the context second. The clearest statement of the solution belongs in the first 20% of the page, not after several subheadings of warm-up.
  • Front-load definitions. The 10–20% zone was the single most productive segment in the CXL sample, so a concept's definition should land in the first paragraph or two.
  • Keep FAQ blocks for specific questions, since self-contained answers are the one pattern that beats the position effect.
  • Measure directly. Google introduced dedicated Search Generative AI performance reports in June 2026, rolled out worldwide by August 31, exposing AI Overview impressions in Search Console rather than forcing reliance on third-party tools.

Front-loading is not the same as truncating. The bottom 40% of pages still produced 21 of 100 citations in the CXL study, and depth, examples and context continue to matter to readers who actually click through.

Why it matters

For developers and technical writers, AI search has become a discovery channel with its own physics. The retrieval layer appears to read from the top and stop once it has enough, so a correct, well-written answer buried under three sections of background may never enter the model's context window at all. Combined with the Ahrefs and BrightEdge findings, classic keyword rankings turn out to be a weak predictor of citation, decoupling "ranks first" from "gets quoted". The shift required is structural, not stylistic: pages that place the answer where both a scanning human and a retrieval system look first — near the top, or inside a self-contained FAQ entry — are the ones AI search surfaces.

  • #ai-search
  • #seo
  • #technical-writing
  • #documentation
  • #google

Related posts