deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Perplexity software citations dominated by obscure sites, including 215,000 machine-made pages

A 380-category test of Perplexity's sonar models found 59.8% of 7,534 citations pointing to domains ranked below Tranco's top 100,000, including a three-site network with 215,128 machine-generated buying guides.

Perplexity software citations dominated by obscure sites, including 215,000 machine-made pages

A stress test of Perplexity's answer engine found that when it recommends software, most of the web pages it retrieves come from obscure, often machine-generated sites rather than established publishers. According to a report published by Trellner that reached the Hacker News front page, 59.8% of 7,534 citations gathered across 380 product categories pointed to domains ranked below position 100,000 in the Tranco top-one-million list, and 23.4% went to domains outside the top million entirely.

How the test was run

On 2 September 2026 the researchers put 380 buyer-intent categories — from broad ones like CRM software to niches such as tools for managing museum collections — to Perplexity's sonar and sonar-pro models through OpenRouter, one prompt per category per model, 760 calls in total. Each prompt requested a ranked top five as JSON, including every product's official homepage domain, and the category list was fixed before any results were seen. Both models were chosen because they disclose the URLs they retrieve. The run produced 3,800 recommendation slots naming 1,807 distinct products and 7,534 citations across 2,055 domains, each checked against Tranco's daily list for 1 September and the Wayback Machine.

Google was deliberately left out: grounding a Gemini model through OpenRouter routes retrieval through OpenRouter's own web-search plugin, so the results would describe that plugin rather than Google's. The report is explicit that only Perplexity was measured and that nothing in it should be read as a claim about other engines.

Where the citations land

The median Tranco rank among the 5,768 citations hitting ranked domains was 71,611, and 751 of the 2,055 cited domains — 36.5% — do not appear in the top million at all. Unranked domains were also markedly younger: their median first Wayback capture is 2020 versus 2011 for ranked domains, and 16.6% of archived unranked domains first appeared in 2025 or later, against 1.6% of ranked ones. Concentration at the top is described as unremarkable — the ten most-cited domains take 17.3% of citations — so the story is not a cartel of famous sites but what fills the remaining four-fifths. Wikipedia, for comparison, was cited three times in 7,534.

The third most-cited source was guideflow.com, a vendor of interactive product demos. It is neither a review publisher nor a directory, and it competes in none of the tested categories, yet its content-marketing blog was cited 194 times across 96 of the 380 categories — a quarter of them — placing it ahead of Gartner. The report stresses that nothing here is deceptive; the finding is what the retrieval layer does with ordinary corporate listicles about markets a company does not operate in.

215,128 generated buying guides

Three sites further down the ranking — wifitalents.com (71 citations, 27 categories), worldmetrics.org (60, 22) and gitnux.org (50, 23) — appear to be a single operation. All three were registered through NameCheap between December 2023 and May 2024, delegate DNS to the same pair of Cloudflare nameservers, and share one page template and navigation. Trellner calls the nameserver overlap circumstantial rather than proof, but notes the sites also run identical six-post blogs that write only about each other and a fourth brand, zipdo.co, on the same nameservers. Their sitemaps list 103,578, 107,083 and 105,541 URLs, of which 215,128 are pages of the form "best [category] software". There are not 215,128 software categories.

The self-description is what stands out: worldmetrics and gitnux title their homepage "Facts & Grounding Page" — grounding being the retrieval step these models perform — and describe themselves as market research companies offering software best lists in a machine-readable record. The report's reading is that these pages are addressed to the software that reads them, not to people. The network also monetises directly, advertising custom market research from €5,000, ready-made reports from €499 and vendor selection from €2,500.

Its output is arbitrary. The same "project estimation software" page fetched from all three brands ranks ten tools with different winners — Gitnux's top pick does not appear in Worldmetrics' top five — credits nine distinct named staff across three pages, and carries unrendered template variables in the bylines. Separately, of the 1,502 vendor homepages the models supplied, 17 domains (1.1%) are dead or unreachable and 92 (6.1%) now redirect elsewhere; asked about research data management, both models named Dryad, but sonar-pro cited datadryad.org while sonar offered dryad.co, a redirect.

Why it matters

Answer engines are becoming a discovery surface for buying decisions, and this data suggests their retrieval layer can be captured by whoever publishes the most keyword-shaped pages. Programmatic SEO at a scale of 215,000 near-identical buying guides, plus homepages explicitly titled for the grounding step, points to an emerging discipline of optimising for the crawlers that answer engines trust — with paid research and report services attached. The contradictory rankings from a single template show the resulting recommendations carry no editorial weight, and the dead or redirected vendor domains show weak citation hygiene. The scope is limited: one engine, one run, 380 categories. But if a three-site network can become a top-ten evidence base within roughly two years of its domains existing, every answer engine built on open-web retrieval shares the same vulnerability.

  • #perplexity
  • #seo
  • #ai-search
  • #search-engines
  • #programmatic-seo

Related posts