· via dev.to (home feed)
ChatGPT and Gemini agree on the same top local business in just 4.2% of queries, study finds
A Steady Demand study of 1,487 local-service queries found ChatGPT and Gemini named the same top business only 4.2% of the time, and repeated prompts cited different sources 60% of the time.

The headline finding
ChatGPT and Gemini rarely reach the same conclusion when asked to recommend a local service business. According to research published by Steady Demand in its AI Citation Ledger and surfaced via dev.to, the two assistants named the same top business in only 4.2% of identical queries.
The study covered 1,487 local-service searches spanning 50 U.S. metropolitan areas and 10 service verticals, using prompts in the mold of "best plumber near me". For each response, the researchers recorded which business was named first and which sources the assistant used to ground its answer.
Different engines, different evidence
The disagreement traces back to how each assistant assembles evidence. Gemini drew roughly 60% of its citations from business websites, while ChatGPT leaned more heavily on Reddit and traditional directories. Across both engines, only about 8% of cited domains overlapped, which means the two systems are effectively reading different versions of the same local market.
The source mix also shifts by industry. The study found distinct citation behavior across verticals such as personal injury law, dentistry and auto repair. National platforms including Angi, the Better Business Bureau and Reddit appeared as common reference points across metros, but regional differences still shaped which businesses and sources showed up.
AI answers are less stable than the local pack
Repeatability is a separate problem. When the researchers re-ran queries word for word, the sources cited aligned only about 40% of the time, a phenomenon the report calls "grounding drift". A single favorable or unfavorable screenshot is therefore weak evidence of anything durable.
The contrast with conventional local search is stark. Google's local pack reproduced its top result about 90% of the time in the study's benchmark, compared with roughly 7% top-match stability for Gemini's AI-generated results. The supplied research did not break out a separate repeatability figure for ChatGPT.
What this means in practice
The report's core message is that "ranking in AI" is not a single settled score a business can measure once. Its suggested monitoring approach has several moving parts:
- Test each assistant separately rather than treating one engine as representative of all AI discovery.
- Use prompts that mirror real customer phrasing, services and locations.
- Repeat the same prompts over time to separate one-off answers from recurring patterns.
- Record cited sources alongside the businesses named, since sources reveal where each engine looks for evidence.
- Segment results by vertical and market, because a tactic that helps one category or metro may not transfer.
The source patterns hint at where effort might pay off. If Gemini favors business websites, clear service pages and consistent location details should make a business easier to ground. If ChatGPT favors directories and community discussion, listing accuracy and clear public information matter more. The study is careful not to promise that any specific tactic, structured data included, guarantees inclusion in either assistant.
Why it matters
As consumers increasingly encounter businesses through AI assistants rather than a map pack or a list of links, discovery is fragmenting by engine. A business that looks dominant in ChatGPT may be nearly invisible in Gemini, and vice versa, because the two systems cite largely non-overlapping sources. A 4.2% agreement rate means that checking one assistant, the equivalent of monitoring a single search engine in an earlier era, materially misreads a brand's AI visibility.
The study also has clear limits worth respecting. It measures citations and named businesses, not conversions, and its scope is U.S. local services across ten categories, so it cannot establish whether AI answers drive leads or how the findings generalize elsewhere. What it does establish is a measurement problem: teams using AI outputs for decisions should log the engine, prompt, location, date, cited sources and repeat runs, and treat any single answer as one data point rather than a stable market ranking.
- #ai-search
- #chatgpt
- #gemini
- #local-seo
- #digital-marketing