Launching soon. Preferium for agencies is in private preview — partner registration isn’t open yet. Join the waitlist →

All articles

Ask Gemini the same question twice and most of the cited sources change

A study of 14,472 local AI citations found that repeating an identical query returned only 26-46% of the same sources, and Gemini named the same top business 7.9% of the time against 90.2% for Google.

Two identical magnifying glasses on stone plinths: one frames a calm blank wall, the other a scattered collage of photos, beside neat and toppled stone stacks.

TL;DR: When you ask an AI assistant a question, it usually goes and reads a few web pages before answering, then tells you which pages it used. Those are its sources. New research says that if you ask the exact same question again a few minutes later, most of those sources are different ones. Not slightly different. Roughly half changed on an immediate repeat, and about three quarters changed by the next day. So a single screenshot of “here is where the AI got its answer” is not a measurement of anything. It is one roll of the dice.

First, what is a “citation” here?

A citation is a link the assistant shows next to its answer, pointing at a page it read while writing that answer. It matters commercially because being one of those links is how a business gets seen inside an AI answer, now that many people never scroll to the ten blue links underneath.

The other term you need is the local pack: the boxed set of three nearby businesses Google shows for something like “plumber near me”. It has been the benchmark for local search for years, and it is stable. That stability is the thing AI answers turn out not to have.

What the study actually did

Steady Demand, a local SEO agency, ran 1,487 local-service queries across all 50 US metro areas and ten categories: plumber, roofer, HVAC, electrician, locksmith, pest control, cleaning, lawyer, dentist, auto repair. That produced 14,472 citations, collected on 27 and 28 July 2026. Gemini was queried through the API as gemini-flash-latest with Google Search grounding switched on. ChatGPT was captured through a scraper hitting its consumer Search mode. Search Engine Land covered the findings on 19 August.

The headline result is the friendly one. Almost 60% of Gemini’s local citations pointed at the business’s own website. Reddit took 13.7%, beating every directory: the local-service directories added up to 10.3% between them. If you have spent two years being told your own site no longer matters, that is a useful correction.

Then the authors did something most citation studies skip. They asked the same questions again.

The repeat test

Six fixed queries, four rounds, same wording every time. 72 calls in total. That is a small sample and the authors say so, along with the rest of their stated limitations: one snapshot in time, API behaviour that may not match what a consumer sees in the app, US only.

The overlap between rounds was measured with a similarity score where 100% means the two lists of cited domains are identical and 0% means they share nothing. The results:

  • Back-to-back, same day: 46.3%
  • Same day, 3.5 hours later: 41.0%
  • Next day, 19 to 20 hours later: 26.5%

Rephrasing the question three different ways landed around 40% too, so wording was not the driver. Neither was time: the cross-day number sits close to the same-day cross-round number.

The mechanism is visible one level down. Before answering, the model writes its own search queries and sends those to Google. On back-to-back repeats those internally generated search strings overlapped by 0.056 on the same scale. Essentially not at all. The randomness is upstream of the citations. Different searches go out, so different pages come back, so different sources get cited. Steady Demand calls it grounding drift.

The comparison that makes it concrete: asked repeatedly, Gemini named the same top business 7.9% of the time. Google’s local pack, asked the same thing, returned the same top listing 90.2% of the time. Instability here is a property of the AI layer, not of local search.

The two engines are not reading the same web

The same 1,487 queries went to both engines. Their cited domains overlapped 8% of the time on average. They recommended the same top business in 4.2% of cases.

That falls out of the source mix. Gemini’s citations were 59.9% business websites. ChatGPT’s were 15.9% business websites and 41.7% social and community forums. One engine reads your site. The other mostly reads people talking about you. Optimising for one of them tells you very little about the other.

What to do about it

Three things follow, and none of them are exotic.

Stop treating a single check as evidence. If one query has a coin-flip chance of returning the same sources, a screenshot in a client deck is decoration. Run the same prompts on a schedule and report the trend across many samples, which is the same argument for measuring rather than estimating we have made before.

Measure each engine separately. At 8% overlap, “AI visibility” as a single number averages away the only thing that is actionable.

And keep your own pages in shape. Gemini’s mix says the business site is the single largest source of local citations, so the ordinary work of accurate, crawlable, well-structured pages is still the work.

Preferium measures four engines on every plan, ChatGPT, Claude, Perplexity and Gemini, with Google AI Overviews and AI Mode tracked separately, six answer surfaces in total. The point of running them continuously rather than on request is exactly this study’s finding: one sample is noise, and repeated samples across engines are the only way a trend line means anything. On the page side, 47 automated checks crawl every page and score the site out of 1000, and fixes are deployed and re-checked in a real browser without a dev queue. More on how the system works.

Key takeaways

  • Repeating an identical query to Gemini returned only 46.3% of the same cited domains immediately, 41.0% after 3.5 hours and 26.5% the next day.
  • The drift starts before the citations: the model’s self-written search strings barely overlapped between repeats.
  • Gemini named the same top business on repeat 7.9% of the time; Google’s local pack managed 90.2%.
  • Gemini and ChatGPT agreed on cited domains 8% of the time and on the top business 4.2% of the time.
  • Gemini’s local citations were 59.9% business websites, 13.7% Reddit; ChatGPT’s were 15.9% business websites and 41.7% forums.
  • The repeat test is 72 calls from one two-day window. Treat the direction as solid and the exact percentages as indicative.
More articles Become a partner