Launching soon. Preferium for agencies is in private preview — partner registration isn’t open yet. Join the waitlist →

All articles

ChatGPT has its own index — and it builds your snippet out of your H1

A study of 1,249 ChatGPT answers captured in July found OpenAI runs a real in-house index that treats unlicensed small sites like licensed publishers, and cuts each snippet at 202 characters starting near the H1.

A black Labrador rests one paw on a single raised paper file among endless rows of documents stretching into shadow.

TL;DR: When you ask ChatGPT something, it usually goes and looks things up on the web before answering. New research says it does that in two different ways: sometimes by reading a copy of Google’s results, and sometimes from a web index OpenAI built itself. That in-house index has a name — labrador — and the useful finding is that it does not favour the big publishers who have signed deals with OpenAI. A small site with no contract is stored and shown the same way. What it stores from your page is tiny: about 202 characters, taken from the visible text near the top, usually starting at your H1 heading. Not your meta description. Your H1.

First, what happens when ChatGPT “searches”?

Two things happen that people mix up.

A retrieval pipeline is the plumbing that fetches web results. ChatGPT doesn’t have one, it has several. Some of them scrape Google’s results page and hand back what Google would have shown. One of them queries OpenAI’s own stored copy of the web — its own index, the way Google has an index. An index is just a big filing cabinet: for every page it knows about, the search engine keeps a title, an address, and a short piece of the text so it can decide later whether that page answers your question.

Until 21 July, ChatGPT’s own server responses were labelling which pipeline fetched each result. The French SEO consultancy Resoneo read those labels across 1,249 real ChatGPT conversations captured in July and published the whole teardown. The independent analyst Suganthan Mohanadasan had been pulling on the same thread in his own source-selection census, and corrected an earlier assumption when the July data came in. The labels are gone now. The plumbing behind them isn’t.

Labrador is a real index, not a members’ club

The earlier theory was that OpenAI’s own index only held content from publishers it had paid — Le Monde, WSJ, Condé Nast and the rest — and everyone else was reached by scraping Google.

That turned out to be wrong. Resoneo found hundreds of outlets with no OpenAI agreement being served from labrador with the same snippet format, the same length, and the same freshness as the licensed partners. Search Engine Journal’s write-up of the findings leads on exactly that point.

There is a contractual difference, but it sits somewhere else. Resoneo found a wordlim value attached to results — a cap on how many words of your page ChatGPT may quote in its answer. Known licensees came in at 100 words, a handful of outlets at 25, and the default for the rest of the web was 200. The unlicensed default is the most generous of the three.

It’s also not Bing wearing a different hat. Resoneo’s argument there is nicely concrete: a quarter of labrador’s stored titles are longer than the 75 characters Bing will display, so they can’t have come from Bing.

The 202-character budget

Here is the part you can act on this week.

Labrador cuts its stored snippet at 202 characters, and it takes them from the page as rendered — not from the meta description tag, which it ignores completely. (The Google-scraping pipe still uses meta descriptions roughly a third of the time, so the tag is not dead. It’s just irrelevant to this one index.)

Those 202 characters break into three parts, at Resoneo’s median measurements: about 7 characters of whatever sits above your H1, then the H1 itself at around 51, then roughly 146 characters of body text. SEJ reports the H1 landed inside the snippet in 83.6% of the cases reviewed — 387 of 463.

So the anchor is the H1. And that puts a price on two common bits of front-end decoration. A kicker line above the heading — the little “GUIDES” or “CASE STUDY” label — ate about 18 characters in Resoneo’s data. An image with descriptive alt text sitting above the heading could eat 50. Every one of those is a character not spent on your actual first sentence.

Then there’s the plainer failure: Resoneo found roughly one page in seven has no H1 at all. If there’s no heading, the 202-character window starts wherever the visible text starts, which might be your cookie banner copy or a breadcrumb trail.

What to check

  • Does the page have exactly one H1, and does it say what the page is about in plain words?
  • What sits between the top of the rendered body and that H1? Kickers, badges, decorative images with long alt text — each one costs you characters.
  • Read the first sentence after the H1 on its own. Does it stand up as a summary, or does it begin with “In this article we’ll look at…”?
  • Don’t rely on your meta description to carry the page’s meaning into AI answers. Write the top of the page as if it is the description, because for this index it is.

None of that is new SEO. It’s the same on-page hygiene that has always mattered, with a newly specific reason to actually go and do it — across every page, not just the ten you look at.

That last bit is the awkward part in practice. Checking one page takes a minute; checking four hundred, then re-checking them after the next template change, is a project nobody’s retainer covers. Preferium runs 47 automated checks across every page on a site, scores it 0–1000, and then fixes what it finds, deploys the change and re-checks the live page in a real browser afterwards. Missing or duplicated H1s and thin, buried page openings are exactly the class of problem that gets found and closed without anyone opening a ticket. More on how the system works.

Key takeaways

  • ChatGPT reaches the web through several pipelines. One is OpenAI’s own index, labrador; others scrape Google’s results.
  • Labrador does not privilege licensed publishers. Unlicensed sites get the same snippet treatment — and the most generous quoting allowance, 200 words against a licensee’s 100.
  • The stored snippet is 202 characters of rendered page text, anchored on your H1. Meta descriptions are ignored by this index.
  • Anything you put above the H1 — kickers, badges, images with long alt text — is subtracted from what ChatGPT stores about the page.
  • About one page in seven has no H1 at all, which leaves the snippet starting wherever the visible text happens to start.
  • OpenAI removed the pipeline labels on 21 July, so this window into the mechanics has closed. The behaviour it documented has not.
More articles Become a partner