Skip to content
PreferiumJoin the waitlist

AI SearchAI Visibility

What a study of ChatGPT snippets found near the page heading

Resoneo examined ChatGPT retrieval in July 2026 and found short snippets often started near the H1. Useful evidence for clearer page openings, with limits on its scope.

A black Labrador rests one paw on a single raised paper file among endless rows of documents stretching into shadow.
Jump to a section

Resoneo’s analysis of ChatGPT retrieval offers a useful reason to inspect the top of a page: some retrieved snippets were short and often included the H1. That is a finding from a captured sample, not a permanent specification for how ChatGPT reads every website.

The retrieval paths in the study

A retrieval pipeline is the process that fetches web results for an answer. The study identified several routes: some retrieved Google results, while another queried OpenAI’s own web index, labelled labrador in the captured responses.

An index stores information about pages, such as their titles, addresses and snippets. In this study, the researchers could identify which retrieval route supplied a result. That made it possible to compare snippet formats across routes.

Until 21 July, ChatGPT’s server responses labelled which pipeline fetched each result. Resoneo read those labels across 1,249 conversations captured in July and published its analysis. Once the labels disappeared, researchers had less direct visibility into the retrieval path.

Independent analyst Suganthan Mohanadasan also investigated source selection in his source-selection census, correcting an earlier assumption when the July data became available.

Licensed and unlicensed sites appeared in the index

The earlier theory was that OpenAI’s own index only held content from publishers it had paid — Le Monde, WSJ, Condé Nast and the rest — and everyone else was reached by scraping Google.

The study found evidence against that assumption. Resoneo found hundreds of outlets with no OpenAI agreement being served from labrador with the same snippet format, the same length, and the same freshness as the licensed partners. Search Engine Journal’s write-up of the findings leads on exactly that point.

There is a contractual difference, but it sits somewhere else. Resoneo found a wordlim value attached to results — a cap on how many words of your page ChatGPT may quote in its answer. Known licensees came in at 100 words, a handful of outlets at 25, and the default for the rest of the web was 200. The unlicensed default is the most generous of the three.

Resoneo also argued that these results did not originate from Bing. One clue was title length: a quarter of labrador’s stored titles exceeded the 75 characters Bing displays. This is part of the researchers’ interpretation of the captured responses, not a published OpenAI specification.

The short snippets near the H1

In the studied results, labrador snippets were cut at 202 characters and drew from page text rather than the meta description. (The Google-scraping pipe still uses meta descriptions roughly a third of the time, so the tag is not dead. It’s just irrelevant to this one index.)

Those 202 characters break into three parts, at Resoneo’s median measurements: about 7 characters of whatever sits above your H1, then the H1 itself at around 51, then roughly 146 characters of body text. SEJ reports the H1 landed inside the snippet in 83.6% of the cases reviewed — 387 of 463.

The H1 was a frequent anchor in this sample, but text before it also occupied part of the snippet. A label above the heading, such as “GUIDES” or “CASE STUDY”, used about 18 characters in Resoneo’s data. Descriptive alt text from an image above the heading could use 50.

Keep that introductory material useful. These measurements are not a reason to remove descriptive alt text needed for accessibility.

Then there’s the plainer failure: Resoneo found roughly one page in seven has no H1 at all. If there’s no heading, the 202-character window starts wherever the visible text starts, which might be your cookie banner copy or a breadcrumb trail.

What to check

  • Does the page have exactly one H1, and does it say what the page is about in plain words?
  • What sits above the H1? Remove redundant labels where they add no value, while keeping meaningful navigation and accessible image descriptions.
  • Read the first sentence after the H1 on its own. Does it stand up as a summary, or does it begin with “In this article we’ll look at…”?
  • Don’t rely on your meta description to carry the page’s meaning into AI answers. Make the visible opening useful on its own too.
Want this running under your brand?Preferium AI Edge is a white-label platform agencies resell to their clients: your brand, your Stripe, your packages and prices. Registration opens to agencies from the waitlist first.Talk to usJoin the waitlistHow white label works

Build your agency on Preferium

Partner registration opens by invitation from the waitlist, and there is no date yet. Read the agency agreement and how partner billing works before you decide.

  1. Connect a client site
  2. Set the control level
  3. Run under your brand

Privacy choices

Choose which optional technologies Preferium AS may use. All of them are off until you choose.

Analytics: Google Analytics 4 counts page views. Google may receive the page address, referrer, network address, browser and device details, and online identifiers. Browser storage: _ga, _ga_*: Up to 730 days. Renewed on activity. The browser may shorten the storage period.

Provider: Google Ireland Limited. Google may transfer data to the United States.

See the cookie notice, the privacy notice and theterms.

Necessary technologies Used for site functions

Preferium AS and Cloudflare deliver the site and protect forms against abuse. A local preference remembers if you pause animation. When optional tracking is available, the site can also remember your documented privacy choices.

Consent receipt
Provider: Preferium AS. HttpOnly receipt that documents and retrieves your consent choice. Name in the browser: __Host-preferium_consent. Storage period: The receipt is valid for up to 180 days without rolling renewal.
Local privacy choices and pending rejections
Provider: Preferium AS. Local storage of privacy choices and pending rejections. The entry alone can never allow optional technologies; a valid server receipt is required. Name in the browser: preferium-consent-v2. Storage period: Until the entry is overwritten or the browser site data is cleared. No automatic timed deletion is configured.
Privacy choice synchronization between tabs
Provider: Preferium AS. The latest message that synchronizes privacy choices and pending rejections between tabs. The message can only close optional technologies and trigger a new server check. Name in the browser: preferium-consent-sync-v1. Storage period: Until the entry is overwritten or the browser site data is cleared. No automatic timed deletion is configured.
Motion pause preference
Provider: Preferium AS. Session storage restores the requested accessibility preference between pages in this tab. Nothing is sent to a server. Name in the browser: preferium-motion-paused. Storage period: Until this browser tab session ends.
Cloudflare Turnstile
Provider: Cloudflare. Abuse protection that loads only on forms where Turnstile is necessary. Storage period: Short-lived control value tied to a form submission.
Analytics

Helps us understand how the site is used, when you consent.

Provider: Google Ireland Limited. The data may include the page address, referrer, network address, browser/device, online identifiers and usage events.

_ga, _ga_*
Processes site usage for aggregated analytics after specific consent. Storage period: Up to 730 days. Renewed on activity. The browser may shorten the storage period.

You can withdraw your choice via Privacy choices. That stops further optional loading but does not recall data already sent to Google. We attempt to delete known first-party values; the browser may prevent deletion of third-party values.

How Google uses and is responsible for data · How Google uses information from partner sites · Google privacy policy