Skip to content
PreferiumJoin the waitlist

AI SearchAI Visibility

Three frontier AI models now search the web about equally well

GPT-6 Astra shipped on 3 September with 91.5% on BrowseComp, a hard web-research test. Claude Opus 5 gets 90.8%. The OpenAI model Astra replaces got 90.4%. Retrieval skill has stopped being the variable.

Article illustration
Three horses, one white, one grey, one black, gallop side by side across a dusty plain, all exactly level with each other.
AI Search · Preferium briefing

TL;DR: OpenAI released a new flagship model, GPT-6 Astra, on 3 September, and talked about it in terms of artificial general intelligence. Buried in the benchmark table is a number that matters more to anyone running a website. On a test that measures how well an AI can dig hard-to-find facts out of the open web, Astra scores 91.5%. Anthropic’s Claude Opus 5 scores 90.8%. The OpenAI model Astra replaces scores 90.4%. Three different systems from two different companies, separated by roughly one point. Being good at looking things up online is no longer what separates one AI assistant from another. So when your brand turns up in one assistant’s answer and not another’s, the model’s ability to search is not the reason any more. Something about your site, or about what that particular company has collected from it, is.

First, what a benchmark actually is

A benchmark is a fixed set of test questions everyone runs their model against, so the scores can be lined up side by side. Like a standardised exam, it’s useful and also gameable, and the company selling the model is usually the one publishing its own score.

The specific exam here is called BrowseComp. It’s a web-research test: the model is given a question whose answer is genuinely hard to find, has to go out and search the live web for it, follow leads across several pages, and come back with the right answer. It’s not a memory test. The model has to actually go and look.

That distinction matters for the rest of this. BrowseComp measures the thing that determines whether an AI assistant can find your page when someone asks it a question in your category.

The numbers

From the benchmark tables published alongside the launch and compiled by Vellum and Yotta Labs:

  • BrowseComp: GPT-6 Astra 91.5%, Claude Opus 5 90.8%, GPT-5.6 Sol 90.4%
  • OSWorld 2.0 (operating a desktop computer): Astra 72.6%, Opus 5 70.2%, Sol 65.7%
  • ScreenSpot-Pro (finding the right thing on a screen): Astra 92.7%, Claude Fable 5 87.3%, Sol 76.9%

Look at the spread. On web research, the gap between the newest OpenAI model and a competitor’s model is 0.7 points. On operating a computer, the same comparison is a 16-point gap on ScreenSpot-Pro.

Web research looks close to finished as a competitive dimension. Computer use clearly isn’t.

What converged retrieval changes

Most agencies still explain a missing citation with some version of “that engine isn’t as good yet.” For a while that was a reasonable guess. It’s now a hard one to defend, at least among the frontier models: they are all within a rounding error of each other at finding things on the open web.

Which leaves the differences that are actually yours to influence, plus the ones that aren’t:

What each company has collected. ChatGPT, Claude, Perplexity and Gemini each run their own crawler, on their own schedule, with their own idea of what’s worth keeping. A page that OpenAI’s crawler swept last week and Anthropic’s crawler hasn’t visited since June will show up in one and not the other. Same model quality, different stock of pages.

Whether the crawler could read the page at all. This is the boring, mechanical half, and it’s where most of the recoverable losses sit. A page that hides its links behind JavaScript, or blocks a specific crawler in robots.txt, or returns a slightly different response to a bot, is invisible to that engine regardless of how clever the model is. We covered a 41-day test where AI crawlers found zero JavaScript-only pages a couple of weeks ago.

Licensing and deals. Some of it you can’t touch, and pretending otherwise is dishonest.

So the question worth asking about a weak engine is what it physically holds of your site, and when it last looked.

The other number worth noting

OpenAI reports its hallucination rate dropping to 4.2%, against 12.2% for the model Astra replaces. Those are the company’s own figures, run at maximum effort settings, so treat them as a direction rather than a measurement. Artificial Analysis, testing independently, also found a substantial reduction in hallucination.

A model less willing to guess is a model more willing to go and fetch. That’s more requests hitting your pages, and it raises the cost of a page that loads slowly, answers vaguely, or contradicts itself across sections.

Don’t buy the framing

OpenAI’s president floated the idea that people will look back on this release as the start of the AGI era, and The New Stack ran that framing in its headline. Artificial Analysis, which runs its own tests rather than reprinting vendor tables, puts Astra at 61 on its Intelligence Index. That’s level with the model it replaces, and behind Claude Fable 5.1 at 66. Astra is also 2.5x more expensive per token than its predecessor.

There are real improvements here in agentic work and in hallucination rates. There isn’t a step change in raw intelligence, and the pricing has gone the wrong way.

What to do

Pull your visibility numbers per engine rather than as one blended score, and treat a gap between two engines as a technical question with a findable answer. Then check the mechanical layer for the engine you’re losing: is its crawler getting a clean, link-rich HTML response from your key pages, or is it getting a shell?

That check is per page and per crawler, and it goes stale every time you deploy. Measured per engine, a gap between two assistants shows up as two different lines instead of one average.

Key takeaways

  • On BrowseComp, a hard web-research test, GPT-6 Astra scores 91.5%, Claude Opus 5 scores 90.8% and GPT-5.6 Sol scores 90.4%. Frontier models now find things on the web about equally well.
  • Model quality no longer explains why your brand appears in one AI assistant and not another. What each company crawled, and whether your pages were readable when it did, does.
  • Computer use is still a real gap between models — a 16-point spread on ScreenSpot-Pro against 0.7 points on BrowseComp.
  • Lower hallucination rates mean more live fetches of your pages, not fewer.
  • Artificial Analysis rates Astra level with the model it replaces on overall intelligence, and behind Claude Fable 5.1, at 2.5x the token price of its predecessor. The AGI framing is marketing.
  • Track AI visibility per engine. A blended score hides exactly the differences you can act on.
Want this running under your brand?Preferium AI Edge is a white-label platform agencies resell to their clients: your brand, your Stripe, your packages and prices. Registration opens to agencies from the waitlist first.Talk to usJoin the waitlistHow white label works

Build your agency on Preferium

Partner registration opens by invitation from the waitlist, and there is no date yet. Read the agency agreement and how partner billing works before you decide.

Connect a client siteSet the control levelRun under your brand

Privacy choices

Choose which optional technologies Preferium AS may use. All of them are off until you choose.

Analytics: Google Analytics 4 counts page views. Google may receive the page address, referrer, network address, browser and device details, and online identifiers. Browser storage: _ga, _ga_*: Up to 730 days. Renewed on activity. The browser may shorten the storage period.

Provider: Google Ireland Limited. Google may transfer data to the United States.

See the cookie notice, the privacy notice and theterms.

Necessary technologies Used for site functions

Preferium AS and Cloudflare deliver the site and protect forms against abuse. A local preference remembers if you pause animation. When optional tracking is available, the site can also remember your documented privacy choices.

Consent receipt
Provider: Preferium AS. HttpOnly receipt that documents and retrieves your consent choice. Name in the browser: __Host-preferium_consent. Storage period: The receipt is valid for up to 180 days without rolling renewal.
Local privacy choices and pending rejections
Provider: Preferium AS. Local storage of privacy choices and pending rejections. The entry alone can never allow optional technologies; a valid server receipt is required. Name in the browser: preferium-consent-v2. Storage period: Until the entry is overwritten or the browser site data is cleared. No automatic timed deletion is configured.
Privacy choice synchronization between tabs
Provider: Preferium AS. The latest message that synchronizes privacy choices and pending rejections between tabs. The message can only close optional technologies and trigger a new server check. Name in the browser: preferium-consent-sync-v1. Storage period: Until the entry is overwritten or the browser site data is cleared. No automatic timed deletion is configured.
Motion pause preference
Provider: Preferium AS. Session storage restores the requested accessibility preference between pages in this tab. Nothing is sent to a server. Name in the browser: preferium-motion-paused. Storage period: Until this browser tab session ends.
Cloudflare Turnstile
Provider: Cloudflare. Abuse protection that loads only on forms where Turnstile is necessary. Storage period: Short-lived control value tied to a form submission.
Analytics

Helps us understand how the site is used, when you consent.

Provider: Google Ireland Limited. The data may include the page address, referrer, network address, browser/device, online identifiers and usage events.

_ga, _ga_*
Processes site usage for aggregated analytics after specific consent. Storage period: Up to 730 days. Renewed on activity. The browser may shorten the storage period.

You can withdraw your choice via Privacy choices. That stops further optional loading but does not recall data already sent to Google. We attempt to delete known first-party values; the browser may prevent deletion of third-party values.

How Google uses and is responsible for data · How Google uses information from partner sites · Google privacy policy