AI SearchAI Visibility
Three frontier AI models now search the web about equally well
GPT-6 Astra shipped on 3 September with 91.5% on BrowseComp, a hard web-research test. Claude Opus 5 gets 90.8%. The OpenAI model Astra replaces got 90.4%. Retrieval skill has stopped being the variable.

TL;DR: OpenAI released a new flagship model, GPT-6 Astra, on 3 September, and talked about it in terms of artificial general intelligence. Buried in the benchmark table is a number that matters more to anyone running a website. On a test that measures how well an AI can dig hard-to-find facts out of the open web, Astra scores 91.5%. Anthropic’s Claude Opus 5 scores 90.8%. The OpenAI model Astra replaces scores 90.4%. Three different systems from two different companies, separated by roughly one point. Being good at looking things up online is no longer what separates one AI assistant from another. So when your brand turns up in one assistant’s answer and not another’s, the model’s ability to search is not the reason any more. Something about your site, or about what that particular company has collected from it, is.
First, what a benchmark actually is
A benchmark is a fixed set of test questions everyone runs their model against, so the scores can be lined up side by side. Like a standardised exam, it’s useful and also gameable, and the company selling the model is usually the one publishing its own score.
The specific exam here is called BrowseComp. It’s a web-research test: the model is given a question whose answer is genuinely hard to find, has to go out and search the live web for it, follow leads across several pages, and come back with the right answer. It’s not a memory test. The model has to actually go and look.
That distinction matters for the rest of this. BrowseComp measures the thing that determines whether an AI assistant can find your page when someone asks it a question in your category.
The numbers
From the benchmark tables published alongside the launch and compiled by Vellum and Yotta Labs:
- BrowseComp: GPT-6 Astra 91.5%, Claude Opus 5 90.8%, GPT-5.6 Sol 90.4%
- OSWorld 2.0 (operating a desktop computer): Astra 72.6%, Opus 5 70.2%, Sol 65.7%
- ScreenSpot-Pro (finding the right thing on a screen): Astra 92.7%, Claude Fable 5 87.3%, Sol 76.9%
Look at the spread. On web research, the gap between the newest OpenAI model and a competitor’s model is 0.7 points. On operating a computer, the same comparison is a 16-point gap on ScreenSpot-Pro.
Web research looks close to finished as a competitive dimension. Computer use clearly isn’t.
What converged retrieval changes
Most agencies still explain a missing citation with some version of “that engine isn’t as good yet.” For a while that was a reasonable guess. It’s now a hard one to defend, at least among the frontier models: they are all within a rounding error of each other at finding things on the open web.
Which leaves the differences that are actually yours to influence, plus the ones that aren’t:
What each company has collected. ChatGPT, Claude, Perplexity and Gemini each run their own crawler, on their own schedule, with their own idea of what’s worth keeping. A page that OpenAI’s crawler swept last week and Anthropic’s crawler hasn’t visited since June will show up in one and not the other. Same model quality, different stock of pages.
Whether the crawler could read the page at all. This is the boring, mechanical half, and it’s where most of the recoverable losses sit. A page that hides its links behind JavaScript, or blocks a specific crawler in robots.txt, or returns a slightly different response to a bot, is invisible to that engine regardless of how clever the model is. We covered a 41-day test where AI crawlers found zero JavaScript-only pages a couple of weeks ago.
Licensing and deals. Some of it you can’t touch, and pretending otherwise is dishonest.
So the question worth asking about a weak engine is what it physically holds of your site, and when it last looked.
The other number worth noting
OpenAI reports its hallucination rate dropping to 4.2%, against 12.2% for the model Astra replaces. Those are the company’s own figures, run at maximum effort settings, so treat them as a direction rather than a measurement. Artificial Analysis, testing independently, also found a substantial reduction in hallucination.
A model less willing to guess is a model more willing to go and fetch. That’s more requests hitting your pages, and it raises the cost of a page that loads slowly, answers vaguely, or contradicts itself across sections.
Don’t buy the framing
OpenAI’s president floated the idea that people will look back on this release as the start of the AGI era, and The New Stack ran that framing in its headline. Artificial Analysis, which runs its own tests rather than reprinting vendor tables, puts Astra at 61 on its Intelligence Index. That’s level with the model it replaces, and behind Claude Fable 5.1 at 66. Astra is also 2.5x more expensive per token than its predecessor.
There are real improvements here in agentic work and in hallucination rates. There isn’t a step change in raw intelligence, and the pricing has gone the wrong way.
What to do
Pull your visibility numbers per engine rather than as one blended score, and treat a gap between two engines as a technical question with a findable answer. Then check the mechanical layer for the engine you’re losing: is its crawler getting a clean, link-rich HTML response from your key pages, or is it getting a shell?
That check is per page and per crawler, and it goes stale every time you deploy. Measured per engine, a gap between two assistants shows up as two different lines instead of one average.
Key takeaways
- On BrowseComp, a hard web-research test, GPT-6 Astra scores 91.5%, Claude Opus 5 scores 90.8% and GPT-5.6 Sol scores 90.4%. Frontier models now find things on the web about equally well.
- Model quality no longer explains why your brand appears in one AI assistant and not another. What each company crawled, and whether your pages were readable when it did, does.
- Computer use is still a real gap between models — a 16-point spread on ScreenSpot-Pro against 0.7 points on BrowseComp.
- Lower hallucination rates mean more live fetches of your pages, not fewer.
- Artificial Analysis rates Astra level with the model it replaces on overall intelligence, and behind Claude Fable 5.1, at 2.5x the token price of its predecessor. The AGI framing is marketing.
- Track AI visibility per engine. A blended score hides exactly the differences you can act on.