Correction: we cannot source the Intelligence Index figures in our Astra verdict, so we are removing them

OpenAI logoOpenAIVerdict changedSeptember 9, 2026Models
What happened
We cannot source the three Artificial Analysis Intelligence Index figures our Astra verdict quoted (61.2 for Astra, 65.7 for Fable 5.1, 63.1 for Opus 5), so we are removing them and the third-place ranking built on them.
Why it matters
That ranking was the lead condition on a buying recommendation, and an unsourceable benchmark inside a published verdict is a fabrication risk.
What to do
Re-run any Astra decision that rested on the third-place ranking. The verdict stays conditional on two constraints: EU data residency blocks Fast mode, and OpenAI disclosed decreased monitorability versus GPT-5.6 Sol.

Astra stays conditional. The rating does not move. What moves is the evidence under it: three benchmark figures we published are coming out, because we cannot source them.

What happened

On September 4 we published our GPT-6 Astra launch coverage and moved Astra from pending to conditional. The lead condition was "third on general intelligence," and it rested on three numbers we attributed to OpenAI's own launch comparison table: an Artificial Analysis Intelligence Index of 61.2 for Astra, 65.7 for Claude Fable 5.1 and 63.1 for Claude Opus 5.

We went looking for that table on September 9 and re-fetched every page involved. It is not where we said it was.

  • Artificial Analysis publishes different numbers in a different shape. Its own Fable 5.1 write-up gives Fable 5.1 a 66 at maximum effort, which it calls the highest score it has measured, and Opus 5 a 63 at maximum effort. GPT-5.6 Sol is 61. Whole numbers, max effort. Astra does not appear in that article at all.
  • OpenAI's GPT-6 Astra model page carries the specs and the rate card and no benchmark comparison of any kind. No index figure, no rival column.
  • OpenAI's API changelog calls Astra "our most capable model, built for the hardest end-to-end work" and publishes no benchmark table.
  • OpenAI's launch blog post returns HTTP 403 to our fetcher. A comparison table may well be on that page. We cannot read it, so we will not cite it. Our source list below still labels that page as carrying benchmark tables, from when we first filed it; we can no longer stand behind that label either.

The middle two are the primaries our own Astra entry cites. Neither carries the table we attributed the numbers to.

So the three figures come out, and the "third on general intelligence" ranking built on them comes out with them. We are not swapping in the Artificial Analysis numbers above as a replacement: those are max-effort scores for two rival models, and there is no published Astra score to rank against them.

Why it matters

An unsourceable benchmark ranking inside a published verdict is the exact failure we exist to avoid, and this one was load-bearing. "Third on general intelligence" was the first condition on a buying recommendation. It is the line that told readers to route general reasoning away from Astra.

A verdict with one fewer reason is worth more than a verdict with one invented reason. The figures also appear in our September 4 article and its summary. This notice retracts them there too.

What changes for you

Astra stays conditional. Two conditions hold it there, and we are explicit about how solid each one is today.

ConditionEvidence status
Fast mode is unavailable under EU data residencyRe-verified today. OpenAI's pricing docs state it in those words and route EU-residency requests to Standard processing
Monitorability decreased versus GPT-5.6 SolCarried forward, not re-verified. It comes from OpenAI's Astra safety overview, which we read on September 4; that page returns 403 to our fetcher today

Re-verified today against OpenAI's own docs and unchanged: $10 input and $50 output per 1M tokens short context, $20/$75 long context, Fast mode at $20/$100 and $40/$150. Context window 1,050,000 tokens, max output 128,000.

If you ruled Astra out because we ranked it third, that reason is gone. Re-run the decision on what survives: a computer-use, terminal and long-context specialist at a frontier price, with no vendor-neutral general-intelligence score published for it at all.

FAQ

Did the verdict change? No. Astra was conditional before this correction and is conditional after it. What changes is the evidence: the standing verdict summary loses the index figures and the third-place ranking, and keeps the two conditions above.

Why not just publish the Artificial Analysis numbers instead? Because they do not answer the question we asked. Artificial Analysis has published no Astra score, so its 66 and 63 rank Fable 5.1 and Opus 5 against each other, not against Astra. Substituting one set of numbers for another would repeat the original mistake with tidier sourcing.

Where did 61.2, 65.7 and 63.1 come from, then? We do not know, and that is the problem. Our own coverage attributed them to OpenAI's launch comparison table. That page is unreadable to us, and none of the OpenAI or Artificial Analysis pages we can read contains them. Until someone can point at a fetchable source, we treat all three as unsupported.

Should I still avoid Astra for general reasoning? That is now a judgement call rather than something our data settles. Claude Opus 5 and Claude Fable 5.1 hold the two highest Artificial Analysis scores we can verify. Astra has none. Absence of a score is not evidence of a low one.

conditionalprevious pick
conditionalnew pick

The Artificial Analysis Intelligence Index figures (61.2 / 65.7 / 63.1) and the third-place general-intelligence ranking could not be sourced against Artificial Analysis or either OpenAI primary cited on the entity, and are removed from the standing verdict. The rating holds at conditional on two constraints: Fast mode is unavailable under EU data residency (re-verified 2026-09-09), and OpenAI disclosed decreased monitorability versus GPT-5.6 Sol (read 2026-09-04, page unreadable today).

What to do

  1. 1 If you ruled Astra out on the strength of a third-place intelligence ranking, re-run that decision. The ranking is withdrawn and we have no replacement figure to offer.
  2. 2 If you quoted our 61.2 / 65.7 / 63.1 figures in an internal model-selection doc, strike them.
  3. 3 If you operate under EU data residency, treat Fast mode as unavailable on Astra and budget the Standard tier at $10/$50 per 1M tokens short context.
  4. 4 Budget long-context Astra work at $20/$75 per 1M tokens, double the short-context rate.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.