Executive summary

GPT‑6 Astra's ARC‑AGI‑3 scores are audited, dollar‑precise, and already public. Those figures are measured, not projected, and merit close scrutiny before any single "Astra score" is quoted.

ARC Prize's results page shows Astra scoring 62.7% on the provider-neutral Standard harness for $26,098. The same page records 99.9% on OpenAI's own Provider Adapter harness for $18,817, a cheaper run producing the higher score. That gap is worth sitting with before treating either number as "Astra's score." A separate figure of 15-20% for o3 at "very high compute" is a scaling projection rather than a benchmark result.

For anyone pricing a capability claim, four things follow:

  • Measured, estimated, and unknown are not the same evidence. ARC Prize itself distinguishes a finished, costed run from an asterisked partial estimate from a compute-scaled projection. Treating all three as "the score" misreads the source.
  • The only precedent for a score surviving into the next benchmark generation is thin and partly unmeasured. o3's ARC-AGI-1-to-2 transition is the sole documented case, and half of it is an estimate.
  • ARC-AGI-4 does not exist. There is no data, measured or estimated, on how the Provider Adapter score would hold up against a harder benchmark, because no one has built one yet.
  • Harness choice is not a neutral variable. R&D World's coverage notes the Provider Adapter preserves OpenAI-specific hidden reasoning state unavailable to competitors, meaning that score describes an OpenAI-only configuration, not a portable capability.

When the next vendor cites an ARC-AGI-3 number in a pitch, the question is which harness produced it, and whether anyone has run the Standard version.

For practitioners

A provider adapter that persists hidden reasoning state between calls is functionally a cache with elevated privileges: it lets the model reuse intermediate work — partial hypotheses about game rules, discarded action sequences — that a stateless harness discards after every call. R&D World's coverage describes it as compaction for longer conversations, which is the same tradeoff behind context-window management in any production agent stack: keep more state, spend more engineering effort maintaining it, and get a result that depends on infrastructure the next vendor over doesn't have access to.

That has a direct procurement implication. Any RFP response, competitive bake-off, or internal eval that cites an ARC-AGI-3 number should specify which harness produced it, the same way a benchmark of database throughput should specify whether the client had a warm connection pool. The choice of whether to grant an agent under test a persistent scratchpad across steps sits underneath any internal eval pipeline built the same way, and the ARC Prize data is a clean illustration of how much that single architectural decision moves the number, independent of anything that changed in the underlying model.

The absence of ARC-AGI-4 is not a caveat to footnote. It is the actual constraint on what can be built against these results. That Provider Adapter score describes performance on a benchmark whose environments are, per ARC Prize's own characterization reported by R&D World, bounded and deterministic. Nothing in ARC Prize's results page or its comparison table says what happens when the task space stops being bounded.

The practical move is to treat harness-conditioned scores the way a systems team treats a benchmark run on a warm cache: real, reproducible, and non-transferable until someone reruns it cold.

Deep dive

Model ARC-AGI-1 ARC-AGI-2 Change
GPT-6 Astra 98.5% 95.0% -3.5 pts
DeepSeek V4 Pro 0813 90.5% 61.3% -29.2 pts
o3-preview-low 75.7% ~4% (est.) -71.7 pts

Scores per ARC Prize's results table; o3's ARC-AGI-2 figure is the asterisked in-progress estimate ARC Prize publishes with its ARC-AGI-2 announcement, not a completed run.

The one generational transition ARC Prize has actually recorded runs the other direction from what a capability curve is supposed to do. ARC Prize's comparison table shows o3-preview-low scoring 75.7% on ARC-AGI-1, then falling to an asterisked, in-progress estimate near 4% on ARC-AGI-2. That is the only documented case in ARC Prize's data of a score crossing a benchmark generation at all. Recompute that gap and it is worth stating plainly: 75.7 percentage points to an estimated 4 is a model that could solve three in four ARC-AGI-1 tasks failing nearly all of ARC-AGI-2's. A separate figure of 15-20% for o3 at "very high compute" is a scaling projection, and ARC Prize does not publish the extrapolation method behind it. That is the entire base rate available for answering the question every Astra number implicitly raises: what happens to a saturating score when the benchmark gets harder.

One data point is a thin foundation for a forecast. The o3 estimate predates the Standard-versus-Provider-Adapter split that defines Astra's ARC-AGI-3 results. Nothing in ARC Prize's results page or its comparison table ties the two trajectories together analytically, and no source claims one predicts the other. Treating o3's ARC-AGI-1-to-2 collapse as a template for what Astra would do on a hypothetical ARC-AGI-4 assumes a consistency across model families and benchmark designs that the data does not establish. It would be equally defensible to argue the opposite: that ARC-AGI-2's design specifically targeted the failure modes of 2024-generation reasoning models, and that a benchmark built against Astra's architecture would degrade its score by a different amount, in a different direction, for reasons unconnected to o3.

What can be said without extrapolating past the chart is narrower and more useful. Astra fell from 98.5% to 95.0% moving from ARC-AGI-1 to ARC-AGI-2, a 3.5-point drop, per ARC Prize's comparison table.

Two frontier labs, the same generational jump, a 3.5-point loss against a 29-point loss. That gap is itself informative in a way a single case is not. It suggests the ARC-AGI-1-to-2 transition punishes something specific to how a model was trained or evaluated, something that does not scale uniformly across providers. What that specific thing is, training data overlap with ARC-AGI-1's public task set, brittleness to the interactive format ARC-AGI-2 introduced, something else, is not addressed by any source here, and assigning a mechanism to it would be reasoning past what ARC Prize has published.

That uncertainty compounds rather than resolves at the ARC-AGI-3-to-4 boundary, because ARC-AGI-3 introduced a structural change ARC-AGI-1-to-2 did not: the harness itself became a variable. Astra's split between the two harnesses, reported on ARC Prize's Astra results page, has no analog in the ARC-AGI-1-to-2 data, where every score in the comparison table is a single number per model with no harness distinction disclosed. A hypothetical ARC-AGI-4 would need to specify not only its task design relative to ARC-AGI-3's bounded, deterministic environments, a limitation ARC Prize itself flagged, per R&D World's coverage, but whether it preserves the harness-dependent scoring that already makes that split a plural rather than a singular.

The forecasting move that is actually available is a floor, not a point estimate. The widest recorded drop in ARC Prize's own table is DeepSeek V4 Pro 0813's 29.2 points.

The number that actually matters for anyone pricing Astra's capability is the 29-point spread that already exists between two models on the transition that has already happened.