SINCE MAY 2026
Aggregating from major public AI benchmarks

ONE SCORE.
INFINITE CLARITY.

AGI Ranker measures how close each frontier AI is to AGI. One transparent score (0-100) per model, distilled from 10 public benchmarks. Score 100 marks the AGI threshold.

Independent verification preferred over lab self-reports. Scores we can't verify are flagged, not invented. Every correction is logged publicly.

TOP MODEL
-
-
FRONTIER MODELS
-
AGI Score ≥ 65
BENCHMARKS AGGREGATED
-
Live benchmarks (TBA excluded)
Distance to AGI
-
pts
Top model's gap to AGI Score = 100
Top Model Breakdown
-
Thinking
44% weight
-
Fluid Reasoning + World Knowledge
Doing
43% weight
-
Agency (tools, planning, coding) + Visual Reasoning
Communicating
13% weight
-
Language Production + Visual Reasoning output
REAL-TIME RANKINGS

The AGI Ranker Leaderboard

100 = AGI
 
# Model AGI Arena Elo Coverage Price ·
Scores normalized to a per-benchmark ceiling: measured human parity where a human study exists, benchmark maximum where none does (see Methodology). See full methodology →
CUSTOMIZE THE FORMULA

Interactive Model Explorer

Adjust domain weights and instantly see how the AGI Score changes. Transparency at its core.

Total: 100%
SELECT MODELS TO COMPARE (max 4)
Domain Proficiency Radar
0 models selected
TRANSPARENT METHODOLOGY

How the AGI Score is Calculated

A transparent, reproducible composite index designed for maximum signal and minimum noise.

1
Data Ingestion & Curation
We aggregate scores from major public AI benchmarks: GPQA Diamond, HLE, ARC-AGI-2, AIME 2025, SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking, LiveBench, LiveBench Agentic Coding, and MMMU-Pro. Each cell is cited so any score can be traced back to its origin. We do not run benchmarks ourselves - we aggregate what others have already published. Source quality is tiered: independent benchmark leaderboards (Tier 1) count fully at 1.00×; third-party evaluators with provider involvement (Tier 2) at 0.85×; self-reports from labs with verified track records (Tier 3) at 0.75×; non-verified labs and commentary content (Tier 4) are excluded from scoring entirely.

Note on the AA Intelligence Index (removed June 2026): We previously carried Artificial Analysis's composite Intelligence Index as a Language signal. We have removed it from the AGI Score. Their v4.1 revision turned it into an explicitly agentic composite (GDPval-AA, Terminal-Bench 2.1, τ³-Banking, SciCode, HLE, GPQA and more) that re-bundles benchmarks we already score directly - keeping it would double-count those signals across components and mislabel agency as language. Language now rests on LiveBench; a dedicated language/writing benchmark is on our roadmap to restore a second independent source. We continue to track the AA Index as an external reference.
2
Normalization to Best-Human
Revised in v2.0.0: each raw benchmark score is rescaled by a ceiling. Where a published human study exists under a protocol comparable to the models’, that measurement is the ceiling and 100 means human parity. Where none exists, the ceiling is the benchmark maximum and 100 means a perfect score, with no human claim attached. Today exactly one scored benchmark carries a measured human ceiling, GPQA Diamond at 0.81, and the other nine are scored against the benchmark maximum, so the AGI Score is currently a mixed scale rather than a pure human comparison. Three further benchmarks passed the same ceiling audit and none of them is on the board: OSWorld (0.72) was retired in v2.0.0, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. We previously described every ceiling as best-human; that was not accurate for most of them, and correcting it is the main reason scores moved in v2.0.0. AGI Score = 100 remains our marker for the genesis of AGI, per our working definition in full: “AGI is an artificial intelligence that surpasses the best human on every purely brain-based intellectual task with no involvement of a physical body.” Scores climb past 100 as the AI grows from genesis into super-human (ASI-direction) territory; we preserve those scores rather than clamping, so the index stays informative in an ASI / post-AGI world.
3
5-Component Aggregation
Benchmarks roll up into 5 cognitive components AGI must master: Agency (35%), Fluid Reasoning (29%), World Knowledge (15%), Visual Reasoning (11%), Language Production (10%). Within each component, the score is the weighted mean of all the model's contributing benchmarks; each benchmark's effective weight is sub_weight × source_tier_multiplier. Sub-weights are fixed per benchmark and reflect each benchmark's relative importance within a component (e.g., ARC-AGI-2 carries 25% of Reasoning, GPQA carries 20%). The saturation rule kicks in dynamically: if the IQR of normalized scores among the top-10 ranked models on a benchmark drops below 5 percentage points, that benchmark's sub-weight is halved and the freed weight is redistributed to non-saturated benchmarks in the same component - keeping the component score responsive to whichever benchmarks still discriminate frontier models. Source tiering applies a multiplier per cell based on source independence (T1 1.00×, T2 0.85×, T3-Verified-Lab 0.75×; T4 excluded from scoring entirely). Component weights are user-adjustable in the Explorer below.
AGI Score = Σ (wᵢ × Cᵢ) / Σ wᵢ
where Cᵢ = component score (shrunk toward median if thin-coverage), wᵢ = component weight
4
Sparse-Data Handling (Asymmetric Pull-Down + Renormalize)
Benchmark coverage is uneven. Within each component we compute the weighted mean over present benchmarks (renormalizing the denominator). Then asymmetric pull-down shrinkage: if a model's coverage in a component is below 60% AND the raw score is above the population median, we pull it toward the median proportionally to coverage shortfall. Low scores with thin coverage stay low - no reward for hiding weaknesses. High scores with thin coverage get discounted toward "typical." Models need ≥3 components and ≥5 benchmarks (with Reasoning + Agency required) to be ranked. Cross-component coverage floor (v1.4): a model with 2+ components below their coverage thresholds gets demoted from RANKED to PROVISIONAL - even if its present cells are exceptional. Thresholds: 30% sub-weight for World Knowledge / Visual Reasoning / Language; 40% for the core components (Fluid Reasoning and Agency, the load-bearing AGI dimensions per the locked definition). This prevents selection-bias rankings: a model can only claim near-AGI position if its evidence is reasonably broad across capability dimensions. Refined 2026-08-07 (MCP-0005): Ranked status requires both core components, Agency and Fluid Reasoning, to meet their existing minimum coverage on the declared base weights; the one-thin-component allowance applies only to non-core components. Official ordinal ranks, the Top 3 and the headline leader are drawn from Ranked models only. A Provisional model keeps its full capability estimate, shown prominently, but receives no official rank until its core evidence meets the standard. The saturation rule separately protects against dead benchmarks: any benchmark whose IQR among the top-10 ranked models drops below 5pp gets its sub-weight halved and redistributed to non-saturated benchmarks in the same component.
5
Specialty Rankings & Custom View
Above the leaderboard, tabs switch the ranking between the canonical AGI Score and task-focused views. Each specialty (currently Coding and Knowledge) re-ranks models using only the benchmarks that test that capability, with within-set sub-weights renormalized to 100%. Specialty scores apply the same source-tier weighting and asymmetric pull-down shrinkage as the main score - so a model with thin coverage in a specialty is pulled toward the population median for that specialty, never the other way. Specialties are shown on a separate 0 to 10 index, not on the AGI scale, because they are a different kind of claim: 10 means a perfect score on every benchmark in that set, and it is not a statement about human performance. Most of these benchmarks have no published human study, so there is no human bar to measure distance to. A specialty index is never AGI - AGI requires the full battery, which only the canonical AGI Score measures. The Custom tab (visible when you adjust the Explorer's weight sliders) shows the AI Score under your settings; the AGI tab always reverts to default weights, keeping the canonical view unambiguous.
Calibration Constants & Operational Discipline
reproducibility addendum
Full reproducibility requires three things beyond the formula: the exact numeric thresholds we apply, the within-set sub-weights used in specialty views, and the discipline we hold against common mis-attribution patterns. All disclosed below.
Constants unchanged since v1.4, reviewed and re-confirmed for v2.0.0. These are operational values, not permanent invariants. Future methodology versions may revise specific constants as the methodology evolves and as benchmark coverage matures; substantive changes will be noted in the version history.
Numeric thresholds
Constant Value Used in
Asymmetric shrinkage coverage threshold (main AGI Score & specialties with 3+ benchmarks) 60% Step 4
Asymmetric shrinkage coverage threshold (specialties with 2 benchmarks, e.g. Reasoning) 85% Step 5
Saturation IQR threshold (top-10 ranked models) 5 pp Step 4
Coverage floor - Fluid Reasoning & Agency (core) 40% Step 4
Coverage floor - Knowledge / Visual Reasoning / Language 30% Step 4
Max thin components allowed in the RANKED tier 1 Step 4
Eval-date grace window vs. model release 30 days below
Specialty sub-weight splits
Within each specialty tab, only the relevant benchmarks contribute, with within-set sub-weights renormalized to sum to 100. The splits are fixed. Minimum cells required for a model to enter a specialty ranking: 2. Reasoning previously ran at 1, on the grounds that it had only two benchmarks; that exception is what let it publish a one-benchmark ranking under a two-benchmark label, and it was withdrawn in v2.0.0 rather than kept. Models below the minimum but with at least one relevant cell appear in a separate "insufficient evidence" section below the ranking.
Note on Coding methodology (v2.0): Real-world coding in 2026 is agentic - it happens via tools like Claude Code, Codex, Cursor, not bare LLM completion, and the composition reflects that. SWE-bench Verified and Terminal-Bench 2.1 are both run under standardized agentic harnesses, and LiveBench Agentic Coding joins as a third independent evaluator. SWE-bench Pro was retired in v2.0.0: none of our cells for it came from a clean principal source, and the owner leaderboard has not been refreshed in months. We would rather drop a benchmark than keep one we cannot stand behind.
Contamination note for SWE-bench Verified: Public-test-set SWE-bench scores can be inflated by training-data leakage. Private-test-set evaluators (e.g., vals.ai) tend to report 5-15pp lower numbers than public-leaderboard aggregators for the same model. We prefer T1 private-test sources where available; see the Corrections Log for the DeepSeek V4 Pro GPQA correction (0.901 self-report -> 0.729 independent T2) as a concrete example of the contamination correction in action.
Pre-release builds: preview, beta and release-candidate models are ranked normally and flagged. They meet the same evidence requirements as any other entry and are not given easier treatment. Holding them off the board would mean omitting models people are already using, and parking them in Provisional would misuse a tier that means thin evidence rather than unsettled product. What does deserve saying is that the weights behind a preview can change before general availability, so the score describes the build that was measured, not a finished product. Those entries carry a Pre-release tag. If a lab ships a materially different build under the same name, we treat that as a new model rather than silently updating the old row.
Source-choice sensitivity, GPQA Diamond and MMMU-Pro: both moved from Artificial Analysis to vals.ai on 2026-07-26, to break a single-evaluator dependency: Artificial Analysis had been supplying 97% of Knowledge and 100% of Visual Reasoning. As with Terminal-Bench we measured the gap before switching rather than after. On GPQA Diamond the two boards agree closely, with a −0.30pp offset and 1.38pp scatter across all 20 models, and vals additionally publishes a standard error per row. On MMMU-Pro vals runs +5.71pp higher, with 1.67pp scatter and vals higher on all eleven comparable models, so scores on that benchmark rise by design rather than by accident. One row is held rather than used: vals reports Grok 4.5 at 61.8 on MMMU-Pro, below Grok 4.3 (83.1), below Grok 4 (76.3) and below xAI’s own non-reasoning variants. A flagship scoring beneath its predecessor and beneath the non-reasoning builds is not a capability measurement, so we cannot say what that row measured and do not use it. Grok 4.5 therefore has no Visual Reasoning cell. Evaluator mixing is not allowed to masquerade as capability, but as of 2026-08-05 it is priced rather than banned outright. When the canonical evaluator does not list a model, a column whose measured cross-evaluator scatter is at most 2pp may admit a fallback cell: the shadow evaluator’s value converted by the column’s measured offset, carrying a widened interval (±5.61pp on GPQA Diamond, ±5.55pp on MMMU-Pro — calibrated so the largest cross-evaluator miss we have ever observed sits inside it), permanently labelled on the cell, and automatically replaced by the canonical value the day it publishes, even when the canonical value is less flattering. Terminal-Bench 2.1, at 4.54pp scatter, does not qualify and stays single-evaluator. Uncorrected mixing — taking vals for eleven models and Artificial Analysis for the twelfth as if the numbers were interchangeable — remains ruled out for exactly the original reason: it would put two measurement stacks in one column and call the difference capability.
Source-choice sensitivity, Terminal-Bench 2.1: we report this separately rather than folding it into the uncertainty band, because it is a different kind of doubt. Terminal-Bench 2.1 is published by two independent evaluators, and they disagree systematically: across the 19 models present on both boards, vals.ai scores 7.77 points lower on average, and lower on 18 of the 19. Matching the declared effort level does not close it; Claude Opus 4.8 is max-against-max and still 12.7 points apart. We source from vals.ai, so our Terminal-Bench column sits about 7.8 points below where it would sit had we chosen the other board. What that is worth on the AGI Score: 0.71 points on average, 1.01 at most, because Terminal-Bench is 25% of Agency and Agency is 35% of the score. Had we chosen the other evaluator, only 2 of 20 models would change rank, and those two are 0.01 apart in any case. The published interval already spans roughly ±9.5 points, so this sits comfortably inside it; folding it in would count the same doubt twice. The remaining per-model disagreement after removing that systematic gap, 4.54 points, is inside the interval, where it belongs.
Harness disclosure for Terminal-Bench 2.1: Agentic-benchmark scores depend on the scaffold used to execute tasks. Every Terminal-Bench 2.1 cell comes from vals.ai using the Terminus 2 harness, pass@1 over 89 tasks. We deliberately use one evaluator rather than blending several, because two evaluators running the same 89 tasks do not agree: across the 19 models present on both the vals.ai and Artificial Analysis boards, vals.ai sits an average of 7.8 points lower, and it is lower on 18 of those 19. Matching the declared effort level does not close the gap. Blending the two would produce a number neither evaluator published, so we pick one, name it, and carry the disagreement into the uncertainty band instead. Native-agent rows (Claude Code, Codex CLI, Cursor CLI) are excluded entirely: they measure a product, not a model. Real users running each model with its own agentic tool may see different relative performance than these scores predict.
Specialty Benchmarks & within-set sub-weights
Coding SWE-bench Verified 40 · Terminal-Bench 2.1 35 · LiveBench Agentic Coding 25
Reasoning ARC-AGI-2 70 · AIME 2025 30
Knowledge HLE 40 · GPQA Diamond 35 · MMMU-Pro 25
Tool Use withdrawn in v2.0.0 — OSWorld, BrowseComp and the two Tau-bench domains all left the board in the v2 evidence review: OSWorld had 2 clean cells out of 22, BrowseComp had no uniform browsing harness across models, and the Tau-bench retail and airline pair is superseded by τ³-Banking. That leaves a single clean non-coding tool-use benchmark, and a one-benchmark ranking is exactly what our two-benchmark minimum exists to prevent. The tab returns when a second one exists.
Value view: how value for money is derived
The Value tab pairs each model's published capability for the selected area (Overall = the AGI Score; otherwise the specialty score above - the exact same number that area's tab shows) with an estimated API cost, so capability can be weighed against price.
Estimated cost is a blended price per 1M tokens across four token classes — new context, cache writes, cache reads and output — at a typical mix for that workload, using each lab’s published list prices. The coding mix is measured from real agentic coding sessions across three vendors: about 97.6% of tokens are context read back out of cache and only 0.24% are output, which is why the cached-input rate dominates a coding bill. How broad that measurement is, precisely. It covers three vendors and three coding agents, but weighted by tokens the sample is about 87% a single agent, so read it as one workload measured carefully rather than an average across the field. Measured per agent, the share of new context ranges from about 1.8% to 7.1% — the agents differ more than the vendors do. We apply one mix to every model on purpose: the Coding score itself ranks models under a single standardised harness, and a per-vendor mix would price each model against a different workload. The other workloads are stated assumptions, not measurements. Where a lab publishes no cached-input rate we leave the model out of a cached workload rather than estimate one. Some providers also bill cache storage by the hour; because a blended per-token price has no time dimension, a model charging hourly storage for the caching mode we price is left out too, rather than have us invent how long a cache is held. No model is currently affected — every provider we verified charges nothing to hold a cache in the mode we price. Prices are standard public list prices at the base context band — no promotional, volume or data-sharing tiers. One thing the base band does not settle is how long a cache is kept: some providers charge more to hold a cache for an hour than for five minutes. That is a property of the workload rather than of the model, so we take the tier the workload actually needs — for coding, measured, that is the one-hour tier, because an agentic session holds its context far longer than five minutes and a five-minute cache would simply expire and be rewritten. For the workloads whose retention we have not measured we keep the base tier rather than guess one. Long-context surcharges, applied only where we could measure them. Most providers charge roughly double once a single prompt crosses a threshold — 200K tokens at xAI and Google, 272K at OpenAI, 512K at MiniMax — and the whole request is repriced when it does. What that costs depends on how often a workload crosses the line, which is a property of the coding agent, not of the model, so it has to be measured per vendor. We have measured it for Grok 4.5 (38.4%) and Grok 4.6 (44.6%) from real billing records, and their coding costs here include that uplift. Grok 4.3, MiniMax M3 and Gemini 3.1 Pro also have a surcharge and are shown without one, because we have no usage of them to measure and will not invent a figure — their true coding cost is higher than shown by an unknown amount. Anthropic, the Gemini Flash models and ChatGPT 5.5 Pro have no surcharge at all, which we verified rather than assumed. It is an estimate of typical cost, not a measured per-task bill.
Value for money is the extra capability a model delivers over the weakest option shown, per estimated dollar: (score − lowest score in view) / cost. We subtract that floor on purpose - a capability score has no true zero (a coding score of 0 is not "no value"), so raw score-per-dollar would over-reward the cheapest model no matter how capable it is. Subtracting the floor measures the extra capability you actually buy.
Each area surfaces three picks: Top (highest capability), Budget (cheapest to run), and Best value (the most extra capability per dollar). The ranking shows the picks rather than a raw value number, which is ambiguous to read in isolation.
Variant-attribution discipline
A cell enters scoring only when the source page's row label identifies the variant unambiguously. The lab's named max-tier configuration (OpenAI Pro, Anthropic thinking, etc.) must appear in the source label. Effort-knob suffixes are not variants.
Accept rows labeled e.g. "Claude Opus 4.7 (thinking)" or "GPT-5.5 Pro".
Reject rows like "GPT-5.5 (xHigh)", "GPT-5.5 (High)", "DeepSeek V4 Pro (Max)" - these are base or default config with an effort knob, not the separately-priced Pro / thinking SKU.
Reject generic family rows where a single row on the source cannot distinguish base from Pro / thinking.
Eval-date sanity check
A cell whose eval_date predates the model's release_date by more than 30 days is rejected as a mis-attribution. The 30-day grace window covers pre-release lab evaluations; anything older almost certainly tested a preceding model that happens to share part of the name.
Two update clocks: evidence vs calibration
Adopted 2026-08-06: this board distinguishes updating the models being measured from recalibrating the measuring instrument. Evidence releases are frequent and run on a locked calibration: they may move only models that received new benchmark results, and the change is traceable to named cells. Calibration releases are scheduled, explicitly labelled, and are the only releases allowed to change benchmark discrimination weights (the saturation mechanism), which can move every model at once. The current calibration is pinned from 2026-08-05.
Why: benchmark reweighting is a legitimate response to benchmarks losing discriminating power at the frontier, but arriving unannounced inside a routine data update it reads as instability rather than method. Announced and labelled, it reads as what it is: a deliberate recalibration, explained in the changelog with before and after.
Active-roster governance
Adopted 2026-08-05: AGI Ranker tracks up to the best 4 models from each lab, including sideways SKUs (mid-ladder and specialized variants), not only flagship successions. New releases enter on arrival. When a lab’s count exceeds 4, the lowest-scoring model rotates off the board. A new entry that cannot yet be scored is protected from rotation until it becomes scoreable, or for 8 weeks, whichever comes first. Reasoning-effort settings never create separate rows.
The rule applies prospectively from its adoption date: models already on the board are grandfathered and rotate only under the rule’s normal operation going forward. The previous latest-two-generations rule governed the roster through v2.0.22 and remains documented in the version history. The specialist-slot principle carries over: a narrow branch earns its slot by qualifying for a leaderboard, not by name.
Version history
v2.6.1 · 2026-08-22 · GLM-5.3 now has its LiveBench results, and its score went down. LiveBench has added GLM-5.3, so it gains two results: 76.1% overall and 60.9% on agentic coding. Both are weaker than what it already had on this board — 95.4% on SWE-bench Verified, 91.7% on GPQA Diamond — so measuring it more made it score slightly lower, 74.25 to 73.77. It stays Provisional; it is still short of evidence in two of the five areas we score. And a correction we found while checking. Rather than copy the two numbers across, we re-read LiveBench’s full results file and checked every model we track against it. Twenty of our stored LiveBench figures turned out to be the number shown on their website, rounded to one decimal place, rather than the figure we are supposed to compute ourselves. That matters because LiveBench publishes only the individual task scores — the category and overall figures you see on their site are worked out in your browser and exist nowhere in their data, so we calculate them from the tasks. Twenty cells had skipped that step. All are now recomputed from the same file. How much it changed: almost nothing. Every correction is in the third decimal place of a percentage. No model changed rank and none changed tier. We fixed it anyway, because leaving GLM-5.3 computed properly while its neighbours sat on rounded website figures is the same inconsistency that caused the much larger problem we corrected earlier today. Three models keep their existing LiveBench figures: Muse Spark, GLM-5.1 and MiniMax M2.7 no longer appear on the board under any name, and a row disappearing is not a reason to change a measurement that was true when it was taken. No price changed.
v2.6.0 · 2026-08-22 · We were rewarding models for tests they had never taken. That is now fixed, and it reorders most of the board. What went wrong. One of the tests we score, τ³-Banking, is punishingly hard — every model that has taken it scores between 12% and 51%. Some models had taken it and some had not, and our Agency score averages over whichever tests a model has results for. So sitting the hard test dragged a model down, and never sitting it cost nothing. Being unmeasured was worth about three and a half points. The clearest symptom: ChatGPT 5.6 Sol beat ChatGPT 5.6 Terra on all nine tests they had both taken, and still ranked below it. Three other pairs had the same problem. Two things were wrong, not one. Ten models had no τ³ result recorded even though one was published — the numbers were in a file already on our own disk. And twelve of the results we did hold were reading the wrong line: they were picked by an old rule we replaced on 7 August and never went back to re-apply. Claude Opus 5 was being scored on its low reasoning setting, 30.3%, when its published maximum is 42.1%. What we would not do. Qwen 3.8 Max has a τ³ score of 51.3% published — one of the best on the board — and we did not take it. Artificial Analysis publishes it as a single unlabelled row, so there is no way to tell which configuration was tested, and we do not score a result we cannot identify. That decision costs Qwen 3.8 Max real points and leaves it the one model still holding an advantage from a missing result. Nemotron 3 Ultra is held for the same reason. The new order. ChatGPT 5.6 Sol rises from 4th to 2nd and now leads Terra 81.36 to 76.26, which is what beating it on every single test should look like. Qwen 3.8 Max rises to 3rd, ChatGPT 5.5 from 7th to 5th, Kimi K3 from 10th to 7th. Terra slips from 3rd to 4th, Gemini 3.7 Flash from 2nd to 6th and Grok 4.6 from 6th to 9th — all three had been carried by a hard result they had never been given. All four of those unfair comparisons are gone. No model changed tier. One more thing worth saying plainly: the reason this survived a routine harvest is that our collector for that source has been failing silently since 19 August — the site changed its data format and our reader kept reporting success. That is being fixed separately, and the results used here were re-read and cross-checked against a fresh capture before being published.
v2.5.1 · 2026-08-22 · The top of this board is now a three-way tie, and you should read it as one. Claude Opus 5 scores 80.34, Gemini 3.7 Flash 80.32 and ChatGPT 5.6 Terra 80.32. The gap from first to third is two hundredths of a point. That is not a difference between three models, and we would rather say so plainly than let an ordered list imply otherwise. What moved. Vals.ai has re-run ChatGPT 5.6 Terra on SWE-bench Verified, the single heaviest test on this board. It scored 75.2% when we read it on 5 August, at 34th place; the same model under the same test harness now scores 95.4%, at 5th. Its AGI Score rises from 77.29 to 80.32 and it moves from 6th to 3rd. Why we believe the new number. A twenty-point jump deserves suspicion, so: it is the same model identifier and the same harness, the result’s own margin of error narrowed from ±1.93 to ±0.94 — which a longer re-run explains and a typo does not — and 95.4% sits between Claude Opus 5 at 97.0% and DeepSeek V4 Pro at 96.4% rather than above everything. The old score was the odd one out. And Qwen 3.8 27B gains its first coding result. 86.0% on SWE-bench Verified, which lifts its score from 67.50 to 72.21. It stays Provisional rather than Ranked: it is still missing enough evidence in two of the five areas we score. How we found the second one. We were told about Qwen. Instead of copying that one row across, we re-read the entire 86-system board and checked every model we track against it — which is the only reason the Terra result surfaced at all. Everything else on that board matches what we already publish, exactly. One caution for anyone reading the table: several systems appear on it twice under different coding harnesses, and the scores differ by several points. We only ever take the standardised harness, so the comparison stays like-for-like. No price changed.
v2.5.0 · 2026-08-22 · CALIBRATION RELEASE Coding cost for the Grok models was about a third too low, and now it is not. The instrument changes, not the evidence: no score, rank or tier moved. Most providers charge roughly double once a single prompt crosses a size threshold, and they reprice the entire request when it happens. We were pricing every model as if that never occurred. How much it actually happens, measured rather than assumed. Across 2,663 model calls of real agentic coding, 38.4% of a Grok 4.5 bill and 44.6% of a Grok 4.6 bill landed above the threshold. The Grok 4.5 figure reconciles to the cent against a real August invoice. Better still, the threshold itself falls out of the billing unprompted: sort those calls by how much context they carried and the price per token sits flat below 200,000 tokens and exactly double above it, which is what xAI publishes. Three models have a surcharge and are still shown without one, deliberately. Grok 4.3, MiniMax M3 and Gemini 3.1 Pro all have a threshold, but we have no usage of them to measure and will not invent the number — so their coding cost here is understated by an unknown amount, and the note under the Value view says so by name. Applying a figure to some models and not others is uncomfortable; the alternative was to keep publishing a Grok cost we could prove was a third too low. Anthropic, the Gemini Flash models and ChatGPT 5.5 Pro have no surcharge at all, verified rather than assumed. And a display bug, fixed. Since v2.2.0 the note under the Value view has read “an estimate at a NaN% in / 0% out mix” — the text was still reading the old two-class mix after we moved to four. It now states the real shares, and rounds to a hundredth where a class is smaller than one percent, which is why output stopped reading as zero. Grok 4.5 coding cost rises from $0.351 to $0.485 per million tokens and Grok 4.6 from $0.546 to $0.789. On value for money that moves Grok 4.5 from 23rd to 25th and Grok 4.6 from 26th to 29th; the best-value pick is unchanged and everything between simply shifts up a place. Every other model is unchanged, and no capability score moved.
v2.4.1 · 2026-08-22 · We were publishing a context window for Inkling that nothing supports, and it is now gone. Our card said one million tokens, inherited from an old data file. Thinking Machines’ own model card states no native context window at all, and the figure on their training platform — 64K, or 256K on an extended variant — is a limit that platform imposes, not a property of the model. The same table lists Kimi K2.6 at 32K where Moonshot itself publishes 262,144, which is how you can tell what that column measures. So the field is blank until somebody publishes a real one. A blank card is honest; a confident wrong number is not. The rest of this release fills gaps rather than fixing them. Context windows go from 11 models to 29, every one read at the lab’s own documentation. Licences go from 1 to 13, architectures from 1 to 9, and parameter counts from 1 to 9. Two models — DeepSeek V4 Pro 0813 and DeepSeek V4 Flash 0731 — are now marked open-weight, because their own repositories publish downloadable weights under a named licence; we simply had not looked before. Round numbers are marked as round. Google publishes 1,048,576 and we store 1,048,576. Anthropic, Z.AI, Xiaomi and NVIDIA publish “1M” and we store 1,000,000, flagged as rounded at the source, so you can tell an exact published figure from a tidy one without us inventing precision in either direction. Qwen 3.8 27B is recorded at its native 262,144 rather than the million it can be extended to: a technique for stretching a window is not a window. Parameter counts come only from what a lab writes down. Several repositories display a parameter figure computed from the weight files themselves. That is a measurement of the artefact, not a claim by the lab, and putting the two in one field would make them look like the same kind of statement. GLM 5.1, GLM 5.2 and MiniMax M2.7 therefore show a licence and no parameter counts. Seven models still show no context window, each for a stated reason. Alibaba publishes none per model for the Max tier; Muse Spark is no longer in Meta’s documentation; GLM 5.1 states none; both Inkling models state none. And the April DeepSeek V4 Pro is left alone on purpose: DeepSeek publishes a dated 0813 repository and an undated one whose figures differ, and this is the exact model identity that was repointed under us once already. An undated repository is not proof of an April checkpoint. No price, score, rank or tier changed.
v2.4.0 · 2026-08-21 · CALIBRATION RELEASE Coding cost for the Claude models was too low, and we can show you the receipt. This changes the instrument, not the evidence. No model gained or lost a benchmark result and no score, rank or tier moved. What was wrong. Anthropic charges two different prices to write something into its cache, depending on how long the cache is kept alive: 1.25× the input price for five minutes, 2× for an hour. We were using the five-minute price because our rule is to publish standard list prices at the base tier. That rule exists to stop us quoting promotional discounts, and it is a good rule. It is not a rule about cache lifetime, and applying it here produced a figure nobody could actually pay: a real August invoice for Claude Opus 5 only balances at the one-hour price, and at the five-minute price it is out by a factor of 2.3. Why the cheaper tier is not the cheaper option. A five-minute cache is not a discount on the same thing — it is a different setting, and for coding it is the wrong one. Agentic coding sits on the same context for far longer than five minutes, so a five-minute cache expires mid-session and has to be written again. Choosing it does not save 40% on writes; it multiplies how many writes you pay for. In the sessions we measured, the one-hour tier was bought 100% of the time, across 170 million cache-write tokens. What moved. Coding cost rises about 12% for Claude Opus 5, Claude Sonnet 5 and Claude Opus 4.8 — the three models whose provider prices a cache write by how long it lasts. Nothing else on the board changed: every other provider publishes a single cache-write price with no duration tiers, so there is no choice to make and none was invented. What that does to the value ranking. Less than you might expect, and we would rather show it than let you wonder. The best value for coding is unchanged, and so is the whole top of that list. Two places swap: Claude Sonnet 5 slips from 17th to 19th on value-per-dollar, behind Gemini 3.1 Pro and ChatGPT 5.6 Terra, and Claude Opus 5 slips from 27th to 28th, behind ChatGPT 5.6 Sol. Capability scores are untouched — only the denominator moved. And only where we measured the hold time. The change applies to the Coding workload alone. Tool Use and Knowledge also read from cache, but we have not measured how long those workloads hold one, and assigning them an hour would be making a number up — the same reason this view already refuses to price hourly cache storage at all. They keep the base tier. This correction runs against our own interest in the sense that matters: it makes the models we measured most carefully look more expensive, not less.
v2.3.3 · 2026-08-21 · We described our own coding measurement as broader than it is, and this corrects it. The Value tab says the coding token mix is measured across three vendors. That is true, and it left a false impression: weighted by tokens, about 87% of the sample is a single coding agent. The mix itself is unchanged and still reproduces real invoices, but it should be read as one workload measured carefully, not as an average across the field. The Value view now says so. What the sample actually shows. Coding agents differ from each other more than the vendors do. The share of a bill that is new context runs from about 1.8% to 7.1% depending on which agent you use — some write to the cache explicitly and some rely on the provider doing it, and the two produce different bills on identical work. Why we still use one mix for every model. Because the Coding score ranks models under a single standardised harness. Giving each vendor its own token mix would price every model against a different workload, which is precisely what a standardised comparison exists to avoid. No price, score, rank or tier changed. This release only changes what we say about the evidence behind a number we already published.
v2.3.2 · 2026-08-21 · Gemini 3.7 Flash is now second, and effectively tied for first. ARC Prize has published its ARC-AGI-2 result — 84.6% at the High reasoning setting — and it is the first abstract-reasoning measurement we have ever had for this model. Its AGI Score moves 79.10 to 80.32 and it passes ChatGPT 5.6 Sol into second place. Sol did not get worse: it sat still at 80.17 and was overtaken. Please do not read the top of this board as a ranking today. Claude Opus 5 leads on 80.34 against 80.32. Two hundredths of a point is not a real difference between two models, and we would rather say so than let the order imply one. Treat the top two as tied. Where the number came from. ARC Prize dates the result 13 August, but the rows are genuinely new: our own captures of that same leaderboard on the 13th, the 14th and the 19th contain no Gemini 3.7 Flash at all. So we re-read the board at source rather than trust a date column, and kept the capture. 84.6% is a striking number for a small, fast model, so we checked where it actually sits before publishing it: below ChatGPT 5.6 Sol at its own High setting (85.4%), and well below the maximum settings of Sol, Opus 5 and Fable 5 (92.5%, 90.4%, 89.2%) — inside the frontier rather than above it, at 25 cents per task against their 74 cents to $2.06. Cheap, not impossible. We took the High row because High is the setting we already use for this model everywhere else on this board, not because it was the best of the three on offer. ARC-AGI-1 and ARC-AGI-3 sit on the same source row and were deliberately not taken. No price changed, and no other model moved in score, tier or rank.
v2.3.1 · 2026-08-21 · Model cards now say what each model can actually take in and give back. Thirty-one of thirty-six models gained input and output modalities, and thirty gained open-weight status — up from two and one. Whether a model reads images, or video, or only text is a basic fact about it, and until now almost none of our cards carried it. Where those facts came from, precisely. Six models were read at the provider’s own documentation. The rest come from a third-party model catalogue, and that deserves an explanation, because in the last release we refused that same catalogue’s context windows. The reason we refused them was that it rounds: it published Kimi K2.6 at 256,000 where Moonshot states 262,144, and ChatGPT 5.6 Sol at 1,000,000 where OpenAI states 1,050,000. That is a flaw in reporting numbers. Whether a model accepts images is not a number and there is nothing to round, and on every one of the six models we had checked independently, the catalogue agreed exactly. So: numbers from the source, categories from the catalogue, and each field on each model records which. Five models still show nothing, deliberately. ChatGPT 5.5 Pro, Claude Opus 4.8, both DeepSeek V4 checkpoints and Gemini 3.7 Flash are missing from the catalogue and we have not read them at source. In every case a sibling model would suggest the answer. A suggestion is not a source, so those cards stay blank. GLM 5.3 shows text-only but no open-weight status: the catalogue calls it proprietary while its own licence field says MIT and its published weights link is dead. That is a contradiction, not a fact. No price, score, rank or tier changed. Licence, architecture and parameter counts remain almost entirely blank — they are published mainly for open-weight models and we have not yet read them at source.
v2.3.0 · 2026-08-21 · Most of the newest models on this board could never show their own specifications. That is now fixed. Every model card has a Model facts section — context window, input and output types, licence, architecture, parameter counts. It was reading those from a data file frozen back at v1.11.27, and a frozen file cannot describe a model that did not exist when it was frozen. The result was that 24 of our 36 models showed no specifications at all, and never would have: Claude Opus 5, ChatGPT 5.6 Sol, Grok 4.6, Gemini 3.7 Flash and GLM 5.3 among them. Our own model registry now supplies those facts, with the frozen file kept only as a fallback for older models it still describes best. Eleven models now publish a verified context window. Each read from the provider’s own documentation: ChatGPT 5.6 Sol at 1,050,000 tokens, Kimi K3 at 1,048,576, Claude Opus 5 and Sonnet 5 at 1,000,000, Grok 4.6 and 4.5 at 500,000, Kimi K2.6 at 262,144. Where we could not verify a window at the source we leave the field blank rather than fill it in — a widely-used third-party tracker had windows for every model on the board, and rounded two of the nine we could check down to a neater number, so we did not take any of them. No price, score, rank or tier changed.
v2.2.1 · 2026-08-21 · CALIBRATION RELEASE Some providers charge rent on a cache, by the hour. We checked who, and built the model to refuse rather than guess. Completing the cost model from v2.2.0: alongside new context, cache writes, cache reads and output there is a fifth charge — storage — billed per token per hour for as long as a cache is held. It is the one charge we cannot estimate, and we are saying so rather than papering over it. The other four are per-token, so a blended per-token price expresses them exactly. Storage depends on how long a cache is kept alive, and nothing we measured tells us that: our billing data gives token counts, not cache lifetimes. We can say each cached token is read back about forty-five times; we cannot say over how many hours. So where a provider charges hourly storage for the caching mode we price, the model now declines to price that model at all, exactly as it declines when a cached-input rate is unpublished. Inventing an hours figure would have invented the largest term. Today it declines nobody, and that is the finding. Every provider we checked charges nothing to hold a cache in the mode we price. Google’s storage fee — up to $4.50 per million tokens per hour, the steepest published — applies only to caches you create and manage yourself; its default caching is automatic and free to hold. Alibaba publishes no storage charge in either mode. Z.AI’s is listed as limited-time free, which is a promise with an expiry date: when it ends, those models will leave the cached workloads until someone measures how long a cache actually lives. We also now record which caching mode each price assumes. Where a provider offers more than one at list price we price the one a capable user would actually use — that is the automatic mode for Google, and the manually-managed one for Alibaba, where it costs half as much. Commercial terms we take as standard; technical setup we assume competent. No price, score, rank or tier changed in this release.
v2.2.0 · 2026-08-21 · CALIBRATION RELEASE We rebuilt how coding cost is estimated, and measured the workload instead of assuming it. This release changes the instrument, not the evidence. No model gained or lost a benchmark result, and no score, rank or tier moved. What changed is the arithmetic behind the Value tab. A modern agentic bill has four token classes, and we were only counting three. New context, cache writes, cache reads and output are priced differently by every provider that separates them, and the old model had no way to express a cache write at all — so it understated every cache-heavy workload. It now counts all four. The coding mix is now measured, and it was not close. We had assumed a coding workload was 90% input, 10% output, with 80% of the input cached. Measured across billing records from three vendors and three different coding agents — 5.8 billion tokens — the real shape is about 2% new context, 97.6% context read back out of cache, and 0.24% output. Output is not a tenth of a coding bill; it is a quarter of one percent. The four-class model reproduces those real invoices to within 0.05%. The practical consequence, if you write code with these models. Almost the entire cost of agentic coding is the cached-input rate. Headline input and output prices barely matter; the cached rate very nearly is the bill. Published cached rates across the board range from about 0.8% to 25% of a model’s own input price, so two models with identical headline pricing can differ several-fold in what coding actually costs. And we stopped guessing where a lab publishes nothing. The old model filled a missing cached rate with “about 10% of input”. Under a workload that is 97.6% cached, that guess would have set almost the whole estimate. Models without a published cached rate are now left out of cached workloads rather than estimated into them. The other four workloads — Overall, Reasoning, Knowledge, Tool Use — carry the same assumptions they always have, restated in the new shape and unchanged in effect. Only the coding mix is measured, and the view now says which is which.
v2.1.1 · 2026-08-21 · We re-checked every price on the board against the provider’s own page, and some of ours were wrong. Prices drive the Value tab, and ours had drifted: a third of the board was priced from third-party trackers rather than the lab, and several entries had not been re-read since early June. Nineteen of thirty-one came back correct. Eleven did not. The largest error was on the cheapest model we list. DeepSeek V4 Flash 0731 was carrying an old price schedule — not a misread tier, a superseded one — which understated it by roughly three to five times. It is still the cheapest model on the board, but it is no longer cheap by as much as we were showing. DeepSeek V4 Pro 0813 turned out to have a published price all along, and now appears. Grok 4.5’s cached rate came down, and MiniMax M2.7 joins the Value view. Seven models gained a cached-input price they were previously missing, including Gemini 3.7 Flash, GLM 5.3 and Qwen 3.8 Max. That matters more than it sounds: for coding work most of what you pay for is context read back out of cache, so the cached rate is close to the whole bill. Where a lab publishes no cached rate we now leave the model out of that view rather than estimate one — only ChatGPT 5.5 Pro is affected. The rule we now price by, stated plainly. Standard public list price: no promotional rate, no volume or subscription tier, no discount for letting a lab train on your data, and the base context band. A discount with an end date is excluded because it expires; a discount with no end date is simply the price. Where a provider charges different rates by time of day, we publish the 24-hour average of its published rates and say so. Nothing here moves a score. No model changed rank, tier or AGI Score — price is not an input to capability. This is the Value tab getting more honest, not the leaderboard changing.
v2.1.0 · 2026-08-20 · We were not reading one of our sources properly, and it cost two models their place. When our primary evaluator for a benchmark has not covered a model, we treat a second independent evaluator’s published row for that model as usable rather than discarding it. That rule is new; the situation was not. Three models had results sitting in plain sight on a board we were skipping because it was not our first choice for that column — which is only a reason to skip it when the first choice actually has the model. What changed on the board. Muse Spark 1.2 joins the ranked table at #9. It had been held back as provisional purely for thin reasoning coverage, and the missing reasoning result existed all along. Qwen 3.8 27B rises from 53.77 to 67.50 and stays provisional. GLM 5.3 is on the board, and in v2.0.45 we said it would not be. Six days ago we wrote that it missed our reasoning-coverage requirement. It missed it by 0.0125, and the result that closes that gap was already in our own records. It enters provisional rather than ranked: it is text-only, so it carries no visual-reasoning result at all. Two things we found and deliberately did not add. A Grok 4.5 visual-reasoning row that our own notes had already flagged as implausible — it scores below far older Grok models — stays out. And Inkling Small has two evaluators disagreeing by nearly six points on the same benchmark, which usually means they measured different settings; until we know which, we are taking neither. The new rule is about coverage gaps, not about accepting whatever number is nearest to hand. One correction to our own precision. Qwen 3.8 27B’s Humanity’s Last Exam score had been read off a chart at 0.339; the published figure is 0.33920296570899, and that is what we now carry.
v2.0.45 · 2026-08-20 · Two models join the board, and four stale “fallback” labels come off it. Gemini 3.7 Flash enters ranked at #3, on seven measured cells drawn from three independent evaluators. Qwen 3.8 27B enters provisional. Both were admitted under the same published entry rules as every other model, on results that already stood on the evaluators’ own boards. A third model was harvested in the same pass and is not here. GLM 5.3 posts strong agentic results, but it is absent from the reasoning columns this board sources, so it does not meet the reasoning-coverage requirement for entry and stays off the board until it does. We would rather show the gap than fill it. [Superseded in v2.1.0 — that reasoning column did exist, on a source we were not reading for models our primary evaluator had not yet covered. GLM 5.3 is on the board.] Refreshed evidence, and some corrections we owed. Several Humanity’s Last Exam figures were still carried from an older projection and now read the evaluator’s current published numbers. MiniMax M3 gained its first result on that benchmark, which lowered its score while qualifying it for a rank — more measured, not better. Four models were still labelled as carrying a converted “fallback” result whose evidence no longer supports the claim, and those labels are gone. Family and naming metadata for Kimi K3 and Inkling, and release details for Inkling Small and DeepSeek V4 Flash 0731, now match the canonical record.
v2.0.44 · 2026-08-14 · The ranked board now shows only ranked models, and every unranked model says why. Hiding unranked models now actually hides them. The default view is meant to be the official board, and it was also carrying models that hold no rank, while the button offered to hide a count it did not match. Both came from the same cause: different parts of the page were each deciding what "ranked" meant and had drifted apart. There is now one definition, and the list, the button, the count and the number beside each model all read from it. A model can be unranked for more than one reason, and we now name the real one. Some models are unranked because their evidence is still thin. Others have plenty of evidence and are unranked because a result we were relying on turned out not to be usable, so we withheld the rank until it is re-checked. Those are different situations, and several were being described as thin coverage, which was wrong. Each unranked model now carries its own explanation, and where more than one thing is holding it back, all of them are stated. No score, rank, benchmark result or methodology changed. The same models are ranked, in the same order, with the same numbers. This release changes what the site tells you about that state, not the state.
v2.0.43 · 2026-08-14 · Grok 4.6 is now officially ranked, at number 5. ARC Prize has published an ARC-AGI-2 result for Grok 4.6, and we read it directly from their leaderboard: 67.1% at the XHigh setting, evaluated 2026-08-11. That is the model's seventh qualifying result. Its score went down, and that is why it can now be ranked. The new result is weaker than Grok 4.6's others, so its AGI Score moves from 78.43 to 76.18. But ranking is not a reward for scoring well, it is a statement that we have measured a model broadly enough to place it against the field. Reasoning was the gap, and this result closes it. A model can become more accurately described and lower at the same time. Ten models shift down one place. None of their scores changed and none of their evidence changed; they simply have one more model above them. The top three and the leading model are unchanged. Nothing was estimated. We took only the ARC-AGI-2 result, at the precision ARC publishes it, and left the other numbers on that row alone.
v2.0.42 · 2026-08-13 · Fixed model deep links so direct visits, refreshes and normal navigation now consistently serve the current AGI Ranker release. Opening a model page from a link, a bookmark or a search result now shows exactly what you see when you open that model from the leaderboard: the same release, the same data, the same score and the same official standing. No score, rank, benchmark result or model changed.
v2.0.41 · 2026-08-13 · Two models join, both Provisional and both unranked. DeepSeek V4 Pro 0813 enters at 74.36 on six directly observed results, and NVIDIA Nemotron 3 Ultra 550B A55B enters at 64.67 on four. Every value is a measurement we read from an evaluator ourselves; nothing was estimated, converted or filled in. Neither receives an official rank. Our coverage rules ask for a breadth of evidence before a model is ranked against the field, and neither model has it yet: Nemotron 3 Ultra launches with four qualifying results, and DeepSeek V4 Pro 0813 clears the count but not the spread across the abilities we measure. Provisional is what we publish when a model is real and measured but not yet comparable, and it is not a placeholder for a score we expect to improve. DeepSeek V4 Pro 0813 is a separate model from DeepSeek V4 Pro, not a replacement for it. Both are published side by side by the evaluators we read, under different identifiers, so we carry both. The existing DeepSeek V4 Pro is untouched. NVIDIA joins as a provider on this one model. No ranked model moved. No existing score, tier or ordinal position changed anywhere on the board.
v2.0.40 · 2026-08-12 · LiveBench remeasured, and we refreshed the whole verifiable cohort at once. LiveBench re-runs models inside a release without changing its version, so the numbers we stored in June are no longer the numbers the board reports today. Rather than take the new rows we wanted and leave the rest stale, we recaptured the entire board in one read and refreshed every model we could verify against it, from the same capture, in one release. Six existing figures moved. Claude Opus 4.8 78.93 to 76.22, Muse Spark 1.1 76.23 to 75.30, Grok 4.5 76.25 to 75.77 and Gemini 3.1 Pro 77.13 to 76.95 all fall; Kimi K3 78.54 to 79.19 and Inkling 71.68 to 71.92 rise. Four down and two up, because a remeasurement is not a correction and we do not get to choose the direction. Grok 4.6 gains its LiveBench evidence on launch day, 77.56 overall and 54.24 on agentic coding, which moves it from 80.57 to 78.43. It stays Provisional with no official rank: the current evidence remains insufficient under our existing coverage rules. Muse Spark 1.2 gains both LiveBench cells and moves 60.55 to 65.76, also Provisional. The official ranking order is unchanged. Every movement here is under half a point and no model changed rank or status. Two things we deliberately did not do. A LiveBench row for Claude Opus 5 is now on the board, but our existing figure for it is tied to a pre-release identity we have never resolved, and two numbers being close is not proof they are the same measurement; it stays held until that is settled. Three older figures we hold, for Muse Spark, GLM 5.1 and MiniMax M2.7, have no row on the current board at all; withdrawing evidence is a separate decision from refreshing it, so they are untouched and flagged. Every value above was recomputed by us from LiveBench’s own task-level data rather than read off their headline column, and no weight, ceiling or threshold changed.
v2.0.39 · 2026-08-12 · Grok 4.6 joins the board on launch day, listed Provisional. xAI released Grok 4.6 today and we have four measured cells for it: GPQA Diamond 94.70 and SWE-bench Verified 95.60 with a 0.92 point error bar under the Mini-SWE-agent harness, both from Vals; Terminal-Bench 2.1 78.28 from Vals under Terminus 2; and Humanity’s Last Exam 42.91 from Artificial Analysis. Its analytical score is 80.57. That is the highest number on the page, and it is not ranked. Official ranking requires five scoring benchmarks and Grok 4.6 has four, so it sits in the Provisional cohort until more of it has been measured. We are not going to move the bar to fit a new arrival, and a high score on a thin profile is not the same claim as a rank. Two figures we could have used and did not. Artificial Analysis publishes Terminal-Bench 2.1 at 88.39 for this model, ten points above the Vals figure we used. Our Terminal-Bench column is built entirely from Vals under a fixed harness, and the two boards carry a measured systematic offset, so mixing them would have inflated the agentic score rather than improved it. Artificial Analysis also publishes a τ³-Banking result, but their board has moved to a successor version while our scoring column is still on the frozen earlier one; that reading is recorded and quarantined rather than scored. No ARC-AGI-2, LiveBench or MMMU-Pro row exists for Grok 4.6 yet. We checked each board directly rather than assuming, and the rows are genuinely not there on launch day. As always, a lab’s own published figure is where we start looking, never what we publish: every number above was read from an independent evaluator’s current board. No other model moved and no weight, ceiling or threshold changed.
v2.0.38 · 2026-08-12 · ARC-AGI-2 sweep: two overdue corrections, one new cell, one source upgrade. ChatGPT 5.5 and ChatGPT 5.5 Pro were reading the wrong row. Both stored their High-effort result when the ARC Prize board also publishes xHigh, which is the setting our method selects. 5.5 moves 83.3 to 85.0 and 5.5 Pro moves 84.6 to 84.2, up for one and down for the other, because we pick by configuration and not by which number flatters. These were identified on 2026-08-08 and then left behind when the integrity release deliberately carried no new coverage; that was our error and it is now closed. Gemini 3.6 Flash gains ARC-AGI-2 at 60.4, its first cell on that benchmark, which lowers its score from 75.46 to 72.30 and moves it to eighth. Gemini 3.5 Flash’s ARC figure now comes from ARC Prize rather than Google. We had a Google-published 72.1 with no row label or effort setting, which our source rules exclude from scoring. ARC Prize now publishes the same 72.1 as a labelled High-effort row, so the number is unchanged but it finally counts, and the score adjusts from 71.37 to 70.38 as a result. No other model moved.
v2.0.37 · 2026-08-11 · Custom weightings now show how much evidence each score rests on. A reader set Agency to 100 and everything else to 0, and found ChatGPT 5.6 Luna scoring above ChatGPT 5.6 Sol, which is the wrong way round. The cause is coverage, not a calculation error. Sol has all four agentic benchmarks including tau3-Banking, where it scored 33.0, its weakest result anywhere. Luna and Terra have no tau3-Banking row at all, so they are scored on the three legs they do have and their weakest agentic evidence is simply absent. Luna beats Sol on Agency largely because Luna was not measured on the benchmark Sol struggled with. Every custom-weighted row now carries a coverage chip showing the share of the weighted evidence that model actually has, amber when it is short, with the weakest component named. Sol reads 100%; Luna and Terra read 70%. No score, rank or methodology changed, and the default view is untouched. This makes an existing limitation visible rather than hiding it: missing benchmarks are left out of the average rather than counted as zero, so a model can score higher by having been measured on fewer of them. Whether the formula itself should handle that is a separate question we are taking up deliberately rather than patching in place.
v2.0.36 · 2026-08-11 · Meta Muse Glimmer joins the board, listed but not scored. It is a 30B open-weights model that runs on consumer hardware, and Artificial Analysis rates it strongly for its size. We have three directly observed cells for it: GPQA-Diamond 83.5, Humanity’s Last Exam 22.0 and MMMU-Pro 74.3. What we do not have is any agentic evidence: vals.ai carries no page for it and LiveBench has not run it. So it gets no AGI Score. Our formula spreads a score across five capability components, and computing one from three of them would produce a number that says more about what is missing than about the model. The cells we do have are shown on its page. A score follows when an agentic benchmark does. Separately, a calibration we had retired was still running. The MCP-0001 fallback conversion was suspended on 2026-08-08 after it failed its own sufficiency test, but the suspension was recorded only in our notes: the benchmark registry still marked it eligible, so new evidence was quietly being adjusted by an offset we had disowned. The policies are now marked suspended in the data, intake refuses to use them, and we checked every affected cell. No published score was wrong: zero converted values were still counting. One DeepSeek cell was carrying the old label on what is in fact a clean, directly read value, and that label is corrected. No other model moved.
v2.0.35 · 2026-08-11 · DeepSeek V4 Flash (the 0420 checkpoint) leaves the board. DeepSeek withdrew it: the company’s pricing page now lists model id deepseek-v4-flash as version DeepSeek-V4-Flash-0731, so the original checkpoint is no longer served and has no current first-party price. We rank models you can actually use, so a model its own maker has retired comes off, on the same principle that kept Fable 5 off the board for being unscorable. DeepSeek V4 Flash 0731 is a different model and stays, ranked, priced and unaffected. The removal changes no other model’s score and no ranking position: the retired entry held no official rank. Its evidence stays in the ledger, as our records are append-only and we do not erase what we once published. No other change in this release.
v2.0.34 · 2026-08-10 · Qwen 3.8 Max rises to #3, and two gaps a reader spotted are closed. vals.ai published GPQA-Diamond 93.69, MMMU-Pro 88.03 and SWE-bench Verified 85.60 for Qwen 3.8 Max, all with standard errors. The GPQA and MMMU cells supersede the Artificial Analysis figures we had been using: an independent evaluator publishing a standard error is stronger evidence than the aggregator row it replaces. Score moves 73.34 to 78.30, taking third place and moving ChatGPT 5.6 Terra to fourth. DeepSeek V4 Flash 0731 gains ARC-AGI-2 at 61.4 (Max effort, the frontier row on the ARC Prize board), captured on 2026-08-08 but held out of the integrity release, which carried no new coverage. It lowers that model’s score from 69.27 to 66.00 while completing its coverage, so it moves from Provisional to Ranked. Qwen 3.8 Max now shows a price, $2.00 in / $6.00 out per million tokens, from Alibaba Cloud Model Studio’s own page, and enters the Value view. DeepSeek V4 Flash (the 0420 checkpoint) still shows no price, and that is correct. DeepSeek’s pricing page now lists model id deepseek-v4-flash as version DeepSeek-V4-Flash-0731: the vendor repointed the endpoint and no longer sells 0420. Grok 4.5 MMMU-Pro remains held at a listed 61.79, far below Grok 4.3’s 83.06 on the same benchmark. We hold rather than publish a figure we believe the source has wrong. No other model moved.
v2.0.33 · 2026-08-08 · Evidence integrity release. Thirteen models keep their score but lose their rank while we re-check a benchmark result. Some results we had been using could not be traced back to a specific evaluator row: the wrong effort setting, or in two cases the wrong model entirely. Where a current, verified row existed we replaced the old value. Where it did not, we withdrew the result rather than keep a number we could not stand behind. A model affected this way now appears under evidence check: its AGI Score is still shown, computed from the evidence that remains, but it receives no ordinal rank, no Top 3 place and no headline position until the evidence is re-verified. Gemini 3.1 Pro (preview) is the clearest case, at 77.87 with no rank. Two results were traced to the wrong model. Artificial Analysis reused the slug deepseek-v4-flash for the newer 0731 checkpoint, and two of our DeepSeek V4 Flash cells were carrying 0731 values. Those are voided, not re-attested. Nothing else moved. Every model not under evidence check has exactly the score it had before, to the digit. The tau3 Banking column is frozen pending a separate calibration decision, and no new benchmark coverage entered in this release.
v2.0.32 · 2026-08-07 · The “Most recent eval” date is removed, because it was stating something false. It showed the newest evaluation date any source had published, but only 65 of our 217 published cells carry a source-published date at all: vals.ai and Artificial Analysis publish none. The figure therefore ignored two thirds of the board, including every cell harvested this week, and it read 2026-07-30 on a day we had scored a new model from evaluator pages read that same morning. A freshness signal that is stale by construction is worse than no signal. Per-cell evaluation dates stay where they mean something: on each model page, next to the value they describe, still driving the stale-cell flag for results older than twelve months. “Last updated” stays too, derived from the deployment timestamp of the underlying data file, so it is true by construction. No score, rank, evidence or methodology change.
v2.0.31 · 2026-08-07 · IndexNow enabled. Every production deploy now notifies Bing-family search engines and Copilot grounding to re-crawl within minutes, so a board that updates daily is indexed daily. No score, rank or methodology change.
v2.0.30 · 2026-08-07 · Claude Sonnet 5 completes its evidence at 80.75, the highest estimate on the board, and does not take #1. Nothing was capped to make that true; three governance rules ship with this release. (1) Two clocks: evidence releases run on a locked benchmark calibration (pinned 2026-08-05); only labelled calibration releases may reweight the instrument, so adding one model can no longer move every other. (2) Ranked-only podium: official ranks, the Top 3 and the headline leader come exclusively from Ranked models; a Provisional estimate is shown in full but receives no ordinal until its evidence meets the Ranked standard. (3) MCP-0005: Ranked requires both core components (Agency and Fluid Reasoning) to meet the existing 40% coverage floor on the declared base weights; the one-thin-component allowance now applies only to non-core components. The evidence, all from direct evaluator reads: Muse Spark 1.2 (Meta, 2026-08-05) enters at 58.87, Provisional, five cells. Claude Sonnet 5 rises 70.85 to 80.75 on six cells and is Provisional under MCP-0005 (Reasoning coverage 37.5% against the 40% floor: no HLE, ARC or AIME cell); its estimate is disclosed beside the podium and it graduates automatically when a Reasoning cell arrives. Gemini 3.6 Flash 75.46, Ranked (its HLE arrived at 38, a low cell an honest aggregate must count). Qwen 3.8 Max 73.99, Ranked (vals Terminal-Bench supersedes the held AA value). Inkling Small 56.55 (vals converts a week-old Terminal hold). Official podium: Claude Opus 5 80.37, ChatGPT 5.6 Sol 79.95, ChatGPT 5.6 Terra 77.18. Every other model is unchanged to the last digit.
v2.0.28 · 2026-08-06 · DeepSeek V4 Flash 0731 appears on the Value view, fixing a gap a reader spotted: it sat in the Coding specialty but not under Value, Coding. The two views have different entry requirements. A specialty tab needs benchmark cells; the Value view additionally needs a price, and the 0731 checkpoint’s site entry still carried none. The cause was a data-plumbing defect, not a policy: the site exporter set pricing only when it first created a model’s entry, so DeepSeek’s first-party price (verified 2026-08-06, /usr/bin/bash.14 in / /usr/bin/bash.28 out per million tokens, cache hits at /usr/bin/bash.0028) was recorded in our registry a day after the entry existed and never reached the published data. The exporter now refreshes pricing from the registry whenever the registry has one, so a price learned late can no longer strand a model off the Value view. At its price, 0731 at 68.02 lands among the strongest value positions on the board. No score, rank or interval changed in this release.
v2.0.27 · 2026-08-06 · The fallback policy keeps its hardest promise: vals published DeepSeek V4 Flash 0731, and the canonical value replaced our fallback cell even though it is lower. When we admitted fallback cells under the signed 2026-08-05 policy, rule 5 said the canonical evaluator’s value takes over the day it publishes, flattering or not. That day was today. vals now lists the checkpoint: GPQA Diamond 89.90±1.69 replaces our converted 90.5 (the fallback prediction missed by 0.60pp against a promised containment of ±5.61, which is the calibration working), SWE-bench Verified 88.80±1.41 is new, and Terminal-Bench 67.04±1.35 supersedes the held Artificial Analysis 78.7, the 11.7-point gap restating the measured evaluator offset we publish for that column. DeepSeek V4 Flash 0731 rises from 58.91 to 68.02 on seven cells, still Provisional. Its identity is also now first-party confirmed: the 2026-07-31 checkpoint officially superseding DeepSeek V4 Flash, with open weights and API pricing on the model page. Two models join under the best-4 rule. Gemini 3.6 Flash enters at 75.06, Provisional on four cells (vals GPQA 93.43±1.33, which would lead the column; vals Terminal-Bench 73.78±1.63; LiveBench global 73.6 and agentic task mean 43.4). Claude Sonnet 5 enters at 70.85, Provisional on three cells (vals Terminal-Bench 74.53±2.08; LiveBench global 76.0 and agentic 59.4). Seven further harvested values are held with their reasons: five vals values that arrived without the board’s own standard error or through a mirror (including an identical SWE-bench figure reported for both new models by the same mirror, which is precisely the artifact the rule exists to catch), and two secondary-sourced HLE values. No existing model’s score moved. Lab self-reports for both models are recorded as discovery only, never scored.
v2.0.26 · 2026-08-06 · The lead changes hands for the first time in the v2 era: Claude Opus 5 at 80.37 over ChatGPT 5.6 Sol at 79.95. No model changed; an evaluator refreshed its tasks, and we promised you this yesterday. The v2.0.25 entry disclosed that LiveBench’s rolling board had moved since our 2026-06-25 capture and that a full single-date refresh would arrive as its own logged release. This is that release: all sixteen remaining Agentic Coding cells rebuilt from the task table captured 2026-08-05, plus the two global-average rows the same capture showed moving. Sol’s agentic tasks now average 56.2 where the June table gave 65.6, and its global average eases from 82.4 to 81.0, taking Sol from 80.61 to 79.95. Claude Opus 5’s refreshed row runs higher (65.2 vs 61.3; LiveBench also moved its Claude rows from xHigh to Max Effort configuration since June, which we record on the cells), lifting it from 80.14 to 80.37. The point estimates crossed; the evidence still does not separate the top two, exactly as every release since v2.0.0 has said, so the honest headline is that the coin flipped within its interval, not that one model overtook another in capability. Eight further models moved, every one by 0.34 or less: ChatGPT 5.5 +0.14, Kimi K3 +0.24, Claude Opus 4.8 −0.30, Muse Spark 1.1 −0.34, Grok 4.5 −0.20, Gemini 3.1 Pro −0.07, Inkling +0.09, and hundredth-scale wobbles from cohort statistics. Every superseded cell keeps its June value and capture date in the ledger. Also: DeepSeek V4 Flash 0731 is now scored, entering at 58.91, Provisional, on five cells: HLE 36.8 and tau3-banking 31.1 from Artificial Analysis (the canonical evaluator for both columns), a GPQA Diamond fallback cell (AA 90.8 converted to 90.5, labelled, ±5.61pp, auto-superseded the day vals lists the checkpoint), and its two LiveBench cells from yesterday. Its Terminal-Bench 78.7 from AA is held under the standing vals-column rule. It ranks directly above the sibling checkpoint it will presumably one day retire.
v2.0.25 · 2026-08-05 · The Agentic Coding task table arrives, four held cells convert, and Qwen 3.8 Max’s first Agency evidence costs it seven points, which is the system working. LiveBench’s Agentic Coding category is rebuilt on this board from the per-task scores (JavaScript, TypeScript, Python), never from the browser-derived category number, and today’s capture of the task table converts every standing agentic hold: Qwen 3.8 Max gains its first Agency cell at a task mean of 64.7 and moves from 79.67 to 72.57, rank 3 to 6, still Provisional and now flagged thin in Agency. Nothing about the model changed; a component that was entirely unmeasured is now measured once, and the score follows the evidence. The interval published this morning, [69.6, 87.5], already contained today’s 72.57. ChatGPT 5.6 Terra adds its agentic cell (54.97) and settles from 78.15 to 77.18 at rank 3; ChatGPT 5.6 Luna adds 48.43 and settles from 76.40 to 74.38, slipping behind ChatGPT 5.5; DeepSeek V4 Flash converts its held 37.6 and rises to 56.42. No other model’s score moved. Also: DeepSeek V4 Flash 0731 joins the roster, the first organic entry under the best-4 rule adopted this morning: LiveBench lists it as a distinct checkpoint (global average 74.2, agentic task mean 46.8), and with two admissible cells it enters as insufficient data, unscored until a third arrives. One disclosure ahead of the next release: LiveBench is a rolling benchmark, and today’s task table shows some older rows have moved since our 2026-06-25 capture (GPT-5.6 Sol’s tasks now average 56.2 against our stored 65.6). Today’s intake deliberately touched no existing cell; a full agentic-column refresh at a single capture date is queued and will be its own logged release, so a change to the leader’s score arrives announced rather than smuggled.
v2.0.24 · 2026-08-05 · Eight held cells convert on direct source reads, and both new OpenAI tiers reach Ranked the same day they arrived. The morning release held ten harvested values because their provenance ran through mirrors, prose or secondary panels rather than the evaluators themselves. Direct reads arrived hours later: the vals.ai model pages for both models, each cell now carrying vals’ own standard error, and Artificial Analysis’s HLE chart read at the source. Every one of the eight direct values matched the relayed figure exactly, which is the hold system doing precisely what it is for: the numbers were probably right all morning, and now they are verified instead of probable. ChatGPT 5.6 Terra rises from 73.28 to 78.15 and flips Provisional to Ranked at #4 (six cells: GPQA Diamond 90.91±1.91, MMMU-Pro 86.47±0.82, SWE-bench Verified 75.20±1.93 join ARC-AGI-2 83.9, LiveBench 77.9, Terminal-Bench 73.41±2.08); its interval narrows from [62.5, 83.9] to [72.9, 82.4]. ChatGPT 5.6 Luna rises from 69.37 to 76.40 and flips to Ranked at #5 (GPQA 91.67±1.74, MMMU-Pro 85.03±0.86, Terminal-Bench 79.03±0.99 join SWE-bench 93.0±1.14, ARC 59.5, LiveBench 73.6); interval [58.9, 79.9] to [71.1, 80.7]. The default Ranked view now reads Sol, Opus 5, Terra, Luna, then ChatGPT 5.5: displacement only, since no existing model’s score moved by a hundredth. The two LiveBench Agentic Coding category values stay held under the task-table rule until the subtask rows are captured. The held records themselves are superseded, not deleted, with links both ways.
v2.0.23 · 2026-08-05 · Two models join, and the roster rule that admits them goes public: best 4 per lab. Ratified in this week’s methodology review: AGI Ranker tracks up to the best 4 models from each lab, including sideways SKUs, with new releases entering on arrival, the lowest-scoring rotating off above 4, an 8-week protection for entries that cannot yet be scored, and reasoning-effort settings never creating separate rows. The rule is prospective; the existing roster is grandfathered. First application: OpenAI’s GPT-5.6 mid and fast tiers. ChatGPT 5.6 Terra enters at 73.28, Provisional on three cells: ARC-AGI-2 83.9 from the ARC Prize board at max effort (the strongest ARC-AGI-2 result we have ever scored), LiveBench global 77.9, and Terminal-Bench 2.1 at 73.41±2.08 from vals. Interval [62.5, 83.9], rank range 1-22. ChatGPT 5.6 Luna enters at 69.37, Provisional on three cells: SWE-bench Verified 93.0±1.14 from vals (below Opus 5 at 97.0 and Sol at 96.2), ARC-AGI-2 59.5, LiveBench global 73.6. Interval [58.9, 79.9], rank range 1-25. Ten further harvested values are recorded but not scored, each held with its reason in the ledger: six vals values that reached us through mirrors or prose rather than a direct board read (mirrors stopped being acceptable provenance after the BenchLM incident), two Artificial Analysis HLE values read from secondary aggregator panels rather than from AA, and two LiveBench Agentic Coding category scores under the standing task-table rule. Every hold converts the moment a direct source read arrives; the same lag resolved for Inkling Small within a week. No existing model’s score moved: both entries are Provisional and sit outside the rankable cohort that drives shrinkage. OpenAI now stands at five tracked models; under the grandfathering clause nothing rotates off today.
v2.0.22 · 2026-08-05 · Qwen 3.8 Max gains its first LiveBench cell; the score barely moves and the interval narrows, which is exactly what a fourth measurement should do. LiveBench now lists Qwen 3.8 Max: global average 78.5, measured by the benchmark's own evaluation harness, no fallback conversion involved. The score moves from 79.78 to 79.67, the interval narrows from [67.9, 88.2] to [69.6, 87.5], the rank range tightens from 1-15 to 1-14, and the model gains its Language component at 78.5. It remains Provisional at rank 3, now on four scored cells. One number from the same page is held rather than used: LiveBench also shows an Agentic Coding category score of 64.6, but that figure is computed in their browser and published nowhere in their data; our Agentic Coding column is rebuilt from the per-task table, under the same rule that held DeepSeek V4 Flash's category row on 2026-08-01. It becomes a scored cell once the subtask rows are captured. No other model's score moved.
v2.0.21 · 2026-08-05 · Fallback admission: when our canonical evaluator has not caught up to a new model, we now convert the shadow evaluator’s value instead of publishing an empty cell — and we charge for it. The rule, decided in this week’s methodology review and implemented under a signed change proposal: a column qualifies only when the two evaluators’ measured disagreement scatter is at most 2 points (GPQA Diamond at 1.66 and MMMU-Pro at 1.67 qualify; Terminal-Bench 2.1 at 4.54 does not and stays single-evaluator). A fallback cell is the shadow value converted by the column’s measured offset, permanently labelled on the model page, given a widened interval of ±5.61pp (GPQA) or ±5.55pp (MMMU-Pro) — calibrated so the largest cross-evaluator miss we have ever observed sits inside it — and automatically replaced by the canonical value the day it publishes, even when that value is lower. Five cells were admitted, and the board moved as follows. Qwen 3.8 Max is now scored: 79.78, Provisional, on three cells with an honestly enormous interval of [67.9, 88.2] and a rank range of 1–15 — the point estimate says third, the evidence says somewhere in the top fifteen, and we print both. Inkling Small rises from 51.15 to 55.84 and reaches Ranked on six cells; DeepSeek V4 Flash rises from 41.52 to 55.57. Counterintuitively, both models’ intervals got narrower, not wider: a fallback cell removes nine to eleven points of missing-coverage uncertainty and adds back only one or two of conversion uncertainty. One displacement with no new evidence behind it, named rather than buried: Muse Spark 1.1 slips below Kimi K3 in the default Ranked view. Its own cells did not change; new MMMU-Pro observations shifted the cohort statistics that drive the saturation check and shrinkage, moving it and three other models by up to 0.28 points. The top two are unchanged. ChatGPT 5.5 (Pro) also gained a fallback cell but remains unscored: fallback admission does not bypass its verification gate. The full impact table, calibration residuals and policy text were signed before any of this touched production.
v2.0.20 · 2026-08-04 · Kimi K3 falls from 3rd to 6th on a new measured result, which is the system working. The ARC Prize board now lists Kimi K3 on ARC-AGI-2: 60.4% at max effort (CoT, $1.59/task, evaluated 2026-07-16; the High and Low effort rows read 55.0 and 12.4, and we take the max row to match the effort convention of every other Kimi K3 cell). ARC-AGI-2 enters the Reasoning component at 60.4 against a component that previously averaged 83.2 without it, so Kimi K3 moves from 76.25 to 72.42 and its Reasoning from 83.2 to 68.5. Nothing about the model changed; what changed is that a component previously carried by knowledge-style benchmarks now includes an abstract-reasoning measurement, and the score follows the evidence. This is the case the no-imputation policy exists for. A gap is never a free pass: while the cell was missing, the interval said so, and now that it is measured, the number moves. Under an imputation regime this cell would have been guessed near the component average and the reveal would have been a shock; here the interval [71.4, 80.8] published with the previous release already contained today’s 72.42. Four other models shifted by up to 0.34 points (Muse Spark 1.1, Muse Spark, Gemini 3.5 Flash, Qwen 3.7 Max): adding an observation to a benchmark changes the cohort statistics that drive the saturation check and shrinkage, and small knock-on movements are the honest consequence. The top two are unchanged and remain not separable from each other.
v2.0.19 · 2026-08-04 · A build defect in the previous entry, live for a few minutes. The v2.0.18 changelog text was assembled by a script that used a regex replacement, and the dollar signs in the new pricing figures collided with the replacement syntax: “$1” inside “$1.00” was interpreted as a pattern reference and swallowed a fragment of adjacent markup into the entry, unbalancing the page and briefly doubling the methodology export. Caught by the size jump in our own build output on the same deploy cycle, repaired, and the tooling now assembles changelog text with plain string operations that have no special characters at all. No score, weight or data cell was involved at any point; the defect was confined to the changelog markup itself.
v2.0.18 · 2026-08-04 · A correction to yesterday’s explanation, and pricing for both Inkling models. The v2.0.17 entry attributed the 5.9-point GPQA Diamond gap between vals.ai (83.59) and Artificial Analysis (89.5) on Inkling Small to an effort confound, resting on the claim that vals publishes no effort for the row. That claim was wrong: vals lists the row’s reasoning effort as 0.99, essentially maximum and the same regime as Artificial Analysis’s xhigh. The convenient explanation is withdrawn. The gap now stands as the largest unexplained evaluator disagreement in our 21-model GPQA comparison, and we would rather publish an unexplained number than a wrong explanation. Nothing about the score changes: the canonical column still decides, and since the vals value is the lower of the two, the published score errs conservative. Also in this release: Thinking Machines publishes first-party serverless pricing on its Tinker platform, verified directly against their documentation: Inkling at $1.00 in / $4.05 out per million tokens, Inkling Small at $0.30 in / $1.20 out, both with cached-input rates and a beta caveat carried in the data. Both models now appear on the Value view. Inkling Small at 51.15 for about $0.53 blended at the standard mix lands among the strongest value positions on the board, which is worth knowing precisely because its score is still Provisional on five cells.
v2.0.17 · 2026-08-04 · Inkling Small rises to 51.15 as vals.ai coverage arrives, and one hold turns out to be permanent. vals.ai now lists Inkling Small on GPQA Diamond at 83.59% (±2.04pp). vals is the canonical evaluator for that column, so the cell scores directly: 41.88 to 51.15, five of ten cells, Knowledge from 31.6 to 59.1. The disagreement with the held value is on the record, not swept. Artificial Analysis publishes 89.5 for the same model on the same benchmark, a 5.9-point gap, larger than any pair in the 21-model comparison we published with the calibration work. The likeliest explanation is effort, not evaluator error: Artificial Analysis labels its Inkling rows as xhigh-effort, vals publishes no effort for the row, and ARC Prize’s six effort rows show this model’s scores collapse as effort drops. We score the canonical column and keep the other value held, and this pair is now the strongest argument on file for evaluators publishing effort labels alongside scores. MMMU-Pro is a different story: vals does not list Inkling Small there and appears unlikely to, so the held 74 from Artificial Analysis is not pending, it is stranded, and the model’s Visual Reasoning stays empty under the current source policy. That is now a named cost of the no-mixing rule rather than a wait. No other model’s score moved.
v2.0.16 · 2026-08-04 · Qwen 3.8 Max joins the roster unscored, and the difference between “benchmarked” and “admissible” is the whole story. Artificial Analysis has evaluated Alibaba’s new flagship on four of our ten benchmarks: HLE 40.3, GPQA Diamond 92.2, MMMU-Pro 83 and Terminal-Bench 2.1 80.9. Under our source policy only the HLE row is admissible, because the other three are columns this board sources from vals.ai, which does not list the model yet. Those three values are held in the ledger with their reason, exactly as in the 2026-08-01 intake. One admissible cell is below the scoring floor, so Qwen 3.8 Max enters as insufficient data rather than with a score that is one cell away from nothing. Cells activate as vals coverage arrives. Also in this release: Inkling Small rises from 27.11 to 41.88. vals.ai now lists it on SWE-bench Verified at 82.2% (±1.71pp, Mini-SWE-agent harness), and since vals is the canonical evaluator for that column the cell scores directly, taking its Agency component from 15.5 to 48.8 on four scored cells. This also confirms what our evaluator calibration study concluded three days ago: the cells held at its intake were our harvest being older than the model, not vals declining coverage. The model was released five days after our previous vals capture, and coverage arrived within a week. No other model’s score moved in this release.
v2.0.15 · 2026-08-03 · The share buttons are now visible. The share row shipped this morning at a size that failed its one job: even someone who knew the buttons were there had to hunt for them. The buttons grow from 36px to 46px with brighter icons, and the row label steps up to match. A feature a visitor cannot see is a feature that does not exist. No score or data changed.
v2.0.14 · 2026-08-03 · The share card is now shareable, and the logo is clickable too. Two follow-ups on yesterday’s card viewer. The model’s logo now opens the card just like its name does - people try the logo first, so it should work. And the card viewer grew a share row: X, LinkedIn, Reddit, WhatsApp, Telegram and Facebook, plus copy-link, download-image, and the device’s native share sheet where the browser offers one (which can share the image file itself). Shares link to the model’s own page, and because that page carries the card as its preview image, the card renders wherever the link lands. No score or data changed.
v2.0.13 · 2026-08-03 · Clicking a model’s name now shows its share card. Every model has had a share card since the per-model pages shipped - the image social platforms render when a model link is posted - but the only people who ever saw one were crawlers, and the model name in the leaderboard did nothing when clicked. The name is now a button: it opens the card full-size over the page, with a route to the full detail view and a link to open the image itself for saving or sharing. Escape, the close button or a click outside dismisses it. The cards were already there; now the humans get to see them too. No score or data changed.
v2.0.12 · 2026-08-03 · The mark carries through, and it is bigger where you meet it first. Two follow-ups on the new brand mark. In the header it rendered noticeably smaller than the wordmark beside it, because the artwork carries its bar-chart element inside the same canvas; it now sits at 56px on desktop and 44px on the compact header, reading at parity with the site name - and it does so without regrowing the header bar that v2.0.8 disciplined, by letting the mark overflow the bar’s padding instead of stretching it. And the social preview cards - the homepage card and all 23 model cards - now carry the real mark instead of the old lettermark placeholder, republished at new versioned URLs so every platform fetches the rebranded card rather than serving its cached copy, which is the lesson v2.0.6 taught us. One correction surfaced by the rebuild: the homepage card still said 21 models tracked; the board has 23 since Inkling Small and DeepSeek V4 Flash joined in v2.0.9. The card now says 23 - the same class of stale rendered artifact v2.0.6 described, caught this time because regenerating forced us to look. No score or data changed.
v2.0.11 · 2026-08-03 · The site has a real mark. The header icon had been a placeholder since launch - a generic infinity glyph in a gradient box - and the browser-tab icon a lettermark circle. Both are replaced by AGI Ranker’s own infinity mark: two interlocked loops, one silver and one gold, with a rising bar chart growing out of the curve. The header uses a transparent render with edges built for our navy theme, so there is no dark fringe around the curves; the iOS home-screen icon uses a version pre-flattened onto the theme colour, because iOS composites transparent icons onto black and would have swallowed the mark. One honest limit: the mark is currently raster from a 960px source, so anything larger than 1024px - print, large banners - needs a vector redraw before it will hold up. No score or data changed.
v2.0.10 · 2026-08-03 · The model comparison section was reworked. Four things, all in the compare panel. It no longer opens pre-filled: a leftover "randomly select 3 models for demo" line from an early build had been seeding three picks on every page load, so the selector read "3 models selected" before you touched it and hit the four-model limit after one or two of your own; it now starts empty. Picks are legible: each of up to four models gets its own colour - indigo, emerald, amber, pink, in the order you pick them - shown as a filled dot and a coloured border on its row and as its line on the radar, so a row always matches its line and two models from the same lab no longer share a colour. The radar never jumps: before you pick anything it draws the empty five-axis web it is about to fill, as the same chart at the same size, so nothing resizes when the first model goes in. And the model list is ordered by AGI Score, highest first, within each tier, with Provisional and awaiting-verification models kept below the ranked ones. A note on this entry: these changes shipped as four small increments a few minutes apart (originally v2.0.10 through v2.0.13). We have folded them into this one entry, and will batch work like this into a single version in future, because Version History is here to tell you what changed since you last looked - not to list cosmetic bumps minutes apart.
v2.0.9 · 2026-08-01 · Two models added, and both of them show why our source policy has a cost. Inkling Small (Thinking Machines) and DeepSeek V4 Flash enter as Provisional on three scored cells each. That is thin, and the Provisional tier says so. They are on the board so that new evidence has somewhere to land. Six further cells were harvested and are not being used. Artificial Analysis publishes GPQA Diamond, MMMU-Pro and Terminal-Bench 2.1 for these models, and this board sources all three from vals.ai, which does not yet list them. Using the Artificial Analysis figures would put two different measurement stacks inside one benchmark column and call the difference capability, which is the practice we ruled out in v2.0.0. The gaps are measured, not hypothetical: vals runs 0.30pp below Artificial Analysis on GPQA Diamond, 5.71pp above on MMMU-Pro and 7.77pp below on Terminal-Bench 2.1, so an Artificial Analysis Terminal-Bench of 78.7 is about 71 on our column. A seventh cell is excluded for a different reason: LiveBench Agentic Coding is rebuilt here from their task table rather than their derived category score, and only the derived number was available. The cost is visible and we would rather show it than hide it. Inkling Small scores 89.5 on GPQA Diamond, above its own parent Inkling at 87.2, and that result cannot be used, so its published 27.11 understates what is known about the model. Every held cell is recorded in the ledger with its value and its reason, so nothing has to be re-harvested when coverage arrives. If vals.ai does not pick these models up, the source policy itself goes back on the table rather than the models being quietly rescored. No existing model moved: both new entries are Provisional and therefore outside the rankable median that drives shrinkage, so all 15 ranked scores are unchanged to twelve decimal places.
v2.0.8 · 2026-07-30 · The compact header now travels with the compact navigation. Since the last release the scrolling section rail appears below 1280px, but the full-size wordmark still scaled up at 768px, so laptops and tablets got the large brand stacked above the rail and a 128px header for no reason. The brand scale-up now happens at the same 1280px as the link bar. The result is one of two states at every width: 97px with the rail, 88px with the link bar. The Contribute button keeps its label from 768px up and back-to-top stays a touch-screen affordance, because those follow available width and input method rather than which navigation is showing.
v2.0.7 · 2026-07-30 · The Key Strength column is gone, because it was not telling you what it claimed to. It named each model’s highest-scoring component, and it read Visual Reasoning for 12 of the 15 ranked models. That is not a finding about those models. Visual Reasoning rests on a single benchmark, MMMU-Pro, scored against the benchmark maximum, and MMMU-Pro results cluster high, so the component runs higher than the others almost everywhere: across the board it averages 86.2 against Reasoning 75.1, Language 73.8, Knowledge 71.0 and Agency 54.0. The column was therefore reporting which component has the most generous scale, not what any model is actually best at, and a label that says the same word about four models in five carries no information while looking like it does. The capability heatmap already shows the full per-component picture without collapsing it to one misleading word. Removing the column also returns a meaningful slice of width to the table, which matters most on a phone. Also in this release: Version History has its own entry in the navigation, at the right-hand end next to the release number, on both the desktop bar and the mobile rail. This log is the part of the site a sceptical reader most wants to reach directly, and until now it could only be found by scrolling to the bottom of the methodology.
v2.0.6 · 2026-07-30 · Our social preview card was still showing pre-v2 numbers, weeks after the numbers changed. Anyone sharing a link to this site saw a card reading 14 live benchmarks, 20 models tracked and “multimodal”. The truth is 10, 21 and Visual Reasoning. The image file on our server was in fact correct and had been regenerated with the rest of the v2 release; what was wrong was our assumption that replacing a file replaces what the world sees. Social platforms cache preview images against the URL, so re-rendering the same filename changes nothing for anyone who has already shared the link, and we had verified the file rather than the card. Every preview image now lives at a versioned URL, the homepage card and all 21 model cards, which forces every platform to fetch it again. The previous files stay where they are, so an existing embed shows an old image rather than a broken one. The card generator carries the version token now, so this is a one-line change next time rather than a rediscovery. This is the same failure as the human-ceiling counts two releases ago: a rendered artifact that nobody re-derived after the thing underneath it changed. The site text was clean, and we checked: no page description, title or preview text anywhere on the site still carries the old figures.
v2.0.5 · 2026-07-29 · Tapping a specialty tab on a phone gave you a wall of text instead of scores. The specialty-index notice rendered 586px tall on a 375px screen, 72% of the viewport, and pushed the leaderboard 742px below the tab bar. A reader tapping Coding to see coding scores got a screenful of methodology and had to scroll to find a single number. Below tablet width these notices now collapse to a summary with a More expander, and the full text is one tap away in place. Collapsed, the notice is 151px and the table sits 307px from the tabs, so it is reachable in one short scroll. Nothing was deleted and nothing was softened. The visible summary is a compression, not a gentler version: it still says the specialty index is a 0 to 10 scale rather than the AGI Score, that 10 means a perfect score on every benchmark in the set, and that this is not a human-parity claim. Those are the three things a reader has to know before reading the number, so they stay visible whether or not anybody taps More. The same treatment went to the Value view note about cost being an estimate rather than a measured bill, which kept its estimate caveat in the collapsed line. Switching tabs always re-collapses, so no tab inherits the previous one's expanded state. The expander is a real button with aria-expanded, works without hover, and the Back to AGI Score control stays visible while collapsed. Desktop is unchanged: it has the room, so it shows the full text and no expander at all.
v2.0.4 · 2026-07-29 · The mobile tap states would not have rendered on an iPhone. The new section pills carry a pressed state, and it verified correctly in desktop Chromium, which is exactly the wrong place to check it. iOS Safari declines to apply the CSS active state to a link unless a touch listener exists on the element or one of its ancestors, so on the only kind of device that has a touch screen the pills would have looked inert when tapped. One empty listener on the page body fixes it for every tap target on the site. Recorded because verifying a touch behaviour in a mouse browser and calling it done is the kind of shortcut that ships a defect, and because the same trap applies to anything interactive we add from here.
v2.0.3 · 2026-07-29 · The site had no mobile navigation at all. On a phone the header rendered a wordmark and a Contribute button and nothing else. The six section links were desktop-only and the version pill was hidden below tablet width, so a visitor on a phone had no route to Corrections, Methodology or anything else, on a page that is otherwise one continuous scroll, and could not see which release they were reading. Small screens now get a compact sticky bar with the version pill visible and a scrolling row of section pills beneath it, carrying the same six anchors as the desktop nav. A row rather than a hamburger: there is no open state to get stuck, nothing to trap keyboard focus, and every section stays one tap away instead of two. Phones also get a back-to-top button once you are far enough down that scrolling back is a chore. Section anchors were landing behind the header at every width, because nothing on the page set a scroll margin; jumping to Corrections put its own heading underneath the sticky bar. Fixed for both layouts. Nothing else moved: the desktop bar is unchanged, and no score, weight or benchmark is touched by this release.
v2.0.2 · 2026-07-29 · We were overstating how much of the scale is anchored to humans. Three places on this site claimed that three or four scored benchmarks carry a measured human ceiling, and named OSWorld, FrontierMath and SimpleBench among them. The correct number is one. Of the ten benchmarks on the board, only GPQA Diamond (0.81) carries a measured human ceiling; the other nine are scored against the benchmark maximum, where 100 means a perfect score and no human claim is made at all. The other three did pass the same ceiling audit, and none of them is on the board: OSWorld (0.72) was retired in this very release, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. The claim was wrong in the AGI definition modal, in the methodology summary and in the full methodology, and the three did not even agree with one another. This is the same class of error as the correction we published two versions ago: a number that survived because nobody re-derived it after the thing underneath it changed. The AGI Score is a mixed scale, and it is a good deal more mixed than we were saying. Also in this release: the DeepSeek apology is signed by Barak Laniado, founder and CEO, and its corrections contact is now an email address you can actually write to rather than a link back into the site. The calibration constants box no longer reads “Current as of v1.4” under a v2 banner; the constants are unchanged and were re-confirmed for v2.0.0, which is what it now says. Three hover styles on the corrections card and twelve layout classes on the apology page were also missing from the compiled stylesheet, which is purged to what the leaderboard uses, so they had been failing silently.
v2.0.1 · 2026-07-29 · The DeepSeek apology gets its own page. The correction we published in v2.0.0, an ARC-AGI-2 score attributed to DeepSeek V4 Pro that carried no source at all, now has a dedicated page at /corrections/deepseek-arc-agi-2, linked from its entry in the log above. An apology buried as one card among several is easy to walk past, and this one should be readable, citable and linkable on its own. The page sets out what we published, why it was wrong, what it did and did not affect, and what changed in the process so that it cannot recur: under methodology v2 a recorded source is a condition of scoring rather than an expectation, and all 156 scored cells carry one. The wording of the apology itself is unchanged and identical in both places. One count in that entry was also wrong. It said three further cells were voided in the same pass and then referred to four of them in the next sentence. Four further cells were voided, five in total including the DeepSeek one. Corrected here rather than quietly.
v2.0.0 · 2026-07-29 · Methodology v2. Every score falls, and no model got worse. Two changes drive it. First, human ceilings. Because a score is normalised as raw divided by a human ceiling, that ceiling is not only an anchor for what counts as human level, it is a volume knob: it multiplies both the level and the spread of a benchmark, and so how hard that benchmark pushes on its component. Most of ours had no measurement behind them. Humanity’s Last Exam sat at 0.50, a number our own internal audit described as a policy floor rather than a finding, and that setting quietly made HLE count double. From now on a ceiling is either a published human result under a protocol comparable to the models’, or it is simply the benchmark maximum and we make no human claim at all. Only four survived the test: GPQA Diamond at 0.81, OSWorld at 0.72 (which we raised, having carried 0.85 against our own recorded measurement), FrontierMath at 0.35, and SimpleBench at 0.837, where we had rounded a human baseline up. Eleven moved to 1.00. Three component scores that previously sat above 100, on a scale whose 100 is meant to be the human mark, no longer do. Second, Agency was rebuilt on four legs across three independent evaluators: SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking and LiveBench Agentic Coding. Terminal-Bench now comes from vals.ai rather than Artificial Analysis, cutting our reliance on any single evaluator. τ³-Banking enters at benchmark maximum, so its true range shows: the best model completes about a third of these stateful banking workflows. Retired: SWE-bench Pro, OSWorld, BrowseComp, Tau-bench retail and airline, Aider Polyglot, Terminal-Bench 2.0. MCP Atlas demoted to informational. The Tool Use tab is withdrawn until a second clean non-coding benchmark exists. Scores fell by about 6 points from the ceiling change and further from the Agency rebuild; ordering shifted only locally, and the top two are unchanged. Third, evaluator concentration. Artificial Analysis had been supplying 45% of the whole board, 97% of Knowledge and 100% of Visual Reasoning, so one evaluator changing terms would have taken two components to zero. GPQA Diamond and MMMU-Pro moved to vals.ai alongside Terminal-Bench 2.1, taking Artificial Analysis to 26% of the board and splitting Knowledge 52/48 between two evaluators. Every switch was measured before it was made rather than after, and each offset is published rather than absorbed: GPQA Diamond −0.30pp with 1.38pp scatter across all 20 models, MMMU-Pro +5.71pp with 1.67pp scatter and vals higher on all eleven comparable models, Terminal-Bench 2.1 −7.77pp with 4.54pp scatter and vals lower on 18 of 19. One row is held rather than used: vals reports Grok 4.5 at 61.8 on MMMU-Pro, below Grok 4.3 at 83.1 and below xAI’s own non-reasoning variants, which is not a capability measurement we can account for, so we do not use it. Grok 4.5 consequently has no Visual Reasoning cell and falls three places. Visual Reasoning is still single-source, having moved from wholly Artificial Analysis to wholly vals.ai: it rests on one benchmark, so this relocates the dependency rather than removing it, and we would rather say so than imply it is solved. Also in this release: the Reasoning specialty tab is withdrawn. It ran on ARC-AGI-2 and AIME 2025, but AIME 2025 is scored for a single model and sits at 98-99% where it appears, so the tab was ranking on ARC-AGI-2 alone while labelled as using two benchmarks. Withdrawn under the same rule as Tool Use rather than kept with a misleading label; it returns when AIME 2026 is harvested. Cursor Composer 2.5 is removed from the board: it is a product system rather than a model, so under v2 it scores no cells at all, and listing it implied we were waiting on evidence that was never coming. Specialty tabs now report a 0 to 10 index rather than a 0 to 100 score. On the AGI Score, 100 means the genesis of AGI; a specialty score never meant that, and showing both on the same scale implied they were the same kind of claim. Ten now means a perfect score on every benchmark in that set, with no human comparison implied, because most of those benchmarks have no published human study to compare against. The Multimodal component is renamed Visual Reasoning. It rests on a single scored benchmark, MMMU-Pro, which is an exam built on diagrams and charts, so calling 11% of the score “multimodal” claimed a breadth of perception we do not measure. Visual Reasoning says what it is: interpreting image data, which underpins reading a chart, extracting a figure from a document, or working from a screenshot. The AGI definition was revised in the same pass and for the same reason. It previously excluded “sensory perceptions whatsoever”, which contradicted our own scoring of a diagram-based benchmark. Rather than bolt vision on, which would have raised the obvious question of why sight and not hearing or smell, we removed the sensory clause and let the physical-body clause carry the weight: what reaches a model as data is in scope, what needs a body to acquire is not. The definition got shorter. The corrections summary counters have also been replaced. They had been hardcoded since May and one of them read “1 model entered scoring” for two months without meaning anything; they are now derived from the scoring ledger on every release.
v1.11.28 · 2026-07-25 · Corrections to our corrections. We audited this log against our own commit history and found that two entries were wrong. Both times we compared two different configurations of the same benchmark and reported the difference as a laboratory overstating its results. DeepSeek did not overstate GPQA Diamond: the 72.9% we published as an independent contradiction is DeepSeek’s own non-thinking-mode figure, reaching us through a secondary aggregator and set against their thinking-mode result. Measured like for like, their self-report sits 1.3 points above independent measurement rather than 17.2 points above. Moonshot did not overstate Humanity’s Last Exam: their 54.0% was explicitly labelled as a with-tools result, and the same post published 36.4% without tools, within half a point of independent measurement. Our source policy had excluded them on the strength of that misreading. Both entries are retracted and both originals are retained rather than deleted. We have also withdrawn a laboratory promotion test that the policy described but that was never run, and corrected an entry which presented source relabelling as new independent measurement. No score, weight, human ceiling or source tier changed in this release.
v1.11.27 · 2026-07-25 · Claude Opus 5 and Anthropic roster correction. Added Claude Opus 5 on five independent, variant-distinct frozen-v1 cells: ARC-AGI-2 90.4%, SWE-bench Verified 97.0%, HLE 52.59%, GPQA Diamond 93.23%, and MMMU-Pro 84.74%. The unchanged eligibility engine places it Provisional at 97.31 because one Agency cell and no accepted LiveBench row leave two thin components. The available LiveBench EAP row remains held pending public-model identity resolution; Terminal-Bench 2.1, ARC-AGI-1 and ARC-AGI-3 remain outside frozen-v1 scoring. Anthropic’s active latest-two roster is now Opus 5 and Opus 4.8. Fable 5 leaves the current roster under an identity-integrity exception because fallback-enabled, fallback-disabled and generic rows cannot be combined into one standalone model; Opus 4.7 thinking and base retire as older generations. All historical evidence remains preserved, with no transfer or imputation. No formula, weight or source tier changed.
v1.11.26 · 2026-07-24 · Grok 4.5 ARC-AGI-2 update. Added ARC Prize’s independently verified Grok 4.5 High result on the 120-task ARC-AGI-2 Public Eval at the official displayed precision of 52.6%. High is the single frozen-v1 scoring configuration; the equal-scoring Medium result and the Low result remain supporting provenance only. ARC-AGI-1 and ARC-AGI-3 remain informational, and no formula, weight, source tier, roster entry or unrelated evidence changed. No imputation.
v1.11.25 · 2026-07-22 · Capability Heatmap isolated from custom weights. The heatmap now reads a separately stored canonical snapshot of default-weight AGI Score, component values, coverage and tier. Explorer sliders and presets continue to recompute AI Score and leaderboard order without changing any heatmap value, order, colour, tooltip or accessible label. This fixes the state-dependent path left open by v1.11.24, whose direct score accessor still pointed to mutable custom-weight state. No score, rank, component, benchmark, evidence, eligibility or methodology changed. No imputation.
v1.11.24 · 2026-07-22 · Capability Heatmap AGI Score binding locked. The rightmost heatmap column now declares its full-precision value, colour and accessible label directly from the same canonical final score field used by the leaderboard, model details and Explorer; its default order is explicitly canonical AGI Score descending. A no-cache v1.11.23 preflight already rendered the canonical values, so this release changes no score, rank, component, benchmark, evidence, eligibility or methodology. It adds a machine-verifiable binding contract and rendered-DOM regression coverage against an Agency-column regression.
v1.11.23 · 2026-07-21 · Inkling and July evidence refresh. Added Thinking Machines’ open-weight Inkling on six independent cells: ARC-AGI-2 36.5%, SWE-bench Verified 77.6%, LiveBench Overall 71.68%, HLE 29.70%, GPQA Diamond 87.17%, and MMMU-Pro 73.47%. The unchanged eligibility engine places Inkling Ranked #17 at 70.97; it remains Coding-ineligible at 1/3 because SWE-bench Pro and Terminal-Bench 2.0 are absent. Refreshed every exact mapped row from the current LiveBench snapshot and current Artificial Analysis HLE, GPQA, and MMMU-Pro tables, including newly available Gemini 3.5 Flash cells; Flash moves from Provisional to Ranked #8. Removed Claude Fable 5’s HLE cell because AA explicitly identifies that row as an Opus 4.8 fallback configuration; routed evidence does not score as pure Fable. Vals SWE-bench Verified and ARC-AGI-2 rows were audited across the active roster. AA Intelligence Index, Terminal-Bench 2.1, SciCode, routed results, and incompatible variants remain unscored. No formula, weight, source tier, or historical artifact changed. No imputation.
v1.11.22 · 2026-07-18 · Kimi K3 SWE-bench Verified update. Added Kimi K3’s independently evaluated 93.4% SWE-bench Verified result from Vals AI using the Mini-SWE-agent harness. The fifth scoreable cell adds Agency evidence and moves K3 from Provisional to Ranked #2 under the existing frozen-v1 eligibility engine; its AGI Score changes from 89.87 to 88.38 because the new Agency component is below its prior four-component aggregate. K3 has one of three Coding cells and remains Coding-ineligible because SWE-bench Pro and Terminal-Bench 2.0 are missing. Terminal-Bench 2.1 remains held as a distinct version, and ARC-AGI-2 remains missing. No formula, weight, source tier, roster, or existing K3 cell changed. No imputation.
v1.11.21 · 2026-07-17 · Kimi K3 Frontier Integration. Kimi K3 joins as Kimi’s latest general-purpose flagship, alongside previous general-purpose generation Kimi K2.6. Kimi K2.7 Code is a specialist coding branch; it remains historically documented but leaves the active latest-two roster because it currently qualifies for neither the main Ranked leaderboard nor the frozen-v1 Coding leaderboard. K3 enters Provisional #2 at 89.87 on four independent T1 cells: LiveBench Overall 77.9%, GPQA Diamond 93.5%, Humanity’s Last Exam 44.3% (AA text-only, no tools), and MMMU-Pro 81.0% (AA multimodal, 10-option, no tools). K2.6 remains Ranked #9 at 79.98 and Coding-eligible. Moonshot’s conflicting 56.0% HLE with-tools self-report is retained as rejected provenance, not averaged or scored. Terminal-Bench 2.1 at 85.0% is held because v1 scores Terminal-Bench 2.0; SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.0 and ARC-AGI-2 remain missing. The specialist-slot rule is provider-neutral. No scoring formula or benchmark weight changed. No imputation.
v1.11.20 · 2026-07-14 · Visual Identity Patch. Added official model-family and provider marks across AGI Ranker’s model surfaces, with local assets, accessible fallbacks and no scoring changes.
v1.11.19 · 2026-07-14 · Frontier fairness refresh. Muse Spark 1.1 gains four scoreable cells from the approved source harvest: LiveBench 76.23%, Scale SWE-bench Pro 61.5%, Terminal-Bench 2.0 80.0%, and OSWorld-Verified 80.8%. It is now Ranked #4 at 86.76 and places #5 in Coding at 79.64. Gemini 3.5 Flash enters as Provisional at 81.75 using LiveBench, SWE-bench Verified, ARC-AGI-2, and OSWorld-Verified. Gemini 3 Pro leaves the active board under the latest-two Google roster rule, with its historical snapshot retained. GPT-5.6 Sol remains #1 at 93.33 after the frozen roster-sensitive scoring mechanics recompute. Terminal-Bench 2.1 and unsupported Gemini benchmark variants remain held outside production. No methodology or weight change. No imputation.
v1.11.18 · 2026-07-12 · Muse Spark 1.1 and Sol pricing. Added Muse Spark 1.1 as a separate Meta model with three scoreable independent cells: SWE-bench Verified 82.0% (vals.ai, Mini-SWE-agent), GPQA Diamond 89.8%, and Humanity’s Last Exam 45.1% (Artificial Analysis thinking configuration; HLE no-tools). It enters as Provisional at an AGI Score of about 87.15. AA Intelligence Index v4.1 at 51 and preliminary LM Arena Text Elo 1490 ±10 are retained as informational evidence only. The existing Muse Spark generation remains separate, and its MMMU-Pro 81% result was not transferred to 1.1. Added official GPT-5.6 Sol API pricing of $5 input, $0.50 cached input, and $30 output per million tokens. No benchmark methodology or weights changed. No imputation.
v1.11.17 · 2026-07-10 · Composer benchmark identity corrected. Removed Cursor Composer 2.5’s 47% result from the SWE-bench Pro column after confirming that the underlying Artificial Analysis result belongs to the distinct SWE-Bench-Pro-Hard-AA system benchmark. Composer remains Coding-eligible through SWE-bench Verified and Terminal-Bench. Because Coding scores renormalise over eligible available cells, its score changes from 70.50 to 80.56 and its rank moves from #10 to #4. No other Coding score values change, though ranks #4–#9 shift down one position. The removed evidence is retained internally for future Coding Systems consideration. The Ranked AGI leaderboard is unchanged. No imputation.
v1.11.16 · 2026-07-09 · GPT-5.6 Sol added. Added GPT-5.6 Sol from six screenshot-confirmed current-schema cells: ARC-AGI-2, GPQA Diamond, HLE, LiveBench, MMMU-Pro and SWE-bench Verified. Sol enters the Ranked leaderboard immediately at #1 (AGI Score about 92.6 after roster re-equilibration) using Max-effort rows across sources. With Sol now Ranked, the GPT-5.4 family was retired under AGI Ranker’s latest-two-OpenAI-generations roster rule. GPT-5.6 Terra/Luna and Muse Spark 1.1 were not added pending separate scoreable coverage. AA Intelligence Index remains informational and unscored.
v1.11.15 · 2026-07-09 · LiveBench official-board refresh. Updated seven LiveBench Global Average cells from the official board and added Grok 4.5’s LiveBench result. Grok 4.5 now clears the Ranked floor at #4. GLM 5.2 also moves from Provisional to Ranked via saturation/coverage recomputation; no new GLM benchmark cell was added. GPT-5.6 Sol is not yet included pending Sol-distinct scoreable coverage.
v1.11.14 · 2026-07-09 · Grok 4.5 added (Provisional). xAI Grok 4.5 enters on four independent T1 cells: AA HLE 0.40 and GPQA Diamond 0.93 (source rows “Grok 4.5 (high)” / effort knob, not a separate SKU), AA MMMU-Pro 0.80, and vals.ai SWE-bench Verified 0.866 (Mini-SWE-agent). Both cores present; stays Provisional (fifth independent cell still required for Ranked; agency and language coverage remain thin). Grok 4.3 retained under the latest-two-versions scope rule. GPT-5.6 family not added (no Sol-distinct row with scoreable coverage). AA Intelligence Index tracked as informational only (not in score). No imputation. (HLE corrected 0.38→0.40 after bar-label re-check.)
v1.11.13 · 2026-07-01 · MiniMax M2.7 LiveBench harvest. MiniMax M2.7 gains LiveBench Global Average 0.6349 from livebench.ai (source row “Minimax M2.7”). Fourth independent T1 cell; stays Provisional (one more independent cell needed for Ranked). No AA Intelligence Index ingested. No imputation.
v1.11.12 · 2026-07-01 · Kimi K2.7 Code harvest + Composer SWE-V. Kimi K2.7 Code enters as Provisional on four independent T1 cells: LiveBench Global Average 0.7189, GPQA Diamond 0.896, AA-standardised no-tools HLE 0.328, and vals.ai SWE-bench Verified 0.782 (Mini-SWE-agent harness). Provisional AGI Score about 75.8 on 4/14 benchmarks across all four components; one more independent cell needed for Ranked. Cursor Composer 2.5 gains vals.ai SWE-bench Verified 0.796 (Cursor CLI harness), lifting its Coding specialty from about 65.6 to 70.5. No AA Intelligence Index ingested. No imputation.
v1.11.11 · 2026-06-29 · GLM 5.2 gains ARC-AGI-2. ARC Prize now lists GLM-5.2 with ARC-AGI-2 at 22.8% (dated 2026-06-13), so the cell is added from the independent leaderboard. GLM 5.2 now appears on the Reasoning tab at about 26.8, while its overall Provisional score settles around 73.5; the lower score is expected because a hard missing reasoning benchmark is now measured rather than absent. It stays Provisional because agency coverage remains just under the core floor and multimodal coverage is still empty. No imputation.
v1.11.10 · 2026-06-27 · GLM lab label + performance housekeeping. GLM 5.1 and GLM 5.2 now display their lab as Z.AI instead of the generic "other" provider bucket everywhere the public data is served. Also replaced the browser-side Tailwind CDN compiler with a precompiled local stylesheet and lazy-load the optional radar-chart library, reducing startup JavaScript without changing scores or methodology.
v1.11.9 · 2026-06-19 · GLM 5.1 demoted to Provisional. Two BenchLM cells are voided: SWE-bench Pro was a Provider-exact relay of Z.AI's own figure (not an independent eval), and Tau-bench Retail no longer exists on BenchLM (replaced by tau-Telecom, a different benchmark we do not score). GLM 5.1 now rests on four independent T1 cells only. Provisional AGI Score about 71.9. No imputation.
v1.11.8 · 2026-06-19 · GLM 5.2 harvest: vals.ai + LiveBench. Five independent T1 cells now ground GLM 5.2: GPQA Diamond 0.86, SWE-bench Verified 0.83, and Terminal-Bench 2.1 0.68 from vals.ai (replacing earlier Artificial Analysis figures on the overlapping benchmarks), HLE 0.40 from our standardized AA source, and LiveBench 0.76. LM Arena Text Elo 1471 is published. Stays Provisional: agency sub-weight coverage sits just below the 40% core floor alongside absent multimodal data. No imputation.
v1.11.7 · 2026-06-17 · GLM 5.2 graduates to Provisional. Same day it launched, Artificial Analysis published independent results, so GLM 5.2 moves from Awaiting Verification to Provisional on three T1 cells: GPQA Diamond 0.89, Terminal-Bench 2.1 0.75, and HLE 0.40. The independent numbers came in below the lab's launch self-reports (which we had excluded) - exactly why we wait for them. It needs a fifth independent cell to reach Ranked. No imputation.
v1.11.6 · 2026-06-17 · GLM 5.2 added (Awaiting Verification). Z.AI released GLM 5.2 on June 16. So far only the lab's own launch-day numbers are public (its model card and a benchmark aggregator relaying the same figures), which we exclude as unverified self-reports. GLM 5.2 is listed as Awaiting Verification - no score - until an independent evaluator publishes results. No imputation.
v1.11.5 · 2026-06-16 · Hero badge. The launch-month pill now reads "Since May 2026" - a founding mark rather than a stale-looking date stamp. No scoring change.
v1.11.4 · 2026-06-16 · Count consistency. The hero subhead and structured data still read "15 benchmarks" after yesterday's removal of the AA Index; both now read 14, matching the live count. No scoring change.
v1.11.3 · 2026-06-16 · AA Intelligence Index removed from the score. Artificial Analysis shipped Intelligence Index v4.1, an explicitly agentic composite (GDPval, Terminal-Bench 2.1, τ³-Banking, SciCode, HLE, GPQA and more). Because it re-bundles benchmarks we already score directly, keeping it as our Language signal would double-count those results and mislabel agency as language - so we removed it from the AGI Score (14 live benchmarks now). Language rests on LiveBench; a dedicated language benchmark is on the roadmap. Most scores rise ~0.3-0.5 (the Index had been a slight drag); the ranking is unchanged. We still track the AA Index as an external reference.
v1.11.2 · 2026-06-16 · Cleaner number type. Scores and stats now render in Inter with tabular figures - more legible in the dense tables than the previous display face, and consistent across the site. Headings and the wordmark are unchanged. No change to any score.
v1.11.1 · 2026-06-16 · Value view polish. Refined the value-for-money UI and method after testing: the ranking now leads with the picks (Top / Best value / Budget) instead of a raw value index that was hard to read at a glance, the cost column is labelled $/1M tokens for clarity, and the methodology page now spells out exactly how value for money is derived. No change to any score.
v1.11.0 · 2026-06-15 · Value view. The Value tab ranks models by capability against an estimated API cost. Pick a capability area - Overall (the AGI Score) or Coding, Reasoning, Knowledge, Tool Use - and the table shows the exact same score that area's own tab shows, next to an estimated cost at that area's typical token mix (cache-aware). Sort by capability or by value for money, and read the picks at a glance: top capability, best value, and budget. A value-frontier graph plots the same models. Also renamed Google to Google DeepMind, and renamed the Agentic specialty to Tool Use (coding is also agentic, so the old name was ambiguous; Tool Use covers computer use, web browsing and tool/function calling beyond code). No change to the canonical AGI Score or any benchmark data.
v1.10.1 · 2026-06-15 · Full data refresh, plus MiniMax M2.7 (Provisional). A complete re-harvest of every model from independent sources corrected several stale or mislabeled cells (a DeepSeek score that was far too low, a Qwen coding number drawn from the wrong benchmark, and more), filled gaps, and upgraded many cells to higher-trust independent measurements. Effect on the board: the top stays a near-tie between ChatGPT 5.5 and Claude Opus 4.8; Qwen 3.7 Max earns a canonical rank as new coverage completes its profile; Gemini 3 Pro and DeepSeek V4 Pro rise on corrected data. MiniMax M2.7 (the prior MiniMax flagship) joins as Provisional. Every cell traces to an independent source. No imputation.
v1.10.0 · 2026-06-15 · MiniMax M3 added, Ranked at #6. The new MiniMax flagship enters scored entirely from independent sources (Artificial Analysis, vals.ai, benchlm, LiveBench) rather than the lab's own numbers. It is strong on multimodal and general knowledge, weaker on language, and its independently-measured score lands well below its launch claims, which is exactly why we wait for third-party data. Priced low, so it places well in the Value view. No imputation.
v1.9.3 · 2026-06-14 · Accessibility 100. Underlined the last in-text link that was distinguishable by color alone. Lighthouse: Accessibility 100, SEO 100, Best Practices 100 (desktop Performance 98).
v1.9.2 · 2026-06-14 · Accessibility and SEO pass. Form controls and the sort menu now carry proper labels, in-text links are underlined rather than color-only, the page has a main landmark for screen readers, and muted text is lightened to readable contrast. Two action links became buttons so search engines crawl every link. Lighthouse accessibility rises from 70 toward the high 90s, with no change to any ranking or score.
v1.9.1 · 2026-06-14 · Copy fix: 15 live benchmarks. The homepage, methodology, and page metadata now read 15 live benchmarks, matching the count the leaderboard has been computing for a while. The static text had lagged the live data by one. No scoring change.
v1.9.0 · 2026-06-12 · Sensitivity bands on every score, plus a Claude Fable 5 update. Each score on the main leaderboard now shows a sensitivity band (for example 87.02 ±3.6): we remove each of a model's benchmarks one at a time, re-run the entire scoring pipeline, and report how far the score moves. It is not a confidence interval - it shows how much a score depends on which benchmarks exist, and it widens honestly when coverage is thin. When two models' bands overlap, read their order as a statistical tie. Separately, Claude Fable 5 gains its Humanity's Last Exam result from our standardized independent source, lifting its provisional score to about 95.8. It stays Provisional by explicit editorial hold, with the reason published in the open dataset and shown on its badge: ARC-AGI-2 has not yet been run on Fable 5, and every ranked model's reasoning score includes that benchmark, so ranking it today would compare unlike baskets. It ranks the day that result publishes. No imputation, and no corner-cutting in either direction.
v1.8.1 · 2026-06-10 · Site visibility upgrade. The full technical methodology now lives at its own address (/methodology) rather than only inside a popup. Every model page now carries a written summary and a static benchmark table with source attribution, readable even without JavaScript - the same numbers as the interactive view. Added robots.txt and per-page structured data so search engines can find and understand every page. The open dataset (models.json) is now formally licensed CC BY 4.0 - cite it freely with attribution.
v1.8.0 · 2026-06-09 · Claude Fable 5 added, on launch day. Anthropic's new generally-available flagship (its Mythos-class model with production safeguards) enters as Provisional. It posts the strongest agentic-coding and knowledge results on the board, but day-one coverage is uneven across components, so it is not yet canonically ranked. Two launch figures arrived above their benchmarks' human ceilings from a single source; we are holding both pending independent confirmation rather than letting them inflate the score. Fable lifts to ranked once independent evaluations complete its profile. No imputation, even for the year's most anticipated launch.
v1.7.3 · 2026-06-09 · ARC-AGI-2 cleanup: one effort-tier fix, one version fix. Completing the High-effort standardization, ChatGPT 5.4's ARC-AGI-2 is now read at the same High tier as the rest of the column. Separately, a GLM ARC-AGI-2 result we had attributed to GLM 5.1 in fact belonged to the previous version, GLM 5; with no GLM 5.1 result published on that benchmark, the cell is removed rather than guessed. The correction lifts GLM 5.1 and takes it off the Reasoning leaderboard, where it no longer has a qualifying result. No imputation.
v1.7.2 · 2026-06-09 · ARC-AGI-2 read at one effort level. Reasoning models can be run at several effort settings, and one model on the board had been carried at a higher tier than the rest. We now standardize ARC-AGI-2 on the High-effort tier so the column compares like-for-like. Claude Opus 4.8 (thinking) gains its ARC-AGI-2 result and joins the Reasoning leaderboard. Effect on the top: ChatGPT 5.5 and Claude Opus 4.8 stay within a fraction of a point, with ChatGPT 5.5 nominally first. No imputation.
v1.7.1 · 2026-06-09 · SWE-bench Verified standardized on one independent source. Every model's SWE-bench Verified result now comes from the same independent third-party leaderboard (vals.ai), replacing a mix that leaned on a benchmark host whose public table has not refreshed since February. The result is consistent, current agentic-coding measurement across the whole column - and it tightens the top of the board: Claude Opus 4.8 (thinking) and ChatGPT 5.5 now sit within 0.2 points, effectively tied, with Opus 4.8 nominally first. No imputation.
v1.7.0 · 2026-06-05 · HLE standardized on one independent source; Opus 4.8 joins the ranked board; Opus 4.6 retired. Humanity's Last Exam is now sourced consistently from Artificial Analysis (independent, no-tools) for every model it covers, replacing a patchwork of sources that disagreed by up to 20 points on the same model. Claude Opus 4.8 enters the ranked leaderboard at #2 now that independent reasoning and language scores complete its coverage. Per our scope rule (latest two versions per lab), Claude Opus 4.6 leaves the board. No imputation.
v1.6.1 · 2026-06-04 · Value view polish. Added a reset for the workload-mix slider (back to the standard 75% input / 25% output), spaced out overlapping model labels on the scatter, and paused weight customization while the Value view is open (weights don't change price-performance).
v1.6.0 · 2026-06-04 · New: the Value view - capability per dollar. A price column (input/output API cost per million tokens) plus a Value tab that plots AGI Score against blended API price and highlights the value frontier - the best score available at each price point. A workload-mix slider weights input vs output cost for your use case. Pricing covers current frontier models; superseded or preview-only models without public pricing are labeled as such. Best value among frontier models - we do not track budget tiers. AGI Score remains the default view.
v1.5.3 · 2026-05-31 · Source upgrade for Claude Opus 4.8 (thinking). Its SWE-bench Verified result is now drawn from an independent third-party evaluation that corroborates the launch-day figure, replacing the lab self-report we carried at launch. Same value, higher-trust source; the model stays Provisional pending independent reasoning data. No imputation.
v1.5.2 · 2026-05-29 · Claude Opus 4.8 (thinking) added the day after launch, as Provisional. We hold its GPQA Diamond and AA Intelligence Index (independent T1) plus launch-day agentic results (SWE-bench Verified/Pro, OSWorld); reasoning-component coverage is still too thin for a canonical rank, so no AGI Score position yet. It lifts once independent reasoning and second-source agentic benchmarks publish. Claude Opus 4.6 stays on the board until then. No imputation.
v1.5.1 · 2026-05-23 · Roster expansion to 18 models. Added Qwen 3.7 Max (Provisional), Cursor Composer 2.5 (System entry - scores reflect the full product harness, not weights alone), and MiMo V2.5 Pro. Lifted Claude Opus 4.7 and Claude Opus 4.6 (thinking) from awaiting-verification to Provisional after new variant-distinct evaluations surfaced.
v1.5.0 · 2026-05-17 · Coding specialty restructured to reflect agentic-coding reality. Real coding in 2026 happens via tools like Claude Code, Codex, Cursor - not bare LLM completion. Composition: SWE-bench Verified 35 / SWE-bench Pro 30 / Terminal-Bench 35 (replacing Aider Polyglot, which had zero variant-distinct roster coverage after Round 14 cleanup). 9 of 15 roster models now have Terminal-Bench data (Round 25 harvest from vals.ai T1 + benchlm.ai T2 second-source). Added contamination note for SWE-bench Verified and harness disclosure for Terminal-Bench on the methodology page. Top of leaderboard shifts: GPT-5.5 takes #1 in Coding by ~0.7pp over Claude Opus 4.7 (thinking) - within statistical noise of the benchmarks' error bars, which is itself worth surfacing honestly. Each tab now answers its question directly.
v1.4.15 · 2026-05-15 · Homepage OG card gains a call-to-action: emerald pill-shaped "See live rankings →" in the bottom-right corner, replacing the previous "Built with obsessive care for truth" tagline. Closes the last issue OpenGraph debugger flagged ("Missing call-to-action in your image"). The CTA is visual-only inside the PNG (not a clickable link - it's part of the social-share image), but it telegraphs the action a viewer should take if they click through.
v1.4.14 · 2026-05-15 · Homepage OG image fixed: regenerated at native 1200x630 / 408KB (was 2400x1260 / 1.14MB - exceeded WhatsApp's <600KB ceiling and was over-spec for OG's recommended dimensions). Same fix already applied to per-model OG cards in v1.4.11; this brings the homepage card into line. Also extended page title from 47 to 51 characters ("AGI Ranker - Open AGI Score for Frontier AI Models") to land in the optimal 50-60 char window for SERP and OG previews.
v1.4.13 · 2026-05-13 · Developer-mode analytics toggle. Visit /?dev=1 on any browser/device to disable Vercel Analytics tracking for that browser (sets the localStorage va-disable flag Vercel respects); /?dev=0 to re-enable. Confirmation toast appears for ~3.5s. URL is cleaned after action so the param doesn't persist on refresh or share. Designed for Barak's own repeat visits to not inflate metrics, without requiring browser DevTools console access (which is unrealistic on mobile).
v1.4.12 · 2026-05-13 · "Last updated" freshness indicator added to the leaderboard header line. Auto-bumps every commit that touches models.json - the date comes from the HTTP Last-Modified header Vercel sets on the file, so no manual maintenance. Complements the existing "Most recent eval" date (which signals source-side freshness) with a "Last updated" date (which signals our integration activity). Visual cue for returning users that the site is actively maintained.
v1.4.11 · 2026-05-13 · Per-model OG cards properly delivered to social-media crawlers. Pre-generated 15 per-model HTML files at /model/{slug}.html with per-model meta tags (title, og:image, og:description, canonical URL). Social-media bots don't execute JavaScript, so the v1.4.10 client-side meta updates never reached them - they saw only the homepage card. With per-model HTML now served via Vercel cleanUrls, OG/Twitter/WhatsApp/Slack previews show the correct model-specific card and copy. Also reduced PNG file size from 1.14MB to ~380KB (native 1200x630 instead of 2x device scale) to fit WhatsApp's <600KB ceiling.
v1.4.10 · 2026-05-13 · Per-model OG cards. Each of the 15 model URLs now has its own social-share preview image at /og/{apiName}.png, showing the model's AGI Score, tier badge, 5-component breakdown, and rank within the canonical view. Social shares of /model/{slug} URLs (X, LinkedIn, Slack, Discord, iMessage) now display model-specific imagery rather than the homepage card. Cards generated via tools/regenerate_model_og_cards.py - re-run whenever scores materially change.
v1.4.9 · 2026-05-13 · Capability heatmap shipped. New section between Corrections Log and By Capability Area showing every scoreable model on every cognitive component in one grid. Color-coded by score (rose → amber → emerald), sorted by AGI Score descending, with PROVISIONAL rows visually flagged. Each model name links to its /model/{slug} detail page. Designed to be screenshot-shareable - one image tells the strengths-and-weaknesses story across the roster.
v1.4.8 · 2026-05-13 · Hero clarity rewrite. The subhead now states what AGI Ranker measures (how close each frontier AI is to AGI) and defines the score scale (0-100, with 100 as the AGI threshold) up-front instead of burying that context in the modal. Added a second smaller paragraph surfacing the three trust differentiators: independent-verification preference, no-imputation policy, public Corrections Log. Improves mass-market readability without weakening the rigor signal.
v1.4.7 · 2026-05-13 · Per-model URL routing shipped. Each model in the roster now has a shareable, SEO-indexable URL at /model/{apiName}. The existing detail view opens automatically when the URL is loaded directly, and clicking "View" on a leaderboard row updates the URL via History API. Document title and meta description update dynamically per model. sitemap.xml added with all 15 model URLs.
v1.4.6 · 2026-05-13 · Corrections Log refreshed to reflect Round 20: source-tier upgrades table now lists the Claude Opus 4.7 (thinking) SWE-V T2→T1 swap alongside the headline DeepSeek correction. New section added for the four Round 20 vals.ai T1 additions (GPT-5.5, GPT-5.4, Gemini 3.1 Pro Preview, Claude Opus 4.6 Thinking on SWE-bench Verified). Summary stats updated: 14 cells corrected, 23 cells from independent verification.
v1.4.5 · 2026-05-13 · Round 20 follow-up: Claude Opus 4.7 (thinking) SWE-bench Verified source-tier upgrade. The vals.ai Settings panel confirmed the unsuffixed "Claude Opus 4.7" row at vals.ai was tested with "Thinking Type: Adaptive" (= Anthropic thinking variant per our convention). Replaced the existing T2 benchlm.ai cell (0.876) with the T1 vals.ai cell (0.820). Lower value, higher trust weight - exactly the methodology pattern: independent T1 verification supersedes T2 aggregation.
v1.4.4 · 2026-05-13 · Round 20 vals.ai harvest integrated. Four T1 SWE-bench Verified cells added: GPT-5.5 (0.826), GPT-5.4 (0.782), Gemini 3.1 Pro Preview (0.788), Claude Opus 4.6 Thinking (0.782). The Coding specialty now has 9 eligible models (up from 6) - ChatGPT 5.5, ChatGPT 5.4, and Gemini 3.1 Pro Preview moved from "insufficient evidence" into the main Coding ranking.
v1.4.3 · 2026-05-13 · Reasoning tab now shows a data-sparsity disclaimer. With only 2 benchmarks (ARC-AGI-2 and AIME 2025) and most models having only one measured, single-cell rankings here can be misleading. We kept minBenchmarks at 1 (rather than raising to 2, which would leave only ~2 models ranked) and added a visible amber disclaimer so users understand the limitation.
v1.4.2 · 2026-05-12 · Specialty eligibility raised to minimum 2 cells for Coding, Knowledge, and Agentic (Reasoning stays at 1 because it has only 2 benchmarks total and is already protected by the 85% shrinkage threshold). Single-cell models on hard benchmarks were being unfairly dragged to the bottom; they now appear in a separate "insufficient evidence" section.
v1.4.1 · 2026-05-12 · Reasoning specialty (and any future 2-benchmark specialty) now uses an 85% shrinkage coverage threshold instead of the flat 60%. Prevents a single high-weight cell from producing an inflated specialty score.
v1.4.0 · 2026-05-12 · Cross-component coverage floor introduced. Three-section visibility layout (RANKED / PROVISIONAL / AWAITING). 40% core-component coverage requirement (Reasoning & Agency) added.
CORRECTIONS LOG

Independent verification, public corrections

Every cell on the leaderboard cites its source. When we find a mis-attribution, an inflated self-report contradicted by independent measurement, or a stale evaluation that predates the model itself, we correct it - and log the change here.

Our error · 2026-07-26
We published a DeepSeek score we could not source
22.8% DeepSeek V4 Pro · ARC-AGI-2 · voided
While auditing every cell for methodology v2, we found an ARC-AGI-2 score attributed to DeepSeek V4 Pro with no source at all: no URL, no evaluator, no record of where it came from. It is also identical, to three decimal places, to GLM 5.2’s ARC-AGI-2 score, which is properly sourced to the ARC Prize leaderboard.
Two different models holding the same number to three decimals, with one of them sourceless, is the same duplication pattern we found in an earlier audit of a secondary aggregator. We cannot establish that DeepSeek V4 Pro was ever evaluated on ARC-AGI-2, so the cell is voided. The record is retained rather than deleted, as with our other corrections.

This did not change any ranking. The cell had already been excluded from scoring before we found it, so no published score moved, on the current board or the previous one. We are logging it anyway. A correction that costs us nothing is the easiest kind to stay quiet about, and staying quiet is how a value with no source survives long enough to matter. We owe DeepSeek an apology. We published a number against their model that we could not substantiate. We should have been more careful before publishing it. We will learn from this and will do better in the future. Four further cells were voided in the same pass for weaker provenance defects: a benchmark whose contest year was never pinned, two sourced to a vendor model page rather than the evaluation board that produced them, and one sourced to a commentary article about a different model entirely. None of the five score on the current board.
Read the full apology and what we changed
Retraction · 2026-07-25
Two of our own corrections were wrong
We audited this log against our commit history. Twice, we compared two different configurations of the same benchmark and reported the difference as a laboratory overstating its results. Neither laboratory overstated anything.
In both cases the number we treated as a contradiction was measured under different conditions to the number we compared it against: thinking mode against non-thinking mode in one case, tool-assisted against no-tools in the other. Both labs had published the comparable figure themselves, clearly labelled. We did not read it carefully enough. Details below. No score on this site has changed as a result, because in both cases the underlying cell had already been replaced with a correct independent measurement. What was wrong was what we said about it.
Retraction · our error
Moonshot · Humanity’s Last Exam
36.4% Moonshot, no tools / 35.9% independent, no tools
Our source policy excluded Moonshot on the stated grounds that their Kimi K2.6 figure had failed verification. It had not. Their headline 54.0% was explicitly labelled as a with-tools result, and the same post published 36.4% for the no-tools configuration, within half a point of independent measurement. We compared their with-tools number against our no-tools number and read the difference as overstatement. The error was entirely ours. We have removed the claim. Moonshot’s evidence is being reassessed cell by cell, and we are replacing laboratory-level trust ratings with per-measurement assessment, because a rating attached to a company is what allowed a configuration mismatch to become an accusation in the first place.
Retraction · our error
DeepSeek V4 Pro · GPQA Diamond
90.1% DeepSeek, Max thinking / 88.8% independent, Max thinking
We previously reported this cell moving from 90.1% to 72.9% and described 72.9% as an independent third-party measurement. It was neither. DeepSeek publishes 72.9% themselves, as their result for non-thinking mode, and the figure reached us through a secondary aggregator that republishes other parties’ numbers. We compared it against their thinking-mode result and read the gap as a 17-point overstatement. Measured like for like, DeepSeek’s self-report sits 1.3 points above independent measurement. The original entry is retained below rather than deleted.
Independent result added
Kimi K3 · SWE-bench Verified
93.4% ±1.11 pp · Mini-SWE-agent
Vals AI’s exact standalone Kimi K3 row supplies one frozen-v1 Coding cell. K3 remains Coding-ineligible at 1/3; SWE-bench Pro and Terminal-Bench 2.0 remain missing, and Terminal-Bench 2.1 is not substituted.
Configuration selected
Kimi K3 · Humanity’s Last Exam
56.0% with tools / 44.3% no tools
Two different measurements, not a correction. Moonshot reports the with-tools configuration and labels it clearly. Our Humanity’s Last Exam slot is text-only and no-tools, so we score the independent no-tools evaluation. The values are not averaged, and neither of them is wrong.
Withdrawn · retained as record
DeepSeek V4 Pro · GPQA Diamond
0.901 0.729 −17 pp
This is what we published, and it was wrong. We described 0.729 as an independent third-party measurement contradicting DeepSeek’s self-report. It was neither independent nor a contradiction. The entry stays here because a corrections log that quietly deletes its own mistakes is worth nothing. See the retraction above.
Correction · our error
Kimi K2.6 “enters scoring”
We described four cells as moving from excluded to independently verified. On the day we published that, all four were the same numbers with rewritten source addresses.
GPQA Diamond, SWE-bench Verified, SWE-bench Pro and MMMU-Pro were relabelled to a secondary aggregator without a single value changing, and our source rules read the new hostname as a higher trust tier. Three of the four were genuinely re-measured by independent evaluators some weeks later. SWE-bench Pro never was. A change of source address is not a change of evidence, and we should not have presented it as one.
Source-tier upgrades
Cells where a previously lower-trust source was replaced by a higher-trust independent measurement. The headline DeepSeek correction and the Kimi K2.6 promotion (both above) are the most consequential examples. Two specific value-changing upgrades worth surfacing in detail:
Model Benchmark Change Note
DeepSeek V4 Pro GPQA Diamond 0.901 (T3) → 0.729 (T2) withdrawn: relayed self-report, not an independent measurement (see retraction)
Claude Opus 4.7 (thinking) SWE-bench Verified 0.876 (T2) → 0.820 (T1) private-test-set evaluation (vals.ai)
This table previously stated that further tier upgrades were “tracked in models.json”. That claim has been withdrawn: models.json records only the current state of each cell, so it cannot evidence a historical change. We are rebuilding this section from our commit history so that every published correction can be traced to the exact release that made it.
Variant-attribution corrections
Cells where the source row label didn't unambiguously identify the variant (Pro / thinking / base). Either re-attributed to the correct column on the source page, or nulled per the variant-distinct evidence policy.
Model Benchmark Change Reason
ChatGPT 5.5 (Pro) BrowseComp 0.844 → 0.901 corrected to actual Pro column on release page
ChatGPT 5.5 (Pro) FrontierMath 0.517 → 0.524 corrected to actual Pro column on release page
ChatGPT 5.4 (Pro) SWE-bench Pro 0.577 → null value was from the GPT-5.4 base column
Claude Opus 4.6 (thinking) SWE-bench Pro 0.534 → null source row didn't distinguish thinking from base
Claude Opus 4.6 (thinking) OSWorld 0.727 → null source row didn't distinguish thinking from base
Claude Opus 4.7 (base) GPQA Diamond 0.942 → null source row didn't distinguish base from thinking
ChatGPT 5.5, 5.5 (Pro), 5.4, 5.4 (Pro) AA Intelligence Index 57-60 → null (4 cells) AA family row not variant-distinct
Claude Opus 4.7 (base) AA Intelligence Index 57 → null AA "(max)" attribution ambiguous
Stale-source removals
Cells where the cited source no longer hosts the value (typically because the source's leaderboard has rotated to newer models). Per "verifiable evidence only," these cells are nulled rather than left citing an unfetchable source. Note that ChatGPT 5.4's SWE-bench Verified cell was subsequently restored from an independent T1 source (see additions table below).
Model Benchmark Change Reason
ChatGPT 5.4 SWE-bench Verified 0.728 → null no longer on swebench.com leaderboard
ChatGPT 5.4 (Pro) SWE-bench Verified 0.728 → null no longer on swebench.com leaderboard
Cells added from independent T1 verification
Cells that were previously null in our schema and have now been populated from a T1 independent third-party evaluation (1.00× trust weight). These additions reflect new measurements arriving in the public ecosystem and do not displace any prior values.
Model Benchmark Change Source
ChatGPT 5.5 SWE-bench Verified null → 0.826 vals.ai T1
ChatGPT 5.4 SWE-bench Verified null → 0.782 vals.ai T1 (restores the previously nulled stale-source cell)
Gemini 3.1 Pro (preview) SWE-bench Verified null → 0.788 vals.ai T1
Claude Opus 4.6 (thinking) SWE-bench Verified null → 0.782 vals.ai T1 (first variant-distinct SWE-V cell for this AWAITING model)
156
Scored cells on the board
113
From independent evaluators
0
Values imputed
The remaining 43 come from the benchmark owners’ own leaderboards (ARC Prize and LiveBench). Independent evaluators here are vals.ai (73 cells) and Artificial Analysis (40). Artificial Analysis supplied 45% of the board before this release and now supplies 26%: Terminal-Bench 2.1, GPQA Diamond and MMMU-Pro all moved to vals.ai. Knowledge is now split 52/48 between two evaluators where it was 97% one of them. Visual Reasoning is still single-source, having moved from wholly Artificial Analysis to wholly vals.ai. It rests on a single benchmark, so it cannot be diversified until a second visual benchmark exists.
Spotted a value that disagrees with the cited source? Or know of a published independent measurement we should be tracking? .
CAPABILITY HEATMAP

Strengths and weaknesses, at a glance

Each cell shows a model's score on one of the five cognitive components. The AGI Score (rightmost column) blends these by parent weights (shown beneath each header). 100 marks the AGI threshold; values above are super-human on that dimension.

A component is not the same thing as a specialty tab, even where the names match. The Reasoning component here blends ARC-AGI-2, GPQA Diamond, HLE, AIME and LiveBench; the Coding specialty tab above uses only its three coding benchmarks. The two answer different questions and will not show the same number for the same model.

Model Agency
35%
Reasoning
29%
Knowledge
15%
Visual Reasoning
11%
Language
10%
AGI Score
Loading...
Scale:
0 to 30
30 to 60
60 to 85
85+
no data
· Provisional rows shown with italic name and amber chip

By Capability Area

Top performers in each of the three public capability framings. The overall AGI Score blends these by their parent weights (Thinking 44 / Doing 43 / Communicating 13).

Thinking
44% weight
Fluid Reasoning + World Knowledge
  1. Loading…
Doing
43% weight
Agency (tools, planning, coding) + Visual Reasoning
  1. Loading…
Communicating
13% weight
Language Production + Visual Reasoning output
  1. Loading…