← Back to New
MODEL RELEASE

Kimi K3 debuts with seven top-three finishes

Across 15 first observations, Kimi K3 places in the top three seven times and leads the newly added AutomationBench-AA board.

Event date Published

A strong but uneven first profile

Kimi K3 enters EveryBench with 15 first observations. Seven are top-three placements, including one first place. The other eight range from fourth to a tie for fifteenth among models with a reported score on each benchmark.

First observations
15
Kimi K3 was absent from the earlier EveryBench data.
Top-three placements
7
Across the 15 benchmark results shown by EveryBench.
First places
1
AutomationBench-AA, a newly added rolling board.

The strongest results cluster around agentic and composite work

Kimi K3 leads AutomationBench-AA at 52.7%, places second on AA-Briefcase at 1547 Elo and Vals Index at 74.7, and places third on the AA Intelligence Index at 57.1 and GDPval at 58.4%. AutomationBench-AA was newly added, so Kimi did not displace an earlier leader.

AutomationBench-AA
52.7% #1 of 34
Ahead of Grok 4.5 at 51.4% and GPT-5.6 Sol at 51.2%.
AA-Briefcase
1547 Elo #2 of 41
Behind Claude Fable 5 at 1583 Elo.
Vals Index
74.7 #2 of 37
Behind Claude Fable 5 at 75.1.
AA Intelligence Index
57.1 #3 of 318
Behind Claude Fable 5 and GPT-5.6 Sol.
GDPval
58.4% #3 of 117
This version of GDPval is fixed rather than continually updated.

Where Kimi K3 sits against nearby leaders

Five benchmark panels compare Kimi K3 with nearby leading models on AutomationBench-AA, AA-Briefcase, Vals Index, the Artificial Analysis Intelligence Index, and GDPval.
PanelItemBenchmark score
AutomationBench-AA (%)Kimi K352.7
AutomationBench-AA (%)Grok 4.551.4
AutomationBench-AA (%)GPT-5.6 Sol51.2
AA-Briefcase (Elo)Claude Fable 51,583
AA-Briefcase (Elo)Kimi K31,547
AA-Briefcase (Elo)GPT-5.6 Sol1,495
Vals IndexClaude Fable 575.1
Vals IndexKimi K374.7
Vals IndexGPT-5.6 Sol73.1
AA Intelligence IndexClaude Fable 559.9
AA Intelligence IndexGPT-5.6 Sol58.9
AA Intelligence IndexKimi K357.1
AA Intelligence IndexClaude Opus 4.855.7
GDPval (%)Claude Fable 563
GDPval (%)GPT-5.6 Sol62.4
GDPval (%)Kimi K358.4
GDPval (%)Claude Sonnet 555.3

Kimi K3 is highlighted in gold. Each panel uses a labeled local scale to make nearby gaps readable; compare positions only within the same benchmark.

Eight new boards split their winners

Six distinct models lead the eight newly added benchmarks. Claude Fable 5 and GPT-5.4 lead two each; GLM-5.2, Claude Opus 4.8, GPT-5.5, and Kimi K3 lead one each. No model wins more than two.

New benchmarks
8
Three coding, three agentic, and two multimodal benchmarks.
Distinct leaders
6
No model leads more than two of the eight.