Claude Fable 5 leads on first SciCode (AA) results in EveryBench
Artificial Analysis's first-hand SciCode evaluation now covers 310 models; Claude Fable 5 is 1.3 points ahead of second place.
Event date Published
SciCode (Artificial Analysis) joins EveryBench
Artificial Analysis's first-hand SciCode evaluation is now available on EveryBench, with 310 models receiving initial pass@1 scores on 288 subproblems across 80 scientific-computing problems in 16 disciplines. This is the benchmark's first appearance in the EveryBench dataset; it is reported as a separate entity from the citation-only SciCode column.
- Leader's margin
- 1.3 points Claude Fable 5 ahead of Gemini 3.1 Pro; the #2-to-#3 gap is just 0.2 points.
- Models with first results
- 310 Initial pass@1 scores from Artificial Analysis's first-hand run.
- Three-way tie
- 56.1% GPT-5.6 Sol, GPT-5.5, and Gemini 3 Pro share rank 6.
Claude Fable 5 leads the first results
Claude Fable 5 tops the debut SciCode (AA) leaderboard at 60.2%, 1.3 points ahead of Gemini 3.1 Pro at 58.9%. Kimi K3 (58.7%) and Muse Spark 1.1 (58.2%) are close behind, and GPT-5.4 rounds out the top five at 56.6%. These are first results on a single first-hand run, not before/after improvements, and the top cluster is tightly packed.
- Claude Fable 5
- 60.2% #1 of 310 Leads the first SciCode (AA) results.
- Gemini 3.1 Pro
- 58.9% #2 of 310 1.3 points behind the leader.
- Kimi K3
- 58.7% #3 of 310 Close behind second place.
- Muse Spark 1.1
- 58.2% #4 of 310 Fourth in the debut leaderboard.
- GPT-5.4
- 56.6% #5 of 310 Top five in the first results.
Top 10 models on SciCode (AA)
| Rank | Item | Value |
|---|---|---|
| 1 | Claude Fable 5 | 60.19 |
| 2 | Gemini 3.1 Pro | 58.91 |
| 3 | Kimi K3 | 58.68 |
| 4 | Muse Spark 1.1 | 58.22 |
| 5 | GPT-5.4 | 56.6 |
| 6 | GPT-5.6 Sol | 56.13 |
| 7 | GPT-5.5 | 56.13 |
| 8 | Gemini 3 Pro | 56.13 |
| 9 | GPT-5.2-Codex | 54.63 |
| 10 | Claude Opus 4.7 | 54.51 |
Top 10 of 310 models on the first SciCode (AA) results. GPT-5.6 Sol, GPT-5.5, and Gemini 3 Pro tie at 56.1%. These are first results from Artificial Analysis's first-hand run, not before/after improvements.
What to read from this debut
The top of the SciCode (AA) leaderboard is tightly packed: five models sit within four points of each other, and three more tie at 56.1%. That clustering, plus the fact that every leader solves only about six in ten subproblems, is the main takeaway. It is too early to treat this as a settled coding ranking; the benchmark has just been added to EveryBench and these are its first 310 results.