← Back to New
SCORE RESULT

First independent results land for Inkling and Grok 4.5

LiveBench profiles Thinking Machines' debut model mid-pack at its highest effort setting, and Grok 4.5 enters DeepSWE effectively even with Claude Sonnet 5.

Event date Published

Inkling's first LiveBench profile: mid-pack overall, math-leaning

Inkling, Thinking Machines' open-weights reasoning model released July 15, posted its first LiveBench results, evaluated at its highest reasoning-effort configuration. It debuts twenty-first of 32 models overall at 71.7 on a board led by GPT-5.6 Sol at 82.4. Its highest score is Mathematics at 88.4, which still ranks eighteenth of 32 there; its best category rank is eleventh, on Instruction Following. Lower-effort configurations could score differently.

LiveBench overall
71.7 #21 of 32
Board leader GPT-5.6 Sol at 82.4.
LiveBench Mathematics
88.4 #18 of 32
Inkling's highest score, not a board standout.
LiveBench Instruction Following
70.1 #11 of 32
Its best category rank.
LiveBench Reasoning
78.4 #24 of 32
Logic and inference problems requiring multi-step deduction.
LiveBench Coding
71.0 #25 of 32
Code generation and completion tasks with verifiable solutions.
LiveBench Agentic Coding
47.7 #17 of 32
Coding tasks run through an agentic harness.
LiveBench Data Analysis
72.8 #18 of 32
Table-join and format-conversion tasks, answered without writing code.
LiveBench Language
73.5 #26 of 32
English reading comprehension of recent text.

Grok 4.5 enters DeepSWE even with Claude Sonnet 5

xAI's Grok 4.5 received its first DeepSWE result: 53.76% of tasks resolved, ninth of the 16 models on the board and effectively even with Claude Sonnet 5 at 53.85% — a gap under 0.1 points. The board leader, GPT-5.6 Sol, resolved 72.7% of tasks. This is one agent-run coding benchmark of 16 models, not a broad coding-quality ranking.

Grok 4.5
53.76% #9 of 16
Effectively even with Claude Sonnet 5.
Claude Sonnet 5
53.85% #8 of 16
The gap to Grok 4.5 is under 0.1 points.
GPT-5.6 Sol
72.7% #1 of 16
Board leader.