Project Benchmarks

click on snapshot for details

    model data provided by Surge AI Leipzig-100 Updated Leipzig-100
  1. Sep 1 September 1, 2025
  2. Nov 1 November 1, 2025
  3. Nov 26 November 26, 2025
  4. Mar 23 March 23, 2026
  5. Apr 11 April 11, 2026
  6. Apr 30 April 30, 2026
  7. May 26 May 26, 2026
  8. Sep 17 September 17, 2026

The Leipzig Benchmark

100
research-level problems
49
contributing researchers
25
subfields included

Model scores

Model Name Effort Web Search Code Execution Failures Correct Answer
GPT-5.6 Sol Pro max
128k tokens
3 91%
GPT-6 Astra max
128k tokens
91%
GPT-5.6 Sol max
128k tokens
90%
GPT-6 Astra Pro max
128k tokens
4 88%
Claude Fable 5.1 high
128k tokens
79%
Claude Opus 5 high
128k tokens
2 75%
GPT-5.6 Sol xhigh
128k tokens
16 68%
GPT-5.5 xhigh
128k tokens
2 44%
Gemini 3.8 Flash* default
65,536 tokens
41 41%
GPT-5.4 xhigh
128k tokens
14 28%
Gemini 3.1 Pro default
64k tokens
15%
Claude Opus 4.6 default
15k tokens, 80k thinking
1 14%
Gemini 3 Pro default
50k tokens
14%
Claude Opus 4.7 default
128k tokens
18 13%
DeepSeek V4 Pro default
384k tokens
13 10%
DeepSeek-V3.2 default
66k tokens
27 8%
Grok-4.3 high
no output cap
6%
Grok-4.20 default
256k tokens
5%
Based on all 100 problems, one attempt per model. Bold entries means that this is the maximal possible option. Claude ran at high effort on purpose, with an advisory 100k budget under its 128k cap, because max produced the same answers on twice the tokens.

* Gemini 3.8 Flash was run without the code interpreter, and was allowed a second attempt. With code execution enabled it very often errors with TOO_MANY_TOOL_CALLS. This is a known issue for the current Gemini models. Raising the thinking level to high makes it worse: the same 100 problems then score 33 instead of 41.

Coverage overlap

ModelGPT-5.6 Sol-MaxGPT-6 AstraClaude Opus 5Claude Fable 5.1Only this model
GPT-5.6 Sol-Max908672770
GPT-6 Astra869172742
Claude Opus 5727275700
Claude Fable 5.1777470791
The diagonal is each model's own score; off-diagonal cells count problems solved by both. 66 of the 100 problems are solved by all four, 96 by at least one, and 4 by none.

Contributing Subfields

Algebraic Geometry 36
Algebraic Combinatorics 22
Matroid Theory 18
Enumerative Combinatorics 15
Representation Theory 15
Combinatorics 14
Discrete Geometry 10
Algebra 7
Commutative Algebra 5
Graph Theory 5
Algebraic Statistics 4
Homological Algebra 4
Polytope Theory 4
Tropical Geometry 4
Analysis 3
Arithmetic Geometry 3
Knot theory 3
Number Theory 3
Topology 3
Complex Analysis 2
Euclidean Geometry 2
Metric Geometry 1
Probability Theory 1
Real Algebraic Geometry 1
Theoretical Computer Science 1
Contributed by 49 researchers at the Benchmarks in Leipzig event.