An estimation benchmark for language models

A set of estimation problems, made of a mix of novel and contamination-resistant questions, intended to be challenging for both humans and models alike.

Why this benchmark

Estimation problems are interesting because they test capabilities that matter in the real world: recursive nature of subproblems, creativity in solutions, need for reasonable abstractions, broad knowledge, telling apart known from unknown, commonsense, etc.

Estimating a quantity you've never seen is recursive problem solving under uncertainty. A hard number splits into smaller ones, and those split again, until you reach quantities you can actually anchor. There's no fixed procedure for it, so reaching an answer means finding an angle: which quantity to lean on, which route to take across.

For a model to be proficient at this task, it will need to be fluent in a variety of operations — recall, decomposition, triangulation — and to know when to use which.

The benchmark is designed to test these capabilities, and to provide a clear signal of which models are better at them.

Model scores over time

mean Cramér vs. release date
0.00.51.01.52.0May 2023Dec 2024Jul 2026release date →↑ better · mean Cramér-logbest so farGpt 4o3 MiniLlama 4 Mavericko3qwen3 235b a22b Thinking 2507Gpt 5 MiniDeepseek v3.2Claude Sonnet 4.6Gemini 3.1 Pro PreviewDeepseek v4 ProGpt 5.5Gpt 5.5 ProGrok 4.3Claude Opus 4.8Claude Fable 5Claude Fable LatestGlm 5.2Claude Sonnet 5
  • OpenAI
  • Anthropic
  • Google DeepMind
  • DeepSeek
  • Meta
  • Other

Each dot is a model on the open edition — its mean Cramér-log (lower is better) against its release date (from the OpenRouter catalog, so ≈ when it was published); the dashed line traces the best score reached so far. See the full standings →.

Explore the bench

leaderboard, questions, methods & review