Atlas/Benchmarks/Leaderboard
OenoBench Leaderboard
OenoBench evaluates 16 model configurations across 3,266 multiple-choice questions covering viticulture, winemaking, wine business, the world’s wine regions, grape varieties, and producers. Use the filters to slice the leaderboard by domain, difficulty tier, or whether the question is answerable from parametric knowledge alone (closed-book) or requires contextual reasoning. The other tabs surface three further analyses computed from the same run: reasoning-mode lift, self-preference bias, and cost efficiency.
Step plot. The corpus is ordered by the closed-book split: the first 1,601 questions are answerable from parametric knowledge alone, the remaining 1,665 require contextual reasoning. The pen holds 0.974 in the first regime and 0.703 in the second — a drop of 0.271 for o3 — inside a ±1 binomial standard-error band of 0.0159 and 0.0457 at a reference window of 100. The dashed rules are the unweighted mean over all configurations (16): 0.896 and 0.570, a drop of 0.326. Accuracy runs from 0.30 to 1.00. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.
Dumbbell chart. The configurations (16) are ranked first to last by overall accuracy, one row each. A rule joins the row's accuracy on questions that require contextual reasoning (n 1,665) to its accuracy on closed-book questions (n 1,601), on a scale from 0.30 to 1.00, with the drop between the two printed at the right of the row. The smallest drop is GPT-5 at 0.266; the largest is DeepSeek-V3 at 0.396. The overall leader is o3.
- 1o3OpenAIeffort83.6%
- 2GPT-5OpenAI82.8%
- 3Gemini 2.5 Pro (thinking)Googlethinking82.6%
- 4Gemini 2.5 ProGoogle81.7%
- 5Claude Opus 4.7Anthropic81.0%
- 6Claude Opus 4.7 (thinking)Anthropicthinking81.0%
- 7GPT-5 miniOpenAI78.4%
- 8DeepSeek-R1DeepSeekthinking77.1%
- 9Gemini 2.5 FlashGoogle75.1%
- 10DeepSeek-V3DeepSeek70.3%
- 11Mistral Large 2411Mistral AI69.1%
- 12Qwen 2.5 72BAlibaba67.4%
- 13Llama 3.3 70BMeta67.1%
- 14Llama 3.1 8BMeta60.5%
- 15Qwen 2.5 7BAlibaba57.0%
- 16Claude Haiku 4.5Anthropic53.3%
Every figure above is read from a single versioned run file, published under CC BY 4.0: download the full results (JSON, 23 KB). Run files accumulate rather than being replaced, so this URL keeps resolving after later runs are published — cite the dated file, not the leaderboard page.
There is no companion paper. This leaderboard, the run file it is built from, and the OenoBench page — corpus construction, multi-model generation, AI validation, and the bias-mitigation framework — are the published record of the benchmark.