Skip to content
VinumExMachina Atlas of AI in Wine
Select language: Русский

AtlasBenchmarksLeaderboard

OenoBench Leaderboard

Section § 3.2
Status published
Updated
Languages EN · RU
Measure 72 ch · 1 min

OenoBench evaluates 16 model configurations across 3,266 multiple-choice questions covering viticulture, winemaking, wine business, the world’s wine regions, grape varieties, and producers. Use the filters to slice the leaderboard by domain, difficulty tier, or whether the question is answerable from parametric knowledge alone (closed-book) or requires contextual reasoning. The other tabs surface three further analyses computed from the same run: reasoning-mode lift, self-preference bias, and cost efficiency.

Fig.02
Fig. 02Where the instrument dropsreading 1 of 3 · o3 · two regimes · n 1,601 / 1,665

Step plot. The corpus is ordered by the closed-book split: the first 1,601 questions are answerable from parametric knowledge alone, the remaining 1,665 require contextual reasoning. The pen holds 0.974 in the first regime and 0.703 in the second — a drop of 0.271 for o3 — inside a ±1 binomial standard-error band of 0.0159 and 0.0457 at a reference window of 100. The dashed rules are the unweighted mean over all configurations (16): 0.896 and 0.570, a drop of 0.326. Accuracy runs from 0.30 to 1.00. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.

Δ 0.2710.300.400.500.600.700.800.901.0001,0001,6012,0003,0003,266CLOSED-BOOK (PARAMETRIC) · n 1,601CONTEXTUAL REASONING · n 1,6650.9740.7030.8960.570ordinate: accuracy · abscissa: corpus ordered by closed-book split
±1 SE, binomial, at a reference window of 100mean of all configurations (16)solid pen: o3 · OpenAI · effort
Finding 1. Every configuration collapses on questions that require contextual reasoning rather than recall. o3 loses 0.271; the mean across all configurations (16) loses 0.326. The noise band widens with the drop — the contextual regime is measured less precisely at the same reference window. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.
Fig.03
Fig. 03The gap does not close with rankall 16 configurations · sorted by overall accuracy

Dumbbell chart. The configurations (16) are ranked first to last by overall accuracy, one row each. A rule joins the row's accuracy on questions that require contextual reasoning (n 1,665) to its accuracy on closed-book questions (n 1,601), on a scale from 0.30 to 1.00, with the drop between the two printed at the right of the row. The smallest drop is GPT-5 at 0.266; the largest is DeepSeek-V3 at 0.396. The overall leader is o3.

0.300.400.500.600.700.800.901.00ACCURACYΔ COLLAPSECONFIGURATIONo30.271GPT-50.266Gemini 2.5 Pro (thinking)0.290Gemini 2.5 Pro0.300Claude Opus 4.70.319Claude Opus 4.7 (thinking)0.323GPT-5 mini0.319DeepSeek-R10.337Gemini 2.5 Flash0.343DeepSeek-V30.396Mistral Large 24110.358Qwen 2.5 72B0.381Llama 3.3 70B0.364Llama 3.1 8B0.304Qwen 2.5 7B0.311Claude Haiku 4.50.331
Closed-book (parametric) · n 1,601Contextual reasoning · n 1,665
Ranked first to last, the bar never shortens by more than a few points. The smallest gap belongs to GPT-5 (0.266); the largest to DeepSeek-V3 (0.396).
OenoBench · v2026-05-0416 model configurations · 3,266 questions across six wine domains
Wine domain
Question difficulty
Closed-book vs contextual
ModelAccuracy
  1. 1
    o3OpenAIeffort
    83.6%
  2. 2
    GPT-5OpenAI
    82.8%
  3. 3
    Gemini 2.5 Pro (thinking)Googlethinking
    82.6%
  4. 4
    Gemini 2.5 ProGoogle
    81.7%
  5. 5
    Claude Opus 4.7Anthropic
    81.0%
  6. 6
    Claude Opus 4.7 (thinking)Anthropicthinking
    81.0%
  7. 7
    GPT-5 miniOpenAI
    78.4%
  8. 8
    DeepSeek-R1DeepSeekthinking
    77.1%
  9. 9
    Gemini 2.5 FlashGoogle
    75.1%
  10. 10
    DeepSeek-V3DeepSeek
    70.3%
  11. 11
    Mistral Large 2411Mistral AI
    69.1%
  12. 12
    Qwen 2.5 72BAlibaba
    67.4%
  13. 13
    Llama 3.3 70BMeta
    67.1%
  14. 14
    Llama 3.1 8BMeta
    60.5%
  15. 15
    Qwen 2.5 7BAlibaba
    57.0%
  16. 16
    Claude Haiku 4.5Anthropic
    53.3%

Every figure above is read from a single versioned run file, published under CC BY 4.0: download the full results (JSON, 23 KB). Run files accumulate rather than being replaced, so this URL keeps resolving after later runs are published — cite the dated file, not the leaderboard page.

There is no companion paper. This leaderboard, the run file it is built from, and the OenoBench page — corpus construction, multi-model generation, AI validation, and the bias-mitigation framework — are the published record of the benchmark.