Spec sheet · run v2026-05-04 · OenoBench release_v1.2
How much does a language model actually know about wine?
A bilingual reference on artificial intelligence across viticulture, winemaking, and the wine business — plus OenoBench, the wine knowledge benchmark for large language models.
Instrument record
- Run
- v2026-05-04
- Released
- 2026-05-04
- Paper
- OenoBench release_v1.2
- Questions
- 3,266
- Evaluations
- 52,256
- Configurations
- 16
- Evaluation spend
- $98.33
Domain tags sum to 3,329 and difficulty tags to 3,329 against a corpus of 3,266 — unreconciled, reported as published.
Measured answer
o3 · OpenAI · best of 16 configurations · 3,266 questions
Chromatogram. 6 peaks, one per wine domain, on a shared baseline with 5 dashed drop lines marking the integration boundaries between them. Each peak's area is proportional to the number of questions carrying that domain tag, and each is labelled with its domain, its question count and its share. The largest is Wine regions at 1,108 questions (33.3%); the smallest is Winemaking at 188 (5.6%). Shares are of the 3,329 domain tags the run publishes, not of its 3,266 questions: 63 questions carry more than one tag. The vertical axis is unlabelled detector signal.
Contents
§1 About the project
Scope, method, and how the atlas is written.
§2 Keynote
The long read: what AI in wine actually does, by the numbers.
§3 Benchmarks
- 3.1 Leaderboard v2026-05-04
§5 Research
What is being measured next, and an open door to collaborate.
§6 About the author
Who compiles it, and on what evidence.