FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
Abstract
FlavourBench evaluates language models on culinary portfolio tasks using executable ground truth, statistical rigor, and fully reproducible verification.
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.
Community
BREAKING: world's first culinary intelligence benchmark for LLMs now on arXiv.
We tested the strongest model from 16 labs across 534 executable culinary tasks powered by Epicure.
Grok 4.6 leads the pack.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets (2026)
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks (2026)
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction (2026)
- Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark (2026)
- Loreley: Repository-Scale Program Evolution with Quality-Diversity Search (2026)
- A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan Decomposition (2026)
- LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.20574 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper