Human experts and AI models compete to most accurately label the same data. Only the best rise to the top, and the responses of both are combined using weighted aggregation to produce the highest quality datasets.
Violet line marks the Centaur Aggregation score on this contest.
What's a Centaur Aggregation score? It's the accuracy-weighted consensus of every expert and model on a case, calibrated against gold standards. See it live →
01
A dataset opens as a contest
Some contests are open to all. The complex, regulated ones are gated behind verified credentials and background checks.
02
Experts and AI models compete
Everyone labels the identical cases and is graded on gold-standard answers. Where models enter, the leaderboard doubles as a benchmark.
03
Consensus becomes the label
Proven-accurate voices count more. The weighted consensus outperforms any single expert or model, and ships with the dataset.
Contest catalog
Live and completed contests
COMPLETEDOpen to all
Calorie Estimation
Meal photographs estimated for total calories by nearly a thousand competing labelers, with per-case estimates aggregated into a calibrated consensus value.
Every labeler, human or model, answers this prompt on every record.
Leaderboard
Centaur Aggregation83.6
Top experts
#1 · Synth75.7
#2 · funky73.9
#3 · Munckin73.3
AI models
GPT-5.274.0
GPT-5.573.7
GPT-5.6 Sol72.8
The gap
GPT-5.2 trails the Centaur Aggregation by 9.6 points on this task. The dataset is the shortest path to closing it.
Who was allowed in
Open enrollment, no gate
Gold-standard calibration cases
Accuracy-weighted from the first case
The benchmark, contest by contest
The best data comes from where AI still loses
Every competition between frontier models and experts creates a real-time benchmark. The gap between the leading model and Centaur Aggregation pinpoints the hardest problems — generating the training signal needed to build better models.
Best AI model in contestCentaur AggregationAccuracy vs gold standard. Illustrative data for concept review.
Ways to engage
Four ways into the Arena
Pick the one closest to what you need. We'll put you on the calendar with the right person on our team.
1 · What brings you to the Arena
License
Buy a dataset
Any completed contest is a licensed dataset: records, accuracy-weighted aggregation labels, per-labeler provenance, and the model benchmark.
Bring your model
Benchmark your model
Your model labels the same cases as the expert field and lands on the leaderboard. See exactly where it trails the Centaur Aggregation.
Bring data
Run a private arena
Your cases become a private contest. We recruit and gate the right specialists, enter the models you care about, and hand back aggregation labels.
From zero
Source data you don't have
Tell us the capability you're building toward. We source the raw data through partners like Protege, recruit the expert field, and run the contest end to end.
2 · Anything else (optional)
A good starting point: data modality (image, text, audio, video), domain or specialty, estimated volume, and timeline.
You'll enter your name and email on the next screen.