Competition is the best quality control

Human experts and AI models compete to most accurately label the same data. Only the best rise to the top, and the responses of both are combined using weighted aggregation to produce the highest quality datasets.

100,000+
Experts competing
1,000+
Contests run
Every label
Accuracy-weighted
CENTAUR AGGREGATION
0.0
Best score on Skin Cancer Classification
Top human annotator · Disposition92.0
Top AI · GPT-567.9
Violet line marks the Centaur Aggregation score on this contest.

What's a Centaur Aggregation score? It's the accuracy-weighted consensus of every expert and model on a case, calibrated against gold standards. See it live →

01

A dataset opens as a contest

Some contests are open to all. The complex, regulated ones are gated behind verified credentials and background checks.

02

Experts and AI models compete

Everyone labels the identical cases and is graded on gold-standard answers. Where models enter, the leaderboard doubles as a benchmark.

03

Consensus becomes the label

Proven-accurate voices count more. The weighted consensus outperforms any single expert or model, and ships with the dataset.

Contest catalog

Live and completed contests

COMPLETEDOpen to all

Calorie Estimation

Meal photographs estimated for total calories by nearly a thousand competing labelers, with per-case estimates aggregated into a calibrated consensus value.

Request Dataset AccessSample records available under NDA
Sample case preview
Replay the contest
3,490
Records
402k
Annotations
961
Labelers
5
AI models
Schema
image_refjpegMeal photograph
calorie_estimateintIndividual kcal estimate
consensus_kcalfloatAggregated kcal estimate
opinion_countintReads collected per case
consensus_scorefloatWeighted agreement, 0 to 1
Task prompt

How many calories do you think are in this meal?

Every labeler, human or model, answers this prompt on every record.

Leaderboard
Centaur Aggregation83.6
Top experts
#1 · Synth75.7
#2 · funky73.9
#3 · Munckin73.3
AI models
GPT-5.274.0
GPT-5.573.7
GPT-5.6 Sol72.8
The gap

GPT-5.2 trails the Centaur Aggregation by 9.6 points on this task. The dataset is the shortest path to closing it.

Who was allowed in
Open enrollment, no gate
Gold-standard calibration cases
Accuracy-weighted from the first case
The benchmark, contest by contest

The best data comes from where AI still loses

Every competition between frontier models and experts creates a real-time benchmark. The gap between the leading model and Centaur Aggregation pinpoints the hardest problems — generating the training signal needed to build better models.

5060708090100Calorie Estimation74.083.6+9.6Skin Cancer Classification67.994.3+26.4Real or Synthetic Chest X-ray68.498.5+30.1
Best AI model in contestCentaur AggregationAccuracy vs gold standard. Illustrative data for concept review.
Ways to engage

Four ways into the Arena

Pick the one closest to what you need. We'll put you on the calendar with the right person on our team.

1 · What brings you to the Arena
2 · Anything else (optional)

A good starting point: data modality (image, text, audio, video), domain or specialty, estimated volume, and timeline.

You'll enter your name and email on the next screen.