How it works
Create
Pick the prompts and models for your dataset.
Text-to-image
Image-to-image
Bulk upload datasets
Add larger benchmark sets without manual entry.
TXT upload
One prompt per line.
CSV + ZIP
Reference images and questions.
Generate
Pay once to generate a dataset — every model runs every prompt.
“A Victorian airship drifting above a frozen ocean at sunrise.”
“A lone samurai crossing a wide and mysterious field.”
“A giant turtle carrying an ancient forest on its huge shell.”
Each model generates every prompt.
Compare
Launch one or more evaluations from your dataset to collect pairwise votes.
Customize evaluation questions
Shows the default question to raters.
Shows your custom question to raters.
Compare outputs for this prompt: “A whale swimming through clouds.”
Analyze
Rank models with confidence from collected votes.
Higher Elo score is better
Rankings are calculated from pairwise human preferences.