How it works

Create

Pick the prompts and models for your dataset.

Prompt
"An astronaut walking through a flooded cathedral."

Text-to-image

Generate images from text alone.
Text-to-image example
Text-to-image example

Image-to-image

Use a reference image plus text guidance.
Image-to-image input
Image-to-image output
Guidance
"Add dark hair and outfit. Change to nighttime scene."

Bulk upload datasets

Add larger benchmark sets without manual entry.

TXT upload

One prompt per line.

Download sample
prompts.txt
a cat sitting on a windowsill
a futuristic city skyline
a bowl of ramen, overhead

CSV + ZIP

Reference images and questions.

benchmark.csv
prompt,image,question
redesign this logo,logo.png,...
put a hat on this llama,llama.jpg,...

Generate

Pay once to generate a dataset — every model runs every prompt.

Prompts
Prompt 1 of 3

“A Victorian airship drifting above a frozen ocean at sunrise.”

Models
flux
gpt image
seedream
imagen
nano banana
Outputs

Each model generates every prompt.

Generated frozen ocean airship output
Generated frozen ocean airship output alternate
Generated frozen ocean airship output variation
Generated frozen ocean airship output detail
Generated frozen ocean airship output duplicate
Generated frozen ocean airship output alternate duplicate
Generated frozen ocean airship output variation duplicate
Generated frozen ocean airship output detail duplicate
3
prompts
5
models
12
outputs

Compare

Launch one or more evaluations from your dataset to collect pairwise votes.

Customize evaluation questions

Control what raters are asked when comparing images. Choose a preset or write your own.
Default

Shows the default question to raters.

Which image best matches the prompt?
Custom

Shows your custom question to raters.

Which image best reflects {prompt}?
Questions without {prompt} won't include the prompt text.
Comparison option 1
Image 1
Comparison option 2
Image 2

Compare outputs for this prompt: “A whale swimming through clouds.”

Prev
Vote
Next

Analyze

Rank models with confidence from collected votes.

Higher Elo score is better

flux 2 max
1750
seedream v4.5
1700
qwen image
1650
gpt image 1 mini
1600
nano banana
1550
grok imagine
1525
Elo score
1,000 1,250 1,500
Votes collected
1,284
Across all comparisons
Comparisons
642
Head-to-head matchups
Confidence interval
95%
Model ranking confidence

Rankings are calculated from pairwise human preferences.

Create your own benchmark

Generate a dataset, then launch one or more evaluations to rank the models.