nanobench
How it works
All evals
My evals
Navigation
How it works
All evals
My evals
Legal
Terms of service
Privacy policy
Create an eval
Generate images from your prompts and models, then evaluate them — or just keep the images.
1
Name & prompts
Eval name
A name for this evaluation. Save it to reference and share it later.
Prompts
First prompt
Question:
uploading…
✕
Add prompt
Upload prompts
Upload a .txt (one prompt per line), or a .csv with optional image and question columns — select its .zip of reference images together with the CSV.
Continue
2
Models
Recommended
Plausible price-to-quality frontier, budget to flagship
Frontier
The current top tier, price no object
Budget
Best models at $0.02/image or less
Open weights
Best models you can self-host
Over the years
One flagship per era, 2022 to today
Clear all
No matching models
Browse all models
All models
Select at least 2 models to compare.
Close
Open weights only
Model
Publisher
Elo score
Released
Open weights
Cost
No models match these filters.
model
selected
Select
Images per model per prompt
Number of images to generate for each model/prompt combination. More images yields a more accurate estimate of model performance.
Continue
3
Evaluation
Eval type
Head-to-head comparison
Raters pick the better of two images, pair by pair. You get an Elo ranking of your models.
Checklist
Raters check which of your statements are true of each image. You get a pass rate for every model and statement.
Images only
No raters — you get the finished images to browse and download.
Rater
Human raters
Claude Opus 4.5
Qwen3 VL 235B
Who judges the image comparisons — crowdsourced human raters or an AI vision model.
Question
Overall preference (default)
Prompt adherence
Image quality
Aesthetics
Realism
Custom…
Raters will be asked:
Select the image you prefer as a completion of the prompt: “{prompt}”
Raters will be asked:
Select the image that more accurately and completely depicts the prompt: “{prompt}”
Raters will be asked:
Select the image with higher overall quality — sharper detail and fewer artifacts or distortions.
Raters will be asked:
Select the image you find more beautiful or visually appealing.
Raters will be asked:
Select the image that looks more like a real photograph or human-made artwork.
Raters answer your question for every image pair.
Write
{prompt}
to insert the prompt text into your question.
Comparisons per model
15
30
Custom…
How many head-to-head comparisons each model participates in. More comparisons yields higher confidence in the final ranking.
Checklists are always rated by crowdsourced human raters.
Statements
Raters see “Check each statement that is true of this image” followed by your statements — phrase each so that checked means the model did well. Write
{prompt}
to insert the prompt text.
Add statement
+
Only prompts with statements are rated (
of
prompts).
Prompt-specific statements
Optional — statements that apply to a single prompt only.
Add
Responses per image
3
5
Custom…
How many raters check the statements for each image. More responses yields tighter confidence intervals.
You can still add an evaluation later — Clone eval on the finished page reopens this form with every image reused, so you only pay for the rating.
Continue
4
Summary
Prompts
·
models
Image generation
Images
The original images have expired — every image will be regenerated.
Total
List this eval publicly
When unchecked, this eval will only be accessible to people who have the link.
I agree to the
Terms of service
and
Privacy policy
.