Create an eval

Generate images from your prompts and models, then evaluate them — or just keep the images.
A name for this evaluation. Save it to reference and share it later.
Upload a .txt (one prompt per line), or a .csv with optional image and question columns — select its .zip of reference images together with the CSV.
No matching models
Number of images to generate for each model/prompt combination. More images yields a more accurate estimate of model performance.
Who judges the image comparisons — crowdsourced human raters or an AI vision model.
Raters will be asked: Select the image you prefer as a completion of the prompt: “{prompt}” Raters will be asked: Select the image that more accurately and completely depicts the prompt: “{prompt}” Raters will be asked: Select the image with higher overall quality — sharper detail and fewer artifacts or distortions. Raters will be asked: Select the image you find more beautiful or visually appealing. Raters will be asked: Select the image that looks more like a real photograph or human-made artwork. Raters answer your question for every image pair.
Write {prompt} to insert the prompt text into your question.
How many head-to-head comparisons each model participates in. More comparisons yields higher confidence in the final ranking.
Checklists are always rated by crowdsourced human raters.
Raters see “Check each statement that is true of this image” followed by your statements — phrase each so that checked means the model did well. Write {prompt} to insert the prompt text.
Only prompts with statements are rated ( of prompts).
Optional — statements that apply to a single prompt only.
How many raters check the statements for each image. More responses yields tighter confidence intervals.
You can still add an evaluation later — Clone eval on the finished page reopens this form with every image reused, so you only pay for the rating.

Prompts
·
models
Image generation
Images
The original images have expired — every image will be regenerated.
Total