Text rendering: literary passages (pairwise evaluation)

Comparison
10 prompts
·
13 models
Raters are asked: Select the image you prefer as a completion of the prompt. Prioritise the accuracy of the text: {prompt}
Completed

Add more comparisons

More comparisons tighten the confidence intervals on each model's Elo score.
Total
additional comparisons

Elo scores

Each bar spans a 95% credible interval; the dot marks the mean. Hover for exact values.

Cost vs. Elo score