This tool uses vision-language models, through OpenRouter, Gemini, and OpenAI, to automatically judge pairs of AI-generated images. The judgment is driven by a written rubric, so the quality of the rubric decides the quality of the result.
I improved the rubric by testing it against live, real examples, and widened it to cover anatomy, water and reflections, document text, chart inspection, and tonal intent.
Why the weights matter
The same two images can produce a different winner depending on what the rubric cares about. These scores are sample numbers for illustration.