AI evaluation

VLM Image-Pair Judge

A Python tool that uses vision-language models to judge pairs of AI-generated images against a written rubric.

PythonOpenRouterGeminiOpenAI

This tool uses vision-language models, through OpenRouter, Gemini, and OpenAI, to automatically judge pairs of AI-generated images. The judgment is driven by a written rubric, so the quality of the rubric decides the quality of the result.

I improved the rubric by testing it against live, real examples, and widened it to cover anatomy, water and reflections, document text, chart inspection, and tonal intent.

Why the weights matter

The same two images can produce a different winner depending on what the rubric cares about. These scores are sample numbers for illustration.