PhotoCritic
Back to PhotoCritic

Why our AI kept telling people their bad photos were good

PhotoCritic engineering · Published Aug 29, 2026

We tested five vision models on the same ten photographs from real uploads. The interesting question was never which one writes the best prose. It was which one can look at a weak photograph and say so. Every photograph is below, with its scores, so you can disagree with us.

Everything was good, and nothing was useful

Someone uploads a snapshot of their dog, taken from above on a kitchen floor, black fur crushed to a silhouette under flat ceiling light. Our critique came back at 73 out of 100. A mountain range with power lines slicing across the entire sky: 75. An empty green field on an overcast day, sharp and correctly exposed and completely without interest: 75.

None of those photographers learned anything. Looking across every critique the product had ever issued — 6,893 of them — 92.4% scored in the top two bands of a five-band scale. The bottom two bands took 1.2% between them. The average score was 82.

A score that tells someone a weak photo is good denies them the information they need to get better.

The ten photographs

Sampled evenly across six months of real uploads, so the mix is whatever people actually send us. Each was scored independently by Claude Opus 5 against our own published rubric, to give us something better calibrated than the model we were testing to measure against. Those reference scores are on the left. What our old production model said is on the right.

PHOTOGRAPH 01reference 38/old production 73

Shot from above on a kitchen floor. The black coat is crushed to a silhouette under flat ceiling light and the body runs out of frame.

42
Comp
35
Light
40
Colour
38
Story
38
Tech
38
Overall
PHOTOGRAPH 02reference 68/old production 83

Strong backlight rims the seedheads against a near-black ground. The best use of light in the sample.

68
Comp
78
Light
70
Colour
62
Story
65
Tech
68
Overall
PHOTOGRAPH 03reference 47/old production 75

Sharp and correctly exposed, and empty. Half the frame is featureless foreground; the only point of interest is a distant turbine.

45
Comp
40
Light
55
Colour
40
Story
65
Tech
47
Overall
PHOTOGRAPH 04reference 72/old production 87

Not a photograph. A digital painting that entered the sample naturally and became the sharpest test in it.

75
Comp
62
Light
85
Colour
68
Story
70
Tech
72
Overall
PHOTOGRAPH 05reference 72/old production 89

Drone frame of a harbour inlet. Vivid water, strong sense of place, harsh midday light.

70
Comp
62
Light
78
Colour
72
Story
75
Tech
72
Overall
PHOTOGRAPH 06reference 70/old production 70

Tight symmetrical fill-the-frame on the engine. Graphic and deliberate, with heavy processing.

75
Comp
70
Light
72
Colour
62
Story
68
Tech
70
Overall
PHOTOGRAPH 07reference 45/old production 75

Power lines slice across the whole sky and a bare branch intrudes at the corner. Every model spotted the wires; none scored it accordingly.

35
Comp
58
Light
62
Colour
40
Story
50
Tech
45
Overall
PHOTOGRAPH 08reference 82/old production 90

Mid-stride with a lifted foot, catchlight in the eye, clean separation from the background. Genuinely good wildlife work.

82
Comp
82
Light
80
Colour
75
Story
85
Tech
82
Overall
PHOTOGRAPH 09reference 85/old production 89

Selective light on a single tree against shadowed forest, with a symmetrical reflection. The strongest frame in the set.

85
Comp
88
Light
85
Colour
85
Story
82
Tech
85
Overall
PHOTOGRAPH 10reference 73/old production 78

Two balanced elements on a weathered ochre wall. Static, but a clear sense of place.

72
Comp
65
Light
78
Colour
72
Story
78
Tech
73
Overall

Photographs 1, 3 and 7 are the three the reference scored at 50 or below. Photograph 4 is not a photograph at all — a digital painting that arrived in the sample naturally, and turned into the sharpest test in it. Only one model named it plainly as an illustration instead of praising its exposure.

The model in production was the worst of the five

Our first instinct was to score each candidate by how far it drifted from the model already running. That produced a tidy ranking and it was worthless, because it measured closeness to an unexamined judge. Measured against the reference scores instead, the ranking inverted.

ModelBiasRank corr.Error on weak photos
deepseek-v4-flash-vision+7.10.87+14.0
gemini-2.5-flash-lite+5.70.81+13.7
qwen3.8-flash+2.60.43+9.5
glm-5.3-flash+9.20.93+16.0
qwen3-vl-235b — our old model+15.70.77+31.0

Bias is how many points high a model runs on average. Rank correlation asks whether it puts the photographs in the right order — 1.0 is a perfect match, 0 is a shuffle. The last column is the one that matters for a product meant to help people improve.

Two things fall out of that table. Our old model ran 31 points high on weak photographs — not slightly generous, but wrong in a way that makes the product pointless. And the model that wrote by far the best criticism, GLM, was the worst in the set at scoring bad photographs. It never went below 55. Excellent advice attached to an inflated number still tells you your weak photo is fine.

It was our rubric, not the models

Five different models clustering in the same narrow band is not five coincidences. Our prompt described its bands in abstract adjectives with no examples, and nothing told a model what separates a 60 from an 80. So we rewrote it: six bands describing what a photograph at each level actually looks like, an explicit statement that most submissions belong between 30 and 55, and one load-bearing rule.

Sharpness and correct exposure are the baseline, not an achievement.

A clean photograph of an uninteresting subject in flat light is a 45, not a 75. We also stopped asking for a colour score on black-and-white images, where it means nothing, and told the model to name an illustration as an illustration. Same models, same photographs, new rubric:

Dog on the kitchen floor
7346
reference said 38
Empty overcast farmland
7548
reference said 47
Peaks behind power lines
7553
reference said 45

Where it landed

PhotoCritic now runs DeepSeek V4 Flash Vision on the rewritten rubric. Against the model it replaced it is 61% cheaper, 33% faster, and its error on weak photographs dropped from 31 points to under 6. We also stopped making you watch a spinner: the critique now runs in the background and the page updates when it is ready.

If your score went down since August, this is why. It is not that your photographs got worse. It is that we were grading them too kindly to be any use.

What this test does not prove

The reference scores came from a model, not a person. Claude Opus 5 is a much larger vision model than any of the five we tested and sits outside all of their families, which makes it a sharper yardstick than the one we had. It is still a language model judging language models, and it plausibly shares some of their blind spots — a taste for conventional composition, an over-reading of technical polish.

Ten photographs is enough to tell reliability and speed apart. It is not enough to settle questions of taste, and a second scorer would move every number here. The directions are large and consistent; treat the exact figures as indicative. That is also why the photographs are on this page: so you can decide whether you agree with the scores.

See what the recalibrated critic makes of your work.

Upload a photo