Why our AI kept telling people their bad photos were good
PhotoCritic engineering · Published Aug 29, 2026
We tested five vision models on the same ten photographs from real uploads. The interesting question was never which one writes the best prose. It was which one can look at a weak photograph and say so. Every photograph is below, with its scores, so you can disagree with us.
Everything was good, and nothing was useful
Someone uploads a snapshot of their dog, taken from above on a kitchen floor, black fur crushed to a silhouette under flat ceiling light. Our critique came back at 73 out of 100. A mountain range with power lines slicing across the entire sky: 75. An empty green field on an overcast day, sharp and correctly exposed and completely without interest: 75.
None of those photographers learned anything. Looking across every critique the product had ever issued — 6,893 of them — 92.4% scored in the top two bands of a five-band scale. The bottom two bands took 1.2% between them. The average score was 82.
A score that tells someone a weak photo is good denies them the information they need to get better.
The ten photographs
Sampled evenly across six months of real uploads, so the mix is whatever people actually send us. Each was scored independently by Claude Opus 5 against our own published rubric, to give us something better calibrated than the model we were testing to measure against. Those reference scores are on the left. What our old production model said is on the right.
Shot from above on a kitchen floor. The black coat is crushed to a silhouette under flat ceiling light and the body runs out of frame.
Strong backlight rims the seedheads against a near-black ground. The best use of light in the sample.
Sharp and correctly exposed, and empty. Half the frame is featureless foreground; the only point of interest is a distant turbine.
Not a photograph. A digital painting that entered the sample naturally and became the sharpest test in it.
Drone frame of a harbour inlet. Vivid water, strong sense of place, harsh midday light.
Tight symmetrical fill-the-frame on the engine. Graphic and deliberate, with heavy processing.
Power lines slice across the whole sky and a bare branch intrudes at the corner. Every model spotted the wires; none scored it accordingly.
Mid-stride with a lifted foot, catchlight in the eye, clean separation from the background. Genuinely good wildlife work.
Selective light on a single tree against shadowed forest, with a symmetrical reflection. The strongest frame in the set.
Two balanced elements on a weathered ochre wall. Static, but a clear sense of place.
Photographs 1, 3 and 7 are the three the reference scored at 50 or below. Photograph 4 is not a photograph at all — a digital painting that arrived in the sample naturally, and turned into the sharpest test in it. Only one model named it plainly as an illustration instead of praising its exposure.
The model in production was the worst of the five
Our first instinct was to score each candidate by how far it drifted from the model already running. That produced a tidy ranking and it was worthless, because it measured closeness to an unexamined judge. Measured against the reference scores instead, the ranking inverted.
| Model | Bias | Rank corr. | Error on weak photos |
|---|---|---|---|
| deepseek-v4-flash-vision | +7.1 | 0.87 | +14.0 |
| gemini-2.5-flash-lite | +5.7 | 0.81 | +13.7 |
| qwen3.8-flash | +2.6 | 0.43 | +9.5 |
| glm-5.3-flash | +9.2 | 0.93 | +16.0 |
| qwen3-vl-235b — our old model | +15.7 | 0.77 | +31.0 |
Bias is how many points high a model runs on average. Rank correlation asks whether it puts the photographs in the right order — 1.0 is a perfect match, 0 is a shuffle. The last column is the one that matters for a product meant to help people improve.
Two things fall out of that table. Our old model ran 31 points high on weak photographs — not slightly generous, but wrong in a way that makes the product pointless. And the model that wrote by far the best criticism, GLM, was the worst in the set at scoring bad photographs. It never went below 55. Excellent advice attached to an inflated number still tells you your weak photo is fine.
It was our rubric, not the models
Five different models clustering in the same narrow band is not five coincidences. Our prompt described its bands in abstract adjectives with no examples, and nothing told a model what separates a 60 from an 80. So we rewrote it: six bands describing what a photograph at each level actually looks like, an explicit statement that most submissions belong between 30 and 55, and one load-bearing rule.
Sharpness and correct exposure are the baseline, not an achievement.
A clean photograph of an uninteresting subject in flat light is a 45, not a 75. We also stopped asking for a colour score on black-and-white images, where it means nothing, and told the model to name an illustration as an illustration. Same models, same photographs, new rubric:
Where it landed
PhotoCritic now runs DeepSeek V4 Flash Vision on the rewritten rubric. Against the model it replaced it is 61% cheaper, 33% faster, and its error on weak photographs dropped from 31 points to under 6. We also stopped making you watch a spinner: the critique now runs in the background and the page updates when it is ready.
If your score went down since August, this is why. It is not that your photographs got worse. It is that we were grading them too kindly to be any use.
What this test does not prove
The reference scores came from a model, not a person. Claude Opus 5 is a much larger vision model than any of the five we tested and sits outside all of their families, which makes it a sharper yardstick than the one we had. It is still a language model judging language models, and it plausibly shares some of their blind spots — a taste for conventional composition, an over-reading of technical polish.
Ten photographs is enough to tell reliability and speed apart. It is not enough to settle questions of taste, and a second scorer would move every number here. The directions are large and consistent; treat the exact figures as indicative. That is also why the photographs are on this page: so you can decide whether you agree with the scores.
See what the recalibrated critic makes of your work.
Upload a photo