OutfitGen

OutfitGen Research

The benchmark for AI photo editing models, on the edits people actually make

OutfitGen runs the leading image-editing models in production and re-tests them every few weeks. These are the write-ups: the same photos and instructions sent to every model, judged with the names hidden, with every output published so the call can be checked by anyone.

Published studies
1
Models compared
4
Edit cases
11
Raw outputs published
55

Latest study

September 14, 2026 · Version 1.0

Which AI image model keeps your face? Nano Banana 2 vs GPT Image 2.5 Flare vs Seedream 5 Pro

Four AI image-editing model configurations given the same 11 photo edits, three runs each, judged blind: 132 outputs scored on identity preservation, scene preservation, reference-garment transfer and relative speed.

Read the study
Source photo: man in his forties at a three-quarter angle with rectangular glasses, a cheek mole and asymmetric eyebrows, wearing a henley in a kitchenNano Banana 2 result: glasses, cheek mole and crooked smile unchanged, henley replaced by a navy blazer and white shirt, kitchen unchangedGPT Image 2.5 Flare at medium: the same glasses, mole and uneven eyebrows, navy blazer over a white open-collar shirt, head angle unchanged

How every study works

The rules are the same for every study, so results stay comparable from one model generation to the next.

01

Identical inputs, production path

Every model receives the same photo and the same instruction, sent through the request path OutfitGen uses for paying customers. No model gets a tuned prompt.

02

Repeated, because these models are not deterministic

At least three runs per model per case. A single output can flatter or damn a model by luck, and a two-run read has already produced a false result in our own work.

03

Judged blind, then unblinded

Contact sheets are shuffled per case and verdicts are written before the key is opened. The unblinded verdicts, and every image behind them, are published.

04

Faces built to expose drift

Portrait cases use subjects with specific identity markers at non-frontal angles, because ordinary stock faces are exactly what edit models regenerate cleanly.

05

Stated, dated, versioned

Each study carries its design table, its limitations, a version and a changelog, and a stable URL to cite. A refreshed finding is a new version, never a silent edit.

06

Proven on real traffic before it ships

A model that wins a study earns a live trial on a slice of OutfitGen generations. Refusal behaviour on real customer photos only shows there, and it has reversed a study result before.

Why OutfitGen publishes this

The model behind an AI photo editor changes what the product can do, and the market moves every few weeks. OutfitGen chooses its Standard and Pro models by running exactly these studies, so publishing them costs nothing extra and lets anyone check the work: researchers comparing models, builders choosing one, and people deciding which tool to trust with their own photo.

Studies are free to cite and reuse with attribution. Each one carries a stable URL, a version and a citation block. Questions, corrections and requests for a model to be included go to the address on the about page.

See the winning models on your own photo

OutfitGen runs them for you. Upload a photo, describe the change, no account needed to start.

Try the AI Clothes Changer