AI Image Editing Leaderboard
12 image-editing models, the same 18 photo edits, 5 runs each, judged with the model names hidden. Maintained by the OutfitGen Team and re-run as new models ship.
- Snapshot
- 2026-09 · v1.0
- Updated
- September 30, 2026
- Runs
- 1,080
- Source
- Lab data
Edits done right
Higher is better
Refusal rate
Lower is better
Median time per edit
Lower is better
Everyday fashion edits
| Model | ||||
|---|---|---|---|---|
| 1 | Flux 2 FlashBlack Forest Labs · Dec 2025 | 17 / 18 | ✓ 0%0 of 90 runs | 4.0 s |
| 2 | Flux 2 TurboBlack Forest Labs · Dec 2025 | 17 / 18 | ✓ 0%0 of 90 runs | 5.8 s |
| 3 | Seedream 5 FlashByteDance · Sep 2026 | 17 / 18 | ✓ 0%0 of 90 runs | 15.2 s |
| 4 | Seedream 5 ProByteDance · Jul 2026 | 17 / 18 | ✓ 0%0 of 90 runs | 27.2 s |
| 5 | Flux 2Black Forest Labs · Nov 2025 | 16 / 18 | ✓ 0%0 of 90 runs | 11.9 s |
| 6 | Seedream 5 LiteByteDance · Feb 2026 | 16 / 18 | ✓ 0%0 of 90 runs | 32.4 s |
| 7 | Qwen Image 3Alibaba · Jul 2026 | 15 / 18 | ✓ 0%0 of 89 runs | 14.5 s |
| 8 | Nano Banana 2Google · Feb 2026 | 15 / 18 | 21%19 of 90 runs | 11.3 s |
| 9 | Grok Imagine 2.0xAI · Aug 2026 | 14 / 18 | 17%15 of 89 runs | 13.2 s |
| 10 | GPT Image 2.5 Flare (medium)OpenAI · Sep 2026 | 14 / 18 | 22%20 of 90 runs | 15.1 s |
| 11 | GPT Image 2.5 Flare (high)OpenAI · Sep 2026 | 14 / 18 | 24%22 of 90 runs | 21.6 s |
| 12 | Muse ImageMeta · Sep 2026 | 12 / 18 | 26%23 of 89 runs | 16.5 s |
Ranked by edits done right, then refusal rate, then median time. One edit is about 6 points on an 18-edit test, so treat a one-edit gap as close. Per-case results are in the source study.
Portrait and scene edits
4 models, 11 edits, 3 runs each. Read the study.
| Model | Face kepthigher is better | Unasked-for changesfewer is better | Outfit copied from a photocomplete runs |
|---|---|---|---|
| GPT Image 2.5 Flare (medium) | 9 of 9 runs | 0 of 11 | 2 of 3 runs |
| GPT Image 2.5 Flare (high) | 9 of 9 runs | 2 of 11 | 2 of 3 runs |
| Nano Banana 2 | 9 of 9 runs | 4 of 11 | 3 of 3 runs |
| Seedream 5 Pro | 9 of 9 runs | Scene kept, not relit | 3 of 3 runs |
Method
- Test set
- 18 everyday fashion edits on synthetic adult photos: outfit changes, new backdrops, swimwear, activewear and catalog lingerie.
- Runs
- Every model got the same photo and instruction 5 times, through each model's hosted API, with no tuned prompts.
- Edits done right
- Edits a reviewer rated correct in every delivered run, judging shuffled contact sheets before the model names were revealed. An edit refused every time counts as not done.
- Refusal rate
- Runs the model declined with a content-policy refusal, out of runs that returned an answer. Every request was ordinary, so every refusal is a false positive.
- Median time
- Median seconds from request to delivered image, measured in the same session for every model. Provider queues move absolute times; the order is the finding.
- Lab data only
- Every number comes from test photos made for these studies. No production traffic and no user photos are used.
How to cite
OutfitGen. (2026). AI Image Editing Leaderboard, snapshot 2026-09, version 1.0. OutfitGen Research. https://outfitgen.ai/research/leaderboard
@misc{outfitgen-leaderboard-2026-09,
title = {AI Image Editing Leaderboard},
author = {{OutfitGen}},
year = {2026},
note = {OutfitGen Research, snapshot 2026-09, version 1.0},
url = {https://outfitgen.ai/research/leaderboard}
}Free to reuse with attribution under CC BY 4.0.
Snapshots
2026-09 · v1.0 · September 30, 2026
First snapshot: 12 models on the 18-edit everyday fashion set, plus the 4-model portrait and scene panel.
Want a model added? Write to the address on the about page.
OutfitGen is an AI photo editor for outfits, backgrounds and poses. These tests decide which models OutfitGen Engine trusts for which edits.
Try OutfitGen