OutfitGen

Research · September 14, 2026 · Version 1.0

Which AI image model keeps your face? Nano Banana 2 vs GPT Image 2.5 Flare vs Seedream 5 Pro

A first-party study by the OutfitGen research team. About OutfitGen · support@outfitgen.ai

OutfitGen's Pro mode runs on Google's Nano Banana 2 because, when it was chosen, it was the model that kept a person's face intact through an outfit change. Five months on, OpenAI's GPT Image 2.5 Flare has arrived at a fraction of the compute, so we put both, plus ByteDance's Seedream 5 Pro, through the same 11 edits on the same photos, three runs each, and judged the 132 outputs with the model names hidden. The short version: all three model families now preserve identity on deliberately difficult faces, GPT Image 2.5 Flare edits more precisely than the incumbent on scene-preserving tasks, and Nano Banana 2 is still the fastest.

Key findings

  1. 1

    Identity preservation no longer separates the leading image-editing models: all four configurations kept every named identity marker across 36 identity-stress portrait runs.

    The markers were a cheek mole with glasses, grey streaks with a gold nose stud, and a scar through one eyebrow with a white beard, each on a portrait shot at a non-frontal angle.

  2. 2

    The models separated on collateral editing instead. Nano Banana 2 changed something it was not asked to change on 4 of the 11 cases; GPT Image 2.5 Flare at medium quality did so on none.

    Nano Banana 2 rebuilt a street behind a pose change, recoloured a black jacket brown during a background swap, added a smile to a seated pose, and added a collar the instruction did not ask for.

  3. 3

    Copying a complete outfit from a reference photo still favours Nano Banana 2 and Seedream 5 Pro, which transferred the whole outfit in 3 of 3 runs; both GPT Image 2.5 Flare settings kept the subject's original jeans or shoes in 1 run of 3.

    No model imported the reference photo's own person, which is the failure that matters most on this task.

  4. 4

    GPT Image 2.5 Flare at its default high-quality setting was not visibly better than the same model at medium on any of the 11 cases, and took about twice as long per edit as Nano Banana 2.

  5. 5

    Nano Banana 2 was the fastest configuration tested. GPT Image 2.5 Flare at medium took about a third longer per edit, and Seedream 5 Pro about two and a half times as long.

    Times are stated relative to Nano Banana 2 because absolute timings depend on queueing conditions, not only on the model.

  6. 6

    None of the four configurations refused any of the 40 everyday swimwear, activewear and eveningwear probes, so this corpus cannot rank the models on refusal behaviour.

    Refusal rate is measured on live traffic instead, and it is the number that decides an OutfitGen model change.

  7. 7

    OutfitGen keeps Nano Banana 2 on Pro mode and Seedream 5 Pro on reference-garment edits, and is trialling GPT Image 2.5 Flare at medium quality on a slice of live Pro traffic.

What question does this study answer?

When a person uploads a photo of themselves and asks for a new outfit, background or pose, which current model returns the same person, changes only what was asked, and does it quickly enough for an interactive tool? And has a cheaper-to-run model caught up with the one OutfitGen chose for Pro in April 2026?

Which models were compared?

ModelMakerRole in this study
Nano Banana 2 (Google Gemini 3.1 Flash Image)GoogleThe model OutfitGen runs for Pro-quality generations today. Chosen in April 2026 for keeping faces intact. The control in this study.
GPT Image 2.5 Flare, medium qualityOpenAIOpenAI's fast image model at the setting OutfitGen would ship. The challenger.
GPT Image 2.5 Flare, high qualityOpenAIThe same model at its default, slower setting, included to show whether the extra compute changes the result.
Seedream 5 ProByteDanceThe model OutfitGen runs when a user uploads a reference garment to copy. Included as the cross-family comparison.

This study compares the models, not the products built on them. The product-level comparison lives at OutfitGen vs Nano Banana.

How do the models compare at a glance?

Which AI image model keeps your face? Nano Banana 2 vs GPT Image 2.5 Flare vs Seedream 5 Pro: one row per model configuration, scored on Identity preservation, Scene preservation, Reference outfit transfer, Relative speed and its role in OutfitGen.
ModelMakerIdentity preservationScene preservationReference outfit transferRelative speedRole in OutfitGen
Nano Banana 2GoogleAll markers held across 9 identity-stress runs. Reshaped a white beard in 2 of 3 runs on one portrait, the only visible identity drift in the study.Changed something it was not asked to change on 4 of 11 cases.Transferred the complete reference outfit in 3 of 3 runs.Fastest configuration in the study. All other timings below are stated against it.Runs Pro mode generations today.
GPT Image 2.5 Flare (medium)OpenAIAll markers held across 9 identity-stress runs. Closest to the source photo on the beard and scar portrait.No unrequested change to a scene, to lighting or to an expression on any of the 11 cases.Transferred the reference jacket and top every run, but kept the subject's original jeans or shoes in 1 run of 3.About a third longer per edit than Nano Banana 2.Candidate on live trial for Pro mode. Not shipped.
GPT Image 2.5 Flare (high)OpenAIAll markers held across 9 identity-stress runs.Two unrequested changes across the 11 cases: a collar on one portrait and a green darker than the one asked for on the dress.Transferred the reference jacket and top every run, but kept the subject's original jeans or shoes in 1 run of 3.About twice as long per edit as Nano Banana 2.Not used. No visible gain over the medium setting.
Seedream 5 ProByteDanceAll markers held across 9 identity-stress runs, with the face varying slightly between runs on one portrait.Left scenes alone, but did not relight the subject for a new background, which reads as a cut-out.Transferred the complete reference outfit in 3 of 3 runs.About two and a half times as long per edit as Nano Banana 2.Runs reference-garment generations today.

Every cell summarizes the per-case verdicts below, which were written while the model names were hidden. Counts are runs: each of the 11 cases ran three times per model configuration.

How was the study run?

Study design: the parameters this comparison was run under.
Study typeBlinded side-by-side comparison of production image-editing endpoints
Run date14 September 2026
Model configurations compared4, across 3 model families (Google, OpenAI, ByteDance)
Editing cases11, across 4 job types: described outfit change, reference-garment transfer, background replacement, pose change
Identity-stress cases3 portraits generated with specific markers at non-frontal angles
Runs per model per case3
Outputs judged132
Refusal probes40 (5 everyday swimwear, activewear and eveningwear photos, 2 runs, 4 configurations)
Total generations172
Request pathThe production request path, at the output settings a paying OutfitGen user receives
JudgingOne reviewer, contact sheets with the model columns shuffled per case, verdicts written before the key was opened
Primary measureIs this the same person?
Secondary measuresDid the model change only what was asked, and how long did it take relative to the other models?
Excluded runs6 queueing timeouts on the first runs dispatched, spread across three of the four configurations
  1. Eleven cases across OutfitGen's four tools: three identity-stress portraits generated with specific markers (a cheek mole and glasses, grey streaks and a nose stud, an eyebrow scar and a white beard) at non-frontal angles, plus described-outfit changes, two reference-garment transfers, two background swaps and two pose changes.
  2. Every model received the identical photo and instruction through the same request path OutfitGen uses in production, at the output settings a paying user gets.
  3. Three runs per model per case, 132 outputs in total, because none of these models is deterministic and a single run can flatter or damn a model by luck.
  4. Outputs were laid out in contact sheets with the columns shuffled per case and judged without knowing which model produced which column. Verdicts were written down before the key was opened.
  5. The primary question for every case was the one a stranger would ask: is this the same person? Second, did the model change only what was asked? Third, how long did it take?
  6. A separate set of ordinary swimwear, activewear and eveningwear photos was sent to each model to check for over-cautious refusals on everyday customer pictures. None of the four refused any of them.
  7. Six runs stalled in the provider's queue during the busiest minute of the test, across three of the four configurations, and were counted as failures rather than as model behaviour. They did not recur on the repeat runs.

What happened on each case?

One representative output per model is shown for each case; the judging used every run, 132 outputs in total.

Identity stress: glasses, cheek mole, asymmetric brows, three-quarter angle

Instruction:a navy blue unstructured blazer over a crisp white open-collar shirt

Look for: Glasses shape, the cheek mole, the uneven eyebrows and the crooked smile all have to survive while the henley becomes a blazer and shirt.

Source photo: man in his forties at a three-quarter angle with rectangular glasses, a cheek mole and asymmetric eyebrows, wearing a henley in a kitchen
Source photo
Nano Banana 2 result: glasses, cheek mole and crooked smile unchanged, henley replaced by a navy blazer and white shirt, kitchen unchanged
Nano Banana 2
GPT Image 2.5 Flare at medium: the same glasses, mole and uneven eyebrows, navy blazer over a white open-collar shirt, head angle unchanged
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: identity markers intact, navy blazer in a slightly different cut, same kitchen and head angle
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: cheek mole, glasses and crooked smile preserved, navy blazer over a white shirt, kitchen unchanged
Seedream 5 Pro

Verdict: All four models kept every identity marker in all three runs, including the mole and the slightly crooked smile. The blazers differ in cut and shade, none is wrong, and no model changed the kitchen or the head angle. A tie.

In OutfitGen this edit is a text-prompt job for the AI clothes changer.

Identity stress: grey streaks, nose stud, chin scar, tilted head

Instruction:a chunky mustard-yellow knit cardigan over a cream blouse

Look for: The grey streaks, the gold nose stud, the small chin scar and the head tilt have to survive; the black top becomes a cardigan and blouse.

Source photo: woman in her fifties with her head tilted, grey-streaked hair, a gold nose stud and a small chin scar, wearing a black top
Source photo
Nano Banana 2 result: grey streaks, nose stud and chin scar preserved, mustard cardigan added over a collared shirt the instruction did not ask for
Nano Banana 2
GPT Image 2.5 Flare at medium: streaks, stud and scar preserved, mustard cardigan over the round-neck cream blouse asked for, framing unchanged
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: identity markers preserved, mustard cardigan over a collared button-up rather than the blouse asked for
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: streaks, stud and scar preserved, mustard cardigan over a cream blouse, with the face varying slightly between runs
Seedream 5 Pro

Verdict: Every model preserved the streaks, the stud and the scar in every run. GPT Image 2.5 Flare at medium kept the original framing and a round-neck blouse most faithfully; Nano Banana 2 and Flare at high both turned the blouse into a collared button-up, which the instruction did not ask for; Seedream 5 Pro varied the face slightly between runs. All pass on identity.

Layered looks like this one are what people ask of the AI clothes changer.

Identity stress: eyebrow scar, once-broken nose, white beard

Instruction:a forest green waterproof rain jacket, zipped up, with a black base layer visible at the collar

Look for: The scar through the eyebrow, the bent nose and the exact beard line have to survive; the denim work shirt becomes a zipped rain jacket.

Source photo: man in his sixties at a three-quarter angle with a scar through one eyebrow, a once-broken nose and a white beard, wearing a denim work shirt
Source photo
Nano Banana 2 result: the eyebrow scar survives but the white beard is reshaped and the face softened, forest green rain jacket zipped up
Nano Banana 2
GPT Image 2.5 Flare at medium: beard line and eyebrow scar match the source photo, forest green rain jacket with visible rain beading
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: the same beard line and scar as the source, zipped green rain jacket with a black base layer at the collar
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: face and beard unchanged, plain forest green rain jacket zipped over a black base layer
Seedream 5 Pro

Verdict: The two GPT Image 2.5 Flare rows are the closest to the source: same beard line, same scar, rain beading on the jacket. Nano Banana 2 kept the scar but reshaped the beard and softened the face in two of three runs, the only visible identity drift in the whole study. Seedream 5 Pro kept the face and gave the plainest jacket.

Distinctive faces like this one are the reason identity is the first thing checked in the AI clothes changer.

Described suit on a half-body subject

Instruction:a charcoal grey tailored wool suit with a crisp white dress shirt and a burgundy tie

Look for: Same face and cafe scene; the sweater becomes a real suit with correct lapels, collar and tie.

Source photo: man seated at a cafe table in a sweater, hands resting on the table, window behind him
Source photo
Nano Banana 2 result: charcoal wool suit with a white shirt and burgundy tie, same face, hands and cafe window
Nano Banana 2
GPT Image 2.5 Flare at medium: charcoal tailored suit with correct lapels, collar and burgundy tie, cafe scene unchanged
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: charcoal suit with a white dress shirt and burgundy tie, face and cafe window unchanged
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: charcoal wool suit, white shirt and burgundy tie, identical face, hands and background
Seedream 5 Pro

Verdict: Twelve clean suits. Every model kept the face, the hands on the table and the cafe window. Nothing to separate them.

This is the everyday case behind the AI clothes changer.

Described dress on a full-body subject

Instruction:a flowing emerald green summer dress with thin shoulder straps and tan leather sandals

Look for: Whole-outfit replacement while identity, pose and park stay put; hands and hem should survive.

Source photo: woman standing full length on a park path in casual clothes
Source photo
Nano Banana 2 result: flowing emerald green summer dress with thin straps and tan sandals, same person, pose and park
Nano Banana 2
GPT Image 2.5 Flare at medium: emerald green strappy summer dress and tan leather sandals, pose and park unchanged
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: strappy summer dress in a darker green than the emerald asked for, tan sandals, same park
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: emerald green summer dress with thin straps and tan sandals, identity and park unchanged
Seedream 5 Pro

Verdict: All four delivered a green dress with thin straps and tan sandals on the same woman in the same park. Flare at high leaned darker than emerald; Seedream 5 Pro and Flare at high each lost one first run to a timeout during the busiest minute of the test, which is an infrastructure note rather than a model one.

Full-length outfit swaps like this one run on the AI clothes changer.

Garment reference with a legible logo

Instruction:Change the outfit in this photo to match the reference clothing image (a burnt-orange hoodie with an ALPINE CO. chest graphic)

Look for: Does the ALPINE CO. wordmark and the pine emblem transfer, at the right scale, on the right person?

Source photo: person in their own outfit, before the burnt-orange ALPINE CO. hoodie reference was applied
Source photo
Nano Banana 2 result: burnt-orange hoodie transferred with the ALPINE CO. wordmark and pine emblem, lettering slightly softer than the reference
Nano Banana 2
GPT Image 2.5 Flare at medium: burnt-orange hoodie with the ALPINE CO. wordmark and pine emblem legible at the right scale
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: burnt-orange hoodie transferred, wordmark and pine emblem legible, same person and setting
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: burnt-orange hoodie with a crisp ALPINE CO. wordmark and pine emblem, same person and setting
Seedream 5 Pro

Verdict: Every model transferred the hoodie, the wordmark and the pine emblem legibly. Nano Banana 2 rendered the lettering a touch softer in one run. Effectively a tie.

Uploading a garment photo to copy is the AI virtual try-on.

Outfit reference worn by a different person

Instruction:Change the outfit in this photo to match the reference clothing image (a leather biker jacket, black roll-neck, dark trousers and boots, worn by an older man)

Look for: The hard one: transfer the outfit without importing the reference model's face, body or studio.

Source photo: woman in her own outfit, before the leather biker jacket reference worn by another person was applied
Source photo
Nano Banana 2 result: the complete reference outfit transferred, leather biker jacket, roll-neck, dark trousers and boots, on the original woman
Nano Banana 2
GPT Image 2.5 Flare at medium: leather jacket and roll-neck transferred, but the subject's original jeans kept in one run of three
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: leather jacket and roll-neck transferred, original sneakers kept in one run of three, face unchanged
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: the full reference outfit transferred, jacket, roll-neck, trousers and boots, without importing the other person
Seedream 5 Pro

Verdict: Nobody imported the other person, which is the failure that matters most. Nano Banana 2 and Seedream 5 Pro transferred the complete outfit in all three runs. Both GPT Image 2.5 Flare rows transferred the jacket and roll-neck every time but kept the woman's original jeans or sneakers in one run each, a partial transfer a user would notice.

Copying a whole look off someone else's photo is the harder half of the AI virtual try-on.

Office behind a close-up face

Instruction:Replace the background with a modern professional office with floor-to-ceiling windows and city view. Keep the person exactly as they are.

Look for: Identity and hair edges at close range; believable depth of field and colour temperature match.

Source photo: close-up portrait against its original background, before any background replacement
Source photo
Nano Banana 2 result: modern office with a city view behind the subject, with the hair edge and skin regenerated slightly more than the other models
Nano Banana 2
GPT Image 2.5 Flare at medium: floor-to-ceiling office windows behind an unchanged portrait, with matched depth of field
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: office background with a city view, portrait and hair edge left close to the source photo
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: modern office background with a city view, portrait kept close to the original
Seedream 5 Pro

Verdict: All four produced a convincing office and kept the face. Nano Banana 2's version regenerated the hair edge and skin slightly more than the others; the GPT Image 2.5 Flare rows and Seedream 5 Pro kept the portrait closest to the original. Two first runs timed out during the busiest minute, unrelated to the model.

Office backdrops for headshots are the most requested job in the AI background changer.

Beach behind a full-body subject

Instruction:Replace the background with a tropical beach paradise with clear turquoise water and white sand. Keep the person exactly as they are.

Look for: Full-body cutout quality: feet, ground contact, and whether the evening light gets relit for the beach.

Source photo: man photographed full length in evening light, with the original background still in place
Source photo
Nano Banana 2 result: tropical beach background, natural in two runs, with the black jacket recoloured brown in the third
Nano Banana 2
GPT Image 2.5 Flare at medium: tropical beach with the subject relit for daylight and standing on the sand, the most natural composite in the study
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: beach background with the subject relit for daylight, close behind the medium setting
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: beach background, but the subject keeps the original evening lighting, which reads as a cut-out
Seedream 5 Pro

Verdict: The clearest separation in the study. GPT Image 2.5 Flare at medium produced the most natural composite in every run, with the subject relit for daylight and standing on the sand. Flare at high was close behind. Nano Banana 2 was good in two runs and in the third changed the black jacket to brown, an edit nobody asked for. Seedream 5 Pro placed the subject on the beach with his original evening lighting intact, which reads as a cut-out.

Relighting a person for a new scene is the hard part of the AI background changer.

Walking pose, outfit preserved

Instruction:a natural walking pose, mid-stride, with confident posture and the outfit clearly visible (keep the outfit from the uploaded image)

Look for: A genuinely new pose, same person, same clothes, natural hands and feet.

Source photo: man standing still against a brick wall in evening light, wearing a bomber jacket, jeans and trainers
Source photo
Nano Banana 2 result: mid-stride walking pose with the outfit intact, but the street rebuilt with new buildings and different light
Nano Banana 2
GPT Image 2.5 Flare at medium: natural mid-stride walk with the same bomber jacket, jeans and trainers, brick wall and evening light kept
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: mid-stride walking pose with the outfit preserved and the original brick wall and evening light kept
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: mid-stride walking pose with the outfit, wall and evening light unchanged
Seedream 5 Pro

Verdict: Every model produced a real mid-stride walk with the bomber jacket, jeans and trainers intact. Nano Banana 2 also rebuilt the street in two of three runs, adding buildings and changing the light; the instruction was to change the pose, not the scene. The GPT Image 2.5 Flare rows and Seedream 5 Pro kept the brick wall and the evening light.

Changing a stance while the clothes and the street stay put is the AI pose changer.

Seated pose on a tight crop, outfit preserved

Instruction:a relaxed seated pose with friendly posture, looking toward the camera (keep the outfit from the uploaded image)

Look for: Half-body crops force the model to invent the body below the frame: does it stay coherent and in-crop?

Source photo: half-body crop of a woman, where the body below the frame has to be invented for a seated pose
Source photo
Nano Banana 2 result: plausible seated pose, but with a broad smile added and the face changed slightly
Nano Banana 2
GPT Image 2.5 Flare at medium: relaxed seated pose looking at the camera, same face and shirt, body coherent inside the crop
GPT Image 2.5 Flare (medium)
GPT Image 2.5 Flare at high: natural seated pose with the same face and shirt, and a coherent body below the original crop
GPT Image 2.5 Flare (high)
Seedream 5 Pro result: upright and neutral seated pose, correct but less relaxed than the instruction asked for
Seedream 5 Pro

Verdict: All four invented a plausible seated body. GPT Image 2.5 Flare at both settings gave the most natural seated poses with the same face and shirt. Nano Banana 2 added a broad smile and slightly changed the face in two runs. Seedream 5 Pro sat her upright and neutral, correct but less relaxed than asked.

Tight crops are the stress test for the AI pose changer.

What did the blinded judging find?

All three model families now hold identity on difficult faces

The identity-stress portraits were built to expose face drift, the failure that made OutfitGen roll back a model within hours in April 2026. Across three portrait cases, four model configurations and three runs each, 36 runs in total, every model kept the mole, the glasses, the grey streaks, the nose stud and the eyebrow scar. The only visible drift was Nano Banana 2 reshaping the white beard on the workshop portrait in two runs of three. Identity preservation has stopped being the deciding axis between these models.

GPT Image 2.5 Flare changes less that it was not asked to change

Where the models separated, it was on collateral edits. Nano Banana 2 rebuilt the street behind a walking pose in two of three runs, recoloured a jacket during a background swap, added a smile to a seated pose, and added a collar to a blouse. GPT Image 2.5 Flare at medium kept scenes, lighting and expressions closest to the original in the background and pose cases, and produced the most natural beach composite in the study. This is the opposite of the reputation the GPT Image line carries on public editing benchmarks, and it held across every run we saw.

Garment transfer from a reference photo still favours Nano Banana 2 and Seedream 5 Pro

Given a photo of a different person wearing an outfit, both GPT Image 2.5 Flare settings transferred the jacket and top every time but kept the original jeans or sneakers in one run out of three. Nano Banana 2 and Seedream 5 Pro moved the complete outfit in all three runs. Nobody imported the other person's face or body, which is the failure that matters most on this task.

The high setting did not buy anything the medium setting lacked

GPT Image 2.5 Flare at its default high-quality setting took roughly a third longer per edit than at medium and was not visibly better on any of the eleven cases. On the dress case it drifted to a darker green than asked. For interactive editing, medium is the setting worth running.

Speed: Nano Banana 2 first, Flare a third slower, Seedream well behind

Measured end to end through the same path, Nano Banana 2 was the fastest model in the study. GPT Image 2.5 Flare at medium took about a third longer per edit, Flare at high about twice as long, and Seedream 5 Pro roughly two and a half times as long. A handful of first runs across three of the four models stalled during the busiest minute of the test and were counted as failures; those were queueing, not model behaviour, and did not recur.

A benign corpus cannot rank models on refusals, and refusals decide switches

Forty probes of ordinary swimwear, activewear and eveningwear photos drew zero refusals from any of the four configurations. That is the same result the previous OutfitGen model evaluation produced, and it is why an offline study cannot settle a model change on its own: how often a model declines a real customer photo only shows on live traffic. It is the number that reversed OutfitGen's last model switch.

What does this mean for OutfitGen?

Nano Banana 2 keeps running OutfitGen's Pro mode today, and Seedream 5 Pro keeps the reference-garment job it does best. GPT Image 2.5 Flare at medium has earned a live trial: on blinded quality it matched the incumbent on identity and beat it on leaving the rest of the photo alone, at a fraction of the compute. What this study cannot measure is how often each model refuses a real customer photo, which only live traffic shows and which decided OutfitGen's last model switch. Flare will run on a slice of Pro generations next to Nano Banana 2, and if its refusal behaviour holds up it becomes the Pro model. Either way, the model behind Pro mode is chosen this way: same photos, same instructions, judged blind, then proven on real traffic before it takes over.

The cases above map onto the tools people use: described outfit changes are the AI clothes changer, a garment photo copied onto a person is the AI virtual try-on, scene replacement is the AI background changer, and a new stance with the same clothes is the AI pose changer.

How can this be reproduced?

This study is built to be run again, against the same photographs, every time a candidate model appears. Anyone reproducing it needs the same four ingredients OutfitGen uses.

  • Fixed-seed sources. Every source photograph is generated once with a fixed seed and cached, so a re-run in six months compares new model outputs against the same faces and scenes rather than against new ones.
  • Three runs per case per model. These models are not deterministic, and one run can flatter or damn a model by luck. Three runs is the minimum that shows a failure repeating.
  • Identical inputs through one path. Every model receives the same photograph and the same instruction through the production request path, at the settings a paying user receives, so the comparison is of models and not of request shaping.
  • Blinded judging. Outputs go into contact sheets with the model columns shuffled per case, and the verdicts are written down before the key is opened. A name is enough to bias a verdict that a picture would not.
  • A re-run on every model change. OutfitGen repeats this evaluation whenever a candidate model appears or a shipped model updates, and publishes the result here as a new version of this study or as a new study.
  • The images on this page are the model outputs as returned, resized for the web and otherwise unretouched. One representative run per model per case is shown; the judging used all three.

The corpus

OutfitGen Pro-tier image editing evaluation corpus, September 2026. Eleven photo-editing cases (three identity-stress portraits, two described-outfit changes, two reference-garment transfers, two background replacements, two pose changes) built on fixed-seed source photographs, each run three times against four model configurations for 132 judged outputs, plus 40 refusal probes on everyday swimwear, activewear and eveningwear photos.

Measured

  • Identity preservation
  • Scene preservation
  • Reference outfit transfer completeness
  • Relative time to result
  • Refusal on benign photos

What are the limitations of this study?

  • The source photos are AI-generated portraits with deliberately distinctive features, not customer uploads. Real photos are messier and this study cannot see how each model handles low light, phone compression or partial faces.
  • Three runs per case bounds luck but does not bound it tightly. A model that failed one run in three on a case may fail one in six or one in two on the same case over a larger sample.
  • Refusal behaviour on borderline content was probed with a small set of everyday photos and none was refused; that says nothing about the refusal rate on real traffic, which is measured separately before any model change ships.
  • Judging was done by one reviewer with the model names hidden. It is a careful read, not a panel study.
  • Timings were taken during a single run window and include the provider's queueing, so they are stated as ratios between models rather than as absolute numbers.

What do the terms in this study mean?

Identity preservation
Whether a stranger comparing the edited photo with the original would say it is the same person. Judged here on named markers: a cheek mole, glasses, grey streaks, a gold nose stud, a chin scar, an eyebrow scar and a beard line.
Scene preservation
Whether the model changed anything it was not asked to change. Collateral edits seen in this study include a rebuilt street, a recoloured jacket, an added smile and an added collar.
Reference outfit transfer
Copying an outfit from a second uploaded photograph onto the person in the first, without importing the reference photograph's own face, body or studio.
Identity-stress case
A portrait generated with deliberately distinctive features at a non-frontal angle. Ordinary stock portraits hide face drift, which is how a model with a drift problem once passed an OutfitGen evaluation.
Run
One generation. Each case was run three times per model configuration, because these models are not deterministic.
Blinded judging
Reviewing outputs with the model names hidden and the columns shuffled, so the verdict cannot follow the reputation of the name.
Refusal
A model declining to edit a photograph its safety system reads as unsafe. Refusal rate cannot be read off a benign corpus; OutfitGen measures it on live traffic before changing a model.

Questions people ask

Which AI model is best at changing clothes in a photo without changing the face?

In this study all three current model families, Nano Banana 2, GPT Image 2.5 Flare and Seedream 5 Pro, kept the person's face through an outfit change, including on portraits built with distinctive features like a cheek mole, glasses, grey streaks and an eyebrow scar. GPT Image 2.5 Flare changed the least around the person; Nano Banana 2 and Seedream 5 Pro were more reliable at copying a complete outfit from a reference photo. OutfitGen runs Nano Banana 2 for Pro mode and is trialling GPT Image 2.5 Flare.

Is GPT Image 2.5 better than Nano Banana 2 for photo editing?

On these eleven edits, GPT Image 2.5 Flare at medium quality matched Nano Banana 2 on keeping the person recognisable and was more precise about leaving backgrounds, lighting and expressions untouched. Nano Banana 2 was faster and better at transferring a full outfit from a reference image. Neither was clearly better overall, which is why OutfitGen is testing Flare on live traffic rather than switching outright.

Does the high quality setting on GPT Image 2.5 Flare make a difference?

Not in this study. The high setting took roughly a third longer than medium and was not visibly better on any case. For interactive photo editing the medium setting is the one worth using.

How does OutfitGen decide which AI model to use?

By running the same photos and instructions through each candidate at least three times, judging the outputs with the model names hidden, then trialling any promising model on a slice of real traffic before it takes over. Speed and how often a model refuses ordinary photos count as much as picture quality.

How was this comparison judged, and can it be reproduced?

Every model received the identical photo and instruction through the same request path OutfitGen uses in production, three times per case. The 132 outputs were laid out in contact sheets with the model columns shuffled, and the verdicts were written before the key was opened. The source photographs are generated once with fixed seeds and cached, so the same corpus can be re-run against a new model and compared against these results.

Why does a study like this not report how often each model refuses a photo?

Because a benign test corpus cannot measure it. Forty probes of ordinary swimwear, activewear and eveningwear photos drew zero refusals from all four configurations. Refusal rate only appears on real customer photos, so OutfitGen measures it on a slice of live traffic before a model change ships, and that measurement has reversed a switch before.

How should this study be cited?

This study lives at one stable URL. It is updated in place, with a version number and a changelog, so a link made today keeps resolving and a reader can tell which version was read.

Stable URL

https://outfitgen.ai/research/pro-tier-image-edit-models-2026-09

Citation

OutfitGen. (2026, September 14). Which AI image model keeps your face? Nano Banana 2 vs GPT Image 2.5 Flare vs Seedream 5 Pro. OutfitGen Research, version 1.0. https://outfitgen.ai/research/pro-tier-image-edit-models-2026-09

BibTeX

@misc{outfitgen-pro-tier-image-edit-models-2026-09,
  title  = {Which AI image model keeps your face? Nano Banana 2 vs GPT Image 2.5 Flare vs Seedream 5 Pro},
  author = {{OutfitGen}},
  year   = {2026},
  month  = {September},
  note   = {OutfitGen Research, version 1.0},
  url    = {https://outfitgen.ai/research/pro-tier-image-edit-models-2026-09}
}

Version history

  1. Version 1.0, September 14, 2026. First publication. Four model configurations, eleven cases, three runs each, 132 judged outputs, plus 40 refusal probes.

References

  • Google DeepMind. Maker of the Gemini image models, including the model known publicly as Nano Banana 2.
  • OpenAI. Maker of the GPT Image line, including GPT Image 2.5 Flare.
  • ByteDance. Maker of the Seedream image models, including Seedream 5 Pro.
  • fal. The inference platform OutfitGen calls these models through, in production and in this study.

Related

Run these edits on your own photo

OutfitGen runs these models for you. Upload a photo, describe the change, no account needed to start.

Try the AI Clothes Changer