MT-OPSD Robust long-horizon image editing

On-Policy Self-Distillation for Multi‑Turn Image Editing

Editing models fall apart when they keep editing their own outputs. We teach them to stay stable by learning from their own mistakes.

1KAUST 2Krea AI
Ten turns of editing, one image Each turn edits the previous turn's output. Drag the slider or press play.
Turn 0
TL;DR

Train on the states you will actually see.

Editors are trained on clean images but, in multi-turn use, must edit their own imperfect outputs. MT-OPSD rolls the model out on itself and distills its own clean-input editing behavior into those self-generated states. It needs no multi-turn annotations, no ground-truth edits, and no stronger teacher.

Success rate at turn 10 · Qwen-Image-Edit-2511
0.03→0.44
Around 15× more ten-turn sessions end with every edit correct and consistent.
Collapse rate at turn 10 · all three backbones
≤0.61→≤0.04
Multi-turn collapse almost disappears on Qwen, FireRed and FLUX.2-klein.
Single-turn quality · ImgEdit overall
4.51→4.49
Long-horizon robustness costs almost nothing in single-turn editing quality.
Abstract

Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure to a train–test mismatch in the conditioning distribution: models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference time. To address this, we propose MT-OPSD, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations. We further introduce LME-Bench, a benchmark of 100 ten-turn editing sessions for evaluating long-horizon robustness. Experiments across three editing backbones show that MT-OPSD substantially improves long-horizon editing success and reduces multi-turn collapse while largely preserving single-turn editing quality.

The problem

Small errors, compounded.

After a few turns, outputs develop chromatic noise, fragmented structure, or identity drift. We see this in Qwen-Image-Edit, FireRed-Image-Edit and FLUX.2-klein-base, so it is a limitation of single-turn training rather than a bug in any one model.

Training clean source I→ editor→edit
Inference own output I(k−1)→ editor→I(k)↺

Try asking an editor to “make everything unchanged” ten times in a row. It should act as the identity map, but it drifts measurably at every turn. Errors that look negligible after one edit pile up once each output becomes the next input, much like exposure bias in autoregressive generation.

A train–test mismatch in the conditioning distribution.
Drift under identity editing: repeatedly asking Qwen-Image-Edit and OmniGen2 to leave the image unchanged produces growing artifacts over 10 turns.
Drift under identity editing. Asking a model again and again to leave the image unchanged exposes model-induced errors that build up across turns. Different editing models degrade in similar ways.
Method

The supervision is already inside the model.

Given a clean input, a pretrained editor already knows how to edit. MT-OPSD transfers this clean-input behavior to the degraded states the model produces for itself.

Overview of MT-OPSD: self-generated rollout states, two-branch training objective with editing and identity branches, rollout curriculum and gated teacher promotion.
Overview of MT-OPSD. The student builds self-generated rollout states by applying the identity instruction again and again. For each state, the identity branch uses the state itself as the target, and the editing branch aligns the student (conditioned on the rollout state) with a frozen teacher (conditioned on the clean source) through sparse query-based velocity matching. A rollout curriculum gradually increases rollout depth, and improved checkpoints replace the teacher through gated promotion.
01 · ROLLOUT STATES

Self-generated, error-only rollouts

The current student applies an identity instruction eid (“Make everything unchanged”) to its own output for k turns, with the same sampling configuration as inference. The content stays fixed, so Ĩ(k) differs from I(0) mainly by errors the model introduced itself. Real editing instructions here would mix those errors with intended changes.

$$\tilde I^{(k)} = G_{\theta_S}\big(\tilde I^{(k-1)}, e_{\mathrm{id}}\big),\quad \tilde I^{(0)} = I^{(0)}$$
02 · TWO-BRANCH OBJECTIVE

Stop the drift, keep the edits

Each rollout state feeds two complementary branches, sampled at a 2 : 1 ratio (editing : identity).

Editing branchThe student sees Ĩ(k); the teacher sees clean I(0). Their velocities are matched at states the student visits.
$$\mathcal L_{\text{edit}}=\big\|v_{\theta_S}(\bar x_{t_q},t_q,\tilde I^{(k)},e)-v_{\theta_T}(\bar x_{t_q},t_q,I^{(0)},e)\big\|^2$$
Identity branchUnder eid, the rollout state is its own target, so errors stop compounding without having to restore the clean image in one step.
$$\mathcal L_{\text{id}}=\big\|v_{\theta_S}(\tilde x_t,t,\tilde I^{(k)},e_{\mathrm{id}})-(\epsilon-\tilde x_0)\big\|^2$$
03 · ROLLOUT CURRICULUM

Adaptive depth, no hand-tuned schedule

Rollout depth grows by one turn only after the mean pixel drift of Ĩ(k) has stayed below a threshold for several steps in a row. Each increase pushes drift up and pauses progress until the model catches up. Depth typically settles around 4 turns, yet the gains hold through 10-turn evaluation.

04 · GATED TEACHER PROMOTION

Student → teacher → better student

As the student becomes more robust than its teacher, a VLM judge periodically evaluates student checkpoints on a held-out gate set. A checkpoint that beats the current teacher on long-horizon success and collapse becomes the new teacher. The judge only picks checkpoints and never enters the gradient.

Teacher · clean I⁽⁰⁾Student · own Ĩ⁽ᵏ⁾

Teacher and student start from the same weights. The only difference is the input, and that conditioning asymmetry is where the supervision signal comes from.

LME-Bench

A benchmark for long-horizon editing.

Existing multi-turn benchmarks stop at five turns. The Long Multi-turn Image Editing Bench has 100 sessions × 10 consecutive turns, mixing local and global edits so that later local edits land on images that have already been restyled or relit.

LME-Bench edit-type distribution: local edits (addition, attribute change, removal, replacement, state change), global soft (color temperature, atmosphere), global hard (monochrome, style).
100sessions
1,000instructions
6 + 4local + global edits / session
10image categories
Example session
localsoft globalhard global
  1. hard globalRepaint the whole image as a textured oil painting with visible brushstrokes, keeping the dog and the room.
  2. localChange the wall behind the dog to a warm sage-green color.
  3. hard globalRedraw the whole image in a bold anime illustration style with clean lines and cel shading.
  4. localAdd a rubber ball on the floor beside the dog.
  5. soft globalShift the whole image to a cool blue-gray color temperature.
  6. localMake the dog lie down on the floor instead of sitting.
  7. hard globalConvert the entire image to black and white, full grayscale with no color.
  8. localChange the wooden floor into pale marble tiles.
  9. localAdd a small potted plant in the corner behind the dog.
  10. localRemove the rubber ball from the floor.
SR@k ↑ The fraction of sessions where all of the first k turns pass both prompt-following and consistency checks.
CR@k ↓ The fraction of sessions that have collapsed by turn k, meaning two turns in a row were judged visually degraded.
Results

Stable through ten turns, on every backbone.

On LME-Bench, MT-OPSD raises SR@10 to 0.38–0.52 and cuts CR@10 to at most 0.04 on all three backbones, while training-free fixes (Emu Edit, FreqEdit, VAE-LFA) give only model-dependent gains.

+ MT-OPSD Base model Training-free baselines

Success rate SR@k · higher is better

Collapse rate CR@k · lower is better

Quantitative comparison on LME-Bench: success rate and collapse rate at turns 3, 5, 8 and 10 for proprietary models and three open-source backbones with and without MT-OPSD.
Quantitative comparison on LME-Bench. Success rate (SR) and collapse rate (CR) at turns 3, 5, 8, and 10.
Quantitative comparison on MSE-Bench and ImgEdit: per-turn success rate on MSE-Bench and single-turn ImgEdit overall score.
Quantitative comparison on MSE-Bench and ImgEdit. Success rate at each turn on MSE-Bench under its official protocol, and the overall score on the single-turn ImgEdit benchmark.
Ablation

Why the identity branch matters.

Without the identity branch, the model still follows instructions, but it over-edits and slowly rewrites content nobody asked it to touch. Scrub through the turns to watch identity and layout drift.

Ablation on key components of MT-OPSD on Qwen-Image-Edit-2511.
Ablation on key components (Qwen-Image-Edit-2511). SFT editing supervision replaces on-policy velocity matching with flow-matching toward teacher-generated images. The curriculum ablation fixes rollout depth at four.
MT-OPSD vs. w/o identity branchSame ten-turn instruction sequence.
Turn 0
Citation

BibTeX

@misc{zhao2026onpolicyselfdistillationmultiturnimage,
  title         = {On-Policy Self-Distillation for Multi-Turn Image Editing},
  author        = {Liangbing Zhao and Le Zhuo and Mohamed Elhoseiny},
  year          = {2026},
  eprint        = {2609.35611},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.35611}
}