The paper was selected as an Oral Presentation at ECCV 2026.
Overview
News
Training and inference code, training data, benchmark, and model weights were released.
DyRef was accepted by ECCV 2026.
Highlights
OmniRef-Bench
A benchmark for complex MRIG across diverse combinations of reference-image types and quantities. We find that current open-source models suffer a sharp performance drop on OmniRef-Bench as the number and diversity of reference images increase.
Superior Performance
By fine-tuning Qwen-Image-Edit-2511 on only approximately 14K training samples, our method achieves performance comparable to the closed-source Nano Banana Pro on OmniRef-Bench, while outperforming Seedream 4.5.
Generalization Across Diverse Tasks
The improvements of our method extend beyond OmniRef-Bench. It also delivers consistent gains on multi-reference image generation benchmarks, including OmniContext and MultiBanana, as well as single-image editing benchmarks, including DreamBench++ and ImgEdit.
Results Gallery
Reference Cases
Case 1
Subject + Style + BackgroundPrompt. A woman wearing a wide-brimmed hat in reference image 1 stands beside a brown yak grazing on lush green grass in reference image 2. With visual aesthetics matching reference image 3. Against the backdrop shown in reference image 4.
Case 2
Subject + Style + BackgroundPrompt. A black motorcycle helmet in image 1 resting on rich soil beside a giant radish in image 2, with a tuna fish in image 3 leaping from the ocean. With visual aesthetics matching image 4. Set against the background from image 5.
Case 3
Subject + Style + BackgroundPrompt. A heavy-duty electric drill in reference image 1 resting on a sunlit hill where a giraffe stands tall in reference image 2, a neon-colored volleyball in reference image 3 nestled in the grass nearby, and a zippered pencil case in reference image 4 placed casually beside them. Stylistically resembling reference image 5. Set against the background from reference image 6.
Case 4
Subject + PosePrompt. A shiba inu in reference image 1 stands protectively near an elderly man with glasses in reference image 2 and a woman in a yellow sweater in reference image 3, holding a Damascus steel knife in reference image 4 and a black baseball bat in reference image 5. Tension fills a dimly lit urban alleyway, cinematic composition with the dog as the focal point, muted tones with the yellow sweater as a striking accent, shallow depth of field. The elderly man adopts the body position from reference image 6.
Case 5
Subject + Pose + LightingPrompt. A young woman in reference image 1 kneeling on the grass points at a flower while a boy in reference image 2 crouching beside her listens, and an elderly man in reference image 3 bending forward, mirroring the pose in reference image 4. Make sure their whole bodies are visible. Use the lighting approach of reference image 5.
Case 6
Subject + Pose + Background + StylePrompt. A woman in image 1 crouching with a red frisbee in image 4 while a boy in image 2 kneeling beside her watches a brown dog in image 3 mid-jump trying to catch it, with visual aesthetics matching image 5. Make sure their whole bodies are visible. Set against the background from image 6. The woman mimicking the posture shown in image 7.
Method
Resources
Citation
@article{huang2026scaling,
title={Scaling Multi-Reference Image Generation with Dynamic Reward Optimization},
author={Huang, Wenwang and Fu, Yusen and Wang, Junjie and Huang, Mengfei and Li, Yulin and Liu, Gan and Cai, Jing and He, Yancheng and Tian, Zhuotao},
journal={arXiv preprint arXiv:2606.26947},
year={2026}
}