Key facts
TL;DR: iADD fine-tunes diffusion models with reinforcement learning while keeping diversity: it updates a sparse, incrementally chosen set of timesteps and uses Feynman-Kac branch/resample during training, avoiding the mode collapse and reward hacking of DDPO-style methods.
- Task: RL fine-tuning / reward alignment of diffusion models (diffusion policy optimization, RLHF for diffusion).
- Method: sparse incremental timestep updates + training-time discrete Feynman-Kac (FK) branching; no preference data and no test-time scaling.
- Rarity (images): 66.85% vs 24.81% for DDPO.
- With GRPO: iADD+GRPO raises CLIPScore 0.3886 → 0.4126 and rarity 61.67% → 88.33% over DanceGRPO.
- Domains: Stable Diffusion v1.4 + LoRA (CLIPScore), vanishing-point correction, 3D indoor scene synthesis (MiDiffusion, 3D-FRONT).
- Paper: arXiv:2610.01789
Motivation
Higher reward without giving up diversity.
RL fine-tuning of diffusion models (DDPO and its variants) tends to trade diversity and visual quality for reward. iADD reaches a better point on this tradeoff curve, and holds up on rare, hard-to-align prompts where the pretrained model's probability mass is thin.
Abstract
Reinforcement learning based post-training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimization do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that only-latter-timestep updates of a diffusion model may be harmful for diversity, contrary to the conclusions presented in previous work. Additionally, we propose an incremental Feynman–Kac training procedure, grounded in strong theoretical foundations, to achieve the best-yet alignment–diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches on three different tasks, and provide strong ablations for each component, validating performance gains in both alignment and diversity.
Method
Theory
Why sparse, early timesteps
For a fixed budget of optimized timesteps, we show uniform (early-inclusive) timestep sampling causes a strictly larger output perturbation than a late-only strategy (Generative Reach), and that sparse updates better preserve the prior's probability volume than updating all timesteps (Volume Preservation). Together these motivate an incremental curriculum: start with a sparse set of early, high-leverage timesteps and grow the active set as training stabilizes.
Component 1
Incremental timestep update
Each training stage samples and updates a randomly selected subset of denoising timesteps, with the subset size increasing across stages — from a handful of high-leverage steps to nearly the full trajectory.
Component 2
Feynman–Kac path guidance
A discrete-time Feynman–Kac particle filter branches and resamples partial denoising trajectories during training (not just at inference), using reward-derived potentials to prune unpromising paths and branch promising ones — improving sample efficiency and rare-event discovery.
Interactive: FK branch & resample
Interactive: sparse incremental curriculum
Video
Results
Setup. We validate iADD on three tasks: text-to-image generation (Stable Diffusion v1.4 + LoRA, reward = CLIPScore), vanishing-point correction in generated images (reward = geometric consensus of converging line segments), and 3D indoor scene generation (continuous-domain MiDiffusion on 3D-FRONT, reward = spatial / collision constraints). Baselines: the frozen pretrained model, DDPO, and B²-DiffuRL (updates only late timesteps). We also compare against recent inference-time guidance (AIG) and GRPO-style methods (DanceGRPO, BranchGRPO), and test an iADD+GRPO hybrid.
Alignment, diversity and rarity across tasks
| Method | Image generation | 3D indoor scene | Vanishing point | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Reward ↑ | IS ↑ | Rarity ↑ | AUC ↑ | Collision ↓ | CKL ↓ | Reward ↑ | IS ↑ | Rarity ↑ | |
| Pretrained (SD) | 0.3272 | 1.3514 | – | – | 53.87% | 30.48% | 0.6105 | 1.2735 | – |
| DDPO | 0.3417 | 1.2786 | 24.81% | 53.8% | 63.39% | 7.726% | 0.6815 | 1.2600 | 21.33% |
| B²-DiffuRL | 0.3452 | 1.2728 | 24.22% | 58.7% | 64.06% | 6.278% | 0.6486 | 1.2550 | 31.11% |
| iADD (ours) | 0.3553 | 1.3241 | 66.85% | 77.7% | 59.20% | 6.13% | 0.7763 | 1.2881 | 95.11% |
Each method evaluated at its best reward–diversity operating point. iADD gets the strongest reward–diversity AUC and rarity across all three tasks while matching or beating DDPO / B²-DiffuRL on raw reward.
Comparison with GRPO and inference-time guidance (Template 1: animal activities)
| Method | CLIP ↑ | IS ↑ | Rarity ↑ | AUC ↑ | LPIPS ↑ |
|---|---|---|---|---|---|
| AIG (inference-time) | 0.3526 | 1.3147 | 17.50% | – | 0.7531 |
| iADD | 0.3830 | 1.3035 | 78.33% | 83.7% | 0.7693 |
| DanceGRPO | 0.3886 | 1.2750 | 61.67% | 39.3% | 0.7624 |
| BranchGRPO | 0.3854 | 1.2917 | 56.67% | 43.73% | 0.7447 |
| iADD+GRPO | 0.4126 | 1.2942 | 88.33% | 73.3% | 0.7778 |
iADD's trajectory guidance and sparse curriculum are complementary to the underlying policy optimizer: combining them with GRPO raises CLIPScore and rarity further while improving LPIPS diversity.
More from the paper (supplementary)
Volume preservation, measured
Log-determinant of the full Jacobian at matched compute (stage 27); less negative is better (paired difference 1605 nats, 437 SEM, 24 prompt–seed pairs).
| Update | log |det JT→0| | SEM |
|---|---|---|
| Dense | −27,568 | 658 |
| Sparse | −25,963 | 584 |
KL regularization (DPOK-style)
Diversity at fixed reward 0.38; reward at fixed diversity 1.285.
| Method | Reward | Div. | Rarity |
|---|---|---|---|
| SD | 0.3565 | 1.3067 | – |
| DPOK | 0.3582 | 1.2855 | 28.33 |
| B²Diff-RL + KL | 0.3646 | 1.2784 | 27.5 |
| Ours + KL | 0.3800 | 1.3002 | 55.83 |
Training cost
GPU hours to reach the same mean reward (CLIP 0.388).
| Method | GPU hours |
|---|---|
| DDPO | 2.55 |
| DDPO + incremental | 2.46 |
| B²-DiffuRL | 3.42 |
| B²-DiffuRL + incremental | 2.76 |
| FK + incremental | 9.55 |
| Branch-FK + incremental (iADD) | 7.43 |
FK scale λFK
Effect on Inception Score; the paper uses λFK = 2 with the max potential.
| Method | IS |
|---|---|
| λFK = 10 | 1.2830 |
| λFK = 2 | 1.3035 |
Qualitative results
FAQ: RL fine-tuning of diffusion models without losing diversity
iADD is a diffusion policy optimization method for fine-tuning diffusion models with reinforcement learning (RLHF for diffusion). It builds on DDPO, DPOK and B2-DiffuRL, brings Feynman-Kac (FK) steering / sequential Monte Carlo ideas into training, and combines with GRPO methods such as DanceGRPO and BranchGRPO.
What is iADD?
iADD is a reinforcement learning fine-tuning method for diffusion models that combines an incremental, sparse timestep-update curriculum with discrete-time Feynman-Kac branching during training, to improve alignment and diversity together.
Why does RL fine-tuning of diffusion models cause mode collapse?
Optimizing a reward over all denoising steps pushes the model away from its prior and toward a few high-reward modes. iADD's sparse updates preserve more of the prior's probability volume (log-det -25,963 sparse vs -27,568 dense).
Why not update only late timesteps like B2-DiffuRL?
iADD's Proposition 1 shows uniform timestep sampling gives more generative reach than late-only updates; in ablations, updating any N timesteps beats updating the last N on Inception Score.
How is iADD different from Feynman-Kac (FK) steering?
FK steering applies Feynman-Kac guidance at inference time. iADD uses it during training: low-potential branches are pruned for alignment and promising ones are branched for diversity, so no test-time scaling is needed.
DDPO vs iADD: what changes?
DDPO fine-tunes the full denoising trajectory. iADD updates sparse, incrementally chosen timesteps and adds training-time FK branch/resample, reaching 66.85% rarity vs 24.81% for DDPO on images.
Does iADD work with GRPO (DanceGRPO)?
Yes. iADD+GRPO raises CLIPScore from 0.3886 (DanceGRPO) to 0.4126 and rarity from 61.67% to 88.33%.
Does iADD reduce reward hacking?
In 3D scenes it has fewer collisions at matched reward than DDPO and B2-DiffuRL; on images, baselines drift to cartoonish, over-saturated outputs.
Does iADD need preference data or extra inference compute?
No. It optimizes a reward function (no human preference pairs, unlike Diffusion-DPO) and adds no test-time scaling.
What tasks was iADD tested on?
Text-to-image generation with Stable Diffusion v1.4 + LoRA and a CLIPScore reward, vanishing-point correction, and 3D indoor scene synthesis (MiDiffusion on 3D-FRONT).
BibTeX
@misc{neupane2026iadd,
title = {iADD: Improving Alignment and Diversity in Diffusion Policy Optimization},
author = {Neupane, Ashok Prasad and Adhikari, Saugat and Paudel, Pramish
and Chhatkuli, Ajad and Paudel, Danda Pani},
year = {2026},
eprint = {2610.01789},
archivePrefix = {arXiv},
primaryClass = {cs.LG}
}