iADD: Improving Alignment and Diversity
in Diffusion Policy Optimization

Ashok Prasad Neupane1* Saugat Adhikari1* Pramish Paudel3* Ajad Chhatkuli2,3 Danda Pani Paudel2,3

1Pulchowk Campus, IOE, Tribhuvan University, Lalitpur, Nepal   2NAAMII, Kathmandu, Nepal   3INSAIT, Sofia University “St. Kliment Ohridski”, Bulgaria
*Equal contribution

Key facts

TL;DR: iADD fine-tunes diffusion models with reinforcement learning while keeping diversity: it updates a sparse, incrementally chosen set of timesteps and uses Feynman-Kac branch/resample during training, avoiding the mode collapse and reward hacking of DDPO-style methods.

  • Task: RL fine-tuning / reward alignment of diffusion models (diffusion policy optimization, RLHF for diffusion).
  • Method: sparse incremental timestep updates + training-time discrete Feynman-Kac (FK) branching; no preference data and no test-time scaling.
  • Rarity (images): 66.85% vs 24.81% for DDPO.
  • With GRPO: iADD+GRPO raises CLIPScore 0.3886 → 0.4126 and rarity 61.67% → 88.33% over DanceGRPO.
  • Domains: Stable Diffusion v1.4 + LoRA (CLIPScore), vanishing-point correction, 3D indoor scene synthesis (MiDiffusion, 3D-FRONT).
  • Paper: arXiv:2610.01789

Motivation

Higher reward without giving up diversity.

RL fine-tuning of diffusion models (DDPO and its variants) tends to trade diversity and visual quality for reward. iADD reaches a better point on this tradeoff curve, and holds up on rare, hard-to-align prompts where the pretrained model's probability mass is thin.

Alignment-diversity trade-off curves: Inception Score vs CLIP Score for DDPO, B2-DiffuRL and iADD
(a) iADD reaches a higher CLIP Score at any given diversity (Inception Score) level than DDPO and B²-DiffuRL, which become increasingly cartoonish and over-saturated as reward increases.
Bar chart comparing alignment performance on common and rare prompts across methods at constant Inception Score
(b) Alignment on common vs. rare prompts at constant diversity. Rare prompts are conditions that inherently yield low-reward outcomes under the pretrained model.

Abstract

Reinforcement learning based post-training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimization do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that only-latter-timestep updates of a diffusion model may be harmful for diversity, contrary to the conclusions presented in previous work. Additionally, we propose an incremental Feynman–Kac training procedure, grounded in strong theoretical foundations, to achieve the best-yet alignment–diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches on three different tasks, and provide strong ablations for each component, validating performance gains in both alignment and diversity.

Method

iADD method overview: Feynman-Kac particle sampling branches high-potential trajectories and prunes low-potential ones from Gaussian noise to the final sample, alongside a progressive incremental timestep update schedule used during RL training
Overview of FK sampling and incremental timestep training. (A) FK sampling: starting from Gaussian noise at t = T, partial trajectories are evaluated at discrete checkpoints based on their potentials. FK branching duplicates high-potential particles and prunes low-potential ones, redirecting stochastic denoising paths toward higher expected reward. (B) Incremental timestep updates: training uses a progressive, stage-wise curriculum in which the number of updated denoising timesteps grows over the course of fine-tuning.

Theory

Why sparse, early timesteps

For a fixed budget of optimized timesteps, we show uniform (early-inclusive) timestep sampling causes a strictly larger output perturbation than a late-only strategy (Generative Reach), and that sparse updates better preserve the prior's probability volume than updating all timesteps (Volume Preservation). Together these motivate an incremental curriculum: start with a sparse set of early, high-leverage timesteps and grow the active set as training stabilizes.

Component 1

Incremental timestep update

Each training stage samples and updates a randomly selected subset of denoising timesteps, with the subset size increasing across stages — from a handful of high-leverage steps to nearly the full trajectory.

Component 2

Feynman–Kac path guidance

A discrete-time Feynman–Kac particle filter branches and resamples partial denoising trajectories during training (not just at inference), using reward-derived potentials to prune unpromising paths and branch promising ones — improving sample efficiency and rare-event discovery.

Interactive: FK branch & resample

k particles branch from one prefix; low-weight ones are resampled away.

This diagram is interactive and needs JavaScript. See panel (A) of the overview figure above for the static version.

Feynman-Kac branch and resample diagram A denoising trajectory from noise at the left to the sample at the right. One shared prefix branches into k particles at the 5th step; at steps 8, 12 and 16 particles are scored and low-weight ones are replaced by copies of high-weight ones.
1 Shared prefix 2 Branch at tb 3 Score drafts 4 Resample 5 Best vs worst → training pair

shared prefix active particle pruned best worst resample point

Schematic. See Sec. 4.3, Fig. 3(A).

Interactive: sparse incremental curriculum

Only a growing subset of timesteps gets gradient updates; early steps matter most.

This widget is interactive and needs JavaScript. In the paper, the number of updated timesteps grows over four stages: 5, 10, 15 and 20 of T = 20 (see panel (B) of the overview figure).

Incremental timestep curriculum A strip of 20 denoising timesteps from t=T on the left to t=1 on the right. Highlighted timesteps receive gradient updates at the selected stage. A decaying shaded area behind the strip marks the schematic generative reach of each timestep, highest for early steps.
Any-N random steps, early ones included. Late-only last N steps only (B²-DiffuRL).

updated this stage frozen generative reach (schematic)

Schematic. See Sec. 4.2, Fig. 3(B).

Video

A walkthrough of iADD: the alignment-diversity tradeoff problem, the incremental timestep curriculum, and Feynman–Kac branch/resample sampling.

Results

Setup. We validate iADD on three tasks: text-to-image generation (Stable Diffusion v1.4 + LoRA, reward = CLIPScore), vanishing-point correction in generated images (reward = geometric consensus of converging line segments), and 3D indoor scene generation (continuous-domain MiDiffusion on 3D-FRONT, reward = spatial / collision constraints). Baselines: the frozen pretrained model, DDPO, and B²-DiffuRL (updates only late timesteps). We also compare against recent inference-time guidance (AIG) and GRPO-style methods (DanceGRPO, BranchGRPO), and test an iADD+GRPO hybrid.

Alignment, diversity and rarity across tasks

Method Image generation 3D indoor scene Vanishing point
Reward ↑IS ↑Rarity ↑AUC ↑ Collision ↓CKL ↓ Reward ↑IS ↑Rarity ↑
Pretrained (SD)0.32721.3514––53.87%30.48%0.61051.2735–
DDPO0.34171.278624.81%53.8%63.39%7.726%0.68151.260021.33%
B²-DiffuRL0.34521.272824.22%58.7%64.06%6.278%0.64861.255031.11%
iADD (ours)0.35531.324166.85%77.7%59.20%6.13%0.77631.288195.11%

Each method evaluated at its best reward–diversity operating point. iADD gets the strongest reward–diversity AUC and rarity across all three tasks while matching or beating DDPO / B²-DiffuRL on raw reward.

Comparison with GRPO and inference-time guidance (Template 1: animal activities)
MethodCLIP ↑IS ↑Rarity ↑AUC ↑LPIPS ↑
AIG (inference-time)0.35261.314717.50%–0.7531
iADD0.38301.303578.33%83.7%0.7693
DanceGRPO0.38861.275061.67%39.3%0.7624
BranchGRPO0.38541.291756.67%43.73%0.7447
iADD+GRPO0.41261.294288.33%73.3%0.7778

iADD's trajectory guidance and sparse curriculum are complementary to the underlying policy optimizer: combining them with GRPO raises CLIPScore and rarity further while improving LPIPS diversity.

More from the paper (supplementary)

Volume preservation, measured

Log-determinant of the full Jacobian at matched compute (stage 27); less negative is better (paired difference 1605 nats, 437 SEM, 24 prompt–seed pairs).

Updatelog |det JT→0|SEM
Dense−27,568658
Sparse−25,963584

KL regularization (DPOK-style)

Diversity at fixed reward 0.38; reward at fixed diversity 1.285.

MethodRewardDiv.Rarity
SD0.35651.3067–
DPOK0.35821.285528.33
B²Diff-RL + KL0.36461.278427.5
Ours + KL0.38001.300255.83

Training cost

GPU hours to reach the same mean reward (CLIP 0.388).

MethodGPU hours
DDPO2.55
DDPO + incremental2.46
B²-DiffuRL3.42
B²-DiffuRL + incremental2.76
FK + incremental9.55
Branch-FK + incremental (iADD)7.43

FK scale λFK

Effect on Inception Score; the paper uses λFK = 2 with the max potential.

MethodIS
λFK = 101.2830
λFK = 21.3035

Qualitative results

Qualitative comparison on common and rare image prompts across Stable Diffusion, DDPO, B2-DiffuRL and iADD
Common and rare prompts. On rare prompts, DDPO and B²-DiffuRL either miss the prompt or look cartoonish; iADD keeps both semantic alignment and visual quality.
3D scene synthesis qualitative comparison: collision reward and spatial reward (TV stand facing bed) across DDPO, B2-DiffuRL and iADD
3D scene synthesis. DDPO and B²-DiffuRL satisfy the target constraint but introduce collisions; iADD avoids them.
Vanishing point correction qualitative comparison with red off-vanishing intersections and green vanishing points across methods
Vanishing-point correction. Red: off-vanishing intersections. Green: the vanishing point. iADD converges perspective lines while preserving texture.

FAQ: RL fine-tuning of diffusion models without losing diversity

iADD is a diffusion policy optimization method for fine-tuning diffusion models with reinforcement learning (RLHF for diffusion). It builds on DDPO, DPOK and B2-DiffuRL, brings Feynman-Kac (FK) steering / sequential Monte Carlo ideas into training, and combines with GRPO methods such as DanceGRPO and BranchGRPO.

What is iADD?

iADD is a reinforcement learning fine-tuning method for diffusion models that combines an incremental, sparse timestep-update curriculum with discrete-time Feynman-Kac branching during training, to improve alignment and diversity together.

Why does RL fine-tuning of diffusion models cause mode collapse?

Optimizing a reward over all denoising steps pushes the model away from its prior and toward a few high-reward modes. iADD's sparse updates preserve more of the prior's probability volume (log-det -25,963 sparse vs -27,568 dense).

Why not update only late timesteps like B2-DiffuRL?

iADD's Proposition 1 shows uniform timestep sampling gives more generative reach than late-only updates; in ablations, updating any N timesteps beats updating the last N on Inception Score.

How is iADD different from Feynman-Kac (FK) steering?

FK steering applies Feynman-Kac guidance at inference time. iADD uses it during training: low-potential branches are pruned for alignment and promising ones are branched for diversity, so no test-time scaling is needed.

DDPO vs iADD: what changes?

DDPO fine-tunes the full denoising trajectory. iADD updates sparse, incrementally chosen timesteps and adds training-time FK branch/resample, reaching 66.85% rarity vs 24.81% for DDPO on images.

Does iADD work with GRPO (DanceGRPO)?

Yes. iADD+GRPO raises CLIPScore from 0.3886 (DanceGRPO) to 0.4126 and rarity from 61.67% to 88.33%.

Does iADD reduce reward hacking?

In 3D scenes it has fewer collisions at matched reward than DDPO and B2-DiffuRL; on images, baselines drift to cartoonish, over-saturated outputs.

Does iADD need preference data or extra inference compute?

No. It optimizes a reward function (no human preference pairs, unlike Diffusion-DPO) and adds no test-time scaling.

What tasks was iADD tested on?

Text-to-image generation with Stable Diffusion v1.4 + LoRA and a CLIPScore reward, vanishing-point correction, and 3D indoor scene synthesis (MiDiffusion on 3D-FRONT).

BibTeX

@misc{neupane2026iadd,
  title   = {iADD: Improving Alignment and Diversity in Diffusion Policy Optimization},
  author  = {Neupane, Ashok Prasad and Adhikari, Saugat and Paudel, Pramish
             and Chhatkuli, Ajad and Paudel, Danda Pani},
  year    = {2026},
  eprint  = {2610.01789},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG}
}