iARCS: Iterative Agentic RL for
Controllable 3D Scene Generation

Saugat Adhikari1* Ashok Prasad Neupane1* Pramish Paudel3* Ajad Chhatkuli2,3 Danda Pani Paudel2,3

1Pulchowk Campus, IOE, Tribhuvan University   2NAAMII, Nepal   3INSAIT, Sofia University “St. Kliment Ohridski”
*Equal contribution

Motivation

A scene generator learns a broad prior over rooms from data. For a task, we want the part of that prior that satisfies the constraint, and we want to sample from it with the same diversity as the original dataset.

Prompt

“Robot grasping scene where all support surfaces are within a 1.0 m vertical reach limit”

Top-down view of a language-driven layout with furniture placed outside the room boundary

LLM / language-driven

Holodeck, SAGE, InstructScene

  • Meets the prompt
  • Learned from data

Plans each scene separately. Every new scene costs more LLM calls, layouts are not learned from real rooms, and furniture often ends up outside the room: 43.39% of objects for InstructScene.

Shown: InstructScene sample.

MiDiffusion bedroom sample with a tall wardrobe and colliding furniture

Data-driven generators

ATISS, DiffuScene, MiDiffusion

  • Meets the prompt
  • Learned from data

They learn the spread of real rooms, flaws included: 42% of objects in real 3D-FRONT bedrooms already collide. They cannot take a prompt, so hard constraints are rarely met. MiDiffusion meets this one in 4.53% of scenes.

Shown: MiDiffusion sample.

iARCS bedroom where the desk and nightstands are all below 1 meter

iARCS (ours)

Diffusion prior + agentic RL

  • Meets the prompt
  • Learned from data

Starts from the prior learned from data and fine-tunes it toward the part that meets the prompt. Samples stay diverse and no LLM runs at sampling time. 73.83% of scenes meet the prompt.

Intuition

  • Rooms the model learned from data
  • Rooms that meet the prompt
  • Base model sample
  • iARCS sample
  1. 1A diffusion model learns a broad prior from data.
  2. 2The prompt marks a small region. Few samples land there.
  3. 3iARCS fine-tunes the model toward it. Samples stay spread out.

Learn a broad prior from data, then sample from the part of it that meets the prompt.

iARCS scenes for three prompts: TV visible from bed for a farsighted person, a bedroom study zone, and support surfaces under 1 meter for robot grasping
iARCS results for three different prompts.

Abstract

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-phase strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.

Method

iARCS overview: an LLM writes a task reward, a diffusion RL loop fine-tunes the scene generator with DDPO, and a reflection loop revises the reward from training statistics
An LLM writes reward code from the prompt. DDPO fine-tunes the diffusion model on it. Every few stages the LLM looks at the reward statistics and rewrites the reward.

Phase 1

Universal rewards

RL with collision and boundary rewards fixes physical errors in the pretrained model.

Phase 2

Task rewards

The LLM breaks the prompt into geometric checks and writes Python reward code. A LoRA adapter is trained on this reward.

Loop

Reflection

Every 10 RL stages the LLM checks how the reward is doing, using past rewards and their statistics, and rewrites it.

Agentic reward synthesis for the prompt about a farsighted person viewing a 4K TV from bed: scene context, reasoning and decomposition, executable reward function
Example: turning a prompt into reward code.

Results

Setup. Trained and evaluated on 3D-FRONT bedrooms, living rooms and dining rooms, with furniture from 3D-FUTURE. Base model: MiDiffusion. RL: DDPO with LoRA. Reward LLM: Gemini 3.6 Flash. Dataset: 12,000 iARCS scenes on Hugging Face (4,000 per room).

Scene synthesis

Method Colobj ↓ Colscene ↓ Rout ↓ Rreach ↑ Rwalkable ↑ CLIP-FID ↓ SCA ↓
Ground truth43.1280.082.6771.510.851––
ATISS71.7782.3519.8665.240.8142.4682.83
MiDiffusion62.9891.4917.4577.350.8292.5077.29
InstructScene52.9588.0443.3969.220.8435.6083.97
PhyScene53.8788.1924.2970.550.8492.4878.62
iARCS (ours)50.0474.8810.9980.390.8962.4576.49

Average over bedroom, living room and dining room. All values in % except Rwalkable and CLIP-FID.

Per-room breakdown
RoomMethod Colobj ↓Colscene ↓Rout ↓ Rreach ↑Rwalkable ↑CLIP-FID ↓SCA ↓
BedroomGround truth42.0072.045.7982.800.841––
ATISS72.3953.6116.3560.270.8391.5279.24
MiDiffusion52.6781.675.8985.700.8061.4065.99
InstructScene43.4874.0737.9880.140.7985.2380.29
PhyScene45.3576.0420.3274.720.8131.4964.55
iARCS (ours)40.4564.632.9487.820.8611.6066.61
Living roomGround truth39.9379.001.4971.860.882––
ATISS67.8195.3530.1267.230.7923.0589.12
MiDiffusion64.9793.1031.9075.420.8563.2685.92
InstructScene63.2195.8346.0763.670.8855.6987.73
PhyScene60.6293.9229.8572.140.8713.3686.81
iARCS (ours)58.3365.5020.7980.170.9553.1084.12
Dining roomGround truth47.4389.200.7359.870.829––
ATISS75.1298.1013.1068.230.8122.8280.12
MiDiffusion71.3199.7014.5570.920.8242.8379.96
InstructScene52.1794.2246.1163.860.8465.8883.88
PhyScene55.6494.6222.7064.790.8642.6084.51
iARCS (ours)51.3494.509.2373.180.8732.6578.74
Side-by-side bedrooms from ATISS, MiDiffusion and iARCS with collisions, out-of-bound objects and unreachable areas marked
Red: collisions. Blue: objects out of bounds. Purple: unreachable areas.

Task-specific constraints

TaskMethodSuccess ↑SCA ↓ Colobj ↓Colscene ↓Rout ↓ Rreach ↑Rwalkable ↑
1. Robot grasping, support surfaces ≤ 1.0 m3D-FRONT*5.6797.3450.6883.722.7982.260.786
MiDiffusion4.5386.3054.2988.806.1975.400.711
iARCS73.8377.9234.7759.285.8480.370.744
2. TV visible from bed for a farsighted viewer3D-FRONT*6.0997.0238.7970.041.7186.380.810
MiDiffusion5.2289.3456.1078.985.7879.320.733
iARCS66.0285.7046.4677.153.4586.770.790
3. Bedroom with a functional study zone3D-FRONT*1.0497.5856.8688.673.3375.170.728
MiDiffusion0.8785.4072.0098.525.1069.320.704
iARCS33.3977.9752.3879.394.5974.750.738

3D-FRONT*: real scenes that already meet the task (for reference). MiDiffusion: the base model without RL. Bold: iARCS beats the base model.

Bar chart of human preference: iARCS preferred 71%, 63% and 52% on tasks 1 to 3, 61% overall
Human preference: 21 people chose between iARCS and real 3D-FRONT scenes for each task.

iARCS data helps other models

Model (training data) Colobj ↓ Colscene ↓ Rout ↓ Rreach ↑ Rwalkable ↑ CLIP-FID ↓ SCA ↓
ATISS (3D-FRONT)72.3953.6116.3560.270.8391.5279.24
ATISS (3D-FRONT + MiDiffusion)71.5451.3917.8260.920.8391.5478.82
ATISS (3D-FRONT + iARCS)69.6050.2813.2164.590.8621.5278.07
MiDiffusion (3D-FRONT)52.6781.675.8985.700.8061.3465.99
MiDiffusion (3D-FRONT + MiDiffusion)48.2675.653.7890.000.8131.3766.86
MiDiffusion (3D-FRONT + iARCS)41.4963.613.1292.520.8271.3466.10

Each model is retrained with 4,000 extra scenes, either from plain MiDiffusion or from iARCS. iARCS scenes help more.

Bar chart comparing MiDiffusion trained on 3D-FRONT with MiDiffusion trained on 3D-FRONT plus iARCS data
MiDiffusion with and without iARCS data.
Bedroom from MiDiffusion trained on 3D-FRONT next to one from MiDiffusion trained with iARCS augmentation
Fewer collisions and more walkable space with iARCS data.

Does the reflection loop matter?

MethodTask 1Task 2Task 3
3D-FRONT*5.676.091.04
MiDiffusion4.535.220.87
Initial reward, frozen28.4 ± 5.122.7 ± 4.39.6 ± 2.7
Final reward, frozen at step 069.1 ± 4.661.8 ± 5.232.2 ± 4.1
Iterative, no memory35.3 ± 6.841.4 ± 7.112.8 ± 5.6
iARCS (full)73.83 ± 3.266.02 ± 3.833.39 ± 3.5

Task success (%), mean ± std over 3 seeds. Without the reflection loop or the memory, success drops sharply.

Why two phases?

Line plots over RL stages: two-phase training reaches higher task success with lower out-of-bound and collision rates than one-phase training on three tasks
Fixing physical errors first (2-phase) gives higher task success and fewer collisions than training on the task right away (1-phase).

More examples

Additional bedroom comparison between ATISS, MiDiffusion and iARCS with collision and walkability scores
Same floor plan, three methods.

BibTeX

@article{adhikari2026iarcs,
  title   = {iARCS: Iterative Agentic RL for Controllable 3D Scene Generation},
  author  = {Adhikari, Saugat and Neupane, Ashok Prasad and Paudel, Pramish
             and Chhatkuli, Ajad and Paudel, Danda Pani},
  journal = {arXiv preprint arXiv:2608.06161},
  year    = {2026}
}