Personal Knowledge Base

A long-term research and learning notebook for posts, notes, papers, projects, and research directions.

Skip to content
← Back to notes

Experiment Log: Conditioning a Video Diffusion Model

An early record of how affordance conditioning changes temporal consistency in generated driving scenes.

1 min read

Hypothesis

Explicit affordance maps should make the generated motion easier to control than text-only conditioning, especially around interactions between agents.

First observation

The conditioning signal improves coarse scene layout before it improves fine motion. This suggests evaluating control fidelity at multiple temporal scales instead of using a single visual similarity score.

Next run

The next experiment will separate lane geometry, agent occupancy, and motion intent so that each source of improvement can be measured independently.

Related Posts