A song's loudness draws a shape, and Wan 2.1 is only allowed to paint inside it. Chained 66 frames at a time, it grows flowers to music for as long as the track lasts. The model never hears a note.
No audio conditioning, no music embedding, no fine-tune. The track is analyzed once with
librosa — RMS loudness resampled to 30 fps — and that curve drives the radius of a
white circle on black. That video becomes the VACE mask: Wan generates only where it's white.
Displacing the circle's edge with Perlin noise makes it writhe instead of pulse — same signal, far
more organic once petals fill it in.
“The music draws a shape, and the model is only allowed to paint inside the shape. Everything blooming and closing — that's just loudness, turned into permission.”
Wan's trained context is 81 frames — about 2.7 seconds. The extend graph stitches past it: every
pass carries the last 15 frames of the previous render in front of 66 new ones, and forces the
mask black over those 15, so VACE treats them as fixed context and regenerates only the new frames.
A SimpleMath node seeks the mask video to 81 + (n−1)·66,
which keeps the audio locked to wall-clock time across any number of passes.
The whole thing is a single ComfyUI graph — 109 nodes across nine groups
(Loaders, Init, Extend, and the rest). Everything that
matters is a handful of numbers; the mask and the seek math do the work.
| control | value | what it sets |
|---|---|---|
| Frame Count | 81 | the init render — exactly Wan's trained context |
| Frame overlap | 15 | context frames carried into each following pass |
| per sequence | (a − 1) · 66 | new frames per pass, and the mask seek offset |
| W × H | 720 × 720 | 960² for the render above — same graph, one flag |
| WanVideoSampler | 6 steps · cfg 1.0 | shift 10, euler — bought by the LoRA stack |
| WanVideoLoraSelect | ×4 | lightx2v step/cfg distill, CausVid, detailz, sh4rpn3ss |
| VideoCombine | 30 fps · crf 18 | h264 out, one segment per pass |