Wan Flower
Wan Flower · a ComfyUI workflow walkthrough

The audio is the mask.

A song's loudness draws a shape, and Wan 2.1 is only allowed to paint inside it. Chained 66 frames at a time, it grows flowers to music for as long as the track lasts. The model never hears a note.

The whole demo — play it with sound. Streaming the original file, uncompressed. 2,325 frames: 77 s at 960², 34 passes chained end to end, against the track that shaped every bloom. Wan 2.1 T2V-14B + VACE · 6 steps · 81 frames in 188 s per pass on an RTX PRO 6000.
The trick

Audio never touches the model

No audio conditioning, no music embedding, no fine-tune. The track is analyzed once with librosa — RMS loudness resampled to 30 fps — and that curve drives the radius of a white circle on black. That video becomes the VACE mask: Wan generates only where it's white. Displacing the circle's edge with Perlin noise makes it writhe instead of pulse — same signal, far more organic once petals fill it in.

“The music draws a shape, and the model is only allowed to paint inside the shape. Everything blooming and closing — that's just loudness, turned into permission.”

track.mp3 librosa RMS 30 fps mask.mp4 radius ∝ loudness + Perlin edge WanVideo VACE Encode white = paint black = keep WanVideoSampler T2V-14B · 6 steps cfg 1.0 · 4 LoRAs 81 f one prompt — the only text in the graph
Loudness → geometry → permission-to-generate. Swap the prompt and the same mask grows flames, water, anything.
Plain circle mask frame
plain amplitude discmask_circle
Perlin-distorted mask frame
perlin-displaced edge…_distorted
The chain

Going longer than the model's context

Wan's trained context is 81 frames — about 2.7 seconds. The extend graph stitches past it: every pass carries the last 15 frames of the previous render in front of 66 new ones, and forces the mask black over those 15, so VACE treats them as fixed context and regenerates only the new frames. A SimpleMath node seeks the mask video to 81 + (n−1)·66, which keeps the audio locked to wall-clock time across any number of passes.

init 81 f pass 1 66 new pass 2 66 new pass n … 66 new last 15 f carried over — mask forced black there, so VACE treats them as fixed mask.mp4 — seek to 81 + (n−1)·66 so the audio stays wall-clock locked
Each pass costs 66 fresh frames. The only bridge between passes is 15 frames of pixels.
The workflow

109 nodes, one prompt

The whole thing is a single ComfyUI graph — 109 nodes across nine groups (Loaders, Init, Extend, and the rest). Everything that matters is a handful of numbers; the mask and the seek math do the work.

controlvaluewhat it sets
Frame Count81 the init render — exactly Wan's trained context
Frame overlap15 context frames carried into each following pass
per sequence(a − 1) · 66 new frames per pass, and the mask seek offset
W × H720 × 720 960² for the render above — same graph, one flag
WanVideoSampler6 steps · cfg 1.0 shift 10, euler — bought by the LoRA stack
WanVideoLoraSelect×4 lightx2v step/cfg distill, CausVid, detailz, sh4rpn3ss
VideoCombine30 fps · crf 18 h264 out, one segment per pass
Download the ComfyUI graph 124 KB JSON · drop it on the canvas