Sound
Plan the sound, design the voices, then voice it, score it and clean it — in that order, so the last step is what makes the result sound designed rather than assembled.
Plan Sound free
Takes two things: the clips, and the script. The clips say how long there is to say each line in — a shot's planned length is an intention and the generated clip is what exists — and the script says what is said. From those it:
- Casts a voice for each speaking character — a preset from the nine the model ships with, so two characters in a scene never share one — and writes a one-line voice description for each, which is what Design Voices renders.
- Fits each line to the clip it is spoken over, the reading speed being the slack that makes the words land inside it — and tells you which lines do not fit at all, which is a shot to lengthen rather than a line to rush.
- Reads each scene's mood off its own panels, which is what Generate SFX is prompted with.
It costs nothing: no GPU and no credits.
Design Voices GPU
Optional, and the step that makes a cast sound like itself. Every speaker on the cast carries a plain voice description Plan Sound wrote — open the speaker's chip on the Sound board and rewrite it: “a low, tired male voice, mid-forties, dry humour, slight rasp”. Describe timbre, age, pace and mood; never name a real person. Run Design Voices and each speaker says an audition line in the designed voice; the take is filed as that speaker's reference clip.
Design once, clone always. A description rendered twice is two different voices, so a description is never what a line is read from — the clip is. From the moment a speaker has one, Generate Voice-over clones every line of theirs from it, and a line re-spoken a week later matches the film. Re-run the step to re-roll a voice you do not like.
Or bring the voice. On the same chip, upload a 3–10 second sample of clean speech and it goes in the same slot, no design needed. Add the transcript of the sample if you know it — the clone is sharper for it. You are responsible for having the right to use a voice you upload; the Terms say so.
Generate Voice-over GPU
Qwen3-TTS reads the film's script aloud, one file per cue, in the cast's voices — a preset per speaker, or a clone of the speaker's reference clip where one exists — at the rate Plan Sound fitted to each shot. Give it the marked script; it speaks the lines, not the timestamps. A per-cue voice set on the board wins over the cast. Because a line is a file, one bad reading is one line to fix: open it and re-say it, in the same voice at the same cue.
Ten languages: English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Every preset and every clone speaks all of them.
Generate SFX GPU
MMAudio watches each generated clip and generates scene-synced effects and ambience for it — one effect track per clip — prompted by the mood Plan Sound cued for the scene. A film with dialogue and no room tone sounds unfinished.
Clean Sound GPU
A mastering pass over every audio clip of the film, not a denoiser on one of them: hiss and artefacts out, levels matched, so the voices and effects sit together cleanly. Run it last, over everything.
Where the mix is
Not here. A mix is only correct for the clips it was made from, so it lives in Deliver, next to the render that consumes it, and is made fresh each time.
Last updated 2026-09-17. Credit figures are from the studio's own reference runs; the estimate on the run button is what you are charged.