One project, six stages
Source, Art, Story, Video, Sound, Deliver. Every stage reads what the one before it made, so nothing is re-described and nothing drifts.
Shotink is a production studio, not a prompt box. Nineteen steps across six stages take a comic or manga page from ink to a finished, voiced, scored anime scene — and every one of them can be run on its own, fixed by hand, and run again.
Source, Art, Story, Video, Sound, Deliver. Every stage reads what the one before it made, so nothing is re-described and nothing drifts.
Each stage opens with a planning pass that reads your pages and writes its notes down. The generation steps read those notes instead of guessing.
Paint a region and describe the repair, crop a panel, clone a background, re-say a line, re-run a shot. The result replaces the file in place, undoable.
Heavy steps rent a dedicated GPU for your job, with your model cache on a private volume. You see the estimate before you run and pay for the minutes used.
Steps are listed in the order you would normally run them. GPU steps rent a card for your job; AI steps run on the server against a language model and cost a few credits; free steps cost nothing.
Bring the comic in and break each page down into single panels, with the speech bubbles removed.
Drop in pages as images, a folder or a zip. Pages are kept in reading order and nothing is altered.
Reads each whole page: transcribes every bubble, caption and sound effect, describes the scene, and reads the gutters — the action the reader infers between panels. Afterwards the assistant can discuss your comic from these notes instead of re-reading the images.
A trained ensemble of panel detectors crops every panel from the page in reading order, combining several models so a border one misses another catches.
A second ensemble finds every balloon and knocks it out to transparency, leaving a see-through hole where the lettering was. Paint over any it missed.
Repair the panels, repaint them in an anime style, and shape them ready for video.
Looks at each panel as a picture: who is in it, how it is framed, its own light and colour, and what has to be repaired before it can be a frame of film. It also decides whether a panel is a frame at all — a sliver of border or a sound effect over blank ground is marked not animatable, and every later step skips it. One click overrules it.
Flux.1 Fill paints artwork back into the holes the bubbles left — one hole at a time, each from a crop around it and its own sentence from the panel notes, so what belongs behind one hole cannot end up in another.
Restyles every panel into one of 60+ looks — studio anime, watercolour, painted, ink, stop-motion, classic manga masters — on Flux.2 Klein. The style is how the picture is drawn; each panel's own light and mood ride separately, so a cheerful page stays cheerful in a dark-action style. It is also the last step that can remove a border or a drawn-in sound effect.
Extends each plate outward until it is exactly the canvas the video model generates at, so every shot starts from the same aspect ratio. It only ever adds a margin: your plate is composited back over the result, so the picture itself cannot change.
Read the pages as a story, then write the one script the film is made from — a narration or a dialogue, both cued.
Reads your imported pages — the originals, where the reading order and the whole moment are — and writes one story analysis each: what the page does, and every line said on it with a description of who says it. Descriptions survive restyling where a name the comic never printed would not.
Arranges what Plan Story read into one .srt you can open and edit. Narration is one voice; dialogue has the characters speaking, each cue naming who. Switch a finished script between the two without touching its timings. With nothing analysed it writes from your prompt; with nothing at all it hands you a timed outline to write over.
Cut a film out of the finished plates: plan the shots against the story, then generate them.
The only step that reads every plate at once. It decides which panels are worth a shot and drops the rest with a reason, groups consecutive panels that are one continuous moment into a single shot, fixes the film's lighting, palette and cast once so a jacket cannot change colour between cuts, and — where you wrote a script — sizes each shot to the line spoken over it. Everything lands on a storyboard you can edit.
MiniMax H3 renders each shot as one continuous take, with your plate pinned as the first frame so the shot starts exactly on your art. Every plate in a shot goes to the model as reference art; the storyboard's direction does the rest, including the gutter — the action the comic never drew. Re-run one shot with a better description and it replaces the clip in place.
Plan the sound, design the voices, then voice it, score it and clean it — in that order, so the last step makes it sound designed rather than assembled.
Takes the clips and the script and casts a voice for every speaking character — the same voice all film, and never two characters sharing one in a scene. Fits each line to the clip it is spoken over and flags the ones that do not fit, which is a shot to lengthen rather than a line to rush. Cues a mood per scene for the effects pass.
Every speaker gets a one-line voice description from the plan — rewrite it in your own words — and this renders each one into a reference clip: an audition take in that voice. Or upload a sample of a voice you have the right to use. From then on the reading is cloned from the clip, so a character sounds the same across the film and across re-takes.
Qwen3-TTS reads the script aloud, one file per cue: a preset voice per speaker, or a clone of the speaker's reference clip, at the rate the plan fitted to each shot. A line is a file, so a bad reading is fixed by saying that one line again.
MMAudio watches each clip and lays scene-synced ambience and impact sound beneath it, prompted by the mood Plan Sound cued. A film with dialogue and no room tone sounds unfinished.
A mastering pass over every audio clip of the film: hiss and artefacts out, levels matched, so voices and effects sit together cleanly.
Combine everything into the one file you take away. Nothing is baked in until you ask for it, so a clip or a track changed upstream is picked up by the next render.
Lay the generated clips out in order with their transitions. The head and tail of every clip are trimmed automatically, and the cut is laid fresh each time from whatever the clips are now.
Every voice-over at its cue and every effect under the shot it was made for, balanced and loudness-normalised to a broadcast target. Play it against the picture before you render.
Joins picture and soundtrack into one finished video, burns in the film's own script as subtitles if you ask, and optionally upscales the finished film with ESRGAN — once, on the final cut, rather than on twenty separate clips. Export the film alone, or the whole project as a zip with a manifest.
A tool appears on the rail of the step whose mistakes it fixes, and nowhere else — so the rail is short enough to be read.
The cursor back to normal, and the way to pick a region for the other tools.
One brush whose meaning is decided by where you use it: on a panel it knocks out a balloon the detector missed; on a plate it marks the region to redraw.
Cut a border or a neighbour out of a panel the splitter took too much of.
Paint over what came out wrong and say what should be there in plain words. Two models to choose from — an instruction editor that can hold “give the hand five fingers”, and an inpainter for pure repairs.
Copy a region of the image over another. When the right answer is “more of the wall six inches to the left”, copying beats imagining. No GPU, no queue.
One panel, one look, in place — before committing a folder of forty-nine to it. It shares the picker with Apply Anime Style.
Pages, panels and plates with the tool rail for that step, a per-file edit history you can walk back, and the file's own notes beside it.
The film's .srt as cues you can rewrite, retime and reorder. Switch narration to dialogue or back; mark the script the film is planned against.
Watch a shot, describe what is wrong, and render it again. Trim the head and tail. The new take replaces the old in place, undoable.
One file per line, so one bad reading is one line to say again — in the same voice, at the same cue.
Chat about the comic from the notes the planning passes wrote — who is in a panel, what a page does, why a shot was dropped. It can read any file's insights and propose the next step; it never spends a GPU without you.
Point it at a stage, or a run of stages, and it works through the steps end to end: inspect, choose the next step, run it, judge it, repeat. Sample three panels before committing two hundred, park at each stage boundary, stop at any time. Every guard is enforced in code, not in a prompt.
Nothing here is a black box: every step names its model and its settings, and you can change either.
Create an account, import a chapter, and run the planning passes for a few credits before you spend a minute of GPU.