Pattern chooser
Every tried-and-tested recipe in the engine, as a card. Filter by lane, stage, tier or capability to find the path for the work in front of you, then add the ones you want to your plan. Jump in at any stage; nothing here is a fixed order.
Lyric analysis pass
Turn lyrics into themes, mood, story seed and 8-12 panel beats before any spend.
Models: claude CLI or local Qwen
Inputs: A vault note at status ingested.
Outputs: Analysis merged into the note; status becomes analysed.
Risk: Cheap to redo. Read the story seed and beats and fix by hand.
Recon draft clip
Fast image-to-video test on the cheapest tier; keep the good frames as LoRA training data.
Models: HunyuanVideo 1.5
Inputs: A keyframe or panel plus a short motion prompt.
Outputs: A rough clip; salvaged stills feed the cast foundry.
Risk: Accept jank. Batch overnight on interruptible GPUs; cap spend per song.
Model bake-off
Run the same shot across several models to see which wins on your own material.
Models: HunyuanVideo 1.5, Wan 2.2, LTX-2.3
Inputs: One fixed shot spec.
Outputs: A side-by-side sheet; the winner gets promoted for that use.
Risk: Only promote a new model to Hero after it beats the incumbent here.
Character LoRA foundry
Train a reusable LoRA for a hero recurring character; validate before it's allowed downstream.
Models: FLUX.1-dev or Qwen-Image base
Inputs: A curated reference set, ideally seeded from your own art and photos.
Outputs: A versioned .safetensors in the asset registry, with trigger word and sample grid.
Risk: A bad LoRA poisons every downstream job. Test-sheet it before use; version it.
Singer / band likeness LoRA
Same foundry, aimed at real people (singers, band, supports) for on-model performance shots.
Models: FLUX.1-dev or Qwen-Image base
Inputs: Consented reference photos of the person.
Outputs: A versioned likeness LoRA in the registry.
Risk: Keep consent and any private likenesses local; mark the asset private.
Reference-image character
Hold a character across shots with multi-reference, no training. Good for occasional cast.
Models: FLUX.2 multi-ref, Qwen-Image-Edit
Inputs: 1-10 reference images of the character.
Outputs: On-model frames without a trained LoRA.
Risk: Weaker lock than a LoRA under heavy motion; fine for lighter appearances.
Comic pre-viz panels
Render 8-12 stills from the panel beats. The second cheap gate before any video spend.
Models: Qwen-Image, FLUX.2 Klein
Inputs: A briefed note with panel beats and its cast.
Outputs: A panel set; status becomes panels. Approve the strongest.
Risk: Stills are cheap to re-roll; do the visual arguing here, not in video.
Keyframe selection
Promote the approved panels to start and end frames for video shots.
Models: -
Inputs: Approved panels.
Outputs: Keyframe pairs that bookend each motion.
Risk: Pick pairs that imply believable motion between them.
Pose capture + retarget
Capture a move from any footage and retarget it onto your character.
Models: DWPose, AnimateAnyone/UniAnimate-class
Inputs: A reference clip (phone footage, a performance) plus a target character.
Outputs: A pose-driven clip of your character performing the move.
Risk: Capture and pose maps run local; only the driven render costs GPU.
Beat-synced motion
Snap a captured pose track to the song's beat grid so the movement lands on the music.
Models: librosa, Demucs
Inputs: A pose track and the song's audio.
Outputs: A beat-aligned motion track ready to drive generation.
Risk: Runs local on audio analysis; no GPU cost until render.
Duplicate / mirror choreography
Apply one captured motion to several cast members, or mirror it, before rendering.
Models: -
Inputs: A captured motion track and target cast.
Outputs: Group or symmetric choreography as data.
Risk: Pure data manipulation; free until it hits the generator.
Controlled hero shot
Pose, depth and reference conditioned video for a deliberately blocked hero shot.
Models: Wan-VACE-class controllable video
Inputs: Keyframes, a pose or depth map, cast and location assets.
Outputs: A directed shot with intended blocking and continuity.
Risk: The most demanding path; prove the shot at draft before rendering premium.
Location plate + lock
Establish a consistent environment and reuse it across every shot in a scene.
Models: Qwen-Image, SAM 2
Inputs: A location description or plate image.
Outputs: A registered location asset shots can pull by name.
Risk: Lock it once; reuse buys both consistency and cheaper shots.
Object / prop transfer
Detect, lift and place a consistent object or prop across shots.
Models: SAM 2, Grounding DINO, inpaint
Inputs: A reference of the object and the target frames.
Outputs: The prop held consistent through the scene.
Risk: Detection and matting run local; only the placement render costs GPU.
Beat-synced lyric montage
Cut approved clips to the beat and render for each platform. The forgiving, shareable lane.
Models: librosa, ffmpeg
Inputs: Approved clips and the song audio.
Outputs: A cut, rendered at 9:16, 1:1 and 16:9.
Risk: Local assembly; the cheapest way to a publishable, community-shareable piece.
Vertical micro-drama cut
A short vertical story with dialogue voice and lip-sync, cut for phone-first viewing.
Models: Qwen3-TTS, LatentSync/InfiniteTalk, ffmpeg
Inputs: A story seed, cast, and approved shots.
Outputs: A 9:16 micro-drama.
Risk: Lip-sync is a real gate; vet on a Recon pass first.
TTS voiceover
Narration or spoken-word voiceover from a script. The non-music backbone.
Models: Qwen3-TTS, Kokoro-82M
Inputs: A script and a chosen voice.
Outputs: A voice track ready to cut against picture.
Risk: Runs light and local; Kokoro is CPU-viable for fast drafts.
Voice clone (character / singer)
A cloned voice for a recurring character or singer, registered like a cast asset.
Models: Qwen3-TTS, Chatterbox v3
Inputs: A few seconds of consented reference audio.
Outputs: A reusable voice profile in the asset registry.
Risk: Keep consent and private voices local; mark the asset private.
Speech-to-text transcribe
Transcribe dialogue or voiceover to timed text. Feeds captions, edits and search.
Models: Whisper large v3, faster-whisper
Inputs: Any audio or video with speech.
Outputs: A timed transcript (word and segment timestamps).
Risk: Local and cheap; distil/faster variants run near realtime.
Auto captions / subtitles
Generate SRT from transcription and burn styled captions for phone-first viewing.
Models: Whisper large v3, ffmpeg
Inputs: A clip with speech, or an existing transcript.
Outputs: An SRT and a caption-burned render.
Risk: Always proofread auto-captions before publish; names and lyrics trip them.
Cultural-translation dub
Transcribe, translate, re-voice and re-sync into another language. For the cultural-translation work.
Models: Whisper large v3, Qwen3-TTS, LatentSync
Inputs: A finished clip and a target language.
Outputs: A dubbed, lip-synced version.
Risk: Translation needs a human check for meaning and register, not just literal words.