Model picker
Open-weight models grouped by what they do, sourced from Hugging Face. The teal chip on each card is the task; the line under it is what it's best at. Status: active is in the working stack, recon is being tested, watching is on the radar. Filter to narrow; the recon watcher keeps this fresh.
Video
Text and image to video, the hero outputWan 2.2 A14B
text/image-to-videoBest open quality with a clean commercial licence; huge LoRA ecosystem.
HunyuanVideo 1.5
text/image-to-videoBest quality per GB on a rented 4090/5090; the draft-tier workhorse.
LTX-2.3
text/image-to-video + audioNative synced audio+video, native 4K and vertical 1080x1920.
Kandinsky 5.0 Video
text-to-videoLite 2B is a fast, cheap draft option; worth a Recon bake-off.
Control and motion
Pose, depth, segmentation, directed shotsWan-VACE
controllable video (pose/depth/ref)All-in-one control for directed hero shots; the core of the control rig.
DWPose
pose estimationCapture pose from any footage; runs local, no GPU spend.
SAM 3.1
segmentation (image/video)Promotes SAM 2. Promptable concept segmentation (text/exemplar picks all matching instances at once) plus faster video tracking; detect, lift and lock objects across shots.
Depth Anything v2
depth estimationDepth maps for control-conditioned generation.
Image
Keyframes, comic panels, character consistencyQwen-Image + Edit 2512
text-to-image + editCharacter consistency and best open text/lettering; 2512 edit build tops the 2511 you had.
Z-Image Turbo
text-to-imageNear-flagship quality at tiny VRAM (~2.3s/image on a 4090, 8 steps); the best efficiency-per-licence play. Verified on ComfyUI Wiki + HF.
FLUX.2 Klein 4B
text-to-imageThe only permissive FLUX.2 weight; commercial-safe drafts.
FLUX.2 dev
text-to-imageQuality leader with 10 native reference images; needs a paid licence here.
Voice (TTS)
Text to speech and voice cloningQwen3-TTS
text-to-speech + cloning3-second-reference cloning, commercial-safe default voice engine.
Chatterbox Multilingual v3
text-to-speech + cloningBeat ElevenLabs in blind tests; note output watermarking.
Kokoro-82M
text-to-speech (fixed voices)Fastest narration voice; cannot clone.
Speech to text
Transcription and the caption backboneNVIDIA Parakeet TDT 1.1B
speech-to-text (English)Fastest high-quality open ASR; best throughput for bulk autocaptioning.
Whisper large v3
speech-to-text (multilingual)Now the multilingual fallback (99+ languages), not the English accuracy leader. Widely deployed.
faster-whisper (distil)
speech-to-textNear-realtime STT for bulk captioning on modest hardware.
Lip-sync and avatars
Audio-driven talking performanceInfiniteTalk
audio-driven performanceUnlimited-length audio-driven video with head and body sync.
LatentSync 1.x
lip-sync dubbingBest pure lip fidelity for dubbing existing footage.
MuseTalk 1.5
realtime lip-syncRealtime-class, the cheap bulk dubbing pass.
Language
Ideation and scriptingQwen3.6-27B
ideation / scriptingStrongest single-24GB-card ideation model for local scripting. Verify the exact 27B dense build exists (mainline Qwen3.6 reads as MoE).
GLM-5.2
agentic scriptingHighest-ranked open-weight model on Artificial Analysis (Intelligence Index 51); for rented big GPUs, not a desktop card. Verified.
Embeddings
Semantic search over the vaultQwen3-Embedding-0.6B
semantic searchCheap local semantic search over the lyric/script vault.
Audio
Stem and music analysis for timing, energy, beat-syncAll-In-One
structure + beats + tempoOne call gives timestamped sections (intro/verse/chorus/bridge/outro) plus beats, downbeats and tempo as JSON. The backbone of audio ingestion.
WhisperX
lyric-to-audio alignmentRun on the isolated vocal stem with your lyric text as the transcript for word/line timestamps; syncs visuals to the words.
Essentia + MTG models
mood / energy / danceabilityHuman-readable energy/mood/danceability/arousal-valence per stem and over time; drives visual intensity. Watch the NC weights if commercialised.
Beat This!
beat / downbeat trackingSOTA beat grid without DBN post-processing; handles tempo and metre changes for tight visual cuts.
CLAP (LAION)
text-audio embeddingsQuery stems by natural-language mood ("dark", "euphoric"); the commercial-safe embedding option vs the NC MERT/Essentia weights.
Demucs (htdemucs_ft)
stem separation (fallback)4-stem separator for non-Suno material or re-splitting a bounce. Suno already gives you stems, so this is the fallback.
madmom
key / chord / beat referenceLong-standing beat/downbeat and ~90% chord/key detection; the harmony tag layer, lowest priority for visuals.
3D and worlds
Meshes, environments and explorable worldsTRELLIS.2
image-to-3D asset (PBR mesh)Strongest new open image-to-3D; full PBR + transparency from one backbone. Clean MIT.
Step1X-3D
image-to-3D asset (controllable)The cleanest commercial licence for textured 3D; strong geometry-texture alignment.
Hunyuan3D 2.1
image-to-3D asset (PBR)Production image-to-3D with PBR materials; flag the region and MAU licence limits.
PartCrafter
part-aware image-to-3DMultiple separable 3D parts from one image, no pre-segmentation; low VRAM, clean MIT.
TripoSR
fast image-to-3DSub-second single-image-to-mesh; the cheap MIT draft-tier baseline.
Direct3D-S2
high-res image-to-3DGigascale image-to-3D up to 1024 resolution via sparse attention; permissive MIT.
HunyuanWorld-Voyager
explorable world (RGBD)Camera-controlled RGBD world with real-time 3D reconstruction; exports point clouds for virtual sets. Flag licence.
NVIDIA Cosmos
world foundation modelFine-tunable world-foundation models; commercial licence but 80GB-class, more physical-AI than cinematic.
Matrix-Game 3.0
real-time interactive worldReal-time streaming interactive world with long-horizon memory; permissive MIT.
Prometheus
feed-forward 3D sceneSeconds-fast text-to-3D-scene as pixel-aligned Gaussians; MIT. For whole environments, not single props.
gsplat + Nerfstudio
gaussian-splat trainingThe main commercially-usable open splat-training stack; turn real footage into 3D sets you can light.
COLMAP
photogrammetry / camera solveThe de-facto SfM camera-pose and point-cloud step feeding every splat trainer; commercial-friendly BSD.