Model picker

Open-weight models grouped by what they do, sourced from Hugging Face. The teal chip on each card is the task; the line under it is what it's best at. Status: active is in the working stack, recon is being tested, watching is on the radar. Filter to narrow; the recon watcher keeps this fresh.

Category
3D type
Status
Licence
Content
Sort

Video

Text and image to video, the hero output
video

Wan 2.2 A14B

text/image-to-video

Best open quality with a clean commercial licence; huge LoRA ecosystem.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 24 (FP8) / 40-80 (FP16) GB · HF likes: 1900
Hugging Face →
video

HunyuanVideo 1.5

text/image-to-video

Best quality per GB on a rented 4090/5090; the draft-tier workhorse.

activecommercialunfiltered
Licence: Tencent Community (>100M-MAU clause; NOT Apache, verified)
VRAM: 14-24 GB · HF likes: 1400
Hugging Face →
video

LTX-2.3

text/image-to-video + audio

Native synced audio+video, native 4K and vertical 1080x1920.

activerestrictedunfiltered
Licence: LTX-2 Community (free commercial under US$10M/yr)
VRAM: 16-80 GB · HF likes: 2600
Hugging Face →
video

Kandinsky 5.0 Video

text-to-video

Lite 2B is a fast, cheap draft option; worth a Recon bake-off.

watchingrestrictedunfiltered
Licence: open (check per-weight)
VRAM: 12 (Lite) / 40+ (Pro) GB · HF likes: 700
Hugging Face →

Control and motion

Pose, depth, segmentation, directed shots
control

Wan-VACE

controllable video (pose/depth/ref)

All-in-one control for directed hero shots; the core of the control rig.

reconcommercial
Licence: Apache 2.0
VRAM: 24-80 GB · HF likes: 900
Hugging Face →
control

DWPose

pose estimation

Capture pose from any footage; runs local, no GPU spend.

activecommercial
Licence: Apache 2.0
VRAM: low / local GB · HF likes: 400
Hugging Face →
control

SAM 3.1

segmentation (image/video)

Promotes SAM 2. Promptable concept segmentation (text/exemplar picks all matching instances at once) plus faster video tracking; detect, lift and lock objects across shots.

activerestricted
Licence: Meta SAM Licence (permissive, not OSI; check terms)
VRAM: low / local GB · HF likes: 3600
Hugging Face →
control

Depth Anything v2

depth estimation

Depth maps for control-conditioned generation.

activecommercial
Licence: Apache 2.0
VRAM: low / local GB · HF likes: 1500
Hugging Face →

Image

Keyframes, comic panels, character consistency
image

Qwen-Image + Edit 2512

text-to-image + edit

Character consistency and best open text/lettering; 2512 edit build tops the 2511 you had.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 12-24 GB · HF likes: 2100
Hugging Face →
image

Z-Image Turbo

text-to-image

Near-flagship quality at tiny VRAM (~2.3s/image on a 4090, 8 steps); the best efficiency-per-licence play. Verified on ComfyUI Wiki + HF.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 5-16 GB · HF likes: 5000
Hugging Face →
image

FLUX.2 Klein 4B

text-to-image

The only permissive FLUX.2 weight; commercial-safe drafts.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 8-10 GB · HF likes: 1800
Hugging Face →
image

FLUX.2 dev

text-to-image

Quality leader with 10 native reference images; needs a paid licence here.

watchingrestrictedunfiltered
Licence: NON-COMMERCIAL (paid licence for commercial)
VRAM: 32 (FP8) - 64 GB · HF likes: 5400
Hugging Face →

Voice (TTS)

Text to speech and voice cloning
tts

Qwen3-TTS

text-to-speech + cloning

3-second-reference cloning, commercial-safe default voice engine.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 4-8 GB · HF likes: 1200
Hugging Face →
tts

Chatterbox Multilingual v3

text-to-speech + cloning

Beat ElevenLabs in blind tests; note output watermarking.

activecommercialunfiltered
Licence: MIT
VRAM: 6-8 GB · HF likes: 1600
Hugging Face →
tts

Kokoro-82M

text-to-speech (fixed voices)

Fastest narration voice; cannot clone.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: cpu-ok GB · HF likes: 4100
Hugging Face →

Speech to text

Transcription and the caption backbone
stt

NVIDIA Canary-Qwen 2.5B

speech-to-text (English)

activerestricted
Licence: NVIDIA open (verify exact terms)
VRAM: low GB · HF likes: 800
Hugging Face →
stt

NVIDIA Parakeet TDT 1.1B

speech-to-text (English)

Fastest high-quality open ASR; best throughput for bulk autocaptioning.

reconrestricted
Licence: NVIDIA open (verify)
VRAM: low GB · HF likes: 600
Hugging Face →
stt

Whisper large v3

speech-to-text (multilingual)

Now the multilingual fallback (99+ languages), not the English accuracy leader. Widely deployed.

activecommercial
Licence: Apache 2.0
VRAM: 6-10 GB · HF likes: 4600
Hugging Face →
stt

faster-whisper (distil)

speech-to-text

Near-realtime STT for bulk captioning on modest hardware.

activecommercial
Licence: MIT
VRAM: 2-6 GB · HF likes: 900
Hugging Face →

Lip-sync and avatars

Audio-driven talking performance
lipsync

InfiniteTalk

audio-driven performance

Unlimited-length audio-driven video with head and body sync.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 24 GB · HF likes: 800
Hugging Face →
lipsync

LatentSync 1.x

lip-sync dubbing

Best pure lip fidelity for dubbing existing footage.

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 8-12 GB · HF likes: 1100
Hugging Face →
lipsync

MuseTalk 1.5

realtime lip-sync

Realtime-class, the cheap bulk dubbing pass.

activecommercialunfiltered
Licence: MIT-style
VRAM: 6-8 GB · HF likes: 1300
Hugging Face →

Language

Ideation and scripting
llm

Qwen3.6-27B

ideation / scripting

Strongest single-24GB-card ideation model for local scripting. Verify the exact 27B dense build exists (mainline Qwen3.6 reads as MoE).

activecommercialunfiltered
Licence: Apache 2.0
VRAM: 16-17 (Q4) GB · HF likes: 2200
Hugging Face →
llm

GLM-5.2

agentic scripting

Highest-ranked open-weight model on Artificial Analysis (Intelligence Index 51); for rented big GPUs, not a desktop card. Verified.

reconcommercialunfiltered
Licence: MIT/Apache
VRAM: rented multi-GPU GB · HF likes: 900
Hugging Face →

Embeddings

Semantic search over the vault
embeddings

Qwen3-Embedding-0.6B

semantic search

Cheap local semantic search over the lyric/script vault.

activecommercial
Licence: Apache 2.0
VRAM: low GB · HF likes: 700
Hugging Face →

Audio

Stem and music analysis for timing, energy, beat-sync
audio

All-In-One

structure + beats + tempo

One call gives timestamped sections (intro/verse/chorus/bridge/outro) plus beats, downbeats and tempo as JSON. The backbone of audio ingestion.

activecommercial
Licence: MIT
VRAM: GPU rec, CPU ok (~73s/33min on 4090) GB · HF likes: 900
Repo →
audio

WhisperX

lyric-to-audio alignment

Run on the isolated vocal stem with your lyric text as the transcript for word/line timestamps; syncs visuals to the words.

activecommercial
Licence: BSD (Whisper MIT)
VRAM: ~10 (large-v3); CPU possible GB · HF likes: 3500
Repo →
audio

Essentia + MTG models

mood / energy / danceability

Human-readable energy/mood/danceability/arousal-valence per stem and over time; drives visual intensity. Watch the NC weights if commercialised.

activerestricted
Licence: AGPL-3.0 code; most model weights CC-BY-NC (flag)
VRAM: CPU / light GB · HF likes: 1200
audio

Beat This!

beat / downbeat tracking

SOTA beat grid without DBN post-processing; handles tempo and metre changes for tight visual cuts.

reconcommercial
Licence: MIT
VRAM: small; CPU fine GB · HF likes: 400
Repo →
audio

CLAP (LAION)

text-audio embeddings

Query stems by natural-language mood ("dark", "euphoric"); the commercial-safe embedding option vs the NC MERT/Essentia weights.

reconcommercial
Licence: Apache 2.0
VRAM: modest GB · HF likes: 800
Hugging Face →
audio

Demucs (htdemucs_ft)

stem separation (fallback)

4-stem separator for non-Suno material or re-splitting a bounce. Suno already gives you stems, so this is the fallback.

activecommercial
Licence: MIT
VRAM: 3-8 GB · HF likes: 3000
Repo →
audio

madmom

key / chord / beat reference

Long-standing beat/downbeat and ~90% chord/key detection; the harmony tag layer, lowest priority for visuals.

reconrestricted
Licence: BSD (some algos carry commercial caveats)
VRAM: CPU GB · HF likes: 1400
Repo →

3D and worlds

Meshes, environments and explorable worlds
worlds

TRELLIS.2

image-to-3D asset (PBR mesh)

Strongest new open image-to-3D; full PBR + transparency from one backbone. Clean MIT.

activecommercial
Licence: MIT
VRAM: 24 GB · HF likes: 3000
Hugging Face →
worlds

Step1X-3D

image-to-3D asset (controllable)

The cleanest commercial licence for textured 3D; strong geometry-texture alignment.

activecommercial
Licence: Apache 2.0
VRAM: 27 GB · HF likes: 1200
Hugging Face →
worlds

Hunyuan3D 2.1

image-to-3D asset (PBR)

Production image-to-3D with PBR materials; flag the region and MAU licence limits.

activerestricted
Licence: Tencent Community (EXCLUDES EU/UK/South Korea; 1M-MAU cap)
VRAM: 10-29 GB · HF likes: 4000
Hugging Face →
worlds

PartCrafter

part-aware image-to-3D

Multiple separable 3D parts from one image, no pre-segmentation; low VRAM, clean MIT.

activecommercial
Licence: MIT
VRAM: 8 GB · HF likes: 800
Hugging Face →
worlds

TripoSR

fast image-to-3D

Sub-second single-image-to-mesh; the cheap MIT draft-tier baseline.

activecommercial
Licence: MIT
VRAM: ~6 GB · HF likes: 2500
Hugging Face →
worlds

Direct3D-S2

high-res image-to-3D

Gigascale image-to-3D up to 1024 resolution via sparse attention; permissive MIT.

reconcommercial
Licence: MIT
VRAM: 10 (512) - 24 (1024) GB · HF likes: 700
Hugging Face →
worlds

HunyuanWorld-Voyager

explorable world (RGBD)

Camera-controlled RGBD world with real-time 3D reconstruction; exports point clouds for virtual sets. Flag licence.

reconrestricted
Licence: Tencent HunyuanWorld Community (EXCLUDES EU/UK/South Korea; 1M-MAU cap)
VRAM: 60-80 GB · HF likes: 1500
Hugging Face →
worlds

NVIDIA Cosmos

world foundation model

Fine-tunable world-foundation models; commercial licence but 80GB-class, more physical-AI than cinematic.

reconrestricted
Licence: NVIDIA Open Model (commercial ok)
VRAM: 39-80+ GB · HF likes: 2000
Hugging Face →
worlds

Matrix-Game 3.0

real-time interactive world

Real-time streaming interactive world with long-horizon memory; permissive MIT.

watchingcommercial
Licence: MIT
VRAM: unstated GB · HF likes: 900
Hugging Face →
worlds

Prometheus

feed-forward 3D scene

Seconds-fast text-to-3D-scene as pixel-aligned Gaussians; MIT. For whole environments, not single props.

reconcommercial
Licence: MIT
VRAM: 48 (tested A6000) GB · HF likes: 400
Repo →
worlds

gsplat + Nerfstudio

gaussian-splat training

The main commercially-usable open splat-training stack; turn real footage into 3D sets you can light.

activecommercial
Licence: Apache 2.0
VRAM: 8-24 GB · HF likes: 3500
Repo →
worlds

COLMAP

photogrammetry / camera solve

The de-facto SfM camera-pose and point-cloud step feeding every splat trainer; commercial-friendly BSD.

activerestricted
Licence: BSD-3-Clause
VRAM: cpu + optional gpu GB · HF likes: 3000
Repo →