A unified audio-video generation model.
114B parameters. Only 6B activated per token.

MAGI 2 is a unified model that generates video and audio together. Text, visuals and sound live in a single token sequence, so speech, music, motion and emotion stay in sync.
Text, video and audio are processed by one backbone instead of separate towers.
Dialogue, music, sound effects and visuals are generated together.
MagiMoE keeps capacity at 114B while activating only 6B parameters per token.
Compressing video into generative, composable world knowledge.
Efficiency and expressiveness designed together - from architecture to training systems to data.
What MAGI-2 Preview brings to audio-video generation.
Precise dialogue, singing and emotion with coordinated lip-sync, expressions and body language.
Visuals generated from music with rhythm-aware scene flow.
MagiMoE: 12 heads x 256-dim routed subspaces, top-6 of 256 experts per head.
One Transformer backbone processes text, video and audio together.
Head Parallel over NVLink and InfiniBand keeps 114B training practical.
A growing community of creators making ads, films, MVs and more.
Questions about MAGI 2. Learn more at sand.ai.
Read the MAGI-2 Preview announcement and explore what unified audio-video generation can do.