Step-Video-T2V
Open in the interactive archive →
Overview
30B-parameter text-to-video foundation model, up to 204 frames. Deep-compression Video-VAE (16x16 spatial, 8x temporal). Bilingual (EN/ZH) encoders, 3D full-attention DiT, Video-DPO. MIT license. Turbo + TI2V (image-to-video) variants.
Generates
Video generation
Details
- Vendor
- StepFun
- Type
- Video Generation
- Release year
- 2025
- Jurisdiction
- China
- CEO
- Jiang Daxin
- Key staffers
- Jiang Daxin (Founder, CEO; former Microsoft/MSRA)
- Deployment
- Local
- Open source
- Yes
- Free tier
- Yes
- Stock status
- Private
- X
- @StepFun_AI
Known safety issues
None publicly documented
Last verified 2026-07-06 · Primary source ↗
Cite this
“Step-Video-T2V.” AIDb — The AI Database. Verified 2026-07-06. https://theaidatabase.com/m/step-video-t2v-stepfun.html