bytedance/Lance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing. observed · 2026-08-28
Health v2 · maintenance only
55/100
- Activity 92
- Release rhythm 35
- Longevity 7
Flags: no_releases young
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 110
- days_rel: n/a
- days_push: 50
- n_releases_24m: 0
Adoption not part of the score
1329 stars · 95 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Lance is a 3B-parameter native unified multimodal model from ByteDance for image and video understanding, generation, and editing, trained from scratch with a multi-task synergy recipe. It is released as a research artifact with inference and fine-tuning code, model checkpoints on Hugging Face, and vLLM-Omni support.
Use cases
- generate videos from text prompts
- generate images from text descriptions
- edit images and videos with natural language instructions
- understand and reason about image and video content
- fine-tune a small unified multimodal model on custom data
- run multimodal generation on a limited GPU budget
When to choose
- you need a single compact model handling image/video understanding, generation, and editing
- you want a research baseline for unified multimodal modeling under constrained compute
- you want to experiment with fine-tuning a 3B multimodal model
When to avoid
- you need production-grade, polished output quality for commercial media generation
- you need high-resolution or high-framerate video (training capped at 768x768 images and 480p 12 FPS video)
- you only need a specialized single-task model like pure text-to-image
Facets
library · maturity active
machine-learning deep-learning image-processing video-processing llm-inference artificial-intelligence machine-learning computer-vision image-processing python multimodal-model unified-model text-to-video text-to-image image-editing video-editing research-project bytedance 3b-parameters video gpu linux
2 sources
- readme: https://github.com/bytedance/Lance · fetched 2026-08-28 · 313f76d11796
- homepage: https://lance-project.github.io · fetched 2026-08-29 · 37466d4a72e0
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| bytedance/Lance | main | 55 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem