bytedance/Sa2VA
Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS) observed · 2026-08-28
Health v2 · maintenance only
70/100
- Activity 96
- Release rhythm 52
- Longevity 43
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 1
- age_days: 604
- days_rel: 320
- days_push: 29
- n_releases_24m: 2
Adoption not part of the score
1666 stars · 131 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understanding of images and videos. It includes the core Sa2VA model plus extensions like SAMTok, VRT, Pixel-SAIL, and SaSaSa2VA, with pretrained checkpoints on Hugging Face.
Use cases
- referring image and video segmentation with a multimodal LLM
- grounded conversation about images and videos
- chat with an AI about specific pixels or regions in an image
- generate segmentation masks from natural language instructions
- visual prompting and region-level question answering
- evaluate grounded visual reasoning with VRT-Bench
- train a multimodal LLM that outputs segmentation masks
- video object segmentation benchmark submission
When to choose
- you need unified referring segmentation and visual chat in one model
- you want pixel-level grounding on top of InternVL or Qwen-VL backbones
- you are doing research on grounded multimodal understanding
- you need pretrained checkpoints for image and video segmentation tasks
When to avoid
- you need a production-ready, latency-optimized vision API
- you lack a GPU or cannot run large multimodal models
- you only need simple image classification or object detection
- you want a no-code GUI tool rather than a Python research codebase
Facets
library · maturity active
machine-learning computer-vision deep-learning nlp image-processing video-processing computer-vision large-language-models machine-learning artificial-intelligence python multimodal-llm segmentation sam2 grounded-understanding referring-segmentation visual-prompting research-code image-chat video-chat mask-tokens gpu linux
1 source
- readme: https://github.com/bytedance/Sa2VA · fetched 2026-08-28 · b80c36c49b30
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| bytedance/Sa2VA | main | 70 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem