Ross ROSS = Recommend OSS · open-source software intelligence for agents

bytedance/Sa2VA

Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS) observed · 2026-08-28

github.com/bytedance/Sa2VA · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

70/100

  • Activity 96
  • Release rhythm 52
  • Longevity 43
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 1
  • age_days: 604
  • days_rel: 320
  • days_push: 29
  • n_releases_24m: 2

Full methodology

Adoption not part of the score

1666 stars · 131 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Sa2VA is a family of research models and codebases from ByteDance that combine SAM-2 with multimodal LLMs for pixel-level grounded understanding of images and videos. It includes the core Sa2VA model plus extensions like SAMTok, VRT, Pixel-SAIL, and SaSaSa2VA, with pretrained checkpoints on Hugging Face.

Use cases

  • referring image and video segmentation with a multimodal LLM
  • grounded conversation about images and videos
  • chat with an AI about specific pixels or regions in an image
  • generate segmentation masks from natural language instructions
  • visual prompting and region-level question answering
  • evaluate grounded visual reasoning with VRT-Bench
  • train a multimodal LLM that outputs segmentation masks
  • video object segmentation benchmark submission

When to choose

  • you need unified referring segmentation and visual chat in one model
  • you want pixel-level grounding on top of InternVL or Qwen-VL backbones
  • you are doing research on grounded multimodal understanding
  • you need pretrained checkpoints for image and video segmentation tasks

When to avoid

  • you need a production-ready, latency-optimized vision API
  • you lack a GPU or cannot run large multimodal models
  • you only need simple image classification or object detection
  • you want a no-code GUI tool rather than a Python research codebase

Facets

library · maturity active

machine-learning computer-vision deep-learning nlp image-processing video-processing computer-vision large-language-models machine-learning artificial-intelligence python multimodal-llm segmentation sam2 grounded-understanding referring-segmentation visual-prompting research-code image-chat video-chat mask-tokens gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
bytedance/Sa2VAmain70

For agents

markdown · JSON · MCP: product_card(name="bytedance/Sa2VA")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem