Ross ROSS = Recommend OSS · open-source software intelligence for agents

IDEA-Research/Grounded-SAM-2

Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2 observed · 2026-08-28

github.com/IDEA-Research/Grounded-SAM-2 · homepage · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

37/100

  • Activity 51
  • Release rhythm 8
  • Longevity 54
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 762
  • days_rel: 630
  • days_push: 295
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

3708 stars · 430 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Grounded SAM 2 is a foundation-model pipeline that combines open-set detectors (Grounding DINO, Grounding DINO 1.5/1.6, Florence-2, DINO-X) with SAM 2 to detect, segment, and track arbitrary objects in images and videos from text prompts. It provides simple demo implementations and visualization built on the supervision library.

Use cases

  • segment any object in an image using a text prompt
  • track objects across video frames with text-based grounding
  • auto-annotate datasets with masks for training segmentation models
  • detect and segment dense small objects in high-resolution 4K images with SAHI
  • visualize detection, segmentation, and tracking results on videos
  • run zero-shot open-vocabulary segmentation benchmarks

When to choose

  • you need text-prompted open-vocabulary detection and segmentation without training a custom model
  • you want to segment and track objects in videos with a proven research pipeline
  • you need automated mask annotation for dataset labeling
  • you want simple reference implementations combining Grounding DINO/Florence-2/DINO-X with SAM 2

When to avoid

  • you need a production-ready optimized inference service rather than demo notebooks
  • you lack GPU resources, since the models are heavy to run locally
  • you only need classic closed-set detection with fixed classes, where a lightweight YOLO-style detector suffices
  • you need real-time performance on edge devices

Facets

library · maturity active

computer-vision image-processing machine-learning deep-learning computer-vision image-processing artificial-intelligence deep-learning python cross-platform object-detection image-segmentation video-object-tracking open-vocabulary-detection sam2 grounding-dino florence-2 dino-x zero-shot jupyter-notebooks video gpu

6 sources

Member repositories

RepositoryRoleHealth v2
IDEA-Research/Grounded-SAM-2main37

For agents

markdown · JSON · MCP: product_card(name="IDEA-Research/Grounded-SAM-2")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem