Ross ROSS = Recommend OSS · open-source software intelligence for agents

IDEA-Research/Grounded-Segment-Anything

Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything observed · 2026-08-28

github.com/IDEA-Research/Grounded-Segment-Anything · homepage · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 0
  • Release rhythm 35
  • Longevity 88

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1245
  • days_rel: n/a
  • days_push: 727
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

17710 stars · 1594 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Grounded-Segment-Anything (Grounded SAM) combines Grounding DINO with Segment Anything to detect and segment arbitrary objects from text prompts, and integrates Stable Diffusion, Recognize Anything, and other models for extended visual tasks. It is a research demo and toolkit distributed as Jupyter notebooks and Python code for open-world detection, segmentation, automatic dataset annotation, and image editing.

Use cases

  • segment any object in an image using a text prompt
  • automatically annotate and label datasets for object detection training
  • zero-shot open-vocabulary object detection and segmentation
  • inpaint or edit images with Stable Diffusion using detected masks
  • generate captions and tags for images with Recognize Anything and BLIP
  • build pipelines connecting open-world vision models

When to choose

  • you need text-prompted detection and segmentation without training a custom model
  • you want to auto-label image datasets for downstream training
  • you need zero-shot segmentation benchmarks like SegInW
  • you want a composable base for combining vision foundation models

When to avoid

  • you need real-time production inference with low latency
  • you want a polished end-user application rather than research code
  • you need SAM 2 tracking, in which case Grounded-SAM-2 is the successor
  • you lack a GPU or cannot run large vision models locally

Facets

library · maturity active

computer-vision image-processing machine-learning data-generation computer-vision image-processing machine-learning artificial-intelligence python cross-platform grounded-sam segment-anything grounding-dino open-vocabulary-detection open-vocabulary-segmentation automatic-labeling stable-diffusion image-editing zero-shot gpu

6 sources

Member repositories

RepositoryRoleHealth v2
IDEA-Research/Grounded-Segment-Anythingmain30

For agents

markdown · JSON · MCP: product_card(name="IDEA-Research/Grounded-Segment-Anything")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem