# IDEA-Research/Grounded-Segment-Anything

Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything

Repository: https://github.com/IDEA-Research/Grounded-Segment-Anything
Canonical: https://ross.abutalabs.com/products/grounded-segment-anything
Homepage: https://arxiv.org/abs/2401.14159
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Topics: open-vocabulary-detection, open-vocabulary-segmentation, data-generation, automatic-labeling-system, caption, speech, 3d-whole-body-pose-estimation, image-editing
Last push: 2024-09-05T06:07:32+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 88
- inputs: {"age_days": 1245, "days_push": 727, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 17710, forks 1594 (observed 2026-08-28T04:11:20.201769+00:00)

## What it is
Grounded-Segment-Anything (Grounded SAM) combines Grounding DINO with Segment Anything to detect and segment arbitrary objects from text prompts, and integrates Stable Diffusion, Recognize Anything, and other models for extended visual tasks. It is a research demo and toolkit distributed as Jupyter notebooks and Python code for open-world detection, segmentation, automatic dataset annotation, and image editing.

## Use cases
- segment any object in an image using a text prompt
- automatically annotate and label datasets for object detection training
- zero-shot open-vocabulary object detection and segmentation
- inpaint or edit images with Stable Diffusion using detected masks
- generate captions and tags for images with Recognize Anything and BLIP
- build pipelines connecting open-world vision models

## When to choose
- you need text-prompted detection and segmentation without training a custom model
- you want to auto-label image datasets for downstream training
- you need zero-shot segmentation benchmarks like SegInW
- you want a composable base for combining vision foundation models

## When to avoid
- you need real-time production inference with low latency
- you want a polished end-user application rather than research code
- you need SAM 2 tracking, in which case Grounded-SAM-2 is the successor
- you lack a GPU or cannot run large vision models locally

## Facets
- artifact type: library
- maturity: active
- function: computer-vision, image-processing, machine-learning, data-generation
- domain: computer-vision, image-processing, machine-learning, artificial-intelligence
- platform: python, cross-platform
- tags: grounded-sam, segment-anything, grounding-dino, open-vocabulary-detection, open-vocabulary-segmentation, automatic-labeling, stable-diffusion, image-editing, zero-shot, gpu

## Member repositories
- IDEA-Research/Grounded-Segment-Anything (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:11:20.201769+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:02:29.152577+00:00, confidence not recorded.
  - readme: https://github.com/IDEA-Research/Grounded-Segment-Anything (fetched 2026-08-28T04:11:20.201769+00:00, sha 3a3933de7887)
  - homepage: https://arxiv.org/abs/2401.14159 (fetched 2026-08-29T08:00:42.879420+00:00, sha 9a6bd885f6d5)
  - site_page: https://info.arxiv.org/about/donate.html (fetched 2026-08-29T08:00:42.889316+00:00, sha cca9c3a11c56)
  - site_page: https://info.arxiv.org/about/ourmembers.html (fetched 2026-08-29T08:00:42.893813+00:00, sha 47cbc55ff1de)
  - site_page: https://info.arxiv.org/about (fetched 2026-08-29T08:00:42.896186+00:00, sha a1f16f915a9a)
  - site_page: https://info.arxiv.org/labs/index.html (fetched 2026-08-29T08:00:42.891738+00:00, sha b14a8d05a0ec)
- Data as of 2026-08-30T08:39:29.467469+00:00.
