IDEA-Research/Grounded-Segment-Anything
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything observed · 2026-08-28
Health v2 · maintenance only
30/100
- Activity 0
- Release rhythm 35
- Longevity 88
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1245
- days_rel: n/a
- days_push: 727
- n_releases_24m: 0
Adoption not part of the score
17710 stars · 1594 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Grounded-Segment-Anything (Grounded SAM) combines Grounding DINO with Segment Anything to detect and segment arbitrary objects from text prompts, and integrates Stable Diffusion, Recognize Anything, and other models for extended visual tasks. It is a research demo and toolkit distributed as Jupyter notebooks and Python code for open-world detection, segmentation, automatic dataset annotation, and image editing.
Use cases
- segment any object in an image using a text prompt
- automatically annotate and label datasets for object detection training
- zero-shot open-vocabulary object detection and segmentation
- inpaint or edit images with Stable Diffusion using detected masks
- generate captions and tags for images with Recognize Anything and BLIP
- build pipelines connecting open-world vision models
When to choose
- you need text-prompted detection and segmentation without training a custom model
- you want to auto-label image datasets for downstream training
- you need zero-shot segmentation benchmarks like SegInW
- you want a composable base for combining vision foundation models
When to avoid
- you need real-time production inference with low latency
- you want a polished end-user application rather than research code
- you need SAM 2 tracking, in which case Grounded-SAM-2 is the successor
- you lack a GPU or cannot run large vision models locally
Facets
library · maturity active
computer-vision image-processing machine-learning data-generation computer-vision image-processing machine-learning artificial-intelligence python cross-platform grounded-sam segment-anything grounding-dino open-vocabulary-detection open-vocabulary-segmentation automatic-labeling stable-diffusion image-editing zero-shot gpu
6 sources
- readme: https://github.com/IDEA-Research/Grounded-Segment-Anything · fetched 2026-08-28 · 3a3933de7887
- homepage: https://arxiv.org/abs/2401.14159 · fetched 2026-08-29 · 9a6bd885f6d5
- site_page: https://info.arxiv.org/about/donate.html · fetched 2026-08-29 · cca9c3a11c56
- site_page: https://info.arxiv.org/about/ourmembers.html · fetched 2026-08-29 · 47cbc55ff1de
- site_page: https://info.arxiv.org/about · fetched 2026-08-29 · a1f16f915a9a
- site_page: https://info.arxiv.org/labs/index.html · fetched 2026-08-29 · b14a8d05a0ec
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| IDEA-Research/Grounded-Segment-Anything | main | 30 |
For agents
markdown · JSON · MCP: product_card(name="IDEA-Research/Grounded-Segment-Anything")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem