NVlabs/describe-anything
[ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 28
- Release rhythm 35
- Longevity 36
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 516
- days_rel: n/a
- days_push: 433
- n_releases_24m: 0
Adoption not part of the score
1514 stars · 94 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Describe Anything Model (DAM) is a vision-language model that generates detailed descriptions of user-specified regions in images and videos, indicated via points, boxes, scribbles, or masks. The repository provides the model implementation, Gradio demos, SAM integration, and the DLC-Bench evaluation benchmark.
Use cases
- generate detailed captions for a specific region of an image
- describe objects in a video by annotating one frame
- caption image regions using masks from SAM
- evaluate localized captioning models with DLC-Bench
- build an app that describes clicked or boxed image areas
- run a local gradio demo for region-based image description
When to choose
- you need fine-grained, localized image or video descriptions rather than whole-image captions
- you want a research-grade vision-language model with a permissive Apache-2.0 license
- you need a benchmark for evaluating detailed localized captioning
When to avoid
- you only need simple whole-image captioning without region specification
- you lack a GPU or cannot run large multimodal models locally
- you need a production-ready hosted API rather than a research codebase
Facets
library · maturity active
machine-learning computer-vision image-processing video-processing nlp computer-vision large-language-models artificial-intelligence image-processing python windows vision-language-model region-captioning image-captioning video-captioning segmentation gradio-demo benchmark iccv-2025 multimodal video linux macos gpu
2 sources
- readme: https://github.com/NVlabs/describe-anything · fetched 2026-08-28 · 13e2a6651342
- homepage: https://describe-anything.github.io/ · fetched 2026-08-29 · 3f9b27456586
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| NVlabs/describe-anything | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="NVlabs/describe-anything")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem