# ttengwang/Caption-Anything

Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything

Repository: https://github.com/ttengwang/Caption-Anything
Canonical: https://ross.abutalabs.com/products/caption-anything
Language: Python
License: BSD-3-Clause
License Family: permissive
Topics: chatgpt, controllable-generation, segment-anything, controllable-image-captioning, image-captioning
Last push: 2023-08-29T05:26:45+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 88
- inputs: {"age_days": 1244, "days_push": 1100, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1777, forks 104 (observed 2026-08-28T04:05:34.722453+00:00)

## What it is
Caption-Anything combines Segment Anything image segmentation, visual captioning, and ChatGPT to generate tailored captions for any object in an image. It supports visual controls (mouse clicks) and language controls (length, sentiment, factuality, language) plus an interactive chat about selected objects.

## Use cases
- generate captions for specific objects in an image
- caption images in different styles or languages
- click on an object and get a description of it
- chat with an AI about a selected region of an image
- control caption length, sentiment, and factuality
- segment an image and describe each object

## When to avoid
- you need video captioning or tracking (see Track-Anything instead)
- you need a production-ready API rather than a research demo
- you cannot run GPU-heavy models like Segment Anything and captioning models

## Facets
- artifact type: application
- maturity: maintenance
- function: image-processing, nlp, llm-inference, chatbot, ui-components
- domain: computer-vision, image-processing, large-language-models, artificial-intelligence
- platform: python
- tags: segment-anything, image-captioning, chatgpt, controllable-generation, multimodal, gradio-demo, linux, web-server, gpu

## Member repositories
- ttengwang/Caption-Anything (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:34.722453+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:25:14.781732+00:00, confidence not recorded.
  - readme: https://github.com/ttengwang/Caption-Anything (fetched 2026-08-28T04:05:34.722453+00:00, sha 2909ee7f8339)
- Data as of 2026-08-30T08:39:29.467469+00:00.
