# fpgaminer/joycaption

JoyCaption is an image captioning Visual Language Model (VLM) being built from the ground up as a free, open, and uncensored model for the community to use in training Diffusion models.

Repository: https://github.com/fpgaminer/joycaption
Canonical: https://ross.abutalabs.com/products/joycaption
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Topics: captioning, vlm, joycaption
Last push: 2026-02-24T20:48:35+00:00

## Health v2 (maintenance only)
Score: 53/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 69, release rhythm 35, longevity 49
- inputs: {"age_days": 690, "days_push": 190, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1244, forks 69 (observed 2026-08-28T04:04:06.886327+00:00)

## What it is
JoyCaption is an open, free, and uncensored image captioning Visual Language Model (VLM) with released weights and training scripts. It generates descriptive captions for images, primarily to support training and finetuning of diffusion models.

## Use cases
- generate captions for images to train diffusion models
- automatically caption datasets for stable diffusion finetuning
- caption NSFW and SFW images without censorship
- describe anime, furry, and digital art images
- replace paid captioning services like ChatGPT for image description
- run an uncensored vision language model locally

## When to choose
- you need descriptive image captions for text-to-image model training
- you want an open, uncensored alternative to GPT-4o for captioning
- you need coverage of diverse image styles including anime and digital art
- you want to inspect or reproduce the model training pipeline

## When to avoid
- you need a general-purpose chat or reasoning VLM
- you require a small CPU-only model for edge devices
- you need a fully moderated, safety-filtered captioning service
- you want a hosted API rather than self-hosted inference

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, image-processing, nlp, llm-inference
- domain: machine-learning, artificial-intelligence, image-processing, large-language-models
- platform: python, cross-platform
- tags: image-captioning, vision-language-model, vlm, diffusion-model-training, uncensored, stable-diffusion, llava, gpu

## Member repositories
- fpgaminer/joycaption (main) score 53

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:06.886327+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T05:08:14.870680+00:00, confidence not recorded.
  - readme: https://github.com/fpgaminer/joycaption (fetched 2026-08-28T04:04:06.886327+00:00, sha 47f5631c9840)
- Data as of 2026-08-30T08:39:29.467469+00:00.
