Ross ROSS = Recommend OSS · open-source software intelligence for agents

Vision-CAIR/MiniGPT-4

Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/) observed · 2026-08-28

github.com/Vision-CAIR/MiniGPT-4 · homepage · Python · BSD-3-Clause (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 0
  • Release rhythm 35
  • Longevity 88

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1236
  • days_rel: n/a
  • days_push: 730
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

25627 stars · 2868 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a single projection layer. It enables multimodal chat capabilities like detailed image description, visual question answering, and grounded image-text generation.

Use cases

  • generate detailed descriptions of images with an open-source model
  • build a visual chatbot that answers questions about images
  • run multimodal vision-language research experiments
  • caption images using a GPT-4-like open model
  • fine-tune a vision-language model on custom image-text data
  • create websites or stories from handwritten drafts and images

When to choose

  • you want an open-source, computationally efficient GPT-4-style multimodal model
  • you need a research baseline for vision-language multi-task learning
  • you want to fine-tune only a small projection layer instead of a full LLM
  • you need reproducible code from a published paper with pretrained checkpoints

When to avoid

  • you need a production-grade, actively maintained multimodal model
  • you lack a GPU or cannot run large language models locally
  • you want the latest multimodal capabilities rather than a 2023-era research model
  • you need commercial licensing beyond BSD-3-Clause terms for model weights

Facets

library · maturity maintenance

machine-learning deep-learning llm-inference llm-training image-processing nlp artificial-intelligence large-language-models computer-vision deep-learning machine-learning python vision-language-model multimodal image-captioning visual-question-answering research-code vicuna gradio-demo gpu linux docker

2 sources

Member repositories

RepositoryRoleHealth v2
Vision-CAIR/MiniGPT-4main30

For agents

markdown · JSON · MCP: product_card(name="Vision-CAIR/MiniGPT-4")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem