Vision-CAIR/MiniGPT-4
Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/) observed · 2026-08-28
Health v2 · maintenance only
30/100
- Activity 0
- Release rhythm 35
- Longevity 88
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1236
- days_rel: n/a
- days_push: 730
- n_releases_24m: 0
Adoption not part of the score
25627 stars · 2868 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Official code for MiniGPT-4 and MiniGPT-v2, vision-language models that align a frozen visual encoder with a frozen LLM (Vicuna) via a single projection layer. It enables multimodal chat capabilities like detailed image description, visual question answering, and grounded image-text generation.
Use cases
- generate detailed descriptions of images with an open-source model
- build a visual chatbot that answers questions about images
- run multimodal vision-language research experiments
- caption images using a GPT-4-like open model
- fine-tune a vision-language model on custom image-text data
- create websites or stories from handwritten drafts and images
When to choose
- you want an open-source, computationally efficient GPT-4-style multimodal model
- you need a research baseline for vision-language multi-task learning
- you want to fine-tune only a small projection layer instead of a full LLM
- you need reproducible code from a published paper with pretrained checkpoints
When to avoid
- you need a production-grade, actively maintained multimodal model
- you lack a GPU or cannot run large language models locally
- you want the latest multimodal capabilities rather than a 2023-era research model
- you need commercial licensing beyond BSD-3-Clause terms for model weights
Facets
library · maturity maintenance
machine-learning deep-learning llm-inference llm-training image-processing nlp artificial-intelligence large-language-models computer-vision deep-learning machine-learning python vision-language-model multimodal image-captioning visual-question-answering research-code vicuna gradio-demo gpu linux docker
2 sources
- readme: https://github.com/Vision-CAIR/MiniGPT-4 · fetched 2026-08-28 · df2a5b07b302
- homepage: https://minigpt-4.github.io · fetched 2026-08-29 · d35c20cdc67e
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Vision-CAIR/MiniGPT-4 | main | 30 |
For agents
markdown · JSON · MCP: product_card(name="Vision-CAIR/MiniGPT-4")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem