# X-PLUG/mPLUG-Owl

mPLUG-Owl: The Powerful Multi-modal Large Language Model  Family

Repository: https://github.com/X-PLUG/mPLUG-Owl
Canonical: https://ross.abutalabs.com/products/mplug-owl
Homepage: https://www.modelscope.cn/studios/damo/mPLUG-Owl
Language: Python
License: MIT
License Family: permissive
Topics: chatbot, chatgpt, large-language-models, llama, multimodal, damo, mplug, instruction-tuning, pretraining, mplug-owl, huggingface, pytorch, transformer, alpaca, visual-recognition, gpt, gpt4, gpt4-api, dialogue, video
Last push: 2025-04-02T12:30:45+00:00

## Health v2 (maintenance only)
Score: 36/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 14, release rhythm 35, longevity 87
- inputs: {"age_days": 1227, "days_push": 518, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2539, forks 189 (observed 2026-08-28T04:06:59.267889+00:00)

## What it is
mPLUG-Owl is a family of open-source multimodal large language models that combine visual encoders with LLMs (built on LLaMA) for image and video understanding plus dialogue. It includes mPLUG-Owl, Owl2, and Owl3 with pretrained weights and training/inference code released on HuggingFace.

## Use cases
- chat with an AI about images
- build a multimodal chatbot that understands pictures
- run visual question answering on images and videos
- fine-tune a vision-language model on my own data
- understand long sequences of images with an LLM
- describe video content with a large language model
- deploy a GPT-4-like multimodal assistant locally

## When to choose
- you need an open-source multimodal LLM for image or video understanding
- you want to fine-tune or research vision-language models
- you need Chinese and English multimodal dialogue support
- you want CVPR-validated architectures with released weights

## When to avoid
- you only need text-only LLM inference
- you lack GPU resources for 7B-scale models
- you need a production-ready commercial API rather than research code
- you need audio or speech modality support

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, llm-inference, chatbot, image-processing, video-processing, nlp
- domain: large-language-models, artificial-intelligence, computer-vision, machine-learning
- platform: python, cross-platform
- tags: multimodal, vision-language-model, instruction-tuning, huggingface, pytorch, llama, image-understanding, video-understanding, mplug-owl, modelscope, natural-language-processing, gpu

## Member repositories
- X-PLUG/mPLUG-Owl (main) score 36

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:59.267889+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:25:09.879215+00:00, confidence not recorded.
  - readme: https://github.com/X-PLUG/mPLUG-Owl (fetched 2026-08-28T04:06:59.267889+00:00, sha 0cf8abbdd149)
  - homepage: https://www.modelscope.cn/studios/damo/mPLUG-Owl (fetched 2026-08-29T10:07:14.164832+00:00, sha 473f667bb4f3)
- Data as of 2026-08-30T08:39:29.467469+00:00.
