# OpenBMB/MiniCPM-V

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Repository: https://github.com/OpenBMB/MiniCPM-V
Canonical: https://ross.abutalabs.com/products/minicpm-v
Language: Python
License: Apache-2.0
License Family: permissive
Topics: minicpm, minicpm-v, multi-modal, minicpm-o
Last push: 2026-08-26T09:50:07+00:00

## Health v2 (maintenance only)
Score: 61/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 8, longevity 67
- inputs: {"age_days": 947, "days_push": 7, "days_rel": 463, "gap_med": null, "n_releases_24m": 1}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 26240, forks 2055 (observed 2026-08-28T04:11:38.926292+00:00)

## What it is
MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deployable on phones and edge devices. The repository provides model weights, inference code, and edge adaptation tooling for iOS, Android, and HarmonyOS.

## Use cases
- run a vision-language model on a phone
- understand images and video with a small LLM
- build an on-device multimodal chat assistant
- real-time streaming video and speech interaction
- deploy an efficient MLLM on edge devices
- caption or answer questions about images offline

## When to choose
- you need multimodal (image/video/audio) understanding on resource-constrained devices
- you want an open Apache-2.0 small MLLM with strong benchmarks
- you target mobile platforms like iOS, Android, or HarmonyOS

## When to avoid
- you need the highest possible accuracy from frontier-scale models
- you only need text-only LLM inference
- you cannot run any local inference and prefer cloud APIs

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, llm-inference, computer-vision, speech-recognition, rag
- domain: artificial-intelligence, large-language-models, computer-vision, image-processing, mobile-development
- platform: python, cross-platform
- tags: multimodal, vision-language-model, on-device-ai, edge-deployment, video-understanding, omnimodal, video, android, ios, mobile, gpu

## Member repositories
- OpenBMB/MiniCPM-V (main) score 61

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:11:38.926292+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T16:55:54.625112+00:00, confidence not recorded.
  - readme: https://github.com/OpenBMB/MiniCPM-V (fetched 2026-08-28T04:11:38.926292+00:00, sha 6ad8c606a54c)
- Data as of 2026-08-30T08:39:29.467469+00:00.
