Ross ROSS = Recommend OSS · open-source software intelligence for agents

OpenBMB/MiniCPM-V

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone observed · 2026-08-28

github.com/OpenBMB/MiniCPM-V · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

61/100

  • Activity 99
  • Release rhythm 8
  • Longevity 67
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 947
  • days_rel: 463
  • days_push: 7
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

26240 stars · 2055 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

MiniCPM-V and MiniCPM-o are a series of small multimodal large language models for efficient image, video, and audio understanding, deployable on phones and edge devices. The repository provides model weights, inference code, and edge adaptation tooling for iOS, Android, and HarmonyOS.

Use cases

  • run a vision-language model on a phone
  • understand images and video with a small LLM
  • build an on-device multimodal chat assistant
  • real-time streaming video and speech interaction
  • deploy an efficient MLLM on edge devices
  • caption or answer questions about images offline

When to choose

  • you need multimodal (image/video/audio) understanding on resource-constrained devices
  • you want an open Apache-2.0 small MLLM with strong benchmarks
  • you target mobile platforms like iOS, Android, or HarmonyOS

When to avoid

  • you need the highest possible accuracy from frontier-scale models
  • you only need text-only LLM inference
  • you cannot run any local inference and prefer cloud APIs

Facets

library · maturity active

machine-learning llm-inference computer-vision speech-recognition rag artificial-intelligence large-language-models computer-vision image-processing mobile-development python cross-platform multimodal vision-language-model on-device-ai edge-deployment video-understanding omnimodal video android ios mobile gpu

1 source

Member repositories

RepositoryRoleHealth v2
OpenBMB/MiniCPM-Vmain61

For agents

markdown · JSON · MCP: product_card(name="OpenBMB/MiniCPM-V")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem