# OpenBMB/VisCPM

[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列

Repository: https://github.com/OpenBMB/VisCPM
Canonical: https://ross.abutalabs.com/products/viscpm
Language: Python
License Family: other
Topics: diffusion-models, large-language-models, multimodal, transformers
Last push: 2024-06-13T14:02:54+00:00

## Health v2 (maintenance only)
Score: 29/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 82
- inputs: {"age_days": 1160, "days_push": 811, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1062, forks 88 (observed 2026-08-28T04:03:26.093556+00:00)

## What it is
VisCPM is a family of open-source bilingual (Chinese/English) multimodal large models built on the 10B CPM-Bee language model, comprising VisCPM-Chat for image-to-text multimodal conversation and VisCPM-Paint for text-to-image generation. It fuses a Muffin visual encoder and a Diffusion-UNet decoder and achieved spotlight status at ICLR 2024.

## Use cases
- build a multimodal chatbot that understands images in Chinese and English
- generate images from Chinese text prompts
- run bilingual vision-language question answering
- research cross-lingual transfer in multimodal models
- deploy a web demo for image conversation
- integrate multimodal model inference via API

## When to choose
- you need strong Chinese-language multimodal understanding or generation
- you want open-source image-to-text and text-to-image models from one family
- you are researching bilingual generalization in multimodal LLMs

## When to avoid
- you need the latest state-of-the-art multimodal performance (the team recommends MiniCPM-V or OmniLMM)
- you require a permissive license for commercial use (no license is specified)
- you lack GPU resources for 10B-parameter model inference

## Facets
- artifact type: library
- maturity: maintenance
- function: machine-learning, deep-learning, llm-inference, image-processing, chatbot, stable-diffusion
- domain: large-language-models, computer-vision, artificial-intelligence, image-processing
- platform: python
- tags: multimodal, text-to-image, bilingual, chinese, vision-language-model, cpm-bee, iclr-2024, natural-language-processing, gpu, linux, docker

## Member repositories
- OpenBMB/VisCPM (main) score 29

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:26.093556+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:56:42.588234+00:00, confidence not recorded.
  - readme: https://github.com/OpenBMB/VisCPM (fetched 2026-08-28T04:03:26.093556+00:00, sha b7b713a8245e)
- Data as of 2026-08-30T08:39:29.467469+00:00.
