# DAMO-NLP-SG/VideoLLaMA3

Frontier Multimodal Foundation Models for Image and Video Understanding

Repository: https://github.com/DAMO-NLP-SG/VideoLLaMA3
Canonical: https://ross.abutalabs.com/products/videollama3
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Last push: 2025-08-14T06:52:30+00:00

## Health v2 (maintenance only)
Score: 37/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 36, release rhythm 35, longevity 42
- inputs: {"age_days": 591, "days_push": 384, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1179, forks 89 (observed 2026-08-28T04:03:53.170860+00:00)

## What it is
VideoLLaMA 3 is a frontier multimodal foundation model for image and video understanding, released with checkpoints, inference code, and demos. It is a research model from DAMO NLP-SG built on vision-language modeling for video and image QA.

## Use cases
- understand and answer questions about videos
- video question answering with an LLM
- image understanding with a multimodal model
- analyze long videos frame by frame
- build a video chatbot
- run a vision-language model on GPU

## When to choose
- you need open-weights video/image understanding with an LLM
- you want a research-grade multimodal model with demos and checkpoints
- you need Apache-2.0 licensed video LLM code

## When to avoid
- you need lightweight CPU-only inference
- you need production video analytics pipelines rather than a research model
- you need audio understanding (see VideoLLaMA 2)

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, llm-inference, computer-vision, video-processing, image-processing, nlp
- domain: artificial-intelligence, large-language-models, computer-vision, image-processing, deep-learning
- platform: python, cross-platform
- tags: multimodal, video-llm, vision-language-model, video-understanding, foundation-model, huggingface, research, video, natural-language-processing, gpu, linux

## Member repositories
- DAMO-NLP-SG/VideoLLaMA3 (main) score 37

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:53.170860+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:25:26.105763+00:00, confidence not recorded.
  - readme: https://github.com/DAMO-NLP-SG/VideoLLaMA3 (fetched 2026-08-28T04:03:53.170860+00:00, sha 7922261ae20f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
