Ross ROSS = Recommend OSS · open-source software intelligence for agents

PKU-YuanGroup/Video-LLaVA

【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection observed · 2026-08-28

github.com/PKU-YuanGroup/Video-LLaVA · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

27/100

  • Activity 0
  • Release rhythm 35
  • Longevity 74

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1045
  • days_rel: n/a
  • days_push: 638
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3500 stars · 255 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Video-LLaVA is a large vision-language model that aligns image and video representations into a unified visual space before projection into the language model, enabling joint image and video understanding. It is an EMNLP 2024 research codebase with pretrained checkpoints, training scripts, and demo spaces.

Use cases

  • answer questions about videos with an llm
  • video question answering model
  • chat with images and videos using a vision-language model
  • run multimodal llm inference on video
  • fine-tune a video-language model with instruction tuning
  • unified image and video understanding model

When to choose

  • you need a single model that handles both image and video inputs
  • you want an open-source research baseline for video-language alignment
  • you need video question answering or captioning with an LLM

When to avoid

  • you need production-grade, commercially supported multimodal APIs
  • you lack a GPU for inference or training
  • you only need text-only LLM capabilities

Facets

library · maturity active

machine-learning deep-learning llm-inference video-processing image-processing nlp large-language-models computer-vision artificial-intelligence python multimodal vision-language-model video-understanding instruction-tuning video-qa research video natural-language-processing gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
PKU-YuanGroup/Video-LLaVAmain27

For agents

markdown · JSON · MCP: product_card(name="PKU-YuanGroup/Video-LLaVA")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem