mbzuai-oryx/Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models. observed · 2026-08-28
Health v2 · maintenance only
45/100
- Activity 35
- Release rhythm 35
- Longevity 85
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1203
- days_rel: n/a
- days_push: 394
- n_releases_24m: 0
Adoption not part of the score
1506 stars · 128 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Video-ChatGPT is a video conversation model that combines large language models with a pretrained visual encoder adapted for spatiotemporal video representation, enabling detailed conversation about videos. It also ships a quantitative evaluation benchmarking framework (VCGBench) for assessing video-based conversational models.
Use cases
- chat with a model about a video's content
- generate detailed video descriptions and captions
- zero-shot video question answering
- benchmark video conversational models quantitatively
- build a multimodal video understanding assistant
- evaluate temporal understanding of video LLMs
When to choose
- you need a research-grade video-language conversation model
- you want a standardized benchmark for video chat models
- you need zero-shot video QA on datasets like MSVD, MSRVTT, TGIF, or ActivityNet
When to avoid
- you need production-ready, low-latency video analysis at scale
- you lack GPU resources for inference and training
- you only need image (not video) understanding
Facets
library · maturity active
machine-learning video-processing chatbot llm-inference benchmarking artificial-intelligence computer-vision large-language-models python multimodal vision-language-model video-understanding video-qa video-conversation evaluation-benchmark research video natural-language-processing gpu linux
2 sources
- readme: https://github.com/mbzuai-oryx/Video-ChatGPT · fetched 2026-08-28 · f8dd666e4b74
- homepage: https://mbzuai-oryx.github.io/Video-ChatGPT · fetched 2026-08-29 · 1f7658b62290
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| mbzuai-oryx/Video-ChatGPT | main | 45 |
For agents
markdown · JSON · MCP: product_card(name="mbzuai-oryx/Video-ChatGPT")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem