ByteDance-Seed/Seed1.5-VL resource
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks. observed · 2026-08-28
Health v2 · maintenance only
31/100
- Activity 26
- Release rhythm 35
- Longevity 34
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 479
- days_rel: n/a
- days_push: 445
- n_releases_24m: 0
Adoption not part of the score
1586 stars · 66 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Seed1.5-VL is ByteDance's vision-language foundation model (532M vision encoder + 20B active-parameter MoE LLM) for general-purpose multimodal understanding and reasoning, state-of-the-art on 38 of 60 public benchmarks. This repository is the official cookbook with code samples and best practices for using the model via its hosted API on Volcano Engine.
Use cases
- understand images and diagrams with a vision-language model
- extract text from images with ocr
- build gui agents that control apps from screenshots
- analyze video content with a multimodal model
- ground and locate objects in 2d and 3d scenes
- solve visual reasoning puzzles with an ai model
When to choose
- you want to integrate a top-performing vision-language model via API
- you need cookbook examples for grounding, video understanding, or GUI agents
- you want to evaluate a strong multimodal model against benchmarks
When to avoid
- you need to run the model weights locally - the repo only provides API usage examples
- you want a self-hosted or fine-tunable open-weights model
- you need a model without dependency on Volcano Engine's paid API
Facets
learning-resource · maturity active
machine-learning nlp computer-vision ocr sdk large-language-models computer-vision artificial-intelligence tutorials python cloud vision-language-model multimodal cookbook moe foundation-model gui-agents video-understanding visual-grounding ai-agents web-server
2 sources
- readme: https://github.com/ByteDance-Seed/Seed1.5-VL · fetched 2026-08-28 · d54d298ba9ea
- homepage: https://seed.bytedance.com/en/tech/seed1_5_vl · fetched 2026-08-29 · 605f70939a47
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| ByteDance-Seed/Seed1.5-VL | main | 31 |
For agents
markdown · JSON · MCP: product_card(name="ByteDance-Seed/Seed1.5-VL")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem