# TencentARC/SEED-Voken

SEED-Voken: A Series of Powerful Visual Tokenizers

Repository: https://github.com/TencentARC/SEED-Voken
Canonical: https://ross.abutalabs.com/products/seed-voken
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2025-11-25T07:28:22+00:00

## Health v2 (maintenance only)
Score: 48/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 54, release rhythm 35, longevity 58
- inputs: {"age_days": 812, "days_push": 281, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1021, forks 46 (observed 2026-08-28T04:03:15.746699+00:00)

## What it is
SEED-Voken is a collection of visual tokenizers (Open-MAGVIT2 and IBQ) that convert images and videos into discrete tokens for autoregressive visual generation. It provides pretrained tokenizers with large codebooks and high codebook utilization for image and video generation research.

## Use cases
- tokenize images into discrete codes for autoregressive image generation
- train a visual tokenizer with a large codebook for image synthesis
- convert video frames into tokens for video generation models
- reproduce Open-MAGVIT2 or IBQ research results
- use pretrained tokenizers for text-conditional image generation

## When to choose
- you need state-of-the-art visual tokenization for autoregressive image or video generation
- you want large codebook sizes with high codebook utilization
- you are doing research on image tokenization or visual generation

## When to avoid
- you need a production-ready image generation service rather than a tokenizer
- you work outside PyTorch/GPU research environments
- you need diffusion-based generation rather than token-based autoregressive generation

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, image-processing, serialization
- domain: deep-learning, computer-vision, image-processing, large-language-models
- platform: python
- tags: visual-tokenizer, image-tokenization, autoregressive-generation, vq-quantization, video-tokenizer, research, gpu

## Member repositories
- TencentARC/SEED-Voken (main) score 48

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:15.746699+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:08:49.110710+00:00, confidence not recorded.
  - readme: https://github.com/TencentARC/SEED-Voken (fetched 2026-08-28T04:03:15.746699+00:00, sha e044eb184148)
- Data as of 2026-08-30T08:39:29.467469+00:00.
