Ross ROSS = Recommend OSS · open-source software intelligence for agents

Picsart-AI-Research/StreamingT2V

[CVPR 2025] StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text observed · 2026-08-28

github.com/Picsart-AI-Research/StreamingT2V · homepage · Python observed · 2026-08-28

Health v2 · maintenance only

31/100

  • Activity 13
  • Release rhythm 35
  • Longevity 64

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 898
  • days_rel: n/a
  • days_push: 524
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1630 stars · 157 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

StreamingT2V is a research codebase (CVPR 2025) implementing an autoregressive technique that turns short text-to-video diffusion models like SVD and ModelScope into long, temporally consistent video generators. It supports generating videos of 200 to 1200+ frames with smooth transitions, rich motion, and text/image alignment.

Use cases

  • generate long videos from a text prompt
  • extend a short text-to-video model to minutes-long output
  • generate video from an image with consistent motion
  • produce temporally consistent AI video without hard cuts
  • run diffusion-based video generation on a single GPU
  • experiment with autoregressive video diffusion research

When to choose

  • you need long (8s to 2min+) AI-generated videos with smooth transitions
  • you want to build on SVD or ModelScope as a base video model
  • you have a CUDA GPU with at least 24 GB VRAM and can run Python research code
  • you're reproducing or extending CVPR 2025 video generation research

When to avoid

  • you need a production-ready product with a polished UI or API
  • you only need short 16-24 frame clips, where standard T2V models suffice
  • you lack a high-VRAM GPU (default config needs 60 GB)
  • you require a permissive or clearly defined license - the repo has none
  • you need real-time or low-latency video generation

Facets

library · maturity active

video-processing machine-learning deep-learning llm-inference artificial-intelligence deep-learning image-processing python text-to-video image-to-video diffusion-models long-video-generation autoregressive stable-video-diffusion cvpr-2025 research-code huggingface video linux gpu

2 sources

Member repositories

RepositoryRoleHealth v2
Picsart-AI-Research/StreamingT2Vmain31

For agents

markdown · JSON · MCP: product_card(name="Picsart-AI-Research/StreamingT2V")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem