GetStream/Vision-Agents
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency. observed · 2026-08-28
Health v2 · maintenance only
84/100
- Activity 99
- Release rhythm 97
- Longevity 27
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 4.0
- age_days: 387
- days_rel: 20
- days_push: 7
- n_releases_24m: 59
Adoption not part of the score
8100 stars · 680 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (OpenAI, Gemini, Deepgram, ElevenLabs, YOLO, Roboflow) and uses Stream's WebRTC edge network for sub-500ms join and sub-30ms audio/video latency.
Use cases
- build a real-time voice agent that talks to users in the browser
- create an AI sports or golf coach using YOLO pose detection and Gemini
- build a smart security camera with face recognition and package detection
- create a phone support agent that answers inbound Twilio calls with RAG
- build a live sports commentator that tracks players and generates play-by-play
- create an interactive lip-synced avatar that sees and hears users
- run object detection models on live video frames in real time
When to choose
- you need real-time multimodal agents that process live audio and video with low latency
- you want to mix vision models like YOLO or Roboflow with LLMs like Gemini or OpenAI in one pipeline
- you want provider-agnostic STT, TTS, LLM, and vision plugins with a consistent interface
- you need client SDKs across React, iOS, Android, Flutter, React Native, or Unity
- you want to build telehealth, live coaching, or voice support agents quickly
When to avoid
- you only need offline or batch video analysis with no real-time interaction
- you want to avoid Stream's edge network or paid services and have no WebRTC infrastructure of your own
- you need a simple text-only chatbot without voice or video
- you require a non-Python backend language for your agent logic
Facets
framework · maturity active
agent-framework speech-recognition tts computer-vision llm-inference rag chatbot sdk artificial-intelligence computer-vision speech-processing large-language-models python cross-platform cloud voice-ai video-ai realtime-agents webrtc stt tts multimodal stream-video yolo telehealth ai-agents video real-time web-server docker
5 sources
- readme: https://github.com/GetStream/Vision-Agents · fetched 2026-08-28 · a26f242a36f6
- homepage: https://visionagents.ai · fetched 2026-08-29 · 499f0e4b66f8
- site_page: https://visionagents.ai/introduction/quickstart · fetched 2026-08-29 · 2146b53149f6
- registry_pypi: https://pypi.org/pypi/vision-agents/json · fetched 2026-08-29 · 12f26363d708
- site_page: https://visionagents.ai/integrations/introduction-to-integrations · fetched 2026-08-29 · 9d58993b2abc
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| GetStream/Vision-Agents | main | 84 |
For agents
markdown · JSON · MCP: product_card(name="GetStream/Vision-Agents")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem