Ross ROSS = Recommend OSS · open-source software intelligence for agents

GetStream/Vision-Agents

Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency. observed · 2026-08-28

github.com/GetStream/Vision-Agents · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

84/100

  • Activity 99
  • Release rhythm 97
  • Longevity 27
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 4.0
  • age_days: 387
  • days_rel: 20
  • days_push: 7
  • n_releases_24m: 59

Full methodology

Adoption not part of the score

8100 stars · 680 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

An open-source Python framework by Stream for building low-latency real-time voice and video AI agents. It provides 35+ provider plugins (OpenAI, Gemini, Deepgram, ElevenLabs, YOLO, Roboflow) and uses Stream's WebRTC edge network for sub-500ms join and sub-30ms audio/video latency.

Use cases

  • build a real-time voice agent that talks to users in the browser
  • create an AI sports or golf coach using YOLO pose detection and Gemini
  • build a smart security camera with face recognition and package detection
  • create a phone support agent that answers inbound Twilio calls with RAG
  • build a live sports commentator that tracks players and generates play-by-play
  • create an interactive lip-synced avatar that sees and hears users
  • run object detection models on live video frames in real time

When to choose

  • you need real-time multimodal agents that process live audio and video with low latency
  • you want to mix vision models like YOLO or Roboflow with LLMs like Gemini or OpenAI in one pipeline
  • you want provider-agnostic STT, TTS, LLM, and vision plugins with a consistent interface
  • you need client SDKs across React, iOS, Android, Flutter, React Native, or Unity
  • you want to build telehealth, live coaching, or voice support agents quickly

When to avoid

  • you only need offline or batch video analysis with no real-time interaction
  • you want to avoid Stream's edge network or paid services and have no WebRTC infrastructure of your own
  • you need a simple text-only chatbot without voice or video
  • you require a non-Python backend language for your agent logic

Facets

framework · maturity active

agent-framework speech-recognition tts computer-vision llm-inference rag chatbot sdk artificial-intelligence computer-vision speech-processing large-language-models python cross-platform cloud voice-ai video-ai realtime-agents webrtc stt tts multimodal stream-video yolo telehealth ai-agents video real-time web-server docker

5 sources

Member repositories

RepositoryRoleHealth v2
GetStream/Vision-Agentsmain84

For agents

markdown · JSON · MCP: product_card(name="GetStream/Vision-Agents")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem