Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.
Installation
LLM
Text-only language model with streaming and function calling.VLM
Vision language model with automatic video frame buffering and function calling. Supports models like Qwen2-VL.Next Steps
Build a Voice Agent
Get started with voice
Build a Video Agent
Add video processing