Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.
Installation
Detection (Cloud)
Detection (Local)
Runs on-device without API calls. RequiresHF_TOKEN for model access.
VLM (Cloud)
Visual question answering or automatic captioning.VLM (Local)
Cloud vs Local
Local models require
HF_TOKEN for HuggingFace authentication. CUDA
recommended for best performance.Next Steps
Build a Voice Agent
Get started with voice
Build a Video Agent
Add video processing