Skip to main content
Fast-Whisper is a high-performance local STT using CTranslate2. Provides 2-4x faster inference than standard Whisper with support for CPU and GPU acceleration.
Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.

Installation

Quick Start

Fast-Whisper runs locally. No API key required. Models download automatically on first use.

Parameters

Model Sizes

Optimization

Next Steps

Build a Voice Agent

Get started with voice

Build a Video Agent

Add video processing