Skip to main content
Fish Audio provides speech-to-text with automatic language detection. Buffers audio per participant (minimum 1 second) before sending to the API for accurate transcription.
Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.
Fish Audio also provides high-quality text-to-speech with prosody control and voice cloning. You can use both in the same agent.

Installation

Quick Start

Set FISH_API_KEY in your environment or pass api_key directly.

Parameters

Next Steps

Fish Audio TTS

Text-to-speech with prosody control

Build a Voice Agent

Get started with voice