Skip to main content
Fish Audio provides high-quality text-to-speech with fine-grained prosody control, voice cloning support, and multiple backend models. Ideal for multilingual applications.
Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.
Fish Audio also provides speech-to-text with automatic language detection. You can use both in the same agent.

Installation

Quick Start

Set FISH_API_KEY in your environment or pass api_key directly.

Basic Usage

Prosody Control

The S2-Pro model (default) supports inline control tags for natural prosody:

Selecting a Model

Parameters

Next Steps

Fish Audio STT

Speech-to-text with auto language detection

Build a Voice Agent

Get started with voice