Skip to main content
Cartesia provides low-latency text-to-speech with the Sonic model. Designed for real-time voice applications with natural-sounding speech synthesis.
Vision Agents uses Stream Video for real-time WebRTC transport by default. External WebRTC transports are supported as well. Most AI providers offer free tiers to get started.
Cartesia also provides low-latency speech-to-text. You can use both in the same agent.

Installation

Quick Start

Set CARTESIA_API_KEY in your environment or pass api_key directly.

Parameters

Next Steps

Build a Voice Agent

Get started with voice

Build a Video Agent

Add video processing