#1, then #1 again.

Introducing Sonic-3.6 and Ink-2.

Build your entire voice stack with one model provider - the only one ranked #1 on both speech and transcription. Don't compromise on quality or speed.

The full stack for interactive intelligence.

Ink speech-to-text and Sonic text-to-speech bracket your agent config, LLM, API, guardrails, and tools in one real-time pipeline

Co-designed end to end for voice agents

The only STT and TTS optimized across the full real-time pipeline.

One API, no assembly required

Ship both models in one integration — less vendor stitching, more building.

The tightest loop in voice

Hit sub-90ms TTS and 100ms transcript latency with native turn detection.

Cartesia’s state-space models bring enterprise-grade speed and quality to our AI Voice Agents… making it possible for businesses to deploy secure, scalable voice agents that can understand, act, and adapt in real time.

Ravi Krishnamurthy, VP Product

At Cartesia, we believe the tradeoffs that define today’s voice AI

Speed versus Naturalness,

Accuracy versus Cost,

are largely architectural in origin, not inevitable.

We’ve spent years building and scaling State Space Models because we believe the right primitives eliminate constraints rather than work around them.

And we built Sonic-3.6 and Ink-2 not by optimizing within accepted limits, but by questioning whether those limits need to exist at all.

Build with the fastest models you can trust.

Our models are designed for live, synchronous interactions, built on State Space Models (SSMs). A new primitive for large-scale foundation models, SSMs deliver ultra-low latency, long-context reasoning, and greater efficiency at scale.

Architecting AI that learns and interacts like humans.

Status