A Guide to Choosing Voice AI Models

A Guide to Choosing Voice AI Models

A practical framework for evaluating ASR, TTS, and turn-detection models against your real-world use case — not lab conditions.

[Product]

A Beginner's Guide to Voice AI Terminology

A Beginner's Guide to Voice AI Terminology

A plain-language glossary of the 20 key terms you need to understand, evaluate, and build conversational voice AI.

[Product]

GDPR logo next to text, Cartesia is now GDPR compliant

Cartesia achieves GDPR compliance

Cartesia's text-to-speech platform is now GDPR compliant, reaffirming our commitment to data protection, privacy, and responsible human-centered AI.

[News]

Introducing Line: The Modern Voice Agent Development Platform

Introducing Line: The Modern Voice Agent Development Platform

Introducing Line by Cartesia: the modern voice agent development platform. Line was built to be code-first, because best-in-class products are built in code.

[Product]

Hierarchical modeling

Hierarchical modeling

Introducing H-Nets, hierarchical architectures that learn directly from raw data by grouping inputs into higher-level concepts, overcoming tokenization limits.

[Research]

Introducing Ink: speech-to-text models for real-time conversation

Introducing Ink: speech-to-text models for real-time conversation

Today we’re introducing Ink, a new family of streaming speech-to-text (STT) models for developers building real-time voice applications. Ink-Whisper is the fastest, most affordable STT model–designed for enterprise-grade voice agents.

[Product]

Introducing Organizations and Dashboards

Introducing Organizations and Dashboards

We’re building Cartesia for developers scaling voice AI. Today, we’re introducing two features to make collaboration and visibility easier: Organizations and Dashboards.

[Product]

Introducing Professional Voice Cloning

Introducing Professional Voice Cloning

Introducing Professional Voice Cloning (PVC), professional-quality voice clones trained on Sonic, now available on Startup+ plans.

[Product]

Cartesia Python SDK v2.0.0

Cartesia Python SDK v2.0.0

We are excited to announce the release of v2.0.0 of our Python SDK, polishing the developer experience when using Cartesia's AI voice capabilities with Python.

[Engineering]

Cartesia Named to 7th Annual Enterprise Tech 30 List Presented by Wing Venture Capital

Cartesia Named to 7th Annual Enterprise Tech 30 List Presented by Wing Venture Capital

Cartesia, a leading provider of generative audio models,  today announced it has been named to the seventh annual Enterprise Tech 30—a definitive list of the most promising, private enterprise tech companies across all stages of maturity.

[News]

Introducing Narrations: create and edit long-form audio content with precision

Introducing Narrations: create and edit long-form audio content with precision

Today we're excited to introduce Narrations, a platform that enables creators to transform written content into polished audio productions with unprecedented control and efficiency.

[Product]

Series A and the future of voice AI

Series A and the future of voice AI

We’re thrilled to announce our $64 million Series A led by Kleiner Perkins. The new funding will help us expand our team and invest in research to build the next generation of models.

[News]

Llamba: scaling distilled recurrent models for efficient language processing

Llamba: scaling distilled recurrent models for efficient language processing

The next few years will usher in a new era of on-device AI. On-device models will power a wide range of applications.

[Research]

How to build a voice AI agent with Cartesia

How to build a voice AI agent with Cartesia

A step-by-step guide to building a natural, low-latency voice AI agent with Cartesia's Sonic text-to-speech, from architecture to production deployment.

[Engineering]

State of voice AI 2024

State of voice AI 2024

In our first 2024 State of Voice report, we highlight the key infrastructure breakthroughs and emerging use cases driving the industry forward, and look ahead to what’s next in 2025.

[Research]

Announcing our seed round

Announcing our seed round

We’re excited to announce our $27M seed round, led by Index Ventures with participation from Lightspeed, Factory, Conviction, General Catalyst, A*, SV Angel, and 90 amazing angel investors.

[News]

‘Tis the Hackathon season at Cartesia

‘Tis the Hackathon season at Cartesia

October 2024 was a busy month for Cartesia. We brought together over 2,000 builders across San Francisco and gave away $20,000 in prizes for the most innovative ideas built on Sonic.

[Engineering]

Introducing voice changer: transform audio your way

Introducing voice changer: transform audio your way

Voice Changer transforms the voice of any audio clip while preserving its original delivery, emotion, and prosody — with studio-quality voices or clones.

[Product]

Introducing our next 8 languages on Sonic Multilingual

Introducing our next 8 languages on Sonic Multilingual

Today, we're excited to announce the Alpha Release of our next 8 languages—Hindi, Italian, Korean, Dutch, Polish, Russian, Swedish, and Turkish—on Sonic Multilingual.

[Product]

The on-device intelligence update

The on-device intelligence update

At Cartesia, our mission is to build the next generation of AI: ubiquitous, interactive intelligence that runs wherever you are.

[Research]

Announcing Sonic: a low‑latency voice model for lifelike speech

Announcing Sonic: a low‑latency voice model for lifelike speech

We're releasing Sonic, our low-latency voice model that generates lifelike speech today.

[Research]

Based: Simple linear attention language models balance the recall‑throughput tradeoff

Based: Simple linear attention language models balance the recall‑throughput tradeoff

Based is 56% and 44% faster at processing prompts than FlashAttention-2 and Mamba respectively. Based achieves 24x higher text generation throughput than FlashAttention-2.

[Research]

Mamba‑3B-SlimPJ: State-space models rivaling the best Transformer architecture

Mamba‑3B-SlimPJ: State-space models rivaling the best Transformer architecture

We're releasing the strongest Mamba language model yet, Mamba-3B-SlimPJ, in partnership with Cartesia & Together under an Apache 2.0 license.

[Research]

Get started today

Talk to an expert. Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.

Contact Sales

Start building. Access our models via API and bring an agent into production with our robust SDKs and developer tools.

Try Cartesia