AI Voice Agents Mid-2026: gpt-realtime vs Modular Stack Decision Guide
A year ago, building a voice agent meant wiring Whisper to an LLM to a TTS API, and praying the round-trip didn’t kill the conversation. In mid-2026, two very different stacks have quietly won — and most teams are about to bet on the wrong one without realizing it.
OpenAI’s Realtime API, with the gpt-realtime model released to GA earlier this year, has turned “speech-to-speech in one model” into the default mental model for anyone new to voice. But the harder-won lesson from production teams is that monolithic speech-to-speech (S2S) is not the universal answer. Modular stacks — best-in-class STT, your favorite LLM, and a TTS that matches your brand — are still the right call for a large chunk of real workloads. The market did not converge. It split.
This guide walks through what actually changed in the last six months, what is still overhyped, and a decision framework for picking your stack. It is not a tutorial — there are plenty of those — it is closer to a field report from teams shipping voice in production.
What actually changed since late 2025
Three shifts moved the goalposts for voice-agent builders. None of them are about flashy demos. They are about the parts that quietly break in production.
Speech-to-speech models finally sound human
OpenAI’s gpt-realtime went GA with measurable jumps in three places that actually matter: audio quality, instruction following, and function calling. On internal benchmarks, the new model scores 82.8% on Big Bench Audio reasoning (up from 65.6% on the December 2024 version), 30.5% on MultiChallenge instruction adherence (up from 20.6%), and 66.5% on ComplexFuncBench function calling (up from 49.7%). It also reads alphanumerics — phone numbers, VINs, account IDs — more reliably in Spanish, Chinese, Japanese, and French. Two new voices (Marin and Cedar) shipped alongside, and existing voices got an expressiveness pass.
Equally important: OpenAI cut prices 20% at GA. Audio input now runs $32 per 1M tokens ($0.40 for cached input), output $64 per 1M tokens. For a one-minute call of back-and-forth, that is roughly a quarter to a third of what you would have paid 18 months ago. The “uncanny valley” speech problem is largely solved at the frontier — most customers can no longer tell from voice alone whether they are talking to a person or an agent.
STT quietly had its own GPT moment
The bigger story for most builders is that speech-to-text got really good. Cartesia’s Ink-2, released July 9 2026, currently tops the Artificial Analysis streaming leaderboard. On AppTek — a 14-accent call-center benchmark — Ink-2 hits 8% word error rate, versus Deepgram Flux’s 10% and ElevenLabs Scribe v2’s 12%. On Cartesia’s internal production-audio benchmark (real call-center calls with background noise and degraded network audio), Ink-2 is at 6.5%, versus 9.2% and 9.4% for the others.
The more important improvement is what Ink-2 calls semantic endpointing. Old-school voice agents ended turns based on silence thresholds, which meant cutting people off mid-address or sitting through dead air. Ink-2 reads meaning, not silence — it knows the user is still saying an email address and waits. It emits three events natively (turn.start, turn.eager_end, turn.end) without any separate VAD layer. That is the difference between a polished agent and one that feels like a 2023 IVR.
Modular stacks got composable
The single-model S2S pattern is seductive, but the modular STT→LLM→TTS pattern has its own counter-offensive. The pieces are now independently best-in-class. Deepgram shipped Flux Multilingual in April — 10 languages in one conversational STT model, with monolingual-grade accuracy. ElevenLabs v3 went GA (after alpha) with audio tags like [whispers] and [sighs] inline in the script, plus a new Text-to-Dialogue endpoint that handles speaker transitions and emotional shifts for you. Hume released Voice Control on EVI 2, exposing ten voice dimensions as continuous sliders — assertiveness, buoyancy, confidence, enthusiasm, nasality, relaxedness, smoothness, tepidity, tightness — without going anywhere near voice cloning.
The pieces snap together well. You can pipe Cartesia Ink-2 → Claude or GPT → ElevenLabs or Hume in an afternoon, and tune each layer to the failure mode of the others. That tuneability is what the monolithic stack still cannot give you.
The honest tradeoffs of each stack
A speech-to-speech model like gpt-realtime is doing five jobs at once: turn-taking, STT, intent understanding, response generation, and TTS. That buys you three things and costs you two.
You gain latency. The whole pipeline is one model hop. You can hit roughly 300–500 ms time-to-first-byte in a well-tuned deployment, which feels conversational. You gain natural turn-taking because the model knows when to stop talking the way humans do — something the modular stack still has to engineer carefully. You lose the seam between STT and LLM, which means emotion and “mm-hmm” cues survive transcription instead of being flattened.
You give up voice customization. You can pick from OpenAI’s preset voices or write voice-direction prompts, but you cannot clone a brand voice or adjust nasality by 30%. You give up per-language independent tuning, which matters if you need native-speaker voices per locale. And you give up price predictability — audio token billing is harder to forecast per minute than per-character TTS pricing.
A modular stack inverts that tradeoff. You can pick the most natural-sounding voice on the market for your brand. You can use a smaller, cheaper LLM (Haiku, Flash, or a fine-tuned 8B open model) for the intent layer. You can swap any piece independently — which matters when Ink-2 or ElevenLabs ships a 30% improvement overnight and you want to capture it next week. You also get clearer per-component observability, which is the difference between “the call felt weird” and “the LLM hit a 1.8-second function call.”
The cost is engineering. You are integrating three or four vendors, each with their own SDK quirks and auth, and you are on the hook for the latency budget across all of them. The teams who win with modular are the ones who treat the voice pipeline as a production system — with tracing, eval sets, regression tests for prompt changes, and a clear owner.
Decision framework: which stack when
Rather than abstract principles, here is the rule of thumb I use when a team asks me which way to go.
| Pick this | If your situation looks like this |
|---|---|
| Speech-to-speech (gpt-realtime, Gemini Live) | Mostly conversation with light tool use (one or two function calls per minute). You care more about naturalness than brand voice. You do not have a voice team and do not plan to hire one. Latency is a hard product constraint — telephony, IVR replacement, real-time tutoring. |
| Modular (Cartesia Ink-2 + LLM + ElevenLabs / Hume) | Voice is part of your brand. You need voice cloning or fine-grained emotional control. You need to swap STT or TTS independently for regional cost, regulatory, or latency reasons. You want a smaller, cheaper LLM at the intent layer. You need on-prem or a specific cloud region for compliance. |
| Hybrid: start S2S, carve out modular over time | You need a working v1 in days, not weeks. You expect to learn where the S2S stack breaks for your specific use case and pull components out as you go. This is the path most production voice agents I know have actually taken. |
There is a third path a lot of teams miss: start with S2S, then carve out components as you learn what fails. The conversational fluency of gpt-realtime is a great way to get to a working v1 in days. Once you see where it breaks — usually brand voice, complex multi-step tool chains, or a regulated workflow — you can pull TTS out, or the LLM out, or both. This is the path most production voice agents I know of have actually taken.
The real cost comparison (and why per-minute math is misleading)
The official pricing does not tell you what a call actually costs. Here is a back-of-envelope for a 3-minute customer-support call with moderate back-and-forth:
- OpenAI Realtime API, gpt-realtime: roughly 60–90K input tokens plus 30–45K output tokens of audio per call. At $32/$64 per 1M, that is $1.90–$5.75 per call, depending on conversation density. Add SIP/phone costs on top if you are connecting to the PSTN.
- Modular (Cartesia Ink-2 + Claude Sonnet + ElevenLabs Flash): roughly 800 transcribed tokens/sec, ~500 LLM tokens/sec, and ~700 characters of TTS per minute. At current per-unit pricing, you land at $0.20–$0.40 per minute for typical traffic, or $0.60–$1.20 for a 3-minute call. This is 3–5× cheaper than S2S for the same call.
The catch: that $0.60 does not include the engineering hours to keep the modular stack under 800 ms latency. Modulate those costs based on whether you already have a voice engineer on payroll.
If you are shipping under 10K minutes per month, the per-call delta does not justify the engineering. If you are at 1M+ minutes per month, the difference is real money — and the modular stack pays for a dedicated voice engineer in the first quarter.
The turn-taking problem nobody puts in the pitch deck
The most common production failure for voice agents in 2026 is not transcription error. It is turn-taking. Voice agents interrupt customers mid-address, sit through two seconds of dead air after a user finishes, or fail to interrupt the user when the user starts rambling. The failure happens at the VAD or endpointing layer, not the LLM.
If you go with gpt-realtime, you inherit the model’s turn-taking behavior, which is genuinely good but not customizable. If you go modular, your turn-taking quality is determined by your STT provider’s endpointing. This is why Ink-2’s semantic endpointing matters more than its 8% WER — it is solving the right problem. The metric you actually want is F1 against a human-labeled set of real customer calls, with separate precision (do not cut people off) and recall (do not leave dead air).
Practical takeaway: do not eval STT on accuracy alone. Eval it on a labeled set of real customer calls where someone has tagged where the speaker actually finished. Precision and recall predict whether your agent feels real far better than WER does.
What to skip in 2026
A few things are still being sold hard that I would push back on.
Multilingual S2S as a solved problem
It is not. Cross-language S2S works for high-resource languages in quiet conditions, but falls apart for code-switching, regional accents, or noisy environments. For real multilingual deployments, you want language-aware STT (Flux Multilingual, Ink-2 multilingual variants when they ship) feeding into a language-strong LLM and a TTS with native-speaker voices per locale.
Voice cloning as a feature
The legal and disclosure burden has gotten serious under the EU AI Act and a handful of US state laws. Hume’s Voice Control and ElevenLabs’ Voice Design get you 80% of the brand-voice benefit without cloning. Default to those unless you have a specific, lawful use case.
Latency under 200 ms
Anything under 250 ms time-to-first-byte is bragging rights on a leaderboard, not a customer experience. Human conversational turn-taking sits around 200–300 ms in fast conversation and 500–700 ms in normal conversation. Your budget should target ~500 ms end-to-end for the agent. Spending engineering hours to push from 350 ms to 250 ms is a waste unless you are shipping real-time language tutoring.
Self-hosted TTS for privacy
Occasionally right (regulated industries, true on-prem requirements), but most teams overestimate the privacy benefit. If your STT and TTS happen on a vendor’s GPU, the audio is already in someone else’s cloud. The real privacy question is end-to-end data handling, not whether the inference is in your data center.
Practical checklist before you commit
- What is the longest expected call? (Drives token cost and how much context the S2S model needs to retain.)
- Do you need voice cloning, or just a brand voice? (Drives you toward ElevenLabs / Hume vs. OpenAI presets.)
- How many concurrent calls at peak? (Drives you away from vendors with per-account rate limits.)
- What is your language coverage requirement? (Drives you toward modular if you need 5+ languages with quality.)
- Is latency or voice quality more important to your user? (Drives you toward S2S or modular respectively.)
- Do you have a voice engineer, or are you a full-stack team learning voice for the first time? (Drives you toward S2S.)
- What data residency rules apply? (Drives you away from vendors with limited regional coverage.)
If you answered “yes” to two or more of voice cloning, multilingual, on-prem, or high concurrent volume, go modular. Otherwise, start with S2S and revisit later.
Common mistakes to avoid
Five pitfalls I see in most voice-agent builds.
- Choosing the vendor’s demo voice as your production voice. Their voice has been optimized for showcase conversations. Yours needs to be tested against your actual call flow. Build a 50-call eval set of real customer interactions and re-record the voice against that.
- Optimizing for WER instead of task completion. Lower word error rate matters less than whether the agent completes the customer’s task. Eval on the latter. A 7% WER model that fails to fill in the right form field is worse than a 10% WER model that does.
- Forgetting the agent’s voice is part of your brand. Companies spend $200K on a brand identity, then ship a default Cartesia or OpenAI voice that sounds like a generic SaaS chatbot. Voice is one of the strongest signals of brand trust on a phone call. Spend a week picking.
- Underestimating the latency tax of function calls. When your voice agent makes a database query or hits an internal API, the call latency spikes. Most teams do not measure this in their eval set. Build turn latency distributions, not just averages.
- Skipping the conversation design work. “Just hook up an LLM to a phone number” is not a strategy. The difference between a voice agent customers trust and one they hang up on is almost always the system prompt, the few-shot examples, and the escalation rules. Treat conversation design as engineering, not as a content task.
When this is — and isn’t — the right decision
The voice-agent space in mid-2026 is not a horse race anymore. It is a clear split between two valid stacks, each with honest tradeoffs. The mistake is not picking the “wrong” one. The mistake is picking based on the demo, without doing the work to understand which failure modes your specific use case will hit.
If you are shipping voice this quarter, start with the simplest stack that handles your hard constraint. S2S for conversational fluency out of the box. Modular for brand voice and swap-ability. Then commit to building an eval set of real calls before you scale. The teams that win are the ones who treat voice as a system to be measured, not a feature to be launched.
The pitch-deck answer is “use the latest realtime model.” The honest answer is: it depends on what you are trying to ship, who it is for, and how much of your time you can spend on voice engineering. The good news is that in 2026, both stacks are good enough to ship. The bad news is that picking between them still requires actually knowing your product.
FAQ
Is OpenAI gpt-realtime production-ready?
Yes. It went GA in 2025 with a 20% price cut, measurable benchmark gains over the preview model, and SIP support for connecting to the public phone network. It is the fastest path to a working voice agent for most teams. Remote MCP support also means you can plug your existing MCP servers into voice directly without writing glue code.
Do I still need VAD if I use Cartesia Ink-2?
No. Ink-2 emits semantic turn events (turn.start, turn.eager_end, turn.end) natively, so you do not need a separate voice activity detection layer. This eliminates a class of integration bugs that has plagued modular voice stacks since 2024.
Can I clone a voice legally in 2026?
In the EU, the AI Act requires clear disclosure of synthetic voices and imposes obligations on providers. In the US, state laws vary but the trend is toward requiring consent for clones. Hume’s Voice Control and ElevenLabs’ Voice Design give you most of the brand-voice benefit without the cloning risk. Default to those.
How much latency budget do I actually need?
Human conversation runs at 200–700 ms turn-around depending on context. Target ~500 ms end-to-end for the agent. Pushing below 300 ms is rarely worth the engineering cost unless you are shipping real-time tutoring or accessibility-focused use cases where faster is genuinely better.
What is the cheapest realistic stack?
Cartesia Ink-2 + a small open LLM (Qwen 3, Llama 4 8B, or a fine-tune) + ElevenLabs Flash or Cartesia Sonic TTS. You can hit $0.20–0.40 per minute of conversation. The cost is engineering time, not API spend.

