Decagon Introduces Voice 3 Agent With Chord Speech Model

decagon-introduces-voice-3-agent-with-chord-speech-model

Source: Unite.AI

Decagon introduced Voice 3, its voice AI agent for customer calls, and Chord, the first speech model from Decagon Labs, at the company’s Decagon Dialogues 2026 event on October 1, 2026, saying the agent supports more than 70 languages and runs on a new duplex architecture.

In the Voice 3 announcement, authored by Product Manager Quique Lores and Senior Product Marketing Manager Ariana Xiang, Decagon described Voice 3 as its most advanced voice AI agent yet. The agent speaks through Chord, and it launched alongside three other products at Decagon Dialogues 2026.

In the Decagon Dialogues 2026 post, co-founders Jesse Zhang, Decagon’s chief executive, and Ashwin Sreenivas, its president, named the other three releases: Personal Agent Gateway, which identifies customers’ personal AI agents and gives them a dedicated channel to engage with a business; Agent Modules, which extend an agent across journeys beyond support; and Duet Apprentice, which lets the company’s Duet system learn a business from internal resources and escalated customer conversations.

Enterprises first put Decagon agents to work resolving support tickets, the co-founders wrote, and those agents now qualify leads, onboard customers, collect payments, and grow relationships. The four releases expand how customers reach what Decagon calls an AI concierge and what it can do for them.

Voice 3 and the Chord Speech Model

Chord is the first voice model trained by Decagon Labs, and the company says it was post-trained on real-world customer experience conversations rather than for the broad range of uses most voice models serve, from audiobook narration to video-game voiceovers. Decagon says the model shapes speech phrase by phrase, slowing down for a confirmation code or phone number and then returning to a conversational pace, instead of relying on one global speed setting. Existing voice customization controls carry over: voice selection, pronunciation of business-specific terms, and delivery tuned to the business.

Voice 3 supports more than 70 languages without requiring teams to build a separate agent for each one, according to the announcement. It detects the caller’s language and switches automatically, even when a caller moves between languages mid-sentence, and Decagon says its locale-specific voices reflect how people speak in each market and are validated by native speakers before they ship.

Decagon states that Chord is trained on licensed data and consented voice talent, never on customer-owned data. Christian Niedworok, Lead of Digital Service Communication at Deutsche Telekom, said in the announcement that years of IVR systems have trained customers to speak in short fragments just to reach a human, and that Decagon’s voice sounds like it is actually listening and keeps the conversation moving instead of going quiet while it works. Chime Chief Operating Officer Janelle Sallenave said Decagon Voice lets Chime combine high performance and seamless brand customization with cross-channel memory, keeping every interaction connected and true to its member-first values.

Training Pipeline and Base Model

A companion post by Samuel Zhang, a member of Decagon’s technical staff focused on agent orchestration, describes how Chord was built. Its training pipeline is designed to preserve the texture of real conversation, including shifting pace, natural emphasis, pauses, and filler words, rather than scrubbing recordings into the clean, even read the post says makes voices sound synthetic. Two methods support that approach: a verbatim transcription process that preserves pacing, emphasis, and disfluencies, and a recording method that combines the content of a script with the naturalness of free-flowing conversation rather than a performative read from voice actors.

For the base model, Decagon post-trained a tokenizer-free diffusion model. The post argues that discretizing audio into tokens discards information at every tokenization step, while a token-free representation stays close to lossless and preserves the fine detail that separates a voice that is merely intelligible from one that sounds present. It also makes the model far more steerable, the post says, allowing Decagon to shape emotion, pace, and emphasis with precision. In production, Chord is served on Modal.

Duplex Architecture

Decagon says most voice agents run a cascaded pipeline in which speech-to-text transcribes the caller, a large language model decides how to respond, and text-to-speech speaks the reply, with each stage adding delay and the coordination between stages making natural turn-taking difficult. Voice 3 replaces that arrangement with a duplex architecture that runs two layers in parallel: a low-latency conversational model handles listening and speaking, from answers to progress updates, while a more powerful model manages reasoning, tool calling, and guardrail enforcement behind the conversation.

According to the company, the agent processes incoming audio even while it is speaking, so it talks through a caller saying “mhm” but yields on a genuine interruption. It narrates its progress during long lookups, answers follow-up questions while a task is still running, and helps with smaller requests in between, rather than leaving the caller in silence or on hold.

Company-Reported Evaluation Results

Zhang’s post reports on customers in telecom, financial services, and travel and hospitality that switched from an off-the-shelf voice to a voice on Chord, comparing the same programs before and after with only the voice changed. Resolution rate increased in every case, Decagon reports: up 6.1 points for the telecom customer, 7.1 points for the financial services provider, and 2.6 points for the travel brand. Barge rates, which Decagon defines as the share of calls where the caller’s only goal is to reach a human, decreased by 14.5, 6.1, and 5.3 points across the three customers.

In a blind listening test across three voices, listeners heard the same audio samples spoken once by Chord and once by the voice talent it was modeled on, and were asked to pick which was the AI. Decagon reports that 45.7% of listeners believed the real human voice was the AI, a result the company describes as close enough to a 50-50 coin flip that the two were effectively indistinguishable. The Voice 3 announcement summarizes the result as approximately 90% of users unable to tell.

Decagon also reports a preference test that put Chord head-to-head against three other speech models it describes as leading, with the same words spoken by each and listeners asked to choose the voice they would most want to talk to as a customer with a problem. Across 185 listeners, the company says Chord placed first overall at 36.2% and reached as high as 48.4% in individual rounds.

Looking ahead, Decagon says the harder and more valuable problem is how a voice behaves inside a real conversation, including whether it earns trust over five minutes and holds up when a call gets messy. The company describes Chord as the latest expression of the core bet at Decagon Labs: that for customer interactions, a model fine-tuned for a specific task outperforms a general-purpose frontier model built for everything.