CAMBRIDGE, United Kingdom, Sept. 17, 2026 (GLOBE NEWSWIRE) -- Speechmatics, the voice AI company building technology to understand every voice, today launches Agent STT – a speech-to-text API built specifically for production voice agents. Powered by its new Linden model, Agent STT is designed to catch the small, high-consequence errors that can derail an otherwise well-built call.
For voice agents, average transcription accuracy can hide the mistakes that matter most. If speech-to-text changes a digit in an account number, misses a negation or fails to capture a one-word confirmation, every component downstream can work correctly and still produce the wrong outcome.
Agent STT has been built around these production failures, focusing on semantic accuracy and the conversational context an agent needs to act reliably.
Linden finalizes spoken segments in under 350ms in Speechmatics’ internal testing and supports 55+ languages. Agent STT also includes custom vocabulary of up to 1,000 terms, live speaker diarization and Speaker ID, plus conversational events alongside transcripts.
In the Pipecat STT benchmark at time of its launch, Linden recorded a 1.05% pooled semantic error rate with a 369ms median finalization time, placing it on the speed-accuracy Pareto frontier. Of the 23 streaming models tested, no other model was both faster and more accurate.
“Speechmatics has always looked to solve the hard problems in voice, and Agent STT is a good example of that. It focuses on the small recognition errors that can have an outsized impact on voice agents in production. Pipecat lets developers choose the strongest components for their use case, and semantic error rate is central to the open source benchmark we maintain. Speechmatics Agent STT powered by Linden gives developers and enterprises another strong option for the speech layer.”
Mark Backman, VP Product, Daily, which maintains Pipecat and the open source STT benchmark
Developers can integrate Agent STT directly via API or through platforms including Pipecat and LiveKit, allowing teams to add the speech layer to existing voice agent stacks without rebuilding their pipeline.
“Some of the most exciting innovation in voice AI is happening across the partner ecosystem, where it’s becoming dramatically easier to build and ship sophisticated agents. With Agent STT, we’re bringing a new level of accuracy and reliability to that stack without adding complexity. We found that customers who went to production rated accuracy as the key decider once a certain latency threshold had been met. Because however advanced the agent, if it misses a digit, a name or a single ‘not’, the whole interaction can go wrong.”
Ricardo Herreros-Symons, Chief Strategy & Revenue Officer at Speechmatics
Agent STT is available now, with launch pricing from $0.30 per hour with further volume discounts making it $0.16 per hour.
Start building with Agent STT: GitHub | LiveKit | Pipecat
About Speechmatics
Speechmatics is on a mission to build technology to ‘understand every voice’. Globally, developers and enterprises use its APIs to power real-time speech applications in highly regulated industries that across 55+ languages. Its solutions are available via the cloud, on premise and on device.
Learn more at speechmatics.com.
About Pipecat
Pipecat is the most widely used open source ecosystem for building voice and multimodal AI agents. Maintained by Daily with support from its developer community, Pipecat connects developers with models and services across speech, language, vision and real-time communications.
Learn more at pipecat.ai
Mieke Smith
Mieke.smith@speechmatics.com
01223 794497
