August brought two useful open voice-stack releases: IBM announced Granite Speech 5.0 TurboCTC on August 25, and Daily introduced Pipecat PhoneLLM Alpha 1 on August 27, 2026. They occupy different layers. Granite converts speech to text; PhoneLLM supplies the language-model component of a voice agent.
The IBM announcement describes compact English recognition models. Daily’s release post describes a model paired with transcription and text-to-speech systems. Neither component alone provides a complete calling service.
The pipeline has more than one latency budget
Transcription, language-model processing, tools, and speech output all contribute to the time a caller waits. Network transport and deciding when a person has finished speaking also affect the experience.
Measure the complete turn from the user’s speech to a useful spoken answer. A fast batch transcription result on a powerful GPU is not the same metric as live call responsiveness. An inexpensive language model can still produce an expensive call if it repeatedly asks unnecessary questions or triggers slow tools.
Granite’s two variants have different licenses
IBM describes 470M-parameter English recognition models. The standard TurboCTC variant is Apache 2.0 licensed; the -nc variant carries CC-BY-NC-SA-4.0 terms. The announcement notes that this encoder-only design gives up capabilities found in some earlier language-model-equipped speech models, including translation.
A commercial team should choose the exact artifact intentionally. Do not assume the best-scoring variant in a chart has the same commercial permissions as another similarly named model. Inspect the official model card and preserve the revision evaluated.
PhoneLLM remains an alpha language model
Daily positions PhoneLLM for low-latency, multi-turn phone workloads. The model card identifies BSD 2-Clause terms for its contribution and explains inherited NVIDIA Nemotron obligations, including notice requirements when redistributing the underlying work.
The alpha label is a reason to test carefully, not to assume the model is unusable. Evaluate tool calls, conversation style, and handling of missing information. A phone-oriented model still needs an application enforcing permissions, confirmations, and truthful task status.
Where an open stack is worth the effort
Controlled deployment can be useful for teams with clear data, customization, or cost requirements. It also transfers serving responsibility to the operator. Include monitoring, scaling, updates, and recovery in the comparison with a managed service.
Nerova’s assessment is that these releases make a complete open voice pipeline more practical to investigate. The evaluation should preserve each model’s role and license, and measure the caller’s actual task. Component speed and open weights are inputs to that decision, not proof of a reliable voice agent.