Phoneme decoding
Also: phone decoding, phonemic decoding, sub-word speech decoding
By neuraspeak editorial · Updated 2026-10-07 · 1 min read
Phoneme decoding is the approach, used by leading speech neuroprostheses, of translating brain activity into a sequence of phonemes, the small units of sound that distinguish words, and then assembling those phonemes into words and sentences with a language model.
Why decode phonemes instead of whole words?
English has tens of thousands of words but only about 40 phonemes. A decoder that recognizes whole words must see each word during training, which limits the vocabulary. The 2021 UCSF study classified among 50 words (Moses et al., NEJM). A phoneme decoder only has to learn a small set of sounds, and any word can be spelled from them. This allowed Willett et al. (Nature 2023) to decode from a 125,000-word vocabulary with a 23.8% word error rate, compared with 9.1% on a 50-word vocabulary.
How does phoneme decoding work in practice?
Speech motor cortex encodes the movements of the lips, tongue, jaw, and larynx that produce each sound. A recurrent neural network reads short windows of neural activity and outputs a probability for each phoneme, plus silence, at each time step; Willett et al. and Card et al. used 80-millisecond steps. Because the network is not told exactly when each sound starts, it is trained with a technique called connectionist temporal classification (CTC), which aligns predicted and target phoneme sequences automatically. Metzger et al. (Nature 2023) similarly trained models to predict phone probabilities from ECoG signals. A language model then converts the phoneme stream into the most plausible sentence.
Why it matters for speech BCI
Phoneme decoding is the main reason modern speech neuroprostheses can offer large, open vocabularies instead of fixed word lists. Card et al. (NEJM 2024) reached 90.2% accuracy on a 125,000-word vocabulary on the second day of use, after 1.4 additional hours of training, and 97.5% accuracy sustained over 8.4 months. Some newer systems skip text altogether and decode sound features directly to synthesize voice in real time (Wairagkar et al., Nature 2025).
Questions
Can phoneme decoding handle names and new words?+
In principle yes, because any word can be built from phonemes. In practice the language model must include the word in its vocabulary, or the system tends to substitute a similar, more common word.
Is phoneme decoding the same as speech recognition?+
It is similar in structure to automatic speech recognition, with phoneme models plus a language model, but the input is neural activity rather than audio.
Related: Neural decoderSpeech neuroprosthesisUtah arrayElectrocorticography (ECoG)
Sources
- Willett et al., A high-performance speech neuroprosthesis, Nature 620, 1031-1036 · 2023
- Metzger et al., A high-performance neuroprosthesis for speech decoding and avatar control, Nature 620, 1037-1046 · 2023
- Card et al., An Accurate and Rapidly Calibrating Speech Neuroprosthesis, New England Journal of Medicine 391, 609-618 · 2024
- Moses et al., Neuroprosthesis for Decoding Speech in a Paralyzed Person with Anarthria, New England Journal of Medicine · 2021
Facts on this page are drawn from the sources listed and were accurate to the best of our knowledge on the update date shown. Company and product names are trademarks of their owners and are used for identification only. Nothing here is medical, investment or legal advice. Companies and individuals can request a correction via the corrections page.