Researchers put a single cortical implant to the test, asking whether attempted speech and body gestures could be decoded together and translated into a personalized full-body avatar.

Study: Simultaneous speech and gesture decoding for multimodal communication in paralysis. Image Credit: Kittyfly
In a recent study published in the journal Nature Neuroscience, a group of researchers investigated whether a single high-density electrocorticography (ECoG) implant could simultaneously decode attempted speech and gestures and support real-time, expressive communication through a personalized virtual avatar in people with severe paralysis.
Background
How can brain activity reveal attempts to speak or gesture when severe paralysis limits movement? Stroke and neurodegenerative diseases can impair speech and body movements, limiting social participation and quality of life.
Brain-computer interfaces (BCIs) aim to restore lost functions by translating neural activity into commands for external devices. Earlier BCI research has decoded speech and other movements separately. Natural communication often combines spoken language with gestures. The sensorimotor cortex (SMC) contains representations of multiple body regions, while ECoG can record neural activity across this territory.
About the study
Three participants with severe speech and motor impairments were enrolled in the BCI Restoration of Arm and Voice (BRAVO) clinical trial. Two had brainstem stroke, and one had amyotrophic lateral sclerosis (ALS). Each received a 253-electrode high-density ECoG array implanted subdurally over the left hemisphere, targeting speech and motor areas.
Neural activity was recorded during multiple body movements and used for initial multi-effector mapping and decoding in all three participants; after Bravo-3 withdrew, the main concurrent speech-and-gesture and avatar experiments involved Bravo-1r and Bravo-6. Attempt strategies reflected residual abilities: Bravo-1r and Bravo-3 silently attempted speech, whereas Bravo-6 attempted speech with vocalization; body movements were attempted, minimally attempted, or imagined according to participant capability.
In the copy task, targets were presented as text, movement animations, or both, followed by a go cue and a 2-second intertrial interval. A cued conversation paradigm was also used with Bravo-1r and Bravo-6.
Signals were processed to extract high-gamma activity (HGA; 70–150 Hz) and low-frequency signals (LFSs; 1–100 Hz). Participant-specific convolutional recurrent neural network decoders were trained on isolated trials, simultaneous trials, or both. A rest class used true-rest and opposite-modality trials. Performance was assessed offline with cross-validation and statistical comparisons between decoding conditions. Real-time parallel decoders controlled a personalized full-body virtual avatar. Participants provided informed consent, and institutional approval was obtained.

a, Trial structure for three task paradigms - S-only (text stimulus), G-only (animated avatar stimulus) and simultaneous S + G (text and animated avatar stimulus). Stimulus presentation is followed by a countdown to a go-cue, when the participant is instructed to attempt the target. b, Highlighted electrodes (e58, e151 and e164) overlaid on Bravo-6’s brain. c, HGA (mean ± s.e.m.) for the same three example electrodes during S-only trials (blue), G-only trials (orange) and S + G trials (green) aligned to the go-cue. d, Spatial maps of relative HGA during S-only, G-only and S + G trials for Bravo-6. Circled electrodes correspond to those labeled in b and plotted in c. e, Venn diagrams of the overlap between electrodes with significant neural activity during S-only trials and G-only trials (Bravo-6, top; Bravo-1r, bottom). f, The overlap map highlights electrodes significantly active during both speech and gesture conditions for Bravo-6. Higher overlap scores indicate electrodes with stronger relative HGA modulation during both behaviors, whereas lower overlap scores indicate weaker modulation in one or both conditions. Nonsignificant electrodes are shown as empty outlined circles.
Study results
A single ECoG array captured neural representations of multiple effectors. HGA associated with hand, arm, head, eye, and speech-related orofacial movements showed broad somatotopic organization across the SMC, with the degree of separation varying among participants. Speech- and gesture-related activity was strongest in the precentral gyrus, with speech activity more ventral and gesture activity more dorsal.
Decoder performance depended on behavioral context. Models trained only on isolated speech or gesture trials did not fully generalize to concurrent speech-and-gesture attempts. At the largest shared training size, models trained on concurrent trials performed better on concurrent tests, whereas models trained in isolation performed better on isolated tests.
Hybrid models trained on both isolated and concurrent data maintained high accuracy across contexts and outperformed models trained on only one context. Electrode contribution patterns changed with training context, with models generally shifting reliance toward electrodes showing lower speech–gesture overlap.
In the Bravo-6 leave-one-out analysis, median gesture accuracy was 65.8% for seen pairs and 66.7% for unseen pairs; speech accuracy was 79.3% and 80.0%, respectively. Within this analysis, neither speech nor gesture accuracy differed significantly between pairings considered natural or unnatural in conversation. During longitudinal testing, accuracy remained above chance across days to months, with stability varying across participants and modalities; newer training data generally improved subsequent performance.
Before cross-modality training, median opposite-modality false positive rates (FPRs) were 30.6% for Bravo-6 and 76.0% for Bravo-1r. Cross-modality training reduced both to 0.0% while maintaining low FPR during true rest. Models trained on both isolated and concurrent behaviors also maintained low false negative rates (FNRs) in both contexts, reducing the tendency to misclassify attempted speech or gestures as rest.
Using the full available datasets with cross-modality training, mean offline concurrent decoding accuracy for Bravo-6 was 68.8% for gestures and 77.5% for speech. For Bravo-1r, corresponding accuracies were 88.3% and 84.0%. In the gesture-only copy task, accuracy was 81.8% for Bravo-6 and 61.7% for Bravo-1r.
During real-time concurrent decoding in Bravo-6, accuracy was 66.0% for gestures and 70.0% for speech. In the conversation paradigm, Bravo-6 achieved average accuracies of 85.0% for gestures and 75.0% for speech, while Bravo-1r achieved median accuracies of 100.0% for both.
This proof-of-concept conversational test was limited to five four-trial blocks for Bravo-6 and three three-trial blocks for Bravo-1r. Decoded speech appeared as text, while decoded gestures animated the avatar, enabling speech alone, gestures alone, or both during virtual interaction.
Conclusions
The findings provide proof-of-concept evidence that a single high-density ECoG implant can decode speech and gestures concurrently in two people with severe paralysis. The array captured distinct but partially overlapping neural representations across speech and body movements.
Training models with both isolated and concurrent behaviors improved generalization across contexts, while cross-modality training reduced false activations. The system also supported real-time control of a personalized full-body avatar during copy and conversational tasks.
These results provide a proof of concept for multi-effector BCIs supporting more expressive communication. The early feasibility study involved three participants overall, with concurrent speech-and-gesture decoding tested in two using restricted vocabularies and structured tasks, and conversational testing limited to a small number of trials.
Further research is needed to determine whether speech and gestures can be reliably decoded when used together.
Journal reference:
- Brosler, S. C., Liu, J. R., Silva, A. B., Hallinan, I. P., Kurtz-Miott, C. M., Dunkel Wilker, J. F., Tu-Chan, A., Ganguly, K., & Chang, E. F. (2026). Simultaneous speech and gesture decoding for multimodal communication in paralysis. Nature Neuroscience, 1-11. DOI: 10.1038/s41593-026-02446-2, https://www.nature.com/articles/s41593-026-02446-2