BEGIN:VCALENDAR VERSION:2.0 X-WR-CALNAME:EventsCalendar PRODID:-//hacksw/handcal//NONSGML v1.0//EN CALSCALE:GREGORIAN BEGIN:VTIMEZONE TZID:America/New_York LAST-MODIFIED:20240422T053451Z TZURL:https://www.tzurl.org/zoneinfo-outlook/America/New_York X-LIC-LOCATION:America/New_York BEGIN:DAYLIGHT TZNAME:EDT TZOFFSETFROM:-0500 TZOFFSETTO:-0400 DTSTART:19700308T020000 RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU END:DAYLIGHT BEGIN:STANDARD TZNAME:EST TZOFFSETFROM:-0400 TZOFFSETTO:-0500 DTSTART:19701101T020000 RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU END:STANDARD END:VTIMEZONE BEGIN:VEVENT CATEGORIES:College of Arts and Sciences,College of Engineering,Thesis/Disse rtations DESCRIPTION:Advisor: Dr. Ashokkumar R. Patel - Department of Computer & Inf ormation ScienceĚýCommittee Members:Dr. Yuchou Chang – Department of Com puter & Information ScienceDr. Debarun Das – Department of Computer & In formation ScienceĚýAbstract:Inner speech - the silent production of words in the mind, without any movement or sound — is an appealing control sig nal for a brain–computer interface (BCI), because the command is the tho ught. For someone who has lost the ability to speak or move, a decoder tha t reads intended words directly would be far more natural than the indirec t mental tasks most BCIs rely on. Reading inner speech from scalp electroe ncephalography (EEG) is, however, extremely hard: the signals are weak and non-stationary, the neural traces of covert language are faint, and repor ted four-class accuracies in the literature rarely climb much past the low thirties. In this regime, the meaningful question is not whether a model reaches high accuracy; none reliably do but whether a design choice yields a real, statistically reliable signal, assessed honestly.ĚýThis thesis co mpares three standard architectures - a compact convolutional network (EEG Net), an LSTM recurrent network, and a self-attention Transformer - agains t a hybrid model that feeds a shared convolutional front end into parallel recurrent and self-attention branches and fuses them before classificatio n. All four are evaluated identically on the public "Thinking Out Loud" in ner-speech dataset (Nieto et al., 2022), under subject-dependent five-fold cross-validation on the four-class directional-word task (Up, Down, Right , Left), with on-the-fly augmentation and a fixed seed. The unit of statis tical analysis is the subject (n = 10). The hybrid model attains the highe st mean accuracy, 28.5% (SD 3.1%), and is the only model whose accuracy is statistically significantly above the 25% chance level (Wilcoxon signed-r ank p = 0.014; one-sample t-test p = 0.006). The three baselines do not re ach significance against chance. In direct paired comparisons the hybrid i s not significantly better than any individual baseline, and an ablation s hows that each single branch performs at roughly the level of its correspo nding baseline, with the two-branch fusion adding a small, non-significant improvement. A subject-independent analysis falls to chance, consistent w ith the well-documented failure of cross-subject generalization for inner speech. These results are consistent with prior decoding studies on this d ataset. The contribution is therefore a controlled, like-for-like benchmar k of four architectures under one protocol, with rigorous statistical asse ssment: it shows that on this difficult task the hybrid is the only archit ecture to clear chance, while the architectures are otherwise statisticall y indistinguishable from one another. For further questions, please contac t Professor Ashokkumar R. Patel at ashok.patel@umassd.edu\nEvent page: htt ps://www.umassd.edu/events/cms/8-19-26-inner-speech-decoding-from-eeg-a-co mparative-study.php X-ALT-DESC;FMTTYPE=text/html:
Advisor: Dr. Ashokkumar R. Pate
l - Department of Computer & Information Science
Ěý
Committee Me
mbers:
Dr. Yuchou Chang – Department of Computer & Information Scie
nce
Dr. Debarun Das – Department of Computer & Information Science<
br />Ěý
Abstract:
Inner speech - the silent production of words
in the mind\, without any movement or sound — is an appealing control si
gnal for a brain–computer interface (BCI)\, because the command is the t
hought. For someone who has lost the ability to speak or move\, a decoder
that reads intended words directly would be far more natural than the indi
rect mental tasks most BCIs rely on. Reading inner speech from scalp elect
roencephalography (EEG) is\, however\, extremely hard: the signals are wea
k and non-stationary\, the neural traces of covert language are faint\, an
d reported four-class accuracies in the literature rarely climb much past
the low thirties. In this regime\, the meaningful question is not whether
a model reaches high accuracy\; none reliably do but whether a design choi
ce yields a real\, statistically reliable signal\, assessed honestly.
Ěý
This thesis compares three standard architectures - a compact con
volutional network (EEGNet)\, an LSTM recurrent network\, and a self-atten
tion Transformer - against a hybrid model that feeds a shared convolutiona
l front end into parallel recurrent and self-attention branches and fuses
them before classification. All four are evaluated identically on the publ
ic "Thinking Out Loud" inner-speech dataset (Nieto et al.\, 2022)\, under
subject-dependent five-fold cross-validation on the four-class directional
-word task (Up\, Down\, Right\, Left)\, with on-the-fly augmentation and a
fixed seed. The unit of statistical analysis is the subject (n = 10). The
hybrid model attains the highest mean accuracy\, 28.5% (SD 3.1%)\, and is
the only model whose accuracy is statistically significantly above the 25
% chance level (Wilcoxon signed-rank p = 0.014\; one-sample t-test p = 0.0
06). The three baselines do not reach significance against chance. In dire
ct paired comparisons the hybrid is not significantly better than any indi
vidual baseline\, and an ablation shows that each single branch performs a
t roughly the level of its corresponding baseline\, with the two-branch fu
sion adding a small\, non-significant improvement. A subject-independent a
nalysis falls to chance\, consistent with the well-documented failure of c
ross-subject generalization for inner speech. These results are consistent
with prior decoding studies on this dataset. The contribution is therefor
e a controlled\, like-for-like benchmark of four architectures under one p
rotocol\, with rigorous statistical assessment: it shows that on this diff
icult task the hybrid is the only architecture to clear chance\, while the
architectures are otherwise statistically indistinguishable from one anot
her.
For further questions\, please contact Professor Ashokkumar R . Patel at ashok.patel@umassd.edu
Event page:
DTSTAMP:20260725T174253 DTSTART;TZID=America/New_York:20260819T130000 DTEND;TZID=America/New_York:20260819T140000 LOCATION:Zoom (please contact: pnadipalli@umassd.edu or ashok.patel@umassd. edu for Zoom information) SUMMARY;LANGUAGE=en-us:Inner Speech Decoding from EEG: A Comparative Study of Deep Learning Architectures for Brain–Computer Interfaces UID:d1376068ee08b54e62b17dc429cd276d@www.umassd.edu END:VEVENT END:VCALENDAR