Metaheuristic-Guided Neural Architecture Search for Cross-Session EEG Speech Imagery Decoding

Authors

  • Mahmoud Rahimi * Department of Computer Engineering, Sharif University of Technology, Tehran, Iran.
  • Sara Al-Farsi Department of Computer Science and Software Engineering, College of Information Technology, UAE University, Al Ain, United Arab Emirates.

https://doi.org/10.48313/maa.v1i3.101

Abstract

Brain-Computer Interfaces (BCIs) based on Speech Imagery (SI) hold transformative potential for restoring communication in individuals with locked-in syndrome, Amyotrophic Lateral Sclerosis (ALS), and severe dysarthria. However, practical deployment of Speech Imagery BCIs is critically limited by the cross-session transfer problem: Electroencephalography (EEG) signal distributions shift substantially between recording sessions due to electrode repositioning, impedance fluctuations, and cognitive state variations, causing within-session classification accuracies of 60–70% to collapse to 30–40% in cross-session evaluation. Existing deep learning architectures for EEG decoding are hand-designed and optimized for within-session performance, leaving the cross-session generalization problem largely unaddressed. We propose Metaheuristic-guided Neural Architecture Search (NAS) for Speech Imagery (MetaNAS-SI), a novel framework that automatically discovers optimal neural network architectures explicitly optimized for cross-session EEG Speech Imagery decoding. MetaNAS-SI introduces a hybrid evolutionary search combining aging evolution for macro-architecture exploration with Differential Evolution (DE) for micro-architecture optimization, a cross-session fitness function based on Leave-One-Session-Out (LOSO) evaluation, a searchable Session-Invariant Feature Alignment Module (SIFAM) that performs domain adaptation via maximum mean discrepancy within candidate architectures, and an EEG-specific search space incorporating physiologically motivated temporal convolutions, Common Spatial Pattern (CSP)-inspired spatial filters, and multi-scale temporal attention. MetaNAS-SI was evaluated on three Speech Imagery datasets: KaraOne (5-class), the coretto imagined speech dataset (4-class), and an in-house Farsi-Arabic dataset (6-class, 5 sessions per subject). Cross-session accuracy reached 45.8% on KaraOne (vs. 36.4% for the best baseline), 52.3% on Coretto (vs. 43.8%), and 48.1% on the Farsi-Arabic dataset (vs. 35.2%). Discovered architectures were 3–15× smaller than hand-designed alternatives while achieving 18–38% relative improvement in cross-session accuracy. Ablation analysis confirmed the critical contributions of the cross-session fitness function and SIFAM module. Pareto-optimal architecture selection enables deployment-aware trade-offs between accuracy, model size, and inference latency for real-time BCI applications.

Keywords:

Speech imagery, Electroencephalography decoding, Neural architecture search, Metaheuristic optimization, Cross-session transfer, Domain adaptation, Evolutionary optimization

References

  1. [1] McFarland, D. J., & Wolpaw, J. R. (2011). Brain-computer interfaces for communication and control. Communications of the ACM, 54(5), 60–66. https://doi.org/10.1145/1941487.1941506

  2. [2] Birbaumer, N., Murguialday, A. R., & Cohen, L. (2008). Brain-computer interface in paralysis. Current opinion in neurology, 21(6), 634—638. https://doi.org/10.1097/wco.0b013e328315ee2d

  3. [3] Lotte, F., Bougrain, L., Cichocki, A., Clerc, M., Congedo, M., Rakotomamonjy, A., & Yger, F. (2018). A review of classification algorithms for EEG-based brain–computer interfaces: A 10 year update. Journal of neural engineering, 15(3), 31005. https://doi.org/10.1088/1741-2552/aab2f2 1741-2552/aab2f2

  4. [4] Makeig, S., Debener, S., Onton, J., & Delorme, A. (2004). Mining event-related brain dynamics. Trends in cognitive sciences, 8(5), 204–210. https://doi.org/10.1016/j.tics.2004.03.008

  5. [5] Hickok, G., & Poeppel, D. (2007). The cortical organization of speech processing. Nature reviews neuroscience, 8(5), 393–402. https://doi.org/10.1038/nrn2113

  6. [6] Martin, S., Iturrate, I., Millán, J. del R., Knight, R. T., & Pasley, B. N. (2018). Decoding inner speech using electrocorticography: Progress and challenges toward a speech prosthesis. Frontiers in neuroscience, 12, 1–10. https://doi.org/10.3389/fnins.2018.00422

  7. [7] DaSalla, C. S., Kambara, H., Sato, M., & Koike, Y. (2009). Single-trial classification of vowel Speech Imagery using common spatial patterns. Neural networks, 22(9), 1334–1339. https://doi.org/10.1016/j.neunet.2009.05.008

  8. [8] Brigham, K., & Kumar, B. V. K. V. (2010). Imagined speech classification with EEG signals for silent communication: A preliminary investigation into synthetic telepathy. 2010 4th international conference on bioinformatics and biomedical engineering (pp. 1–4). IEEE. https://doi.org/10.1109/ICBBE.2010.5515807

  9. [9] Nguyen, C. H., Karavas, G. K., & Artemiadis, P. (2017). Inferring imagined speech using EEG signals: A new approach using Riemannian manifold features. Journal of neural engineering, 15(1), 16002. https://doi.org/10.1088/1741-2552/aa8235

  10. [10] Shenoy, P., Krauledat, M., Blankertz, B., Rao, R. P. N., & Müller, K. R. (2006). Towards adaptive classification for BCI. Journal of neural engineering, 3(1), R13. https://doi.org/10.1088/1741-2560/3/1/R02

  11. [11] Kang, H., Nam, Y., & Choi, S. (2009). Composite common spatial pattern for subject-to-subject transfer. IEEE signal processing letters, 16(8), 683–686. https://doi.org/10.1109/LSP.2009.2022557

  12. [12] Jayaram, V., Alamgir, M., Altun, Y., Scholkopf, B., & Grosse-Wentrup, M. (2016). Transfer learning in brain-computer interfaces. IEEE computational intelligence magazine, 11(1), 20–31. https://doi.org/10.1109/MCI.2015.2501545

  13. [13] Lawhern, V. J., Solon, A. J., Waytowich, N. R., Gordon, S. M., Hung, C. P., & Lance, B. J. (2018). EEGNet: A compact convolutional neural network for EEG-based brain–computer interfaces. Journal of neural engineering, 15(5), 56013. https://doi.org/10.1088/1741-2552/aace8c

  14. [14] Schirrmeister, R. T., Springenberg, J. T., Fiederer, L. D. J., Glasstetter, M., Eggensperger, K., Tangermann, M., … & Ball, T. (2017). Deep learning with convolutional neural networks for EEG decoding and visualization. Human brain mapping, 38(11), 5391–5420. https://doi.org/10.1002/hbm.23730

  15. [15] Song, Y., Zheng, Q., Liu, B., & Gao, X. (2023). EEG conformer: Convolutional transformer for EEG decoding and visualization. IEEE transactions on neural systems and rehabilitation engineering, 31, 710–719. https://doi.org/10.1109/TNSRE.2022.3230250

  16. [16] Barret, Z., Le Quoc, V. (2017). Neural architecture search with reinforcement learning. International conference on learning representatoins (Vol. 1, pp. 1–16). ICLR. https://arxiv.org/pdf/1611.01578

  17. [17] Liu, H., Simonyan, K., & Yang, Y. (2018). Darts: Differentiable architecture search. https://doi.org/10.48550/arXiv.1806.09055

  18. [18] Real, E., Aggarwal, A., Huang, Y., & Le, Q. V. (2019). Regularized evolution for image classifier architecture search. Proceedings of the AAAI conference on artificial intelligence (pp. 4780–4789). PKP Publishing Services Network. https://doi.org/10.1609/aaai.v33i01.33014780

  19. [19] Sun, J., Xie, J., & Zhou, H. (2021). EEG classification with transformer-based models. 2021 IEEE 3rd global conference on life sciences and technologies (Lifetech) (pp. 92–93). IEEE. https://doi.org/10.1109/LifeTech52111.2021.9391844

  20. [20] Zhu, Q., Zhao, X., Zhang, J., Gu, Y., Weng, C., & Hu, Y. (2023). EEG2VEC: Self-supervised electroencephalographic representation learning. https://doi.org/10.48550/arXiv.2305.13957

  21. [21] Zhao, S., & Rudzicz, F. (2015). Classifying phonological categories in imagined and articulated speech. 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 992–996). IEEE. 10.1109/ICASSP.2015.7178118

  22. [22] Coretto, G. A. P., Gareis, I. E., & Rufiner, H. L. (2017). Open access database of eeg signals recorded during imagined speech. 12th international symposium on medical information processing and analysis (Vol. 10160, p. 1016002). SPIE. https://doi.org/10.1117/12.2255697

  23. [23] Torres-García, A. A., Reyes-García, C. A., Villaseñor-Pineda, L., & García-Aguilar, G. (2016). Implementing a fuzzy inference system in a multi-objective EEG channel selection model for imagined speech classification. Expert systems with applications, 59, 1–12. https://doi.org/10.1016/j.eswa.2016.04.011

  24. [24] Iqbal, S., Khan, Y. U., & Farooq, O. (2015). EEG based classification of imagined vowel sounds. 2015 2nd international conference on computing for sustainable global development (INDIACom) (pp. 1591–1594). IEEE. https://arxiv.org/pdf/1904.05746

  25. [25] Lee, S. H., Lee, M., Jeong, J. H., & Lee, S. W. (2019). Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery. 2019 IEEE international conference on systems, man and cybernetics (SMC) (pp. 4409–4414). IEEE. https://doi.org/10.1109/SMC.2019.8914645

  26. [26] Nieto, N., Peterson, V., Rufiner, H. L., Kamienkowski, J. E., & Spies, R. (2022). Thinking out loud, an open-access EEG-based BCI dataset for inner speech recognition. Scientific data, 9(1), 52. https://doi.org/10.1038/s41597-022-01147-2

  27. [27] Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., … & Keutzer, K. (2019). FBNet: Hardware-aware efficient convnet design via differentiable neural architecture search. 2019 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 10726–10734). IEEE. https://doi.org/10.1109/CVPR.2019.01099

  28. [28] He, H., & Wu, D. (2020). Transfer learning for brain–computer interfaces: A euclidean space data alignment approach. IEEE transactions on biomedical engineering, 67(2), 399–410. https://doi.org/10.1109/TBME.2019.2913914

  29. [29] Barachant, A., Bonnet, S., Congedo, M., & Jutten, C. (2013). Classification of covariance matrices using a Riemannian-based kernel for BCI applications. Neurocomputing, 112, 172–178. https://doi.org/10.1016/j.neucom.2012.12.039

  30. [30] Özdenizci, O., Wang, Y., Koike-Akino, T., & Erdoğmuş, D. (2020). Learning invariant representations from EEG via adversarial inference. IEEE access, 8, 27074–27085. https://doi.org/10.1109/ACCESS.2020.2971600

  31. [31] Long, M., Cao, Y., Wang, J., & Jordan, M. (2015). Learning transferable features with deep adaptation networks. Proceedings of the 32nd international conference on machine learning (Vol. 37, pp. 97–105). Lille, France: PMLR. https://proceedings.mlr.press/v37/long15.html

  32. [32] Zhao, H., Zheng, Q., Ma, K., Li, H., & Zheng, Y. (2020). Deep representation-based domain adaptation for nonstationary EEG classification. IEEE transactions on neural networks and learning systems, 32(2), 535–545. https://doi.org/10.1109/TNNLS.2020.3010780

  33. [33] Zanini, P., Congedo, M., Jutten, C., Said, S., & Berthoumieu, Y. (2018). Transfer learning: A riemannian geometry framework with applications to brain–computer interfaces. IEEE transactions on biomedical engineering, 65(5), 1107–1116. https://doi.org/10.1109/TBME.2017.2742541

  34. [34] Rodrigues, P. L. C., Jutten, C., & Congedo, M. (2019). Riemannian procrustes analysis: Transfer learning for brain–computer interfaces. IEEE transactions on biomedical engineering, 66(8), 2390–2401. https://doi.org/10.1109/TBME.2018.2889705

  35. [35] Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., & Le, Q. V. (2019). MnasNet: Platform-aware neural architecture search for mobile. 2019 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 2815–2823). IEEE. https://doi.org/10.1109/CVPR.2019.00293

  36. [36] Bender, G., Kindermans, P.-J., Zoph, B., Vasudevan, V., & Le, Q. (2018). Understanding and simplifying one-shot architecture search. Proceedings of the 35th international conference on machine learning (Vol. 80, pp. 550–559). PMLR. https://proceedings.mlr.press/v80/bender18a/bender18a.pdf

  37. [37] Guo, Z., Zhang, X., Mu, H., Heng, W., Liu, Z., Wei, Y., & Sun, J. (2020). Single path one-shot neural architecture search with uniform sampling. Computer vision- ECCV 2020 (pp. 544–560). Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-58517-4_32

  38. [38] Holland, J. (1975). Adaptation in natural and artificial systems. University of Michigan Press, Ann Arbor, MI. https://doi.org/10.7551/mitpress/1090.001.0001

  39. [39] Eberhart, R., & Kennedy, J. (1995). Particle swarm optimization. Proceedings of the IEEE international conference on neural networks (Vol. 4, pp. 1942–1948). IEEE. https://doi.org/10.1109/ICNN.1995.488968

  40. [40] Storn, R., & Price, K. (1997). Differential evolution – A simple and efficient heuristic for global optimization over continuous spaces. Journal of global optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328

  41. [41] Kirkpatrick, S., Gelatt, C. D., & Vecchi, M. P. (1983). Optimization by simulated annealing. Science, 220(4598), 671–680. https://doi.org/10.1126/science.220.4598.671

  42. [42] Dorigo, M., & Gambardella, L. M. (1997). Ant colony system: A cooperative learning approach to the traveling salesman problem. IEEE transactions on evolutionary computation, 1(1), 53–66. https://doi.org/10.1109/4235.585892

  43. [43] Stanley, K. O., Clune, J., Lehman, J., & Miikkulainen, R. (2019). Designing neural networks through neuroevolution. Nature machine intelligence, 1(1), 24–35. https://doi.org/10.1038/s42256-018-0006-z

  44. [44] Stanley, K. O., & Miikkulainen, R. (2002). Evolving neural networks through augmenting topologies. Evolutionary computation, 10(2), 99–127. https://doi.org/10.1162/106365602320169811

  45. [45] Real, E., Moore, S., Selle, A., Saxena, S., Suematsu, Y. L., Tan, J., …& Kurakin, A. (2017). Large-scale evolution of image classifiers. Proceedings of the 34th international conference on machine learning (Vol. 70, pp. 2902–2911). PMLR. https://proceedings.mlr.press/v70/real17a.html

  46. [46] Aler, R., Galván, I. M., & Valls, J. M. (2012). Applying evolution strategies to preprocessing EEG signals for brain–computer interfaces. Information sciences, 215, 53–66. https://doi.org/10.1016/j.ins.2012.05.012

  47. [47] Zhang, R., Xu, P., Liu, T., Zhang, Y., Guo, L., Li, P., & Yao, D. (2013). Local temporal correlation common spatial patterns for single trial EEG classification during motor imagery. Computational and mathematical methods in medicine, 2013(1), 591216. https://doi.org/10.1155/2013/591216

  48. [48] Ilyas, M. Z., Saad, P., & Ahmad, M. I. (2015). A survey of analysis and classification of eeg signals for brain-computer interfaces. 2015 2nd international conference on biomedical engineering (ICOBE) (pp. 1–6). IEEE. https://doi.org/10.1109/ICoBE.2015.7235129

  49. [49] Salmelin, R., & Hari, R. (1994). Spatiotemporal characteristics of sensorimotor neuromagnetic rhythms related to thumb movement. Neuroscience, 60(2), 537–550. https://doi.org/10.1016/0306-4522(94)90263-1

  50. [50] Crone, N. E., Hao, L., Hart, J., Boatman, D., Lesser, R. P., Irizarry, R., & Gordon, B. (2001). Electrocorticographic gamma activity during word production in spoken and sign language. Neurology, 57(11), 2045–2053. https://doi.org/10.1212/WNL.57.11.2045

  51. [51] Bastiaansen, M., & Hagoort, P. (2006). Oscillatory neuronal dynamics during language comprehension. In Event-related dynamics of brain oscillations (Vol. 159, pp. 179–196). Elsevier. https://doi.org/10.1016/S0079-6123(06)59012-0

  52. [52] Pfurtscheller, G., & Lopes Da Silva, F. H. (1999). Event-related EEG/MEG synchronization and desynchronization: Basic principles. Clinical neurophysiology, 110(11), 1842–1857. https://doi.org/10.1016/S1388-2457(99)00141-8

  53. [53] Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132–7141). IEEE. https://doi.org/10.1109/CVPR.2018.00745

  54. [54] Shaw, P., Uszkoreit, J., & Vaswani, A. (2018). Self-attention with relative position representations. Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: Human language technologies, volume 2 (short papers) (pp. 464–468). New Orleans, Louisiana: Association for Computational Linguistics. https://doi.org/10.18653/v1/N18-2074

  55. [55] Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE transactions on evolutionary computation, 6(2), 182–197. https://doi.org/10.1109/4235.996017

  56. [56] Ramoser, H., Müller-Gerking, J., & Pfurtscheller, G. (2000). Optimal spatial filtering of single trial EEG during imagined hand movement. IEEE transactions on rehabilitation engineering, 8(4), 441–446. https://doi.org/10.1109/86.895946

  57. [57] Ang, K. K., Chin, Z. Y., Zhang, H., & Guan, C. (2008). Filter bank common spatial pattern (FBCSP) in brain-computer interface. 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence) (pp. 2390–2397). IEEE. https://doi.org/10.1109/IJCNN.2008.4634130

  58. [58] Ingolfsson, T. M., Hersche, M., Wang, X., Kobayashi, N., Cavigelli, L., & Benini, L. (2020). EEG-tcnet: An accurate temporal convolutional network for embedded motor-imagery brain–machine interfaces. 2020 IEEE international conference on systems, man, and cybernetics (SMC) (pp. 2958–2965). IEEE. https://doi.org/10.1109/SMC42975.2020.9283028

  59. [59] Kingma, D. P., & Ba, J. L. (2014). Adam: Amethod for stochastic optimization. Proceedings of the 3rd international conference on learning representations (ICLR) (pp. 1–15). ICLR. https://jjcurtin.quarto.pub/010_neural_networks-6260/pdfs/kingma_adam_optimizer.pdf

  60. [60] Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. Proceedings of the 34th international conference on machine learning (Vol. 70, pp. 3319–3328). PMLR. https://doi.org/10.5555/3305890.3306024

  61. [61] Mersov, A. M., Jobst, C., Cheyne, D. O., & De Nil, L. (2016). Sensorimotor oscillations prior to speech onset reflect altered motor networks in adults who stutter. Frontiers in human neuroscience, 10, 1–16. https://doi.org/10.3389/fnhum.2016.00443

Published

2024-09-19

How to Cite

Rahimi, M. ., & Al-Farsi, S. . (2024). Metaheuristic-Guided Neural Architecture Search for Cross-Session EEG Speech Imagery Decoding. Metaheuristic Algorithms With Applications, 1(3), 269-292. https://doi.org/10.48313/maa.v1i3.101

Similar Articles

31-40 of 60

You may also start an advanced similarity search for this article.