Article Contents
REVIEW   Open Access     Cite

Generative AI for brain-computer interfaces decoding: Advances, challenges and future

    Show all affliationsShow less
More Information
  • Corresponding author: sq.wang@siat.ac.cn
  • DownLoad: Full size image
    1. Generative BCI decoding reconstructs semantically rich content, including text and images, from brain signals.

      Generative AI boosts BCI decoding through data augmentation and sensor optimization.

      Key challenges include BCI decoding reliability, cross-subject generalizability, and fairness.

  • Brain-computer interfaces (BCIs) establish direct communication pathways between the human brain and external devices, enabling novel modes of human-machine interaction and providing essential technological foundations for innovative clinical therapies. BCI decoding can be divided into discriminative and generative approaches. Discriminative decoding models are primarily used for predefined tasks such as emotion recognition, limited performance in reconstructing high-dimensional semantic content such as natural language or visual images. Recent advances in generative artificial intelligence (AI), including diffusion models and autoregressive Transformers, have brought generative decoding to the forefront of BCI research by enabling the reconstruction of semantically rich content from neural activity. Current studies further demonstrate that decoding visually evoked neural activity into images can achieve structural similarity index measure (SSIM) scores ranging approximately from 0.26 to 0.43. Furthermore, generative AI enhances BCI decoding through data augmentation and sensor optimization, thereby improving both performance and generalizability. Here, we review recent progress in generative AI-enabled BCI decoding, focusing on methodological advances in language and visual decoding. We also summarize its roles in data augmentation and sensor optimization, outline key challenges in decoding accuracy, cross-subject generalization and fairness, and suggest future directions such as multimodal integration and fairness-aware frameworks.
  • 加载中
  • [1] Xu S., Liu Y., Lee H., et al. (2024). Neural interfaces: Bridging the brain to the world beyond healthcare. Exploration 4:20230146. DOI:10.1002/exp.20230146

    View in Article CrossRef Google Scholar

    [2] Chaudhary U., Birbaumer N. and Ramos-Murguialday A. (2016). Brain-computer interfaces for communication and rehabilitation. Nat. Rev. Neurol. 12:513−525. DOI:10.1038/nrneurol.2016.113

    View in Article CrossRef Google Scholar

    [3] Schwemmer M.A., Skomrock N.D., Sederberg P.B., et al. (2018). Meeting brain-computer interface user performance expectations using a deep neural network decoding framework. Nat. Med. 24:1669−1676. DOI:10.1038/s41591-018-0171-y

    View in Article CrossRef Google Scholar

    [4] Degenhart A.D., Bishop W.E., Oby E.R., et al. (2020). Stabilization of a brain-computer interface via the alignment of low-dimensional spaces of neural activity. Nat. Biomed. Eng. 4:672−685. DOI:10.1038/s41551-020-0542-9

    View in Article CrossRef Google Scholar

    [5] Ding Y., Udompanyawit C., Zhang Y., et al. (2025). EEG-based brain-computer interface enables real-time robotic hand control at individual finger level. Nat. Commun. 16:1−20. DOI:10.1038/s41467-025-61064-x

    View in Article CrossRef Google Scholar

    [6] Kunz E.M., Krasa B.A., Kamdar F., et al. (2025). Inner speech in motor cortex and implications for speech neuroprostheses. Cell 188:4658−4673. DOI:10.1016/j.cell.2025.06.015

    View in Article CrossRef Google Scholar

    [7] Wei Q., Cui M., Liu Z., et al. (2025). Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI. Innov. Med. 3:100145. DOI:10.59717/j.xinn-med.2025.100145

    View in Article CrossRef Google Scholar

    [8] Schalk G., McFarland D.J., Hinterberger T., et al. (2004). BCI2000: A general-purpose brain-computer interface (BCI) system. IEEE Trans. Biomed. Eng. 51:1034−1043. DOI:10.1109/tbme.2004.827072

    View in Article CrossRef Google Scholar

    [9] Littlejohn K.T., Cho C.J., Liu J.R., et al. (2025). A streaming brain-to-voice neuroprosthesis to restore naturalistic communication. Nat. Neurosci. 28:902−912. DOI:10.1038/s41593-025-01905-6

    View in Article CrossRef Google Scholar

    [10] Santhanam G., Ryu S.I., Yu B.M., et al. (2006). A high-performance brain-computer interface. Nature 442:195−198. DOI:10.1038/nature04968

    View in Article CrossRef Google Scholar

    [11] Wairagkar M., Card N.S., Singer-Clark T., et al. (2025). An instantaneous voice-synthesis neuroprosthesis. Nature 644:145−152. DOI:10.1038/s41586-025-09127-3

    View in Article CrossRef Google Scholar

    [12] Willsey M.S., Shah N.P., Avansino D.T., et al. (2025). A high-performance brain-computer interface for finger decoding and quadcopter game control in an individual with paralysis. Nat. Med. 31:96−104. DOI:10.1038/s41591-024-03341-8

    View in Article CrossRef Google Scholar

    [13] Del Pozo-Banos M., Alonso J.B., Ticay-Rivas J.R., et al. (2014). Electroencephalogram subject identification: A review. Expert Syst. Appl. 41:6537−6554. DOI:10.1016/j.eswa.2014.05.013

    View in Article CrossRef Google Scholar

    [14] Hu Y., Sun L., Mao X., et al. (2024). EEG data augmentation method for identity recognition based on spatial-temporal generating adversarial network. Electronics 13:4310. DOI:10.3390/electronics13214310

    View in Article CrossRef Google Scholar

    [15] Jolfaei A., Wu X.W. and Muthukkumarasamy V. (2013). On the feasibility and performance of pass-thought authentication systems. 2013 Fourth International Conference on Emerging Security Technologies pp:33–38. DOI:10.1109/est.2013.12

    View in Article Google Scholar

    [16] Horikawa T. and Kamitani Y. (2017). Generic decoding of seen and imagined objects using hierarchical visual features. Nat. Commun. 8:15037. DOI:10.1038/ncomms15037

    View in Article CrossRef Google Scholar

    [17] Fares A., Zhong S.h. and Jiang J. (2019). EEG-based image classification via a regionlevel stacked bi-directional deep learning framework. BMC Med. Inform. Decis. Making 19:268. DOI:10.1186/s12911-019-0967-9

    View in Article CrossRef Google Scholar

    [18] Mukherjee P., Das A., Bhunia A.K. et al. (2019). Cogni-net: Cognitive feature learning through deep visual perception. 2019 IEEE International Conference on Image Processing pp:4539–4543. DOI:10.1109/icip.2019.8803717

    View in Article Google Scholar

    [19] Song Y., Wang Y., He H., et al. (2025). Recognizing natural images from EEG with language-guided contrastive learning. IEEE Trans. Neural Netw. Learn. Syst. 36:15896−15910. DOI:10.1109/tnnls.2025.3562743

    View in Article CrossRef Google Scholar

    [20] Patel P., B S. and Annavarapu R.N. (2025). Application of supervised machine learning models in human emotion classification using Tsallis entropy as a feature. J. Big Data 12:126. DOI:10.1186/s40537-025-01177-8

    View in Article CrossRef Google Scholar

    [21] Fu K., Du C., Wang S., et al. (2022). Multi-view multi-label fine-grained emotion decoding from human brain activity. IEEE Trans. Neural Netw. Learn. Syst. 35:9026−9040. DOI:10.1109/tnnls.2022.3217767

    View in Article CrossRef Google Scholar

    [22] Huang Z., Du C., Li C., et al. (2025). Identifying the hierarchical emotional areas in the human brain through information fusion. Inf. Fusion 113:102613. DOI:10.1016/j.inffus.2024.102613

    View in Article CrossRef Google Scholar

    [23] Jin M., Du C., He H., et al. (2024). PGCN: Pyramidal graph convolutional network for EEG emotion recognition. IEEE Trans. Multimedia 26:9070−9082. DOI:10.1109/tmm.2024.3385676

    View in Article CrossRef Google Scholar

    [24] Silversmith D.B., Abiri R., Hardy N.F., et al. (2021). Plug-and-play control of a brain-computer interface through neural map stabilization. Nat. Biotechnol. 39:326−335. DOI:10.1038/s41587-020-0662-5

    View in Article CrossRef Google Scholar

    [25] Huang D., Qian K., Fei D.Y., et al. (2012). Electroencephalography (EEG)-based brain-computer interface (BCI): A 2-D virtual wheelchair control based on event-related desynchronization/synchronization and state control. IEEE Trans. Neural Syst. Rehabil. Eng. 20:379−388. DOI:10.1109/tnsre.2012.2190299

    View in Article CrossRef Google Scholar

    [26] Huang S., Wang Y. and Luo H. (2025). CCSUMSP: A cross-subject Chinese speech decoding framework with unified topology and multi-modal semantic pre-training. Inf. Fusion 119:103022. DOI:10.1016/j.inffus.2025.103022

    View in Article CrossRef Google Scholar

    [27] Zhao X., Sun J., Wang S. et al. (2024). Mapguide: A simple yet effective method to reconstruct continuous language from brain activities. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies pp:3822–3832. DOI:10.18653/v1/2024.naacl-long.211

    View in Article Google Scholar

    [28] de la Torre-Ortiz C., Spapé M.M., Ravaja N., et al. (2024). Cross-subject EEG feedback for ′ implicit image generation. IEEE Trans. Cybern. 54:6105−6117. DOI:10.1109/tcyb.2024.3406159

    View in Article CrossRef Google Scholar

    [29] Meng L. and Yang C. (2023). Dual-guided brain diffusion model: Natural image reconstruction from human visual stimulus fMRI. Bioengineering 10:1117. DOI:10.3390/bioengineering10101117

    View in Article CrossRef Google Scholar

    [30] Deng C., Li X. and Dai J. (2023). Challenges for translating implantable brain-computer interface to medical device. Innov. Med. 1:100040. DOI:10.59717/j.xinn-med.2023.100040

    View in Article CrossRef Google Scholar

    [31] Tang J., LeBel A., Jain S., et al. (2023). Semantic reconstruction of continuous language from non-invasive brain recordings. Nat. Neurosci. 26:858−866. DOI:10.1038/s41593-023-01304-9

    View in Article CrossRef Google Scholar

    [32] Ozdenizci O., Wang Y., Koike-Akino T. et al. (2019). Transfer learning in brain-computer interfaces with adversarial variational autoencoders. 9th International IEEE EMBS Conference on Neural Engineering pp:207–210. DOI:10.1109/ner.2019.8716897

    View in Article Google Scholar

    [33] Scotti P., Banerjee A., Goode J., et al. (2023). Reconstructing the mind’s eye: fMRI-toimage with contrastive learning and diffusion priors. Advances in Neural Information Processing Systems 36:24705−24728. DOI:10.48550/arXiv.2305.18274

    View in Article CrossRef Google Scholar

    [34] Jeong J.H., Cho J.H., Lee B.H., et al. (2022). Real-time deep neurolinguistic learning enhances noninvasive neural language decoding for brain-machine interaction. IEEE Trans. Cybern. 53:7469−7482. DOI:10.1109/tcyb.2022.3211694

    View in Article CrossRef Google Scholar

    [35] Antonello R., Vaidya A. and Huth A. (2023). Scaling laws for language encoding models in fMRI. Advances in Neural Information Processing Systems 36:21895−21907. DOI:10.48550/arXiv.2305.11863

    View in Article CrossRef Google Scholar

    [36] Ye Z., Ai Q., Liu Y., et al. (2025). Generative language reconstruction from brain recordings. Commun. Biol. 8:346. DOI:10.1038/s42003-025-07731-7

    View in Article CrossRef Google Scholar

    [37] Metzger S.L., Littlejohn K.T., Silva A.B., et al. (2023). A high-performance neuroprosthesis for speech decoding and avatar control. Nature 620:1037−1046. DOI:10.1038/s41586-023-06443-4

    View in Article CrossRef Google Scholar

    [38] Liu Y., Zhao Z., Xu M., et al. (2023). Decoding and synthesizing tonal language speech from brain activity. Sci. Adv. 9:eadh0478. DOI:10.1126/sciadv.adh0478

    View in Article CrossRef Google Scholar

    [39] Tian Z., Quan R., Ma F., et al. (2025). BRAINGUARD: Privacy-preserving multisubject image reconstructions from brain activities. Proceedings of the AAAI Conference on Artificial Intelligence 39:14414−14422. DOI:10.1609/aaai.v39i13.33579

    View in Article CrossRef Google Scholar

    [40] Song Y., Liu B., Li X. et al. (2024). Decoding natural images from EEG for object recognition. 12th International Conference on Learning Representations pp:15475–15492. DOI:10.48550/arXiv.2308.13234

    View in Article Google Scholar

    [41] Xia W., De Charette R., Oztireli C. et al. (2024). Dream: Visual decoding from reversing human visual system. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision pp:8226–8235. DOI:10.1109/wacv57701.2024.00804

    View in Article Google Scholar

    [42] Kumar S., Sumers T.R., Yamakoshi T., et al. (2024). Shared functional specialization in transformer-based language models and the human brain. Nat. Commun. 15:5523. DOI:10.1038/s41467-024-49173-5

    View in Article CrossRef Google Scholar

    [43] Ko W., Jeon E., Yoon J.S., et al. (2022). Semi-supervised generative and discriminative adversarial learning for motor imagery-based brain-computer interface. Sci. Rep. 12:4587. DOI:10.1038/s41598-022-08490-9

    View in Article CrossRef Google Scholar

    [44] Wang S., Zhou T., Shen Y., et al. (2025). Generative AI enables EEG super-resolution via spatio-temporal adaptive diffusion learning. IEEE Trans. Consum. Electron. 71:1034−1045. DOI:10.1109/tce.2025.3528438

    View in Article CrossRef Google Scholar

    [45] Vetter J., Macke J.H. and Gao R. (2024). Generating realistic neurophysiological time series with denoising diffusion probabilistic models. Patterns 5:101047. DOI:10.1016/j.patter.2024.101047

    View in Article CrossRef Google Scholar

    [46] Kwon J. and Im C.H. (2022). Novel signal-to-signal translation method based on StarGAN to generate artificial EEG for SSVEP-based brain-computer interfaces. Expert Syst. Appl. 203:117574. DOI:10.1016/j.eswa.2022.117574

    View in Article CrossRef Google Scholar

    [47] Yao W., Lyu Z., Mahmud M., et al. (2025). CATD: Unified representation learning for EEGto-fMRI cross-modal generation. IEEE Trans. Med. Imaging 44:2757−2767. DOI:10.1109/tmi.2025.3550206

    View in Article CrossRef Google Scholar

    [48] Li Y., Wang Y., Lei B., et al. (2025). SCDM: Unified representation learning for EEG-tofNIRS cross-modal generation in MI-BCIs. IEEE Trans. Med. Imaging 44:2384−2394. DOI:10.1109/tmi.2025.3532480

    View in Article CrossRef Google Scholar

    [49] Hu M., Chen J., Jiang S., et al. (2022). E2SGAN: EEG-to-SEEG translation with generative adversarial networks. Front. Neurosci. 16:971829. DOI:10.3389/fnins.2022.971829

    View in Article CrossRef Google Scholar

    [50] Antoniades A., Spyrou L., Martin-Lopez D., et al. (2018). Deep neural architectures for mapping scalp to intracranial EEG. Int. J. Neural Syst. 28:1850009. DOI:10.1142/s0129065718500090

    View in Article CrossRef Google Scholar

    [51] Zhang R., Zeng Y., Tong L., et al. (2022). ERP-WGAN: A data augmentation method for EEG single-trial detection. J. Neurosci. Meth. 376:109621. DOI:10.1016/j.jneumeth.2022.109621

    View in Article CrossRef Google Scholar

    [52] Fahimi F., Dosen S., Ang K.K., et al. (2020). Generative adversarial networks-based data augmentation for brain-computer interface. IEEE Trans. Neural Netw. Learn. Syst. 32:4039−4051. DOI:10.1109/tnnls.2020.3016666

    View in Article CrossRef Google Scholar

    [53] Tang Y., Chen D., Liu H., et al. (2022). Deep EEG superresolution via correlating brain structural and functional connectivities. IEEE Trans. Cybern. 53:4410−4422. DOI:10.1109/tcyb.2022.3178370

    View in Article CrossRef Google Scholar

    [54] Gayon-Lombardo A., Mosser L., Brandon N.P., et al. (2020). Pores for thought: generative adversarial networks for stochastic reconstruction of 3D multi-phase electrode microstructures with periodic boundaries. npj Comput. Mater. 6:82. DOI:10.1038/s41524-020-0340-7

    View in Article CrossRef Google Scholar

    [55] Sanchez-Lengeling B. and Aspuru-Guzik A. (2018). Inverse molecular design using machine learning: Generative models for matter engineering. Science 361:360−365. DOI:10.1126/science.aat2663

    View in Article CrossRef Google Scholar

    [56] Manica M., Born J., Cadow J., et al. (2023). Accelerating material design with the generative toolkit for scientific discovery. npj Comput. Mater. 9:69. DOI:10.1038/s41524-023-01028-1

    View in Article CrossRef Google Scholar

    [57] Dan Y., Zhao Y., Li X., et al. (2020). Generative adversarial networks (GAN) based efficient sampling of chemical composition space for inverse design of inorganic materials. npj Comput. Mater. 6:84. DOI:10.1038/s41524-020-00352-0

    View in Article CrossRef Google Scholar

    [58] Guenther F.H., Brumberg J.S., Wright E.J., et al. (2009). A wireless brain-machine interface for real-time speech synthesis. PloS one 4:e8218. DOI:10.1371/journal.pone.0008218

    View in Article CrossRef Google Scholar

    [59] Tang J. and Huth A.G. (2025). Semantic language decoding across participants and stimulus modalities. Curr. Biol. 35:1023−1032. DOI:10.1016/j.cub.2025.01.024

    View in Article CrossRef Google Scholar

    [60] Komeiji S., Mitsuhashi T., Iimura Y., et al. (2024). Feasibility of decoding covert speech in ECoG with a transformer trained on overt speech. Sci. Rep. 14:11491. DOI:10.1101/2024.02.05.578911

    View in Article CrossRef Google Scholar

    [61] Fan C., Hahn N., Kamdar F., et al. (2023). Plug-and-play stability for intracortical braincomputer interfaces: a one-year demonstration of seamless brain-to-text communication. Advances in Neural Information Processing Systems 36:42258−42270. DOI:10.48550/arXiv.2311.03611

    View in Article CrossRef Google Scholar

    [62] Metzger S.L., Liu J.R., Moses D.A., et al. (2022). Generalizable spelling using a speech neuroprosthesis in an individual with severe limb and vocal paralysis. Nat. Commun. 13:6510. DOI:10.1038/s41467-022-33611-3

    View in Article CrossRef Google Scholar

    [63] Proix T., Delgado Saa J., Christen A., et al. (2022). Imagined speech can be decoded from low-and cross-frequency intracranial EEG features. Nat. Commun. 13:48. DOI:10.1038/s41467-021-27725-3

    View in Article CrossRef Google Scholar

    [64] Qi Y., Zhu X., Xiong X., et al. (2025). Human motor cortex encodes complex handwriting through a sequence of stable neural states. Nat. Hum. Behav. 9:1260−1271. DOI:10.1038/s41562-025-02157-x

    View in Article CrossRef Google Scholar

    [65] Miliotou E., Kyriakis P., Hinman J.D., et al. (2023). Generative decoding of visual stimuli. International Conference on Machine Learning 202:24775−24784. https://proceedings.mlr.press/v202/miliotou23a/miliotou23a.pdf.

    View in Article Google Scholar

    [66] Liu A., Jing H., Liu Y. et al. (2024). Hidden states in LLMs improve EEG representation learning and visual decoding. European Conference on Artificial Intelligence pp:2130–2137. DOI:10.3233/faia240732

    View in Article Google Scholar

    [67] Huang S., Sun L., Yousefnezhad M., et al. (2023). Functional alignment-auxiliary generative adversarial network-based visual stimuli reconstruction via multi-subject fMRI. IEEE Trans. Neural Syst. Rehabil. Eng. 31:2715−2725. DOI:10.1109/tnsre.2023.3283405

    View in Article CrossRef Google Scholar

    [68] Shimizu S., Ota A. and Nakane A. (2025). Reviving intentional facial expressions: An interface for ALS patients using brain decoding and image-generative AI. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems pp:1–10. DOI:10.1145/3706599.3719775

    View in Article Google Scholar

    [69] Mahajan G., Jeevan R., Divija L. et al. (2024). Deciphering EEG waves for the generation of images. 2024 12th International Winter Conference on Brain-Computer Interface pp:1–6. DOI:10.1109/bci60775.2024.10480501

    View in Article Google Scholar

    [70] Yeung J., Luo A.F., Sarch G. et al. (2025). Reanimating images using neural representations of dynamic stimuli. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:5331–5343. DOI:10.1109/cvpr52734.2025.00502

    View in Article Google Scholar

    [71] Allen E.J., St-Yves G., Wu Y., et al. (2022). A massive 7t fmri dataset to bridge cognitive neuroscience and artificial intelligence. Nat. Neurosci. 25:116−126. DOI:10.1038/s41593-021-00962-x

    View in Article CrossRef Google Scholar

    [72] Horikawa T., Tamaki M., Miyawaki Y., et al. (2013). Neural decoding of visual imagery during sleep. Science 340:639−642. DOI:10.1126/science.1234330

    View in Article CrossRef Google Scholar

    [73] Breedlove J.L., St-Yves G., Olman C.A., et al. (2020). Generative feedback explains distinct brain activity codes for seen and mental images. Curr. Biol. 30:2211−2224. DOI:10.1016/j.cub.2020.04.014

    View in Article CrossRef Google Scholar

    [74] Ritchie J.B., Kaplan D.M. and Klein C. (2019). Decoding the brain: Neural representation and the limits of multivariate pattern analysis in cognitive neuroscience. British J. Philos. Sci. 70:581−607. DOI:10.1093/bjps/axx023

    View in Article CrossRef Google Scholar

    [75] Wang S., Jiang C., Yu Y., et al. (2025). Tellurium nanowire retinal nanoprosthesis improves vision in models of blindness. Science 388:eadu2987. DOI:10.1126/science.adu2987

    View in Article CrossRef Google Scholar

    [76] Kravitz D.J., Saleem K.S., Baker C.I., et al. (2011). A new neural framework for visuospatial processing. Nat. Rev. Neurosci. 12:217−230. DOI:10.1038/nrn3008

    View in Article CrossRef Google Scholar

    [77] Kwon M., Han S., Kim K., et al. (2019). Super-resolution for improving EEG spatial resolution using deep convolutional neural network—feasibility study. Sensors 19:5317. DOI:10.3390/s19235317

    View in Article CrossRef Google Scholar

    [78] Bao G., Zhang Q., Gong Z., et al. (2025). Wills Aligner: Multi-subject collaborative brain visual decoding. Proceedings of the AAAI Conference on Artificial Intelligence 39:14194. DOI:10.1609/aaai.v39i13.33554

    View in Article CrossRef Google Scholar

    [79] Li D., Du C., Wang S., et al. (2021). Multi-subject data augmentation for target subject semantic decoding with deep multi-view adversarial learning. Inf. Sci. 547:1025−1044. DOI:10.1016/j.ins.2020.09.012

    View in Article CrossRef Google Scholar

    [80] Ozcelik F. and VanRullen R. (2023). Natural scene reconstruction from fMRI signals using generative latent diffusion. Sci. Rep. 13:15666. DOI:10.1038/s41598-023-42891-8

    View in Article CrossRef Google Scholar

    [81] Takagi Y. and Nishimoto S. (2023). High-resolution image reconstruction with latent diffusion models from human brain activity. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition pp:14453–14463. DOI:10.1109/cvpr52729.2023.01389

    View in Article Google Scholar

    [82] Du C., Fu K., Li J., et al. (2023). Decoding visual neural representations by multimodal learning of brain-visual-linguistic features. IEEE Trans. Pattern Anal. Mach. Intell. 45:10760−10777. DOI:10.1109/tpami.2023.3263181

    View in Article CrossRef Google Scholar

    [83] Shen G., Zhao D., He X., et al. (2024). Neuro-vision to language: Enhancing brain recording-based visual reconstruction and language interaction. Advances in Neural Information Processing Systems 37:98083−98110. DOI:10.52202/079017-3113

    View in Article CrossRef Google Scholar

    [84] Zhang A., Wang B., Wu X. et al. (2025). A novel multimodal method for decoding speech perception from brain activities. ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–5. DOI:10.1109/icassp49660.2025.10889637

    View in Article Google Scholar

    [85] Luo S., Angrick M., Coogan C., et al. (2023). Stable decoding from a speech BCI enables control for an individual with ALS without recalibration for 3 months. Adv. Sci. 10:2304853. DOI:10.1002/advs.202304853

    View in Article CrossRef Google Scholar

    [86] Xi N., Zhao S., Wang H. et al. (2023). Unicorn: Unified cognitive signal reconstruction bridging cognitive signals and human language. Proceedings of the Annual Meeting of the Association for Computational Linguistics pp:13277–13291. DOI:10.18653/v1/2023.acl-long.741

    View in Article Google Scholar

    [87] Guo J., Yi C., Li F. et al. (2024). MindLDM: Reconstruct visual stimuli from fMRI using latent diffusion model. 2024 IEEE International Conference on Computational Intelligence and Virtual Environments for Measurement Systems and Applications pp:1–6. DOI:10.1109/civemsa58715.2024.10586647

    View in Article Google Scholar

    [88] Wilson G.H., Stavisky S.D., Willett F.R., et al. (2020). Decoding spoken english from intracortical electrode arrays in dorsal precentral gyrus. J. Neural Eng. 17:066007. DOI:10.1088/1741-2552/abbfef

    View in Article CrossRef Google Scholar

    [89] Ferrante M., Boccato T., Ozcelik F. et al. (2023). Multimodal decoding of human brain activity into images and text. Advances in Neural Information Processing Systems vol. 243 pp:87–101. https://proceedings.mlr.press/v243/ferrante24a.

    View in Article Google Scholar

    [90] Scotti P.S., Tripathy M., Villanueva C.K.T. et al. (2024). MindEye2: Shared-subject models enable fMRI-to-image with 1 hour of data. Proceedings of the 41st International Conference on Machine Learning pp:44038–44059. DOI:10.48550/arXiv.2403.11207

    View in Article Google Scholar

    [91] Wang S., Liu S., Tan Z. et al. (2024). Mindbridge: A cross-subject brain decoding framework. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:11333–11342. DOI:10.1109/cvpr52733.2024.01077

    View in Article Google Scholar

    [92] Gong Z., Bao G., Zhang Q., et al. (2024). NeuroClips: Towards high-fidelity and smooth fMRI-to-video reconstruction. Advances in Neural Information Processing Systems 37:51655−51683. DOI:10.52202/079017-1636

    View in Article CrossRef Google Scholar

    [93] Vansteensel M.J., Pels E.G., Bleichner M.G., et al. (2016). Fully implanted brain-computer interface in a locked-in patient with ALS. N. Engl. J. Med. 375:2060−2066. DOI:10.1056/nejmoa1608085

    View in Article CrossRef Google Scholar

    [94] Willett F.R., Avansino D.T., Hochberg L.R., et al. (2021). High-performance brain-to-text communication via handwriting. Nature 593:249−254. DOI:10.1038/s41586-021-03506-2

    View in Article CrossRef Google Scholar

    [95] Martin S., Brunner P., Iturrate I., et al. (2016). Word pair classification during imagined speech using direct brain recordings. Sci. Rep. 6:25803. DOI:10.1038/srep25803

    View in Article CrossRef Google Scholar

    [96] Herff C., Heger D., De Pesters A., et al. (2015). Brain-to-text: decoding spoken phrases from phone representations in the brain. Front. Neurosci. 8:141498. DOI:10.3389/fnins.2015.00217

    View in Article CrossRef Google Scholar

    [97] Moses D.A., Leonard M.K., Makin J.G., et al. (2019). Real-time decoding of question-andanswer speech dialogue using human cortical activity. Nat. Commun. 10:3096. DOI:10.1038/s41467-019-10994-4

    View in Article CrossRef Google Scholar

    [98] Moses D.A., Metzger S.L., Liu J.R., et al. (2021). Neuroprosthesis for decoding speech in a paralyzed person with anarthria. N. Engl. J. Med. 385:217−227. DOI:10.1056/nejmoa2027540

    View in Article CrossRef Google Scholar

    [99] Duraivel S., Rahimpour S., Chiang C.H., et al. (2023). High-resolution neural recordings improve the accuracy of speech decoding. Nat. Commun. 14:6938. DOI:10.1038/s41467-023-42555-1

    View in Article CrossRef Google Scholar

    [100] Zhang D., Wang Z., Qian Y., et al. (2024). A brain-to-text framework for decoding natural tonal sentences. Cell Rep. 43:114924. DOI:10.1016/j.celrep.2024

    View in Article CrossRef Google Scholar

    [101] Makin J.G., Moses D.A. and Chang E.F. (2020). Machine translation of cortical activity to text with an encoder-decoder framework. Nat. Neurosci. 23:575−582. DOI:10.1038/s41593-020-0608-8

    View in Article CrossRef Google Scholar

    [102] Sun P., Anumanchipalli G.K. and Chang E.F. (2020). Brain2Char: a deep architecture for decoding text from brain recordings. J. Neural Eng. 17:066015. DOI:10.1088/1741-2552/abc742

    View in Article CrossRef Google Scholar

    [103] Simanova I., Hagoort P., Oostenveld R., et al. (2014). Modality-independent decoding of semantic information from the human brain. Cereb. Cortex 24:426−434. DOI:10.1093/cercor/bhs324

    View in Article CrossRef Google Scholar

    [104] Defossez A., Caucheteux C., Rapin J., et al. (2023). Decoding speech perception ′ from non-invasive brain recordings. Nat. Mach. Intell. 5:1097−1107. DOI:10.1038/s42256-023-00714-5

    View in Article CrossRef Google Scholar

    [105] Chen Q., Wang Y., Wang F., et al. (2025). Decoding text from electroencephalography signals: a novel hierarchical gated recurrent unit with masked residual attention mechanism. Eng. Appl. Artif. Intel. 139:109615. DOI:10.1016/j.engappai.2024.109615

    View in Article CrossRef Google Scholar

    [106] Wang Z. and Ji H. (2022). Open vocabulary electroencephalography-to-text decoding and zero-shot sentiment classification. Proceedings of the AAAI Conference on Artificial Intelligence 36:5350−5358. DOI:10.1609/aaai.v36i5.20472

    View in Article CrossRef Google Scholar

    [107] Duan Y., Zhou J., Wang Z., et al. (2023). Dewave: Discrete encoding of EEG waves for EEG to text translation. Advances in Neural Information Processing Systems 36:9907−9918. DOI:10.48550/arXiv.2309.14030

    View in Article CrossRef Google Scholar

    [108] Chen X., Du C., Liu C. et al. (2025). BP-GPT: Auditory neural decoding using fmriprompted llm. ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–5. DOI:10.1109/icassp49660.2025.10890142

    View in Article Google Scholar

    [109] Liu H., Hajialigol D., Antony B. et al. (2024). EEG2Text: Open vocabulary EEG-to-text translation with multi-view transformer. 2024 IEEE International Conference on Big Data pp:1824–1833. DOI:10.1109/bigdata62323.2024.10825980

    View in Article Google Scholar

    [110] Zhou J., Duan Y., Chang Y.C., et al. (2024). BELT: bootstrapped EEG-to-language training by natural language supervision. IEEE Trans. Neural Syst. Rehabil. Eng. 32:3278−3288. DOI:10.1109/tnsre.2024.3450795

    View in Article CrossRef Google Scholar

    [111] Amrani H., Micucci D. and Napoletano P. (2024). Deep representation learning for open vocabulary electroencephalography-to-text decoding. IEEE J. Biomed. Health Inform. pp:1–12. DOI:10.1109/JBHI.2024.3416066

    View in Article Google Scholar

    [112] Tao Y., Liang Y., Wang L. et al. (2025). See: Semantically aligned EEG-to-text translation. ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–5. DOI:10.1109/icassp49660.2025.10889658

    View in Article Google Scholar

    [113] Wang J., Song Z., Ma Z. et al. (2024). Enhancing EEG-to-text decoding through transferable representations from pre-trained contrastive EEG-text masked autoencoder. Proceedings of the Annual Meeting of the Association for Computational Linguistics pp:7278–7292. DOI:10.18653/v1/2024.acl-long.393

    View in Article Google Scholar

    [114] Luo K. (2024). Real-time open-vocabulary sentence decoding from MEG signals using transformers. 2024 International Conference on Image Processing, Computer Vision and Machine Learning pp:460–463. DOI:10.1109/icicml63543.2024.10957766

    View in Article Google Scholar

    [115] Boyko M., Druzhinina P., Kormakov G. et al. (2024). MEGFormer: enhancing speech decoding from brain activity through extended semantic representations. International Conference on Medical Image Computing and Computer-Assisted Intervention pp:281–290. DOI:10.1007/978-3-031-72069-7_27

    View in Article Google Scholar

    [116] Wang B., Xu X., Zhang L. et al. (2024). Semantic reconstruction of continuous language from meg signals. ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing pp:2190–2194. DOI:10.1109/icassp48485.2024.10448281

    View in Article Google Scholar

    [117] Guo Y., Dong Y., Ng M.K.P. et al. (2025). A pre-trained framework for multilingual brain decoding using non-invasive recordings. arXiv Preprint arXiv:2506.03214. DOI:10.48550/arXiv.2506.03214

    View in Article Google Scholar

    [118] Silva A.B., Liu J.R., Metzger S.L., et al. (2024). A bilingual speech neuroprosthesis driven by cortical articulatory representations shared between languages. Nat. Biomed. Eng. 8:977−991. DOI:10.1038/s41551-024-01207-5

    View in Article CrossRef Google Scholar

    [119] Huang W., Yang P., Tang Y., et al. (2024). From sight to insight: A multi-task approach with the visual language decoding model. Inf. Fusion 112:102573. DOI:10.1016/j.inffus.2024.102573

    View in Article CrossRef Google Scholar

    [120] Han J., Gong K., Zhang Y. et al. (2024). Onellm: One framework to align all modalities with language. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:26584–26595. DOI:10.1109/cvpr52733.2024.02510

    View in Article Google Scholar

    [121] Chen J., Qi Y., Wang Y., et al. (2025). Mindgpt: Interpreting what you see with non-invasive brain recordings. IEEE Trans. Image Process. 34:3281−3293. DOI:10.1109/tip.2025.3572784

    View in Article CrossRef Google Scholar

    [122] Tang J., Yang Y., Zhao Q., et al. (2024). Visual guided dual-spatial interaction network for fine-grained brain semantic decoding. IEEE Trans. Instrum. Meas. 73:2534214. DOI:10.1109/tim.2024.3480232

    View in Article CrossRef Google Scholar

    [123] Ikegawa Y., Fukuma R., Sugano H., et al. (2024). Text and image generation from intracranial electroencephalography using an embedding space for text and images. J. Neural Eng. 21:036019. DOI:10.1088/1741-2552/ad417a

    View in Article CrossRef Google Scholar

    [124] Willett F.R., Kunz E.M., Fan C., et al. (2023). A high-performance speech neuroprosthesis. Nature 620:1031−1036. DOI:10.1101/2023.01.21.524489

    View in Article CrossRef Google Scholar

    [125] Ramsey N.F., Salari E., Aarnoutse E.J., et al. (2018). Decoding spoken phonemes from sensorimotor cortex with high-density ECoG grids. Neuroimage 180:301−311. DOI:10.1016/j.neuroimage.2017.10.011

    View in Article CrossRef Google Scholar

    [126] Angrick M., Herff C., Mugler E., et al. (2019). Speech synthesis from ECoG using densely connected 3d convolutional neural networks. J. Neural Eng. 16:036019. DOI:10.1088/1741-2552/ab0c59

    View in Article CrossRef Google Scholar

    [127] Herff C., Diener L., Angrick M., et al. (2019). Generating natural, intelligible speech from brain activity in motor, premotor, and inferior frontal cortices. Front. Neurosci. 13:1267. DOI:10.3389/fnins.2019.01267

    View in Article CrossRef Google Scholar

    [128] Berezutskaya J., Freudenburg Z.V., Vansteensel M.J., et al. (2023). Direct speech reconstruction from sensorimotor brain activity with optimized deep learning models. J. Neural Eng. 20:056010. DOI:10.1088/1741-2552/ace8be

    View in Article CrossRef Google Scholar

    [129] Anumanchipalli G.K., Chartier J. and Chang E.F. (2019). Speech synthesis from neural decoding of spoken sentences. Nature 568:493−498. DOI:10.1038/s41586-019-1119-1

    View in Article CrossRef Google Scholar

    [130] Shigemi K., Komeiji S., Mitsuhashi T. et al. (2023). Synthesizing speech from ECoG with a combination of transformer-based encoder and neural vocoder. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–5. DOI:10.1109/icassp49357.2023.10097004

    View in Article Google Scholar

    [131] Ticha M.B.B., Ran X., Roussel P. et al. (2024). A vision transformer architecture for overt speech decoding from ECoG data. 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society pp:1–4. DOI:10.1109/embc53108.2024.10781877

    View in Article Google Scholar

    [132] Chen X., Wang R., Khalilian-Gourtani A., et al. (2024). A neural speech decoding framework leveraging deep learning and speech synthesis. Nat. Mach. Intell. 6:467−480. DOI:10.1038/s42256-024-00824-8

    View in Article CrossRef Google Scholar

    [133] Meng K., Goodarzy F., Kim E., et al. (2023). Continuous synthesis of artificial speech sounds from human cortical surface recordings during silent speech production. J. Neural Eng. 20:046019. DOI:10.1088/1741-2552/ace7f6

    View in Article CrossRef Google Scholar

    [134] Wu X., Wellington S., Fu Z., et al. (2024). Speech decoding from stereoelectroencephalography (sEEG) signals using advanced deep learning methods. J. Neural Eng. 21:036055. DOI:10.1088/1741-2552/ad593a

    View in Article CrossRef Google Scholar

    [135] Angrick M., Ottenhoff M.C., Diener L., et al. (2021). Real-time synthesis of imagined speech processes from minimally invasive recordings of neural activity. Commun. Biol. 4:1055. DOI:10.1038/s42003-021-02578-0

    View in Article CrossRef Google Scholar

    [136] Angrick M., Luo S., Rabbani Q., et al. (2024). Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS. Sci. Rep. 14:9617. DOI:10.1038/s41598-024-60277-2

    View in Article CrossRef Google Scholar

    [137] Card N.S., Wairagkar M., Iacobacci C., et al. (2024). An accurate and rapidly calibrating speech neuroprosthesis. N. Engl. J. Med. 391:609−618. DOI:10.1056/nejmoa2314132

    View in Article CrossRef Google Scholar

    [138] Wu H., Cai C., Ming W., et al. (2024). Speech decoding using cortical and subcortical electrophysiological signals. Front. Neurosci. 18:1345308. DOI:10.3389/fnins.2024.1345308

    View in Article CrossRef Google Scholar

    [139] Feng C., Cao L., Wu D., et al. (2025). Acoustic inspired brain-to-sentence decoder for logosyllabic language. Cyborg Bionic Syst. 6:0257. DOI:10.34133/cbsystems.0257

    View in Article CrossRef Google Scholar

    [140] Accou B., Vanthornhout J., hamme H.V., et al. (2023). Decoding of the speech envelope from EEG using the VLAAI deep neural network. Sci. Rep. 13:812. DOI:10.1038/s41598-022-27332-2

    View in Article CrossRef Google Scholar

    [141] Van Dyck B., Yang L. and Van Hulle M.M. (2023). Decoding auditory EEG responses using an adapted wavenet. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–2. DOI:10.1109/icassp49357.2023.10095420

    View in Article Google Scholar

    [142] Fan C., Zhang S., Zhang J. et al. (2025). SSM2Mel: State space model to reconstruct Mel spectrogram from the EEG. ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing pp:1–5. DOI:10.1109/icassp49660.2025.10888785

    View in Article Google Scholar

    [143] Qi D., Kong L., Yang L. et al. (2023). Audiodiffusion: Generating high-quality audios from EEG signals: Reconstructing audio from EEG signals. 2023 4th International Symposium on Computer Engineering and Intelligent Communications pp:344–348. DOI:10.1109/isceic59030.2023.10271237

    View in Article Google Scholar

    [144] Xiong W., Ma L. and Li H. (2025). Synthesizing intelligible utterances from EEG of imagined speech. Front. Neurosci. 19:1565848. DOI:10.3389/fnins.2025.1565848

    View in Article CrossRef Google Scholar

    [145] Lee J.W., Lee S.H., Lee Y.E. et al. (2023). Sentence reconstruction leveraging contextual meaning from speech-related brain signals. 2023 IEEE International Conference on Systems, Man, and Cybernetics pp:3721–3726. DOI:10.1109/smc53992.2023.10394348

    View in Article Google Scholar

    [146] Park J.Y., Tsukamoto M., Tanaka M., et al. (2025). Natural sounds can be reconstructed from human neuroimaging data using deep neural network representation. PLoS Biol. 23:e3003293. DOI:10.1371/journal.pbio.3003293

    View in Article CrossRef Google Scholar

    [147] Craik A., Dial H. and Contreras-Vidal J.L. (2025). Continuous and discrete decoding of overt speech with scalp electroencephalography (EEG). J. Neural Eng. 22:026017. DOI:10.1088/1741-2552/ad8d0a

    View in Article CrossRef Google Scholar

    [148] Dado T., Papale P., Lozano A., et al. (2024). Brain2GAN: Feature-disentangled neural encoding and decoding of visual perception in the primate brain. PLoS Comput. Biol. 20:e1012058. DOI:10.1371/journal.pcbi.1012058

    View in Article CrossRef Google Scholar

    [149] Doerig A., Kietzmann T.C., Allen E., et al. (2025). High-level visual representations in the human brain are aligned with large language models. Nat. Mach. Intell. 7:1220−1234. DOI:10.1038/s42256-025-01072-0

    View in Article CrossRef Google Scholar

    [150] Wang H., Zhang Q., Wang L. et al. (2025). Neurons: Emulating the human visual cortex improves fidelity and interpretability in fMRI-to-video reconstruction. Proceedings of the IEEE/CVF International Conference on Computer Vision pp:18367–18376. DOI:10.48550/arXiv.2503.11167

    View in Article Google Scholar

    [151] Xia R., Yin C. and Li P. (2024). Decoding the echoes of vision from fmri: Memory disentangling for past semantic information. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing pp:2040–2052. DOI:10.18653/v1/2024.emnlp-main.122

    View in Article Google Scholar

    [152] Meng L. and Yang C. (2024). Semantics-guided hierarchical feature encoding generative adversarial network for visual image reconstruction from brain activity. IEEE Trans. Neural Syst. Rehabil. Eng. 32:1267−1283. DOI:10.1109/tnsre.2024.3377698

    View in Article CrossRef Google Scholar

    [153] Dado T., Guçlütürk Y., Ambrogioni L., et al. (2022). Hyperrealistic neural decoding for reconstructing faces from fMRI activations via the GAN latent space. Sci. Rep. 12:141. DOI:10.1038/s41598-021-03938-w

    View in Article CrossRef Google Scholar

    [154] Quan R., Wang W., Tian Z. et al. (2025). Psychometry: An omnifit model for image reconstruction from human brain activity. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:233–243. DOI:10.1109/cvpr52733.2024.00030

    View in Article Google Scholar

    [155] Wu H., Li Q., Zhang C. et al. (2025). Bridging the vision-brain gap with an uncertainty aware blur prior. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:2246–2257. DOI:10.1109/cvpr52734.2025.00215

    View in Article Google Scholar

    [156] Miyawaki Y., Uchida H., Yamashita O., et al. (2008). Visual image reconstruction from human brain activity using a combination of multiscale local image decoders. Neuron 60:915−929. DOI:10.1016/j.neuron.2008.11.004

    View in Article CrossRef Google Scholar

    [157] Du C., Du C., Huang L., et al. (2018). Reconstructing perceived images from human brain activities with bayesian deep multiview learning. IEEE Trans. Neural Netw. Learn. Syst. 30:2310−2323. DOI:10.1109/tnnls.2018.2882456

    View in Article CrossRef Google Scholar

    [158] Güçlütürk Y., Güçlü U., Seeliger K., et al. (2017). Reconstructing perceived faces from brain activations with deep adversarial neural decoding. Advances in Neural Information Processing Systems 30:4249−4260.

    View in Article Google Scholar

    [159] VanRullen R. and Reddy L. (2019). Reconstructing faces from fMRI patterns using deep generative neural networks. Commun. Biol. 2:193. DOI:10.1038/s42003-019-0438-y

    View in Article CrossRef Google Scholar

    [160] Du C., Du C., Huang L., et al. (2020). Conditional generative neural decoding with structured cnn feature prediction. Proceedings of the AAAI Conference on Artificial Intelligence 34:2629−2636. DOI:10.1609/aaai.v34i03.5647

    View in Article CrossRef Google Scholar

    [161] Naselaris T., Prenger R.J., Kay K.N., et al. (2009). Bayesian reconstruction of natural images from human brain activity. Neuron 63:902−915. DOI:10.1016/j.neuron.2009.09.006

    View in Article CrossRef Google Scholar

    [162] Beliy R., Gaziv G., Hoogi A., et al. (2019). From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI. Advances in Neural Information Processing Systems 32:6517−6527. DOI:10.48550/arXiv.1907.02431

    View in Article CrossRef Google Scholar

    [163] Shen G., Horikawa T., Majima K., et al. (2019). Deep image reconstruction from human brain activity. PLOS Comput. Biol. 15:e1006633. DOI:10.1371/journal.pcbi.1006633

    View in Article CrossRef Google Scholar

    [164] Luo J., Cui W., Liu J., et al. (2022). Visual image decoding of brain activities using a dual attention hierarchical latent generative network with multiscale feature fusion. IEEE Trans. Cogn. Develop. Syst. 15:761−773. DOI:10.1109/tcds.2022.3181469

    View in Article CrossRef Google Scholar

    [165] Akamatsu Y., Harakawa R., Ogawa T., et al. (2020). Brain decoding of viewed image categories via semi-supervised multi-view bayesian generative model. IEEE Trans. Signal Process. 68:5769−5781. DOI:10.1109/tsp.2020.3028701

    View in Article CrossRef Google Scholar

    [166] Li D., Du C. and He H. (2020). Semi-supervised cross-modal image generation with generative adversarial networks. Pattern Recognit. 100:107085. DOI:10.1016/j.patcog.2019.107085

    View in Article CrossRef Google Scholar

    [167] Gaziv G., Beliy R., Granot N., et al. (2022). Self-supervised natural image reconstruction and large-scale semantic classification from brain activity. NeuroImage 254:119121. DOI:10.1016/j.neuroimage.2022.119121

    View in Article CrossRef Google Scholar

    [168] Zhou Q., Du C., Li D., et al. (2025). Interpretable visual neural decoding with unsupervised semantic disentanglement. Mach. Intell. Res. 22:553−570. DOI:10.1007/s11633-023-1484-y

    View in Article CrossRef Google Scholar

    [169] Tirupattur P., Rawat Y.S., Spampinato C. et al. (2018). Thoughtviz: Visualizing human thoughts using generative adversarial network. Proceedings of the 26th ACM International Conference on Multimedia pp:950–958. DOI:10.1145/3240508.3240641

    View in Article Google Scholar

    [170] Palazzo S., Spampinato C., Kavasidis I. et al. (2017). Generative adversarial networks conditioned by brain signals. Proceedings of the IEEE international conference on computer vision pp:3410–3418. DOI:10.1109/iccv.2017.369

    View in Article Google Scholar

    [171] Kavasidis I., Palazzo S., Spampinato C. et al. (2017). Brain2image: Converting brain signals into images. Proceedings of the 25th ACM International Conference on Multimedia pp:1809–1817. DOI:10.1145/3123266.3127907

    View in Article Google Scholar

    [172] Ahmadieh H., Gassemi F. and Moradi M.H. (2024). Visual image reconstruction based on EEG signals using a generative adversarial and deep fuzzy neural network. Biomed. Signal Proces. 87:105497. DOI:10.1016/j.bspc.2023.105497

    View in Article CrossRef Google Scholar

    [173] Fares A., Zhong S.h. and Jiang J. (2020). Brain-media: A dual conditioned and lateralization supported GAN (DCLS-GAN) towards visualization of image-evoked brain activities. Proceedings of the 28th ACM International Conference on Multimedia pp:1764–1772. DOI:10.1145/3394171.3413858

    View in Article Google Scholar

    [174] Huang W., Yan H., Wang C., et al. (2020). Perception-to-image: Reconstructing natural images from the brain activity of visual perception. Ann. Biomed. Eng. 48:2323−2332. DOI:10.1007/s10439-020-02502-3

    View in Article CrossRef Google Scholar

    [175] Lu Y., Du C., Zhou Q. et al. (2023). Minddiffuser: Controlled image reconstruction from human brain activity with semantic and structural diffusion. Proceedings of the 31st ACM International Conference on Multimedia pp:5899–5908. DOI:10.1145/3581783.3613832

    View in Article Google Scholar

    [176] Lu Y., Du C., Wang C. et al. (2025). Animate your thoughts: Reconstruction of dynamic natural vision from human brain activity. 13th International Conference on Learning Representations pp:23255–23293. DOI:10.48550/arXiv.2405.03280

    View in Article Google Scholar

    [177] Li J., Li D., Xiong C. et al. (2022). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. International Conference on Machine Learning pp:12888–12900. DOI:10.48550/arXiv.2201.12086

    View in Article Google Scholar

    [178] Luo A., Henderson M., Wehbe L., et al. (2023). Brain diffusion for visual exploration: Cortical discovery using large scale generative models. Advances in Neural Information Processing Systems 36:75740−75781.

    View in Article Google Scholar

    [179] Ferrante M., Boccato T., Passamonti L., et al. (2024). Retrieving and reconstructing conceptually similar images from fmri with latent diffusion models and a neuro-inspired brain decoding model. J. Neural Eng. 21:046001. DOI:10.1088/1741-2552/ad593c

    View in Article CrossRef Google Scholar

    [180] Chen Z., Qing J., Xiang T. et al. (2023). Seeing beyond the brain: Conditional diffusion model with sparse masked modeling for vision decoding. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition pp:22710–22720. DOI:10.1109/cvpr52729.2023.02175

    View in Article Google Scholar

    [181] Bai Y., Wang X., Cao Y.P. et al. (2024). Dreamdiffusion: High-quality EEG-to-image gener ation with temporal masked signal modeling and CLIP alignment. Computer Vision–ECCV 2024: 18th European Conference pp:472–488. DOI:10.1007/978-3-031-72751-1_27

    View in Article Google Scholar

    [182] Liu Q., Zhu H., Chen N., et al. (2024). Mind-bridge: reconstructing visual images based on diffusion model from human brain activity. Signal Image Video Process. 18:953−963. DOI:10.1007/s11760-024-03207-z

    View in Article CrossRef Google Scholar

    [183] Zhou Q., Du C., Wang S. et al. (2024). CLIP-MUSED: CLIP-guided multi-subject visual neural information semantic decoding. 12th International Conference on Learning Representations pp:43466–43482. DOI:10.48550/arXiv.2402.08994

    View in Article Google Scholar

    [184] Ma Y., Liu Y., Chen L., et al. (2025). BrainCLIP: Brain representation via CLIP for generic natural visual stimulus decoding. IEEE Trans. Med. Imaging 44:3962−3972. DOI:10.1109/TMI.2025.3537287

    View in Article CrossRef Google Scholar

    [185] Bhalerao S.V. and Pachori R.B. (2024). Automated classification of cognitive visual objects using multivariate swarm sparse decomposition from multichannel EEG-MEG signals. IEEE Trans. Hum.-Mach. Syst. 54:455−464. DOI:10.1109/thms.2024.3395153

    View in Article CrossRef Google Scholar

    [186] Gong Z., Zhang Q., Bao G., et al. (2025). Mindtuner: Cross-subject visual decoding with visual fingerprint and semantic correction. Proceedings of the AAAI Conference on Artificial Intelligence 39:14247−14255. DOI:10.1609/aaai.v39i13.33560

    View in Article CrossRef Google Scholar

    [187] Luo A.F., Henderson M.M., Tarr M.J. et al. (2024). BrainSCUBA: Fine-grained natural language captions of visual cortex selectivity. 12th International Conference on Learning Representations pp:1744–1780. DOI:10.48550/arXiv.2310.0442

    View in Article Google Scholar

    [188] Zhao Y., Dong G., Zhu L., et al. (2025). Memory recall: Retrieval-augmented mind reconstruction for brain decoding. Inf. Fusion 123:103280. DOI:10.1016/j.inffus.2025.103280

    View in Article CrossRef Google Scholar

    [189] Huo J., Wang Y., Wang Y. et al. (2024). Neuropictor: Refining fMRI-to-image reconstruction via multi-individual pretraining and multi-level modulation. European Conference on Computer Vision pp:56–73. DOI:10.1007/978-3-031-72983-6_4

    View in Article Google Scholar

    [190] Liu M. and Kobayashi I. (2024). Do feature representations from different language models affect accuracy of brain encoding models’ predictions? 2024 IEEE International Conference on Systems, Man, and Cybernetics pp:2766–2771. DOI:10.1109/smc54092.2024.10831584

    View in Article Google Scholar

    [191] Singh A.K., Wang Y.K., King J.T., et al. (2020). Extended interaction with a BCI video game changes resting-state brain activity. IEEE Trans. Cogn. Develop. Syst. 12:809−823. DOI:10.1109/tcds.2020.2985102

    View in Article CrossRef Google Scholar

    [192] Nishimoto S., Vu A.T., Naselaris T., et al. (2011). Reconstructing visual experiences from brain activity evoked by natural movies. Curr. Biol. 21:1641−1646. DOI:10.1016/j.cub.2011.08.031

    View in Article CrossRef Google Scholar

    [193] Hanke M., Baumgartner F.J., Ibe P., et al. (2014). A high-resolution 7-Tesla fMRI dataset from complex natural stimulation with an audio movie. Sci. Data 1:14003. DOI:10.1038/sdata.2014.3

    View in Article CrossRef Google Scholar

    [194] Sun J., Li M. and Moens M.F. (2025). Neuralflix: A simple while effective framework for semantic decoding of videos from non-invasive brain recordings. Proceedings of the AAAI Conference on Artificial Intelligence 39:7096−7104. DOI:10.1609/aaai.v39i7.32762

    View in Article CrossRef Google Scholar

    [195] Fosco C., Lahner B., Pan B. et al. (2024). Brain netflix: Scaling data to reconstruct videos from brain signals. European Conference on Computer Vision pp:457–474. DOI:10.1007/978-3-031-73347-5_26

    View in Article Google Scholar

    [196] Simonyan K. and Zisserman A. (2014). Two-stream convolutional networks for action recognition in videos. Advances in Neural Information Processing Systems 27:568−576. DOI:10.48550/arXiv.1406.2199

    View in Article CrossRef Google Scholar

    [197] Carreira J. and Zisserman A. (2017). Quo vadis, action recognition? a new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp:6299–6308. DOI:10.1109/cvpr.2017.502

    View in Article Google Scholar

    [198] Chen Z., Qing J. and Zhou J.H. (2023). Cinematic mindscapes: High-quality video re construction from brain activity. Advances in Neural Information Processing Systems 36:24841−24858. DOI:10.48550/arXiv.2305.11675

    View in Article CrossRef Google Scholar

    [199] Liu X.H., Liu Y.K., Wang Y., et al. (2024). EEG2video: Towards decoding dynamic visual perception from EEG signals. Advances in Neural Information Processing Systems 37:72245−72273. DOI:10.52202/079017-2306

    View in Article CrossRef Google Scholar

    [200] Lin J., Chen H., Fan Y. et al. (2025). Multi-layer visual feature fusion in multimodal llms: 1167 Methods, analysis, and best practices. Proceedings of the IEEE/CVF Conference on 1168 Computer Vision and Pattern Recognition pp:4156−4166. DOI:10.1109/cvpr52734.2025.00393

    View in Article CrossRef Google Scholar

    [201] Zhang B., Fang Y., Ren T. et al. (2022). Multimodal analysis for deep video understanding with video language transformer. Proceedings of the 30th ACM International Conference on Multimedia pp:7165–7169. DOI:10.1145/3503161.3551600

    View in Article Google Scholar

    [202] Holz F.G., Le Mer Y., Muqit M.M. et al. (2025). Subretinal photovoltaic implant to restore vision in geographic atrophy due to AMD. N. Engl. J. Med. 394:232-242. DOI:10.1056/nejmoa2501396

    View in Article CrossRef Google Scholar

    [203] Elnabawy R.H., Abdennadher S., Hellwich O., et al. (2022). PVGAN: a generative adversarial network for object simplification in prosthetic vision. J. Neural Eng. 19:056007. DOI:10.1088/1741-2552/ac8acf

    View in Article CrossRef Google Scholar

    [204] Chen X., Wang F., Fernandez E., et al. (2020). Shape perception via a high-channelcount neuroprosthesis in monkey visual cortex. Science 370:1191−1196. DOI:10.1126/science.abd7435

    View in Article CrossRef Google Scholar

    [205] Luo T.j., Fan Y., Chen L., et al. (2020). EEG signal reconstruction using a generative adversarial network with wasserstein distance and temporal-spatial-frequency loss. Front. Neuroinf. 14:15. DOI:10.3389/fninf.2020.00015

    View in Article CrossRef Google Scholar

    [206] Aznan N.K.N., Atapour-Abarghouei A., Bonner S. et al. (2019). Simulating brain signals: Creating synthetic EEG data via neural-based generative models for improved SSVEP classification. 2019 International Joint Conference on Neural Networks pp:1–8. DOI:10.1109/ijcnn.2019.8852227

    View in Article Google Scholar

    [207] Xie J., Chen S., Zhang Y., et al. (2021). Combining generative adversarial networks and multi-output CNN for motor imagery classification. J. Neural Eng. 18:046026. DOI:10.1088/1741-2552/abecc5

    View in Article CrossRef Google Scholar

    [208] Mao M., Komes D., Zhao S., et al. (2025). Non-invasive neuromodulation assisted by exogenous stimuli-responsive nanoplatforms for Alzheimer’s disease and Parkinson’s disease therapy. Innov. Med. 3:100121. DOI:10.59717/j.xinn-med.2025.100121

    View in Article CrossRef Google Scholar

    [209] Yuste R., Goering S., Arcas B.A.Y., et al. (2017). Four ethical priorities for neurotechnologies and AI. Nature 551:159−163. DOI:10.1038/551159a

    View in Article CrossRef Google Scholar

    [210] Bahador N., Jokelainen J., Mustola S., et al. (2021). Reconstruction of missing channel in electroencephalogram using spatiotemporal correlation-based averaging. J. Neural Eng. 18:056045. DOI:10.1088/1741-2552/ac23e2

    View in Article CrossRef Google Scholar

    [211] An D., Guo Y., Lei N. et al. (2020). AE-OT: A new generative model based on extended semi-discrete optimal transport. 8th International Conference on Learning Representations pp: 8996–9014. https://openreview.net/pdf?id=HkldyTNYwH.

    View in Article Google Scholar

    [212] Anderson A.J., McDermott K., Rooks B., et al. (2020). Decoding individual identity from brain activity elicited in imagining common experiences. Nat. Commun. 11:5916. DOI:10.1038/s41467-020-19630-y

    View in Article CrossRef Google Scholar

    [213] Li C., Qian X., Wang Y. et al. (2024). Enhancing cross-subject fMRI-to-video decoding with global-local functional alignment. European Conference on Computer Vision pp:353–369. DOI:10.1007/978-3-031-73010-8_21

    View in Article Google Scholar

    [214] Wen S., Yin A., Furlanello T., et al. (2023). Rapid adaptation of brain-computer interfaces to new neuronal ensembles or participants via generative modelling. Nat. Biomed. Eng. 7:546−558. DOI:10.1038/s41551-021-00811-z

    View in Article CrossRef Google Scholar

    [215] Karpowicz B.M., Ali Y.H., Wimalasena L.N., et al. (2025). Stabilizing brain-computer interfaces through alignment of latent dynamics. Nat. Commun. 16:4662. DOI:10.1038/s41467-025-59652-y

    View in Article CrossRef Google Scholar

    [216] Ahuja C. and Sethia D. (2025). SS-EMERGE - self-supervised enhancement for multidimension emotion recognition using GNNs for EEG. Sci. Rep. 15:14254. DOI:10.1038/s41598-025-98623-7

    View in Article CrossRef Google Scholar

    [217] Chen R.J., Wang J.J., Williamson D.F., et al. (2023). Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng. 7:719−742. DOI:10.1038/s41551-023-01056-8

    View in Article CrossRef Google Scholar

    [218] Yang Y., Zhang H., Gichoya J.W., et al. (2024). The limits of fair medical imaging AI in real-world generalization. Nat. Med. 30:2838−2848. DOI:10.1038/s41591-024-03113-4

    View in Article CrossRef Google Scholar

    [219] Yao W., Chen X. and Wang S. (2025). Empowering functional neuroimaging: A pretrained generative framework for unified representation of neural signals. arXiv Preprint arXiv:2506.02433. DOI:10.48550/arXiv.2506.02433

    View in Article Google Scholar

    [220] Liu R., Chen Y., Li A., et al. (2024). Aggregating intrinsic information to enhance BCI performance through federated learning. Neural Netw. 172:106100. DOI:10.1016/j.neunet.2024.106100

    View in Article CrossRef Google Scholar

    [221] Dai J., Xu H., Chen T., et al. (2025). Artificial intelligence for medicine 2025: Navigating the endless frontier. Innov. Med. 3:100120. DOI:10.59717/j.xinn-med.2025.100120

    View in Article CrossRef Google Scholar

    [222] Huang T., Xu H., Wang H., et al. (2023). Artificial intelligence for medicine: Progress, challenges, and perspectives. Innov. Med. 1:10030. DOI:10.59717/j.xinn-med.2023.100030

    View in Article CrossRef Google Scholar

  • Cite this article:

    Guo Y., Ren G., Ma S., et al. (2026). Generative AI for brain-computer interfaces decoding: Advances, challenges and future. The Innovation Medicine 4:100193. https://doi.org/10.59717/j.xinn-med.2026.100193
    Guo Y., Ren G., Ma S., et al. (2026). Generative AI for brain-computer interfaces decoding: Advances, challenges and future. The Innovation Medicine 4:100193. https://doi.org/10.59717/j.xinn-med.2026.100193

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(6)     Tables(2)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(21221) PDF downloads(3001)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint