Medical Physics, AI and Neurotechnology · University of Rome Tor Vergata

fismed@uniroma2.it

FISMED / PEOPLE

Tommaso Boccato

AI Research Scientist · Tether Evo

BACKGROUND

Biography

Tommaso Boccato is an AI Research Scientist at Tether Evo. His research concerns biologically inspired neural architectures and the decoding of visual and speech information from brain activity. He holds a bachelor's degree in Information Engineering and a master's degree in ICT for Internet and Multimedia from the University of Padova. His work with the Medical Physics, AI and Neurotechnology group has examined complex network topologies, neuromorphic computing and generative models for brain decoding.

RESEARCH & EXPERTISE

Research interests

  • Brain decoding
  • Speech brain–computer interfaces
  • Neuromorphic computing
  • Neural network topology

RESEARCH OUTPUT

Selected publications

  1. Cross-subject decoding of human neural data for speech brain computer interfaces.Journal of neural engineering · 2026
    Objective.Brain-to-text systems have recently achieved impressive performance when trained on single-participant data, but remain limited by uninvestigated cross-subject generalization.Approach.We present the first neural-to-phoneme decoder trained jointly on the two largest intracortical speech datasets (Willettet al2023Nature6201031-6; Cardet al2024New Engl. J. Med.391609-18), introducing day- and dataset-specific affine transforms to align neural activity into a shared space. Additionally, a hierarchical GRU decoder with intermediate CTC supervision and feedback connections is designed to address the conditional-independence assumption of standard CTC loss.Main results.Our model matches or outperforms within-subject baselines while being trained across participants, and adapts to unseen subjects using only a linear transform or brief fine-tuning. On an independent inner-speech dataset (Kunzet al2025Cell1884658-4673.e17), our approach shows some initial evidence of generalization, by training only subject-, day-specific transforms.Significance.These results demonstrate the feasibility of cross-subject pretraining as a promising direction toward more scalable speech Brain Computer Interfaces.
  2. Evidence for compositionality in fMRI visual representations via Brain Algebra.Communications biology · 2025
    Electrophysiological and neuroimaging studies have revealed how the brain encodes various visual categories and concepts. An open question is how combinations of multiple visual concepts are represented in terms of the component brain patterns: are brain responses to individual concepts composed according to algebraic rules? To explore this, we generated "conceptual perturbations" in neural space by averaging fMRI responses to images with a shared concept (e.g., "winter" or "summer"). After thresholding to ensure specificity, we applied these perturbations to the neural pattern associated with a base image, forming new brain patterns that incorporate the added concept. These modified brain patterns were then decoded into images using a pretrained fMRI-to-image decoding model. Qualitative and quantitative inspection of the resulting images provides insight into how the brain might combine visual concepts. For example, adding a "winter" perturbation to the brain pattern of a man on a skateboard yields a new pattern representing a man on a snowboard in a winter scene-even when the perturbation modifies only a small subset of voxels. Our findings reveal that compositional processes in neural representations may lead to predictable perceptual outcomes, as interpreted by our decoding model. This suggests that the brain's combinatory encoding of concepts may follow a systematic, algebraic-like process-what we term "brain algebra." Although our study is model-driven, it opens avenues for future empirical work into the mechanisms of compositionality in the brain.
  3. Through their eyes: Multi-subject brain decoding with simple alignment techniques.Imaging neuroscience (Cambridge, Mass.) · 2024
    To-date, brain decoding literature has focused on single-subject studies, that is, reconstructing stimuli presented to a subject under fMRI acquisition from the fMRI activity of the same subject. The objective of this study is to introduce a generalization technique that enables the decoding of a subject's brain based on fMRI activity of another subject, that is, cross-subject brain decoding. To this end, we also explore cross-subject data alignment techniques. Data alignment is the attempt to register different subjects in a common anatomical or functional space for further and more general analysis. We utilized the Natural Scenes Dataset, a comprehensive 7T fMRI experiment focused on vision of natural images. The dataset contains fMRI data from multiple subjects exposed to 9,841 images, where 982 images have been viewed by all subjects. Our method involved training a decoding model on one subject's data, aligning new data from other subjects to this space, and testing the decoding on the second subject based on information aligned to the first subject. We also compared different techniques for fMRI data alignment, specifically ridge regression, hyper alignment, and anatomical alignment. We found that cross-subject brain decoding is possible, even with a small subset of the dataset, specifically, using the common data, which are around 10 % of the total data, namely 982 images, with performances in decoding comparable to the ones achieved by single-subject decoding. Cross-subject decoding is still feasible using half or a quarter of this number of images with slightly lower performances. Ridge regression emerged as the best method for functional alignment in fine-grained information decoding, outperforming all other techniques. By aligning multiple subjects, we achieved high-quality brain decoding and a potential reduction in scan time by 90 % . This substantial decrease in scan time could open up unprecedented opportunities for more efficient experiment execution and further advancements in the field, which commonly requires prohibitive (20 hours) scan time per subject.
  4. Decoding visual brain representations from electroencephalography through knowledge distillation and latent diffusion models.Computers in biology and medicine · 2024
    Decoding visual representations from human brain activity has emerged as a thriving research domain, particularly in the context of brain-computer interfaces. Our study presents an innovative method that employs knowledge distillation to train an EEG classifier and reconstruct images from the ImageNet and THINGS-EEG 2 datasets using only electroencephalography (EEG) data from participants who have viewed the images themselves (i.e. "brain decoding"). We analyzed EEG recordings from 6 participants for the ImageNet dataset and 10 for the THINGS-EEG 2 dataset, exposed to images spanning unique semantic categories. These EEG readings were converted into spectrograms, which were then used to train a convolutional neural network (CNN), integrated with a knowledge distillation procedure based on a pre-trained Contrastive Language-Image Pre-Training (CLIP)-based image classification teacher network. This strategy allowed our model to attain a top-5 accuracy of 87%, significantly outperforming a standard CNN and various RNN-based benchmarks. Additionally, we incorporated an image reconstruction mechanism based on pre-trained latent diffusion models, which allowed us to generate an estimate of the images that had elicited EEG activity. Therefore, our architecture not only decodes images from neural activity but also offers a credible image reconstruction from EEG only, paving the way for, e.g., swift, individualized feedback experiments.
  5. Retrieving and reconstructing conceptually similar images from fMRI with latent diffusion models and a neuro-inspired brain decoding model.Journal of neural engineering · 2024
    Objective.Brain decoding is a field of computational neuroscience that aims to infer mental states or internal representations of perceptual inputs from measurable brain activity. This study proposes a novel approach to brain decoding that relies on semantic and contextual similarity.Approach.We use several functional magnetic resonance imaging (fMRI) datasets of natural images as stimuli and create a deep learning decoding pipeline inspired by the bottom-up and top-down processes in human vision. Our pipeline includes a linear brain-to-feature model that maps fMRI activity to semantic visual stimuli features. We assume that the brain projects visual information onto a space that is homeomorphic to the latent space of last layer of a pretrained neural network, which summarizes and highlights similarities and differences between concepts. These features are categorized in the latent space using a nearest-neighbor strategy, and the results are used to retrieve images or condition a generative latent diffusion model to create novel images.Main results.We demonstrate semantic classification and image retrieval on three different fMRI datasets: Generic Object Decoding (vision perception and imagination), BOLD5000, and NSD. In all cases, a simple mapping between fMRI and a deep semantic representation of the visual stimulus resulted in meaningful classification and retrieved or generated images. We assessed quality using quantitative metrics and a human evaluation experiment that reproduces the multiplicity of conscious and unconscious criteria that humans use to evaluate image similarity. Our method achieved correct evaluation in over 80% of the test set.Significance.Our study proposes a novel approach to brain decoding that relies on semantic and contextual similarity. The results demonstrate that measurable neural correlates can be linearly mapped onto the latent space of a neural network to synthesize images that match the original content. These findings have implications for both cognitive neuroscience and artificial intelligence.

LABORATORIES & RESEARCH

Related research

EDUCATION & MENTORING

Teaching archive

Back to researchers and staff