Medical Physics, AI and Neurotechnology · University of Rome Tor Vergata

fismed@uniroma2.it

FISMED / PEOPLE

Matteo Ferrante

Postdoctoral Researcher · Lead AI Research Scientist · University of Rome Tor Vergata · Tether Evo · Silenzio

BACKGROUND

Biography

Matteo Ferrante is a postdoctoral researcher at the University of Rome Tor Vergata, Lead AI Research Scientist at Tether Evo and CEO of Silenzio. His research focuses on NeuroAI and brain–computer interfaces, particularly the relationship between neural activity and representations learned by artificial neural networks. He develops methods for decoding visual, linguistic and other information from brain recordings, including generative and contrastive learning approaches.

He trained in physics and biomedical physics at the University of Pavia and undertook doctoral research in the National PhD in Artificial Intelligence, Health and Life Sciences. He continues to collaborate with Tor Vergata researchers on brain decoding, neural representations and computational methods for biomedical imaging.

RESEARCH & EXPERTISE

Research interests

  • NeuroAI
  • Brain–computer interfaces
  • Generative brain decoding
  • Representation learning
  • Biomedical machine learning

RESEARCH OUTPUT

Selected publications

  1. NeuroFusion: A Unified Framework for Generalized Visual Stimulus Decoding from fMRI Across Datasets and Subjects.Neuroinformatics · 2026
    Recent advancements in neural decoding have shown promising results in reconstructing visual experiences from brain activity. However, existing approaches focus primarily on decoding within a single dataset or subject, which limits generalization across various sources of neuroimaging. In this work, we propose a novel framework for the decoding of visual stimuli between subjects and between data sets, integrating neural recordings from multiple publicly available fMRI datasets. To address inherent intersubject and interdataset variability, we introduce a contrastive learning-based alignment strategy using image embeddings from a pre-trained IP-Adapter model. Our approach learns a shared latent space by aligning subject-specific neural representations with image features, enabling generalized decoding across both subjects and datasets. In addition, we propose a simple yet effective data augmentation method using ridge regression. This method synthesizes realistic fMRI-like signals from novel images by predicting voxel activity and injecting learned noise distributions, thus enhancing training diversity and model robustness. To the best of our knowledge, while several recent studies have explored cross-subject decoding, we extend recent cross-subject decoding efforts by training a single unified framework jointly across multiple public fMRI datasets and subjects, enabling cross-dataset transfer in addition to cross-subject generalization. We distinguish this multi-dataset unified training setting, where each dataset contributes training data, from a stricter leave-one-dataset-out transfer setting in which the target dataset is excluded from source pretraining and used only for lightweight alignment-layer adaptation. Empirically, our unified model achieves strong semantic reconstruction across datasets (e.g., up to 94.8% CLIP similarity on NSD (AUG) and 0.403 SSIM on BOLD5000 after lightweight finetuning), demonstrating robust cross-subject and cross-dataset transfer.
  2. A modular semantic-structural pipeline for visual decoding from primate spiking data via selective temporal integration.Imaging neuroscience (Cambridge, Mass.) · 2026
    Characterizing the information content of intracortical signals during visual processing is a central challenge in systems neuroscience. We address the problem of decoding visual information from high-density intracortical recordings in primates, using the THINGS Ventral Stream Spiking Dataset. We systematically evaluate the effects of model architecture, training objectives, and data scaling on decoding performance. Results show that decoding accuracy is jointly driven by non-linearity and selective temporal aggregation, rather than heavier sequence modelling in this data regime. A simple model combining temporal attention with a shallow MLP achieves up to 70% top-1 image retrieval accuracy, outperforming linear baselines as well as recurrent and convolutional approaches. Scaling analyses reveal predictable diminishing returns with increasing input dimensionality and dataset size. Building on these findings, we design a modular generative decoding pipeline that combines low-resolution latent reconstruction with semantically conditioned diffusion, generating plausible images from 200 ms of brain activity. This framework provides principles for brain-computer interfaces and semantic neural decoding.
  3. Cross-subject decoding of human neural data for speech brain computer interfaces.Journal of neural engineering · 2026
    Objective.Brain-to-text systems have recently achieved impressive performance when trained on single-participant data, but remain limited by uninvestigated cross-subject generalization.Approach.We present the first neural-to-phoneme decoder trained jointly on the two largest intracortical speech datasets (Willettet al2023Nature6201031-6; Cardet al2024New Engl. J. Med.391609-18), introducing day- and dataset-specific affine transforms to align neural activity into a shared space. Additionally, a hierarchical GRU decoder with intermediate CTC supervision and feedback connections is designed to address the conditional-independence assumption of standard CTC loss.Main results.Our model matches or outperforms within-subject baselines while being trained across participants, and adapts to unseen subjects using only a linear transform or brief fine-tuning. On an independent inner-speech dataset (Kunzet al2025Cell1884658-4673.e17), our approach shows some initial evidence of generalization, by training only subject-, day-specific transforms.Significance.These results demonstrate the feasibility of cross-subject pretraining as a promising direction toward more scalable speech Brain Computer Interfaces.
  4. R&B - rhythm and brain: Cross-subject decoding of music from human brain activity.Neural networks : the official journal of the International Neural Network Society · 2026
    Music is a universal phenomenon that influences human experiences across cultures. We investigate whether music can be decoded from human brain activity measured with fMRI, by modeling mappings between neural data and latent representations of musical stimuli. Our approach integrates functional and anatomical alignment techniques to facilitate cross-subject decoding. Starting from the GTZan fMRI dataset, where five participants listened to 540 musical tracks from 10 genres, we used the CLAP model to extract latent representations of the musical stimuli and developed voxel-wise encoding models to identify brain regions responsive to these stimuli, by applying a threshold to the correlation between predicted and actual brain activity. Our decoding pipeline, primarily retrieval-based, employs a linear map to project back brain activity to the corresponding CLAP features. This enables us to retrieve the musical stimuli most similar to those that originated the fMRI data. Our results demonstrate state-of-the-art identification accuracy, outperforming existing approaches.

LABORATORIES & RESEARCH

Related research

EDUCATION & MENTORING

Teaching archive

Back to researchers and staff