Doctoral Researcher · University of Rome Tor Vergata
BACKGROUND
Biography
Matteo Ciferri is a doctoral researcher at the University of Rome Tor Vergata. He holds a master's degree in Management Engineering from Sapienza University of Rome, with training in data science and optimisation. His research develops models linking brain activity to representations of sound, language and vision. He works with functional MRI, MEG and intracortical recordings, studying how training objectives, temporal dynamics and model complexity affect neural encoding and the retrieval or reconstruction of sensory stimuli.
Characterizing the information content of intracortical signals during visual processing is a central challenge in systems neuroscience. We address the problem of decoding visual information from high-density intracortical recordings in primates, using the THINGS Ventral Stream Spiking Dataset. We systematically evaluate the effects of model architecture, training objectives, and data scaling on decoding performance. Results show that decoding accuracy is jointly driven by non-linearity and selective temporal aggregation, rather than heavier sequence modelling in this data regime. A simple model combining temporal attention with a shallow MLP achieves up to 70% top-1 image retrieval accuracy, outperforming linear baselines as well as recurrent and convolutional approaches. Scaling analyses reveal predictable diminishing returns with increasing input dimensionality and dataset size. Building on these findings, we design a modular generative decoding pipeline that combines low-resolution latent reconstruction with semantically conditioned diffusion, generating plausible images from 200 ms of brain activity. This framework provides principles for brain-computer interfaces and semantic neural decoding.
Music is a universal phenomenon that influences human experiences across cultures. We investigate whether music can be decoded from human brain activity measured with fMRI, by modeling mappings between neural data and latent representations of musical stimuli. Our approach integrates functional and anatomical alignment techniques to facilitate cross-subject decoding. Starting from the GTZan fMRI dataset, where five participants listened to 540 musical tracks from 10 genres, we used the CLAP model to extract latent representations of the musical stimuli and developed voxel-wise encoding models to identify brain regions responsive to these stimuli, by applying a threshold to the correlation between predicted and actual brain activity. Our decoding pipeline, primarily retrieval-based, employs a linear map to project back brain activity to the corresponding CLAP features. This enables us to retrieve the musical stimuli most similar to those that originated the fMRI data. Our results demonstrate state-of-the-art identification accuracy, outperforming existing approaches.
Brain decoding aims to reconstruct external stimuli from brain activity, providing insights into the neural representation of cognitive experiences. Music decoding from functional magnetic resonance imaging (fMRI) is particularly challenging due to the complexity of auditory processing and the temporal limitations of fMRI signals. In this study, we introduce a novel decoding framework that improves the alignment between fMRI activity and latent musical representations extracted using a pre-trained multimodal model (CLAP). We propose a dual-loss approach combining Optimal Transport and Contrastive Learning to enhance feature mapping and retrieval accuracy. The first loss ensures structural consistency between brain-predicted and true musical embeddings, while the contrastive loss refines the embedding space by maximizing similarities between corresponding pairs and minimizing non-correspondences. Using fMRI data from five subjects listening to music tracks from the GTZAN dataset, our method achieves improved decoding performance, surpassing traditional regression-based approaches from 22.1% top-1 accuracy to 29.3%. These results highlight the potential of integrating Optimal Transport and Contrastive Learning to improve brain decoding performance, paving the way for extending the approach to different sensory domains and applications in Brain-Computer Interfaces (BCI).Clinical relevance- This study could have clinical implications for understanding auditory processing disorders and developing neurorehabilitation strategies. By elucidating how the brain encodes complex auditory stimuli, this approach may contribute to BCI applications for speech and music perception restoration in individuals with hearing impairments or neurological conditions affecting auditory cognition.
Reconstructing music directly from brain activity provides insight into the neural representations underlying auditory processing and paves the way for future brain-computer interfaces. We introduce a fully data-driven pipeline that combines cross-subject functional alignment with bayesian decoding in the latent space of a diffusion-based audio generator. Functional alignment projects individual fMRI responses onto a shared representational manifold, increasing the performance of cross-participant accuracy with respect to anatomically normalized baselines. A bayesian search over latent trajectories then selects the most plausible waveform candidate, stabilizing reconstructions against neural noise. Crucially, we bridge CLAP's multi-modal embeddings to music-domain latents through a dedicated aligner, eliminating the need for hand-crafted captions and preserving the intrinsic structure of musical features. Evaluated on ten diverse genres, the model achieves a cross-subject-averaged identification accuracy of [Formula: see text], and produces audio that human listeners recognize above chance in 85.7% of trials. Voxel-wise analyses locate the predictive signal within a bilateral circuit spanning early auditory, inferior-frontal, and premotor cortices, consistent with hierarchical and sensorimotor theories of music perception. The framework establishes a principled bridge between generative audio models and cognitive neuroscience.