Medical Physics, AI and Neurotechnology · University of Rome Tor Vergata

fismed@uniroma2.it

FISMED / PEOPLE

Matteo Ciferri

Doctoral Researcher · University of Rome Tor Vergata

BACKGROUND

Biography

Matteo Ciferri is a doctoral researcher at the University of Rome Tor Vergata. He holds a master's degree in Management Engineering from Sapienza University of Rome, with training in data science and optimisation. His research develops models linking brain activity to representations of sound, language and vision. He works with functional MRI, MEG and intracortical recordings, studying how training objectives, temporal dynamics and model complexity affect neural encoding and the retrieval or reconstruction of sensory stimuli.

RESEARCH & EXPERTISE

Research interests

  • Multimodal brain decoding
  • Neural encoding
  • Auditory and language processing
  • Representation learning

RESEARCH OUTPUT

Selected publications

  1. A modular semantic-structural pipeline for visual decoding from primate spiking data via selective temporal integration.Imaging neuroscience (Cambridge, Mass.) · 2026
    Characterizing the information content of intracortical signals during visual processing is a central challenge in systems neuroscience. We address the problem of decoding visual information from high-density intracortical recordings in primates, using the THINGS Ventral Stream Spiking Dataset. We systematically evaluate the effects of model architecture, training objectives, and data scaling on decoding performance. Results show that decoding accuracy is jointly driven by non-linearity and selective temporal aggregation, rather than heavier sequence modelling in this data regime. A simple model combining temporal attention with a shallow MLP achieves up to 70% top-1 image retrieval accuracy, outperforming linear baselines as well as recurrent and convolutional approaches. Scaling analyses reveal predictable diminishing returns with increasing input dimensionality and dataset size. Building on these findings, we design a modular generative decoding pipeline that combines low-resolution latent reconstruction with semantically conditioned diffusion, generating plausible images from 200 ms of brain activity. This framework provides principles for brain-computer interfaces and semantic neural decoding.
  2. R&B - rhythm and brain: Cross-subject decoding of music from human brain activity.Neural networks : the official journal of the International Neural Network Society · 2026
    Music is a universal phenomenon that influences human experiences across cultures. We investigate whether music can be decoded from human brain activity measured with fMRI, by modeling mappings between neural data and latent representations of musical stimuli. Our approach integrates functional and anatomical alignment techniques to facilitate cross-subject decoding. Starting from the GTZan fMRI dataset, where five participants listened to 540 musical tracks from 10 genres, we used the CLAP model to extract latent representations of the musical stimuli and developed voxel-wise encoding models to identify brain regions responsive to these stimuli, by applying a threshold to the correlation between predicted and actual brain activity. Our decoding pipeline, primarily retrieval-based, employs a linear map to project back brain activity to the corresponding CLAP features. This enables us to retrieve the musical stimuli most similar to those that originated the fMRI data. Our results demonstrate state-of-the-art identification accuracy, outperforming existing approaches.
  3. Optimal Transport and Contrastive Learning for Brain Decoding of Musical Perception.Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference · 2025
    Brain decoding aims to reconstruct external stimuli from brain activity, providing insights into the neural representation of cognitive experiences. Music decoding from functional magnetic resonance imaging (fMRI) is particularly challenging due to the complexity of auditory processing and the temporal limitations of fMRI signals. In this study, we introduce a novel decoding framework that improves the alignment between fMRI activity and latent musical representations extracted using a pre-trained multimodal model (CLAP). We propose a dual-loss approach combining Optimal Transport and Contrastive Learning to enhance feature mapping and retrieval accuracy. The first loss ensures structural consistency between brain-predicted and true musical embeddings, while the contrastive loss refines the embedding space by maximizing similarities between corresponding pairs and minimizing non-correspondences. Using fMRI data from five subjects listening to music tracks from the GTZAN dataset, our method achieves improved decoding performance, surpassing traditional regression-based approaches from 22.1% top-1 accuracy to 29.3%. These results highlight the potential of integrating Optimal Transport and Contrastive Learning to improve brain decoding performance, paving the way for extending the approach to different sensory domains and applications in Brain-Computer Interfaces (BCI).Clinical relevance- This study could have clinical implications for understanding auditory processing disorders and developing neurorehabilitation strategies. By elucidating how the brain encodes complex auditory stimuli, this approach may contribute to BCI applications for speech and music perception restoration in individuals with hearing impairments or neurological conditions affecting auditory cognition.
  4. Reconstructing music perception from brain activity using a prior guided diffusion model.Scientific reports · 2025
    Reconstructing music directly from brain activity provides insight into the neural representations underlying auditory processing and paves the way for future brain-computer interfaces. We introduce a fully data-driven pipeline that combines cross-subject functional alignment with bayesian decoding in the latent space of a diffusion-based audio generator. Functional alignment projects individual fMRI responses onto a shared representational manifold, increasing the performance of cross-participant accuracy with respect to anatomically normalized baselines. A bayesian search over latent trajectories then selects the most plausible waveform candidate, stabilizing reconstructions against neural noise. Crucially, we bridge CLAP's multi-modal embeddings to music-domain latents through a dedicated aligner, eliminating the need for hand-crafted captions and preserving the intrinsic structure of musical features. Evaluated on ten diverse genres, the model achieves a cross-subject-averaged identification accuracy of [Formula: see text], and produces audio that human listeners recognize above chance in 85.7% of trials. Voxel-wise analyses locate the predictive signal within a bilateral circuit spanning early auditory, inferior-frontal, and premotor cortices, consistent with hierarchical and sensorimotor theories of music perception. The framework establishes a principled bridge between generative audio models and cognitive neuroscience.

LABORATORIES & RESEARCH

Related research

Back to researchers and staff