Researchers at Princeton University, supervised by Professor Ken Norman, have demonstrated an AI system capable of reconstructing images from fMRI brain activity within seconds of a participant viewing them. This development represents a notable advancement in brain decoding, moving from methods that took hours or days to near real-time reconstruction. The study, titled "Real-time Reconstruction of Human Visual Perception from fMRI," shows that the system can also predict brain activity based on a viewed image.
The experiment involved participants lying in a 3-Tesla MRI scanner and viewing hundreds of natural images, such as animals, people, and outdoor scenes. Each image was displayed for a few seconds while the scanner measured activity across the brain's visual regions. The AI system first underwent an hour of training data collection for each participant, during which it learned how that individual's brain responded to different images. After this training phase, the model could reconstruct what the participant was seeing in a subsequent scanning session.
The AI system does not directly map brain activity to pixels. Instead, it translates brain signals into a high-level representation of the image, essentially a semantic description of its visual features. A generative AI then converts this representation back into an image. The reconstructed images are not photographic duplicates of the originals but rather AI-generated approximations. For instance, if a participant viewed a skier in a red jacket on a snowy mountain, the reconstruction might show a person-shaped figure in red surrounded by snowy terrain. These reconstructions, while fuzzy, maintain the broad structure and meaning of the original scene.
The primary breakthrough of this research lies in its speed. Previous methods for reconstructing images from fMRI data could take significantly longer. The Princeton system can produce image reconstructions approximately 15 seconds after an image is viewed. This delay is primarily due to the biological nature of the blood-oxygen signal measured by fMRI, which peaks several seconds after a stimulus and can take up to 20 seconds to dissipate. The researchers addressed computational challenges by combining a cloud-based real-time fMRI framework with a streamlined version of the MindEye2 architecture.
The fMRI signal itself presents difficulties because it is indirect and noisy, measuring blood flow changes rather than direct neuronal activity. The AI models used for reconstruction are complex, containing hundreds of millions of parameters, requiring substantial computing resources for rapid operation. The research was funded by Princeton's Office of the Dean for Research Innovation Fund for New Industrial Collaborations, in partnership with Sophont. Earlier stages of the work received support from the National Institute of Mental Health.
