7 ms·
fMRI-to-image with contrastive learning and diffusion priors
- woeirua 3y agoTechniques like this give me hope that some day we will be able to objectively diagnose mental illness, and monitor the efficacy of treatment.
- robg 3y agoAwesome, predicting words from fMRI has been around for a while and visual cortex can be mapped well. That said, and coming from a background in neuroimaging 20 years ago, what’s the applicability? MRI hasn’t gotten that much more cost effective for more widespread uses. Magnets are expensive.
- aledalgrande 3y agoPeople with disabilities could benefit greatly from this.
- ggm 3y agoAs long as they want to talk about london buses, steam trains, surfing and football.
- deleted 3y ago[deleted]
- 01100011 3y agoYeah, the first thing that comes to mind(har har) when I see this is that we'd be better off trying to develop better scanning technology. You can't exactly walk around town with an MRI strapped to your skull.
- seydor 3y agoreading suspect's mind
- tmabraham 3y agoWe think it could be useful for clinical research and maybe even diagnostics. For example, you could imagine a person with depression(or other neurological disorders) may have a different perception of the same image than a healthy person. Now with the much higher fidelity that both more powerful MRI machines and better generative AI tools can provide, this may now be a very promising direction for future research.
- 2ap 3y agoI work in pediatrics and am an academic investigating MRI of kids in various diseases. When I saw this work, I did wonder about us being better able to functionally map where things are going wrong in the pathways of neurodisability. I wondered if this would have applications in being able to do that - for example being able to say that someone could process the image. Do you think it could have this type of application? One thing which would be a deal breaker at the moment is the amount of time participants spend in the scanner. But if we wanted to (for example) see if a child could perceive simple objects, would that be doable do you think?
- hospadar 3y agoThis is SO COOL. I'd guess (I did analysis for an fMRI lab for a year so I'm not a pro but not totally talking out of my orifice) that detecting images like this is among the easier things you could do (it probably wouldn't be so easy to do things like "guess the words I'm thinking of") and I suspect other sensory stuff might be harder but I have little knowledge there. One of the biggest issues with any attempt to extract information from an fMRI scan is resolution, both spatial and temporal - this study used 1.8mm voxels which is a TON of neurons (also recall that fMRIs scan blood flow, not neuron activity - we just count on those things being correlated). Temporally, fMRI sample frequency are often <1hz. I didn't see that they mentioned a specific frequency, but they showed images to the subject for 3 seconds at a time so I'd guess that's designed to ensure you get a least a frame or three while the subject is looking at the image. You can sort of trade voxel size for sample frequency - so you can get more voxels, or more samples, but not both. So detecting things that happen quickly (like, say, moving images or speech) would probably be quite hard (even if you could design an ai thingey that could do it, getting the raw data at the resolution you'd need is not currently possible with existing scanners) Also, not all brain functions are as clearly localized as vision - the visual cortex areas in the back of the brain map pretty directly to certain kinds of visual stimulus, while other kinds of stimulus and activity are much less localize (there isn't a clear "lighting up" of an area). You can get better resolution if you only scan part of the brain (i.e. the visual cortex) (I don't know if that's what they did for this study), but that's obviously only useful for activity happening in a small part of the brain. ANYWAY SO COOL!!! I wonder if you could use this to draw people's faces with a subject who is imagining looking at a face? fMRI police sketch? How do brains even work!?
- atom_101 3y agoWe did have a face reconstruction project planned. It is on the back-burner for now. That one will be based on something like the Celeb-A dataset instead of the Natural Scenes Dataset (images from MS-COCO) used here.
- tmabraham 3y agoYeah using data from a 7T MRI giving higher spatial resolution definitely helps! The fMRI dataset includes signal from the whole brain but we only use the data from the visual cortex for this study.
- ilaksh 3y agoHuman communication will change dramatically once useful invasive brain-computer interfaces are available. People will suddenly realize that the reason language is primarily serial is simply due to the fact that it must be conveyed by a series of sounds. There will likely be a new type of visual language used via BCI "telepathy". It may have some ordering but will not rely so heavily on serializing information, since the world is quite multidimensional.
- SamPatt 3y agoIn a popular sci-fi (avoiding spoliers), the alien race has transparent skulls, and their visible thoughts are broadcast to anyone within visual range. It does seem more efficient than sound.
- flangola7 3y agoI want to know the scifi You may base64 to spoiler proof it
- lithiumii 3y agoI think it's "VGhlIFRocmVlLUJvZHkgUHJvYmxlbQ==" The aliens cannot lie to each other (they don't even have the idea of a lie), because their thoughts are transparent to each other.
- esafak 3y agoSounds like a recipe for conflict.
- inglor_cz 3y agoIt would certainly be a very violent shift towards a very different societal equilibrium. Few, if any, people who currently have power, would be able to absorb the sheer amount of hitherto hidden distrust or resentment that their subordinates harbor towards them. Interestingly, there might be two very different end stages. Either a very open society where people at the top are selected to be non-narcissist and stoic, or a very closed and oppressive society where the absolute ruler is kept in power by a bunch of truly zombified and obedient warriors whose loyalty is real and unshakeable, and who will kill anyone whose brain entertains any rebel ideas too much.
- ImaCake 3y agoAside from helping those with disabilities, what is stopping authorities using this as a lie detector? I assume the tech isn’t quite there yet.
- FPGAhacker 3y agoNot yet. > Models were trained separately for every participant and are not generalizable across people.
- ImaCake 3y agoAh right, so the models are subject to some serious overfitting then. Good proof of concept, but not useful in practice yet.
- collsni 3y agoThis is mind reading, we are in the future.
- QuantumG 3y agoFound the sucker.
- deleted 3y ago[deleted]
- theaiquestion 3y agoI think the method of merging the pipelines via img2img should use controlnet. Possibly needing to be finetuned specifically for this, although existing controlnet models might work fine for this. This is exactly what you'd want to use controlnet for - mapping semantic information onto the perceived structure.
- tmabraham 3y agoYes we've been looking into ControlNet as well, and I think there is one recent fMRI-to-image paper that also has tried ControlNet. Maybe we'll use ControlNet in MindEye v2 :)
- atom_101 3y agoYes controlnet will be used in the next version. For this one we couldn't get it working in time.
- satvikpendem 3y agoWasn't there something similar a few months ago on HN and where the top comment talked about how it's not as impressive as it sounds [0]? The main issue is that this type of methodology is pulling from a pool of images, not literally reconstructing what image was seen in the brain directly. > I immediately found the results suspect, and think I have found what is actually going on. The dataset it was trained on was 2770 images, minus 982 of those used for validation. I posit that the system did not actually read any pictures from the brains, but simply overfitted all the training images into the network itself. For example, if one looks at a picture of a teddy bear, you'd get an overfitted picture of another teddy bear from the training dataset instead. > The best evidence for this is a picture(1) from page 6 of the paper. Look at the second row. The building generated by 'mind reading' subject 2 and 4 look strikingly similar, but not very similar to the ground truth! From manually combing through the training dataset, I found a picture of a building that does look like that, and by scaling it down and cropping it exactly in the middle, it overlays rather closely(2) on the output that was ostensibly generated for an unrelated image. > If so, at most they found that looking at similar subjects light up similar regions of the brain, putting Stable Diffusion on top of it serves no purpose. At worst it's entirely cherry-picked coincidences. > 1. https://i.imgur.com/ILCD2Mu.png https://i.imgur.com/ILCD2Mu.png > 2. https://i.imgur.com/ftMlGq8.png https://i.imgur.com/ftMlGq8.png [0] https://news.ycombinator.com/item?id=35012981 https://news.ycombinator.com/item?id=35012981
- jph00 3y agoYes there was. However this is a different paper, describing a different method, applied to a different dataset, with different results. As the abstract says, "In particular, MindEye can retrieve the exact original image even among highly similar candidates indicating that its brain embeddings retain fine-grained image-specific information. This allows us to accurately retrieve images even from large-scale databases like LAION-5B. We demonstrate through ablations that MindEye's performance improvements over previous methods result from specialized submodules for retrieval and reconstruction, improved training techniques, and training models with orders of magnitude more parameters." Note that LAION-5B has five billion images.
- 3y ago
- mometsi 3y agoFor context, early vision is easier to map than you might expect. Here's a radiograph of the primary visual cortex created in 1982 by projecting a pattern onto a macaque's retina: https://web.archive.org/web/20100814085656im_/http://hubel.med.harvard.edu/book/114.jpg https://web.archive.org/web/20100814085656im_/http://hubel.m... An injection of radioactive sugar lets you see where the neurons were firing away and metabolizing the sugar. (https://pubmed.ncbi.nlm.nih.gov/7134981/ https://pubmed.ncbi.nlm.nih.gov/7134981/)
- dr_dshiv 3y agoBut can brain activity be mapped anywhere like this with fmri? I doubt it. But yes — it is cool that the brain keeps spatial proportions of reality in the brain map! Very unlike latent space.
- echo_time 3y agoIn some ways, absolutely not - precision is a huge challenge with an indirect method like fMRI - but this example is over a decade old now: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3130346/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3130346/ Fig4 shows the letter M on the cortical surface, where the stimulus accounted for the effects of foveal magnification (foveal vision gets more cortical space). Keep in mind that we now, in theory, have stronger magnets, better head coils (the part that picks up the image information), and better sequences (the software that manipulates the magnets to produce the images) so we could do even better than that these days.
- caycep 3y agothe best recons I've seen so far are from Jack Gallant and Alex Huth's labs (at least what was shown publicly at SfN)
- parth17291 3y agoThis is cool . We are going towards future like shown in inception to plant an idea.
- umvi 3y agoWhat if a copyrighted image or video can be recovered from your brain using external tech like this? That's not fair to the rights holders, what we really need is technology that can clean such illegal memories from the brain; a brainwasher if you will
- puchatek 3y agoAs someone suffering from intrusive thoughts I do not look forward to a future where other people can see what I sometimes see in my head.
- wiz21c 3y agoAs a person who is normal by any measuring standard, I do not look forward to a future where other people can see what I sometimes see in my head. You'd be quite surprised.
- radicaldreamer 3y agoThey should dose people with DMT in the fMRI and run it through the model.
- deleted 3y ago[deleted]
- mikeiz404 3y agoSome important points from the article under the limitations section - Each participant in the dataset spent up to 40 hours in the MRI machine to gather sufficient training data. - Models were trained separately for every participant and are not generalizable across people. Image limitations: MindEye is limited to the kinds of natural scenes used for training the model. For other image distributions, additional data collection and specialized generative models would be needed. - … Non-invasive neuroimaging methods like fMRI not only require participant compliance but also full concentration on following instructions during the lengthy scan process. …
- bertil 3y agoFor anyone who hasn’t been in an MRI: 40 hours is a lot. Those things are tight; not “look where you are going” tight, but you absolutely need to tell people beforehand that they will feel very uncomfortable inside, remind them that they can get out at any time, and show them how because they will not like being in there. I wouldn’t spend 20 minutes in one if it were not important. I’d seriously push back on an hour. 40 hours is something I’d only do if that is absolutely necessary.
- smusamashah 3y agoThere was also DreamDiffusion recently. https://arxiv.org/abs/2306.16934 https://arxiv.org/abs/2306.16934