6 ms·
Biomni: A General-Purpose Biomedical AI Agent
- freedomben 1y agoAwesome! This is the type of stuff I'm most excited about with AI - improvements to medical research and capabilities. AI can be awesome at identifying patterns in data that humans can't, and there has to be troves of data out there full of patterns that we aren't catching. Of course there's also the possibility of engineering new drugs/treatments and things, which is also super exciting.
- panabee 1y agoAgreed. There is deep potential for ML in healthcare. We need more contributors advancing research in this space. One opportunity as people look around: many priors merit reconsideration. For instance, genomic data that may seem identical may not actually be identical. In classic biological representations (FASTA), canonical cytosine and methylated cytosine are both collapsed into the letter "C" even though differences may spur differential gene expression. What's the optimal tokenization algorithm and architecture for genomic models? How about protein binding prediction? Unclear! There are so many open questions in biomedical ML. The openness-impact ratio is arguably as high in biomedicine as anywhere else: if you help answer some of these questions, you could save lives. Hopefully, awesome frameworks like this lower barriers and attract more people.
- govideo 1y agoI'd love to hear more of our thoughts re open questions in biomedical ML. You sound like you have a crisp, nuanced grasp the landscape, which is rare. That would be very helpful to me, as an undergrad in CS (with bio) trying to crystalize research to pursue in bio/ML/GenAI. Thank you.
- panabee 1y agoThanks, but no one truly understands biomedicine, let alone biomedical ML. Feynman's quote -- "A scientist is never certain" -- is apt for biomedical ML. Context: imagine the human body as the most devilish operating system ever: 10b+ lines of code (more than merely genomics), tight coupling everywhere, zero comments. Oh, and one faulty line may cause death. Are you more interested in data, ML, or biology (e.g., predicting cancerous mutations or drug toxicology)? Biomedical data underlies everything and may be the easiest starting point because it's so bad/limited. We had to pay Stanford doctors to annotate QA questions because existing datasets were so unreliable. (MCQ dataset partially released, full release coming). For ML, MedGemma from Google DeepMind is open and at the frontier. Biology mostly requires publishing, but still there are ways to help. After sharing preferences, I can offer a more targeted path.
- govideo 1y agoML first, then Bio and Data. Of course, interconnectedness runs high (eg just read about ML for non-random missingness in med records) and that data is the foundational bottleneck/need across the board. Interesting anecdote abt Stanford doctors annotating QA question! Each of your comments get my mind going... I'm going to think about them more and may ping you on other channels, per your profile. Thanks!
- panabee 1y agoMore like alarming anecdote. :) Google did a wonderful job relabeling MedQA, a core benchmark, but even they missed some (e.g., question 448 in the test set remains wrong according to Stanford doctors). For ML, start with MedGemma. It's a great family. 4B is tiny and easy to experiment with. Pick an area and try finetuning. Note the new image encoder, MedSigLIP, which leverages another cool Google model, SigLIP. It's unclear if MedSigLIP is the right approach (open question!), but it's innovative and worth studying for newcomers. Follow Lucas Beyer, SigLIP's senior author and now at Meta. He'll drop tons of computer vision knowledge (and entertaining takes). For bio, read 10 papers in a domain of passion (e.g., lung cancer). If you (or AI) can't find one biased/outdated assumption or method, I'll gift a $20 Starbucks gift card. (Ping on Twitter.) This matters because data is downstream of study design, and of course models are downstream of data. Starbucks offer open to up to three people.
- AIorNot 1y agovery cool -passed on to my friend who is working a Crispr lab
- Edmond 1y agoThis is nice, a lot of possibilities regarding AI use for scientific research. There is also the possibility of building intelligent workspaces that could prove useful in aiding scientific research: https://news.ycombinator.com/item?id=44509078 https://news.ycombinator.com/item?id=44509078
- VerminOctopus1 1y agoNot to take away from this or its usefulness (not my intent), but it is wild to me how many pieces of software of this type are being developed. We’re seeing endless waves of specialized wrappers around LLM API calls. There’s very little innovation happening beyond specializing around particular niches and invoking LLMs in slightly different ways with carefully directed context and prompts.
- gronky_ 1y agoI see it a bit differently - LLMs are an incredible innovation but it’s hard to do anything useful with them without the right wrapper. A good wrapper has deep domain knowledge baked into it, combined with automation and expert use of the LLM. It maybe isn’t super innovative but it’s a bit of an art form and unlocks the utility of the underlying LLM
- tedy1996 1y agoHow is something that cant admit it doesnt know, and hallucinates a good innovation?
- knowaveragejoe 1y agoModern LLMs frequently do state that they "don't know", for what it's worth. Like everything, it highly depends on the question.
- mrlongroots 1y agoExactly. To present a potential usecase: there's a ridiculous and massive backlog in the Indian judicial system. LLMs can be let loose on the entire workflow: triage cases (simple, complicated, intractable, grouped by legal principles or parties), pull up related caselaw, provide recommendations, throw more LLMs and more reasoning at unclear problems. Now you can't do this with just a desktop and chatgpt, you need a systemic pipeline of LLM-driven workflows, but doing that unlocks potentially billions of dollars of value that is otherwise elusive.
- andy99 1y agoI'm sure they've thought of this but curious how it fared on evaluations for supporting biological threats, ie elevating threat actor capabilities with respect to making biological weapons. I'm personally sceptical that LLMs can currently do this (and it's based on Claude that does test this) but still interesting to see.
- greazy 1y agoCreating a biological weapon requires a whole bunch of unique and specialised skills, equipment, safety measures (so you don't infect/kill yourself/your people) and even multidisciplinary skill sets. Take for example the Kameido (Japan) incident by the Aum Shinrikyo cult/religious group [1]. Same group which committed the Sarin attack [2]. > The use of an attenuated B. anthracis strain, low spore concentrations, ineffective dispersal, a clogged spray device, and inactivation of the spores by sunlight are all likely contributing factors to the lack of human cases. Now you may say, that's bacteria, what about viruses? A similar set of problems would arise, how do you successfully grow virus to high titers? Even vaccine companies struggle to do this with certain viruses. Then the issue of dispersal, infectivity and mortality arise (too quick, it kills the host without spreading and authorities will notice, too slow, same problem: authorities will notice). I haven't even mentioned biological engineering which requires years of technical knowledge and laboratory experience combined with a intimate knowledge of the organism you're working with. What worries me the most is nature springing a new influenza subtype. Our farming practices, especially in developing countries, is bound to breed a new subtype. It happened in 2009 (H1N1pdm) and it is bound to happen again. We got lucky with H1N1pdm. 1. https://pmc.ncbi.nlm.nih.gov/articles/PMC3322761/ https://pmc.ncbi.nlm.nih.gov/articles/PMC3322761/ 2. https://en.wikipedia.org/wiki/Tokyo_subway_sarin_attack https://en.wikipedia.org/wiki/Tokyo_subway_sarin_attack
- spwa4 1y ago> Creating a biological weapon requires a whole bunch of unique and specialised skills, equipment, safety measures I just tell some investors our god tells me to do that. > too quick, it kills the host without spreading and authorities will notice, too slow, same problem: authorities will notice The current authorities are Trump's authorities and don't believe in vaccines, have said in an interview Covid is either a Jewish or Chinese conspiracy (and "they made themselves immune"), and that the "disease epidemic" needs to end (this last one is easy to misunderstand. Kennedy Jr. doesn't believe any particular disease is an epidemic. Epidemics of that sort don't exist according to him. That people believe they get sick, THAT is the epidemic that must be stopped)
- deepdarkforest 1y agoInteresting. It's just an agent loop with access to python exec and web search as standard, BUT with premade, curated, 150 tools like analyze_circular_dichroism_spectra, with very specific params that just execute a hardcoded python function. Also with easy to load databases that conform to the tools' standards. The argument is that if you just ask claude code to do niche biomed tasks, it will not have the knowledge to do it like that by just searching pubmed and doing RAG on the fly, which is fair, given the current gen of LLM's. It's an interesting approach, they show some generalization on the paper(with well known tidy datasets), but real life data is messier, and the approach here(correct me if im wrong) is to identify the correct tool for a task, and then use the generic python exec tool to shape the data into the acceptable format if needed, try the tool and go again. It would be useful to use the tools just as a guidance to inform a generic code agent imo, but executing the "verified" hardcoded tools narrows the error scope, as long as you can check your data is shaped correctly, the analysis will be correct. Not sure how much of an advantage this is in the long term for working with proprietary datasets, but it's an interesting direction
- epistasis 1y agoThis is great, I've been on the waitlist for their website for a while and am now excited to be able to try it out!
- teenvan_1995 1y agoI wonder if giving 150+ tools is really a good idea considering context limitations. Need to check out if this works IRL.
- Herring 1y agoThere's an inner ToolRetriever which is a LLM call to select the most relevant tools/data/libraries.
- dmezzetti 1y agoVery interesting work! If biomedical research and paper analysis is of interest to you, I've been working on a set of open source projects that enable RAG over medical literature for a while. PaperAI: https://github.com/neuml/paperai https://github.com/neuml/paperai PaperETL: https://github.com/neuml/paperetl https://github.com/neuml/paperetl There is also this tool that annotates papers inline. AnnotateAI: https://github.com/neuml/annotateai https://github.com/neuml/annotateai
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- joelthelion 1y agoThis is really cool, but I think the big question is whether it works and whether it's useful to a professional. Is there anyone in the field who could comment on this?
- Domainzsite 1y ago[dead]
- dbcooper 1y agoAnyone have a spare invite?
- b0a04gl 1y ago[dead]