6 ms·
Training mRNA Language Models Across 25 Species for $165
We built an end-to-end protein AI pipeline covering structure prediction, sequence design, and codon optimization. After comparing multiple transformer architectures for codon-level language modeling, CodonRoBERTa-large-v2 emerged as the clear winner with a perplexity of 4.10 and a Spearman CAI correlation of 0.40, significantly outperforming ModernBERT. We then scaled to 25 species, trained 4 production models in 55 GPU-hours, and built a species-conditioned system that no other open-source project offers. Complete results, architectural decisions, and runnable code below.
- HocusLocus 6mo agogray goo of the future
- deleted 6mo ago[deleted]
- simianwords 6mo agoWhat makes these Domain specific models work when we don’t have good domain models for health care, chemistry, economics and so on
- colechristensen 6mo ago>we don’t have good domain models for health care, chemistry, economics and so on Who says we don't?
- simianwords 6mo agoExamples please?
- colechristensen 6mo agoNo, it's really simple to search for domain specific models being used "in production" all over the place
- simianwords 6mo agoI didn’t find a single one that outperforms a general model.
- colechristensen 6mo agoOk, alphafold.
- simianwords 6mo agoIt’s not a large language model
- khalic 6mo ago> In Progress: CodonJEPA JEPA is going to break the whole industry :D
- digdugdirk 6mo agoCan you explain this? I haven't heard of JEPA, and from a quick search it seems to be vision/robotics based?
- lukeinator42 6mo agohttps://openreview.net/pdf?id=BZ5a1r-kVsf https://openreview.net/pdf?id=BZ5a1r-kVsf
- khalic 6mo agoIt’s a self supervised learning architecture, and it’s pretty much universal. The loss function runs on embeddings, and some other smart architectural choices allover. Worth diving into for a few hours, Yann LeCun gives some interesting talks about it
- rubicon33 6mo agoCan someone explain what one might use this model for? As a developer with a casual interest in biology it would be fun to play with but honestly not sure what I would do
- colechristensen 6mo agoYou can get your feet wet with genetic engineering for surprisingly little money. This guy shows a lot of how it's done: https://www.youtube.com/@thethoughtemporium https://www.youtube.com/@thethoughtemporium Basically you can design/edit/inject custom genes into things and see real results spending on the scale of $100-$1000.
- someuser54541 6mo agoIs there something like this in text/readable format?
- _zoltan_ 6mo agoMy main concern is using fungi. If it ends up in my lungs I'm most likely screwed, right?
- nurettin 6mo agoYes, but most students produce their best work while infected.
- colechristensen 6mo agoThis is the classic meme https://www.reddit.com/r/labrats/comments/mmv2ig/lab_strains_unite/#lightbox https://www.reddit.com/r/labrats/comments/mmv2ig/lab_strains... Lab strains of things tend to be extremely sensitive and not human adapted. You shouldn't study and modify human-infecting organisms in your basement anyway. While you shouldn't ignore protective equipment and proper procedure... paranoia about infecting yourself with a lab leak isn't warranted.
- 6mo ago
- yieldcrv 6mo agoDistributing the load on this will probably be infinitely more useful than “folding at home”
- skyskys 6mo agohmmmm seems like some fake hype.
- deleted 6mo ago[deleted]
- seamossfet 6mo agoThe problem with models like this is they're built on very little actual training data we can trace back to verifiable protein data. The protein data back, and other sources of training data for stuff like this, has a lot of broken structures in them and "creative liberties" taken to infer a structure from instrument data. It's a very complex process that leaves a lot for interpretation. On top of that, we don't have a clear understanding on how certain positions (conformations) of a structure affect underlying biological mechanisms. Yes, these models can predict surprisingly accurate structures and sequences. Do we know if these outputs are biologically useful? Not quite. This technology is amazing, don't get me wrong, but to the average person they might see this and wonder why we can't go full futurism and solve every pathology with models like these. We've come a long way, but there's still a very very long way to go.
- stardust2 6mo agoHow do we get more verifiable protein data? So even if we had better data, we don't yet understand how the structure impacts the biology?
- colingauvin 6mo agoHN's blindspots never cease to amaze me. I am a structural biologist working in pharmaceutical design and this type of thing could be wildly useful (if it works).
- justinclift 6mo agoBlind spot?
- dhruv3006 6mo agoInteresting work - Looks like AI for science is having it's day right now.
- nradclif 6mo ago"Complete results, architectural decisions, and runnable code below." This is a weird post, there doesn't seem to be any "below" here. Another comment linked the article: https://huggingface.co/blog/OpenMed/training-mrna-models-25-species https://huggingface.co/blog/OpenMed/training-mrna-models-25-...
- justinclift 6mo agoYeah. Things like "Complete results, architectural decisions, and runnable code below." is literally how AI outputs stuff, so I'd expect the post was AI written too. :(
- jazzpush2 6mo agoA Codon-based model is cool. I know NVIDIA is building quite a large one. At GTC they showed an SAE they built on a smaller version of it, allowing you to see what their model learned: https://research.nvidia.com/labs/dbr/blog/sae/ https://research.nvidia.com/labs/dbr/blog/sae/
- agenexus 5mo ago[flagged]
- maziyar 6mo agofull article: https://huggingface.co/blog/OpenMed/training-mrna-models-25-species https://huggingface.co/blog/OpenMed/training-mrna-models-25-...
- xyz100 6mo agoWhat makes this dataset or problem worth solving compared to other health datasets? Would the results on this task be broadly useful to health?
- CyberDildonics 6mo agoWhat other "datasets" are you talking about? How do you "solve a dataset" ?
- xyz100 6mo agoYou solve a dataset when you learn what there is to learn about the phenomenon of interest. The limit of such phenomenon is “cure all disease”, and clearly this is not solving that.
- CyberDildonics 6mo agoWhat are you talking about? "the phenomenon of interest"? There is nothing you wrote in either comment that makes sense. What is a "dataset" that has been "solved" and what did the program do that 'solved' it?
- xyz100 5mo agoMNIST (the number classification task) has been “solved” a billion times and it is hard to imagine any subsequent advances there as scores using a variety of methods have hit the saturation point of accuracy. Any further improvements are likely overfitting to noise. Therefore, we know that it is easy to detect handwritten numbers. However, we may not know how to detect other things as well, like reading an MRI. Those datasets/tasks are clearly different and require different techniques. Training an LLM is likewise different.