7 ms·
Show HN: Visual intuitive explanations of LLM concepts (LLM University)
Hi HN,
We've just published a lot of original, visual, and intuitive explanations of concepts to introduce people to large language models.
It's available for free with no sign-up needed and it includes text articles, some video explanations, and code examples/notebooks as well. And we're available to answer your questions in a dedicated Discord channel.
You can find it here: https://llm.university/
Having written https://jalammar.github.io/illustrated-transformer/, I've been thinking about these topics and how best to communicate them for half a decade. But this project is extra special to me because I got to collaborate on it with two of who I think of as some of the best ML educators out there. Luis Serrano of https://www.youtube.com/@SerranoAcademy and Meor Amer, author of "A Visual Introduction to Deep Learning" https://kdimensions.gumroad.com/l/visualdl
We're planning to roll out more content to it (let us know what concepts interest you). But as of now, it has the following structure (With some links for highlighted articles for you to audit):
---
Module 1: What are Large Language Models
- Text Embeddings (https://docs.cohere.com/docs/text-embeddings)
- Similarity between words and sentences (https://docs.cohere.com/docs/similarity-between-words-and-sentences)
- The attention mechanism
- Transformer models (https://docs.cohere.com/docs/transformer-models HN Discussion: https://news.ycombinator.com/item?id=35576918)
- Semantic search
---
Module 2: Text representation
- Classification models (https://docs.cohere.com/docs/classification-models)
- Classification Evaluation metrics (https://docs.cohere.com/docs/evaluation-metrics)
- Classification / Embedding API endpoints
- Semantic search
- Text clustering
- Topic modeling (goes over clustering Ask HN posts https://docs.cohere.com/docs/clustering-hacker-news-posts)
- Multilingual semantic search
- Multilingual sentiment analysis
---
Module 3: Text generation
- Prompt engineering (https://docs.cohere.com/docs/model-prompting)
- Use case ideation
- Chaining prompts
---
A lot of the content originates from common questions we get from users of the LLMs we serve at Cohere. So the focus is more on application of LLMs than theory or training LLMs.
Hope you enjoy it, open to all feedback and suggestions!
- stclaus 3y agoLooks great, thanks! It would be useful to add chapters indicators / links to jump directly to a specific news in the audio
- jfarmer 3y ago> We've just published a lot of original, visual, and intuitive explanations of concepts to introduce people to large language models. Kinda frustrating that the main link dumps me onto what reads like a university syllabus, and nothing original, visual, or intuitive. If I click through the sections in order, there are 5 "preamble" sections describing logistical and other meta-information about the course. All text. The first pedagogical image I see this this, which tbh doesn't make any sense to me: https://files.readme.io/329efd5-image.png https://files.readme.io/329efd5-image.png "Where would you put the word apple?" The image alone doesn't work without reading the supporting text very closely. I also have to have a pretty sophisticated understanding to get the idea that I can represent words as points in a plane. Representing the words as icons is fundamentally confusing, too, I think. After all, maybe I say the word "apple" should go in "d" because it has at least two senses: a fruit and a machine. Oh, sorry, you failed your first quiz! "You can't fail the quiz, you're not being graded." Then why call it a quiz? Why use classroom metaphors unless you want students to fall back on classroom behaviors? Of course, you know the #1 student classroom behavior: not reading the syllabus. But if I have no trouble with that level of abstraction, what's with the cutesty way of describing the problem? Get rid of all this chocolate-covered broccoli. Just say and show what you mean. Computers like numbers. Vectors are lists of numbers. Vectors come with concepts like length and distance. We want to transform words into vectors so that words we think of as similar are close together as vectors. There are many ways to translate words into vectors. Here are 5-10 examples of how we might do that. What are some pros/cons? What relationship(s) do they make clear or obscure? Get them thinking about what it means to embed things and why we'd want to embed words one way vs. another. That'll pay dividends. Having them remember "where the apple icon goes" isn't going to be something they'll benefit from reflecting on in any future experience.
- jayalammar 3y agoThe landing page is technically the course overview. I'd love to hear what you think would've made it more engaging for you. We can probably pull up some of the visuals to it as a preview. Let me see what we can do on that front.
- 3y ago
- ZeroCool2u 3y agoThis looks like a pretty great resource and I'm looking forward to checking it out. My only ask is that since it's the type of site I'd probably be looking at for quite a while it'd be nice if it had a dark mode.
- sva_ 3y agoInteresting, just yesterday I was googling something about transformers and had arrived on your page.
- toppy 3y agoJay, I liked your tutorial on Transformer models. Helped me a lot when I read it in 2020. One of the best resources on a topic then. Thanks for your work! Fingers crossed for your new endeavour.
- jayalammar 3y agoThank you so much (and others for your kind messages). Glad you found them useful! Writing is the best way for me to learn, I find.
- senttoschool 3y agoLooks great. Thank you.
- HarHarVeryFunny 3y agoI'm not sure how much is actually known to write about, but what I'd like to see explained is how transformer-based LLMs/AI really work - not at the mechanistic level of the architecture, but in terms of what they learn (some type of world model ? details, not hand waving!) and how do they utilize this when processing various types of input ? What type of representations are being used internally in these models ? We've got token embeddings going in, and it seems like some type of semantic embeddings internally perhaps, but exactly what ? OTOH it's outputting words (tokens) with only a linear layer between the last transformer block and the softmax, so what does that say about the representations at that last transformer block ?
- uoaei 3y ago> not at the mechanistic level of the architecture, but in terms of what they learn (some type of world model ? details, not hand waving!) https://imgs.xkcd.com/comics/tasks.png https://imgs.xkcd.com/comics/tasks.png
- HarHarVeryFunny 3y agoSure - but it's still the interesting part! I'm sure some of key players know at least a little, but they don't seem inclined to share. In his Lex Fridman interview Sam Altam said something along the lines of "a LOT of knowledge went into designing GPT-4", and there's a time gap between GPT-3 (2020) and GPT-4 (2022) where it seems they spent a lot of time probably trying to understand it, among other things. It seems the way values are looked up via query/key and added must constrain representations quite a bit, and comparing internal activations for closely related types of input might be one way to start to understand what's going on. A high level understanding of what the model has learnt may be the last thing to fall, but understanding the internal representations would go a long way towards that.
- quickthrower2 3y agoAre you saying no one really knows how these things work? I am very curious about if you can “peer into the weights”. I have seen simple examples of that with image recognition but only for early layers.
- kfarr 3y agoThis is pretty excellent material, even just spending 10 minutes I have learned more than most random blog posts in the past few months.
- abrinz 3y agoNice work! Minor nitpick: The intercom button obscures the topic expansion button for the final appendix in the nav menu. Maybe move intercom to the bottom right instead?
- beeburrt 3y agoYou know what would be helpful? A little tag or something at the beginning of each section that says about how long it's going to take. From what I've seen so far, it looks awesome. I'm excited to dive in. Thanks!
- jwilber 3y agoLove these. I’ve also made some visual explanations for ml for Amazon, available at https://mlu-explain.github.io/ https://mlu-explain.github.io/ Big fan of your early work, Jay, a big inspiration for me!
- jayalammar 3y agoThat's beautiful! Hope you're getting to do more of these!
- jwilber 3y agoThanks, am definitely trying to squeeze them out as I’m able!
- axpy906 3y agoYou sir get an up vote for simply being Jay on HN. Thank you for all you do.
- coolandsmartrr 3y agoHi Jay, I really loved your [explainer on AI Art](https://www.youtube.com/watch?v=MXmacOUJUaw https://www.youtube.com/watch?v=MXmacOUJUaw), and I've already added more of your videos and articles on my watch-later read-later lists! Can't wait to spend more time with them this weekend. Thank you for creating such wonderful resources!
- 40fishes 3y agoLooks really helpful. Joined the community as well.