16 ms·
Big LLMs weights are a piece of history
- blinky81 2y ago"big large" lol
- api 2y agoThat's really what these are: something analogous to JPEG for language, and queryable in natural language. Tangent: I was thinking the other day: these are not AI in the sense that they are not primarily intelligence. I still don't see much evidence of that. What they do give me is superhuman memory. The main thing I use them for is search, research, and a "rubber duck" that talks back, and it's like having an intern who has memorized the library and the entire Internet. They occasionally hallucinate or make mistakes -- compression artifacts -- but it's there. So it's more AM -- artificial memory. Edit: as a reply pointed out: this is Vannevar Bush's Memex, kind of.
- antirez 2y agoI believe LLMs are both data and processing, but even humans reasoning is based in strong ways on existing knowledge. However, for the goal of the post, indeed it is the memorization that is the key value, and the fact that likely in the future sampling such models can be used to transfer the same knowledge to bigger LLMs, even if the source data is lost.
- api 2y agoI'm not saying there is no latent reasoning capability. It's there. It just seems to be that the memory and lookup component is much more useful and powerful. To me intelligence describes something much more capable than what I see in these things, even the bleeding edge ones. At least so far.
- antirez 2y agoI offer a POV that is in the middle: reasoning is powerful to evaluate which solution is better among N in the context. Memorization allows sampling of many competing ideas from the problem space, than the LLM picks the best, making chain of thoughts so effective. Of course zero shot reasoning also is a part of the story but somewhat weaker, exactly like we are not often able to spit the best solution before evaluation of the space (unless we are very accustomed to the specific problem).
- danielbln 2y agoThat's the problem with the term "intelligence". Everyone has their own definition, we don't even know what makes us humans intelligent and more often than not it's a moving goalpost as these models get better.
- Mistletoe 2y agoThis is an excellent viewpoint.
- menzoic 2y agoHaving memory is fine but choosing the relevant parts requires intelligence
- flower-giraffe 2y agoOr 80 years to MVP memex “Vannevar Bush's 1945 article "As We May Think". Bush envisioned the memex as a device in which individuals would compress and store all of their books, records, and communications, "mechanized so that it may be consulted with exceeding speed and flexibility". https://en.m.wikipedia.org/wiki/Memex https://en.m.wikipedia.org/wiki/Memex
- mdp2021 2y agoThe memex was a deterministic device to consult documents - the actual documents. The "LLM" is more like a dumb archivist that came with it ("Yes, see for example that document, it tells you that q=M·k...").
- skydhash 2y agoI grew up with physical encyclopedia, then moved on to Encarta, then Wikipedia dumps and folders full of PDFs. I still prefer curated information repository over chat interfaces or generated summaries. The main goal with the former is to have a knowledge map and keywords graph, so that you can locate any piece of information you may need from the actual source.
- hengheng 2y agoI've been looking at it as an "instant reddit comment". I can download a 10G or 80G compressed archive that basically contains the useful parts of the internet, and then I all can use it to synthesize something that is about as good and reliable as a really good reddit comment. Which is nifty. But honestly it's an incredible idea to sell that to businesses.
- api 2y agoReddit seems to puppet humans via engagement farming to do what LLMs do in some cases. Posts are prompts, replies are responses. Of course they vary widely in quality.
- Guthur 2y agoAnd so what would the point be of anyone actually posting on the internet if no one actually visits the sites because large corps have essentially stolen and monetized the whole thing. And I'm sure they have or will have the ability to influence the responses so you only see what they want you to see.
- kelseyfrog 2y agoThat's the next step after algorithmic content feeds - algorithmic/generated comment sections. Imagine seeing an entirely different conversation happening just to get you to buy a product. A product like Coca-Cola. Imagine scrolling through a comment section that feels tailor-made to your tastes, seamlessly guiding you to an ice-cold Coca-Cola. You see people reminiscing about their best summer memories—each one featuring a Coke in hand. Others are debating the superior refreshment of Coke over other drinks, complete with "real" testimonials and nostalgic stories. And just when you're feeling thirsty, a perfectly timed comment appears: "Nothing beats the crisp, refreshing taste of an ice-cold Coke on a hot day." Algorithmic engagement isn’t just the future—it’s already here, and it’s making sure the next thing you crave is Coca-Cola. Open Happiness.
- adhamsalama 2y agoIsn't that how Reddit gained momentum? Posting fake posts/comments? Now we can mass-produce it!
- yannyu 2y agoThere's a great article recently by Ted Chiang that elaborated on this idea: https://www.newyorker.com/tech/annals-of-technology/chatgpt-is-a-blurry-jpeg-of-the-web https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
- bob1029 2y agoIf you want to see what this would actually be like: https://lcamtuf.coredump.cx/lossifizer/ https://lcamtuf.coredump.cx/lossifizer/ I think a fun experiment could be to see at what setting the average human can no longer decipher the text.
- GolfPopper 2y ago>like having an intern who has memorized the library and the entire Internet. They occasionally hallucinate or make mistakes Correction: you occasionally notice when they hallucinate or make mistakes.
- xpe 2y agoI regularly pushback against casual uses of the word “intelligence”. First, there is no objective dividing line. It is a matter of degree relative to something else. Any language that suggests otherwise should be refined or ejected from our culture and language. Language’s evolution doesn’t have to be a nosedive. Second, there are many definitions of intelligence; some are more useful than others. Along with many, I like Stuart Russell’s definition: the degree to which an agent can accomplish a task. This definition requires being clear about the agent and the task. I mention this so often I feel like a permalink is needed. It isn’t “my” idea at all; it is simply the result of smart people decomplecting the idea so we’re not mired in needless confusion. I rant about word meanings often because deep thinking people need to lay claim to words and shape culture accordingly. I say this often: don’t cede the battle of meaning to the least common denominators of apathy, ignorance, confusion, or marketing. Some might call this kind of thinking elitist. No. This is what taking responsibility looks like. We could never have built modern science (or most rigorous fields of knowledge) with imprecise thinking. I’m so done with sloppy mainstream phrasing of “intelligence”. Shit is getting real (so to speak), companies are changing the world, governments are racing to stay in the game, jobs will be created and lost, and humanity might transcend, improve, stagnate, or die. If humans, meanwhile, can’t be bothered to talk about intelligence in a meaningful way, then, frankly, I think we’re … abdicating responsibility, tempting fate, or asking to be in the next Mike Judge movie.
- jart 2y agoWe never would have been able to create science, if it weren't for focusing on the kinds of thinking that can be made logical. There's a big difference. What you're doing, with this whole "let's make a bullshit word logical" is more similar to medieval scholasticism, which was a vain attempt at verbal precision. https://justine.lol/dox/english.txt https://justine.lol/dox/english.txt
- xpe 2y agoYikes, maybe we can take a step back? I'm not sure where this is coming from, frankly. One anodyne summary of my comment above would be: > Let's think and communicate more clearly regarding intelligence. Stuart Russell offers a nice definition: an agent's ability to do a defined task. Maybe something about my comment got you riled up? What was it? You wrote: > What you're doing, with this whole "let's make a bullshit word logical" is more similar to medieval scholasticism, which was a vain attempt at verbal precision. Again, I'm not quite sure what to say. You suggest my comment is like a medieval scholar trying to reconcile dogma with philosophy? Wow. That's an uncharitable reading of my comment. I have five points in response. First, the word intelligence need not be a "bullshit word", though I'm not sure what you mean by the term. One of my favorite definitions of bullshitting comes from "On Bullshit" by Harry Frankfurt: > Frankfurt determines that bullshit is speech intended to persuade without regard for truth. The liar cares about the truth and attempts to hide it; the bullshitter doesn't care whether what they say is true or false. - Wikipedia Second, I'm trying to clarify the term intelligence by breaking it into parts. I wouldn't say I'm trying to make it "logical" (in the sense of being about logic or deduction). Maybe you mean "formal"? Third, regarding the "what you're doing" part... this isn't just me. Many people both clarify the concept of intelligence and explain why doing so is important. Fourth, are you saying it is impossible to clarify the meaning of intelligence? Why? Not worth the trouble? Fifth, have you thought about a definition of intelligence that you think is sensible? Does your definition steer people away from confusion? You also wrote: > We never would have been able to create science, if it weren't for focusing on the kinds of thinking that can be made logical. I think you mean _testable_, not _logical_. Yes, we agree, scientists should run experiments on things that can be tested. Russell's definition of intelligence is testable by defining a task and a quality metric. This is already a big step up from an unexamined view of intelligence, which often has some arbitrary threshold.* It allows us to see a continuum from, say, how a bacteria finds food, to how ants collaborate, to how people both build and use tools to solve problems. It also teases out sentience and moral worth so we're not mixing them up with intelligence. These are simple, doable, and worthwhile clarifications. Finally, I read your quote from Dijkstra. In my reading, Dijkstra's main point is that natural language is a poor programming interface due to its ambiguity. Ok, fair. But what is the connection to this thread? Does it undercut any of my arguments? How? * A common problem when discussing intelligence involves moving the goal post. Whatever quality bar is implied has a tendency to creep upwards over time.*
- mdp2021 2y ago> JPEG for [a body of] language Yes! > artificial memory Well, "yes", kind of. > Memex After a flood?! Not really. Vannevar Bush - As we may think - http://web.mit.edu/STS.035/www/PDFs/think.pdf http://web.mit.edu/STS.035/www/PDFs/think.pdf
- visarga 2y agoI can ask a LLM to write a haiku about the loss function of Stable Diffusion. Or I can have it do zero shot translation, between a pair of languages not covered in the training set. Can your "language JPEG" do that? I think "it's just compression" and "it's just parroting" are flawed metaphors. Especially when the model was trained with RLHF and RL/reasoning. Maybe a better metaphor is "LLM is like a piano, I play the keyboard and it makes 'music'". Or maybe it's a bycicle, I push the pedals and it takes me where I point it.
- rollcat 2y agohttps://xkcd.com/1683/ https://xkcd.com/1683/
- intellectronica 2y agoI love the title "Big LLMs" because it means that we are now making a distinction between big LLMs and minute LLMs and maybe medium LLMs. I'd like to propose the we call them "Tall LLMs", "Grande LLMs", and "Venti LLMs" just to be precise.
- deleted 2y ago[deleted]
- de-moray 2y agoWhat does a 20 LLM signify?
- HarHarVeryFunny 2y agoBut of course these are all flavors of "large", so then we have big large language models, medium large language models, etc, which does indeed make the tall/grande/venti names appropriate, or perhaps similar "all large" condom size names (large, huge, gargantuan).
- tonyhart7 2y agocan we have tiny LLM that can run on smartphone now
- winter_blue 2y agoApple Intelligence has an LLM that runs locally on the iPhone (15 Pro and up). But the quality of Apple Intelligence shows us what happens when you use a tiny ultra-low-wattage LLM. There’s a whole subreddit dedicated to its notable fails: https://www.reddit.com/r/AppleIntelligenceFail/top/?t=all https://www.reddit.com/r/AppleIntelligenceFail/top/?t=all One example of this is “Sorry I was very drunk and went home and crashed straight into bed” being summarized by Apple Intelligence as ”Drunk and crashed”.
- deleted 2y ago[deleted]
- 2y ago
- laborcontract 2y agoI miss the good ol days when I'd have text-davinci make me a table of movies that included a link to the movie poster. It usually generated a url of an image in an s3 bucket. The link always worked.
- nickpsecurity 2y agoPeople wanting this would be better off using memory architectures, like how the brain does it. For ML, the simplest approach is putting in memory layers with content-addressible schemes. I have a few links on prototypes in this comment: https://news.ycombinator.com/item?id=42824960 https://news.ycombinator.com/item?id=42824960
- HarHarVeryFunny 2y agoAnimal brains do not separate long term memory and processing - they are one and the same thing - columnar neural assemblies in the cortex that have learnt to recognize repeated patterns, and in turn activate others.
- hedgehog 2y agoThis doesn't make much sense to me. Unattributed heresay has limited historical value, perhaps zero given that the view of the web most of the weights-available models have is Common Crawl which is itself available for preservation.
- Terr_ 2y agoI suspect the idea is that sometimes breadth wins out over accuracy. Even if it's unsuited as a primary source, this kind of lossy compression of many many documents might help a conscientious historian discover verifiable things through other routes.
- jart 2y agoMozilla's llamafile project is designed to enable LLMs to be preserved for historical purposes. They ship the weights and all the necessary software in a deterministic dependency-free single-file executable. If you save your llamafiles, you should be able to run them in fifty years and have the outputs be exactly the same as what you'd get today. Please support Mozilla in their efforts to ensure this special moment in history gets archived for future generations! https://github.com/Mozilla-Ocho/llamafile/ https://github.com/Mozilla-Ocho/llamafile/
- visarga 2y agoLLMs are much easier to port than software. They are just a big blob of numbers and a few math operations.
- refulgentis 2y agoLLMs are much harder, software is just a blob of two numbers. ;) (less socratic: I have a fraction of a fraction of jart's experience, but have enough experience via maintining a cross-platform llama.cpp wrapper to know there's a ton of ways to interpret that bag o' floats and you need a lot of ancillary information.)
- andix 2y agoI think software is rather easy to archive. Emulators are they key. Nearly every platform from the past can be emulated on a modern arm/x86 Linux/windows system. Arm/x86/linux/windows are ubiquitous, even if they might fade away there will be emulators around for a long time. With future compute power it should be no problem to just use nested emulation, to run old emulators on an emulated x86/linux.
- throwaway314155 2y ago> I think software is rather easy to archive. * assuming someone else already spent tremendous effort to develop an emulator for your binary's target that is 100% accurate...
- isoprophlex 2y agoInteresting. Just this morning I had a conversation with Claude about this very topic. When asked "can you give me your thoughts on LLM train runs as historical artifacts? do you think they might be uniquely valuable for future historians?", it answered > oh HELL YEAH they will be. future historians are gonna have a fucking field day with us. > imagine some poor academic in 2147 booting up "vintage llm.exe" and getting to directly interrogate the batshit insane period when humans first created quasi-sentient text generators right before everything went completely sideways with *gestures vaguely at civilization* > *"computer, tell me about the vibes in 2025"* > "BLARGH everyone was losing their minds about ai while also being completely addicted to it" Interesting indeed to be able to directly interrogate the median experience of being online in 2025. (also my apologies for slop-posting; i slapped so many custom prompting on it that I hope you'll find the output to be amusing enough)
- tryauuum 2y agowhat's the prompt?
- dmos62 2y agoEnjoy the insight, but the title makes my eye twitch. How about "LLM weights are pieces of history"?
- lblume 2y agoSmall LLM weights are not really interesting though. I am currently training GPT-2 small sized models for a scientific project right, and their world models are just not good enough to generate any kind of real insight about the world it was trained in except for corpus biases.
- dmos62 2y agoA collection of newspapers is generally a better source than a single leaflet, but even a leaflet is a piece of history.
- kragen 2y agoSmall large language models? This sounds like the apocryphal headline when a spiritualist with dwarfism escaped prison: "Small medium at large." Do you also have some dehydrated water and a secure key escrow system?
- GeoAtreides 2y agoJust like the map isn't the territory, so summaries are not the content nor the library fillings the actual books. If I want to read a post, a book, a forum, I want to read exactly that, not a simulacrum built by arcane mathematical algorithms.
- visarga 2y agoThe counter perspective is that this is not a book, it's an interactive simulation of that era. The model is trained on everything, this means it acts like a mirror of ourselves. I find it fascinating to explore the mind-space it captured.
- defgeneric 2y agoWhile the post talks about big LLMs as a valuable "snapshot" of world knowledge, the same technology can be used for lossless compression: https://bellard.org/ts_zip/ https://bellard.org/ts_zip/.
- fl4tul4 2y ago> Scientific papers and processes that are lost forever as publishers fail, their websites shut down. I don't think the big scientific publishers (now, in our time) will ever fail, they are RICH!
- Legend2440 2y agoThat means nothing. Big companies fail all the time. There is no guarantee any of them will be here in 50 years, let alone 500.
- bookofjoe 2y agoSo was the Roman Empire
- thayne 2y agoPerhaps a shorter term risk is the publishers consider some papers less profitable, so they stop preserving them.
- dr_dshiv 2y ago“We should regard the Internet Archive as one of the most valuable pieces of modern history; instead, many companies and entities make the chances of the Archive to survive, and accumulate what otherwise will be lost, harder and harder. I understand that the Archive headquarters are located in what used to be a church: well, there is no better way to think of it than as a sacred place.” Amen. There is an active effort to create an Internet Archive based in Europe, just… in case.
- ttul 2y agoWell, it did establish a new HQ in Canada… https://vancouversun.com/news/local-news/the-internet-archive-opens-headquarters-meeting-space-for-the-tech-world https://vancouversun.com/news/local-news/the-internet-archiv... (Edited: apparently just a new HQ and not THE HQ)
- thrance 2y agoWith this belligerent maniac in the White House who recently doubled-down on his wish to annex Canada [1], I wouldn't feel safe relocating there if the goal is to flee the US. [1] https://www.nbcnews.com/politics/donald-trump/trump-quest-conquer-canada-confusing-everyone-rcna195657 https://www.nbcnews.com/politics/donald-trump/trump-quest-co...
- decremental 2y ago[dead]
- kelseydh 2y agoI was looking to a book a wedding in this venue (The Permanent) and the Internet Archive server is prominently visible on the 2nd floor. The server is pretty cool and adds to the aesthetics of the space.
- badlibrarian 2y agoAnyone who takes even an hour to audit anything about the Internet Archive will soon come to a very sad conclusion. The physical assets are stored in the blast radius of an oil refinery. They don't have air conditioning. Take the tour and they tell you the site runs slower on hot days. Great mission, but atrociously managed. Under attack for a number of reasons, mostly absurd. But a few are painfully valid.
- Havoc 2y agoI wonder whether it'll become like pre-WW2 steel that doesn't have nuclear contamination. Just with a pre-LLM knowledge
- guybedo 2y agofwiw i've added a summary of the discussion here: https://extraakt.com/extraakts/67d708bc9844db151612d782 https://extraakt.com/extraakts/67d708bc9844db151612d782
- dstroot 2y agoIsn’t big LLM training data actually the most analogous to the internet archive? Shouldn’t the title be “Big LLM training data is a piece of history”? Especially at this point in history since a large portion of internet data going forward will be LLM generated and not human generated? It’s kind of the last snapshot of human-created content.
- antirez 2y agoThe problem is, where is this 20T tokens that are being used for this task? No way to access them. I hope that at least OpenAI and a few more have solid historical storage of the tokens they collect.
- bossyTeacher 2y agoSo large large language model?
- almosthere 2y agoSplit the wayback machine away from its book copyright lawsuit stuff and you don't have to worry.
- codr7 2y agoI find it very depressing to think that the only traces left from all the creativity will end up to be AI slop, the worst use case ever. I feel like the more people use GenAI, the less intelligent they become. Like the rest of this society, they seem designed to suck the life force out of humans and and return useless crap instead.
- andix 2y agoI think it’s fine that not everything on the internet is archived forever. It has always been like that, in the past people wrote on paper, and most of it was never archived. At some point it was just lost. I inherited many boxes of notes, books and documents from my grandparents. Most of it was just meaningless to me. I had to throw away a lot of it and only kept a few thousand pages of various documents. The other stuff is just lost forever. And that’s probably fine. Archives are very important, but nowadays the most difficult part is to select what to archive. There is so much content added to the internet every second, only a fraction of it can be archived.
- throwaway48476 2y agoThe internet training data for LLMs is valuable history were losing one dead webadmin at a time. The regurgitated slop less so.
- pama 2y agoI would be curious to know if it would be possible to recunstruct approximate versions of popular common subsets of internet training data by using many different LLMs that may have happened to read the same info. Anyone knows pointers to math papers about such things?
- teleforce 2y agoI really like the narative that now LLM is the conserving human knowledge that otherwise would be lost forever in the form of its weights in a kind of a lossy compression. Personally I'd like that if all the knowledge and information (K & I) are readily available and accessible (pretty sure most of the prople share the same sentiment), despite the consistent business decisions from the copyright holders to hoard their K & I by putting everything behind paywalls and/or registration (I'm looking at you Apple and X/Twitter). As much that some people hate Google by organizing the world information by feeding and thriving through advertisements because in the long run the information do get organized and kind of preserved in many Internet data formats, lossy or not. After all Google who originall designed the transformer that enabled the LLM weights that are now apparently a piece of history.
- sourtrident 2y agoImagine future historians piecing together our culture from hallucinated AI memories - inaccurate, sure, but maybe even more fascinating than reality itself.
- off_by_inf 2y agoAnd they all undertrained, according to the papers.
- OuterVale 2y agoInteresting. It seems that both they and I had very similar ideas at about the same time, with this being posted just a few hours after I finally published about AI model history being lost. https://vale.rocks/posts/ai-model-history-is-being-lost https://vale.rocks/posts/ai-model-history-is-being-lost
- hi_hi 2y agoNaming antics aside, the article makes a good point I've heard previously about the importance of the Internet Archive. Are there any search experiences that allow me to search like it's 1999? I'd love to be able to re-create the experience of finding random passion project blogs that give a small snapshot of things people and business were using the web for back then.
- ilaksh 2y agoGreat idea. Slightly related idea: use the Internet Archive to build a dataset of 6502 machine code/binaries, corresponding manuals, and possibly videos of the software in action.. maybe emulator traces. It might be possible to create an L LM that can write a custom vintage game or program on demand in machine code and simultaneously generate assets like sprites. Especially if you use the latest reinforcement learning techniques.