5 ms·
Show HN: A 6M-token movable window on a single 46GB GPU
- Wetime 2mo ago[dead]
- SwellJoe 2mo agoYou couldn't be bothered to write a coherent summary of what this actually is and what it does, you just let the AI write some random noise, eh?
- Wetime 2mo agoThis paper show a new tool called Galahad.Normally, AI has to think and guess the answer every time, which costs time and money. The knowledge of the model grows next to it not the model itself and no it is not the same as cache No fine-tuning needed It gives the exact same right answer every time, costs zero extra tokens, and saves lots of energy
- SwellJoe 2mo agoSo, you don't know what it is, either.
- Wetime 2mo agoWe handeld llm like a human brain we decopelled knowledge from the memory and build a memory layer that makes redoing things free and fast, so the llm can once it learned something solves it for free the next time
- saidnooneever 2mo agothis sounds like ur explaining caching
- SwellJoe 2mo agoSounds like they're explaining magic. Because current models cannot learn and they cannot remember.
- saidnooneever 2mo agothere is RL for LLMs which actually changes the weights but its more specialization than learning and wont counter the probabalistic nature of the thing
- SwellJoe 2mo agoTraining is not something you can just bolt on, and it generally requires even larger hardware than inference for a given model, and a huge amount of time. If you're aiming for "free", a OP claims, RL aint it.
- saidnooneever 2mo agohow is it not the same as a cache it its exact description matches the description of a cache?
- Wetime 2mo agoA cache remembers answers (only useful for the exact same question again). We remember the proven method and redo the work on every new question, so it solves ones it's never seen, which a cache simply can't.
- flowerbreeze 2mo agoMaybe it all works, but the paper is not trivial to decipher and the GitHub repository does not seem to exist. It doesn't seem to define what are the inputs to the system (what is a query? UTF-8 text? tokens?) and what are the outputs. It'd really help if the algorithm was written out step by step with all the expected type information included. At first I thought it was similar to something I've built before as a long-term slowly degrading cache for augmenting an FFN by caching well-learned answers, answering by performing a beam search in the key space resulting in located key accuracy measure (how well it corresponds to the input query) and answer confidence (has it been a long time since verification?), but that's not quite it? It feels similar in some way, but is it?
- himata4113 2mo ago"No implementation detail, algorithm, or configuration is contained in this document by design." + odd page cuts, it's as-if no human has ever looked at this before uploading it.
- SwellJoe 2mo agoThese LLMs are absolute poison for some folks.
- stephantul 2mo ago100% generated. I skimmed the paper, and came out with a feeling of still not knowing what this is about.
- Wetime 2mo agoWe handeld llm like a human brain we decoupled knowledge from the memory and build a memory layer that makes redoing things free and fast, so the llm can once it learned something solves it for free the next time
- array4277 2mo agoOnly a single 46Gb GPU? Wow AI sure is amazing tech.