6 ms·
Just to think what this will look like in a couple of years.
by buddhistdude 3mo ago
Just to think what this will look like in a couple of years.
- OGWhales 3mo agoHopefully like this (but smarter): https://chatjimmy.ai/ https://chatjimmy.ai/
- nomel 3mo agoThis is genuinely confusing to my senses. The future is going to be so strange/neat/me unemployed.
- razodactyl 3mo agoYeah. It keeps catching me off guard that it answered me already.
- falcor84 3mo ago> strange/neat/me unemployed I'm not sure if that's what you were going for, but I read it as if it were written by The Board in the game Control, and found myself with the appropriate level of existential dread.
- matheusmoreira 3mo agoThe future is totally illegible to me. I love these AI models, but I feel like I'm going to be jobless within 10 years. Anomie is at an all time high right now.
- the_af 3mo ago10 years? An optimist, I see.
- niyazpk 3mo agoWow.. what?! How is this so fast?! Where can I read more?
- dmd 3mo agohttps://taalas.com/ https://taalas.com/
- fcsp 3mo agoFunnily enough, pasting your comment straight into Jimmy leads to a... Funnily suboptimal answer that does not answer the question. As someone else already contributed, this is driven by a Canadian startup taalas that basically makes chips that are llms, so everything is very fast but also, baked into the chip. Once this kind of stuff is a commodity in like 10 years, our world will be very, very different.
- hajile 3mo agoTaalas HC1 AI uses Llama 3.1 8B, but takes up a massive 53B transistors and 815mm2 on TSMC N6 (nearly at the reticle limit of 858mm2). N2 is a little less than 3x as dense (110MTr/mm2 vs 313MTr/mm2). This chip would still be 272mm2 on N2 which is an eye-watering $30k/wafer and bigger than a 9950x or Nvidia 5070. This just isn't feasible. Some of the latest-gen LLMs seem to have 5-10T parameters or about 1000x more. I don't know that taping out just one chip makes economic sense let alone the 300-1000 chips required for a cutting-edge model. Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips. There are a TON of uses for an 8B parameter models on the edge, but this is WAY too big to put on the edge of anything. Something like a 10mm2 100m parameter voice model might be feasible on the edge, but only for expensive devices, but most of those are TSMC 28nm (up to 29MTr/mm2) or GF FDX22 (up to 40MTR/mm2) which would increase the AI chip to the point where it would absolutely dominate the BOM.
- HaloZero 3mo agothe flash models have fallen in size at least between deep seek models. Is there a limit to the shrinking capacity of the models?
- plaguuuuuu 3mo ago[dead]
- kkotak 3mo agoWhy is the insane speed of 13KTPS of this site is not more on the the top of the AI conversations?
- Ey7NFZ3P0nzAe 3mo agoIt's pretty well known by now.
- chromadon 3mo agoI asked it for a block of C++ code and it hit 14,189 tok/s. I assume it cached someone else's session?
- fcsp 3mo agoNo - it's custom silicon https://news.ycombinator.com/item?id=48693490 https://news.ycombinator.com/item?id=48693490
- mlrtime 3mo agoBecause I just tested it and it took 3-4 clarifications before it actually gave a correct response vs gemini/google search. It's not great, but good. I'd rather wait 3x as long.
- mike_hearn 3mo agoBecause there's been nothing to discuss since their announcement. Their API access immediately closed due to overwhelming demand and they didn't fab newer models than Llama3 yet. Probably they will make bank selling to HFT for a while.
- vitorgrs 3mo agoNot opening here... HN killed?
- Bombthecat 3mo agoWhat How? Which model is behind it?
- victorbjorklund 3mo agoDamn that is crazy.
- archon810 3mo agoThis is the reaction every time it's posted, and deservedly so.
- dirasieb 3mo agohugged to death?
- jeingham 3mo agoThis caused me to have some sense what blistering fast AI actually is. What it means for the future is a question that remains.
- refulgentis 3mo agoImagine a Beowulf cluster of these…
- noisy_boy 3mo agoThat's a name I haven't heard in a while.
- mlrtime 3mo agoFirst post?
- cactusplant7374 3mo agoThe user has many comments and updoots if you look at their profile.
- refulgentis 3mo agoWe’re being silly and spamming Slashdot spam comments :p (“imagine a Beowulf cluster” and “first post?”)
- addaon 3mo agoMe too!
- chromadon 3mo agoI always think of Furbies because of that geocities (memories!) site.
- senectus1 3mo agoprobably something like this https://sb0xw.csb.app/ https://sb0xw.csb.app/
- alienbaby 3mo agoI started with a 2400baud modem, I've seen how this goes
- accrual 3mo agoSometimes I visualize a setup like this [0], based on 2D art by Simon Stålenhag. Someone has their home robot sitting on a desk connected to their old PC with thick cabling, dumping endless lines of each subsystem's <think> logs to diagnosis why it did something weird earlier in the day. Systems pushing 750+ tokens per second per subsystem might even be considered on the slow side for realtime tasks by then. [0] https://www.therookies.co/entries/39513 https://www.therookies.co/entries/39513
- bredren 3mo agoProbably will not be looking at text like this in a few years.
- cactusplant7374 3mo agoProbably not. Everyone will still need a lot of reasoning tokens and tool calls. Running the tests for every round is tiring but must be done.