8 ms·
A 30B Qwen model walks into a Raspberry Pi and runs in real time
- yjftsjthsd-h 9mo agoIn case anyone else clicked in wondering what counts as "real time" for this: > On a Pi 5 (16GB), Q3_K_S-2.70bpw [KQ-2] hits 8.03 TPS at 2.70 BPW and maintains 94.18% of BF16 quality. And they talk about other hardware and details. But that's the expanded version of the headline claim.
- CSSer 9mo agoSomeone should make a version of the Hacker News homepage that is just LLM extracts of key article details like this.
- mschuster91 9mo agoPlease not. There were some bots (or karma-farming users) doing this and yuck, was it annoying.
- boothby 9mo agoCounterpoint: if somebody builds that elsewhere, that's one fewer person posting slop on HN proper
- Imustaskforhelp 9mo agohttps://chatgpt.com/share/695d9ac2-c314-8011-8938-b0d7de70592a https://chatgpt.com/share/695d9ac2-c314-8011-8938-b0d7de7059... You can paste any article and chatgpt (took the most laymen AI thing) and just writing summarize this article https://byteshape.com/blogs/Qwen3-30B-A3B-Instruct-2507/ https://byteshape.com/blogs/Qwen3-30B-A3B-Instruct-2507/ can give you insights about it. Although I am all for freedom, one forgets that this is one of the few places left on internet where discussions feel meaningful and I am not judging you if you want AI but do it at your own discretion using chatbots. If you want, you can even hack around a simple extension (tampermonkey etc.) where you can have a button which can do this for you if you really so desire. Ended up being bored and asked chatgpt to do this but chatgpt is having something wrong, it got just blinking mode so I asked claude web (4.5 sonnet) to do it and I ended up building it with tampermonkey script. Created the code. https://github.com/SerJaimeLannister/tampermonkey-hn-summarize-ai/wiki/ https://github.com/SerJaimeLannister/tampermonkey-hn-summari... I was just writing this comment and I just got curious I guess so in the end ended up building it. Although Edit: Thinking about it, I felt that we should read other people's articles as well. I just created this tool not out of endorsement of idea or anything but just curiosity or boredom but I think that we should probably read the articles themselves instead of asking chatgpt or LLM's about it. There is this quote which I remembered right now If something is worth talking/discussing about, its worth writing If something is worth writing, then its worth reading. Information that we write is fundamentally subjective (our writing style etc with our biases etc.), passing it through a black box which will try to homogenify all of it just feels like it misses the point.
- 6510 9mo ago<s>I'm not entirely sure but I think</s> if the file name ends with .user.js like HN%20ChatGPT%20Summarize.user.js it will prompt to install when opening the raw file. haha, like so works too https://raw.githubusercontent.com/SerJaimeLannister/tampermonkey-hn-summarize-ai/refs/heads/main/HN%20ChatGPT%20Summarize.js#.user.js https://raw.githubusercontent.com/SerJaimeLannister/tampermo...
- Imustaskforhelp 9mo agoAlright so I did change the name of the file from HN ChatGPT Summarize.js to hn-summarize-ai.user.js Is this what you are talking about? If you need any cooperation from my side lemme know, I don't know too much about tampermonkey but I end up using it for my mini scripts because its way much easier to deal with compared to building pure extensions themselves and these have their own editors as well so I just copy paste for a faster way to prototype with stuff like this
- Alex2037 9mo ago>we should read other people's articles sure, and reading a LLM summary allows one to decide whether the full article is worth reading or not.
- jake_wls 9mo ago[flagged]
- Imustaskforhelp 9mo agoFair I guess, I think I myself might use it when sometimes the articles are more dense than my liking perhaps. As I said, I just built it out of curiosity but also a solution to their problem because I didnt like the idea of having an AI generated summary in the comments.
- bigyabai 9mo agoSeems like a bad habit for media literacy.
- kadoban 9mo agoI mean, they didn't bury it far in the article, it's like a two second skim into it and it's labelled with a tl;dr. Not a bad idea in general but you don't even need it for this one.
- Aurornis 9mo agoIf you read a lot of comment sections, there are bot accounts showing up on LLM that try to do this constantly. Their output is not great so they get downvoted and spotted quickly.
- jacquesm 9mo agoIf you spot any that live longer than a few comments please pass that info to Dan & Tom.
- grosswait 9mo agoNot sure if it is still updating https://hackyournews.com/ https://hackyournews.com/
- ukuina 9mo agoThanks for pointing this out, https://hackyournews.com https://hackyournews.com should be up and running again!
- Latitude7973 9mo agoIs this your project? It would be great to bolster it with links to comment sections and the current points tally.
- ukuina 8mo agoYes! And yes, there's a link to the Comment section if you click on the "Comments" summary header. Up-to-date comment tallies are hard, since the summaries are only updated a few times a day.
- make3 8mo agoI wonder what "94.18% of quality" means
- kristianp 8mo agoAlso this is the model name: Qwen3-30B-A3B-Instruct-2507 I tried the q4 quantization when it came out and didn't find it to be great for my coding use case.
- Havoc 9mo agoIs there something different about that accuracy measure? i.e. relative to perplexity term usually used Going from BF16 to 2.8 and losing only ~5% sounds odd to me.
- kouteiheika 9mo agoIt's accuracy across GSM8K, MMLU, IFEVAL and LiveCodeBench. They detail their methodology here: https://byteshape.com/blogs/Qwen3-4B-I-2507/ https://byteshape.com/blogs/Qwen3-4B-I-2507/
- jmward01 9mo agoThere is a huge market segment waiting here. At least I think there is. Well, at least people like me want this. Ok, tens of dollars can be made at least. It is just missing a critical tipping point. Basically, I want an alexa like device for the home backed by local inference and storage with some standardized components identified: - the interactive devices - all the alexa/google/apple devices out there are this interface, also, probably some TV input that stays local and I can voice control. That kind of thing. It should have a good speaker and voice control. It probably should also do other things like act as a wifi range extender or be the router. That would actually be good. I would buy one for each room so no need for crazy antennas if they are close and can create true mesh network for me. But I digress. - the home 'cloud' server that is storage and control. This is a cheap CPU, a little ram and potentially a lot of storage. It should hold the 'apps' for my home and be the one place I can back-up everything about my network (including the network config!) - the inference engines. That is where this kind of repo/device combo comes in. I buy it and it knows how to advertise in a standard way its services and the controlling node connects it to the home devices. It would be great to just plug it in and go. Of course all of these could be combined but conceptually I want to be able to swap and mix and match at these levels so options here and interoperability is what really matters. I know a lot of (all of) these pieces exist, but they don't work well together. There isn't a simple standard 'buy this turn it on and pair with your local network' kind of plug and play environment. My core requirements are really privacy and that it starts taking over the unitaskers/plays well together with other things. There is a reason I am buying all this local stuff. If you phone home/require me to set up an account with you I probably don't want to buy your product. I want to be able to say 'Freddy, set timer for 10 mins' or 'Freddy, what is the number one tourist attraction in South Dakota' (wall drugs if you were wondering)
- protocolture 9mo agoKeen for this also. Been having issues getting a smooth voice experience from HA to ChatGPT. I dont like the whole wakeword concept for the receiver either. I think theres work to be done on the whole stack.
- 6510 9mo ago
- geerlingguy 9mo agoI've just tried replicating this on my Pi 5 16GB, running the latest llama.cpp... and it segfaults: ./build/bin/llama-cli -m "models/Qwen3-30B-A3B-Instruct-2507-Q3_K_S-2.70bpw.gguf" -e --no-mmap -t 4 ... Loading model... -ggml_aligned_malloc: insufficient memory (attempted to allocate 24576.00 MB) ggml_backend_cpu_buffer_type_alloc_buffer: failed to allocate buffer of size 25769803776 alloc_tensor_range: failed to allocate CPU buffer of size 25769803776 llama_init_from_model: failed to initialize the context: failed to allocate buffer for kv cache Segmentation fault I'm not sure how they're running it... any kind of guide for replicating their results? It does take up a little over 10 GB of RAM (watching with btop) before it segfaults and quits. [Edit: had to add -c 4096 to cut down the context size, now it loads]
- batch12 9mo agoCould they have added some swap?
- geerlingguy 9mo agoNo, just updated the parent comment, I added -c 4096 to cut down the context size, and now the model loads. I'm able to get 6-7 tokens/sec generation with 10-11 tokens/sec prompt processing with their model. Seems quite good, actually—much more useful than llama 3.2:3b, which has comparable performance on this Pi.
- layoric 9mo agoThanks for posting the performance numbers from your own validation. 6-7 tokens/sec is quite remarkable for the hardware.
- geerlingguy 9mo agoSome more benchmarking, and with larger outputs (like writing an entire relatively complex TODO list app) it seems to go down to 4-6 tokens/s. Still impressive.
- lostmsu 9mo agoGPT-OSS-20B is only 11.2GB. Should fit in any 16GB machine with descent context without any quality degradation.
- jareds 9mo agoIs there a good place for easy comparisons of different models? I know gpt-oss-20b and gpt-oss-120b have different numbers of parameters, but don't know what this means in practice. All my experience with AI has been with larger models like Gemini and GPT. I'm interested in running models on my own hardware but don't know how small I can go and still get useful output both for simple things like fixing spelling and grammar, as well as complex things like programming.
- jdright 9mo agohttps://swe-rebench.com/ https://swe-rebench.com/
- ekidd 9mo agoOne easy way to test different models is purchase $20 worth of tokens from one of the Open Router-like sites. This will let you asks tons of questions and try out lots of models. Realistically, the biggest models you can run at a reasonable price right now are quantized versions of things like the Qwen3 30B A3B family. A 4-bit quantized version fits in roughly 15GB of RAM. This will run very nicely on something like an Nvidia 3090. But you can also use your regular RAM (though it will be slower). These models aren't competitive with GPT 5 or Opus 4.5! But they're mostly all noticeably better than GPT-4o, some by quite a bit. Some of the 30B models will run as basic agentic coders. There are also some great 4B to 8B models from various organizations that will fit on smaller systems. A 8B model, for example, can be a great translator. (If you have a bunch of money and patience, you can also run something like GPT OSS 120B or GLM 4.5 Air locally.)
- cmrdporcupine 9mo agoThis is the answer. There's a half dozen sites that let you run these models by the token, and actually $20 is excessive. $5 will get you a long long way.
- kouteiheika 9mo ago> (If you have a bunch of money and patience, you can also run something like GPT OSS 120B or GLM 4.5 Air locally.) Don't need patience for these, just money. A single RTX 6000 Pro runs those great and super fast.
- anonzzzies 9mo agoWe need custom inference chips at scale for this imho. Every computer (whatever formfactor/board) should have an inference unit on it so at least inference is efficient and fast and can be offloaded while the cpu is doing something else.
- fouc 9mo agoI can't believe this was downvoted. It makes a lot of sense that it would be highly useful to have mass custom inference chips.
- bigyabai 8mo agoIt's quite easy to understand. The tech industry has gone through 4-5 generations of obsolete NPU hardware that was dead-on-arrival. Meanwhile, there are still GPUs from 2014-2016 that run CUDA and are more power efficient than the NPUs. The industry has to copy CUDA, or give up and focus on raster. ASIC solutions are a snipe chase, not to mention small and slow.
- chvid 9mo agoLook at the specs of this Orange Pi 6+ board - dedicated 30 TPU NPU. https://boilingsteam.com/orange-pi-6-plus-review/ https://boilingsteam.com/orange-pi-6-plus-review/
- baq 9mo agoAt this point of the timeline compute is cheap, it’s RAM which is basically unavailable.
- Aurornis 9mo agoThe bottleneck in common PC hardware is mostly memory bandwidth. Offloading the computation part to a different chip wouldn’t help if memory access is the bottleneck. There have been a lot of boards and chips for years with dedicated compute hardware, but they’re only so useful for these LLM models that require huge memory bandwidth.
- 9mo ago
- syntaxing 9mo agoI feel like calling it a “30B” model is slightly disingenuous. It’s a 30B-A3B. So only 3B parameters is active at a given time. While still impressive nevertheless, being able to get 8T/s for a “A3B” compared to a dense 30B is very different.
- throwaway894345 9mo agoWhat does it mean that only 3B parameters are active at a time? Also any indication of whether this was purely CPU or if it’s using the Pi’s GPU?
- kouteiheika 9mo ago> What does it mean that only 3B parameters are active at a time? In a nutshell: LLMs generate tokens one at a time. "only 3B parameters active a a time" means that for each of those tokens only 3B parameters need to be fetched from memory, instead of all of them (30B).
- tgv 9mo agoThen I don't understand why it would matter. Or does it really mean that for each input token 10% of the total network runs, and then another 10% for the next token, rather than running each 10 batches of 10% for each token? If so, any idea or pointer to how the selection works?
- kouteiheika 9mo agoYes, for each token only, say, 10% of the weights are necessary, so you don't have to fetch the remaining 90% from memory, which makes inference much faster (if you're memory bound; if you're doing single batch inference then you're certainly memory bound). As to how the selection works - each mixture-of-experts layer in the netwosk has essentially a small subnetwork called a "router" which looks at the input and calculates the scores for each expert; then the best scoring experts are picked and the inputs are only routed to them.
- 9mo ago
- tgtweak 9mo agoSo basically the quantization in a byteshape model is per-tensor and can be variable and is an "average" in the final result? The results look good - curious why this isn't more prevalent! Would also love to better understand what factors into "accuracy" since there might be some nuance there depending on the measure.
- kouteiheika 9mo ago> Would also love to better understand what factors into "accuracy" since there might be some nuance there depending on the measure. It's accuracy across GSM8K, MMLU, IFEVAL and LiveCodeBench. They detail their methodology here: https://byteshape.com/blogs/Qwen3-4B-I-2507/ https://byteshape.com/blogs/Qwen3-4B-I-2507/
- TheRealPomax 9mo agoLLMs are, by definition, real time at any speed. 50,000 tokens per second? Real time. Only 0.0002 tokens per minute? Still real time. Eight tokens per second is "real time" in that sense, but that's also the kind of speeds that we used to mock old video games for, when they would show "computers" but the text would slowly get printed to a screen letter for letter or word for word.
- kouteiheika 9mo agoIn this context by "real time" people usually mean "as fast as I can read the reply", so, 0.0002 tokens per minute would not be considered "real time".
- rurban 9mo agoReal time typically means guaranteed reaction time below 30ms, because slower reactions will make the body through up.
- baq 9mo agoReal time is defined as ‘no slower than some critical speed’, in case of conversation with humans this should be around 10 tok/s including speech synthesis.
- runlaszlorun 9mo agoI'm just pulling stuff out of my butt here but is this an area where an fpga might be worth it for a particular type of model?
- syx 9mo agoI can’t wait to get home and try this on my Pi. Past few months, I’ve been building a fully local agent [0] that runs inference entirely on a Raspberry Pi, and I’ve been extensively testing a plethora of small, open models as part of my studies. This is an incredibly exciting field, and I hope it gains more attention as we shift away from massive, centralized AI platforms and toward improving the performance of local models. For anyone interested in a comparative review of different models that can run on a Pi, here’s a great article [1] I came across while working on my project. [0] https://github.com/syxanash/maxheadbox https://github.com/syxanash/maxheadbox [1] https://www.stratosphereips.org/blog/2025/6/5/how-well-do-llms-perform-on-a-raspberry-pi-5 https://www.stratosphereips.org/blog/2025/6/5/how-well-do-ll...
- cess11 9mo agoWe're approaching the point where we could have sophisticated sound controlled sex toys.
- oliwary 9mo agoDoes anyone have any use-cases for long-running, interesting tasks that do not require perfect accuracy? Seems like this would be the sweet spot for running local models on low-powered hardware.
- baq 9mo agolooking at the dump of your home assistant server and sensor data and saying 'hmm that's interesting' if it notices something interesting
- MORPHOICES 9mo ago[dead]
- nl 9mo agoI've been super impressed by qwen3:0.6b (yes, 0.6B) running in Ollama. If you have very specific, constrained tasks it can do quite a lot. It's not perfect though. https://tools.nicklothian.com/llm_comparator.html?gist=fcae9ffde8adeda3c2bb4adfd1b92456 https://tools.nicklothian.com/llm_comparator.html?gist=fcae9... is an example conversation where I took OpenAI's "Natural language to SQL" prompt[1], send it to Ollama:qwen3:0.6b and the asked Gemini Flash 3 to compare what qwen3:0.6b did vs what Flash did. Flash was clearly correct, but the qwen3:0.6b errors are interesting in themselves. [1] https://platform.openai.com/docs/examples/default-sql-translate https://platform.openai.com/docs/examples/default-sql-transl...
- Aurornis 9mo agoI’ve experimented with several of the really small models. It’s impressive that they can produce anything at all, but in my experience the output is basically useless for anything of value.
- nl 9mo agoYes, I thought that too! But qwen3:0.6b (and to some extent gemma 1b) has made me reevaluate. They still aren't useful like large LLMs, but for things like summarization, and other tasks where you can give them structure but want the sheen of natural language they are much better than things like the Phi series were.
- redman25 9mo agoThat's interesting. For what projects would you want the "sheen of natural language" though?
- nl 8mo agoSay I want to auto-bookmark a bunch of tabs and need a summary of each one. Using the title is a mechanical solution, but a nice prompt and a small model can summarize the title and contents into something much more useful.
- shnpln 9mo agoHow would this run on a Jetson Orin Nano dev kit?
- zoe_mnode 9mo agoThis is impressive work on quantization. It really validates the hypothesis that software optimization (and proper memory management) matters more than just raw FLOPs.
- cwoolfe 8mo agoHow can I use ByteShape to run LLMs faster on my 32GB MacBook M1 Max? Or has Ollama already optimized that?
- nunodonato 8mo agodon't use ollama. llama.cpp is better because ollama has an outdated llama.cpp
- nunodonato 8mo agoHad to try this on my laptop (vs my non-byteshape qwen3-30-a3b) Original: 11tok/s Byteshape: 16tok/s Quite a nice improvement!