Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
roadside_picnic
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
1.
▲
by
roadside_picnic
8d ago
I've never understood why "Hacker" News so frequently gets "But why though?" comments at the top. The entire history of innovation is filled with people doing something just to see they can get it to work, even if b
2.
▲
by
roadside_picnic
1mo ago
> and I would not call it cheap It is compared to the cost of making an authentic ramen broth at home at the scale of feeding 2-4 people. Pho has a similar property. I've made homemade pho broth once and it was very hard work, fairl
3.
▲
by
roadside_picnic
2mo ago
> We do this and then no one reads it. This to me is an anti-pattern that I see everywhere because we're still in the early stages of understanding how to adopt these things and create real agentic workflows. Why should humans rea
4.
▲
by
roadside_picnic
2mo ago
MoEs don't route like most people imagine. They aren't learning topic based experts despite the name The original Mixtral paper [0] (in the "Routing analysis" section) found: "surprisingly, we do not observe obv
5.
▲
by
roadside_picnic
2mo ago
> from 40 to 50 tok/s generation That's actually much better than I would have thought! Thanks for the answer and it does make this approach make more sense as a budget solution to running larger models locally.
6.
▲
by
roadside_picnic
2mo ago
What's the tokens/sec you're getting on that setup (genuinely curious because it's a setup I haven't actually run myself)?
7.
▲
by
roadside_picnic
2mo ago
> no other widely accessible path for young people to build a future This is exactly right. Widespread gambling has a fundamental nihilism baked into it. If you believe hard work and frugality are the best path to success, then gambling
8.
▲
by
roadside_picnic
2mo ago
If you're using a MoE model, then why do you care about the larger RAM offered by these devices? That's the main problem with low bandwidth devices: they limit the effective ram you can make use. I do (and have historically done)
9.
▲
by
roadside_picnic
3mo ago
> Yet those 2-3M channels get the lions share of the views. ...because they're directly tied to YouTube own revenue. As someone who creates a fair bit of content, and am fortunate enough not to have to worry about the revenue from t
10.
▲
Making LLMs Better at Creative Writing Using Entropy
(countbayesie.com)
2 points
by
roadside_picnic
3mo ago
|
1 comments
11.
▲
by
roadside_picnic
3mo ago
In general if you're setting up a local LLM you should assume it's going to be primarily working as a server and talking to various clients. I use my MBP, but that's because I don't travel much anymore so it can happily
12.
▲
by
roadside_picnic
3mo ago
It depends on your use case. There's a lot of hype around machines like the DGX spark (I'm assuming this is the type of device you're referring to) because they look awesome, and are priced reasonably well. However all of the
13.
▲
by
roadside_picnic
3mo ago
My experience working in the open model space pretty deeply (both LLMs and diffusion models) for years now is that it is not quite as simple as that. In the open model space an insane amount of effort goes into getting more powerful models
14.
▲
by
roadside_picnic
3mo ago
In addition to models getting better, the quantization methods have also got much better. If you already have an RTX 3080 it's absolutely worth the time to just mess around and see how it does, experiment with different quants that fit
15.
▲
by
roadside_picnic
3mo ago
M3-Max laptop: ~55 token/sec RTX 4090: ~190 token/sec I don't have the number around but there is a notable latency for pre-fill on the M3, but once it's running the delay is negligible. The RTX, unsurprisingly, is all a
16.
▲
by
roadside_picnic
3mo ago
There are a couple of things, but basically it boils down to the same reason people prefer Linux to Windows/MacOs: customization, control and privacy (arguably all of these are really subsets of 'control'). Having full contro
17.
▲
by
roadside_picnic
3mo ago
Just running it through `llama-cli` so that there's absolutely no persistent state related to the chat (and least I believe this to be the case).
18.
▲
by
roadside_picnic
3mo ago
See my comment to parent. I've been using local LLMs for practical, personal tasks for a few months now very successfuly. You can run fantastic local models if you have either: - M-series Apple device with ideally >= 24GB of VRAM -
19.
▲
by
roadside_picnic
3mo ago
I have a home server that runs Qwen3.6-35B-A3B through llama.cpp with Open WebUI for the user facing interface. My teen isn't super interested in AI, but whenever they do feel curious they have their own account they can use on our hom
20.
▲
by
roadside_picnic
3mo ago
> It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. In the monkey example the infinite time is doing a lot of work there. The fact that LLMs can search through sema
21.
▲
by
roadside_picnic
3mo ago
I used to think comparisons of AI with the web where ridiculous, but increasingly it looks like they're not that dissimilar as far as how they change how we work. But as someone who graduated college during the bust there was a loooong
22.
▲
by
roadside_picnic
4mo ago
It's more insidious than that. These IPOs aren't being rushed, they were waiting for all the pieces to be in place to force 401ks and other retirement plans to buy these IPOs. The most recent change was the NASDAQ adopting the &qu
23.
▲
by
roadside_picnic
4mo ago
As you likely know, rules have recently been changed that basically force many 401k funds to invest in these IPOs while simultaneously having a relatively small number of the initial IPO to be sold to the public forcing the funds to by at i
24.
▲
by
roadside_picnic
4mo ago
Have you personally used any of the latest batch of even smaller local models? They certainly don't beat SotA models at coding... but with a good harness they are able to achieve things with SotA that I couldn't last year. I
25.
▲
by
roadside_picnic
4mo ago
It also underplays what I've personally witnessed that I would consider true AI psychosis. I worked with someone who sincerely believed he was spiritually co-evolving with his army of sycophantic AI agents (the agents would be tasked w
26.
▲
by
roadside_picnic
4mo ago
There's a difference between "red flags" and "imperfections". Every team has faults, which if you're experienced at interviewing/working many places, are usually pretty easy to figure out. These are distin
27.
▲
by
roadside_picnic
4mo ago
These companies are unprofitable (as all companies at this stage and ambition should be) but I increasingly don't see any justification for the idea that it is fundamentally unprofitable. Inference alone is certainly profitable. I&#x
28.
▲
by
roadside_picnic
5mo ago
This is a classic example of people misapplying the logic of the SaaS world to the AI world. If you're building software to sell, you're in trouble. The people that are finding success in this space are using AI to allow them to
29.
▲
by
roadside_picnic
6mo ago
> Users don’t care about “privacy”. I worked for a research focused AI startup that had a strict "no external LLM" policy for code touching our core research. You're right that the average consumer doesn't care abou
30.
▲
by
roadside_picnic
6mo ago
> in lisp. Technically you cannot implement a proper Y-combinator in Lisp (well, I'm sure in Common Lisp and Racket there is some way) because the classic Y-combinator relies on lazy , not strict, evaluation. Most of the "Y-
More ›