Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joefourier
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
joefourier
6d ago
The app not existing would unironically be better in so many instances, though. The web version doesn’t take up 1GB of disk space, install persistent services, send you push notifications by default, and it’s trivial to block ads in compari
2.
▲
by
joefourier
20d ago
> It's especially palpable just looking at the last few years of LLM's: a frontier model with all the world knowledge you can stuff in it and every tool at its disposal has always performed the best at all tasks. Suggesting oth
3.
▲
by
joefourier
1mo ago
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
4.
▲
by
joefourier
2mo ago
You cannot just "try all possible optimisations". It takes time, effort, and money that could otherwise be spent elsewhere (especially for training, where each training run is especially costly, and optimisations might be promisin
5.
▲
by
joefourier
2mo ago
Anthropic's API has two nines availability and Claude Code is a TUI made with React that can regularly consume more than 1GB of RAM, and the codebase is utter slop. They couldn't fix the flickering bug for over a year! And yet,
6.
▲
by
joefourier
2mo ago
> Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? I'm not sure
7.
▲
by
joefourier
3mo ago
I stopped using Cursor because of how terribly optimised it is (worse than VSCode despite being a fork). It would routinely take up 50% of the CPU resources on my MacBook M4 and gigabytes of RAM for absolutely no reason. I switched to Zed,
8.
▲
by
joefourier
3mo ago
Why is it unethical? I'm both a freelance engineer and a business owner that sells software, and I've both sold my labour for equity/revenue share, and for a flat hourly rate. If I charge a client $50k for some software and t
9.
▲
by
joefourier
3mo ago
> My response is “great! Let businesses take a lesson here: give all your employees a chunk of the company. Let’s all share in the success!” Don't >95% of tech companies offer stock options or equity, from startups to FAANG?
10.
▲
by
joefourier
3mo ago
Yeah most of the performance increases have mostly been from architectural improvements like reduced precision tensor cores. AFAIK FP4 is basically the limit for floating point matmuls, after which you need to switch to integer addition if
11.
▲
by
joefourier
3mo ago
Demand is so high and supply so low customers will go to anyone that has any gear, period. Anthropic is paying xAI for GPUs from 2022, not the latest Nvidia release.
12.
▲
by
joefourier
3mo ago
And yet Anthropic is paying xAI over a billion dollars a month for those out of date GPUs in their first datacentre (H100s being nearly 4 years old at this point). Even A100s are still barely available on the major clouds despite being 6 ye
13.
▲
by
joefourier
4mo ago
Do you think the work will still apply to speculative/alternative decoding methods like MTP and block diffusion, which are making batch=1 decoding less memory bound? Kernel launch overhead and memory transfer become less and less signi
14.
▲
by
joefourier
4mo ago
You'd be surprised, people are somehow buying Tesla P40s and M40s on eBay for almost $300 and $180 respectively (M40 being the same gen as GTX 950). Google Colab still offers T4s and it's taken them years to add modern GPUs. Hope
15.
▲
by
joefourier
4mo ago
Outside of training the biggest LLMs at big labs, GPU lifespan isn't as short as the OP made it out to sound. A100s are 6 years old and still a reliable work-horse, and the 80GB version hasn't depreciated that much on the used mar
16.
▲
by
joefourier
4mo ago
> And why is V100 even used? V100 is four generations old and not even supported anymore. It wouldn’t surprise me that due to bureaucratic processes, it’s still somehow the most readily available GPU for Apple researchers despite being a
17.
▲
by
joefourier
4mo ago
Not a single new 64GB GPU, but multiple used GPUs. They’ve significantly increased in price (so much for hardware depreciation…) but you can still get a modded 22GB 2080 ti for $320, or a Mi50 32GB for ~$450 each (used to be $150 a few mont
18.
▲
by
joefourier
4mo ago
What quant? You should have no problem running it at Q4 with 256K context, Q5 or Q6 even although maybe not at full context. I can run Q4 on a 4090 with just 24GB VRAM.
19.
▲
by
joefourier
4mo ago
Who is going to buy a $4299 M5 Max MBP with 64GB of RAM just to run Gemma 4 31b? Firstly you don't need 64GB for that model. Secondly if you want a machine that sits in the corner and does nothing but LLM inference, you don't buy
20.
▲
by
joefourier
4mo ago
Why didn't you take into account batching, input tokens, different costs of electricity, and the fact that a laptop can still hold a decent % of its resale value, and is useful for many other tasks than running an LLM?
21.
▲
by
joefourier
4mo ago
> I'd compare it to OpenAI 5 years ago except I think even then OpenAI had way more! Say what? 5 years ago OpenAI had received around $139 million in funding, and they’d just come out with GPT3 with 175B parameters, a 2048 context
22.
▲
by
joefourier
4mo ago
It's incredibly common all over Europe, not just Switzerland. Not only the metros but the trams and even buses often rely on this system where there's no turnstile or barrier, you just walk in. Not sure it's about being a hig
23.
▲
by
joefourier
4mo ago
Personally I feel like it would be less undignified and infantilising to have a machine take care of my basic bodily functions than a human being. There's no feeling of judgement or being shamed in front of someone else, and the machin
24.
▲
by
joefourier
4mo ago
> I also used dedicated servers in the late ’90s (and they still offer great value today). But before AWS, provisioning new hardware typically took days, not minutes. VPSes and non-custom configs for dedicated servers were pretty instant
25.
▲
by
joefourier
4mo ago
> Cloud computing was an absolutely mind blowing revolution - suddenly your startup could run its own computer systems in minutes without need to install and run your own systems in a data center. This was an absolute game changer, and I
26.
▲
by
joefourier
4mo ago
> In every country, men commit almost all violent crimes. In school, boys physically bully other boys. Hence the physical punishment for them. As I've said, and @echoangle repeated, caning is used for cyberbullying, which girls do t
27.
▲
by
joefourier
4mo ago
Boys and girls being different does not mean one sex deserves corporal punishment and one does not. Girls are equally capable of cyberbullying (which is covered by this law), why should they only get detention while a 9 year old boy has to
28.
▲
by
joefourier
4mo ago
> There he was, a short, beefy guy with a goatee and a Red Sox cap and a thick Boston accent, and I suddenly learned that I didn’t have the slightest idea what to say to someone like him. So alien was his experience to me, so unguessable
29.
▲
by
joefourier
4mo ago
Calling anything "large" in computing is problematic since hardware keeps improving. GPT-1 was an LLM in 2017 and had 117M parameters, when did it stop being large? GPT would have been a better term than LLM, but unfortunately bec
30.
▲
by
joefourier
5mo ago
> Local models sound great until you realize you dont get alot of the features that we implicitly expect from hosted models. Many things would require additional investment into the operations and setup to get to a comparable system. We
More ›