Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mechagodzilla
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
mechagodzilla
1mo ago
Yes. I can split a model like this across 3 GPUs (a 1080 with 8GB and two Titan Vs with 12GB), and it's much faster than running it on 36 CPU cores. As long as it fits in aggregate VRAM, it seems very advantageous to do so.
2.
▲
by
mechagodzilla
4mo ago
"Household Debt Service Payments as a Percent of Disposable Personal Income": https://fred.stlouisfed.org/series/TDSP
3.
▲
by
mechagodzilla
4mo ago
Relatively accurate (I think they're pretty close like 70%+ of the time). Significantly better than just projecting that the next five years will be like the last 5 years.
4.
▲
by
mechagodzilla
9mo ago
You haven't really been getting 'human written and thoughtful content' for a vast swath of search topics for probably 15-20 years now. You get SEO-hyper-optimized (probably LLM-generated for anything in the last 3 years) blog
5.
▲
by
mechagodzilla
9mo ago
It's not really an apples-to-apples comparison - I enjoy playing around with LLMs, running different models, etc, and I place a relatively high premium on privacy. The computer itself was $2k about two years ago (and my employer reimbu
6.
▲
by
mechagodzilla
9mo ago
I've been running the 'frontier' open-weight LLMs (mainly deepseek r1/v3) at home, and I find that they're best for asynchronous interactions. Give it a prompt and come back in 30-45 minutes to read the response. I&
7.
▲
by
mechagodzilla
9mo ago
1-2 tokens/sec is perfectly fine for 'asynchronous' queries, and the open-weight models are pretty close to frontier-quality (maybe a few months behind?). I frequently use it for a variety of research topics, doing feasibilit
8.
▲
by
mechagodzilla
9mo ago
You can keep scaling down! I spent $2k on an old dual-socket xeon workstation with 768GB of RAM - I can run Deepseek-R1 at ~1-2 tokens/sec.
9.
▲
by
mechagodzilla
10mo ago
Kodak didn't really have the option to compete. Their business was largely film, which just disappeared completely, and even digital cameras got replaced pretty quickly with phones. There was nothing to pivot too for Kodak.
10.
▲
by
mechagodzilla
10mo ago
Did you have to do anything special to get the SSD to play nice with OS9? I tried adding one to a 300MHz G3 iMac and it took forever to initialize on boot and would randomly stall a lot.
11.
▲
by
mechagodzilla
1y ago
If anyone can pay-as-you-go use a fully automated factory, and the factories are interchangeable, it seems like the value of capital is nearly zero in your envisioned future. Anyone with an idea for soup can start producing it with world cl
12.
▲
by
mechagodzilla
1y ago
Buy a used workstation with 512GB of DDR4 RAM. It will probably cost like $1-1.5k, and be able to run a Q4 version of the full deepseek 671B models. I have a similar setup with dual-socket 18 core Xeons (and 768GB of RAM, so it cost about $
13.
▲
by
mechagodzilla
1y ago
Yeah, it was just a giant HP workstation - I currently have 3 graphics cards in it (but only 40GB total of VRAM, so not very useful for deepseek models).
14.
▲
by
mechagodzilla
1y ago
I use a dual-socket 18-core (so 36 total) xeon with 768GB of DDR4, and get about 1.5-2 tokens/sec with a 4-bit quantized version of the full deepseek models. It really is wild to be able to run a model like that at home.
15.
▲
by
mechagodzilla
1y ago
Interns and new grads have always been a net-negative productivity-wise in my experience, it's just that eventually (after a small number of months/years) they turn into extremely productive more-senior employees. And interns and
16.
▲
by
mechagodzilla
1y ago
I have a $2k used dual-socket xeon with 768GB of DDR4 - It runs at about 1.5 tokens/sec for the 4-bit quantized version.
17.
▲
by
mechagodzilla
1y ago
I've done all of those except tend livestock and build a house, but I could probably figure those out with some effort.
18.
▲
by
mechagodzilla
1y ago
So applying this to China and the USA - the USA has a median household income of ~$80k, and China, in terms of purchasing power parity, has a median household income of $32k. China has a workforce of ~775M people vs 163M people in the USA.
19.
▲
by
mechagodzilla
1y ago
Ha! When I was first learning to program in high school, I wrote a 'distributed monkeys-on-typewriters' simulator. I somehow acquired a stack of surplus Pentium 100s that I had running in an unused closet at the school, communicat
20.
▲
by
mechagodzilla
1y ago
I think we can already get open-weight frontier class models today. I've run Deepseek R1 at home, and it's every bit as good as any of the ChatGPT models I can use at work.
21.
▲
by
mechagodzilla
1y ago
That's been the 'endgame' of technology improvements since the industrial revolution - there are many industries that mechanized, replaced nearly their entire human workforce, and were never terribly profitable. Consider farm
22.
▲
by
mechagodzilla
1y ago
None of that means that the current companies will be profitable or that their valuations are anywhere close to justified though. The future could easily be "Open-weight models are moderately useful for some niches, no-name cloud provi
23.
▲
by
mechagodzilla
1y ago
After being a professional programmer for ~20 years, and recently playing around with leetcode - my main issue with leetcode is that there's almost no overlap between leetcode problems and the problems I actually encounter in the wild.
24.
▲
by
mechagodzilla
1y ago
I married at 22 and moved to nyc. I lived through meeting peers that thought we were freaks (“you two are definitely gonna get divorced”), to the gradual normalization as we got older and more of our friends married. I knew lots of people t
25.
▲
by
mechagodzilla
2y ago
But the open models create a (rapidly rising!) 'floor' - models that are worse/less capable than the best open-weight models effectively have zero economic value to their creators, even if they just spent $10B to create them.
26.
▲
by
mechagodzilla
2y ago
And the Llama series, and they have half a dozen closed competitors with similar performance/capability. It really seems impossible to justify these evaluations from any kind of economic perspective.
27.
▲
by
mechagodzilla
2y ago
I think that's the right interpretation, but that's pretty weak for a company that's nominally worth $150B but is currently bleeding money at a crazy clip. "We spent years and billions of dollars to come up with somethin
28.
▲
by
mechagodzilla
2y ago
If you don't need to pay for the model development costs, I think running inference will just be driven down to the underlying cloud computing costs. The actual requirement to passably (~4-bit quantization) run Deepseek v3/r1 at h
29.
▲
by
mechagodzilla
2y ago
Or, like Meta, they make their money elsewhere and just seem interested in wrecking the economics of LLMs. As soon as an open-weight model is released, it basically sets a global floor that says "Models with similar or worse performanc
30.
▲
by
mechagodzilla
2y ago
It does seem like it will be very, very hard for the companies training their own models to recoup their investment when the capabilities of open-weight models catch up so quickly - general purpose LLMs just seem destined to be a cheap comm
More ›