Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Chamix
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
Chamix
1mo ago
Naturally, he has a whole detailed post/trace about it on gwern net! https://gwern.net/twitter An interesting case of echo chamber formation in that its pragmatic to be scared of overtly critiquing him on twitter lest
2.
▲
by
Chamix
1mo ago
Ha, haha, across hundreds of personal discussions I've been involved with on lesswrong/lighthaven, twitter, Wikipedia talk/editing, SF parties etc I think his most distinguishing feature has always been his abiding and unabas
3.
▲
by
Chamix
4mo ago
I note that (though summarized), this is ~100k tokens. Anyone who routinely works with Codex (or any agentic harness really) can tell you how trivial it is to eat up 100k tokens doing complex work. I've personally had plenty of codex 5
4.
▲
by
Chamix
6mo ago
I appreciate the detailed comment! I took the day off and am bored so have a brain dump of a reply - basically I think we are talking past each other on two major points: 1. All the discussion about model size is CRITICALLY bisected into ta
5.
▲
by
Chamix
6mo ago
Sorry if that was unclear, I did mean 100Bs as in the next order of magnitude. Even GPT-4 had ~220B active params, though the trend has been towards increased sparsification (lower activation:total ratio). GPT 4.5 is the only publicly facin
6.
▲
by
Chamix
6mo ago
What do you think labs are doing with the minimum 10TB memory in NvLink 72 systems that were publicly reported to all start coming online in November/December of last year? And why would this 1 TB -> 10 TB jump matter so much for An
7.
▲
by
Chamix
6mo ago
I assure you, the number of people paying to use Qwen3-Max or other similar proprietary endpoints is far less than 1.6 billion.
8.
▲
by
Chamix
6mo ago
I generally agree, back of the napkin math shows H20 cluster of 8gpu * 96gb = 768gb = 768B parameters on FP8 (no NVFP4 on Hopper), which lines up pretty nicely with the sizes of recent open source Chinese models. However, I'd say its r
9.
▲
by
Chamix
6mo ago
Try 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gpu rubin clusters (and equivalent-ish world size for those re
10.
▲
by
Chamix
7mo ago
You, know, it sure does add some additional perspective to the original Anthropic marketing materia... ahem, I mean article, to learn that the CCC compiled runtime for SQLite could potentially run up to 158,000 times slower than a GCC compi
11.
▲
by
Chamix
2y ago
Indeed, and the difference could in essence be achieved yourself with a different system prompt on 4o. What exactly is 4.5 contributing here in terms of a more nuanced intelligence? The new RLHF direction (heavily amplified through scaling
12.
▲
by
Chamix
2y ago
It's interesting to compare the cost of that original gpt-4 32k(0314) vs gpt-4.5: $60/M input tokens vs $75/M input tokens $120/M output tokens vs $150/M output tokens
13.
▲
by
Chamix
2y ago
Forgive me if I'm missing your existing realization (I did a quick check of your HN, reddit, twitter, LW), but I think the big deal with Sohu (wrt Etched) is that they have pivoted from the "all model parameters hard etched onto t
14.
▲
by
Chamix
2y ago
I was thinking about the llm writing tool from Janus.
15.
▲
by
Chamix
2y ago
4chan already has a torrent out, of course.
16.
▲
by
Chamix
3y ago
The little secret is that the training run (meaning, creating the raw autocompleting multimodal token weights) for 5 ran in parallel with 4.
17.
▲
by
Chamix
3y ago
Luckily Eliezer has written hundreds of approachable essays on the development of his epistemic processes over at lesswrong.com so you too can learn rationality and derive the killeveryonism conclusion yourself. (/s since this is the
18.
▲
by
Chamix
3y ago
Fair enough, shame "Large Tokenized Models" etc never entered the nomenclature.
19.
▲
by
Chamix
3y ago
You are conflating Illya's belief in the transformer architecture (with tweaks/compute optimizations) being sufficient for AGI with that of LLMs being sufficient to express human-like intelligence. Multi-modality (and the swath
20.
▲
by
Chamix
3y ago
The issue, as pointed above, is primarily bandwidth (at inference), not addressable memory. Put simply, the best bandwidth stack we currently have is on-package HBM -> NVLink, -> Mellanox InfiniBand, and for inference speed you really
21.
▲
by
Chamix
3y ago
Agriculture > Manufacturing -> Service -> Content based economy, turns out youth have the head start, as always.
22.
▲
by
Chamix
3y ago
Used it for weeks internally, finally gave up after deciding it I more than once felt the day's use was a net productive loss even when working with internal Amazon packages (that it should have had a training advantage on vs copilot).
23.
▲
by
Chamix
3y ago
Spoiler: He did not say that.
24.
▲
by
Chamix
4y ago
Won nearly every single coding jam (8/9) he was in, highest code force rating of all time, 6x consecutive IOI gold, yea he’s the cut and dry winner.
25.
▲
by
Chamix
4y ago
Tech survey results, ops load via CTI tickets, CR/LOC stats, accolades, COEs, promotion rate, turnover/attrition.... There are so many data sources that are all queryable via API. There are a couple basic greasemonkey scripts alre
26.
▲
by
Chamix
4y ago
Oh my friend, there is a tool called Lift used in OLR/Talent Review by every single 50+ group at Amazon that quite literally will show your face up on a screen for everyone at the OLR meeting to discuss and rank vs your peers. If they
27.
▲
by
Chamix
4y ago
For fun, there are currently precisely 312 employees under DynamoDB directors/VPs that are categorized at "Software Development". Including systems engineers ups it to 348.
28.
▲
by
Chamix
4y ago
The news here is that amazon is going to start the 2 month+ WARN period for a much larger set of orgs this week. Amazon Care/Health Service layoffs are old news.
29.
▲
by
Chamix
4y ago
Unfortunately development has been paused for over two years now. But yes, outside of Arpia + some other total conversions for older EVs, this is really the only thing out there that directly scratches the EV itch. The closest contenders IM
30.
▲
by
Chamix
5y ago
Instead of getting good at 1 big tech job making 300k in 20-30 hours a week you get good at 2-3 and make 800k+ in 60 hours a week. Though I think the actual best way to optimize labor-to-income is F500 Cloud SRE via B2B contract, but its ha
More ›