Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
boroboro4
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
boroboro4
6d ago
I think the biggest architectural change here is them doing different compute for prefill & decode, with pretty much architecture from this microsoft research work from 2024 https://arxiv.org/abs/2405.05254 , very e
2.
▲
by
boroboro4
2mo ago
For production loads? No one will run DS without expert parallelism so it’s not important for model to fit on one gpu.
3.
▲
by
boroboro4
2mo ago
The issue is it’s cpu compute which is underutilized in gpu clusters anyway, so practically it’s not really 1/1000.
4.
▲
by
boroboro4
3mo ago
It's very unclear what's special in Rubin to be optimized for inference? I can see disaggregated bit (with having separate prefill and decoding nodes), but what else?
5.
▲
by
boroboro4
3mo ago
What does it mean it's better at nvfp4 training? What's different between training and inference to make this true?
6.
▲
by
boroboro4
3mo ago
It’s crazy something which basically changes the way US government works on a such a deep level is not even on a main page of hackernews.
7.
▲
by
boroboro4
3mo ago
The core insight there is to separate value semantics (no identity) from reference (itself) semantics (nullability). While this particular change can bring very limited amount of improvements it’s still does some - probably smaller to no ob
8.
▲
by
boroboro4
3mo ago
There is obvious proxy to the amount of training data - revenue. And I think anthropic is way ahead of them.
9.
▲
by
boroboro4
4mo ago
DOJ puts an accusation with clarifying text in semi private document? They don’t do this, they do much worse things (and get, rightly, much worse response). This document isn’t great, but comparing it to Trump administration actions isn’t g
10.
▲
by
boroboro4
4mo ago
I somewhat agree that the whole document is sloppy, but I also think this conversation overemphasizes what is wrong with it. It is not as if the document simply accuses the journalist and leaves it there; it actually elaborates on what it m
11.
▲
by
boroboro4
4mo ago
But this particular one does?..
12.
▲
by
boroboro4
4mo ago
This is what great reporting looks like: well-written, transparent, and rigorous. It’s sad to see how hatred toward progressives can distort people’s judgment.
13.
▲
by
boroboro4
7mo ago
It's unclear why the most probable next token given the context "please pick random number" won't be distributed uniformly across all the possible numbers (in the end it's totally possible for LLM return 10 logits o
14.
▲
by
boroboro4
7mo ago
My bad. However billionaires don’t own tiny part of US wealth, more like 5%-10%. And top 1% (and grandparent was talking about rich people) own 1/3 of US wealth.
15.
▲
by
boroboro4
7mo ago
Total US wealth is ~170T so obviously it will be enough to cover federal and state government for a year (and more like 20 years). Even considering obvious issue of wealth going down like crazy in such hypothetical scenario in its ends this
16.
▲
by
boroboro4
7mo ago
While I mostly agree with you, it worth noting modern llms are trained on 10-20-30T of tokens which is quite comparable to their size (especially given how compressible the data is)
17.
▲
by
boroboro4
8mo ago
This https://github.com/niklas-heer/speed-comparison/blob/master/... and this https://github.com/niklas-heer/speed-comparison/blob/master/... from parent comment
18.
▲
by
boroboro4
8mo ago
It's not an issue of warmup time, it's an issue of jit compilation. On my server (AMD EPYC 7252): 1) base time of the java program from the repo is 3.23s (which is ~2 worse than the one in linked page, so I assume my cpu is about
19.
▲
by
boroboro4
8mo ago
This one doesn’t even have warmup for Java, which makes results complete non sense. Those benchmarks should be just forbidden for their misleading nature.
20.
▲
by
boroboro4
8mo ago
He’s not a king to do whatever he promised as is, he’s bound by laws and constitution (which are passed by congress). Also as you were corrected there is constant goalpost moving in terms of whom exactly should be deported and how. If you’r
21.
▲
by
boroboro4
8mo ago
You can find stats including pending charges: https://bsky.app/profile/reichlinmelnick.bsky.social/post/3m... the main uptick in recent arrests is mostly people without any criminal charges including pending.
22.
▲
by
boroboro4
9mo ago
There might be more than one reason for an ongoing crisis, and different takes on who’s responsible. However Maduro is responsible of huge number of refugees fleeing Venezuela, and we (and some other countries around) have some obligations
23.
▲
by
boroboro4
9mo ago
Obviously the one which sets the law, also the one which has first article dedicated to it.
24.
▲
by
boroboro4
9mo ago
Language is just a form, what exactly is encoded inside of the model can be very different. And to encode logical reasoning inside of the weights with activation functions is more than possible. Models solving IMO level problems imo proves
25.
▲
by
boroboro4
9mo ago
This being said in this setup of 2-4 h100 you’ll be able to generate with batch size of somewhere around 128 ie its 128 humans and not one. And just like that difference in efficiency isn’t that high anymore.
26.
▲
by
boroboro4
9mo ago
Issue with your original statement is that it implies negative correlation between intellect and being good politician. All the issues you describe (both being self serving and power seeking) apply to anyone regardless to their intellect (a
27.
▲
by
boroboro4
9mo ago
In my opinion even charisma implies higher than average intellect which is already something. We should put more pressure on elected politicians around competence and integrity, sure, but it doesn’t mean random person is going to be better.
28.
▲
by
boroboro4
9mo ago
Would you rather be treated (medically) by the first 2000 people? Do you think code will be written better by the first 2000? I get being unhappy about current political class, but this kind of claims is wild to me.
29.
▲
by
boroboro4
9mo ago
Mechanics is exactly the same - it's not Tesla revenues driving returns for investors, it's new investors putting their money into the stock at very high price.
30.
▲
by
boroboro4
9mo ago
From the comment below: > Groq raised $750 million at a valuation of about $6.9 billion three months ago. Investors in the round included Blackrock and Neuberger Berman, as well as Samsung, Cisco, Altimeter and 1789 Capital, where Donald
More ›