Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
germanjoey
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
germanjoey
3mo ago
IMO "bugs per commit" is even worse than that, because, in addition to what you say, it also hides the extraordinary spike of commit activity of a project that had previously been stable. [0] It is the exact metric you'd choo
2.
▲
by
germanjoey
1y ago
TBH, the 2x-4x improvement over a naive implementation that they're bragging about sounded kinda pathetic to me! I mean, it depends greatly on the kernel itself and the target arch, but I'm also assuming that the 2x-4x number is t
3.
▲
by
germanjoey
2y ago
This is really incredible, thank you!
4.
▲
by
germanjoey
2y ago
Sambanova's RDU is a dataflow processor being used for ML/AI workloads! It's amazing and actually works.
5.
▲
by
germanjoey
2y ago
Pretty amazing speed, especially considering this is bf16. But how many racks is this using? The used 4 racks for 70B, so this, what, at least 24? A whole data center for one model?!
6.
▲
by
germanjoey
2y ago
the title says "Cerebras Trains Llama Models"...
7.
▲
by
germanjoey
2y ago
They said in the announcement that they've implemented speculative decoding, so that might have a lot to do with it. A big question is what they're using as their draft model; there's ways to do it losslessly, but they could
8.
▲
by
germanjoey
2y ago
Simply increasing processing power for the AI isn't enough. Gameplay mechanics are intimately related to the capabilities of the AI. For example, when they redesigned combat around the 1-Unit-Per-Tile (1UPT) mechanic for CIV 5, this cr
9.
▲
by
germanjoey
2y ago
How are you verifying accuracy for your JAX port of Llama 3.1? IMHO, the main reason to use pytorch is actually that the original model used pytorch. What can seem to be identical logic between different model versions may actually cause mo
10.
▲
by
germanjoey
2y ago
Looks like some kind of power play... Originally discussed here: https://news.ycombinator.com/item?id=41234180
11.
▲
Sambanova breaks 1000 tokens/SEC on LLama3 8B
(twitter.com)
7 points
by
germanjoey
2y ago
|
0 comments
12.
▲
by
germanjoey
2y ago
Is there a demo of a model visualized using this somewhere? Even if it's just a short video... it's hard to tell what it's like from screenshots.
13.
▲
by
germanjoey
2y ago
cost effective in what sense? groq doesn't achieve high efficiency, only low latency. but that's not done in a cost-effective way. compare sambanova achieving the same performance with 8 chips instead of 568, and with higher preci
14.
▲
Try SambaNova chat: 1T param LLM, 500 tokens/SEC
(coe-1.cloud.snova.ai)
1 points
by
germanjoey
2y ago
|
1 comments
15.
▲
by
germanjoey
2y ago
We're showing off our 1.05T param Composition of Experts LLM! It's 150 experts running on 1 node consisting of 8 SN40L RDU chips. Each of our nodes has a huge amount of DDR attached, in addition to copious amounts of on-chip HBM a
16.
▲
by
germanjoey
3y ago
Sambanova just launched something similar to what you're describing. It's a demo of their new chip running a 1T param MoE model 150 7B llama2s, each retrained to be an expert in a different topic. So one of them is a "law&qu
17.
▲
SambaNova launches new SN40L chip; demo of 1T param CoE LLM
(sambanova.ai)
3 points
by
germanjoey
3y ago
|
0 comments
18.
▲
by
germanjoey
4y ago
welp, This report focuses on the capabilities, limitations, and safety properties of GPT-4. GPT-4 is a Transformer-style model [33 ] pre-trained to predict the next token in a document, using both publicly available data (such as internet d
19.
▲
by
germanjoey
4y ago
How big is this model? (i.e., how many parameters?) I can't find this anywhere.
20.
▲
by
germanjoey
4y ago
I worked with the author for a couple of years, pre- and post- acquisition, and I have to admit that he drove me somewhat crazy sometimes too. Leaving that aside, I also had an immense amount of personal respect for him as I could see how m
21.
▲
by
germanjoey
4y ago
What's the new performance process?
22.
▲
by
germanjoey
4y ago
> You don't introduce more coupling, you don't the coupling that already exists. This is true at the code level. But at the system-design level, this documentation is the extra coupling. I feel like it's important to und
23.
▲
by
germanjoey
4y ago
It is interesting reading that second paragraph many years later. Most of the things that Steve Yegge brags about that Google "does right" (e.g. how they do recruiting, their engineering "standards", SREs running product
24.
▲
by
germanjoey
4y ago
Great post; this is how I felt about it too.
25.
▲
by
germanjoey
4y ago
The article (or, rather, the commentary in the link above on the article) talks about the fallacious notion of "market cap" in regards to cryptocurrencies. That is to say, e.g., multiplying the number of bitcoins in existence time
26.
▲
by
germanjoey
5y ago
It means you play through the four characters in order in four successive runs, e.g. ironclad -> silent -> defect -> watcher, rather than e.g. just playing the watcher (the character generally considered the easiest to win with at
27.
▲
by
germanjoey
5y ago
...and what service is that?
28.
▲
by
germanjoey
5y ago
You're severely underselling Google's incapability, e.g. https://i.imgur.com/UoIZSU2.png
29.
▲
by
germanjoey
5y ago
"whitespacer" reminds me of Damian Conway's classic "Acme::Bleach" perl module... https://metacpan.org/pod/Acme::Bleach
30.
▲
by
germanjoey
5y ago
Many manga fans have a love/hate relationship with mangadex. On one hand, it's provided hosting for countless hours of entertainment over the years. Their "v3" version of the site was basically perfect from a usability p
More ›