Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stygiansonic
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
stygiansonic
3mo ago
From a brief reading of what Fusion does: https://openrouter.ai/docs/guides/features/plugins/fusion Looks like Fusion calls a bunch of models and then uses an LLM to synthesize the results, and pass to a
2.
▲
by
stygiansonic
6mo ago
See also: https://en.wikipedia.org/wiki/Metcalf_sniper_attack (Perpetrators also not caught)
3.
▲
Family of Tumbler Ridge shooting victim suing OpenAI
(cbc.ca)
9 points
by
stygiansonic
6mo ago
|
0 comments
4.
▲
by
stygiansonic
8mo ago
The jury found that Ding stole trade secrets relating to the hardware infrastructure and software platforms that allow Google’s supercomputing data center to train and serve large AI models. The trade secrets contained detailed informa
5.
▲
Apple picks Gemini to power Siri
(cnbc.com)
1038 points
by
stygiansonic
8mo ago
|
651 comments
6.
▲
by
stygiansonic
8mo ago
Neat experiment that gives a mechanistic interpretation of temperature. I liked the reference to the "anomalous" tokens being near the centroid, and thus having very little "meaning" to the LLM.
7.
▲
US authorities shut down major China-linked GPU smuggling operation
(justice.gov)
2 points
by
stygiansonic
9mo ago
|
0 comments
8.
▲
Anthropic to buy $30B in Azure capacity in deal with Microsoft, Nvidia
(cnbc.com)
3 points
by
stygiansonic
10mo ago
|
1 comments
9.
▲
by
stygiansonic
1y ago
That paper doesn’t seem to be about security vulnerabilities in MiG but rather using it to improve workload efficiency
10.
▲
by
stygiansonic
1y ago
Wonder why they haven’t gotten an in house pizzeria yet to reduce the signal on this side channel leak
11.
▲
by
stygiansonic
1y ago
From the article it appears to be something they invented: > Gemma 3n leverages a Google DeepMind innovation called Per-Layer Embeddings (PLE) that delivers a significant reduction in RAM usage. Like you I’m also interested in the archit
12.
▲
by
stygiansonic
1y ago
Interesting! If Knuth is not the original author then they’ve been lost to the sands of time
13.
▲
by
stygiansonic
1y ago
Great article and nice explanation. I believe this describes “Algorithm R” in this paper from Vitter, who was probably the first to describe it: https://www.cs.umd.edu/~samir/498/vitter.pdf
14.
▲
by
stygiansonic
1y ago
The article mentions this union, not sure if it meets your definition of success: https://www.alphabetworkersunion.org/our-wins
15.
▲
by
stygiansonic
2y ago
When subtlety proves too constraining, competitors may escalate to overt cyberattacks, targeting datacenter chip-cooling systems or nearby power plants in a way that directly—if visibly—disrupts development. Should these measures falter, s
16.
▲
Superintelligence Strategy
(nationalsecurity.ai)
33 points
by
stygiansonic
2y ago
|
62 comments
17.
▲
by
stygiansonic
2y ago
Sorry to hear this A lot of my teenage years were spent building and playing with PCs and a lot of the knowledge and interest came from reading each and every issue of boot and maximum pc
18.
▲
by
stygiansonic
2y ago
+1 Jumping into an unknown codebase (which may be a library you depend on) and being able to quickly investigate, debug, and root cause an issue is an extremely invaluable skill in my experience Acting as if this isn’t useful in the real wo
19.
▲
by
stygiansonic
2y ago
I wrote about something similar, which was motivated by an issue I saw caused by an (incorrect) expectation that a Java hashmap iteration order would be random: https://peterchng.com/blog/2022/06/17/what-
20.
▲
by
stygiansonic
2y ago
Yeah, ops comment makes it seem like they are building racks of RTX 4090s, when this isn’t remotely true. Tensor Core performance is far different on the data center class devices vs consumer ones.
21.
▲
by
stygiansonic
2y ago
Thanks for writing this. Is this concept (dice room puzzle, doomsday argument) at all related to the st Petersburg paradox? https://en.m.wikipedia.org/wiki/St._Petersburg_paradox
22.
▲
by
stygiansonic
2y ago
Reminiscent of a scene from Billions: https://www.reddit.com/r/Billions/comments/czlg4u/need_help_...
23.
▲
by
stygiansonic
2y ago
Thanks - I added my contact info (I don’t comment a lot on HN, mostly just read) but will drop you a line
24.
▲
by
stygiansonic
2y ago
This is probably using their excess capacity, but not necessarily that their GPUs are idle. For LLMs/large models the huge cost is memory ops to load each layer weights during the forward pass. This is why doing inference at batch size
25.
▲
by
stygiansonic
2y ago
A simplified explanation, which I think I heard from Karpathy, is that transformer models only do computation when they generate (decode) a token. So generating more tokens (using CoT) gives the model more time to “think”. Obviously this do
26.
▲
by
stygiansonic
2y ago
Op probably referring to an M series MacBook since it has a unified memory architecture and the same memory space used by both cpu and gpu
27.
▲
by
stygiansonic
3y ago
This is a much better set of general terms vs the original, which currently is heavily oriented towards LLMs/Gen AI
28.
▲
by
stygiansonic
3y ago
It seems like he (re)joined OpenAI almost exactly 1 year ago: https://twitter.com/karpathy/status/1623476659369443328
29.
▲
by
stygiansonic
3y ago
Also who would cover the legal liability in case of damages?
30.
▲
by
stygiansonic
3y ago
https://developer.apple.com/contact/request/download/alterna... One of the requirements for third party App Store use by developers is maintaining a standby letter of credit of 1M euros from an A rated financ
More ›