Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
calebkaiser
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
calebkaiser
5d ago
Which hyperscalers are doing that? I've heard of one startup trying this, XFRA, currently in early pilot phases. But I'd be pretty surprised to hear that AWS or Google are paying people to host GPU clusters in their home.
2.
▲
by
calebkaiser
6d ago
Yeah, in essence. This is actually a pretty cool part of working in Lean. It's a somewhat normal convention to write something in a human readable way and then write a second optimized implementation with some kindness of correctness t
3.
▲
by
calebkaiser
6d ago
I'd wager that the vast majority of the ML research community, especially anyone interested in "AGI", is familiar with Hofstader's work. And I don't think anyone working on contemporary language models would argue t
4.
▲
by
calebkaiser
11d ago
Love seeing projects like this. The performance benchmarks are nice to see. Have you done any benchmarks against approaches like OpenAI's Symphony for things like token usage or task completion?
5.
▲
by
calebkaiser
12d ago
There will be companies doing this. I'm saying the labs are well positioned to be those companies, as they effectively are those companies right now. The same dynamics that define the public cloud ecosystem are at play here. What AWS s
6.
▲
by
calebkaiser
12d ago
I support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking a
7.
▲
by
calebkaiser
15d ago
I haven't read this in depth yet, though I plan to. If this general line of research is interesting to you, I'd recommend checking out some of the lines of research it touches upon--they're really rich and fascinating, and so
8.
▲
by
calebkaiser
15d ago
It gets extremely blurry, because people commonly refer to any model that uses a component associated with the Transformer architecture as a Transformer (i.e. using some kind of QKV-esque attention mechanism). I think it's easier to th
9.
▲
by
calebkaiser
20d ago
I think maybe just click around Huggingface's site a bit? They don't run a compute-intensive business on low monthly plans. You're describing frontier labs that sell access to their enormous models with subsidized subscriptio
10.
▲
by
calebkaiser
20d ago
I just feel like you're saying things based on a general vibe about "AI companies" but didn't pause to look up the particular company being discussed. Huggingface reached profitability 2 years ago. They've reportedl
11.
▲
by
calebkaiser
20d ago
I'm curious why?
12.
▲
by
calebkaiser
21d ago
It's not too dissimilar from GitHub, but geared towards ML. They have a 9/mo pro plan for individual users for upgraded storage/usage, and an enterprise version of Hub that larger orgs can pay for. I think the enterprise has
13.
▲
by
calebkaiser
21d ago
Solid point. I'm sure Nvidia went into this deal expecting completely flat growth and no other benefits to their core business. Sorta like how Meta never increased Instagram's revenue from $0 and is still waiting for it to pay off
14.
▲
by
calebkaiser
21d ago
Yeah, great point. Nvidia could potentially be overpaying. Not sure how that equates to Huggingface being "a file download mirror with a couple of side features dangling off".
15.
▲
by
calebkaiser
21d ago
I dunno. You can argue over whether they're overpaying, but it's not like Huggingface is Clinkle. They hit $150 million in ARR this year, they have tons of runway, and according to reports, have just started to even burn the money
16.
▲
by
calebkaiser
21d ago
They hit 150 million in annual recurring revenue this year. Pretty nice dangling side features apparently.
17.
▲
by
calebkaiser
21d ago
What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They l
18.
▲
by
calebkaiser
22d ago
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020. If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done ci
19.
▲
by
calebkaiser
28d ago
Sorry--I was unclear. I was specifically responding to the OP's "CEOs these days want to join the AI hype wave" claim. My point was that the "hype wave" OP is referencing is really about deep learning, not what we t
20.
▲
by
calebkaiser
28d ago
1. That's not what progression-free survival means. 2. I don't think AI was mentioned once in this press release. 3. This drug was approved for trial in 2014, so if it had something to do with deep learning, that would actually be
21.
▲
by
calebkaiser
1mo ago
I dunno. There is a lot of valid and interesting criticism to write about digital marketing. Lots of people have attempted to study the efficacy of digital advertising, and I'd love to see a deep dive into the "MarTech" ecosy
22.
▲
by
calebkaiser
1mo ago
It's a totally reasonable question, and one that everyone asks when they're learning about CUDA. The frustrating answer to your last question is that lots of companies have shipped GPU dev environments that can theoretically be us
23.
▲
by
calebkaiser
2mo ago
So OpenAI would be proving to the government that they should be trusted to govern models, by failing to govern their models? Plenty of shady stuff goes on in marketing, and I'm sure that OpenAI is going to opportunistically grab any p
24.
▲
by
calebkaiser
2mo ago
I think the commenter is saying that OpenAI committed a false flag operation, not Tailscale. For Tailscale, this may very well be marketing, but it would be strangely self defeating for OpenAI to do something like what is being suggested. S
25.
▲
by
calebkaiser
2mo ago
I'm typically very skeptical of most content marketing/corporate PR. In the case of this incident, I struggle to see the clear upshot for OpenAI. It seems pretty unlikely they'd ever okay this intentionally as some sort of ma
26.
▲
by
calebkaiser
2mo ago
Implementing models directly from papers is typically pretty doable (and is of course more straightforward when the full implementation is open sourced). Often there is some amount of specific knowledge, like particular hyperparameters, tha
27.
▲
by
calebkaiser
2mo ago
OpenAI's head of strategic futures publicly stated that you can't explain the quality of the newest Kimi via distillation. Further, you can just read the papers released alongside most open models. Plenty of hugely influential res
28.
▲
by
calebkaiser
2mo ago
Here's his original post https://x.com/i/status/2078133895766114412
29.
▲
by
calebkaiser
2mo ago
The guy who said those things just joined OpenAI in the last month. He was previously a senior policy advisor, specifically on AI, to the current administration in the White House. I don't know that the US government has a unified posi
30.
▲
by
calebkaiser
2mo ago
The conversation going on in the industry is a bit broader than that. OpenAI's head of strategic futures just said last week that open models are inherently decelerationist, ungovernable, and will slow development on the frontier. You&
More ›