Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mcbuilder
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
mcbuilder
7d ago
Distillation is not infiltrating systems and exfiltrating weights.
2.
▲
by
mcbuilder
7d ago
I had to do a double take myself, scrolled down and arrived at a similar number. Wish I built this a few years ago!
3.
▲
by
mcbuilder
7d ago
Look what's coming out of China, they are catching up on performance and surpassing the US in efficiency. They're on a different level when it comes to open releases of weights. I don't, to me the entire premise is a bit flaw
4.
▲
by
mcbuilder
7d ago
One of the most biased claims IMO in AI 2027 is that a huge portion of the geopolitical and existential risk argument is hinged on the notion that China just steals the US frontier weights.
5.
▲
by
mcbuilder
7d ago
I feel that quote more likely is pointed at "music appreciation" as a written work, versus understanding the theory. I know a few musicians who are happy to use terminology from theory, "I love to write songs with diminished
6.
▲
by
mcbuilder
13d ago
And I didn't touch it, because it's your call to make...
7.
▲
by
mcbuilder
14d ago
I mean CoT came out of research circles not marketing
8.
▲
by
mcbuilder
14d ago
Well at least DeepSeek is up. :)
9.
▲
by
mcbuilder
17d ago
One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.
10.
▲
by
mcbuilder
24d ago
If I'm about to step into it 5 seconds after you leave, yes it does.
11.
▲
by
mcbuilder
24d ago
honestly "product guy" is enough of a slur for most technical folks. What you are saying just sounds nasty.
12.
▲
by
mcbuilder
26d ago
`emacs -nw` with DOOM Emacs for me. :)
13.
▲
by
mcbuilder
29d ago
Small # of parameters means no memory bottleneck, which means blazing fast performance.
14.
▲
by
mcbuilder
1mo ago
Nah, that's the same sort of thinking that makes people type "make no mistakes", I don't make my model roll play, etc. I believe that the longer the system prompt and the more you cram in it the worse the model does. You
15.
▲
by
mcbuilder
1mo ago
It mostly hurts people in countries with weak purchasing power. DS was the main game in down for them. Personally, I don't think we've seen the total end of dirt cheap LLMs, it's just a frontier lab doesn't want to be in
16.
▲
by
mcbuilder
1mo ago
Yeah, I feel like we bigly lose to China if we don't.
17.
▲
by
mcbuilder
2mo ago
Great distribution and great community!
18.
▲
by
mcbuilder
2mo ago
I love deepseek's (Pro) writing. I feel it's more nuanced and natural.
19.
▲
by
mcbuilder
2mo ago
I don't think it's necessarily "wiser" to go closed source. All of AI is built on mostly openness, at least on the software side. There are other ways to compete, it's just the model itself will be a commodity.
20.
▲
by
mcbuilder
2mo ago
Also base models versions are very useful for researchers, since they allow for a cold start to post training.
21.
▲
by
mcbuilder
2mo ago
Compared to the Opus 5 "model card", which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
22.
▲
by
mcbuilder
2mo ago
Non homogenized, lightly flash pasteurized milk has been my favorite for years. Just tastier, luckily I grew up in a little strange town where it was norm, and I don't have the weird modern health movement baggage to go with it.
23.
▲
by
mcbuilder
2mo ago
Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchm
24.
▲
by
mcbuilder
2mo ago
Just like we trained on human language to create LLMs, we can train on human keystrokes with a similar algorithm and spit out believable (at least statistically) "human" keystrokes generated by machine.
25.
▲
by
mcbuilder
2mo ago
We have probably hit a limit to scaling LLMs through raw parameter count alone, at least we're not seeing the exponential pace. I personally think we'll end up with a nice sigmoid curve plateauing in the sub 10T parameter regime.
26.
▲
by
mcbuilder
3mo ago
I consider myself to be in that cohort as well. :)
27.
▲
by
mcbuilder
3mo ago
LRMs are plateauing for sure, not that there won't be gains to be had in the future, but it's not like the era of rapid progress that was the past year any more.
28.
▲
by
mcbuilder
3mo ago
Yeah, it's actually the case. Researchers have shown that the models response doesn't always follow from the reasoning. Whether you consider that an internal language or not really depends on what you're speculating the neura
29.
▲
by
mcbuilder
3mo ago
I'm a white dude from Iowa, working in top levels of AI/ML. I'm in the minority at work/conferences. I hardly ever even interview homegrown US job candidates. I'm just saying, that the reason I think you see more pe
30.
▲
by
mcbuilder
3mo ago
I've always found this line of reasoning troubling and uninformed. Chinese models first of all can be hosted on your own hardware, I'd argue they are way more transparent than US companies, by well releasing stuff. Second, the &qu
More ›