Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
carterschonwald
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
carterschonwald
26d ago
perhaps, but what public activity isn't activism? eg, i have very interesting empirical evidence that recent Anthropic models are specifically trained to refuse to critique the whitehouse cabinet and elected officials, and that this is
2.
▲
by
carterschonwald
26d ago
… that dude probably would have preferred not having that in their life
3.
▲
by
carterschonwald
28d ago
good.
4.
▲
by
carterschonwald
29d ago
ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task. wish i was joking.
5.
▲
by
carterschonwald
1mo ago
what sorts of background texts/sources of info did you use? edit: i think i have the dover book your readme mentions :)
6.
▲
by
carterschonwald
1mo ago
this confirms what i had determined empirically: even handedness directions got baked into the weights post opus 4.7. theres two reasons this is deeply bad 1) knowingly pursuing a policy that foreseeably causes the deaths of many thousands
7.
▲
Silent Data Corruption in PyTorch
(github.com)
2 points
by
carterschonwald
1mo ago
|
1 comments
8.
▲
by
carterschonwald
1mo ago
for those here using pytorch, theres some really nasty data corruption for computations that arent using powers of two as the coordinate axes. in gradient calcs ive had examples with multiple orders of magnitude of remative or absolute erro
9.
▲
by
carterschonwald
2mo ago
im building my own harness and inference tool chain for much of these reasons. theres so much to do that makes a big difference for users. hoping to get things into shape for early alpha as a saas in the next two months. heres the easiest b
10.
▲
by
carterschonwald
2mo ago
ummm, there is no substantive censorship with deep seek aside from first oarty hosting by deepseek for cya. trust me, and the suppression on deepseek hosted deepseek is pretty thin if you’re sophisticated and explain ethics justificstions.
11.
▲
by
carterschonwald
2mo ago
exactly. you certainly know more about the brain than i :)
12.
▲
by
carterschonwald
2mo ago
thx for the kind response! at the very least i have tools that let me easily hit better perf for fancy dense memory layouts, and the same tooling lets me experiment with frankly wildly wacky sparse and structured memory formats. the perfor
13.
▲
by
carterschonwald
2mo ago
i definitely will be doing some drop of some faster attention kernels in the next few weeks. like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen.
14.
▲
by
carterschonwald
2mo ago
amusingly ive been working on ultra sparse llm inference/ training/ model design because nature loaths a dense graph/matrix and cause i think it shoukd be possible. i actually stood up a 20-25 percent faster than sota causal
15.
▲
by
carterschonwald
2mo ago
i think the line is: expressing that you reputationally certify its correct and its worth the time
16.
▲
by
carterschonwald
2mo ago
this mirrors my approx experience.
17.
▲
by
carterschonwald
2mo ago
i've had trouble finding any anecdotes or data about how to actually set/explore logit sampler settings
18.
▲
by
carterschonwald
2mo ago
5.6 sol is pretty good, its definitely very very well tuned, i'm not quite sure which of the others i should use near term, but theyre doing great work
19.
▲
by
carterschonwald
2mo ago
this so true. its really hard to make sure a model isnt going off the rails if i dont see full cot. the fact that oai and anthropic models hide it now has made them less reliable. which is a shame
20.
▲
by
carterschonwald
2mo ago
its a wrapper around mlx, so thats gonna be the portability bottleneck
21.
▲
by
carterschonwald
2mo ago
not sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out
22.
▲
by
carterschonwald
2mo ago
one of the most hilarious examples of how the current wh admins false protectionism, in my mind, is the crucible steel bankruptcy in jan/feb 2025. its a technological tragedy because it was the only facility im aware of globally that c
23.
▲
by
carterschonwald
2mo ago
ive had doom loops on release day with opus 4.6. quantization aint the culprit ;)
24.
▲
by
carterschonwald
2mo ago
even before the llm era sites would flag me as a bot for opening 15 links to read later. its fucking infuriating now
25.
▲
by
carterschonwald
2mo ago
this is pretty cool. i think part of the root cause is current rlhf post training design around confidence and optics rather than cooperative transparent honesty. though its kinda an expensive hypothesis to dig into as a private individual
26.
▲
by
carterschonwald
2mo ago
This paper points at an idea, but its really only legible if you have a more developed version of the idea already. I really should write more
27.
▲
by
carterschonwald
2mo ago
i literally had a co2 sensor for my engineering team last fall cause the space was so poorly ventilated. just measuring it continuously radically changed how everyone approached using the space packing wise and ventilation. smelled better
28.
▲
by
carterschonwald
3mo ago
good. Of course the precise language of the ruling matters, but good.
29.
▲
by
carterschonwald
3mo ago
so the most notoriously patent oriented tech firm is buying this up. lol ;) good for the founders. also explains why my resume got dropped on the floor as a desk reject :p
30.
▲
by
carterschonwald
3mo ago
.... i thought this was more widely known, granted i did write up a pretty wacky doc explaining way more fun experiments than these, and i have a fix that even prevents role collapse in my harness on github
More ›