Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
asaiacai
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Distributed PyTorch Profiling
(zhenyu.github.io)
1 points
by
asaiacai
15d ago
|
0 comments
2.
▲
by
asaiacai
23d ago
we're only calling it an "incident" now i see. smh
3.
▲
Visualizing Combos in Judo in R
(r-bloggers.com)
37 points
by
asaiacai
23d ago
|
3 comments
4.
▲
by
asaiacai
5mo ago
this list has a lot of false positives.
5.
▲
by
asaiacai
6mo ago
lmao
6.
▲
by
asaiacai
6mo ago
it's the perfect drug. You don't know how to code something up. Ask AI to implement it. It's broken? Ask AI to fix it for you. Will people become unable to fix things without it?
7.
▲
by
asaiacai
6mo ago
have you guys tested this against any resource/batch managers in k8s (i.e. kueue, volcano, apache yunikorn)? seems like this is a good fit for people who have already have a large cluster. does it handle autoscaling environments well i
8.
▲
by
asaiacai
7mo ago
its cool to see the iterative improvements to your model laid out, but for everything that workedm i imagine there were at least a million other things you also tried but didnt work out. whats your process of trying these different techniqu
9.
▲
How Upc Barcodes Work
(craigball.net)
2 points
by
asaiacai
8mo ago
|
0 comments
10.
▲
Kamaji: Containerized Control Planes for K8s
(kamaji.clastix.io)
1 points
by
asaiacai
8mo ago
|
0 comments
11.
▲
by
asaiacai
9mo ago
echoing this sentiment. I really do believe at some level there is strength/wisdom in being able to step away from a problem and return to it from a new perspective despite of what narratives are being pushed online by hustle
12.
▲
Debugging TLS failures in distroless containers
(lucabaggi.com)
28 points
by
asaiacai
9mo ago
|
0 comments
13.
▲
Glibc malloc performance degradation with CPU affinity masks
(bugs.launchpad.net)
1 points
by
asaiacai
11mo ago
|
0 comments
14.
▲
When Python can't thread: a deep-dive into the GIL's impact
(pythonspeed.com)
3 points
by
asaiacai
11mo ago
|
0 comments
15.
▲
by
asaiacai
1y ago
MFU is probably the best but requires application logic. You can export metrics at the infra level like SM efficiency. We explain it a bit how we used it to do some optimization. https://www.trainy.ai/blog/gpu-utilizati
16.
▲
by
asaiacai
2y ago
John Ousterhout's "A Philosophy of Software Design" I liked. It was supposed to be assigned reading for Berkeley's data structures class CS61B, and I don't think I really internalized the lessons within, but after r
17.
▲
by
asaiacai
2y ago
power is also a good proxy. For example, we've had distributed runs that we monitored on WandB where one of our workers died in the middle and the rest were basically stalling on the dead worker. On WandB, we were only logging GPU stat
18.
▲
Skypilot on Kubernetes
(blog.skypilot.co)
1 points
by
asaiacai
2y ago
|
0 comments
19.
▲
by
asaiacai
2y ago
totally agreed. A lot of our findings during this process is that there's still a lot of alpha in finding the right kernels for the job/model. We're hoping that in the future `torch.compile` will become more mature because cu
20.
▲
GPU utilization is a misleading metric. DCGM and Konduktor
(trainy.ai)
4 points
by
asaiacai
2y ago
|
1 comments
21.
▲
by
asaiacai
2y ago
I really hope this makes it easier to install/upgrade NVIDIA drivers on Linux. It's a nightmare to figure out version mismatches between drivers, utils, container-runtime...
22.
▲
Show HN: OSS tool to finetune and serve LLMs on different clouds
(llm-atc.readthedocs.io)
12 points
by
asaiacai
3y ago
|
0 comments