Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
technoabsurdist
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
technoabsurdist
2mo ago
hi I work at wafer. yes we ran benchmarks. for example our Kimi K3 is live on open router and in order to host there you have to run accuracy checks. like tau/gpqa. and then u must pass test regarding thinking, coherency, and tool call
2.
▲
by
technoabsurdist
3mo ago
sounds good feedback taken, thanks beffjezos
3.
▲
by
technoabsurdist
3mo ago
hi yes it’s not optimized for single stream it’s optimized for total node throughput
4.
▲
by
technoabsurdist
3mo ago
this is exactly our thesis at wafer :) thank you for the support
5.
▲
by
technoabsurdist
3mo ago
AMD MI355X uses 1,400W per GPU and NVIDIA B200 uses 1,200W. So AMD uses about 16% more power.
6.
▲
by
technoabsurdist
3mo ago
yes it is 213 tok/s single stream (so per user)
7.
▲
by
technoabsurdist
3mo ago
hi i work at wafer. no the margins are lower averaging at about ~40%. utilization is one of the highest order bits in determining margins here, yes.
8.
▲
AI Could Democratize One of Techs Most Valuable Resources
(wired.com)
6 points
by
technoabsurdist
5mo ago
|
2 comments
9.
▲
Show HN: Wafer – Profile, inspect assembly, and iterate on CUDA within your IDE
(wafer.ai)
3 points
by
technoabsurdist
9mo ago
|
1 comments
10.
▲
Show HN: GPU Profiling That's Useful in 60 Seconds
(keysandcaches.com)
1 points
by
technoabsurdist
1y ago
|
0 comments
11.
▲
Show HN: We made PyTorch profiling usable for ML engineers
(herdora.mintlify.app)
3 points
by
technoabsurdist
1y ago
|
0 comments
12.
▲
Chip Benchmark: Hardware-Centric Performance Insights for AI Workloads
(herdora.com)
1 points
by
technoabsurdist
1y ago
|
0 comments
13.
▲
by
technoabsurdist
1y ago
^ We currently just have llama3.1-8b, so we'll be working on adding more models across more hardware options!
14.
▲
ChipBenchmark: Open-Source Benchmarking for LLM Performance Across Hardware
(chipbenchmark.com)
3 points
by
technoabsurdist
1y ago
|
2 comments
15.
▲
by
technoabsurdist
1y ago
We just launched Chip Benchmark, an open-source tool for hardware-centric benchmarking of open-weight LLMs across accelerators like NVIDIA A100/H100/L40S and AMD MI300X. It measures throughput, latency, and time-to-first-token wit
16.
▲
Using AMD MI300X for High-Throughput, Low-Cost LLM Inference
(herdora.com)
8 points
by
technoabsurdist
1y ago
|
0 comments
17.
▲
Profile CUDA kernels with one command, zero GPU setup
(github.com)
3 points
by
technoabsurdist
1y ago
|
1 comments
18.
▲
by
technoabsurdist
1y ago
We've been doing lots of GPU kernel profiling and optimization on cloud infrastructure, but without local GPU hardware, that meant constant SSH juggling: upload code, compile remotely, profile kernels, download results, repeat. Or, wor
19.
▲
Show HN: Profile GPU Kernels with One Command, Zero GPU Setup
(github.com)
2 points
by
technoabsurdist
1y ago
|
0 comments
20.
▲
Show HN: Chisel – Profile GPU Kernels Without a GPU (Nvidia and AMD)
(github.com)
3 points
by
technoabsurdist
1y ago
|
0 comments
21.
▲
by
technoabsurdist
1y ago
oh yeah, in my experience anything below ROCm6.x really sucks. I tried to run qwen2.5-32B on ROCm5.x and it was running at <15tok/s lol. Have you tried running any sort of LLM inference on your MI25, or what NN workloads are you run
22.
▲
Ask HN: Is anyone using AMD GPUs for their AI workloads?
6 points
by
technoabsurdist
1y ago
|
2 comments
23.
▲
Show HN: Chisel – GPU development through MCP
(github.com)
1 points
by
technoabsurdist
1y ago
|
0 comments
24.
▲
Show HN: Chisel – Profile AMD MI300X kernels locally
(github.com)
2 points
by
technoabsurdist
1y ago
|
0 comments
25.
▲
by
technoabsurdist
1y ago
correct github link: https://github.com/Herdora/chisel
26.
▲
Show HN: Chisel – AMD GPU development that feels local but runs in the cloud
(pypi.org)
2 points
by
technoabsurdist
1y ago
|
2 comments
27.
▲
Crackd – unbiased technical skill comparison among students
(crackd.io)
2 points
by
technoabsurdist
1y ago
|
1 comments
28.
▲
by
technoabsurdist
1y ago
We built Crackd to solve a fundamental problem: helping students understand where they stand technically, and helping them learn how to become the absolute best. Technical students have no reliable way to know how good they are. Grades don&
29.
▲
Show HN: Molecule-Rs – A Fast Protein Visualization Engine Written in Rust
(molecule-rs.vercel.app)
4 points
by
technoabsurdist
1y ago
|
0 comments
30.
▲
Show HN: Molecule-Rs – A Fast Protein Visualization Engine Written in Rust
(technoabsurdist.github.io)
2 points
by
technoabsurdist
1y ago
|
1 comments
More ›