Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dikobraz
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
dikobraz
7mo ago
LLM inference throughput benchmark for RTX PRO 6000 SE vs H100, H200, and B200 GPUs, based on the vllm serve and vllm bench serve benchmarking tools, to understand the cost-efficiency of various datacenter GPU options. Benchmarking Setup: T
2.
▲
by
dikobraz
7mo ago
I spent ten years inside three Big Tech companies, most of it at Apple. For a long time, I thought of corporations as ruthless business machines — designed to maximize revenue, or shareholder value, depending on my level of cynicism at the
3.
▲
Tech Bro Saga: big tech critique essay series
1 points
by
dikobraz
7mo ago
|
0 comments
4.
▲
Essay: Why Big Tech Leaders Destroy Value – When Identity Outlives Purpose
(medium.com)
2 points
by
dikobraz
8mo ago
|
1 comments
5.
▲
by
dikobraz
8mo ago
Over my ten-year tenure in Big Tech, I’ve witnessed conflicts that drove exceptional people out, hollowed out entire teams, and hardened rifts between massive organizations long after any business rationale, if there ever was one, had faded
6.
▲
Why Big Tech Performance Reviews Aren't Meritocratic
(medium.com)
5 points
by
dikobraz
8mo ago
|
1 comments
7.
▲
by
dikobraz
8mo ago
No matter how they’re designed—manager discretion, calibration committees, or opaque algorithms—performance reviews in big tech reliably produce results that are neither meritocratic nor humane. In practice, compensation and promotions stil
8.
▲
by
dikobraz
9mo ago
I haven’t worked at Google, but I doubt they’re immune. Lack of “important work” is one mechanism, but in the case I’m describing the issue was overlapping legitimacy: multiple orgs had good reasons to own the same scope. Hardware felt inse
9.
▲
by
dikobraz
11mo ago
I present an LLM inference throughput benchmark for RTX4090 / RTX5090 / PRO6000 GPUs based on vllm serving and vllm bench serve client benchmarking tool. The hardware configurations used: - 1x4090, 2x4090, 4x4090 - 1x5090; 2x5090;
10.
▲
How to Give Your RTX 4090 Nearly Infinite Memory for LLM Inference
(medium.com)
2 points
by
dikobraz
1y ago
|
1 comments
11.
▲
by
dikobraz
1y ago
We explored a network-attached KV-cache for consumer GPUs to offset their limited VRAM. It doesn’t make RTX cards run giant models efficiently. Still, for workloads that repeatedly reuse lengthy prefixes—such as chatbots, coding assistants,
12.
▲
Any solution for "reset bug" on Nvidia GPUs?
(medium.com)
2 points
by
dikobraz
1y ago
|
1 comments
13.
▲
by
dikobraz
1y ago
I am working on a platform for GPU rental and have recently encountered an extremely annoying issue. On all machines with RTX 5090 and RTX PRO 6000 GPUs, the cards occasionally become completely unresponsive — usually after a few days of VM
14.
▲
Limited-time GPU firepower Dirt-cheap LLM Inference: Llama 4, DeepSeek 0528
(cloudrift.ai)
3 points
by
dikobraz
1y ago
|
1 comments
15.
▲
by
dikobraz
1y ago
We’ve got a temporarily underutilized 64 x AMD MI300X cluster, so instead of letting it sit idle, we’re opening it up for LLM inference. Running: LLaMA 4 Maverick, DeepSeek V3, R1, and R1-0528. Want another open model? Let us know. We are h
16.
▲
by
dikobraz
2y ago
Hi HN, We're launching https://www.neuralrack.ai/ , an affordable datacenter for AI applications. Neuralrack specializes in affordable but powerful GPUs like the RTX 4090, which can outclass professional GPUs in specifi
17.
▲
NeuralRack – Affordable GPUs
(neuralrack.ai)
3 points
by
dikobraz
2y ago
|
2 comments