Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dhruvdh
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
dhruvdh
8mo ago
Try `uvx pocket-tts serve`
2.
▲
Accelerating LLM Inference with Parallel Draft Models (PARD)
(amd.com)
1 points
by
dhruvdh
1y ago
|
0 comments
3.
▲
by
dhruvdh
1y ago
El Capitan can also do FP8. HPC requires double precision generally but people are trying to make low precision work.
4.
▲
by
dhruvdh
2y ago
To be fair, you can buy ~3 of these for the price Nvidia charges for 24GB/32GB models.
5.
▲
by
dhruvdh
2y ago
To add, AMD only makes _parts_ of an MI300X server. It's like asking a tire manufacturer to give you a car for free.
6.
▲
by
dhruvdh
2y ago
I wish more people would just try to do things just like this and blog about their failures. > The published version of a proof is always condensed. And even if you take all the math that has been published in the history of mankind, it’
7.
▲
by
dhruvdh
2y ago
Disappointed that there wasn’t anything on inference performance in the article at all. That’s what the major customers have announced they use it for.
8.
▲
by
dhruvdh
2y ago
Which algorithm you pick for what shape of matrices is different and not straightforward to figure out. AMD currently wants you to “tune” ops and likely search for the right algorithm for your shapes while Nvidia has accurate heuristics for
9.
▲
Open-sourcing Three EXAONE 3.5 Models: 2.4B, 7.8B, 32B
(lgresearch.ai)
13 points
by
dhruvdh
2y ago
|
4 comments
10.
▲
by
dhruvdh
2y ago
> despite them being fabless That's not how it works. You need to pump money into fabs to get them working, and Intel doesn't have money. If AMD had fabs to light up their money, they would also have a much lower valuation. The
11.
▲
by
dhruvdh
2y ago
> Performance per watt was better for Intel No, not its not even close. AMD is miles ahead. This is a Phoronix review for Turin (current generation): https://www.phoronix.com/review/amd-epyc-9965-9755-benchmark... Y
12.
▲
by
dhruvdh
2y ago
Is AMD behind hyperscaler in-house efforts? Outside of Google I don't think so.
13.
▲
by
dhruvdh
2y ago
Oh, maybe also change the title? I flagged it because of the title/url not matching.
14.
▲
by
dhruvdh
2y ago
I don't think having a common ancestry for the ISA means much, or even having the same ISA. Anyway, I don't understand what you want from me or are arguing about. They were trying to win the datacenter CPU market and not the GPU m
15.
▲
by
dhruvdh
2y ago
Those are Vega, not CDNA. It wouldn't surprise me if those are rebranded consumer chips, though I haven't checked.
16.
▲
by
dhruvdh
2y ago
And yet Meta is using MI300X exclusively for all live inference on Llama 405B. Clearly there are workloads AMD wins at, and just going Nvidia by default for everything without considering AMD is suboptimal.
17.
▲
by
dhruvdh
2y ago
You know AMD primarily sells CPUs right? For datacenter GPUs, they're going from ~500M-750M in 2023 full year (can't find proper numbers), to 4.5B+ full year 2024. In GPUs, it's almost like they're entering a new market.
18.
▲
by
dhruvdh
2y ago
Batching is how you get ~350 tokens/sec on Qwen 14b on vLLM (7900XTX). By running 15 requests at once. Also, there is a Dockerfile.rocm at the root of vLLM's repo. How is it a pain?
19.
▲
by
dhruvdh
2y ago
Why would you use this over vLLM?
20.
▲
by
dhruvdh
2y ago
What's the point of the 8000 LOC limit? Has anyone worked in a project with a LOC limit? Why was the limit in place?
21.
▲
by
dhruvdh
2y ago
The MacBook NPU is 3x slower than the 45 TOPS threshold required for Copilot+ PC branding.
22.
▲
by
dhruvdh
2y ago
Yeah, and any updates to the model that make Recall attractive, cannot go back in time and reanalyze the past to be more useful. At least they will learn a lot from this, in 6-12 months and perhaps a new generation of NPUs we might have som
23.
▲
by
dhruvdh
2y ago
The NPU runs this Silica model at 1.5 watts. MacBooks cannot even drive multiple monitors in this price range.
24.
▲
by
dhruvdh
2y ago
The FPGA being used is I believe one of the lowest speced SKUs. AWS instance prices are more of a supply/demand/availability thing, it would be more interesting to compare from a total cost of ownership / perf-power-area pres
25.
▲
by
dhruvdh
2y ago
I don't know what you are trying to say here. If one system doesn't need to move as much data because it is more flexible, that is a good thing. What do we gain by making it "fair"?
26.
▲
by
dhruvdh
2y ago
I would imagine the importance of weights depends on the prompt. How do you decide which weights are important?
27.
▲
by
dhruvdh
2y ago
There is a VitisAI execution provider for ONNX, and you can use ONNX backends for inference frameworks that support it. More info here - https://ryzenai.docs.amd.com/en/latest/ But regardless, 16 TOPs is no good f
28.
▲
by
dhruvdh
2y ago
You can't buy these pro variants from Microcenter for example, but you can buy them from pre-built OEM desktops. Mostly meant for enterprise customers who buy in bulk, I think.
29.
▲
by
dhruvdh
3y ago
It's not like Microsoft is working on "Windows AI Studio" [1], or released Orca, or Phi. It's not like there's any talk of AI PCs with mandatory TOPs requirements for Windows 12. Big bad Microsoft coming for your l
30.
▲
by
dhruvdh
3y ago
The seven conjectures the paper presents are as listed below. If you find yourself intuitively agreeing with the conjectures, you might find the arguments and the limited support the paper presents helpful in reinforcing your intuition; and
More ›