Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
junrushao1994
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
junrushao1994
3y ago
This is great! Have you guys considered integrating with one of the existing systems?
2.
▲
by
junrushao1994
3y ago
Yeah thanks for sharing! This is definitely super valuable data and insights :) Regarding exllama-V2, MLC/TVM does benchmark against it: - Single GPU: https://github.com/mlc-ai/llm-perf-bench#int4-quantized-sing...
3.
▲
Scaling LLama2-70B with Multiple Nvidia/AMD GPU
(blog.mlc.ai)
13 points
by
junrushao1994
3y ago
|
6 comments
4.
▲
by
junrushao1994
3y ago
Machine Learning Compilation (MLC) now supports compiling LLMs to multiple GPUs. For Llama2-70B, it runs 4-bit quantized Llama2-70B at: - 34.5 tok/sec on two NVIDIA RTX 4090 at $3k - 29.9 tok/sec on two AMD Radeon 7900XTX at $2k -
5.
▲
by
junrushao1994
3y ago
Ah please help us by submitting a PR! I noticed the rust build failed last night but didn’t get a chance to look into it
6.
▲
by
junrushao1994
3y ago
This is a particular unique course offering introduction on ML compilation and deployment :)
7.
▲
by
junrushao1994
3y ago
As of today performance in WebGPU isn't as competitive yet, but there are really quite a lot of low-hanging fruits for WebGPU to pick up.
8.
▲
by
junrushao1994
3y ago
That's a great idea! We should dig around and see if there's any plugin to use
9.
▲
by
junrushao1994
3y ago
This is amazing to hear Steven! (Sorry I locked myself out of discord a couple of days ago...) I'm sure there's bunch of features missing like biased sampling you mentioned, and more than happy to merge PRs if you'd love to :
10.
▲
by
junrushao1994
3y ago
True and there are some other issues to be addressed. Those two particular issue is on our roadmap. Regarding quantization, we wanted to develop a code path that absorbs any quantization formats, for example, those from GGML or GPTQ, so tha
11.
▲
by
junrushao1994
3y ago
LLM decoding is dominated by memory bandwidth, and 3090Ti and 4090 happen to have the identical theoretical memory bandwidth
12.
▲
by
junrushao1994
3y ago
We haven't done any comparison them yet, but generally we believe Vulkan as a more generic cross-vendor API should be slower than ROCm. Same for CUDA vs Vulkan.
13.
▲
by
junrushao1994
3y ago
Well, I'm very much into true open source, and my belief is that any contributor is automatically part of the team :)
14.
▲
by
junrushao1994
3y ago
Generally speaking I expect Vulkan to be slower than ROCm given it's designed for generic gaming across GPU vendors, so the takeaway is, whenever ROCm is available and usable, we should use ROCm. And it's the same for CUDA vs Vulk
15.
▲
by
junrushao1994
3y ago
> Can you comment on how difficult it was to achieve this, and what the relative advantages b/w cards? Thanks for asking! I personally believe TVM Unity is a proper software stack for ML compilation (MLC), and its existing optimizat
16.
▲
by
junrushao1994
3y ago
Really depends on how good ROCm support for WSL2 is. Our team don't have a windows machine so could not verify ourselves, but if you got ROCm set up properly on WSL2, MLC LLM should work out of the box
17.
▲
by
junrushao1994
3y ago
ROCm has improved a lot over the past few months, and now ROCm 5.6 seems to work out of box by just following this tutorial: https://rocm.docs.amd.com/en/latest/deploy/linux/installer/i... . TVM Unit
18.
▲
by
junrushao1994
3y ago
yeah we tried out popular solutions like exllama and llama.cpp among others that support inference of 4bit quantized models
19.
▲
by
junrushao1994
3y ago
tbh im not sure what amds plan is on ROCm support on consumer devices, but i dont really think amd is being fraudulent or something. Both rocm and vulkan are supported in MLC LLM as mentioned in our blog post. we are aware that rocm is not
20.
▲
by
junrushao1994
3y ago
One of the authors here. Glad it’s on HackerNews! There are two points I personally wanted to make through this project: 1) With a sufficiently optimized software stack, AMD GPUs can be sufficiently cost-efficient to use in LLM serving; 2)
21.
▲
by
junrushao1994
3y ago
I don't think TVM advertised a lot on its full capabilities, for example, high-perf codegen for dynamic shapes without auto-tuning, or auto-tuning-based codegen, at least in the past few years, and that might be one of the factors it d
22.
▲
by
junrushao1994
3y ago
Yeah I believe countless new research and product ideas will be built on top of the open-source Llama-2
23.
▲
MLC LLM: 70B Llama-2-4bit on MacBook at 50%-80% speed of A100
(twitter.com)
12 points
by
junrushao1994
3y ago
|
3 comments
24.
▲
Running RedPajama and other open LLMs on phones, browsers and AMD/NV/Intel GPUs
(mlc.ai)
11 points
by
junrushao1994
3y ago
|
0 comments
25.
▲
MLC: Bringing Hardware Accelerated Language Models to Consumer Devices
(mlc.ai)
8 points
by
junrushao1994
3y ago
|
0 comments
26.
▲
by
junrushao1994
3y ago
Thanks for sharing! Sometimes LLMs do generate some weird stuff, but if the issue persists, please do report this to our github issues!
27.
▲
by
junrushao1994
3y ago
TVM Unity, the compiler used by MLC-LLM, does support CPU and SIMD instructions on each CPU backend via LLVM, but we haven't tried it out yet. I believe llama.cpp is the best option out of box at the moment.
28.
▲
by
junrushao1994
3y ago
As long as there is a Vulkan SDK for your AMD APU (likely there is), MLC-LLM can use TVM Unity to generate code for it
29.
▲
by
junrushao1994
3y ago
Thanks for the feedback! This is definitely something we need to do. To share some data, currently the default model is Vicuna-7b, aggressively quantized to 2.9G. We are expanding the coverage to more models, particularly, Dolly and StableL
30.
▲
by
junrushao1994
3y ago
This is our latest project on making LLMs accessible to everyone. With this project, users no longer need to spend a fortune on huge VRAM, top-of-the-line GPUs, or powerful workstations to run LLMs at an acceptable speed. A consumer-grade G
More ›