Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mezark
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
mezark
2mo ago
lol - co-founder of Doubleword here. honoured you think we have sophisticated enough marketing to 'newsjack'. What actually happened is my cofounder wrote it over the weekend because he's a mega-nerd and put it live yesterday
2.
▲
by
mezark
2mo ago
And a PhD in Quantum Computing! I'm a physicist so a fan of bra-ket tbh
3.
▲
by
mezark
2mo ago
Another banger
4.
▲
by
mezark
3mo ago
really like the framing of this post :)
5.
▲
What happens when you run a CUDA kernel?
(fergusfinn.com)
294 points
by
mezark
3mo ago
|
32 comments
6.
▲
A running list of reasons to move to open source
(whyopensource.ai)
6 points
by
mezark
3mo ago
|
0 comments
7.
▲
by
mezark
3mo ago
I love this blog
8.
▲
by
mezark
3mo ago
(As someone who cares a lot about philosophy of consciousness / & cogsci) The whole point of consciousness being a 'hard problem' is that we just cannot make claims like 'X is not conscious'
9.
▲
by
mezark
4mo ago
we think so - but haven't tested it ourselves
10.
▲
by
mezark
4mo ago
Hi! Co-founder of Doubleword here - we've hugely increased the number of models that we offer (partly thanks to work that we've done on hotswapping https://blog.doubleword.ai/fast-sglang-starts . We're kind of
11.
▲
by
mezark
4mo ago
We at doubleword are bullish for AMD for low-interactivity inference - it does just take a bigger lift on the software side...
12.
▲
Moe inference optimizations: 15% lower expert load by request reordering
(blog.doubleword.ai)
3 points
by
mezark
4mo ago
|
0 comments
13.
▲
by
mezark
4mo ago
If you're talking about UK sovereign LLM inference you need to mention Doubleword... very serious inference optimization lab in london with public endpoints for OS models
14.
▲
Tensor Network Attention
(mainlymatmul.com)
2 points
by
mezark
4mo ago
|
0 comments
15.
▲
Redundant Information in LLM Weights
(fergusfinn.com)
5 points
by
mezark
4mo ago
|
0 comments
16.
▲
Tans: Precomputing RANS
(fergusfinn.com)
3 points
by
mezark
5mo ago
|
0 comments
17.
▲
Also-RANS: Asymmetric Numeral Systems for Entropy Coding
(fergusfinn.com)
25 points
by
mezark
5mo ago
|
0 comments
18.
▲
70x faster cold(ish) starts for SGLang
(fergusfinn.com)
4 points
by
mezark
5mo ago
|
0 comments
19.
▲
QueueSpec – drafting speculation tokens while a request queues
(blog.doubleword.ai)
1 points
by
mezark
8mo ago
|
0 comments
20.
▲
ZeroDP: Just-in-Time Weight Offloading over NVLink for Data Parallelism
(mainlymatmul.com)
1 points
by
mezark
8mo ago
|
0 comments
21.
▲
Parallel Primitives for Multi-Agent Workflows
(fergusfinn.com)
1 points
by
mezark
8mo ago
|
0 comments
22.
▲
New fastest AI Model Gateway – 450x less overhead than LiteLLM
(github.com)
2 points
by
mezark
11mo ago
|
0 comments
23.
▲
Should GPUs Make Free Trade Agreements?
(doubleword.ai)
3 points
by
mezark
1y ago
|
1 comments
24.
▲
by
mezark
1y ago
We look at how comparative advantage from economics applies to LLM inference - some GPUs are relatively better at FLOPs, others at memory bandwidth. What happens if you let each do what it’s best at?
25.
▲
by
mezark
2y ago
Huge congrats - and when you look at the latency graphs as well it really shows the value of these specialised systems!
26.
▲
Controlled generation of OS LLMs – without impacting latency
(youtube.com)
7 points
by
mezark
3y ago
|
1 comments
27.
▲
by
mezark
3y ago
TitanML Takeoff Inference Server demonstrating controlled generation
28.
▲
Takeoff Inference Server Is Now Open Source
(github.com)
3 points
by
mezark
3y ago
|
1 comments
29.
▲
by
mezark
3y ago
Drop in replacement for HF's TGI server. The fastest and easiest way to inference LLMs locally Github: https://github.com/titanml/takeoff Docs: https://docs.titanml.co/docs/titan-takeoff/
30.
▲
by
mezark
3y ago
Hey there - TitanML is these guys: https://www.titanml.co/ . I think the impressive thing isn't actually whether the model is good (although it is a good model especially when fine-tuned) - but how fast this model runs
More ›