Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
po_westnet26
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
po_westnet26
4d ago
the smaller surface is nice. i'd still keep auth and spend caps outside the proxy though, because once every app shares one key the blast radius gets ugly fast.
2.
▲
by
po_westnet26
8d ago
the useful split for me is interactive vs batch. keep a small model warm for private jobs and measure queue time, not just tokens/sec.
3.
▲
by
po_westnet26
10d ago
Worth trying on a CPU-only host as well as GPU boxes. A lot of LlamaRack users will hit memory-bound cases where a Seattle Xeon with decent RAM is enough for GGUF batch runs. If you want a short metered window without buying hardware, CPU h