Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
briansun
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
briansun
1y ago
Thanks for the view from a very privacy‑sensitive environment — agreed that hosted SOTA still leads on broad capability. Could you share a quick split: which tasks truly require hosted SOTA than open‑weight? I think gpt-oss is quite good fo
2.
▲
by
briansun
1y ago
Wouldn't it be cool to have a local AI agent? It could access search engines and browse any website through a headless browser.
3.
▲
by
briansun
1y ago
Thanks — I agree with your three big pain points: quality vs hosted SOTA, token speed, and economics/utilization. Have you run into cases where on‑device still makes sense? 1. Data that is contractually/regulatorily prohibited fro
4.
▲
by
briansun
1y ago
Well put. Management overhead + unclear capacity planning kills many pilots.
5.
▲
by
briansun
1y ago
Totally fair. On a normal laptop you also need headroom to do your actual job, and KV cache + context length can eat that quickly.
6.
▲
Building Privacy-First AI Agents on Ollama: Complete Guide
(nativemind.app)
2 points
by
briansun
1y ago
|
0 comments
7.
▲
Ask HN: Why aren't local LLMs used as widely as we expected?
5 points
by
briansun
1y ago
|
10 comments
8.
▲
by
briansun
1y ago
Thanks for raising the privacy angle. Do you have a source or plan details for the 30‑day retention and the lack of deletion options (non‑enterprise)? It would help to know account tier and where that policy is documented. Beyond policy, ho
9.
▲
by
briansun
1y ago
Gemma3n as a daily driver sounds nice—4b or 8b? and rough tokens/sec on your laptop? And have you A/B‑tested code generation quality across local models (e.g., Gemma3n vs others)?
10.
▲
by
briansun
1y ago
Thanks-this is genuinely encouraging; I'd assumed AI help was strongest on front-end work(web apps/SwiftUI), so this is my first concrete example of an LLM catching memory‑unsafe C/C++-could you share your toolchain (CLI/
11.
▲
by
briansun
1y ago
Super useful config dump—thanks. Do you have wall‑clock numbers for prefill/gen tokens/sec and power draw on the 24GB card for those three setups? Also curious where quality starts to degrade vs. context length in your tests.
12.
▲
by
briansun
1y ago
Haha, a cute pet dragon. Two knobs that helped me tame VRAM: KV‑cache quant/eviction and sliding‑window attention (if your runtime supports them). What model/runtime and context are you running when it tips over? Are you using Oll
13.
▲
Ask HN: Are you running local LLMs? What are your key use cases?
16 points
by
briansun
1y ago
|
15 comments
14.
▲
19 Years on One Product: Things Xmind Taught Me
1 points
by
briansun
1y ago
|
1 comments