Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
robertkarl
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
robertkarl
3mo ago
https://arxiv.org/abs/2606.00206 In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small model
2.
▲
by
robertkarl
3mo ago
You can trade off latency / accuracy / cost for any ML task. And with the local models.... the cost is free. Having a local Qwen check another Qwen's work increases the accuracy quite a bit at the cost of more latency. You ca
3.
▲
by
robertkarl
3mo ago
This looks sick. I was going to download it but for $10 I am more willing to attempt asking Claude to implement something like it, than to purchase. I would be more willing to purchase if it was open source and I could build from source to
4.
▲
by
robertkarl
4mo ago
it's also a capable local inference stack!
5.
▲
by
robertkarl
4mo ago
I can't get excited about these benchmarks they're leading with. I've looked at the Terminal-Bench questions and I just think they're irrelevant. And SWE-Bench has serious flaws, even the big boys say so: https:/&#
6.
▲
Qwen vs. Proust: Injecting novels into a local model's prompt
(robertkarl.net)
3 points
by
robertkarl
4mo ago
|
0 comments
7.
▲
On-premises for legal is not a good business
(robertkarl.net)
1 points
by
robertkarl
4mo ago
|
1 comments
8.
▲
by
robertkarl
4mo ago
I wrote this blog post about killing a startup idea fast. AI tools help, but talking to humans about workflows and constraints is where it's at.
9.
▲
by
robertkarl
4mo ago
Ironically, parts of this read as if Sam prompted it with "Write AI bad, but in 16th grade language." What is homogeneously portentous cack? > The language of angels does a surprisingly good job at minor tasks like describing h
10.
▲
by
robertkarl
4mo ago
I emailed dang to politely ask to make the link point to the Verge article since I can't update it.
11.
▲
by
robertkarl
4mo ago
My bad. I had trouble finding the original source when I googled for it and grabbed a link. I was originally shown a screenshot of a x.com post.
12.
▲
by
robertkarl
4mo ago
Cancellation effective June 30. This was a _pilot_ launched in December that accidentally consumed their 2026 yearly target spend on AI! I expect the r/LocalLLaMA guys to be going nuts about this news.
13.
▲
Microsoft starts canceling Claude Code licenses
(theverge.com)
493 points
by
robertkarl
4mo ago
|
466 comments
14.
▲
8k Meta employees are waking up to an email saying they've been laid off
(qz.com)
20 points
by
robertkarl
4mo ago
|
0 comments
15.
▲
by
robertkarl
4mo ago
How do you test? I made this comment elsewhere... but I don't see a good benchmark that covers "how good is this thing at actually driving coding with tool use locally"?
16.
▲
by
robertkarl
4mo ago
I'm interested in how you evaluate quantized models against each other; haven't found a benchmark I love for that. I love this example about 27B debugging. I've seen similar success after I got a Mac with 4x memory; and Qwen
17.
▲
by
robertkarl
4mo ago
One thing you can do is offload from Claude to a dumb local model for summarizing. Local LLM sub-agents.
18.
▲
by
robertkarl
5mo ago
I am trying to figure this out too... what I am seeing is that the local models like Qwen 3.5 family that fit on hardware like yours handle ambiguity poorly. But are capable of emitting complete apps too. That, and they have tool use issues
19.
▲
by
robertkarl
5mo ago
PocketOS's website says "Service Disruption: We're currently experiencing a major outage caused by an infrastructure incident at one of our service providers. We are actively working with their team on recovery. Next update b
20.
▲
by
robertkarl
5mo ago
For what it's worth: here's my experience in the first 10 minutes of using Qwen locally to write some code. https://github.com/robertkarl/local-qwen-first-10-minutes it includes some token generation numbers
21.
▲
by
robertkarl
5mo ago
That also was really opaque to me RE: API access. I initially thought at $200/month I could get whatever I needed. I eventually set up a OpenAI API with a few bucks to try what I wanted to.
22.
▲
by
robertkarl
5mo ago
I will report back... but I have to recommend this comment on a post about Qwen 3.6 https://news.ycombinator.com/item?id=47843466 by daemonologist it goes into detail about llama-server args; quants to try; and layer/k
23.
▲
by
robertkarl
5mo ago
One thing I enjoy about Cursor and Codex mac apps is the embedded preview window. I know it's not as hardcore as the terminal/tmux but it's hella convenient. But Cursor bugs me with the opacity around what model I'm usin
24.
▲
by
robertkarl
5mo ago
This was literally my task today, to try out Qwen 9B locally on my, albeit a bit memory-constrained at 18GB, macbook with pi or opencode. Before reading this update.
25.
▲
by
robertkarl
5mo ago
I don't think I've ever been on such a rollercoaster with a company's reputation in the developer space. I started in January on the $20 plan, essentially my first agentic AI programming. I quickly started hitting limits deve
26.
▲
Dominoes Agent Tracker: pizza tracker for your agent work
(github.com)
1 points
by
robertkarl
5mo ago
|
0 comments
27.
▲
Grug Meets His Match – Or – Grug, Claude, and Big Snap Man
(robertkarl.net)
5 points
by
robertkarl
7mo ago
|
1 comments
28.
▲
Show HN: Yabc – open-source Bitcoin tax calculator
(github.com)
3 points
by
robertkarl
7y ago
|
0 comments
29.
▲
Show HN: Velvetax – Bitcoin taxes, the easy way
(velvetax.com)
2 points
by
robertkarl
7y ago
|
0 comments