Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rohansood15
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
rohansood15
16d ago
Wishing you guys the best, but this isn't a very compelling answer. Would be helpful to provide more concrete or quantifiable examples. In the early stages when pitching to users, it's preferable to state exactly what can be done
2.
▲
by
rohansood15
21d ago
Because labs can learn to optimize inference post launch, plus can move to use bigger/better clusters depending on demand. It is not impossible to imagine Qwen cuts prices further with QAT/MTP-like improvements.
3.
▲
by
rohansood15
21d ago
Given the new architecture, speeds are harder to estimate. Max would be ~40 tok/s.
4.
▲
by
rohansood15
21d ago
This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.
5.
▲
by
rohansood15
21d ago
128GB, 4-bit quantized.
6.
▲
by
rohansood15
21d ago
For world knowledge, you'd want it to find and reference the source material to be sure. At that point, it doesn't matter if the knowledge is embedded.
7.
▲
by
rohansood15
21d ago
Didn't expect it to beat 3.8 27B so cleanly. Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.
8.
▲
by
rohansood15
1mo ago
So an engineer who knows your codebase and tests can sneak in malicious code/backdoors because you're high trust. And I am guessing you'll extend that high trust to your agents/'orbs' next. And I am sure you&#x
9.
▲
by
rohansood15
2mo ago
Companies have learned their lessons on stickiness with cloud providers. Every enterprise has a multi-provider strategy now.
10.
▲
by
rohansood15
2mo ago
I agree with you, but Codex is open source.
11.
▲
by
rohansood15
2mo ago
I feel similarly. But texting became the primary form of inter-personal communication in the last couple of decades, in part because that was the primary modality technology could handle. So, now that we can talk to computers, maybe we will
12.
▲
by
rohansood15
3mo ago
Or Codex models are more efficient that Claude. Plus the two things you mentioned.
13.
▲
by
rohansood15
3mo ago
Cyber-attacks to start, and real-world terrorist attacks/bombings (inc. chemical/biological weapons) later.
14.
▲
Fable situation update from David Sacks
(twitter.com)
11 points
by
rohansood15
3mo ago
|
26 comments
15.
▲
by
rohansood15
3mo ago
I mean, we all pay via CC so it's bit like they can't know who you are if they wanted to.
16.
▲
by
rohansood15
3mo ago
It is only abuse flagged data and there too for OpenAI they're not sharing that data with them. But for Anthropic they are.
17.
▲
by
rohansood15
3mo ago
Pretty sure this doesn't work for any regulated enterprise or government client. But AWS knows this, so I am curious why they'd agree to it.
18.
▲
by
rohansood15
3mo ago
I can't find it. Can you state your performance versus comparable 3-bit quantization from Unsloth/Bartowski? Edit: I appreciate that you seem to have open-sourced the quantization pipeline. This is not to question your work, but t
19.
▲
by
rohansood15
3mo ago
Have you benchmarked against other 3-bit dynamic quants like Unsloth? I am sorry but this framing against a full precision, newer, smaller MoE just seems misleading. Also, Gemma-4-26B-A4B is not the SOTA for edge. Even at launch, that woul
20.
▲
by
rohansood15
4mo ago
Anthropic better get that IPO out soon. Their incredible revenue run-up was basically a result of botched Gemini releases and OpenAI having their hands-tied behind their Azure backs. Anthropic models were quite literally the only viable ser
21.
▲
by
rohansood15
4mo ago
Are you comparing single-user requests or multiple concurrent requests when you say comparable to rented GPU? Most of the cost efficiencies kick in with concurrent/batch requests. A single H100 node can provide like 5k input + 2k outpu
22.
▲
by
rohansood15
4mo ago
So ask for it. Seems like your issue isn't immigration, it is abuse. The recent changes don't do much to fix that, imo.
23.
▲
by
rohansood15
4mo ago
Let's say the government can't care for 100M people because of lack of doctors. Now they could train one over 10 years, or you could have one of the smartest doctors in the world come be 100M+1. Would you take that? Now expand tha
24.
▲
by
rohansood15
4mo ago
Subjective, but if we compare to compute not everyone needs the most expensive laptops or super computers for their work. I think frontier models will be invaluable for scientific research, defense, financial analysis and such. But the aver
25.
▲
by
rohansood15
4mo ago
Which part of this is a 'prediction'?
26.
▲
by
rohansood15
4mo ago
I don't think I follow?
27.
▲
by
rohansood15
4mo ago
Thanks for the info. Daniel fixed it - and no it wasn't an LLM error. :P
28.
▲
by
rohansood15
4mo ago
I used the same assumptions as the original HN post https://news.ycombinator.com/item?id=48168198
29.
▲
by
rohansood15
4mo ago
The title auto-corrected, my post was 'less' not 'more'.
30.
▲
by
rohansood15
4mo ago
Nope, HN changed the title. https://imgur.com/a/UgJqWEh
More ›