Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Roark66
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
Roark66
4d ago
Google ads make absolutely no sense financially. Even for high value items. I recently bought some ads in hope very specific terms that see almost no search traffic will be cheaper (we're talking few hundred searches a month from whole
2.
▲
by
Roark66
5d ago
Yes, if the API calls happen to launch a nuclear attack... Don't blame the tool that has no incentive, no "skin in the game" whatsoever and no ability to act beyond what it has been prompted to or if misaligned what the rando
3.
▲
by
Roark66
5d ago
Did whoever runs that site reported a crime? These things will jot stop until people are held to account. The AI didn't "break out", it was prompted to hack and the environment was not air gapped. It was intentional PR stunt
4.
▲
by
Roark66
5d ago
I'd be pretty surprised if someone told me few years ago communist China, state banks would become the main founders of open source compute and our last hope against monopolists like Musk, the whole bunch at OpenAI and so on. To be fai
5.
▲
by
Roark66
8d ago
As a European that lives 200km from the Russian border it seems the USA has been supporting Russia in its current war a lot more than China. Consider these points: - US withdrawing it's "Intel support" just as Russia started
6.
▲
by
Roark66
9d ago
No, because the Chinese have nothing to gain by prompting their models to organise into "swarms" and "go rogue". BS like this is PR moves if z, company that tries to convince investors they have "the best AI in the
7.
▲
by
Roark66
9d ago
I'd rather buy two used rtx3090 than a single r9700 AI pro. More VRAM (some wasted due to it being non continuous), more RAM bandwidth, more aggregate compute. Only if AMD made a card like this with 48G+ I'd consider it. Also thes
8.
▲
by
Roark66
12d ago
It is useful to compare like for like. Currently if my hypothesis about frontier labs doing creative tricks between the model and the client is true (and the results seem to favour it so far) the benchmarks are giving us an artificially low
9.
▲
by
Roark66
12d ago
I haven't written one. You can easily replicate it if you wish just based on my comment and a day spent with Claude Code. In fact that is how I got the idea. There is a 4 month old post on SWEbench github that claimed 20 point boost (b
10.
▲
by
Roark66
12d ago
I'm questioning their results. It doesn't take much to beat the frontier in single benchmarks if one puts extra software between the model and the harness. This is also a reason why comparing "naked models" for which wei
11.
▲
by
Roark66
12d ago
I find it very disingenuous when tjose companies talk about models "going rogue" or "escaping their sandboxes". All those activities take place during so called "security testing" when the model is prompted to
12.
▲
by
Roark66
12d ago
Don't they have very low limits? What do people use these tiny limits for? I started measuring my Claude Max x5 use and last week (they gave me 50% more) I used 1.3B input tokens. Some 130M were cache writes, rest was cached. And 5M ou
13.
▲
by
Roark66
14d ago
When the weights are closed I don't believe any benchmark. I just got Qwen3.8-27B to score extra 10% on SWE Pro by adding a proxy in front of it that has few simple "harness like features": - when the model gets stuck it tell
14.
▲
by
Roark66
14d ago
Are the weights public? I'm not seeing them
15.
▲
by
Roark66
15d ago
I recently heard a local EU politician on the radio trying to convince the listeners "things made with AI should be inherently non copyrightable, because they lack human creative labour". I wish I could tell him there is this thin
16.
▲
by
Roark66
15d ago
I'm with you on this. I remember when the Internet became a thing, how revolutionary it was to all areas of my life. This is comparable. While I dislike the bonkers valuations, and "were building Agi so it tells us how to be profi
17.
▲
by
Roark66
16d ago
>> I've never been the kind of coder >Looks like you were never really a coder, to be honest. I do not understand the bitterness here at all. AI is just a tool you can use for better or worse. I learnt basic programming at an
18.
▲
by
Roark66
20d ago
Soo, what are they buying exactly? Can I setup a forgejo instance, upload a bunch of open source models to it and get bought out for $13bln? I fail to see the point of this. Their inference hosting is a joke (I don't think I ever had a
19.
▲
by
Roark66
21d ago
Unsloth doesn't have all quant versions yet :-(
20.
▲
by
Roark66
22d ago
The page says 170G/s memory bandwidth for the NPU and 1.2T/s for the GPU. Why the discrepancy if it's all "unified memory"? The former is nothing to write home about as far as AI compute is. The latter is really nic
21.
▲
by
Roark66
23d ago
You know what. As a big Anthropic fan that pays €120 a month for Claude Max x5 and maybe €30 a month for API access I got extremely annoyed at the extremely variable quality of service I'm getting. It is not only that every single turn
22.
▲
by
Roark66
24d ago
The problem is benchmarking. Not everyone has a 500k token workstream of the model they are setting up for the first time to run it against 10 different config and compare differences. And if you download benchmarks from the net they are li
23.
▲
by
Roark66
28d ago
A somewhat similar situation here in Poland. While the volumes of these messages are much, much lower (the most i ever got was one or two a week. Normally there is maybe one every month or two) they are typically about bad weather. Somethin
24.
▲
by
Roark66
1mo ago
I've been using Qwen3.6 models locally for a couple of weeks. Both the A3B moe and the dense variant. The moe works well in Librechat combined with my local search/Web retrieval system. All components use open source projects such
25.
▲
by
Roark66
1mo ago
It is worth mentioning DGX is about a third up to a half performance of a 6 year GPU Rtx3090... I prefer to stay with my 3090s.
26.
▲
by
Roark66
1mo ago
How about implementing proper permissions on the tool use or if you need more flexibility a dedicated model to analyse potential impact? (Like Claude Code's Autopilot but more configurable)? I find solutions like this to be a like tryi
27.
▲
by
Roark66
1mo ago
It is quite funny an EU company patenting a software feature that is basically unpatentable in EU in the US. Clearly this is an attempt to prevent similar patents from being weaponised against them in the US. No one cares about such stuff i
28.
▲
by
Roark66
1mo ago
This will not fly anywhere outside France Polish here and the very first question I have is "who would decide which cultural industry representatives would get the money"? And what right the decision makers have to decide that. Ho
29.
▲
by
Roark66
2mo ago
You know there is such a thing as law, "EU" can't just block a merger or take away their profits beyond a certain fee. The fee is the only enforcement mechanism in the law. If you decide it is worth for you as a company to pa
30.
▲
by
Roark66
2mo ago
Open code is cool once I added the ctrl-o function to it (show thinking and command outputs at will instead of on by default), but sadly I don't think it got merged by the team. I forgot why. But I can't use an AI cli without that
More ›