Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
glub
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
I gave Microsoft for Startups the wrong address on purpose
(lubaretsi.com)
3 points
by
glub
23h ago
|
0 comments
2.
▲
by
glub
3d ago
Dario: Mr. Trump, we don't want to advance anymore, but China very bad, if we stop, China take over. Please stop China, then we rest. Trump: you're not in the EU, get to work. Jokes aside, I'm very happy that Dario's att
3.
▲
by
glub
3d ago
> That the answer is to just leave AI labs to carry on as they are? Yes, and no. I do not think there's a need for special treatment of LLMs. But the labs certainly shouldn't be allowed to go around hacking things on the intern
4.
▲
by
glub
3d ago
They may be genuinely concerned, but that's beside the point. You may think a technology holds too much power for someone to wield it, and therefore come to conclusion that you, the benevolent, the fluffy, the unicorn, with your 3 unic
5.
▲
by
glub
3d ago
>In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. This my experience also. I had an issue with my Unraid server, so I had an agent running on my machine fig
6.
▲
by
glub
3d ago
I'm not saying that we should ban any models or that such bans would be effective - for the reasons that you've outlined that they're counter productive, and as a principle, I don't think government should have any say i
7.
▲
by
glub
3d ago
You're right. It's my opinion that if your sandbox has a path to the internet, it is not a sandbox, it's a gimmick. And the 2 other incidents with OAI/ANT had the same issue, but it's even funnier - sandbox in those
8.
▲
by
glub
3d ago
By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.
9.
▲
by
glub
3d ago
Open research and open weights from China are not contributions to China only. If you can secure compute, there's a whole lot you can do as a US firm with this research and weights. So it's a simple strategy: 1. Ban big players fr
10.
▲
by
glub
3d ago
And they've been trained on user data where users have been trying to set up effective coordination flows since the very first harness.
11.
▲
by
glub
3d ago
Even if you take out the LLMs out of the equation, it's at the very least a negligence. Model didn't escape a sandbox, as there was no sandbox.
12.
▲
by
glub
4d ago
I think what matters in this case is how proactive and greedy the model is. GPT models are extremely proactive and gredy. So when Fable mentions something that may affect some obscure component of the system, GPT will start digging the code
13.
▲
by
glub
4d ago
oh-my-pi. I was using my own homegrown (mega slop) harness for a while, but it distracted me from working on my actual projects, and I realized oh-my-pi was doing the same things I've been doing, including advisor, native server-side c
14.
▲
by
glub
4d ago
I'm the opposite. Every time I've let Sol/Astra be decisive, I ended up with an overengineered mess. I much prefer getting alerted when there's more than 1 approach to the problem and it's discovered mid-implementat
15.
▲
by
glub
4d ago
I'm surprised Sol and Astra are leading "Unverified assumption" metric and Fable is better there. I run Fable as my main model with Sol as advisor that watches every turn. Fable likes to throw around assumptions that it didn&
16.
▲
by
glub
4d ago
> This year it became common for people to entirely delegate coding to AI This has been the case for around 2 years now, more reliably - a year. We've mostly stayed there since then. Saying that more people started doing it isn'
17.
▲
by
glub
4d ago
As long as user provides inputs and LLMs stay LLMs, you can waltz through any guardrail. Fable is the extreme case, but it's not that hard if you know what you're doing and know how LLMs and their guardrails work. Am I saying that
18.
▲
by
glub
4d ago
Tell me you haven't tried letting Astra go without telling me. Astra can confidently one-shot 500k lines of slop, with 800k lines of tests covering it, without testing a single intended product requirement, and none of it actually work
19.
▲
by
glub
4d ago
Anthropic could serve Opus 4.5 from a year ago under opus:latest and most heavy users would probably have no idea. Some of them would probably even prefer it. Yes, I do use them, quite heavily. The only difference at this point is in benchm
20.
▲
by
glub
4d ago
The attention is shifting towards RL, harnesses, and memory systems from the pretrains of more intelligent and capable base models. So extracting additional capabilities from what we already have. That is a much easier catch up game. GLM 5.
21.
▲
by
glub
4d ago
Dario, Sam, and Elon are all on the same page on this. So it's either they truly think AI is going to kill us all, or there's some other motives at play here. I don't think these people could possibly agree on the color of th
22.
▲
by
glub
4d ago
I don't think people care as much as we'd like them to care about dangers of technologies. It takes a single step outside of technological bubble to see that their opinion of SOTA LLMs is vastly different. To them, AI means ChatGP
23.
▲
by
glub
4d ago
So what's the fuss about then? Is the idea that a Chinese company will release a model that will have no safeguards? For what purpose? Basic safeguards are all that's required, and they've been there in every usable model sin
24.
▲
by
glub
4d ago
I love watching NileRed/NileBlue on youtube - crazy chemist that does a lot of insane stuff. While it's obvious that he's very smart and probably way above the average, his education is still nothing that probably millions of
25.
▲
by
glub
4d ago
And all of this information has been available on the open web way before LLMs. I'd wager it's likely easier for an average person to do this with Tor browser than it is to get an LLM to help them with it. Even ones that Dario cal
26.
▲
by
glub
4d ago
> We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top > Transparency. Regardless
27.
▲
by
glub
5d ago
I have several $200 subscriptions as a developer/founder. I used to blow through all of their limits when the limits were quite high. As I progressively learned the limitations, and what to make of them to get useful results, I may be
28.
▲
by
glub
7d ago
I recently attempted to get on Azure because of their startup package, and I fail to understand how Microsoft holds anything together at all. Docs are a mess with contradictory information, coordination between teams is a mess, azure UI is
29.
▲
Don't Let Architecture Astronauts Scare You
(joelonsoftware.com)
4 points
by
glub
7d ago
|
1 comments
30.
▲
by
glub
9d ago
> But no, let's in fact choose the stupidest possible way of doing it — by rendering everything as HTML commit by commit and then parsing it. This drives me crazy with so-called SOTA LLMs that have "achieved AGI". Fable, S
More ›