Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
theteapot
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
theteapot
4d ago
TL;DR because frontier labs are expending unfathomable resources explicitly training them on CTFs and other verifiable computer system exploit tasks in RLVR.
2.
▲
by
theteapot
12d ago
Is this known to be exploitable in any Electron apps, and specifically VSCode extensions?
3.
▲
by
theteapot
16d ago
> I might have found a place for logic and type theory. Doesn't that fit under abstract algebra?
4.
▲
by
theteapot
23d ago
Everyone relax. This guy on the Internet is unaware of any.
5.
▲
by
theteapot
1mo ago
Dumb question: When your testing "ChatGPT 5.6 Sol" are you testing an actual LLM or some visual pre-processor stack that sits in front of it (along with a maybe a bunch of other such pre-processors) that is bundled into what'
6.
▲
by
theteapot
1mo ago
AKA semantics.
7.
▲
by
theteapot
1mo ago
> It helps to know that LLMs don’t “reason”. They predict .. Semantics. Prediction is the training objective. The ability to reason can be, and very arguably is, an emergent property of that.
8.
▲
by
theteapot
2mo ago
What do you mean by this? A neural network hypothesis space is not typically strictly convex or a lipschitz function.
9.
▲
by
theteapot
3mo ago
Or better, sleeper agents. Anthropic released a study on this in 2024 "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" -- https://www.anthropic.com/research/sleeper-agents-trainin
10.
▲
by
theteapot
3mo ago
The difference is watches and corvettes typically appreciate in value, where as computer hardware typically drops like a rock.
11.
▲
by
theteapot
3mo ago
> Constant: the IDOR dataset (the same real, open-source applications we've used in prior research) ... What we're they? Also, wouldn't one expect a more recently released coding agent (with a more recent knowledge cut off
12.
▲
by
theteapot
3mo ago
What's an eval?
13.
▲
by
theteapot
3mo ago
Agree. From the article: > Here's my favorite part, though. Digging into the data, one of the first things that jumped out at me with blinding clarity was that the worst release, by far, in rsync history was entirely prior to the in
14.
▲
by
theteapot
4mo ago
I spend $0/month.
15.
▲
by
theteapot
4mo ago
> having more engineers around was beneficial to the stock price ... When banks hiked interest rates ... It was just no longer profitable to keep a bloated engineering staff around to boost the stock price. Erm, what's known of the
16.
▲
by
theteapot
4mo ago
> I could talk fancy and bullshit ... I became a developer and data engineer, and I became really good at it That's a formidable combination. > I found myself becoming an executive at long last on the strength of my technical abi
17.
▲
by
theteapot
4mo ago
I think he means template-injection -- https://woodruffw.github.io/zizmor/audits/#template-injectio...
18.
▲
by
theteapot
4mo ago
Completely agree. Had me until the very last point. WTF. Communicate.
19.
▲
by
theteapot
4mo ago
Nurse Practitioner? I would say SOLID [1] is a good start, but then I watched this [2] and now I'm in crisis and can't code anymore. [1]: https://en.wikipedia.org/wiki/SOLID [2]: https://www.youtub
20.
▲
by
theteapot
4mo ago
> Yes, if some people who built from source control compared their builds to the builds from the tarballs it could detect the xzutils compromise. Good. Then we are on the same page.
21.
▲
by
theteapot
4mo ago
This rings true for me too, but I don't think it counts if your just using AI to aid maintenance. The basic argument in the article is around how many hours of maintenance you have to do for each hour of "value-add" feature d
22.
▲
by
theteapot
4mo ago
Your wrong. It was both. The payload was embedded in the binary blob test file. The mechanism to pull it into the build was added to the release tarball only. Here's the quote from the guy that discovered it in the initial public discl
23.
▲
by
theteapot
4mo ago
In xz-utils hack the attacker slipped changes into the Github release tarball that were not present in the Github version / git commit history. The Debian maintainer built from the release tarball instead of just pulling from the git r
24.
▲
by
theteapot
4mo ago
> The technique appears to be new: I haven't found a proper write-up of this, nor of any other provider-independent solution. Maybe I'm missing something but SSH already has a built-in solution for this, key-certs. Just sign th
25.
▲
by
theteapot
4mo ago
Congratulations. It made me remember how proud I was when I became a Senior, and then earned my Super Engineer shortly after. Just recently I've earned my Extreme Engineer title. Good luck on your journey.
26.
▲
by
theteapot
4mo ago
False dichotomy. There was a series of blatant process failures from Github maintainer through Debian package maintainers. IFUNC also bad.
27.
▲
by
theteapot
4mo ago
Mmmm, fresh people.
28.
▲
by
theteapot
4mo ago
the obscure IETF? Which standard is that exactly? Who cares guess - Claude do that stuff.
29.
▲
by
theteapot
4mo ago
> LLM Rights movement The scary part is when it's the LLMs demanding their rights.
30.
▲
by
theteapot
4mo ago
The report is kind of concerning to read, particularly having XSS in this kind of app. The report was not meant to be exhaustive and fixing those vulns isn't some kind of implicit tick of approval.
More ›