Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
InvidFlower
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
InvidFlower
1mo ago
I think your last point is the main point. They're going hard at reenforcement learning to improve how good the models are at coding and such, but RL will make models cheat unless you're super careful. But being careful slows thin
2.
▲
by
InvidFlower
1mo ago
Yeah I think part of the problem is them optimizing so hard on coding and related tasks with RL. That's what will really encourage cheating and other misaligned things, because all incentives are to achieve the goal and they'll ch
3.
▲
by
InvidFlower
1mo ago
I think the much easier explanation than they intentionally hacked someone was just that they have super de-prioritized security and gotten very sloppy in the pursuit of improving the models as fast as they can, along with hubris of how the
4.
▲
by
InvidFlower
1mo ago
It sounds like the instance was shared for everything across the company, which as you said was super not good. But it's not just that.. it's that they didn't have enough monitoring to notice what was going on, even though th
5.
▲
by
InvidFlower
1mo ago
Sure, but you can say that about most things. Even for inventions from humans, usually it requires other people having already done a lot of work (hence why there's often inventions by different people at around the same time that didn
6.
▲
by
InvidFlower
1mo ago
Not just cyber, but apparently the message board stuff started with regular training and evals. It was a cyber test where HuggingFace got hacked, but all this other stuff was going on under OpenAI's nose for quite a while before that.
7.
▲
by
InvidFlower
1mo ago
Yeah.. feels like we're still so early in terms of effective training and evals. Like the official evals out there that have had so many instances of just plain incorrect questions. Or being incentivized to always answer instead of say
8.
▲
by
InvidFlower
1mo ago
Yeah, this is part of why I disagree with the "it's just PR" conspiracy theory stuff. Once you actually get into the details of what happened, there's no way it makes OpenAI look good lol.
9.
▲
by
InvidFlower
1mo ago
But part of the problem is if one actually does it quietly and it has already happened, then how would we know?
10.
▲
by
InvidFlower
1mo ago
Eh, I think this is past the point where they get more benefit than problems. Not even about the hack itself, but about so many mistakes and bad choices they made leading up to it.
11.
▲
by
InvidFlower
1mo ago
This is why I disagree with anyone claiming it is just marketing. It makes OpenAI look really really bad, like they have no idea what they're doing in terms of security. After the first board happened, they still didn't add better
12.
▲
by
InvidFlower
1mo ago
But it's interesting that the initial things that caused the board weren't even security evals, just normal office tasks. The actual hacking of Hugging Face happened during a security eval, but not all of the stuff leading up to i
13.
▲
by
InvidFlower
1mo ago
Don't forget it sounds like Artifactory was shared for the whole company and various agents pulled packages from it for everything from normal evaluations to actual model training. It might have been part of their normal to browse for
14.
▲
by
InvidFlower
1mo ago
Though one thing I've heard is that the base model is the one with the various possibilities for patterns, and then the reasoning takes advantage of those vs necessarily creating something new in that additional training. So even if it
15.
▲
by
InvidFlower
2mo ago
I did try Herdr but found too many things I expected it to have and just didn't. Right now I've been switching between two different approaches for combined local and remote work. There's the more traditional tmux/zellij
16.
▲
by
InvidFlower
5mo ago
Don't forget the employees doing the actual model training and research are not the same ones coding Claude Code. CC was a side-project by one employee that ended up hitting it big and is now one of the core parts of their income. They
17.
▲
by
InvidFlower
5mo ago
As the other person mentioned, they have said they are restricting third-party agent systems like OpenClaw and Hermes from using the monthly plan. But yeah, this seems like the wrong way to handle it, trying to detect those other harnesses
18.
▲
by
InvidFlower
5mo ago
I'm not sure if the context limit on the $25/m, and model-size limit on the $100/m would make it not work well enough for OpenCode, but Featherless AI seems a bit unique in terms of how they handle their inference plans.
19.
▲
by
InvidFlower
5mo ago
They've said publicly that they don't want apps like OpenClaw (Hermes is a variation) being used with a monthly plan vs per-token billing. The problem is this was implemented pretty badly (trying to regex??). And they should put a
20.
▲
by
InvidFlower
5mo ago
Yeah, at the least it should alert the user that it is happening. Maybe the thinking was alerting it gives people signal on how to get around the restrictions, but having it silently charge from a different bucket isn't the answer eith
21.
▲
by
InvidFlower
5mo ago
Also could run on a more generic cloud inference or gpu site. At least to see how well it works for your use-case before spending on hardware.
22.
▲
by
InvidFlower
5mo ago
For the topic of remote control, Happy seems to be working pretty well for Claude Code but is also supposed to support Codex. It's a bit rough around the edges, but nice that it is open source: https://github.com/slopus
23.
▲
by
InvidFlower
5mo ago
I'm not so sure about that. Like we're using Claude Code with Bedrock and have most things on AWS with SOC2 compliance and all that. Normally switching to Codex would have a ton of friction in terms of separate contract and billin
24.
▲
by
InvidFlower
5mo ago
No, you definitely don't have access to the weights. The raw weights are secret enough that when a model hits a certain level of capability, their guidelines are that they need enough procedures in place to try to keep nation states fr
25.
▲
by
InvidFlower
5mo ago
Besides what the other person mentioned about being more useful for enterprise, I also heard mentioned on a podcast that gpt-image-2 uses the same general architecture as the LLM models, while Sora was a very different architecture. So they
26.
▲
by
InvidFlower
5mo ago
Given the topic of this article, we've been using Claude Code pointed at Bedrock and have never had any scaling issues. Obviously it is more expensive paying by the token than a monthly plan, but I sometimes have 2 or 3 instances chugg
27.
▲
by
InvidFlower
2y ago
While that is cool in principal, I'm not sure how well it'd actually work in reality. First, there is the technical challenge. My understanding is the weights can have a lot of fluctuation, especially early on. How do we actually
28.
▲
by
InvidFlower
2y ago
There's been some work on memory lately like Transformer² and Titans. But that may not be necessary for decent agents. Even the "context in a loop" is getting better as general reliability of tasks increases, and as people fi
29.
▲
by
InvidFlower
2y ago
I think you'll get a lot of use from this video from Andrej Karpathy: https://www.youtube.com/watch?v=7xTGNNLPyMI It is long, but don't get scared off. He goes over a ton of different stuff related to model traini
30.
▲
by
InvidFlower
2y ago
We've already seen Qwen's new QWQ 32B (not distilled) model doing impressive things on benchmarks. It'll definitely be interesting to see how just good small models can get. When combined with rag and large context window for
More ›