Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy12_
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
andy12_
6d ago
They need to be loaded into shared memory. The weights might fit in global memory if the VRAM is big enough, but they still need to be moved to shared memory for computation.
2.
▲
by
andy12_
7d ago
Meanwhile, my job commute is a 30 minute walk to the train station or... a 30 minute bus trip to the station (yeah, taking a bus literally saves no time at all). Plus a 50 minute train ride plus another 30 minute walk. Honestly, I would 100
3.
▲
by
andy12_
13d ago
That's a terrible metric, because people going on a vacation probably aren't going there purposefully to commit crimes. What you want to do is look for increases in crime in a given place during holidays https://coolidg
4.
▲
by
andy12_
14d ago
> It sounds more like the models did close to what they were told to do Absolutely not. If I tell a kid to "Get good grades on the next math test" I don't expect the kid to try to kidnap their teacher to extract the next q
5.
▲
by
andy12_
14d ago
I don't want someone to blame. I want agents to be aligned by default. Their good behavior shouldn't depend on all users at all times using them correctly, because everyone will not just[1] use them correctly at all times. > If
6.
▲
by
andy12_
15d ago
> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. Great, so we can basically ignore AI alignment altogether and
7.
▲
by
andy12_
1mo ago
> I think the fundamental principle is that this approach messes with the distribution in ways that deviate from the trained model. But it doesn't! The distribution doesn't change at all. The only thing that changes is that sa
8.
▲
The Obsessed Encoder
(enigma.inc)
1 points
by
andy12_
2mo ago
|
0 comments
9.
▲
Nvidia Vera Rubin NVL72 measured to have 10x TPS per megawatt than Blackwell
(twitter.com)
4 points
by
andy12_
2mo ago
|
0 comments
10.
▲
by
andy12_
2mo ago
The automated AI pipeline also had an automatic grading model to try to reduce false positives. But anyway, my point was that in that case the prompt involved was indeed pretty much "hey, ChatGPT, solve an unsolved problem, thanks.&quo
11.
▲
by
andy12_
2mo ago
> but it is worth noting that this wasn't a matter of "ChatGPT, solve this unsolved problem. Make no mistakes." It wasn't the case for this, but when OpenAI disproved the Unit Distance Conjecture, it was really done a
12.
▲
by
andy12_
2mo ago
When Google Maps routes me using a smaller secondary road instead of the main road that I would otherwise have used , I've always wondered whether that significantly changes the amount of traffic that smaller road sees. It's funny
13.
▲
by
andy12_
2mo ago
> Even interns can understand ambiguous asks with a bit of help This is not a case of an ambiguous task. This is literally trying to judge a model based on information it cannot possibly know, like trying to judge someone based on whethe
14.
▲
by
andy12_
2mo ago
It's pretty much confirmed by OpenAI here [1]. > We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. And Gem
15.
▲
by
andy12_
2mo ago
No, GPTCyber is specifically trained for cybersecurity, and GPT-5.5-pro is just an ensemble of many subagents, not an actual model. Mythos is simply a much bigger model in terms of parameters and I don't think OpenAI will have anything
16.
▲
by
andy12_
2mo ago
I think what's unexpected is that it seems that some cases of model errors are truly caused by the model being misaligned? In the "Catching a model fabricating data" example I would have thought that it was just the model bei
17.
▲
by
andy12_
3mo ago
I think it makes more sense to make it so that major versions are different pretraining runs, and minor versions are simply the same pretraining run that was finetuned to different degrees. But it seems that that isn't cool anymore.
18.
▲
by
andy12_
3mo ago
I mean it as in, train a model across different clusters instead of a centralized cluster. It's been shown that it's possible to train 10B models this way. If more research effort was put into this, that would be great I don'
19.
▲
by
andy12_
3mo ago
To be fair. There is a security concern angle: even open-source models could be trained as sleeper agents that act adversarially (for example, adding backdoors) when used in specific national companies in specific settings. This is very dif
20.
▲
by
andy12_
3mo ago
I'm from Spain and I also hate these projects with passion. Creating models that speak multiple languages is a solved problem. Having each European Nation train its own useless "sovereign model" in its own language is a total
21.
▲
by
andy12_
3mo ago
This is making me extremely depressed. If this was coming from Anthrohpic I would just need to wait for OpenAI to drop a similar model. But if this comes from the US government, they will do the same to OpenAI when the moment comes. Similar
22.
▲
by
andy12_
3mo ago
I don't know if you are aware, but some people reported in Twitter that Fable 5 may flag the message regardless of content if it knows (from either pretraining knowledge or memories) that you work in either of those fields. I don'
23.
▲
by
andy12_
3mo ago
> Performance on benchmarks has practically leveled off Ehm, no? DeepSWE[1] for example shows that new models like gpt-5.5 continue to show big improvements compared to older models. > Also prices are going up. Prices for frontier int
24.
▲
by
andy12_
4mo ago
Claude can indeed decide to terminate conversations on its own using a special tool[1] if it feels "uncomfortable" with how the conversation is going. Also, very famously, in the middle of recording Computer Use demos, Claude stop
25.
▲
by
andy12_
4mo ago
You don't get it. A human set up a software system allowing spicy autocomplete to solve open math problems if the appropriate keyword appears in its output.
26.
▲
by
andy12_
4mo ago
I skimmed through the paper completely expecting polite prompts to do better, and when I saw table 2 I lost it hahahahaha. The rude prompts are specially funny. I mean: > You poor creature, do you even know how to solve this? > Hey go
27.
▲
by
andy12_
4mo ago
Someone blatantly copied their tutorials but ChatGPT is to blame, somehow? The accusation here isn't even that ChatGPT learned from their tutorials and then generated them verbatim. The accusation is that someone copied the whole artic
28.
▲
by
andy12_
4mo ago
> Was the question asked by a mathematician? As per the report, the prompt used to solve the problem is AI-written and the solution was initially graded by an AI grading pipeline. They don't say this explicitly, but it seems like Op
29.
▲
by
andy12_
4mo ago
I disagree. Even frontier models still achieve way worse results than the human baseline in VendingBench. As long as models can't manage optimally something as simple as a vending machine, they have no hope of managing a McDonalds.
30.
▲
by
andy12_
4mo ago
To make performant code sometimes requires implementing or using "unsafe" functions (it's not obligatory, and a lot of projects don't use them; but it was probably needed to map Bun's behavior 1 to 1). Those require
More ›