5 ms·
> Preliminary trials with Claude Mythos Preview showed that it would not provide an apples-to-apples comparison with other models because of how we had set up t
by bob778 3mo ago
> Preliminary trials with Claude Mythos Preview showed that it would not provide an apples-to-apples comparison with other models because of how we had set up the experiment and how the model was served.
What does this mean? My guess is they couldn’t co-locate Mythos close enough to reduce latency?
(I’m assuming this experiment pre-dates the export controls)
- georgemcbay 3mo ago> My guess is they couldn’t co-locate Mythos close enough to reduce latency? I doubt network latency is the reason. Even when connecting from literally across the world network latency is lost in the noise of overall response latency of even fast models. The overall response latency of the model very well could have been the difference, though. AFAIK Mythos is structured to do relatively slow "deep thinking".
- bannable 3mo agoDepending on the timeline, it could be that they're not allowed to access Mythos because of something like non-US citizens on the team or the lack of some way for them to meet the constraint DOD has them under.
- georgemcbay 3mo agoI strongly suspect if that was the case they would have just directly mentioned that Mythos couldn't be used because of that reason, it would be less confusing and less suspect messaging than saying it wasn't an "apples-to-apples comparsion".
- daveguy 3mo agoBecause this was a staged demo, not an experiment. Mythos performed more poorly but they don't want to admit it. The phrase "because of how we had set up the experiment" means "we didn't have experimental controls and got a bunch of bullshit noise that we cherry picked." At least that's my guess.