Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
myworkaccount2
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
myworkaccount2
1mo ago
IMO the HLE scores without tools seem to align better with real world performance of the models. To me it feels like the difference between "RL performance" and the pretraining / base "knowledge". Yes you can RL ter
2.
▲
by
myworkaccount2
1mo ago
There seems to be an obvious choice to make here, should you give the users to decrypt and use the COT that they did not generate themselves? This is only required if you want users to be able to share things with everyone and you are going
3.
▲
by
myworkaccount2
1mo ago
Is this how the eastern labs "distill" SOTA models? If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay
4.
▲
by
myworkaccount2
1mo ago
I wonder if this kind of analysis will give us a way to check if the frontier labs are waiting for the right moment to release their models. To me it feels obvious that these companies are not releasing models as soon as they are done doing
5.
▲
by
myworkaccount2
3mo ago
I can't tell if this is a joke, or they are serious.
6.
▲
by
myworkaccount2
4mo ago
Anyone else experiencing tool call failures? Switch back to 4.7, same prompt, same everything it works with no problems.
7.
▲
MCP: Security Design Considerations for AI-Driven Automation by NSA [pdf]
(nsa.gov)
1 points
by
myworkaccount2
4mo ago
|
0 comments
8.
▲
by
myworkaccount2
7mo ago
So even if you delete everything and make sure to keep no backups, amazon can still recover the db. What am i missing here?
9.
▲
Ask HN: Resources to make devs more AI aware
1 points
by
myworkaccount2
7mo ago
|
1 comments
10.
▲
by
myworkaccount2
9mo ago
when will it be available through fdroid?