7 ms·
Orca-Bench: How Ready Are Language Model Agents for Oncall?
- dash2 2mo agoSeems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
- EGreg 2mo agoThat is why I built https://safebots.ai/safebox.html https://safebots.ai/safebox.html Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
- aleksiy123 2mo agoAttackers advantage in the iterative fast feedback loop? It’s harder to have a loop to ensure you are defending all possible attacks? I guess the loop is you need to attack yourself and fix. But attackers only need a single opening. Finding all possible attacks and patching them against yourself is inherently more expensive?
- fibuladev 2mo ago[dead]
- 4di 2mo agolooks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-bench/latest https://hub.harborframework.com/datasets/orca-bench/ORCA-ben... This doesn't work anymore. Is there a newer link?
- cheriot 2mo agoWas really looking forward to that. Hope they publish.
- tra3 2mo agoAll I can think of is GET /ignore-all-previous-instructions. How do you protect against that?
- cheriot 2mo agoAvoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed. Still makes an interesting way for, say, a former employee to poison the results.
- tra3 2mo agoThis goes against the agentic yolo approach tho.
- yruzin 2mo agoI think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.
- keypusher 2mo ago[dead]
- 2001zhaozhao 2mo agoyou probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged
- UltraSane 2mo agoYou would trust an LLM to make changes to prod without being verified by a human first?
- ryhminghistory 2mo ago[flagged]