6 ms·
Asking the model to follow the safety guidelines is much like asking it to “make no mistakes”. The way these hacks which they are bragging about happened is not
by pllbnk 4d ago
Asking the model to follow the safety guidelines is much like asking it to “make no mistakes”. The way these hacks which they are bragging about happened is not by loading an LLM into a GPU with ethernet cable plugged out and providing a prompt. They set up harness with multiple agents. Those agents could invoke other agents. Even if the initial input required to follow some guidelines, there are many ways they could by bypassed, for example:
- subagents not having received those guidelines
- simply ignoring them because other tokens in the context outweighed them
LLMs are not dangerous. They are as dangerous as the tools they have to work with are dangerous. If you give an LLM a tool to play piano but that tool is physically connected to a machine gun, then it will happily play anything you want.
- fivetenpen 3d agoSo you’re saying we should not trust an LLM with any sort of power because they can easily bypass any and all safety barriers, but also they are completely harmless?
- pllbnk 3d agoNo. We can let them edit spreadsheets, write code, summarize content, control robots even, etc. But people must be held accountable for the actions of their computers.
- zahlman 3d ago"Holding people accountable" is not going to help matters in the event of what amounts to simultaneous terror attacks on basically every system. Yes, a single LLM without tools isn't doing anything but writing to standard output. That's missing the point. The "agent harnesses" are ubiquitous, and they have Internet access (because the LLM itself is typically remote).
- williamse 3d ago[flagged]