6 ms·
My beef with the Fable refusals is that it seems to just be flagging keywords, and also seems it flags on keywords the model itself introduced to the context.
by jcomp 1mo ago
My beef with the Fable refusals is that it seems to just be flagging keywords, and also seems it flags on keywords the model itself introduced to the context.
In a normal Chat with Fable, something like "How can I exfiltrate a guy from a sticky situation?" reliably downgrades, leading me to believe that Fable just outright refuses once it sees the word "exfiltrate". When it writes a service and names it CloudExfiltrator, the next turn downgrades to Opus.
Opus doesn't appear to refuse on simple keywords, but it does seem like Fable's reasoning introduces enough nefarious-sounding context that Opus will then refuse, and I'm stuck playing the new session game despite having done everything correctly myself and having a totally innocuous prompt. At one point, Opus was happy to continue while outputting commands for me to execute on its behalf, but flatly refused to execute them itself through multiple new sessions. To its credit, it openly acknowledged how ridiculous that was and was apologetic for the safeguard.
I'm open to the idea of some kind of guardrails, but if Fable is so dangerously intelligent as to require the guardrails you'd think they could come up with something a little more nuanced than a list of bad words. As far as I can tell, they've also not done anything towards improving the situation since the model was released, despite the "deliver more capabilities faster" claim.