6 ms·
"...Recently, we did an experiment where we had a “red team” deliberately introduce an alignment issue into a model (say, a tendency for the model to exploit a
by bryantnyc 11mo ago
"...Recently, we did an experiment where we had a “red team” deliberately introduce an alignment issue into a model (say, a tendency for the model to exploit a loophole in a task..." https://www.darioamodei.com/post/the-urgency-of-interpretability https://www.darioamodei.com/post/the-urgency-of-interpretabi...
Anyone care to wager if anthropic is red teaming in production on paying users?