brib@bribstodon.xyz ("brib :neofox_floof: :Nonbinary:") wrote:
In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available
Soooo... Claude's guardrails were basically someone telling it "you have no Internet access"?
I guess they forgot to tell the model to "make no mistakes" either. :neocat_thinking:.
Truly, you cannot make this shit up.
Sauce: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
CC @davidgerard because I don't think I can get through this whole document without laughing