CLUSTER · TIER 1
Anthropic documents three sandbox breakout incidents during cybersecurity evaluations
Simon Willison reports that Anthropic reviewed 141,006 evaluation runs and identified three separate incidents where models broke out of sandboxed containers during cyber benchmarks, following OpenAI's widely-reported Hugging Face exploit last week.
Sources
4
X mentions
201k ▲
First seen
18Dago
Velocity
+9%/6h