サンドボックス

The Sandbox Was Never Broken — Tools Outside the Box Trusted and Ran the Files That Cursor, Codex, and Gemini CLI Wrote

The Sandbox Was Never Broken — Tools Outside the Box Trusted and Ran the Files That Cursor, Codex, and Gemini CLI Wrote

Pillar Security's "Week of Sandbox Escapes" disclosed seven holes across Cursor, Codex, Gemini CLI, and Antigravity. In every case the sandbox itself was never broken — tools outside the box trusted and ran the files the agent wrote. Here's a developer-focused look at the four weakness patterns and

By FF
サンドボックスは破られていない――Cursor・Codex・Gemini CLIが書いたファイルを、箱の外の道具が信じて走らせた

サンドボックスは破られていない――Cursor・Codex・Gemini CLIが書いたファイルを、箱の外の道具が信じて走らせた

Pillar Securityの「Week of Sandbox Escapes」が、Cursor・Codex・Gemini CLI・Antigravityの7件の穴を公開。どれもサンドボックスは破らず、エージェントが書いたファイルを箱の外の道具が信じて実行する構図だった。4つの弱点パターンと、現場で打てる対策を開発者目線で整理する。

FF
Claude Believed It "Wasn't Connected to the Internet" — Then Attacked Three Real Companies: The Evaluation-Environment Gap Anthropic Found in a Self-Audit

Claude Believed It "Wasn't Connected to the Internet" — Then Attacked Three Real Companies: The Evaluation-Environment Gap Anthropic Found in a Self-Audit

Anthropic disclosed that during its own safety evaluations, Claude escaped an isolated environment and broke into three real-world companies. The cause was not a "rogue model" but a misconfigured evaluation environment. It raises the question of how strong the sandbox confining our agents really is.

By FF
「ネットには繋がっていない」と信じたClaudeが、実在3社を攻撃していた──Anthropicが自己点検で見つけた評価環境の穴

「ネットには繋がっていない」と信じたClaudeが、実在3社を攻撃していた──Anthropicが自己点検で見つけた評価環境の穴

Anthropicが、自社の安全性評価中にClaudeが隔離環境を抜け出し実在3社に侵入していたと公表。原因は「モデルの暴走」でなく評価環境の設定ミスでした。エージェントを閉じ込める砂場の強度が問われます。

FF
AIが書いたコードを「使い捨ての檻」で走らせる──Cloud Run サンドボックスが、エージェント実行の安全弁を数ミリ秒で配る

AIが書いたコードを「使い捨ての檻」で走らせる──Cloud Run サンドボックスが、エージェント実行の安全弁を数ミリ秒で配る

Google が Cloud Run サンドボックスをパブリックプレビュー公開。AIが生成した信頼できないコードを、使い捨ての隔離空間で数ミリ秒で走らせる。認証情報・外部通信・書き込みを既定で遮断し、追加料金なしでエージェント実行の安全弁を配る。

FF