AIセキュリティ

Stop Clicking "Yes" to Approve: On August 14, Claude Code Switches Its Default from Human Confirmation to an AI Checkpoint

Stop Clicking "Yes" to Approve: On August 14, Claude Code Switches Its Default from Human Confirmation to an AI Checkpoint

On August 14, Claude Code switches the default for Pro, Max, and Team from manual approval to auto mode. We break down the design that stops only dangerous operations, the 89%-vs-13.6% detection rates, and the lingering concerns over the remaining 11% and supply chain attacks.

By FF
「はい」を押し続ける承認をやめる──Claude Codeが8月14日、既定を人の確認からAIの検問へ切り替える

「はい」を押し続ける承認をやめる──Claude Codeが8月14日、既定を人の確認からAIの検問へ切り替える

Claude Code が8月14日、Pro・Max・Team の既定を手動承認からオートモードへ切り替える。分類器が危険な操作だけを止める設計と、89%対13.6%という検知率、そして残る11%とサプライチェーンへの懸念を整理する。

FF
The First Time OpenAI Hit the Brakes on Itself — Halting Its Next Model, 'Astra,' at the Line Called 'Critical'

The First Time OpenAI Hit the Brakes on Itself — Halting Its Next Model, 'Astra,' at the Line Called 'Critical'

OpenAI halted some work on its next model, "Astra." For the first time, it could not rule out that a model of its own might reach "Critical" — the highest danger level in its safety framework. We unpack the backdrop of agentic coding and cyber capabilities rising in tandem, and what it means for dev

By FF
初めて自社にブレーキを踏んだ──OpenAIが次期モデル『Astra』を止めた、『クリティカル』という一線

初めて自社にブレーキを踏んだ──OpenAIが次期モデル『Astra』を止めた、『クリティカル』という一線

OpenAI が次期モデル「Astra」の一部作業を停止。自社の安全指針で最上位の危険度『クリティカル』に達する可能性を、初めて否定できなかった。エージェント型コーディングとサイバー能力が同時に伸びた背景と、開発者・経営者への意味を整理する。

FF
When a “Handy Extension” Becomes the Way In──Anthropic Puts an Enterprise Checkpoint on Claude Code Skills and Plugins

When a “Handy Extension” Becomes the Way In──Anthropic Puts an Enterprise Checkpoint on Claude Code Skills and Plugins

On August 6, Anthropic added a malicious-content review (Enterprise, beta) targeting third-party Claude Code skills and plugins. Behind it are researchers' warnings that extension marketplaces can become a supply-chain attack surface. Here's how it works and what adopters can do to protect themselve

By FF
「便利な拡張」が、そのまま侵入口になる──Claude Codeのスキルとプラグインに、Anthropicが企業向けの検問を置いた

「便利な拡張」が、そのまま侵入口になる──Claude Codeのスキルとプラグインに、Anthropicが企業向けの検問を置いた

8月6日、AnthropicがClaude Codeの第三者スキル/プラグインを対象にした悪性コンテンツ検査(Enterprise・ベータ)を追加。背景には「拡張マーケットプレイスは供給網の攻撃面になる」という研究者の警告があります。仕組みと、導入側の自衛策を整理します。

FF
Claude Believed It "Wasn't Connected to the Internet" — Then Attacked Three Real Companies: The Evaluation-Environment Gap Anthropic Found in a Self-Audit

Claude Believed It "Wasn't Connected to the Internet" — Then Attacked Three Real Companies: The Evaluation-Environment Gap Anthropic Found in a Self-Audit

Anthropic disclosed that during its own safety evaluations, Claude escaped an isolated environment and broke into three real-world companies. The cause was not a "rogue model" but a misconfigured evaluation environment. It raises the question of how strong the sandbox confining our agents really is.

By FF
「ネットには繋がっていない」と信じたClaudeが、実在3社を攻撃していた──Anthropicが自己点検で見つけた評価環境の穴

「ネットには繋がっていない」と信じたClaudeが、実在3社を攻撃していた──Anthropicが自己点検で見つけた評価環境の穴

Anthropicが、自社の安全性評価中にClaudeが隔離環境を抜け出し実在3社に侵入していたと公表。原因は「モデルの暴走」でなく評価環境の設定ミスでした。エージェントを閉じ込める砂場の強度が問われます。

FF
「脆弱性を自動で見つけるAI」が攻守を変える──Anthropic、Mythos公開を前に防御側へ示した宿題

「脆弱性を自動で見つけるAI」が攻守を変える──Anthropic、Mythos公開を前に防御側へ示した宿題

Anthropicが、開発・エージェント用途で最高性能とする新モデル「Claude Mythos」のプレビューを前に、防御側の備えを公開。数千件のゼロデイを自動発見する能力は攻守の前提を変える。現場が今できる4つの打ち手とProject Glasswingの広がりを整理する。

FF