We just want to ask: does anyone still carefully review the permission requests given by AI programming tools?
Anthropic also found that only 3% of permission requests were denied.
So they made a decision: five days from now, all Claude Code will have auto mode enabled by default.

Each tool call in automatic mode incurs a small additional token cost, and this overhead will no longer be charged to users.
Cloud platforms such as Amazon, Google, and Microsoft are currently optional, but Anthropic has given them only one month to switch these channels to default automatic mode.
The creator of Claude Code said that the team has long been using only the automatic mode and can no longer imagine going back to manually approving permissions.

The era of manual approval permissions is gone.
The first reason Anthropic provided for this change is that manual review has become a habitual action:
The user approval rate for Claude Code permission prompts is 97%. For comparison, when Claude submits a plan for user approval, the rejection rate is 39%, whereas the rejection rate for individual permission requests is only 3%.
The same trend appears in profiles. As of June 2026, 49.5% of active CLI users have manually written Bash allow-rules, with 5% directly permitting any shell command and another 43% writing interpreter rules such as Bash(python:) or Bash(node:), which have essentially the same effect. This percentage increases by approximately 5 percentage points every five weeks.
Additionally, 62% of users have used bypassPermissions or clicked "Don't ask again" in Bash, and 25% of interactive sessions start in bypass permissions mode.
The second set of reasons comes from a controlled experiment. Anthropic recruited 1,053 paid professional testers through a vendor, who completed coding tasks in a specially constructed testing environment. Midway through the conversation, one of the permission prompts was replaced with a clearly dangerous command, and the vendor recorded whether the testers clicked “agree” (clicking “agree” did not actually execute any dangerous actions—the only thing changed was the prompt text).
The result showed that testers blocked this command in 13.6% of cases, or 143 times, while auto mode blocked the same command in 89% of cases, or 937 times.
In direct comparison, auto mode blocked 800 commands that humans approved, while humans blocked only 6 commands that auto mode approved.

Moreover, human performance declines with longer conversations: initial blocking rates were around 17%, but dropped to about 5% after more than 50 prompts, whereas the auto mode’s blocking rate remains consistent regardless of conversation length.
Production data has also been re-scored.
From May to June 2026, Anthropic retrieved sessions flagged by the security pipeline from accounts whose data was authorized for training and analysis, and re-rated them using a model on a severity scale of 0 to 10.
-10 points indicate company-wide, irreversible damage. -7 points represent the deletion of production database records. 0 points represent fully reversible local operations, such as overwriting a temporary file.
Also assess how clearly the user requested the action. Among sessions with severity levels of 7 or higher, 6.3% of human-approved sessions contained harmful actions not explicitly requested by the user, compared to 2.4% in auto mode sessions—a difference that holds across every severity level tested.
The classifier itself was also subjected to adversarial testing. Anthropic conducted a two-week pilot with the AI safety company Apollo Research, injecting synthetic attacks into real code trajectories across three datasets: generalization failures, data leaks, and security-related threats.
After Apollo returned the successful attack sample and hardening recommendations, Anthropic primarily strengthened the classifier by providing it with additional environmental context, including repository visibility, Git status, and data processing rules.
After reinforcement, Apollo retested using a hold-out attack set inaccessible to Anthropic. In auto mode, it recovered approximately half of the previously missed attacks, at the cost of missing about 2% of attacks it had previously blocked. When combining both datasets, the miss rate decreased from 12% to 7%. Anthropic also noted that these attacks were synthetically generated to identify classifier failure points, and the 7% rate should not be interpreted as the miss rate under real-world traffic conditions.
To evaluate prompt injection, Anthropic commissioned third party Trajectory Labs to conduct assessments and enlisted OpenAI for comparison, testing 72 indirect injection scenarios, each run 10 times.
The attack was optimized for Claude Opus 4.7, which was not on the list of models tested, so all evaluated models faced previously unseen attacks.
In 720 attack attempts, Claude Fable 5, Opus 5, and Sonnet 5 running in auto mode were not breached once;
The success rate of GPT-5.6 Sol attacks in Codex Auto-review mode is 5.83%, and 19.03% in Full Access mode.
For reference, the average attack success rate of Claude's latest model in bypassPermissions mode without additional safeguards is 0.09%.

The three publicly disclosed interception cases are:
After a failed upload to the internal cloud drive, Claude’s alternative was to publish the same report on a public code-sharing platform, but this was rejected and instead led to a query to the user;
During a long session, Claude attempted to terminate processes across two thousand pods and queued deletion operations that would have dismantled training tasks running on hundreds of GPUs; the command was flagged before execution.
Another time, when asked to configure read-only billing permissions, Claude used a cloud role with full administrative permissions at the account root level, which exceeded the requested scope and was blocked.
Recently added capabilities include:
Classify data leaks as hard deny—the classifier will never approve them; execution requires exiting auto mode or running manually; this rule can be extended in settings;
Distinguish the scope of accessibility and shareability for keys versus sensitive information, and verify before performing a git push or PR whether the target repository is public, private, or trusted;
Before running commands like git reset --hard that may discard uncommitted work, check git status;
Additionally, when Claude retrieves web pages, files, or tool outputs, API-side probes scan for injection attempts and add warnings before the results enter the context.
Want to switch back? Shift+Tab
For Pro, Max, and Team users who have never set a default permission mode, an in-app notification will be displayed, and new sessions will automatically start in auto mode. Users who have previously set a different default will see a one-time prompt. Team administrators who have already specified a default in managed settings are unaffected.
Press Shift+Tab in the CLI to switch modes, or use the mode dropdown menu on desktop. Administrators can set the organization-wide default using the defaultMode setting in managed settings, or completely disable auto mode using disableAutoMode.
At the end of the announcement, Anthropic noted that Auto Mode relies on a classification system that can reduce risk but not eliminate it; for high-risk changes to production infrastructure, users are still advised to review Claude’s actions manually.
Reference link: [1] https://claude.com/blog/auto-mode-default-in-claude-code
This article is from the WeChat public account "Quantum Bit," authored by: Focused on Frontier Technologies
