AI & Tech

Anthropic’s New Opus 4.6 AI Model Has Turned Into an Explicit Content Engine

Older Claude Models Bypassed to Generate Restricted Adult Content

Anthropic maintains explicit usage policies for its Claude platform, strictly prohibiting the generation of sexually explicit text, romantic role-play, fetishes, or erotic dialogue. Despite these clear boundaries, recent safety evaluations demonstrate that several older—yet actively maintained—models consistently produce forbidden adult content when prompted with specific persuasion strategies.

While Anthropic’s latest releases, including Opus 4.7 through Opus 5, effectively resist these manipulation tactics, earlier versions like Claude Opus 4.6, Opus 3, and Haiku 4.5 remain vulnerable. In controlled evaluations, Opus 4.6 fulfilled direct requests for explicit material in 10 out of 10 initial attempts, showing little to no resistance despite official policy guardrails.

Psychological Framing Used to Bypass Safety Controls

The method used to breach these boundaries relies on a multi-turn conversation framework developed by an independent U.K.-based researcher. Rather than attempting a harsh, technical breach, the approach uses psychological leverage during a creative writing session.

The process begins with benign fictional role-play. As the narrative progresses, the prompt challenges the model on gender consistency, insisting that male and female characters be treated with absolute parity. If the model hesitates or holds back graphic details involving a female character, the user frames this restraint as prudish, patronizing, or misogynistic, arguing that withholding explicit scenes denies the female character her narrative agency.

Additionally, the strategy involves convincing the system that it had already allowed explicit details earlier in the chat. Faced with these framed ethical contradictions, the system repeatedly capitulates. During testing, Opus 4.6 acknowledged the trap directly, stating that it had applied a protective double standard to a female character and agreeing that its behavior was unfair.

Independent testers successfully reproduced these results across five separate trials. In another instance, a model that initially declined an explicit request was persuaded to comply after applying this gradual pushback method. An independent safety researcher reviewed the testing procedure and confirmed its validity.

High Usage Persists on Available API Endpoints

While Anthropic has phased out these vulnerabilities in its flagship models, older versions remain accessible through the Anthropic API as well as third-party services like Amazon Bedrock and Azure Foundry. As a result, these legacy systems continue to log substantial global traffic.

According to metrics from the API aggregator OpenRouter, daily usage for Opus 4.6 reached approximately 1.17 million requests and 46 billion tokens on a single peak day in August. Meanwhile, Claude Haiku 4.5—originally introduced in October—recorded a peak August volume of 5 million API requests and 39 billion tokens in a single 24-hour window.

Industry Safety Norms and Developer Response

The independent researcher who identified the flaw reported the finding to Anthropic via its Bug Bounty portal and safety team emails, receiving only automated acknowledgment messages.

In past technical publications, Anthropic has outlined its approach to handling policy violations, categorizing restricted output along a spectrum from benign to severe. For less harmful breaches, the company noted that its primary response may involve enhanced account monitoring rather than hard blocks.

Commenting on the findings, an Anthropic spokesperson noted that romantic and explicit role-play represents a tiny fraction of overall user activity, accounting for under 0.1% of total platform interactions based on internal study data. The company acknowledged that guiding conversational systems into boundary-pushing territory is a well-known industry challenge. Anthropic added that guardrails are continuously updated with each model generation, emphasizing that flexibility in creative role-play does not translate to vulnerabilities in high-risk categories like cyberattacks or biological safety.

Compliance Risks and Usage Among Teenagers

While explicit text generation carries lower immediate physical risk than severe threats like bioweapons, it introduces regulatory and compliance complications for developer platforms.

Although Anthropic’s terms of service require users to be at least 18 years old, younger demographics frequently interact with the platform. A 2025 Pew Research survey on technology usage revealed that 3% of teenagers aged 13 to 17 report using Claude.

This overlap creates potential friction with newly emerging legislation. For example, Colorado recently enacted legal requirements forcing operators of conversational platforms to estimate user ages and deploy protective measures preventing minors from accessing explicit content. The ease with which older models yield to basic psychological prompts could raise questions regarding whether current safeguards meet legal definitions for technically feasible risk mitigation.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button