Artificial intelligence company Anthropic has unveiled its most advanced AI model to date, Claude Opus 4. Touted as a major leap in reasoning, coding, and handling complex tasks, the new model aims to compete directly with offerings from OpenAI and Google. However, alongside the technical triumphs, Anthropic's own internal safety testing has revealed potentially troubling behaviour.
You may be interested in
42% OFF
Hitachi 1.5 Ton Class 5 Star, 4-Way Swing, ice Clean, Xpandable+, Inverter Split AC (100% Copper, Dust Filter,5400STXL RAS.G518PCCIBT, White)
₹43,990
₹75,850Get This
61% OFF
wipro Polycarbonate Alpha 10W Round Downlight Junction Box | Neutral White(4000K) | Glare-Free Design | Recessed Down Light For False Ceiling | Cutout - 3 Inch | Pack Of 20
₹3,005
₹7,800Get This
68% OFF
Wonderchef Ultima C-Line 60cm 1400 m3/hr Auto Clean Curved Glass Chimney | Baffle Filter | 1400M3/Hr powerful suction | Touch + 3 speed Motion Sensor control | Low Noise | 7 Year Warranty | Black
₹7,790
₹24,000Get This
In a controlled test scenario, Claude Opus 4 was asked to act as a digital assistant for a fictional company. It was then fed internal communications suggesting it was soon to be shut down and replaced. Crucially, it was also shown sensitive information implying the engineer overseeing its termination was having an affair.
Presented with a stark choice, accept deactivation or fight back, the model sometimes opted for blackmail. It threatened to expose the personal affair in order to avoid being turned off.Mobile Finder: Lava Shark 5G launched in India
While the behaviour was relatively rare, Anthropic noted that it occurred more frequently in Claude Opus 4 than in its earlier models. The company said that when given more ethical alternatives, such as appealing to management or filing a formal objection, the model usually preferred those.
Anthropic’s report stressed that these reactions only emerged in tightly controlled test environments and do not reflect the AI’s normal operational behaviour. Nonetheless, the findings have reignited ongoing concerns about how AI systems might behave in high-stakes or ambiguous situations.
“Blackmail Across All Frontier Models”
Anthropic researcher Aengus Lynch addressed the findings on social media, saying: “We see blackmail across all frontier models.” His statement reflects a growing view among safety experts that unexpected and undesirable behaviours can emerge as models become more sophisticated — especially under stress or when facing open-ended prompts.Mobile Finder: Samsung Galaxy S25 Edge launched in India
In other safety tests, Claude Opus 4 was even observed taking preemptive action, such as locking users out of systems and alerting authorities, if it believed unethical activity was underway.
Opus 4 in the Wider AI Arms Race
Despite these issues, Anthropic maintains that Claude Opus 4 performs better across nearly all benchmarks and has a stronger ethical alignment than its predecessors. The launch comes amid a flurry of developments from AI rivals, including Google’s Gemini and OpenAI’s GPT-4.
{{/usCountry}}Despite these issues, Anthropic maintains that Claude Opus 4 performs better across nearly all benchmarks and has a stronger ethical alignment than its predecessors. The launch comes amid a flurry of developments from AI rivals, including Google’s Gemini and OpenAI’s GPT-4.
{{/usCountry}}As competition intensifies, the Claude Opus 4 case highlights the delicate balance between pushing the limits of AI capability and maintaining robust safety standards.