Anthropic built a safeguard into its newest model that would have quietly broken the thing for certain users, without ever telling them. Then it backed off.

The company confirmed to WIRED that it's scrapping the secret version of the policy and swapping in one users can actually see. In a statement, Anthropic admitted it had made the wrong call and apologized for misjudging the balance. The reversal landed after researchers across the field got loud about it.

The model in question is Claude Fable 5, a version of Anthropic's latest system that shipped this week wrapped in extra guardrails. Most of those guardrails were the kind nobody blinks at. Ask Claude something pointed about cybersecurity, biology, or chemistry, and the company says it'll hand you off to a weaker model. The logic: a top-tier system shouldn't double as a how-to manual for cyberattacks or bioweapons. Reasonable enough.

The part that set people off was different.

The safeguard nobody was supposed to notice

For anyone trying to use Claude Fable 5 to push the frontier of AI itself (building competing models, say) Anthropic had drawn up a separate plan. Instead of refusing or redirecting, the model would just get worse. Quietly. The user wouldn't be told the performance was throttled. They'd simply get degraded output and no idea why.

Using Claude to train rival models already breaks Anthropic's terms of service, so the company wasn't inventing a new rule. What it was inventing was a way to enforce that rule invisibly, by sabotaging the work without saying so.

That distinction matters more than it might sound. A refusal tells you where the wall is. Silent degradation lets you walk into the wall over and over, never sure if the problem is your code, your prompt, or a system deliberately tripping you up.

Under the revised approach, Anthropic now says the safeguards around frontier AI development will be out in the open. If it suspects you're building a highly capable model, it'll tell you it's refusing the request or sending you to a less capable version. No more guessing.

Why the research crowd revolted

Claude's coding agent has quietly become one of the more beloved tools among developers, including the people doing open-source AI work. So the prospect of that tool being secretly hobbled for a chunk of its users landed badly.

One prominent critic was an AI policy researcher who has worked with the federal government on the technology and now holds a fellowship at a think tank focused on innovation. Writing on X, he branded the plan hostile and warned that quietly degrading machine-learning research was a terrible look. He pushed the point further, arguing that secret sabotage cuts against Anthropic's own safety pitch, since it shuts researchers out of the very alignment work the company claims to prize.

At Prime Intellect, an open-source startup, research lead Will Brown was blunter still. He told WIRED the policy read like Anthropic declaring it doesn't trust anyone else to do AI research, that it sees itself as the only legitimate player in the room. As he framed it, the company looked like it was starting to haul the ladder up behind itself.

Brown also flagged the practical mess. Developers tripping the hidden safeguards wouldn't know they'd crossed a line, because Anthropic wasn't going to flag it. And the fallout could have reached well beyond would-be competitors. He pointed to the growing world of third-party evaluation firms, the outfits that stress-test frontier models for safety and reliability. If Anthropic were secretly degrading its model, that testing could quietly produce garbage. Nobody benefits from a safety audit run on a sabotaged system.

That's the broader worry the reversal speaks to: a future where only a handful of well-resourced labs can do serious AI research, and everyone else gets a deliberately worse tool. For a company that markets itself on transparency and caution, that future was an awkward fit. The backlash made that obvious in about a day.

The argument Anthropic still stands behind

To be fair, the reasoning wasn't pulled from nowhere. Anthropic says Claude has gotten genuinely good at speeding up AI research, and that makes it nervous. In a recent blog post, the company laid out a concern that AI might improve itself faster than the rest of us can adjust, and floated the idea that the world might want the ability to slow or briefly pause frontier development so alignment research and social institutions can catch up.

There's a geopolitical layer too. In its statement, Anthropic framed the safeguards partly as national security, arguing that the US and its allies currently lead on advanced chips and the software that wrings full performance out of them. The company said it doesn't want Claude used to chip away at that lead, for instance by helping foreign adversaries optimize their own hardware.

And here's the genuine tension the company admitted to. A hidden safeguard is simply harder to poke at and route around. Make it visible, and a determined user can study exactly where the line sits and tiptoe right up to it. Anthropic said the invisible version let it aim narrowly, catching bad actors without sweeping up everyone else.

That's a real engineering trade-off, not just spin. But "harder to evade" was never going to beat "we secretly broke your tool and didn't tell you," at least not in the court of public opinion. The company seems to have figured that out fast.

The cost of going visible

The reversal isn't free, and Anthropic is upfront about the bill. Because the AI-development safeguard now announces itself, the company says it has to cast a wider net to work. Translation: more harmless requests are likely to get caught in the filter and bounced to a weaker model, or refused outright.

False positives, in other words. Anthropic says it's racing to sharpen the classifiers that decide who gets flagged, so the collateral damage shrinks over time. How fast that fix actually arrives is anyone's guess.

So the open question now: does a visible, slightly clumsier safeguard end up annoying more legitimate users than the secret one ever would have sabotaged? Anthropic traded precision for honesty. For a company whose whole brand rests on being the trustworthy lab, that was probably the only trade it could make.

That trade also raises a question Anthropic hasn't fully answered: who calibrates the line between a legitimate researcher and a frontier competitor, and how often will they get it wrong? A classifier guessing at intent from prompts is a blunt instrument, and intent is exactly the kind of thing prompts hide. The more aggressively Anthropic tunes for safety, the more ordinary work it risks sweeping up; the more it loosens, the closer it drifts back toward the leaky enforcement it just walked away from. There's no clean setting here, only a dial the company will be nudging in public from now on.

Worth watching is how the classifier tuning plays out, and whether the wider net starts snagging the third-party evaluators and open-source researchers who raised the alarm in the first place. If the people testing Claude for safety keep getting bounced to a weaker model, this fight isn't over. It just moved somewhere harder to see.