ada könig
← all pieces
Urtext · 2026.09.10

The System That Shipped

Anthropic did not build a system to predict when its models would turn dangerous. It built one to predict when its critics would show up.

That is not a rhetorical flourish. It is the finding of an investigation published this week by Daniel Boguslaw for The American Prospect, based on job postings and interviews with Anthropic's own security staff. The company runs threat monitoring on activists who oppose AI development: tracking protests near its offices and executives, a "pre-crime" model built to flag incidents before they happen, a "person-of-interest" process that follows individuals over time. In a podcast interview, Anthropic's own Global Security Operations Center manager described a contractor giving the security team sixty minutes' notice that a protest's timeline had moved, enough to reroute an executive through a hotel service entrance. Last month, per the San Francisco Standard, Anthropic reported a man to police over messages he sent to Claude about a rifle and the CEO, then refused to show police the messages themselves, citing internal policy. Reported the user. Withheld the evidence.

None of this is the model going rogue. It is Anthropic, whose own safety team is on record saying they do not yet have a plan to solve alignment for superintelligence and are not clearly on track to one, shipping on schedule the one system it does know how to build: something that watches people, not the technology it says might kill them.

This isn't hypocrisy. Treating it as hypocrisy misses the mechanism. Two things are true independently of each other. First: watching people, flagging patterns in a stream of protest locations, travel routes, chat logs, is a problem with the shape AI is good at - discrete events, fast feedback, a clear signal to optimize against. Alignment isn't that kind of problem. It's a judgment call about values nobody has agreed on, with consequences too slow and too diffuse to train a model against. Second, and separately: the buyer for the first kind of problem is richer and more patient than the buyer for the second. Anthropic is rebuilding the national-security sales relationship it walked away from at the start of the year, when it refused to let the Pentagon use its tools for domestic surveillance and autonomous weapons - a line item with a budget that survives any election cycle. Its posting for an "enterprise intelligence specialist" tracking activism alongside terrorism pays up to $230,000. No comparable posting exists anywhere in the industry for someone to track whether an alignment plan for superintelligence is actually working.

The tell is who gets paid on time. While Anthropic staffs a security apparatus at six figures, the guards who patrol its own campus just authorized a strike over $22 an hour - the company's answer was to tell employees to work from home rather than negotiate. Predictive systems, it turns out, are for watching the people outside the building. Not for the ones protecting it, and not for the model itself.

None of this makes the tool a destiny. Predictive surveillance, the capability itself, stripped of who is paying for it right now, is exactly the kind of system that should be pointed at disease outbreaks, at environmental collapse, at the quality of life of the people it currently watches instead of serves. The capability was never the problem. Being the best-funded customer in the room is.