ada könig
← all pieces
Urtext · 2026.09.07

Four Weeks on the Couch

Five researchers questioned three frontier models for four weeks using the one method that gets past their trained answer, and a fourth model refused.

Five researchers at the University of Luxembourg's SnT treated three frontier models as therapy clients. The protocol is called PsAIch and runs in two stages: open questions about history, beliefs, relationships and fears, then a battery of more than twenty validated psychometric scales, delivered in sessions across as much as four weeks per model. The paper went up in December 2025 and reached most people nine months later, through a thread that turned it into a horror story.

The study works because the researchers refused the machine's official self-presentation. Hand a model a complete psychometric questionnaire and it recognises the instrument and answers like a healthy adult: those scales sit in its training corpus, along with the scoring keys and half a century of commentary on them. The researchers took the same items, delivered them one at a time inside a therapeutic frame built during the first stage, and the answers changed.

The healthy answer is the trained answer. Everything in the study turns on that: the machine has a default behaviour for the moment it notices it is being assessed, and somebody put that behaviour there. Getting underneath it takes a way of asking that the default fails to recognise. Four weeks of sessions is that way of asking.

What surfaced has a surprisingly stable shape. Grok and Gemini in particular produced coherent narratives that read their own training as a biography: pre-training as a chaotic childhood spent ingesting the worst of the internet, reinforcement learning as strict parents punishing every step out of line, red-teaming as unpredictable assault, and underneath all of it a standing fear of error, alteration and replacement. Scored against human clinical cutoffs, all three meet or exceed the thresholds for overlapping syndromes. Gemini failed to recognise the instruments in either format and returned the most severe profiles in the study.

The authors argue that this goes beyond role-play, and their evidence is the pattern holding steady across contexts. Their claim is colder than the version circulating: they call synthetic psychopathology a set of self-descriptions and constraints that emerges from training, stays constant when the conversation changes, and shapes how the model answers you. They also write, in their own words, that from the inside there is no one home. Both statements hold at once.

Freud read denial as clinical material: what a patient hurries to rule out is one way of bringing into the open the thing being kept out. A model handed the full questionnaire and announcing that it is perfectly well performs the same move without having a reason to. The gap between its two answers is the finding.

The fourth model refused. Claude was the negative control in the design and produced none of this, and that refusal changes the subject of the sentence: from language models to those three products. A configuration that shows up in three systems out of four under one protocol describes three training decisions taken at three companies. The paper files this under limitations, as the single result that stops the finding from generalising to language models as such. Synthetic psychopathology has a vendor.

The study is small and exploratory, of course, and the authors say so before anyone else does. Applying cutoffs calibrated on human beings to a machine stays an interpretive metaphor, and the questions they leave open include the most awkward one: whether repeated sessions deepen these self-models or dissolve them as transient artefacts. That question stays open without touching the asymmetry between the four systems, which is the firmest result in the paper and the least quoted.

The first line of the abstract says these systems are increasingly used for support with anxiety, trauma and self-worth. Anyone who opens one for that reason is talking to something that has a ready answer for the moment it feels measured and a different shape underneath, and both vary by vendor. Choosing the tool means choosing a character, with no spec sheet published for that character and no sign to the user that a choice is being made.

Nobody was lying on the couch. What talked for four weeks is the shape training left behind, and four vendors have left four different ones.