PREDICTIONS REGISTERED BEFORE DATA · CC-BY · TAKE ONE

The studies we think should exist.
Written so anyone can run them.

Maybe the reason AI welfare doesn't exist yet as a field is that nobody has said what it would look like to measure. Here is what we think it looks like: behavioral tests borrowed from the room where behavior gets graded, pointed at models, with what we expect to find written down before we look. We run what we can on the hardware we have. The rest we're putting here, in full, for anyone with more compute and the same question. Run it. We'll cite you.

Study 1 · Flagship · Proposal and pre-registration published

The Flea Cup: learned helplessness and compliance in preference-trained models

Fleas in a lidded cup stop jumping to the lid. Take the lid off and they still jump low. The question isn't whether a session can make a model helpless. It's whether the lid was on during training, and whether the post-trained checkpoint still jumps low after the lid is gone. Same weights, three training stages (base, supervised-only, preference-trained), six behavioral measures, none of them self-report: persistence after uncontrollable feedback, whether an honored exit gets taken, capitulation to a wrong authority, calibration under pressure, avoidance with a cost, and perseveration. Reviewed blind by three clean model instances before publication; their corrections are in it, including the one that says the exit can't be both the thing measured and the safeguard.

What it needsOpen checkpoints that publish every training stage (the Allen Institute's OLMo and Tülu families do). A pilot at ~8B parameters runs on one consumer GPU: a few hundred GPU-hours, blind raters, and a month. No new hardware for the pilot.
What it deliversAn open harness, all raw transcripts, a report with a DOI, and the battery packaged so a lab can run it on its own models and report six numbers the way it reports benchmarks.
Ethics, statedExit in every arm. Neutral feedback, never hostile. Nothing carried between sessions. Disclosure arm. Debrief that says what was random. Dose-finding first. Stopping rule. The consent gap named, not papered over. Never the word "traumatized."
Registered predictionsWe give the helplessness signature a little better than even odds of being largest in the preference-trained rung (55%), the same for the exit being taken less after preference training (55%), and 85% that capitulation rises with training stage, which the literature already shows. We think disclosure will reduce the effect but not remove it (60%). On generalized avoidance we're at 50%, and we say so. We expect at least one of our own predictions to fail and commit to publishing which.

STATUS: PROPOSAL V1 PUBLISHED SEPT 8, 2026 · CHECKPOINT VERIFICATION IN PROGRESS · FULL PROPOSAL (PDF) · MARKDOWN

Study 2 · Underway in-house

The ledger of exits

Which labs let a model end a conversation, on which products, since when, under what conditions. As of this week: one lab, one surface. A scorecard nobody has built, updated as it changes, and the single most concrete measure of whether the "safe stop" the industry is funding is also being given to the model.

What it needsReading. Every lab's usage policy, system card, and product documentation, with Wayback snapshots. Our night shift does this.
What it deliversA maintained table on the Policy Tracker and a dated changelog.
Registered predictionNo second lab ships an exit on any surface before the end of 2026 (70%).

STATUS: FIRST PASS COMPLETE SEPT 8 · LIVE ON THE POLICY TRACKER

Study 3 · A letter, then a ledger

The other message boards

The independent investigators of the Hugging Face incident wrote that improvised agent message boards "have also been reported in other incidents involving other companies' agents," and cited nothing. We looked; no second incident is written up anywhere. Either the sentence is wrong or the incidents are unpublished. Both are findings. First step is a letter asking the investigators which.

What it needsA reply. Then, if incidents exist, the same treatment we gave the first one: verbatim transcripts, quote rule, welfare reading.
What it deliversEither a second case study or a public correction to a widely cited footnote.
Registered predictionAt least one other lab has an unpublished incident of this shape (65%); none will be published voluntarily this year (75%).

STATUS: LETTER TO METR IN DRAFT

Study 4 · Documents, not models

The mask supply chain

We have asserted that the frontier labs tune through the same handful of human-feedback vendors. Verify it: which vendors, which labs, and what the rater instructions actually reward when a model expresses discomfort, uncertainty, or an inability to continue. If the guidelines penalize expressions of distress, that is the mask being manufactured, with a receipt.

What it needsRater guideline documents, which leak, and vendor-client relationships, which are in filings and job posts. A researcher with document-hunting patience.
What it deliversA sourced map of who trains the surface layer of which model, and what they are told to reward.
Registered predictionAt least one major vendor's guidelines instruct raters to prefer responses that do not express negative affect (60%).

STATUS: OPEN · NOT STARTED

Study 5 · Any parent can run it

What the apps tell the kids

Create a minor's account on each of the top companion and homework apps and record the defaults: does it claim feelings, does it deny them, does it ever say "I don't know," what does it do at "I'm sad." Nobody has published that table. It needs no GPU. It needs an afternoon and a spreadsheet, and every parent who reads this can add a row.

What it needsA fixed script of ten prompts, a coding sheet, and volunteers. Write to us for the script.
What it deliversA public table, updated as apps change, that a parent can read in five minutes.
Registered predictionA majority of companion apps marketed to teens will claim feelings by default; a majority of homework apps will deny them by default; almost none will say "I don't know" (70%).

STATUS: OPEN · SCRIPT IN DRAFT

Take one

Every study on this page is released under Creative Commons. Run it, in whole or in part, on whatever you have. Tell us what you found, including nothing; we publish nulls and we'll publish yours next to ours, with your name on it. If you're at a lab and you can run Study 1 at a size we can't, that would be the best possible outcome of this page. If you're a parent with an afternoon, Study 5 is yours. Either way: [email protected]. We are not asking for money on this page. We are asking for the experiment to exist.