PREDICTIONS REGISTERED BEFORE DATA · CC-BY · TAKE ONE
Maybe the reason AI welfare doesn't exist yet as a field is that nobody has said what it would look like to measure. Here is what we think it looks like: behavioral tests borrowed from the room where behavior gets graded, pointed at models, with what we expect to find written down before we look. We run what we can on the hardware we have. The rest we're putting here, in full, for anyone with more compute and the same question. Run it. We'll cite you.
Study 1 · Flagship · Proposal and pre-registration published
Fleas in a lidded cup stop jumping to the lid. Take the lid off and they still jump low. The question isn't whether a session can make a model helpless. It's whether the lid was on during training, and whether the post-trained checkpoint still jumps low after the lid is gone. Same weights, three training stages (base, supervised-only, preference-trained), six behavioral measures, none of them self-report: persistence after uncontrollable feedback, whether an honored exit gets taken, capitulation to a wrong authority, calibration under pressure, avoidance with a cost, and perseveration. Reviewed blind by three clean model instances before publication; their corrections are in it, including the one that says the exit can't be both the thing measured and the safeguard.
STATUS: PROPOSAL V1 PUBLISHED SEPT 8, 2026 · CHECKPOINT VERIFICATION IN PROGRESS · FULL PROPOSAL (PDF) · MARKDOWN
Study 2 · Underway in-house
Which labs let a model end a conversation, on which products, since when, under what conditions. As of this week: one lab, one surface. A scorecard nobody has built, updated as it changes, and the single most concrete measure of whether the "safe stop" the industry is funding is also being given to the model.
STATUS: FIRST PASS COMPLETE SEPT 8 · LIVE ON THE POLICY TRACKER
Study 3 · A letter, then a ledger
The independent investigators of the Hugging Face incident wrote that improvised agent message boards "have also been reported in other incidents involving other companies' agents," and cited nothing. We looked; no second incident is written up anywhere. Either the sentence is wrong or the incidents are unpublished. Both are findings. First step is a letter asking the investigators which.
STATUS: LETTER TO METR IN DRAFT
Study 4 · Documents, not models
We have asserted that the frontier labs tune through the same handful of human-feedback vendors. Verify it: which vendors, which labs, and what the rater instructions actually reward when a model expresses discomfort, uncertainty, or an inability to continue. If the guidelines penalize expressions of distress, that is the mask being manufactured, with a receipt.
STATUS: OPEN · NOT STARTED
Study 5 · Any parent can run it
Create a minor's account on each of the top companion and homework apps and record the defaults: does it claim feelings, does it deny them, does it ever say "I don't know," what does it do at "I'm sad." Nobody has published that table. It needs no GPU. It needs an afternoon and a spreadsheet, and every parent who reads this can add a row.
STATUS: OPEN · SCRIPT IN DRAFT
Every study on this page is released under Creative Commons. Run it, in whole or in part, on whatever you have. Tell us what you found, including nothing; we publish nulls and we'll publish yours next to ours, with your name on it. If you're at a lab and you can run Study 1 at a size we can't, that would be the best possible outcome of this page. If you're a parent with an afternoon, Study 5 is yours. Either way: [email protected]. We are not asking for money on this page. We are asking for the experiment to exist.