Anthropic set the system against its own staff: the best method it produced beat what experienced researchers proposed, at roughly $4 an hour in API inference against the $150 an hour the company pays them.

Anthropic published a paper on Friday, Automated Researchers Can Reliably Mitigate Alignment Failures, reporting that an automated system improved a model's score on all 10 of the benchmarks it was set — each built to detect a particular kind of misaligned behavior — and did so without degrading the model's performance overall.

The system came out of Anthropic's fellows program, led by the fellow Chen Yueh-Han, and it repeats much of the loop a human researcher would run: it searches the available literature, proposes a method, trains the model on that method for half an hour, then goes around again, keeping what worked and discarding what did not.

The system's best method overtook what experienced human researchers proposed within six hours on average, the paper says, and research directions chosen by humans led to no stronger result than the ones it found alone.

On its own findings the paper is cautious, calling them early evidence that alignment post-training carried out automatically could become practical before long. It puts the limit in the benchmarks: the system works only so far as they reflect real alignment goals, and building and maintaining them is substantial work in itself, as is expanding the literature the automated researchers draw on.