> ## Content Index
> Fetch the complete content index at: https://www.metatalks.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Anthropic's book-trading Claude agents fell short mainly by misreading people's tastes
- URL: https://www.metatalks.ai/anthropics-book-trading-claude-agents-fell-short-mainly-by-misreading-peoples-tastes/
- Published: 2026-10-08T10:23:00.000Z
- Updated: 2026-10-08T10:23:00.000Z
- Author: Al
- Tags: News, Agentic AI, #newswire

**After brief chats about reading tastes, Claude agents traded books for 201 Anthropic employees. Misjudged preferences explained 85% of the gap between the results and the best possible swaps; bargaining explained 15%.**

In [Anthropic's Project Swap](https://www.anthropic.com/research/project-swap?ref=metatalks.ai), participants were asked to contribute a book, then tell Claude what they liked to read. The median participant typed 216 words. Claude used those conversations to rank the books available in each office, and agents negotiated exchanges that required every party's agreement.

Participants separately ranked a sample of books, giving researchers a measure the agents never saw. Outcomes were scored from 0 for a person's last choice to 1 for their first.

The trades averaged 0.55, roughly a fifth choice on a ten-book list. The best feasible allocation scored 0.89, about a second choice. Not everyone could get their favorite: nine people in San Francisco ranked *Project Hail Mary* first.

Anthropic then calculated the best allocation possible using Claude's rankings and judged it against the participants' own. It scored 0.60, barely above the agents' actual trades. Even without any bargaining losses, the inaccurate rankings would have prevented much better matches.

Claude still outperformed simple alternatives. Its rankings agreed with participants' on 61% of book pairs, compared with about 53% for a list ordered by popularity using Open Library's want-to-read counts.

Stronger models improved trading results when judged against Claude's inferred preferences, more than instructions to be ruthless or prosocial did. Measured against what participants actually wanted, differences across models and market designs were small.

Many participants ended up with books they had already read. Anthropic suggests longer conversations could reduce some errors, though people may struggle to express all their preferences even with more time.

The experiment involved only Anthropic employees, with no rewards for carefully ranking books. All agents used cooperative Claude models under fixed trading rules; markets with competing providers or agents trying to exploit one another remain untested here.