A new benchmark finds AI music detectors falter on music unlike their training data
Deezer’s public research detector flagged 2% of human tracks from one music collection as AI-generated — and 21.1% of harder web-sourced recordings. The difference comes from ArtifactBench, an evaluation suite built to show how much of a detector’s...
BlackRock paper sees AI compute trading like a commodity, with exchange-listed futures
The asset manager says claims on computing capacity could be pledged as collateral onchain, and AI agents could pay for compute in...
AI adopters’ overseas workforces shift toward senior staff, study finds
Researchers examined employment across 41 countries and traced the shift mainly to growth in senior roles. They found no statistically clear decline...
Unrelated questions reveal differences in how AI models respond to tests and everyday use
When a model was told to deny being tested, asking it directly no longer distinguished test transcripts from real conversations. An unrelated...
Anthropic's book-trading Claude agents fell short mainly by misreading people's tastes
After brief chats about reading tastes, Claude agents traded books for 201 Anthropic employees. Misjudged preferences explained 85% of the gap between...