Deezer’s public research detector flagged 2% of human tracks from one music collection as AI-generated — and 21.1% of harder web-sourced recordings.

The difference comes from ArtifactBench, an evaluation suite built to show how much of a detector’s measured performance depends on whether its test music overlaps what it was trained on. The paper on arXiv says commonly used benchmarks make that overlap hard to see: they pool recordings but provide only partial information about which came from related generators or were derived from the same originals. ArtifactBench groups recordings that share source material, so their shared content stays explicit in the evaluation.

The suite tests publicly available detectors under a common protocol, with detector versions fixed. It keeps calibration data separate from final testing, and counts a detector’s failure to return a result separately from a wrong answer — a distinction that can change a model’s measured ranking.

The Deezer figures come from two sets of human-made music: the Free Music Archive, and a separate set of difficult web-sourced recordings the paper uses to test detectors outside familiar material.

The paper reports that ArtifactNet achieved a balanced accuracy — equal weight to each class — of 0.918, compared with 0.776 for the Deezer model, on the 562 test tracks for which every compared detector returned a result.

ArtifactNet failed to return a score for 17 of the 579 test tracks. Treating every failure as an error reduced its balanced accuracy to 0.865.

The paper’s author, Heewon Oh, also authored the earlier ArtifactNet paper. The Deezer model tested was a public research implementation, not the company’s production detector.