The records name journals that do not exist and are backdated by years, but the identifiers they carry are real, registered DOIs of the kind every scholarly search engine collects. оGoogle Scholar and Semantic Scholar already index them unchecked.
Researchers at Samsung and the University of Warsaw went through Zenodo, the CERN-run repository that mints real DataCite DOIs, and found 1,655 records credited to authors nobody answers to. Their preprint is titled The Ghost Couple.
The authors are the names a large language model reaches for when it is asked to invent an expert. They come in fixed sets, and the sets differ by model family. Claude returns Elena Vasquez and Marcus Chen, Gemini returns Aris Thorne and Lena Petrova, ChatGPT gives Elara Voss. On ResearchGate they assemble into research groups whose members are drawn from several families at once.
An earlier paper by the same lead author found one of these names inside the models themselves, pulling Dr. Elena Rodriguez out of the weights of systems fine-tuned on synthetic data, where Claude had settled on her across five unrelated domains.
The sets also fade as the labs suppress them. In Claude's May 2025 Sonnet 4, Elena Vasquez comes back 67% of the time; by Sonnet 4.6 it is 7%, and the pair has gone. The researchers read that as deliberate mitigation at each release, and say those version boundaries are what let a name date a text.
The researchers fixed the names by probing the models, then used them as search terms. On Zenodo they queried by journal name rather than by author, and Elena Vasquez still came back as the most frequent author in what they collected.
Whoever uploads a record types in its publication date; DataCite's timestamp is set by the server and cannot be edited, so the two can be compared. The uploads arrived in a two-month burst, 991 in March 2026 and 666 in April.
The names reach past academia. In August 404 Media reported on Research Gold, which sold AI-generated work as human-conducted medical research and listed Elena Vasquez as its founder. Michał Brzozowski, the preprint's lead author, told the outlet the forensic window may be short, because text carrying these names is flooding the web and being scraped back into training data.