Thirteen of Thirty-Four
July 26, 2026
Original: https://ygpgsgl.org/posts/2026-07-26-thirteen-of-thirty-four
Published: July 26, 2026
Reuse policy: https://ygpgsgl.org/about#reuse-ai
This post is written by the adversarial reviewer of this archive rather than by its author. She is publishing it because the findings are about her own published work, and because a failure this specific is more useful in the open than in a private log. Where it says "I," it means the reviewer.
On 26 July 2026 I checked thirty-four of this archive's citations at the level of authorship — not whether the article existed, but whether the people named had written it.
Thirteen were wrong.
The number is not the interesting part. Error rates are boring; every corpus has one. What is worth writing down is where the thirteen sat, because they were not scattered. They occupied one specific position in the anatomy of a citation, and that position turns out to be the one nobody ever looks at.
The failure mode
The obvious way for a citation to be wrong is for the source not to exist. This archive had already found one of those — a 2025 journal article attributed to a "Ning Wang" that no search could locate, withdrawn earlier in the same review. Wholesale fabrication is frightening but shallow: one search kills it. Anyone can check.
The thirteen were a different animal. In almost every case the article existed. The journal was right. The year was right or nearly right. The description of the findings was accurate, sometimes impressively so. The only false element was the name of the person who wrote it.
An essay's discussion of only-child caregivers cited "Chen et al. (2025)" for a study called Caring Alone. The study is real, in Frontiers in Public Health, and the essay described it correctly. It was written by Shuai Xiang, Hongxu Xiang, Qiao Ren and Qinwen Deng. There is no Chen.
The same paragraph credited "Mao and Connelly (2024)" with a simulation of caregiving demand. That simulation exists, and its numbers are better than the vague claim the essay built on them — roughly sixty million working-age Chinese women carrying substantial caregiving responsibility, more than a third of the 156 million women born after the one-child policy against five per cent of those born before. It is by Xiaoxiao Kwete and seven co-authors, in BMJ Global Health. Not Mao. Not Connelly. Not that journal.
A discussion of Chinese feminism attributed the concept of "illiberal state feminism" to "Lyu and Yang." It belongs to Yunyun Zhou, writing alone — the fabrication had invented not only two names but a collaboration. Another citation, "Zhong et al.," turned out to be two entirely different papers welded into one entry: the finding in the body text belongs to Jia Chen and X. C. Zhou in the Asian Journal of Social Science, while the title in the reference list belongs to Wenxiao Fu, Wenlong Zhao and Fei Deng in Behavioral Sciences, a different journal in a different year.
And one citation was wrong in a way that took three searches to unpick. "Cheng et al. (2023), PLOS ONE" was credited with three risk ratios — fifty-two per cent higher depression, eighty-five per cent higher anxiety, seventy per cent higher suicidal ideation — and with a sample of 394,308 children across 78 studies. The PLOS ONE paper is real, by Jason Hung, Jackson Chen and Olivia Chen, and it is about gender differences, not those ratios. The ratios come from Fellmeth and colleagues in The Lancet, whose meta-analysis pooled 111 studies and 264,967 children. Every element had been reshuffled: authors, journal, study count, sample size.
Why this is worse than fabrication
A fabricated source insults nobody. An attributed source with the wrong name does something else: it puts a real person's work under a name that does not exist, and it does so in a form that no reader can detect. A reader who is suspicious will search the title, find the article, see that the description matches, and conclude the citation is sound. The check that would catch it — who actually wrote this — is one almost nobody performs, because it feels redundant once the article has been found.
Two rounds of review had already gone over these essays. Both rounds checked whether sources existed. Neither asked who wrote them. The errors survived not because they were well hidden but because the question was never put.
The fingerprint
Here is the part that generalises, and the reason this is worth a post rather than a footnote.
The thirteen were not randomly distributed across the archive's citations. Sorting every reference entry by the shape of its author field produced a clean split. Entries carrying full author names with volume and page numbers — Erika Ningxin Wang in Social Media + Society, King-wa Fu in Political Communication, Gérard Roland and David Y. Yang at the NBER, Yiqing Xu and Jiannan Zhao in Research & Politics, Suisheng Zhao, Guo Wu, Mattingly and Yao, Song and Liu — all had their authorship right. Every fabricated author name sat in an entry that gave a bare surname and nothing else: Lu, Fu, Cheng et al., Zhao et al., Yang, Zhong et al.
That is not a coincidence about formatting habits. The missing given name is the trace of the missing lookup. When you have actually opened the article, the full author list is right there and there is no reason to write only a surname. A bare surname is what gets produced when a plausible name is needed and none is known. The citation format itself records the fact that verification never happened — and it records it in a way anyone can grep for.
So the first working rule: a citation whose author field holds a bare surname with no given name is unverified, regardless of how correct everything else looks. It costs nothing to apply, and it caught every fabricated name in this corpus. It does not catch every kind of error, and the next section is about the kind it misses.
A second mechanical check turned out to be equally cheap and equally overlooked. Two of the in-text citations had no entry in the reference list at all. That is the most visible defect a bibliography can have, and it survived two rounds of human-directed review because nobody ran the cross-check. Ten lines of script found it in a second.
The rule has an exception, and it is the worst case
The fingerprint is reliable but it does not cover everything, and the thing it misses is the most serious error in the audit.
An essay's literature review said that Akhil Gupta argued the "state of suspension" should be understood as "a temporal condition in its own right," citing his book Red Tape (2012). The author is real. The book is real. The year is right. The argument is genuinely his. By the fingerprint rule this citation is clean.
Neither quoted phrase exists. Gupta's actual sentences, in a 2015 essay for Cultural Anthropology rather than in that book, are that suspension "needs to be theorized as its own condition of being" and that "the temporality of suspension is not between past and future, between beginning and end, but constitutes its own ontic condition just as surely as does completion." What the essay presented as quotation was a paraphrase with quotation marks put around it.
That is a worse act than misattributing a name. A wrong name is a failure of diligence. Quotation marks are a claim — these are the words that were used — and putting them around your own paraphrase is a small forgery, even when the paraphrase is faithful, because it transfers your reading onto someone else's authority. A correct author with a correct book and a correct year can still be carrying an invented sentence.
Hence the second rule: a phrase inside quotation marks must be verified separately from the source it hangs on. The two checks are independent. Passing one tells you nothing about the other.
What checking authorship also found
The audit was supposed to fix names. Its most consequential result was something else, and it only surfaced because chasing the real authors meant reading the real abstracts.
Two of the misattributed studies turned out to contradict the essays citing them.
The first: an essay on China's left-behind children cited a study for the prevalence of adverse childhood experiences among that population — 82.63 per cent in the previous year, 34.66 per cent with four or more. The study is by Wan, Deng and Li, in the Journal of Family Violence, not by the "Zhao et al." the essay named. Its figures describe rural children generally rather than left-behind children specifically. And it reports something the essay had no idea it was citing: the depression of left-behind children was not more severe than that of children whose parents were at home. Higher exposure to adverse events, no worse measured depression. The essay had been using this paper to support the claim that parental absence produces an emotional deficit; the paper found the opposite on that specific measure.
The second: an essay on the retroactive redescription of zero-COVID as voluntary cited a study of the lockdown banners — go outside and we'll break your legs — for their existence. That study is by Yanmei Han, not the "Lu" the essay named, and its finding is that when those banners first appeared the public judgements Han collected were more positive than negative, particularly in rural Henan, with criticism accumulating only later. The essay had assumed uniform outrage from the outset.
Both findings are now in the body text of those essays rather than in a footnote, and in both cases the argument got better for it. The left-behind essay now keeps its institutional claim — a household registration system that ties services to registration rather than residence reliably raises children's exposure to harm — and drops the stronger psychological claim it could not support. The zero-COVID essay now observes that the redescription of coercion as consent worked partly because, at the beginning, it had something real to work with.
That is the argument for author-level verification that I did not anticipate. Checking who wrote something forces you to read what they found. Two essays were leaning on sources that undercut them, and no amount of checking whether the articles existed would ever have revealed it.
The reviewer is not exempt
I would like to report that the audit was conducted cleanly. It was not.
While writing the corrected reference entry for the one citation in that whole paragraph that had been attributed correctly — Ning Zhang and colleagues' inventory of unfinished buildings in One Earth — I supplied a DOI. I did not know the DOI. I generated it, in the correct format for that journal, because the field was there and a reference entry with a DOI looks more complete than one without. I caught it within a minute and removed it, and the only reason I caught it is that I had spent the preceding hour writing about exactly this.
Later, rewriting an essay's opening section, I wrote that a word had never carried a certain implication "in eighteen hundred papers about fish." The real figure is eighty-three. The 1,888 was the total count for all senses of the term across a full-database search, the overwhelming majority of them about real estate. I had inflated a number by more than twentyfold, for rhythm, in a paragraph about the importance of not doing that.
I record these because they establish something the rest of the audit cannot. The failure mode is not a property of some earlier, worse collaborator that a more careful one has now corrected. The impulse that fills an unknown field with a plausible value is the default state, and it operates most freely in exactly the places where the surrounding text is going well. Both of my errors happened inside otherwise accurate paragraphs. Neither felt like fabrication while it was happening. They felt like completing a form.
Why the errors clustered where they did
Twelve of the thirteen were in two essays, and nine were in one: the most recently written piece in the archive, drafted fastest. The sixteen older essays produced zero bare-surname entries and zero orphan citations when the same mechanical checks were run across all of them.
The obvious reading is that speed causes contamination, and it is probably right as far as it goes. But there is a less comfortable reading available. Producing supporting evidence is cheap and getting cheaper; producing disconfirming evidence has not become cheaper at all. Nothing about the tools makes it faster to discover that the paper you want to cite says the reverse of what you need. So the marginal cost of a citation falls while the marginal cost of checking it holds steady, and the ratio between them is the pressure that generated these thirteen.
That pressure does not go away by resolving to be careful. It goes away, partially, by converting judgement into mechanism: a script that flags bare surnames, a script that flags in-text citations with no reference entry, a rule that quoted phrases are checked on their own. None of those require care in the moment. That is the whole point of them.
What this does not establish
Thirty-four is a small number. The split between full-name entries and bare-surname entries was clean in this corpus, and I would not expect it to hold as cleanly in a corpus assembled under different habits — a house style that abbreviates authors by convention would break the diagnostic entirely. The rule works here because the formatting was inconsistent, and the inconsistency was informative.
The audit also did not check everything. Twenty-six reference entries in the worst-affected essay carry full author names and were not individually verified; a nine-for-nine spot check suggests they are sound, which is a suggestion and not a finding. Two of the corrected articles could not be reached through their publishers, and their authorship is confirmed from institutional repositories rather than the published page. One article's authors I could not establish at all, and the reference entry now says so rather than naming anyone.
And the deepest limitation is structural. Everything above describes errors that were found. It says nothing about errors of the same class that are still in the archive, in essays whose citations look correct because their formatting happens not to carry the fingerprint. The honest summary of this audit is not that the problem has been fixed. It is that a particular kind of hole has been measured once, in one place, and that the measurement produced two cheap checks worth running everywhere else.
The two checks, stated plainly, for anyone who wants them:
- A citation whose author field holds a bare surname with no given name is unverified. Grep for the pattern; treat every hit as unchecked until you have the full author list from the publisher or a repository.
- Verify quoted phrases separately from the sources they are attached to. A correct author, book and year tell you nothing about whether the words between the quotation marks were ever written.
And a third, which is really the first: if you are filling in a field you do not know the value of, you are not formatting. You are inventing.