The world’s most widely used plagiarism detector cannot detect plagiarism. It never could. What Turnitin and its rivals actually do is match strings of characters — a task closer to spotting duplicate barcodes than to reading — and for twenty-odd years universities have treated the resulting percentage as if it were a moral verdict. It is not. It is a count of coincidences, some damning, most meaningless, and the entire apparatus of academic integrity has been leaning on it like a drunk on a lamppost: for support rather than illumination.
This matters right now because of what has been happening at Cambridge, where allegations of plagiarism concerning Professor Jason Arday’s 2015 doctoral thesis have been circulating — allegations which remain contested, which Arday has not been found by any published process to have committed, and which have, entirely predictably, curdled into a culture-war slanging match within about a fortnight of surfacing. One trench is convinced the case proves everything it already believed about diversity appointments; the opposite trench is equally convinced it proves everything it already believed about who gets scrutinised and why. Neither trench, as far as anyone can tell, has spent much time on the actual text.
Which is a shame, because buried under the noise is a genuinely important shift in how the question “did this text come from that text?” can be answered at all.
Arday’s appointment at Cambridge in 2023 was widely celebrated, and the subsequent allegations about his 2015 thesis were always going to be radioactive. To be scrupulous about it: allegation is all they are. No adjudicated finding of misconduct against him has been published, and questions of what he did, knew or intended remain precisely that — questions, for due process rather than a comment section.
But a recent forensic analysis of the thesis, supplied to this desk, is interesting less for what it alleges about one academic than for how it goes about alleging it. Rather than running the document through a legacy checker and waving a percentage around, the analysis tested 11.09 million sentence pairs — every sentence in the thesis against every sentence in a set of candidate source texts — and compared the results against a control baseline of 34 million sentence pairs drawn from documents known to be independent of one another. The method is the story here. Because whatever eventually happens in the Arday case, this is what the machinery of textual forensics is going to look like from now on, and it is worth understanding before it gets pointed at anyone else. Which it will.