$17 Million for AI Access. $1.1 Million to Detect It.
Universities are spending millions to give students AI tools, then spending millions more to catch them using those tools. Nothing about this is new. The answer is not better detection.
In 2024, the California State University system paid OpenAI $17 million to give all 460,000 students, 63,000 faculty, and staff across 23 campuses access to an education-specific version of ChatGPT. They have since renewed, at $13 million a year.
In 2025, the same system paid Turnitin $1.1 million for plagiarism and AI detection, including $163,000 specifically for the AI-detection add-on. Total plagiarism-tech spend since 2019: $6 million.
Same institution. Back-to-back years. Two bets running in opposite directions: hand them the tools, then scan their papers for evidence they used them.
No published source we can find has framed these two numbers as a single institution's contradiction.
And here is the thing that keeps nagging at me: none of this is new. A student can paste AI output into a paper exactly as easily as they can paste a paragraph from a website. Plagiarism has always existed. We have always known how to solve provenance when it actually matters: proctored exams, cameras, controlled environments. Those tools still work. Nothing has changed except the fear.
The money is everywhere
CSU is not alone. Arizona State's OpenAI contracts grew from $300,000 to $2.1 million across three contracts. The University of Colorado signed a $2 million-a-year deal in February 2026, drawing enough faculty backlash to delay the rollout. The University of Maine system committed $1.39 million, starting this month. OpenAI has sold over 700,000 ChatGPT Edu licenses to at least 35 public universities.
The University of Michigan invested $20 million in OpenAI, now worth roughly $2 billion, plus $180 million across two funds managed by Sam Altman's Hydrazine Capital. A university with a financial stake in the company whose product it is also policing on its own campuses.
Google pledged $1 billion over three years for AI training in higher education. The University of Manchester became the first university worldwide to give all 65,000 staff and students access to Microsoft 365 Copilot.
On the other side of the ledger, California institutions alone spent $15 million on detection tools across 57 campuses, tracked by CalMatters through public records requests. UC Berkeley signed a ten-year, roughly $1.2 million Turnitin contract. The LA Community College District pays $265,000 a year. Turnitin, acquired for $1.75 billion in 2019, reports estimated revenue of over $200 million annually and maintains a database of 1.9 billion student papers.
The money flows both ways, and the amounts are not close.
Detection fails where it matters
To be fair: on unedited, fully AI-generated text, current detection tools perform reasonably well, catching 90 to 95 percent. If every student who used AI simply pasted raw ChatGPT output and submitted it unchanged, the tools would mostly work.
Nobody does that. Students edit. They blend. They rewrite sections and keep others. And that is exactly where detection falls apart, which is exactly where the stakes are highest.
A 2023 study published in Patterns (Cell Press) by Liang, Yuksekgonul, Mao, Wu, and Zou tested AI detectors against non-native English writing and found a 61.3% average false-positive rate. Of the TOEFL essays they tested, 97.8% were flagged by at least one detector as AI-generated. When those same essays were rewritten with more sophisticated vocabulary, the false-positive rate dropped from 61.3% to 11.6%. The sample was small (91 TOEFL essays), but the finding has been replicated: Giray (2024), Hadra (2026), and Pratama (2025) all confirm the ESL bias is real and systematic.
Read that finding one more time. The detectors punish authentic non-native voice and reward students who learn to game the prose style. The tool is training students to sound less like themselves.
Independent testing has found a false-positive rate of roughly 4% on human-written text. Vanderbilt calculated that even a 1% rate meant roughly 750 of its 75,000 submitted papers could be wrongly flagged in a single term. A 2023 study in the International Journal for Educational Integrity tested 14 detection tools. None scored above 80% accuracy. Only five crossed 70%.
Turnitin has shipped updates since these studies: AIR-1 in 2024 for paraphrase detection, AIW-2 in 2025 claiming better ESL handling, and a February 2026 model update. None of these have been independently verified to close the ESL gap. Turnitin's own data still shows a 6 to 9 percent false-positive rate for non-native speakers, compared with 1 to 4 percent for native speakers.
The people behind the percentages
False positives are not abstractions. They are specific students, with names, in specific classrooms.
At Adelphi University, student Orion Newby submitted a paper for World Civilizations 1. Turnitin flagged it as fully AI-written. Two other detection tools said it was human-written. Newby, who has documented learning differences, had written the paper with help from university-provided tutors. A New York state Supreme Court ruled the finding "without valid basis and devoid of reason" and ordered Adelphi to expunge the charge. A student using the support system his university gave him, punished by a tool his university also gave him.
At a high school in Palo Alto, Turnitin scored a student's essay on The Crucible as 76% likely AI-generated. The family filed a federal civil rights complaint in May 2026, submitting 1,162 pages of evidence including drafts and full revision history, alleging false accusation, punishment without parental notification, and "malicious targeting, confirmation bias." The case is pending.
At Yale's School of Management, an EMBA student was suspended for a full year after a professor ran exam answers through GPTZero. The student, a French national and non-native English speaker, alleges coercion toward a false confession and denial of due process. Yale contends the punishment was for "not being forthcoming" to the Honor Committee rather than for the detection flag itself. A federal judge denied the student's injunction to return to campus. The case is still active.
A UK survey of 2,373 students, conducted by YouGov and commissioned by Studiosity, found 75% of those using AI felt stressed their work would be wrongly flagged. Fifty-two percent cited "being accused of cheating when I did nothing wrong" as a specific source of stress. International students were twice as likely to report "a lot" of stress. The methodology is sound. The funder, Studiosity, sells AI study-help tools, which is worth noting, but the polling firm is independent and the numbers have been reported across Times Higher Education and Inside Higher Ed without challenge.
The pattern is consistent: the students most likely to be falsely flagged are those writing in a second language, those with unconventional styles, and those without the resources to fight the accusation. Being punitive has never worked in these spaces. It does not work now.
The institutions reading the evidence
Vanderbilt disabled AI detection in August 2023, stating plainly: "We do not believe that AI detection software is an effective tool." That policy has not been reversed.
The University of Waterloo disabled Turnitin's AI detection in September 2025 after internal testing found the tool flagged human-written text as 100% AI-generated. Curtin University in Australia followed in January 2026, citing a need for "trust and clarity" over automated detection. The Australian Catholic University abandoned it in March 2025. Indiana University's Kelley School of Business prohibits faculty from using any AI detection tool. The University of Cape Town discontinued detection in October 2025 and adopted an AI in Education Framework focused on assessing the process of learning, not policing the product.
As of early 2026, over 50 universities across the US, Canada, the UK, Australia, and South Africa have disabled or formally restricted AI detection tools. Proposed legislation in multiple states and federal guidance now caution against using AI detection output as the sole basis for academic misconduct proceedings.
A Berkeley study of 31,692 course syllabi found faculty are broadly moving away from outright bans on AI use. UNESCO's AI Competency Framework for Students positions them as "AI co-creators and responsible citizens" rather than suspects.
The direction is clear. The question is what replaces detection.
Why is AI so polarizing
Cell phones were going to ruin attention spans. Wearable tech was going to make us all cyborgs. VR is still waiting for its moment after a few false starts. Every pervasive technology follows the same arc: fear, stigma, normalization, and then the quiet realization that it was never the technology that mattered, it was what people did with it.
AI is following the same path, but the fear is sharper this time, because it touches something people care about more than their screen time: their work. Their identity as makers. The question every creative quietly asks now is whether anyone will believe the work is really theirs. And honestly, that should not be a new fear. Attribution and trust have always come from the same place: do you trust who worked on this? Different people have different thresholds for that. A reader, a professor, a client, a collaborator, they all have their own trust ladder. They always have.
What is new is that nobody has given them a way to climb it. A checkbox that says "AI was used" tells you nothing. A blanket disclaimer at the bottom of a syllabus tells you nothing. You cannot apply your own judgment to a yes-or-no binary, because a binary carries no information. You would only believe someone about a private claim if they told you the specifics, the what and the how, not just a yes or no.
The only thing I can think of that actually solves this: show the work, abstracted by the act. Who devised it. Who authored it. Who reviewed it. Who prepared it for release. Let every person who encounters the work apply their own trust ladder to those specifics and make their own call. That is what attribution frameworks are for.
What if the answer is not detection at all
The alternative is deceptively simple: instead of scanning finished work for statistical traces of AI, ask people to say how they used it. Specifically.
Not a checkbox. Not a disclaimer. A record: what was conceived by the student, what an AI helped draft, who reviewed the work, what was rewritten by the student's own hand. The detail is what makes the claim credible. A yes-or-no binary never could.
DARP is one framework for exactly this kind of record, designed to describe the layers of contribution in human-AI work. It is open, CC-BY licensed, and free. There are others. The point is not which framework. The point is that frameworks for this exist and the institutions spending millions on detection have not tried them.
The research community is arriving at the same idea independently. A 2026 paper on arXiv frames AI misuse in education as a measurement and visibility problem, arguing that the gap between what a student submits and what can be attributed to their own learning is the real issue to solve. The paper's conclusion aligns closely: attribution is a better answer than detection.
The phrase "attribution over detection" as a position, as a policy stance, is open territory. Nobody in higher education has staked it out. The idea is not new, Yale, Cornell, and others have gestured toward it, but nobody has named it, standardized it, or built a framework around it. The institutions caught in the spending paradox have not found the words for what they are looking for.
A note on Turnitin
This is not a hit piece on Turnitin. They built a real product that addressed a real problem, and a federal court affirmed their right to build it. But it is worth noticing the incentive structure.
Turnitin maintains a database of 1.9 billion student papers, collected largely under institutional mandates with limited student choice. They were acquired for $1.75 billion. They have since added ExamSoft (exam software, 2020) and ProctorExam (remote proctoring, 2021). They now cover the full pipeline: write the test, proctor the test, check the test. Their revenue grows when institutional anxiety about academic integrity grows. There is no business incentive to declare the problem solved.
That is not fraud. It is a misaligned incentive. And it is worth naming, because the institutions writing the checks should understand what they are buying and what problem it is actually structured to solve.
A student hands in a paper. At the bottom, five lines: what she came up with, what an AI helped her draft, who reviewed it, what she rewrote herself. Her professor reads the record and knows exactly what they are looking at.