The short version
- A detector returns a probability about text, not a finding about a person. Turnitin states plainly that it "does not make a determination of misconduct" and tells instructors to use the score "to initiate a conversation, not to draw a conclusion".
- The score has published limits the letter usually omits: it covers long-form prose only, nothing under 300 words is scored, and Turnitin marks its own figures under 20% because false positives are more common there.
- At a public institution, discipline is state action. Goss v. Lopez requires notice of the charge, an explanation of the evidence, and a chance to respond — but not counsel, cross-examination or witnesses.
- At a private institution there is no constitutional floor. The handbook is the contract, and courts ask whether the institution substantially observed its own published procedure, so quoting its clause numbers beats arguing about the science.
Two questions are tangled together here. One is measurement: what does the number measure, and how much weight will it bear. The other is process: what must the institution do before it acts. Only the second is law.
What the number in the letter actually is
Turnitin's AI writing indicator, which most of these letters rest on, is a prediction with limits the vendor publishes and the letter almost never carries. Four of them decide how much the number can be made to bear.
What the AI writing indicator actually reports
The AI writing report behind the percentage
Two of those turn straight into questions. If the flagged work runs under 300 words there is no score to discuss, because Turnitin raised its own minimum from 150 words in May 2023 on the basis that accuracy "increases with a little more text". And if the figure carries an asterisk, the vendor is saying that its own testing found "a higher incidence of false positives" in that range — a concession from the company, not an argument from you.
Why "less than 1%" is not the reassurance it sounds like
Turnitin's published document-level false positive rate is under 1% where 20% or more AI writing is detected, checked against 800,000 samples written before ChatGPT existed. As a claim about one paper that is comforting. As a claim about an institution it is not.
Vanderbilt did that arithmetic in public when it disabled the detector in August 2023: it had submitted around 75,000 papers to Turnitin in 2022, so "around 750 student papers could have been incorrectly labeled". It also objected that Turnitin "gives no detailed information as to how it determines if a piece of writing is AI-generated or not".
The bias finding, and the honest argument about it
The most cited study is Liang and others, "GPT detectors are biased against non-native English writers", in Patterns in July 2023. Seven commercial detectors were run over 91 TOEFL essays by non-native speakers and 88 US eighth-grade essays. The eighth-grade essays were classified accurately; more than half the TOEFL essays were labelled AI-generated, at an average false-positive rate of 61.3%.
The mechanism is the useful part. Most detectors lean on perplexity — how predictable the next word is — and a narrower vocabulary is more predictable. The effect ran both ways: improving the word choice in the TOEFL essays cut the false-positive rate to 11.6%, and simplifying the eighth-grade essays raised their misclassification. Prompting ChatGPT to "elevate the provided text by employing literary language" dropped detection of genuinely AI-written essays to near zero.
Turnitin's answer deserves stating fairly. Its detector was not among the seven tested, the TOEFL essays were all under 150 words — below its own floor — and its own evaluation found the difference in false-positive rate between learner and native-speaker writing "small, and not statistically significant" above 300 words, but significantly larger below it. That concedes the mechanism rather than denying it.
The wider picture is messier, and not uniformly in your favour. Weber-Wulff and others tested 12 public tools plus Turnitin and PlagiarismCheck in 2023 and found them "neither accurate nor reliable", with "a main bias towards classifying the output as human-written" — the commonest error is missing AI, not inventing it. OpenAI withdrew its own classifier in July 2023 "due to its low rate of accuracy", having caught 26% of AI text while flagging human text 9% of the time.
The distinction that decides what process you can demand
Everything above concerns the evidence. What you may demand in response turns on one fact most students never learn: whether the institution is an arm of the state.
Which body of rules governs the hearing
Who runs the institution?
A public school, college or university
Discipline is state action. The Fourteenth Amendment applies, and Goss v. Lopez sets a floor no handbook can write away.
A private school or university
No constitutional due process. The handbook and honour code are the contract, and courts ask whether the institution substantially observed it.
At a public institution: how thin the floor really is
Goss v. Lopez, 419 U.S. 565 (1975), is the anchor. For a suspension of ten days or fewer, due process requires "oral or written notice of the charges against him and, if he denies them, an explanation of the evidence the authorities have and an opportunity to present his version". The middle clause matters most here: an explanation of the evidence means the report, not the sentence quoting a number from it.
Read the rest before building hopes on it. The Court declined to require "the opportunity to secure counsel, to confront and cross-examine witnesses supporting the charge, or to call his own witnesses", and the hearing may be an informal conversation the same day. Longer suspensions or expulsions "may require more formal procedures" — so the severity of the sanction, not of the allegation, is what buys a real hearing.
The academic side of the line is thinner again. In Board of Curators of the University of Missouri v. Horowitz, 435 U.S. 78 (1978), the Court held that dismissals "for academic (as opposed to disciplinary) cause do not necessitate a hearing", because "misconduct is a very different matter from failure to attain a standard of excellence in studies". An AI allegation is a factual determination about conduct, so it belongs on the disciplinary side. A school routing it through an academic-judgment channel with no hearing has made a classification worth contesting in writing.
One further limit surprises people: the machinery starts only once a protected interest is deprived. In Harris v. Adams (D. Mass., November 2024) a student received zeroes on two of six components, a Saturday detention and an initial rejection from the National Honor Society. None of that was "total exclusion from the educational process" under Goss. A grade penalty alone may trigger no constitutional process at all, which pushes you back onto the institution's own policy.
Put your account of how you wrote it in writing
A dated statement of what you did, on which day, in which order, keyed to the files you still hold. It becomes part of the record the decision-maker reads. A conversation in an office does not.
At a private institution: the handbook is the contract
A private college is not the government, so the Fourteenth Amendment does not reach it. Contract does. Courts describe the relationship as having a "strong, albeit flexible, contractual flavor" and require handbook promises to be "substantially observed" — general compliance rather than letter-perfect, with ambiguities read the way a reasonable student would. Several states add a common-law duty on private associations to treat members with at least minimal fairness, barring discipline that is arbitrary and capricious.
That changes the register. Your strongest argument is rarely that the science is contested; it is that clause 4.2 promised a step and the step did not happen. Find the published procedure in the version in force that term, and check whether it is incorporated into your enrolment agreement. AI policies have been rewritten every year since 2023, and the one that binds you is the one published when you submitted.
What your own institution's policy decides
Read these out of the policy before the meeting
- The standard of proof, and who carries it. UNC Charlotte requires a preponderance and presumes a student "not responsible until determined otherwise"; Pittsburgh requires clear and convincing evidence. Find which phrase yours uses.
- Whether AI use was permitted for this assignment. Where brainstorming is allowed and drafting is not, the score cannot tell them apart.
- What you are entitled to see: the highlighted sentences, the asterisk if there is one, the word count scored, and whatever else is in the file.
- The appeal grounds and the clock. Grounds are narrow — procedural error, new evidence, disproportionate sanction — and the window is short: UNC Charlotte allows five days from the notice of outcome.
- Who decides, whether you may bring an adviser, and whether the adviser may speak.
Where the institution will not hand over the file, your right to inspect your education records is a separate route on its own timetable. A dated written request fixes the moment you asked, which matters if the answer arrives after the hearing.
The evidence that actually answers a score
A detector score is evidence about text. What rebuts it is evidence about process, and that decays: version history is overwritten, library and browser records roll off, and a file re-saved after the accusation looks worse than one left alone.
- 1
Stop editing and freeze the file
Export the version history before anything overwrites it — File then Version history in Google Docs, the version list or OneDrive history in Word — with the dates visible.
- 2
Gather the trail around the document
Outlines, notes, library loans, downloaded PDFs, drafts emailed to yourself, messages to a tutor. Timestamps you did not create for this purpose carry more weight.
- 3
Get the written standard and the full report
The policy in force that term, the assignment brief and its AI instructions, and the report itself rather than the figure quoted from it. Ask in writing.
- 4
Answer in writing before you answer in person
A dated, factual account of how the work was produced, keyed to the files. That is what an appeal panel reads months later.
- 5
Be able to talk about the work
Turnitin tells instructors to ask the student to explain the essay, the process and the key takeaways. Reread your sources: this is the part that persuades a person in a room.
| Evidence | What it shows | What it does not show |
|---|---|---|
| Detector score | A prediction that the prose resembles model output | That a particular person used a particular tool |
| Version history | When text appeared, in what size chunks, over how long | Where pasted text came from; a paste is not an AI paste |
| Drafts, notes and sources | That the argument was built over time from material you handled | That the final wording is yours |
| Talking through the work | Command of the argument and the choices behind it | Nothing alone — an articulate student can still have used AI |
Where a case actually sits
Is there evidence beyond the score?
Was AI use permitted for this assignment?
Banned outright
Permitted in part
The score alone
One number, contested
Everything rests on a probability the vendor says is not a determination. The standard of proof does all the work.
The wrong question
An accurate score still cannot say which use it caught. Permitted brainstorming and banned drafting produce the same flag.
Score plus corroboration
The hard case
Pasted blocks, minutes rather than hours in the document, citations to books that do not exist. Findings are sustained here.
About the line, not the fact
The dispute is whether the use crossed the permitted boundary — a reading of the policy, not of the text.
What a sustained accusation looks like
Harris v. Adams is worth reading because the school did not stop at the flag. Turnitin marked parts of a submitted script as AI-generated. The teacher then ran a revision-history extension, which showed large portions pasted in and roughly 52 minutes spent in the document where others spent seven to nine hours. She ran the work through two further detection tools. The first and third footnotes cited books that do not exist.
On that record the court found "nothing in the preliminary factual record to suggest that HHS officials were hasty", and noted that the student and his parents "were afforded prompt notice of the school's findings and were given an opportunity to be heard". Read it as a template: the flag started the inquiry, four independent things finished it. If your letter contains only the flag, say so — and be ready for the file to hold the rest.
One caution about the reply. Draft your appeal with a chatbot and you hand the panel a document that reads exactly like the thing you are denying; the limits of AI legal advice bite hardest on procedure, because handbook deadlines are local and unguessable. Where AI use was permitted and you used it, read who owns AI-generated output first.
What the score is for
Some of these accusations are correct, and nothing here is a method for beating a fair one. A student who pasted in generated prose and is now assembling a paper trail is doing something panels have seen before.
But the vendor's own position is the most useful material available, and it sits on its own site: Turnitin "does not make a determination of misconduct", there is "no 'right' or 'target' score", and an instructor should "use the information to initiate a conversation, not to draw a conclusion". The company selling the number says the number does not decide.
Which leaves the decision where it always was: a person applying an institutional standard to a record. You cannot control what they conclude. You can control whether the record holds anything beyond the number — the report itself, the policy actually in force, the history of the file, and an account you wrote down before anyone asked for one.
Sources
- Goss v. Lopez, 419 U.S. 565 (1975) (Cornell LII)
- Board of Curators of the Univ. of Missouri v. Horowitz, 435 U.S. 78 (1978) (Cornell LII)
- Harris v. Adams, No. 1:24-cv-12437 (D. Mass. Nov. 2024) — memorandum and order
- Turnitin — Understanding false positives within our AI writing detection capabilities
- Turnitin — AI writing detection update from the Chief Product Officer (300-word floor, 20% asterisk)
- Turnitin — the sentence-level false positive rate
- Turnitin — research on bias against English language learners
- Liang et al., "GPT detectors are biased against non-native English writers", Patterns 4(7) 2023
- Weber-Wulff et al., "Testing of detection tools for AI-generated text", Int. J. Educ. Integrity 19:26 (2023)
- OpenAI — New AI classifier for indicating AI-written text (withdrawn 20 July 2023)
- Vanderbilt University — Guidance on AI detection and why we are disabling Turnitin's AI detector
- FIRE — Procedural fairness at private universities
- UNC Charlotte — Code of Student Academic Integrity (preponderance standard, five-day appeal)
- University of Pittsburgh — Academic Integrity Code (clear and convincing standard)
- 34 CFR § 99.10 — right to inspect and review education records (Cornell LII)
General information, not legal advice. This guide explains how these documents and rules generally work. Law varies by jurisdiction and changes, and none of it is applied to your circumstances here. For anything consequential, consult a licensed attorney where you are.
Frequently asked
Can a school punish me based only on an AI detector score?
Whether it may depends on the standard of proof in its own policy, not on the detector. Turnitin states that it does not make a determination of misconduct and that instructors should use the score to start a conversation rather than draw a conclusion. A policy requiring clear and convincing evidence is hard to satisfy with a single probability; one requiring a preponderance is easier, but still needs someone to weigh it.
What does a Turnitin AI percentage actually mean?
It is the predicted share of qualifying long-form prose the model attributes to AI, not a share of your whole submission. Tables, bullet points, code and bibliographies are excluded. Documents under 300 words are not scored at all, and scores under 20% carry an asterisk because Turnitin found more false positives in that range.
Do I have a right to a hearing before I am punished?
At a public institution, Goss v. Lopez requires notice of the charge, an explanation of the evidence and a chance to give your side before a short suspension — but not counsel, cross-examination or witnesses, with more formal procedures only where the sanction is heavier. At a private institution the right comes from the handbook, and courts ask whether the institution substantially observed the procedure it published.
Are AI detectors biased against non-native English speakers?
A 2023 study in Patterns found seven detectors misclassified more than half of a set of TOEFL essays as AI-generated, at an average false-positive rate of 61.3%, against near-perfect accuracy on US eighth-grade essays. Turnitin was not among the tools tested, the essays were all short, and Turnitin's own evaluation found no statistically significant difference above its 300-word floor. The finding is strongest for short text.
What evidence best answers an AI accusation?
Evidence about how the document was produced rather than about the text: exported version history showing text appearing over time, dated drafts and outlines, library and download records, and messages about the assignment. Export the history before editing anything further, because it is overwritten. Being able to explain the argument and the sources in person carries real weight alongside it.