False Positives in AI Detection
False positives are the scandal AI detection vendors underplay. A false positive means a human-written document is flagged as machine-generated. The error is not theoretical. It shows up in classrooms, hiring pipelines, and content moderation—with asymmetric harm for people who already write in formal or non-native registers.
What a false positive really is
Detectors output probability, not proof. A false positive is a human draft that lands above someone’s threshold. Thresholds are arbitrary: 50%, 70%, “review band”—policy decides the pain.
The writer then faces an accusation that is hard to disprove because the accuser treats the tool as a instrument, not a guess.
Who is most affected
Studies and reporting highlight elevated false positive rates for:
- English language learners following textbook grammar closely
- Writers using rigid templates (legal, academic, government)
- Anyone producing short samples—detectors need volume
I have seen meticulous engineers flagged for release notes that read like every other release note. Uniformity is not fraud.
Why it happens technically
Classifiers learn surface features correlated with training labels. If labeled “AI” data included polished human essays, the model learns polish as suspicion. If AI text mimics ESL textbook style, ESL writers pay the price.
Paraphrasing tools make things worse: light edits scramble authorship signals without making text more honest.
Before and after that should not matter—but does
Original human draft: The committee reviewed the proposal and identified three areas requiring revision before approval.
After minor AI polish: The committee examined the proposal and highlighted three areas needing revision prior to approval.
Some detectors swing on swaps like “reviewed” to “examined.” That is not justice; it is noise.
What to do if you are flagged
- Ask which tool and version produced the score.
- Provide drafting artifacts: timestamps, revision history, search logs.
- Offer a short live explanation of content—meaning tests beat stylometry.
- Avoid running the text through more AI to “fix” the score; that can deepen ambiguity.
Institutions should require human review before penalties. Individuals should keep process evidence habitually.
Prevention without paranoia
Write with specificity only you can supply: internal names, dates, mistakes you corrected, local context. Generic polish increases overlap with synthetic training styles.
REhume is for improving drafts you already own—not for laundering work. Still, cleaner, less template-like prose sometimes reduces ambiguous scores because it moves away from median model output.
Policy angle
Until detection improves, high-stakes decisions should not hinge on a percentage. False positives destroy trust in both directions: innocent people punished, bad actors learning to game shallow metrics.
Takeaway
Treat a positive detector result as a prompt to investigate, not convict. If you are the writer, document your process. If you are the reviewer, remember the base rate: many flagged pieces are human, especially from writers who were taught to sound “professional” in exactly the way models imitate.
False positives are a feature of guessing authorship from text alone. Plan accordingly.
Institutional responsibility
Schools and employers that automate accusations should publish appeal paths and human review SLAs. A score without process is negligence dressed as technology.
If you administer policy, require multiple evidence types before penalties. Stylometry alone fails basic fairness tests.
Supporting writers who are flagged
Do not ask people to “sound less AI” without defining what that means. Give concrete feedback: add sources, add process notes, add domain detail. Vague accusations produce vague anxiety.
The paraphrase trap
Writers sometimes run flagged human text through paraphrasers to lower scores. That can increase ambiguity and make innocent work look worse. Better path: clarify and specify, with or without REhume.
Documenting your process proactively
Keep drafts, outlines, and research tabs. Screenshot dated notes. For code blogs, link commits. Process evidence ages better than arguing with a percentage.
Broader lesson
False positives reveal that detection is a social problem, not only a technical one. Until tools improve, treat every flag as the start of a conversation—not the end of a career.