Yes. Turnitin extracts the text from any PDF that contains a real text layer and runs it through exactly the same AI writing detection it applies to Word documents. The file format is irrelevant; the words inside it are everything. The only PDFs Turnitin can't analyze are ones it rejects outright — and submitting one of those creates a different, usually worse, problem.
That's the whole answer in three sentences. But the question behind the question — "is there a file format that shields my writing from AI detection?" — deserves a longer, more honest treatment, because the answer to that is no, and understanding why will save you from a family of tricks that reliably backfire.
Why the PDF question comes up at all
The theory goes something like this: Word documents are "native" text, so maybe Turnitin reads them more deeply. A PDF is a fixed, printed-looking thing — maybe the detector treats it as an image, or skims it, or can't parse it properly. Some version of this idea circulates on every student forum, usually alongside its cousins: submit a screenshot, use a weird font, swap in lookalike characters.
The theory fails on the first fact of how Turnitin works. Every submission, regardless of format, goes through text extraction before anything else happens. A .docx file, a PDF, a plain .txt file — all of them get reduced to the same thing: a stream of words. The Similarity Report runs on that stream. The AI writing report, the feature Turnitin launched on April 4, 2023 inside the existing Similarity Report, runs on that stream too. By the time detection happens, Turnitin has no idea and no interest in what container the words arrived in.
This is worth internalizing because it generalizes. AI detectors — Turnitin's included — measure statistical properties of text: how predictable each next word is (perplexity) and how much sentence rhythm varies (burstiness). We've written a full explainer on how detectors actually work, but the short version is that the analysis operates on language, not files. Changing the envelope doesn't change the letter.
What Turnitin actually accepts
Since we're talking formats, here are the real rules, straight from Turnitin's own file requirements documentation. Two different sets of requirements matter, and people constantly confuse them.
For the Similarity Report (the plagiarism side), Turnitin accepts Microsoft Word (.doc/.docx), PDF, plain text, RTF, HTML, OpenOffice (.odt), Google Docs submitted via Google Drive, Hangul (.hwp), PowerPoint (.pptx), Corel WordPerfect, and Adobe PostScript. Files must be under 100MB, under 800 pages, and contain at least 20 words.
For the AI writing report, the requirements are tighter. Per Turnitin's AI-report file requirements page, the file must be .docx, .pdf, .txt, or .rtf; it must contain at least 300 words of prose in a long-form writing format; it must not exceed 30,000 words; and the supported languages are English, Spanish, and Japanese.
| Requirement | Similarity Report | AI Writing Report |
|---|---|---|
| PDF accepted? | Yes (with selectable text) | Yes (with selectable text) |
| Other formats | .doc/.docx, .txt, .rtf, .html, .odt, .hwp, .pptx, Google Docs, WordPerfect, PostScript | .docx, .pdf, .txt, .rtf only |
| Minimum length | 20 words | ~300 words of continuous prose |
| Maximum | 800 pages / 100MB | 30,000 words |
| Languages | Many | English, Spanish, Japanese |
| Scanned/image PDFs | Rejected | Rejected |
Notice what's on both lists. PDF isn't an exotic edge case Turnitin handles awkwardly. It's a first-class, explicitly supported format for both reports. If anything, PDF is the format Turnitin's documentation discusses most, because it's the one with a genuine failure mode: the missing text layer.
Text-layer PDFs vs. scanned PDFs: the one real distinction
Not all PDFs are the same under the hood, and this is the only place where format genuinely matters.
A text-layer PDF — the kind you get from "Save as PDF" in Word, Google Docs, or LaTeX — stores actual characters. You can select the text with your cursor, copy it, search it. Turnitin extracts this text trivially and analyzes it in full. This describes the overwhelming majority of PDFs students submit.
A scanned or image-based PDF is different. It's photographs of pages wearing a PDF costume. There are no characters inside, just pixels. Turnitin's documentation is blunt about these: the system will not accept "PDF image files, forms, portfolios, files that do not contain highlightable text (e.g., scanned documents that are images), or documents containing multiple embedded files." Turnitin even suggests a self-test: copy a chunk of your PDF and paste it into Notepad or TextEdit. If nothing appears, the file has no selectable text and will be rejected.
So there's the loophole, right? Scan your essay, submit the image PDF, and Turnitin can't read it?
Slow down. Walk through what actually happens.
First, in most configurations the submission simply fails or the report errors out, and you get told to resubmit in a readable format. You've bought yourself nothing but a deadline problem. Second, if the assignment somehow accepts the file, your instructor opens a Similarity Report showing 0% matched text and no AI score on a ten-page essay — which reads less like a technical hiccup and more like a flare fired over your paper. Instructors have seen this move. Many syllabi now explicitly require machine-readable submissions for exactly this reason. Third, the fix is trivial on their end: modern OCR (optical character recognition — software that converts images of text back into text) is built into Adobe Acrobat and a dozen free tools. An instructor who suspects something can reconstruct your text in about ninety seconds and run it through whatever checker they like, now with the added context that you appeared to be hiding it.
The scanned-PDF play doesn't defeat detection. It defers detection while advertising intent. That's a bad trade in any integrity process, because intent is the thing academic misconduct panels care about most.
Character tricks: why they look worse than the thing they hide
The scanned PDF has more sophisticated siblings, and they deserve a section because they're the purest illustration of why format games fail.
The classic versions: replace Latin letters with visually identical Cyrillic or Greek ones (the "а" in Cyrillic looks like the "a" you're reading now); insert white-colored characters or zero-width spaces between letters to break up words; hide a wall of white-on-white text to dilute a similarity score. PDFs get singled out for these tricks because the format makes text encoding less visible to a casual reader — what prints identically can be encoded very differently.
Here's what actually happens when these files hit Turnitin.
Turnitin has shipped explicit countermeasures. Its Similarity Report has a Flags tab that opens an "Integrity Flags for Review" panel, alerting instructors to hidden text, replaced characters and "further text manipulations" in any document of at least 500 characters. Turnitin is careful to add that "a flag is not necessarily an indicator of a problem." The company built this because the tricks were common enough to build for. A flagged submission doesn't show a lower score; it shows the instructor a notice that says, in effect, someone tried something here.
Even setting the countermeasures aside, think about the mechanics. Text with substituted characters is gibberish to any text processor — spellcheckers, screen readers, search, and yes, extraction pipelines. If the trick "works," your essay's extracted text is corrupted nonsense, which produces bizarre report behavior an instructor will investigate. If it doesn't work, the substitutions are detected and flagged. There is no branch of this decision tree where you end up better off.
And now weigh the two possible offenses. An elevated AI score is ambiguous evidence — Turnitin itself displays scores of 1–19% as an asterisk rather than a number, its own acknowledgment that low-range scores aren't reliable enough to print, and false positives are a documented reality the company openly discusses. An instructor looking at a 34% AI score has genuine uncertainty, and honest institutions treat the score as a conversation starter, not a verdict. But hidden white text? Cyrillic character swaps? Those are not ambiguous. They don't happen by accident. They convert "maybe this student used AI, let's talk" into "this student deliberately attempted to deceive an academic integrity system" — a distinct and usually more serious offense category at most universities, regardless of whether the underlying essay was written by you, ChatGPT, or your grandmother.
This is what we mean by "backfire." The trick's failure mode isn't detection of AI. It's documentation of deception.
What the AI report does with your PDF, precisely
Assuming a normal, readable PDF — which, again, is what you should submit — here's the pipeline your text goes through.
Turnitin extracts the text and segments it into chunks of roughly a few hundred words. Each segment is scored by a classifier trained to estimate whether the prose is machine-typical: highly predictable word-by-word, evenly rhythmic sentence-to-sentence. The document-level percentage your instructor sees is the share of the text the model believes is AI-generated. Turnitin claims 98% accuracy with under a 1% false positive rate, but — and this qualifier does a lot of work — both numbers apply only to documents where more than 20% of the text is flagged. Below that, scores show as an asterisk. The system also needs roughly 300 words of continuous prose to run at all, which is why it exists for essays but not for lab-notebook fragments or bulleted slide decks (we've covered what happens when you submit a PowerPoint separately — the answer differs more than you'd expect).
None of these numbers change by a single decimal based on file format. A .docx and a PDF containing the same 2,000 words produce the same analysis. Turnitin's AI-report documentation does make one mildly interesting recommendation: that instructors have all students submit in the same format for consistency of processing. That's an admission that extraction across formats can differ in tiny ways — line-break handling, footnote placement — but it's consistency housekeeping, not a detection gap.
The scale here is worth a moment, too. Turnitin reported screening over 200 million papers in the AI detector's first year (April 2023–April 2024), with about 11% showing 20%+ AI writing. Whatever share of those 200 million arrived as PDFs — plausibly most, given how common PDF submission portals are — they were all processed the same way. There is no format-shaped hole in a system operating at that volume.
The strongest case for format games — and why it still fails
Fairness demands we give the other side its best argument. Here it is: text extraction is genuinely imperfect. Two-column layouts can scramble reading order. Equations, tables, and footnotes extract messily. A PDF generated by an obscure tool might garble ligatures ("fi" becoming a single glyph that extracts as nothing). Couldn't a sufficiently mangled-but-legitimate PDF muddy the analysis?
Marginally, sometimes, yes. And it doesn't matter, for two reasons.
First, the mess cuts randomly, not in your favor. Garbled extraction produces garbled statistics, which can just as easily raise suspicion — weird segments, error states, an instructor squinting at a report that doesn't look like the other twenty-nine. You're not steering the outcome; you're rolling dice while attached to your name and student ID.
Second, and more fundamentally: the human reads the human-readable version. Your instructor doesn't grade the extracted text stream. They grade the PDF as rendered — and increasingly, they read it with the question "does this sound like this student?" in mind. Every format trick targets the machine layer and leaves the human layer untouched. But the human layer is where the judgment actually happens. An essay that reads like unedited ChatGPT — the confident emptiness, the tidy topic sentences, the total absence of your voice from class — gets flagged by a person no matter what the software says. Turnitin's own chief product officer has said the system deliberately leaves roughly 15% of AI text unflagged to keep false accusations down (BestColleges, April 2023); instructors know the score is a floor, not a ceiling, and read accordingly.
Format games are the weakest possible strategy because they attack the only layer of the system that doesn't matter.
What actually determines whether writing gets flagged
Strip away the container and you're left with the real variables, which have nothing to do with file extensions.
Raw AI output is the case detectors handle best. It is also the only case where the independent record is relatively kind to them — every tool in the largest peer-reviewed test still scored below 80% overall. Pasting ChatGPT into a document and exporting to PDF is not a strategy; it's a delay.
Genuinely human writing is usually fine — with documented exceptions. The false-positive risk concentrates in specific groups: non-native English speakers (a 2023 study in Patterns found seven detectors falsely flagged an average of 61.22% of human-written TOEFL essays, though Turnitin wasn't among the seven tested), students taught rigid essay formulas, technical writers, and heavy self-editors. If you're in one of those groups, your protection isn't a file format — it's process evidence: drafts, version history, notes.
The middle is where honesty gets complicated. Writing drafted with AI assistance and then substantially rewritten in your own voice is the hard case for every detector — the Washington Post's April 2023 testing showed Turnitin struggling most with exactly these blended documents, and Turnitin acknowledged it. Where your course allows AI assistance, that editing work is legitimate; where the course bans AI entirely, no amount of rewriting makes it compliant, and we'll say that plainly rather than pretend otherwise.
If you want to know how your writing reads to a detector before an instructor does, run it through an AI detector yourself and look at which sentences trip the classifier — that tells you something about your prose's statistical texture, which is actionable, unlike anything about file formats. Tools like HumanFlow's AI humanizer exist for the legitimate middle ground — reworking permitted AI-assisted drafts toward your own register — but it doesn't promise to beat any detector, because nobody can honestly promise that, and this site refuses to publish bypass rates for exactly that reason.
For the full picture of what Turnitin's detector can and can't do — scores, thresholds, accuracy claims, institutional pushback — start with our Turnitin AI detection hub.
FAQ
Can Turnitin detect AI in a PDF file? Yes. Turnitin extracts text from any PDF with a selectable text layer and runs the identical AI analysis it applies to Word documents. PDF is an explicitly supported format for both the Similarity Report and the AI writing report. The format provides no protection whatsoever.
Will Turnitin accept a scanned PDF of my essay? No. Turnitin's file requirements explicitly reject image-based PDFs and any file without highlightable text. Scanned documents typically fail at submission or produce an error, and you'll be asked to resubmit in a readable format — often with your instructor's attention now fully engaged.
Do character-substitution tricks (lookalike letters, white text) work on Turnitin? No, and they're dangerous. Turnitin's Similarity Report includes flags that alert instructors to hidden text and replaced characters . Getting caught converts an ambiguous AI-score conversation into a clear-cut deception charge, which most universities treat as the more serious offense.
Is a PDF safer than a Word document for AI detection? Neither is "safer." Both formats are reduced to the same extracted text before any analysis runs, so identical words produce identical results. Turnitin's only related guidance is that instructors should collect one consistent format across a class for processing consistency.
How many words does a PDF need for Turnitin's AI report to run? At least 300 words of continuous prose, per Turnitin's AI-report file requirements, with a 30,000-word maximum. Shorter documents get a Similarity Report (minimum 20 words) but no AI score. The detector supports English, Spanish, and Japanese.
Can my instructor tell if I converted my essay to PDF to avoid detection? Converting a normal essay to a normal PDF is completely unremarkable — millions of students do it, and it changes nothing about detection. What instructors notice is the failure modes of trickery: rejected files, empty reports, garbled extractions, and flag notices.
Does Turnitin's AI detection work differently on PDFs than the plagiarism check does? The extraction step is shared, but the requirements differ. The Similarity Report accepts more formats and needs only 20 words; the AI report accepts .docx, .pdf, .txt, and .rtf, needs about 300 words of long-form prose, and works in three languages. A readable PDF sails through both.
Key facts
- Turnitin accepts PDFs for both the Similarity Report and the AI writing report; AI-report formats are limited to .docx, .pdf, .txt, and .rtf (Turnitin file requirements documentation).
- The AI writing report requires at least ~300 words of long-form prose, caps at 30,000 words, and supports English, Spanish, and Japanese (Turnitin AI-report file requirements).
- General submission limits: under 100MB, under 800 pages, minimum 20 words; image-only/scanned PDFs are rejected outright (Turnitin file requirements).
- Turnitin's AI indicator launched April 4, 2023, and screened 200M+ papers in its first year; ~11% showed 20%+ AI writing (Turnitin first-anniversary release, April 2024).
- Turnitin claims 98% accuracy and <1% false positives, but only for documents where more than 20% of text is flagged; 1–19% scores display as an asterisk (Turnitin AI writing FAQ).
- Liang et al., Patterns, 2023: seven detectors falsely flagged an average of 61.22% of human-written TOEFL essays by non-native speakers; Turnitin was not among the tools tested (Cell Press).
- The Washington Post's April 2023 test found Turnitin weakest on blended human/AI drafts, which Turnitin acknowledged (Washington Post, Geoffrey Fowler).
Sources
- Turnitin — "File requirements for submitting your assignment to Turnitin" (Turnitin Guides help center)
- Turnitin — "File requirements for an AI writing report" (Turnitin Guides help center)
- Turnitin — AI writing detection FAQ / transparency page (accuracy claims, asterisk policy, launch date)
- Turnitin — first-anniversary AI detection data release, April 2024 (200M papers, prevalence figures)
- Liang, W. et al. — "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023
- Fowler, G. — Washington Post test of Turnitin's AI detector, April 2023
- Turnitin Guides — "Integrity Flags in the new enhanced Similarity Report" and "Flags in the Similarity Report"