HumanFlow is not one of the two tools on this page. We sell a rewriter and run a detector of our own, so weigh this accordingly — our methodology sets out how we source what appears below and what we refuse to claim.
Compilatio's AI testing covers 24 languages; Turnitin's indicator covers three. In the one peer-reviewed study covering both, Turnitin scored higher — and machine-translated text was where Compilatio fell away.
Last reviewed 16 August 2026 · The HumanFlow team
Why people compare these two
This is the pair that decides itself by geography more than by merit. Compilatio is a French company selling to European institutions, particularly French-speaking ones; Turnitin is the global incumbent. Most people encountering this comparison are not choosing between them at all — they are working out which one their institution runs and what that means for their work.
It is also the most useful pair on this site for anyone writing in a language other than English, because it is the only one where multilingual coverage is a genuine axis of difference rather than a footnote.
Most readers of this page are not choosing between these two, which shapes what is useful to say. Institutional detection is a procurement decision made once, often years ago, and the person affected by it had no part in that. So the page is written to be read in both directions: by someone weighing a purchase, and by someone working out what the report on their desk actually establishes.
Side by side
No prices. Neither of these is a tool most readers choose on cost, and for one of them there is no public price at all — detector pricing, where it could be captured, is on our detector pricing page.
Dimension
Compilatio
Turnitin
Who it is sold to
Institutions. A French company selling to universities rather than to individuals; its Magister+ product bundles plagiarism checking, AI detection and altered-text detection into one teacher-facing report.
Institutions only. A student or an individual instructor cannot buy it; it arrives inside the systems a university already licenses.
How you get at it
No public self-serve checker of the kind GPTZero or Scribbr offer.
No public checker. If your institution has not enabled the indicator for your class, there is no way to see your own score.
What it needs
Its testing covers 24 languages, including Arabic, Croatian, Czech, Danish, Dutch, Finnish, French, German, Greek, Hindi, Hungarian, Italian, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish and Ukrainian.
At least 300 words of prose, up to 30,000. .docx, .pdf, .txt or .rtf under 100MB, in English, Spanish or Japanese.
What independent testing found
Between 66.67% and 77.78% in Weber-Wulff et al. (2023), depending on which of that paper's four scoring approaches you use. Machine-translated documents were a notable weak point.
Scored highest of the fourteen tools in Weber-Wulff et al. (2023) — in a study whose own conclusion was that detection tools “are neither accurate nor reliable”, with every tool below 80% accuracy.
What the vendor claims
A reliability rate of 94–99% on academic content with a false-positive rate below 1%, measured across roughly 7,400 texts in 24 languages.
Aims to keep false positives under 1% above a 20% detected share. Below 20% it attributes no score at all, showing an asterisk, because its own testing found more false positives in that band.
What the vendor says about its limits
Its documentation states that no AI detector can be 100% reliable and that it is always up to the examiner to interpret the information; its blog adds that the final judgement should always rest with educators and students.
States its indicator is not intended as the sole basis for an academic misconduct finding.
What happens to your text
Not established from its published pages by us.
Submissions may be retained in its repository depending on the institution's configuration, which the student does not set.
Every line above is summarised from our own examination of each tool, where the studies and vendor documents behind it are quoted and linked: Compilatio and Turnitin.
What actually separates them
Language coverage is the real distinction and it is not close. Compilatio's AI testing spans 24 languages — Arabic, Croatian, Czech, Danish, Dutch, Finnish, French, German, Greek, Hindi, Hungarian, Italian, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish and Ukrainian among them. Turnitin's AI indicator covers English, Spanish and Japanese. In a category that is overwhelmingly English-only, that is a serious differentiator.
It is also, awkwardly, where the independent testing found Compilatio weakest. Machine-translated documents were a notable weak point in Weber-Wulff et al. (2023) — so the breadth of language support and the reliability of the verdict in those languages are two different claims, and only the first is established.
Both are institutional products, which shapes what you can do about a result. Neither offers the kind of public self-serve checker GPTZero or Scribbr do. Turnitin requires 300 to 30,000 words of prose in .docx, .pdf, .txt or .rtf under 100MB. Compilatio's Magister+ bundles plagiarism checking, AI detection and altered-text detection into a single teacher-facing report, which means an instructor sees one document covering three different kinds of claim — convenient to read and easy to conflate.
On how they handle uncertainty, both are more careful than the category average and they express it differently. Turnitin builds it into the product: no score at all between 0% and 20%, an asterisk instead, because its own testing found more false positives there. Compilatio states it in words — its documentation says no AI detector can be 100% reliable and that it is always up to the examiner to interpret the information, with its blog adding that final judgement should rest with educators and students.
There is a structural point underneath the language comparison that applies well beyond these two. Detectors are trained predominantly on English text, and a model's sense of what is statistically ordinary is a sense of what is ordinary in the languages it saw most. Supporting a language and performing well in it are therefore separate claims, and Compilatio's own independent result — weakest on machine-translated documents — is exactly the shape that distinction predicts.
What the evidence supports, and what it does not
Weber-Wulff et al. (2023) is the rare study that tested both. Turnitin scored highest of the fourteen tools examined; Compilatio landed between 66.67% and 77.78%, depending on which of that paper's four scoring approaches you use. The study's own conclusion was that detection tools “are neither accurate nor reliable”, with every tool below 80% accuracy — so Turnitin's win is a win inside a set the authors judged unfit for the purpose.
Compilatio's own figure is far higher: 94–99% reliability on academic content with a false-positive rate below 1%, across roughly 7,400 texts in 24 languages. The gap between that and the independent result is explained by its own methodology note, which says the test used “simple questions without specific instructions on writing style”. That is undisguised model output — the easiest case a detector ever meets. The independent study also tested obfuscated and machine-translated text, where performance fell away. A student who has made any attempt to disguise AI writing is not in the population Compilatio measured.
Neither party lied, and that is the point worth taking from this page. Two honest measurements of different things produce numbers thirty points apart, and the conditions attached to a figure matter more than the figure.
The gap between the two Compilatio figures is the most instructive thing on this page and it generalises to every vendor number in this category. A detector tested on undisguised model output measures the easiest case it will ever meet; a detector tested on obfuscated and translated text measures something closer to the population it is actually deployed against. Neither party lied, and a thirty-point spread came out of the difference in what was tested.
Where each one came from
Compilatio is a French company selling to European institutions, and almost everything distinctive about it follows from that. It grew up in a market where academic integrity software has to work across languages as a baseline rather than as an expansion, where the customer is frequently a public university with procurement rules, and where the dominant vendor was American. Testing in 24 languages is not a feature decision so much as a condition of competing at all.
Turnitin's history runs the other way. It became the default in English-speaking higher education first, and its AI indicator covers English, Spanish and Japanese — three languages, from a company with the resources to support thirty. That is a statement about where the customers are rather than about what is technically possible, and it is the single most consequential difference between these two for anyone outside the anglophone world.
The two also package differently in a way that reflects who reads the output. Compilatio's Magister+ bundles plagiarism checking, AI detection and altered-text detection into one teacher-facing report — convenient, and a genuine risk, because three different kinds of claim arrive in one document and are easy to read as one verdict. Turnitin keeps the similarity score and the AI indicator visually separate, which is the safer arrangement and still gets conflated constantly.
If one of these has produced a result about you
Neither of these is something you can run on your own work, so the first practical step with both is finding out what your institution's policy actually says. Both are institutional products with no public self-serve checker, and any third-party site offering to show you your Turnitin or Compilatio score is showing you a different tool's opinion of your text.
If you wrote in a language other than English, that belongs in the conversation early. Compilatio's own independent testing showed machine-translated documents as a notable weak point, and Turnitin's indicator does not cover most languages at all — so a result on a translated or non-English document is weaker evidence than the same result on English prose, and the vendors' own material supports saying so.
Both vendors have published statements you can quote, and quoting the vendor is more effective than quoting us. Turnitin states its indicator is not intended as the sole basis for an academic misconduct finding. Compilatio's documentation states that no AI detector can be 100% reliable and that it is always up to the examiner to interpret the information, with its blog adding that final judgement should rest with educators and students.
If a Compilatio report is in front of you, check which part of it the accusation rests on. A Magister+ report can carry a similarity score, an AI indication and an altered-text flag together, and they are three different claims with three different evidential weights.
What we could not establish
Weber-Wulff et al. (2023) tested both, which makes this the rare pair with a shared measurement — and the result has a wide band on one side. Compilatio scored between 66.67% and 77.78% depending on which of that paper's four scoring approaches is used, and we do not have a basis for preferring one of those approaches over another.
Compilatio's own figure of 94–99% reliability comes with a methodology note saying the test used simple questions without specific instructions on writing style. That tells us the AI text was undisguised. It does not tell us the corpus size per language, the human baseline, or how the roughly 7,400 texts were distributed across 24 languages — so the multilingual claim is broad rather than deep as far as we can verify it.
Turnitin publishes no per-language accuracy figures for Spanish or Japanese, and no independent study we hold has tested it in either. Its three-language support is a documented capability and not a documented performance.
Neither vendor's pricing is public. Turnitin returns 403 to automated requests and Compilatio sells through institutional quotes, so no figure for either appears on this site.
Which to pick
Pick Compilatio if you are a European institution, particularly a francophone one, or you assess in languages Turnitin's indicator does not cover. Bundled altered-text detection and a 24-language test base have no equivalent in Turnitin's offering.
Pick Turnitin if you assess mainly in English, Spanish or Japanese and want the tool with the widest peer-reviewed testing behind it — plus a suppression band that refuses to show a number the vendor does not trust.
What neither score proves
A detector reports how statistically machine-typical a piece of prose reads. It has no access to how the text was produced, so it cannot establish authorship in either direction — a flag is not evidence of AI use, and a clean result is not a clearance. If you have been accused on the strength of one, drafting history is what actually answers it. And if your institution requires you to disclose AI assistance, disclose it — nothing on this page changes that obligation.
Compilatio vs Turnitin: common questions
Is Compilatio as accurate as Turnitin?
In the one peer-reviewed study covering both, no. Weber-Wulff et al. (2023) placed Turnitin highest of fourteen tools and Compilatio between 66.67% and 77.78% depending on the scoring approach used. Compilatio's own testing reports 94–99% reliability, measured on undisguised model output. Both numbers are real measurements of different conditions.
Which languages does each one cover for AI detection?
Compilatio's testing covers 24 languages, including Arabic, Czech, Dutch, Finnish, French, German, Greek, Hindi, Hungarian, Italian, Polish, Portuguese, Romanian, Russian, Spanish, Swedish and Ukrainian. Turnitin's AI indicator covers English, Spanish and Japanese. That is the widest genuine gap between the two.
Why is there such a gap between Compilatio's figure and the independent one?
Because of what the AI text looked like. Compilatio's methodology note says its test used simple questions with no instructions on writing style — undisguised model output, the easiest case for any detector. The independent study also tested obfuscated and machine-translated text, where Compilatio's performance fell away. The two studies measured different populations of text.
Can Compilatio detect paraphrased or humanized AI text?
Less well than its headline figure implies, and it says so itself. Compilatio acknowledges that detection struggles with paraphrased AI text and with AI-generated content corrected by humans. In the peer-reviewed testing its accuracy on obfuscated content fell substantially. That is the pattern across every detector examined rather than a fault peculiar to this one — Weber-Wulff and colleagues measured 26% accuracy across all fourteen tools on machine-paraphrased text.
Can I check my own work in either one?
Not directly. Both sell to institutions rather than to individuals, and neither offers the kind of open self-serve checker consumer tools do. Any third-party site claiming to show you your Turnitin or Compilatio score is showing you a different tool's opinion of your text.
Should either score be used to accuse a student?
Both vendors say no, on their own. Turnitin states its indicator is not intended as the sole basis for an academic misconduct finding. Compilatio's documentation states that no AI detector can be 100% reliable and that it is always up to the examiner to interpret the information, with its blog adding that the final judgement should rest with educators and students.