humanflow

Compilatio, reviewed

Compilatio says its detector is 94–99% reliable. Peer-reviewed testing put it under 78%. Unusually, the company publishes enough about its own method to explain the gap without anyone having to accuse anyone of anything — and the explanation is the most useful thing on this page.

Last reviewed 7 August 2026 · The HumanFlow team

What it is

A French academic-integrity company selling to institutions rather than to individuals. Its teacher-facing product, Magister+, runs plagiarism checking, AI detection and altered-text detection together and returns one combined report. There is a student-facing side too. If you are at a European university, and particularly a French-speaking one, this is a plausible candidate for the tool your work goes through.

What Compilatio publishes about itself

More than most, and it deserves credit for it. Its reliability note gives a 94–99% reliability rate on academic content and a false-positive rate below 1% — meaning that of 100 passages it calls human, fewer than one is really AI. The measurement covers roughly 7,400 texts across 24 languages, split evenly: 3,700 human, 3,700 AI.

A published sample size and an explicit language list are rare in this category. Most vendors give a percentage and nothing you could use to interrogate it.

The sentence that explains everything

In the same note, describing how the AI half of that sample was produced, Compilatio says the test used “simple questions without specific instructions on writing style”.

That is undisguised model output — the easiest case a detector will ever be handed. No paraphrasing, no humanizer, no instruction to write like a tired undergraduate at 2am. A 94–99% figure measured on that population is a real result about that population. It is not a result about a student who has made any attempt at all to disguise what they did.

Compilatio also states plainly that performance varies with “document length, language, and the nature of the analyzed content”. All three of those move in the wrong direction in real coursework.

What the independent testing found

Compilatio was one of the fourteen tools in Weber-Wulff et al. (2023), whose overall verdict was that the detectors were “neither accurate nor reliable”, with all fourteen scoring below 80% and only five above 70%.

Compilatio landed in that band: 66.67% to 77.78%, depending on which of the paper’s four scoring approaches is applied. It did well on human-written English text — consistent with its low false-positive claim — and degraded on machine-translated documents and on obfuscated content. The paper also records a straightforwardly broken case, where the tool returned “NaN% reliability” on ChatGPT-generated text containing code.

So the two figures are not in conflict once you see what each measured. Compilatio tested clean AI text and got 94–99%. Independent researchers tested clean text and text someone had tried to disguise, and got under 78%. The second population is the one that turns up in a misconduct hearing.

The 24 languages, which cut both ways

Most detectors in this category are English-only, and quietly so. Compilatio testing across Arabic, Hindi, Ukrainian, Greek, Slovenian and nineteen others is a real difference, and for a non-English institution it may be the deciding one.

It is also where the independent testing found the softest ground: machine-translated documents degraded its performance notably. Read those two facts together and the practical reading is that breadth of language coverage is not the same as equal accuracy across languages — which matters most for exactly the students who already carry the heaviest false-positive risk, as the non-native-speaker research sets out.

Where the company is careful

We spend a lot of this section criticising vendor claims, so it is worth recording where one gets it right. Compilatio’s own documentation says “no AI detector can be 100% reliable” and that “it is always up to the examiner to interpret this information to validate or impute potential fraud”. Its blog goes further, acknowledging that detection struggles with paraphrased AI text and with AI text corrected by a human, that some detectors work better in some languages than others, that results are “probabilities, not certainties”, and that “the final judgement should always rest with educators and students”.

That is a more careful public position than several better-known competitors hold. It also sits awkwardly beside a 94–99% headline, and the headline is the part that ends up in an email to a student.

Where we stand

We sell a detector and a humanizer, so Compilatio is a competitor and this page is interested testimony. We have published no measured pass rate against it and will not before our benchmark produces one under a method published in advance.

If a Compilatio score has been used against you, the specific and checkable point to raise is the one the company itself documents: the headline reliability figure was measured on AI text generated from simple prompts with no style instruction, and the peer-reviewed result on a harder sample is under 78%. What to do next and how to prove you wrote it apply the same as for any other tool.

Compared against

  • Compilatio vs Turnitin Compilatio against Turnitin on what independent testing found, who each is sold to, and what happens to the text you submit.

Related

Common questions

How accurate is Compilatio's AI detector?
It depends entirely on who ran the test and on what. Compilatio publishes a reliability rate of 94–99% on academic content with a false-positive rate below 1%, measured across roughly 7,400 texts in 24 languages. Peer-reviewed testing by Weber-Wulff et al. (2023) put it between 66.67% and 77.78%, depending on which of that paper's four scoring approaches you use. Both numbers are real. The reason they differ is on this page, and it is not that either party lied.
Why is there such a gap between Compilatio's figure and the independent one?
Because of what the AI text looked like. Compilatio's own methodology note says its test used 'simple questions without specific instructions on writing style' — that is, undisguised model output, the easiest case a detector ever meets. The independent study also tested obfuscated and machine-translated text, where Compilatio's performance fell away. A student who has made any attempt to disguise AI writing is not in the population Compilatio measured.
Does Compilatio detect AI in languages other than English?
It is one of the few that seriously tries. Its testing covers 24 languages including Arabic, Croatian, Czech, Danish, Dutch, Finnish, French, German, Greek, Hindi, Hungarian, Italian, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish and Ukrainian. That is a genuine differentiator in a category that is mostly English-only. It is also, in the independent testing, where the tool degraded — machine-translated documents were a notable weak point.
Is Compilatio used by universities?
Yes. It is a French company selling to institutions rather than individuals, and its Magister+ product bundles plagiarism checking, AI detection and altered-text detection into a single report for teachers. If you are in a European institution, particularly a French-speaking one, this is a realistic candidate for the tool checking your work.
Can Compilatio detect paraphrased or humanized AI text?
Less well than its headline figure implies, and its own blog says so. Compilatio acknowledges that detection struggles with 'paraphrased AI text' and 'AI-generated content corrected by humans'. In the peer-reviewed testing its accuracy on obfuscated content fell substantially. That is the general pattern across every detector we have examined, not a fault peculiar to this one.
Should a Compilatio score be used to accuse a student?
Compilatio's own answer is no, on its own. Its documentation states that 'no AI detector can be 100% reliable' and that 'it is always up to the examiner to interpret this information to validate or impute potential fraud', with its blog adding that 'the final judgement should always rest with educators and students'. That is a more careful position than several competitors take, and it is the vendor's, not ours.

Sources

  1. 1.How reliable is the Compilatio AI detector? Compilatio Helpdesk, 2026
  2. 2.Are AI detectors accurate? Power and limits of AI detection Compilatio (blog), 2026
  3. 3.Testing of detection tools for AI-generated text Weber-Wulff et al. — International Journal for Educational Integrity 19:26, 2023
  4. 4.Simple techniques to bypass GenAI text detectors: implications for inclusive education Perkins, Roe, Vu, Postma, Hickerson, McGaughran & Khuat — Int. J. of Educational Technology in Higher Education 21:53, 2024