What it is
A French academic-integrity company selling to institutions rather than to individuals. Its teacher-facing product, Magister+, runs plagiarism checking, AI detection and altered-text detection together and returns one combined report. There is a student-facing side too. If you are at a European university, and particularly a French-speaking one, this is a plausible candidate for the tool your work goes through.
What Compilatio publishes about itself
More than most, and it deserves credit for it. Its reliability note gives a 94–99% reliability rate on academic content and a false-positive rate below 1% — meaning that of 100 passages it calls human, fewer than one is really AI. The measurement covers roughly 7,400 texts across 24 languages, split evenly: 3,700 human, 3,700 AI.
A published sample size and an explicit language list are rare in this category. Most vendors give a percentage and nothing you could use to interrogate it.
The sentence that explains everything
In the same note, describing how the AI half of that sample was produced, Compilatio says the test used “simple questions without specific instructions on writing style”.
That is undisguised model output — the easiest case a detector will ever be handed. No paraphrasing, no humanizer, no instruction to write like a tired undergraduate at 2am. A 94–99% figure measured on that population is a real result about that population. It is not a result about a student who has made any attempt at all to disguise what they did.
Compilatio also states plainly that performance varies with “document length, language, and the nature of the analyzed content”. All three of those move in the wrong direction in real coursework.
What the independent testing found
Compilatio was one of the fourteen tools in Weber-Wulff et al. (2023), whose overall verdict was that the detectors were “neither accurate nor reliable”, with all fourteen scoring below 80% and only five above 70%.
Compilatio landed in that band: 66.67% to 77.78%, depending on which of the paper’s four scoring approaches is applied. It did well on human-written English text — consistent with its low false-positive claim — and degraded on machine-translated documents and on obfuscated content. The paper also records a straightforwardly broken case, where the tool returned “NaN% reliability” on ChatGPT-generated text containing code.
So the two figures are not in conflict once you see what each measured. Compilatio tested clean AI text and got 94–99%. Independent researchers tested clean text and text someone had tried to disguise, and got under 78%. The second population is the one that turns up in a misconduct hearing.
The 24 languages, which cut both ways
Most detectors in this category are English-only, and quietly so. Compilatio testing across Arabic, Hindi, Ukrainian, Greek, Slovenian and nineteen others is a real difference, and for a non-English institution it may be the deciding one.
It is also where the independent testing found the softest ground: machine-translated documents degraded its performance notably. Read those two facts together and the practical reading is that breadth of language coverage is not the same as equal accuracy across languages — which matters most for exactly the students who already carry the heaviest false-positive risk, as the non-native-speaker research sets out.
Where the company is careful
We spend a lot of this section criticising vendor claims, so it is worth recording where one gets it right. Compilatio’s own documentation says “no AI detector can be 100% reliable” and that “it is always up to the examiner to interpret this information to validate or impute potential fraud”. Its blog goes further, acknowledging that detection struggles with paraphrased AI text and with AI text corrected by a human, that some detectors work better in some languages than others, that results are “probabilities, not certainties”, and that “the final judgement should always rest with educators and students”.
That is a more careful public position than several better-known competitors hold. It also sits awkwardly beside a 94–99% headline, and the headline is the part that ends up in an email to a student.
Where we stand
We sell a detector and a humanizer, so Compilatio is a competitor and this page is interested testimony. We have published no measured pass rate against it and will not before our benchmark produces one under a method published in advance.
If a Compilatio score has been used against you, the specific and checkable point to raise is the one the company itself documents: the headline reliability figure was measured on AI text generated from simple prompts with no style instruction, and the peer-reviewed result on a harder sample is under 78%. What to do next and how to prove you wrote it apply the same as for any other tool.
Compared against
- Compilatio vs Turnitin — Compilatio against Turnitin on what independent testing found, who each is sold to, and what happens to the text you submit.
Related
- SafeAssign, reviewed — the institutional tool that does not detect AI at all
- Copyleaks, reviewed
- Are AI detectors accurate?
- How detectors work