What is the RAID benchmark?
RAID is a shared public benchmark for evaluating machine-generated text detectors, introduced in a paper at ACL 2024, built specifically to test how detectors hold up under adversarial conditions rather than how they score on clean text.
Last reviewed 15 August 2026 · The HumanFlow team
It matters because detector vendors quote it. Grammarly and QuillBot both cite RAID results for their detectors, so a reader comparing tools will meet the name without any of those pages explaining what it is.
The design point is robustness. Reporting accuracy on unmodified generated text is easy and flatters everyone; RAID evaluates across many generators, domains and deliberate adversarial modifications, which is where detectors that look strong in a vendor's own test tend to come apart.
Being a shared benchmark with a public leaderboard is the substantive difference from a vendor's internal figure. Anyone can submit, the conditions are the same for every entrant, and the results sit somewhere you can go and read rather than in a marketing page.
It is still a benchmark rather than a guarantee. A leaderboard position is a measurement on that dataset under those conditions, and it does not tell you the false positive rate on the kind of writing you personally produce — which remains the number nobody publishes.
When this answer changes
Leaderboards move. A claim of first place is a claim about a date, and any vendor quoting one without a capture date is quoting a position that may no longer hold.
A benchmark can only test what it contains. As new models appear, results measured against an earlier generation describe an earlier problem.
Where to go next
- How accurate detectors are — The published figures and what each test actually measured.
- Precision and recall — The two numbers a single accuracy figure hides, and why benchmarks report both.
- Confidence intervals — Why a headline rate from a small test can be much less certain than it sounds.
Sources
One of our direct answers.