humanflow

Research

One published review, one protocol with no results behind it yet, and one study we have not started. Each row below says which it is, because a research page that blurs the difference is not a research page.

Published

AI detector accuracy: the research, reviewed

Six studies on detector accuracy — what each tested, what it found, and the four things the literature does not establish. Every finding carries the source it came from.

Protocol published, results pending

The public benchmark protocol

The four conditions a benchmark has to meet before we would publish one: a released corpus, named detector versions with test dates, identical settings across tools, and raw per-sample results. The protocol is public; the results are not, because the test has not been run.

Not started

False-positive replication on ESL writing

The study we most want to exist: an independent replication of the non-native-speaker false-positive finding, on a corpus we publish, against named current detector versions. Nothing has been run and nothing is scheduled.

Why there is so little here

Because original research is slow and the alternative is worse. The obvious page for a company like ours is a benchmark showing our own tool performing well, and the reason we have not published one is set out in full on our methodology: a benchmark without a released corpus, named detector versions and raw results is a press release, and we would rather have an empty research section than that.

In the meantime the useful work is reviewing what other people have measured properly, and being precise about what those measurements do and do not support. That is what the literature review is for.

If you are citing any of it, how to cite this site has the formats — and leads with the advice to cite the primary source instead wherever we are only reporting somebody else's finding. That is not modesty. A summary can be wrong, ours has been, and the original is one click away on every page here.