200 human-written source documents across five categories, 40 per cell. Every document is written by a person, with authorship attested, and none has appeared on the public web — a corpus a detector may have trained on measures the wrong thing.
The fifth cell exists because of the published false-positive research. If detectors misclassify non-native English writing at the rates Liang and colleagues reported, then a benchmark drawn only from native-speaker prose will overstate how well everything works.