Detectors misfire on non-native English writers
Liang et al., published in Patterns (2023), ran 91 TOEFL essays written by real people through seven detectors. They flagged an average of 61.22% as AI-generated; 89 of the 91 were flagged by at least one tool. The same detectors read US 8th-grade essays almost perfectly.