One question per page, answered in the first sentence, with the source underneath it and an honest note on when the answer stops holding. 22 so far.
These are the questions that did not belong anywhere else. Before a question gets a page here we check that no page on this site already answers it — a sweep of the FAQs across the site found a dozen planned questions that were already answered on the hub that owns the subject, and those were dropped rather than built. Two URLs competing to answer one question serve nobody.
If the question you have is about a specific detector, the AI detection hub is the better starting point. If it is about a term you have seen in a report, try the glossary.
No. A low perplexity score establishes that your writing was predictable, which is a property of the prose rather than a fact about its author.
No, and the clearest statement of that comes from the vendors: Turnitin states its AI indicator is not intended as the sole basis for an academic misconduct finding.
Whatever your institution's policy says it is — there is no shared definition across institutions, and the tools people assume are exempt are often the ones a policy names explicitly.
Cite the company that made the model as the author, not the model itself — APA's template is: OpenAI. (2023). ChatGPT (Mar 14 version) [Large language model]. https://chat.openai.com/chat
Increasingly yes — and the honest complication is that the answer changed without most users noticing, because the product added generative features to a tool they had used for years as a spellchecker.
Unevenly — one detector performed very well on medium-to-long passages in the most recent systematic comparison, while others were placed in a tier the authors called unsuitable for short text and susceptible to humanizers.
They can provide positive evidence that a participating model or application produced something — and neither can establish the opposite, because the absence of a watermark or a credential is the normal state of almost all text.
Vanderbilt disabled it in 2023 and published its reasoning; Yale, Georgetown, the University of Pittsburgh, Johns Hopkins, the University of Alabama and Curtin have since turned it off, and Washington State University cancelled its contract outright in February 2026.
No. A similarity score counts how much of your text matches documents in a database, while an AI score estimates how machine-like your writing reads — the first points at something you can open and read, the second does not.
No. Plagiarism is presenting someone else's ideas or work as your own, so rewording a passage without citing it changes the wording and leaves the offence exactly where it was.
Frequently not. Research presented at NeurIPS 2023 showed that running generated text through a paraphraser drove several detectors' accuracy down sharply, though the same paper proposed a retrieval-based defence that continued to work.
Because detectors measure how predictable text is, and writing carefully in a second language produces predictable text — in peer-reviewed testing, seven detectors flagged 61.22% of TOEFL essays by non-native speakers as AI-generated while classifying US 8th-grade essays almost perfectly.
There is no safe percentage, because the number is not a measurement of wrongdoing that a threshold could clear you of — no published figure separates acceptable from unacceptable, and any number circulating as one was invented by somebody.
Yes, substantially. These tools work from statistical patterns across a document, so short passages give them very little to work with — Turnitin requires a minimum of 300 words of prose before it will produce an AI writing report at all.
They can, and for a reason that has nothing to do with you: quoted passages and reference lists are highly predictable text, and predictability is what these tools measure.
RAID is a shared public benchmark for evaluating machine-generated text detectors, introduced in a paper at ACL 2024, built specifically to test how detectors hold up under adversarial conditions rather than how they score on clean text.
It depends entirely on what you are asking them to do — rewriting text genuinely changes its rhythm and vocabulary, and no tool, including this one, can promise you a particular detector outcome.
They are two unrelated products with almost the same name — GPTZero is the detector built by Edward Tian and marketed to institutions, while ZeroGPT is a separate free tool from a different operator, and the pair are confused constantly.
Less reliably than on English, and the vendors themselves say so — language support varies by product, independent testing outside English is much thinner, and the failure mode that flags careful writing gets worse rather than better.
Almost every case in this area is heard by an institutional panel rather than a court, and those panels are not bound by rules of evidence — but detector vendors themselves state that a score is not proof of misconduct, which is the more useful fact to have in hand.
Nobody outside the model developers knows, and none of them has published an explanation — the observation that generated text uses em dashes heavily is widely reported and the cause is not documented anywhere we can point you to.
Name behaviours rather than technologies, state explicitly what a detection score may and may not be used for, and put the operative rule at course level where the assignment is actually set — those three decisions determine whether a policy is usable.