humanflow

The AI detection glossary

What the words in a detection report actually mean, defined so each one stands on its own. Every entry carries a worked example with real text, because a definition you cannot see applied is a definition you cannot use. 50 terms so far.

These pages exist because the vocabulary is doing real damage. A student told their essay scored 87% is rarely told what was measured, and the number sounds like a probability of guilt when it is nothing of the kind. Most of the terms below describe how predictable a piece of writing is — which is a property of the prose, not a fact about its author.

We are adding terms as each one gets a real example rather than publishing a list of stubs. If a term you need is missing, tell us and it goes to the front of the queue.

A

  • Academic integrityThe standard of honest scholarly practice an institution sets and enforces.
  • Academic misconductThe umbrella term for breaches of an institution's academic integrity rules.
  • AI detectorA category name for tools that estimate whether text was machine-generated.
  • AI disclosure statementA written declaration of how AI tools were used in producing a piece of work.
  • AI slopLow-effort generated content published in volume with no one accountable for it.
  • AI writing scoreThe share of a document a detector flagged as likely machine-written.
  • AppealA formal challenge to a decision, heard on stated grounds rather than reargued.

B

  • BurstinessHow much sentence length and complexity vary across a passage.

C

  • C2PAAn open standard attaching signed provenance data to digital content.
  • Confidence intervalThe range a measured rate could plausibly be, given how little was tested.
  • Contract cheatingPaying a third party to produce work you submit as your own.

D

  • Detection thresholdThe score above which a tool reports a document as AI-written.
  • Document-level detectionReporting one score for a whole document rather than for its parts.
  • Due processThe procedural protections you are owed before a decision is made against you.

F

G

  • Generative AISystems that produce new content rather than classify existing content.
  • GPTGenerative pre-trained transformer — the architecture behind most text models.

H

  • HedgingLanguage that softens a claim to match the strength of the evidence behind it.
  • Honour codeAn institution-wide commitment to academic honesty, often signed and student-enforced.
  • Humanized text detectionDetection aimed specifically at text that has been through a rewriting tool.

I

  • Invisible characterUnicode characters that occupy no visible space but are present in the text.

L

  • Large language modelA model trained to predict the next piece of text, used to generate prose.
  • Lexical diversityHow varied a text's vocabulary is, and why the usual measure of it cannot be compared.

N

  • N-gramA run of n consecutive words, used to compare texts for overlap.
  • NominalizationA verb turned into a noun, usually needing an empty verb to prop it up.

O

P

  • Passive voiceA construction where the thing acted on becomes the subject of the sentence.
  • PerplexityA measure of how surprised a language model is by a piece of text.
  • PlagiarismUsing someone else's work or ideas without crediting them.
  • Precision and recallTwo separate measures that a single accuracy figure averages away.
  • ProctoringSupervision of an assessment to confirm the candidate is working unaided.
  • PromptThe instruction and context given to a model to produce a response.

R

  • ReadabilityHow much effort a text takes to read, estimated from sentence and word length.
  • RegisterThe level of formality a text adopts for its audience and situation.
  • Retrieval-augmented generationFetching real source material and giving it to a model before it answers.

S

  • Sentence-level detectionReporting a score per sentence rather than one figure for the document.
  • Similarity scoreThe share of a document that matches text in a database of sources.
  • StylometryIdentifying an author from measurable habits in how they write.
  • Syllabus AI policyThe course-level statement setting what AI use is permitted in that module.
  • SynthIDGoogle DeepMind's watermarking system for AI-generated content.

T

  • TemperatureA setting controlling how much randomness a model uses when choosing each word.
  • TokenThe unit of text a language model actually reads — usually a word fragment.
  • Transition wordA word or phrase signalling how one sentence relates to the one before it.

V

  • Viva voceAn oral examination used to establish whether you can account for your own work.

W

  • WatermarkingHiding a detectable signal inside AI output at the moment it is generated.

Where to go from here

If you are here because something you wrote was flagged, the glossary is the wrong place to start — read what to do when you are falsely accused first, then come back for the vocabulary you need to argue with. If you are here to understand the tools, start with how detectors work and how accurate they are.