This site tells you our rewrites pass AI detectors 85-95% of the time on text up to 3,000 characters. A number like that only means something if the test behind it is visible. This page is that test: what we run against, what counts as a pass, how often we re-run, and how to reproduce the whole thing yourself.

What we test against

Four detectors, chosen because they are the ones schools and publishers actually run: GPTZero, Turnitin AI, Originality.ai, and Winston AI. Detector products ship updates constantly, so every run logs the detector version and the date. A pass rate from six months ago is not evidence about today's detector — that is why we re-run monthly instead of publishing one number forever.

Sample set

Each run uses at least 150 AI-generated samples, split across five categories that map to how people actually use the tool:

All samples run 1,500-3,000 characters — the range the free tool processes in one pass — generated with ChatGPT, Gemini, and Claude using ordinary prompts, no adversarial tricks, raw output only.

Pass criteria

Two stages, deliberately conservative:

  1. Sanity check — the raw AI sample must be flagged as AI by the detector. If the baseline is missed, the sample is discarded. We never count an easy baseline as a win.
  2. Scoring — the humanized rewrite goes through the same detector. A pass is the detector's own interface reporting the text as not AI-generated or below its threshold. "Mixed" and "uncertain" verdicts count as failures.

Pass rate = passes ÷ (samples × detectors). Borderline texts are never re-scored until they pass.

Monthly regression schedule

One full run every month, plus an unscheduled run after any major detector release. Results go straight into the log below:

RunDateDetectorsSamplesPass rate
First full runNovember 2026GPTZero, Turnitin, Originality.ai, Winston≥150Published with the run

The 85-95% figure comes from our internal calibration while building the engine. From November 2026 onward, every month's run is published here under this protocol — including runs where the number moves down.

Reproduce it yourself

  1. Generate 150 short-to-medium texts across the five categories with any mainstream chatbot.
  2. Score every raw sample through each detector; drop any that are not flagged as AI.
  3. Rewrite each sample with our AI humanizer — or your own method. The protocol does not care which tool produced the rewrite.
  4. Score every rewrite on the same detectors and record verdicts exactly as the interface shows them.
  5. Divide passes by total scores, then compare with our log.

Limitations, stated plainly

Related: free AI detector — check any rewrite before you submit · GPTZero bypass and Turnitin bypass — per-detector rewrite tuning