Paste the same paragraph into two online detectors and you can get 12 percent from one and 60 percent from the other. That is not a malfunction. These are separately built systems, trained on different data, combining different signals, and applying cutoffs each company chose for itself, so identical text producing different numbers is the expected outcome rather than a surprising one. Neither number can be checked against a public reference, since there is no single benchmark against which all commercial detectors report comparable results. This article is not an accuracy ranking of the two products. It uses them as the example most people search for, and covers where the divergence comes from, what an accuracy claim does and does not tell you, and how to read a detection score without either trusting it or dismissing it.
Where the divergence comes from
Three stages sit between your text and the number on screen, and each one can move the result.
What the system learned. Detectors are trained systems. Vendors generally use different or undisclosed training corpora, and outsiders cannot determine how much those datasets overlap. The models included, the proportion of human to machine text, the subject areas, and the writing levels represented all shape what a tool treats as unusual.
What it looks at. Approaches vary across the category and include trained classifiers, token-level or sentence-level probability estimates, stylometric features, and combinations of several signals. Some early and simpler tools relied heavily on predictability measures. Most commercial products do not publish their feature set, so blanket statements about how detectors work, including confident ones, tend to outrun the evidence.
Where the line is drawn. A model produces a continuous output. Turning that into a percentage, or into a flag, requires a cutoff that somebody chose.
Two reasonable engineering teams making different calls at each stage will produce different numbers from the same paragraph without either being broken.
Training data shapes the answer
This is the stage that explains the most, and it is also where the category's known failure mode lives.
A tool performs best on distributions close to what it was trained on. Public evaluations tend to concentrate on a small number of widely used models, and vendors do not disclose their training corpora, so how well any product handles a given model family or language is largely unknown from outside. That includes systems that appear less often in benchmarks, such as DeepSeek, Qwen, Kimi, and Doubao. Unfox tests against those families and across 17 languages, which is a statement about coverage rather than a claim to be right more often.
The same mechanism produced the most serious documented problem in this field. Detector bias research published in Patterns found that widely used detectors flagged human-written essays by non-native English speakers at a far higher rate than a comparable native-speaker set. Writing that sat outside the pattern the tools had learned was misread. There are documented cases of human writing being flagged across the category, and the cost falls on individuals rather than vendors.
Two products, two companies, one letter apart
Before the mechanics, the practical point that trips up most people searching this comparison: ZeroGPT and GPTZero are separate products built by separate teams. The names differ by the position of one word.
That similarity shows up in the material written about them. Published summaries of their basic public facts contradict each other. On free-tier limits alone, third-party sources variously report a character cap per scan, a much smaller cap, and no cap at all. Reported accuracy figures for the same product range widely depending on who ran the test and what they tested with.
We are not going to add another set of numbers to that pile. What matters more is knowing where to look.
Current limits, pricing, and feature availability change often enough that each vendor's own pricing and documentation pages are the only reliable source. Check there rather than trusting a comparison article, including this one. And confirm which of the two products a review is actually discussing before you rely on it, because the similar names make these comparisons easy to misread.
Thresholds are a policy choice
Behind every percentage is a decision about which error to prefer.
Lower the cutoff and a tool catches more machine text while flagging more human writing. Raise it and human writing is safer while more machine text passes. There is no setting that escapes the trade. There are only settings that decide who bears the cost of being wrong.
Turnitin made its choice visible: it does not attribute a score between 1 and 19 percent, on the grounds that false positives concentrate there. That is a defensible policy, and it is a policy rather than a discovery. Turnitin's published reporting rules set out the rest.
Many free tools print a figure down to a single percentage point across the whole range. A number reported to that resolution can create an impression of precision the underlying estimate may not support, which matters when a result reads as 7 percent rather than as nothing at all.
How to read an accuracy claim
You will see figures of 99 percent and above across this industry, including on our own site. Here is what such a number does and does not tell you.
Many vendor accuracy claims are self-reported and measured on a test set the vendor assembled. Academic benchmarks for machine-generated text detection do exist, but there is no single public benchmark that commercial detectors are all required to report against, so vendor figures are frequently not comparable with one another.
The composition of a test set moves the result substantially. Which models generated the machine samples, the ratio of human to machine text, the length of each sample, and how much human editing was applied all change the outcome. A tool tuned for one mix can look excellent on that mix and ordinary on another.
A single accuracy figure also hides the distinction that matters most to an individual. Overall accuracy can be dominated by easy cases when a test set contains large amounts of unedited model output and clearly human prose. What you probably want to know is the false positive rate, meaning how often human writing gets flagged, and that number is reported far less often than headline accuracy is.
So when you meet an accuracy claim from any vendor in this space, the useful questions are what produced it, on what mix of text, and whether anyone outside the company can reproduce it.
Accuracy figures degrade outside the test set
Reported figures describe performance under the conditions of a specific test, and rewriting is the condition that most clearly falls outside them.
In a recursive paraphrasing study from Maryland and Harvard, researchers repeatedly paraphrased AI-generated text and measured the effect on several detection approaches. For one watermarking scheme, the true positive rate measured at a 1 percent false positive rate fell from 99.8 percent to 9.7 percent after five rounds of recursive paraphrasing. That result concerns a specific watermarking method under a specific metric and a specific attack, so it is not a statement about ZeroGPT, GPTZero, or detectors in general. What it does show is that a figure produced on clean model output cannot be assumed to hold for text that has been rewritten. What rewriting does to a detection result covers the practical side of that.
How to read a result you cannot verify
Look at the flagged passages, not the total. A percentage is an aggregate of many small judgments and gives you nothing to act on. A highlighted sentence is something you can read and decide about.
Treat clusters and scatter differently, while remembering that neither pattern establishes who wrote anything. A cluster is simply easier to inspect as a coherent passage than a few isolated flags across ten pages.
Get a second reading before drawing conclusions. Running a passage through another tool will not tell you which one is right, but it will tell you whether the signal survives a different model's assumptions. Criteria for choosing a checker covers what to compare. Unfox is an AI detector that reports at the sentence level, and the same instruction applies to our output as to anyone else's: read the sentences.
And do not treat a score from any tool as a finding about a person. It is the failure this whole category keeps repeating, and the vendor documentation says so more often than the marketing does.
FAQ
Why do ZeroGPT and GPTZero give different results?
They are separately built systems with different training data, different combinations of signals, and different cutoffs. Any of those alone would produce different numbers from identical text, and all three apply at once. Neither number can be checked against a public reference.
Is GPTZero more accurate than ZeroGPT?
Neither claim can be verified from outside. Vendor accuracy figures are usually self-reported on internally assembled test sets, and no single public benchmark covers commercial detectors in a comparable way. A more answerable question is which tool shows you the underlying passages.
What does a 99 percent accuracy claim actually mean?
Less than it appears. Ask which models were in the test set, how much human editing the samples had, how long they were, and whether the false positive rate is reported separately. A test set weighted toward easy cases produces a high overall figure that is still compatible with being wrong about a specific real document.
Can AI detectors be wrong about human writing?
Yes, and the pattern is documented. Research in Patterns found widely used detectors flagged more than 61 percent of TOEFL essays by non-native English writers as AI-generated, while classifying a native-speaker comparison set far more accurately. Careful, conventional human prose is the hardest case across this category.
Are ZeroGPT and GPTZero the same company?
No. They are separate products from separate teams, and the names differ only by word order. The similarity makes third-party comparisons easy to misread, so check which product a review is actually discussing before you rely on it.
Which detector should I trust?
None as a verdict, and several as a second reading. Look at flagged passages rather than totals, and check whether a signal holds up in a second tool. A single percentage is one model's estimate presented with more confidence than the underlying method supports.
Does paraphrasing lower detection scores?
Research shows repeated paraphrasing can substantially reduce detection performance for some methods, including a watermarking scheme whose true positive rate at a 1 percent false positive rate fell from 99.8 percent to 9.7 percent after five rounds. Those findings are method-specific and should not be read as a general rule about commercial detectors.

