Method
How the numbers were measured.
A single accuracy figure tells a buyer nothing about where it came from.
Here are the conditions. Put the same questions to anyone else you are evaluating.
0. Three conditions, reported separately
A single accuracy figure hides the condition it was measured under. We split ours into
three.
| Condition | Result | Sample |
Known generators, unedited the condition the industry
quotes |
98.2 percent 95 percent interval
93.7 to 99.5 | 111 tracks |
A generator we have never seen what happens when a new tool
launches | 84.0 percent |
150 tracks |
Edited tracks pitch, reverb, re-encoding |
71.0 percent | 221 tracks |
| Human-made music | 1.5 percent wrongly flagged |
282 tracks |
The samples are not large. The known-generator figure rests on
111 tracks, roughly ten per generator, so we publish the interval alongside it.
Collecting more is the next job, not a footnote we are hiding.
1. Measured on generators held out of training
Score on a generator you trained on and the number goes as high as you like. The moment that
actually hurts is when a new tool launches, so we remove a generator from training
entirely and measure on it. That figure is 84.0%.
2. The false-positive rate is fixed first
Detection rate rises as far as you want if you lower the threshold. So we fix the
false-positive rate on human music at 1.5% and read detection at that point.
Either number alone is meaningless.
3. Splits are by song, not by file
If edits of the same song land on both sides of the split, the score inflates. We group by
original recording — and we verify the path resolution, not just the list, because we
have had that leak silently.
4. Edited material is reported separately
We do not average unedited and edited together. Unedited 89.7%, edited
71.0%, stated apart. What you meet in practice is usually edited.
5. Reproducibility is checked
The same file must give the same answer. Stages with randomness run on fixed seeds, and when
we move to different server hardware we confirm that no verdict flips before migrating.
Every figure here comes from our own evaluation
data. We make no claim of validation by any outside body.
What to ask any vendor
This table does not compare companies. It separates claims from measurements, and you
can put these questions to anyone — including us.
| Ask this | The usual answer | Ours |
| Which generator was that accuracy measured on? |
A named generator and a high number |
84.0 percent on generators held out of training |
| What is your false-positive rate? | Not published |
1.5 percent, at a stated operating point |
| How much edited material do you miss? |
“Robust to editing” |
We miss roughly 29.0 percent |
| Can you name which tool made it? | A generator name |
Only when the file declares it. From audio alone we have not measured it, so we do not
ship it |
| Same file twice — same answer? | — |
Yes; reproducibility verified with fixed seeds |
| Where does my audio end up? | — |
Deleted within 4 hours; we keep the hash and the verdict |