Method

How the numbers were measured.

A single accuracy figure tells a buyer nothing about where it came from. Here are the conditions. Put the same questions to anyone else you are evaluating.

0. Three conditions, reported separately

A single accuracy figure hides the condition it was measured under. We split ours into three.

ConditionResultSample
Known generators, unedited
the condition the industry quotes
98.2 percent
95 percent interval 93.7 to 99.5
111 tracks
A generator we have never seen
what happens when a new tool launches
84.0 percent 150 tracks
Edited tracks
pitch, reverb, re-encoding
71.0 percent221 tracks
Human-made music1.5 percent wrongly flagged 282 tracks

The samples are not large. The known-generator figure rests on 111 tracks, roughly ten per generator, so we publish the interval alongside it. Collecting more is the next job, not a footnote we are hiding.

1. Measured on generators held out of training

Score on a generator you trained on and the number goes as high as you like. The moment that actually hurts is when a new tool launches, so we remove a generator from training entirely and measure on it. That figure is 84.0%.

2. The false-positive rate is fixed first

Detection rate rises as far as you want if you lower the threshold. So we fix the false-positive rate on human music at 1.5% and read detection at that point. Either number alone is meaningless.

3. Splits are by song, not by file

If edits of the same song land on both sides of the split, the score inflates. We group by original recording — and we verify the path resolution, not just the list, because we have had that leak silently.

4. Edited material is reported separately

We do not average unedited and edited together. Unedited 89.7%, edited 71.0%, stated apart. What you meet in practice is usually edited.

5. Reproducibility is checked

The same file must give the same answer. Stages with randomness run on fixed seeds, and when we move to different server hardware we confirm that no verdict flips before migrating.

Every figure here comes from our own evaluation data. We make no claim of validation by any outside body.

What to ask any vendor

This table does not compare companies. It separates claims from measurements, and you can put these questions to anyone — including us.

Ask thisThe usual answerOurs
Which generator was that accuracy measured on? A named generator and a high number 84.0 percent on generators held out of training
What is your false-positive rate?Not published 1.5 percent, at a stated operating point
How much edited material do you miss? “Robust to editing” We miss roughly 29.0 percent
Can you name which tool made it?A generator name Only when the file declares it. From audio alone we have not measured it, so we do not ship it
Same file twice — same answer? Yes; reproducibility verified with fixed seeds
Where does my audio end up? Deleted within 4 hours; we keep the hash and the verdict