Technology

How it decides

We look for traces in the audio that a listener will never hear. Here is what the system actually looks at.

Never just one detector

A single detector is easy to fool. So we run several that look at different things, then combine what they say: the texture of the sound, how the music develops over time, whether those properties stay consistent from one section to the next. They disagree with each other often — which is exactly why they are worth running separately. The combination beats every one of them alone.

Tags in the file are a hint, not the answer

Some generators leave their name in the file they produce. When that happens we mark the result grade A. The verdict does not lean on it, though: we compared the analysis of tagged and untagged tracks to check. Stripping the tag does not change the verdict.

Why a grade instead of a probability

“87 percent likely AI” looks precise but leaves you nothing to act on. Where is the line? What is the number based on? So we give a grade instead. Grade A is a fact you can re-verify yourself by opening the file. Grade B is a judgement from the sound alone, and it gets less certain the more a track has been edited. Those are different kinds of evidence, and collapsing them into one number hides that from you.

Speed and retention

2.65 seconds for one track; about 4.0 hours for a 12,480‑track catalog. Files are kept only while they are being analysed and deleted within 4 hours. One stage runs on an outside GPU service, so audio leaves our machines for that step — if your contract does not allow that, tell us and we will keep it in‑house instead, more slowly.

What we do not do

We do not analyse video. We do not do voice deepfakes. And we cannot tell you which tool made a track unless the file says so — working that out from the sound alone is something we have never measured, so we do not sell it.

Same length, same loudness — one of them is generated

generated
spectrogram
human
spectrogram
Waveform and spectrogram computed from two real recordings in our evaluation set. Standard transforms, no post-processing — the point is that you cannot see it by eye.

Generators in our training set

Tools keep launching that are not on this list. That is why the score on a generator we never trained on matters more than the list itself.

SunoUdioRiffusionStable AudioMusicGenAudioLDMSonautoElevenLabs MusicLyria

We do not publish the implementation. The more precisely we described it, the more it would read as instructions for getting around it. What we do publish in full is how the accuracy was measured — see Method and Known limits.