I think of this one as the project that failed, and the write-up below will show that this is not really true. But it is how I remember it, because the first version was confidently wrong, and confidently wrong is the worst kind.
Phase one, December 2025
Twelve tracks, six AI and six human. MFCCs with mean and standard deviation over twenty coefficients, spectral centroid, bandwidth, rolloff, zero crossing rate, all computed over ten second clips. Logistic regression on a 75/25 split with a StandardScaler. A notebook of distributions and a two dimensional feature space that looked pleasingly separable.
Three things were wrong, and I want to name them in order of how much they should have embarrassed me.
- The split leaked. It was random over clips. With roughly fifteen clips per song, near-identical windows of the same track landed on both sides of the split, so the reported accuracy was measuring memorisation of tracks, not discrimination between classes. This is the canonical grouped-split mistake and I made it anyway.
- The dataset was far too small, and one of the six human files turned out to be a byte-identical duplicate of another. Five human tracks. A model can learn five of anything.
- The environment could not run. The virtualenv was a Windows one, with a
Scripts/directory and.exeentry points, so nothing reproduced on the machine I now use.
None of this is exotic. That is the point. The failure modes that get you are the ones you already know about, applied to your own work when you are excited.
The music side of this is not incidental for me. I spend a lot of time inside music and I have opinions about it, and I wanted to know whether the thing I could sometimes hear, the slightly too clean, slightly too centred quality of generated tracks, was a real acoustic signature or a story I was telling myself. A detector that generalises is a way of asking that question without my own taste in the loop.
The rebuild, August 2026
Some context on where my head was. This was one of the first machine learning projects I had picked up in a while, after spending a couple of years with my brain switched over from engineering to design for the masters. I was lucky with the timing: I was in a multimodal ML class when I came back to this, and it let me pick things up faster than I would have alone. Even so, I had to relearn more than I expected, go deeper into audio classification than I had ever gone, and, for the first time, learn what it actually takes to run a model on a phone rather than a laptop: quantisation, export, the front end, all the parts of on-device ML that are still genuinely hard and that nobody's tutorial covers past the toy example.
Data that actually has both classes as audio
Surveying Hugging Face for corpora was its own lesson. SONICS has 32 GB of Suno
and Udio fakes and no real audio at all, only YouTube metadata for the human side.
FakeMusicCaps is half a megabyte of metadata. The one that ships both classes
as waveforms is SleepyJesse/ai_music_large: 20,139 tracks, each carrying an
ai_generated flag and a source string naming the generator, with
the human side drawn from the Free Music Archive.
Seventy three gigabytes is more than I wanted on disk, so ingest streams it: download one parquet shard, decode, slice into ten second clips, compute log-mels, append to a memmapped cache, delete the shard. The whole corpus lands as a 17 GB cache inside about a gigabyte of working space, checkpointed per shard and resumable. One detail I am fond of: the corpus is stored ordered by class, so a front to back ingest sees nothing but human tracks until it is half done. Shards are visited interleaved, first, last, second, second-last, so both classes are present from the first minutes and stopping early still leaves something balanced.
One front end, shared by construction
Every clip is a 128 by 431 log-mel spectrogram, 22,050 Hz, n_fft 2048, hop 512, stored as float16 at 108 KB per clip. Mels are standardised per clip, which removes absolute loudness as a feature, and that matters because AI tracks are typically normalised in a way that human uploads are not.
The mel transform is one torchaudio module, MelFrontend, and the
same module object is used by the cache builder, the training loop and the exported Core ML
graph. The features the phone computes are the features the model trained on by
construction, rather than by a careful reimplementation in Swift that drifts the first time
someone touches it.
Splits that cannot lie
Grouped by track_uid, stratified by label and source. A holdout
split withholds an entire generator, Suno, so there is a number for "a synthesiser the
model has never heard." A probe split holds eleven tracks of my own, cached
separately, because they are the closest thing I have to the deployment distribution.
And I kept the old baseline script, fitting the same model under grouped versus random
splitting, purely to quantify how inflated the phase one number had been.
The shortcut I was worried about
The human class is entirely FMA MP3s and the AI class comes from generators, so the two
classes carry different codec bandwidths. A convolutional net will happily learn the
spectral cutoff instead of anything musical, and the TISMIR literature on this task found
exactly that in their classifiers. I mitigated it three ways: a bandwidth_jitter
augmentation that randomly discards top mel bands, per-source metrics reported on every
split, and the non-FMA probe set. Whether it worked is a question the robustness sweep
answers below.
The model
AudioCNN, in the PANNs CNN10 family: four double-conv blocks with batch norm,
64 to 512 channels, frequency averaged out, then attention pooling over time so the
classifier can weight the frames that matter rather than averaging them flat. 4.82 million
parameters. It trains on the Apple Silicon GPU through MPS, which is the whole reason the
rebuild happened on this machine at all.
Results
The run that mattered was called overnight. 14,543 training tracks, 81,037
clips, batch 64, learning rate 1.5e-4 auto-scaled from the reference batch. Early stopped at
epoch 39; the best checkpoint was epoch 31. The decision threshold, 0.745, was chosen by
maximising balanced accuracy on validation and then applied unchanged to every other split,
which is the only honest way to report it.
| Split | Measures | Clip AUC | Clip acc | Track AUC | Track acc |
|---|---|---|---|---|---|
| val | model selection | 0.9866 | 0.9581 | 0.9618 | 0.8829 |
| test | unseen tracks, seen generators | 0.9844 | 0.9544 | 0.9584 | 0.8793 |
| holdout | unseen generator (Suno) | n/a | 0.8000 | n/a | 0.8196 |
| probe | my own 11 tracks | 0.7411 | 0.6951 | 0.8000 | 0.5455 |
The holdout is all AI, so AUC is undefined there and "accuracy" means recall. Validation and test agree to two decimals, so this is not overfit to the selection set. Against the earlier epoch 8 checkpoint, test AUC moved from 0.9613 to 0.9844.
Training was not smooth. A loss spike at epoch 17 dropped validation AUC from 0.975 to 0.568, and it took about six epochs to recover before going on to its best score. Gradient clipping at 5.0 is almost certainly too loose, and I have left that as a note to my future self rather than pretending the curve was clean.
Generalisation to a generator it never heard
Suno, withheld entirely: 80.0% of clips and 82.0% of tracks. What I find more interesting than the number is that it did not move. Test AUC went from 0.96 to 0.98 between checkpoints and the Suno recall stayed put. Fitting the known generators better bought nothing on the unknown one, which is the kind of result that tells you where the actual ceiling is.
The finding I did not expect
I ran the model across eleven audio transformations on 595 balanced clips, expecting to find the codec shortcut and expecting to find the detector fragile. I found something more specific than either.
Filtering causes calibration drift, not signal loss.
Look at the two worst rows. Lowpass at 1 kHz and highpass at 2 kHz both drive accuracy at the fixed threshold to chance, 0.496 and 0.498. If you only reported accuracy you would conclude the detector had been destroyed. But AUC in those rows is still 0.95 and 0.94. The ranking survived. What moved was the operating point: mean P(AI) swings from 0.303 under highpass to 0.857 under 16 kHz resampling, and a fixed cut-off has no defence against that. Re-choose the threshold on the transformed data and accuracy comes back to 0.87 to 0.90. Lowpass pushes scores up, highpass pushes them down, and the model is still telling the two classes apart underneath.
The literature reads this symptom differently. The TISMIR study saw detectors label everything AI under highpass and everything human under lowpass, and reported it as failure. It may well have been failure for their models. Here the discriminative signal is intact and only the calibration moves, and those need different fixes. You do not retrain for a calibration problem. You calibrate.
Two questions I had carried since phase one closed at the same time. MP3 re-encoding, which overwrites the bitrate fingerprints TISMIR found their classifiers exploiting, costs 0.006 AUC and a point and a half of accuracy here, so the codec shortcut is largely ruled out. And resampling does not break this detector: IRCAM Amplify reportedly misclassified every Suno sample when resampled to 22.05 kHz, while this model holds AUC 0.955 through a 16 kHz round trip and 0.841 through 8 kHz, which discards everything above 4 kHz.
The other thing the threshold was hiding
Per-source recall at the chosen threshold: FMA human 0.973, Udio 0.975, MusicSet 0.522. An earlier version of my notes called MusicSet "near chance," and that was wrong in a way worth being precise about.
MusicSet is separated from human, median 0.764 against 0.195. It is just wide, tenth percentile 0.232, ninetieth 0.950, where Udio sits between 0.884 and 0.964. A global threshold tuned on an aggregate that Udio and FMA dominate lands squarely in MusicSet's middle and halves it. That is a threshold policy problem. Per-source thresholds, or a calibrated score with the decision deferred to the product, would fix it without touching the model. The lesson generalises past music: a single number for "accuracy" can hide an entire subpopulation being cut in half.
Getting it onto the phone, and two bugs parity did not catch
The exported Core ML model takes raw audio and computes its own mel front end, so there is no Swift feature code to keep in sync with Python. The parity check against the PyTorch graph passed at 1e-7. That number is true, and it hid two real bugs.
- fp16 overflow in the exported graph. The squared STFT magnitude for loud music reaches about 2.4e5, past the float16 maximum of 65,504. It saturates to infinity, and every loud clip returned the same probability on device. Quiet clips were fine, which is exactly why nobody noticed. The fix attenuates the waveform by 1/32 before the transform and scales the log epsilon identically, so the change is a constant log-domain offset that per-clip standardisation removes. Cached features stayed valid to 7e-7.
- The parity check was too weak to see it. It probed with quiet synthetic noise. It now probes across five amplitudes and real music, and warns loudly on mismatch. Export defaults to fp32, because on real audio fp16 was systematically low by up to 5.5e-2 where fp32 was off by 1.1e-5.
Then, adding a microphone path, a third one. Core ML parity was 1.7e-4, so the model was
faithful, and the audio reaching it was not. The app's decoded RMS on a test track was
0.1568 against a Python reference of 0.1436, and P(AI) read 0.500 where Python said 0.728.
Two causes, both train/serve skew in the front end: AVAudioConverter's default
resampler uses a gentler anti-alias filter than the soxr_hq that built the
cache, leaving extra high frequency content, and letting it downmix to mono applies its own
channel weighting rather than the plain channel mean numpy performs. Per-clip
RMS differed by 7 to 18 percent depending on how wide the stereo image was. Fixed by
resampling at maximum quality with channels intact and averaging them by hand. RMS now
matches training to 0.06 percent.
A numerically perfect export is not the same as a correct pipeline. Parity has to be checked on decoded audio, end to end.
The residual that remains is about 50 ms: AVFoundation keeps roughly 1,153 samples of MP3 encoder delay that libsndfile strips, so clip windows on the phone sit slightly apart from the ones in the cache. That one is inherent to using two decoders and I have decided to live with it, but I wanted it written down rather than discovered by someone else.
On an actual phone, with actual songs
Everything above is offline evaluation, and I want the next paragraph to sit right beside those numbers so nobody reads one without the other.
The first time I tested it on my own phone, with music I chose, it was bad. The first two songs I tried were human and it called both correctly, and for about a minute I thought it was working. After that it did not really identify anything. Verdicts drifted, tracks I knew the provenance of came back wrong, and there was no consistent signal I could point to and say, that is the model working. On the bench it has a 0.98 AUC. In my hand it was closer to a coin.
I am not going to explain that away, and I am not going to pretend the offline numbers are the real result either. My own test is also a result. It is the one closest to how anyone would actually use this, it is n of not very many, and it says the deployed system does not work well enough to trust. Some of the gap is the front end skew I found afterwards. Some is that my music is not FMA and not Udio. Some of it I do not know yet, and that is the honest state of it.
What is still weak
- On my own eleven tracks, the probe, the score went from 8 of 11 to 6 of 11 between checkpoints, almost entirely because the threshold moved from 0.685 to 0.745 and two tracks now sit just under it. Eleven is a very noisy n, but it is the metric closest to real use, and it is the one that regressed.
- The human class is still entirely FMA. That is a distribution, not the world.
- 8 kHz resampling is the one transform that genuinely costs signal rather than calibration.
- Clips are ten seconds. SONICS argues long range structure, on the order of two minutes, is where a lot of the signal lives, and I have not tested that.
What I actually learned
The technical list is real: grouped splits, held-out generators, shortcut audits, a shared front end, calibration versus discrimination, end to end parity. I can now recognise each of those failure modes by feel, which I could not before, and that recognition is worth more to me than the AUC.
The less technical lesson is that "the project failed" was a story about phase one that I kept telling through phase two. The rebuild has a 0.98 test AUC and an 80 percent hit rate on a generator it never trained on, and then it did badly on my phone. Both of those are true. Calling the whole thing a failure was a calibration error of my own; calling it a success would be another one.
Where this started, and the questions I am left holding
AI music is not new and it is not leaving. This project began before any of the code, in a research study I ran at CMU on how listeners perceive AI-generated music against music made by people. The study was about perception and labelling. At some point I thought: I am technical, I could try to build the label. So I did, and now I know what that costs.
What I believe after building it is that detection from the finished audio is the wrong end to start from. Watermarks exist now, and they are the right idea, because they are built in at generation rather than reverse engineered from the surface. A detector like mine is always going to be chasing a distribution that moves every time a generator ships. Someone will do this better than I did. I am not the best at any single part of it. But being able to go from a research question to a model to an app on my phone, and to find out in my own hand that it does not work yet, is the thing I like most about working across disciplines. You do not just get to ask the question. You get to find out.
And the question has moved while I was building. Friends say a generated track has no soul in it. A family member says they can tell because it just does not sound as good. I used to agree. It is 2026 and we are past the uncanny valley now, and I do not think the ear alone can be relied on any more, mine included. So these are the questions I am actually holding, without answers:
- Should we be afraid of this, in music specifically? Does it matter, and to whom?
- What does it do to a community of listeners, and to the listener alone with headphones on, to not know?
- What does it do to artists? Tools have always been there: the DJ pad, Logic, sample packs, presets. Very little has been made from scratch for a long time. Is a generated track just the next tool, or is it a different kind of thing?
- AI augments me as a developer and I can point at exactly how. How does it augment an artist? Is the augmentation in the lyrics, the arrangement, the sound? Where is the line past which it stops being augmentation?
- What does the next few years of this actually look like, and who decides?
I built a detector because I wanted a yes or no. What I got was a better set of questions and a phone that shrugs at my playlist. I will take it.