BlogResearch
Does MP3 hurt music transcription? We measured WAV vs MP3
We ran the same recordings through three transcription engines as lossless audio, MP3 and AAC, beside an inaudible control. What changed, what didn’t, and which file to upload.
The short answer
Not at 320 kbit/s. On October 5, 2026 we transcribed the same recordings as lossless audio and after a round trip through MP3 and AAC, with three engines: Spotify’s Basic Pitch on 12 guitar recordings, a solo-piano model on 50 piano performances and Mirelo on a band song. MP3 at 320 kbit/s moved each engine’s onset F1 by at most 0.002 on a scale of 0 to 1, a handful of notes in a few thousand, and kept every file’s length to the sample. MP3 at 192 kbit/s and AAC at 256 kbit/s moved scores by up to 0.008, up as often as down, and the AAC files came back up to 733 samples longer.
Why would MP3 change the notes?
MP3 and AAC throw away detail the ear is unlikely to notice, mostly quiet sounds next to louder ones and high frequencies. Models such as Basic Pitch read a time-frequency picture of the audio, and nothing guarantees they ignore the same detail a listener does. A note near a model’s threshold can appear or vanish when the spectrum shifts a little.
Earlier work points to a small effect. Urbano and colleagues found MFCC features stable under lossy encoding from about 160 kbit/s, and chroma features, which follow pitch classes, hardly affected by codec or bitrate at all (ISMIR 2014). They measured features, not transcribed notes, so we measured the notes.
We had a practical reason too. Lossless copies of recordings and of their playback made up 77% of our score storage, and we wanted to keep smaller copies only if they cost nothing in accuracy.
How we tested
Every recording went through each codec with FFmpeg 8.1.1 (LAME at a constant 320 or 192 kbit/s, FFmpeg’s own AAC encoder at 256 kbit/s), then was decoded by the same code that reads an upload. A control changed the lossless audio by 0.01 dB, which nobody can hear, to show how much the measures move with no codec at all. All three engines are deterministic: the same input gives the same notes.
| Engine | Recordings | Reference notes |
|---|---|---|
| Basic Pitch, as ScoreStarling runs it for one instrument | 12 GuitarSet recordings (a solo and an accompaniment from each of six players), the first 60 seconds of a Slakh2100 band song, and a 58-second band piece we wrote and synthesized | The datasets’ annotations |
| Kong and colleagues’ solo-piano model | 25 piano pieces (24 public-domain scores and one we wrote), each played twice with different rubato and rendered with a SoundFont piano and sustain pedal: 50 performances | The notes played |
| Mirelo’s Audio-to-MIDI (a2m-1.1) | The same 60-second Slakh band song, lossless and MP3 320 only | The song’s MIDI |
Onset F1 counts a transcribed note as right when its pitch is within 50 cents and its start within 50 ms of a reference note, and it punishes missed and extra notes alike; 1.0 is perfect. Our accuracy article explains the measure in full.
What changed for one instrument?
Almost nothing. Across the 12 guitar recordings Basic Pitch’s onset F1 stayed between 0.761 and 0.764 in every format, while the stereo file shrank to between 14% and 23% of the WAV.
| Measure | Lossless | Control | MP3 320 | AAC 256 | MP3 192 |
|---|---|---|---|---|---|
| Onset F1, guitar | 0.7626 | 0.7627 | 0.7611 | 0.7644 | 0.7616 |
| Onset and offset F1, guitar | 0.5224 | 0.5222 | 0.5212 | 0.5228 | 0.5210 |
| Onset F1, Slakh song | 0.5204 | 0.5210 | 0.5204 | 0.5240 | 0.5227 |
| Onset F1, our band piece | 0.7136 | 0.7136 | 0.7123 | 0.7089 | 0.7054 |
| Notes found, of 3,893 | 2,741 | 2,742 | 2,739 | 2,744 | 2,737 |
| Found notes that changed | — | 1 | 10 | 17 | 14 |
| Key as annotated | 10 of 12 | 10 of 12 | 10 of 12 | 10 of 12 | 10 of 12 |
| Tempo unchanged | — | 14 of 14 | 14 of 14 | 14 of 14 | 14 of 14 |
| Length change, samples | — | 0 | 0 | +16 to +733 | 0 |
| Signal-to-noise, dB (median) | — | 58.4 | 53.2 | 39.4 | 30.3 |
| Size, mono / stereo | 100% | 100% | 45% / 23% | 23% / 17% | 27% / 14% |
The lossless score is 0.763, not the 0.744 in our October 3 article: since October 4, ScoreStarling writes one instrument as one line, which changed five of the solos. Both runs use Basic Pitch 0.4.0.
And for solo piano?
The piano model heard the same notes in every format. Its onset F1 against the notes played was 0.9427 lossless, 0.9426 for the control, 0.9428 for MP3 320 and 0.9433 for MP3 192, and it found about 92.2% of the written notes in every format.
The step that reads beats and bar lines from those notes was less steady, and not because of the codec. It placed the bar lines differently in 5 of the 50 performances after the inaudible control, and in 7 after MP3 320, four of them the same performances as the control. That step is sensitive to any change at all, so we treat it as noise in the measure rather than a cost of MP3. AAC was not run on piano.
And for a full band?
One paid comparison with Mirelo, on the 60-second band song: 842 notes and drum strokes came back for both the lossless file and the MP3, and 829 of them (98.5%) started within 50 ms of the lossless run’s.
| Measure | Lossless | MP3 320 |
|---|---|---|
| Onset F1, pitched notes (703 in the song) | 0.6707 | 0.6687 |
| Drum onset F1 (423 strokes) | 0.6983 | 0.7023 |
| Bass / guitar / piano onset F1 | 0.455 / 0.681 / 0.827 | 0.471 / 0.678 / 0.832 |
| Tempo, BPM | 133.0 | 133.0 |
Why did any notes change?
Because any change to the input moves a few notes that sit right at a model’s threshold. Of about 2,740 notes Basic Pitch found, the inaudible control moved one, MP3 320 moved 10, AAC 17 and MP3 192 14, never more than 7 in one recording. MP3 320 came out a hair lower in three of the five comparisons (by 0.0013 to 0.0020), level in one and a hair higher in one. AAC and MP3 192 went up on the guitar or the Slakh song and down on our synthesized band piece, by up to 0.008. A codec that cost real accuracy would show up as a clear drop everywhere, and none did.
What a codec can change for certain is timing at the edges. AAC encoders add silent priming samples at the start and pad the end to a whole frame; Apple documents a delay of 2112 samples. Our AAC files decoded 16 to 733 samples longer. FFmpeg writes MP3 with a Xing/LAME header that records the encoder delay, and our MP3 files decoded to the exact original length. That matters when a score is played in sync with its recording.
Which file should you upload?
- The file you already have. Converting a WAV to MP3 first gains nothing, and converting an MP3 back to WAV restores nothing.
- If you must compress, MP3 at 320 kbit/s matched the original most closely by signal-to-noise and kept the exact length.
- Low bitrates are unmeasured. We stopped at 192 kbit/s, so check what your recording app saves.
- Recording quality matters far more than the format: one clear part, little reverb, no background music. Our voice memo guide covers what to check.
ScoreStarling acts on this result: since October 5, 2026 a WAV, AIFF or FLAC upload is kept as a 320 kbit/s MP3 and transcribed from that copy, which stores a stereo recording at 23% of its WAV size. MP3, M4A, Ogg and video files are kept as they came.
What this test can’t tell you
- One band song for Mirelo, and synthesized piano rather than recorded pianos.
- No singing. Our real recorded takes are private and were already AAC files, so they could not serve as a lossless reference.
- One encoder per format and no bitrate below 192 kbit/s.
- Onset F1 says whether a note was heard, not whether the written score reads well.
Sources
- What Is the Effect of Audio Quality on the Robustness of MFCCs and Chroma Features? — Urbano, Bogdanov, Herrera, Gómez and Serra, ISMIR 2014
- GuitarSet 1.1.0 — Xi, Bittner, Pauwels, Ye and Bello, Zenodo (CC BY 4.0)
- Slakh2100 — Manilow, Wichern, Seetharaman and Le Roux, Zenodo (CC BY 4.0)
- A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation — Bittner et al., Spotify, 2022 (Basic Pitch)
- High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times — Kong, Li, Song, Wan and Wang, 2020
- mir_eval.transcription — mir_eval documentation
- FFmpeg formats: the mp3 muxer — FFmpeg documentation
- TN2258: AAC Audio, Encoder Delay and Synchronization — Apple
Questions and answers
Should I upload WAV or MP3 for transcription?
Either. In our October 5, 2026 test an MP3 at 320 kbit/s scored within 0.002 onset F1 of the WAV it came from, with three different engines. Upload the file your recorder or editor gives you, and skip converting it. An MP3 at 192 kbit/s and an AAC file at 256 kbit/s stayed within 0.008; lower bitrates were not tested.
Does a lower MP3 bitrate change the notes?
A few, in both directions. At 192 kbit/s, 14 of about 2,740 matched notes changed across 14 recordings, against 10 at 320 kbit/s and 1 for an inaudible volume change. Onset F1 moved by up to 0.008 at 192 kbit/s and by at most 0.002 at 320 kbit/s. Bitrates below 192 kbit/s were not measured.
Why does an AAC file come out slightly longer than the WAV?
AAC encoders add silent priming samples at the start and pad the end to a whole frame; Apple documents a priming delay of 2112 samples. If the player or decoder does not trim them, the audio is a little longer. In our test AAC files decoded 16 to 733 samples longer. MP3 written by FFmpeg with its Xing/LAME header decoded to the exact length.
Does ScoreStarling convert my WAV to MP3?
Yes. Since October 5, 2026 a WAV, AIFF or FLAC upload is kept as a 320 kbit/s MP3, the setting measured here, and transcribed from that copy. MP3, M4A, Ogg and video uploads are kept as they came.