Question
What "accurate" actually means in key detection
Accuracy has no fixed meaning in key detection until somebody says what was measured, on which music, and what counted as right. Change only the counting rule and the same tool on the same songs will produce wildly different figures, which is why an accuracy number with no published test behind it tells you nothing you can use.
We build a key reader, so we have an obvious interest in how this word gets used. We also publish no accuracy figure of our own, and the section near the end says why.
The short version
- Accuracy compared to what?
- To a human label. Somebody decided the correct key of each test track, and every score is really a score against those people.
- What counts as right?
- That is a choice, not a fact. Exact match, relative match, same note set and same-notes-different-center are four different rules with four different answers.
- Right on which music?
- A test set is a sample of the world, not the world. Guitar records, dance records and modern trap are not equally easy, and a score is only about what was in the box.
- So what is worth trusting?
- Behavior you can check yourself in a few seconds, on your own music, rather than a figure you cannot reproduce.
Every accuracy score is a score against people
To measure whether a detector is right, you need a set of tracks whose keys are already known. Somebody had to write those keys down, which means every accuracy figure in this category is a comparison against human judgment.
That is fine and it is the only thing available. It is also worth remembering, because trained musicians do not always agree with each other about the key of a modal record, a track that never states a full chord, or a section that sits comfortably in two places.
Where the humans disagreed, a perfect score is not possible, and any tool claiming one has told you something about its test set rather than about itself.
Four rules for counting a hit
These all get called accuracy, and the gap between the loosest and the strictest is very large.
- Exact match. The tool must name the same tonic and the same quality as the label. The strictest rule, and the lowest number.
- Relative counted as correct. Returning the relative minor for a major label scores as a hit. Defensible for a writer, because the notes are identical, and it lifts the figure a long way.
- Any near miss forgiven. The fifth above, the fifth below and the parallel key all counted as partial credit. This is common in academic scoring and it is not what a producer means by right.
- Fit for the job. Did the answer let you pitch the sample correctly, or not. The only rule that matches why anybody opened the tool, and the hardest one to put in a table.
Notice that a single tool, unchanged, moves between those four rules without a line of code being touched. Our page on relative major and minor covers why the second rule is the one that swings the figure most.
The counting rule moves the number more than the algorithm does
This is the part that makes cross-product accuracy claims almost impossible to compare.
Two tools of genuinely similar quality can publish figures far apart because one scored itself strictly on hard material and the other scored itself loosely on easy material. Neither of them lied, and the two numbers still cannot be put next to each other.
So when you see a figure with no test set named, no labeling method described and no scoring rule stated, the honest reading is that it is a claim rather than a measurement.
A test set is a sample of the world, not the world
Accuracy is always accuracy on something. Move the something and the figure moves.
- Genre. Records built on full chord progressions are far easier to read than records built on an 808 and a melody.
- Era and production. Heavy limiting, saturation and wide stereo processing all change what a detector actually sees.
- Tuning. A set of tracks all at standard pitch hides a whole class of failure that shows up the moment something is a half step down.
- Length. Scoring whole finished records is a different task from scoring eight-bar sections, and the second one is what most people actually do.
What a real accuracy claim would have to include
Four things, and a claim missing any of them cannot be checked by the person reading it.
- The test set. What music, how much of it, and how it was chosen.
- The labels. Who decided the correct answers, and what happened where they disagreed.
- The scoring rule. Which of the four rules above, stated plainly.
- Reproducibility. Whether anybody outside the company could run the same test and get the same figure.
Why we do not print a number either
We measure our own reader constantly, and we do not publish the result. That is a deliberate position rather than an omission.
A figure you cannot reproduce is not evidence, and publishing one would only add another uncheckable number to a category that already has several. We would rather be judged on behavior a reader can test in ten seconds on their own music.
It also keeps us honest in the other direction. Once a company has printed a number, every design decision quietly starts serving that number instead of serving the person using the tool.
What to look at instead
Accuracy is not the only property that matters, and on a working session it is not even the most useful one.
- Does it tell you when it is unsure? A tool that is right most of the time and never flags the rest still costs you the verification on every single read.
- Does it refuse? Returning nothing on material with no key is a correctness property that no accuracy figure captures. Our page on why drums have no key explains it.
- Does it answer the question you asked? A correct label for a whole record is the wrong answer about the section you are sampling.
- Is it wrong in a survivable way? The relative minor costs you nothing. A key a semitone away costs you the take.
Frequently asked
How accurate is key detection in general?
Accurate enough to be worth using on ordinary records with clear chords, and unreliable on sparse, modal, heavily processed or off-tuned material. Any single percentage stated without a test set behind it is describing a marketing position rather than that range.
Why do two key detectors disagree about the same song?
Usually because the song genuinely supports more than one answer, most often a relative pair that shares every note. Different analysis windows and different tie-breaking rules then push the two tools to opposite sides of a close call.
Is a higher accuracy claim a better product?
Not on its own. Without the test set, the labeling method and the scoring rule, a higher figure may only mean an easier test or a looser definition of correct.
What should I test before trusting a detector?
Give it material with no key, give it a section rather than a whole record, and give it something tuned below standard. How it behaves on those three tells you more than any published figure.
Session™ Key
A key and BPM reader for Mac that marks every read with how sure it is, so a strong reading and a coin flip never look the same. It reads the audio already playing, and the audio never leaves the Mac.
$39 once, no subscription. Available now for macOS 14.4 or later on Apple Silicon.