Summary
musWM, AnalysisGNN and AugmentedNet are set side by side on one corpus and aligned by bar and
beat. The measures below are kept apart because each answers a different question. Theory conformance asks whether musWM’s own label is the one the textbook derives from the same notes and root. Textbook grades ask, for each engine, how many of its labels at the compared positions are written as the textbook writes musWM’s reading of the chord — a test of naming, not of the reading, and not a head-to-head accuracy. The independent note check asks, from the written notes alone, whether each engine’s root and bass are the ones in the score. Agreement between engines asks how often all three engines give the same label where all three produce one; it is shown below and row by row on the Alignment and Disagreements pages.
1 · Theory conformance
—
Does musWM’s label follow the textbook? Measured against
Kostka, Payne & Almén, Tonal Harmony, 8th edition, over the corpus shown. No human annotation
involved. → Theory
2 · Textbook grades
—
One point or none per engine, fixed in advance, against the textbook label. → Ratings
3 · Independent note check
—
Root and bass read from the written notes, in bars holding one unambiguous chord. → Method
Preliminary results. Every figure on this site is recomputed from the corpus pack at each build. On 15 September 2026 the pack was rebuilt after an alignment fault was found and corrected, and the textbook criterion was audited against the book — see
Corrections. The textbook grades measure tolerant agreement with a label derived from musWM’s own reading of the chord; they are not an accuracy against an independent reference.
What this site does not claim. Roman numeral analysis has no single correct answer that all
of these could be scored against. As the When in Rome authors put it, “analysis corpora
are problematic in several ways that make the term ‘ground truth’ almost never
appropriate” (Gotham et al., 2023). Yet while a passage may allow more than one defensible
reading, what is untenable is always obvious to a music analyst: a label that contradicts the
sounding notes, the bass or the key cannot be right under any reading. Each number above is
therefore named for what it actually measures. What the figures support: musWM’s labels are transparent and follow the textbook’s naming rules (99.6%, a consistency check). In unambiguous bars its root and bass match the score at 99.3% and 99.8%. That makes it an interpretable reference to compare against AnalysisGNN and AugmentedNet, not a validated arbiter.
Roman numerals where all three read the same key
Agreement between engines (all positions)
Textbook grades (locked)
Independent note check
Pairwise agreement (rows where both engines produced a label)
Bar-number mapping
The bar number reported by the engines is not always the written bar number.
This mapping is applied when a score is opened.
Most frequent disagreement triples