Summary

musWM, AnalysisGNN and AugmentedNet are set side by side on one corpus and aligned by bar and beat. The measures below are kept apart because each answers a different question. Theory conformance asks whether musWM’s own label is the one the textbook derives from the same notes and root. Textbook grades ask, for each engine, how many of its labels at the compared positions are written as the textbook writes musWM’s reading of the chord — a test of naming, not of the reading, and not a head-to-head accuracy. The independent note check asks, from the written notes alone, whether each engine’s root and bass are the ones in the score. Agreement between engines asks how often all three engines give the same label where all three produce one; it is shown below and row by row on the Alignment and Disagreements pages.

1 · Theory conformance Does musWM’s label follow the textbook? Measured against Kostka, Payne & Almén, Tonal Harmony, 8th edition, over the corpus shown. No human annotation involved. → Theory
2 · Textbook grades One point or none per engine, fixed in advance, against the textbook label. → Ratings
3 · Independent note check Root and bass read from the written notes, in bars holding one unambiguous chord. → Method
Preliminary results. Every figure on this site is recomputed from the corpus pack at each build. On 15 September 2026 the pack was rebuilt after an alignment fault was found and corrected, and the textbook criterion was audited against the book — see Corrections. The textbook grades measure tolerant agreement with a label derived from musWM’s own reading of the chord; they are not an accuracy against an independent reference.
What this site does not claim. Roman numeral analysis has no single correct answer that all of these could be scored against. As the When in Rome authors put it, “analysis corpora are problematic in several ways that make the term ‘ground truth’ almost never appropriate” (Gotham et al., 2023). Yet while a passage may allow more than one defensible reading, what is untenable is always obvious to a music analyst: a label that contradicts the sounding notes, the bass or the key cannot be right under any reading. Each number above is therefore named for what it actually measures. What the figures support: musWM’s labels are transparent and follow the textbook’s naming rules (99.6%, a consistency check). In unambiguous bars its root and bass match the score at 99.3% and 99.8%. That makes it an interpretable reference to compare against AnalysisGNN and AugmentedNet, not a validated arbiter.

Roman numerals where all three read the same key

Loading…

Agreement between engines (all positions)

Textbook grades (locked)

Independent note check

Loading…

Alignment check

Loading…

Pairwise agreement (rows where both engines produced a label)

Bar-number mapping

The bar number reported by the engines is not always the written bar number. This mapping is applied when a score is opened.

Most frequent disagreement triples