Textbook grades
Each compared position yields one verdict per engine: 1 if its Roman numeral matches the textbook label at that bar and beat, 0 if it does not. An engine’s score is the share of its verdicts that are 1. The textbook label is derived from musWM’s own key, root, bass and notes, so the score says whether a label is written as the book writes that reading; it does not say whether the reading itself is right, and it is not a head-to-head accuracy. Root and bass are checked against the written notes separately, on the independent note check.
What the figures support. musWM’s labels are transparent and follow the textbook’s naming rules (99.6%, a consistency check). In unambiguous bars its root and bass match the score at 99.3% and 99.8%. That makes it an interpretable reference to compare against AnalysisGNN and AugmentedNet, not a validated arbiter.
All grades come from a single source: the textbook. Each compared cell is fixed at 1 or 0 by the criterion below, on Alignment and Disagreement rows alike, and no grade can be altered by hand.
Per engine
Accuracy by agreement class
Both rows use textbook grades. On identical rows the engines wrote the same label and so share a score by definition; the different row is where they part company.