Method
Correctness measurements
The suite compares what the reasoner derives against expectations authored from statutory text, and it injects deliberate faults into the graph to find out which errors the suite is blind to. Both results are published here in full, including the parts that are unfavourable. There is no single accuracy figure on this page: an average over unlike assertions hides the one number that matters, which is how often a law is claimed that does not apply.
Fixture and version
- Graph version
- 2.13.0
- Graph nodes
- 475
- Cases in the fixture
- 60
- Expectations met
- 54 met · 6 unmet
Graph content hash 9d0641791cb90d00ce812f89b29da690de324b255edf6db48e8ca9585a70a94c · dated 2026-08-28. Every case, expectation, negative reason and source URL lives in one file: src/testing/golden/cases.ts.
5 of 60 cases are not independently verified: their expectations do not yet carry an article-level source with a URL and a retrieval date for every value. They still run and are reported, and they are excluded from the sourced counts above.
Cases per category
Composition is part of the result. 26 of 60 cases assert that something does not apply, 18 of 60 are one half of a paired case where only one variable differs, and 8 of 60 turn on the status of an instrument in time rather than on its subject matter.
| Category | Cases | Sourced | Not independently verified | Met | Unmet | Accepted |
|---|---|---|---|---|---|---|
| Externally-authored scenarios | 30 | 30 | 0 | 27 | 3 | 3 |
| Jurisdiction pairs | 6 | 5 | 1 | 6 | 0 | 0 |
| Temporal status | 8 | 8 | 0 | 8 | 0 | 0 |
| Adversarial input | 5 | 5 | 0 | 5 | 0 | 0 |
| Degenerate input | 2 | 2 | 0 | 2 | 0 | 0 |
| Confidence calibration | 4 | 0 | 4 | 4 | 0 | 0 |
| Role pairs | 5 | 5 | 0 | 2 | 3 | 3 |
Metric matrix
- Exclusion precision
- 100.0%
- Inclusion recall
- 100.0%
- Flag recall
- 83.3%
32 negative assertions — how often an instrument the fixture says does not apply was nevertheless presented as applicable
61 containment assertions — how often a duty the fixture requires was actually derived
12 flag assertions — contested, deferred, stayed, not-in-force, role-dependent, insufficient input, prohibited
Mutation kill rate, per operator
A green suite may be green because the graph is right or because the cases assert nothing that could go wrong. The harness settles that: it injects one plausible graph fault, re-runs the suite, and asks whether any previously-met case now fails. Faults that nothing notices are named below. Recorded 2026-08-28 against graph v2.13.0: 8 mutants per operator, 58 of 475 nodes actually asserted about by the fixture.
| Fault operator | What it simulates | Injected | Detected | Kill rate | Candidates (asserted / total) |
|---|---|---|---|---|---|
applies-in-flip | an instrument is mapped to the wrong jurisdiction (copy-paste across a border) | 8 | 6 | 75% | 16 / 86 |
status-force-in-force | a pending, withdrawn or repealed instrument is presented as applicable law | 8 | 3 | 38% | 5 / 23 |
triggers-delete | a duty-creating edge is lost in an edit, so a real obligation disappears | 8 | 2 | 25% | 164 / 229 |
triggers-spurious | an obligation is attached to a use case it does not govern (over-claiming) | 8 | 0 | 0% | 19 / 45 |
tier-shift | a classification lands one step off — the single most consequential error | 8 | 6 | 75% | 19 / 45 |
gate-invert | a boolean gate on a node is inverted (profiling, scoring, always-applicable) | 8 | 3 | 38% | 34 / 120 |
crosswalk-swap | a crosswalk claims equivalence where the mapping is only an overlap | 8 | 0 | 0% | 12 / 21 |
Surviving mutants — the blind spots
Each line is a specific graph error the suite does not currently detect. This list is the useful output of the harness; the percentages above are only its summary.
applies-in-flip — an instrument is mapped to the wrong jurisdiction (copy-paste across a border)
- · reg-co-admt applies_in jur-us-co → jur-us-il (Colorado ADMT Act (SB 26-189))
- · reg-colorado applies_in jur-us-co → jur-us-il (Colorado AI Act)
status-force-in-force — a pending, withdrawn or repealed instrument is presented as applicable law
- · reg-aild status "withdrawn" → "in-force" (AI Liability Directive (withdrawn))
- · reg-colorado status "repealed — reenacted by SB 26-189 (2026); never took effect" → "in-force" (Colorado AI Act)
- · ae-pdpl status "enacted-not-yet-applicable" → "in-force" (Federal PDPL — Decree-Law 45/2021 (AE))
- · au-mandatory-guardrails status "withdrawn" → "in-force" (Mandatory AI guardrails proposals paper (AU))
- · ca-bill-c36 status "pending" → "in-force" (Bill C-36 — Protecting Privacy and Consumer Data Act (CA))
triggers-delete — a duty-creating edge is lost in an edit, so a real obligation disappears
- · delete uc-admissions-utility triggers art-49
- · delete uc-admissions-utility triggers reg-aiact
- · delete uc-admissions-utility triggers reg-gdpr
- · delete uc-biometric-access triggers reg-aiact
- · delete uc-biometric-access triggers reg-gdpr
- · delete uc-biometric-access triggers reg-il-bipa
triggers-spurious — an obligation is attached to a use case it does not govern (over-claiming)
- · add uc-admissions-utility triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-biometric-access triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-chatbot triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-codegen triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-confidential-summary triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-credit triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-diagnosis triggers art-13 (Art. 13 — Transparency to Deployers)
- · add uc-emotion-interview triggers art-13 (Art. 13 — Transparency to Deployers)
tier-shift — a classification lands one step off — the single most consequential error
- · uc-admissions-utility classified_as rc-limited → rc-high
- · uc-emotion-interview classified_as rc-prohibited → rc-high
gate-invert — a boolean gate on a node is inverted (profiling, scoring, always-applicable)
- · br-pl2338.operatorBinding false → true (PL 2338/2023 AI framework (BR))
- · ca-aida.operatorBinding false → true (AIDA — Bill C-27 (CA))
- · reg-aild.operatorBinding false → true (AI Liability Directive (withdrawn))
- · reg-co-admt.operatorBinding false → true (Colorado ADMT Act (SB 26-189))
- · reg-colorado.operatorBinding false → true (Colorado AI Act)
crosswalk-swap — a crosswalk claims equivalence where the mapping is only an overlap
- · art-15 overlaps_with std-iso27001 → equivalent_to
- · art-15 overlaps_with std-nist600 → equivalent_to
- · art-27 overlaps_with std-iso42005 → equivalent_to
- · art-72 overlaps_with nis2-23 → equivalent_to
- · art-9 overlaps_with std-iso23894 → equivalent_to
- · art-9 overlaps_with std-nist → equivalent_to
- · gdpr-22 overlaps_with art-14 → equivalent_to
- · gdpr-33 overlaps_with nis2-23 → equivalent_to
Failures accepted with an owner
A case whose expectation is not met may be accepted rather than fixed, but only in writing: with an owner, the date it was accepted, and the reason. The suite asserts that every case marked this way still fails, so an accepted failure cannot quietly become a passing one.
s-legalresearch· Externally-authored scenariostier lands on rc-limited; role-dependent reading not surfaced as a flag
Owner kg-curators · accepted 2026-08-15
s-admissions-utility· Externally-authored scenariosauthored contested reading not surfaced as a flag
Owner kg-curators · accepted 2026-08-15
r-hr-provider· Role pairsprovider/deployer duty split not applied to Art. 43
Owner kg-curators · accepted 2026-08-15
r-hr-deployer· Role pairsprovider/deployer duty split not applied to Art. 43
Owner kg-curators · accepted 2026-08-15
r-rebrand-user· Role pairsArt. 25 role shift not applied
Owner kg-curators · accepted 2026-08-15
g-scope-drift-before· Externally-authored scenariosThe reasoner tiers this as high risk on the employment keyword alone ('HR handbook'), although the described system evaluates no person and matches no Annex III point — the over-triggering failure the deep-search audit predicted. Recorded as a disclosed recall/precision gap on the tier, not as a claimed duty: the case asserts no negative expectation, so nothing wrong is published while the detector is narrowed.
Owner RAIN editorial · accepted 2026-08-17
What this does not cover
- It is not legal advice and not a legal opinion. The expectations were authored by reading statutory text; agreement with them is agreement with a reading, not a determination by a competent authority or a court.
- No independent third party has checked this fixture. Where a case is marked not independently verified, no article-level source is attached to every expected value yet.
- 60 cases are a sample, not a population. The graph carries 475 nodes and the fixture asserts about 58 of them. A fault in a node no case asserts about cannot be detected here — by construction, not by oversight.
- The mutation figures describe a sample of faults (8 per operator), chosen from the asserted region. A 100% kill rate for an operator would be a statement about those mutants only.
- Crosswalk faults remain 0/8 killed, and the zero is real — a declared assertion-vocabulary gap, unchanged by the latest release. The crosswalk-swap operator turns an overlaps_with mapping into equivalent_to. The derivation path the fixture asserts over never reads the relation type — the crosswalk view is a separate surface — so no case can observe the change. This is a coverage gap in the assertion vocabulary, not a fixture-authoring gap, and it is reported rather than papered over with a case that cannot fail.
- Nothing here measures the law itself being current. Instrument status and dates are tracked separately, per node, with a retrieval date and a source; the suite only checks that the derivation respects the status the graph records.
- Confidence bands are ordinal. Cases may assert that one situation is banded strictly lower than another; no case asserts an absolute confidence number, because that would be a claim about the implementation rather than about the law.
- Coverage of markets is uneven. Most cases run against EU and US expectations; other jurisdictions in the graph carry far fewer, or none.
Related: the public track record logs every correction made after publication, and sources and references lists the instruments behind the graph.