Turn an AI use case into its full regulatory footprint — every domain it touches, from AI law and data protection to cyber, product safety and sector rules — with the obligations, the architecture and the evidence you owe, in about two minutes.
Community-curated knowledge graph — every claim carries its citation across law, engineering and governance. Every change traceable →
RAIN will run a real comparative study of how AI use cases in regulated industries get scoped today — and this protocol is published before the first participant is recruited.
Pilot published · the human study has not started.A machine-conducted pilot source audit ran on 9 August 2026 and its results are published below, clearly labelled. The pre-registered human study — participants, panel, blinded raters — has not begun recruiting, and this protocol was published before it does.Protocol published 9 August 2026 · graph v2.13.0 · jump to the pilot results ↓ · timestamped in the changelog →
The question
For AI use cases in regulated industries, how do three scoping methods compare on the five pain points decision-makers actually have?
Arm A
Conventional research
Legal texts, firm memos, internal precedent and general web search — what a compliance or legal team does today with no dedicated tooling.
Arm B
Open checkers and databases
EU AI Act Service Desk, FLI Compliance Checker, appliedAI use-case database, CSA AICM, NIST AI RMF playbook — the free public instruments practitioners already reach for.
Arm C
RAI·N·avigator
The knowledge-graph traversal: risk class → obligations → control objectives → architecture → evidence → market layer, with citations and verification dates.
Coverage
“What all applies?” — obligations found across every regulatory layer, not just the AI Act.
Currency
“Is it still true?” — withdrawn instruments, wrong standard status, superseded deadlines.
Translation gap
From legal obligation to a concrete technical control, component and evidence artifact.
Speed & cost
Minutes and euros to a defensible scoping answer.
Decision-readiness
Can a decision-maker actually say go / no-go / redesign at the end?
Design summary
Cases
30 AI use cases, stratified by risk class (prohibited / high / limited / minimal) and by sector (financial services, health, public sector, HR, industrial). None are drawn from RAIN's own catalog; any conceptual overlap with an existing profile is documented case by case and reported.
Participants
12–18 practitioners (compliance officers, in-house counsel, AI architects, risk managers) in a within-subject latin square, so every participant works in all three arms on different cases and arm order is balanced. Hard cap of 90 minutes per case.
Gold standard
Built independently and WITHOUT RAIN by a named expert panel: two qualified lawyers plus one compliance architect. The panel's reference answer per case is published with the results.
Scoring
Two independent raters, blinded to arm and to which tool produced an answer, scoring against a rubric that is published before data collection. Inter-rater reliability (Krippendorff's α) is reported; disagreements are adjudicated by the panel and the adjudication rate is reported.
Endpoints
Pre-registered endpoints and how each is measured
Endpoint
How it is measured
Obligation recall
Share of gold-standard obligations found, reported per regulatory layer (AI, data protection, cyber & resilience, product safety & liability, financial, sector).
Precision
Share of stated obligations that are actually applicable. Over-inclusion counts as an error — naming everything is not coverage.
Currency errors
Count of statements relying on withdrawn instruments, wrong standard status, or superseded dates and deadlines.
Chain depth
0–5: risk class → obligations → control objectives → architecture → evidence → vendor/market layer. How far the answer actually gets.
Time
Minutes to a scoping answer the participant is willing to hand to a decision-maker.
Cost
Loaded practitioner time plus any tool or subscription cost, per case.
Go/no-go quality
Panel rating of the final recommendation: defensible, defensible-with-gaps, or not defensible.
Hypotheses
H1Coverage: arm C finds more gold-standard obligations outside the AI Act than arms A and B, with the largest gap in the cyber & resilience and product-liability layers.
H2Currency: arm C produces fewer currency errors than arms A and B, because status and verification dates are rendered next to every claim.
H3Translation: arm C reaches a greater chain depth — specifically, more cases reach the evidence and architecture steps at all.
H4Speed: arm C reaches a decision-ready answer in less practitioner time per case than arm A, and no slower than arm B.
H5Precision: arm C shows no precision disadvantage versus conventional research. A graph can over-include; if it does, we publish that result and fix the graph.
Bias controls
Conflict of interest, stated plainly: RAIN runs this study on its own product. That is a real bias and no amount of method removes it — these six mitigations are how we make the result checkable anyway.
Public pre-registration: this protocol was published before recruitment and is timestamped in the changelog. Any deviation is documented as a deviation.
Independent, named expert panel builds the gold standard without access to RAIN.
Independent, blinded raters score against a rubric fixed before data collection.
Raw data, rubric and scoring sheets are published with the results.
Publication regardless of outcome — including a null or negative result for RAIN.
External co-investigators are invited, with the right to publish a dissenting reading alongside ours.
A mandatory limitations section: sample size, practitioner self-selection, prompt and search skill, and the conflict of interest named above.
Participate
Three roles. Each request runs through the normal invite path, with a study marker the curator sees on review.
Practitioner participant
Work 2–3 cases across the three arms, max 90 minutes each. Compensated in contribution credits plus named credit in the publication (opt-out available).
Read this label before the numbers.This pilot was conducted by AI agents auditing information sources — not by human practitioners. Time, cost and human-performance endpoints were NOT measured; they are reserved for the pre-registered human study below. Conducted 2026-08-09 against graph v2.3.1.
Five use cases were scoped three ways and scored against primary-source reference dossiers: EU credit scoring; German ambient clinical documentation; CV screening under French law plus NYC LL144; agentic trade-settlement reconciliation for an EU broker with a US SEC/FINRA leg; and an EU insurance information / first-notice-of-loss chatbot. Three endpoints only — instrument coverage against the reference dossier, currency of the instruments cited, and chain depth reached.
Web research34%
10 / 29 instruments — stale top results in 3/5 cases
Free checkers17%
5 / 29 instruments — single-act scope by design
RAINavigator62%
18 / 29 instruments — full chain to design & evidence
Conventional web research
10 of 29 reference instruments (34%). Confirmed stale top results in 3 of 5 cases.
Best open checker per case
5 of 29 reference instruments (17%). Legitimately correct risk-class answers in 4 of 5 cases, including the valuable negative finding.
RAI·N·avigator graph
18 of 29 reference instruments (62%). Full six-stage chain depth in 4 of 5 cases.
Machine-conducted pilot source audit, Aug 2026 — method, limitations and our own 12 found weaknesses: /study
Currency. The graph carried the Digital-Omnibus dates, the AILD and SEC-PDA withdrawals and the EN 18286 status correctly. It also carried two deficiencies of its own, both found by the audit: the Colorado node was stale, and DORA's in-force date was absent. Both are fixed in the release logged below.
Fairness note. The free checkers do what they claim, and they do it well — scoring them on chain stages they never claimed is a statement about scope, not about quality.
The audit scored our own graph too. Found and queued:
11 missing instruments — CCD2, EBA loan-origination guidelines, EIOPA AI Opinion, IDD, MDCG 2025-6, the German NIS2 transposition plus national health layer, MiFID II / CSDR / T+1, FINRA 3110, SR 11-7, Illinois AIVIA, and GDPR Art. 22 case law.
1 stale node — Colorado: the old 30 June 2026 effective date survived a court stay and a repeal-and-replace.
1 structural gap — non-AI-Act regulations impose no control objectives, so UC4's control stage came back empty. Now a permanent integrity warning.
1 profile-match limit — the generic chatbot profile lacks the insurance sector layer.
The human-study results will appear here.Nothing to see yet is the point of pre-registration. When data collection closes, this section fills with the endpoint tables, the raw data, the rubric, the inter-rater reliability and the limitations — whatever the numbers say.