How Scanara's rules are built and verified
Every rule is anchored to the regulation's exact article, paragraph, and point — and gated through six verification layers before it ships.
Measured on September 9, 2026 — re-verified before every release.
How rules are built
Scanara's rule pipeline parses the full text of Regulation (EU) 2024/1689 and anchors every rule to its exact article, paragraph, and point. A single mapping file classifies every obligation in the regulation as one of three things: detectable in code, detectable in a policy document, or not machine-detectable and left to a guided assessment.
The release gate ladder
Every ruleset release runs through six gates in series, from broad structural checks down to an exact diff against known-good output. A release does not ship unless every gate passes.
L2.5
Catalog parity
Checks that the article catalog, the obligation map, and any overlap-suppression rules stay in sync with each other.
L1/L2
Corpus structure
Nine structural checks across the rule corpus: article coverage, no duplicate patterns, complete metadata, valid enum values, unique rule IDs, and stub-rule limits.
L3
Per-obligation fixtures
For every code-detectable obligation, a positive fixture must trigger at least one finding and a negative fixture must trigger zero — in every supported language.
L4
KPI thresholds
17 coverage, completeness, and determinism thresholds — see the table below.
L5
Golden snapshots
An exact diff of every finding against a committed expected result, across 27 benchmark applications, pinned to a specific ruleset version.
R1
Remediation gate
Confirms that a bare fixture triggers an unsatisfied finding, and that the same fixture with the recommended remediation applied is recognized as satisfied.
How Scanara evaluates documentation with OPA and Rego
Scanara's document-level compliance checks run on Open Policy Agent (OPA). OPA evaluates technical documentation against EU AI Act obligations. This follows the same policy-as-code approach used throughout the OPA ecosystem: declarative rules, deterministic evaluation, and an auditable result every time. This runs alongside Scanara's source-code scanning. Code scanning checks what the code does. OPA checks what the documentation says.
Each documentation file goes through the same steps. First, Scanara parses it and detects its language (8 EU languages are supported). Then it normalizes the file into a structured input document: sections, headings, extracted text. OPA evaluates this structured document against a compiled Rego policy bundle.
Scanara generates these policies from a canonical mapping between EU AI Act article text and machine-checkable requirements. No policy is hand-written per repository. Every rule traces back to one specific article, paragraph, and point. Evaluation produces one finding per obligation: satisfied, missing, or partial. These findings feed into the compliance report together with the source-code findings.
The example below shows the pattern. It's simplified for public documentation. The production ruleset covers dozens of obligation categories, plus multi-language detection, severity weighting, and cross-document consistency checks not shown here.
package eu_ai_act.example.article_11
# Illustrative pattern only — simplified for public documentation.
# Scanara's production ruleset covers dozens of obligation categories with
# multi-language detection, severity weighting, and cross-document
# consistency checks not shown here.
default satisfied := false
# A document satisfies this (illustrative) Article 11 check when its parsed
# sections include the required topic, with non-empty content.
satisfied if {
some section in input.document.sections
section.topic == "risk_management_process"
count(trim_space(section.content)) > 0
}
finding := {
"article": "11",
"obligation": "technical_documentation.risk_management_process",
"status": "satisfied",
} if satisfied
finding := {
"article": "11",
"obligation": "technical_documentation.risk_management_process",
"status": "missing",
"severity": "high",
} if not satisfiedTogether, the two engines automatically detect 66 EU AI Act articles: 44 code-detectable and 22 document-level via OPA. Coverage spans all 8 EU languages. Findings from both engines feed directly into the generated compliance dossier. The evidence trail traces back to an actual scan result, not a self-reported checklist.
Current measured values
| Metric | Threshold | Current |
|---|---|---|
| Article coverage (119 articles of the regulation) | 100% | 100% |
| Real-pattern coverage of code-eligible articles | 100% | 100% (45/45) |
| Document/policy articles with OPA coverage | 100% | 100% (22/22) |
| Code-detectable obligations with at least one rule | 100% | 100% (80/80) |
| Rules carrying risk-level metadata | 100% | 100% (545/545) |
| Rules carrying actor metadata | 100% | 100% (545/545) |
| Rules carrying structured remediation guidance | 100% | 100% (545/545) |
| Spread in rule count across the 12 supported languages | ≤ 1 | 0 |
| Duplicate rule IDs | 0 | 0 |
| Stub rules (no code pattern) per article | ≤ 1 | 0 |
| L4 KPI thresholds passing | 17/17 | 17/17 |
What these numbers do not measure
These are coverage, completeness, and determinism metrics — they show how much of the regulation is mapped and how consistently the pipeline behaves across languages and releases. There is no precision, recall, or false-positive-rate KPI, because that requires a human-labelled ground-truth corpus we don't have. The structural evidence against false positives is this: every per-obligation negative fixture must produce zero findings, and a dedicated minimal-risk benchmark application — with every high-risk, GPAI, and limited-risk rule suppressed — must also produce zero findings.
Scanara catches code-level and document-level compliance gaps. For complex legal interpretations, we recommend pairing with specialized legal counsel.