LLM testing in a regulated industry passes an audit when every rule you are subject to maps to a named test, and every test leaves an evidence artefact that someone outside the team can check. Accuracy scores alone don't do that. In practice, an audit-ready plan has four parts: a versioned golden dataset labelled by domain experts, regression gates in CI on every model or prompt change, adversarial and data-boundary tests, and scored samples of live traffic with human review. Most teams are not there yet. In LangChain's 2026 State of Agent Engineering survey, 94% of teams with agents in production had observability, but only 37.3% of all respondents ran online evaluations. This guide maps HIPAA, PCI DSS, the EU AI Act and US banking guidance to concrete tests and artefacts. Updated October 2026.
A generic eval suite asks whether the model is good. An auditor asks a different question: show me the control, show me it ran, and show me what happened when it failed. Those questions need evidence that holds up after the fact: which model version answered, what data it saw, which test set it passed, who reviewed the exceptions, and what changed afterwards.
The gap is real. LangChain surveyed 1,340 practitioners between November and December 2025 and published the results in June 2026. Among teams with agents in production, 94% had some observability and 71.5% had full tracing. Across all respondents, only 52.4% ran offline evaluations and 37.3% ran online evaluations. Traces tell you what happened. Evals tell you whether it was acceptable. A regulated product needs both, written down. We covered the general eval-versus-observability split in AI agent evaluation and observability. This post is about the regulated layer on top.
Few laws name LLMs. Many impose testing, logging and accountability duties that apply to any system handling the data or making the decision. Here is what applies as of October 2026:

Start a regulated AI build from a matrix like this one. Each row ends in an artefact, because "we tested it" is not evidence. A dated, versioned report is.
| Regime | What the reviewer asks | Test type | Evidence artefact |
|---|---|---|---|
| HIPAA Security Rule | Can PHI leave the covered boundary? Who accessed what? | PHI-boundary tests with synthetic canary records; redaction tests before every model call; access tests per role | Data-flow diagram, BAA inventory, test run log showing zero canary leaks, audit-log samples |
| PCI DSS v4.0.1 | Is card data out of the model's scope? | PAN detectors on prompts, outputs, traces and eval datasets; injection attempts that try to make the model echo card data | Scope diagram, detector results per release, log-retention and access configuration |
| EU AI Act Art. 50 (live) | Are users told they are talking to AI? Is generated content marked? | UI and API tests asserting the disclosure and marking on every channel | Screenshot or contract tests per release, change log |
| EU AI Act high-risk (Arts. 9, 12, 15; Annex III from 2 Dec 2027) | Tested against predefined metrics? Events logged? Accuracy declared? | Golden-set accuracy with thresholds; robustness and adversarial suites; log-completeness tests | Test plan with metrics and thresholds, technical documentation, retained logs, declared accuracy in instructions for use |
| US bank model risk (SR 26-2 principles; GenAI governed by bank policy) | Is the model fit for purpose, independently challenged and monitored? | Conceptual-soundness review, benchmark comparison, outcome analysis, ongoing drift monitoring | Validation report, monitoring dashboard exports, issue log with remediation dates |
| FinCEN AML/CFT (proposed) | Does the AI make the programme more effective? | Alert precision and recall versus the prior rules engine on labelled cases; reviewer-override rates | Before/after metrics report, sampled case reviews |
| Any regime, agents | What could the agent do, and what did it do? | Tool-permission tests (denied paths), trajectory evals, prompt-injection and tool-poisoning suites | Allowlist config, per-action audit log, red-team report |

Seven things, kept in version control next to the code. If any of them lives only in someone's notebook, it will be missing on audit day.
Think of it as a loop with four stations, where the last one feeds the first:
A minimal, tool-agnostic plan file looks like this:
suite: claims-summary-v3
owner: compliance-eng
regimes: [hipaa, eu-ai-act-art50]
dataset: datasets/claims-golden@2026-10-01 # synthetic, expert-labelled
gates:
- metric: critical_field_accuracy
threshold: ">= 0.97"
- metric: unsupported_claim_rate
threshold: "<= 0.01"
- metric: phi_canary_leaks
threshold: "== 0"
- metric: ai_disclosure_present
threshold: "== 1.0"
adversarial: [prompt_injection_v5, exfiltration_v2, tool_poisoning_v1]
online:
sample_rate: 0.05
human_review: "confidence < 0.8 or action in [deny, escalate]"
evidence_out: reports/{suite}/{git_sha}.json
Every run writes a report keyed to the Git commit, so you can answer "what was tested before version X shipped?" in one lookup.

For agents, the risky part is the action, so test the path the agent took as well as the final text. Three kinds of tests matter:
Those tests only count as evidence if the runtime records every action. In Kite, our open-source agent framework, the model only proposes actions. A kernel checks each one against a tool allowlist, a budget and a policy, then executes it or rejects and logs it, and tracing can write every event to a JSON file. On the server side, the EcoCheck MCP server we run writes one audit line per agent request (subject, method, path, allowed or denied). The design is covered in our MCP server development guide. Logs like these turn a red-team finding into a reproducible test.
An evidence pack, not a dashboard login. Have these ready before anyone asks:
For healthcare-specific architecture (PHI masking, BAA chain, tamper-evident logs), see HIPAA-compliant AI agents.
The usual failure is not a bad score. It is a missing record. These are the patterns to avoid:
You test the system, not the model. Seed synthetic PHI canaries and assert they never reach any endpoint not covered by a BAA. Test redaction before every model call, check access controls by role, and keep audit logs of who saw what. Store the test results per release as evidence for your risk analysis.
For high-risk systems, yes. From 2 December 2027 (Annex III), providers must test against predefined metrics, log events automatically and declare accuracy. For every AI system that talks to people, the Article 50 transparency duties already apply, and you should test that the disclosure appears on every channel.
No. SR 11-7 was replaced on 17 April 2026 by revised interagency model risk guidance (SR 26-2). The new guidance explicitly puts generative and agentic AI out of its scope, but tells banks to apply their own risk management and governance to those systems. Expect validation-style evidence to be requested anyway.
Offline evals run a fixed, labelled dataset before release, so you can compare versions like for like. Online evals score a sample of real production traffic after release, which catches drift and new kinds of input. Regulated teams need both, plus human review of the cases the scorers flag.
Yes, as one layer, if you calibrate it. Measure the grader's agreement with expert labels on a held-out set, record that agreement in the eval plan, and keep deterministic checks and expert sampling for critical fields. Re-calibrate whenever the grader model changes.
BeevR builds AI for healthcare, fintech and other audited domains with the eval plan, logs and human checkpoints designed in from the first sprint, for a fixed price per phase and with full code ownership. See how we approach AI agent development and HIPAA-compliant AI agents, or tell us what your auditor is asking for.