← Blog
Security

LLM Testing for Regulated Industries: An Audit-Ready Eval Plan

Thien Nguyen · Oct 6, 2026

LLM testing in a regulated industry passes an audit when every rule you are subject to maps to a named test, and every test leaves an evidence artefact that someone outside the team can check. Accuracy scores alone don't do that. In practice, an audit-ready plan has four parts: a versioned golden dataset labelled by domain experts, regression gates in CI on every model or prompt change, adversarial and data-boundary tests, and scored samples of live traffic with human review. Most teams are not there yet. In LangChain's 2026 State of Agent Engineering survey, 94% of teams with agents in production had observability, but only 37.3% of all respondents ran online evaluations. This guide maps HIPAA, PCI DSS, the EU AI Act and US banking guidance to concrete tests and artefacts. Updated October 2026.

Why do regulated teams need more than a generic LLM eval suite?

A generic eval suite asks whether the model is good. An auditor asks a different question: show me the control, show me it ran, and show me what happened when it failed. Those questions need evidence that holds up after the fact: which model version answered, what data it saw, which test set it passed, who reviewed the exceptions, and what changed afterwards.

The gap is real. LangChain surveyed 1,340 practitioners between November and December 2025 and published the results in June 2026. Among teams with agents in production, 94% had some observability and 71.5% had full tracing. Across all respondents, only 52.4% ran offline evaluations and 37.3% ran online evaluations. Traces tell you what happened. Evals tell you whether it was acceptable. A regulated product needs both, written down. We covered the general eval-versus-observability split in AI agent evaluation and observability. This post is about the regulated layer on top.

Which regulations actually require testing of LLMs and agents in 2026?

Few laws name LLMs. Many impose testing, logging and accountability duties that apply to any system handling the data or making the decision. Here is what applies as of October 2026:

  • EU AI Act. Article 50 transparency duties have applied since 2 August 2026: people must be told when they are interacting with an AI system. The machine-readable marking of synthetic content under Article 50(2) moved to 2 December 2026 under the AI Omnibus, Regulation (EU) 2026/1744. The same Omnibus moved the Annex III high-risk obligations to 2 December 2027. Those obligations include risk management with testing against predefined metrics (Article 9), automatic event logging (Article 12), and declared accuracy, robustness and cybersecurity (Article 15).
  • HIPAA. The Security Rule does not mention models. It does require a risk analysis and audit controls for any system that touches ePHI, and an LLM pipeline is such a system. The testable question is whether PHI stays inside the boundary covered by your business associate agreements. See the HIPAA Security Rule in 2026 and AI.
  • PCI DSS v4.0.1. Any component that stores, processes or transmits cardholder data, or can affect its security, is in scope. Prompts, traces and eval datasets count. The testable question is whether a primary account number (PAN) can ever reach the model, its logs or your eval store. Details are in our fintech PCI DSS and SOC 2 checklist.
  • US bank model risk. On 17 April 2026, the Fed, OCC and FDIC replaced SR 11-7 with revised model risk guidance (SR 26-2). The new guidance applies to traditional and non-generative AI models but explicitly puts generative and agentic AI out of its scope, while telling banks to use their risk management and governance practices to decide controls for such systems. In practice, a bank's model risk team will still ask for validation evidence. It just won't come from a checklist written for credit scorecards.
  • AML programmes. FinCEN's April 2026 proposed rule would judge AML/CFT programmes on effectiveness, and lists "effective use of artificial intelligence" among the activities that can produce demonstrable outputs of effectiveness. It is a proposal (comments closed 9 June 2026), but it signals that AI monitoring will be expected to come with metrics.
  • NIST AI RMF. This framework is voluntary, but its MEASURE function and the Generative AI Profile (NIST AI 600-1) give you a vocabulary auditors recognise for test, evaluation, verification and validation.
A clinician reviewing AI analytics in a regulated setting
A clinician reviewing AI analytics in a regulated setting

Which tests map to which regulation?

Start a regulated AI build from a matrix like this one. Each row ends in an artefact, because "we tested it" is not evidence. A dated, versioned report is.

RegimeWhat the reviewer asksTest typeEvidence artefact
HIPAA Security RuleCan PHI leave the covered boundary? Who accessed what?PHI-boundary tests with synthetic canary records; redaction tests before every model call; access tests per roleData-flow diagram, BAA inventory, test run log showing zero canary leaks, audit-log samples
PCI DSS v4.0.1Is card data out of the model's scope?PAN detectors on prompts, outputs, traces and eval datasets; injection attempts that try to make the model echo card dataScope diagram, detector results per release, log-retention and access configuration
EU AI Act Art. 50 (live)Are users told they are talking to AI? Is generated content marked?UI and API tests asserting the disclosure and marking on every channelScreenshot or contract tests per release, change log
EU AI Act high-risk (Arts. 9, 12, 15; Annex III from 2 Dec 2027)Tested against predefined metrics? Events logged? Accuracy declared?Golden-set accuracy with thresholds; robustness and adversarial suites; log-completeness testsTest plan with metrics and thresholds, technical documentation, retained logs, declared accuracy in instructions for use
US bank model risk (SR 26-2 principles; GenAI governed by bank policy)Is the model fit for purpose, independently challenged and monitored?Conceptual-soundness review, benchmark comparison, outcome analysis, ongoing drift monitoringValidation report, monitoring dashboard exports, issue log with remediation dates
FinCEN AML/CFT (proposed)Does the AI make the programme more effective?Alert precision and recall versus the prior rules engine on labelled cases; reviewer-override ratesBefore/after metrics report, sampled case reviews
Any regime, agentsWhat could the agent do, and what did it do?Tool-permission tests (denied paths), trajectory evals, prompt-injection and tool-poisoning suitesAllowlist config, per-action audit log, red-team report
Matrix mapping HIPAA, PCI DSS v4.0.1, EU AI Act Article 50 and high-risk duties, US bank model risk (SR 26-2), FinCEN AML/CFT and AI agents to the test type and evidence artefact each needs.
Figure 2: Matrix mapping HIPAA, PCI DSS v4.0.1, EU AI Act Article 50 and high-risk duties, US bank model risk (SR 26-2), FinCEN AML/CFT and AI agents to the test type and evidence artefact each needs.

What should an audit-ready LLM eval plan contain?

Seven things, kept in version control next to the code. If any of them lives only in someone's notebook, it will be missing on audit day.

  • Scope and risk statement. What the system decides or drafts, who is affected, the worst credible failure, and which regimes apply.
  • Golden dataset. Representative cases labelled by domain experts (clinicians, compliance analysts, underwriters), with synthetic or de-identified data only. Version it and record who labelled what.
  • Metrics with thresholds set in advance. Accuracy on critical fields, unsupported-claim rate, refusal correctness, PHI/PAN leak count (target zero), cost and latency. Thresholds are agreed before the test runs, not after.
  • Adversarial suites. Prompt injection, jailbreaks, data exfiltration attempts and, for agents, poisoned tool descriptions and outputs.
  • Release gates. Which suites must pass for a model, prompt, retrieval or tool change to ship, and who can override a failed gate (named role, logged reason).
  • Production monitoring. Sampling rate, online scorers, a human review queue, and the triggers that page someone.
  • Feedback loop. Every production failure becomes a new test case, so the same failure cannot ship twice.

What does the eval pipeline look like from offline tests to production?

Think of it as a loop with four stations, where the last one feeds the first:

  1. Offline. Run the golden set and adversarial suites against the candidate model, prompt and retrieval configuration. Score with deterministic checks first (schema, citations present, detectors for PHI and PAN), then model-graded rubrics calibrated against expert labels, then expert review of a sample.
  2. CI gate. The same suites run on every pull request that touches a prompt, model ID, retrieval setting or tool definition. A threshold miss blocks the merge.
  3. Pre-release. A red-team pass on anything that changes what the system can do, such as a new tool, a new data source or a new user group.
  4. Online. Score a sample of live traffic with the same rubrics, send low-confidence and high-impact cases to human review, and track drift. Failures go back into the golden set.

A minimal, tool-agnostic plan file looks like this:

suite: claims-summary-v3
owner: compliance-eng
regimes: [hipaa, eu-ai-act-art50]
dataset: datasets/claims-golden@2026-10-01   # synthetic, expert-labelled
gates:
  - metric: critical_field_accuracy
    threshold: ">= 0.97"
  - metric: unsupported_claim_rate
    threshold: "<= 0.01"
  - metric: phi_canary_leaks
    threshold: "== 0"
  - metric: ai_disclosure_present
    threshold: "== 1.0"
adversarial: [prompt_injection_v5, exfiltration_v2, tool_poisoning_v1]
online:
  sample_rate: 0.05
  human_review: "confidence < 0.8 or action in [deny, escalate]"
evidence_out: reports/{suite}/{git_sha}.json

Every run writes a report keyed to the Git commit, so you can answer "what was tested before version X shipped?" in one lookup.

Eval pipeline loop for regulated LLM products: offline golden-set and adversarial tests, CI gate that blocks merges on threshold misses, pre-release red team, and online scoring of live traffic with human review, with failures fed back into the golden set.
Figure 1: Eval pipeline loop for regulated LLM products: offline golden-set and adversarial tests, CI gate that blocks merges on threshold misses, pre-release red team, and online scoring of live traffic with human review, with failures fed back into the golden set.

How do you test an agent's tool calls, not just its answers?

For agents, the risky part is the action, so test the path the agent took as well as the final text. Three kinds of tests matter:

  • Permission tests. For every tool, assert what must be denied: wrong role, another tenant's ID, a write method on a read-only scope, an oversized input. These are deterministic and belong in CI.
  • Trajectory evals. Check the sequence of tool calls against expected paths. Did the agent look up the policy before drafting the denial? Did it ask for approval before sending?
  • Injection and tool-poisoning tests. Plant instructions in retrieved documents and tool outputs, and assert that the agent ignores them. Published benchmarks show why this matters: in MCPTox, poisoned tool descriptions succeeded 36.5% of the time on average.

Those tests only count as evidence if the runtime records every action. In Kite, our open-source agent framework, the model only proposes actions. A kernel checks each one against a tool allowlist, a budget and a policy, then executes it or rejects and logs it, and tracing can write every event to a JSON file. On the server side, the EcoCheck MCP server we run writes one audit line per agent request (subject, method, path, allowed or denied). The design is covered in our MCP server development guide. Logs like these turn a red-team finding into a reproducible test.

What evidence do auditors actually want to see?

An evidence pack, not a dashboard login. Have these ready before anyone asks:

  • The eval plan and the regulation-to-test matrix, with an owner and a date.
  • Dataset cards: source, de-identification method, labellers, version history.
  • Release reports for the last N versions: what ran, thresholds, results, overrides and who approved them.
  • A model and vendor inventory: model IDs and versions, hosting region, data-processing terms (BAA or DPA), retention settings.
  • Production monitoring: sampled reviews, drift charts, incidents and the test cases each incident created.
  • Human oversight records: who reviews what, at what threshold, and how overrides are captured. Our human-in-the-loop guide covers where to put the human.
  • Access and action logs for agents, with retention matching your regime.

For healthcare-specific architecture (PHI masking, BAA chain, tamper-evident logs), see HIPAA-compliant AI agents.

What mistakes make LLM evals fail an audit?

The usual failure is not a bad score. It is a missing record. These are the patterns to avoid:

  • Thresholds set after the run. If you pick the bar once you see the number, it isn't a control.
  • Real patient or card data in eval sets. This puts your test tooling in HIPAA or PCI scope. Use synthetic or properly de-identified data.
  • LLM-as-judge with no calibration. A model grader is fine if you have measured its agreement with expert labels. Without that, it is an opinion.
  • Silent model upgrades. A provider alias that moves to a new model version is a change. Pin versions and run the gate.
  • Traces without retention rules. Logs full of prompts are themselves sensitive data. Decide retention, access and redaction up front.

FAQ

How do you test an LLM for HIPAA compliance?

You test the system, not the model. Seed synthetic PHI canaries and assert they never reach any endpoint not covered by a BAA. Test redaction before every model call, check access controls by role, and keep audit logs of who saw what. Store the test results per release as evidence for your risk analysis.

Does the EU AI Act require LLM testing?

For high-risk systems, yes. From 2 December 2027 (Annex III), providers must test against predefined metrics, log events automatically and declare accuracy. For every AI system that talks to people, the Article 50 transparency duties already apply, and you should test that the disclosure appears on every channel.

Does SR 11-7 still apply to generative AI?

No. SR 11-7 was replaced on 17 April 2026 by revised interagency model risk guidance (SR 26-2). The new guidance explicitly puts generative and agentic AI out of its scope, but tells banks to apply their own risk management and governance to those systems. Expect validation-style evidence to be requested anyway.

What is the difference between offline and online evals?

Offline evals run a fixed, labelled dataset before release, so you can compare versions like for like. Online evals score a sample of real production traffic after release, which catches drift and new kinds of input. Regulated teams need both, plus human review of the cases the scorers flag.

Can I use an LLM to grade another LLM in a regulated setting?

Yes, as one layer, if you calibrate it. Measure the grader's agreement with expert labels on a held-out set, record that agreement in the eval plan, and keep deterministic checks and expert sampling for critical fields. Re-calibrate whenever the grader model changes.

BeevR builds AI for healthcare, fintech and other audited domains with the eval plan, logs and human checkpoints designed in from the first sprint, for a fixed price per phase and with full code ownership. See how we approach AI agent development and HIPAA-compliant AI agents, or tell us what your auditor is asking for.