From Sidhant Tamrkar | Product & Market Analysis
HIPAA, DORA and AI Agents: Controls You Have to Evidence, Not Implement
On this page
Only 6.5% of the registers filed in the European supervisors' DORA dry run passed every data quality check. HIPAA AI projects and DORA compliance work fail at the same seam, and it is not the seam vendors sell against. Most of those firms had the controls. What they could not produce, on demand and in machine-readable form, was the proof.
Key takeaways
- The evidence fails long before the control does. Of 947 registers analysed in the European Supervisory Authorities' DORA dry run, 6.5% passed all 116 data quality checks, and missing mandatory information accounted for 86% of every error found.
- Your AI vendor is already a regulated third party. An AI provider serving a financial entity is an ICT third-party service provider under DORA, so the contract terms and the register entry apply from the first production call.
- Healthcare cannot yet assemble the record. 22% of hospital leaders surveyed by Black Book Research in late 2025 said they were highly confident of producing a complete AI audit trail within 30 days.
- The 2026 delays cut one way only. The EU deferred high-risk AI Act duties to December 2027, but DORA has applied since January 2025 and the existing HIPAA Security Rule is being enforced now.
What "evidenced" means when a regulator asks
A control is implemented when it works. It is evidenced when you can hand someone a dated artefact proving it worked, for a named system, without asking the vendor for help.
Generic AI governance advice stops at the first definition. Sector regulators start at the second. That gap is why compliant-looking programmes fail examinations they were, in substance, ready to pass.
Three properties separate the two. An evidenced control is dated, attributable to a person or a system, and reproducible next quarter. I would treat any AI control you cannot export as a file as not existing. That sounds harsh until you watch a team try to rebuild a 90 day access log from a chat interface.
HIPAA: the agent creates ePHI you did not plan for
Healthcare teams scope an AI project around the output. The regulator scopes it around the data touched along the way, which is a much larger set.
An ambient documentation agent is the clearest case. In one encounter it holds a live audio stream, an interim transcript, a machine-drafted note and metadata about clinician, patient and visit. Each of those is electronic protected health information. Each needs safeguards, retention rules and a place in your inventory.
The proposed Security Rule turns judgement into a schedule
The HIPAA Security Rule has been mostly unchanged since 2013. The Office for Civil Rights published a proposed overhaul on 6 January 2025, and the detail matters even though the rule is not final. It would remove the addressable and required distinction, making specifications mandatory with limited exceptions. It would require a technology asset inventory and network map maintained at least once every 12 months and on change.
The rest is a calendar. Multi-factor authentication and encryption of ePHI at rest and in transit, both with limited exceptions. Vulnerability scanning every six months. Penetration testing and a compliance audit every 12 months. A 72 hour target to restore certain systems after a loss. Notification no later than 24 hours after activating a contingency plan.
The rule is not close to final. HHS moved it to the long-term actions section of its regulatory agenda in July 2026, naming July 2027 as the target, after more than 4,000 comments. My position is to build to the proposed schedule anyway. It is the cheapest available specification of what a healthcare examiner will eventually want, and none of it is wasted if the rule changes shape.
The business associate chain now includes the model provider
A signed business associate agreement must exist before the first request carrying patient data. There is no grace period and no pilot exemption. Agent architectures make the chain longer than the contract usually admits. Your scribe vendor calls a model provider, which runs in a region you never named. An observability tool captures prompts in between. Every hop is a subprocessor and belongs in the paperwork covered in the guide to AI clauses in data processing agreements.
Two clauses are worth fighting for. First, an explicit ban on using your data to train or fine-tune models, stated for every tier of the chain. Second, a right to receive the subprocessor list on change rather than on request. The proposed rule would go further and require written verification of a business associate's safeguards, certified by a subject matter expert, every 12 months.
What OCR is actually penalising
Not model choice. Not prompt design. Risk analysis. The risk analysis initiative had produced 12 enforcement actions by early March 2026. The alleged failure is nearly always the same: no accurate and thorough assessment of risks to ePHI across all systems.
Agents get missed here for a structural reason. They are bought as a feature of an existing platform rather than as a new system, so nobody opens a new risk analysis entry. The vendor call becomes a data flow the last assessment never contemplated.
The fix takes an afternoon. Add the agent to the asset inventory as its own line, with its own data flow, review date and owner. Do that before production, and use the procurement checklist for AI agents to catch the questions that only have answers before signature.
DORA: your AI vendor is an ICT third-party provider
Financial entities keep asking when an AI-specific rule will arrive. For operational resilience it already has, and it is not called an AI rule.
DORA has applied in full since 17 January 2025. It regulates the financial entity, not the vendor, but it reaches the vendor through contract. A model provider, an agent platform or an evaluation tool supplying a regulated firm is an ICT third-party service provider, and the obligations follow automatically.
Article 30 splits your contract book in two
Every ICT contract needs a baseline set of terms under Article 30(2) of DORA. A full description of the services and any subcontracting. The locations where services run and data is processed. Data availability and integrity provisions, and data return on termination. Service levels, incident assistance, cooperation with authorities and termination rights.
Contracts supporting a critical or important function take a heavier second layer under Article 30(3). Quantitative and qualitative performance targets. Notice of developments that could affect delivery. Contingency measures and security standards. Participation in threat-led penetration testing. Unrestricted rights of access, inspection and audit. An exit strategy with a transition period long enough to actually move.
The audit rights clause is the one term I would not trade, whatever the vendor says about model confidentiality. Everything else on that list can be satisfied by a report somebody else wrote. That clause is the only one that lets you check.
The register is where the paperwork fails
Article 28 requires a register of information covering every contractual arrangement for ICT services. It is filed in machine-readable form, and supervisors use it to map concentration risk across the EU financial sector.
Read the dry run result as a warning about that filing, not about resilience. 6.5% of analysed registers passed all 116 checks, and 86% of the errors were missing mandatory information. These are firms with real vendor management functions failing on fields, not on judgement.
The concentration point is now concrete. On 18 November 2025 the European supervisors designated the first critical ICT third-party providers, a list of 19 spanning core infrastructure through to business and data services. If your agent runs on a designated provider, supervisors already see that dependency in aggregate. The questions in the analysis of AI data residency across the EU and India stop being theoretical at that point.
The four hour clock does not care that a model caused it
Once an incident is classified as major, DORA starts a fixed cascade. Initial notification within 4 hours of classification, and no later than 24 hours after detection. An intermediate report within 72 hours. A final report within one month.
Agent incidents are hard to classify quickly, which is the practical risk. A server outage is obvious. An agent producing subtly wrong output at scale for nine days is a judgement call, and the clock starts when you make it. Write the classification thresholds down before you need them. Rehearse them against the detection tooling argued for in the piece on monitoring agent drift, and against the escalation path in the agent runbook template.
The control map, side by side
Most of the underlying work is shared. What differs is the artefact each regulator expects, and teams that build one artefact per regulator do the work twice.
| Control area | HIPAA evidence | DORA evidence |
|---|---|---|
| System inventory | Asset inventory and network map showing the agent and its data flows | Register entry per arrangement, machine readable, with criticality |
| Risk assessment | Risk analysis naming the agent, threats to ePHI and assigned risk levels | ICT risk framework covering the full agent lifecycle |
| Third-party contract | Business associate agreement, subprocessors, training ban, verification | Article 30(2) terms, plus Article 30(3) where the function is critical |
| Access control | Unique IDs, role-based access, periodic reviews of human and service accounts | Identity and access controls evidenced per critical function |
| Incident handling | Contingency plan, 24 hour activation notice as proposed, breach assessment | Classification record, then 4 hour, 72 hour and one month reports |
| Testing | Scans every six months, penetration test every 12 months, as proposed | Resilience testing, including threat-led testing where in scope |
| Human oversight | Workforce procedures and sanction policy covering agent-assisted work | Governance with a management body carrying final responsibility |
The HIPAA column mixes current requirements with items from the January 2025 proposed rule, marked as proposed. The DORA column reflects obligations in force since 17 January 2025. A working map, not legal advice.
Read the rows, not the columns. Six of the seven can be satisfied by one artefact each, produced once and formatted twice. The row that genuinely differs is incident handling, because the DORA cascade has no HIPAA equivalent and its deadlines are shorter than most internal escalation paths.
The artefacts that survive an audit
An examiner does not read your governance deck. They ask for a small number of records and check whether they are current, complete and consistent with each other.
| Artefact | What it must contain | Refresh | Weak version |
|---|---|---|---|
| Agent inventory entry | Owner, purpose, data classes, systems called, criticality | On change, yearly | A row in a software list with no data classes |
| Data flow record | Every hop, including model provider, region, retention, deletion | On architecture change | A design-time diagram never revised |
| Model and prompt log | Model identifier, version, system prompt hash, effective dates | Continuous | Whatever the vendor console shows today |
| Human review record | What was reviewed, by whom, what was overridden and why | Continuous | An assurance that a human is in the loop |
| Vendor file | Contract, subprocessors, attestations, exit plan, audit rights | At renewal and on change | An order form and a security marketing page |
| Test evidence | Dates, scope, findings, remediation and retest | To the regulatory calendar | A vendor certificate covering the platform |
| Incident timeline | Detection, classification, decision maker, reports with timestamps | Per incident | A ticket thread nobody has turned into a narrative |
Cadences reflect the proposed HIPAA schedule and DORA practice. Your competent authority or auditor may expect a shorter cycle for critical functions.
The model and prompt log is the one most teams skip, and the one that makes every other record checkable. Without it, an incident timeline cannot say which system was running when the failure happened. The staged view in the agent governance maturity model is a reasonable way to place yourself honestly.
What 2026 changed, and what it did not
The EU delayed its AI-specific rules. Under the digital omnibus agreed by the Council and Parliament in May 2026, obligations for standalone Annex III high-risk systems moved from 2 August 2026 to 2 December 2027. Annex I systems moved to 2 August 2028. Creditworthiness assessment and life and health insurance pricing sit in Annex III, so this lands squarely on financial services. The Article 5 prohibitions, the Article 4 literacy duty and the Article 50 transparency rules were not deferred, and those are unpacked in the EU AI Act transparency checklist.
The United States delayed its healthcare security rule, as described above. Germany's supervisor moved the other way. BaFin published guidance on ICT risks in the use of artificial intelligence on 30 January 2026, aimed at banks under the capital requirements regime and insurers under Solvency II. It is formally non-binding. It says AI systems belong inside existing ICT risk, testing and third-party frameworks across the full lifecycle, and supervisors tend to examine against their own published expectations.
So the AI-specific rules slipped and the sector rules did not move at all. The delay is not relief, and treating it as relief is the mistake I expect most teams to make this year. Every obligation in this post was already in force before the omnibus, and none of it was touched by the deferral.
Where this argument is weakest
Three honest problems with everything above.
A control map is not a legal opinion
The table above is an operating aid built from published regulatory text and supervisory material. It is not advice, and it cannot be. Scope turns on facts a general map cannot see. Whether a function is critical or important. Whether an entity is in scope for threat-led testing. Whether a use case is high risk under Annex III. The most common error I see is a team mapping controls correctly and misclassifying the function they sit under.
One of the three headline figures is a small sample
The 22% audit trail figure comes from a market research survey of 182 US hospital leaders, fielded between 15 October and 8 November 2025. It measures self-reported confidence, not tested capability. It is not peer reviewed and not a regulatory count. I use it because it points the same way as the DORA dry run, which is a supervisory measurement of nearly 1,000 entities. Directionally consistent, not equal in weight.
Evidence discipline has a real cost
Every artefact in the second table costs engineering time and creates a maintenance obligation. For a low-risk internal agent that never touches regulated data, a full evidence pack is overhead, and I would not build one. The competing view is defensible: over-documenting early agents slows the learning that makes later ones safe. What tips the balance is scope creep, because internal tools acquire regulated data quietly and a retrospective build costs far more. If the agent could plausibly touch patient or customer financial data within two quarters, build the record now.
A 30 day sequence for one agent
Pick your highest-exposure production agent. Not the portfolio, one agent. Doing this properly once teaches you more than a programme plan for twenty.
Days 1 to 5. Write the inventory entry and the data flow record, covering every hop, region and retention period. Expect to find at least one undocumented subprocessor, because that is what happens every time.
Days 6 to 12. Pull the vendor file. Check the contract against the Article 30(3) list if DORA applies, or against business associate requirements if HIPAA does. Record the gaps as gaps rather than fixing them yet.
Days 13 to 20. Turn on the model and prompt log if it does not exist, and start the human review record. These are engineering work and take the longest. Nothing downstream is checkable without them.
Days 21 to 30. Run one tabletop exercise. Assume the agent has been producing wrong output for nine days. Time how long classification takes and whether you could file within 4 hours of it. That number is your real readiness, and it is usually a surprise.
The one test worth running
Ask for a 90 day access and version history for one production agent, in writing, with a 48 hour deadline. Whatever comes back is your evidence position. If nothing comes back, you have found your first project.
Frequently asked questions
Does HIPAA apply to AI agents?
Yes, whenever the agent creates, receives, maintains or transmits electronic protected health information. An ambient documentation agent generates several artefacts in one encounter: the audio stream, an interim transcript, a machine-drafted note and visit metadata. Each of those is ePHI that you are responsible for safeguarding. The vendor is a business associate, so a signed agreement has to be in place before the first request carrying patient data is sent.
Does DORA apply to AI vendors?
DORA regulates financial entities rather than AI vendors directly, but it reaches vendors through contract. An AI or large language model provider supplying a financial entity is an ICT third-party service provider, so Articles 28 to 30 apply. The arrangement must appear in your register of information. If it supports a critical or important function, the contract also needs the extra terms in Article 30(3), including audit rights and a workable exit plan.
What evidence do regulators ask for on AI systems?
Artefacts carrying dates, owners and versions. In practice that means an inventory entry for the agent, a data flow record and the risk assessment covering it. It also means the contract and subprocessor list, model and prompt version history, access and human review logs, and incident timelines. A policy document on its own is not evidence. Neither is a screenshot from a vendor dashboard that you cannot export or reproduce next quarter.
Has the EU AI Act high-risk deadline been delayed?
Yes. Under the digital omnibus agreed by the Council and the European Parliament in May 2026, obligations for standalone Annex III high-risk systems moved from 2 August 2026 to 2 December 2027. Annex I systems moved to 2 August 2028. The prohibitions in Article 5, the AI literacy duty in Article 4 and the Article 50 transparency rules were not deferred alongside them.
When will the new HIPAA Security Rule take effect?
Not soon. The proposed rule was published on 6 January 2025 and drew more than 4,000 comments. HHS then moved it to the long-term actions section of its regulatory agenda in July 2026, naming July 2027 as the target for final action. That is a target rather than a commitment. The existing Security Rule still applies in full, and the Office for Civil Rights is actively enforcing it.
How do you report an AI incident under DORA?
Through the same route as any other major ICT incident. Once you classify the incident as major, the initial notification is due within 4 hours, and no later than 24 hours after you detected it. An intermediate report follows within 72 hours and a final report within one month. The cause being a model rather than a server changes nothing at all about the clock.
Where to start this week
Start with the artefact you are least confident about, which for most teams is the version log. Open your production agent's configuration and ask one question: can you state, with a timestamp, which model version and which system prompt were live 30 days ago? If the answer is no, that is the first build, and it takes a sprint rather than a quarter.
Then put one line on the agenda of your next risk committee. Name the agent, name its criticality rating, and name the person who owns its record. Most regulated AI programmes have neither of those names written down anywhere, and the distance between an agent nobody owns and a finding nobody expected is short.
References
- HHS Office for Civil Rights, HIPAA Security Rule NPRM fact sheet, 6 January 2025. Used for every proposed requirement and cadence.
- HHS Office for Civil Rights, Resolution agreements and civil money penalties, accessed 31 August 2026. Used for the 2026 enforcement record.
- Clark Hill, HIPAA Security Rule update delayed until 2027, 2026. Used for the Unified Agenda move and the July 2027 target.
- European Supervisory Authorities, DORA dry run exercise summary report, December 2024. Used for the 6.5%, 116 checks and 86% figures.
- EBA, EIOPA and ESMA, Designation of critical ICT third-party providers, 18 November 2025. Used for the CTPP list.
- EUR-Lex, Regulation (EU) 2022/2554 (DORA). Used for the Article 26, 28 and 30 obligations.
- Gibson Dunn, EU AI Act omnibus agreement, postponed high-risk deadlines, May 2026. Used for the deferral dates.
- Black Book Research, reported by Becker's Hospital Review, December 2025. Survey of 182 US hospital leaders, 15 October to 8 November 2025.
The weakest source here is the Black Book survey: 182 self-reported responses from a market research firm, neither peer reviewed nor a regulatory count, so the 22% is directional only. BaFin's January 2026 guidance is cited in the body and is expressly non-binding. Every other figure traces to a regulator's own publication.
Related reading