From Sidhant Tamrkar | Product & Market Analysis

HIPAA, DORA and AI Agents: Controls You Have to Evidence, Not Implement

On this page

Only 6.5% of the registers filed in the European supervisors' DORA dry run passed every data quality check. HIPAA AI projects and DORA compliance work fail at the same seam, and it is not the seam vendors sell against. Most of those firms had the controls. What they could not produce, on demand and in machine-readable form, was the proof.

Key takeaways

  • The evidence fails long before the control does. Of 947 registers analysed in the European Supervisory Authorities' DORA dry run, 6.5% passed all 116 data quality checks, and missing mandatory information accounted for 86% of every error found.
  • Your AI vendor is already a regulated third party. An AI provider serving a financial entity is an ICT third-party service provider under DORA, so the contract terms and the register entry apply from the first production call.
  • Healthcare cannot yet assemble the record. 22% of hospital leaders surveyed by Black Book Research in late 2025 said they were highly confident of producing a complete AI audit trail within 30 days.
  • The 2026 delays cut one way only. The EU deferred high-risk AI Act duties to December 2027, but DORA has applied since January 2025 and the existing HIPAA Security Rule is being enforced now.
6.5%Share of DORA registers that passed every data quality check in the supervisors' dry run. Source: ESAs, December 2024.
22%Hospital leaders highly confident of producing a 30 day AI audit trail. Source: Black Book Research, reported by Becker's, December 2025.
12OCR enforcement actions under its risk analysis initiative, as of March 2026. Source: HHS OCR resolution agreements.

What "evidenced" means when a regulator asks

A control is implemented when it works. It is evidenced when you can hand someone a dated artefact proving it worked, for a named system, without asking the vendor for help.

Generic AI governance advice stops at the first definition. Sector regulators start at the second. That gap is why compliant-looking programmes fail examinations they were, in substance, ready to pass.

Three properties separate the two. An evidenced control is dated, attributable to a person or a system, and reproducible next quarter. I would treat any AI control you cannot export as a file as not existing. That sounds harsh until you watch a team try to rebuild a 90 day access log from a chat interface.

HIPAA: the agent creates ePHI you did not plan for

Healthcare teams scope an AI project around the output. The regulator scopes it around the data touched along the way, which is a much larger set.

An ambient documentation agent is the clearest case. In one encounter it holds a live audio stream, an interim transcript, a machine-drafted note and metadata about clinician, patient and visit. Each of those is electronic protected health information. Each needs safeguards, retention rules and a place in your inventory.

The proposed Security Rule turns judgement into a schedule

The HIPAA Security Rule has been mostly unchanged since 2013. The Office for Civil Rights published a proposed overhaul on 6 January 2025, and the detail matters even though the rule is not final. It would remove the addressable and required distinction, making specifications mandatory with limited exceptions. It would require a technology asset inventory and network map maintained at least once every 12 months and on change.

The rest is a calendar. Multi-factor authentication and encryption of ePHI at rest and in transit, both with limited exceptions. Vulnerability scanning every six months. Penetration testing and a compliance audit every 12 months. A 72 hour target to restore certain systems after a loss. Notification no later than 24 hours after activating a contingency plan.

The rule is not close to final. HHS moved it to the long-term actions section of its regulatory agenda in July 2026, naming July 2027 as the target, after more than 4,000 comments. My position is to build to the proposed schedule anyway. It is the cheapest available specification of what a healthcare examiner will eventually want, and none of it is wasted if the rule changes shape.

The business associate chain now includes the model provider

A signed business associate agreement must exist before the first request carrying patient data. There is no grace period and no pilot exemption. Agent architectures make the chain longer than the contract usually admits. Your scribe vendor calls a model provider, which runs in a region you never named. An observability tool captures prompts in between. Every hop is a subprocessor and belongs in the paperwork covered in the guide to AI clauses in data processing agreements.

Two clauses are worth fighting for. First, an explicit ban on using your data to train or fine-tune models, stated for every tier of the chain. Second, a right to receive the subprocessor list on change rather than on request. The proposed rule would go further and require written verification of a business associate's safeguards, certified by a subject matter expert, every 12 months.

What OCR is actually penalising

Not model choice. Not prompt design. Risk analysis. The risk analysis initiative had produced 12 enforcement actions by early March 2026. The alleged failure is nearly always the same: no accurate and thorough assessment of risks to ePHI across all systems.

Agents get missed here for a structural reason. They are bought as a feature of an existing platform rather than as a new system, so nobody opens a new risk analysis entry. The vendor call becomes a data flow the last assessment never contemplated.

The fix takes an afternoon. Add the agent to the asset inventory as its own line, with its own data flow, review date and owner. Do that before production, and use the procurement checklist for AI agents to catch the questions that only have answers before signature.

Nearly everyone had the controls. Almost nobody had the record. European supervisors' DORA dry run: 947 registers scored against 116 data quality checks 6.5% passed all 116 checks 93.5% failed at least one check Where the errors came from 86% missing mandatory information 14% other 1,039 entities took part voluntarily, before reporting became mandatory. Source: ESAs dry run summary report, December 2024.
Notice what the failure was. Not a missing control, a missing field. That is the shape of most regulated AI findings.

DORA: your AI vendor is an ICT third-party provider

Financial entities keep asking when an AI-specific rule will arrive. For operational resilience it already has, and it is not called an AI rule.

DORA has applied in full since 17 January 2025. It regulates the financial entity, not the vendor, but it reaches the vendor through contract. A model provider, an agent platform or an evaluation tool supplying a regulated firm is an ICT third-party service provider, and the obligations follow automatically.

Article 30 splits your contract book in two

Every ICT contract needs a baseline set of terms under Article 30(2) of DORA. A full description of the services and any subcontracting. The locations where services run and data is processed. Data availability and integrity provisions, and data return on termination. Service levels, incident assistance, cooperation with authorities and termination rights.

Contracts supporting a critical or important function take a heavier second layer under Article 30(3). Quantitative and qualitative performance targets. Notice of developments that could affect delivery. Contingency measures and security standards. Participation in threat-led penetration testing. Unrestricted rights of access, inspection and audit. An exit strategy with a transition period long enough to actually move.

The audit rights clause is the one term I would not trade, whatever the vendor says about model confidentiality. Everything else on that list can be satisfied by a report somebody else wrote. That clause is the only one that lets you check.

The register is where the paperwork fails

Article 28 requires a register of information covering every contractual arrangement for ICT services. It is filed in machine-readable form, and supervisors use it to map concentration risk across the EU financial sector.

Read the dry run result as a warning about that filing, not about resilience. 6.5% of analysed registers passed all 116 checks, and 86% of the errors were missing mandatory information. These are firms with real vendor management functions failing on fields, not on judgement.

The concentration point is now concrete. On 18 November 2025 the European supervisors designated the first critical ICT third-party providers, a list of 19 spanning core infrastructure through to business and data services. If your agent runs on a designated provider, supervisors already see that dependency in aggregate. The questions in the analysis of AI data residency across the EU and India stop being theoretical at that point.

The four hour clock does not care that a model caused it

Once an incident is classified as major, DORA starts a fixed cascade. Initial notification within 4 hours of classification, and no later than 24 hours after detection. An intermediate report within 72 hours. A final report within one month.

Agent incidents are hard to classify quickly, which is the practical risk. A server outage is obvious. An agent producing subtly wrong output at scale for nine days is a judgement call, and the clock starts when you make it. Write the classification thresholds down before you need them. Rehearse them against the detection tooling argued for in the piece on monitoring agent drift, and against the escalation path in the agent runbook template.

Two regulators, two clocks, one incident Deadlines that begin the moment a failure is classified or a plan is activated DORA, MAJOR ICT INCIDENT 4h Initial notice 24h Outer limit from detection 72h Intermediate report 1 month Final report HIPAA SECURITY RULE, AS PROPOSED 24h Notice after contingency plan activation 72h Restore certain systems and data Proposed, not yet final.
The DORA clock starts at classification, not detection. Deciding what counts as major is therefore a control in its own right.

The control map, side by side

Most of the underlying work is shared. What differs is the artefact each regulator expects, and teams that build one artefact per regulator do the work twice.

Agent controls and the evidence each regime expects
Control areaHIPAA evidenceDORA evidence
System inventoryAsset inventory and network map showing the agent and its data flowsRegister entry per arrangement, machine readable, with criticality
Risk assessmentRisk analysis naming the agent, threats to ePHI and assigned risk levelsICT risk framework covering the full agent lifecycle
Third-party contractBusiness associate agreement, subprocessors, training ban, verificationArticle 30(2) terms, plus Article 30(3) where the function is critical
Access controlUnique IDs, role-based access, periodic reviews of human and service accountsIdentity and access controls evidenced per critical function
Incident handlingContingency plan, 24 hour activation notice as proposed, breach assessmentClassification record, then 4 hour, 72 hour and one month reports
TestingScans every six months, penetration test every 12 months, as proposedResilience testing, including threat-led testing where in scope
Human oversightWorkforce procedures and sanction policy covering agent-assisted workGovernance with a management body carrying final responsibility

The HIPAA column mixes current requirements with items from the January 2025 proposed rule, marked as proposed. The DORA column reflects obligations in force since 17 January 2025. A working map, not legal advice.

Read the rows, not the columns. Six of the seven can be satisfied by one artefact each, produced once and formatted twice. The row that genuinely differs is incident handling, because the DORA cascade has no HIPAA equivalent and its deadlines are shorter than most internal escalation paths.

The artefacts that survive an audit

An examiner does not read your governance deck. They ask for a small number of records and check whether they are current, complete and consistent with each other.

Seven records to hold for every production agent
ArtefactWhat it must containRefreshWeak version
Agent inventory entryOwner, purpose, data classes, systems called, criticalityOn change, yearlyA row in a software list with no data classes
Data flow recordEvery hop, including model provider, region, retention, deletionOn architecture changeA design-time diagram never revised
Model and prompt logModel identifier, version, system prompt hash, effective datesContinuousWhatever the vendor console shows today
Human review recordWhat was reviewed, by whom, what was overridden and whyContinuousAn assurance that a human is in the loop
Vendor fileContract, subprocessors, attestations, exit plan, audit rightsAt renewal and on changeAn order form and a security marketing page
Test evidenceDates, scope, findings, remediation and retestTo the regulatory calendarA vendor certificate covering the platform
Incident timelineDetection, classification, decision maker, reports with timestampsPer incidentA ticket thread nobody has turned into a narrative

Cadences reflect the proposed HIPAA schedule and DORA practice. Your competent authority or auditor may expect a shorter cycle for critical functions.

The model and prompt log is the one most teams skip, and the one that makes every other record checkable. Without it, an incident timeline cannot say which system was running when the failure happened. The staged view in the agent governance maturity model is a reasonable way to place yourself honestly.

What 2026 changed, and what it did not

The EU delayed its AI-specific rules. Under the digital omnibus agreed by the Council and Parliament in May 2026, obligations for standalone Annex III high-risk systems moved from 2 August 2026 to 2 December 2027. Annex I systems moved to 2 August 2028. Creditworthiness assessment and life and health insurance pricing sit in Annex III, so this lands squarely on financial services. The Article 5 prohibitions, the Article 4 literacy duty and the Article 50 transparency rules were not deferred, and those are unpacked in the EU AI Act transparency checklist.

The United States delayed its healthcare security rule, as described above. Germany's supervisor moved the other way. BaFin published guidance on ICT risks in the use of artificial intelligence on 30 January 2026, aimed at banks under the capital requirements regime and insurers under Solvency II. It is formally non-binding. It says AI systems belong inside existing ICT risk, testing and third-party frameworks across the full lifecycle, and supervisors tend to examine against their own published expectations.

So the AI-specific rules slipped and the sector rules did not move at all. The delay is not relief, and treating it as relief is the mistake I expect most teams to make this year. Every obligation in this post was already in force before the omnibus, and none of it was touched by the deferral.

The sector rules are already live. The AI rules moved. Dates in force or formally scheduled as at 31 August 2026 Deferred by the 2026 digital omnibus Jan 2025 DORA applies in full Nov 2025 19 critical providers named Jan 2026 BaFin AI guidance Aug 2026 AI Act transparency duties Jul 2027 HIPAA rule target Dec 2027 AI Act Annex III Aug 2028 AI Act Annex I Blue markers are obligations already in force. Red markers are scheduled or targeted dates.
The two dates that already bind you are on the left. Planning around the ones on the right is how programmes arrive late to obligations that never moved.

Where this argument is weakest

Three honest problems with everything above.

A control map is not a legal opinion

The table above is an operating aid built from published regulatory text and supervisory material. It is not advice, and it cannot be. Scope turns on facts a general map cannot see. Whether a function is critical or important. Whether an entity is in scope for threat-led testing. Whether a use case is high risk under Annex III. The most common error I see is a team mapping controls correctly and misclassifying the function they sit under.

One of the three headline figures is a small sample

The 22% audit trail figure comes from a market research survey of 182 US hospital leaders, fielded between 15 October and 8 November 2025. It measures self-reported confidence, not tested capability. It is not peer reviewed and not a regulatory count. I use it because it points the same way as the DORA dry run, which is a supervisory measurement of nearly 1,000 entities. Directionally consistent, not equal in weight.

Evidence discipline has a real cost

Every artefact in the second table costs engineering time and creates a maintenance obligation. For a low-risk internal agent that never touches regulated data, a full evidence pack is overhead, and I would not build one. The competing view is defensible: over-documenting early agents slows the learning that makes later ones safe. What tips the balance is scope creep, because internal tools acquire regulated data quietly and a retrospective build costs far more. If the agent could plausibly touch patient or customer financial data within two quarters, build the record now.

A 30 day sequence for one agent

Pick your highest-exposure production agent. Not the portfolio, one agent. Doing this properly once teaches you more than a programme plan for twenty.

Days 1 to 5. Write the inventory entry and the data flow record, covering every hop, region and retention period. Expect to find at least one undocumented subprocessor, because that is what happens every time.

Days 6 to 12. Pull the vendor file. Check the contract against the Article 30(3) list if DORA applies, or against business associate requirements if HIPAA does. Record the gaps as gaps rather than fixing them yet.

Days 13 to 20. Turn on the model and prompt log if it does not exist, and start the human review record. These are engineering work and take the longest. Nothing downstream is checkable without them.

Days 21 to 30. Run one tabletop exercise. Assume the agent has been producing wrong output for nine days. Time how long classification takes and whether you could file within 4 hours of it. That number is your real readiness, and it is usually a surprise.

The one test worth running

Ask for a 90 day access and version history for one production agent, in writing, with a 48 hour deadline. Whatever comes back is your evidence position. If nothing comes back, you have found your first project.

Frequently asked questions

Does HIPAA apply to AI agents?

Yes, whenever the agent creates, receives, maintains or transmits electronic protected health information. An ambient documentation agent generates several artefacts in one encounter: the audio stream, an interim transcript, a machine-drafted note and visit metadata. Each of those is ePHI that you are responsible for safeguarding. The vendor is a business associate, so a signed agreement has to be in place before the first request carrying patient data is sent.

Does DORA apply to AI vendors?

DORA regulates financial entities rather than AI vendors directly, but it reaches vendors through contract. An AI or large language model provider supplying a financial entity is an ICT third-party service provider, so Articles 28 to 30 apply. The arrangement must appear in your register of information. If it supports a critical or important function, the contract also needs the extra terms in Article 30(3), including audit rights and a workable exit plan.

What evidence do regulators ask for on AI systems?

Artefacts carrying dates, owners and versions. In practice that means an inventory entry for the agent, a data flow record and the risk assessment covering it. It also means the contract and subprocessor list, model and prompt version history, access and human review logs, and incident timelines. A policy document on its own is not evidence. Neither is a screenshot from a vendor dashboard that you cannot export or reproduce next quarter.

Has the EU AI Act high-risk deadline been delayed?

Yes. Under the digital omnibus agreed by the Council and the European Parliament in May 2026, obligations for standalone Annex III high-risk systems moved from 2 August 2026 to 2 December 2027. Annex I systems moved to 2 August 2028. The prohibitions in Article 5, the AI literacy duty in Article 4 and the Article 50 transparency rules were not deferred alongside them.

When will the new HIPAA Security Rule take effect?

Not soon. The proposed rule was published on 6 January 2025 and drew more than 4,000 comments. HHS then moved it to the long-term actions section of its regulatory agenda in July 2026, naming July 2027 as the target for final action. That is a target rather than a commitment. The existing Security Rule still applies in full, and the Office for Civil Rights is actively enforcing it.

How do you report an AI incident under DORA?

Through the same route as any other major ICT incident. Once you classify the incident as major, the initial notification is due within 4 hours, and no later than 24 hours after you detected it. An intermediate report follows within 72 hours and a final report within one month. The cause being a model rather than a server changes nothing at all about the clock.

Where to start this week

Start with the artefact you are least confident about, which for most teams is the version log. Open your production agent's configuration and ask one question: can you state, with a timestamp, which model version and which system prompt were live 30 days ago? If the answer is no, that is the first build, and it takes a sprint rather than a quarter.

Then put one line on the agenda of your next risk committee. Name the agent, name its criticality rating, and name the person who owns its record. Most regulated AI programmes have neither of those names written down anywhere, and the distance between an agent nobody owns and a finding nobody expected is short.

References

  1. HHS Office for Civil Rights, HIPAA Security Rule NPRM fact sheet, 6 January 2025. Used for every proposed requirement and cadence.
  2. HHS Office for Civil Rights, Resolution agreements and civil money penalties, accessed 31 August 2026. Used for the 2026 enforcement record.
  3. Clark Hill, HIPAA Security Rule update delayed until 2027, 2026. Used for the Unified Agenda move and the July 2027 target.
  4. European Supervisory Authorities, DORA dry run exercise summary report, December 2024. Used for the 6.5%, 116 checks and 86% figures.
  5. EBA, EIOPA and ESMA, Designation of critical ICT third-party providers, 18 November 2025. Used for the CTPP list.
  6. EUR-Lex, Regulation (EU) 2022/2554 (DORA). Used for the Article 26, 28 and 30 obligations.
  7. Gibson Dunn, EU AI Act omnibus agreement, postponed high-risk deadlines, May 2026. Used for the deferral dates.
  8. Black Book Research, reported by Becker's Hospital Review, December 2025. Survey of 182 US hospital leaders, 15 October to 8 November 2025.

The weakest source here is the Black Book survey: 182 self-reported responses from a market research firm, neither peer reviewed nor a regulatory count, so the 22% is directional only. BaFin's January 2026 guidance is cited in the body and is expressly non-binding. Every other figure traces to a regulator's own publication.

RR
Sidhant Tamrkar
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading