From Sidhant Tamrkar | Product & Market Analysis
SOC 2 Auditors Are Asking About AI. The Criteria Did Not Change
On this page
There is no SOC 2 for AI. The AICPA has published zero AI-specific trust services criteria, and the operative standard is still the 2017 set whose points of focus were last revised in November 2022, before generative AI reached the enterprise. Your AI tooling is in scope anyway, through 5 criteria series that were already there. What changed is the evidence an auditor samples, not the rulebook.
Key takeaways
- No new criteria exist, and none are needed. The 33 common criteria under Security are technology-neutral. Risk assessment, logical access, monitoring, change management and vendor management reach an AI feature without a single word being added.
- The gate is 9 artefacts, and most teams can produce 6 of them. The two that fail are the dated AI system inventory and evidence that somebody actually read the model vendor's own report.
- Your model provider is usually a carved-out subservice organisation. Carving it out does not remove the work. It converts the provider's controls into assumptions you have to state and defend.
- The premise of this post is the weakest part of it. No published dataset counts how many Type II reports contain AI-related exceptions. The trend is real in practitioner accounts and unmeasured in the public record.
What auditors are actually asking about AI
The question arriving in fieldwork is not philosophical. It is a sampling request. Show me the AI systems in scope, show me who can reach them, show me what you did when one of them behaved badly.
That request is not new in shape. It is the same request an auditor has made about any production system for a decade. It lands differently because most teams adopted AI tooling outside the process that produces audit evidence.
A support chatbot went live because a product manager wanted it. A model API key was created in an afternoon. A drafting assistant spread across the sales team through a personal account. None of that passed through change management, access review or vendor onboarding, which is precisely where the evidence lives.
So the audit conversation feels like a new demand. It is an old demand meeting a class of software that skipped the queue.
Why nothing in the criteria had to change
The Security category, mandatory in every SOC 2, contains 33 common criteria across nine series labelled CC1 to CC9. They describe control objectives, not technologies. That is deliberate, and it is why the framework has outlived several waves of infrastructure.
The AICPA revised the points of focus in November 2022 to reflect newer infrastructure and threats, as EY summarised at the time. The criteria themselves were untouched. Points of focus are illustrative considerations, not requirements, and an auditor can apply a criterion to a technology no point of focus mentions.
CC3 puts new technology in scope by construction
The risk assessment series requires management to identify risks to its objectives and to assess changes that could significantly affect the system of internal control. Deploying a model that touches customer data is exactly such a change.
This is the criterion that does the work. If your risk assessment does not name AI, the failure is not that you lack an AI control. It is that your risk assessment is out of date, which is a finding against a criterion written in 2017.
CC6 covers who and what can reach the model
Logical access criteria govern identification, authentication, authorisation and the removal of access. A model endpoint is an information asset. An API key is a credential. An admin console for a fine-tuning job is privileged access.
Auditors are asking who holds those keys, whether the list was reviewed, and what happened to the key belonging to the engineer who left in March. None of that requires an AI-specific criterion.
CC9.2 is the vendor clause auditors now read closely
The risk mitigation series requires an assessment and management of risks arising from vendors and business partners. A model provider is a vendor with unusually deep access to whatever you send it.
The practical test is narrow. Can you produce a dated assessment of that provider, the report you relied on, and evidence that a named person read it. Most teams have the first, some have the second, and very few have the third.
The 9 artefacts an AI-aware auditor asks for
This is the pack. It is assembled from the criteria above rather than from any AI framework, which is why it survives an auditor who has never audited an AI feature before.
| Artefact | Criteria it answers | What good looks like |
|---|---|---|
| Inventory of every AI system and model in the product and the business. | CC3.2, CC3.4. | A dated register with owner, purpose, model, provider and data class. |
| Risk assessment that names AI risks explicitly. | CC3.2. | Named risks, rated, with the control that addresses each. |
| Change record for the release that shipped the AI feature. | CC8.1. | Ticket, approver, test evidence, deployment date. |
| Access list for model endpoints, API keys and admin consoles. | CC6.1, CC6.2, CC6.3. | Current holders, a review with a date, and revocation on exit. |
| Data flow showing what customer data leaves your boundary. | CC6.7, CC9.2. | A diagram plus the provider's retention and training terms in writing. |
| Vendor file for each model provider. | CC9.2. | Their current report, its period, and who reviewed it. |
| Logging and monitoring output for prompts and responses. | CC7.2. | Redaction applied before logging, with retention stated. |
| Incident procedure naming an AI failure as an in-scope event. | CC7.3, CC7.4. | One worked example, even a low severity one. |
| Acceptable use policy and training records for AI tools. | CC1.4, CC2.2, CC5.3. | Acknowledgement dates per employee, not a policy page nobody signed. |
Criteria references are to the 2017 trust services criteria for Security. Your auditor may map an artefact to a neighbouring criterion. The mapping matters far less than the artefact existing with a date on it.
The two artefacts most teams cannot produce
In my experience the pack fails in the same two places. The inventory does not exist, because AI adoption was decentralised and nobody was asked to write it down. And the vendor file contains a downloaded report that no named person has demonstrably read.
Both are cheap to fix and neither is fixable retroactively with credibility. A register created the week before fieldwork is visibly a register created the week before fieldwork. I would build both now, even if your next audit window is a year away, because the value is in the dates.
Your model vendor is a subservice organisation
If a model provider performs part of the service you commit to your customers, it is a subservice organisation in SOC 2 terms. Your description has to say so, and you have to pick a method.
Carve-out is the default, and it moves the work to you
Under the carve-out method the provider's controls sit outside your audit. Your report names the provider, describes what it does, and lists the complementary subservice organisation controls you assume it operates. Your auditor tests only your side of that line.
That sounds like a reduction in scope. It is a transfer. Every assumption you publish is a claim you are making about somebody else's controls, and the only evidence supporting it is their report and your review of it.
| Method | What it means in practice | Where it costs you |
|---|---|---|
| Carve-out. | Provider named, its controls excluded, assumed controls listed in your report. | You carry the assumptions and must evidence that you monitored them. |
| Inclusive. | The provider's relevant controls are tested inside your engagement. | Needs the provider's cooperation, which a large model vendor will not give a small customer. |
| Where inclusive genuinely wins. | A small provider you have real bargaining power over, or an affiliate you control. | Nowhere, for a frontier model API. Carve-out is the only realistic path. |
What to ask a model vendor for, in writing
Four things, and they are all published or obtainable. The current SOC 2 Type II report and the period it covers. Whether an ISO/IEC 42001 certificate exists. The retention period for prompts and outputs on your plan. And a written statement that your data is not used for training.
The larger providers publish most of this. Anthropic, for example, lists SOC 2 Type I and Type II, ISO 27001:2022 and ISO/IEC 42001:2023 in its own documentation. Treat that as the starting point of a vendor file, not the end of one. What matters at audit is the review you performed, not the certificate the vendor holds. The contract terms behind those answers are worked through in the piece on the AI contract clauses a finance lead should insist on.
Where a provider will not answer, that silence is itself the finding. Record it and rate the risk rather than leaving the question open. The same discipline applies to the remedies you negotiate, which is the subject of the analysis of agent liability caps and what they actually cover.
The Type II window is where AI roadmaps break
A Type I report describes controls at a point in time. A Type II tests whether they operated across a period, commonly 3 to 12 months. That distinction is where AI feature velocity collides with audit mechanics.
Ship a model-backed feature in month 10 of a 12 month window and the control governing it has two months of operating evidence. Not twelve. Your auditor handles that with a partial period description, a scope note, or an exception, and none of the three reads well to a procurement team scanning for clean opinions.
There are three honest options and no clever fourth one. Ship before the window opens. Scope the feature out of the description deliberately and say so. Or accept the partial period and explain it. I would take the second option more often than teams do, because a narrow description that is true beats a broad one carrying an exception.
This is also the practical argument for treating AI features as a release class with its own gate. The failure patterns that show up in production deployments are catalogued in the breakdown of how agent pilots actually fail, and most of them produce audit evidence problems as a side effect.
Discovery comes before policy
Every readiness plan I have seen starts with writing an AI acceptable use policy. That is the wrong first step, and it is wrong for a reason auditors will find.
A policy asserts a boundary. If you have not discovered what is actually running, the policy documents an intention rather than a control, and the first sample against CC6 or CC7 will show tools nobody governs. The cost of that gap is quantified in the analysis of what shadow AI adds to a breach.
Do the discovery first. Expense reports, identity provider logs, browser extension inventories, and a direct question to every team lead. The method matters less than the output, which is a list with names against it. The same problem in its non-AI form is the subject of the piece on application sprawl and rationalisation, and the discovery techniques transfer directly.
Then write the policy against what you found. It will be shorter, more specific and defensible in a way a generic template never is.
What ISO 42001 adds that SOC 2 does not
The two are frequently confused in procurement, and the difference is not subtle. SOC 2 is an attestation over controls relevant to security and related criteria during a period. ISO/IEC 42001 certifies a management system for artificial intelligence itself.
Both want governance, defined responsibilities, risk assessment and monitoring. A team that has assembled the 9 artefacts above has done a meaningful share of the groundwork for either.
The divergence is in subject matter. ISO 42001 asks about AI impact assessment, lifecycle management and the effects of a system on the people subject to it. SOC 2 does not ask those questions, because they were never inside its scope.
In the 2026 AI Index, 36% of surveyed organisations cited ISO/IEC 42001 as an influence on their responsible AI practice, with the NIST AI Risk Management Framework at 33%. Those are influence figures, not certification counts, and the gap between the two is large.
My view is that ISO 42001 is a buyer-driven purchase, not a risk-driven one. Add it when your deals start naming it, and not before. If your obligations are regulatory rather than commercial, the relevant work is different again and is set out in the EU AI Act transparency checklist.
Where this argument is weakest
Three places, and the first one undermines the headline.
The case that this is auditor theatre
If the criteria did not change, a sceptic can reasonably say nothing changed. Auditors have always been able to sample any in-scope system, and calling the AI version of that a shift is partly a marketing frame, one that a large compliance software industry has a direct commercial interest in amplifying.
I think the sceptic is half right. The obligation is old. What is new is that a class of system with unusual data access spread through companies faster than any governance process, and the concentration of that gap is a real change in risk even when the rulebook is identical.
Nobody has published exception rates
This is the honest hole. There is no public dataset counting how many SOC 2 Type II reports issued in 2026 contain AI-related exceptions, or how many auditors added AI sampling procedures. SOC 2 reports are distributed under non-disclosure agreements, so the population is structurally unobservable.
What exists is practitioner reporting and vendor commentary, and vendors selling readiness services are not neutral witnesses. The AICPA's own Journal of Accountancy reported in November 2025 that AI system evaluation is an emerging service line for CPAs, quoting practitioners at the AICPA and EY. That supports direction. It does not measure frequency, and I will not present it as though it does.
Criteria coverage is not the same as adequate coverage
The argument that existing criteria reach AI is true and incomplete. They reach access, change and vendor risk. They do not reach model bias, training data provenance, or whether an output was fit to act on. A clean SOC 2 tells a buyer nothing about those, and a buyer who reads it as AI assurance has misread it.
Frequently asked questions
Does SOC 2 cover AI?
Yes, through criteria that were written before generative AI existed. The AICPA has issued no AI-specific trust services criteria. Your AI tooling is in scope because the risk assessment, logical access, monitoring, change management and vendor criteria are technology-neutral by design. An auditor does not need a new criterion to ask how you control a model. CC3, CC6, CC7, CC8 and CC9 already reach it.
Is there a SOC 2 for AI?
No. There is no separate SOC 2 for AI and no AI module inside the framework. The operative standard is the 2017 trust services criteria, whose points of focus were last revised in November 2022. Vendors selling an AI SOC 2 are selling readiness work against the existing criteria. That work can be useful. The certificate it implies does not exist.
What AI evidence do SOC 2 auditors ask for?
Nine artefacts cover most of it. A dated inventory of AI systems, a risk assessment that names AI risks, the change record for the release that shipped the feature, an access list for model endpoints and keys, a data flow showing what leaves your boundary, a vendor file for each model provider, monitoring output, an incident procedure that names AI failures, and acceptable use training records.
Do I need ISO 42001 if I already have SOC 2?
Not usually, and the two answer different questions. SOC 2 tests whether controls over security and related criteria operated during a period. ISO/IEC 42001 certifies a management system for AI itself, including impact assessment and lifecycle governance. In the 2026 AI Index survey, 36% of organisations cited ISO 42001 as an influence on their responsible AI practice. I would only add it when buyers start asking for it by name.
Is my LLM provider a subservice organisation in my SOC 2?
Usually yes, if the provider performs part of the service you commit to your customers. Most reports carve the provider out. That means your description names the provider and the controls you assume of it, your auditor tests only your side, and the assumption itself becomes your evidence obligation. You need the provider's current report, its period, and a record that somebody read it.
Can I add an AI feature during my SOC 2 Type II window?
You can, and it will cost you evidence. A Type II report covers a period, so a control that started in month 10 of a 12 month window has two months of operating evidence, not twelve. Auditors handle that with a partial period description or an exception. Neither reads well in procurement. Either ship before the window opens, or scope the feature out of the description deliberately.
Where to start this week
Open a spreadsheet and write down every AI system your company uses or ships. One row each, with an owner, the model behind it, the provider, and the most sensitive class of data it touches. Put today's date at the top. That single sheet answers more audit questions than any policy you could write this month, and it is the input to everything else in the pack.
Then pick your largest model provider and file four things against it: the current report, its period, the retention terms on your plan, and a line recording who read them and when. If any of the four is missing, the gap is now documented rather than discovered by somebody else during fieldwork.
Related on the cost side
Evidence work has a running cost, and so does the tooling underneath it. What AI spend actually looks like once it is tracked properly is worked through in the piece on forecasting AI spend.
References
- AICPA, 2017 Trust Services Criteria with 2022 Revised Points of Focus (TSP section 100). Used for the five categories, the common criteria structure and the absence of any AI-specific criterion.
- EY, To the Point: AICPA revises guidance on applying its Trust Services Criteria and SOC 2 Description Criteria, 2 November 2022. Used for the date and scope of the revision.
- Stanford HAI, AI Index Report 2026, chapter 3: Responsible AI. Used for the responsible AI policy figures, the incident counts and the framework influence percentages.
- Journal of Accountancy, A new frontier: CPAs as AI system evaluators, November 2025. Used for the emergence of AI evaluation as a CPA service line and the SOC report comparison.
- ISO, ISO/IEC 42001:2023, Artificial intelligence management system. Used for the scope of the standard and its publication year.
- Anthropic, What certifications has Anthropic obtained. Used as a worked example of what a model provider publishes.
The weakest thing about this source base is that the central claim, that AI questions are appearing more often in Type II fieldwork, has no measured source anywhere. SOC 2 reports are distributed under non-disclosure agreements, so no public dataset of exceptions exists. Treat the trend as practitioner reporting, and treat the criteria mapping, which comes from the AICPA document itself, as the part you can check.
Related reading