From Sanskriti Khandelwal | Product & Market Analysis
AI Training Programmes That Changed Behaviour, Not Just Completion Rates
On this page
Most AI training programmes produce a completion certificate and no measurable change. Dose is part of the difference. In BCG's 2025 survey of more than 10,600 employees, 79% of those who received over five hours of AI training were regular users, against 67% below that line. The programmes that changed behaviour were attached to a specific workflow and measured against a baseline recorded before anyone was trained.
Key takeaways
- Five hours is a floor, not a target. BCG found 79% regular AI use above that line and 67% below it, across 10,600 employees in 11 countries. A 12-point gap is real, and it is smaller than most training budgets assume.
- The same tool, measured two ways, gave answers a factor of two apart. A UK cross-government trial of 20,000 licences reported 26 minutes saved per user per day. HMRC's own evaluation of the same assistant in the same quarter reported around 60 minutes per week.
- Training pays most where the skill floor is lowest. A peer-reviewed study of 5,179 support agents found a 14% average productivity gain, 34% for novices, and minimal effect on the most experienced staff.
- The blocker is rarely motivation. The UK Skills for AI review, drawing on 536 employer responses, found the barriers are time, cost, staff pressure and unclear provision, and that short modules of 30 to 90 minutes work better than courses.
What AI fluency means once you have to measure it
Fluency is a slippery word. Vendors use it to mean familiarity with a product. Learning teams often use it to mean a completed module. Neither definition produces something you can observe six weeks later.
The most useful public definition breaks the skill into four parts. Anthropic's AI Fluency framework, built with Rick Dakan of Ringling College and Joseph Feller of University College Cork, names them delegation, description, discernment and diligence. Delegation is deciding what the model should do at all. Description is communicating the task. Discernment is judging what comes back. Diligence is using it responsibly and owning the output.
This post is written for the person who owns delivery, not the person who owns the licence spend. That is the reader who has to answer for whether the team's output changed, and who therefore needs a definition that survives contact with a metric.
The competency most programmes skip
Nearly every AI training session covers description. Prompting demonstrates well on a screen, it produces an immediate result, and it fills an hour comfortably.
Discernment is the one that gets cut. Judging whether an output is correct, complete and safe to send is slower to teach, harder to script and impossible to demonstrate with a single clever example. It is also the competency that stops a wrong number reaching a client.
I would build a programme backwards from discernment. Give people 40 real outputs from their own workflow, half of them subtly wrong, and have them mark the set. That exercise generates a score you can repeat later, which is more than most curricula manage. The same logic applies when you are hiring, which is covered in the piece on interview questions that test AI fluency rather than tool familiarity.
The five-hour threshold and what it actually buys
BCG has run an annual employee survey on AI at work since 2023. The 2025 edition covered more than 10,600 respondents across 11 countries, which makes it the largest repeated read on this question in public.
Its finding on training is specific. Regular usage is sharply higher for employees who receive at least five hours of training and have access to in-person sessions and coaching. Reporting on the survey put the split at 79% regular users above five hours against 67% below. In the same survey, 18% of people who described themselves as regular AI users said they had received no training at all.
Twelve points is a smaller effect than it sounds
Read the bars honestly. Moving someone from four hours of training to six hours does not turn a refuser into a power user. It moves the odds by roughly 12 points, in a self-reported survey, with no control for the kind of organisation that bothers to fund six hours.
The 2026 edition sharpens the point. Across roughly 12,000 respondents, 72% said skill expectations had shifted and only 36% felt adequately upskilled. That is a supply problem in hours, and hours are the cheapest thing to fix. The reason most programmes still fail is not the hour count.
My position is that the five-hour figure is a floor to clear, then ignore. Spend the design effort on what those hours contain and what happens in the eight weeks afterwards. The gap between people who feel confident and people who actually use the tool is examined further in the analysis of the gap between reported trust in AI tools and reported usage.
Two government trials, one tool, two different answers
The clearest natural experiment in this field is public and largely unread. In late 2024 the UK government deployed the same AI assistant twice, under two different evaluation designs, and published both.
The cross-government experiment run by the Government Digital Service put 20,000 licences across 12 organisations and reported an average saving of 26 minutes per user per day. Satisfaction averaged 7.7 out of 10 and 82% of users said they would not want to return to working without it.
HMRC evaluated the same tool over the same months, and reached a different number.
| Design | Cross-government experiment (GDS) | HMRC Phase 3 evaluation |
|---|---|---|
| Licences | 20,000 across 12 organisations. | 3,000 randomly allocated, plus 500 for volunteers and accessibility needs. |
| Period | 30 September to 31 December 2024. | September to December 2024. |
| Evidence base | 7,115 survey responses. Usage data for 14,500 users, and 5 focus groups. | 1,364 survey responses at a 45% response rate. Usage monitoring, 35 focus group participants, 49 task exercise participants. |
| Headline saving | 26 minutes per user per day. | 2% to 3% of the working week, around 60 minutes, then cut by roughly 20%. |
| Satisfaction | 7.7 out of 10. 82% would not go back. | 7.1 out of 10. 61% would be disappointed to lose the licence. |
Neither evaluation used a control group. Both rely substantially on self-reported time savings. The HMRC report applies an explicit downward adjustment for non-users and survey response bias; the cross-government report does not.
Why the two numbers differ
HMRC applied a stricter method. It reduced its survey figures by around 20% to account for people who held a licence and did not use it, and for the tendency of survey respondents to be enthusiasts. It also reported that 46% of non-users cited security and data privacy concerns, which is a finding no enthusiasm-weighted average will surface.
The cross-government report contains the sentence that matters most for training design. It found a correlation between a user's confidence and familiarity with AI and the time they saved, and concluded that this highlights the need for change management programmes covering communications, engagement and training.
That is a correlation, not a causal claim, and the report is careful to say so. It still points at the practical question. If confidence predicts saving, the programme's job is confidence in a specific task, not awareness in general. Applying HMRC's 20% haircut to any vendor-supplied saving figure is the single cheapest piece of discipline available, and it belongs alongside the tests set out in the forensic read of published AI case studies.
Training pays most where the skill floor is lowest
Two studies with real experimental design tell you where to spend a limited budget, and they point the same way.
The first is Generative AI at Work, published in the Quarterly Journal of Economics in May 2025. Brynjolfsson, Li and Raymond studied the staggered rollout of an AI assistant to 5,179 customer support agents. Access raised issues resolved per hour by 14% on average. For novice and lower-skilled workers the gain was 34%. For experienced, high-performing agents it was minimal.
The mechanism the authors describe is the interesting part. The tool spread the practices of the best agents to everyone else, moving new staff down the experience curve faster. That is a training effect delivered by a product rather than a course.
The teamwork result changes who you enrol
The second study is The Cybernetic Teammate, a pre-registered field experiment with 776 professionals at Procter and Gamble, run in summer 2024 and later published in Organization Science. Participants worked on real product development problems, randomly assigned to work alone or in pairs, with or without AI.
Individuals working with AI matched the performance of two-person teams working without it. The AI also cut across functional silos: without it, research staff proposed technical solutions and commercial staff proposed commercial ones, and with it both groups produced more balanced proposals.
The practical read for a training programme is uncomfortable. If one trained person with a model performs like an untrained pair, the unit of enrolment should not be the department. It should be the workflow that currently needs two people to cover two kinds of expertise.
What the UK evidence review found actually works
In June 2026 the UK government published the output of its Skills for AI programme, What Works for AI Upskilling in the UK. It draws on 23 workshops with around 150 organisations, 10 case studies and a UK employer survey of 536 responses.
Two findings are worth acting on immediately. Over 44% of surveyed organisations report using AI tools daily, while most remain early in adoption, and staff largely learn through trial and error, peer support, online videos and prompts built into the tools. Formal training is not where the learning is happening.
The barriers are also not what training vendors assume. The review names limited time, staff pressure, cost, unclear provision and fear of failing at technical material. Motivation is not on the list. If your programme's core problem statement is engagement, you are solving for something the evidence does not show is broken.
Short and stackable beats a day off site
The review's design guidance is summarised as PRIMES: practical, reachable, integrated, modular, expandable and sustainable. Stripped of the acronym, the operative recommendation is that very short modules of 30 to 90 minutes are more practical than longer courses.
That directly contradicts the standard corporate shape, which is a half day with a vendor and a follow-up email. It also fits what the five-hour threshold implies: five hours accumulated across six weeks in the flow of work, not five hours in a room.
| Dimension | One-off session | Workflow-tied cohort |
|---|---|---|
| Cost per head | Low, and it scales to full headcount. | 3 to 5 times higher, and it does not scale cheaply. |
| Elapsed time | One afternoon. | 5 to 8 weeks, including the baseline period. |
| What it produces | Attendance records, awareness, a signed policy acknowledgement. | A measured change in one named workflow, with a before and after. |
| Where it fails | No observable behaviour change, and no way to prove there was any. | Covers a fraction of the workforce, so the rest stay untrained. |
| Where it genuinely wins | Policy, consent, data handling and legal exposure, delivered to everyone in a week. | Anything where you will later be asked what changed. |
The last row is the honest split. Run both. The mistake is running only the first and describing it as capability building.
Most programmes measure the wrong four things
Learning teams have had a four-level evaluation model since the 1950s: reaction, learning, behaviour, results. Reaction and learning are cheap to collect. Behaviour is the level everyone says they want and almost nobody instruments.
AI training makes this worse, because the tools generate usage telemetry that looks like behaviour data and is not. Weekly active users tells you a licence was opened. It does not tell you a workflow changed.
| Level | What most programmes record | What to instrument instead |
|---|---|---|
| Reaction | A satisfaction score, which is almost always above 4 out of 5. | Whether attendees recommend it to a named peer on their own team within two weeks. |
| Learning | A quiz, or a completion flag. | A marked exercise on 40 real outputs from the team's own workflow, scored against a rubric. |
| Behaviour | Nothing, or licence-level active users. | Use inside the one named workflow, counted at week 8, not week 1. |
| Results | A vendor multiplier applied to headcount. | One business metric already on an existing dashboard, with the pre-training baseline attached. |
The time saved has nowhere to go
Here is the finding that should reshape the closing session of every AI programme. BCG's 2026 survey found that 66% of regular AI users receive limited or no guidance on what to do with the time they save.
Twenty minutes returned to an unstructured day disappears without trace. Twenty minutes redirected into a named backlog, a second review pass or one more client conversation shows up in a number somebody already reports. The redirection is a management decision, and it is the step most programmes leave undesigned.
Budget pressure makes this urgent. The Wharton Human-AI Research and GBK Collective survey of more than 800 senior leaders, published in October 2025, found training investment slipping by 8 percentage points year on year and confidence in upskilling programmes down 14 points, even as 46% named delivering effective training as a top barrier. Programmes that cannot show a result are losing their funding, which is the correct outcome for programmes that cannot show a result. Function-level payback estimates are worked through in the breakdown of where AI payback appears by business function.
Where this argument is weakest
This section exists because the evidence base here is thinner than the confidence with which it is usually quoted, including by me.
The most quoted number in learning and development is not evidence
The claim that only 10% of training transfers to the job appears in hundreds of articles and several academic papers. It traces back to a 1982 article by David Georgenson in the Training and Development Journal, and Robert Fitzpatrick's examination of the original found that Georgenson used it as a rhetorical opener, quoting what training directors say, with no evidence or authority cited for the figure.
I nearly opened this post with that number. It is the exact failure the sourcing rules exist to catch: a striking statistic restated so often it acquires the texture of a finding. Six restatements is one source, and in this case the one source was a rhetorical question. The figure should be retired.
Almost everything here is self-reported
Of the six main sources in this post, only the Quarterly Journal of Economics study and the Procter and Gamble experiment have anything resembling a control. The BCG surveys, the UK employer survey and both government evaluations rest on people estimating their own behaviour and their own saved time.
Self-reported time saving is the least reliable measure in the set. People who like a tool overestimate; people who never opened it do not respond to the survey. The 2.2 times gap between two UK evaluations of the same product is the size of that problem, measured. Treat every number in this post as directional, and treat your own baseline as the only figure you can defend.
A 90-day design that produces evidence
The following is deliberately narrow. It trains fewer people than a company-wide rollout and produces something a finance director will accept.
Weeks 1 to 4 are the baseline. Pick one workflow with a number the business already tracks, and record that number weekly with nobody trained. Cases closed, drafts produced, cycle time, first-response quality. Do not announce the programme yet, because announcing it moves the baseline.
Weeks 5 to 10 are the cohort. Six sessions of 60 to 90 minutes, one per week, each built around the team's own live work rather than demonstration examples. Total contact time lands above the five-hour threshold. One session is the marked discernment exercise. One session is the redirection decision: what the recovered time is now for, agreed with the manager who owns the workflow.
Weeks 11 to 13 are the measurement. Same metric, same collection method, no change to definitions. Report the delta, the number of people who used the tool in that workflow in week 12, and the honest confound list. If the cohort's manager also changed something else, say so.
Two rules make this hold. Cap the cohort at one team, because a mixed cohort gives you a mixed baseline. And do not let the programme own the metric. If the workflow's own manager does not already report that number, pick a different workflow. The failure patterns that follow from a central team owning delivery are set out in the analysis of when an AI centre of excellence becomes the bottleneck, and the quieter forms of refusal are covered in the piece on passive resistance to AI rollouts.
Frequently asked questions
How long should an AI training programme be?
There is no measured optimum, but there is a measured floor. BCG's 2025 survey of more than 10,600 employees found 79% of those with over five hours of AI training were regular users, against 67% below that line. Treat five hours as a minimum, spread across several short sessions rather than one day. The UK Skills for AI review found modules of 30 to 90 minutes more practical than longer courses.
Does AI training actually change behaviour or just completion rates?
It depends entirely on what the programme is attached to. Training tied to one named workflow, with a baseline recorded before the cohort starts, produces a number you can check afterwards. A standalone awareness session produces attendance data and nothing else. The UK evidence review found staff mostly learn through trial and error, peers and online videos, which is what fills the gap when formal training is generic.
Who benefits most from AI training?
The least experienced people in the role. The peer-reviewed study of 5,179 customer support agents published in the Quarterly Journal of Economics found a 14% average gain in issues resolved per hour, rising to 34% for novice and lower-skilled workers, with minimal effect on the most experienced. If your budget covers only part of the workforce, enrol the bottom of the skill distribution first.
How do you measure the return on an AI training programme?
Pick one workflow, record its baseline for four weeks before training, then measure the same thing for eight weeks afterwards. Use a unit the business already tracks, such as cases closed, drafts produced or cycle time. Discount self-reported savings: HMRC reduced its own survey figures by roughly 20% to account for non-users and response bias, which is a reasonable default for any vendor number.
Why do AI time savings not show up in business results?
Because nobody decides what happens to the time. BCG's 2026 survey found 66% of regular AI users receive limited or no guidance on what to do with the time they save. Twenty minutes returned to an unstructured day disappears. Twenty minutes redirected into a named backlog, a second review pass or an extra client call becomes visible in a metric somebody already reports each week.
What should an AI fluency curriculum cover?
Four things, using Anthropic's public framework as a spine: delegation, deciding what the model does at all; description, communicating the task; discernment, judging the output; and diligence, using it responsibly. Most programmes cover description, because prompting demonstrates well, and skip discernment, which is the competency that stops bad output reaching a customer. Add your own data handling and disclosure rules on top.
Where to start this week
Open your last AI training deck and find the slide that names a workflow. If there is no such slide, that is the finding, and it explains the missing behaviour change without any further analysis.
Then pick one team and start the four-week baseline before you book a single session. Recording a number you already collect, for a month, with nobody trained, costs nothing and is the only thing that will let you answer the question you will be asked in the next budget round. Performance expectations that follow from all this are worked through in the piece on rewriting review criteria for AI-assisted work.
One test before you buy
Ask any training vendor for one outcome number with a stated sample, a stated period and a stated exclusion list. Then apply HMRC's 20% haircut to it. If the number survives both, the vendor has measured something. If the answer is a productivity multiplier, read what the deployments with negative return had in common first.
References
- BCG, AI at Work: Why Strategy Matters More Than Tools, fourth edition, June 2026, approximately 12,000 respondents. Used for the 72% and 36% upskilling figures and the 66% guidance figure.
- BCG, AI at Work 2025: Momentum Builds, but Gaps Remain, 26 June 2025, more than 10,600 respondents across 11 countries. Used for the five-hour threshold and the training adequacy figure.
- UNLEASH, BCG's AI at Work 2025 report: four takeaways for HR leaders, 2025. Used for the 79% against 67% split and the 18% no-training figure. Secondary reporting of the BCG survey; upgrade if BCG publishes that cut directly.
- GOV.UK, Microsoft 365 Copilot Experiment: cross-government findings report, June 2025. Used for the 26 minutes per day figure, satisfaction scores and the confidence correlation.
- GOV.UK, Evaluating the impact of Microsoft Copilot in HMRC, Phase 3 evaluation. Used for the 60 minutes per week figure, the 20% adjustment and the non-user barriers.
- GOV.UK, Skills for AI: what works for AI upskilling in the UK, executive summary, 10 June 2026. Used for the evidence base, the barriers list, PRIMES and the module length guidance.
- Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics, May 2025. Used for the 14% and 34% productivity figures across 5,179 agents.
- Dell'Acqua and others, The Cybernetic Teammate, NBER Working Paper 33641, 2025, later published in Organization Science. Used for the 776-participant field experiment at Procter and Gamble.
The weakest thing about this source base: six of the eight references rest on self-reported behaviour, and self-reported time saving is the least reliable measure in the set. Only the Quarterly Journal of Economics study and the Procter and Gamble experiment carry a control condition. Figures are current as of 1 September 2026.
Related reading