How do you measure whether AI training is working?

Compare behaviour against a pre-training baseline: how often people use AI, output quality, time spent on target tasks and error rates.

Why attendance and satisfaction scores fall short

Most training reports count who turned up and how they rated the session. Those numbers show that training happened. They say nothing about whether anyone works differently a month later. A course every attendee rated highly can change nothing, while a session people found uncomfortable can lead to lasting change.

What to measure instead

Focus on behaviour and outcomes in the tasks the training targeted.

  • Regular use: How many people use approved AI tools on the target tasks each week, and whether that number holds steady or falls away.

  • Quality of output: Whether reviewers see better first drafts, fewer rewrites or more consistent standards, for example in bid responses or customer support replies.

  • Time on target tasks: How long specific tasks take, such as preparing a monthly management report or answering a standard supplier query.

  • Errors and rework: Whether mistakes fall, or whether new types of error appear, such as AI-generated figures nobody checked.

  • Judgement: Whether people can explain when they would and would not use AI for a given task.

Do not measure raw usage on its own. High use with poor review is a risk, not a success.

Capture a baseline first

Without a starting point, you cannot show change. Before training begins:

  • Choose three to five target tasks: Pick tasks that matter to the business and that the training will address directly.

  • Record current time and quality: A simple log over two or three weeks, kept by a few people in each team, is often enough.

  • Note current tool use: Ask who already uses AI, for what and with which tools, including anything unapproved, and make it safe to answer honestly.

  • Agree a review date: Set the point at which you will measure again, so the comparison is fair.

Measuring after training

Repeat the same measures at the agreed point, then look for patterns rather than single numbers. If time on invoice queries has dropped but errors have risen, the training needs to cover review more thoroughly. If use has faded in one team, talk to the manager before assuming the training failed.

Allow for confidence developing at different speeds for different people, as set out in how long it takes a team to become confident with AI. A snapshot taken too early can understate progress.

Keeping it proportionate

Measurement should not become a bigger project than the training. For a small team, a short survey, a handful of timed tasks and a conversation with each manager may be enough. Larger programmes justify more structured tracking.

Be clear about what training data can and cannot prove. It can show changes in behaviour and task outcomes. Linking those changes to financial returns needs a broader view of costs and benefits, covered in how to measure AI ROI.

Reporting results honestly

Report what did not work alongside what did. Leaders and staff trust findings more when they include the gaps, and those gaps tell you where to focus next. Share results with the people who took part, too. Seeing that their feedback shaped the next step encourages them to keep using what they learned.

Want to talk this through?