“What changed?”
It’s the question behind every programme budget. And it’s often the hardest one for L&D to answer, because the honest reply is: people came, they rated it well, and we don’t really know what happened next.
What does the Kirkpatrick model ask us to measure?
Four things, in order: how people reacted, what they learned, whether they behaved differently at work, and whether that led to the results the organisation wanted.
Most people in L&D know the model. It’s been around since the 1950s, and later versions add return on investment as a fifth level. The first two levels happen in or near the training room. The last two happen at work. That’s where the trouble starts.
Why does most evaluation stop at the happy sheet?
Because reaction and learning are easy to collect, and behaviour and results aren’t.
In a 2025 ATD survey, 93% of organisations collected satisfaction data, while 90% said isolating the impact of training was a challenge. In the CIPD’s UK survey, only about a quarter of L&D professionals agreed their organisation assesses the impact of learning.
The problem is that the easy measures don’t tell you much about the hard ones. A large review of training studies found that how much people enjoyed a course was essentially unrelated to whether they used what they learned. A good score on the day shows you ran a good day. It doesn’t show that anyone’s behaviour changed.
Why is behaviour change so hard to evidence?
Because it happens after the session, at work, and the people who designed the learning often can’t see it.
It looks different depending on where you sit.
If you’re an external provider, your contact often ends when the programme does. The client commissioned the sessions. The leaders go back to their jobs, and you rarely get to follow up with them, let alone their teams. You’re asked to show impact in a place you no longer have access to.
If you’re in an internal L&D team, you can reach the leaders, but the data is hard to get. Line managers are busy, performance data sits in other systems, and a follow-up survey three months later competes with everything else in people’s inboxes.
I’ve been on both sides of this. In both, the evidence that mattered was being created at work every day, and almost none of it was being captured.
What evidence can you realistically gather?
Evidence recorded by the leader, close to the moment, and repeated over time.
It closes the biggest gap: it moves the record from the training room to the workplace, and from once to many times. There’s also a benefit beyond measurement. In a field study, trainees who spent 15 minutes a day writing down what they’d learned did better on their final assessment than those who didn’t. Recording practice helps people learn from it, as well as showing what happened.
This is how the Trust Leader Hub gathers evidence:
- The Trust Map, revisited. A leader maps one working relationship before they start. After practising, they return to the same relationship and questions and score it again. That gives a dated before and after, based on the changes they’ve observed in how the other person responds.
- Activity Reflections. After each activity, and after at least two days of practice at work, the leader records what they did and what happened.
- Check-ins at 14, 30 and 60 days. After a Behaviour Playlist, leaders are prompted to look back on what has shifted.
- Impact in three places. Leaders can describe the impact for themselves, for the people they work with, and for their organisation.
- A Cohort Impact Report. The person supporting a group can download a group-level view, without individual rows.
What about self-reporting?
It’s a real limit, and it’s worth being clear about.
Everything above is reported by the leader. People don’t always see themselves as others do: research comparing self-ratings with ratings from managers and peers finds only moderate agreement. So we describe this evidence as reported movement, not proof that a relationship changed.
What makes it useful is that it’s specific, dated and repeated. A leader who writes “I raised the delay on Tuesday before I was asked, and she thanked me for the warning” has given you something you can follow up. Patterns across a group tell you more than any single entry. And the strongest cases make an excellent starting point for a short success-case conversation with the leader and their manager, a well-established way of finding out what’s actually working.
So what should you measure?
Measure what happens at work, as close to the behaviour as you can get.
The organisation hasn’t paid for the training. It’s paid for the impact. You can get much closer to that impact than a satisfaction score: what leaders practised, what they noticed, how the people around them responded, and what changed over the following weeks. Gather that where the work happens, and the question “what changed?” becomes one you can answer.
