what does better look like?

Leadership development agreed on how to measure itself in 1959, and has been stopping at the easy part ever since.

Ask a room of executives whether their leaders could be better and every hand goes up, which tells you a good deal about the honesty of executives and almost nothing about what they mean by it. Better travels from the brief into the proposal and out again into the board update six months later, without anyone along the way being obliged to say what it would look like if it turned up.

The people doing the commissioning are usually among the most thoughtful in the building, so what defeats them here is a problem of definition, and its outline has been sitting in the industry's own literature for the better part of seventy years.

Where the evidence gets expensive.

In 1959, Donald Kirkpatrick began publishing a series of four articles in the Journal of the American Society of Training Directors on how to evaluate a training programme. He described four steps, which the industry has since come to call levels. Reaction asked whether people enjoyed it, learning whether they took anything in, behaviour whether they did anything differently afterwards, and results whether the organisation changed as a consequence.

Nearly seven decades on it remains the most widely used evaluation framework in corporate and government learning, and for most of that time practitioners have been observing the same thing about it, which is that each step up costs considerably more than the one below. Reaction can be collected in the room before people have found their coats. Learning takes a short assessment against a baseline. Behaviour is where it turns expensive, because time has to pass and somebody has to go and ask people who were never in the session, and results are harder again, since you are by then attributing a change in the business to a change in a person, which is the sort of claim that attracts scrutiny.

The industry does what any industry does when evidence is free at one end and costly at the other. It reports what it can readily gather: satisfaction scores, completion rates, a self-assessment that moved a few points, confidence recorded at the close of the programme while everyone is still in the room.

All of that is proxy. Between them those measures establish that people had a reasonable time and could recall the material, which leaves untouched the question of whether anything shifted where the organisation actually lives, in meetings and in the conversations people are willing to have with each other on a Tuesday afternoon.

Marked by the candidate.

Better gets defined by the provider, measured by the provider and reported by the provider, with no external check anywhere in the arrangement. There is no route through it to a no.

A claim with no path to being disproved is a weak claim in good clothes, and buyers have begun to notice. The organisations worth listening to on this are the ones sitting a year or two past a well-regarded engagement, holding a folder of positive feedback and a leadership team behaving exactly as it did before. Anger would almost be healthier. What they have instead is a quiet conclusion that this is simply how it goes.

Why the new language never survives the meeting.

Chris Argyris named the mechanism underneath this in the 1970s.

He drew a distinction between the theory a person espouses and the theory they actually use. Espoused theory is what we say we believe about how to lead, that dissent is valuable and that the quietest person in the room may be holding the most important information. Theory-in-use is what governs behaviour when the stakes rise and attention narrows. The two are frequently in conflict, and the person holding them is generally the last to know.

Leadership development, almost without exception, works on espoused theory. It delivers a model and a vocabulary, and it delivers them through instruction, so a leader walks out with an upgraded account of how good leadership works while their theory-in-use sits precisely where it was, having been formed by something other than instruction and remaining mostly out of reach of it. Then comes a meeting where a decision is expensive and a senior person has already signalled a preference, and the older theory takes the wheel.

Addition works when there is room for it. In most leaders the room is already taken.

A definition that could come out wrong.

If better is going to mean anything, it has to be able to fail. Three parts to that, and any provider should be willing to sit the test, ourselves included.

It has to be observable, described in behaviour that a person could witness. It has to be detectable by someone who was outside the session and was never told a session took place. And it has to arrive early enough to be attributed, because a change surfacing eighteen months later belongs to eighteen months of other things too.

Run those filters over the vague version of better and it collapses into a list you can look for:

  • A decision that had been circling for months gets made

  • The thing everyone in the organisation knew and nobody said gets said in the room, by the person closest to it

  • People stop holding the pre-meeting before the meeting

  • A team stops routing around a particular leader in order to get work done

  • A leader stops doing the job of the two levels below them

  • Bad news travels upward faster than it used to

No survey detects any of that. A chief of staff notices it without looking, then mentions it to somebody without meaning to.

What arrives when the drag comes out.

Removal is the half of the definition that is easier to describe. The more interesting question is what fills the space.

An organisation carrying less of its defensive layer often gets louder before it gets smoother, and what changes is where the noise sits. Disagreement moves forward in time, closer to the point where it is cheap to resolve. Because the objections that would ordinarily have surfaced later in a corridor have surfaced earlier in a room, decisions get made once and stay made. Information moves at the speed of the problem, which is usually faster than the reporting line allows, and people spend a smaller portion of the working week managing how things look and a larger portion doing the work they were hired for.

Every one of those can be checked by somebody who was never sold anything.

The demand to make.

The next time anyone proposes to make your leaders better, ourselves included, set the contents of the programme and the pedigree of the people delivering it to one side, and say this. Show me what better looks like. What will be observably different, who will notice it without being briefed, and by when.

Anyone who cannot answer is selling you the espoused version. Anyone who can has agreed in advance to be measured on something that might come back against them.

Good leadership and great leadership are not the same thing. The distance between them is made of behaviour that has survived every attempt to add something on top of it, and behaviour like that comes out by being released, which is a different job from teaching.

Sources.

Kirkpatrick, D. L. (1959 to 1960). Techniques for Evaluating Training Programs, a series of four articles in the Journal of the American Society of Training Directors, setting out the four steps now generally known as reaction, learning, behaviour and results. Kirkpatrick cited Raymond Katzell's earlier evaluation questions as a starting point, and the steps were popularised as "levels" by the industry rather than by Kirkpatrick himself.

Argyris, C. & Schön, D. A. (1974). Theory in Practice: Increasing Professional Effectiveness. San Francisco: Jossey-Bass. On espoused theory and theory-in-use.

Argyris, C. (1990). Overcoming Organizational Defenses: Facilitating Organizational Learning. Boston: Allyn and Bacon. On organisational defensive routines.