Fortune reported last week that Wonder, Marc Lore’s nine billion dollar food hall business, has handed its promotion decisions to a model. Every six months each employee is rated by at least a dozen colleagues on performance, behaviour and leadership. The system folds those ratings together with a value above replacement score, measuring how hard the person would be to replace at their level, and returns a verdict on whether they go up. A manager can override it, provided he makes a convincing argument that it missed something. Lore says there are a handful of disagreements each period, and that as the model gets smarter there are fewer of them every time.
Underneath it sits a belief most capable men hold without having examined it, because it has served them well for twenty years. A good decision is one you can defend. Objectivity is not a tradeoff, it is an upgrade.
The argument about this splits the way it always splits. One side says data finally takes politics out of advancement, and it has a point, because the manager who promotes the man he gets on with is not a hypothetical, and the people he passes over usually have least recourse. The other says a model cannot see what a person sees, and that judgement does not survive being turned into a score. Both sides are arguing about whether the algorithm is accurate. Neither asks what a system of this kind needs in order to keep working.
Start with what it does well, because that part is real. Handling a decision in the moment with a process is more efficient than handling it with judgement, every time. It is faster, it is consistent, it does not get tired and it does not have a favourite. Often that is worth having. But a process is a contained thing by construction. It can only hold what was written into it, and it cannot hold the context that was actually in the room, which is why efficiency now and efficiency later are not the same purchase.
I ran a standardised progression for sizing traders up. We knew when we built it that it would not fit everybody, and it did not. Early on every trainee went onto it, and that was right, because there is no way to tell yet whose trading sits outside the standard. Later you could see it plainly. The schedule holding one man back below what he was already doing well, pushing another faster than he should have been pushed. Those were the ones who needed something built for them, and their trading was not wrong, it was just not the shape the standard was cut for. Forcing them into it would have done them a disservice and cost us the traders. The standard was not the problem and I would build it again. What made it work was that somebody was watching for where it stopped fitting.
Which is what makes the falling override count the wrong thing to celebrate. Exactly two things produce it. The model is getting better, or nobody is looking any more. Those two draw an identical line on a chart, and nothing in the company can tell them apart, because the looking never shows up anywhere. Nobody is assigned to it, and it produces no output on the days it finds nothing. And once the burden of proof has moved from the manager who wants to promote onto the manager who wants to disagree, not looking is also the cheaper option.
That is the durable shape and it has nothing to do with software. A rule is only ever as good as the attention still being paid to where it stops fitting. Setting a default is not making a recommendation, it is deciding who carries the burden of proof, and everything downstream then arrives looking like agreement when what you are reading is the price of speaking. The attention that keeps a rule honest is invisible, unpaid and voluntary, and things with those three properties do not survive a busy year.
And there is a second edge to it, and it runs the wrong way. A standard is at its most accurate when a man is least distinguishable, which is early, and it gets less accurate as he becomes himself. That is the opposite of how firms scale one, since scale pushes standardisation upward. The men most likely to be mispriced are the ones it will be trusted with last.
None of which says the old arrangement was good. It was not, in plenty of places, and the bias the system is aimed at is paid for by people who never see the room in which they lost. The standard is the right place to start, and this is not an argument for a manager’s taste. The point is narrower. The case for the change was made in fairness, the payment is being taken in attention, and nobody ever has to put those two side by side because they never come up in the same conversation.
So the question this puts to you is not whether the machine is fair. It is which of the rules you run your own operation on you have not heard an argument against in a year, and whether anybody is still looking for where that rule has stopped fitting, or whether that job quietly stopped being anybody’s.