• Operator Playbook
  • Current
  • Advanced
  • 14 min read

What Happens When You Pay for a Metric

Wells Fargo's staff opened roughly two million accounts nobody asked for, and the metric they were chasing moved by less than a rounding step: the clearest evidence on record that a number can be thoroughly corrupt and still look completely healthy.

Organization Design · Services

Key takeaways

  • Wells Fargo's own board investigation found approximately 5,300 employees terminated for sales-practices violations between January 2011 and September 2016, and the board learned that number from the regulators' settlements, not from its own management.
  • Stripping out the improper accounts barely touched the number they were chasing: the board's forensic reviewers found removing all unfunded accounts moved the reported cross-sell metric by at most 0.02 in any quarter, from 6.16 to 6.14 in the December 2013 quarter.
  • The board report's finding on motive is the counterintuitive one: sales pressure and goals, not incentive compensation directly, were the primary drivers of misconduct, and investigated employees only infrequently cited the bonus.
  • Piece rates work when failure is attributable. Safelite's 44% output-per-worker gain (Lazear, American Economic Review 2000) came with a rule that an installer who broke a windshield reinstalled it on his own time and paid for the glass, and about half the gain was better hiring, not more effort.

The number stayed clean while the company burned

The standard telling of the Wells Fargo sales-practices scandal is about greed and quotas. The genuinely useful part is quieter and appears in a footnote of the bank's own investigation: the metric everybody was cheating to move barely moved.

The scale first. On 8 September 2016 the Consumer Financial Protection Bureau fined the bank $100 million, alongside $35 million from the Office of the Comptroller of the Currency and $50 million from the Los Angeles City Attorney, $185 million in total, covering conduct back to 1 January 2011. Employees had opened roughly 1.5 million unauthorised deposit accounts and about 565,000 credit card accounts. In February 2020 the Securities and Exchange Commission added a $500 million payment, part of a combined $3 billion resolution with the SEC and the Department of Justice, for misleading investors between 2012 and 2016 about the very cross-sell strategy the accounts were inflating.

Now the footnote. The independent directors' investigation, run by Shearman & Sterling with FTI Consulting doing the forensic work, tested what the reported cross-sell metric would have looked like with the improper accounts stripped out. Removing every unfunded account identified across May 2011 to July 2015 changed the reported metric by at most 0.02 in any quarter. Their worked example: the December 2013 quarter goes from 6.16 to 6.14. Adding the potential simulated-funding accounts and potentially unauthorised credit cards took the maximum impact to 0.04. Across the whole period, backing all of it out flipped the direction of the reported trend in exactly two quarters.

Sit with the implication. Two million accounts nobody asked for. Roughly 5,300 people fired. Two federal regulators, a city attorney, the SEC and the DOJ. And the dashboard number the whole apparatus existed to raise was, to two decimal places, essentially unchanged. The metric was not a liar because it was wrong. It was a liar because it was insensitive: so coarse that industrial-scale gaming disappeared inside its rounding. Every organisation running a headline ratio should ask the same question before it attaches money to one: how much misconduct would it take to move this number by an amount I would notice? If the answer is “more than I could survive,” the number cannot serve as its own alarm.

One more detail from the same report, because it is the piece most retellings drop. Investigated employees only infrequently referenced incentive compensation as a motivating factor; the investigators concluded that sales pressure and goals, rather than incentive compensation directly, were the primary motivators. The plans were internally known as 50/50 plans, set with the expectation that only half the regions could meet them, and performance against them drove rankings, promotion and daily “Motivator” reports circulated down to district level. You do not need a commission cheque to induce this. A ranked list, published often enough, will do it.

Why the theory says this is unavoidable

The behaviour has a formal explanation that predates every example in this briefing, and knowing it saves you from believing that a smarter metric would have prevented any of it.

Bengt Holmström and Paul Milgrom set it out in 1991. Their model starts from an assumption so mild it is almost a definition: the agent's cost depends on the total attention they spend across all their tasks. It follows immediately that raising the pay on any one task pulls attention away from every other one. Incentive pay does not only motivate effort. It allocates effort, and the allocation is the part nobody budgets for.

Two results fall out. The first is startling on first reading and obvious on the second: an optimal contract can be a flat wage, no performance component at all. Their general statement is the sentence to keep: the desirability of providing incentives for any one activity decreases with the difficulty of measuring performance in any other activities that make competing demands on the agent's time and attention. This is why incentive clauses are far rarer in real employment contracts than one-dimensional theory predicts. It is not managerial timidity. It is correct.

The second result is that job design is itself an incentive instrument. If you can separate the measurable task from the unmeasurable one and give them to different people, you can pay steeply for the first without damaging the second. Their own prediction, borne out in the field: franchisees face very steep performance incentives while managers of otherwise identical company-owned stores often receive no incentive pay at all, because the franchisee owns the asset whose long-run condition cannot be measured, and the salaried manager does not.

Their illustration was schooling: pay teachers on test scores and you buy tested skills at the expense of untested ones. Their footnote cites a real case: a ninth-grade teacher in Greenville, South Carolina caught in 1989 passing statewide basic-skills test answers to her geography students to improve her own performance rating. The mechanism has not changed in nearly forty years because it is not a flaw in any particular scheme. It is what paying for a proxy means.

The practical translation for an operator: before attaching money to any measure, write down the tasks that compete with it for the same person's attention, and ask which of them you can measure. If the answer is “none of them,” you are not designing an incentive. You are choosing which parts of the job get abandoned.

The four ways a metric gets hit without the goal being met

Gaming is not one behaviour. It is four, they have different tells, and only one of them is what most people picture.

Reclassification changes what counts, not what happens. The cleanest documented case is the Department of Veterans Affairs. In May 2014 the VA Office of Inspector General reviewed a sample of 226 new-patient primary care appointments at the Phoenix Health Care System. VA national data, as reported by Phoenix, showed those 226 veterans waiting an average of 24 days, with 43% waiting more than 14 days. The OIG's own reconstruction of the same 226 cases found an average wait of 115 days, with an estimated 84% waiting more than 14 days. The mechanism was mundane: schedulers opened the system, found the first available slot, asked the veteran if it was acceptable, then entered that slot's date as the veteran's “desired date of care,” which made the measured wait zero. The OIG also identified 1,700 veterans waiting for primary care who appeared on no wait list at all. Nobody falsified a medical record. They edited the field the metric read.

Suppression removes the input. Here the regulator wrote the rule down. OSHA's March 2012 memorandum, signed by Deputy Assistant Secretary Richard E. Fairfax, addresses employer safety incentive programmes directly, including, in its own words, the arrangement where a team of employees might be awarded a bonus if no one from the team is injured over some period. The problem is structural: when reporting an injury costs the reporter and their whole crew a payment, the programme is not buying safety, it is buying silence, and where the incentive is large enough to have dissuaded a reasonable worker from reporting, OSHA treats it as retaliation under section 11(c). The agency's recommended alternative is to pay for participation (hazard identification, safety committee service, near-miss reports) because those are inputs a worker can only increase by doing more, not by seeing less.

Substitution defeats the countermetric. This is the failure mode that catches sophisticated designers, because they did add a quality gate. Wells Fargo balanced sales with customer survey scores. The board report records employees entering fake customer phone numbers, or substituting their own email addresses, so the survey never reached a customer who might score them badly, and in one branch a manager falsifying numbers and instructing staff to do the same, with at least 192 customer phone numbers deleted. A countermetric collected through a channel the measured person controls is not a countermetric. It is a second thing to game.

Timing moves the event across a boundary. Quarter-end discounting to pull orders forward, deferring a repair past a reporting date, closing a ticket and reopening it. Timing games are the least destructive and the easiest to detect: they show up as a spike in the last days of a period and a hole at the start of the next, and any distribution-by-day chart will find them. They are also the most reliable early warning that the other three are happening, because a team willing to move a date is already treating the metric as the target.

When paying for output genuinely works

None of this argues against performance pay. It argues for a set of conditions, and there is a well-documented case where they were all met.

Edward Lazear studied Safelite Glass, which moved its windshield installers from hourly wages to piece rates across 1994 and 1995, at roughly $20 per unit installed, with a guarantee of about $11 an hour so that no one's floor fell. Output per worker rose 44%. Pay for a given worker rose about 10%; the firm and the workforce split the gain. Output variance rose too, which is what you expect when ambitious people are finally allowed to differentiate themselves.

Two details in that study matter more than the headline. First, only about half of the 44% came from existing workers producing more. The rest came from sorting: the firm attracted and retained more productive installers, and lost the ones for whom the deal was worse. If you are modelling a switch to performance pay, you are forecasting a change in who works there, not only a change in how hard they work, and the sorting half takes quarters to arrive.

Second, and decisively for this briefing: the obvious failure mode of piece rates is that quality collapses, and at Safelite it did not, because of the specific way quality was policed. Most defects surfaced quickly as broken windshields, and the guilty installer could be identified. That installer had to reinstall on his own time and pay the company for the replacement glass before any paying job was assigned to him. The cost of the defect landed on the person whose speed caused it. Notably, the firm had first tried a peer-pressure version, assigning the redo to a random worker in the shop, unpaid, and moved away from it.

That is the general rule, and it is a narrow one. Pay for a measured output when the output is countable without judgement, when a failure is attributable to a specific person, and when the cost of that failure can be routed back to them at roughly its true size. Windshields satisfy all three. Almost nothing in professional, clinical or relationship work satisfies any of them, which is precisely why the account-opening metric could be corrupted for five years without the corruption becoming visible in the number, and why the theory in the previous section says the flat wage is not a cop-out.

Illustrative only, to show the design working in a normal business. Twelve inside-sales reps on $4,000 base plus $150 per closed deal, closing 240 deals a month: $48,000 of base, $36,000 of commission, $84,000 total. Thirty per cent of those deals cancel within 90 days, or 72 deals, and each cancellation burns about $400 of onboarding and support, so $28,800 a month evaporates while the rep keeps the full $150. Restructure to $80 at close and $100 at day 90 if the account is still active. At today's 70% survival the expected payout is 80 + 0.70 × 100 = $150, unchanged, which is the only version that survives contact with a sales floor. If survival rises to 85%, cancellations halve to 36, the onboarding burn falls to $14,400, a $14,400 saving, while commission rises to 240 × $165 = $39,600, up $3,600. Net $10,800 a month.

And the honest caveat, because it is bigger than the gain: reps will hit the new number partly by declining marginal buyers. If closes fall 8%, from 240 to 221, and a surviving deal is worth $2,400 of first-year revenue, that is $45,600 of revenue not booked, four times the saving. Deferred commission pays only when the deals it deters were genuinely worth less than they cost to serve. Measure that before you ship the plan, not after the quarter it breaks.

Designing a plan that survives contact with the people paid by it

Assume competent, decent people who will nonetheless respond to the incentive as written. Every rule below exists because a specific, documented failure would have been caught by it.

Set the quota where most people can reach it. Wells Fargo's plans were built so that only about half the regions would make them, and the investigators found leadership continued to defend targets that regional heads were telling them generated products customers neither needed nor used. A target most of the population cannot meet honestly does not select for excellence; it selects for whoever is most willing to bend, and it teaches everyone else that the honest route does not reach the number.

Audit against a record the measured person cannot edit. The VA's schedulers were typing into the same field the metric read; the fraud was invisible from inside the system and obvious the moment an outside reviewer reconstructed the same 226 cases from the underlying records. Pick a small sample every month, twenty is usually plenty, and reconcile the scored number against an independent trace: bank deposits, a customer callback from a list the rep did not supply, the raw log rather than the summary table. The sample size matters far less than the independence of the source.

Refuse to accept a countermetric collected through the measured channel. If the quality gate is a survey and the salesperson supplies the contact details, you have built the Wells Fargo phone-number problem. Draw the contact list from the system of record, not from the person being scored.

Cap the upside and defer part of the payment. A cap removes the far tail where the return on cheating goes vertical. Deferral, as in the worked example above, buys you the only thing a close-date metric cannot see: whether the thing you paid for lasted.

Pay for inputs where outputs are unmeasurable. This is OSHA's recommendation in the safety case and it generalises cleanly. If the outcome you want cannot be observed without the cooperation of the person you are paying, buy the behaviour instead (hazard reports filed, discovery calls completed, code reviewed) and accept that you are measuring something smaller and more honest.

Finally, watch the distribution, not the average. Gaming almost always shows up in shape before it shows up in level: a pile of results just above the threshold, a hole just below it, a spike in the last three days of the period. That signature was available in the Wells Fargo data for years. The board only learned the scale of the terminations, approximately 5,300 people since January 2011, through the September 2016 settlements with the Los Angeles City Attorney, the OCC and the CFPB. By that point the bank had also clawed back and forfeited more than $180 million from senior leaders and terminated five Community Bank executives for cause. All of that was the price of a number nobody had thought to audit against anything but itself.

Put it to work

Before shipping any bonus plan, write down the three things the metric does not capture and who protects them. Sample twenty scored records against an independent source: a system the paid person cannot edit. Set the quota where most people can hit it, not half. Pay part of it late, tied to survival. Publish the audit result to the people being paid.

Sources & references

Linked entries open the named source directly. Entries without a link say exactly what kind of reference they are — and how to check them yourself.

Educational note: This briefing is general business education, not financial, legal, tax, or investment advice. Figures and rules change and vary by situation — verify current specifics with primary sources and qualified professionals before acting.