- Research Paper
- Current
- Advanced
- 14 min read
Price, Choice Architecture and What Actually Replicates
Price format and menu design really do move behaviour, but half the famous findings shrank when tested properly, and the half that survived is now the half regulators are writing rules about.
Consumer Psychology · Consumer
Key takeaways
- Left-digit thinking survives at industrial scale: across more than 22 million wholesale used-car transactions, sale prices drop discontinuously at 10,000-mile odometer thresholds, with smaller drops at 1,000-mile marks.
- Nudge effects shrink outside journals. Two US government nudge units averaged a 1.4 percentage point take-up effect (an 8.1% increase) against 8.7 points (33.5%) in academic papers, and publication bias plus low power accounts for the full gap.
- The headline nudge meta-analysis reported Cohen's d = 0.45 across 455 effect sizes; a bias-corrected reanalysis of the same data put the combined estimate at d = 0.04, with evidence against an effect for information and assistance nudges.
- Hiding fees works, which is exactly why it is becoming illegal: a StubHub field experiment found upfront fee disclosure reduced both quantity and quality of purchases, and the FTC's fee rule took effect 12 May 2025 for live-event tickets and short-term lodging.
A price is a comparison, not a computation
People do not evaluate prices the way a spreadsheet does. They read the left-hand digits, anchor on whatever number is nearby, and stop. This is not a laboratory artefact. It is visible in some of the largest transaction datasets ever assembled, involving buyers spending real money on expensive things.
The strongest field evidence comes from the used-car market. Lacetera, Pope and Sydnor analysed more than 22 million wholesale transactions and found large, discontinuous drops in sale price at 10,000-mile odometer thresholds, with smaller drops at 1,000-mile marks. A car with 79,900 miles and a car with 80,100 miles are, mechanically, the same car. The market prices them differently because the leftmost digit changed. Their modelling attributes the pattern to partial inattention by final customers rather than to dealer behaviour, which matters: this is buyers, in a market with large financial stakes and fully observable information, systematically failing to read past the first digit.
The pricing tactic built on this is charm pricing, and it has been tested directly. Anderson and Simester ran three field experiments with a retailer in which price endings were manipulated, published in Quantitative Marketing and Economics in 2003. Use of a $9 price ending increased demand in all three experiments. Two qualifications in their results are more useful than the headline. The effect was stronger for new items than for items the retailer had sold in previous years, and it was weaker when the retailer simultaneously used sale cues. The authors' interpretation is the operationally important one: the $9 ending helps most when customers have incomplete information about the product and are using price format itself as a signal, which is also why no retailer applies it to everything.
The practical reading is narrow and honest. Price endings are cheap to change and worth getting right on new or unfamiliar items, especially where the customer has no reference price. They are not a growth strategy, they interact badly with discount signalling, and the effect is smallest exactly where your customers know your category best.
What survived contact with the real world
The single most reliable finding in this literature is not about persuasion at all. It is about structure: what is pre-selected, what is required, what order things appear in. Change the structure and behaviour moves; change the wording and it mostly does not.
The canonical demonstration is retirement saving. Madrian and Shea studied a large US corporation that switched its 401(k) from requiring employees to elect participation to enrolling them automatically unless they opted out, published in the Quarterly Journal of Economics in 2001. None of the economics changed. Participation was significantly higher under automatic enrolment, and, in the more interesting half, a substantial fraction of automatically enrolled participants kept both the default contribution rate and the default fund allocation, an outcome very few employees hired under the old regime had ever chosen for themselves. The authors attribute this to inertia and to employees reading the default as advice.
Two decades on, the practice has scaled and the numbers held. Vanguard reports that the share of plans adopting automatic enrolment grew from 10% in 2006 to 61% in 2024, across a base of nearly five million defined-contribution participants. Reporting on the 2026 edition of the same study, the National Association of Plan Advisors quotes the split directly: automatically enrolled employees had an overall participation rate of 94% in 2025, compared with 64% for employees in plans with voluntary enrolment.
Defaults hold up because they exploit something more durable than a cognitive quirk. Choosing has a cost (attention, deliberation, the small anxiety of being responsible for a decision) and a default lets people avoid paying it while still ending up somewhere reasonable. That mechanism does not fatigue and does not depend on the person failing to notice anything.
The operator's version is the same physics on a smaller stage: which plan is pre-selected, whether the annual option is shown first, whether a service renews automatically, whether the quantity field starts at one or at the case pack. The design constraint is ethical and commercial at once. A default the customer would have chosen anyway is a courtesy and reduces support load. A default the customer would not have chosen, that is hard to reverse, is a refund queue and, increasingly, a regulatory matter.
The findings that shrank
Three of the most repeated claims in consumer psychology are much weaker than their popular versions, and an operator who builds a pricing page on them is building on sand. This section exists because the honest version is more useful than the confident one.
Start with the decoy, or attraction effect: adding a deliberately inferior third option to make one of the original two look better. Frederick, Lee and Baskin tested its limits in the Journal of Marketing Research in 2014 and concluded that the phenomenon may be restricted to stylised representations in which every product dimension is expressed as a number, such as a toaster oven with a durability rating of 7.2 and a cleaning rating of 5.5. Such effects, they report, do not typically occur when consumers actually experience the product, or when even one attribute is represented perceptually, such as hotel rooms whose quality is shown in a photograph. They question the practical significance of the effect. Real pricing pages have logos, photographs, feature lists and prior brand knowledge. They are, almost by construction, the condition in which the effect fails.
Next, choice overload, the idea that offering more options reduces purchase. Scheibehenne, Greifeneder and Todd meta-analysed 63 conditions from 50 published and unpublished experiments covering 5,036 participants, in the Journal of Consumer Research in 2010, and found a mean effect size of virtually zero with considerable variance between studies. They identified some potential preconditions but no sufficient conditions. The correct summary is not "more choice is fine" and not "less is more." It is that the number of options is not, on its own, a lever with a predictable sign.
Then the broadest claim of all. Mertens and colleagues published a meta-analysis of choice architecture interventions in PNAS covering 214 publications, 455 effect sizes and 2,149,683 participants, reporting an overall Cohen's d of 0.45 with a 95% confidence interval of [0.39, 0.52], and noting moderate publication bias. Maier, Bartoš, Stanley, Shanks, Harris and Wagenmakers reanalysed the same dataset using robust Bayesian meta-analysis and reported a bias-corrected combined estimate of d = 0.04 with a 95% credible interval of [0.00, 0.14]. Their category-level results are the part worth memorising: strong evidence against an effect for information interventions and for assistance interventions, and undecided evidence for structure interventions.
Read those two papers together and a pattern falls out that matches everything above. The category that survives correction is structure: defaults, ordering, what is required versus optional. The categories that do not survive are the ones that try to persuade or to help people decide better. That is the same conclusion the retirement-savings evidence reaches from the other direction, and it is a genuinely useful prior for deciding what to build.
Calibrate your expectations before you run anything
The gap between a published effect and the effect you will get is not a rounding error. It has been measured, and it is roughly a factor of six.
DellaVigna and Linos assembled 126 randomised controlled trials covering more than 23 million individuals, run by two large US nudge units, and compared them with academic meta-analyses of the same class of intervention. Nudge-unit trials produced an average take-up effect of 1.4 percentage points, an 8.1% increase over control. Academic journal studies produced an average of 8.7 percentage points, a 33.5% increase. Their conclusion is blunt: publication bias in the academic journals, exacerbated by low statistical power, can account for the full difference. They also note that practitioners forecast outcomes better than academic researchers did.
This single comparison should reset how you plan a test. If you design an experiment with enough traffic to detect a 30% relative lift, and the true effect is 8%, you will run the test, see nothing, and conclude the change does not work, when in fact your instrument was never capable of seeing it. Underpowered tests do not produce cautious conclusions; they produce confident wrong ones in both directions, because the only results large enough to reach significance are the ones inflated by noise.
The reason the gap exists is worth understanding rather than just accepting. Academic studies are typically smaller, are run where the researcher expects an effect, and only reach print when they find one. Nudge units run everything they are asked to run, at population scale, and report the nulls. The nudge-unit number is the honest estimate of what a well-designed intervention does to a real population. The journal number is the honest estimate of what gets published.
There is one more calibration in the same direction. Effects measured on a captive administrative population, everyone who must file a form, transfer poorly to a commercial funnel where the customer can simply leave. Your baseline is not a 23% take-up rate on a government letter. It is a 2% checkout conversion on a page the visitor did not have to open.
Fee architecture: the strongest lever and the one with a legal boundary
The most powerful piece of choice architecture in commerce is not the decoy tier. It is where in the flow the mandatory fees appear, and the evidence on it is uncomfortable, because it shows that the profitable design is the deceptive one.
Blake, Moshary, Sweeney and Tadelis ran a large field experiment on StubHub, published in Marketing Science in 2021, comparing back-end fee disclosure with upfront all-in pricing. Making the full purchase price salient reduced both the quantity of goods purchased and the quality purchased. The effect on quality accounted for at least 28% of the overall revenue decline: shown the true total, buyers did not merely buy less often. They traded down. The effects persisted beyond the first purchase and were present even among experienced users, and click-stream data showed that obfuscation made price comparison difficult and led consumers to spend more than they otherwise would. Sellers, in turn, responded to the increased obfuscation by listing higher-quality tickets.
That is as clean a result as this literature produces, and it is precisely why the practice is being legislated out of existence. The Federal Trade Commission's Rule on Unfair or Deceptive Fees was published in the Federal Register on 10 January 2025 and took effect on 12 May 2025. It applies to live-event tickets and short-term lodging, and it specifies that it is an unfair and deceptive practice to offer, display or advertise any price for those products without clearly, conspicuously and prominently disclosing the total price. It also requires certain disclosures before a consumer consents to pay, and prohibits misrepresenting the nature, purpose, amount or refundability of a fee.
The boundary is moving outward. In March 2026 the Commission published a proposed rule on unfair or deceptive rental housing fee practices, addressing advertised rent that omits mandatory fees, fees imposed without express informed consent, and misleading descriptions of what a fee is for. In April 2026 it published a proposed rule covering fees and charges for food and grocery items ordered through online delivery platforms. Both are proposals soliciting comment, not final rules, but the direction of travel is unambiguous, and it points at exactly the categories where drip pricing is currently most profitable.
The operator's position should therefore be strategic rather than reluctant. If a meaningful share of your revenue depends on the customer not seeing the total until checkout, you have a revenue line with a regulatory expiry date, and the StubHub result tells you roughly what removing it costs: fewer orders, and a mix shift down. Better to run that transition on your own schedule, price the total honestly, and rebuild the lost margin in the base price, where it is defensible, than to have it removed for you in a compliance sprint.
Testing this on your own traffic without fooling yourself
Everything above is a prior. The only thing that settles a question about your customers is a test on your customers, and most operators do not have the traffic to run the test they think they are running. It is better to know that before you spend two months on it.
Illustrative only, with round numbers chosen for legibility. A software tool has three plans at $19, $49 and $99 a month. One thousand visitors reach the pricing page each month and 6.0% buy, so 60 sales. The mix is 55% on $19, 35% on $49 and 10% on $99, giving average revenue per new customer of $37.50 and $2,250 of new monthly recurring revenue.
You pre-select the $49 plan and label it recommended. Suppose the mix moves to 40% / 48% / 12%. Average revenue per new customer rises to $43.00, and at the same 6.0% conversion that is 60 sales and $2,580, a 14.7% gain. But a default that pushes people up the menu can also push people out of it. If conversion falls from 6.0% to 5.5%, you get 55 sales at $43.00, or $2,365: still ahead, by 5.1%. Now find the break-even. At $43.00 per sale you need $2,250 ÷ $43.00 = 52.3 sales, which is a 5.23% conversion rate. The entire $5.50 gain in average revenue is erased by 0.77 percentage points of lost conversion.
Here is the part that decides whether you can run this as an experiment. Detecting a move from 6.0% to 5.2% conversion, at 80% power and a 5% two-sided significance level, needs roughly 13,000 visitors per arm. At 1,000 visitors a month across both arms, that is more than two years. You cannot A/B test the conversion side of this change. You can measure the mix side more cheaply, because mix is a proportion among buyers and moves further, but with 60 buyers a month even that is slow.
So the honest operating procedure is layered. Where you have the traffic, test, and size the test for the effect you actually expect, closer to the nudge units' 8% relative than the journals' 33%. Where you do not, do not pretend: make the change deliberately, define the metric and the review date in advance, hold everything else constant, and compare a clean pre-period against a clean post-period while knowing that seasonality and any concurrent marketing change are confounds you have not controlled for.
Three failure modes are worth naming because they are common and expensive. Testing price on existing customers who can see both prices converts a pricing test into a fairness incident. Peeking at a running test and stopping when it looks good inflates your false-positive rate far above the 5% you nominally chose. And reading a revenue-per-visitor test the same way you would read a conversion test ignores that revenue is far noisier than a binary outcome, so it needs a larger sample, not a smaller one, to say anything at all.
Put it to work
Audit your pricing page this week: check what is pre-selected, what order the tiers appear in, and whether every mandatory fee is visible before checkout. Fix the fee disclosure first, since it is now regulated in lodging and ticketing. Then size one test properly before running it; if you cannot reach the sample, change the design and measure mix instead.
Sources & references
Linked entries open the named source directly. Entries without a link say exactly what kind of reference they are — and how to check them yourself.
- Lacetera, Pope & Sydnor, "Heuristic Thinking and Limited Attention in the Car Market" — NBER Working Paper 17030 (22 million wholesale transactions, price discontinuities at 10,000-mile thresholds)
- Anderson & Simester, "Effects of $9 Price Endings on Retail Sales: Evidence from Field Experiments," Quantitative Marketing and Economics 1(1): 93–110 (2003)
- Madrian & Shea, "The Power of Suggestion: Inertia in 401(k) Participation and Savings Behavior," Quarterly Journal of Economics 116(4): 1149–1187 (2001) — NBER Working Paper 7682
- Vanguard — How America Saves 2025: key trends and insights (automatic enrolment adoption 10% in 2006 to 61% in 2024)
- National Association of Plan Advisors, reporting Vanguard's How America Saves 2026 (94% participation under automatic enrolment vs 64% under voluntary enrolment, 2025)
- Frederick, Lee & Baskin, "The Limits of Attraction," Journal of Marketing Research 51(4): 487–507 (2014) — indexed record with abstract
- Scheibehenne, Greifeneder & Todd, "Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload," Journal of Consumer Research 37(3): 409–425 (2010)
- Mertens, Herberz, Hahnel & Brosch, "The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains," PNAS 119(1) (2022)
- Maier, Bartoš, Stanley, Shanks, Harris & Wagenmakers, "No evidence for nudging after adjusting for publication bias," PNAS (2022)
- DellaVigna & Linos, "RCTs to Scale: Comprehensive Evidence from Two Nudge Units," Econometrica (2022) — NBER Working Paper 27594
- Blake, Moshary, Sweeney & Tadelis, "Price Salience and Product Choice," Marketing Science 40(4): 619–636 (2021) — NBER Working Paper 25186
- Federal Trade Commission — Trade Regulation Rule on Unfair or Deceptive Fees, final rule published 10 January 2025, effective 12 May 2025
- Federal Trade Commission — Rule on Unfair or Deceptive Rental Housing Fee Practices, proposed rule published 13 March 2026
- Federal Trade Commission — Rule on Unfair or Deceptive Fees in Online Food Delivery Services, proposed rule published 16 April 2026
- Worked pricing-page example is illustrative — The $19 / $49 / $99 plans, 1,000 monthly visitors, 6.0% conversion and the two mix splits are made-up round numbers chosen to make the arithmetic legible. The $37.50 and $43.00 averages, the 14.7% and 5.1% gains, the 5.23% break-even conversion rate and the ~13,000 visitors per arm all follow from those figures and no others.
Educational note: This briefing is general business education, not financial, legal, tax, or investment advice. Figures and rules change and vary by situation — verify current specifics with primary sources and qualified professionals before acting.