Case Study

The Automation That Only Worked Once It Was Cut to One Form

What happened

Leadership approved an AI assistant for the whole new-business submission process: read the application, summarise the loss runs, draft the carrier email. Nine months and $410,000 later it demoed well and could not be trusted with live work, so the project was stopped. An operations manager then spent two weeks with a stopwatch on 220 inbound certificates of insurance, and the second attempt shipped in three months for $38,000: read one document type and post six fields. The weekly audit caught eleven certificates posted with the wrong liability limit, from a carrier whose scanned template put two fields in an order the model had learned to expect.

Anonymized composite: a 70-person US commercial-lines insurance brokerage with four offices, drawn from a pattern common to back-office automation. The operating figures are the composite’s own illustrative arithmetic and are labeled as such throughout; the industry context is real and cited: adoption rates from the US Census Bureau’s Business Trends and Outlook Survey, and the certificate form itself from ACORD, the insurance industry’s standards-setting body.

  • Illustrative composite — not a real company
  • Insurance brokerage
  • Commercial-lines brokerage
  • Moderate risk
  • Turnaround
  • Advanced

The case, start to finish

The first project could not fail, because nobody had written down what succeeding would look like.

Four capabilities, one budget

This is an anonymized composite: a 70-person US commercial-lines insurance brokerage with four offices. The operating figures are the composite's own illustrative arithmetic; the industry context is real and cited.

Leadership approved an AI assistant for the whole new-business submission: read the client application, summarize the loss runs, draft the carrier email, answer the account manager's questions. Four capabilities, one project, one budget. Over nine months $410,000 went out in vendor licenses, an integration contractor and two senior account managers' time. The assistant demoed well every month and never reached production.

The reason it failed was not accuracy. Nobody had written down what working would look like, so nobody could say whether it was working. Underneath sat a defect the model could never fix: the assistant's output had to be reviewed by a licensed person before it went to a carrier, so the check became the job and a faster draft shortened nothing. A workflow whose bottleneck is verification cannot be sped up by generating more, faster.

Two weeks with a stopwatch

In month ten the project was stopped. Before anything else was built, an operations manager spent two weeks timing 220 inbound certificates of insurance by hand, producing the number the first project never had: 6.5 minutes of handling each, 2,400 a month, 260 staff-hours.

That reading only turns into money because of a second number most firms never establish. In the composite, the assistants who handle certificate traffic cost about $34 an hour fully loaded: salary, payroll taxes, benefits, software seat, a share of the floor. The nine-month project never worked that out, which is why its business case could not be settled by arithmetic.

Note what it could never be. A brokerage earns commission on placed premium, so automation could only show up as cost saving or capacity, never revenue. Any business case promising revenue was unfalsifiable from the start.

Attempt two shipped in months eleven to thirteen: read one document type, the ACORD 25 certificate of liability insurance, and post six fields into the agency management system. $38,000 to build, $1,900 a month to run. Anything the model was unsure of went to a human queue with the PDF open at the right page. It was chosen not because it was the most valuable thing on the list, but because it was the only one whose success or failure could be settled within a month.

The configuration that looked best was the wrong one

In month fourteen the weekly audit caught eleven certificates posted with the wrong liability limit. One carrier sends a scanned template where the each-occurrence and general-aggregate boxes sit side by side, and at a flat 0.80 threshold the model was confidently wrong.

So thresholds were set per field: 0.97 for dollar limits, 0.90 for dates, 0.85 for the named insured, and anything on an unseen template goes to the queue whatever the score. The auto-post rate fell from 84% to 71%. That fall is the fix, not a regression.

At 0.80 flat, the system used only about 23.4 staff-hours a month, a better headline than the version they kept, and it put eleven wrong liability limits into client files. In insurance that is a coverage dispute waiting for a claim. The safe configuration costs about 14 hours a month more: 71% posts automatically, the remaining 696 take 2.9 minutes each because the fields arrive pre-filled and the human is checking rather than typing, plus 4.1 hours of audit sampling. That is 37.7 hours against 260, roughly 222 hours saved a month, about $7,600 at $34 an hour or $5,700 net of running cost, and it cleared the $38,000 build cost just under seven months from go-live.

The hours have to be claimed in advance

The detail most likely to be skipped is why the saving was real. Nobody was made redundant: the 222 hours a month were assigned before launch to pulling renewal preparation forward. Hours nobody has claimed do not become money. They become slack, and a year later no line on any statement shows the project worked.

So the recipe is boring on purpose. Pick one document, one direction, a handful of fields. Measure the before with a stopwatch, and know your fully loaded hourly cost. Route everything uncertain to a person with the source open at the right page. Set thresholds per field, and keep the audit sample once the system works. Then decide in writing, before go-live, what the saved hours are for.

You are not scaling down your ambition. You are buying the ability to tell whether you are right, which the $410,000 version never had in nine months.

Timeline

  • Month 0 Leadership approves "an AI assistant for the whole new-business submission": read the client application, summarize the loss runs, draft the carrier email, answer the account manager’s questions. Four capabilities, one project, one budget.
  • Months 1–9 $410,000 goes out the door on vendor licenses, an integration contractor, and the internal time of two senior account managers. The assistant demos well every month and never reaches production. The killer is not accuracy: nobody had written down what "working" would look like, so nobody can say whether it is working.
  • Month 10 The project is stopped. Before anything else is built, an operations manager spends two weeks with a stopwatch on 220 inbound certificates of insurance and establishes the number the first project never had: 6.5 minutes of human handling each, 2,400 a month, 260 staff-hours.
  • Months 11–13 Attempt two ships: read one document type, the ACORD 25 certificate of liability insurance, and post six fields into the agency management system. $38,000 to build, $1,900 a month to run. Anything the model is unsure of goes to a human queue with the PDF already open at the right page.
  • Month 14 The weekly audit catches it: 11 certificates posted with the wrong liability limit. One carrier sends a scanned template where "each occurrence" and "general aggregate" sit in adjacent boxes, and the model was confidently wrong at a single 0.80 threshold.
  • Month 15 Thresholds are set per field: 0.97 for dollar limits, 0.90 for dates, 0.85 for the named insured. Any certificate on an unseen template goes to the queue whatever the score. The auto-post rate falls from 84% to 71%. That fall is the fix.
  • Month 20 Cumulative net saving passes the $38,000 build cost, just under seven months from go-live. Nobody is made redundant; the 222 hours a month were assigned before launch to pulling renewal preparation forward, which is the only reason they turned into anything.

You're in the owner's chair

Nine months and $410,000 into an AI assistant meant to handle the whole submission workflow, it still is not in production. The demos are genuinely better every month. Your board wants a decision on Friday. What do you bring them?

  • Extend six months — the integrations are built and the model keeps improving
  • Keep the scope and switch vendors — the concept is right, this implementation is wrong
  • Kill it, spend two weeks measuring one task by hand, then automate that task alone

Staff-hours a month spent getting certificates into the system

  • Before — every certificate handled by hand: 260 hours/month
  • After the $410,000 whole-submission assistant: 260 hours/month
  • Six-field extractor at a flat 0.80 threshold: 23.4 hours/month
  • Six-field extractor with per-field thresholds — the version kept: 37.7 hours/month

Illustrative composite figures. The second bar is unchanged because that project never reached production. The third bar is the one to distrust: it is the cheapest and it is the configuration that posted 11 wrong liability limits into client files. The version they kept is deliberately the second-best number on this chart.

Business model

A brokerage sells advice and access: it places commercial insurance with carriers on behalf of businesses, and lives on commission. Everything else it does is administration that the client will not pay extra for and cannot be asked to skip. That is exactly the kind of work automation should be aimed at, and exactly the kind nobody has ever timed.

Revenue model

Commission on placed premium, plus fees on larger accounts. Automation moves nothing on this line. That matters: this project could only ever show up as a cost saving or as capacity, never as revenue, so any business case that promised revenue was going to be unfalsifiable from the start.

Cost structure

People, almost entirely. In the composite, the service assistants who handle certificate traffic cost about $34 an hour fully loaded: salary, payroll taxes, benefits, software seat, and a share of the floor. That single number is what turns a stopwatch reading into money, and it is the number the nine-month project never bothered to establish.

Strategic challenge

The first project failed for a reason that had nothing to do with the model. Its output had to be checked by a licensed person before it went to a carrier, so the check became the job, and a faster draft did not shorten the check. A workflow whose bottleneck is verification cannot be sped up by generating more, faster.

Key decision

Stop, measure, and then pick the smallest task with a countable before-number. The six-field extractor was chosen not because it was the most valuable thing on the list but because it was the only thing on the list whose success or failure could be settled by arithmetic within a month.

What worked

Illustrative composite arithmetic: 2,400 certificates a month at 6.5 minutes each is 260 staff-hours. Afterwards, 71% post automatically; the remaining 696 take 2.9 minutes each because the fields arrive pre-filled and the human is checking rather than typing, a total of 33.6 hours. A 5% audit sample of the auto-posted batch adds about 4.1 hours. Total: 37.7 hours against 260, a saving of roughly 222 hours a month, about $7,600 at $34 an hour, or about $5,700 after the $1,900 running cost.

What failed

The first configuration, and it failed in the most flattering direction. At a flat 0.80 confidence threshold the system auto-posted 84% of certificates and used only about 23.4 staff-hours a month, a better headline than the version they kept. It also put 11 wrong liability limits into client files, which in insurance is not a rounding error but a coverage dispute waiting for a claim. The safe configuration is 14 hours a month more expensive and it is not close to a hard decision.

Risk factors

A pilot judged on auto-post rate rather than on errors escaping to clients; carriers changing a template without telling anyone, which turns a solved document back into an unsolved one overnight; the audit sample being quietly dropped once the system "works"; and the freed hours evaporating into slack because nobody decided in advance what they were for.

Lesson summary

Narrow scope is not a modest version of a big project; it is a different and better project, because narrowness is what makes the result checkable. Pick one document, one direction, and a handful of fields, measure the before with a stopwatch, and route everything uncertain to a person with the source open. Then decide, in writing and before go-live, what the saved hours will be spent on. Hours that nobody has claimed do not become money.

Key data

  • $410,000 over nine months, never in production First attempt
  • $38,000 to build, $1,900/month to run Second attempt
  • 2,400 a month Certificates handled
  • 6.5 minutes each — 260 staff-hours/month Measured handling time, before
  • 84% (unsafe) → 71% (kept) Auto-post rate
  • 11 Wrong liability limits posted before the audit caught them
  • ~222 a month, ~$5,700 net at $34/hour Staff-hours saved
  • Just under 7 months Payback on the build
  • 33.9% (Census BTOS, May 2026) US finance and insurance firms reporting AI use

Sources & basis

The business in this story is a stand-in, not a company you can look up. This case is an illustrative composite: the operator, the people and most of the dollar figures represent a pattern rather than reporting one firm's history. What the list below cites is the other half, the documented industry data and public reporting the composite was assembled from, including any real company whose published figures the case draws on by name. The mechanism and the arithmetic are real even where the business is not.

  1. U.S. Census Bureau, Business Trends and Outlook Survey — AI use by businesses, data collected December 14, 2025 to May 3, 2026: overall use between 17% and 20% of firms, 33.9% in finance and insurance, 37% among firms with 250+ employees View source ↗
  2. ACORD (Association for Cooperative Operations Research and Development) — the non-profit, industry-owned global standards-setting body for insurance, publisher of the standardised forms used across the industry since 1970 View source ↗
  3. Illustrative composite — the volumes, timings, costs, thresholds and savings above are the composite’s own arithmetic and describe no single identifiable firm