Case Study
The AI Project That Died of Paperwork
What happened
IBM was contracted in June 2012 to deliver a Watson-based advisor for one disease in six months for a fixed fee of $2.4 million. About eight months in the scope grew to five more leukemia types, by amendment rather than a rebid. MD Anderson and IBM announced the Oncology Expert Advisor publicly in October 2013, and the scope later widened again to lung cancer, a different disease with different data and different clinicians. In 2016 the hospital replaced its medical records system, and the audit recorded flatly that the advisor had never been updated to work with it.
DOCUMENTED: the Oncology Expert Advisor (OEA), an IBM Watson-based clinical decision-support system built for the University of Texas MD Anderson Cancer Center, as recorded in the University of Texas System Administration System Audit Office's Special Review of Procurement Procedures Related to the UTMDACC Oncology Expert Advisor Project (November 2016), reported in detail by Forbes, The Register and PCWorld in February 2017. The audit deliberately declined to judge the technology, which is the whole point of this case.
- Real company — documented history
- Healthcare
- Enterprise AI program
- High risk
- Failure
- Advanced
The case, start to finish
Every amendment was small enough to sign. Nobody ever had to sign the whole number.
A six-month contract for one disease
In June 2012 IBM was contracted to build a Watson-based clinical decision-support system for the University of Texas MD Anderson Cancer Center. The scope was specific: one disease, myelodysplastic syndrome, in six months, for a fixed fee of $2.4M. The buyer was not a business but a capital project inside a nonprofit academic medical center, funded substantially by philanthropy, and the eventual user would be an oncologist asked to trust the output with a patient's life.
What the University of Texas System's audit office documented four years later, in a November 2016 special review, is one of the most useful failure records in enterprise technology, precisely because of what it refuses to say. Its scope was procurement, not science, and it states that its results should not be read as an opinion on the scientific basis or functional capabilities of the system.
Two good decisions that never met
The scope moved twice. About eight months in it expanded to five additional leukemia types, by amendment rather than a rebid or a re-baseline. Around two years in it expanded again to lung cancer, a different disease with different data and different clinicians. In October 2013 the project was announced publicly and the press treated it as the arrival of AI in cancer care.
Then in 2016 MD Anderson replaced its ClinicStation medical-records system with Epic. Read that sentence next to the previous paragraph and the whole case is visible. One team had committed to an advisor built on ClinicStation. Another team, doing necessary work for good reasons, replaced ClinicStation. Neither decision was wrong on its own. The failure was that no forum existed in which both were ever on the same agenda.
This is the part that generalizes furthest. Applied AI is mostly not reasoning. It is access: the system has to reach live records, parse them the way the organization actually stores them, and keep doing both while the organization changes its mind about storage. Every one of those dependencies belongs to somebody else, with their own priorities and their own good reasons to move. The project had an owner for the model and no owner at all for the question of what the model would still be able to read in two years.
How $62 million passes through a building unnoticed
By the time the audit published, the original $2.4M contract had been amended twelve times to $39.2M, and total project spend exceeded $62M, including over $21M to PricewaterhouseCoopers. The audit found six non-competitive procurements totaling $51.4M, two of them lacking formal justification and approval, and $11.59M spent from donated funds that had not yet been received.
Two findings explain the mechanism. Fees were, in the audit's words, consistently set just below the amount that would have required Board approval. And invoices were paid in full regardless of whether contracted services were delivered as agreed. A ceiling that only bites above a certain figure teaches people where the figure is, and across twelve separate amendments nobody was ever required to defend the whole number at once.
The ending should frighten anyone running a technology program. IBM ended support for the pilot and demonstration systems on 1 September 2016. The audit recorded flatly that the system had not been updated to integrate with the current records platform, that it was not in clinical use, and that it had not been piloted outside MD Anderson. IBM itself agreed the system was not ready for human investigational or clinical use. No patient was ever treated using it. After four years and more than $62M, nobody had learned whether the reasoning was any good, because it never got to read a real record.
Name the data, name its owner, re-baseline
For a smaller organization the specifics change and the shape does not. Before an AI or automation program starts, write down what data it reads, which team owns that data, and what that team's roadmap says for the next two years. Get the roadmap in writing, not because anybody is untrustworthy, but because otherwise the two plans will only ever meet by accident.
Then re-baseline. Each time the scope or the underlying systems change, recost the whole program from zero rather than pricing the increment. Scope creep is not really about scope, it is about never having to defend a total. If a number you would refuse today is a number you approved in eleven pieces, the approval process is decorative.
And notice what was missing that would have caught this early and cheaply. There was no paying customer, so nothing outside the building ever told the project it was drifting. When a program has no external feedback, internal governance is the only signal it gets, which makes routing around governance the most expensive kind of efficiency available.
Timeline
- June 2012 IBM is contracted to deliver a Watson-based advisor for ONE disease (myelodysplastic syndrome, a form of leukemia) in six months, for a fixed fee of $2.4M.
- ~8 months in Scope expands to five additional leukemia types. The contract is amended rather than rebid or re-baselined.
- October 2013 MD Anderson and IBM publicly announce the Oncology Expert Advisor. The press treats it as the arrival of AI in cancer care.
- ~2 years in Scope expands again, this time to lung cancer: a different disease, different data, different clinicians.
- 2016 MD Anderson replaces its ClinicStation medical-records system with Epic. The audit later records, flatly: “OEA has not been updated to integrate with the current system.”
- September 1, 2016 IBM ends support for the pilot and demonstration systems.
- November 2016 The audit publishes. The original $2.4M contract has been amended twelve times to $39.2M; total project spend exceeds $62M. “The system is not in clinical use and has not been piloted outside of M.D. Anderson.”
You're in the owner's chair
Two years into a $2.4M fixed-fee AI pilot that has already been amended nine times, your hospital's IT department announces it is replacing the medical-records system your AI reads. The AI team says they will “handle the integration later.” You control the budget. What do you do?
- Stop. Re-baseline the program against the new record system first.
- Approve the amendment, but cap it and make the vendor absorb the integration risk.
- Keep going — the model is the hard part and it is nearly done; integration is just plumbing.
One six-month contract, twelve amendments
- Original fixed fee: 6 months, one disease: 2.4 $M
- IBM contract total after 12 amendments: 39.2 $M
- Total project spend, all vendors: 62 $M
Figures per the University of Texas System Audit Office's November 2016 special review, as reported in February 2017. No single approval ever covered the whole number: each amendment only ever covered a slice, which is exactly how a number this size passes through a building unnoticed.
Business model
Not a business but a capital project inside a nonprofit academic medical centre, funded substantially by philanthropy, buying a system that would read a patient's record and recommend treatment. The customer was an oncologist who would have to trust it with a life, which sets a bar no demo can clear.
Revenue model
There wasn't one, and that is a finding rather than a footnote. No revenue meant no market feedback: nothing outside the building ever told this project it was drifting. The only available signals were internal governance signals, and those were precisely the ones the spending pattern routed around.
Cost structure
Over $62M across roughly four years: nearly $40M to IBM and over $21M to PricewaterhouseCoopers, against an original fixed fee of $2.4M. The audit found six non-competitive procurements totalling $51.4M, two of them lacking formal justification and approval as exclusive acquisitions, and $11.59M spent from donated funds that had not yet actually been received.
Strategic challenge
The hard part was never the model. A clinical decision-support system is a few percent reasoning and the overwhelming rest plumbing: it must read the institution's real records, in the institution's real format, kept current as the institution updates them. That plumbing belongs to a different department, with its own roadmap, its own budget and its own perfectly sensible reasons to change it. The AI project owned the reasoning. Nobody owned the sentence “this system reads a record platform we are about to retire.”
Key decision
Two decisions, made years apart by different people, never reconciled. One team committed to a Watson-based advisor built on ClinicStation. Another team, doing good and necessary work, replaced ClinicStation with Epic. Neither decision was wrong on its own. The failure was that no forum existed where both were on the same agenda.
What worked
Remarkably little, and the audit is scrupulous about why. Its scope was procurement, not science, and it says so: “Results stated herein should not be interpreted as an opinion on the scientific basis or functional capabilities of the system in its current state.” Nobody ever demonstrated that the model could not work. It simply never received a fair test: IBM itself agreed the system was “not ready for human investigational or clinical use, and its use in the treatment of patients is prohibited.”
What failed
Three failures, in ascending order of how much they should frighten you. SCOPE: one disease, six months, $2.4M became seven-plus cancers over four years at $39.2M through twelve amendments, with no re-baselining, so nobody was ever required to defend the whole number at once. GOVERNANCE: the audit found fees “consistently set just below the amount that would have required Board approval,” and that “invoices were paid in full regardless of whether contracted services were delivered as agreed upon.” A control that fires at a threshold does not teach discipline; it teaches the threshold. DATA: the electronic-health-record migration removed the ground the project was standing on. After four years and $62M, the system could not read the hospital's records.
Risk factors
An AI program whose inputs are owned by a team that does not report to it; fixed-fee pilots that amend their way into programs; approval thresholds that invite structuring rather than scrutiny; philanthropic funding spent before receipt; vendor and consultant economics that reward amendments; and a project with no paying customer to tell it, early and cheaply, that it is wrong.
Lesson summary
This project did not fail a benchmark. It failed a change-control meeting. Before an AI program starts: name the data it reads, name the team that owns that data, get their roadmap in writing, and re-baseline scope and cost from zero every time either one changes. The audit's own disclaimer is the lesson: you can spend over $62M and four years and still never find out whether the technology worked.
Key data
- $2.4M fixed fee, 6 months, one disease Original IBM contract
- 12, reaching $39.2M Amendments
- over $62M across ~4 years Total project spend
- over $21M Paid to PricewaterhouseCoopers
- $51.4M across six; two unjustified Non-competitive procurements
- $11.59M Spent from gifts not yet received
- zero — “not in clinical use” Patients treated using the system
Sources & basis
The company here is real and named, and nothing about it was invented to make the story land. The list below is where each fact came from — public filings, court records, published reporting — so you can open a source and check it against the sentence that used it.
- University of Texas System Administration, System Audit Office: Special Review of Procurement Procedures Related to the UTMDACC Oncology Expert Advisor Project (November 2016). The primary source for every figure and quoted finding in this case; the UT System document portal blocks automated retrieval, so the audit language below is quoted via the outlets that reported it.
- The Register, February 20, 2017 — quotes the audit findings on scope, invoice payment and fees set below the Board approval threshold View source ↗
- PCWorld, February 2017 — source for the June 2012 contract and its twelve extensions, the leukemia-to-lung-cancer scope expansion, the $51.4M in non-competitive procurements, the Epic integration finding and the “not in clinical use” language View source ↗
- Forbes (Matthew Herper), February 19, 2017 — source for the ~$39.2M/$21.2M vendor split and the audit's disclaimer on scientific and functional capabilities View source ↗