Hospitals Need Two AI Ledgers: Value and Failure
If a hospital can count AI hours saved but cannot count errors, near misses, and patient concerns, its oversight system is incomplete.
Hospital AI business cases increasingly arrive with a value estimate: hours saved, labor avoided, or throughput gained. The missing counterpart is a consistent count of AI errors, near misses, and patient concerns.
Our position is simple: a hospital needs two AI ledgers, kept with equal discipline. One counts value: cost, return, and the conditions required to break even. The other counts failure: errors, near misses, failed escalations, performance drift, and patient concerns. The first ledger governs capital allocation and operating performance. The second protects patients, preserves trust, and creates evidence-based triggers for pausing a system that is not working. Run only the first and you are approving investments with half the evidence.
The value ledger
Accuracy claims and hours-saved claims do not establish financial value. FLARE, an August 24 arXiv preprint, combines activity-based costing, fuzzy logic, and ROI analysis to model human verification, infrastructure, and recurring operating costs. In its illustrative AI-assisted stroke-imaging pathway, break-even occurred at roughly 3,992 patients per year and first-year ROI turned positive near 5,000. Those are modeled estimates, not results achieved by a hospital. The lesson is more important than the numbers: viability depended on patient volume, verification burden, infrastructure, and workflow design. Accuracy alone did not decide the economics.
Measurement is starting to institutionalize. On August 24, HIMSS and five Asia-Pacific health systems signed founding agreements to develop an AI outcomes framework to measure AI's impact, value, and return on investment across institutions. It is a commitment to build a standard, not a finished one: formal launch is planned for 2027 and peer-reviewed publication for 2028. Your board should not wait for it.
Until then, require eight inputs in every AI business case:
- Current workflow cost
- Addressable labor or expense
- Acquisition and integration cost
- Training and change-management cost
- Human verification requirements
- Monitoring and remediation cost
- Minimum volume or utilization needed to break even
- Verified post-deployment financial results
If a proposal cannot supply them, it is a demo, not a business case.
The failure ledger
On August 25, ECRI expanded its Problem Reporting Network to capture AI errors, malfunctions, misleading outputs, and near misses in patient care. In its survey of 124 respondents, mostly quality, safety, risk, and compliance leaders, 31 percent reported encountering a suspected incorrect or misleading AI output, 9 percent said an error reached a patient or influenced care, and 35 percent were unsure whether they had encountered an AI error. The sample is small and perception-based, not a national harm rate. It still exposes a surveillance problem. Hospitals count medication errors and falls because a reporting path exists. AI needs one that is equally clear.
The first move may not require a new platform. Add an AI category to the safety-event reporting system you already run, then define routing, ownership, and escalation. Give staff incident types they can recognize:
- Incorrect or fabricated information, or a clinically significant omission
- Failed escalation or unsafe autonomous action
- Automation bias or improper data access
- Incorrect routing or task execution
- Performance deterioration or disparate performance across patient groups
The failure ledger also has to hear from patients. A systematic review in Nature Health examined 330 medical AI studies. Satisfaction and perceived benefit were commonly assessed, but only 16.7 percent examined trust, only 10.9 percent examined safety, and just 3.9 percent considered patient factors during design and development. Those percentages describe what researchers measured, not how deployed systems perform. The imbalance matters: trust and safety receive far less attention than satisfaction and perceived benefit.
Patients have expectations regardless. In a Pew Research Center survey of 3,488 U.S. adults, 72 percent said disclosure of healthcare AI use is extremely or very important. Eighty-one percent wanted disclosure when AI analyzes medical scans or makes a diagnosis, and 72 percent said the same for AI note-taking. Meanwhile, 53 percent felt they had little or no say over AI in their care, and 46 percent did not know whether it had already been used on them.
These are attitudes, not outcomes. The failure ledger is not only a bug log. It also records failures of the operating model, including disclosure, escalation, and accountability. An AI use that materially affects care but is not disclosed, or a concern that never gets logged, should be treated as an oversight failure even when the algorithm performs as designed. Matching disclosure to the consequence of each use case is an operating-model recommendation, not a statement of legal requirements.
Governance is how the value ledger stays honest
The standard objection is that measurement taxes the return AI was supposed to deliver. We hold the opposite view. The failure ledger is what makes the value ledger believable. Rework, remediation, workarounds, quiet abandonment, and eroded patient trust are real costs, and they arrive in next year's operating budget whether or not anyone counted them. A hospital that counts failures is not slowing its AI program down. It is finding those costs early, and earning the credibility to defend the investments that survive.
Five questions for your next AI governance meeting
- Do we know the complete lifecycle cost of every production AI system?
- Have we established measurable financial and operational baselines?
- Can employees report an AI error or near miss through our existing safety system?
- Do our disclosure and consent practices reflect the consequences of each AI use case?
- Who has the authority to pause or remove an AI system after deployment?
If the answer is yes to all five, both ledgers are taking shape. If you can count the hours saved but not the errors, near misses, or patient concerns, the system is incomplete, and the gap will surface on its own schedule, not yours.
If your team is building or revising its AI oversight process, reply and we will send you the Two-Ledger checklist.
Sources
- FLARE preprint (Idoko, Paudel, Bento, Souza, Ginde; arXiv, August 24, 2026): https://arxiv.org/abs/2608.23643
- HIMSS AI Outcomes Framework MOU announcement (HIMSS, August 24, 2026; himss.org page is access-limited to automated retrieval): https://www.himss.org/news-center/himss-signs-two-memorandums-of-understanding-at-himss26-apac-advancing-global-collaboration-to-improve-care/
- First-party verification of the HIMSS announcement (Healthcare IT News, a HIMSS company, August 24, 2026): https://www.healthcareitnews.com/news/asia/himss-announces-initiative-measure-impact-ai-healthcare
- ECRI, "ECRI Expands Problem Reporting Network and Calls for More Data on AI Errors in Patient Care" (August 25, 2026): https://home.ecri.org/blogs/ecri-news/ecri-expands-problem-reporting-network-and-calls-for-more-data-on-ai-errors-in-patient-care
- Nature Health, "Patient factors in medical artificial intelligence: a systematic review" (August 26, 2026): https://www.nature.com/articles/s44360-026-00190-2
- Pew Research Center, "Americans want transparency when AI is used in their healthcare" (Yam, Pasquini, Kikuchi; August 25, 2026): https://www.pewresearch.org/short-reads/2026/08/25/americans-want-transparency-when-ai-is-used-in-their-healthcare/
