Evaluating Healthcare AI Vendors: Buy Outcomes, Not Software

A point-solution company approached the chair of vascular surgery at a health system with a tool that tracked aneurysms over time. The pitch landed. The clinical case was sound, the top-line case was reasonable, and the surgical group wanted it. When the request reached the system's AI governance group, though, the answer was neither yes nor no. It was a question: before onboarding another vendor, would the chair give the enterprise AI platform already under contract six to eight weeks to solve the same problem? The platform's forward-deployed team built the aneurysm-tracking workflow, the vascular group got what it needed, and by one estimate it arrived eight to twelve months ahead of what a fresh procurement cycle would have delivered.

Nishith Khandwala, CEO of Bunkerhill Health, told that story on a Becker's Healthcare panel of four healthcare AI founders, and it reframes the problem facing anyone assessing healthcare AI vendors. The hard question is rarely whether a given product works in a demo. It is whether you need a new vendor at all to get the outcome you are buying.

The Governance Question Behind Every Point Solution Pitch

Health systems are not short on options, and that abundance is itself the constraint. Steve Liou, founder of Clarium, described speaking with the CIO of a leading system who reported running roughly 1,800 different applications and wanting to consolidate aggressively. That figure covers the full application estate rather than AI tools alone, but it explains why the buying reflex has shifted. The question is no longer "does this work?" so much as "does this outcome justify another vendor relationship, another integration, another security review, and another line in the governance queue?" Liou expects a wave of point-solution consolidation over the next three to five years, with systems retaining their core systems of record and concentrating new spend on a small group of partners that can show return inside the first twelve to eighteen months.

The available evidence rewards that caution, although not in the way the headlines suggest. MIT's NANDA initiative, in a 2025 study of enterprise AI that circulated widely, found that roughly 95% of generative AI pilots produced no measurable impact on profit and loss. The more useful finding in the same study tends to get dropped: deployments built through partnerships with specialized vendors succeeded around 67% of the time, while internally built efforts succeeded roughly a third as often (MIT NANDA, 2025). Buying AI is not what fails. Buying AI without a defined, owned outcome is what fails, and that distinction should shape how you run an evaluation.

Breadth of a Platform, Depth of a Point Solution

Khandwala describes Bunkerhill's ambition as offering "the breadth of a platform but the depth of a point solution," and the phrase works better as an evaluation standard than as positioning. Point solutions genuinely do win on depth. Anyone who has spent time in early-stage software knows that a team aimed at one narrow problem will usually outperform a broader platform on that problem. Yet no health system wants to integrate, govern, secure, and support a thousand narrow tools, which is why nearly every successful point solution eventually grows into something wider. Both halves of that tension are true at once, and a good evaluation acknowledges it rather than pretending one side wins.

The practical question is which direction the vendor in front of you is traveling, and how honestly they can describe where they are. Sam Schwager, CEO of SuperDial, put the test bluntly: ask a vendor what its company is genuinely the best in the world at. If a vendor cannot answer very specifically, naming the workflow where it has more operating experience and more training data than anyone else, then the promise of fast time to value on your first use case is hard to believe. Vendor bloat is a real cost and consolidation is a reasonable goal, but a platform assembled from claims rather than from one proven core tends to be shallow everywhere.

What makes platform depth possible is less glamorous than the model architecture. Khandwala noted that none of Bunkerhill's forward-deployed staff are engineers by training. They are former startup operators and consultants who configure a composable platform for each new use case, which is how a platform reaches point-solution depth without becoming a consultancy. That detail points at the part of the purchase most evaluations underweight.

Forward Deployment Is Part of the Product, Not a Courtesy

Every founder on the panel converged on the same point from a different direction. Bo Gu, a former cardiothoracic surgeon and hospital executive who now leads Youlify, said forward deployment succeeds or fails on language: whether the vendor's team understands what a medical necessity denial subcategory is, what utilization management involves, and what a biller is actually doing inside a pivot table. He estimated that his team spends about 90% of its forward-deployment time listening rather than building.

Schwager described the same function as configuration rather than customization. SuperDial's forward-deployed engineers and RCM subject-matter experts meet customers to identify integration needs and specific behavioral requirements, then configure an existing base rather than build something bespoke. His framing of what is being sold is worth quoting directly: "We're not just delivering software; we're delivering outcomes and labor via these AI agents. In the same way you'd want new employees to be trained and ramp up on your systems and configurations, you require the same thing from an AI workforce."

Liou grounded it in your team's calendar. Every health system IT group is running a hundred projects alongside a major EHR or ERP upgrade, which means it does not have the bandwidth to implement a new platform. Clarium assigns a dedicated team of forward-deployed engineers and subject-matter experts to each onboarding and reports average deployment in eight to ten weeks. Whether or not that timeline generalizes, the underlying claim should be contractual rather than aspirational: implementation labor is something you are paying for, so it belongs in the statement of work with named owners and dates. In revenue cycle specifically, where the work is a dense mix of payer rules, denial subcategories, and system-of-record quirks, that domain requirement is especially unforgiving. Our companion piece on AI agents in revenue cycle management walks through what deployment looks like workflow by workflow.

Five Questions Worth Asking Before You Sign

What measurable outcome will you commit to, in writing?

Liou's recommendation is to demand that vendors show measurable return inside a predefined window and to walk away from anyone who will not commit to those results in writing. The commitment matters more than the number, because a vendor willing to put a figure and a date in a contract has usually done the work to know what it can deliver.

How quickly will we see it?

Gu observed that health system conversations rarely turn on who a vendor's investors are. What the CFO wants to know is how many dollars come back for each dollar spent, and when. Ask for the expected value curve by month, not just the twelve-month figure, and ask which milestone signals the deployment is off track.

Who handles implementation and change management?

Liou and Khandwala both put change management on the same footing as the technology itself. Ask who is assigned, whether they are employees of the vendor or subcontractors, what happens at four in the morning when something breaks, and who your team calls. Khandwala made the build-versus-buy version of this point sharply: you cannot page a researcher or a physician builder at 4 a.m. and ask them to restore a production system.

Can this expand past the first use case?

Gu's first question as a former health system executive was whether he could foresee a long-term future with a vendor: whether they could adapt as his pain points moved, build additional agents, and absorb work that business process outsourcers were doing. A vendor that can only ever do the thing you are buying today becomes a consolidation target in your own application estate within a few years.

Which of your healthcare customers can verify the results?

Schwager's version is specific enough to be useful: ask whether there are a handful of other leaders in your function that the vendor has delivered for and that you can call, ideally ones they have worked with for at least a year. This is the fastest way to find out whether a company has real domain expertise in its DNA or has arrived recently with a general-purpose model and a healthcare landing page.

Make Experiments Cheap and Reversible

Reference checks have a limit, and Khandwala was candid about it. In categories where the underlying capability is only a couple of years old, no vendor can honestly offer thirty customers with four years of history, because the capability did not exist. If you screen exclusively for long track records in genuinely new categories, you will screen out the entire category and wait three to five years for problems you could address now.

His alternative is to shift the risk control from diligence to structure. Set up your AI strategy so the cost of iteration is low and the speed of iteration is high, which in practice means avoiding three-to-five-year contracts at six or seven figures annually with no exit. Bunkerhill sells AI credits that health systems can direct toward whichever use cases they choose, so an experiment that does not work can simply stop consuming credits. The specific mechanism matters less than the principle: when every decision feels like a one-way door, an organization cannot afford to be nimble, and the pace of the technology punishes organizations that cannot move. Ask any vendor how a use case gets discontinued, and how much that costs you.

A Healthcare AI Vendor Evaluation Checklist

The purpose of a checklist is not to score vendors into a ranking. It is to make sure the parts of the purchase that are easy to leave vague, and that later determine whether the deployment works, get written down while you still have leverage. Take this into your next vendor conversation and fill in the middle column with their actual answers.

Evaluation area What to ask for What a strong answer includes Warning sign
Outcome definition The specific metric this deployment will move A named metric, a baseline, a target, and a measurement method both sides agree on Activity metrics (tasks touched, calls placed) offered instead of results
Time to value A month-by-month value curve with a go-live date A defined window, contractual willingness to be measured against it, and an early milestone that flags trouble "Most customers see value in the first year"
Core competency The one or two workflows the vendor is best in the world at A specific workflow backed by operating volume and proprietary data A capability list with no center of gravity
Forward deployment Named implementation owners and their backgrounds Domain experts plus technical staff, employed by the vendor, with a written project plan Implementation described as "lightweight" or handed to your IT team
Domain fluency A working conversation with the team who will deploy People who can discuss denial subcategories, utilization management, or your specific operational vocabulary unprompted Generalists who need your team to explain the workflow
Integration and writeback Read and write behavior with your systems of record Named EHR and ERP integrations, writeback support, and a clear data flow diagram Exports, spreadsheets, or a portal your staff must check
Auditability A sample of the artifact your team will actually review Structured output, source evidence, timestamps, and a retrievable audit trail Summaries that a human still has to verify from scratch
Data readiness The vendor's assessment of your data before deployment An honest scoping of what must be cleaned, unified, or enriched first Assurance that data quality will not matter
Escalation and support The on-call path and the human handoff rules A defined severity model, response times, and rules for when work routes to a person A shared inbox and best-effort response
Security and compliance Current attestations and a signed BAA HIPAA compliance, SOC 2 Type II or equivalent, and documentation your security team can review Certifications described as "in progress" without dates
Expansion path The second and third use cases, priced A roadmap tied to your priorities, with incremental rather than net-new commercial terms Expansion requiring a new contract and a new implementation
Exit terms How a use case gets discontinued Short initial terms, usage-based or credit-based pricing, and a documented offboarding path Multi-year lock-in with no termination for non-performance
References Customers in your function, live for a year or more Direct conversations with peers who own the same metric you do Logos on a slide, or references who piloted but never scaled

Buy Outcomes, Not Software

Liou's summary of what he tells executives comes in three parts, and they sequence well. Fix your data before deploying AI, because putting models on top of fragmented data produces expensive noise faster. Buy outcomes, not software, and walk away from vendors who will not commit to measurable results in a defined window. Separate signal from noise by concentrating on companies that have built real relationships with your peers and can produce case studies rather than demos.

Gu would add a fourth, and it is the one that survives contact with a live deployment. Trust in healthcare is not a handshake; it is whether a vendor tells you the truth when the metrics look bad, when an HL7 feed breaks, or when a cloud provider goes down. That is not a diligence question you can ask directly, since every vendor will answer it well. It is a reason to structure the first engagement small enough that you learn the answer cheaply, on a use case where being wrong is survivable.

The practical consequence for the next twelve months is that the evaluation and the contract should carry more of the weight than the demo. A demo tells you a product can work under conditions the vendor chose. A written outcome, a named implementation team, an auditable output, and a low cost of reversal tell you what happens under yours.

Sources

  • MIT NANDA Initiative. "The GenAI Divide: State of AI in Business 2025." Reported in Fortune, August 18, 2025. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
  • Becker's Healthcare Podcast. "Building AI Solutions That Work for Healthcare." Panel with Bo Gu (Youlify), Nishith Khandwala (Bunkerhill Health), Sam Schwager (SuperDial), and Steve Liou (Clarium), hosted by Scott Becker, August 12, 2026. https://podcast.show/beckershealthcarepodcast/beckershealthcarepodcast-526/
  • Bunkerhill Health. "Bunkerhill Health Raises $55 Million to Help Health Systems Turn Their Best Ideas into Reality." July 16, 2026. https://www.bunkerhillhealth.com/resources/series-b-announcement
  • Clarium. "Clarium Raises $27M Series A to Scale AI-Powered Supply Chain Resiliency Technology to Leading Health Systems." May 2025. https://www.clariumhealth.com/newsroom/clarium-raises-27m-series-a-to-scale-ai-powered-supply-chain-resiliency-technology-to-leading-health-systems
  • Centers for Medicare & Medicaid Services. "National Health Expenditures 2024 Highlights." CMS, 2026. https://www.cms.gov/files/document/highlights.pdf

Run a pilot on a real workflow.

Bring a representative batch, define the output schema, and validate ROI with your payer mix in 30 to 90 days.