
Why Northwestern Medicine Chose Abridge as its AI Partner After a Rigorous Multi-Vendor Evaluation
With Hannah Koczka, Vice President, Ventures & Innovation – Northwestern Memorial HealthCare
Choosing a clinical AI platform is one of the most consequential technology decisions health systems are making today. For Northwestern Medicine (NM), the question wasn't simply which solution performed best. It was whether any platform could deliver enough measurable value to justify switching from an existing enterprise deployment.
To answer that question, NM conducted a rigorous, multi-vendor evaluation spanning clinical, operational, financial, and partnership outcomes.
In this conversation, Hannah Koczka, Vice President of Ventures & Innovation at Northwestern Memorial HealthCare, shares the evaluation framework NM used to compare ambient AI platforms, what metrics mattered most, and why the organization ultimately selected Abridge.
About Northwestern Medicine
In Conversation with Hannah Koczka
Northwestern Medicine evaluated multiple ambient AI platforms. Why was it important to take such a rigorous approach?
We evaluate AI the same way we evaluate any strategic investment—we want evidence. This wasn't simply about choosing a documentation tool. We already had ambient technology in place. The question was whether another solution could create enough additional value to justify making a change. NM was seeking a long-term strategic partner to collaborate with us and be part of our healthcare transformation journey.
Switching enterprise technology isn't something you do lightly. There are implementation costs, change management considerations, and switching costs for clinicians who have already built workflows around an existing platform. If we're going to ask thousands of clinicians to change the way they work, we need clear evidence that the long-term value outweighs the disruption.
That's why we built an evaluation framework that looked across multiple dimensions. Clinical experience matters. Financial outcomes matter. Operational impact matters. And just as importantly, the partnership matters.
I actually think one of the most valuable things we can share is how we made our decision. Rather than relying on demonstrations or anecdotal feedback, we measured objective outcomes and paired those with structured clinician feedback to determine which solution would have the greatest long-term impact.
What metrics did Northwestern evaluate?
We intentionally looked well beyond documentation quality. Our evaluation framework included four broad categories:
But the numbers only told part of the story. We also gathered structured feedback from clinicians throughout the evaluation. Some of our clinicians had exposure to just one platform; others used all three side by side. That gave us a direct, head-to-head comparison rather than relying on separate, disconnected pilots.
Ultimately, it was the combination of quantitative performance data and frontline clinician feedback that gave us confidence in the results.
What ultimately differentiated Abridge?
We also observed stronger physician productivity, improvements in RAF score optimization, and importantly, no increase in denial rates. Maintaining documentation quality while improving efficiency was a critical part of our evaluation.
Just as importantly, clinician feedback aligned with the data. Many physicians described the notes as more natural, organized, and requiring less editing than other solutions they had used. Several also commented that the product improved throughout the evaluation based on clinician feedback, reinforcing confidence in both the technology and the partnership.
What did clinicians tell you during the evaluation?
The clinician feedback closely mirrored what we were seeing in the quantitative data. One physician shared that Abridge cut down on charting time, captured histories articulately, and is extremely efficient at improving patient notes access.
Another told us: “My production and efficiency increased rather dramatically when using the AI. This improved not only my daily work satisfaction but also many of my patients' satisfaction, as they felt more of their history was captured in the note, which they read via their MyChart.”
When you're seeing measurable operational improvements alongside that level of clinician enthusiasm, it's a powerful validation that the technology is making a meaningful difference.
One physician assistant in Otolaryngology expressed, “I liked being able to give the patient undivided attention without needing to stress over taking notes and writing down important information. I feel like they make the history much more consistent with the patient’s report.”
A Primary Care clinician stated, “It just works. Impressed with how it does generate an accurate summary and organized HPI and A&P, easy and quick to learn.”
Partnership was also part of your evaluation. Why?
Technology is only part of a successful implementation. We evaluated how each company partnered with Northwestern Medicine just as carefully as we evaluated the product itself.
We looked at responsiveness to feedback, collaboration with our clinicians, executive engagement, communication quality, governance, and whether partners consistently followed through on the feedback we provided. Those things become incredibly important because enterprise AI isn't a one-time purchase—it's an ongoing collaboration.
We also spent considerable time evaluating the product roadmap. We weren't looking for an ambient documentation vendor. We were looking for a long-term AI partner. A lot of the earlier generation of clinical AI was narrowly focused on quality — is the note better, is it more accurate. That's valuable, but it's not the whole picture. We're just as focused on productivity and on the patient experience, and I don't think those are the same thing as quality.
The question we keep coming back to is: now that documentation is better, what does that actually unlock? How do you build on top of that?
Northwestern seems to have a different philosophy around AI implementation. How do you approach adoption?
One lesson we've learned is that successful AI adoption begins long before deployment. Sometimes organizations identify a technology first and then ask clinicians to adapt afterward. We try to do the opposite.
We bring the right clinical stakeholders into the discussion from the very beginning. We're really only implementing technologies where we believe we've already built partnerships with our clinical leaders. That approach has been one of the biggest drivers of adoption across our innovation portfolio.
Today, after 3 months, we’re already reaching the point where clinician demand exceeds our available licenses. The conversation isn't whether clinicians want the technology—it's how we thoughtfully deploy licenses where they'll create the greatest value and ensure they're being used consistently.
Beyond documentation, what excites you most about the future?
Ambient documentation is really the foundation. What excites us is everything that can be built on top of that foundation. We're interested in how AI can support clinicians before the visit, during the encounter, and afterward.
That includes preparing clinicians before appointments, streamlining post-visit administrative work, supporting payer collaboration, identifying treatment opportunities, and helping match patients with clinical trials.
Today, most patients never get evaluated for a trial they might qualify for—not because they wouldn't be a fit, but because screening is manual and time-consuming for already-stretched clinicians. It shouldn't be 100%, but there's certainly a broad population that should be getting screened that currently isn't.
That's why the roadmap mattered so much during our evaluation. We weren't simply selecting the best ambient documentation technology available today. We were selecting a partner capable of helping us solve much larger challenges over time.
What advice would you give other health systems evaluating AI?
Start with the problems you're trying to solve—not the technology. Bring clinicians into the evaluation process early. Measure outcomes that matter across clinical, operational, financial, and user experience dimensions. Evaluate the partnership as carefully as you evaluate the product.
At NM, we've found that taking a rigorous, evidence-based approach upfront leads to stronger adoption, better implementation outcomes, and a much higher likelihood that successful pilots ultimately scale across the organization. When you do that well, you're not simply implementing software. You're building the foundation for long-term transformation.
This interview was edited for length and clarity.


