Best AI for Medical Practice: A Job-by-Job Buyer's Map
The shift this year is from chatbots to agents that take actions
Within weeks of each other, two vendors shipped agents rather than chatbots. Fierce Healthcare reported that Assort Health rolled out an AI agent dedicated to automating referrals for specialty practices and health systems. Clearwave announced what it calls an agentic patient-engagement platform aimed at the industry’s most manual work.
What changed isn’t the underlying models so much as the packaging: vendors are shipping software that does multi-step tasks — call the payer, read the response, update the record, escalate the exception — rather than software that answers questions. That’s the working definition of an agent, and it’s why “which AI should we buy” is now a harder question than it was two years ago.
Six categories, six different decisions
1. Ambient clinical documentation. The most mature category: a scribe listens to the visit and drafts the note, the clinician edits and signs. Buy off-the-shelf — it’s a crowded, competitive market and building your own makes no sense. The caveats are real, though: drafts can omit findings or assert details that weren’t said, and accuracy degrades in noisy rooms, with multiple speakers, or when a visit is heavily interrupted. Whatever the draft says, the clinician’s signature carries the accountability, so review time is part of the cost, not an optional extra. See our vendor-neutral scribe comparison.
2. Front-office voice and text agents. Answering the phone, booking and rescheduling, reminders, waitlist backfill. This is where agentic vendors are pushing hardest. Quality varies enormously by how well the agent is wired into your scheduling rules (provider templates, appointment types, insurance restrictions) — a demo on a clean sandbox tells you very little.
Accessibility deserves its own line item here. Speech recognition still degrades on accented English, on older-adult speech patterns, and on speech disabilities, and an IVR-first flow can strand patients with limited English proficiency who would have been fine with a person. Build in a fast, obvious zero-out-to-human path, and track failures and opt-outs broken down by language and, where you can observe it, by accent — during the shadow pilot, not after go-live. Related: scheduling automation.
3. Revenue-cycle agents. Eligibility verification, claim scrubbing before submission, denial triage and appeal drafting. High-value, and increasingly available as add-ons inside practice management platforms. Deeper: automating medical claims and billing.
4. Referral and prior-auth packet prep. Assembling the clinical documentation a payer or specialist requires, in the same format every time. This is the best fit for the “skill” pattern below, and the category the Assort Health launch targets. See prior authorization automation with AI agents.
5. A general AI assistant connected to your own systems. Claude, ChatGPT, or a comparable assistant given governed access to your practice management data, scheduling, and documents — a general-purpose analyst your office manager talks to, not a point solution.
6. Custom agents. Built for a workflow no vendor sells, because it’s specific to your specialty, your payer mix, or your legacy system.
What MCP actually is, and why category 5 got interesting
MCP — the Model Context Protocol — is an open standard for giving an AI assistant secure, scoped access to a specific system’s data and tools. Instead of pasting a patient ledger into a chat window, you run an MCP server that exposes a small, defined set of operations: look_up_appointment, list_open_slots, get_claim_status. The assistant can only do what the server exposes, and every call can be logged.
A skill is the companion concept: a packaged, reusable instruction set that teaches the assistant to do one job the same way every time — “prepare a GI referral packet” — including which fields to pull, what format the receiving practice wants, and when to stop and ask a human.
Put together: your office manager asks the assistant to prep tomorrow’s referral packets. The assistant uses MCP to pull the chart data it’s permitted to see, follows the skill to assemble each packet, and drops drafts in a review queue. A human sends them. That last sentence is not optional.
An agent that can read your schedule is a productivity tool. An agent that can change your schedule without review is an incident waiting for a root-cause analysis.
Whatever platform you run — athenahealth, Epic, Dentrix — the practical question is what its API actually exposes, which is usually the binding constraint, not the AI. Our EHR integration primer covers that terrain.
Where ChatGPT and other consumer assistants fit — and don’t
ChatGPT is a generative AI product built on a large language model. It is not, by itself, a clinical system.
On HIPAA: under the HHS Office for Civil Rights, a vendor that creates, receives, maintains, or transmits PHI on your behalf is a business associate, and you need a business associate agreement in place. Consumer-tier chat products are generally not offered under a BAA; some vendors offer BAAs for specific enterprise or API tiers. That varies and changes — verify against the vendor’s current documentation and your privacy officer before any PHI touches a tool. “HIPAA compliant” isn’t a property a chatbot has; it’s a property of your configuration, your BAA, and your access controls together. Our HIPAA-compliant AI workflow guide goes deeper.
A reasonable middle path: use a general assistant freely for de-identified operational work — drafting policies, summarizing payer bulletins, building schedule templates — and require a BAA-covered, access-controlled path for anything patient-specific.
Which roles actually change
This is opinion, stated plainly rather than dressed up as research. The tasks most exposed to agents are high-volume, rule-bound, and text-based: eligibility checks, appointment confirmations, claim status inquiries, first-pass denial triage, packet assembly. The least exposed work needs hands, judgment under ambiguity, or a relationship — clinical assessment, chair-side care, behavioral-health rapport, complex payer negotiation, the front-desk call about the upset patient in the lobby.
The usual framing is that this is a shift from doing the queue to supervising it, and for many practices it will be. But it would be dishonest to stop there: some practices will use these tools to leave an open front-desk seat unfilled, or to cut hours outright. If that’s the plan, own it — give staff advance notice and a real path into exception handling, quality review, and payer work that pays better than answering phones, rather than letting people find out the week it happens.
Will AI replace primary care doctors? No — and the framing misses the point. Diagnosis and treatment carry legal accountability that sits with a licensed clinician, and the FDA regulates certain clinical decision support software as a medical device. The realistic near-term change is that documentation and administrative burden shrink.
A 30-day way to decide
-
Pick one measurable queue
Not “AI for the practice.” One queue: outbound eligibility checks, inbound reschedule calls, or denials under a dollar threshold. Count today’s volume and touch time for two weeks first. -
Check what your existing platform already ships
Practice management vendors are adding agentic features fast. Free-with-your-license and mediocre often beats excellent-and-siloed. Ask for the feature list in writing. -
Run a shadow pilot
Let the tool draft, but keep a human sending. Track the correction rate — how often staff had to fix the output — because that number, not the demo, is the real efficiency signal. -
Model the economics with your own numbers
Hours saved per week × your loaded staff rate, plus revenue you capture that you were previously losing (filled cancellations, denials worked that used to be written off). Then subtract everything: subscription and setup, any interface or integration fee your PM vendor charges, the weekly staff time spent working the correction queue, and the BAA, security-review, and legal effort. Reviewer time is the line item that most often flips a positive model negative. Don’t accept a vendor’s ROI slide as your baseline. -
Only then ask whether to build
If no vendor covers the workflow, or the integration cost of three vendors exceeds one custom build, that’s when MCP and a custom server earn their place.
The card below is a formula card, not a data card — the first two values are placeholders you fill in with your own measurements, not benchmarks.
The practices that get the most out of this wave aren’t the ones that bought the most AI. They’re the ones that measured one painful queue, put an agent behind a human reviewer, and only widened the gate when the correction rate earned it. As of 2026, that’s still the whole playbook. For a broader sequencing view, start with the practice automation audit.
Not sure where to start?
Get a free automation audit: we map your scheduling, intake, insurance, billing, and patient communication and show you what's worth automating — before you spend a dollar.
Get a free automation audit