*Editor’s note: This piece was reviewed by Dr. Heather Signorelli, DO, as physician-reviewed operational guidance. It is not medical, legal, or compliance advice, and NatRevMD does not endorse any specific AI vendor. Verify any workflow against your own HIPAA and payer obligations.
Every week someone asks us which AI tool they should buy for the practice. The honest answer is that the question is slightly wrong. There is no single best tool. There are jobs, and different tools are better at different jobs, and the smartest small practices use two of them for different reasons rather than betting the whole office on one.
So instead of crowning a winner, we ran the four tools most billing offices are actually considering against the work those offices actually do: drafting appeals, explaining denial codes, checking a payer policy, writing patient communications. Here is what held up. And before the comparison, the rule that overrides all of it: none of this involved a shred of PHI, and neither should your use, unless you have a signed BAA for the specific tool and tier. Everything below assumes de-identified inputs.
The jobs we tested against
We did not score these tools on trivia or creative writing. We scored them on the tasks a $150K-a-month practice repeats hundreds of times a month: turning a denial code into a plan, drafting an appeal skeleton, summarizing a public payer policy, rewriting a collections letter to sound human, and researching a current payer rule. Different jobs, different strengths.
The strong drafters: general-purpose assistants
The two big general-purpose assistants are the workhorses for drafting. Give them a de-identified denial reason and they produce a clean appeal skeleton, a patient-friendly explanation, or a rewritten letter faster than any human on your team. This is where most of your daily value lives.
The difference between them is a matter of feel more than capability. One tends to be a little more concise and businesslike out of the box; the other is often a little more careful and cautious about hedging when it is unsure. For billing work, both are more than good enough, and the honest recommendation is to try the same prompt in both and keep whichever one your team likes reading. The tool your staff actually enjoys using is the one that gets used.
Where all general-purpose assistants share a weakness: they will confidently invent a policy number or a “timely filing limit” that sounds authoritative and is completely made up. Treat every specific claim of fact as unverified until you check it against the payer’s own material. Use them to draft and structure, not to be the final source of truth on a rule.
The researcher: answer engines with citations
For the job of “what does this payer’s current policy actually say,” a citation-first answer engine is a different and useful animal. Instead of generating a fluent guess, it searches, pulls sources, and shows you where the answer came from. For payer research, that is exactly the behavior you want, because it lets you click through and confirm rather than trust.
It is less useful as a drafter. The output tends to be more report-like and less flexible than a general assistant when you want a warm patient letter or a nuanced appeal. So the pattern that works is to use the answer engine to find and verify the rule, then hand the confirmed facts to a general assistant to write the letter. Research in one, draft in the other.
The integrated option: the assistant living in your other software
The fourth category is the AI built into a productivity suite you may already pay for. Its advantage is not raw quality; it is location. If your office already lives in a particular email and documents ecosystem, having a capable assistant right there in the sidebar removes friction, and friction is the main reason staff stop using a tool. For drafting inside documents and email, it is a reasonable default simply because it is already open.
The same caution applies as everywhere else. Convenience is not compliance. The fact that a tool is bundled with software you trust does not mean the tier you are on will sign a BAA or keep your inputs out of training. Check the terms for the exact product and plan you have, not the brand in general.
What actually works: a two-tool stack
After all the testing, the recommendation for a typical independent practice is boring and effective. Pick one strong general-purpose assistant as your drafting workhorse, and one citation-first answer engine as your research and verification tool. That covers the overwhelming majority of billing and front-desk jobs. You do not need four subscriptions. You need one good drafter and one good fact-finder, and the discipline to use each for what it is good at.
A few honest caveats. Tool capabilities change fast, so treat any specific ranking as a snapshot rather than gospel. Free tiers behave differently from paid tiers on quality, limits, and data handling. And the biggest variable is not the tool at all. It is the prompt. A mediocre tool with a sharp, well-structured prompt beats a great tool with a lazy one, every time.
The part nobody wants to hear
The tool matters less than your process around it. The practices that get value from AI are not the ones that picked the “right” product. They are the ones that de-identify by default, keep a library of tuned prompts, and treat every AI output as a draft to be checked rather than an answer to be trusted. Get those habits right and any of the four tools will serve you. Get them wrong and the best tool on the market will still get you in trouble.
We keep our prompt library tuned across these tools so it works no matter which one your team lands on, and it is all built to run on de-identified inputs. If you want a head start on the process rather than the product, that is where to begin.
We test these tools the same way we run everything else in our approach to AI in medical billing: against real results.
Get the AI Kit → https://eligibility.natrevmd.com/natrevmd-ai-kit-tool


