HomeBlogAI for private practice
AI CPT coding in physical therapy and where the line is
CMS blames 1.5% of improper payments to PTs on incorrect coding and 88.6% on thin documentation. That gap changes what a coding AI is actually for.

The short version
- CMS traced 88.6% of improper payments to PTs in private practice to insufficient documentation and 1.5% to incorrect coding. The code is rarely what failed.
- AI is dependable on the arithmetic: total timed minutes against units billed, how units split between codes, and edit pairs that need a modifier.
- Asked to generate codes from notes, ChatGPT-4 matched human coders exactly in 14% of 191 surgical cases and failed to apply modifiers at all.
- In that same study the models underbilled against human coders. An AI that promises to find your missed revenue is promising something the evidence has not shown.
- The coding questions are case questions. A scribe whose only input is one visit's audio cannot answer them; one writing inside the record already holds the plan of care, the payer and the prior episodes.
Medicare’s own error data carries an inconvenient number for anyone selling AI CPT coding to a physical therapy practice. Incorrect coding accounted for 1.5% of improper payments to physical therapists in private practice. Insufficient documentation accounted for 88.6%. Both sit on CMS’s compliance page for PTs in private practice, refreshed in February 2026 for the 2024 reporting period, next to a 15.8% improper payment rate and a projected $659.2 million.
Put those two percentages side by side and the pitch inverts. The code your therapist picked was rarely the problem. The record behind it was. So the useful question is not whether AI can pick a code. It is whether AI can tell you, before the claim goes out, that the note supports the code on it.
Can AI choose your CPT codes?
It can produce a suggestion. It should not be the thing that decides, and that is not squeamishness about robots.
A CPT code is the compressed form of a decision your therapist already made in the room: what impairment she targeted, what skill the intervention took, how the patient responded. The software was not there. It reads what got written down, a lossy copy of the visit, and works backward. When the note is thin, a code generator does not stall. It guesses fluently.
The measured results are unflattering. A University of Cincinnati otology team ran 191 operative notes through ChatGPT and compared the output to what human coders had actually billed, publishing in the Annals of Otology, Rhinology and Laryngology in May 2026. ChatGPT-3.5 matched exactly in 22% of cases, ChatGPT-4 in 14%. The pattern inside those averages is the part worth stealing. On cochlear implantation, the high-volume, unmistakable one, sensitivity ran 94% and 96%. On cartilage grafting, a finer distinction in the same notes, it was 4.2% and zero. Both models, the authors write, “failed to apply modifiers.” A Mount Sinai group had benchmarked the same class of tool two years earlier in NEJM AI, under a title needing no summary: “Large Language Models Are Poor Medical Coders”.
Those are surgical notes, not your Tuesday caseload, and an otology sensitivity figure is not a PT number. Read the shape instead: strong on the obvious call, collapsed on the fine one. In outpatient PT, the fine one is where every argument with a payer lives.
Where AI CPT coding earns its keep
On the arithmetic. Deterministic, unforgiving, and the part humans get wrong at 5:40pm.
Medicare’s unit math is a lookup table. The Medicare Claims Processing Manual, chapter 5 sets one unit at 8 through 22 timed minutes, two at 23 through 37, three at 38 through 52, and caps the day’s units by total treatment minutes. Run two codes past 15 minutes each and the manual says to give more units to the service that took the most time. CMS spells out the classic trap in its outpatient rehabilitation documentation fact sheet: four distinct 8-minute treatments total 32 minutes, which supports two units, not four. The NCCI Policy Manual says it colder. A practitioner “is not permitted to perform multiple services, each for the minimal reportable time, and report each of these as separate UOS.”
None of that needs judgment. It needs someone to add up minutes and compare them to units on every visit without getting tired, the one job software does better than you. Do it inside the visit instead of in an appeal six weeks later and you are ahead of the failure Medicare reviewers cite most often.
One catch decides whether a tool is worth anything: it has to know which rule the payer is on. APTA notes that CPT’s own convention only asks you to pass the midpoint, “7 minutes and 31 seconds” for a 15-minute unit, with no requirement to total the day first. Medicare’s 8-minute rule does require it. Apply one rule to every payer and you are confidently wrong across half the book.
97110 or 97530 is a documentation question
Every vendor demo reaches for this pair. Be exact about what settles it.
Under LCD L33631, the Medicare contractor policy revised effective April 2026, therapeutic exercise restores “strength, endurance, range of motion and flexibility where loss or restriction is a result of a specific disease or injury and has resulted in a functional limitation.” Therapeutic activities use functional activities “to restore functional performance in a progressive manner,” and the LCD’s own examples are bending, lifting, carrying, reaching, transfers and bed mobility. Both demand the therapist’s skill to design the activity and instruct the patient in it.
An AI reading “ther ex 15 min, HEP reviewed” cannot tell you which one happened, because the note does not say. Neither can a reviewer. The billing and coding article attached to that LCD puts the duty plainly: the record should “identify each specific skilled intervention/modality provided to justify coding.” Timed Code Treatment Minutes and Total Treatment Time both have to be there to justify the units. The Medicare Benefit Policy Manual wants each intervention written in language “that can be compared with the billing on the claim to verify correct coding.”
Which is why the honest job here has little to do with picking between the two codes. What matters is noticing that whichever one your therapist chose, the note does not yet say enough to defend it, and saying so while she is still in the room. Description first, code second. That is the order documentation has to run in, and most software has it backward.
Modifier 59 is where a suggestion becomes exposure
Modifier 59 unbundles an edit pair, and CMS says flatly that providers “often use” it incorrectly. The April 2026 booklet on proper use of modifiers 59, XE, XP, XS and XU gives the rule in one line: “Medical documentation must support the use of the modifier.” It also kills the shortcut everyone reaches for. Two code descriptors being different is not a reason.
Therapy gets a narrow carve-out. The NCCI manual allows modifier 59 on timed codes when two separate and distinct services ran in separate and distinct time blocks, with the day’s total-minutes math still applying on top. That is a documentation test with a stopwatch attached. An AI that appends 59 because a pair tripped an edit has not coded anything. It has made a claim, in your therapist’s name, about how the visit was sequenced. The chart may not back that up.
Does AI catch undercoding?
Sometimes. The direction of the catch matters more than the hit rate. Undercoding in outpatient PT is usually a documentation failure with a billing symptom: skilled, progressive functional retraining happens, three words get written about it, and the biller codes what three words can carry. The money went missing in the note.
A tool that says “your documentation describes more than you billed” is doing real work, provided the fix is the therapist adding what she actually did. A tool that quietly proposes the higher-paying code across your caseload is optimizing a number, and the NCCI manual names where that ends. “Providers/suppliers must avoid upcoding. A HCPCS/CPT code may be reported only if all services described by that code have been performed.” In the otology study the models leaned the other way, assigning fewer work RVUs than the human coders did. Found revenue is not what the published evidence shows.
Why a scribe outside your EHR cannot do this
Here is the part the market skips. Every question that decides a code is a question about the case, not about the visit.
Whether an activity is 97530 rather than 97110 turns on whether it ties to a functional goal in the plan of care. Whether the day’s units hold turns on total treatment time across the whole session. Whether modifier 59 is defensible turns on how the visit was sequenced. Whether Medicare’s 8-minute rule applies at all, or CPT’s midpoint convention, turns on which payer you are billing.
A bolt-on scribe cannot answer any of those, because its entire input is the audio of one visit. It hears a room. It does not know the goals you set eight weeks ago, the authorization that caps this episode at twelve visits, or what the last episode’s baselines were. So it does the only thing available to it: it guesses a code from words, which is exactly the task the published evidence says these models are worst at.
Aurora starts from the other end. It writes inside the record, so the plan of care, prior episodes and their baselines, the payer and its authorization limits, medications and risk flags are already in hand before the first word is drafted. That is the difference between transcription and context, and it is why the check is possible at all. A tool that knows the goal can tell you the note describes functional retraining while the claim says therapeutic exercise. A tool that only heard the room cannot.
Which is the honest version of the pitch. Aurora is not better at guessing codes. It is built so that nobody has to guess: the record it draws on is wide enough to test the code against the documentation while the clinician is still in the room, and that is the 88.6% problem, not the 1.5% one.
Where the line sits
Draw it at authorship. The therapist writes the note and owns what it says. The biller submits the claim and owns what it claims. AI sits between them as a check that never gets bored, and never becomes the author of either. APTA’s 2024 policy on artificial intelligence backs integration “that reduces administrative burden,” and its September 2025 practice advisory on ambient scribes covers the documentation responsibilities that come with it. Burden is the target. Judgment is not.
Three questions sort the vendors. Does it check the code your therapist chose, or choose for her? Does it show the line in the note that supports or contradicts that code, or only a confidence score? Does it know which payer’s timed-code rule applies? Whoever fails the third will be wrong quietly, for months.
Description first, code second, check third. That is how Orion is built. Aurora drafts the PT-specific note while you treat, with the case already in front of it, so the description exists before anyone thinks about a code. Focus on your patient, not your keyboard. A timed code missing its minutes gets flagged inside the visit, while the clinician is still there, and the claim scrub checks units against documented treatment time on the way out. Aurora is included on both plans, and nothing in that chain picks a code for a therapist. The same three questions run through choosing an AI scribe; the liability version is answered at who is responsible if an AI scribe makes a mistake.
Pull ten Medicare visits this week and read the notes against the claims. If a stranger would land on the code you billed, your coding is fine, whatever software you run. If she would not, automated code selection does not fix it. It just makes the wrong answer arrive faster.
Because you read about AI
The note is done before the patient leaves the room.
Aurora is the ambient scribe built into Orion: tap once, treat, and sign a PT-specific SOAP note in under two minutes. Included on both plans.
