HomeBlogClinical documentation
Physical therapy outcome measures: what to use, what counts
Much of what ranks here still teaches a Medicare program that ended in 2019. The measures worth using, what a change means, and where reporting stands now.

The short version
- Functional limitation reporting ended for dates of service on and after January 1, 2019. Any guide still walking you through G-codes and severity modifiers is describing a discontinued program.
- Medicare recommends standardized outcome tools but does not require them. What it requires is documentation that establishes progress through objective measurements.
- An MCID is not a constant. The NDI's published important-change values span 5 to 19 points out of 50 depending on the study and population.
- If MIPS applies to you, quality reporting requires at least one outcome measure, and CMS's PT and OT specialty set is built on Functional Status Change measures for knee, hip, low back, shoulder and neck.
Medicare does not require the LEFS. Or the Oswestry, or the Neck Disability Index, or any other instrument. Chapter 15 of the Medicare Benefit Policy Manual says so directly: standardized assessment instruments and outcome tools “are not required, but their use will enhance the justification for needed therapy.” What the manual does demand is narrower and harder to fake. Documentation “should establish through objective measurements that the patient is making progress toward goals.”
That one distinction untangles most of the confusion around outcome measures in physical therapy. It even explains the largest single source of it: a Medicare reporting program that ended in 2019 and still fills page one of the search results. Here is the current picture, instrument by instrument, with the psychometrics sourced.
Which outcome measure should you use?
Match the instrument to the body region you are treating and the decision the score needs to inform. The list below covers most outpatient PT caseloads, and every change threshold on it comes from a validation study or systematic review, quoted with the population it was derived in.
| Measure | Built for | The score | Meaningful change, and in whom |
|---|---|---|---|
| LEFS | Lower extremity, any condition | 20 items, 0 to 80; higher is better function | MCID 9 points in outpatient lower-extremity musculoskeletal patients; a 2016 review pooled the MDC90 at 6 |
| QuickDASH | Arm, shoulder, hand | 11 items, 0 to 100; higher is worse | MCID 15.91 points, MDC90 12.85, in upper-limb musculoskeletal patients in physical therapy |
| NDI | Neck | 10 items, out of 50 | MDC around 5 in uncomplicated neck pain, up to 10 in cervical radiculopathy |
| ODI (Oswestry) | Low back | 0 to 100 | Consensus minimal important change of 10 points, or 30% improvement from baseline |
| TUG | Basic mobility and fall risk | Seconds to rise, walk 3 meters, turn, return, sit | Identified community-dwelling older adults prone to falls with 87% sensitivity and 87% specificity |
| 30-second chair stand | Lower-body strength, adults 60+ | Stands completed in 30 seconds | Validated against maximum leg press (r = .78 in men, .71 in women); performance drops by decade of age |
| PROMIS Physical Function | Function across conditions | T-score; 50 is the US general population mean | Read in population terms: every 10 points is one standard deviation |
Where those numbers come from: Binkley and colleagues’ original LEFS study and Mehta’s 2016 systematic review of it, Franchignoni’s MCID study of the DASH and QuickDASH, MacDermid’s systematic review of the NDI, the Ostelo and de Vet international consensus on low back pain change scores, Podsiadlo and Richardson’s original TUG paper plus Shumway-Cook’s fall-prediction study, Jones, Rikli and Beam’s chair-stand validation, and the PROMIS Wave 1 calibration, which set every PROMIS T-score to a mean of 50 and a standard deviation of 10 against the general population.
One licensing note before you print anything. The DASH and QuickDASH are copyrighted by the Institute for Work & Health, which allows use without charge for clinical practice, non-commercial research and other not-for-profit purposes, and forbids altering the forms even slightly.
Do therapists actually use these? Mostly yes. A 2025 survey of 514 APTA members in PLOS One found 97% used performance-based tests and 83% used self-report surveys. Fewer than half agreed the profession administers them in a standardized way, which is worth remembering the next time someone quotes you a clinic-to-clinic benchmark.
What does an MCID actually tell you?
Three different thresholds hide inside the phrase “the score improved,” and mixing them up is the main way outcome data gets oversold.
The minimal detectable change (MDC) is a statistical floor: the smallest change bigger than the instrument’s own measurement error. The minimal clinically important difference (MCID) is a judgment call anchored to patients: the smallest change patients themselves experience as worthwhile. De Vet and Terwee’s methods paper draws the line cleanly. Some popular calculation methods, they note, “have been merely focused on minimally detectable changes” while claiming to measure importance, and MIC values vary with the anchor chosen, the definition of minimal importance, and the disease being studied. Statistical significance is a third thing entirely. It describes groups in a trial, not the patient in front of you.
So an MCID is a local estimate, not a constant. The NDI shows how local: MacDermid’s review found published important-change values ranging from 5 to 19 points out of 50 across studies. Franchignoni’s QuickDASH figure of 15.91 points came with the caveat that it marks the lower boundary of a range whose upper end, proposed by the DASH’s own website, sits at 20. And the Ostelo consensus for low back pain concluded that a 30% improvement from baseline often says more than any fixed point value, because a 10-point gain means something different at ODI 60 than at ODI 20.
The practical reading order: a change smaller than the MDC is noise. A change past the MDC is real. A change past an MCID derived in a population like your patient is worth writing in the assessment line. Quote an MCID to a payer as if it were physics and you are one literature search away from being contradicted.
Is functional limitation reporting still required?
No. CMS discontinued functional limitation reporting “effective for dates of service on and after January 1, 2019.” That took the nonpayable G-codes and severity modifiers off therapy claims and retired their documentation requirements with them. The program ran from January 1, 2013 through December 31, 2018, and the 2019 fee schedule rulemaking (CMS-1693-F) ended it.
This matters because a large share of the outcome-measure guidance still ranking today was written between those two dates. If an article walks you through reporting G-codes at evaluation, at every tenth visit and at discharge, it is describing a program that has been dead for more than seven years. Nothing replaced it claim-for-claim. What exists now is smaller and lives in three different places.
Where is outcome reporting actually required now?
For Medicare, the record itself is the first requirement. Chapter 15’s progress report, due at least every 10 treatment days, must show the “extent of progress (or lack thereof) toward each goal,” and the manual lists objective measurements as the preferred way to show it. The full anatomy of daily notes and progress reports is in our SOAP notes guide. The short version: an outcome measure administered on schedule is progress evidence a reviewer can verify, and our Medicare audits guide shows what reviewers do when the chart cannot supply it.
MIPS is the second place, and only if you clear its low-volume threshold or opt in. Quality reporting there requires six measures including at least one outcome measure, per CMS’s 2026 quality guide, and the PT and OT specialty measure set is built on Functional Status Change measures for knee, hip, low back, shoulder and neck impairments. Whether the program applies to your practice at all is threshold math we worked through in the MIPS guide; run it before paying anyone for a registry.
Commercial payers are the third, and the least documented. Aetna, Cigna and UnitedHealthcare all publish medical policies that test PT notes for measurable functional progress on their own clocks, which we compared line by line in our medical necessity guide. What those public policies rarely do is name a required instrument. Plan-level scoring requirements, where they exist, tend to live in provider manuals and network portals rather than published policy, so the only reliable answer comes from your contract and your network rep, not from a blog, this one included.
When should you administer outcome measures?
At evaluation, at every progress-report interval, and at discharge. The evaluation score is the baseline your plan-of-care goals should reference, so that “LEFS 34 to 58” can stand as a goal a reviewer can check. Medicare’s own cadence then does the scheduling for you: re-administer within each progress-report period, at least every 10 treatment days, and the discharge note, which the manual treats as the final progress report, carries the closing score. A measure administered at eval and never again is a baseline without a trend, which is to say, not evidence.
The whole system only works if capturing the score costs nothing at the visit, and that is a documentation problem before it is a measurement problem. In Orion the measure lives where the rest of the visit lives: Aurora drafts the note from the visit and the therapist confirms and signs it, and reporting reads those same live records rather than a copy you re-key later.
The one program that made outcome reporting feel like busywork ended in 2019. What is left is the useful half: a number, taken on a schedule, that tells you whether the plan of care is working before anyone else asks.
Because you read about documentation
Focus on your patient, not your keyboard.
Aurora drafts the note while you treat; you review and sign. PT-specific, included, no per-note metering.
