Every carrier and TPA is currently being shown software that reads medical records. Some of it is genuinely good. The question worth asking is not whether a model can read a bill — it can — but whether a finding produced without a professional behind it can survive being challenged.
That is a different question, and it does not get easier as the technology improves. It is worth separating the two clearly before deciding how much of a review to hand over.
1. A model cannot be deposed
When a reduction is challenged — in negotiation, in mediation, or in a courtroom — someone has to say: I reviewed this record. Here is my reasoning. Here is my licence.
A model output cannot say that. It has no licence to put at risk, no professional body to answer to, and no availability for examination. This is not a gap that a better model closes; it is a structural feature of what a model is. The finding may well be correct, and it will still be the wrong kind of thing to put in front of someone whose job is to test who produced it and on what basis.
A reduction nobody will stand behind is not a position. It is a suggestion with a number attached.
2. Errors stop being random and start being correlated
This is the risk that gets least attention and deserves the most.
A human reviewer makes mistakes, and those mistakes scatter. One reviewer misreads a chart on a Tuesday; a different reviewer misses something else on a different claim. The errors are independent, which means they partly cancel and none of them describes the programme as a whole.
An automated system fails differently. If it is systematically miscalibrated on a particular treatment pattern, a particular provider type, or a particular kind of documentation, it makes the same error on every claim that fits the pattern. The mistakes do not scatter. They accumulate quietly across the book and surface together — usually when someone with an incentive to look goes looking.
A human review programme's worst case is a handful of weak files. A fully automated programme's worst case is a systematic defect spanning every claim it touched, discovered by opposing counsel rather than by you.
3. Some questions have no retrievable answer
Automation is genuinely strong at questions with a verifiable answer sitting in the document. Is this page a bill or a clinical note? Is this line billed twice? Is a required record missing? Those have right answers, and finding them is exactly the kind of work that should not consume a professional's afternoon.
Causality is not that kind of question. The record does not contain the answer. A mechanism of injury has to be weighed against a treatment history, a pre-existing condition has to be assessed for whether the incident meaningfully changed it, and a course of care has to be judged against what the documentation actually supports. The answer does not exist in the file waiting to be extracted — it comes into existence when a qualified person forms an opinion and commits to it.
Medical necessity and relatedness work the same way. They are contested by design, which is precisely why they are where the money in a demand package sits.
4. The carrier holds the risk, not the vendor
When a settlement position is attacked, it is the adjuster defending it and the carrier carrying the exposure. Whoever supplied the analysis is not in the room.
That asymmetry is worth being deliberate about. A review programme is not only buying reductions; it is buying the ability to hold those reductions under pressure. Any part of the process that produces a number nobody can account for transfers risk toward the party who has to defend it — and does so invisibly, at the moment the reduction is applied rather than the moment it is questioned.
Where automation does belong
None of the above is an argument against using the technology. It is an argument about which tasks it should be accountable for.
A demand package arrives as hundreds of pages of mixed material: billing statements, clinical notes, correspondence, and a good deal that is neither. Sorting that, separating the bills from everything else, structuring them so they validate cleanly, ordering treatment into a usable chronology, and surfacing duplicates — that is mechanical work with a checkable right answer. It is slow, it is unrewarding, and it is the part of a review most improved by a machine doing it.
The distinction that matters is between surfacing and concluding. Narrowing what a professional has to look at is a legitimate thing for software to do. Deciding what the record means is not, because the value of that decision is inseparable from who made it.
That is the line DemandPro holds, and will keep holding as mechanical work moves to software: every clinical judgment and every pricing decision is made by a certified professional who will stand behind it. Automation is for the work that has a right answer. Judgment stays with the people whose licences are on it.
The short version
Ask a vendor what happens when a reduction is challenged and there is no person who formed the opinion. The answer to that question tells you more about a review programme than any accuracy figure — and unlike an accuracy figure, it does not change when the next model ships.
Unfamiliar with a term used here? The demand review glossary defines the vocabulary plainly, and reviewing a demand package sets out the framework a review follows.
