What medical fax OCR does
Most outside records still reach hospitals and practices by fax. A fax arrives as an image: a picture of a page. Optical character recognition, or OCR, reads that picture and produces text, so the page can be searched and processed by software.
In healthcare, that matters because a faxed record is otherwise invisible to every other system. An image of a referral packet can be stored and viewed, but no software can tell you whether it contains an ECG from the last 90 days.
Where OCR stops
OCR answers the question “what characters are on this page?” Clinical intake asks a different question: “is the fact my specialist needs in this packet, and where?” Four gaps sit between those questions.
1. Text isn’t meaning
An 846-page packet of OCR text is still 846 pages. A coordinator still has to find the ejection fraction, the troponin trend, or the pathology report. Search helps only if you know the exact words the other facility used.
2. Fax quality varies
Faxes are low-resolution, skewed, and sometimes handwritten. OCR on a poor page produces text that looks confident but can be wrong. A pipeline that treats every page the same has no way to say “I’m not sure about this one.”
3. Cost scales with paper
Modern OCR increasingly means running large AI models over every page. At health-system volume, running a full model pass over cover sheets, duplicates, and blank pages adds up fast.
4. Nobody can check the result
If software tells a nurse “LVEF 35%,” the nurse needs to see where that came from before acting on it. OCR output, and especially AI-generated summaries of it, often loses the link back to the page.
What to look for instead
Solving the inbound fax problem takes OCR plus four things.
Confidence routing
Decide what each page needs before paying for it. Easy pages go on a fast path, pages that need a closer read get a targeted pass, and low-confidence pages go to a person. InfoNotData calls these lanes Easy, Medium, and Hard. See how routing works.
Checklist extraction
Extract only what the receiving specialty needs, such as an ECG within 90 days for a cardiology consult, instead of processing everything. See specialty checklists.
Page-level citations
Every extracted value should carry the page and line it came from, so anyone can verify it in one click. See source citations.
Honest uncertainty
The system should say what it couldn’t find and what it isn’t sure about. A flagged gap is useful. A confident guess is dangerous.
Questions to ask a fax OCR vendor
- Does it generate text about patients, or only extract values?
- Can every extracted value be traced to a specific page and line?
- What happens to a page the system isn’t confident about?
- Does it run the same expensive processing on every page?
- Do we need new fax numbers, or changes to our EHR?
- Are the performance figures measured results, or projections? Ask to see the assumptions.
InfoNotData was built around those questions. Read why.