August 27, 2026
Field Notes: Getting Structured Fields Out of Ugly PDFs
What works when the PDF is a scan at three degrees of rotation: confidence thresholds, review queues, and why the review queue is the product.
By Ian Phillips, Founder & CEO, Phillips Data Solutions
Field notes on document extraction, because the demos all use clean PDFs and real ones aren't.
The document you actually get
Vendor demos extract from a born-digital PDF with selectable text and a consistent layout. What arrives in a real inbox is a phone photo of a printed spec sheet, rotated about three degrees, with a coffee ring over the part number.
Both are "PDFs". They are not the same problem. Budget for the second one, and ask for a representative sample before quoting β if half the volume is scans, that changes the approach.
What works
Confidence scoring per field, not per document. A document isn't 92% correct. Field by field, the part number might be certain and the delivery date a guess. Score each one and act on each one separately.
A threshold with a queue underneath it. Above the threshold, write through. Below it, the document goes to a person with only the uncertain fields highlighted. Reviewing three flagged fields takes fifteen seconds; re-reading the whole document takes three minutes.
Keep the original linked to every value. Store the source document and, ideally, the page. When someone challenges a number six weeks later, the answer is one click instead of a mailbox search.
The thing I got wrong first
I set one global confidence threshold. That's the wrong shape, because the cost of an error isn't uniform: a wrong delivery date is an apology, a wrong quantity or price is a real financial problem.
Per-field thresholds, set by what a mistake costs. Money and quantity fields sit high enough that they route to review more often than anything else, and that's the correct trade.
The review queue is the product
The extraction is the easy half. The queue is where a document pipeline either gets adopted or abandoned.
It needs to show the uncertain field next to the region of the document it came from, accept a correction in one keystroke, and move to the next item without a page reload. Get that wrong and people go back to typing, because your tool is slower than the thing it replaced.
Corrections are also training data. Log what the human changed and you get a ranked list of which fields and which layouts need work β far more useful than an overall accuracy percentage.
What we don't automate
Anything where the document is the decision. Contracts, anything with legal exposure, anything where a human needs to have read it and be accountable for having read it. Extraction can pull the key terms into a summary to make that reading faster; it shouldn't replace it.
Where the payoff is
Any process where the same values get typed into more than two systems, and any queue where turnaround time is what you compete on. That's the shape of our own quoting workflow, where customer turnaround went from days to minutes β the bottleneck was never the decision, it was the retyping in front of it.
Volume matters less than you'd think. Ours handles two to three dense technical documents a week, and the economics don't change if that becomes thirty: the labour cost of a document pipeline is flat, which is why low-volume-but-painful processes are often better candidates than high-volume-but-tolerable ones.
The service version is on the document automation page; the filing half is SharePoint automation.
If you've got a document type that's eating a person's morning, send one representative example to a discovery call and we'll tell you what's extractable and what still needs eyes on it.
Free checklist
Custom AI App Readiness Checklist
Ten questions that tell you whether a workflow is ready for a custom AI build β the same filter we run before taking a project.
Instant access β no spam, unsubscribe anytime.
Scope your custom AI build in a free discovery call
Bring the workflow thatβs eating your teamβs hours β weβll tell you in 30 minutes whether itβs a build, a buy, or a not-yet.
Scope My Build β FreeKeep reading
Automating a LinkedIn Company Page: What the API Allows β and What It Doesn't
Read article βWe Replaced HubSpot With a Custom CRM Built on Claude Code
Read article βBuild vs. Buy AI in 2026: A Decision Framework With the Numbers
Read article βRelated services: Custom Apps Β· Workflow Consulting