PDF Security Blog

Medical Claims Fraud Detection: Pharmacy and Prior-Auth PDFs

HTPBE Team··12 min read
Medical Claims Fraud Detection: Pharmacy and Prior-Auth PDFs

This article is a snapshot – content was accurate as of September 2026 (code examples tested against the API as of August 2026). The product evolves actively; specific counts, examples, and detection rules may have changed since publication – see the changelog for the current state.

A pharmacy benefit manager opens a desk audit on a retail pharmacy: forty claims, ninety days, one high-cost therapeutic class. The pharmacy responds the way every pharmacy responds — a zip of PDFs. Scanned hardcopy prescriptions, a signature log, wholesaler invoices, a handful of prescriber clarification notes. The audit analyst reads the documents against the claim lines: drug, strength, quantity, days supply, date written, refills authorized.

On one file, the quantity on the hardcopy reads 180. The claim was billed at 180. The prescriber wrote 30.

Nothing in that packet contradicts itself, because the packet is the pharmacy’s own account of what happened. The analyst is comparing a document against a claim, and both sides of that comparison came from the same submitter. The one question nobody in the workflow asks is whether the PDF itself was written once and left alone, or opened afterwards and re-saved.

That question has an answer more often than the workflow assumes, and it lives in the file, not in the content.

Scope: payment integrity, not credentialing

This article is about documents that arrive after a service was rendered or before it is approved — the paperwork that flows through pharmacy audit, prior authorization and claims adjudication at a payer, PBM or health plan. Provider-side credentialing fraud — forged licenses, altered board certificates, tampered CE records — is a different intake pipeline with a different owner, and we cover it separately in healthcare document fraud and medical credentials.

Two boundaries before anything else, because they are the ones that get blurred in this vertical:

  • Structural PDF analysis makes no clinical judgment. It cannot tell you whether a therapy was appropriate, whether a diagnosis supports a code, or whether a prior-authorization request should be approved. It answers one narrow question about one file.
  • It is not identity verification. It does not confirm that a prescriber is real, that an NPI is active, or that a pharmacy is licensed. Those are registry lookups and they remain your job.

What it adds is a file-integrity layer running alongside the clinical and claims review you already do, on the same PDF, at intake.

Three document surfaces on the payer side

1. Pharmacy audit response packets

PBM provider manuals routinely require pharmacies to retain and produce prescription hardcopies, signature logs, purchase invoices and documentation for any quantity adjustment or override, and to hand them over for desk and on-site audits on short notice. In practice that means a batch of PDFs, produced by the audited party, describing the audited party’s own conduct.

The fraud pattern is mundane, not exotic forgery. A hardcopy is scanned, the scan is opened in an editor, and a digit changes — a quantity, a refill count, a date written, a directions line. Or a missing hardcopy is reconstructed after the audit letter arrives, backdated, and dropped into the packet alongside genuine ones.

Both tend to leave file-level traces that have nothing to do with the visible pixels. A document that was edited after it was produced can carry structural evidence of that second write — evidence that is independent of how convincing the visible page looks. Which findings come back is reported per file as named modification markers.

The biggest caveat in this article belongs here: most audit hardcopies are scans, and scans have a ceiling. A scanned document has already lost its issuer provenance — the scanner is the producer, not the prescriber’s system. Under the published verdict rules a scan cannot return intact; inconclusive is the normal outcome for this class, and it means the file cannot carry the evidence either way.

2. Prior-authorization request packets

Prior authorization is where the document layer is getting harder. Under the CMS Interoperability and Prior Authorization final rule (CMS-0057-F), impacted payers — Medicare Advantage organizations, state Medicaid and CHIP fee-for-service programs, Medicaid and CHIP managed care plans, and QHP issuers on the federally facilitated exchanges — take on new prior-authorization obligations, and all of them except QHP issuers on the FFEs must return decisions within 72 hours for expedited requests and seven calendar days for standard ones.

Two boundaries on that rule matter for anyone reading this from the pharmacy side. It covers medical items and services; CMS excluded drugs from both the prior-authorization API and the process requirements, on the grounds that drug prior-authorization standards and timeframes work differently. So pharmacy-benefit prior authorization sits outside this rule’s timeframes entirely — other federal and state requirements may still apply, and it runs on plan and PBM policy alongside them. And on the medical-benefit side, faster mandated turnaround does not reduce the volume of supporting documentation. It compresses the window in which anyone looks at it.

A prior-authorization packet is a document bundle: the PA form, a letter of medical necessity, chart notes, prior therapy documentation, sometimes lab values as a PDF export. A denial in this workflow is consequential, so the temptation on the submitting side is to make the record say what the criteria need it to say. In the version we care about, that is done by editing a real document rather than inventing one from nothing — a step-therapy note whose date moves, a chart excerpt where one value is patched, a letter assembled from pages of two different source files.

Assembly and overlay patterns of this kind are the sort of thing a structural check is aimed at surfacing, and they come back as named modification markers attached to the verdict. None of that says the request should be denied. It says a human should read this packet before the clock runs out on it, rather than after.

3. Itemized bills and appeal documentation

Itemized bills, treatment summaries and appeal attachments follow the same inflation pattern as any altered invoice: a genuine document from a genuine provider, with a number changed. We have covered that pattern and its structural signature in detail under medical bill tamper detection and in insurance claims fraud and altered PDFs, and there is no reason to repeat it here. The only payer-specific note worth adding is sequencing — on the medical side the itemized bill usually shows up in an appeal, after an initial adjudication, which is exactly the moment a plan is least likely to route a document to fresh review and most likely to read only the disputed line.

Why the file layer is separate from the claim layer

Every fraud, waste and abuse control a plan runs today reads claim data. Edits, aberrancy models, peer-group outliers, prescriber pattern analysis, network links — all of it operates on structured fields that the claim already contains. None of it opens the attached PDF as a file.

A PDF does not overwrite itself when it is edited. It appends, and it keeps a record of having appended, along with a record of which software did the appending and when. That is the mechanism, and the xref table forensics article walks the byte-level detail if you want it. The operational point for a payment-integrity team is narrower: the evidence that a supporting document was rewritten sits in a place your adjudication engine has never looked, and it survives regardless of how plausible the claim line looks.

It also cuts the other way, which is why this is a layer and not an oracle. A perfectly clean file can support a completely fraudulent claim. Structural integrity is a statement about the file, not about the truth of what is printed on it.

Reading the three verdicts in a payer workflow

The API returns one of three verdicts, and in this vertical the middle one carries most of the operational weight.

modified — the file was written, then written again. Named modification markers come back with it. This is a routing signal: pull the document into human review, and where the document is a copy of something an issuer holds, request the original from the issuer — the prescriber’s system, the dispensing system, the clinic.

inconclusive — the file’s origin is consumer software, an online editor, a scan, or another class with no institutional pipeline behind it, so there is no provenance to check against. This does not mean the analysis failed, and in pharmacy audit work you will see a lot of it. It means the file cannot carry the evidence either way. A scanned hardcopy is the normal case. A prior-authorization letter of medical necessity that a clinic swears came out of its EHR but returns inconclusive is a different conversation — not an accusation, a question. We have written a fuller treatment of what inconclusive actually means.

intact — no evidence of post-creation modification, and the file’s origin is consistent with institutional generation. This is the verdict most likely to be over-read. It is not proof that the document is authentic, that the content is true, or that the claim is payable. It is the absence of one specific kind of evidence.

Routing, never automatic denial

This matters more in healthcare than in any other vertical we write about, so it goes here and not in a limitations section at the end.

A modified verdict is evidence that a file changed after it was created. It is not a finding of fraud, and it is not a basis for denying a prior-authorization request, rejecting a claim, recouping from a pharmacy or terminating a network contract on its own. Documents get re-saved for entirely ordinary reasons — a front-counter workflow that reprints and re-scans, a practice management system that rewrites a PDF on upload, a staff member who rotates a crooked scan and saves it.

The API ships that constraint in the payload rather than leaving the integrator to infer it. Every result carries a usage_caution object with safe_for_automated_adverse_decision: false and a recommended_action of route_to_human_review, request_from_issuer or no_action. Read that field and route on it. In a benefit-decision workflow the correct downstream action is a human reviewing the document, a request for the original from the issuing system, or a normal audit follow-up under your existing provider-manual process — with the appeal rights and turnaround obligations that already apply to that process.

Integration

One POST with a URL to the document, one GET for the result:

curl -X POST https://api.htpbe.tech/v1/analyze \
  -H "Authorization: Bearer $HTPBE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://plan.example.com/pa/packet-88213.pdf" }'
# → { "id": "3f9c8b7a-2e1d-4c5f-9b8e-7a6d5c4b3a21" }

curl https://api.htpbe.tech/v1/result/3f9c8b7a-2e1d-4c5f-9b8e-7a6d5c4b3a21 \
  -H "Authorization: Bearer $HTPBE_API_KEY"

The result endpoint is where the verdict lives:

{
  "id": "3f9c8b7a-2e1d-4c5f-9b8e-7a6d5c4b3a21",
  "status": "modified",
  "modification_confidence": "high",
  "creator": "Clinical Document System",
  "producer": "PDF Editor",
  "xref_count": 2,
  "has_incremental_updates": true,
  "has_digital_signature": false,
  "modification_markers": ["HTPBE_EDITING_TOOL_FINGERPRINT", "HTPBE_MULTIPLE_REVISION_LAYERS"],
  "usage_caution": {
    "safe_for_automated_adverse_decision": false,
    "recommended_action": "route_to_human_review",
    "message": "..."
  }
}

The routing logic on your side is a switch on status, honoring usage_caution.safe_for_automated_adverse_decision as the constant false it always ships as. modified goes to a payment-integrity queue with the marker list attached so the reviewer knows what to look at. inconclusive on a document class you expect to be institutional goes to the same queue at lower priority. inconclusive on a scanned audit hardcopy is the expected baseline for the class — route it through your existing audit sampling.

Every account gets a test key that returns deterministic synthetic results for a fixed set of scenarios, so you can build and test the routing branches before a single real document moves. Plans, limits and key generation are on the API page.

Documents in this vertical carry patient information, so where the analysis runs is a question your security and privacy review will ask before your engineering team does. The hosted API and a self-hosted, on-premise deployment are both available. Which one you can use against your own HIPAA obligations — including whether a Business Associate Agreement is required and available for the hosted path — is a determination for your own compliance function, and we’ll confirm our current BAA position directly if you ask; we are not in a position to draw that conclusion for you here. The API page links the on-premise deployment documentation.

What this does not solve

  • Content truth. A prescription PDF that was generated once, never touched, and describes a drug the patient never received will come back intact. Nothing about file structure speaks to whether the underlying event happened.
  • Fabricated-from-scratch documents. A letter of medical necessity typed in Word and exported once has no editing history to find. It returns inconclusive because its origin is consumer software — a useful signal when the document claims to come from a clinical system, and no signal at all when it does not.
  • Scans, in general. Covered above. The scan ceiling is real and it is the dominant document form in pharmacy audit response.
  • Anything about people. Prescriber identity, patient eligibility, pharmacy licensure, NPI validity. Different problem, different tools.
  • Improper payment in the regulatory sense. CMS is explicit that its improper-payment measurement is not a measure of fraud — in FY2025, 77.17% of Medicaid improper payments were the result of insufficient documentation, which CMS notes is generally not indicative of fraud or abuse. Structural PDF analysis speaks to whether a file was edited, which is a narrower question than either.

Where it fits

If you run pharmacy audit, prior-authorization intake or claims payment integrity, the honest pitch is small: you already collect these PDFs, a structural check is one API call per file, and it surfaces a category of evidence that no part of your current stack examines. It will not adjudicate anything for you, and in a workflow where the majority of documents are scans, a large share of results will be inconclusive by design.

What it changes is which documents a human looks at first, and whether the packet that arrived in response to an audit letter gets examined as a set of files rather than only as a set of claims.

Start with the API documentation and pricing — test keys work before you commit to anything, and the analyze call is a single POST.

Share This Article

Found this article helpful? Share it with others to spread knowledge about PDF security and fraud detection.

https://htpbe.tech/blog/medical-pharmacy-document-fraud-detection

Secure your workflow

Create your account – check PDFs on the web or with an API key, both ready on signup.
No sales call. Cancel any time.