PDF Security Blog

PDF Integrity Report: July 2026

HTPBE Team··11 min read
PDF Integrity Report: July 2026

This article is a snapshot — content was accurate as of August 2026. The product evolves actively; specific counts, examples, and detection rules may have changed since publication — see the changelog for the current state.

Every month we look at aggregate, anonymized data from checks processed by HTPBE? and write up what the structural signals tell us about the state of PDF tampering. No file contents, no personally identifiable information — only the structural and metadata patterns the algorithm uses to classify documents.

This report is about proportions and movement, not raw counts. What share of documents came back flagged, which signals fired more or less often than the month before, which origins shifted, and what the recurring tampering shapes looked like. Those are the numbers that mean something; an absolute file count for a single month is noise by comparison.

July is the month where that distinction earns its keep. The flagged share moved up sharply — and almost none of that movement is about documents getting more tampered.


What the Denominator Is, Before Any Number

Every share below is a share of processed checks — completed analyses, not unique documents and not unique users. That denominator includes our own internal self-check runs, owner testing, fixture regressions, and retries; it is not deduped. On top of that, the classifier itself changed repeatedly within July (see the algorithm section — it was our second-heaviest release month on record), so the same file submitted early and late in the month can land on different sides of the line. A July detector-output number is therefore a blend of a moving population and a moving ruler.

So this is not a population fraud rate, and no single figure here should be read as one. It is a description of what reached our pipeline and how our pipeline classified it. Keep that in view for the headline, which needs it most.


The Shape of the Verdicts

The flagged share rose to approaching six in ten, reversing June's dip back below half — but the two things that moved it are both instrumentation, not tampering. First, July was a record release cadence: twenty-seven algorithm versions shipped across the month, several of them redefining the "modified" and confidence classes mid-month, which mechanically expands what gets flagged. Second, the traffic mix flipped back to API-heavy testing after June's web-dominated month, and API traffic skews toward files that are already suspected — integration tests against known-bad documents and uploads that appear to be testing whether a fake gets caught. A wider net over a more pre-selected population lifts the flagged share on its own. Neither lever says anything about how often documents in the world are being altered.

VerdictDirection vs. June
Not flagged▼ eased to just over two in five
High-confidence modification► flat, around three in ten
Certain modification▲ up — but a calibration effect, see below

Same forensic questions, a heavier and differently-aimed instrument, a different-looking headline. Read the flagged share every month as a statement about who submitted and how the pipeline changed — and in July, more than any month we have published, both of those changed at once.

A note on the "certain" tier, which also moved up this month: treat it as a calibration reading, not a trend. "Certain" describes how confidently the engine made the modification call — never certainty about intent or fraud — and this month's shift is a direct product of the same two forces above. A heavy run of coverage-broadening releases adds converging second signals to files that previously scraped a "high," and an API-skewed population stacks more unambiguous evidence per file. The confidence mix followed the releases and the population; it is not an independent finding.


Source & Origin Mix

The submission channel flipped back to a near-even split, with the API nominally the largest — a near-reversal of June, when roughly four in five checks came through the browser-based free checker. Read this as processed checks including internal testing, not as organic API-customer growth: our own automated self-check and fixture runs go through the API path, and July was a heavy build month, so a chunk of that API weight is us exercising the engine, not the market discovering it.

On origin, consumer-software exports overtook institutional documents as the largest class, with institutional slipping from the plurality it held in June. This is a classifier output on a testing-inclusive population, and it is not "more consumer fraud." Consumer-software and scanned files both fall into a "Cannot Verify" bucket, where the structural layer deliberately returns a conservative inconclusive verdict rather than forcing an intact-or-modified call — a larger consumer-software slice means more files we decline to certify either way, not more tampering. Several July releases also re-routed origin classification (sharpening scan and fake-scanner recognition in both directions), so part of the reshuffle is the ruler moving, not the population.

Scanned documents held at roughly a ninth of submissions — essentially flat, within noise, and itself touched by the reclassification work above. Treat that as marginal, not a decline. A scan can still never earn an "intact" verdict here: capture-origin formats simply carry too little structural history to certify either way, so they route to a conservative inconclusive rather than a clean pass.


Signals That Moved

The cleanest reads this month are the ones that do not depend on the detector at all — pure structural composition of what showed up.

Missing creation date eased again. Files arriving with no creation timestamp at all slipped from about a fifth to about a sixth. This is a detector-independent carrier — it is just what the submitted files contained — and it reflects the changed submission mix, not a real-world change in how documents are authored. Still worth watching, but the direction this month is down.

The suppressed slices stayed suppressed: signed documents, post-signature edits, signature removals, embedded files and JavaScript were each too thin a base this month to quote a rate, so we keep them qualitative. The one signed-document pattern worth repeating is the standing one — a signature valid in the viewer does not guarantee the bytes were not altered, because incremental updates appended after signing fall outside the signed scope. Integrity checked at the structural layer, not the signature-validation layer, is what catches that.


Incremental Updates: Two Rates, Kept Apart

The incremental-update signal splits into two separate numbers this month, and they should not be blurred together.

The first is prevalence — how many of all files carried incremental updates at all. That ran about one in ten, a structural share that ticked up modestly and is safe to read at face value. The average revision chain on those files also shortened, to under three appends, down from a little over three in June — fewer post-write layers stacked per file.

The second is the flag rate among them — of the files that did carry incremental updates, how many came back flagged. In July that sat at very nearly all of them. Read this as a July snapshot only: it was lower in June, so this is not a "near-total all along" continuity claim, and it is not a real-world trend. The mechanism is unchanged — incremental updates let content be appended after the original write, and while legitimate workflows produce them (signature application, annotation, form-fill), those clean cases are a small minority of the population that reaches a tamper-detection tool.


Representative Cases

These are composite, anonymized illustrations of the recurring shapes the engine resolved this month — not specific files. Each maps to the structural markers that actually drove the verdict, and each describes a modification, not proven fraud.

The reconstructed statement (verdict: modified). A "bank statement" looks like one coherent document. Structurally, its pages were assembled from more than one separate source rather than produced as a single original — so it can be reported as an assembled file, but never certified as an untouched original. July broadened this multi-source coverage to catch a wider range of assembled packages that previously slipped through.

The rebuilt-elsewhere document (verdict: modified). A file presented as an untouched institutional original was, structurally, re-saved and rebuilt by a second tool after its original creation, rather than issued and released once by an institution. The public outcome is exactly that: rebuilt by a second tool after the file was first created. It reads as an original; the structure records the later rebuild.

The cover-and-replace patch (verdict: modified). A page reads correctly, line by line. Structurally, the page shows content added after the original was produced — for example, a blackout or cover box placed over part of the page. That is reported as a modification regardless of why it was applied: the fact of the edit is what we record, not the intent behind it.

The render dressed as a scan (verdict: modified). A file arrives looking like a scanner or OCR capture, but structurally it is a software-rendered document presented as a capture origin it cannot genuinely have. Genuine device scans and ordinary phone captures are unaffected — those route to a conservative inconclusive and are never flagged on this basis; only a render claiming to be a scan trips it.


Algorithm Development

July was a very heavy release month — twenty-seven versions shipped, second only to May's twenty-nine and well above June's sixteen. The work spanned new coverage and false-positive reduction at once, described here in outcome terms only.

  • New and broadened detection categories — a software render dressed as a scanner or OCR capture; content added over the original page, such as blackout or cover boxes; fabricated and assembled documents dressed up as institutional originals; documents rebuilt by a second tool after their original creation; broadened multi-source page assembly reaching a wider range of assembled packages; and strengthened detection of post-creation edits. Several of these closed cases that had previously passed, certified as originals.
  • False-positive reduction and confidence demotion — releases that pushed the other way, narrowing misfires on genuine templated documents, signed-document workflows, and minor date discrepancies, and demoting the confidence of borderline calls rather than flagging them outright.

Here is the part that matters for the headline: wider coverage pushes the flagged share up. A share of the documents flagged in July would have passed under the algorithm as it stood on the first of the month. That compounds the traffic-mix effect from the top of the report — and it is precisely why the rise in the flagged share is a statement about our instrument and our intake, not a fraud trend.


PDF Version Landscape

Concentration kept loosening. PDF 1.7 fell from roughly 45% to about a third of the sample, with 1.4 a close second, and 1.3 and 1.6 tied for third. PDF 1.5 took a smaller slice, and PDF 2.0, despite nearly a decade of availability, stayed a rounding-error share. Like the missing-creation-date read, this is a detector-independent carrier — it reflects the changed submission mix this month, not a shift in real-world version adoption.


Summary

July 2026, in relative terms:

  • The flagged share rose to approaching six in ten — but the drivers are a record twenty-seven-version release cadence broadening what gets caught and a traffic mix that flipped back to API-heavy testing, not more tampering.
  • The channel split returned to near-even with the API nominally largest — processed checks including internal testing, not organic API growth — and consumer-software origin overtook institutional inside a "Cannot Verify" population, which is not a consumer-fraud signal.
  • Scanned share held at roughly a ninth, flat and reclassification-affected; the "certain" tier moved up as a calibration effect of the releases and population, not a standalone trend.
  • Incremental-update files ran about one in ten of all files, with shorter chains, and their July flag rate was near-total — read as a snapshot, not a continuity claim.
  • The clean, detector-independent carriers pointed down and loosened: missing creation dates eased from about a fifth to about a sixth, and PDF 1.7 loosened from roughly 45% to about a third — both reflecting the submission mix, not real-world adoption.

Every pattern here comes from the same forensic engine that teams run on their own intake stream through the PDF tamper detection API. If you want to run a single document through the same analysis by hand, the free checker does it in the browser.


This report covers checks processed by HTPBE? in July 2026. We analyze only file structure, never document content; web uploads may be retained in anonymized form to improve detection. All figures are aggregate and anonymized.

Share This Article

Found this article helpful? Share it with others to spread knowledge about PDF security and fraud detection.

https://htpbe.tech/blog/pdf-integrity-report-july-2026

Secure your workflow

Create your account — API key on signup, free test environment on every plan.
From $15/mo. No sales call. Cancel any time.