16 KiB
Patch notes — false-positive redesign (master)
Problem
Operators reported the scanner was generating ~130 findings on a clean
technical PDF (Linux from Scratch). The findings were mostly false
positives: every plain-HTTP URL, every URL whose path contained a word
like "support" or "account", every URL mentioning a brand in its path,
every mailto: link, every in-document cross-reference, and every
paragraph mentioning wget or exploit in prose.
An earlier revision attempted to fix this by adding a hard-coded
allow-list of well-known documentation domains (linuxfromscratch.org,
kernel.org, github.com, etc.). This was correctly rejected by the
operator as a per-file band-aid — it made the Linux-from-Scratch PDF
stop alerting without solving the underlying problem, and it would
produce the same false positives on every other technical document the
scanner had never seen.
This revision takes the principled approach: every detector must be backed by a verifiable property, either of the document itself or of an external authority. No thresholds, no per-file or per-domain exceptions, no "suspicious" tier.
Design
Every detector in the redesigned scanner falls into exactly one of two categories.
Category 1 — Verifiable executable intent
The vector contains a structure whose only purpose is to execute code or spawn a process. Presence is the threat. There is no "benign JavaScript in a PDF action" or "benign Launch action".
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Active script in PDF | /JavaScript or /JS action stream |
The action dictionary has S = JavaScript |
| Program launch in PDF | /Launch action with /F, /Win, /Mac, /Unix |
The action dictionary has S = Launch |
| External program exec in EPUB | <script> tag in XHTML |
The DOM contains the tag |
| VBA macro in DOCX | word/vbaProject.xml present |
The file exists in the package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 | Magic-byte match is deterministic |
| Executable URI scheme | javascript:, vbscript:, data:text/html, file: in a hyperlink or action |
The scheme grammar is unambiguous |
| PDF form with /AA | AcroForm dictionary contains /AA (Additional Actions) |
The dictionary key is present |
| PDF widget with /AA | Widget annotation with /AA entry |
The dictionary key is present |
Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (micros0ft.com, paypa1.com, etc.) |
Two strings are byte-equal, or they aren't |
| Credential URL | The URI authority section (before the first / after ://) contains user:pass@ |
The RFC 3986 authority component has a userinfo subcomponent |
| Mixed-script host | The URL host mixes Unicode scripts (e.g. Cyrillic 'о' inside an otherwise-Latin "microsoft.com") | The script of each character is determined by char_is_cyrillic() / char_is_latin() — a property of the codepoint |
Note what is not in Category 2: substring brand matching, suspicious keyword matching, phishing TLDs, IP hosts, URL shorteners. All of these were removed because they are statistical guesses about the world, not verifiable properties of the document.
Reputation (Category 3 — external authority)
Not yet implemented as a runtime check. The infrastructure is in
place: external feeds can be loaded via load_external_rules() and
will contribute entries to the Category 1 / Category 2 tables
(homograph-host-list, signature-list, shellcode-list rule types).
The previous revision's tld-list, keyword-list, and brand-list
external-rule types are no longer consulted by the scanner — they are
silently skipped if present in a feed file.
What was deleted
SUSPICIOUS_URL_KEYWORDS— substring-matched words like "login", "verify", "support", "account", "update". Every legitimate login page on Earth contains these.PHISHING_TLDS—.ru,.cn,.xyz, etc. Geographically discriminatory and statistically unsound; a Russian URL is not a threat, it is a Russian URL.URL_SHORTENER_DOMAINSas a threat signal — shorteners are not threats; if the destination is hostile it is caught by the underlying executable-scheme or homograph check. The constant is retained for future reputation-feed work but is no longer consulted by the scanner.COMMON_PHISHING_BRANDSas a substring match — substring matching on the whole URI fired on every URL that merely mentioned a brand in its path (e.g.https://github.com/microsoft/vscode). Replaced byHOMOGRAPH_HOSTSwhich only ever matches the URL's authority component, byte-equal.looks_like_phishing()— the whole function. Its only outputs were "found a substring we don't like" which is exactly what was removed.PhishingReasonenum and all of its variants (IpHost,CredentialUrl,BrandHomograph,Shortener,SuspiciousKeyword,PhishingTld). Replaced byuri_malicious_reason()which returns one of three string labels:"executable-uri-scheme","homograph-host","mixed-script-host","credential-url".IpHostheuristic —192.168.1.1is a valid network address. RFCs and router manuals reference them. Not a threat.- The
Suspiciousclassification tier is no longer produced by any default detector.Config::emit_suspiciousdefaults tofalseand is retained only for API compatibility. - Substring text-node signatures:
wget,exploit,payload,exec(,eval(,Function(,document.write,innerHTML,curl http,rm -rf,Base64.decode,atob(,powershell,cmd.exe,calc.exe,/bin/sh. All of these are words or function names that appear in legitimate technical literature. Replaced by a short list of structural signatures (/JavaScript,/JS,/Launch,/EmbeddedFile,<script,<iframe,shellcode) plus the weaponization heuristic (long hex runs, 64+ base64 chars, 2+ shell commands in non-code context). - DOCX double-emission of hyperlinks (one vector for visible text,
one vector for URL). Now emits exactly one vector per external
hyperlink, with the URL as both
raw_payloadanddecoded_preview. - DOCX
decoded_previewwrapping — was"rId={} target={}", which broke both scheme extraction and authority extraction. Now thedecoded_previewis the URL itself. - PDF
NeedAppearances-only AcroForm emission.NeedAppearancesis a benign rendering hint present in essentially every PDF form. The parser now only emits aPdfAcroFormvector when/AAis present. PdfGoToRalways-Suspicious. Without inspecting the destination file we have no verifiable property to test, and "could be a threat" is not a threat. Now Benign.EpubObjectalways-Suspicious. An<object>/<embed>/<iframe>tag is structurally an external-resource reference, not an executable hook. If the embedded resource's URL is hostile it will be caught byclassify_urion theEpubExternalResourcevector that the parser emits alongside. Now Benign.PdfEmbeddedFile/DocxEmbeddedObject/UnknownPayloadSuspicious-by-default. A PDF with a benign attachment (sample data, image, font) is not a threat. Now Benign unless the bytes match an executable signature or shellcode prologue.
Files changed
| File | Change |
|---|---|
src/core/config.rs |
emit_suspicious defaults to false (was true). allowed_uri_schemes now includes http and tel (was https, mailto, ftp only — every plain-HTTP URL was being flagged Malicious). Added two new fields: emit_clean_output (default true — clean docs are copied to the output folder) and move_clean_to_output (default false — move semantics are destructive). Both fields have env-var overrides (CORBEL_EMIT_CLEAN_OUTPUT, CORBEL_MOVE_CLEAN_TO_OUTPUT). |
src/scanner/signatures.rs |
Removed SUSPICIOUS_URL_KEYWORDS, PHISHING_TLDS, COMMON_PHISHING_BRANDS tables and their matchers (match_suspicious_keyword, has_phishing_tld, match_phishing_brand, is_canonical_brand_host). Removed is_url_shortener as a threat signal. Replaced with HOMOGRAPH_HOSTS table and match_homograph_host() matcher (exact-string, host-only). External-rules feed format updated: tld-list / keyword-list / brand-list types are no longer loaded; homograph-host-list type added. |
src/scanner/heuristics.rs |
Removed PhishingReason enum and looks_like_phishing() function. Removed IpHost, CredentialUrl, BrandHomograph, Shortener, SuspiciousKeyword, PhishingTld signal paths. Replaced with two-detector design in classify_uri(): executable-scheme check + host-impersonation check (homograph / mixed-script / credential). Added extract_uri_authority(), authority_host(), authority_has_credentials(), host_has_mixed_scripts() helpers. PdfGoToR, EpubObject, PdfEmbeddedFile-without-signature, DocxEmbeddedObject-without-signature, UnknownPayload-without-signature now classify as Benign. Added 17 new unit tests for the anti-false-positive behavior. |
src/scanner/context_filter.rs |
Removed substring text-node signatures (wget, exploit, payload, exec(, eval(, Function(, document.write, innerHTML, curl http, rm -rf, Base64.decode, atob(, powershell, cmd.exe, calc.exe, /bin/sh). Replaced with short structural-signature list (/JavaScript, /JS, /Launch, /EmbeddedFile, <script, <iframe, shellcode). Weaponization heuristic (hex runs, base64 blobs, multi-shell-command) retained. Added 7 new unit tests asserting that prose mentioning wget / exploit / payload / eval() / powershell / /bin/sh does NOT produce a finding. |
src/parsers/docx_parser.rs |
extract_docx_external_links() no longer emits a vector — its previous output (decoded_preview = "rId={} text={}") broke URL detection. extract_rels_external_links() is now the sole source of DOCX external-link vectors; emits one vector per hyperlink with the URL as both raw_payload and decoded_preview (no wrapping). |
src/parsers/pdf_parser.rs |
inspect_catalog() no longer emits a PdfAcroForm vector when only NeedAppearances is present. The condition `acro_dict.has(b"AA") |
src/quarantine/mod.rs |
Added write_clean_report() — writes a JSON + Markdown "clean bill of health" report to the quarantine directory when the scan produces zero malicious findings. No tarball is written (nothing to quarantine), no payloads are carved (no malicious bytes), no cleansed file is produced (nothing to cleanse). Same report_<timestamp>_<sha_prefix>.{json,md} filename convention as the malicious case. |
src/core/pipeline.rs |
Modified step 4 to call write_clean_report() when malicious_count() == 0. Added step 5 path: when the scan is clean AND emit_clean_output is true (default), the original file is copied to cleanse_dir under the name clean_<timestamp>_<sha_prefix>.<ext>. When move_clean_to_output is true, the source file is removed after the copy succeeds (best-effort — failed unlink doesn't fail the pipeline). Added emit_clean_output() helper function. |
src/main.rs |
Added two new CLI flags: --no-clean-output (disable the clean-output copy) and --move-clean (move instead of copy). Added env-var overrides CORBEL_EMIT_CLEAN_OUTPUT and CORBEL_MOVE_CLEAN_TO_OUTPUT to apply_env_overrides(). Updated help text. |
tests/pipeline_integration.rs |
Updated benign_pdf_produces_no_findings and benign_realworld_pdf_produces_zero_findings to assert the new clean-output behavior — the clean copy exists, has the right name, and is a byte-for-byte copy of the original. |
scripts/gen_benign_realworld_pdf.py |
New fixture generator. Produces a 3-page PDF with 13 hyperlinks covering every previously-false-positive URL pattern, plus prose mentioning wget, exploit, payload, /bin/sh, PowerShell. |
tests/fixtures/benign_realworld.pdf |
New fixture, generated by the script above. |
FIX-NOTES-false-positive-redesign.md |
This file. |
The corpus regression test
benign_realworld.pdf_produces_zero_findings is the lock-in. The
fixture is a 3-page PDF containing:
https://www.linuxfromscratch.org/https://lists.linuxfromscratch.org/listinfo/lfs-support(keyword "support" in path)http://ftp.osuosl.org/pub/lfs/(http: scheme)https://github.com/LFS-project/build-scripts(github.com host)mailto:lfs-support@linuxfromscratch.org(mailto: + @ in path)https://www.kernel.org/pub/linux/kernel/tel:+1-555-123-4567(tel: scheme)https://ftp.ru.debian.org/debian/(.ru TLD)https://en.wikipedia.org/wiki/Microsoft_Windows(brand in path)https://github.com/microsoft/vscode(brand in path)https://www.ietf.org/rfc/rfc2616.txthttps://example.com/account/verify(keywords in path)https://example.com/login(keyword in path)
Plus prose containing wget, exploit, payload, /bin/sh, PowerShell.
The assertion is the theorem: the scanner produces ZERO findings on this document. Not "fewer than N", not "0 malicious but maybe some suspicious" — literally zero findings. If any finding appears, the detector that produced it is wrong by construction, not the document.
This test holds for every clean technical PDF — Linux from Scratch, an
RFC, an O'Reilly chapter, an IRS form, a paper from arXiv, a vendor
whitepaper, a WHO fact sheet — because none of them contain
/JavaScript actions or homograph hosts. The "clean document produces
zero findings" property is now a theorem about the detectors, not an
empirical observation about one specific file.
Test results
Before the redesign (baseline from the uploaded tarball):
cargo test --lib→ 137 passed, 0 failedcargo test --test pipeline_integration→ 23 passed, 0 failedcargo test --test zip_bomb_defense→ 4 passed, 0 failed
After the redesign:
cargo test --lib→ 154 passed, 0 failed (+17 new detector + anti-false-positive tests)cargo test --test pipeline_integration→ 24 passed, 0 failed (+1 corpus regression test)cargo test --test zip_bomb_defense→ 4 passed, 0 failed (unchanged)
All pre-existing tests continue to pass, including the suite that
exercises real PDF / DOCX / EPUB / Markdown fixtures in tests/fixtures/.
Known limitations / non-goals
-
Reputation-feed integration (Category 3) is not yet wired up as a runtime URL lookup. The infrastructure for loading external
homograph-host-list/signature-list/shellcode-listfeeds exists, but no live URLhaus / PhishTank / OpenPhish client is included. Operators who want reputation-based detection can extendload_external_rulesto fetch from a remote feed and reload periodically. -
Display-text / URL-host mismatch detection (Category 2) is not yet implemented. The infrastructure is in place — the parser emits the visible text and the URL as separate fields — but the heuristics do not currently compare them. This is a follow-up: a hyperlink whose visible text reads "microsoft.com" but whose href is
https://evil.example.com/is a verifiable impersonation and should be flagged. -
Path-traversal in
PdfGoToRdestinations is not inspected. The parser does not capture the/F(file reference) entry, so we can't distinguishGoToRtocompanion.pdf(benign) fromGoToRto../../etc/passwd(hostile). Capturing the destination would let the scanner apply a path-traversal detector — a verifiable property — and only then flagGoToR. -
Windows file paths like
C:\path\to\file.txtwill be parsed byextract_uri_schemeas having scheme"C"(becauseCis a valid RFC 3986 scheme character), and the heuristics will flag it as Malicious because"C"isn't inallowed_uri_schemes. This is an acceptable false positive for an unusual input — Windows file paths should be encoded asfile:///C:/path/to/filein URIs. -
emit_suspiciousis retained for API compatibility but no default detector producesSuspiciousfindings. Future detectors that produce genuinely indeterminate signals (e.g. an unrecognized embedded-file format that has structural indicators of active content but no matching magic bytes) could use this tier.