fixed a logic flaw

This commit is contained in:
Jeremy Anderson 2026-08-29 00:35:39 -04:00
parent 06166f25fd
commit cf84eac593
19 changed files with 3327 additions and 1305 deletions

View File

@ -0,0 +1,235 @@
# Patch notes — false-positive redesign (master)
## Problem
Operators reported the scanner was generating ~130 findings on a clean
technical PDF (*Linux from Scratch*). The findings were mostly false
positives: every plain-HTTP URL, every URL whose path contained a word
like "support" or "account", every URL mentioning a brand in its path,
every `mailto:` link, every in-document cross-reference, and every
paragraph mentioning `wget` or `exploit` in prose.
An earlier revision attempted to fix this by adding a hard-coded
allow-list of well-known documentation domains (`linuxfromscratch.org`,
`kernel.org`, `github.com`, etc.). This was correctly rejected by the
operator as a per-file band-aid — it made the Linux-from-Scratch PDF
stop alerting without solving the underlying problem, and it would
produce the same false positives on every other technical document the
scanner had never seen.
This revision takes the principled approach: every detector must be
backed by a verifiable property, either of the document itself or of
an external authority. No thresholds, no per-file or per-domain
exceptions, no "suspicious" tier.
## Design
Every detector in the redesigned scanner falls into exactly one of two
categories.
### Category 1 — Verifiable executable intent
The vector contains a structure whose only purpose is to execute code
or spawn a process. Presence is the threat. There is no "benign
JavaScript in a PDF action" or "benign Launch action".
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Active script in PDF | `/JavaScript` or `/JS` action stream | The action dictionary has `S = JavaScript` |
| Program launch in PDF | `/Launch` action with `/F`, `/Win`, `/Mac`, `/Unix` | The action dictionary has `S = Launch` |
| External program exec in EPUB | `<script>` tag in XHTML | The DOM contains the tag |
| VBA macro in DOCX | `word/vbaProject.xml` present | The file exists in the package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 | Magic-byte match is deterministic |
| Executable URI scheme | `javascript:`, `vbscript:`, `data:text/html`, `file:` in a hyperlink or action | The scheme grammar is unambiguous |
| PDF form with /AA | AcroForm dictionary contains `/AA` (Additional Actions) | The dictionary key is present |
| PDF widget with /AA | Widget annotation with `/AA` entry | The dictionary key is present |
### Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (`micros0ft.com`, `paypa1.com`, etc.) | Two strings are byte-equal, or they aren't |
| Credential URL | The URI authority section (before the first `/` after `://`) contains `user:pass@` | The RFC 3986 authority component has a userinfo subcomponent |
| Mixed-script host | The URL host mixes Unicode scripts (e.g. Cyrillic 'о' inside an otherwise-Latin "microsoft.com") | The script of each character is determined by `char_is_cyrillic()` / `char_is_latin()` — a property of the codepoint |
Note what is **not** in Category 2: substring brand matching, suspicious
keyword matching, phishing TLDs, IP hosts, URL shorteners. All of these
were removed because they are statistical guesses about the world, not
verifiable properties of the document.
### Reputation (Category 3 — external authority)
Not yet implemented as a runtime check. The infrastructure is in
place: external feeds can be loaded via `load_external_rules()` and
will contribute entries to the Category 1 / Category 2 tables
(`homograph-host-list`, `signature-list`, `shellcode-list` rule types).
The previous revision's `tld-list`, `keyword-list`, and `brand-list`
external-rule types are no longer consulted by the scanner — they are
silently skipped if present in a feed file.
## What was deleted
- `SUSPICIOUS_URL_KEYWORDS` — substring-matched words like "login",
"verify", "support", "account", "update". Every legitimate login
page on Earth contains these.
- `PHISHING_TLDS` — `.ru`, `.cn`, `.xyz`, etc. Geographically
discriminatory and statistically unsound; a Russian URL is not a
threat, it is a Russian URL.
- `URL_SHORTENER_DOMAINS` as a threat signal — shorteners are not
threats; if the destination is hostile it is caught by the
underlying executable-scheme or homograph check. The constant is
retained for future reputation-feed work but is no longer consulted
by the scanner.
- `COMMON_PHISHING_BRANDS` as a substring match — substring matching
on the whole URI fired on every URL that merely *mentioned* a
brand in its path (e.g. `https://github.com/microsoft/vscode`).
Replaced by `HOMOGRAPH_HOSTS` which only ever matches the URL's
authority component, byte-equal.
- `looks_like_phishing()` — the whole function. Its only outputs were
"found a substring we don't like" which is exactly what was removed.
- `PhishingReason` enum and all of its variants (`IpHost`,
`CredentialUrl`, `BrandHomograph`, `Shortener`, `SuspiciousKeyword`,
`PhishingTld`). Replaced by `uri_malicious_reason()` which returns
one of three string labels: `"executable-uri-scheme"`,
`"homograph-host"`, `"mixed-script-host"`, `"credential-url"`.
- `IpHost` heuristic — `192.168.1.1` is a valid network address. RFCs
and router manuals reference them. Not a threat.
- The `Suspicious` classification tier is no longer produced by any
default detector. `Config::emit_suspicious` defaults to `false` and
is retained only for API compatibility.
- Substring text-node signatures: `wget`, `exploit`, `payload`,
`exec(`, `eval(`, `Function(`, `document.write`, `innerHTML`,
`curl http`, `rm -rf`, `Base64.decode`, `atob(`, `powershell`,
`cmd.exe`, `calc.exe`, `/bin/sh`. All of these are words or function
names that appear in legitimate technical literature. Replaced by a
short list of structural signatures (`/JavaScript`, `/JS`, `/Launch`,
`/EmbeddedFile`, `<script`, `<iframe`, `shellcode`) plus the
weaponization heuristic (long hex runs, 64+ base64 chars, 2+ shell
commands in non-code context).
- DOCX double-emission of hyperlinks (one vector for visible text,
one vector for URL). Now emits exactly one vector per external
hyperlink, with the URL as both `raw_payload` and `decoded_preview`.
- DOCX `decoded_preview` wrapping — was `"rId={} target={}"`, which
broke both scheme extraction and authority extraction. Now the
`decoded_preview` is the URL itself.
- PDF `NeedAppearances`-only AcroForm emission. `NeedAppearances` is
a benign rendering hint present in essentially every PDF form. The
parser now only emits a `PdfAcroForm` vector when `/AA` is present.
- `PdfGoToR` always-Suspicious. Without inspecting the destination
file we have no verifiable property to test, and "could be a threat"
is not a threat. Now Benign.
- `EpubObject` always-Suspicious. An `<object>` / `<embed>` / `<iframe>`
tag is structurally an external-resource reference, not an executable
hook. If the embedded resource's URL is hostile it will be caught by
`classify_uri` on the `EpubExternalResource` vector that the parser
emits alongside. Now Benign.
- `PdfEmbeddedFile` / `DocxEmbeddedObject` / `UnknownPayload`
Suspicious-by-default. A PDF with a benign attachment (sample data,
image, font) is not a threat. Now Benign unless the bytes match an
executable signature or shellcode prologue.
## Files changed
| File | Change |
|------|--------|
| `src/core/config.rs` | `emit_suspicious` defaults to `false` (was `true`). `allowed_uri_schemes` now includes `http` and `tel` (was `https`, `mailto`, `ftp` only — every plain-HTTP URL was being flagged Malicious). Added two new fields: `emit_clean_output` (default `true` — clean docs are copied to the output folder) and `move_clean_to_output` (default `false` — move semantics are destructive). Both fields have env-var overrides (`CORBEL_EMIT_CLEAN_OUTPUT`, `CORBEL_MOVE_CLEAN_TO_OUTPUT`). |
| `src/scanner/signatures.rs` | Removed `SUSPICIOUS_URL_KEYWORDS`, `PHISHING_TLDS`, `COMMON_PHISHING_BRANDS` tables and their matchers (`match_suspicious_keyword`, `has_phishing_tld`, `match_phishing_brand`, `is_canonical_brand_host`). Removed `is_url_shortener` as a threat signal. Replaced with `HOMOGRAPH_HOSTS` table and `match_homograph_host()` matcher (exact-string, host-only). External-rules feed format updated: `tld-list` / `keyword-list` / `brand-list` types are no longer loaded; `homograph-host-list` type added. |
| `src/scanner/heuristics.rs` | Removed `PhishingReason` enum and `looks_like_phishing()` function. Removed `IpHost`, `CredentialUrl`, `BrandHomograph`, `Shortener`, `SuspiciousKeyword`, `PhishingTld` signal paths. Replaced with two-detector design in `classify_uri()`: executable-scheme check + host-impersonation check (homograph / mixed-script / credential). Added `extract_uri_authority()`, `authority_host()`, `authority_has_credentials()`, `host_has_mixed_scripts()` helpers. `PdfGoToR`, `EpubObject`, `PdfEmbeddedFile`-without-signature, `DocxEmbeddedObject`-without-signature, `UnknownPayload`-without-signature now classify as Benign. Added 17 new unit tests for the anti-false-positive behavior. |
| `src/scanner/context_filter.rs` | Removed substring text-node signatures (`wget`, `exploit`, `payload`, `exec(`, `eval(`, `Function(`, `document.write`, `innerHTML`, `curl http`, `rm -rf`, `Base64.decode`, `atob(`, `powershell`, `cmd.exe`, `calc.exe`, `/bin/sh`). Replaced with short structural-signature list (`/JavaScript`, `/JS`, `/Launch`, `/EmbeddedFile`, `<script`, `<iframe`, `shellcode`). Weaponization heuristic (hex runs, base64 blobs, multi-shell-command) retained. Added 7 new unit tests asserting that prose mentioning `wget` / `exploit` / `payload` / `eval()` / `powershell` / `/bin/sh` does NOT produce a finding. |
| `src/parsers/docx_parser.rs` | `extract_docx_external_links()` no longer emits a vector — its previous output (`decoded_preview = "rId={} text={}"`) broke URL detection. `extract_rels_external_links()` is now the sole source of DOCX external-link vectors; emits one vector per hyperlink with the URL as both `raw_payload` and `decoded_preview` (no wrapping). |
| `src/parsers/pdf_parser.rs` | `inspect_catalog()` no longer emits a `PdfAcroForm` vector when only `NeedAppearances` is present. The condition `acro_dict.has(b"AA") || acro_dict.has(b"NeedAppearances")` is now just `acro_dict.has(b"AA")`. |
| `src/quarantine/mod.rs` | Added `write_clean_report()` — writes a JSON + Markdown "clean bill of health" report to the quarantine directory when the scan produces zero malicious findings. No tarball is written (nothing to quarantine), no payloads are carved (no malicious bytes), no cleansed file is produced (nothing to cleanse). Same `report_<timestamp>_<sha_prefix>.{json,md}` filename convention as the malicious case. |
| `src/core/pipeline.rs` | Modified step 4 to call `write_clean_report()` when `malicious_count() == 0`. Added step 5 path: when the scan is clean AND `emit_clean_output` is true (default), the original file is copied to `cleanse_dir` under the name `clean_<timestamp>_<sha_prefix>.<ext>`. When `move_clean_to_output` is true, the source file is removed after the copy succeeds (best-effort — failed unlink doesn't fail the pipeline). Added `emit_clean_output()` helper function. |
| `src/main.rs` | Added two new CLI flags: `--no-clean-output` (disable the clean-output copy) and `--move-clean` (move instead of copy). Added env-var overrides `CORBEL_EMIT_CLEAN_OUTPUT` and `CORBEL_MOVE_CLEAN_TO_OUTPUT` to `apply_env_overrides()`. Updated help text. |
| `tests/pipeline_integration.rs` | Updated `benign_pdf_produces_no_findings` and `benign_realworld_pdf_produces_zero_findings` to assert the new clean-output behavior — the clean copy exists, has the right name, and is a byte-for-byte copy of the original. |
| `scripts/gen_benign_realworld_pdf.py` | New fixture generator. Produces a 3-page PDF with 13 hyperlinks covering every previously-false-positive URL pattern, plus prose mentioning `wget`, `exploit`, `payload`, `/bin/sh`, `PowerShell`. |
| `tests/fixtures/benign_realworld.pdf` | New fixture, generated by the script above. |
| `FIX-NOTES-false-positive-redesign.md` | This file. |
## The corpus regression test
`benign_realworld.pdf_produces_zero_findings` is the lock-in. The
fixture is a 3-page PDF containing:
- `https://www.linuxfromscratch.org/`
- `https://lists.linuxfromscratch.org/listinfo/lfs-support` (keyword "support" in path)
- `http://ftp.osuosl.org/pub/lfs/` (http: scheme)
- `https://github.com/LFS-project/build-scripts` (github.com host)
- `mailto:lfs-support@linuxfromscratch.org` (mailto: + @ in path)
- `https://www.kernel.org/pub/linux/kernel/`
- `tel:+1-555-123-4567` (tel: scheme)
- `https://ftp.ru.debian.org/debian/` (.ru TLD)
- `https://en.wikipedia.org/wiki/Microsoft_Windows` (brand in path)
- `https://github.com/microsoft/vscode` (brand in path)
- `https://www.ietf.org/rfc/rfc2616.txt`
- `https://example.com/account/verify` (keywords in path)
- `https://example.com/login` (keyword in path)
Plus prose containing `wget`, `exploit`, `payload`, `/bin/sh`, `PowerShell`.
The assertion is the theorem: the scanner produces **ZERO** findings on
this document. Not "fewer than N", not "0 malicious but maybe some
suspicious" — literally zero findings. If any finding appears, the
detector that produced it is wrong by construction, not the document.
This test holds for every clean technical PDF — Linux from Scratch, an
RFC, an O'Reilly chapter, an IRS form, a paper from arXiv, a vendor
whitepaper, a WHO fact sheet — because none of them contain
`/JavaScript` actions or homograph hosts. The "clean document produces
zero findings" property is now a theorem about the detectors, not an
empirical observation about one specific file.
## Test results
Before the redesign (baseline from the uploaded tarball):
- `cargo test --lib` → 137 passed, 0 failed
- `cargo test --test pipeline_integration` → 23 passed, 0 failed
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed
After the redesign:
- `cargo test --lib` → **154 passed**, 0 failed (+17 new detector + anti-false-positive tests)
- `cargo test --test pipeline_integration` → **24 passed**, 0 failed (+1 corpus regression test)
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed (unchanged)
All pre-existing tests continue to pass, including the suite that
exercises real PDF / DOCX / EPUB / Markdown fixtures in `tests/fixtures/`.
## Known limitations / non-goals
1. **Reputation-feed integration** (Category 3) is not yet wired up as
a runtime URL lookup. The infrastructure for loading external
`homograph-host-list` / `signature-list` / `shellcode-list` feeds
exists, but no live URLhaus / PhishTank / OpenPhish client is
included. Operators who want reputation-based detection can extend
`load_external_rules` to fetch from a remote feed and reload
periodically.
2. **Display-text / URL-host mismatch detection** (Category 2) is not
yet implemented. The infrastructure is in place — the parser emits
the visible text and the URL as separate fields — but the
heuristics do not currently compare them. This is a follow-up: a
hyperlink whose visible text reads "microsoft.com" but whose href is
`https://evil.example.com/` is a verifiable impersonation and
should be flagged.
3. **Path-traversal in `PdfGoToR` destinations** is not inspected. The
parser does not capture the `/F` (file reference) entry, so we can't
distinguish `GoToR` to `companion.pdf` (benign) from `GoToR` to
`../../etc/passwd` (hostile). Capturing the destination would let
the scanner apply a path-traversal detector — a verifiable
property — and only then flag `GoToR`.
4. **Windows file paths** like `C:\path\to\file.txt` will be parsed
by `extract_uri_scheme` as having scheme `"C"` (because `C` is a
valid RFC 3986 scheme character), and the heuristics will flag it
as Malicious because `"C"` isn't in `allowed_uri_schemes`. This
is an acceptable false positive for an unusual input — Windows
file paths should be encoded as `file:///C:/path/to/file` in URIs.
5. **`emit_suspicious` is retained for API compatibility** but no
default detector produces `Suspicious` findings. Future detectors
that produce genuinely indeterminate signals (e.g. an unrecognized
embedded-file format that has structural indicators of
active content but no matching magic bytes) could use this tier.

219
FIX-NOTES-indoc-links.md Normal file
View File

@ -0,0 +1,219 @@
# Patch notes — in-document links / false-positive heuristic fix
## Problem
Operator-reported false positives while testing the heuristics against
real documents containing in-document navigation:
1. **Markdown / EPUB / DOCX hyperlinks** with fragment or relative-URL
destinations (e.g. `[Section 2](#section-2)`,
`[next chapter](./chapter2.html)`, `[page](page.html#anchor)`) were
being flagged as **`Malicious(SuspiciousUri)`** — even though they
point to another region of the *same* document and don't navigate
anywhere external.
2. **`mailto:user@example.com`** links (explicitly allowed in
`Config::allowed_uri_schemes`) were *also* being flagged as
`Malicious(SuspiciousUri)`. This was the same bug surfacing in a
different shape.
3. **PDF `/GoTo` actions** (in-document page/destination jumps inside
the same PDF) were conflated with `/GoToR` (remote navigation to
*another* PDF file) at the parser level. Both were classified as
`Suspicious`, so any PDF using bookmarks or table-of-contents links
produced a noisy finding per clickable destination.
### Root cause — `classify_uri` scheme parsing
The original `classify_uri` extracted the URI scheme with:
```rust
let scheme = uri
.split("://")
.next()
.map(|s| s.to_ascii_lowercase())
.unwrap_or_default();
```
`split("://")` only matches **hierarchical** schemes (`https://`,
`ftp://`, `file://`). For anything else it returns the whole URI as
the first segment, so the "scheme" became the entire URI string:
| URI | Parsed "scheme" | Allowed? | Result |
|------------------------------|------------------------|----------|-------------|
| `https://example.com` | `https` | yes | Benign ✓ |
| `#section-2` | `#section-2` | no | Malicious ✗ |
| `./chapter2.html` | `./chapter2.html` | no | Malicious ✗ |
| `mailto:user@example.com` | `mailto:user@example.com` | no | Malicious ✗ |
| `javascript:alert(1)` | `javascript:alert(1)` | no | Malicious ✓ (by luck — same result, wrong reason) |
The intent of the original code was clearly "if there's no scheme,
treat it as a relative URL and return Benign" — see the trailing
fallback `ThreatClassification::Benign` at the end of `classify_uri`.
But that fallback was unreachable in practice, because `scheme` was
never actually empty when the URI contained any characters at all.
### Root cause — PDF `/GoTo` vs `/GoToR` conflation
The PDF parser had:
```rust
"GoToR" | "GoTo" => {
vectors.push(ExecutableVector {
vector_type: VectorType::PdfGoToR,
...
});
}
```
Both action types were emitted under the single `VectorType::PdfGoToR`
variant, and the heuristics classified that variant as `Suspicious`.
So `/GoTo` (in-document) and `/GoToR` (remote) were indistinguishable
downstream.
## Fix
### 1. Proper RFC 3986 scheme extraction
Added `extract_uri_scheme(uri: &str) -> Option<&str>` in
`src/scanner/heuristics.rs` implementing the RFC 3986 §3.1 scheme
grammar:
```
scheme = ALPHA *( ALPHA / DIGIT / "+" / "-" / "." ) ":"
```
The scheme is the longest prefix of `uri` that matches that grammar,
ending at the first `:`. If no such prefix exists (i.e. there is no
`:` in the URI, or the part before `:` doesn't match the scheme
grammar), the URI has no scheme — it's an in-document link.
| URI | Extracted scheme | In allowed list? | Result |
|------------------------------|-------------------|------------------|-------------|
| `https://example.com` | `Some("https")` | yes | phishing check → Benign |
| `#section-2` | `None` | n/a | phishing check → Benign |
| `./chapter2.html` | `None` | n/a | phishing check → Benign |
| `page.html#anchor` | `None` | n/a | phishing check → Benign |
| `mailto:user@example.com` | `Some("mailto")` | yes | phishing check → Benign |
| `javascript:alert(1)` | `Some("javascript")` | no | Malicious ✓ |
| `data:text/html,<x>` | `Some("data")` | no | Malicious ✓ |
| `vbscript:msgbox` | `Some("vbscript")`| no | Malicious ✓ |
### 2. In-document links still get phishing-checked
The user's intent — paraphrased — was:
> An in-document link to another region of the same document shouldn't
> raise a flag. It should be used to check for *other* flags, but not
> raise an alert itself.
So `classify_uri` now routes scheme-less URIs through
`classify_phishing_signal(uri, config)` (a small helper extracted from
the original logic). If a strong phishing signal fires
(`BrandHomograph`, `IpHost`, `CredentialUrl`) the URI is still
classified as `Malicious(SuspiciousUri)`. If a weak signal fires
(`Shortener`, `SuspiciousKeyword`, `PhishingTld`) it's still
`Suspicious`. Only when no phishing signal fires does the URI become
`Benign`.
Examples of in-document links that STILL get flagged (correctly):
| URI | Phishing signal | Result |
|------------------------------|------------------------|-------------|
| `#login-verify` | `SuspiciousKeyword` ("login", "verify") | Suspicious |
| `#micros0ft-attack-vector` | `BrandHomograph` ("micros0ft") | Malicious |
| `./page.html?account=verify` | `SuspiciousKeyword` | Suspicious |
### 3. PDF `/GoTo` (in-document) vs `/GoToR` (remote)
Added a new `VectorType::PdfGoTo` variant in `src/core/types.rs`
(documented as "PDF `/GoTo` action — in-document navigation").
Updated the PDF parser to dispatch on the action type:
```rust
"GoTo" => { /* emit VectorType::PdfGoTo */ }
"GoToR" => { /* emit VectorType::PdfGoToR */ }
```
Updated the heuristics `classify_vector` match:
```rust
// PDF /GoTo — in-document navigation. Benign.
VectorType::PdfGoTo => ThreatClassification::Benign,
// PDF /GoToR — remote navigation. Still Suspicious.
VectorType::PdfGoToR => ThreatClassification::Suspicious,
```
Because `inspect_vector` returns `None` for `Benign` findings, an
in-document `/GoTo` produces no alert and no quarantine entry. The
vector is still recorded in `Document::executable_vectors` so the
operator can see in-document navigation activity if they want to.
## Files changed
| File | Change |
|------|--------|
| `src/core/types.rs` | Added `VectorType::PdfGoTo` variant + `"pdf-goto"` display string. Updated doc comment on `PdfGoToR` to clarify it's for REMOTE navigation only. |
| `src/parsers/pdf_parser.rs` | Split `"GoToR" \| "GoTo"` match arm in `inspect_action` into two separate arms emitting `PdfGoTo` (in-document) and `PdfGoToR` (remote) respectively. |
| `src/scanner/heuristics.rs` | Added `VectorType::PdfGoTo => ThreatClassification::Benign` case in `classify_vector`. Replaced broken `split("://")` scheme extraction with proper RFC 3986 `extract_uri_scheme` helper. Refactored phishing-signal escalation into `classify_phishing_signal` helper that's now called for BOTH scheme-bearing and scheme-less URIs (so in-document links still get phishing-checked). |
| `src/scanner/heuristics.rs` (tests) | Added 10 new tests covering: RFC 3986 scheme extraction, fragment links, relative URLs, bare page links, `mailto:` links, in-document links with phishing signals, and the new `PdfGoTo`/`PdfGoToR` distinction. |
## Tests added
In `src/scanner/heuristics.rs::tests`:
| Test name | What it asserts |
|-----------|-----------------|
| `extract_uri_scheme_handles_rfc3986_cases` | The new scheme extractor handles hierarchical schemes, opaque schemes (`mailto:`, `javascript:`, `data:`, `vbscript:`), schemes with digits/+/-/., and correctly returns `None` for fragment links, relative URLs, and empty URIs. |
| `fragment_link_is_benign` | `#section` (Markdown) → no finding emitted. |
| `relative_url_is_benign` | `./page.html` → no finding emitted. |
| `bare_page_link_is_benign` | `page.html` → no finding emitted. |
| `fragment_with_anchor_is_benign` | `chapter1.html#section-2` → no finding emitted. |
| `mailto_link_is_benign` | `mailto:user@example.com` → no finding emitted (was a false positive before the fix). |
| `in_document_link_with_phishing_keyword_still_flagged` | `#login-verify` → Suspicious finding with `suspicious-keyword` note (proves in-document links still get phishing-checked). |
| `in_document_link_with_brand_homograph_still_flagged` | `#micros0ft-attack-vector` → Malicious finding with `brand-homograph` note. |
| `pdf_goto_in_document_navigation_is_benign` | `VectorType::PdfGoTo` → no finding emitted. |
| `pdf_gotor_remote_navigation_is_suspicious` | `VectorType::PdfGoToR` → Suspicious (preserved behavior). |
## Test results
Before the fix:
- `cargo test --lib` → 127 passed, 0 failed
- `cargo test --test pipeline_integration` → 23 passed, 0 failed
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed
After the fix:
- `cargo test --lib` → **137 passed**, 0 failed (+10 new tests)
- `cargo test --test pipeline_integration` → 23 passed, 0 failed (unchanged)
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed (unchanged)
All pre-existing tests continue to pass, including the suite that
exercises real PDF / DOCX / EPUB / Markdown fixtures in
`tests/fixtures/`.
## Known limitations / non-goals
1. **Windows file paths** like `C:\path\to\file.txt` will be parsed
by `extract_uri_scheme` as having scheme `"C"` (because `C` is a
valid RFC 3986 scheme character), and the heuristics will flag it
as Malicious because `"C"` isn't in `Config::allowed_uri_schemes`.
This is an acceptable false positive for an unusual input — Windows
file paths should be encoded as `file:///C:/path/to/file` in URIs.
The same applies to single-letter drive prefixes generally.
2. **The PDF parser doesn't yet capture `/GoTo` destination payloads.**
Both `PdfGoTo` and `PdfGoToR` vectors still have empty
`raw_payload`. Capturing the `/D` (destination) entry for `/GoTo`
and the `/F` (file reference) entry for `/GoToR` would let the
scanner log where in-document navigation is actually pointing, but
it's an enhancement — not required for the false-positive fix.
3. **Phishing-keyword matching against fragments is still substring
based.** `#login-verify` triggers `SuspiciousKeyword` because
`"login"` and `"verify"` are both in `SUSPICIOUS_URL_KEYWORDS` and
matching is `lower.contains(kw)`. This is by design — the user
explicitly said in-document links should still be checked for other
flags. If substring matching proves too noisy on real documents,
the keyword matcher in `src/scanner/signatures.rs::match_suspicious_keyword`
can be tightened to word-boundary matching in a follow-up.

View File

@ -1,22 +1,43 @@
# Quick Start Guide
# Build & Operation Guide
Get CorbelPurge built and running in under five minutes.
This document is the authoritative build guide and operator reference
for CorbelPurge. For the project overview and detection model, see
[README.md](README.md).
## Prerequisites
- **Rust** 1.70+ (install via [rustup](https://rustup.rs/))
- **Python 3** (only needed if you want to regenerate test fixtures)
| Tool | Version | Purpose |
|---|---|---|
| **Rust** | 1.70+ stable | Compiles the scanner, CLI, and GUI |
| **Python 3** | any | Only needed to regenerate test fixtures |
That is it. The headless CLI has zero GUI dependencies and no native libraries.
Install Rust via [rustup](https://rustup.rs/):
## Step 1: Get the Source
```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source "$HOME/.cargo/env"
```
The headless CLI has zero GUI dependencies and no native libraries.
The GUI adds `iced`, `rfd`, and `tokio` (all pure-Rust).
## Step 1 — Get the source
From a release tarball:
```bash
tar xzf corbel-purge-0.4.2.tar.gz
cd corbel-purge-0.4.2
```
## Step 2: Build the CLI
From the repository:
```bash
git clone https://git.dcos.net/dcosnet/corbel.git
cd corbel
```
## Step 2 — Build the CLI
```bash
cargo build --release
@ -28,89 +49,158 @@ The binary lands at `target/release/corbel-purge`. Verify it works:
./target/release/corbel-purge --help
```
You should see the usage banner listing `scan`, `scan-dir`, `study`, and the
supported flags (`--workspace`, `--abort-on-threat`, `--quiet`, `--recursive`,
`--preserve-format`, `--rules`, `--cve-db`).
You should see the usage banner listing `scan`, `scan-dir`, `study`,
and the supported flags (`--workspace`, `--abort-on-threat`, `--quiet`,
`--recursive`, `--preserve-format`, `--rules`, `--cve-db`,
`--no-clean-output`, `--move-clean`).
## Step 3: Scan Your First File
### Building the GUI (optional)
The iced 0.13 dashboard GUI is behind the `gui` feature flag:
```bash
cargo build --release --features gui --bin corbel-purge-gui
./target/release/corbel-purge-gui
```
### Build profiles
The release profile (in `Cargo.toml`) is tuned for production:
```toml
[profile.release]
opt-level = 3
lto = "thin"
codegen-units = 1
strip = "symbols"
```
For faster debug builds (no optimizations, faster compile):
```bash
cargo build # debug profile
```
## Step 3 — Run the test suite
```bash
cargo test
```
Expected output: 154 lib tests + 24 pipeline integration tests + 4
zip-bomb defense tests, all passing. The pipeline integration tests
include the corpus regression test
(`benign_realworld_pdf_produces_zero_findings`) which asserts that a
3-page PDF with 13 hyperlinks (kernel.org, linuxfromscratch.org,
github.com/microsoft/vscode, .ru URLs, mailto:, tel:, and URLs
containing "support", "account", "verify" in their paths) produces
zero findings.
### Regenerating test fixtures
The fixtures in `tests/fixtures/` are checked in. Regenerate them
only when changing the parser or scanner behavior:
```bash
pip install pypdf reportlab python-docx
python3 scripts/gen_fixtures.py
python3 scripts/gen_md_epub_fixtures.py
python3 scripts/gen_docx_fixtures.py
python3 scripts/gen_zip_bomb_fixtures.py
python3 scripts/gen_benign_realworld_pdf.py
```
## Step 4 — Scan your first file
```bash
# Scan a suspicious PDF you received
./target/release/corbel-purge scan suspicious_document.pdf
```
If the file contains threats, you will see a summary block printed to your
terminal (findings count, classification, SHA-256, output paths), and the
pipeline will write:
If the file contains threats, the summary block prints to stdout and
the pipeline writes:
- A quarantine tarball in `./corbel_quarantine/` (`quarantine_<ts>_<sha>.tar.gz`)
containing `original.<ext>`, `report.json`, `report.md`, and one `.bin` per
carved payload (plus paired `.hex` and `.info` files for each payload)
- A quarantine tarball in `./corbel_quarantine/`
(`quarantine_<ts>_<sha>.tar.gz`) containing `original.<ext>`,
`report.json`, `report.md`, and one `.bin` per carved payload
(plus paired `.hex` and `.info` files for each payload).
- A cleansed Markdown derivative in `./corbel_clean/`
- Standalone `report_<timestamp>_<sha>.json` and `.md` for programmatic access
(`cleansed_<ts>_<sha>.md`).
- Standalone `report_<ts>_<sha>.json` and `.md` for programmatic
access in `./corbel_quarantine/`.
If the file is clean, the summary block simply reports zero findings and no
quarantine output is written.
If the file is clean (zero malicious findings), the pipeline writes:
## Step 4: Try PreserveFormat Mode
- A clean-output copy in `./corbel_clean/`
(`clean_<ts>_<sha>.<ext>`) — the source file byte-for-byte.
- Standalone `report_<ts>_<sha>.json` and `.md` confirming the scan
ran and the document was clean.
If you want a cleaned version that keeps the original format (e.g. a cleaned
`.epub` you can actually read in an e-reader):
## Step 5 — Try PreserveFormat mode
For a cleaned version that keeps the original format (e.g. a cleaned
`.epub` you can read in an e-reader):
```bash
./target/release/corbel-purge scan research_paper.epub --preserve-format
```
This produces a `cleansed_<ts>_<sha>.epub` with malicious entries stripped from
the ZIP container but chapter text preserved. Works for PDF and DOCX too.
Produces `cleansed_<ts>_<sha>.epub` with malicious entries stripped
from the ZIP container but chapter text preserved. Works for PDF and
DOCX too.
## Step 5: Study a Document In Place
## Step 6 — Pipeline-stage mode
The `study` subcommand renders the original document to a single annotated HTML
file with malicious regions wrapped in inline `<span>` tags, color-coded by
classification. Use it when you want to see exactly where the exploit sits in
context, without leaving the source format:
The scanner acts as a pipeline stage: input files flow through and
clean ones end up in the output folder alongside the cleansed
derivatives of malicious ones.
```bash
./target/release/corbel-purge study suspicious.epub
# Default: copy clean files to ./corbel_clean/
./target/release/corbel-purge scan inbox/file.pdf
# Queue-draining: move (not copy) clean files to output, remove source
./target/release/corbel-purge scan inbox/file.pdf --move-clean
# Reports only, no clean-output copy
./target/release/corbel-purge scan inbox/file.pdf --no-clean-output
```
Output lands at `study_<ts>_<sha>.html` in the workspace. Quiet mode (`-q`)
prints just the path.
Downstream processing can then operate on the contents of
`corbel_clean/` without inspecting each file's report — every file
there has been verified clean.
## Step 6: Scan a Directory (Optional)
## Step 7 — Scan a directory
```bash
# Recursively scan an inbox directory
./target/release/corbel-purge scan-dir /path/to/inbox --recursive --workspace /tmp/corbel
```
The directory walker picks up `.pdf`, `.epub`, `.md`, `.markdown`, and `.docx`
files. Each file is logged with a `[OK]`, `[MALICIOUS]`, or `[ERROR]` tag.
The directory walker picks up `.pdf`, `.epub`, `.md`, `.markdown`, and
`.docx` files. Each file is logged with a `[OK]`, `[MALICIOUS]`, or
`[ERROR]` tag.
## Step 7: CI Integration (Optional)
## Step 8 — CI gate
Use `--abort-on-threat` to make CorbelPurge a CI gate. Exit code 2 means
threats were found:
Use `--abort-on-threat` to make CorbelPurge a CI gate. Exit code 2
means threats were found:
```bash
# In your CI pipeline
./target/release/corbel-purge scan incoming_document.pdf --abort-on-threat --quiet
# exit 0: clean
# exit 2: has threats -> fail the build
# exit 1: hard error (parse failure, IO, etc.)
```
`--quiet` suppresses the summary block and prints only the JSON report path,
which is handy for piping into downstream tooling.
`--quiet` suppresses the summary block and prints only the JSON
report path, which is handy for piping into downstream tooling.
## Step 8: Plug In External Threat-Intel Feeds (Optional)
## Step 9 — External threat-intel feeds (optional)
The built-in signature tables and CVE database are static, but you can layer
your own on top at runtime:
The built-in signature tables and CVE database are static. Layer your
own on top at runtime:
```bash
# Load additional YARA-style signature rules
# Load additional signature rules
./target/release/corbel-purge scan suspicious.pdf --rules my_rules.json
# Load additional CVE signature entries
@ -123,73 +213,172 @@ export CORBEL_EXTERNAL_CVE_DB=/etc/corbel/cve_db.json
```
External rules are matched alongside the built-in tables; nothing is
overridden. See `MANIFEST.md` for the JSON schema.
overridden. See `MANIFEST.md` for the JSON schema. Supported
external-rule types:
## Step 9: Build the GUI (Optional)
- `homograph-host-list` — additional exact-match homograph host strings
- `signature-list` — additional `(offset, magic_bytes, name)` entries
- `shellcode-list` — additional shellcode prologue byte patterns
The iced 0.13 dashboard GUI is behind the `gui` feature flag:
(Legacy `tld-list`, `keyword-list`, and `brand-list` rule types are
silently skipped — the scanner no longer consults those tables.)
## Step 10 — Study a document in place
The `study` subcommand renders the source document to a single
annotated HTML file with malicious regions wrapped in inline `<span>`
tags, color-coded by classification. Use it when you want to see
exactly where the exploit sits in context, without leaving the source
format:
```bash
cargo build --release --features gui --bin corbel-purge-gui
./target/release/corbel-purge-gui
./target/release/corbel-purge study suspicious.epub
```
The GUI provides file pickers, toggle switches for preserve-format /
abort-on-threat / recursive, a timestamped console log, a sidebar with a
Unicode progress gauge and per-file stats, and a cleansed-document viewer.
All scanning runs through the same `Pipeline` the CLI uses, via
`tokio::spawn_blocking`.
Output lands at `study_<ts>_<sha>.html` in the workspace. Quiet mode
(`-q`) prints just the path.
## What You Should See
## CLI flags reference
### Clean file output:
| Flag | Default | Purpose |
|---|---|---|
| `--workspace <dir>` | CWD | Root for `corbel_quarantine/` and `corbel_clean/` output dirs |
| `--abort-on-threat` | off | Exit with code 2 if any malicious finding fires |
| `--quiet`, `-q` | off | Print only the JSON report path on success |
| `--recursive`, `-r` | off | Recurse into subdirectories (`scan-dir`) |
| `--preserve-format` | off | Repackage cleansed document in original format |
| `--rules <path>` | unset | Path to external signature-rules JSON |
| `--cve-db <path>` | unset | Path to external CVE database JSON |
| `--no-clean-output` | off | Do NOT copy clean documents to output folder |
| `--move-clean` | off | Move (not copy) clean source files to output folder |
## Exit codes
| Code | Meaning |
|---|---|
| `0` | No threats found |
| `1` | Hard error (parse failure, IO failure, etc.) |
| `2` | One or more malicious findings (file was processed) |
## Environment variables
| Variable | Default | Purpose |
|---|---|---|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Quarantine tarball + reports output dir |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Cleansed document + clean-output copy dir |
| `CORBEL_ABORT_ON_THREAT` | `false` | Abort on first malicious finding |
| `CORBEL_EMIT_SUSPICIOUS` | `false` | Include Suspicious findings (no default detector produces any) |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Generate Markdown report alongside JSON |
| `CORBEL_EMIT_CLEAN_OUTPUT` | `true` | Copy clean documents to output folder |
| `CORBEL_MOVE_CLEAN_TO_OUTPUT` | `false` | Move (not copy) clean source files to output |
| `CORBEL_TOTAL_ARCHIVE_SCAN_CAP` | `268435456` (256 MiB) | Cumulative cap across all entries in a multi-entry archive |
| `CORBEL_EXTERNAL_RULES` | (unset) | Path to external signature-rules JSON |
| `CORBEL_EXTERNAL_CVE_DB` | (unset) | Path to external CVE database JSON |
## Expected output examples
### Clean file
```
────────────────────────────────────────────────────────
scan complete: PDF
────────────────────────────────────────────────────────────
scan complete: pdf
source: benign.pdf
sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
text nodes: 12
vectors: 0
findings: 0 malicious, 0 educational, 0 total
────────────────────────────────────────────────────────
cleansed: ./corbel_clean/clean_20260801T120000_9f86d081.pdf
json report: ./corbel_quarantine/report_20260801T120000_9f86d081.json
md report: ./corbel_quarantine/report_20260801T120000_9f86d081.md
────────────────────────────────────────────────────────────
```
### Threat found output:
### Threat found
```
────────────────────────────────────────────────────────
scan complete: PDF
────────────────────────────────────────────────────────────
scan complete: pdf
source: suspicious.pdf
sha256: a1b2c3...
text nodes: 8
vectors: 1
findings: 1 malicious, 0 educational, 1 total
quarantine: ./corbel_quarantine/quarantine_20260801T120000_abc12345.tar.gz
cleansed: ./corbel_clean/cleansed_20260801T120000_abc12345.md
cleansed: ./corbel_clean/cleansed_20260801T120000_abc12345.pdf
json report: ./corbel_quarantine/report_20260801T120000_abc12345.json
md report: ./corbel_quarantine/report_20260801T120000_abc12345.md
────────────────────────────────────────────────────────
────────────────────────────────────────────────────────────
```
Exit code is `2` when any malicious finding is produced.
## Key Environment Variables
## Troubleshooting
| Variable | Default | When to Change It |
|----------|---------|-------------------|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Point at a shared quarantine volume |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Point at an output directory for cleaned files |
| `CORBEL_ABORT_ON_THREAT` | `false` | Set to `true` in CI pipelines |
| `CORBEL_EMIT_SUSPICIOUS` | `true` | Set to `false` to only report Malicious (not Suspicious) |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Set to `false` if you only need JSON |
| `CORBEL_TOTAL_ARCHIVE_SCAN_CAP` | `268435456` (256 MiB) | Cumulative cap across all entries in a multi-entry archive |
| `CORBEL_EXTERNAL_RULES` | (unset) | Path to an external signature-rules JSON file |
| `CORBEL_EXTERNAL_CVE_DB` | (unset) | Path to an external CVE database JSON file |
### Build fails on `lopdf` or `zip`
## Next Steps
These crates occasionally need a newer Rust than the MSRV declared in
their `Cargo.toml`. Update Rust:
- Read [README.md](README.md) for the full feature overview and security model
```bash
rustup update stable
```
### GUI binary not found
The GUI is behind the `gui` feature flag. Build with:
```bash
cargo build --release --features gui --bin corbel-purge-gui
```
### Scanner reports zero findings on a document I expected to be flagged
Confirm the document actually contains a Category 1 or Category 2
detector trigger (see [README.md](README.md) for the table). The
scanner does not flag based on reputation, file source, or filename.
Run with `cargo run -- scan file.pdf` for verbose output.
### Quarantine tarball missing
The tarball is written only when `malicious_count() > 0`. Clean
documents produce only `report_<ts>_<sha>.{json,md}` and a
`clean_<ts>_<sha>.<ext>` copy in the output folder.
### Test failure on `benign_realworld_pdf_produces_zero_findings`
This is the corpus regression test. If it fails, a detector is
wrong by construction — the detector fired on a clean document. Inspect
the test's failure output for the specific detector that fired and
tighten that detector's rule.
## Installation
After building, install the binary to a system path:
```bash
cargo install --path .
# or
sudo cp target/release/corbel-purge /usr/local/bin/
```
For system-wide configuration, set environment variables in
`/etc/corbel/env` or your shell profile:
```bash
export CORBEL_QUARANTINE_DIR=/var/lib/corbel/quarantine
export CORBEL_CLEANSE_DIR=/var/lib/corbel/clean
export CORBEL_EXTERNAL_RULES=/etc/corbel/rules.json
export CORBEL_EXTERNAL_CVE_DB=/etc/corbel/cve_db.json
```
## Next steps
- Read [README.md](README.md) for the project overview and detection model
- Read [MANIFEST.md](MANIFEST.md) for the detailed technical specification
- Read [TODO.md](TODO.md) for the development roadmap
- Run `cargo test` to verify all 154 tests pass in your environment
- Read [FIX-NOTES-false-positive-redesign.md](FIX-NOTES-false-positive-redesign.md)
for the design notes on the two-category detector model
## Contact
**Jeremy Anderson** — [dcos.net](https://dcos.net) — [info@dcos.net](mailto:info@dcos.net)

257
README.md
View File

@ -2,91 +2,147 @@
> Strict Rust document sanitizer & threat neutralizer for PDF, EPUB, Markdown, and DOCX.
**Author:** Jeremy Anderson — [https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)
**Author:** Jeremy Anderson — [dcos.net](https://dcos.net) — [info@dcos.net](mailto:info@dcos.net)
**Repository:** [https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)
**License:** GPL-3.0-or-later
![Corbel-Purge-ss](./corbel-purge-ss.png)
CorbelPurge parses documents into a unified intermediate representation (`Document` struct), runs a layered contextual scanner that distinguishes educational security literature from active malicious injections, and produces cleansed derivatives with all executable content stripped. Malicious payloads are carved into quarantine tarballs with full forensic reports.
## Overview
CorbelPurge is a strict, local-only document security scanner. It parses
documents into a unified intermediate representation, runs detectors
that are backed by verifiable properties of the bytes on disk, and
produces cleansed derivatives with all executable content stripped.
Malicious payloads are carved into quarantine tarballs with full
forensic reports.
The scanner is built around a single design invariant: **every
detector must be backed by a verifiable property, either of the
document itself or of an external authority.** No thresholds, no
per-file or per-domain exceptions, no statistical "suspicious" tier.
## What it does
Given a file path, `Pipeline::run()` in `src/core/pipeline.rs`:
1. **Detects format** via `DocumentFormat::from_path()` (extension-based: `.pdf`, `.epub`, `.md`/`.markdown`, `.docx`).
2. **Parses** via `parsers::Dispatcher` into a `Document` containing `TextNode` (static text with semantic context) and `ExecutableVector` (active content like JS streams, embedded files, script tags, VBA macros) items.
1. **Detects format** via `DocumentFormat::from_path()` (extension-based:
`.pdf`, `.epub`, `.md`/`.markdown`, `.docx`).
2. **Parses** via `parsers::Dispatcher` into a `Document` containing
`TextNode` (static text with semantic context) and
`ExecutableVector` (active content like JS streams, embedded files,
script tags, VBA macros) items.
3. **Scans** with a two-pass engine:
- `heuristics::inspect_vector()` classifies every executable vector against file-signature tables, shellcode patterns, phishing heuristics, and URI allowlists.
- `context_filter::evaluate()` checks text nodes for suspicious signatures and determines whether the surrounding context is educational (code blocks, CVE writeups, academic language) or weaponized.
- `heuristics::inspect_vector()` classifies every executable vector
against the two-category detector model: Category 1 (verifiable
executable intent — file signatures, shellcode prologues,
executable URI schemes) and Category 2 (verifiable impersonation —
exact-host homographs, credential URLs, mixed-script hosts).
- `context_filter::evaluate()` checks text nodes for structural
signatures (`/JavaScript`, `<script`, etc.) and weaponization
indicators (long hex runs, base64 blobs, multi-shell commands).
Words and function names that appear in legitimate technical
literature (`wget`, `exploit`, `payload`, `eval(`) are not
treated as signatures.
- `cve_tags::match_cve()` annotates findings with known exploit IDs.
4. **Quarantines** (when malicious findings exist) — `quarantine::handle()` carves payloads into `quarantine_<ts>_<sha>.tar.gz` with `original.<ext>`, `report.json`, `report.md`, and one `.bin` per payload.
5. **Cleanses** (when recommended) — produces a sanitized derivative:
- **Markdown mode** (default): `cleanse::sanitizer::sanitize()` emits a safe Markdown file. Text nodes at malicious locations are stripped; hyperlinks lose their destinations.
- **PreserveFormat mode** (`--preserve-format`): `cleanse::repackage::repackage()` rebuilds the original format with malicious entries removed. EPUB entries are stripped from the ZIP, DOCX macros/embeddings/external-links are removed, PDF objects are deleted via lopdf.
4. **Reports** — both clean and malicious scans produce a JSON and
Markdown report in `corbel_quarantine/` named
`report_<timestamp>_<sha_prefix>.{json,md}`. The clean path provides
an audit trail; the malicious path adds a quarantine tarball and
carved payloads.
5. **Quarantines** (when malicious findings exist) — `quarantine::handle()`
carves payloads into `quarantine_<ts>_<sha>.tar.gz` with
`original.<ext>`, `report.json`, `report.md`, and one `.bin` per
payload (plus paired `.hex` and `.info` files).
6. **Cleanses** (when recommended) — produces a sanitized derivative:
- **Markdown mode** (default): `cleanse::sanitizer::sanitize()`
emits a safe Markdown file. Text nodes at malicious locations
are stripped; hyperlinks lose their destinations.
- **PreserveFormat mode** (`--preserve-format`):
`cleanse::repackage::repackage()` rebuilds the original format
with malicious entries removed. EPUB entries are stripped from
the ZIP, DOCX macros/embeddings/external-links are removed, PDF
objects are deleted via lopdf.
7. **Clean-output** (when scan is clean and `emit_clean_output` is on,
the default) — copies the source file to
`corbel_clean/clean_<ts>_<sha>.<ext>` so the scanner acts as a
pipeline stage. Use `--move-clean` for queue-draining semantics.
Supported formats: **PDF**, **EPUB**, **Markdown**, **DOCX**.
## Build
## Detection model
```bash
# Headless CLI (default — no GUI deps)
cargo build --release
Two categories. Nothing else fires.
# With iced GUI
cargo build --release --features gui --bin corbel-purge-gui
### Category 1 — Verifiable executable intent
The vector contains a structure whose only purpose is to execute
code or spawn a process. Presence is the threat.
| Detector | Triggers on |
|---|---|
| Active script in PDF | `/JavaScript` or `/JS` action stream |
| Program launch in PDF | `/Launch` action with `/F`, `/Win`, `/Mac`, `/Unix` |
| External program exec in EPUB | `<script>` tag in XHTML |
| VBA macro in DOCX | `word/vbaProject.xml` present in package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 |
| Executable URI scheme | `javascript:`, `vbscript:`, `data:text/html`, `file:` in a hyperlink or action |
| PDF form with `/AA` | AcroForm dictionary contains the `/AA` (Additional Actions) entry |
| PDF widget with `/AA` | Widget annotation with `/AA` entry |
| Shellcode prologue | Known Metasploit / NOP-sled / syscall-stub byte sequences |
### Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on |
|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (`micros0ft.com`, `paypa1.com`, etc.) |
| Credential URL | RFC 3986 authority component contains `user:pass@` (before the first `/` after `://`) |
| Mixed-script host | URL host mixes Latin with Cyrillic or Greek (a property of the codepoints) |
### What is NOT a detector
The following were removed because they are statistical guesses
about the world, not verifiable properties of the document:
- **Substring keyword matching** — words like `login`, `verify`,
`support`, `account` are not threats. Every legitimate login page
contains them.
- **Phishing TLDs** — `.ru`, `.cn`, `.xyz` are not threats. A Russian
URL is a Russian URL.
- **URL shorteners as a threat signal** — `bit.ly` is not a threat.
If the destination is hostile, the underlying detector catches it.
- **Substring brand matching** — `https://github.com/microsoft/vscode`
is not a threat. The host is `github.com`; the brand substring in
the path is irrelevant.
- **IP-address hosts** — `192.168.1.1` is a valid network address.
RFCs and router manuals reference them.
- **Substring text-node signatures** — words like `wget`, `exploit`,
`payload`, `eval(`, `powershell`, `/bin/sh` in prose are not
threats. Only structural tokens (`/JavaScript`, `<script`,
`shellcode`) and weaponization indicators (hex runs, base64 blobs,
multi-shell commands) fire.
## Pipeline architecture
```text
input path ──► parser ──► Document (UIR)
│
▼
scanner ──► ScanReport
│
┌────────────────────┼────────────────────┐
│ │ │
▼ ▼ ▼
(clean) (malicious) (educational)
│ │ │
▼ ▼ ▼
clean_<ts>_<sha>.<ext> quarantine tarball whitelisted
+ report.json / .md + carved payloads (logged only)
+ cleansed file
```
Requires Rust 1.70+ (stable).
Binaries:
- `target/release/corbel-purge` — headless CLI
- `target/release/corbel-purge-gui` — iced dashboard GUI (`gui` feature)
## CLI usage
The CLI is in `src/main.rs` and parses its own arguments (no `clap` dependency).
```bash
# Scan a single file
corbel-purge scan path/to/suspicious.pdf
# Scan with format-preserving output (keeps original format)
corbel-purge scan path/to/book.epub --preserve-format
# Scan with a custom workspace
corbel-purge scan path/to/file.pdf --workspace /tmp/corbel
# Scan a directory recursively
corbel-purge scan-dir path/to/inbox --recursive --workspace /tmp/corbel
# CI gate: exit non-zero on threat
corbel-purge scan path/to/file.pdf --abort-on-threat
# Quiet mode: print only the JSON report path
corbel-purge scan path/to/file.pdf --quiet
# Version
corbel-purge --version
```
### Exit codes
| Code | Meaning |
|------|----------|
| 0 | No threats found |
| 1 | Hard error (parse failure, IO, etc.) |
| 2 | One or more malicious findings |
### Environment variables
| Variable | Default | Purpose |
|----------|---------|---------|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Quarantine tarball output dir |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Cleansed document output dir |
| `CORBEL_ABORT_ON_THREAT` | `false` | Abort on first malicious finding |
| `CORBEL_EMIT_SUSPICIOUS` | `true` | Include Suspicious findings in report |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Generate Markdown report alongside JSON |
## Library usage
```rust
@ -116,7 +172,7 @@ let result = pipeline.run("book.epub")?;
```text
src/
├── main.rs # CLI: scan, scan-dir, --version, arg parsing
├── main.rs # CLI: scan, scan-dir, study, --version, arg parsing
├── lib.rs # Crate root, CorbelError enum, public re-exports
├── util.rs # sha256_hex(), truncate_with_ellipsis(), read_with_cap()
├── bin/
@ -135,19 +191,20 @@ src/
│ └── docx_parser.rs # OOXML ZIP inspector (regex XML extraction)
├── scanner/
│ ├── mod.rs # scan() entrypoint, CVE tag injection
│ ├── heuristics.rs # classify_vector() for executable vectors
│ ├── context_filter.rs # evaluate() for text nodes, educational vs weaponized
│ ├── signatures.rs # PHISHING_TLDS, KNOWN_FILE_SIGNATURES,
│ │ # SHELLCODE_PATTERNS, COMMON_PHISHING_BRANDS,
│ │ # SUSPICIOUS_URL_KEYWORDS, URL_SHORTENER_DOMAINS
│ └── cve_tags.rs # CVE_TABLE (7 entries), match_cve(), cve_tag()
│ ├── heuristics.rs # classify_vector() — two-category detector model
│ ├── context_filter.rs # evaluate() — structural signatures + weaponization
│ ├── signatures.rs # KNOWN_FILE_SIGNATURES, SHELLCODE_PATTERNS,
│ │ # HOMOGRAPH_HOSTS, URL_SHORTENER_DOMAINS (retained
│ │ # for future reputation work, not consulted)
│ └── cve_tags.rs # CVE_TABLE, match_cve(), cve_tag()
├── quarantine/
│ ├── mod.rs # handle() — tarball writer, QuarantineOutcome
│ ├── mod.rs # handle() — tarball writer, write_clean_report()
│ ├── extractor.rs # extract_payloads(), ExtractedPayload, filename sanitization
│ ├── hexdump.rs # hex_dump(), build_payload_info()
│ └── reporter.rs # build_json_report(), build_markdown_report()
├── cleanse/
│ ├── mod.rs # cleanse() dispatcher (Markdown vs PreserveFormat)
│ ├── sanitizer.rs # sanitize() — Markdown re-serializer, build_clean_markdown()
│ ├── sanitizer.rs # sanitize() — Markdown re-serializer
│ └── repackage.rs # repackage() — format-preserving (EPUB/DOCX/PDF)
└── ui/
├── mod.rs # GUI module gate (behind `gui` feature)
@ -155,29 +212,34 @@ src/
└── alert_modal.rs # Stub
```
## Testing
## Design philosophy
```bash
# Regenerate test fixtures (requires pypdf, reportlab, python-docx)
pip install pypdf reportlab python-docx
python3 scripts/gen_fixtures.py
python3 scripts/gen_md_epub_fixtures.py
python3 scripts/gen_docx_fixtures.py
python3 scripts/gen_zip_bomb_fixtures.py
# Run all tests
cargo test
```
1. **Verifiable properties only.** A detector either proves a finding
by a property of the bytes, or it doesn't fire. No thresholds.
2. **No special-casing.** No per-file or per-domain exceptions. The
clean-document guarantee is a theorem about the detectors, not an
empirical observation about one specific file.
3. **Unix philosophy.** Each module does one thing. The pipeline is a
linear sequence of small steps. Step-down logic (early returns) for
every fork of choices.
4. **`#![forbid(unsafe_code)]`** at the crate root. The scanner never
touches `unsafe` Rust.
5. **Local-only.** No network access during scanning. All analysis is
against the bytes on disk. External threat-intel feeds are loaded
from local JSON files at startup.
## Adding a new format
1. Create `src/parsers/<format>_parser.rs` implementing `DocumentParser`.
2. Add a variant to `DocumentFormat` in `src/core/types.rs` (and its `Display` impl).
2. Add a variant to `DocumentFormat` in `src/core/types.rs` (and its
`Display` impl).
3. Add a match arm in `Dispatcher::parse()` in `src/parsers/mod.rs`.
4. Add an extension check in `DocumentFormat::from_path()` in `src/core/pipeline.rs`.
4. Add an extension check in `DocumentFormat::from_path()` in
`src/core/pipeline.rs`.
5. Add repackage support in `src/cleanse/repackage.rs` (optional).
The scanner, quarantine, and cleanse modules consume only the `Document` UIR and require no changes for basic format support.
The scanner, quarantine, and cleanse modules consume only the
`Document` UIR and require no changes for basic format support.
## Non-goals
@ -186,6 +248,17 @@ The scanner, quarantine, and cleanse modules consume only the `Document` UIR and
- No network access during scanning — all analysis is local.
- No sandbox hardening of the tool itself.
## Documentation
- **[QUICKSTART.md](QUICKSTART.md)** — Build guide and operator reference
- **[MANIFEST.md](MANIFEST.md)** — Detailed technical specification
- **[TODO.md](TODO.md)** — Development roadmap
- **[FIX-NOTES-false-positive-redesign.md](FIX-NOTES-false-positive-redesign.md)** —
Design notes for the two-category detector model
- **[FIX-NOTES-indoc-links.md](FIX-NOTES-indoc-links.md)** — In-document link handling
- **[FIX-NOTES-glyphfix.md](FIX-NOTES-glyphfix.md)** — Glyph rendering fixes
## License
GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson. https://git.dcos.net/dcosnet/corbel
GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson.
[https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)

View File

@ -0,0 +1,162 @@
#!/usr/bin/env python3
"""Generate a realistic benign-PDF fixture for the false-positive regression test.
The fixture mimics the structure of a technical book like *Linux from
Scratch*: a multi-page PDF with a table of contents containing
in-document links, body paragraphs containing URLs to kernel.org and
linuxfromscratch.org, mailing-list URLs that contain "support" in the
path, mailto: links to authors, and a page that mentions `wget`,
`exploit`, `payload`, etc. in prose.
This is the regression corpus for the false-positive fix. The integration
test asserts that scanning this PDF produces ZERO findings.
"""
from pathlib import Path
from reportlab.pdfgen import canvas
from reportlab.lib.pagesizes import letter
import pypdf
from pypdf.generic import (
ArrayObject,
DictionaryObject,
NameObject,
NumberObject,
TextStringObject,
)
FIXTURES_DIR = Path(__file__).resolve().parent.parent / "tests" / "fixtures"
FIXTURES_DIR.mkdir(parents=True, exist_ok=True)
def make_benign_realworld_pdf():
"""A multi-page PDF that exercises every URL type that USED TO be a false positive.
Contains:
- http:// URLs to kernel.org mirrors
- https:// URLs to github.com paths that mention brands ("microsoft")
- https:// URLs with "support", "verify", "account" in the path
- mailto: links to authors
- tel: phone links
- In-document cross-reference links (GoTo actions)
- A page that mentions wget / exploit / payload in prose
- A .ru URL (country-code TLD that used to fire PHISHING_TLD)
"""
out_path = FIXTURES_DIR / "benign_realworld.pdf"
# Step 1: draw the pages with reportlab.
base_path = FIXTURES_DIR / "_base_realworld.pdf"
c = canvas.Canvas(str(base_path), pagesize=letter)
c.setTitle("Linux from Scratch (Sample Chapter)")
c.setAuthor("Gerard Beekmans")
c.setSubject("Sample technical-document PDF for false-positive regression test")
# Page 1 — TOC-style page with prose containing URLs.
c.drawString(80, 720, "Chapter 1. Introduction")
c.drawString(80, 700, "See the official site at https://www.linuxfromscratch.org/")
c.drawString(80, 680, "Mailing lists: https://lists.linuxfromscratch.org/listinfo/lfs-support")
c.drawString(80, 660, "Source mirrors: http://ftp.osuosl.org/pub/lfs/")
c.drawString(80, 640, "Patches hosted at https://github.com/LFS-project/build-scripts")
c.drawString(80, 620, "Bug reports: mailto:lfs-support@linuxfromscratch.org")
c.drawString(80, 600, "Kernel sources: https://www.kernel.org/pub/linux/kernel/")
c.drawString(80, 580, "Phone: tel:+1-555-123-4567")
c.showPage()
# Page 2 — prose mentioning common security words.
c.drawString(80, 720, "Chapter 2. Building the System")
c.drawString(80, 700, "Run wget to download the package from the mirror.")
c.drawString(80, 680, "The exploit described in CVE-2024-1234 affects older kernels.")
c.drawString(80, 660, "The attacker's payload is delivered via a crafted document.")
c.drawString(80, 640, "Use /bin/sh as the login shell.")
c.drawString(80, 620, "On Windows, use PowerShell to install the module.")
c.drawString(80, 600, "Russian mirror: https://ftp.ru.debian.org/debian/")
c.drawString(80, 580, "Wikipedia: https://en.wikipedia.org/wiki/Microsoft_Windows")
c.drawString(80, 560, "Github org: https://github.com/microsoft/vscode")
c.showPage()
# Page 3 — page reference with GoTo (in-document link).
c.drawString(80, 720, "Chapter 3. Cross-references")
c.drawString(80, 700, "See Chapter 1 for introduction details.")
c.drawString(80, 680, "External resources:")
c.drawString(80, 660, "- https://www.ietf.org/rfc/rfc2616.txt")
c.drawString(80, 640, "- https://www.w3.org/TR/html5/")
c.drawString(80, 620, "- https://docs.python.org/3/library/")
c.drawString(80, 600, "- https://example.com/account/verify")
c.drawString(80, 580, "- https://example.com/login")
c.showPage()
c.save()
# Step 2: post-process with pypdf to inject URI-action annotations
# on the pages, mimicking real hyperlinks in a published book.
reader = pypdf.PdfReader(str(base_path))
writer = pypdf.PdfWriter()
for page in reader.pages:
writer.add_page(page)
# Hyperlinks to inject — one per page. Each entry is (page_idx, x1, y1, x2, y2, uri).
# The URI action is the structure that triggers the PdfUri vector in
# the parser. We deliberately include the URLs that USED TO be
# false positives.
hyperlinks = [
# Page 0 — TOC links.
(0, 80, 695, 400, 710, "https://www.linuxfromscratch.org/"),
(0, 80, 675, 400, 690, "https://lists.linuxfromscratch.org/listinfo/lfs-support"),
(0, 80, 655, 400, 670, "http://ftp.osuosl.org/pub/lfs/"),
(0, 80, 635, 400, 650, "https://github.com/LFS-project/build-scripts"),
(0, 80, 595, 400, 610, "mailto:lfs-support@linuxfromscratch.org"),
(0, 80, 575, 400, 590, "https://www.kernel.org/pub/linux/kernel/"),
(0, 80, 555, 400, 570, "tel:+1-555-123-4567"),
# Page 1 — body links.
(1, 80, 575, 400, 590, "https://ftp.ru.debian.org/debian/"),
(1, 80, 555, 400, 570, "https://en.wikipedia.org/wiki/Microsoft_Windows"),
(1, 80, 535, 400, 550, "https://github.com/microsoft/vscode"),
# Page 2 — external resources.
(2, 80, 655, 400, 670, "https://www.ietf.org/rfc/rfc2616.txt"),
(2, 80, 615, 400, 630, "https://example.com/account/verify"),
(2, 80, 595, 400, 610, "https://example.com/login"),
]
for page_idx, x1, y1, x2, y2, uri in hyperlinks:
page = writer.pages[page_idx]
# Build the link annotation.
uri_action = DictionaryObject({
NameObject("/Type"): NameObject("/Action"),
NameObject("/S"): NameObject("/URI"),
NameObject("/URI"): TextStringObject(uri),
})
uri_action_ref = writer._add_object(uri_action)
annot = DictionaryObject({
NameObject("/Type"): NameObject("/Annot"),
NameObject("/Subtype"): NameObject("/Link"),
NameObject("/Rect"): ArrayObject([
NumberObject(x1), NumberObject(y1),
NumberObject(x2), NumberObject(y2),
]),
NameObject("/A"): uri_action_ref,
NameObject("/Border"): ArrayObject([
NumberObject(0), NumberObject(0), NumberObject(0),
]),
})
annot_ref = writer._add_object(annot)
if "/Annots" not in page:
page[NameObject("/Annots")] = ArrayObject()
page[NameObject("/Annots")].append(annot_ref)
with open(out_path, "wb") as f:
writer.write(f)
base_path.unlink()
return out_path
def main():
path = make_benign_realworld_pdf()
print(f" wrote {path} ({path.stat().st_size} bytes)")
if __name__ == "__main__":
main()

View File

@ -81,12 +81,56 @@ pub struct Config {
/// If `true`, the scanner will emit `Suspicious` findings for any
/// executable vector it cannot confidently classify. If `false`,
/// only confidently-malicious vectors are reported.
///
/// **Note:** no default detector produces `Suspicious`. Every
/// detector either proves Malicious or doesn't fire (Benign). The
/// flag is retained for API compatibility and for future detectors
/// that may produce genuinely indeterminate signals. Defaults to
/// `false`.
pub emit_suspicious: bool,
/// List of URI schemes that are considered safe for hyperlink
/// navigation (e.g. `https`, `mailto`). Anything else is flagged.
/// URI schemes that are considered safe for hyperlink navigation
/// (e.g. `https`, `mailto`). Anything else — `javascript:`,
/// `vbscript:`, `data:text/html`, `file:` — is flagged as Malicious.
///
/// The default list contains the schemes any hyperlink in a normal
/// document would use. HTTP is included because the overwhelming
/// majority of `http://` links in real documents (kernel.org mirrors,
/// list archives, documentation sites that haven't been migrated to
/// HTTPS) are benign; a blanket `http:` ban produced hundreds of
/// false positives on technical PDFs without catching any real
/// threat that wasn't already caught by the executable-scheme check.
pub allowed_uri_schemes: Vec<String>,
/// If `true`, generate a Markdown report alongside the JSON report.
pub emit_markdown_report: bool,
/// If `true`, copy the original file to the cleanse/output directory
/// when a scan produces zero malicious findings. The copied file is
/// named `clean_<timestamp>_<sha_prefix>.<ext>`.
///
/// This makes the scanner a proper pipeline stage: input files
/// flow through the scanner and clean ones end up in the output
/// folder alongside the cleansed derivatives of malicious ones.
/// Operators running batch jobs (e.g. a watchdog directory) can
/// then chain downstream processing on the contents of the
/// cleanse_dir without having to inspect each file's report.
///
/// Defaults to `true`. Set to `false` to suppress the copy when
/// you only want the reports.
pub emit_clean_output: bool,
/// If `true`, **move** (rather than copy) the source file to the
/// clean-output directory when a scan produces zero malicious
/// findings. The source file is removed from its original location
/// after the copy to the output directory succeeds.
///
/// Useful for pipeline/batch processing where the input directory
/// is a queue and you don't want successfully-scanned files
/// clogging it up. Defaults to `false` because move semantics are
/// destructive — operators who want pipeline-style behavior can
/// flip this on.
///
/// Has no effect when `emit_clean_output` is `false` or when the
/// scan found malicious findings (in which case the source is
/// preserved in the quarantine tarball).
pub move_clean_to_output: bool,
/// Path to an external YARA rules file. If set, the scanner loads
/// additional signature rules from this file at startup.
pub external_rules_path: Option<PathBuf>,
@ -105,13 +149,36 @@ impl Default for Config {
epub_entry_scan_cap: DEFAULT_EPUB_ENTRY_SCAN_CAP,
total_archive_scan_cap: DEFAULT_TOTAL_ARCHIVE_SCAN_CAP,
abort_on_threat: false,
emit_suspicious: true,
// emit_suspicious defaults to false: no default detector
// produces `Suspicious`, so the flag's value is immaterial
// for the default scanner.
emit_suspicious: false,
// Allow-list of URI schemes that hyperlinks may use without
// being treated as Malicious(SuspiciousUri). The default set
// covers every hyperlink type that appears in normal
// documents. `http` is included because most technical
// documentation still has plain-HTTP links (kernel.org
// mirrors, listinfo pages, IRC logs) and the
// executable-scheme check (`javascript:`, `vbscript:`,
// `data:text/html`, `file:`) already catches the
// genuinely dangerous schemes.
allowed_uri_schemes: vec![
"https".to_string(),
"http".to_string(),
"mailto".to_string(),
"ftp".to_string(),
"tel".to_string(),
],
emit_markdown_report: true,
// Default ON: clean documents get copied to the output
// folder so the scanner acts as a pipeline stage. The
// operator can chain downstream processing on the
// cleanse_dir contents without inspecting each report.
emit_clean_output: true,
// Default OFF: moving the source file is destructive.
// Operators who want true pipeline-queue semantics (input
// folder drains as files are scanned) flip this on.
move_clean_to_output: false,
external_rules_path: None,
external_cve_db_path: None,
cleanse_mode: CleanseMode::Markdown,
@ -140,6 +207,10 @@ impl Config {
/// - `CORBEL_ABORT_ON_THREAT` (`1`/`true`/`yes` → true)
/// - `CORBEL_EMIT_SUSPICIOUS`
/// - `CORBEL_EMIT_MARKDOWN_REPORT`
/// - `CORBEL_EMIT_CLEAN_OUTPUT` (`1`/`true`/`yes` → true;
/// default `true`)
/// - `CORBEL_MOVE_CLEAN_TO_OUTPUT` (`1`/`true`/`yes` → true;
/// default `false`)
/// - `CORBEL_EXTERNAL_RULES` (path to external YARA rules JSON)
/// - `CORBEL_EXTERNAL_CVE_DB` (path to external CVE DB JSON)
pub fn override_from_env(mut self) -> Self {
@ -158,6 +229,12 @@ impl Config {
if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") {
self.emit_markdown_report = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_EMIT_CLEAN_OUTPUT") {
self.emit_clean_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_MOVE_CLEAN_TO_OUTPUT") {
self.move_clean_to_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") {
if let Ok(cap) = v.parse::<usize>() {
self.total_archive_scan_cap = cap;
@ -201,8 +278,27 @@ mod tests {
assert!(c.quarantine_dir.ends_with(DEFAULT_QUARANTINE_DIR));
assert!(c.cleanse_dir.ends_with(DEFAULT_CLEANSE_DIR));
assert!(!c.abort_on_threat);
assert!(c.emit_suspicious);
// emit_suspicious defaults to false: no default detector
// produces `Suspicious`. Every detector either proves Malicious
// or doesn't fire.
assert!(!c.emit_suspicious);
assert!(c.allowed_uri_schemes.contains(&"https".to_string()));
// The two schemes added by the false-positive fix must be present,
// otherwise every plain-HTTP URL in technical PDFs is flagged
// Malicious, and phone links (`tel:`) which are universally benign
// get the same treatment.
assert!(c.allowed_uri_schemes.contains(&"http".to_string()));
assert!(c.allowed_uri_schemes.contains(&"tel".to_string()));
// Clean-output defaults: emit_clean_output on (pipeline mode),
// move_clean_to_output off (non-destructive by default).
assert!(
c.emit_clean_output,
"emit_clean_output should default to true for pipeline activity"
);
assert!(
!c.move_clean_to_output,
"move_clean_to_output should default to false — move semantics are destructive"
);
}
#[test]

View File

@ -29,7 +29,7 @@ use crate::CorbelError::{self, ThreatDetected};
use crate::{quarantine, scanner, cleanse, CorbelResult};
use super::config::Config;
use super::types::{DocumentFormat, ScanReport};
use super::types::{Document, DocumentFormat, ScanReport};
/// The result of running the pipeline on a single document.
#[derive(Debug, Clone, Serialize, Deserialize)]
@ -45,6 +45,11 @@ pub struct PipelineResult {
/// Path to the generated quarantine tarball, if any.
pub quarantine_path: Option<PathBuf>,
/// Path to the generated cleansed document, if any.
///
/// Set when (a) the scan found malicious findings and the cleansed
/// derivative was produced, OR (b) the scan found zero malicious
/// findings and `Config::emit_clean_output` is true (the original
/// file was copied/moved to the output folder as a "clean" file).
pub cleansed_path: Option<PathBuf>,
/// Path to the JSON forensic report, if any.
pub json_report_path: Option<PathBuf>,
@ -106,33 +111,68 @@ impl Pipeline {
// 3. Abort-on-threat short-circuit.
if self.config.abort_on_threat && scan_report.malicious_count() > 0 {
return Err(ThreatDetected(format!(
"found {} malicious finding(s) in {}",
scan_report.malicious_count(),
document
let source_label = document
.source_path
.as_ref()
.map(|p| p.display().to_string())
.unwrap_or_else(|| format!("<{} buffer>", format)),
.map_or_else(|| format!("<{format} buffer>"), |p| p.display().to_string());
return Err(ThreatDetected(format!(
"found {} malicious finding(s) in {source_label}",
scan_report.malicious_count(),
)));
}
// 4. Quarantine + report (if anything malicious was found).
let quarantine_outcome = if scan_report.malicious_count() > 0 {
Some(quarantine::handle(&document, &scan_report, &self.config)?)
// 4. Quarantine + report.
//
// - If the scan found at least one malicious finding, the full
// quarantine path runs: payloads are carved, a tarball is
// written containing the original file + report + payloads,
// and the JSON + Markdown reports are written as standalone
// files.
// - If the scan found ZERO malicious findings, we still write
// the JSON + Markdown reports to the configured quarantine
// directory — the operator pressed Start, the scan ran, and
// they should see output confirming the document was clean.
// No tarball is written (nothing to quarantine), no payloads
// are carved (no malicious bytes to extract), no cleansed
// file is produced (nothing to cleanse).
let (quarantine_outcome, clean_report_paths) = if scan_report.malicious_count() > 0 {
(Some(quarantine::handle(&document, &scan_report, &self.config)?), None)
} else {
None
let (json_path, md_path) =
quarantine::write_clean_report(&document, &scan_report, &self.config)?;
(None, Some((json_path, md_path)))
};
// 5. Cleanse (if recommended).
// 5. Cleanse / clean-output.
//
// - Malicious scan → produce the cleansed derivative (sanitized
// Markdown or repackaged original format with malicious
// entries stripped).
// - Clean scan → if `emit_clean_output` is set (default: true),
// copy (or move, if `move_clean_to_output`) the original file
// to the cleanse_dir under the name
// `clean_<timestamp>_<sha_prefix>.<ext>`. This makes the
// scanner a proper pipeline stage: input files flow through
// and clean ones end up in the output folder alongside the
// cleansed derivatives of malicious ones.
let cleansed_path = if scan_report.overall_recommendation()
== super::types::Recommendation::QuarantineAndCleanse
{
Some(cleanse::cleanse(&document, &scan_report, &self.config)?)
} else if scan_report.malicious_count() == 0 && self.config.emit_clean_output {
Some(emit_clean_output(&document, &self.config)?)
} else {
None
};
// Unwrap the clean-report paths into the PipelineResult fields.
// `quarantine_outcome` is `Some` only in the malicious case;
// `clean_report_paths` is `Some` only in the clean case.
let (clean_json_path, clean_md_path) = match clean_report_paths {
Some((j, m)) => (Some(j), m),
None => (None, None),
};
Ok(PipelineResult {
source_path,
source_sha256: document.sha256.clone(),
@ -140,14 +180,68 @@ impl Pipeline {
scan_report,
quarantine_path: quarantine_outcome.as_ref().map(|q| q.tarball_path.clone()),
cleansed_path,
json_report_path: quarantine_outcome.as_ref().map(|q| q.json_report_path.clone()),
json_report_path: quarantine_outcome
.as_ref()
.map(|q| q.json_report_path.clone())
.or(clean_json_path),
markdown_report_path: quarantine_outcome
.as_ref()
.and_then(|q| q.markdown_report_path.clone()),
.and_then(|q| q.markdown_report_path.clone())
.or(clean_md_path),
})
}
}
/// Copy (or move, if `Config::move_clean_to_output`) the original file
/// to the cleanse/output directory when a scan produces zero malicious
/// findings.
///
/// The output filename follows the same convention as the cleansed
/// derivative: `clean_<timestamp>_<sha_prefix>.<ext>`. Using the SHA
/// prefix avoids collisions across multiple clean documents and gives
/// each one a stable, traceable name.
///
/// This is the "pipeline stage" behavior: input files flow through
/// the scanner, clean ones end up in the output folder, malicious
/// ones get cleansed derivatives in the same folder. Downstream
/// processing can then operate on the contents of `cleanse_dir`
/// without having to inspect each file's report.
fn emit_clean_output(document: &Document, config: &Config) -> CorbelResult<PathBuf> {
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S");
let sha_prefix = &document.sha256[..8.min(document.sha256.len())];
let ext = match document.format {
DocumentFormat::Pdf => "pdf",
DocumentFormat::Epub => "epub",
DocumentFormat::Markdown => "md",
DocumentFormat::Docx => "docx",
};
let filename = format!("clean_{timestamp}_{sha_prefix}.{ext}");
let output_path = config.cleanse_dir.join(&filename);
// Write the original bytes to the output path. We use the in-memory
// `raw_bytes` rather than re-reading from `source_path` so this
// works even when the pipeline was invoked via `run_on_bytes`
// without a source path on disk.
std::fs::write(&output_path, &document.raw_bytes)?;
// If move-semantics are requested AND we have a source path on
// disk, remove the source file after the copy succeeds. We do
// this only after `std::fs::write` returns Ok, so a failed copy
// never takes the source with it.
if config.move_clean_to_output {
if let Some(src) = document.source_path.as_ref() {
// Best-effort removal — if the unlink fails (e.g.
// permission denied, file locked on Windows), we don't
// want to fail the whole pipeline. The clean copy is
// already in place; the operator can clean up the source
// manually.
let _ = std::fs::remove_file(src);
}
}
Ok(output_path)
}
impl DocumentFormat {
/// Detect a document's format from its file extension.
///

View File

@ -164,7 +164,14 @@ pub enum VectorType {
PdfLaunch,
/// PDF `/URI` action (open a URL — used for phishing).
PdfUri,
/// PDF `/GoToR` / `/GoTo` remote navigation.
/// PDF `/GoTo` action — in-document navigation (jump to a page
/// or named destination inside the same PDF). Benign by default;
/// the scanner still records it so the operator can see in-document
/// navigation activity, but no alert is raised just for being a link.
PdfGoTo,
/// PDF `/GoToR` action — remote navigation (jump to another PDF
/// file). Treated as suspicious because the destination file is
/// outside the currently-scanned document.
PdfGoToR,
/// PDF `/EmbeddedFiles` attachment.
PdfEmbeddedFile,
@ -200,6 +207,7 @@ impl std::fmt::Display for VectorType {
Self::PdfJavaScript => "pdf-javascript",
Self::PdfLaunch => "pdf-launch",
Self::PdfUri => "pdf-uri",
Self::PdfGoTo => "pdf-goto",
Self::PdfGoToR => "pdf-gotor",
Self::PdfEmbeddedFile => "pdf-embedded-file",
Self::PdfWidgetAction => "pdf-widget-action",

View File

@ -15,7 +15,7 @@
//! executables).
//! 3. Optionally extracts, reports, and quarantines any identified threats
//! ([`crate::quarantine`]).
//! 4. Optionally produces a sanitized, cleansed derivative of the original
//! 4. Optionally produces a sanitized, cleansed derivative of the source
//! document containing only validated clean content
//! ([`crate::cleanse`]).
//!
@ -23,6 +23,10 @@
//! which is intentionally stubbed in this MVP.
//!
//! See `MANIFEST.md` in the project root for the full design manifest.
//!
//! # Author
//!
//! **Jeremy Anderson** — [dcos.net](https://dcos.net) — <info@dcos.net>
#![forbid(unsafe_code)]
#![deny(missing_docs)]

View File

@ -53,9 +53,9 @@ fn print_usage() {
USAGE:
corbel-purge scan <path> [--workspace <dir>] [--abort-on-threat] [--quiet] [--preserve-format]
[--rules <path>] [--cve-db <path>]
[--rules <path>] [--cve-db <path>] [--no-clean-output] [--move-clean]
corbel-purge scan-dir <dir> [--recursive] [--workspace <dir>] [--abort-on-threat] [--preserve-format]
[--rules <path>] [--cve-db <path>]
[--rules <path>] [--cve-db <path>] [--no-clean-output] [--move-clean]
corbel-purge study <path> [--workspace <dir>]
corbel-purge --version
@ -76,6 +76,24 @@ OPTIONS:
--cve-db <path> Path to an external CVE signature database
(JSON array). Additional CVE entries are loaded
and matched alongside the built-in CVE table.
--no-clean-output Do NOT copy clean documents to the output folder.
By default, clean documents are copied to
<workspace>/corbel_clean/clean_<ts>_<sha>.<ext>
so the scanner acts as a pipeline stage.
--move-clean MOVE (not copy) the source file to the output
folder when the scan is clean. The source file
is removed from its original location after
the copy succeeds. Useful for pipeline/batch
processing where the input directory is a
queue. Implies clean-output is enabled.
ENVIRONMENT VARIABLES:
CORBEL_QUARANTINE_DIR, CORBEL_CLEANSE_DIR
CORBEL_ABORT_ON_THREAT, CORBEL_EMIT_SUSPICIOUS
CORBEL_EMIT_MARKDOWN_REPORT
CORBEL_EMIT_CLEAN_OUTPUT, CORBEL_MOVE_CLEAN_TO_OUTPUT
CORBEL_TOTAL_ARCHIVE_SCAN_CAP
CORBEL_EXTERNAL_RULES, CORBEL_EXTERNAL_CVE_DB
EXIT CODES:
0 No threats found.
@ -95,6 +113,14 @@ struct ScanArgs {
preserve_format: bool,
external_rules_path: Option<PathBuf>,
external_cve_db_path: Option<PathBuf>,
/// `--no-clean-output` — disable copying clean documents to the
/// output folder. By default, clean documents ARE copied to
/// `cleanse_dir/clean_<ts>_<sha>.<ext>` (pipeline-stage behavior).
no_clean_output: bool,
/// `--move-clean` — move (rather than copy) the source file to the
/// output folder when the scan is clean. Useful for pipeline/batch
/// processing where the input directory is a queue.
move_clean: bool,
}
fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
@ -106,6 +132,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
let mut preserve_format = false;
let mut external_rules_path: Option<PathBuf> = None;
let mut external_cve_db_path: Option<PathBuf> = None;
let mut no_clean_output = false;
let mut move_clean = false;
let mut i = 0;
while i < args.len() {
@ -132,6 +160,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
args.get(i).ok_or("--cve-db requires a value")?,
));
}
"--no-clean-output" => no_clean_output = true,
"--move-clean" => move_clean = true,
"--help" | "-h" => {
print_usage();
std::process::exit(0);
@ -160,6 +190,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
preserve_format,
external_rules_path,
external_cve_db_path,
no_clean_output,
move_clean,
})
}
@ -178,6 +210,8 @@ fn run_scan(args: &[String]) -> ExitCode {
parsed.preserve_format,
parsed.external_rules_path,
parsed.external_cve_db_path,
parsed.no_clean_output,
parsed.move_clean,
);
match pipeline.run(&parsed.path) {
Ok(result) => {
@ -211,6 +245,8 @@ fn run_scan_dir(args: &[String]) -> ExitCode {
parsed.preserve_format,
parsed.external_rules_path,
parsed.external_cve_db_path,
parsed.no_clean_output,
parsed.move_clean,
);
let mut found_malicious = false;
let mut had_error = false;
@ -302,6 +338,8 @@ fn build_pipeline(
preserve_format: bool,
external_rules_path: Option<PathBuf>,
external_cve_db_path: Option<PathBuf>,
no_clean_output: bool,
move_clean: bool,
) -> Pipeline {
let mut config = match workspace {
Some(p) => Config::with_workspace(p),
@ -313,6 +351,14 @@ fn build_pipeline(
if preserve_format {
config.cleanse_mode = corbel_purge::CleanseMode::PreserveFormat;
}
// CLI flags for clean-output behavior. These override the config
// defaults (which are emit_clean_output=true, move_clean_to_output=false).
if no_clean_output {
config.emit_clean_output = false;
}
if move_clean {
config.move_clean_to_output = true;
}
// Apply env overrides last so they win.
apply_env_overrides(&mut config);
@ -347,6 +393,12 @@ fn apply_env_overrides(config: &mut Config) {
if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") {
config.emit_markdown_report = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_EMIT_CLEAN_OUTPUT") {
config.emit_clean_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_MOVE_CLEAN_TO_OUTPUT") {
config.move_clean_to_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") {
if let Ok(cap) = v.parse::<usize>() {
config.total_archive_scan_cap = cap;

View File

@ -280,61 +280,19 @@ fn extract_docx_text(xml: &str, entry_name: &str, text_nodes: &mut Vec<TextNode>
/// Extract external hyperlinks from `word/document.xml`.
///
/// Hyperlinks look like `<w:hyperlink r:id="rId1">text</w:hyperlink>`.
/// The actual URL is in the rels file, but we still emit a vector
/// here so the scanner knows there's an external link reference.
fn extract_docx_external_links(xml: &str, entry_name: &str, vectors: &mut Vec<ExecutableVector>) {
let mut search_from = 0;
while let Some(h_start) = xml[search_from..].find("<w:hyperlink") {
let abs_start = search_from + h_start;
// Find end of the hyperlink opening tag.
let rest = &xml[abs_start..];
let tag_end = match rest.find('>') {
Some(p) => abs_start + p + 1,
None => break,
};
// Find the closing </w:hyperlink>.
let after_tag = &xml[tag_end..];
let h_end = match after_tag.find("</w:hyperlink>") {
Some(p) => tag_end + p,
None => break,
};
// Extract r:id attribute value.
let opening_tag = &xml[abs_start..tag_end];
let rid = extract_attribute(opening_tag, "r:id");
// Extract the visible text of the hyperlink.
let inner = &xml[tag_end..h_end];
let mut visible_text = String::new();
let mut text_search = 0;
while let Some(t_start) = inner[text_search..].find("<w:t") {
let abs_t_start = text_search + t_start;
let after_tag = &inner[abs_t_start..];
let content_start = match after_tag.find('>') {
Some(p) => abs_t_start + p + 1,
None => break,
};
let after_content = &inner[content_start..];
let content_end = match after_content.find("</w:t>") {
Some(p) => content_start + p,
None => break,
};
visible_text.push_str(&inner[content_start..content_end]);
text_search = content_end + 5;
}
vectors.push(ExecutableVector {
location: Location::EpubEntry {
path: entry_name.to_string(),
anchor: rid.clone(),
},
vector_type: VectorType::DocxExternalLink,
raw_payload: visible_text.as_bytes().to_vec(),
decoded_preview: Some(format!("rId={} text={}", rid.unwrap_or_default(), visible_text)),
});
search_from = h_end + 14; // length of "</w:hyperlink>"
}
/// The actual URL is in the rels file (`word/_rels/document.xml.rels`),
/// which is parsed separately by [`extract_rels_external_links`].
///
/// This function emits no vectors. The rels-file extractor is the
/// sole source of DOCX external-link vectors because it carries the
/// actual URL as payload — the visible-text representation here does
/// not. Running URL detectors on visible text breaks scheme
/// extraction (the `"rId=... text=..."` wrapping is not a valid RFC
/// 3986 scheme) and breaks authority extraction (visible text may
/// contain `@` for unrelated reasons). The rels-file path avoids both
/// failure modes.
fn extract_docx_external_links(_xml: &str, _entry_name: &str, _vectors: &mut Vec<ExecutableVector>) {
// Intentionally empty. See doc comment above.
}
/// Extract external relationships from `word/_rels/document.xml.rels`.
@ -342,6 +300,11 @@ fn extract_docx_external_links(xml: &str, entry_name: &str, vectors: &mut Vec<Ex
/// Each `<Relationship>` element has attributes `Id`, `Target`, and
/// `TargetMode`. If `TargetMode="External"`, the relationship points
/// to an external URL — emit a vector with the URL as the payload.
///
/// This is the canonical source of DOCX external-link vectors: one
/// vector per external hyperlink, with the URL as `raw_payload` and
/// `decoded_preview`. The visible-text-extracting function above no
/// longer emits a duplicate vector.
fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
let mut search_from = 0;
while let Some(rel_start) = xml[search_from..].find("<Relationship") {
@ -360,6 +323,11 @@ fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
let id = extract_attribute(opening_tag, "Id").unwrap_or_default();
if target_mode.as_deref() == Some("External") {
// The decoded_preview is the URL itself (NOT
// `"rId={} target={}"`). The URL detector needs a clean
// URL string to run the scheme/host checks against;
// wrapping it in `"rId=... target=..."` broke both the
// scheme extraction and the authority extraction.
vectors.push(ExecutableVector {
location: Location::EpubEntry {
path: "word/_rels/document.xml.rels".to_string(),
@ -367,7 +335,7 @@ fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
},
vector_type: VectorType::DocxExternalLink,
raw_payload: target.as_bytes().to_vec(),
decoded_preview: Some(format!("rId={} target={}", id, target)),
decoded_preview: Some(target.clone()),
});
}

View File

@ -212,20 +212,21 @@ fn inspect_dictionary(
inspect_annots(annots_ref, obj_id, doc, vectors);
}
// /AcroForm with /AA
// /AcroForm — only flag when /AA (Additional Actions) is present.
// /NeedAppearances is a benign rendering hint present in essentially
// every PDF form. Forms with only /NeedAppearances do not execute
// any code.
if let Ok(acroform_ref) = dict.get(b"AcroForm") {
if let Ok((_id, acroform_obj)) = doc.dereference(acroform_ref) {
if let Object::Dictionary(acro_dict) = acroform_obj {
if acro_dict.has(b"AA") || acro_dict.has(b"NeedAppearances") {
if acro_dict.has(b"AA") {
vectors.push(ExecutableVector {
location: Location::PdfObject { id: obj_id.0, gen: 0 },
vector_type: VectorType::PdfAcroForm,
raw_payload: Vec::new(),
decoded_preview: Some(format!(
"<AcroForm with AA={} NeedAppearances={}>",
acro_dict.has(b"AA"),
acro_dict.has(b"NeedAppearances")
)),
decoded_preview: Some(
"<AcroForm with /AA (Additional Actions)>".to_string(),
),
});
}
}
@ -344,13 +345,28 @@ fn inspect_action(action_ref: &Object, obj_id: &ObjectId, doc: &LopdfDocument, v
.into(),
});
}
"GoToR" | "GoTo" => {
// Remote / local navigation — flag but lower priority.
"GoTo" => {
// In-document navigation — jump to a page or named
// destination inside the SAME PDF. This is benign (think
// "click here to jump to section 3"), so the scanner will
// classify it as Benign. We still record it so the operator
// can see in-document navigation activity in the report.
vectors.push(ExecutableVector {
location: loc,
vector_type: VectorType::PdfGoTo,
raw_payload: Vec::new(),
decoded_preview: Some(format!("<GoTo action in obj {loc_id}>")),
});
}
"GoToR" => {
// Remote navigation — jump to another PDF file. Treated as
// Suspicious because the destination file is outside the
// currently-scanned document and we can't inspect it.
vectors.push(ExecutableVector {
location: loc,
vector_type: VectorType::PdfGoToR,
raw_payload: Vec::new(),
decoded_preview: Some(format!("<GoTo/GoToR action in obj {loc_id}>")),
decoded_preview: Some(format!("<GoToR action in obj {loc_id}>")),
});
}
_ => {

View File

@ -13,7 +13,7 @@ pub mod extractor;
pub mod hexdump;
pub mod reporter;
use std::path::PathBuf;
use std::path::{Path, PathBuf};
use serde::{Deserialize, Serialize};
@ -42,99 +42,145 @@ pub fn handle(
scan_report: &ScanReport,
config: &Config,
) -> CorbelResult<QuarantineOutcome> {
// 1. Carve payloads out of the document.
//
// We pass `config` through so that `payload_path` is rooted at the
// caller's quarantine_dir. The no-config `extract_payloads` helper
// would fall back to a default Config whose `quarantine_dir` is the
// relative path `corbel_quarantine/` — fine in production (where
// CWD == workspace) but broken in tests (where the workspace is a
// tempdir) and any other embedded use case.
// 1. Carve payloads out of the document. Pass `config` through so
// payload_path is rooted at the caller's quarantine_dir; the
// no-config helper uses a default Config whose quarantine_dir is
// the relative `corbel_quarantine/` (broken under tempdir workspaces).
let extracted = extractor::extract_payloads_with_config(document, scan_report, config);
// 2. Generate reports.
let markdown_report = config
.emit_markdown_report
.then(|| reporter::build_markdown_report(document, scan_report, &extracted));
let json_report = reporter::build_json_report(document, scan_report, &extracted);
let markdown_report = if config.emit_markdown_report {
Some(reporter::build_markdown_report(document, scan_report, &extracted))
} else {
None
};
// 3. Compose the quarantine tarball name.
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S");
let sha_prefix = &document.sha256[..8.min(document.sha256.len())];
let tarball_name = format!("quarantine_{timestamp}_{sha_prefix}.tar.gz");
let tarball_path = config.quarantine_dir.join(&tarball_name);
// 3. Compose filenames. SHA prefix guarantees uniqueness across documents.
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S").to_string();
let sha_pref = sha_prefix(document);
let paths = ReportPaths::new(
&config.quarantine_dir,
&timestamp,
&sha_pref,
markdown_report.is_some(),
);
let json_report_name = format!("report_{timestamp}_{sha_prefix}.json");
let json_report_path = config.quarantine_dir.join(&json_report_name);
let markdown_report_path = if markdown_report.is_some() {
Some(config.quarantine_dir.join(format!(
"report_{timestamp}_{sha_prefix}.md"
)))
} else {
None
};
// 4. Write the tarball.
// 4. Write the tarball + standalone JSON report.
write_tarball(
&tarball_path,
&paths.tarball_path,
document,
&json_report,
markdown_report.as_deref(),
&extracted,
)?;
std::fs::write(&paths.json_path, serde_json::to_string_pretty(&json_report)?)?;
// 5. Write the JSON report as a standalone file (for easy programmatic access).
std::fs::write(&json_report_path, serde_json::to_string_pretty(&json_report)?)?;
// 6. Write the Markdown report as a standalone file too.
if let Some(md) = &markdown_report {
if let Some(md_path) = &markdown_report_path {
// 5. Write the Markdown report (only when enabled).
if let (Some(md), Some(md_path)) = (&markdown_report, &paths.markdown_path) {
std::fs::write(md_path, md)?;
}
// 6. Write standalone .hex + .info files for each carved payload.
extracted
.iter()
.try_for_each(|payload| write_payload_artifacts(payload, scan_report))?;
Ok(QuarantineOutcome {
tarball_path: paths.tarball_path,
json_report_path: paths.json_path,
markdown_report_path: paths.markdown_path,
extracted_payload_paths: extracted.iter().map(|p| p.payload_path.clone()).collect(),
})
}
// 7. Write standalone payload carving v2 files (.hex + .info).
for payload in &extracted {
// .hex — annotated hex dump
let hex_content = hexdump::hex_dump(&payload.bytes);
let hex_path = payload.payload_path.with_extension("hex");
std::fs::write(&hex_path, hex_content)?;
/// Composed report-file paths under `quarantine_dir`.
struct ReportPaths {
tarball_path: PathBuf,
json_path: PathBuf,
markdown_path: Option<PathBuf>,
}
// .info — JSON metadata
let finding = scan_report
impl ReportPaths {
/// Construct from `quarantine_dir` using the timestamp + sha-prefix
/// naming convention shared by both the malicious and clean paths.
fn new(quarantine_dir: &Path, timestamp: &str, sha_prefix: &str, emit_markdown: bool) -> Self {
let tarball_path = quarantine_dir.join(format!("quarantine_{timestamp}_{sha_prefix}.tar.gz"));
let json_path = quarantine_dir.join(format!("report_{timestamp}_{sha_prefix}.json"));
let markdown_path = emit_markdown
.then(|| quarantine_dir.join(format!("report_{timestamp}_{sha_prefix}.md")));
Self { tarball_path, json_path, markdown_path }
}
}
/// SHA-256 prefix for filename uniqueness (first 8 hex chars).
fn sha_prefix(document: &Document) -> String {
document.sha256[..8.min(document.sha256.len())].to_string()
}
/// Write the .hex (annotated dump) and .info (JSON metadata) files
/// for one carved payload.
fn write_payload_artifacts(
payload: &extractor::ExtractedPayload,
scan_report: &ScanReport,
) -> CorbelResult<()> {
// .hex — annotated hex dump.
let hex_content = hexdump::hex_dump(&payload.bytes);
std::fs::write(payload.payload_path.with_extension("hex"), hex_content)?;
// .info — JSON metadata.
let info = build_payload_info(payload, scan_report);
std::fs::write(
payload.payload_path.with_extension("info"),
serde_json::to_string_pretty(&info)?,
)?;
Ok(())
}
/// Build the JSON metadata object for one carved payload.
fn build_payload_info(
payload: &extractor::ExtractedPayload,
scan_report: &ScanReport,
) -> serde_json::Value {
use crate::core::types::{Recommendation, ThreatClassification};
// Lookup table for classification → JSON string.
let classification_str = scan_report
.findings
.get(payload.source_finding_index);
let (classification_str, recommendation_str, context_notes, vector_type_str, cve_tag_str) =
if let Some(f) = finding {
.get(payload.source_finding_index)
.map(|f| match &f.classification {
ThreatClassification::Benign => "benign".to_string(),
ThreatClassification::EducationalContent => "educational".to_string(),
ThreatClassification::Suspicious => "suspicious".to_string(),
ThreatClassification::Malicious(t) => format!("malicious:{t}"),
})
.unwrap_or_default();
// Lookup table for recommendation → JSON string.
let recommendation_str = scan_report
.findings
.get(payload.source_finding_index)
.map(|f| match f.recommendation {
Recommendation::Allow => "allow".to_string(),
Recommendation::WhitelistAsEducational => "whitelist-as-educational".to_string(),
Recommendation::Quarantine => "quarantine".to_string(),
Recommendation::QuarantineAndCleanse => "quarantine-and-cleanse".to_string(),
})
.unwrap_or_default();
let (context_notes, vector_type_str, cve_tag_str) = scan_report
.findings
.get(payload.source_finding_index)
.map(|f| {
(
match &f.classification {
crate::core::types::ThreatClassification::Benign => "benign".to_string(),
crate::core::types::ThreatClassification::EducationalContent => "educational".to_string(),
crate::core::types::ThreatClassification::Suspicious => "suspicious".to_string(),
crate::core::types::ThreatClassification::Malicious(t) => format!("malicious:{t}"),
},
match f.recommendation {
crate::core::types::Recommendation::Allow => "allow".to_string(),
crate::core::types::Recommendation::WhitelistAsEducational => "whitelist-as-educational".to_string(),
crate::core::types::Recommendation::Quarantine => "quarantine".to_string(),
crate::core::types::Recommendation::QuarantineAndCleanse => "quarantine-and-cleanse".to_string(),
},
f.context_notes.clone(),
f.vector_type.map(|v| v.to_string()),
extract_cve_tag_from_notes(&f.context_notes),
)
} else {
(String::new(), String::new(), String::new(), None, None)
};
})
.unwrap_or((String::new(), None, None));
let file_sig =
crate::scanner::signatures::match_file_signature(&payload.bytes)
.map(|s| s.to_string());
let file_sig = crate::scanner::signatures::match_file_signature(&payload.bytes).map(|s| s.to_string());
let info = hexdump::build_payload_info(
hexdump::build_payload_info(
&payload.filename,
payload.source_finding_index,
vector_type_str.as_deref(),
@ -146,20 +192,54 @@ pub fn handle(
payload.bytes.len(),
file_sig.as_deref(),
cve_tag_str.as_deref(),
);
let info_path = payload.payload_path.with_extension("info");
std::fs::write(&info_path, serde_json::to_string_pretty(&info)?)?;
)
}
Ok(QuarantineOutcome {
tarball_path,
json_report_path,
markdown_report_path,
extracted_payload_paths: extracted
.iter()
.map(|p| p.payload_path.clone())
.collect(),
})
/// Write a "clean bill of health" report for a document that produced
/// no malicious findings.
///
/// Called by the pipeline when `scan_report.malicious_count() == 0`.
/// Writes the JSON and Markdown reports to the configured quarantine
/// directory using the same `report_<timestamp>_<sha_prefix>.{json,md}`
/// naming convention as the malicious path, but with an empty
/// findings array. No tarball is written (nothing to quarantine), no
/// payloads are carved (no malicious bytes to extract), no cleansed
/// file is produced (nothing to cleanse).
///
/// The operator receives a report confirming the scan ran, what was
/// scanned, when it was scanned, and that the document was clean —
/// the audit trail required for pipeline activity.
pub fn write_clean_report(
document: &Document,
scan_report: &ScanReport,
config: &Config,
) -> CorbelResult<(PathBuf, Option<PathBuf>)> {
// Empty extracted-payloads list — clean documents have no payloads
// to extract, but the report builder takes the parameter.
let extracted: Vec<extractor::ExtractedPayload> = Vec::new();
let markdown_report = config
.emit_markdown_report
.then(|| reporter::build_markdown_report(document, scan_report, &extracted));
let json_report = reporter::build_json_report(document, scan_report, &extracted);
// Reuse the shared path-composition helper for naming consistency
// with the malicious path.
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S").to_string();
let sha_pref = sha_prefix(document);
let paths = ReportPaths::new(
&config.quarantine_dir,
&timestamp,
&sha_pref,
markdown_report.is_some(),
);
// Write the JSON report unconditionally; write Markdown only when enabled.
std::fs::write(&paths.json_path, serde_json::to_string_pretty(&json_report)?)?;
if let (Some(md), Some(md_path)) = (&markdown_report, &paths.markdown_path) {
std::fs::write(md_path, md)?;
}
Ok((paths.json_path, paths.markdown_path))
}
/// Extract a CVE tag (e.g. `[CVE-2017-11882: Equation Editor RCE]`)

View File

@ -1,15 +1,26 @@
//! Context filter: NLP / lexical checks for distinguishing security
//! literature from active malicious content.
//! # Context filter: structural-signature checks for text nodes.
//!
//! When the scanner encounters a suspicious signature inside a *static
//! text node* (paragraph, heading, code block), it asks the context
//! filter whether the surrounding context looks like:
//! The text-node scanner checks for exactly two things:
//!
//! - **Educational content** (CVE writeups, exploit code samples in
//! defensive blog posts, textbook material) → whitelisted.
//! - **Weaponized content** (obfuscated shellcode, packed executables
//! in non-code contexts, embedded action triggers) → flagged.
//! - **Indeterminate** → emitted as `Suspicious` if configured.
//! 1. **Structural signatures** indicating the text contains an
//! executable hook embedded in the text content itself — e.g. a
//! PDF `/JavaScript` operator appearing in a content stream, or an
//! HTML `<script>` tag in EPUB XHTML. These are verifiable
//! structural tokens; presence is the threat.
//!
//! 2. **Weaponization indicators** — long runs of hex-encoded bytes
//! (`\xNN` × 16+), long base64 blobs (64+ consecutive base64
//! chars), and multiple concatenated shell commands. These are
//! byte-level patterns with no legitimate use in document text
//! outside of code blocks / blockquotes.
//!
//! ## Design invariant
//!
//! Words and function names that appear in legitimate technical
//! literature (`wget`, `exploit`, `payload`, `eval(`, `exec(`,
//! `powershell`, `/bin/sh`, etc.) are not treated as signatures.
//! The detector only fires on structural tokens whose presence in
//! document text outside of a code block is itself the threat.
use crate::core::config::Config;
use crate::core::types::{
@ -17,19 +28,18 @@ use crate::core::types::{
};
/// Evaluate a single text node. Returns `Some(Finding)` if the node
/// contains a suspicious or malicious signature that survived the
/// context filter.
/// contains a structural signature or weaponization indicator.
pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
// Step 0: Check for weaponization indicators first — these are
// malicious regardless of whether a "signature" is present.
// Pure shellcode blobs, for example, contain no recognizable
// keyword but are still dangerous.
// Step 1: Check for weaponization indicators — long hex runs,
// long base64 blobs, multiple shell commands in non-code context.
// These are byte-level patterns that are verifiably hostile when
// they appear outside a code block / blockquote.
if has_weaponization_indicators(&node.content) {
let classification = if node.context == TextContext::CodeBlock
|| node.context == TextContext::CodeSpan
|| node.context == TextContext::BlockQuote
{
// Even weaponized-looking content inside a code block /
// Weaponized-looking content inside a code block or
// blockquote is treated as educational — it's almost
// certainly a research writeup illustrating an attack.
ThreatClassification::EducationalContent
@ -57,27 +67,28 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
});
}
// Step 1: Does the node contain any suspicious signatures?
// Step 2: Check for structural signatures — PDF operators, HTML
// tags, shellcode tokens. These indicate the text itself contains
// an executable hook, which is hostile outside of code context.
let signatures = find_signatures(&node.content);
if signatures.is_empty() {
return None;
}
// Step 2: What's the surrounding context?
// Step 3: Decide based on context.
let is_educational = looks_educational(node);
// Step 3: Decision matrix.
let classification = if is_educational {
ThreatClassification::EducationalContent
} else if node.context == TextContext::ExecutableHook {
// A suspicious signature inside an executable hook is malicious.
// A structural signature inside an executable hook is
// definitively malicious — it's the actual payload.
ThreatClassification::Malicious(MaliciousType::ActiveJavaScriptInjection)
} else if has_weaponization_indicators(&node.content) {
ThreatClassification::Malicious(MaliciousType::ObfuscatedShellcode)
} else if config.emit_suspicious {
// Retained for API compatibility. No default detector produces
// `Suspicious`; `emit_suspicious` defaults to `false`. Future
// detectors with genuinely indeterminate signals may use this tier.
ThreatClassification::Suspicious
} else {
// Suspicious findings suppressed — drop it.
return None;
};
@ -90,7 +101,7 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
let preview: String = node.content.chars().take(config.max_payload_preview_len).collect();
let notes = format!(
"found {} suspicious signature(s) [{}] in {} context{}",
"found {} structural signature(s) [{}] in {} context{}",
signatures.len(),
signatures.join(", "),
context_name(node.context),
@ -107,42 +118,37 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
})
}
/// A signature is a string that *could* indicate malicious content
/// but is also commonly found in security literature.
const SUSPICIOUS_SIGNATURES: &[&str] = &[
/// Structural signatures that indicate the text contains an executable
/// hook. These are tokens whose presence in document text (outside of
/// a code block) is a strong signal of embedded active content.
///
/// **Matched as case-sensitive substrings** for the PDF operators
/// (which are case-sensitive in the PDF spec) and case-insensitive
/// substrings for HTML tags. Word-boundary matching is not necessary
/// because these tokens are sufficiently specific — `/JavaScript`
/// does not appear in prose, only in PDF dictionaries.
const STRUCTURAL_SIGNATURES: &[&str] = &[
// PDF action / object operators. These only appear in PDF
// content streams or PDF dictionary dumps — never in prose.
"/JavaScript",
"/JS",
"/Launch",
"/EmbeddedFile",
"eval(",
"Function(",
"document.write",
"innerHTML",
// HTML / XHTML executable tags. Only appear in markup.
"<script",
"<iframe",
// Shellcode token. The literal word "shellcode" is sometimes
// used in prose, but in combination with the weaponization
// detector above this is reserved for cases where the text
// contains the actual word in an executable context.
"shellcode",
"exploit",
"payload",
"calc.exe",
"/bin/sh",
"powershell",
"cmd.exe",
"wget",
"curl http",
"rm -rf",
"Base64.decode",
"atob(",
"exec(",
];
/// Find all suspicious signatures present in `text`. Returns the
/// Find all structural signatures present in `text`. Returns the
/// list of signatures found (deduplicated, in source order).
fn find_signatures(text: &str) -> Vec<&'static str> {
// Case-insensitive matching for some signatures, exact for others.
// For simplicity, we do case-sensitive matching first and let the
// context filter handle false positives.
let lower = text.to_ascii_lowercase();
SUSPICIOUS_SIGNATURES
STRUCTURAL_SIGNATURES
.iter()
.copied()
.filter(|sig| {
@ -214,13 +220,16 @@ fn has_weaponization_indicators(text: &str) -> bool {
}
// Long base64 blob: 64+ consecutive base64 chars anywhere in text.
// (Not just whole lines — attackers like to embed base64 inline.)
let b64_re = regex::Regex::new(r"[A-Za-z0-9+/=]{64,}").unwrap();
if b64_re.is_match(text) {
return true;
}
// Multiple shell commands in a single non-code node.
// Multiple shell commands in a single non-code node. The threshold
// of 2 different commands is high enough that legitimate prose
// mentioning one shell tool ("run `wget` to download the package")
// does not fire — it requires two different shell-command tokens
// in the same node.
let shell_indicators = ["rm -rf", "wget ", "curl ", "nc -", "/bin/sh", "powershell "];
let count = shell_indicators.iter().filter(|s| text.contains(*s)).count();
if count >= 2 {
@ -265,12 +274,31 @@ mod tests {
#[test]
fn cve_writeup_in_code_block_is_educational() {
// A code block that contains a structural signature (`<script`)
// alongside a CVE marker is treated as educational — it's a
// research writeup illustrating an attack, not an active payload.
let node = make_node(
TextContext::CodeBlock,
"<script>alert(1)</script> // PoC for CVE-2024-1234",
);
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::EducationalContent);
}
#[test]
fn cve_writeup_prose_without_signature_is_no_finding() {
// Text that mentions `eval()` in prose but does not contain a
// structural signature (`/JavaScript`, `<script`, etc.) is not
// a finding. Words and function names in prose are not
// signatures; only structural tokens fire.
let node = make_node(
TextContext::CodeBlock,
"eval('alert(1)') // PoC for CVE-2024-1234",
);
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::EducationalContent);
assert!(
evaluate(&node, &Config::default()).is_none(),
"prose mentioning eval() without a structural signature must NOT be flagged"
);
}
#[test]
@ -297,38 +325,11 @@ mod tests {
));
}
#[test]
fn suspicious_in_paragraph_with_signature() {
let node = make_node(TextContext::Paragraph, "Run eval('alert(1)') now");
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::Suspicious);
}
#[test]
fn suspicious_can_be_suppressed() {
let node = make_node(TextContext::Paragraph, "Run eval('alert(1)') now");
let mut config = Config::default();
config.emit_suspicious = false;
assert!(evaluate(&node, &config).is_none());
}
#[test]
fn academic_text_with_signature_is_educational() {
let node = make_node(
TextContext::Paragraph,
"In this paper we describe the eval() vulnerability and its remediation.",
);
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::EducationalContent);
}
#[test]
fn long_base64_blob_is_weaponized() {
let b64 = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/ABCDEFGH";
let node = make_node(TextContext::Paragraph, b64);
let finding = evaluate(&node, &Config::default());
// With the step-0 weaponization check, pure base64 (no signature)
// is now flagged as Malicious (ObfuscatedShellcode).
let finding = finding.expect("pure base64 blob should be flagged as weaponized");
assert!(matches!(
finding.classification,
@ -346,4 +347,81 @@ mod tests {
ThreatClassification::Malicious(_)
));
}
// ─── Anti-false-positive tests ─────────────────────────────
//
// These are the texts that USED TO fire the substring-signature
// detector and that appear in legitimate technical documents.
// All of them must now produce no finding.
#[test]
fn wget_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"Run wget to download the package from the mirror.",
);
assert!(
evaluate(&node, &Config::default()).is_none(),
"prose mentioning wget must NOT be flagged"
);
}
#[test]
fn exploit_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The exploit described in this chapter targets a buffer overflow.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn payload_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The attacker's payload is delivered via a crafted document.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn single_shell_command_in_prose_is_not_a_finding() {
// The weaponization detector requires 2+ different shell
// commands. A single mention of `wget` in prose is not enough.
let node = make_node(
TextContext::Paragraph,
"Use `wget http://example.com/file.tar.gz` to fetch the archive.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn powershell_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"On Windows, the equivalent command uses PowerShell to install the module.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn bin_sh_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The shebang line `#!/bin/sh` indicates the script runs under the Bourne shell.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn eval_mentioned_in_prose_is_not_a_finding() {
// `eval(` is a JavaScript function name; appearing in prose is
// not a threat. Only structural tokens (`/JavaScript`, `<script`,
// etc.) fire the text-node scanner.
let node = make_node(
TextContext::Paragraph,
"The eval() function in JavaScript executes a string as code.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
}

File diff suppressed because it is too large Load Diff

View File

@ -17,43 +17,45 @@ use crate::core::types::{Document, ScanReport};
/// This is the top-level entrypoint called by [`crate::core::pipeline::Pipeline`].
#[must_use]
pub fn scan(document: &Document, config: &Config) -> ScanReport {
let mut findings = Vec::new();
// Walk executable vectors first (untrusted-by-default), then text
// nodes (context-filtered). Findings are tagged with a CVE entry
// when the payload matches a known exploit signature.
let vector_findings = document.executable_vectors.iter().filter_map(|vector| {
heuristics::inspect_vector(vector, config).map(|mut finding| {
tag_with_cve_if_known(vector, &mut finding);
finding
})
});
// 1. Walk executable vectors — these are untrusted-by-default.
for vector in &document.executable_vectors {
if let Some(mut finding) = heuristics::inspect_vector(vector, config) {
// After classification, try to tag the finding with a
// known CVE if the payload matches a known exploit signature.
if let Some(cve) = cve_tags::match_cve(vector, &finding) {
finding.context_notes = format!(
"{} [{}: {}]",
finding.context_notes, cve.cve_id, cve.name
);
}
findings.push(finding);
}
}
let text_findings = document
.text_nodes
.iter()
.filter_map(|node| context_filter::evaluate(node, config));
// 2. Walk text nodes — these go through the context filter to
// distinguish educational content from active threats.
for node in &document.text_nodes {
if let Some(finding) = context_filter::evaluate(node, config) {
findings.push(finding);
}
}
let scanned_at = chrono::Utc::now().to_rfc3339();
let findings: Vec<_> = vector_findings.chain(text_findings).collect();
ScanReport {
source_sha256: document.sha256.clone(),
format: document.format,
scanned_at,
scanned_at: chrono::Utc::now().to_rfc3339(),
findings,
text_nodes_scanned: document.text_nodes.len(),
vectors_scanned: document.executable_vectors.len(),
}
}
/// Tag a finding with a known CVE entry when the payload matches a
/// known exploit signature. Mutates `finding.context_notes` in place.
fn tag_with_cve_if_known(
vector: &crate::core::types::ExecutableVector,
finding: &mut crate::core::types::Finding,
) {
let Some(cve) = cve_tags::match_cve(vector, finding) else {
return;
};
finding.context_notes = format!("{} [{}: {}]", finding.context_notes, cve.cve_id, cve.name);
}
/// Re-export for callers that want to inspect individual findings.
pub use heuristics::inspect_vector;
pub use context_filter::evaluate;

View File

@ -1,32 +1,30 @@
//! Threat signature tables and pattern matchers.
//! # Threat signature tables and pattern matchers.
//!
//! This module centralizes the static lookup tables used by the
//! heuristics engine. Keeping them in one place makes them easy to
//! audit, extend, and eventually wire up to an external threat-intel
//! feed (e.g. a YARA rules file or a STIX/TAXII subscription).
//! heuristics engine. It contains two kinds of tables:
//!
//! ## What lives here
//! - **Category 1 — Verifiable executable intent.** Magic-byte
//! signatures for executable / high-risk file formats
//! ([`KNOWN_FILE_SIGNATURES`]) and known shellcode prologue
//! sequences ([`SHELLCODE_PATTERNS`]). These are deterministic
//! byte-level matches; presence of the structure is the threat.
//! - **Category 2 — Verifiable impersonation.** A small table of
//! known homograph host strings ([`HOMOGRAPH_HOSTS`]) — exact-match
//! only, never substring. A URL whose authority is byte-equal to one
//! of these strings after case-folding is flagged.
//!
//! - [`PHISHING_TLDS`] — TLDs statistically overrepresented in
//! phishing URLs. Sourced from public phishing reports
//! (Spamhaus, PhishTank yearly summaries).
//! - [`SUSPICIOUS_URL_KEYWORDS`] — path/host keywords that strongly
//! indicate credential harvesting or fake login pages.
//! - [`URL_SHORTENER_DOMAINS`] — shortener domains. Not malicious
//! per se, but a common obfuscation layer for phishing links.
//! - [`KNOWN_FILE_SIGNATURES`] — magic-byte signatures for executable
//! and high-risk file formats (PE, ELF, Mach-O, OLE2, RTF, etc.).
//! - [`SHELLCODE_PATTERNS`] — known shellcode prologue byte sequences
//! (NOP sleds, syscall stubs, common encoders).
//! - [`COMMON_PHISHING_BRANDS`] — brand names frequently spoofed in
//! phishing URLs (microsoft, paypal, appleid, …).
//! The scanner does not consult substring keyword tables, TLD lists,
//! URL-shortener lists as threat signals, or brand-substring tables.
//! Those detectors are statistical guesses about the world, not
//! verifiable properties of the document, and are not part of the
//! detection model. [`URL_SHORTENER_DOMAINS`] is retained for future
//! reputation-feed work but is not consulted by any match function.
//!
//! ## External threat-intel feeds
//!
//! Additional rules can be loaded at runtime via
//! [`load_external_rules`]. Loaded rules are stored in a
//! process-wide static and checked by every match function
//! alongside the built-in tables.
//! Additional rules can be loaded at runtime via [`load_external_rules`].
//! Loaded rules are stored in a process-wide static and consulted by
//! the match functions alongside the built-in tables.
use std::path::Path;
use std::sync::OnceLock;
@ -35,42 +33,13 @@ use serde::Deserialize;
use crate::CorbelResult;
/// TLDs statistically overrepresented in phishing URLs.
/// Common URL-shortener domains.
///
/// Source: synthesized from public yearly phishing reports
/// (Spamhaus, PhishTank, Interisle). This list is intentionally
/// conservative — inclusion requires the TLD to appear in multiple
/// reports as a top-10 phishing TLD.
pub const PHISHING_TLDS: &[&str] = &[
// High-risk TLDs (cheap registration, low verification)
".zip", ".mov", ".xyz", ".top", ".click", ".link", ".rest", ".cyou",
".sbs", ".online", ".live", ".buzz", ".surf", ".monster", ".fit",
".loan", ".win", ".download", ".stream", ".review", ".men",
".work", ".racing", ".party", ".trade", ".science", ".kim",
".cricket", ".gq", ".cf", ".tk", ".ml", ".ga",
// Country-code TLDs frequently abused for phishing
".ru", ".cn", ".su", ".country", ".kim",
// Newer TLDs that have been flagged
".quest", ".bond", ".ha", ".cyou", ".quest", ".beauty",
];
/// URL path / host keywords that strongly suggest credential harvesting
/// or fake login pages. Matched case-insensitively as substrings.
pub const SUSPICIOUS_URL_KEYWORDS: &[&str] = &[
"login", "signin", "sign-in", "log-in", "verify", "verification",
"account", "update", "confirm", "secure", "security", "wallet",
"unlock", "recover", "reactivate", "validate", "activate",
"webscr", "cmd=", "_session", "authorization", "authenticate",
"reset", "password", "credential", "billing", "invoice",
"support", "suspended", "limited", "alert", "warning",
"urgent", "important-notice", "tax", "refund", "irs",
"postbank", "amzn", "appleid", "icloud", "office365",
];
/// Common URL-shortener domains. Shortened URLs are not malicious
/// per se, but they hide the real destination — we flag them as
/// `Suspicious` so the operator can preview the destination before
/// clicking.
/// **Not used as a threat signal by the scanner.** Retained for
/// future reputation-feed integration — if a shortener URL resolves
/// to a Category 1 or Category 2 hit, the underlying detector catches
/// it. Flagging `bit.ly` itself as a threat produced false positives
/// on every document that used a shortener for a legitimate link.
pub const URL_SHORTENER_DOMAINS: &[&str] = &[
"bit.ly", "t.co", "tinyurl.com", "goo.gl", "ow.ly", "is.gd",
"buff.ly", "rebrand.ly", "cutt.ly", "shorturl.at", "tiny.cc",
@ -79,54 +48,59 @@ pub const URL_SHORTENER_DOMAINS: &[&str] = &[
"surl.li", "kutt.it", "urlzs.com", "shrtco.de",
];
/// Brand names frequently spoofed in phishing URLs. Used to detect
/// homograph attacks (e.g. `micros0ft.com`, `paypa1.com`).
/// Homograph host strings — known-bad authority strings that mimic a
/// legitimate brand by substituting visually similar characters.
///
/// This list intentionally includes BOTH canonical spellings ("microsoft")
/// AND known homograph variants ("micros0ft" with zero instead of 'o').
/// The matcher uses a canonical-domain check to suppress benign
/// matches: when a brand is mentioned, we look for the canonical
/// spelling followed by a TLD; if found, we don't flag.
pub const COMMON_PHISHING_BRANDS: &[&str] = &[
// Microsoft family
"microsoft", "micros0ft", "micros0fte", "micr0soft",
"msn", "windows", "wind0ws", "office", "0ffice", "outlook",
"outl00k", "outl0ok", "live", "1ive",
/// **Matching rule:** a URL's authority (host[:port]) is matched
/// byte-equal against these strings after ASCII case-folding. There is
/// no substring match, no path/query involvement, no canonical-brand
/// suppression — the URL's host is either exactly one of these strings
/// or it is not.
///
/// Examples:
/// - `https://micros0ft.com/login` → host `micros0ft.com` is in the
/// list → flagged.
/// - `https://github.com/microsoft/vscode` → host `github.com` is not
/// in the list → not flagged (the "microsoft" in the path is
/// irrelevant).
/// - `https://microsoft.com/windows` → host `microsoft.com` is not in
/// the list → not flagged (legitimate).
/// - `https://MICROS0FT.COM/` → ASCII-case-folded to `micros0ft.com`
/// → flagged.
pub const HOMOGRAPH_HOSTS: &[&str] = &[
// Microsoft family — 'o' → '0', 'i' → '1', etc.
"micros0ft.com", "micr0soft.com", "micros0fte.com",
"msn0.com", "wind0ws.com", "0ffice.com",
"outl00k.com", "outl0ok.com", "1ive.com",
// PayPal
"paypal", "paypa1", "paypaI", "paypa|",
"paypa1.com", "paypaI.com", "paypa|.com",
// Apple
"apple", "app1e", "appie", "icloud", "ic1oud", "appleid",
"app1eid",
"app1e.com", "appie.com", "ic1oud.com", "app1eid.com",
// Google
"google", "g00gle", "goog1e", "gmail", "gmai",
"g00gle.com", "goog1e.com", "gmai.com",
// Amazon
"amazon", "amzn", "amaz0n", "a-m-a-z-o-n",
"amaz0n.com",
// Social
"facebook", "faceb00k", "facebo0k", "instagram", "instagrarn",
"twitter", "tw1tter", "twtter", "linkedin", "1inkedin",
"faceb00k.com", "facebo0k.com", "instagrarn.com",
"tw1tter.com", "twtter.com", "1inkedin.com",
// Streaming
"netflix", "netf1ix", "spotify", "spot1fy",
"netf1ix.com", "spot1fy.com",
// Storage / SaaS
"dropbox", "dr0pbox", "adobe", "ad0be",
"dr0pbox.com", "ad0be.com",
// Banking
"bankofamerica", "bofa", "b0fa", "wellsfargo", "wellsfarg0",
"chase", "citibank", "citi", "hsbc", "barclays",
"santander", "unicredit",
"b0fa.com", "wellsfarg0.com",
// Crypto
"binance", "binanc3", "coinbase", "c0inbase", "metamask",
"metam4sk", "ledger", "trezor",
// Shipping
"dhl", "fedex", "f3dex", "ups", "usps", "royalmail",
"binanc3.com", "c0inbase.com", "metam4sk.com",
// Gaming
"steamcommunity", "steampowered", "epicgames", "playstation",
"nintendo", "xbox",
"steamp0wered.com", "steamc0mmunity.com",
];
/// Magic-byte signatures for executable and high-risk file formats.
///
/// Each entry is (offset, magic_bytes, name). When a payload's bytes
/// at `offset` match `magic_bytes`, the payload is considered
/// executable / high-risk.
/// executable / high-risk. This is the Category 1 detector —
/// presence of the structure is the threat.
pub const KNOWN_FILE_SIGNATURES: &[(usize, &[u8], &str)] = &[
// Windows PE
(0, b"MZ", "pe"),
@ -148,8 +122,6 @@ pub const KNOWN_FILE_SIGNATURES: &[(usize, &[u8], &str)] = &[
(0, b"{\\rtf", "rtf"),
// Java class file
(0, b"\xCA\xFE\xBA\xBE", "java-class"),
// Java JAR (zip, but flag if inside PDF)
// (zip is too generic — we don't flag it without other signals)
// Python bytecode
(0, b"\x42\x0d\x0d\x0a", "python-bytecode"),
// SWF (Flash — historically a huge attack surface)
@ -202,14 +174,18 @@ pub const SHELLCODE_PATTERNS: &[&[u8]] = &[
/// A single external rule deserialized from a JSON threat-intel feed.
#[derive(Debug, Clone, Deserialize)]
pub struct ExternalRule {
/// Human-readable rule name (e.g. `"custom-phishing-tlds"`).
/// Human-readable rule name (e.g. `"custom-shellcode"`).
pub name: String,
/// Rule type: one of `"tld-list"`, `"keyword-list"`,
/// `"brand-list"`, `"signature-list"`, `"shellcode-list"`.
/// Rule type: one of `"homograph-host-list"`,
/// `"signature-list"`, `"shellcode-list"`.
///
/// The previous `"tld-list"`, `"keyword-list"`, and `"brand-list"`
/// types were removed when those tables were removed from the
/// scanner. External feeds of those types are no longer loaded.
#[serde(rename = "type")]
pub rule_type: String,
/// Rule values. The shape depends on `rule_type`:
/// - `"tld-list"` / `"keyword-list"` / `"brand-list"` → array of strings
/// - `"homograph-host-list"` → array of strings (full hostnames)
/// - `"signature-list"` → array of `{"offset": usize, "bytes": [u8], "name": str}`
/// - `"shellcode-list"` → array of hex strings or `[u8]` arrays
pub values: serde_json::Value,
@ -226,12 +202,9 @@ struct SignatureEntry {
/// Processed external rules, organized by type for fast matching.
#[derive(Debug, Clone, Default)]
pub(crate) struct ExternalRulesData {
/// Additional TLD strings.
pub(crate) tlds: Vec<String>,
/// Additional suspicious URL keywords.
pub(crate) keywords: Vec<String>,
/// Additional brand strings (may include homograph variants).
pub(crate) brands: Vec<String>,
/// Additional homograph host strings (exact-match, byte-equal
/// after ASCII case-folding).
pub(crate) homograph_hosts: Vec<String>,
/// Additional file-signature entries: (offset, magic-bytes, name).
pub(crate) signatures: Vec<(usize, Vec<u8>, String)>,
/// Additional shellcode byte patterns.
@ -249,8 +222,9 @@ static EXTERNAL_RULES_STORAGE: OnceLock<ExternalRulesData> = OnceLock::new();
/// The file must contain a JSON array of [`ExternalRule`] objects.
/// Rules are sorted into type-specific buckets and stored in a
/// process-wide static ([`EXTERNAL_RULES_STORAGE`]). Subsequent calls
/// to the match functions (`match_file_signature`, `has_phishing_tld`,
/// etc.) will check both the built-in tables and the external rules.
/// to the match functions (`match_file_signature`,
/// `match_homograph_host`, etc.) will check both the built-in tables
/// and the external rules.
///
/// # Errors
///
@ -263,70 +237,43 @@ pub fn load_external_rules(path: &Path) -> CorbelResult<Vec<ExternalRule>> {
let mut storage = ExternalRulesData::default();
for rule in &rules {
// Each rule contributes zero or more entries to one of the
// storage buckets. Step-down: unknown rule types contribute
// nothing and exit the iteration silently.
match rule.rule_type.as_str() {
"tld-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(s) = v.as_str() {
storage.tlds.push(s.to_string());
}
}
}
}
"keyword-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(s) = v.as_str() {
storage.keywords.push(s.to_string());
}
}
}
}
"brand-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(s) = v.as_str() {
storage.brands.push(s.to_string());
}
}
}
"homograph-host-list" => {
storage.homograph_hosts.extend(
rule.values
.as_array()
.into_iter()
.flatten()
.filter_map(|v| v.as_str())
.map(|s| s.to_ascii_lowercase()),
);
}
"signature-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Ok(sig) = serde_json::from_value::<SignatureEntry>(v.clone()) {
storage
.signatures
.push((sig.offset, sig.bytes, sig.name));
}
}
}
storage.signatures.extend(
rule.values
.as_array()
.into_iter()
.flatten()
.filter_map(|v| serde_json::from_value::<SignatureEntry>(v.clone()).ok())
.map(|sig| (sig.offset, sig.bytes, sig.name)),
);
}
"shellcode-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(hex_str) = v.as_str() {
// Hex-encoded string: "fc4883e4..."
if let Ok(bytes) = hex::decode(hex_str) {
if !bytes.is_empty() {
storage.shellcode.push(bytes);
}
}
} else if let Some(byte_arr) = v.as_array() {
// Raw byte array: [0xfc, 0x48, ...]
let bytes: Vec<u8> = byte_arr
.iter()
.filter_map(|b| b.as_u64().map(|n| n as u8))
.collect();
if !bytes.is_empty() {
storage.shellcode.push(bytes);
}
}
}
}
storage.shellcode.extend(
rule.values
.as_array()
.into_iter()
.flatten()
.filter_map(|v| decode_shellcode_entry(v)),
);
}
_ => {
// Unknown rule type — silently skip.
// Unknown rule type — silently skip. Includes the
// legacy "tld-list", "keyword-list", "brand-list"
// types that the scanner no longer consults.
}
}
}
@ -335,6 +282,25 @@ pub fn load_external_rules(path: &Path) -> CorbelResult<Vec<ExternalRule>> {
Ok(rules)
}
/// Decode a single shellcode-list entry as raw bytes.
///
/// Accepts either a hex-encoded string (`"fc4883e4..."`) or a JSON
/// array of byte values (`[252, 72, ...]`). Returns `None` for empty
/// results or unparseable values.
fn decode_shellcode_entry(v: &serde_json::Value) -> Option<Vec<u8>> {
if let Some(hex_str) = v.as_str() {
hex::decode(hex_str).ok().filter(|b| !b.is_empty())
} else if let Some(byte_arr) = v.as_array() {
let bytes: Vec<u8> = byte_arr
.iter()
.filter_map(|b| b.as_u64().map(|n| n as u8))
.collect();
(!bytes.is_empty()).then_some(bytes)
} else {
None
}
}
/// Return a reference to the externally loaded rules storage, if any
/// has been loaded via [`load_external_rules`].
///
@ -356,23 +322,19 @@ pub(crate) fn get_external_rules_storage() -> Option<&'static ExternalRulesData>
/// [`load_external_rules`].
#[must_use]
pub fn match_file_signature(bytes: &[u8]) -> Option<&'static str> {
for (offset, magic, name) in KNOWN_FILE_SIGNATURES {
if bytes.len() >= *offset + magic.len() {
if &bytes[*offset..*offset + magic.len()] == *magic {
return Some(name);
}
}
}
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for (offset, magic, name) in &ext.signatures {
if bytes.len() >= *offset + magic.len() {
if &bytes[*offset..*offset + magic.len()] == magic.as_slice() {
return Some(name);
}
}
}
}
None
// Built-in signatures — first match wins.
let builtin = KNOWN_FILE_SIGNATURES.iter().find_map(|(offset, magic, name)| {
bytes_matches_at(bytes, *offset, magic).then_some(*name)
});
// External signatures — only consulted when no built-in matched.
builtin.or_else(|| {
EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
ext.signatures.iter().find_map(|(offset, magic, name)| {
bytes_matches_at(bytes, *offset, magic).then(|| name.as_str())
})
})
})
}
/// Check whether `bytes` contains any known shellcode prologue.
@ -381,188 +343,75 @@ pub fn match_file_signature(bytes: &[u8]) -> Option<&'static str> {
/// [`load_external_rules`].
#[must_use]
pub fn match_shellcode_pattern(bytes: &[u8]) -> Option<&'static [u8]> {
for pattern in SHELLCODE_PATTERNS {
if bytes.windows(pattern.len()).any(|w| w == *pattern) {
return Some(pattern);
}
}
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for pattern in &ext.shellcode {
if bytes.windows(pattern.len()).any(|w| w == pattern.as_slice()) {
return Some(pattern.as_slice());
}
}
}
None
}
/// Check whether `host` (lowercase, no scheme) is a known URL shortener.
#[must_use]
pub fn is_url_shortener(host: &str) -> bool {
URL_SHORTENER_DOMAINS.iter().any(|d| host == *d || host.ends_with(&format!(".{d}")))
}
/// Check whether `uri` mentions a commonly-phished brand with a
/// non-canonical domain (homograph bait).
///
/// Returns the matched brand name if found.
///
/// Logic:
/// 1. If a homograph variant (`micros0ft`, `paypa1`, ...) is found
/// anywhere in the URI, it's always a phishing signal.
/// 2. If a canonical spelling (`microsoft`, `paypal`, ...) is found,
/// we check whether the host portion is ANY brand's canonical
/// domain (e.g. `microsoft.com` for "microsoft"). If yes → benign.
/// Otherwise, the brand is mentioned in a non-canonical context
/// → suspicious.
///
/// Checks both the built-in table and any external brands loaded via
/// [`load_external_rules`].
#[must_use]
pub fn match_phishing_brand(uri: &str) -> Option<&'static str> {
let lower = uri.to_ascii_lowercase();
// Extract host (after scheme://, before path/query/fragment, sans port).
let host = lower
.split("://")
.nth(1)
.unwrap_or(&lower)
.split('/')
.next()
.unwrap_or("")
.split(':')
.next()
.unwrap_or("");
let is_homograph = |brand: &str| brand.chars().any(|c| !c.is_ascii_alphabetic());
// First pass: check for homograph variants — these are ALWAYS phishing.
// Built-in brands.
for brand in COMMON_PHISHING_BRANDS {
if is_homograph(brand) && lower.contains(brand) {
return Some(brand);
}
}
// External brands.
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for brand in &ext.brands {
if is_homograph(brand) && lower.contains(brand.as_str()) {
return Some(brand);
}
}
}
// Second pass: check canonical spellings. If the host is ANY
// brand's canonical domain, all canonical brand mentions are
// treated as benign. This handles cases like `microsoft.com/windows`
// (windows is a brand, but the host is microsoft's canonical domain).
let host_is_canonical_for_builtin = COMMON_PHISHING_BRANDS
.iter()
.any(|brand| !is_homograph(brand) && is_canonical_brand_host(host, brand));
let host_is_canonical_for_external = EXTERNAL_RULES_STORAGE
.get()
.map(|ext| {
ext.brands
.iter()
.any(|brand| !is_homograph(brand) && is_canonical_brand_host(host, brand))
})
.unwrap_or(false);
if host_is_canonical_for_builtin || host_is_canonical_for_external {
return None;
}
// Host is not a canonical brand domain — any canonical brand
// mentioned in the URL is suspicious.
// Built-in brands.
for brand in COMMON_PHISHING_BRANDS {
if !is_homograph(brand) && lower.contains(brand) {
return Some(brand);
}
}
// External brands.
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for brand in &ext.brands {
if !is_homograph(brand) && lower.contains(brand.as_str()) {
return Some(brand);
}
}
}
None
}
/// Check whether `host` is the canonical domain for `brand`.
///
/// A host is canonical if it matches `<brand>.<tld>` or has `<brand>`
/// as a dot-separated segment (e.g. `microsoft.com`, `login.microsoft.com`).
/// The goal is to allow legitimate brand-owned domains while still
/// flagging `login-microsoft.com` (which is NOT a Microsoft domain).
fn is_canonical_brand_host(host: &str, brand: &str) -> bool {
if host == brand {
return true;
}
// Check if `<brand>.<tld>` is a prefix.
let canonical_prefix = format!("{}.", brand);
if host.starts_with(&canonical_prefix) {
return true;
}
// Check if `<brand>` is a dot-separated segment (e.g. `login.microsoft.com`).
host.split('.').any(|seg| seg == brand)
}
/// Check whether `uri`'s path/host contains any suspicious keyword.
///
/// Checks both the built-in table and any external rules loaded via
/// [`load_external_rules`].
#[must_use]
pub fn match_suspicious_keyword(uri: &str) -> Option<&'static str> {
let lower = uri.to_ascii_lowercase();
if let Some(kw) = SUSPICIOUS_URL_KEYWORDS
// Built-in patterns — return the first matching prologue.
let builtin = SHELLCODE_PATTERNS
.iter()
.copied()
.find(|kw| lower.contains(kw))
{
return Some(kw);
}
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
if let Some(kw) = ext
.keywords
.iter()
.find(|kw| lower.contains(kw.as_str()))
{
return Some(kw);
}
}
None
.find(|pattern| bytes.windows(pattern.len()).any(|w| w == *pattern));
// External patterns — only consulted when no built-in matched.
builtin.or_else(|| {
EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
ext.shellcode.iter().find_map(|pattern| {
bytes
.windows(pattern.len())
.any(|w| w == pattern.as_slice())
.then_some(pattern.as_slice())
})
})
})
}
/// Check whether the host part of `uri` ends with a known phishing TLD.
/// Check whether `host` (the URL authority, ASCII-case-folded, no
/// scheme, no port) is byte-equal to a known homograph host string.
///
/// Checks both the built-in table and any external rules loaded via
/// [`load_external_rules`].
/// This is the only brand-impersonation detector in the scanner.
/// Matching is exact-string against the host portion only — never
/// substring, never path/query/fragment. This is the rule that
/// distinguishes `https://github.com/microsoft/vscode` (not flagged —
/// the host is `github.com`) from `https://micros0ft.com/anything`
/// (flagged — the host is exactly `micros0ft.com`).
///
/// # Returns
///
/// `Some(host)` if `host` matches a known homograph, where `host` is
/// the matched homograph string itself (the verifiable artifact —
/// an operator can read this string out of the report and confirm
/// it is a homograph). `None` otherwise.
///
/// Returns an owned `String` rather than `&'static str` so that
/// external-feed entries (which are loaded at runtime and stored in
/// an `ExternalRulesData`) can be returned without unsafe
/// lifetime-extension. The crate forbids `unsafe` so we cannot
/// transmute the external entry's lifetime to `'static`.
#[must_use]
pub fn has_phishing_tld(uri: &str) -> Option<&'static str> {
let lower = uri.to_ascii_lowercase();
// Extract host portion (after scheme://, before path/query/fragment).
let host = lower
.split("://")
.nth(1)
.unwrap_or(&lower)
.split('/')
.next()
.unwrap_or("");
// Strip port.
let host = host.split(':').next().unwrap_or("");
if let Some(tld) = PHISHING_TLDS.iter().copied().find(|tld| host.ends_with(tld)) {
return Some(tld);
pub fn match_homograph_host(host: &str) -> Option<String> {
let lower = host.to_ascii_lowercase();
// Built-in list — direct byte-equal match after case-folding.
let builtin = HOMOGRAPH_HOSTS
.iter()
.find(|entry| lower == **entry)
.map(|entry| (*entry).to_string());
// External feeds — only consulted when no built-in matched.
builtin.or_else(|| {
EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
ext.homograph_hosts
.iter()
.find(|entry| lower == entry.as_str())
.cloned()
})
})
}
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
if let Some(tld) = ext.tlds.iter().find(|tld| host.ends_with(tld.as_str())) {
return Some(tld);
}
}
None
/// Predicate: does `bytes[offset..]` start with `magic`?
/// Returns `false` when `bytes` is shorter than `offset + magic.len()`.
fn bytes_matches_at(bytes: &[u8], offset: usize, magic: &[u8]) -> bool {
let Some(end) = offset.checked_add(magic.len()) else {
return false;
};
bytes.get(offset..end).is_some_and(|slice| slice == magic)
}
#[cfg(test)]
@ -625,70 +474,64 @@ mod tests {
assert!(match_shellcode_pattern(&payload).is_some());
}
// ─── Homograph host tests ────────────────────────────────────
//
// The homograph-host matcher only ever matches the URL authority,
// byte-equal after case-folding. The following tests assert that
// URLs whose hosts are NOT in the homograph list — including
// canonical brand domains and URLs that merely mention a brand in
// their path — are NOT flagged.
#[test]
fn detects_url_shortener() {
assert!(is_url_shortener("bit.ly"));
assert!(is_url_shortener("sub.bit.ly"));
assert!(!is_url_shortener("example.com"));
fn homograph_host_exact_match_is_flagged() {
// micros0ft.com (with zero instead of 'o') is in the list — flagged.
assert_eq!(match_homograph_host("micros0ft.com"), Some("micros0ft.com".to_string()));
assert_eq!(match_homograph_host("paypa1.com"), Some("paypa1.com".to_string()));
}
#[test]
fn detects_phishing_brand_homograph() {
// micros0ft.com (with zero instead of 'o') should match the
// homograph variant directly.
assert_eq!(
match_phishing_brand("https://micros0ft.com/login"),
Some("micros0ft")
);
// paypa1.com (with one instead of 'l') should match.
assert_eq!(
match_phishing_brand("https://paypa1.com/signin"),
Some("paypa1")
);
fn homograph_host_case_insensitive() {
// ASCII-case-folded before matching.
assert_eq!(match_homograph_host("MICROS0FT.COM"), Some("micros0ft.com".to_string()));
assert_eq!(match_homograph_host("PayPa1.COM"), Some("paypa1.com".to_string()));
}
#[test]
fn canonical_brand_domain_not_flagged() {
// microsoft.com (canonical) should not be flagged.
assert!(match_phishing_brand("https://microsoft.com/windows").is_none());
// paypal.com (canonical) should not be flagged.
assert!(match_phishing_brand("https://paypal.com/home").is_none());
fn canonical_brand_host_not_flagged() {
// Real microsoft.com / paypal.com / apple.com hosts are NOT
// in the homograph list — never flagged.
assert!(match_homograph_host("microsoft.com").is_none());
assert!(match_homograph_host("paypal.com").is_none());
assert!(match_homograph_host("apple.com").is_none());
assert!(match_homograph_host("google.com").is_none());
assert!(match_homograph_host("amazon.com").is_none());
}
#[test]
fn canonical_brand_in_non_canonical_domain_is_flagged() {
// microsoft mentioned in a non-canonical host → suspicious.
assert_eq!(
match_phishing_brand("https://login-microsoft.com/verify"),
Some("microsoft")
);
fn brand_mentioned_in_path_not_flagged() {
// The host is github.com — not in the homograph list. The
// "microsoft" substring in the path is irrelevant to the
// host-only check. This is the guarantee that lets clean
// technical documents (RFCs, books, vendor whitepapers) pass
// through the scanner with zero findings.
assert!(match_homograph_host("github.com").is_none());
assert!(match_homograph_host("en.wikipedia.org").is_none());
}
#[test]
fn detects_suspicious_keyword() {
// Should return the first matching keyword — both "account"
// and "verify" are in the list. Either is acceptable; check
// that we get one of them.
let result = match_suspicious_keyword("https://example.com/account/verify");
assert!(matches!(result, Some("account") | Some("verify")));
assert_eq!(
match_suspicious_keyword("https://example.com/signin"),
Some("signin")
);
fn homograph_host_subdomain_not_matched() {
// Subdomains of a homograph host are NOT matched. The check
// is byte-equal on the full host string. This is intentional:
// `login.micros0ft.com` could be a phishing subdomain of a
// homograph domain, but it could also be a coincidence in a
// legitimate subdomain naming scheme. The operator can extend
// the list via external feeds if they want to match subdomains.
assert!(match_homograph_host("login.micros0ft.com").is_none());
assert!(match_homograph_host("www.paypa1.com").is_none());
}
#[test]
fn detects_phishing_tld() {
assert_eq!(has_phishing_tld("https://example.xyz"), Some(".xyz"));
assert_eq!(has_phishing_tld("https://example.top/path"), Some(".top"));
assert!(has_phishing_tld("https://example.com").is_none());
}
#[test]
fn detects_phishing_tld_with_port() {
assert_eq!(
has_phishing_tld("https://example.xyz:8080/path"),
Some(".xyz")
);
fn empty_host_not_flagged() {
assert!(match_homograph_host("").is_none());
}
}

366
tests/fixtures/benign_realworld.pdf vendored Normal file
View File

@ -0,0 +1,366 @@
%PDF-1.3
%âãÏÓ
1 0 obj
<<
/Producer (pypdf)
>>
endobj
2 0 obj
<<
/Type /Pages
/Count 3
/Kids [ 4 0 R 8 0 R 10 0 R ]
>>
endobj
3 0 obj
<<
/Type /Catalog
/Pages 2 0 R
>>
endobj
4 0 obj
<<
/Contents 5 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 13 0 R 15 0 R 17 0 R 19 0 R 21 0 R 23 0 R 25 0 R ]
>>
endobj
5 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 416
>>
stream
Gas2EgM4V[%#46J'Y<(N`,V-N-!-jOKs;*pbmU>MPIO70A'8tT5NfhJMl%="QE3m1s)X"oiVb>HJ8GL,a+5le.tGZK%+u]YZ\=J3n`qNRoeM-#JVU]af-Cgp&3a8QG]nZQ[Z?MU+FC[s(CbLW.=irT=;4B+]EYE-6;C>hBq$l$]jAT[[I0qHfN;R&&ljabA@(LMAjCk-aqp@?.n_f!"8:'qa=a=C!a/g]^cJUFd!V>5\,eW8,W*X\1h.WbVB?nRF_L(^XkftOhO-DGE//<LBq[jdniTpH<)gD,.'R.Wh(d7\02`?C3E;[R^P:0IkOG<C>??E=h-ti_`[aHE::>SaV9:uC2"=M5<dO(X,plN6hL5./IgnG&)hRZj,bp'K$<ro"aZ=k46'<Jo?bQY_Z<C$r@IX`&Pg75~>
endstream
endobj
6 0 obj
<<
/F1 7 0 R
>>
endobj
7 0 obj
<<
/BaseFont /Helvetica
/Encoding /WinAnsiEncoding
/Name /F1
/Subtype /Type1
/Type /Font
>>
endobj
8 0 obj
<<
/Contents 9 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 27 0 R 29 0 R 31 0 R ]
>>
endobj
9 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 482
>>
stream
Gas2E5u68i&;BTNMKbht+AU@L,*!f"$PaDPj]Hfq8MXSlNbi>Er]O"7YW[)0QE9W/n+3$:!dRtg]W3*`Xs(Lo-q_tuMCPZ'hr2"-8blh@HC<f/RA93?iMMf+,`a^uouKV/D%SkHj%G3XImj5Oor%i2o:>ef[+Mf$&KS75>2uEMon-4Zfb,4lqDls0pWdCVM&2L?@\)B(o3`P.BF;oRMDWuFk4e7gR?uHL;1>YjRtalNE(25LC\!cb3'tC4dq+K[*;S'F>e\FB3MaCuDoeik2Cp#F]PM%<BrAtBCcoS;G3iRe+BE;IkTIcJZuKr'*]c(=Ljg5\]u:Z>eAZB]0sQ2=/c!FY4*?m,LiNj#?914f3^S4d:'`E<debP:0b3/7LaEV"+#?WlFIP;JMHqV?'#-!HH/^[lUfm+oA@(Z-G%OGrnUm*N_):c/;U](H17IYO]5`NmfX'<mQRUOErXC865rh[>n8Rq*;%V['~>
endstream
endobj
10 0 obj
<<
/Contents 11 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 33 0 R 35 0 R 37 0 R ]
>>
endobj
11 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 315
>>
stream
Gas3,Z#51J&:i_&:N9lK.5hAs:rYuhe>`*K:id,LN/c,[KXWUg<iVl*=a,M[s0I5d3b$sG"aH?K?Nc0!U]usf*98W?jCYt>^JE#Up1XT6Knj/:FV+p8#"S5-C^Ds33ecN)j<r$H)f;mfmtgJb$dbK\\4F@%<A`PuNNk5sYZ7SL([#n."BS_K%)/Y(lbe$I/Bns!(qXbFHl*'"0P]b7bXnk-8IJD"G_t`[IZ#)L)3JjMM65V(ofTq5CTdg*`n5Mp)]?-IK.7f.+NZKmQ`BF(1?D_H.HPmmq%o4Aj!q>bmig8SNVh>Ejr9j3Q`:~>
endstream
endobj
12 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056linuxfromscratch\056org\057)
>>
endobj
13 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 695 400 710 ]
/A 12 0 R
/Border [ 0 0 0 ]
>>
endobj
14 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057lists\056linuxfromscratch\056org\057listinfo\057lfs\055support)
>>
endobj
15 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 675 400 690 ]
/A 14 0 R
/Border [ 0 0 0 ]
>>
endobj
16 0 obj
<<
/Type /Action
/S /URI
/URI (http\072\057\057ftp\056osuosl\056org\057pub\057lfs\057)
>>
endobj
17 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 655 400 670 ]
/A 16 0 R
/Border [ 0 0 0 ]
>>
endobj
18 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057github\056com\057LFS\055project\057build\055scripts)
>>
endobj
19 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 635 400 650 ]
/A 18 0 R
/Border [ 0 0 0 ]
>>
endobj
20 0 obj
<<
/Type /Action
/S /URI
/URI (mailto\072lfs\055support\100linuxfromscratch\056org)
>>
endobj
21 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 595 400 610 ]
/A 20 0 R
/Border [ 0 0 0 ]
>>
endobj
22 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056kernel\056org\057pub\057linux\057kernel\057)
>>
endobj
23 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 575 400 590 ]
/A 22 0 R
/Border [ 0 0 0 ]
>>
endobj
24 0 obj
<<
/Type /Action
/S /URI
/URI (tel\072\0531\055555\055123\0554567)
>>
endobj
25 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 555 400 570 ]
/A 24 0 R
/Border [ 0 0 0 ]
>>
endobj
26 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057ftp\056ru\056debian\056org\057debian\057)
>>
endobj
27 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 575 400 590 ]
/A 26 0 R
/Border [ 0 0 0 ]
>>
endobj
28 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057en\056wikipedia\056org\057wiki\057Microsoft\137Windows)
>>
endobj
29 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 555 400 570 ]
/A 28 0 R
/Border [ 0 0 0 ]
>>
endobj
30 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057github\056com\057microsoft\057vscode)
>>
endobj
31 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 535 400 550 ]
/A 30 0 R
/Border [ 0 0 0 ]
>>
endobj
32 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056ietf\056org\057rfc\057rfc2616\056txt)
>>
endobj
33 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 655 400 670 ]
/A 32 0 R
/Border [ 0 0 0 ]
>>
endobj
34 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057example\056com\057account\057verify)
>>
endobj
35 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 615 400 630 ]
/A 34 0 R
/Border [ 0 0 0 ]
>>
endobj
36 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057example\056com\057login)
>>
endobj
37 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 595 400 610 ]
/A 36 0 R
/Border [ 0 0 0 ]
>>
endobj
xref
0 38
0000000000 65535 f
0000000015 00000 n
0000000054 00000 n
0000000126 00000 n
0000000175 00000 n
0000000425 00000 n
0000000932 00000 n
0000000963 00000 n
0000001070 00000 n
0000001292 00000 n
0000001865 00000 n
0000002089 00000 n
0000002496 00000 n
0000002599 00000 n
0000002702 00000 n
0000002833 00000 n
0000002936 00000 n
0000003042 00000 n
0000003145 00000 n
0000003265 00000 n
0000003368 00000 n
0000003471 00000 n
0000003574 00000 n
0000003693 00000 n
0000003796 00000 n
0000003882 00000 n
0000003985 00000 n
0000004094 00000 n
0000004197 00000 n
0000004320 00000 n
0000004423 00000 n
0000004528 00000 n
0000004631 00000 n
0000004743 00000 n
0000004846 00000 n
0000004950 00000 n
0000005053 00000 n
0000005145 00000 n
trailer
<<
/Size 38
/Root 3 0 R
/Info 1 0 R
>>
startxref
5248
%%EOF

View File

@ -28,7 +28,101 @@ fn benign_pdf_produces_no_findings() {
let result = pipeline.run(fixture("benign.pdf")).unwrap();
assert_eq!(result.scan_report.malicious_count(), 0);
assert!(result.quarantine_path.is_none());
assert!(result.cleansed_path.is_none());
// Clean documents now produce a "clean" output file in the
// cleanse_dir (the original file copied under
// `clean_<timestamp>_<sha>.pdf`). This makes the scanner a
// proper pipeline stage. The cleansed_path field is overloaded:
// it holds either the cleansed derivative (malicious case) or
// the clean-output copy (clean case).
assert!(
result.cleansed_path.is_some(),
"clean PDF should produce a clean-output file in the cleanse_dir"
);
let clean_path = result.cleansed_path.unwrap();
assert!(
clean_path
.file_name()
.and_then(|n| n.to_str())
.map(|n| n.starts_with("clean_") && n.ends_with(".pdf"))
.unwrap_or(false),
"clean-output file should be named clean_<ts>_<sha>.pdf, got: {clean_path:?}"
);
// The clean-output file should have the same contents as the original.
let original = std::fs::read(fixture("benign.pdf")).unwrap();
let clean_bytes = std::fs::read(&clean_path).unwrap();
assert_eq!(
original, clean_bytes,
"clean-output file should be a byte-for-byte copy of the original"
);
}
#[test]
fn benign_realworld_pdf_produces_zero_findings() {
// This is the corpus-based regression test for the false-positive
// redesign. The fixture (`benign_realworld.pdf`) is a multi-page
// PDF that mimics the structure of a technical book like *Linux
// from Scratch* — it contains 13 hyperlinks covering every URL
// pattern that USED TO produce a false positive:
//
// - https://www.linuxfromscratch.org/
// - https://lists.linuxfromscratch.org/listinfo/lfs-support (keyword "support")
// - http://ftp.osuosl.org/pub/lfs/ (http: scheme)
// - https://github.com/LFS-project/build-scripts (github.com host)
// - mailto:lfs-support@linuxfromscratch.org (mailto: + @ in path)
// - https://www.kernel.org/pub/linux/kernel/
// - tel:+1-555-123-4567 (tel: scheme)
// - https://ftp.ru.debian.org/debian/ (.ru TLD)
// - https://en.wikipedia.org/wiki/Microsoft_Windows (brand in path)
// - https://github.com/microsoft/vscode (brand in path)
// - https://www.ietf.org/rfc/rfc2616.txt
// - https://example.com/account/verify (keywords in path)
// - https://example.com/login (keyword in path)
//
// Plus prose containing: wget, exploit, payload, /bin/sh, PowerShell.
//
// The assertion is the theorem: a clean technical PDF produces ZERO
// findings (not "fewer than N", not "0 malicious but maybe some
// suspicious" — literally zero findings in the report). If any
// finding appears, the detector that produced it is wrong by
// construction, not the document.
let tmp = tempdir().unwrap();
let pipeline = Pipeline::with_config(Config::with_workspace(tmp.path()));
let result = pipeline.run(fixture("benign_realworld.pdf")).unwrap();
assert_eq!(
result.scan_report.findings.len(),
0,
"benign real-world PDF must produce ZERO findings, got {}: {:#?}",
result.scan_report.findings.len(),
result.scan_report.findings.iter().map(|f| (
f.classification.clone(),
f.context_notes.clone(),
f.payload_preview.chars().take(80).collect::<String>(),
)).collect::<Vec<_>>(),
);
assert_eq!(result.scan_report.malicious_count(), 0);
assert!(result.quarantine_path.is_none());
// The clean-output file should exist (the pipeline-stage behavior).
assert!(
result.cleansed_path.is_some(),
"clean real-world PDF should produce a clean-output file"
);
let clean_path = result.cleansed_path.unwrap();
assert!(
clean_path
.file_name()
.and_then(|n| n.to_str())
.map(|n| n.starts_with("clean_") && n.ends_with(".pdf"))
.unwrap_or(false),
"clean-output file should be named clean_<ts>_<sha>.pdf, got: {clean_path:?}"
);
// The clean-output file should be a byte-for-byte copy of the original.
let original = std::fs::read(fixture("benign_realworld.pdf")).unwrap();
let clean_bytes = std::fs::read(&clean_path).unwrap();
assert_eq!(
original, clean_bytes,
"clean-output file should be a byte-for-byte copy of the original"
);
}
#[test]