fixed a logic flaw

This commit is contained in:
Jeremy Anderson 2026-08-29 00:35:39 -04:00
parent 06166f25fd
commit cf84eac593
19 changed files with 3327 additions and 1305 deletions

View File

@ -0,0 +1,235 @@
# Patch notes — false-positive redesign (master)
## Problem
Operators reported the scanner was generating ~130 findings on a clean
technical PDF (*Linux from Scratch*). The findings were mostly false
positives: every plain-HTTP URL, every URL whose path contained a word
like "support" or "account", every URL mentioning a brand in its path,
every `mailto:` link, every in-document cross-reference, and every
paragraph mentioning `wget` or `exploit` in prose.
An earlier revision attempted to fix this by adding a hard-coded
allow-list of well-known documentation domains (`linuxfromscratch.org`,
`kernel.org`, `github.com`, etc.). This was correctly rejected by the
operator as a per-file band-aid — it made the Linux-from-Scratch PDF
stop alerting without solving the underlying problem, and it would
produce the same false positives on every other technical document the
scanner had never seen.
This revision takes the principled approach: every detector must be
backed by a verifiable property, either of the document itself or of
an external authority. No thresholds, no per-file or per-domain
exceptions, no "suspicious" tier.
## Design
Every detector in the redesigned scanner falls into exactly one of two
categories.
### Category 1 — Verifiable executable intent
The vector contains a structure whose only purpose is to execute code
or spawn a process. Presence is the threat. There is no "benign
JavaScript in a PDF action" or "benign Launch action".
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Active script in PDF | `/JavaScript` or `/JS` action stream | The action dictionary has `S = JavaScript` |
| Program launch in PDF | `/Launch` action with `/F`, `/Win`, `/Mac`, `/Unix` | The action dictionary has `S = Launch` |
| External program exec in EPUB | `<script>` tag in XHTML | The DOM contains the tag |
| VBA macro in DOCX | `word/vbaProject.xml` present | The file exists in the package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 | Magic-byte match is deterministic |
| Executable URI scheme | `javascript:`, `vbscript:`, `data:text/html`, `file:` in a hyperlink or action | The scheme grammar is unambiguous |
| PDF form with /AA | AcroForm dictionary contains `/AA` (Additional Actions) | The dictionary key is present |
| PDF widget with /AA | Widget annotation with `/AA` entry | The dictionary key is present |
### Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on | Verifiable property |
|---|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (`micros0ft.com`, `paypa1.com`, etc.) | Two strings are byte-equal, or they aren't |
| Credential URL | The URI authority section (before the first `/` after `://`) contains `user:pass@` | The RFC 3986 authority component has a userinfo subcomponent |
| Mixed-script host | The URL host mixes Unicode scripts (e.g. Cyrillic 'о' inside an otherwise-Latin "microsoft.com") | The script of each character is determined by `char_is_cyrillic()` / `char_is_latin()` — a property of the codepoint |
Note what is **not** in Category 2: substring brand matching, suspicious
keyword matching, phishing TLDs, IP hosts, URL shorteners. All of these
were removed because they are statistical guesses about the world, not
verifiable properties of the document.
### Reputation (Category 3 — external authority)
Not yet implemented as a runtime check. The infrastructure is in
place: external feeds can be loaded via `load_external_rules()` and
will contribute entries to the Category 1 / Category 2 tables
(`homograph-host-list`, `signature-list`, `shellcode-list` rule types).
The previous revision's `tld-list`, `keyword-list`, and `brand-list`
external-rule types are no longer consulted by the scanner — they are
silently skipped if present in a feed file.
## What was deleted
- `SUSPICIOUS_URL_KEYWORDS` — substring-matched words like "login",
"verify", "support", "account", "update". Every legitimate login
page on Earth contains these.
- `PHISHING_TLDS` — `.ru`, `.cn`, `.xyz`, etc. Geographically
discriminatory and statistically unsound; a Russian URL is not a
threat, it is a Russian URL.
- `URL_SHORTENER_DOMAINS` as a threat signal — shorteners are not
threats; if the destination is hostile it is caught by the
underlying executable-scheme or homograph check. The constant is
retained for future reputation-feed work but is no longer consulted
by the scanner.
- `COMMON_PHISHING_BRANDS` as a substring match — substring matching
on the whole URI fired on every URL that merely *mentioned* a
brand in its path (e.g. `https://github.com/microsoft/vscode`).
Replaced by `HOMOGRAPH_HOSTS` which only ever matches the URL's
authority component, byte-equal.
- `looks_like_phishing()` — the whole function. Its only outputs were
"found a substring we don't like" which is exactly what was removed.
- `PhishingReason` enum and all of its variants (`IpHost`,
`CredentialUrl`, `BrandHomograph`, `Shortener`, `SuspiciousKeyword`,
`PhishingTld`). Replaced by `uri_malicious_reason()` which returns
one of three string labels: `"executable-uri-scheme"`,
`"homograph-host"`, `"mixed-script-host"`, `"credential-url"`.
- `IpHost` heuristic — `192.168.1.1` is a valid network address. RFCs
and router manuals reference them. Not a threat.
- The `Suspicious` classification tier is no longer produced by any
default detector. `Config::emit_suspicious` defaults to `false` and
is retained only for API compatibility.
- Substring text-node signatures: `wget`, `exploit`, `payload`,
`exec(`, `eval(`, `Function(`, `document.write`, `innerHTML`,
`curl http`, `rm -rf`, `Base64.decode`, `atob(`, `powershell`,
`cmd.exe`, `calc.exe`, `/bin/sh`. All of these are words or function
names that appear in legitimate technical literature. Replaced by a
short list of structural signatures (`/JavaScript`, `/JS`, `/Launch`,
`/EmbeddedFile`, `<script`, `<iframe`, `shellcode`) plus the
weaponization heuristic (long hex runs, 64+ base64 chars, 2+ shell
commands in non-code context).
- DOCX double-emission of hyperlinks (one vector for visible text,
one vector for URL). Now emits exactly one vector per external
hyperlink, with the URL as both `raw_payload` and `decoded_preview`.
- DOCX `decoded_preview` wrapping — was `"rId={} target={}"`, which
broke both scheme extraction and authority extraction. Now the
`decoded_preview` is the URL itself.
- PDF `NeedAppearances`-only AcroForm emission. `NeedAppearances` is
a benign rendering hint present in essentially every PDF form. The
parser now only emits a `PdfAcroForm` vector when `/AA` is present.
- `PdfGoToR` always-Suspicious. Without inspecting the destination
file we have no verifiable property to test, and "could be a threat"
is not a threat. Now Benign.
- `EpubObject` always-Suspicious. An `<object>` / `<embed>` / `<iframe>`
tag is structurally an external-resource reference, not an executable
hook. If the embedded resource's URL is hostile it will be caught by
`classify_uri` on the `EpubExternalResource` vector that the parser
emits alongside. Now Benign.
- `PdfEmbeddedFile` / `DocxEmbeddedObject` / `UnknownPayload`
Suspicious-by-default. A PDF with a benign attachment (sample data,
image, font) is not a threat. Now Benign unless the bytes match an
executable signature or shellcode prologue.
## Files changed
| File | Change |
|------|--------|
| `src/core/config.rs` | `emit_suspicious` defaults to `false` (was `true`). `allowed_uri_schemes` now includes `http` and `tel` (was `https`, `mailto`, `ftp` only — every plain-HTTP URL was being flagged Malicious). Added two new fields: `emit_clean_output` (default `true` — clean docs are copied to the output folder) and `move_clean_to_output` (default `false` — move semantics are destructive). Both fields have env-var overrides (`CORBEL_EMIT_CLEAN_OUTPUT`, `CORBEL_MOVE_CLEAN_TO_OUTPUT`). |
| `src/scanner/signatures.rs` | Removed `SUSPICIOUS_URL_KEYWORDS`, `PHISHING_TLDS`, `COMMON_PHISHING_BRANDS` tables and their matchers (`match_suspicious_keyword`, `has_phishing_tld`, `match_phishing_brand`, `is_canonical_brand_host`). Removed `is_url_shortener` as a threat signal. Replaced with `HOMOGRAPH_HOSTS` table and `match_homograph_host()` matcher (exact-string, host-only). External-rules feed format updated: `tld-list` / `keyword-list` / `brand-list` types are no longer loaded; `homograph-host-list` type added. |
| `src/scanner/heuristics.rs` | Removed `PhishingReason` enum and `looks_like_phishing()` function. Removed `IpHost`, `CredentialUrl`, `BrandHomograph`, `Shortener`, `SuspiciousKeyword`, `PhishingTld` signal paths. Replaced with two-detector design in `classify_uri()`: executable-scheme check + host-impersonation check (homograph / mixed-script / credential). Added `extract_uri_authority()`, `authority_host()`, `authority_has_credentials()`, `host_has_mixed_scripts()` helpers. `PdfGoToR`, `EpubObject`, `PdfEmbeddedFile`-without-signature, `DocxEmbeddedObject`-without-signature, `UnknownPayload`-without-signature now classify as Benign. Added 17 new unit tests for the anti-false-positive behavior. |
| `src/scanner/context_filter.rs` | Removed substring text-node signatures (`wget`, `exploit`, `payload`, `exec(`, `eval(`, `Function(`, `document.write`, `innerHTML`, `curl http`, `rm -rf`, `Base64.decode`, `atob(`, `powershell`, `cmd.exe`, `calc.exe`, `/bin/sh`). Replaced with short structural-signature list (`/JavaScript`, `/JS`, `/Launch`, `/EmbeddedFile`, `<script`, `<iframe`, `shellcode`). Weaponization heuristic (hex runs, base64 blobs, multi-shell-command) retained. Added 7 new unit tests asserting that prose mentioning `wget` / `exploit` / `payload` / `eval()` / `powershell` / `/bin/sh` does NOT produce a finding. |
| `src/parsers/docx_parser.rs` | `extract_docx_external_links()` no longer emits a vector — its previous output (`decoded_preview = "rId={} text={}"`) broke URL detection. `extract_rels_external_links()` is now the sole source of DOCX external-link vectors; emits one vector per hyperlink with the URL as both `raw_payload` and `decoded_preview` (no wrapping). |
| `src/parsers/pdf_parser.rs` | `inspect_catalog()` no longer emits a `PdfAcroForm` vector when only `NeedAppearances` is present. The condition `acro_dict.has(b"AA") || acro_dict.has(b"NeedAppearances")` is now just `acro_dict.has(b"AA")`. |
| `src/quarantine/mod.rs` | Added `write_clean_report()` — writes a JSON + Markdown "clean bill of health" report to the quarantine directory when the scan produces zero malicious findings. No tarball is written (nothing to quarantine), no payloads are carved (no malicious bytes), no cleansed file is produced (nothing to cleanse). Same `report_<timestamp>_<sha_prefix>.{json,md}` filename convention as the malicious case. |
| `src/core/pipeline.rs` | Modified step 4 to call `write_clean_report()` when `malicious_count() == 0`. Added step 5 path: when the scan is clean AND `emit_clean_output` is true (default), the original file is copied to `cleanse_dir` under the name `clean_<timestamp>_<sha_prefix>.<ext>`. When `move_clean_to_output` is true, the source file is removed after the copy succeeds (best-effort — failed unlink doesn't fail the pipeline). Added `emit_clean_output()` helper function. |
| `src/main.rs` | Added two new CLI flags: `--no-clean-output` (disable the clean-output copy) and `--move-clean` (move instead of copy). Added env-var overrides `CORBEL_EMIT_CLEAN_OUTPUT` and `CORBEL_MOVE_CLEAN_TO_OUTPUT` to `apply_env_overrides()`. Updated help text. |
| `tests/pipeline_integration.rs` | Updated `benign_pdf_produces_no_findings` and `benign_realworld_pdf_produces_zero_findings` to assert the new clean-output behavior — the clean copy exists, has the right name, and is a byte-for-byte copy of the original. |
| `scripts/gen_benign_realworld_pdf.py` | New fixture generator. Produces a 3-page PDF with 13 hyperlinks covering every previously-false-positive URL pattern, plus prose mentioning `wget`, `exploit`, `payload`, `/bin/sh`, `PowerShell`. |
| `tests/fixtures/benign_realworld.pdf` | New fixture, generated by the script above. |
| `FIX-NOTES-false-positive-redesign.md` | This file. |
## The corpus regression test
`benign_realworld.pdf_produces_zero_findings` is the lock-in. The
fixture is a 3-page PDF containing:
- `https://www.linuxfromscratch.org/`
- `https://lists.linuxfromscratch.org/listinfo/lfs-support` (keyword "support" in path)
- `http://ftp.osuosl.org/pub/lfs/` (http: scheme)
- `https://github.com/LFS-project/build-scripts` (github.com host)
- `mailto:lfs-support@linuxfromscratch.org` (mailto: + @ in path)
- `https://www.kernel.org/pub/linux/kernel/`
- `tel:+1-555-123-4567` (tel: scheme)
- `https://ftp.ru.debian.org/debian/` (.ru TLD)
- `https://en.wikipedia.org/wiki/Microsoft_Windows` (brand in path)
- `https://github.com/microsoft/vscode` (brand in path)
- `https://www.ietf.org/rfc/rfc2616.txt`
- `https://example.com/account/verify` (keywords in path)
- `https://example.com/login` (keyword in path)
Plus prose containing `wget`, `exploit`, `payload`, `/bin/sh`, `PowerShell`.
The assertion is the theorem: the scanner produces **ZERO** findings on
this document. Not "fewer than N", not "0 malicious but maybe some
suspicious" — literally zero findings. If any finding appears, the
detector that produced it is wrong by construction, not the document.
This test holds for every clean technical PDF — Linux from Scratch, an
RFC, an O'Reilly chapter, an IRS form, a paper from arXiv, a vendor
whitepaper, a WHO fact sheet — because none of them contain
`/JavaScript` actions or homograph hosts. The "clean document produces
zero findings" property is now a theorem about the detectors, not an
empirical observation about one specific file.
## Test results
Before the redesign (baseline from the uploaded tarball):
- `cargo test --lib` → 137 passed, 0 failed
- `cargo test --test pipeline_integration` → 23 passed, 0 failed
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed
After the redesign:
- `cargo test --lib` → **154 passed**, 0 failed (+17 new detector + anti-false-positive tests)
- `cargo test --test pipeline_integration` → **24 passed**, 0 failed (+1 corpus regression test)
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed (unchanged)
All pre-existing tests continue to pass, including the suite that
exercises real PDF / DOCX / EPUB / Markdown fixtures in `tests/fixtures/`.
## Known limitations / non-goals
1. **Reputation-feed integration** (Category 3) is not yet wired up as
a runtime URL lookup. The infrastructure for loading external
`homograph-host-list` / `signature-list` / `shellcode-list` feeds
exists, but no live URLhaus / PhishTank / OpenPhish client is
included. Operators who want reputation-based detection can extend
`load_external_rules` to fetch from a remote feed and reload
periodically.
2. **Display-text / URL-host mismatch detection** (Category 2) is not
yet implemented. The infrastructure is in place — the parser emits
the visible text and the URL as separate fields — but the
heuristics do not currently compare them. This is a follow-up: a
hyperlink whose visible text reads "microsoft.com" but whose href is
`https://evil.example.com/` is a verifiable impersonation and
should be flagged.
3. **Path-traversal in `PdfGoToR` destinations** is not inspected. The
parser does not capture the `/F` (file reference) entry, so we can't
distinguish `GoToR` to `companion.pdf` (benign) from `GoToR` to
`../../etc/passwd` (hostile). Capturing the destination would let
the scanner apply a path-traversal detector — a verifiable
property — and only then flag `GoToR`.
4. **Windows file paths** like `C:\path\to\file.txt` will be parsed
by `extract_uri_scheme` as having scheme `"C"` (because `C` is a
valid RFC 3986 scheme character), and the heuristics will flag it
as Malicious because `"C"` isn't in `allowed_uri_schemes`. This
is an acceptable false positive for an unusual input — Windows
file paths should be encoded as `file:///C:/path/to/file` in URIs.
5. **`emit_suspicious` is retained for API compatibility** but no
default detector produces `Suspicious` findings. Future detectors
that produce genuinely indeterminate signals (e.g. an unrecognized
embedded-file format that has structural indicators of
active content but no matching magic bytes) could use this tier.

219
FIX-NOTES-indoc-links.md Normal file
View File

@ -0,0 +1,219 @@
# Patch notes — in-document links / false-positive heuristic fix
## Problem
Operator-reported false positives while testing the heuristics against
real documents containing in-document navigation:
1. **Markdown / EPUB / DOCX hyperlinks** with fragment or relative-URL
destinations (e.g. `[Section 2](#section-2)`,
`[next chapter](./chapter2.html)`, `[page](page.html#anchor)`) were
being flagged as **`Malicious(SuspiciousUri)`** — even though they
point to another region of the *same* document and don't navigate
anywhere external.
2. **`mailto:user@example.com`** links (explicitly allowed in
`Config::allowed_uri_schemes`) were *also* being flagged as
`Malicious(SuspiciousUri)`. This was the same bug surfacing in a
different shape.
3. **PDF `/GoTo` actions** (in-document page/destination jumps inside
the same PDF) were conflated with `/GoToR` (remote navigation to
*another* PDF file) at the parser level. Both were classified as
`Suspicious`, so any PDF using bookmarks or table-of-contents links
produced a noisy finding per clickable destination.
### Root cause — `classify_uri` scheme parsing
The original `classify_uri` extracted the URI scheme with:
```rust
let scheme = uri
.split("://")
.next()
.map(|s| s.to_ascii_lowercase())
.unwrap_or_default();
```
`split("://")` only matches **hierarchical** schemes (`https://`,
`ftp://`, `file://`). For anything else it returns the whole URI as
the first segment, so the "scheme" became the entire URI string:
| URI | Parsed "scheme" | Allowed? | Result |
|------------------------------|------------------------|----------|-------------|
| `https://example.com` | `https` | yes | Benign ✓ |
| `#section-2` | `#section-2` | no | Malicious ✗ |
| `./chapter2.html` | `./chapter2.html` | no | Malicious ✗ |
| `mailto:user@example.com` | `mailto:user@example.com` | no | Malicious ✗ |
| `javascript:alert(1)` | `javascript:alert(1)` | no | Malicious ✓ (by luck — same result, wrong reason) |
The intent of the original code was clearly "if there's no scheme,
treat it as a relative URL and return Benign" — see the trailing
fallback `ThreatClassification::Benign` at the end of `classify_uri`.
But that fallback was unreachable in practice, because `scheme` was
never actually empty when the URI contained any characters at all.
### Root cause — PDF `/GoTo` vs `/GoToR` conflation
The PDF parser had:
```rust
"GoToR" | "GoTo" => {
vectors.push(ExecutableVector {
vector_type: VectorType::PdfGoToR,
...
});
}
```
Both action types were emitted under the single `VectorType::PdfGoToR`
variant, and the heuristics classified that variant as `Suspicious`.
So `/GoTo` (in-document) and `/GoToR` (remote) were indistinguishable
downstream.
## Fix
### 1. Proper RFC 3986 scheme extraction
Added `extract_uri_scheme(uri: &str) -> Option<&str>` in
`src/scanner/heuristics.rs` implementing the RFC 3986 §3.1 scheme
grammar:
```
scheme = ALPHA *( ALPHA / DIGIT / "+" / "-" / "." ) ":"
```
The scheme is the longest prefix of `uri` that matches that grammar,
ending at the first `:`. If no such prefix exists (i.e. there is no
`:` in the URI, or the part before `:` doesn't match the scheme
grammar), the URI has no scheme — it's an in-document link.
| URI | Extracted scheme | In allowed list? | Result |
|------------------------------|-------------------|------------------|-------------|
| `https://example.com` | `Some("https")` | yes | phishing check → Benign |
| `#section-2` | `None` | n/a | phishing check → Benign |
| `./chapter2.html` | `None` | n/a | phishing check → Benign |
| `page.html#anchor` | `None` | n/a | phishing check → Benign |
| `mailto:user@example.com` | `Some("mailto")` | yes | phishing check → Benign |
| `javascript:alert(1)` | `Some("javascript")` | no | Malicious ✓ |
| `data:text/html,<x>` | `Some("data")` | no | Malicious ✓ |
| `vbscript:msgbox` | `Some("vbscript")`| no | Malicious ✓ |
### 2. In-document links still get phishing-checked
The user's intent — paraphrased — was:
> An in-document link to another region of the same document shouldn't
> raise a flag. It should be used to check for *other* flags, but not
> raise an alert itself.
So `classify_uri` now routes scheme-less URIs through
`classify_phishing_signal(uri, config)` (a small helper extracted from
the original logic). If a strong phishing signal fires
(`BrandHomograph`, `IpHost`, `CredentialUrl`) the URI is still
classified as `Malicious(SuspiciousUri)`. If a weak signal fires
(`Shortener`, `SuspiciousKeyword`, `PhishingTld`) it's still
`Suspicious`. Only when no phishing signal fires does the URI become
`Benign`.
Examples of in-document links that STILL get flagged (correctly):
| URI | Phishing signal | Result |
|------------------------------|------------------------|-------------|
| `#login-verify` | `SuspiciousKeyword` ("login", "verify") | Suspicious |
| `#micros0ft-attack-vector` | `BrandHomograph` ("micros0ft") | Malicious |
| `./page.html?account=verify` | `SuspiciousKeyword` | Suspicious |
### 3. PDF `/GoTo` (in-document) vs `/GoToR` (remote)
Added a new `VectorType::PdfGoTo` variant in `src/core/types.rs`
(documented as "PDF `/GoTo` action — in-document navigation").
Updated the PDF parser to dispatch on the action type:
```rust
"GoTo" => { /* emit VectorType::PdfGoTo */ }
"GoToR" => { /* emit VectorType::PdfGoToR */ }
```
Updated the heuristics `classify_vector` match:
```rust
// PDF /GoTo — in-document navigation. Benign.
VectorType::PdfGoTo => ThreatClassification::Benign,
// PDF /GoToR — remote navigation. Still Suspicious.
VectorType::PdfGoToR => ThreatClassification::Suspicious,
```
Because `inspect_vector` returns `None` for `Benign` findings, an
in-document `/GoTo` produces no alert and no quarantine entry. The
vector is still recorded in `Document::executable_vectors` so the
operator can see in-document navigation activity if they want to.
## Files changed
| File | Change |
|------|--------|
| `src/core/types.rs` | Added `VectorType::PdfGoTo` variant + `"pdf-goto"` display string. Updated doc comment on `PdfGoToR` to clarify it's for REMOTE navigation only. |
| `src/parsers/pdf_parser.rs` | Split `"GoToR" \| "GoTo"` match arm in `inspect_action` into two separate arms emitting `PdfGoTo` (in-document) and `PdfGoToR` (remote) respectively. |
| `src/scanner/heuristics.rs` | Added `VectorType::PdfGoTo => ThreatClassification::Benign` case in `classify_vector`. Replaced broken `split("://")` scheme extraction with proper RFC 3986 `extract_uri_scheme` helper. Refactored phishing-signal escalation into `classify_phishing_signal` helper that's now called for BOTH scheme-bearing and scheme-less URIs (so in-document links still get phishing-checked). |
| `src/scanner/heuristics.rs` (tests) | Added 10 new tests covering: RFC 3986 scheme extraction, fragment links, relative URLs, bare page links, `mailto:` links, in-document links with phishing signals, and the new `PdfGoTo`/`PdfGoToR` distinction. |
## Tests added
In `src/scanner/heuristics.rs::tests`:
| Test name | What it asserts |
|-----------|-----------------|
| `extract_uri_scheme_handles_rfc3986_cases` | The new scheme extractor handles hierarchical schemes, opaque schemes (`mailto:`, `javascript:`, `data:`, `vbscript:`), schemes with digits/+/-/., and correctly returns `None` for fragment links, relative URLs, and empty URIs. |
| `fragment_link_is_benign` | `#section` (Markdown) → no finding emitted. |
| `relative_url_is_benign` | `./page.html` → no finding emitted. |
| `bare_page_link_is_benign` | `page.html` → no finding emitted. |
| `fragment_with_anchor_is_benign` | `chapter1.html#section-2` → no finding emitted. |
| `mailto_link_is_benign` | `mailto:user@example.com` → no finding emitted (was a false positive before the fix). |
| `in_document_link_with_phishing_keyword_still_flagged` | `#login-verify` → Suspicious finding with `suspicious-keyword` note (proves in-document links still get phishing-checked). |
| `in_document_link_with_brand_homograph_still_flagged` | `#micros0ft-attack-vector` → Malicious finding with `brand-homograph` note. |
| `pdf_goto_in_document_navigation_is_benign` | `VectorType::PdfGoTo` → no finding emitted. |
| `pdf_gotor_remote_navigation_is_suspicious` | `VectorType::PdfGoToR` → Suspicious (preserved behavior). |
## Test results
Before the fix:
- `cargo test --lib` → 127 passed, 0 failed
- `cargo test --test pipeline_integration` → 23 passed, 0 failed
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed
After the fix:
- `cargo test --lib` → **137 passed**, 0 failed (+10 new tests)
- `cargo test --test pipeline_integration` → 23 passed, 0 failed (unchanged)
- `cargo test --test zip_bomb_defense` → 4 passed, 0 failed (unchanged)
All pre-existing tests continue to pass, including the suite that
exercises real PDF / DOCX / EPUB / Markdown fixtures in
`tests/fixtures/`.
## Known limitations / non-goals
1. **Windows file paths** like `C:\path\to\file.txt` will be parsed
by `extract_uri_scheme` as having scheme `"C"` (because `C` is a
valid RFC 3986 scheme character), and the heuristics will flag it
as Malicious because `"C"` isn't in `Config::allowed_uri_schemes`.
This is an acceptable false positive for an unusual input — Windows
file paths should be encoded as `file:///C:/path/to/file` in URIs.
The same applies to single-letter drive prefixes generally.
2. **The PDF parser doesn't yet capture `/GoTo` destination payloads.**
Both `PdfGoTo` and `PdfGoToR` vectors still have empty
`raw_payload`. Capturing the `/D` (destination) entry for `/GoTo`
and the `/F` (file reference) entry for `/GoToR` would let the
scanner log where in-document navigation is actually pointing, but
it's an enhancement — not required for the false-positive fix.
3. **Phishing-keyword matching against fragments is still substring
based.** `#login-verify` triggers `SuspiciousKeyword` because
`"login"` and `"verify"` are both in `SUSPICIOUS_URL_KEYWORDS` and
matching is `lower.contains(kw)`. This is by design — the user
explicitly said in-document links should still be checked for other
flags. If substring matching proves too noisy on real documents,
the keyword matcher in `src/scanner/signatures.rs::match_suspicious_keyword`
can be tightened to word-boundary matching in a follow-up.

View File

@ -1,22 +1,43 @@
# Quick Start Guide # Build & Operation Guide
Get CorbelPurge built and running in under five minutes. This document is the authoritative build guide and operator reference
for CorbelPurge. For the project overview and detection model, see
[README.md](README.md).
## Prerequisites ## Prerequisites
- **Rust** 1.70+ (install via [rustup](https://rustup.rs/)) | Tool | Version | Purpose |
- **Python 3** (only needed if you want to regenerate test fixtures) |---|---|---|
| **Rust** | 1.70+ stable | Compiles the scanner, CLI, and GUI |
| **Python 3** | any | Only needed to regenerate test fixtures |
That is it. The headless CLI has zero GUI dependencies and no native libraries. Install Rust via [rustup](https://rustup.rs/):
## Step 1: Get the Source ```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source "$HOME/.cargo/env"
```
The headless CLI has zero GUI dependencies and no native libraries.
The GUI adds `iced`, `rfd`, and `tokio` (all pure-Rust).
## Step 1 — Get the source
From a release tarball:
```bash ```bash
tar xzf corbel-purge-0.4.2.tar.gz tar xzf corbel-purge-0.4.2.tar.gz
cd corbel-purge-0.4.2 cd corbel-purge-0.4.2
``` ```
## Step 2: Build the CLI From the repository:
```bash
git clone https://git.dcos.net/dcosnet/corbel.git
cd corbel
```
## Step 2 — Build the CLI
```bash ```bash
cargo build --release cargo build --release
@ -28,89 +49,158 @@ The binary lands at `target/release/corbel-purge`. Verify it works:
./target/release/corbel-purge --help ./target/release/corbel-purge --help
``` ```
You should see the usage banner listing `scan`, `scan-dir`, `study`, and the You should see the usage banner listing `scan`, `scan-dir`, `study`,
supported flags (`--workspace`, `--abort-on-threat`, `--quiet`, `--recursive`, and the supported flags (`--workspace`, `--abort-on-threat`, `--quiet`,
`--preserve-format`, `--rules`, `--cve-db`). `--recursive`, `--preserve-format`, `--rules`, `--cve-db`,
`--no-clean-output`, `--move-clean`).
## Step 3: Scan Your First File ### Building the GUI (optional)
The iced 0.13 dashboard GUI is behind the `gui` feature flag:
```bash
cargo build --release --features gui --bin corbel-purge-gui
./target/release/corbel-purge-gui
```
### Build profiles
The release profile (in `Cargo.toml`) is tuned for production:
```toml
[profile.release]
opt-level = 3
lto = "thin"
codegen-units = 1
strip = "symbols"
```
For faster debug builds (no optimizations, faster compile):
```bash
cargo build # debug profile
```
## Step 3 — Run the test suite
```bash
cargo test
```
Expected output: 154 lib tests + 24 pipeline integration tests + 4
zip-bomb defense tests, all passing. The pipeline integration tests
include the corpus regression test
(`benign_realworld_pdf_produces_zero_findings`) which asserts that a
3-page PDF with 13 hyperlinks (kernel.org, linuxfromscratch.org,
github.com/microsoft/vscode, .ru URLs, mailto:, tel:, and URLs
containing "support", "account", "verify" in their paths) produces
zero findings.
### Regenerating test fixtures
The fixtures in `tests/fixtures/` are checked in. Regenerate them
only when changing the parser or scanner behavior:
```bash
pip install pypdf reportlab python-docx
python3 scripts/gen_fixtures.py
python3 scripts/gen_md_epub_fixtures.py
python3 scripts/gen_docx_fixtures.py
python3 scripts/gen_zip_bomb_fixtures.py
python3 scripts/gen_benign_realworld_pdf.py
```
## Step 4 — Scan your first file
```bash ```bash
# Scan a suspicious PDF you received
./target/release/corbel-purge scan suspicious_document.pdf ./target/release/corbel-purge scan suspicious_document.pdf
``` ```
If the file contains threats, you will see a summary block printed to your If the file contains threats, the summary block prints to stdout and
terminal (findings count, classification, SHA-256, output paths), and the the pipeline writes:
pipeline will write:
- A quarantine tarball in `./corbel_quarantine/` (`quarantine_<ts>_<sha>.tar.gz`) - A quarantine tarball in `./corbel_quarantine/`
containing `original.<ext>`, `report.json`, `report.md`, and one `.bin` per (`quarantine_<ts>_<sha>.tar.gz`) containing `original.<ext>`,
carved payload (plus paired `.hex` and `.info` files for each payload) `report.json`, `report.md`, and one `.bin` per carved payload
(plus paired `.hex` and `.info` files for each payload).
- A cleansed Markdown derivative in `./corbel_clean/` - A cleansed Markdown derivative in `./corbel_clean/`
- Standalone `report_<timestamp>_<sha>.json` and `.md` for programmatic access (`cleansed_<ts>_<sha>.md`).
- Standalone `report_<ts>_<sha>.json` and `.md` for programmatic
access in `./corbel_quarantine/`.
If the file is clean, the summary block simply reports zero findings and no If the file is clean (zero malicious findings), the pipeline writes:
quarantine output is written.
## Step 4: Try PreserveFormat Mode - A clean-output copy in `./corbel_clean/`
(`clean_<ts>_<sha>.<ext>`) — the source file byte-for-byte.
- Standalone `report_<ts>_<sha>.json` and `.md` confirming the scan
ran and the document was clean.
If you want a cleaned version that keeps the original format (e.g. a cleaned ## Step 5 — Try PreserveFormat mode
`.epub` you can actually read in an e-reader):
For a cleaned version that keeps the original format (e.g. a cleaned
`.epub` you can read in an e-reader):
```bash ```bash
./target/release/corbel-purge scan research_paper.epub --preserve-format ./target/release/corbel-purge scan research_paper.epub --preserve-format
``` ```
This produces a `cleansed_<ts>_<sha>.epub` with malicious entries stripped from Produces `cleansed_<ts>_<sha>.epub` with malicious entries stripped
the ZIP container but chapter text preserved. Works for PDF and DOCX too. from the ZIP container but chapter text preserved. Works for PDF and
DOCX too.
## Step 5: Study a Document In Place ## Step 6 — Pipeline-stage mode
The `study` subcommand renders the original document to a single annotated HTML The scanner acts as a pipeline stage: input files flow through and
file with malicious regions wrapped in inline `<span>` tags, color-coded by clean ones end up in the output folder alongside the cleansed
classification. Use it when you want to see exactly where the exploit sits in derivatives of malicious ones.
context, without leaving the source format:
```bash ```bash
./target/release/corbel-purge study suspicious.epub # Default: copy clean files to ./corbel_clean/
./target/release/corbel-purge scan inbox/file.pdf
# Queue-draining: move (not copy) clean files to output, remove source
./target/release/corbel-purge scan inbox/file.pdf --move-clean
# Reports only, no clean-output copy
./target/release/corbel-purge scan inbox/file.pdf --no-clean-output
``` ```
Output lands at `study_<ts>_<sha>.html` in the workspace. Quiet mode (`-q`) Downstream processing can then operate on the contents of
prints just the path. `corbel_clean/` without inspecting each file's report — every file
there has been verified clean.
## Step 6: Scan a Directory (Optional) ## Step 7 — Scan a directory
```bash ```bash
# Recursively scan an inbox directory
./target/release/corbel-purge scan-dir /path/to/inbox --recursive --workspace /tmp/corbel ./target/release/corbel-purge scan-dir /path/to/inbox --recursive --workspace /tmp/corbel
``` ```
The directory walker picks up `.pdf`, `.epub`, `.md`, `.markdown`, and `.docx` The directory walker picks up `.pdf`, `.epub`, `.md`, `.markdown`, and
files. Each file is logged with a `[OK]`, `[MALICIOUS]`, or `[ERROR]` tag. `.docx` files. Each file is logged with a `[OK]`, `[MALICIOUS]`, or
`[ERROR]` tag.
## Step 7: CI Integration (Optional) ## Step 8 — CI gate
Use `--abort-on-threat` to make CorbelPurge a CI gate. Exit code 2 means Use `--abort-on-threat` to make CorbelPurge a CI gate. Exit code 2
threats were found: means threats were found:
```bash ```bash
# In your CI pipeline
./target/release/corbel-purge scan incoming_document.pdf --abort-on-threat --quiet ./target/release/corbel-purge scan incoming_document.pdf --abort-on-threat --quiet
# exit 0: clean # exit 0: clean
# exit 2: has threats -> fail the build # exit 2: has threats -> fail the build
# exit 1: hard error (parse failure, IO, etc.) # exit 1: hard error (parse failure, IO, etc.)
``` ```
`--quiet` suppresses the summary block and prints only the JSON report path, `--quiet` suppresses the summary block and prints only the JSON
which is handy for piping into downstream tooling. report path, which is handy for piping into downstream tooling.
## Step 8: Plug In External Threat-Intel Feeds (Optional) ## Step 9 — External threat-intel feeds (optional)
The built-in signature tables and CVE database are static, but you can layer The built-in signature tables and CVE database are static. Layer your
your own on top at runtime: own on top at runtime:
```bash ```bash
# Load additional YARA-style signature rules # Load additional signature rules
./target/release/corbel-purge scan suspicious.pdf --rules my_rules.json ./target/release/corbel-purge scan suspicious.pdf --rules my_rules.json
# Load additional CVE signature entries # Load additional CVE signature entries
@ -123,73 +213,172 @@ export CORBEL_EXTERNAL_CVE_DB=/etc/corbel/cve_db.json
``` ```
External rules are matched alongside the built-in tables; nothing is External rules are matched alongside the built-in tables; nothing is
overridden. See `MANIFEST.md` for the JSON schema. overridden. See `MANIFEST.md` for the JSON schema. Supported
external-rule types:
## Step 9: Build the GUI (Optional) - `homograph-host-list` — additional exact-match homograph host strings
- `signature-list` — additional `(offset, magic_bytes, name)` entries
- `shellcode-list` — additional shellcode prologue byte patterns
The iced 0.13 dashboard GUI is behind the `gui` feature flag: (Legacy `tld-list`, `keyword-list`, and `brand-list` rule types are
silently skipped — the scanner no longer consults those tables.)
## Step 10 — Study a document in place
The `study` subcommand renders the source document to a single
annotated HTML file with malicious regions wrapped in inline `<span>`
tags, color-coded by classification. Use it when you want to see
exactly where the exploit sits in context, without leaving the source
format:
```bash ```bash
cargo build --release --features gui --bin corbel-purge-gui ./target/release/corbel-purge study suspicious.epub
./target/release/corbel-purge-gui
``` ```
The GUI provides file pickers, toggle switches for preserve-format / Output lands at `study_<ts>_<sha>.html` in the workspace. Quiet mode
abort-on-threat / recursive, a timestamped console log, a sidebar with a (`-q`) prints just the path.
Unicode progress gauge and per-file stats, and a cleansed-document viewer.
All scanning runs through the same `Pipeline` the CLI uses, via
`tokio::spawn_blocking`.
## What You Should See ## CLI flags reference
### Clean file output: | Flag | Default | Purpose |
|---|---|---|
| `--workspace <dir>` | CWD | Root for `corbel_quarantine/` and `corbel_clean/` output dirs |
| `--abort-on-threat` | off | Exit with code 2 if any malicious finding fires |
| `--quiet`, `-q` | off | Print only the JSON report path on success |
| `--recursive`, `-r` | off | Recurse into subdirectories (`scan-dir`) |
| `--preserve-format` | off | Repackage cleansed document in original format |
| `--rules <path>` | unset | Path to external signature-rules JSON |
| `--cve-db <path>` | unset | Path to external CVE database JSON |
| `--no-clean-output` | off | Do NOT copy clean documents to output folder |
| `--move-clean` | off | Move (not copy) clean source files to output folder |
## Exit codes
| Code | Meaning |
|---|---|
| `0` | No threats found |
| `1` | Hard error (parse failure, IO failure, etc.) |
| `2` | One or more malicious findings (file was processed) |
## Environment variables
| Variable | Default | Purpose |
|---|---|---|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Quarantine tarball + reports output dir |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Cleansed document + clean-output copy dir |
| `CORBEL_ABORT_ON_THREAT` | `false` | Abort on first malicious finding |
| `CORBEL_EMIT_SUSPICIOUS` | `false` | Include Suspicious findings (no default detector produces any) |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Generate Markdown report alongside JSON |
| `CORBEL_EMIT_CLEAN_OUTPUT` | `true` | Copy clean documents to output folder |
| `CORBEL_MOVE_CLEAN_TO_OUTPUT` | `false` | Move (not copy) clean source files to output |
| `CORBEL_TOTAL_ARCHIVE_SCAN_CAP` | `268435456` (256 MiB) | Cumulative cap across all entries in a multi-entry archive |
| `CORBEL_EXTERNAL_RULES` | (unset) | Path to external signature-rules JSON |
| `CORBEL_EXTERNAL_CVE_DB` | (unset) | Path to external CVE database JSON |
## Expected output examples
### Clean file
``` ```
──────────────────────────────────────────────────────── ────────────────────────────────────────────────────────────
scan complete: PDF scan complete: pdf
source: benign.pdf source: benign.pdf
sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08 sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
text nodes: 12 text nodes: 12
vectors: 0 vectors: 0
findings: 0 malicious, 0 educational, 0 total findings: 0 malicious, 0 educational, 0 total
──────────────────────────────────────────────────────── cleansed: ./corbel_clean/clean_20260801T120000_9f86d081.pdf
json report: ./corbel_quarantine/report_20260801T120000_9f86d081.json
md report: ./corbel_quarantine/report_20260801T120000_9f86d081.md
────────────────────────────────────────────────────────────
``` ```
### Threat found output: ### Threat found
``` ```
──────────────────────────────────────────────────────── ────────────────────────────────────────────────────────────
scan complete: PDF scan complete: pdf
source: suspicious.pdf source: suspicious.pdf
sha256: a1b2c3... sha256: a1b2c3...
text nodes: 8 text nodes: 8
vectors: 1 vectors: 1
findings: 1 malicious, 0 educational, 1 total findings: 1 malicious, 0 educational, 1 total
quarantine: ./corbel_quarantine/quarantine_20260801T120000_abc12345.tar.gz quarantine: ./corbel_quarantine/quarantine_20260801T120000_abc12345.tar.gz
cleansed: ./corbel_clean/cleansed_20260801T120000_abc12345.md cleansed: ./corbel_clean/cleansed_20260801T120000_abc12345.pdf
json report: ./corbel_quarantine/report_20260801T120000_abc12345.json json report: ./corbel_quarantine/report_20260801T120000_abc12345.json
md report: ./corbel_quarantine/report_20260801T120000_abc12345.md md report: ./corbel_quarantine/report_20260801T120000_abc12345.md
──────────────────────────────────────────────────────── ────────────────────────────────────────────────────────────
``` ```
Exit code is `2` when any malicious finding is produced. Exit code is `2` when any malicious finding is produced.
## Key Environment Variables ## Troubleshooting
| Variable | Default | When to Change It | ### Build fails on `lopdf` or `zip`
|----------|---------|-------------------|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Point at a shared quarantine volume |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Point at an output directory for cleaned files |
| `CORBEL_ABORT_ON_THREAT` | `false` | Set to `true` in CI pipelines |
| `CORBEL_EMIT_SUSPICIOUS` | `true` | Set to `false` to only report Malicious (not Suspicious) |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Set to `false` if you only need JSON |
| `CORBEL_TOTAL_ARCHIVE_SCAN_CAP` | `268435456` (256 MiB) | Cumulative cap across all entries in a multi-entry archive |
| `CORBEL_EXTERNAL_RULES` | (unset) | Path to an external signature-rules JSON file |
| `CORBEL_EXTERNAL_CVE_DB` | (unset) | Path to an external CVE database JSON file |
## Next Steps These crates occasionally need a newer Rust than the MSRV declared in
their `Cargo.toml`. Update Rust:
- Read [README.md](README.md) for the full feature overview and security model ```bash
rustup update stable
```
### GUI binary not found
The GUI is behind the `gui` feature flag. Build with:
```bash
cargo build --release --features gui --bin corbel-purge-gui
```
### Scanner reports zero findings on a document I expected to be flagged
Confirm the document actually contains a Category 1 or Category 2
detector trigger (see [README.md](README.md) for the table). The
scanner does not flag based on reputation, file source, or filename.
Run with `cargo run -- scan file.pdf` for verbose output.
### Quarantine tarball missing
The tarball is written only when `malicious_count() > 0`. Clean
documents produce only `report_<ts>_<sha>.{json,md}` and a
`clean_<ts>_<sha>.<ext>` copy in the output folder.
### Test failure on `benign_realworld_pdf_produces_zero_findings`
This is the corpus regression test. If it fails, a detector is
wrong by construction — the detector fired on a clean document. Inspect
the test's failure output for the specific detector that fired and
tighten that detector's rule.
## Installation
After building, install the binary to a system path:
```bash
cargo install --path .
# or
sudo cp target/release/corbel-purge /usr/local/bin/
```
For system-wide configuration, set environment variables in
`/etc/corbel/env` or your shell profile:
```bash
export CORBEL_QUARANTINE_DIR=/var/lib/corbel/quarantine
export CORBEL_CLEANSE_DIR=/var/lib/corbel/clean
export CORBEL_EXTERNAL_RULES=/etc/corbel/rules.json
export CORBEL_EXTERNAL_CVE_DB=/etc/corbel/cve_db.json
```
## Next steps
- Read [README.md](README.md) for the project overview and detection model
- Read [MANIFEST.md](MANIFEST.md) for the detailed technical specification - Read [MANIFEST.md](MANIFEST.md) for the detailed technical specification
- Read [TODO.md](TODO.md) for the development roadmap - Read [TODO.md](TODO.md) for the development roadmap
- Run `cargo test` to verify all 154 tests pass in your environment - Read [FIX-NOTES-false-positive-redesign.md](FIX-NOTES-false-positive-redesign.md)
for the design notes on the two-category detector model
## Contact
**Jeremy Anderson** — [dcos.net](https://dcos.net) — [info@dcos.net](mailto:info@dcos.net)

257
README.md
View File

@ -2,91 +2,147 @@
> Strict Rust document sanitizer & threat neutralizer for PDF, EPUB, Markdown, and DOCX. > Strict Rust document sanitizer & threat neutralizer for PDF, EPUB, Markdown, and DOCX.
**Author:** Jeremy Anderson — [https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel) **Author:** Jeremy Anderson — [dcos.net](https://dcos.net) — [info@dcos.net](mailto:info@dcos.net)
**Repository:** [https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)
**License:** GPL-3.0-or-later **License:** GPL-3.0-or-later
![Corbel-Purge-ss](./corbel-purge-ss.png) ![Corbel-Purge-ss](./corbel-purge-ss.png)
CorbelPurge parses documents into a unified intermediate representation (`Document` struct), runs a layered contextual scanner that distinguishes educational security literature from active malicious injections, and produces cleansed derivatives with all executable content stripped. Malicious payloads are carved into quarantine tarballs with full forensic reports. ## Overview
CorbelPurge is a strict, local-only document security scanner. It parses
documents into a unified intermediate representation, runs detectors
that are backed by verifiable properties of the bytes on disk, and
produces cleansed derivatives with all executable content stripped.
Malicious payloads are carved into quarantine tarballs with full
forensic reports.
The scanner is built around a single design invariant: **every
detector must be backed by a verifiable property, either of the
document itself or of an external authority.** No thresholds, no
per-file or per-domain exceptions, no statistical "suspicious" tier.
## What it does ## What it does
Given a file path, `Pipeline::run()` in `src/core/pipeline.rs`: Given a file path, `Pipeline::run()` in `src/core/pipeline.rs`:
1. **Detects format** via `DocumentFormat::from_path()` (extension-based: `.pdf`, `.epub`, `.md`/`.markdown`, `.docx`). 1. **Detects format** via `DocumentFormat::from_path()` (extension-based:
2. **Parses** via `parsers::Dispatcher` into a `Document` containing `TextNode` (static text with semantic context) and `ExecutableVector` (active content like JS streams, embedded files, script tags, VBA macros) items. `.pdf`, `.epub`, `.md`/`.markdown`, `.docx`).
2. **Parses** via `parsers::Dispatcher` into a `Document` containing
`TextNode` (static text with semantic context) and
`ExecutableVector` (active content like JS streams, embedded files,
script tags, VBA macros) items.
3. **Scans** with a two-pass engine: 3. **Scans** with a two-pass engine:
- `heuristics::inspect_vector()` classifies every executable vector against file-signature tables, shellcode patterns, phishing heuristics, and URI allowlists. - `heuristics::inspect_vector()` classifies every executable vector
- `context_filter::evaluate()` checks text nodes for suspicious signatures and determines whether the surrounding context is educational (code blocks, CVE writeups, academic language) or weaponized. against the two-category detector model: Category 1 (verifiable
executable intent — file signatures, shellcode prologues,
executable URI schemes) and Category 2 (verifiable impersonation —
exact-host homographs, credential URLs, mixed-script hosts).
- `context_filter::evaluate()` checks text nodes for structural
signatures (`/JavaScript`, `<script`, etc.) and weaponization
indicators (long hex runs, base64 blobs, multi-shell commands).
Words and function names that appear in legitimate technical
literature (`wget`, `exploit`, `payload`, `eval(`) are not
treated as signatures.
- `cve_tags::match_cve()` annotates findings with known exploit IDs. - `cve_tags::match_cve()` annotates findings with known exploit IDs.
4. **Quarantines** (when malicious findings exist) — `quarantine::handle()` carves payloads into `quarantine_<ts>_<sha>.tar.gz` with `original.<ext>`, `report.json`, `report.md`, and one `.bin` per payload. 4. **Reports** — both clean and malicious scans produce a JSON and
5. **Cleanses** (when recommended) — produces a sanitized derivative: Markdown report in `corbel_quarantine/` named
- **Markdown mode** (default): `cleanse::sanitizer::sanitize()` emits a safe Markdown file. Text nodes at malicious locations are stripped; hyperlinks lose their destinations. `report_<timestamp>_<sha_prefix>.{json,md}`. The clean path provides
- **PreserveFormat mode** (`--preserve-format`): `cleanse::repackage::repackage()` rebuilds the original format with malicious entries removed. EPUB entries are stripped from the ZIP, DOCX macros/embeddings/external-links are removed, PDF objects are deleted via lopdf. an audit trail; the malicious path adds a quarantine tarball and
carved payloads.
5. **Quarantines** (when malicious findings exist) — `quarantine::handle()`
carves payloads into `quarantine_<ts>_<sha>.tar.gz` with
`original.<ext>`, `report.json`, `report.md`, and one `.bin` per
payload (plus paired `.hex` and `.info` files).
6. **Cleanses** (when recommended) — produces a sanitized derivative:
- **Markdown mode** (default): `cleanse::sanitizer::sanitize()`
emits a safe Markdown file. Text nodes at malicious locations
are stripped; hyperlinks lose their destinations.
- **PreserveFormat mode** (`--preserve-format`):
`cleanse::repackage::repackage()` rebuilds the original format
with malicious entries removed. EPUB entries are stripped from
the ZIP, DOCX macros/embeddings/external-links are removed, PDF
objects are deleted via lopdf.
7. **Clean-output** (when scan is clean and `emit_clean_output` is on,
the default) — copies the source file to
`corbel_clean/clean_<ts>_<sha>.<ext>` so the scanner acts as a
pipeline stage. Use `--move-clean` for queue-draining semantics.
Supported formats: **PDF**, **EPUB**, **Markdown**, **DOCX**. Supported formats: **PDF**, **EPUB**, **Markdown**, **DOCX**.
## Build ## Detection model
```bash Two categories. Nothing else fires.
# Headless CLI (default — no GUI deps)
cargo build --release
# With iced GUI ### Category 1 — Verifiable executable intent
cargo build --release --features gui --bin corbel-purge-gui
The vector contains a structure whose only purpose is to execute
code or spawn a process. Presence is the threat.
| Detector | Triggers on |
|---|---|
| Active script in PDF | `/JavaScript` or `/JS` action stream |
| Program launch in PDF | `/Launch` action with `/F`, `/Win`, `/Mac`, `/Unix` |
| External program exec in EPUB | `<script>` tag in XHTML |
| VBA macro in DOCX | `word/vbaProject.xml` present in package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 |
| Executable URI scheme | `javascript:`, `vbscript:`, `data:text/html`, `file:` in a hyperlink or action |
| PDF form with `/AA` | AcroForm dictionary contains the `/AA` (Additional Actions) entry |
| PDF widget with `/AA` | Widget annotation with `/AA` entry |
| Shellcode prologue | Known Metasploit / NOP-sled / syscall-stub byte sequences |
### Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on |
|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (`micros0ft.com`, `paypa1.com`, etc.) |
| Credential URL | RFC 3986 authority component contains `user:pass@` (before the first `/` after `://`) |
| Mixed-script host | URL host mixes Latin with Cyrillic or Greek (a property of the codepoints) |
### What is NOT a detector
The following were removed because they are statistical guesses
about the world, not verifiable properties of the document:
- **Substring keyword matching** — words like `login`, `verify`,
`support`, `account` are not threats. Every legitimate login page
contains them.
- **Phishing TLDs** — `.ru`, `.cn`, `.xyz` are not threats. A Russian
URL is a Russian URL.
- **URL shorteners as a threat signal** — `bit.ly` is not a threat.
If the destination is hostile, the underlying detector catches it.
- **Substring brand matching** — `https://github.com/microsoft/vscode`
is not a threat. The host is `github.com`; the brand substring in
the path is irrelevant.
- **IP-address hosts** — `192.168.1.1` is a valid network address.
RFCs and router manuals reference them.
- **Substring text-node signatures** — words like `wget`, `exploit`,
`payload`, `eval(`, `powershell`, `/bin/sh` in prose are not
threats. Only structural tokens (`/JavaScript`, `<script`,
`shellcode`) and weaponization indicators (hex runs, base64 blobs,
multi-shell commands) fire.
## Pipeline architecture
```text
input path ──► parser ──► Document (UIR)
│
▼
scanner ──► ScanReport
│
┌────────────────────┼────────────────────┐
│ │ │
▼ ▼ ▼
(clean) (malicious) (educational)
│ │ │
▼ ▼ ▼
clean_<ts>_<sha>.<ext> quarantine tarball whitelisted
+ report.json / .md + carved payloads (logged only)
+ cleansed file
``` ```
Requires Rust 1.70+ (stable).
Binaries:
- `target/release/corbel-purge` — headless CLI
- `target/release/corbel-purge-gui` — iced dashboard GUI (`gui` feature)
## CLI usage
The CLI is in `src/main.rs` and parses its own arguments (no `clap` dependency).
```bash
# Scan a single file
corbel-purge scan path/to/suspicious.pdf
# Scan with format-preserving output (keeps original format)
corbel-purge scan path/to/book.epub --preserve-format
# Scan with a custom workspace
corbel-purge scan path/to/file.pdf --workspace /tmp/corbel
# Scan a directory recursively
corbel-purge scan-dir path/to/inbox --recursive --workspace /tmp/corbel
# CI gate: exit non-zero on threat
corbel-purge scan path/to/file.pdf --abort-on-threat
# Quiet mode: print only the JSON report path
corbel-purge scan path/to/file.pdf --quiet
# Version
corbel-purge --version
```
### Exit codes
| Code | Meaning |
|------|----------|
| 0 | No threats found |
| 1 | Hard error (parse failure, IO, etc.) |
| 2 | One or more malicious findings |
### Environment variables
| Variable | Default | Purpose |
|----------|---------|---------|
| `CORBEL_QUARANTINE_DIR` | `./corbel_quarantine` | Quarantine tarball output dir |
| `CORBEL_CLEANSE_DIR` | `./corbel_clean` | Cleansed document output dir |
| `CORBEL_ABORT_ON_THREAT` | `false` | Abort on first malicious finding |
| `CORBEL_EMIT_SUSPICIOUS` | `true` | Include Suspicious findings in report |
| `CORBEL_EMIT_MARKDOWN_REPORT` | `true` | Generate Markdown report alongside JSON |
## Library usage ## Library usage
```rust ```rust
@ -116,7 +172,7 @@ let result = pipeline.run("book.epub")?;
```text ```text
src/ src/
├── main.rs # CLI: scan, scan-dir, --version, arg parsing ├── main.rs # CLI: scan, scan-dir, study, --version, arg parsing
├── lib.rs # Crate root, CorbelError enum, public re-exports ├── lib.rs # Crate root, CorbelError enum, public re-exports
├── util.rs # sha256_hex(), truncate_with_ellipsis(), read_with_cap() ├── util.rs # sha256_hex(), truncate_with_ellipsis(), read_with_cap()
├── bin/ ├── bin/
@ -135,19 +191,20 @@ src/
│ └── docx_parser.rs # OOXML ZIP inspector (regex XML extraction) │ └── docx_parser.rs # OOXML ZIP inspector (regex XML extraction)
├── scanner/ ├── scanner/
│ ├── mod.rs # scan() entrypoint, CVE tag injection │ ├── mod.rs # scan() entrypoint, CVE tag injection
│ ├── heuristics.rs # classify_vector() for executable vectors │ ├── heuristics.rs # classify_vector() — two-category detector model
│ ├── context_filter.rs # evaluate() for text nodes, educational vs weaponized │ ├── context_filter.rs # evaluate() — structural signatures + weaponization
│ ├── signatures.rs # PHISHING_TLDS, KNOWN_FILE_SIGNATURES, │ ├── signatures.rs # KNOWN_FILE_SIGNATURES, SHELLCODE_PATTERNS,
│ │ # SHELLCODE_PATTERNS, COMMON_PHISHING_BRANDS, │ │ # HOMOGRAPH_HOSTS, URL_SHORTENER_DOMAINS (retained
│ │ # SUSPICIOUS_URL_KEYWORDS, URL_SHORTENER_DOMAINS │ │ # for future reputation work, not consulted)
│ └── cve_tags.rs # CVE_TABLE (7 entries), match_cve(), cve_tag() │ └── cve_tags.rs # CVE_TABLE, match_cve(), cve_tag()
├── quarantine/ ├── quarantine/
│ ├── mod.rs # handle() — tarball writer, QuarantineOutcome │ ├── mod.rs # handle() — tarball writer, write_clean_report()
│ ├── extractor.rs # extract_payloads(), ExtractedPayload, filename sanitization │ ├── extractor.rs # extract_payloads(), ExtractedPayload, filename sanitization
│ ├── hexdump.rs # hex_dump(), build_payload_info()
│ └── reporter.rs # build_json_report(), build_markdown_report() │ └── reporter.rs # build_json_report(), build_markdown_report()
├── cleanse/ ├── cleanse/
│ ├── mod.rs # cleanse() dispatcher (Markdown vs PreserveFormat) │ ├── mod.rs # cleanse() dispatcher (Markdown vs PreserveFormat)
│ ├── sanitizer.rs # sanitize() — Markdown re-serializer, build_clean_markdown() │ ├── sanitizer.rs # sanitize() — Markdown re-serializer
│ └── repackage.rs # repackage() — format-preserving (EPUB/DOCX/PDF) │ └── repackage.rs # repackage() — format-preserving (EPUB/DOCX/PDF)
└── ui/ └── ui/
├── mod.rs # GUI module gate (behind `gui` feature) ├── mod.rs # GUI module gate (behind `gui` feature)
@ -155,29 +212,34 @@ src/
└── alert_modal.rs # Stub └── alert_modal.rs # Stub
``` ```
## Testing ## Design philosophy
```bash 1. **Verifiable properties only.** A detector either proves a finding
# Regenerate test fixtures (requires pypdf, reportlab, python-docx) by a property of the bytes, or it doesn't fire. No thresholds.
pip install pypdf reportlab python-docx 2. **No special-casing.** No per-file or per-domain exceptions. The
python3 scripts/gen_fixtures.py clean-document guarantee is a theorem about the detectors, not an
python3 scripts/gen_md_epub_fixtures.py empirical observation about one specific file.
python3 scripts/gen_docx_fixtures.py 3. **Unix philosophy.** Each module does one thing. The pipeline is a
python3 scripts/gen_zip_bomb_fixtures.py linear sequence of small steps. Step-down logic (early returns) for
every fork of choices.
# Run all tests 4. **`#![forbid(unsafe_code)]`** at the crate root. The scanner never
cargo test touches `unsafe` Rust.
``` 5. **Local-only.** No network access during scanning. All analysis is
against the bytes on disk. External threat-intel feeds are loaded
from local JSON files at startup.
## Adding a new format ## Adding a new format
1. Create `src/parsers/<format>_parser.rs` implementing `DocumentParser`. 1. Create `src/parsers/<format>_parser.rs` implementing `DocumentParser`.
2. Add a variant to `DocumentFormat` in `src/core/types.rs` (and its `Display` impl). 2. Add a variant to `DocumentFormat` in `src/core/types.rs` (and its
`Display` impl).
3. Add a match arm in `Dispatcher::parse()` in `src/parsers/mod.rs`. 3. Add a match arm in `Dispatcher::parse()` in `src/parsers/mod.rs`.
4. Add an extension check in `DocumentFormat::from_path()` in `src/core/pipeline.rs`. 4. Add an extension check in `DocumentFormat::from_path()` in
`src/core/pipeline.rs`.
5. Add repackage support in `src/cleanse/repackage.rs` (optional). 5. Add repackage support in `src/cleanse/repackage.rs` (optional).
The scanner, quarantine, and cleanse modules consume only the `Document` UIR and require no changes for basic format support. The scanner, quarantine, and cleanse modules consume only the
`Document` UIR and require no changes for basic format support.
## Non-goals ## Non-goals
@ -186,6 +248,17 @@ The scanner, quarantine, and cleanse modules consume only the `Document` UIR and
- No network access during scanning — all analysis is local. - No network access during scanning — all analysis is local.
- No sandbox hardening of the tool itself. - No sandbox hardening of the tool itself.
## Documentation
- **[QUICKSTART.md](QUICKSTART.md)** — Build guide and operator reference
- **[MANIFEST.md](MANIFEST.md)** — Detailed technical specification
- **[TODO.md](TODO.md)** — Development roadmap
- **[FIX-NOTES-false-positive-redesign.md](FIX-NOTES-false-positive-redesign.md)** —
Design notes for the two-category detector model
- **[FIX-NOTES-indoc-links.md](FIX-NOTES-indoc-links.md)** — In-document link handling
- **[FIX-NOTES-glyphfix.md](FIX-NOTES-glyphfix.md)** — Glyph rendering fixes
## License ## License
GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson. https://git.dcos.net/dcosnet/corbel GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson.
[https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)

View File

@ -0,0 +1,162 @@
#!/usr/bin/env python3
"""Generate a realistic benign-PDF fixture for the false-positive regression test.
The fixture mimics the structure of a technical book like *Linux from
Scratch*: a multi-page PDF with a table of contents containing
in-document links, body paragraphs containing URLs to kernel.org and
linuxfromscratch.org, mailing-list URLs that contain "support" in the
path, mailto: links to authors, and a page that mentions `wget`,
`exploit`, `payload`, etc. in prose.
This is the regression corpus for the false-positive fix. The integration
test asserts that scanning this PDF produces ZERO findings.
"""
from pathlib import Path
from reportlab.pdfgen import canvas
from reportlab.lib.pagesizes import letter
import pypdf
from pypdf.generic import (
ArrayObject,
DictionaryObject,
NameObject,
NumberObject,
TextStringObject,
)
FIXTURES_DIR = Path(__file__).resolve().parent.parent / "tests" / "fixtures"
FIXTURES_DIR.mkdir(parents=True, exist_ok=True)
def make_benign_realworld_pdf():
"""A multi-page PDF that exercises every URL type that USED TO be a false positive.
Contains:
- http:// URLs to kernel.org mirrors
- https:// URLs to github.com paths that mention brands ("microsoft")
- https:// URLs with "support", "verify", "account" in the path
- mailto: links to authors
- tel: phone links
- In-document cross-reference links (GoTo actions)
- A page that mentions wget / exploit / payload in prose
- A .ru URL (country-code TLD that used to fire PHISHING_TLD)
"""
out_path = FIXTURES_DIR / "benign_realworld.pdf"
# Step 1: draw the pages with reportlab.
base_path = FIXTURES_DIR / "_base_realworld.pdf"
c = canvas.Canvas(str(base_path), pagesize=letter)
c.setTitle("Linux from Scratch (Sample Chapter)")
c.setAuthor("Gerard Beekmans")
c.setSubject("Sample technical-document PDF for false-positive regression test")
# Page 1 — TOC-style page with prose containing URLs.
c.drawString(80, 720, "Chapter 1. Introduction")
c.drawString(80, 700, "See the official site at https://www.linuxfromscratch.org/")
c.drawString(80, 680, "Mailing lists: https://lists.linuxfromscratch.org/listinfo/lfs-support")
c.drawString(80, 660, "Source mirrors: http://ftp.osuosl.org/pub/lfs/")
c.drawString(80, 640, "Patches hosted at https://github.com/LFS-project/build-scripts")
c.drawString(80, 620, "Bug reports: mailto:lfs-support@linuxfromscratch.org")
c.drawString(80, 600, "Kernel sources: https://www.kernel.org/pub/linux/kernel/")
c.drawString(80, 580, "Phone: tel:+1-555-123-4567")
c.showPage()
# Page 2 — prose mentioning common security words.
c.drawString(80, 720, "Chapter 2. Building the System")
c.drawString(80, 700, "Run wget to download the package from the mirror.")
c.drawString(80, 680, "The exploit described in CVE-2024-1234 affects older kernels.")
c.drawString(80, 660, "The attacker's payload is delivered via a crafted document.")
c.drawString(80, 640, "Use /bin/sh as the login shell.")
c.drawString(80, 620, "On Windows, use PowerShell to install the module.")
c.drawString(80, 600, "Russian mirror: https://ftp.ru.debian.org/debian/")
c.drawString(80, 580, "Wikipedia: https://en.wikipedia.org/wiki/Microsoft_Windows")
c.drawString(80, 560, "Github org: https://github.com/microsoft/vscode")
c.showPage()
# Page 3 — page reference with GoTo (in-document link).
c.drawString(80, 720, "Chapter 3. Cross-references")
c.drawString(80, 700, "See Chapter 1 for introduction details.")
c.drawString(80, 680, "External resources:")
c.drawString(80, 660, "- https://www.ietf.org/rfc/rfc2616.txt")
c.drawString(80, 640, "- https://www.w3.org/TR/html5/")
c.drawString(80, 620, "- https://docs.python.org/3/library/")
c.drawString(80, 600, "- https://example.com/account/verify")
c.drawString(80, 580, "- https://example.com/login")
c.showPage()
c.save()
# Step 2: post-process with pypdf to inject URI-action annotations
# on the pages, mimicking real hyperlinks in a published book.
reader = pypdf.PdfReader(str(base_path))
writer = pypdf.PdfWriter()
for page in reader.pages:
writer.add_page(page)
# Hyperlinks to inject — one per page. Each entry is (page_idx, x1, y1, x2, y2, uri).
# The URI action is the structure that triggers the PdfUri vector in
# the parser. We deliberately include the URLs that USED TO be
# false positives.
hyperlinks = [
# Page 0 — TOC links.
(0, 80, 695, 400, 710, "https://www.linuxfromscratch.org/"),
(0, 80, 675, 400, 690, "https://lists.linuxfromscratch.org/listinfo/lfs-support"),
(0, 80, 655, 400, 670, "http://ftp.osuosl.org/pub/lfs/"),
(0, 80, 635, 400, 650, "https://github.com/LFS-project/build-scripts"),
(0, 80, 595, 400, 610, "mailto:lfs-support@linuxfromscratch.org"),
(0, 80, 575, 400, 590, "https://www.kernel.org/pub/linux/kernel/"),
(0, 80, 555, 400, 570, "tel:+1-555-123-4567"),
# Page 1 — body links.
(1, 80, 575, 400, 590, "https://ftp.ru.debian.org/debian/"),
(1, 80, 555, 400, 570, "https://en.wikipedia.org/wiki/Microsoft_Windows"),
(1, 80, 535, 400, 550, "https://github.com/microsoft/vscode"),
# Page 2 — external resources.
(2, 80, 655, 400, 670, "https://www.ietf.org/rfc/rfc2616.txt"),
(2, 80, 615, 400, 630, "https://example.com/account/verify"),
(2, 80, 595, 400, 610, "https://example.com/login"),
]
for page_idx, x1, y1, x2, y2, uri in hyperlinks:
page = writer.pages[page_idx]
# Build the link annotation.
uri_action = DictionaryObject({
NameObject("/Type"): NameObject("/Action"),
NameObject("/S"): NameObject("/URI"),
NameObject("/URI"): TextStringObject(uri),
})
uri_action_ref = writer._add_object(uri_action)
annot = DictionaryObject({
NameObject("/Type"): NameObject("/Annot"),
NameObject("/Subtype"): NameObject("/Link"),
NameObject("/Rect"): ArrayObject([
NumberObject(x1), NumberObject(y1),
NumberObject(x2), NumberObject(y2),
]),
NameObject("/A"): uri_action_ref,
NameObject("/Border"): ArrayObject([
NumberObject(0), NumberObject(0), NumberObject(0),
]),
})
annot_ref = writer._add_object(annot)
if "/Annots" not in page:
page[NameObject("/Annots")] = ArrayObject()
page[NameObject("/Annots")].append(annot_ref)
with open(out_path, "wb") as f:
writer.write(f)
base_path.unlink()
return out_path
def main():
path = make_benign_realworld_pdf()
print(f" wrote {path} ({path.stat().st_size} bytes)")
if __name__ == "__main__":
main()

View File

@ -81,12 +81,56 @@ pub struct Config {
/// If `true`, the scanner will emit `Suspicious` findings for any /// If `true`, the scanner will emit `Suspicious` findings for any
/// executable vector it cannot confidently classify. If `false`, /// executable vector it cannot confidently classify. If `false`,
/// only confidently-malicious vectors are reported. /// only confidently-malicious vectors are reported.
///
/// **Note:** no default detector produces `Suspicious`. Every
/// detector either proves Malicious or doesn't fire (Benign). The
/// flag is retained for API compatibility and for future detectors
/// that may produce genuinely indeterminate signals. Defaults to
/// `false`.
pub emit_suspicious: bool, pub emit_suspicious: bool,
/// List of URI schemes that are considered safe for hyperlink /// URI schemes that are considered safe for hyperlink navigation
/// navigation (e.g. `https`, `mailto`). Anything else is flagged. /// (e.g. `https`, `mailto`). Anything else — `javascript:`,
/// `vbscript:`, `data:text/html`, `file:` — is flagged as Malicious.
///
/// The default list contains the schemes any hyperlink in a normal
/// document would use. HTTP is included because the overwhelming
/// majority of `http://` links in real documents (kernel.org mirrors,
/// list archives, documentation sites that haven't been migrated to
/// HTTPS) are benign; a blanket `http:` ban produced hundreds of
/// false positives on technical PDFs without catching any real
/// threat that wasn't already caught by the executable-scheme check.
pub allowed_uri_schemes: Vec<String>, pub allowed_uri_schemes: Vec<String>,
/// If `true`, generate a Markdown report alongside the JSON report. /// If `true`, generate a Markdown report alongside the JSON report.
pub emit_markdown_report: bool, pub emit_markdown_report: bool,
/// If `true`, copy the original file to the cleanse/output directory
/// when a scan produces zero malicious findings. The copied file is
/// named `clean_<timestamp>_<sha_prefix>.<ext>`.
///
/// This makes the scanner a proper pipeline stage: input files
/// flow through the scanner and clean ones end up in the output
/// folder alongside the cleansed derivatives of malicious ones.
/// Operators running batch jobs (e.g. a watchdog directory) can
/// then chain downstream processing on the contents of the
/// cleanse_dir without having to inspect each file's report.
///
/// Defaults to `true`. Set to `false` to suppress the copy when
/// you only want the reports.
pub emit_clean_output: bool,
/// If `true`, **move** (rather than copy) the source file to the
/// clean-output directory when a scan produces zero malicious
/// findings. The source file is removed from its original location
/// after the copy to the output directory succeeds.
///
/// Useful for pipeline/batch processing where the input directory
/// is a queue and you don't want successfully-scanned files
/// clogging it up. Defaults to `false` because move semantics are
/// destructive — operators who want pipeline-style behavior can
/// flip this on.
///
/// Has no effect when `emit_clean_output` is `false` or when the
/// scan found malicious findings (in which case the source is
/// preserved in the quarantine tarball).
pub move_clean_to_output: bool,
/// Path to an external YARA rules file. If set, the scanner loads /// Path to an external YARA rules file. If set, the scanner loads
/// additional signature rules from this file at startup. /// additional signature rules from this file at startup.
pub external_rules_path: Option<PathBuf>, pub external_rules_path: Option<PathBuf>,
@ -105,13 +149,36 @@ impl Default for Config {
epub_entry_scan_cap: DEFAULT_EPUB_ENTRY_SCAN_CAP, epub_entry_scan_cap: DEFAULT_EPUB_ENTRY_SCAN_CAP,
total_archive_scan_cap: DEFAULT_TOTAL_ARCHIVE_SCAN_CAP, total_archive_scan_cap: DEFAULT_TOTAL_ARCHIVE_SCAN_CAP,
abort_on_threat: false, abort_on_threat: false,
emit_suspicious: true, // emit_suspicious defaults to false: no default detector
// produces `Suspicious`, so the flag's value is immaterial
// for the default scanner.
emit_suspicious: false,
// Allow-list of URI schemes that hyperlinks may use without
// being treated as Malicious(SuspiciousUri). The default set
// covers every hyperlink type that appears in normal
// documents. `http` is included because most technical
// documentation still has plain-HTTP links (kernel.org
// mirrors, listinfo pages, IRC logs) and the
// executable-scheme check (`javascript:`, `vbscript:`,
// `data:text/html`, `file:`) already catches the
// genuinely dangerous schemes.
allowed_uri_schemes: vec![ allowed_uri_schemes: vec![
"https".to_string(), "https".to_string(),
"http".to_string(),
"mailto".to_string(), "mailto".to_string(),
"ftp".to_string(), "ftp".to_string(),
"tel".to_string(),
], ],
emit_markdown_report: true, emit_markdown_report: true,
// Default ON: clean documents get copied to the output
// folder so the scanner acts as a pipeline stage. The
// operator can chain downstream processing on the
// cleanse_dir contents without inspecting each report.
emit_clean_output: true,
// Default OFF: moving the source file is destructive.
// Operators who want true pipeline-queue semantics (input
// folder drains as files are scanned) flip this on.
move_clean_to_output: false,
external_rules_path: None, external_rules_path: None,
external_cve_db_path: None, external_cve_db_path: None,
cleanse_mode: CleanseMode::Markdown, cleanse_mode: CleanseMode::Markdown,
@ -140,6 +207,10 @@ impl Config {
/// - `CORBEL_ABORT_ON_THREAT` (`1`/`true`/`yes` → true) /// - `CORBEL_ABORT_ON_THREAT` (`1`/`true`/`yes` → true)
/// - `CORBEL_EMIT_SUSPICIOUS` /// - `CORBEL_EMIT_SUSPICIOUS`
/// - `CORBEL_EMIT_MARKDOWN_REPORT` /// - `CORBEL_EMIT_MARKDOWN_REPORT`
/// - `CORBEL_EMIT_CLEAN_OUTPUT` (`1`/`true`/`yes` → true;
/// default `true`)
/// - `CORBEL_MOVE_CLEAN_TO_OUTPUT` (`1`/`true`/`yes` → true;
/// default `false`)
/// - `CORBEL_EXTERNAL_RULES` (path to external YARA rules JSON) /// - `CORBEL_EXTERNAL_RULES` (path to external YARA rules JSON)
/// - `CORBEL_EXTERNAL_CVE_DB` (path to external CVE DB JSON) /// - `CORBEL_EXTERNAL_CVE_DB` (path to external CVE DB JSON)
pub fn override_from_env(mut self) -> Self { pub fn override_from_env(mut self) -> Self {
@ -158,6 +229,12 @@ impl Config {
if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") { if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") {
self.emit_markdown_report = truthy(&v); self.emit_markdown_report = truthy(&v);
} }
if let Ok(v) = std::env::var("CORBEL_EMIT_CLEAN_OUTPUT") {
self.emit_clean_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_MOVE_CLEAN_TO_OUTPUT") {
self.move_clean_to_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") { if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") {
if let Ok(cap) = v.parse::<usize>() { if let Ok(cap) = v.parse::<usize>() {
self.total_archive_scan_cap = cap; self.total_archive_scan_cap = cap;
@ -201,8 +278,27 @@ mod tests {
assert!(c.quarantine_dir.ends_with(DEFAULT_QUARANTINE_DIR)); assert!(c.quarantine_dir.ends_with(DEFAULT_QUARANTINE_DIR));
assert!(c.cleanse_dir.ends_with(DEFAULT_CLEANSE_DIR)); assert!(c.cleanse_dir.ends_with(DEFAULT_CLEANSE_DIR));
assert!(!c.abort_on_threat); assert!(!c.abort_on_threat);
assert!(c.emit_suspicious); // emit_suspicious defaults to false: no default detector
// produces `Suspicious`. Every detector either proves Malicious
// or doesn't fire.
assert!(!c.emit_suspicious);
assert!(c.allowed_uri_schemes.contains(&"https".to_string())); assert!(c.allowed_uri_schemes.contains(&"https".to_string()));
// The two schemes added by the false-positive fix must be present,
// otherwise every plain-HTTP URL in technical PDFs is flagged
// Malicious, and phone links (`tel:`) which are universally benign
// get the same treatment.
assert!(c.allowed_uri_schemes.contains(&"http".to_string()));
assert!(c.allowed_uri_schemes.contains(&"tel".to_string()));
// Clean-output defaults: emit_clean_output on (pipeline mode),
// move_clean_to_output off (non-destructive by default).
assert!(
c.emit_clean_output,
"emit_clean_output should default to true for pipeline activity"
);
assert!(
!c.move_clean_to_output,
"move_clean_to_output should default to false — move semantics are destructive"
);
} }
#[test] #[test]

View File

@ -29,7 +29,7 @@ use crate::CorbelError::{self, ThreatDetected};
use crate::{quarantine, scanner, cleanse, CorbelResult}; use crate::{quarantine, scanner, cleanse, CorbelResult};
use super::config::Config; use super::config::Config;
use super::types::{DocumentFormat, ScanReport}; use super::types::{Document, DocumentFormat, ScanReport};
/// The result of running the pipeline on a single document. /// The result of running the pipeline on a single document.
#[derive(Debug, Clone, Serialize, Deserialize)] #[derive(Debug, Clone, Serialize, Deserialize)]
@ -45,6 +45,11 @@ pub struct PipelineResult {
/// Path to the generated quarantine tarball, if any. /// Path to the generated quarantine tarball, if any.
pub quarantine_path: Option<PathBuf>, pub quarantine_path: Option<PathBuf>,
/// Path to the generated cleansed document, if any. /// Path to the generated cleansed document, if any.
///
/// Set when (a) the scan found malicious findings and the cleansed
/// derivative was produced, OR (b) the scan found zero malicious
/// findings and `Config::emit_clean_output` is true (the original
/// file was copied/moved to the output folder as a "clean" file).
pub cleansed_path: Option<PathBuf>, pub cleansed_path: Option<PathBuf>,
/// Path to the JSON forensic report, if any. /// Path to the JSON forensic report, if any.
pub json_report_path: Option<PathBuf>, pub json_report_path: Option<PathBuf>,
@ -106,33 +111,68 @@ impl Pipeline {
// 3. Abort-on-threat short-circuit. // 3. Abort-on-threat short-circuit.
if self.config.abort_on_threat && scan_report.malicious_count() > 0 { if self.config.abort_on_threat && scan_report.malicious_count() > 0 {
return Err(ThreatDetected(format!( let source_label = document
"found {} malicious finding(s) in {}",
scan_report.malicious_count(),
document
.source_path .source_path
.as_ref() .as_ref()
.map(|p| p.display().to_string()) .map_or_else(|| format!("<{format} buffer>"), |p| p.display().to_string());
.unwrap_or_else(|| format!("<{} buffer>", format)), return Err(ThreatDetected(format!(
"found {} malicious finding(s) in {source_label}",
scan_report.malicious_count(),
))); )));
} }
// 4. Quarantine + report (if anything malicious was found). // 4. Quarantine + report.
let quarantine_outcome = if scan_report.malicious_count() > 0 { //
Some(quarantine::handle(&document, &scan_report, &self.config)?) // - If the scan found at least one malicious finding, the full
// quarantine path runs: payloads are carved, a tarball is
// written containing the original file + report + payloads,
// and the JSON + Markdown reports are written as standalone
// files.
// - If the scan found ZERO malicious findings, we still write
// the JSON + Markdown reports to the configured quarantine
// directory — the operator pressed Start, the scan ran, and
// they should see output confirming the document was clean.
// No tarball is written (nothing to quarantine), no payloads
// are carved (no malicious bytes to extract), no cleansed
// file is produced (nothing to cleanse).
let (quarantine_outcome, clean_report_paths) = if scan_report.malicious_count() > 0 {
(Some(quarantine::handle(&document, &scan_report, &self.config)?), None)
} else { } else {
None let (json_path, md_path) =
quarantine::write_clean_report(&document, &scan_report, &self.config)?;
(None, Some((json_path, md_path)))
}; };
// 5. Cleanse (if recommended). // 5. Cleanse / clean-output.
//
// - Malicious scan → produce the cleansed derivative (sanitized
// Markdown or repackaged original format with malicious
// entries stripped).
// - Clean scan → if `emit_clean_output` is set (default: true),
// copy (or move, if `move_clean_to_output`) the original file
// to the cleanse_dir under the name
// `clean_<timestamp>_<sha_prefix>.<ext>`. This makes the
// scanner a proper pipeline stage: input files flow through
// and clean ones end up in the output folder alongside the
// cleansed derivatives of malicious ones.
let cleansed_path = if scan_report.overall_recommendation() let cleansed_path = if scan_report.overall_recommendation()
== super::types::Recommendation::QuarantineAndCleanse == super::types::Recommendation::QuarantineAndCleanse
{ {
Some(cleanse::cleanse(&document, &scan_report, &self.config)?) Some(cleanse::cleanse(&document, &scan_report, &self.config)?)
} else if scan_report.malicious_count() == 0 && self.config.emit_clean_output {
Some(emit_clean_output(&document, &self.config)?)
} else { } else {
None None
}; };
// Unwrap the clean-report paths into the PipelineResult fields.
// `quarantine_outcome` is `Some` only in the malicious case;
// `clean_report_paths` is `Some` only in the clean case.
let (clean_json_path, clean_md_path) = match clean_report_paths {
Some((j, m)) => (Some(j), m),
None => (None, None),
};
Ok(PipelineResult { Ok(PipelineResult {
source_path, source_path,
source_sha256: document.sha256.clone(), source_sha256: document.sha256.clone(),
@ -140,14 +180,68 @@ impl Pipeline {
scan_report, scan_report,
quarantine_path: quarantine_outcome.as_ref().map(|q| q.tarball_path.clone()), quarantine_path: quarantine_outcome.as_ref().map(|q| q.tarball_path.clone()),
cleansed_path, cleansed_path,
json_report_path: quarantine_outcome.as_ref().map(|q| q.json_report_path.clone()), json_report_path: quarantine_outcome
.as_ref()
.map(|q| q.json_report_path.clone())
.or(clean_json_path),
markdown_report_path: quarantine_outcome markdown_report_path: quarantine_outcome
.as_ref() .as_ref()
.and_then(|q| q.markdown_report_path.clone()), .and_then(|q| q.markdown_report_path.clone())
.or(clean_md_path),
}) })
} }
} }
/// Copy (or move, if `Config::move_clean_to_output`) the original file
/// to the cleanse/output directory when a scan produces zero malicious
/// findings.
///
/// The output filename follows the same convention as the cleansed
/// derivative: `clean_<timestamp>_<sha_prefix>.<ext>`. Using the SHA
/// prefix avoids collisions across multiple clean documents and gives
/// each one a stable, traceable name.
///
/// This is the "pipeline stage" behavior: input files flow through
/// the scanner, clean ones end up in the output folder, malicious
/// ones get cleansed derivatives in the same folder. Downstream
/// processing can then operate on the contents of `cleanse_dir`
/// without having to inspect each file's report.
fn emit_clean_output(document: &Document, config: &Config) -> CorbelResult<PathBuf> {
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S");
let sha_prefix = &document.sha256[..8.min(document.sha256.len())];
let ext = match document.format {
DocumentFormat::Pdf => "pdf",
DocumentFormat::Epub => "epub",
DocumentFormat::Markdown => "md",
DocumentFormat::Docx => "docx",
};
let filename = format!("clean_{timestamp}_{sha_prefix}.{ext}");
let output_path = config.cleanse_dir.join(&filename);
// Write the original bytes to the output path. We use the in-memory
// `raw_bytes` rather than re-reading from `source_path` so this
// works even when the pipeline was invoked via `run_on_bytes`
// without a source path on disk.
std::fs::write(&output_path, &document.raw_bytes)?;
// If move-semantics are requested AND we have a source path on
// disk, remove the source file after the copy succeeds. We do
// this only after `std::fs::write` returns Ok, so a failed copy
// never takes the source with it.
if config.move_clean_to_output {
if let Some(src) = document.source_path.as_ref() {
// Best-effort removal — if the unlink fails (e.g.
// permission denied, file locked on Windows), we don't
// want to fail the whole pipeline. The clean copy is
// already in place; the operator can clean up the source
// manually.
let _ = std::fs::remove_file(src);
}
}
Ok(output_path)
}
impl DocumentFormat { impl DocumentFormat {
/// Detect a document's format from its file extension. /// Detect a document's format from its file extension.
/// ///

View File

@ -164,7 +164,14 @@ pub enum VectorType {
PdfLaunch, PdfLaunch,
/// PDF `/URI` action (open a URL — used for phishing). /// PDF `/URI` action (open a URL — used for phishing).
PdfUri, PdfUri,
/// PDF `/GoToR` / `/GoTo` remote navigation. /// PDF `/GoTo` action — in-document navigation (jump to a page
/// or named destination inside the same PDF). Benign by default;
/// the scanner still records it so the operator can see in-document
/// navigation activity, but no alert is raised just for being a link.
PdfGoTo,
/// PDF `/GoToR` action — remote navigation (jump to another PDF
/// file). Treated as suspicious because the destination file is
/// outside the currently-scanned document.
PdfGoToR, PdfGoToR,
/// PDF `/EmbeddedFiles` attachment. /// PDF `/EmbeddedFiles` attachment.
PdfEmbeddedFile, PdfEmbeddedFile,
@ -200,6 +207,7 @@ impl std::fmt::Display for VectorType {
Self::PdfJavaScript => "pdf-javascript", Self::PdfJavaScript => "pdf-javascript",
Self::PdfLaunch => "pdf-launch", Self::PdfLaunch => "pdf-launch",
Self::PdfUri => "pdf-uri", Self::PdfUri => "pdf-uri",
Self::PdfGoTo => "pdf-goto",
Self::PdfGoToR => "pdf-gotor", Self::PdfGoToR => "pdf-gotor",
Self::PdfEmbeddedFile => "pdf-embedded-file", Self::PdfEmbeddedFile => "pdf-embedded-file",
Self::PdfWidgetAction => "pdf-widget-action", Self::PdfWidgetAction => "pdf-widget-action",

View File

@ -15,7 +15,7 @@
//! executables). //! executables).
//! 3. Optionally extracts, reports, and quarantines any identified threats //! 3. Optionally extracts, reports, and quarantines any identified threats
//! ([`crate::quarantine`]). //! ([`crate::quarantine`]).
//! 4. Optionally produces a sanitized, cleansed derivative of the original //! 4. Optionally produces a sanitized, cleansed derivative of the source
//! document containing only validated clean content //! document containing only validated clean content
//! ([`crate::cleanse`]). //! ([`crate::cleanse`]).
//! //!
@ -23,6 +23,10 @@
//! which is intentionally stubbed in this MVP. //! which is intentionally stubbed in this MVP.
//! //!
//! See `MANIFEST.md` in the project root for the full design manifest. //! See `MANIFEST.md` in the project root for the full design manifest.
//!
//! # Author
//!
//! **Jeremy Anderson** — [dcos.net](https://dcos.net) — <info@dcos.net>
#![forbid(unsafe_code)] #![forbid(unsafe_code)]
#![deny(missing_docs)] #![deny(missing_docs)]

View File

@ -53,9 +53,9 @@ fn print_usage() {
USAGE: USAGE:
corbel-purge scan <path> [--workspace <dir>] [--abort-on-threat] [--quiet] [--preserve-format] corbel-purge scan <path> [--workspace <dir>] [--abort-on-threat] [--quiet] [--preserve-format]
[--rules <path>] [--cve-db <path>] [--rules <path>] [--cve-db <path>] [--no-clean-output] [--move-clean]
corbel-purge scan-dir <dir> [--recursive] [--workspace <dir>] [--abort-on-threat] [--preserve-format] corbel-purge scan-dir <dir> [--recursive] [--workspace <dir>] [--abort-on-threat] [--preserve-format]
[--rules <path>] [--cve-db <path>] [--rules <path>] [--cve-db <path>] [--no-clean-output] [--move-clean]
corbel-purge study <path> [--workspace <dir>] corbel-purge study <path> [--workspace <dir>]
corbel-purge --version corbel-purge --version
@ -76,6 +76,24 @@ OPTIONS:
--cve-db <path> Path to an external CVE signature database --cve-db <path> Path to an external CVE signature database
(JSON array). Additional CVE entries are loaded (JSON array). Additional CVE entries are loaded
and matched alongside the built-in CVE table. and matched alongside the built-in CVE table.
--no-clean-output Do NOT copy clean documents to the output folder.
By default, clean documents are copied to
<workspace>/corbel_clean/clean_<ts>_<sha>.<ext>
so the scanner acts as a pipeline stage.
--move-clean MOVE (not copy) the source file to the output
folder when the scan is clean. The source file
is removed from its original location after
the copy succeeds. Useful for pipeline/batch
processing where the input directory is a
queue. Implies clean-output is enabled.
ENVIRONMENT VARIABLES:
CORBEL_QUARANTINE_DIR, CORBEL_CLEANSE_DIR
CORBEL_ABORT_ON_THREAT, CORBEL_EMIT_SUSPICIOUS
CORBEL_EMIT_MARKDOWN_REPORT
CORBEL_EMIT_CLEAN_OUTPUT, CORBEL_MOVE_CLEAN_TO_OUTPUT
CORBEL_TOTAL_ARCHIVE_SCAN_CAP
CORBEL_EXTERNAL_RULES, CORBEL_EXTERNAL_CVE_DB
EXIT CODES: EXIT CODES:
0 No threats found. 0 No threats found.
@ -95,6 +113,14 @@ struct ScanArgs {
preserve_format: bool, preserve_format: bool,
external_rules_path: Option<PathBuf>, external_rules_path: Option<PathBuf>,
external_cve_db_path: Option<PathBuf>, external_cve_db_path: Option<PathBuf>,
/// `--no-clean-output` — disable copying clean documents to the
/// output folder. By default, clean documents ARE copied to
/// `cleanse_dir/clean_<ts>_<sha>.<ext>` (pipeline-stage behavior).
no_clean_output: bool,
/// `--move-clean` — move (rather than copy) the source file to the
/// output folder when the scan is clean. Useful for pipeline/batch
/// processing where the input directory is a queue.
move_clean: bool,
} }
fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> { fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
@ -106,6 +132,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
let mut preserve_format = false; let mut preserve_format = false;
let mut external_rules_path: Option<PathBuf> = None; let mut external_rules_path: Option<PathBuf> = None;
let mut external_cve_db_path: Option<PathBuf> = None; let mut external_cve_db_path: Option<PathBuf> = None;
let mut no_clean_output = false;
let mut move_clean = false;
let mut i = 0; let mut i = 0;
while i < args.len() { while i < args.len() {
@ -132,6 +160,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
args.get(i).ok_or("--cve-db requires a value")?, args.get(i).ok_or("--cve-db requires a value")?,
)); ));
} }
"--no-clean-output" => no_clean_output = true,
"--move-clean" => move_clean = true,
"--help" | "-h" => { "--help" | "-h" => {
print_usage(); print_usage();
std::process::exit(0); std::process::exit(0);
@ -160,6 +190,8 @@ fn parse_scan_args(args: &[String]) -> Result<ScanArgs, String> {
preserve_format, preserve_format,
external_rules_path, external_rules_path,
external_cve_db_path, external_cve_db_path,
no_clean_output,
move_clean,
}) })
} }
@ -178,6 +210,8 @@ fn run_scan(args: &[String]) -> ExitCode {
parsed.preserve_format, parsed.preserve_format,
parsed.external_rules_path, parsed.external_rules_path,
parsed.external_cve_db_path, parsed.external_cve_db_path,
parsed.no_clean_output,
parsed.move_clean,
); );
match pipeline.run(&parsed.path) { match pipeline.run(&parsed.path) {
Ok(result) => { Ok(result) => {
@ -211,6 +245,8 @@ fn run_scan_dir(args: &[String]) -> ExitCode {
parsed.preserve_format, parsed.preserve_format,
parsed.external_rules_path, parsed.external_rules_path,
parsed.external_cve_db_path, parsed.external_cve_db_path,
parsed.no_clean_output,
parsed.move_clean,
); );
let mut found_malicious = false; let mut found_malicious = false;
let mut had_error = false; let mut had_error = false;
@ -302,6 +338,8 @@ fn build_pipeline(
preserve_format: bool, preserve_format: bool,
external_rules_path: Option<PathBuf>, external_rules_path: Option<PathBuf>,
external_cve_db_path: Option<PathBuf>, external_cve_db_path: Option<PathBuf>,
no_clean_output: bool,
move_clean: bool,
) -> Pipeline { ) -> Pipeline {
let mut config = match workspace { let mut config = match workspace {
Some(p) => Config::with_workspace(p), Some(p) => Config::with_workspace(p),
@ -313,6 +351,14 @@ fn build_pipeline(
if preserve_format { if preserve_format {
config.cleanse_mode = corbel_purge::CleanseMode::PreserveFormat; config.cleanse_mode = corbel_purge::CleanseMode::PreserveFormat;
} }
// CLI flags for clean-output behavior. These override the config
// defaults (which are emit_clean_output=true, move_clean_to_output=false).
if no_clean_output {
config.emit_clean_output = false;
}
if move_clean {
config.move_clean_to_output = true;
}
// Apply env overrides last so they win. // Apply env overrides last so they win.
apply_env_overrides(&mut config); apply_env_overrides(&mut config);
@ -347,6 +393,12 @@ fn apply_env_overrides(config: &mut Config) {
if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") { if let Ok(v) = std::env::var("CORBEL_EMIT_MARKDOWN_REPORT") {
config.emit_markdown_report = truthy(&v); config.emit_markdown_report = truthy(&v);
} }
if let Ok(v) = std::env::var("CORBEL_EMIT_CLEAN_OUTPUT") {
config.emit_clean_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_MOVE_CLEAN_TO_OUTPUT") {
config.move_clean_to_output = truthy(&v);
}
if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") { if let Ok(v) = std::env::var("CORBEL_TOTAL_ARCHIVE_SCAN_CAP") {
if let Ok(cap) = v.parse::<usize>() { if let Ok(cap) = v.parse::<usize>() {
config.total_archive_scan_cap = cap; config.total_archive_scan_cap = cap;

View File

@ -280,61 +280,19 @@ fn extract_docx_text(xml: &str, entry_name: &str, text_nodes: &mut Vec<TextNode>
/// Extract external hyperlinks from `word/document.xml`. /// Extract external hyperlinks from `word/document.xml`.
/// ///
/// Hyperlinks look like `<w:hyperlink r:id="rId1">text</w:hyperlink>`. /// Hyperlinks look like `<w:hyperlink r:id="rId1">text</w:hyperlink>`.
/// The actual URL is in the rels file, but we still emit a vector /// The actual URL is in the rels file (`word/_rels/document.xml.rels`),
/// here so the scanner knows there's an external link reference. /// which is parsed separately by [`extract_rels_external_links`].
fn extract_docx_external_links(xml: &str, entry_name: &str, vectors: &mut Vec<ExecutableVector>) { ///
let mut search_from = 0; /// This function emits no vectors. The rels-file extractor is the
while let Some(h_start) = xml[search_from..].find("<w:hyperlink") { /// sole source of DOCX external-link vectors because it carries the
let abs_start = search_from + h_start; /// actual URL as payload — the visible-text representation here does
// Find end of the hyperlink opening tag. /// not. Running URL detectors on visible text breaks scheme
let rest = &xml[abs_start..]; /// extraction (the `"rId=... text=..."` wrapping is not a valid RFC
let tag_end = match rest.find('>') { /// 3986 scheme) and breaks authority extraction (visible text may
Some(p) => abs_start + p + 1, /// contain `@` for unrelated reasons). The rels-file path avoids both
None => break, /// failure modes.
}; fn extract_docx_external_links(_xml: &str, _entry_name: &str, _vectors: &mut Vec<ExecutableVector>) {
// Find the closing </w:hyperlink>. // Intentionally empty. See doc comment above.
let after_tag = &xml[tag_end..];
let h_end = match after_tag.find("</w:hyperlink>") {
Some(p) => tag_end + p,
None => break,
};
// Extract r:id attribute value.
let opening_tag = &xml[abs_start..tag_end];
let rid = extract_attribute(opening_tag, "r:id");
// Extract the visible text of the hyperlink.
let inner = &xml[tag_end..h_end];
let mut visible_text = String::new();
let mut text_search = 0;
while let Some(t_start) = inner[text_search..].find("<w:t") {
let abs_t_start = text_search + t_start;
let after_tag = &inner[abs_t_start..];
let content_start = match after_tag.find('>') {
Some(p) => abs_t_start + p + 1,
None => break,
};
let after_content = &inner[content_start..];
let content_end = match after_content.find("</w:t>") {
Some(p) => content_start + p,
None => break,
};
visible_text.push_str(&inner[content_start..content_end]);
text_search = content_end + 5;
}
vectors.push(ExecutableVector {
location: Location::EpubEntry {
path: entry_name.to_string(),
anchor: rid.clone(),
},
vector_type: VectorType::DocxExternalLink,
raw_payload: visible_text.as_bytes().to_vec(),
decoded_preview: Some(format!("rId={} text={}", rid.unwrap_or_default(), visible_text)),
});
search_from = h_end + 14; // length of "</w:hyperlink>"
}
} }
/// Extract external relationships from `word/_rels/document.xml.rels`. /// Extract external relationships from `word/_rels/document.xml.rels`.
@ -342,6 +300,11 @@ fn extract_docx_external_links(xml: &str, entry_name: &str, vectors: &mut Vec<Ex
/// Each `<Relationship>` element has attributes `Id`, `Target`, and /// Each `<Relationship>` element has attributes `Id`, `Target`, and
/// `TargetMode`. If `TargetMode="External"`, the relationship points /// `TargetMode`. If `TargetMode="External"`, the relationship points
/// to an external URL — emit a vector with the URL as the payload. /// to an external URL — emit a vector with the URL as the payload.
///
/// This is the canonical source of DOCX external-link vectors: one
/// vector per external hyperlink, with the URL as `raw_payload` and
/// `decoded_preview`. The visible-text-extracting function above no
/// longer emits a duplicate vector.
fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) { fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
let mut search_from = 0; let mut search_from = 0;
while let Some(rel_start) = xml[search_from..].find("<Relationship") { while let Some(rel_start) = xml[search_from..].find("<Relationship") {
@ -360,6 +323,11 @@ fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
let id = extract_attribute(opening_tag, "Id").unwrap_or_default(); let id = extract_attribute(opening_tag, "Id").unwrap_or_default();
if target_mode.as_deref() == Some("External") { if target_mode.as_deref() == Some("External") {
// The decoded_preview is the URL itself (NOT
// `"rId={} target={}"`). The URL detector needs a clean
// URL string to run the scheme/host checks against;
// wrapping it in `"rId=... target=..."` broke both the
// scheme extraction and the authority extraction.
vectors.push(ExecutableVector { vectors.push(ExecutableVector {
location: Location::EpubEntry { location: Location::EpubEntry {
path: "word/_rels/document.xml.rels".to_string(), path: "word/_rels/document.xml.rels".to_string(),
@ -367,7 +335,7 @@ fn extract_rels_external_links(xml: &str, vectors: &mut Vec<ExecutableVector>) {
}, },
vector_type: VectorType::DocxExternalLink, vector_type: VectorType::DocxExternalLink,
raw_payload: target.as_bytes().to_vec(), raw_payload: target.as_bytes().to_vec(),
decoded_preview: Some(format!("rId={} target={}", id, target)), decoded_preview: Some(target.clone()),
}); });
} }

View File

@ -212,20 +212,21 @@ fn inspect_dictionary(
inspect_annots(annots_ref, obj_id, doc, vectors); inspect_annots(annots_ref, obj_id, doc, vectors);
} }
// /AcroForm with /AA // /AcroForm — only flag when /AA (Additional Actions) is present.
// /NeedAppearances is a benign rendering hint present in essentially
// every PDF form. Forms with only /NeedAppearances do not execute
// any code.
if let Ok(acroform_ref) = dict.get(b"AcroForm") { if let Ok(acroform_ref) = dict.get(b"AcroForm") {
if let Ok((_id, acroform_obj)) = doc.dereference(acroform_ref) { if let Ok((_id, acroform_obj)) = doc.dereference(acroform_ref) {
if let Object::Dictionary(acro_dict) = acroform_obj { if let Object::Dictionary(acro_dict) = acroform_obj {
if acro_dict.has(b"AA") || acro_dict.has(b"NeedAppearances") { if acro_dict.has(b"AA") {
vectors.push(ExecutableVector { vectors.push(ExecutableVector {
location: Location::PdfObject { id: obj_id.0, gen: 0 }, location: Location::PdfObject { id: obj_id.0, gen: 0 },
vector_type: VectorType::PdfAcroForm, vector_type: VectorType::PdfAcroForm,
raw_payload: Vec::new(), raw_payload: Vec::new(),
decoded_preview: Some(format!( decoded_preview: Some(
"<AcroForm with AA={} NeedAppearances={}>", "<AcroForm with /AA (Additional Actions)>".to_string(),
acro_dict.has(b"AA"), ),
acro_dict.has(b"NeedAppearances")
)),
}); });
} }
} }
@ -344,13 +345,28 @@ fn inspect_action(action_ref: &Object, obj_id: &ObjectId, doc: &LopdfDocument, v
.into(), .into(),
}); });
} }
"GoToR" | "GoTo" => { "GoTo" => {
// Remote / local navigation — flag but lower priority. // In-document navigation — jump to a page or named
// destination inside the SAME PDF. This is benign (think
// "click here to jump to section 3"), so the scanner will
// classify it as Benign. We still record it so the operator
// can see in-document navigation activity in the report.
vectors.push(ExecutableVector {
location: loc,
vector_type: VectorType::PdfGoTo,
raw_payload: Vec::new(),
decoded_preview: Some(format!("<GoTo action in obj {loc_id}>")),
});
}
"GoToR" => {
// Remote navigation — jump to another PDF file. Treated as
// Suspicious because the destination file is outside the
// currently-scanned document and we can't inspect it.
vectors.push(ExecutableVector { vectors.push(ExecutableVector {
location: loc, location: loc,
vector_type: VectorType::PdfGoToR, vector_type: VectorType::PdfGoToR,
raw_payload: Vec::new(), raw_payload: Vec::new(),
decoded_preview: Some(format!("<GoTo/GoToR action in obj {loc_id}>")), decoded_preview: Some(format!("<GoToR action in obj {loc_id}>")),
}); });
} }
_ => { _ => {

View File

@ -13,7 +13,7 @@ pub mod extractor;
pub mod hexdump; pub mod hexdump;
pub mod reporter; pub mod reporter;
use std::path::PathBuf; use std::path::{Path, PathBuf};
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
@ -42,99 +42,145 @@ pub fn handle(
scan_report: &ScanReport, scan_report: &ScanReport,
config: &Config, config: &Config,
) -> CorbelResult<QuarantineOutcome> { ) -> CorbelResult<QuarantineOutcome> {
// 1. Carve payloads out of the document. // 1. Carve payloads out of the document. Pass `config` through so
// // payload_path is rooted at the caller's quarantine_dir; the
// We pass `config` through so that `payload_path` is rooted at the // no-config helper uses a default Config whose quarantine_dir is
// caller's quarantine_dir. The no-config `extract_payloads` helper // the relative `corbel_quarantine/` (broken under tempdir workspaces).
// would fall back to a default Config whose `quarantine_dir` is the
// relative path `corbel_quarantine/` — fine in production (where
// CWD == workspace) but broken in tests (where the workspace is a
// tempdir) and any other embedded use case.
let extracted = extractor::extract_payloads_with_config(document, scan_report, config); let extracted = extractor::extract_payloads_with_config(document, scan_report, config);
// 2. Generate reports. // 2. Generate reports.
let markdown_report = config
.emit_markdown_report
.then(|| reporter::build_markdown_report(document, scan_report, &extracted));
let json_report = reporter::build_json_report(document, scan_report, &extracted); let json_report = reporter::build_json_report(document, scan_report, &extracted);
let markdown_report = if config.emit_markdown_report {
Some(reporter::build_markdown_report(document, scan_report, &extracted))
} else {
None
};
// 3. Compose the quarantine tarball name. // 3. Compose filenames. SHA prefix guarantees uniqueness across documents.
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S"); let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S").to_string();
let sha_prefix = &document.sha256[..8.min(document.sha256.len())]; let sha_pref = sha_prefix(document);
let tarball_name = format!("quarantine_{timestamp}_{sha_prefix}.tar.gz"); let paths = ReportPaths::new(
let tarball_path = config.quarantine_dir.join(&tarball_name); &config.quarantine_dir,
&timestamp,
&sha_pref,
markdown_report.is_some(),
);
let json_report_name = format!("report_{timestamp}_{sha_prefix}.json"); // 4. Write the tarball + standalone JSON report.
let json_report_path = config.quarantine_dir.join(&json_report_name);
let markdown_report_path = if markdown_report.is_some() {
Some(config.quarantine_dir.join(format!(
"report_{timestamp}_{sha_prefix}.md"
)))
} else {
None
};
// 4. Write the tarball.
write_tarball( write_tarball(
&tarball_path, &paths.tarball_path,
document, document,
&json_report, &json_report,
markdown_report.as_deref(), markdown_report.as_deref(),
&extracted, &extracted,
)?; )?;
std::fs::write(&paths.json_path, serde_json::to_string_pretty(&json_report)?)?;
// 5. Write the JSON report as a standalone file (for easy programmatic access). // 5. Write the Markdown report (only when enabled).
std::fs::write(&json_report_path, serde_json::to_string_pretty(&json_report)?)?; if let (Some(md), Some(md_path)) = (&markdown_report, &paths.markdown_path) {
// 6. Write the Markdown report as a standalone file too.
if let Some(md) = &markdown_report {
if let Some(md_path) = &markdown_report_path {
std::fs::write(md_path, md)?; std::fs::write(md_path, md)?;
} }
// 6. Write standalone .hex + .info files for each carved payload.
extracted
.iter()
.try_for_each(|payload| write_payload_artifacts(payload, scan_report))?;
Ok(QuarantineOutcome {
tarball_path: paths.tarball_path,
json_report_path: paths.json_path,
markdown_report_path: paths.markdown_path,
extracted_payload_paths: extracted.iter().map(|p| p.payload_path.clone()).collect(),
})
}
/// Composed report-file paths under `quarantine_dir`.
struct ReportPaths {
tarball_path: PathBuf,
json_path: PathBuf,
markdown_path: Option<PathBuf>,
}
impl ReportPaths {
/// Construct from `quarantine_dir` using the timestamp + sha-prefix
/// naming convention shared by both the malicious and clean paths.
fn new(quarantine_dir: &Path, timestamp: &str, sha_prefix: &str, emit_markdown: bool) -> Self {
let tarball_path = quarantine_dir.join(format!("quarantine_{timestamp}_{sha_prefix}.tar.gz"));
let json_path = quarantine_dir.join(format!("report_{timestamp}_{sha_prefix}.json"));
let markdown_path = emit_markdown
.then(|| quarantine_dir.join(format!("report_{timestamp}_{sha_prefix}.md")));
Self { tarball_path, json_path, markdown_path }
} }
}
// 7. Write standalone payload carving v2 files (.hex + .info). /// SHA-256 prefix for filename uniqueness (first 8 hex chars).
for payload in &extracted { fn sha_prefix(document: &Document) -> String {
// .hex — annotated hex dump document.sha256[..8.min(document.sha256.len())].to_string()
}
/// Write the .hex (annotated dump) and .info (JSON metadata) files
/// for one carved payload.
fn write_payload_artifacts(
payload: &extractor::ExtractedPayload,
scan_report: &ScanReport,
) -> CorbelResult<()> {
// .hex — annotated hex dump.
let hex_content = hexdump::hex_dump(&payload.bytes); let hex_content = hexdump::hex_dump(&payload.bytes);
let hex_path = payload.payload_path.with_extension("hex"); std::fs::write(payload.payload_path.with_extension("hex"), hex_content)?;
std::fs::write(&hex_path, hex_content)?;
// .info — JSON metadata // .info — JSON metadata.
let finding = scan_report let info = build_payload_info(payload, scan_report);
std::fs::write(
payload.payload_path.with_extension("info"),
serde_json::to_string_pretty(&info)?,
)?;
Ok(())
}
/// Build the JSON metadata object for one carved payload.
fn build_payload_info(
payload: &extractor::ExtractedPayload,
scan_report: &ScanReport,
) -> serde_json::Value {
use crate::core::types::{Recommendation, ThreatClassification};
// Lookup table for classification → JSON string.
let classification_str = scan_report
.findings .findings
.get(payload.source_finding_index); .get(payload.source_finding_index)
let (classification_str, recommendation_str, context_notes, vector_type_str, cve_tag_str) = .map(|f| match &f.classification {
if let Some(f) = finding { ThreatClassification::Benign => "benign".to_string(),
ThreatClassification::EducationalContent => "educational".to_string(),
ThreatClassification::Suspicious => "suspicious".to_string(),
ThreatClassification::Malicious(t) => format!("malicious:{t}"),
})
.unwrap_or_default();
// Lookup table for recommendation → JSON string.
let recommendation_str = scan_report
.findings
.get(payload.source_finding_index)
.map(|f| match f.recommendation {
Recommendation::Allow => "allow".to_string(),
Recommendation::WhitelistAsEducational => "whitelist-as-educational".to_string(),
Recommendation::Quarantine => "quarantine".to_string(),
Recommendation::QuarantineAndCleanse => "quarantine-and-cleanse".to_string(),
})
.unwrap_or_default();
let (context_notes, vector_type_str, cve_tag_str) = scan_report
.findings
.get(payload.source_finding_index)
.map(|f| {
( (
match &f.classification {
crate::core::types::ThreatClassification::Benign => "benign".to_string(),
crate::core::types::ThreatClassification::EducationalContent => "educational".to_string(),
crate::core::types::ThreatClassification::Suspicious => "suspicious".to_string(),
crate::core::types::ThreatClassification::Malicious(t) => format!("malicious:{t}"),
},
match f.recommendation {
crate::core::types::Recommendation::Allow => "allow".to_string(),
crate::core::types::Recommendation::WhitelistAsEducational => "whitelist-as-educational".to_string(),
crate::core::types::Recommendation::Quarantine => "quarantine".to_string(),
crate::core::types::Recommendation::QuarantineAndCleanse => "quarantine-and-cleanse".to_string(),
},
f.context_notes.clone(), f.context_notes.clone(),
f.vector_type.map(|v| v.to_string()), f.vector_type.map(|v| v.to_string()),
extract_cve_tag_from_notes(&f.context_notes), extract_cve_tag_from_notes(&f.context_notes),
) )
} else { })
(String::new(), String::new(), String::new(), None, None) .unwrap_or((String::new(), None, None));
};
let file_sig = let file_sig = crate::scanner::signatures::match_file_signature(&payload.bytes).map(|s| s.to_string());
crate::scanner::signatures::match_file_signature(&payload.bytes)
.map(|s| s.to_string());
let info = hexdump::build_payload_info( hexdump::build_payload_info(
&payload.filename, &payload.filename,
payload.source_finding_index, payload.source_finding_index,
vector_type_str.as_deref(), vector_type_str.as_deref(),
@ -146,20 +192,54 @@ pub fn handle(
payload.bytes.len(), payload.bytes.len(),
file_sig.as_deref(), file_sig.as_deref(),
cve_tag_str.as_deref(), cve_tag_str.as_deref(),
)
}
/// Write a "clean bill of health" report for a document that produced
/// no malicious findings.
///
/// Called by the pipeline when `scan_report.malicious_count() == 0`.
/// Writes the JSON and Markdown reports to the configured quarantine
/// directory using the same `report_<timestamp>_<sha_prefix>.{json,md}`
/// naming convention as the malicious path, but with an empty
/// findings array. No tarball is written (nothing to quarantine), no
/// payloads are carved (no malicious bytes to extract), no cleansed
/// file is produced (nothing to cleanse).
///
/// The operator receives a report confirming the scan ran, what was
/// scanned, when it was scanned, and that the document was clean —
/// the audit trail required for pipeline activity.
pub fn write_clean_report(
document: &Document,
scan_report: &ScanReport,
config: &Config,
) -> CorbelResult<(PathBuf, Option<PathBuf>)> {
// Empty extracted-payloads list — clean documents have no payloads
// to extract, but the report builder takes the parameter.
let extracted: Vec<extractor::ExtractedPayload> = Vec::new();
let markdown_report = config
.emit_markdown_report
.then(|| reporter::build_markdown_report(document, scan_report, &extracted));
let json_report = reporter::build_json_report(document, scan_report, &extracted);
// Reuse the shared path-composition helper for naming consistency
// with the malicious path.
let timestamp = chrono::Utc::now().format("%Y%m%dT%H%M%S").to_string();
let sha_pref = sha_prefix(document);
let paths = ReportPaths::new(
&config.quarantine_dir,
&timestamp,
&sha_pref,
markdown_report.is_some(),
); );
let info_path = payload.payload_path.with_extension("info");
std::fs::write(&info_path, serde_json::to_string_pretty(&info)?)?; // Write the JSON report unconditionally; write Markdown only when enabled.
std::fs::write(&paths.json_path, serde_json::to_string_pretty(&json_report)?)?;
if let (Some(md), Some(md_path)) = (&markdown_report, &paths.markdown_path) {
std::fs::write(md_path, md)?;
} }
Ok(QuarantineOutcome { Ok((paths.json_path, paths.markdown_path))
tarball_path,
json_report_path,
markdown_report_path,
extracted_payload_paths: extracted
.iter()
.map(|p| p.payload_path.clone())
.collect(),
})
} }
/// Extract a CVE tag (e.g. `[CVE-2017-11882: Equation Editor RCE]`) /// Extract a CVE tag (e.g. `[CVE-2017-11882: Equation Editor RCE]`)

View File

@ -1,15 +1,26 @@
//! Context filter: NLP / lexical checks for distinguishing security //! # Context filter: structural-signature checks for text nodes.
//! literature from active malicious content.
//! //!
//! When the scanner encounters a suspicious signature inside a *static //! The text-node scanner checks for exactly two things:
//! text node* (paragraph, heading, code block), it asks the context
//! filter whether the surrounding context looks like:
//! //!
//! - **Educational content** (CVE writeups, exploit code samples in //! 1. **Structural signatures** indicating the text contains an
//! defensive blog posts, textbook material) → whitelisted. //! executable hook embedded in the text content itself — e.g. a
//! - **Weaponized content** (obfuscated shellcode, packed executables //! PDF `/JavaScript` operator appearing in a content stream, or an
//! in non-code contexts, embedded action triggers) → flagged. //! HTML `<script>` tag in EPUB XHTML. These are verifiable
//! - **Indeterminate** → emitted as `Suspicious` if configured. //! structural tokens; presence is the threat.
//!
//! 2. **Weaponization indicators** — long runs of hex-encoded bytes
//! (`\xNN` × 16+), long base64 blobs (64+ consecutive base64
//! chars), and multiple concatenated shell commands. These are
//! byte-level patterns with no legitimate use in document text
//! outside of code blocks / blockquotes.
//!
//! ## Design invariant
//!
//! Words and function names that appear in legitimate technical
//! literature (`wget`, `exploit`, `payload`, `eval(`, `exec(`,
//! `powershell`, `/bin/sh`, etc.) are not treated as signatures.
//! The detector only fires on structural tokens whose presence in
//! document text outside of a code block is itself the threat.
use crate::core::config::Config; use crate::core::config::Config;
use crate::core::types::{ use crate::core::types::{
@ -17,19 +28,18 @@ use crate::core::types::{
}; };
/// Evaluate a single text node. Returns `Some(Finding)` if the node /// Evaluate a single text node. Returns `Some(Finding)` if the node
/// contains a suspicious or malicious signature that survived the /// contains a structural signature or weaponization indicator.
/// context filter.
pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> { pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
// Step 0: Check for weaponization indicators first — these are // Step 1: Check for weaponization indicators — long hex runs,
// malicious regardless of whether a "signature" is present. // long base64 blobs, multiple shell commands in non-code context.
// Pure shellcode blobs, for example, contain no recognizable // These are byte-level patterns that are verifiably hostile when
// keyword but are still dangerous. // they appear outside a code block / blockquote.
if has_weaponization_indicators(&node.content) { if has_weaponization_indicators(&node.content) {
let classification = if node.context == TextContext::CodeBlock let classification = if node.context == TextContext::CodeBlock
|| node.context == TextContext::CodeSpan || node.context == TextContext::CodeSpan
|| node.context == TextContext::BlockQuote || node.context == TextContext::BlockQuote
{ {
// Even weaponized-looking content inside a code block / // Weaponized-looking content inside a code block or
// blockquote is treated as educational — it's almost // blockquote is treated as educational — it's almost
// certainly a research writeup illustrating an attack. // certainly a research writeup illustrating an attack.
ThreatClassification::EducationalContent ThreatClassification::EducationalContent
@ -57,27 +67,28 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
}); });
} }
// Step 1: Does the node contain any suspicious signatures? // Step 2: Check for structural signatures — PDF operators, HTML
// tags, shellcode tokens. These indicate the text itself contains
// an executable hook, which is hostile outside of code context.
let signatures = find_signatures(&node.content); let signatures = find_signatures(&node.content);
if signatures.is_empty() { if signatures.is_empty() {
return None; return None;
} }
// Step 2: What's the surrounding context? // Step 3: Decide based on context.
let is_educational = looks_educational(node); let is_educational = looks_educational(node);
// Step 3: Decision matrix.
let classification = if is_educational { let classification = if is_educational {
ThreatClassification::EducationalContent ThreatClassification::EducationalContent
} else if node.context == TextContext::ExecutableHook { } else if node.context == TextContext::ExecutableHook {
// A suspicious signature inside an executable hook is malicious. // A structural signature inside an executable hook is
// definitively malicious — it's the actual payload.
ThreatClassification::Malicious(MaliciousType::ActiveJavaScriptInjection) ThreatClassification::Malicious(MaliciousType::ActiveJavaScriptInjection)
} else if has_weaponization_indicators(&node.content) {
ThreatClassification::Malicious(MaliciousType::ObfuscatedShellcode)
} else if config.emit_suspicious { } else if config.emit_suspicious {
// Retained for API compatibility. No default detector produces
// `Suspicious`; `emit_suspicious` defaults to `false`. Future
// detectors with genuinely indeterminate signals may use this tier.
ThreatClassification::Suspicious ThreatClassification::Suspicious
} else { } else {
// Suspicious findings suppressed — drop it.
return None; return None;
}; };
@ -90,7 +101,7 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
let preview: String = node.content.chars().take(config.max_payload_preview_len).collect(); let preview: String = node.content.chars().take(config.max_payload_preview_len).collect();
let notes = format!( let notes = format!(
"found {} suspicious signature(s) [{}] in {} context{}", "found {} structural signature(s) [{}] in {} context{}",
signatures.len(), signatures.len(),
signatures.join(", "), signatures.join(", "),
context_name(node.context), context_name(node.context),
@ -107,42 +118,37 @@ pub fn evaluate(node: &TextNode, config: &Config) -> Option<Finding> {
}) })
} }
/// A signature is a string that *could* indicate malicious content /// Structural signatures that indicate the text contains an executable
/// but is also commonly found in security literature. /// hook. These are tokens whose presence in document text (outside of
const SUSPICIOUS_SIGNATURES: &[&str] = &[ /// a code block) is a strong signal of embedded active content.
///
/// **Matched as case-sensitive substrings** for the PDF operators
/// (which are case-sensitive in the PDF spec) and case-insensitive
/// substrings for HTML tags. Word-boundary matching is not necessary
/// because these tokens are sufficiently specific — `/JavaScript`
/// does not appear in prose, only in PDF dictionaries.
const STRUCTURAL_SIGNATURES: &[&str] = &[
// PDF action / object operators. These only appear in PDF
// content streams or PDF dictionary dumps — never in prose.
"/JavaScript", "/JavaScript",
"/JS", "/JS",
"/Launch", "/Launch",
"/EmbeddedFile", "/EmbeddedFile",
"eval(", // HTML / XHTML executable tags. Only appear in markup.
"Function(",
"document.write",
"innerHTML",
"<script", "<script",
"<iframe", "<iframe",
// Shellcode token. The literal word "shellcode" is sometimes
// used in prose, but in combination with the weaponization
// detector above this is reserved for cases where the text
// contains the actual word in an executable context.
"shellcode", "shellcode",
"exploit",
"payload",
"calc.exe",
"/bin/sh",
"powershell",
"cmd.exe",
"wget",
"curl http",
"rm -rf",
"Base64.decode",
"atob(",
"exec(",
]; ];
/// Find all suspicious signatures present in `text`. Returns the /// Find all structural signatures present in `text`. Returns the
/// list of signatures found (deduplicated, in source order). /// list of signatures found (deduplicated, in source order).
fn find_signatures(text: &str) -> Vec<&'static str> { fn find_signatures(text: &str) -> Vec<&'static str> {
// Case-insensitive matching for some signatures, exact for others.
// For simplicity, we do case-sensitive matching first and let the
// context filter handle false positives.
let lower = text.to_ascii_lowercase(); let lower = text.to_ascii_lowercase();
SUSPICIOUS_SIGNATURES STRUCTURAL_SIGNATURES
.iter() .iter()
.copied() .copied()
.filter(|sig| { .filter(|sig| {
@ -214,13 +220,16 @@ fn has_weaponization_indicators(text: &str) -> bool {
} }
// Long base64 blob: 64+ consecutive base64 chars anywhere in text. // Long base64 blob: 64+ consecutive base64 chars anywhere in text.
// (Not just whole lines — attackers like to embed base64 inline.)
let b64_re = regex::Regex::new(r"[A-Za-z0-9+/=]{64,}").unwrap(); let b64_re = regex::Regex::new(r"[A-Za-z0-9+/=]{64,}").unwrap();
if b64_re.is_match(text) { if b64_re.is_match(text) {
return true; return true;
} }
// Multiple shell commands in a single non-code node. // Multiple shell commands in a single non-code node. The threshold
// of 2 different commands is high enough that legitimate prose
// mentioning one shell tool ("run `wget` to download the package")
// does not fire — it requires two different shell-command tokens
// in the same node.
let shell_indicators = ["rm -rf", "wget ", "curl ", "nc -", "/bin/sh", "powershell "]; let shell_indicators = ["rm -rf", "wget ", "curl ", "nc -", "/bin/sh", "powershell "];
let count = shell_indicators.iter().filter(|s| text.contains(*s)).count(); let count = shell_indicators.iter().filter(|s| text.contains(*s)).count();
if count >= 2 { if count >= 2 {
@ -265,12 +274,31 @@ mod tests {
#[test] #[test]
fn cve_writeup_in_code_block_is_educational() { fn cve_writeup_in_code_block_is_educational() {
// A code block that contains a structural signature (`<script`)
// alongside a CVE marker is treated as educational — it's a
// research writeup illustrating an attack, not an active payload.
let node = make_node(
TextContext::CodeBlock,
"<script>alert(1)</script> // PoC for CVE-2024-1234",
);
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::EducationalContent);
}
#[test]
fn cve_writeup_prose_without_signature_is_no_finding() {
// Text that mentions `eval()` in prose but does not contain a
// structural signature (`/JavaScript`, `<script`, etc.) is not
// a finding. Words and function names in prose are not
// signatures; only structural tokens fire.
let node = make_node( let node = make_node(
TextContext::CodeBlock, TextContext::CodeBlock,
"eval('alert(1)') // PoC for CVE-2024-1234", "eval('alert(1)') // PoC for CVE-2024-1234",
); );
let finding = evaluate(&node, &Config::default()).unwrap(); assert!(
assert_eq!(finding.classification, ThreatClassification::EducationalContent); evaluate(&node, &Config::default()).is_none(),
"prose mentioning eval() without a structural signature must NOT be flagged"
);
} }
#[test] #[test]
@ -297,38 +325,11 @@ mod tests {
)); ));
} }
#[test]
fn suspicious_in_paragraph_with_signature() {
let node = make_node(TextContext::Paragraph, "Run eval('alert(1)') now");
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::Suspicious);
}
#[test]
fn suspicious_can_be_suppressed() {
let node = make_node(TextContext::Paragraph, "Run eval('alert(1)') now");
let mut config = Config::default();
config.emit_suspicious = false;
assert!(evaluate(&node, &config).is_none());
}
#[test]
fn academic_text_with_signature_is_educational() {
let node = make_node(
TextContext::Paragraph,
"In this paper we describe the eval() vulnerability and its remediation.",
);
let finding = evaluate(&node, &Config::default()).unwrap();
assert_eq!(finding.classification, ThreatClassification::EducationalContent);
}
#[test] #[test]
fn long_base64_blob_is_weaponized() { fn long_base64_blob_is_weaponized() {
let b64 = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/ABCDEFGH"; let b64 = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/ABCDEFGH";
let node = make_node(TextContext::Paragraph, b64); let node = make_node(TextContext::Paragraph, b64);
let finding = evaluate(&node, &Config::default()); let finding = evaluate(&node, &Config::default());
// With the step-0 weaponization check, pure base64 (no signature)
// is now flagged as Malicious (ObfuscatedShellcode).
let finding = finding.expect("pure base64 blob should be flagged as weaponized"); let finding = finding.expect("pure base64 blob should be flagged as weaponized");
assert!(matches!( assert!(matches!(
finding.classification, finding.classification,
@ -346,4 +347,81 @@ mod tests {
ThreatClassification::Malicious(_) ThreatClassification::Malicious(_)
)); ));
} }
// ─── Anti-false-positive tests ─────────────────────────────
//
// These are the texts that USED TO fire the substring-signature
// detector and that appear in legitimate technical documents.
// All of them must now produce no finding.
#[test]
fn wget_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"Run wget to download the package from the mirror.",
);
assert!(
evaluate(&node, &Config::default()).is_none(),
"prose mentioning wget must NOT be flagged"
);
}
#[test]
fn exploit_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The exploit described in this chapter targets a buffer overflow.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn payload_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The attacker's payload is delivered via a crafted document.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn single_shell_command_in_prose_is_not_a_finding() {
// The weaponization detector requires 2+ different shell
// commands. A single mention of `wget` in prose is not enough.
let node = make_node(
TextContext::Paragraph,
"Use `wget http://example.com/file.tar.gz` to fetch the archive.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn powershell_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"On Windows, the equivalent command uses PowerShell to install the module.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn bin_sh_mentioned_in_prose_is_not_a_finding() {
let node = make_node(
TextContext::Paragraph,
"The shebang line `#!/bin/sh` indicates the script runs under the Bourne shell.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
#[test]
fn eval_mentioned_in_prose_is_not_a_finding() {
// `eval(` is a JavaScript function name; appearing in prose is
// not a threat. Only structural tokens (`/JavaScript`, `<script`,
// etc.) fire the text-node scanner.
let node = make_node(
TextContext::Paragraph,
"The eval() function in JavaScript executes a string as code.",
);
assert!(evaluate(&node, &Config::default()).is_none());
}
} }

File diff suppressed because it is too large Load Diff

View File

@ -17,43 +17,45 @@ use crate::core::types::{Document, ScanReport};
/// This is the top-level entrypoint called by [`crate::core::pipeline::Pipeline`]. /// This is the top-level entrypoint called by [`crate::core::pipeline::Pipeline`].
#[must_use] #[must_use]
pub fn scan(document: &Document, config: &Config) -> ScanReport { pub fn scan(document: &Document, config: &Config) -> ScanReport {
let mut findings = Vec::new(); // Walk executable vectors first (untrusted-by-default), then text
// nodes (context-filtered). Findings are tagged with a CVE entry
// when the payload matches a known exploit signature.
let vector_findings = document.executable_vectors.iter().filter_map(|vector| {
heuristics::inspect_vector(vector, config).map(|mut finding| {
tag_with_cve_if_known(vector, &mut finding);
finding
})
});
// 1. Walk executable vectors — these are untrusted-by-default. let text_findings = document
for vector in &document.executable_vectors { .text_nodes
if let Some(mut finding) = heuristics::inspect_vector(vector, config) { .iter()
// After classification, try to tag the finding with a .filter_map(|node| context_filter::evaluate(node, config));
// known CVE if the payload matches a known exploit signature.
if let Some(cve) = cve_tags::match_cve(vector, &finding) {
finding.context_notes = format!(
"{} [{}: {}]",
finding.context_notes, cve.cve_id, cve.name
);
}
findings.push(finding);
}
}
// 2. Walk text nodes — these go through the context filter to let findings: Vec<_> = vector_findings.chain(text_findings).collect();
// distinguish educational content from active threats.
for node in &document.text_nodes {
if let Some(finding) = context_filter::evaluate(node, config) {
findings.push(finding);
}
}
let scanned_at = chrono::Utc::now().to_rfc3339();
ScanReport { ScanReport {
source_sha256: document.sha256.clone(), source_sha256: document.sha256.clone(),
format: document.format, format: document.format,
scanned_at, scanned_at: chrono::Utc::now().to_rfc3339(),
findings, findings,
text_nodes_scanned: document.text_nodes.len(), text_nodes_scanned: document.text_nodes.len(),
vectors_scanned: document.executable_vectors.len(), vectors_scanned: document.executable_vectors.len(),
} }
} }
/// Tag a finding with a known CVE entry when the payload matches a
/// known exploit signature. Mutates `finding.context_notes` in place.
fn tag_with_cve_if_known(
vector: &crate::core::types::ExecutableVector,
finding: &mut crate::core::types::Finding,
) {
let Some(cve) = cve_tags::match_cve(vector, finding) else {
return;
};
finding.context_notes = format!("{} [{}: {}]", finding.context_notes, cve.cve_id, cve.name);
}
/// Re-export for callers that want to inspect individual findings. /// Re-export for callers that want to inspect individual findings.
pub use heuristics::inspect_vector; pub use heuristics::inspect_vector;
pub use context_filter::evaluate; pub use context_filter::evaluate;

View File

@ -1,32 +1,30 @@
//! Threat signature tables and pattern matchers. //! # Threat signature tables and pattern matchers.
//! //!
//! This module centralizes the static lookup tables used by the //! This module centralizes the static lookup tables used by the
//! heuristics engine. Keeping them in one place makes them easy to //! heuristics engine. It contains two kinds of tables:
//! audit, extend, and eventually wire up to an external threat-intel
//! feed (e.g. a YARA rules file or a STIX/TAXII subscription).
//! //!
//! ## What lives here //! - **Category 1 — Verifiable executable intent.** Magic-byte
//! signatures for executable / high-risk file formats
//! ([`KNOWN_FILE_SIGNATURES`]) and known shellcode prologue
//! sequences ([`SHELLCODE_PATTERNS`]). These are deterministic
//! byte-level matches; presence of the structure is the threat.
//! - **Category 2 — Verifiable impersonation.** A small table of
//! known homograph host strings ([`HOMOGRAPH_HOSTS`]) — exact-match
//! only, never substring. A URL whose authority is byte-equal to one
//! of these strings after case-folding is flagged.
//! //!
//! - [`PHISHING_TLDS`] — TLDs statistically overrepresented in //! The scanner does not consult substring keyword tables, TLD lists,
//! phishing URLs. Sourced from public phishing reports //! URL-shortener lists as threat signals, or brand-substring tables.
//! (Spamhaus, PhishTank yearly summaries). //! Those detectors are statistical guesses about the world, not
//! - [`SUSPICIOUS_URL_KEYWORDS`] — path/host keywords that strongly //! verifiable properties of the document, and are not part of the
//! indicate credential harvesting or fake login pages. //! detection model. [`URL_SHORTENER_DOMAINS`] is retained for future
//! - [`URL_SHORTENER_DOMAINS`] — shortener domains. Not malicious //! reputation-feed work but is not consulted by any match function.
//! per se, but a common obfuscation layer for phishing links.
//! - [`KNOWN_FILE_SIGNATURES`] — magic-byte signatures for executable
//! and high-risk file formats (PE, ELF, Mach-O, OLE2, RTF, etc.).
//! - [`SHELLCODE_PATTERNS`] — known shellcode prologue byte sequences
//! (NOP sleds, syscall stubs, common encoders).
//! - [`COMMON_PHISHING_BRANDS`] — brand names frequently spoofed in
//! phishing URLs (microsoft, paypal, appleid, …).
//! //!
//! ## External threat-intel feeds //! ## External threat-intel feeds
//! //!
//! Additional rules can be loaded at runtime via //! Additional rules can be loaded at runtime via [`load_external_rules`].
//! [`load_external_rules`]. Loaded rules are stored in a //! Loaded rules are stored in a process-wide static and consulted by
//! process-wide static and checked by every match function //! the match functions alongside the built-in tables.
//! alongside the built-in tables.
use std::path::Path; use std::path::Path;
use std::sync::OnceLock; use std::sync::OnceLock;
@ -35,42 +33,13 @@ use serde::Deserialize;
use crate::CorbelResult; use crate::CorbelResult;
/// TLDs statistically overrepresented in phishing URLs. /// Common URL-shortener domains.
/// ///
/// Source: synthesized from public yearly phishing reports /// **Not used as a threat signal by the scanner.** Retained for
/// (Spamhaus, PhishTank, Interisle). This list is intentionally /// future reputation-feed integration — if a shortener URL resolves
/// conservative — inclusion requires the TLD to appear in multiple /// to a Category 1 or Category 2 hit, the underlying detector catches
/// reports as a top-10 phishing TLD. /// it. Flagging `bit.ly` itself as a threat produced false positives
pub const PHISHING_TLDS: &[&str] = &[ /// on every document that used a shortener for a legitimate link.
// High-risk TLDs (cheap registration, low verification)
".zip", ".mov", ".xyz", ".top", ".click", ".link", ".rest", ".cyou",
".sbs", ".online", ".live", ".buzz", ".surf", ".monster", ".fit",
".loan", ".win", ".download", ".stream", ".review", ".men",
".work", ".racing", ".party", ".trade", ".science", ".kim",
".cricket", ".gq", ".cf", ".tk", ".ml", ".ga",
// Country-code TLDs frequently abused for phishing
".ru", ".cn", ".su", ".country", ".kim",
// Newer TLDs that have been flagged
".quest", ".bond", ".ha", ".cyou", ".quest", ".beauty",
];
/// URL path / host keywords that strongly suggest credential harvesting
/// or fake login pages. Matched case-insensitively as substrings.
pub const SUSPICIOUS_URL_KEYWORDS: &[&str] = &[
"login", "signin", "sign-in", "log-in", "verify", "verification",
"account", "update", "confirm", "secure", "security", "wallet",
"unlock", "recover", "reactivate", "validate", "activate",
"webscr", "cmd=", "_session", "authorization", "authenticate",
"reset", "password", "credential", "billing", "invoice",
"support", "suspended", "limited", "alert", "warning",
"urgent", "important-notice", "tax", "refund", "irs",
"postbank", "amzn", "appleid", "icloud", "office365",
];
/// Common URL-shortener domains. Shortened URLs are not malicious
/// per se, but they hide the real destination — we flag them as
/// `Suspicious` so the operator can preview the destination before
/// clicking.
pub const URL_SHORTENER_DOMAINS: &[&str] = &[ pub const URL_SHORTENER_DOMAINS: &[&str] = &[
"bit.ly", "t.co", "tinyurl.com", "goo.gl", "ow.ly", "is.gd", "bit.ly", "t.co", "tinyurl.com", "goo.gl", "ow.ly", "is.gd",
"buff.ly", "rebrand.ly", "cutt.ly", "shorturl.at", "tiny.cc", "buff.ly", "rebrand.ly", "cutt.ly", "shorturl.at", "tiny.cc",
@ -79,54 +48,59 @@ pub const URL_SHORTENER_DOMAINS: &[&str] = &[
"surl.li", "kutt.it", "urlzs.com", "shrtco.de", "surl.li", "kutt.it", "urlzs.com", "shrtco.de",
]; ];
/// Brand names frequently spoofed in phishing URLs. Used to detect /// Homograph host strings — known-bad authority strings that mimic a
/// homograph attacks (e.g. `micros0ft.com`, `paypa1.com`). /// legitimate brand by substituting visually similar characters.
/// ///
/// This list intentionally includes BOTH canonical spellings ("microsoft") /// **Matching rule:** a URL's authority (host[:port]) is matched
/// AND known homograph variants ("micros0ft" with zero instead of 'o'). /// byte-equal against these strings after ASCII case-folding. There is
/// The matcher uses a canonical-domain check to suppress benign /// no substring match, no path/query involvement, no canonical-brand
/// matches: when a brand is mentioned, we look for the canonical /// suppression — the URL's host is either exactly one of these strings
/// spelling followed by a TLD; if found, we don't flag. /// or it is not.
pub const COMMON_PHISHING_BRANDS: &[&str] = &[ ///
// Microsoft family /// Examples:
"microsoft", "micros0ft", "micros0fte", "micr0soft", /// - `https://micros0ft.com/login` → host `micros0ft.com` is in the
"msn", "windows", "wind0ws", "office", "0ffice", "outlook", /// list → flagged.
"outl00k", "outl0ok", "live", "1ive", /// - `https://github.com/microsoft/vscode` → host `github.com` is not
/// in the list → not flagged (the "microsoft" in the path is
/// irrelevant).
/// - `https://microsoft.com/windows` → host `microsoft.com` is not in
/// the list → not flagged (legitimate).
/// - `https://MICROS0FT.COM/` → ASCII-case-folded to `micros0ft.com`
/// → flagged.
pub const HOMOGRAPH_HOSTS: &[&str] = &[
// Microsoft family — 'o' → '0', 'i' → '1', etc.
"micros0ft.com", "micr0soft.com", "micros0fte.com",
"msn0.com", "wind0ws.com", "0ffice.com",
"outl00k.com", "outl0ok.com", "1ive.com",
// PayPal // PayPal
"paypal", "paypa1", "paypaI", "paypa|", "paypa1.com", "paypaI.com", "paypa|.com",
// Apple // Apple
"apple", "app1e", "appie", "icloud", "ic1oud", "appleid", "app1e.com", "appie.com", "ic1oud.com", "app1eid.com",
"app1eid",
// Google // Google
"google", "g00gle", "goog1e", "gmail", "gmai", "g00gle.com", "goog1e.com", "gmai.com",
// Amazon // Amazon
"amazon", "amzn", "amaz0n", "a-m-a-z-o-n", "amaz0n.com",
// Social // Social
"facebook", "faceb00k", "facebo0k", "instagram", "instagrarn", "faceb00k.com", "facebo0k.com", "instagrarn.com",
"twitter", "tw1tter", "twtter", "linkedin", "1inkedin", "tw1tter.com", "twtter.com", "1inkedin.com",
// Streaming // Streaming
"netflix", "netf1ix", "spotify", "spot1fy", "netf1ix.com", "spot1fy.com",
// Storage / SaaS // Storage / SaaS
"dropbox", "dr0pbox", "adobe", "ad0be", "dr0pbox.com", "ad0be.com",
// Banking // Banking
"bankofamerica", "bofa", "b0fa", "wellsfargo", "wellsfarg0", "b0fa.com", "wellsfarg0.com",
"chase", "citibank", "citi", "hsbc", "barclays",
"santander", "unicredit",
// Crypto // Crypto
"binance", "binanc3", "coinbase", "c0inbase", "metamask", "binanc3.com", "c0inbase.com", "metam4sk.com",
"metam4sk", "ledger", "trezor",
// Shipping
"dhl", "fedex", "f3dex", "ups", "usps", "royalmail",
// Gaming // Gaming
"steamcommunity", "steampowered", "epicgames", "playstation", "steamp0wered.com", "steamc0mmunity.com",
"nintendo", "xbox",
]; ];
/// Magic-byte signatures for executable and high-risk file formats. /// Magic-byte signatures for executable and high-risk file formats.
/// ///
/// Each entry is (offset, magic_bytes, name). When a payload's bytes /// Each entry is (offset, magic_bytes, name). When a payload's bytes
/// at `offset` match `magic_bytes`, the payload is considered /// at `offset` match `magic_bytes`, the payload is considered
/// executable / high-risk. /// executable / high-risk. This is the Category 1 detector —
/// presence of the structure is the threat.
pub const KNOWN_FILE_SIGNATURES: &[(usize, &[u8], &str)] = &[ pub const KNOWN_FILE_SIGNATURES: &[(usize, &[u8], &str)] = &[
// Windows PE // Windows PE
(0, b"MZ", "pe"), (0, b"MZ", "pe"),
@ -148,8 +122,6 @@ pub const KNOWN_FILE_SIGNATURES: &[(usize, &[u8], &str)] = &[
(0, b"{\\rtf", "rtf"), (0, b"{\\rtf", "rtf"),
// Java class file // Java class file
(0, b"\xCA\xFE\xBA\xBE", "java-class"), (0, b"\xCA\xFE\xBA\xBE", "java-class"),
// Java JAR (zip, but flag if inside PDF)
// (zip is too generic — we don't flag it without other signals)
// Python bytecode // Python bytecode
(0, b"\x42\x0d\x0d\x0a", "python-bytecode"), (0, b"\x42\x0d\x0d\x0a", "python-bytecode"),
// SWF (Flash — historically a huge attack surface) // SWF (Flash — historically a huge attack surface)
@ -202,14 +174,18 @@ pub const SHELLCODE_PATTERNS: &[&[u8]] = &[
/// A single external rule deserialized from a JSON threat-intel feed. /// A single external rule deserialized from a JSON threat-intel feed.
#[derive(Debug, Clone, Deserialize)] #[derive(Debug, Clone, Deserialize)]
pub struct ExternalRule { pub struct ExternalRule {
/// Human-readable rule name (e.g. `"custom-phishing-tlds"`). /// Human-readable rule name (e.g. `"custom-shellcode"`).
pub name: String, pub name: String,
/// Rule type: one of `"tld-list"`, `"keyword-list"`, /// Rule type: one of `"homograph-host-list"`,
/// `"brand-list"`, `"signature-list"`, `"shellcode-list"`. /// `"signature-list"`, `"shellcode-list"`.
///
/// The previous `"tld-list"`, `"keyword-list"`, and `"brand-list"`
/// types were removed when those tables were removed from the
/// scanner. External feeds of those types are no longer loaded.
#[serde(rename = "type")] #[serde(rename = "type")]
pub rule_type: String, pub rule_type: String,
/// Rule values. The shape depends on `rule_type`: /// Rule values. The shape depends on `rule_type`:
/// - `"tld-list"` / `"keyword-list"` / `"brand-list"` → array of strings /// - `"homograph-host-list"` → array of strings (full hostnames)
/// - `"signature-list"` → array of `{"offset": usize, "bytes": [u8], "name": str}` /// - `"signature-list"` → array of `{"offset": usize, "bytes": [u8], "name": str}`
/// - `"shellcode-list"` → array of hex strings or `[u8]` arrays /// - `"shellcode-list"` → array of hex strings or `[u8]` arrays
pub values: serde_json::Value, pub values: serde_json::Value,
@ -226,12 +202,9 @@ struct SignatureEntry {
/// Processed external rules, organized by type for fast matching. /// Processed external rules, organized by type for fast matching.
#[derive(Debug, Clone, Default)] #[derive(Debug, Clone, Default)]
pub(crate) struct ExternalRulesData { pub(crate) struct ExternalRulesData {
/// Additional TLD strings. /// Additional homograph host strings (exact-match, byte-equal
pub(crate) tlds: Vec<String>, /// after ASCII case-folding).
/// Additional suspicious URL keywords. pub(crate) homograph_hosts: Vec<String>,
pub(crate) keywords: Vec<String>,
/// Additional brand strings (may include homograph variants).
pub(crate) brands: Vec<String>,
/// Additional file-signature entries: (offset, magic-bytes, name). /// Additional file-signature entries: (offset, magic-bytes, name).
pub(crate) signatures: Vec<(usize, Vec<u8>, String)>, pub(crate) signatures: Vec<(usize, Vec<u8>, String)>,
/// Additional shellcode byte patterns. /// Additional shellcode byte patterns.
@ -249,8 +222,9 @@ static EXTERNAL_RULES_STORAGE: OnceLock<ExternalRulesData> = OnceLock::new();
/// The file must contain a JSON array of [`ExternalRule`] objects. /// The file must contain a JSON array of [`ExternalRule`] objects.
/// Rules are sorted into type-specific buckets and stored in a /// Rules are sorted into type-specific buckets and stored in a
/// process-wide static ([`EXTERNAL_RULES_STORAGE`]). Subsequent calls /// process-wide static ([`EXTERNAL_RULES_STORAGE`]). Subsequent calls
/// to the match functions (`match_file_signature`, `has_phishing_tld`, /// to the match functions (`match_file_signature`,
/// etc.) will check both the built-in tables and the external rules. /// `match_homograph_host`, etc.) will check both the built-in tables
/// and the external rules.
/// ///
/// # Errors /// # Errors
/// ///
@ -263,70 +237,43 @@ pub fn load_external_rules(path: &Path) -> CorbelResult<Vec<ExternalRule>> {
let mut storage = ExternalRulesData::default(); let mut storage = ExternalRulesData::default();
for rule in &rules { for rule in &rules {
// Each rule contributes zero or more entries to one of the
// storage buckets. Step-down: unknown rule types contribute
// nothing and exit the iteration silently.
match rule.rule_type.as_str() { match rule.rule_type.as_str() {
"tld-list" => { "homograph-host-list" => {
if let Some(arr) = rule.values.as_array() { storage.homograph_hosts.extend(
for v in arr { rule.values
if let Some(s) = v.as_str() { .as_array()
storage.tlds.push(s.to_string()); .into_iter()
} .flatten()
} .filter_map(|v| v.as_str())
} .map(|s| s.to_ascii_lowercase()),
} );
"keyword-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(s) = v.as_str() {
storage.keywords.push(s.to_string());
}
}
}
}
"brand-list" => {
if let Some(arr) = rule.values.as_array() {
for v in arr {
if let Some(s) = v.as_str() {
storage.brands.push(s.to_string());
}
}
}
} }
"signature-list" => { "signature-list" => {
if let Some(arr) = rule.values.as_array() { storage.signatures.extend(
for v in arr { rule.values
if let Ok(sig) = serde_json::from_value::<SignatureEntry>(v.clone()) { .as_array()
storage .into_iter()
.signatures .flatten()
.push((sig.offset, sig.bytes, sig.name)); .filter_map(|v| serde_json::from_value::<SignatureEntry>(v.clone()).ok())
} .map(|sig| (sig.offset, sig.bytes, sig.name)),
} );
}
} }
"shellcode-list" => { "shellcode-list" => {
if let Some(arr) = rule.values.as_array() { storage.shellcode.extend(
for v in arr { rule.values
if let Some(hex_str) = v.as_str() { .as_array()
// Hex-encoded string: "fc4883e4..." .into_iter()
if let Ok(bytes) = hex::decode(hex_str) { .flatten()
if !bytes.is_empty() { .filter_map(|v| decode_shellcode_entry(v)),
storage.shellcode.push(bytes); );
}
}
} else if let Some(byte_arr) = v.as_array() {
// Raw byte array: [0xfc, 0x48, ...]
let bytes: Vec<u8> = byte_arr
.iter()
.filter_map(|b| b.as_u64().map(|n| n as u8))
.collect();
if !bytes.is_empty() {
storage.shellcode.push(bytes);
}
}
}
}
} }
_ => { _ => {
// Unknown rule type — silently skip. // Unknown rule type — silently skip. Includes the
// legacy "tld-list", "keyword-list", "brand-list"
// types that the scanner no longer consults.
} }
} }
} }
@ -335,6 +282,25 @@ pub fn load_external_rules(path: &Path) -> CorbelResult<Vec<ExternalRule>> {
Ok(rules) Ok(rules)
} }
/// Decode a single shellcode-list entry as raw bytes.
///
/// Accepts either a hex-encoded string (`"fc4883e4..."`) or a JSON
/// array of byte values (`[252, 72, ...]`). Returns `None` for empty
/// results or unparseable values.
fn decode_shellcode_entry(v: &serde_json::Value) -> Option<Vec<u8>> {
if let Some(hex_str) = v.as_str() {
hex::decode(hex_str).ok().filter(|b| !b.is_empty())
} else if let Some(byte_arr) = v.as_array() {
let bytes: Vec<u8> = byte_arr
.iter()
.filter_map(|b| b.as_u64().map(|n| n as u8))
.collect();
(!bytes.is_empty()).then_some(bytes)
} else {
None
}
}
/// Return a reference to the externally loaded rules storage, if any /// Return a reference to the externally loaded rules storage, if any
/// has been loaded via [`load_external_rules`]. /// has been loaded via [`load_external_rules`].
/// ///
@ -356,23 +322,19 @@ pub(crate) fn get_external_rules_storage() -> Option<&'static ExternalRulesData>
/// [`load_external_rules`]. /// [`load_external_rules`].
#[must_use] #[must_use]
pub fn match_file_signature(bytes: &[u8]) -> Option<&'static str> { pub fn match_file_signature(bytes: &[u8]) -> Option<&'static str> {
for (offset, magic, name) in KNOWN_FILE_SIGNATURES { // Built-in signatures — first match wins.
if bytes.len() >= *offset + magic.len() { let builtin = KNOWN_FILE_SIGNATURES.iter().find_map(|(offset, magic, name)| {
if &bytes[*offset..*offset + magic.len()] == *magic { bytes_matches_at(bytes, *offset, magic).then_some(*name)
return Some(name); });
}
} // External signatures — only consulted when no built-in matched.
} builtin.or_else(|| {
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() { EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
for (offset, magic, name) in &ext.signatures { ext.signatures.iter().find_map(|(offset, magic, name)| {
if bytes.len() >= *offset + magic.len() { bytes_matches_at(bytes, *offset, magic).then(|| name.as_str())
if &bytes[*offset..*offset + magic.len()] == magic.as_slice() { })
return Some(name); })
} })
}
}
}
None
} }
/// Check whether `bytes` contains any known shellcode prologue. /// Check whether `bytes` contains any known shellcode prologue.
@ -381,188 +343,75 @@ pub fn match_file_signature(bytes: &[u8]) -> Option<&'static str> {
/// [`load_external_rules`]. /// [`load_external_rules`].
#[must_use] #[must_use]
pub fn match_shellcode_pattern(bytes: &[u8]) -> Option<&'static [u8]> { pub fn match_shellcode_pattern(bytes: &[u8]) -> Option<&'static [u8]> {
for pattern in SHELLCODE_PATTERNS { // Built-in patterns — return the first matching prologue.
if bytes.windows(pattern.len()).any(|w| w == *pattern) { let builtin = SHELLCODE_PATTERNS
return Some(pattern);
}
}
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for pattern in &ext.shellcode {
if bytes.windows(pattern.len()).any(|w| w == pattern.as_slice()) {
return Some(pattern.as_slice());
}
}
}
None
}
/// Check whether `host` (lowercase, no scheme) is a known URL shortener.
#[must_use]
pub fn is_url_shortener(host: &str) -> bool {
URL_SHORTENER_DOMAINS.iter().any(|d| host == *d || host.ends_with(&format!(".{d}")))
}
/// Check whether `uri` mentions a commonly-phished brand with a
/// non-canonical domain (homograph bait).
///
/// Returns the matched brand name if found.
///
/// Logic:
/// 1. If a homograph variant (`micros0ft`, `paypa1`, ...) is found
/// anywhere in the URI, it's always a phishing signal.
/// 2. If a canonical spelling (`microsoft`, `paypal`, ...) is found,
/// we check whether the host portion is ANY brand's canonical
/// domain (e.g. `microsoft.com` for "microsoft"). If yes → benign.
/// Otherwise, the brand is mentioned in a non-canonical context
/// → suspicious.
///
/// Checks both the built-in table and any external brands loaded via
/// [`load_external_rules`].
#[must_use]
pub fn match_phishing_brand(uri: &str) -> Option<&'static str> {
let lower = uri.to_ascii_lowercase();
// Extract host (after scheme://, before path/query/fragment, sans port).
let host = lower
.split("://")
.nth(1)
.unwrap_or(&lower)
.split('/')
.next()
.unwrap_or("")
.split(':')
.next()
.unwrap_or("");
let is_homograph = |brand: &str| brand.chars().any(|c| !c.is_ascii_alphabetic());
// First pass: check for homograph variants — these are ALWAYS phishing.
// Built-in brands.
for brand in COMMON_PHISHING_BRANDS {
if is_homograph(brand) && lower.contains(brand) {
return Some(brand);
}
}
// External brands.
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for brand in &ext.brands {
if is_homograph(brand) && lower.contains(brand.as_str()) {
return Some(brand);
}
}
}
// Second pass: check canonical spellings. If the host is ANY
// brand's canonical domain, all canonical brand mentions are
// treated as benign. This handles cases like `microsoft.com/windows`
// (windows is a brand, but the host is microsoft's canonical domain).
let host_is_canonical_for_builtin = COMMON_PHISHING_BRANDS
.iter()
.any(|brand| !is_homograph(brand) && is_canonical_brand_host(host, brand));
let host_is_canonical_for_external = EXTERNAL_RULES_STORAGE
.get()
.map(|ext| {
ext.brands
.iter()
.any(|brand| !is_homograph(brand) && is_canonical_brand_host(host, brand))
})
.unwrap_or(false);
if host_is_canonical_for_builtin || host_is_canonical_for_external {
return None;
}
// Host is not a canonical brand domain — any canonical brand
// mentioned in the URL is suspicious.
// Built-in brands.
for brand in COMMON_PHISHING_BRANDS {
if !is_homograph(brand) && lower.contains(brand) {
return Some(brand);
}
}
// External brands.
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() {
for brand in &ext.brands {
if !is_homograph(brand) && lower.contains(brand.as_str()) {
return Some(brand);
}
}
}
None
}
/// Check whether `host` is the canonical domain for `brand`.
///
/// A host is canonical if it matches `<brand>.<tld>` or has `<brand>`
/// as a dot-separated segment (e.g. `microsoft.com`, `login.microsoft.com`).
/// The goal is to allow legitimate brand-owned domains while still
/// flagging `login-microsoft.com` (which is NOT a Microsoft domain).
fn is_canonical_brand_host(host: &str, brand: &str) -> bool {
if host == brand {
return true;
}
// Check if `<brand>.<tld>` is a prefix.
let canonical_prefix = format!("{}.", brand);
if host.starts_with(&canonical_prefix) {
return true;
}
// Check if `<brand>` is a dot-separated segment (e.g. `login.microsoft.com`).
host.split('.').any(|seg| seg == brand)
}
/// Check whether `uri`'s path/host contains any suspicious keyword.
///
/// Checks both the built-in table and any external rules loaded via
/// [`load_external_rules`].
#[must_use]
pub fn match_suspicious_keyword(uri: &str) -> Option<&'static str> {
let lower = uri.to_ascii_lowercase();
if let Some(kw) = SUSPICIOUS_URL_KEYWORDS
.iter() .iter()
.copied() .copied()
.find(|kw| lower.contains(kw)) .find(|pattern| bytes.windows(pattern.len()).any(|w| w == *pattern));
{
return Some(kw); // External patterns — only consulted when no built-in matched.
} builtin.or_else(|| {
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() { EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
if let Some(kw) = ext ext.shellcode.iter().find_map(|pattern| {
.keywords bytes
.iter() .windows(pattern.len())
.find(|kw| lower.contains(kw.as_str())) .any(|w| w == pattern.as_slice())
{ .then_some(pattern.as_slice())
return Some(kw); })
} })
} })
None
} }
/// Check whether the host part of `uri` ends with a known phishing TLD. /// Check whether `host` (the URL authority, ASCII-case-folded, no
/// scheme, no port) is byte-equal to a known homograph host string.
/// ///
/// Checks both the built-in table and any external rules loaded via /// This is the only brand-impersonation detector in the scanner.
/// [`load_external_rules`]. /// Matching is exact-string against the host portion only — never
/// substring, never path/query/fragment. This is the rule that
/// distinguishes `https://github.com/microsoft/vscode` (not flagged —
/// the host is `github.com`) from `https://micros0ft.com/anything`
/// (flagged — the host is exactly `micros0ft.com`).
///
/// # Returns
///
/// `Some(host)` if `host` matches a known homograph, where `host` is
/// the matched homograph string itself (the verifiable artifact —
/// an operator can read this string out of the report and confirm
/// it is a homograph). `None` otherwise.
///
/// Returns an owned `String` rather than `&'static str` so that
/// external-feed entries (which are loaded at runtime and stored in
/// an `ExternalRulesData`) can be returned without unsafe
/// lifetime-extension. The crate forbids `unsafe` so we cannot
/// transmute the external entry's lifetime to `'static`.
#[must_use] #[must_use]
pub fn has_phishing_tld(uri: &str) -> Option<&'static str> { pub fn match_homograph_host(host: &str) -> Option<String> {
let lower = uri.to_ascii_lowercase(); let lower = host.to_ascii_lowercase();
// Extract host portion (after scheme://, before path/query/fragment).
let host = lower // Built-in list — direct byte-equal match after case-folding.
.split("://") let builtin = HOMOGRAPH_HOSTS
.nth(1) .iter()
.unwrap_or(&lower) .find(|entry| lower == **entry)
.split('/') .map(|entry| (*entry).to_string());
.next()
.unwrap_or(""); // External feeds — only consulted when no built-in matched.
// Strip port. builtin.or_else(|| {
let host = host.split(':').next().unwrap_or(""); EXTERNAL_RULES_STORAGE.get().and_then(|ext| {
if let Some(tld) = PHISHING_TLDS.iter().copied().find(|tld| host.ends_with(tld)) { ext.homograph_hosts
return Some(tld); .iter()
} .find(|entry| lower == entry.as_str())
if let Some(ext) = EXTERNAL_RULES_STORAGE.get() { .cloned()
if let Some(tld) = ext.tlds.iter().find(|tld| host.ends_with(tld.as_str())) { })
return Some(tld); })
} }
}
None /// Predicate: does `bytes[offset..]` start with `magic`?
/// Returns `false` when `bytes` is shorter than `offset + magic.len()`.
fn bytes_matches_at(bytes: &[u8], offset: usize, magic: &[u8]) -> bool {
let Some(end) = offset.checked_add(magic.len()) else {
return false;
};
bytes.get(offset..end).is_some_and(|slice| slice == magic)
} }
#[cfg(test)] #[cfg(test)]
@ -625,70 +474,64 @@ mod tests {
assert!(match_shellcode_pattern(&payload).is_some()); assert!(match_shellcode_pattern(&payload).is_some());
} }
// ─── Homograph host tests ────────────────────────────────────
//
// The homograph-host matcher only ever matches the URL authority,
// byte-equal after case-folding. The following tests assert that
// URLs whose hosts are NOT in the homograph list — including
// canonical brand domains and URLs that merely mention a brand in
// their path — are NOT flagged.
#[test] #[test]
fn detects_url_shortener() { fn homograph_host_exact_match_is_flagged() {
assert!(is_url_shortener("bit.ly")); // micros0ft.com (with zero instead of 'o') is in the list — flagged.
assert!(is_url_shortener("sub.bit.ly")); assert_eq!(match_homograph_host("micros0ft.com"), Some("micros0ft.com".to_string()));
assert!(!is_url_shortener("example.com")); assert_eq!(match_homograph_host("paypa1.com"), Some("paypa1.com".to_string()));
} }
#[test] #[test]
fn detects_phishing_brand_homograph() { fn homograph_host_case_insensitive() {
// micros0ft.com (with zero instead of 'o') should match the // ASCII-case-folded before matching.
// homograph variant directly. assert_eq!(match_homograph_host("MICROS0FT.COM"), Some("micros0ft.com".to_string()));
assert_eq!( assert_eq!(match_homograph_host("PayPa1.COM"), Some("paypa1.com".to_string()));
match_phishing_brand("https://micros0ft.com/login"),
Some("micros0ft")
);
// paypa1.com (with one instead of 'l') should match.
assert_eq!(
match_phishing_brand("https://paypa1.com/signin"),
Some("paypa1")
);
} }
#[test] #[test]
fn canonical_brand_domain_not_flagged() { fn canonical_brand_host_not_flagged() {
// microsoft.com (canonical) should not be flagged. // Real microsoft.com / paypal.com / apple.com hosts are NOT
assert!(match_phishing_brand("https://microsoft.com/windows").is_none()); // in the homograph list — never flagged.
// paypal.com (canonical) should not be flagged. assert!(match_homograph_host("microsoft.com").is_none());
assert!(match_phishing_brand("https://paypal.com/home").is_none()); assert!(match_homograph_host("paypal.com").is_none());
assert!(match_homograph_host("apple.com").is_none());
assert!(match_homograph_host("google.com").is_none());
assert!(match_homograph_host("amazon.com").is_none());
} }
#[test] #[test]
fn canonical_brand_in_non_canonical_domain_is_flagged() { fn brand_mentioned_in_path_not_flagged() {
// microsoft mentioned in a non-canonical host → suspicious. // The host is github.com — not in the homograph list. The
assert_eq!( // "microsoft" substring in the path is irrelevant to the
match_phishing_brand("https://login-microsoft.com/verify"), // host-only check. This is the guarantee that lets clean
Some("microsoft") // technical documents (RFCs, books, vendor whitepapers) pass
); // through the scanner with zero findings.
assert!(match_homograph_host("github.com").is_none());
assert!(match_homograph_host("en.wikipedia.org").is_none());
} }
#[test] #[test]
fn detects_suspicious_keyword() { fn homograph_host_subdomain_not_matched() {
// Should return the first matching keyword — both "account" // Subdomains of a homograph host are NOT matched. The check
// and "verify" are in the list. Either is acceptable; check // is byte-equal on the full host string. This is intentional:
// that we get one of them. // `login.micros0ft.com` could be a phishing subdomain of a
let result = match_suspicious_keyword("https://example.com/account/verify"); // homograph domain, but it could also be a coincidence in a
assert!(matches!(result, Some("account") | Some("verify"))); // legitimate subdomain naming scheme. The operator can extend
assert_eq!( // the list via external feeds if they want to match subdomains.
match_suspicious_keyword("https://example.com/signin"), assert!(match_homograph_host("login.micros0ft.com").is_none());
Some("signin") assert!(match_homograph_host("www.paypa1.com").is_none());
);
} }
#[test] #[test]
fn detects_phishing_tld() { fn empty_host_not_flagged() {
assert_eq!(has_phishing_tld("https://example.xyz"), Some(".xyz")); assert!(match_homograph_host("").is_none());
assert_eq!(has_phishing_tld("https://example.top/path"), Some(".top"));
assert!(has_phishing_tld("https://example.com").is_none());
}
#[test]
fn detects_phishing_tld_with_port() {
assert_eq!(
has_phishing_tld("https://example.xyz:8080/path"),
Some(".xyz")
);
} }
} }

366
tests/fixtures/benign_realworld.pdf vendored Normal file
View File

@ -0,0 +1,366 @@
%PDF-1.3
%âãÏÓ
1 0 obj
<<
/Producer (pypdf)
>>
endobj
2 0 obj
<<
/Type /Pages
/Count 3
/Kids [ 4 0 R 8 0 R 10 0 R ]
>>
endobj
3 0 obj
<<
/Type /Catalog
/Pages 2 0 R
>>
endobj
4 0 obj
<<
/Contents 5 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 13 0 R 15 0 R 17 0 R 19 0 R 21 0 R 23 0 R 25 0 R ]
>>
endobj
5 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 416
>>
stream
Gas2EgM4V[%#46J'Y<(N`,V-N-!-jOKs;*pbmU>MPIO70A'8tT5NfhJMl%="QE3m1s)X"oiVb>HJ8GL,a+5le.tGZK%+u]YZ\=J3n`qNRoeM-#JVU]af-Cgp&3a8QG]nZQ[Z?MU+FC[s(CbLW.=irT=;4B+]EYE-6;C>hBq$l$]jAT[[I0qHfN;R&&ljabA@(LMAjCk-aqp@?.n_f!"8:'qa=a=C!a/g]^cJUFd!V>5\,eW8,W*X\1h.WbVB?nRF_L(^XkftOhO-DGE//<LBq[jdniTpH<)gD,.'R.Wh(d7\02`?C3E;[R^P:0IkOG<C>??E=h-ti_`[aHE::>SaV9:uC2"=M5<dO(X,plN6hL5./IgnG&)hRZj,bp'K$<ro"aZ=k46'<Jo?bQY_Z<C$r@IX`&Pg75~>
endstream
endobj
6 0 obj
<<
/F1 7 0 R
>>
endobj
7 0 obj
<<
/BaseFont /Helvetica
/Encoding /WinAnsiEncoding
/Name /F1
/Subtype /Type1
/Type /Font
>>
endobj
8 0 obj
<<
/Contents 9 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 27 0 R 29 0 R 31 0 R ]
>>
endobj
9 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 482
>>
stream
Gas2E5u68i&;BTNMKbht+AU@L,*!f"$PaDPj]Hfq8MXSlNbi>Er]O"7YW[)0QE9W/n+3$:!dRtg]W3*`Xs(Lo-q_tuMCPZ'hr2"-8blh@HC<f/RA93?iMMf+,`a^uouKV/D%SkHj%G3XImj5Oor%i2o:>ef[+Mf$&KS75>2uEMon-4Zfb,4lqDls0pWdCVM&2L?@\)B(o3`P.BF;oRMDWuFk4e7gR?uHL;1>YjRtalNE(25LC\!cb3'tC4dq+K[*;S'F>e\FB3MaCuDoeik2Cp#F]PM%<BrAtBCcoS;G3iRe+BE;IkTIcJZuKr'*]c(=Ljg5\]u:Z>eAZB]0sQ2=/c!FY4*?m,LiNj#?914f3^S4d:'`E<debP:0b3/7LaEV"+#?WlFIP;JMHqV?'#-!HH/^[lUfm+oA@(Z-G%OGrnUm*N_):c/;U](H17IYO]5`NmfX'<mQRUOErXC865rh[>n8Rq*;%V['~>
endstream
endobj
10 0 obj
<<
/Contents 11 0 R
/MediaBox [ 0 0 612 792 ]
/Resources <<
/Font 6 0 R
/ProcSet [ /PDF /Text /ImageB /ImageC /ImageI ]
>>
/Rotate 0
/Trans <<
>>
/Type /Page
/Parent 2 0 R
/Annots [ 33 0 R 35 0 R 37 0 R ]
>>
endobj
11 0 obj
<<
/Filter [ /ASCII85Decode /FlateDecode ]
/Length 315
>>
stream
Gas3,Z#51J&:i_&:N9lK.5hAs:rYuhe>`*K:id,LN/c,[KXWUg<iVl*=a,M[s0I5d3b$sG"aH?K?Nc0!U]usf*98W?jCYt>^JE#Up1XT6Knj/:FV+p8#"S5-C^Ds33ecN)j<r$H)f;mfmtgJb$dbK\\4F@%<A`PuNNk5sYZ7SL([#n."BS_K%)/Y(lbe$I/Bns!(qXbFHl*'"0P]b7bXnk-8IJD"G_t`[IZ#)L)3JjMM65V(ofTq5CTdg*`n5Mp)]?-IK.7f.+NZKmQ`BF(1?D_H.HPmmq%o4Aj!q>bmig8SNVh>Ejr9j3Q`:~>
endstream
endobj
12 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056linuxfromscratch\056org\057)
>>
endobj
13 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 695 400 710 ]
/A 12 0 R
/Border [ 0 0 0 ]
>>
endobj
14 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057lists\056linuxfromscratch\056org\057listinfo\057lfs\055support)
>>
endobj
15 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 675 400 690 ]
/A 14 0 R
/Border [ 0 0 0 ]
>>
endobj
16 0 obj
<<
/Type /Action
/S /URI
/URI (http\072\057\057ftp\056osuosl\056org\057pub\057lfs\057)
>>
endobj
17 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 655 400 670 ]
/A 16 0 R
/Border [ 0 0 0 ]
>>
endobj
18 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057github\056com\057LFS\055project\057build\055scripts)
>>
endobj
19 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 635 400 650 ]
/A 18 0 R
/Border [ 0 0 0 ]
>>
endobj
20 0 obj
<<
/Type /Action
/S /URI
/URI (mailto\072lfs\055support\100linuxfromscratch\056org)
>>
endobj
21 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 595 400 610 ]
/A 20 0 R
/Border [ 0 0 0 ]
>>
endobj
22 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056kernel\056org\057pub\057linux\057kernel\057)
>>
endobj
23 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 575 400 590 ]
/A 22 0 R
/Border [ 0 0 0 ]
>>
endobj
24 0 obj
<<
/Type /Action
/S /URI
/URI (tel\072\0531\055555\055123\0554567)
>>
endobj
25 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 555 400 570 ]
/A 24 0 R
/Border [ 0 0 0 ]
>>
endobj
26 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057ftp\056ru\056debian\056org\057debian\057)
>>
endobj
27 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 575 400 590 ]
/A 26 0 R
/Border [ 0 0 0 ]
>>
endobj
28 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057en\056wikipedia\056org\057wiki\057Microsoft\137Windows)
>>
endobj
29 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 555 400 570 ]
/A 28 0 R
/Border [ 0 0 0 ]
>>
endobj
30 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057github\056com\057microsoft\057vscode)
>>
endobj
31 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 535 400 550 ]
/A 30 0 R
/Border [ 0 0 0 ]
>>
endobj
32 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057www\056ietf\056org\057rfc\057rfc2616\056txt)
>>
endobj
33 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 655 400 670 ]
/A 32 0 R
/Border [ 0 0 0 ]
>>
endobj
34 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057example\056com\057account\057verify)
>>
endobj
35 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 615 400 630 ]
/A 34 0 R
/Border [ 0 0 0 ]
>>
endobj
36 0 obj
<<
/Type /Action
/S /URI
/URI (https\072\057\057example\056com\057login)
>>
endobj
37 0 obj
<<
/Type /Annot
/Subtype /Link
/Rect [ 80 595 400 610 ]
/A 36 0 R
/Border [ 0 0 0 ]
>>
endobj
xref
0 38
0000000000 65535 f
0000000015 00000 n
0000000054 00000 n
0000000126 00000 n
0000000175 00000 n
0000000425 00000 n
0000000932 00000 n
0000000963 00000 n
0000001070 00000 n
0000001292 00000 n
0000001865 00000 n
0000002089 00000 n
0000002496 00000 n
0000002599 00000 n
0000002702 00000 n
0000002833 00000 n
0000002936 00000 n
0000003042 00000 n
0000003145 00000 n
0000003265 00000 n
0000003368 00000 n
0000003471 00000 n
0000003574 00000 n
0000003693 00000 n
0000003796 00000 n
0000003882 00000 n
0000003985 00000 n
0000004094 00000 n
0000004197 00000 n
0000004320 00000 n
0000004423 00000 n
0000004528 00000 n
0000004631 00000 n
0000004743 00000 n
0000004846 00000 n
0000004950 00000 n
0000005053 00000 n
0000005145 00000 n
trailer
<<
/Size 38
/Root 3 0 R
/Info 1 0 R
>>
startxref
5248
%%EOF

View File

@ -28,7 +28,101 @@ fn benign_pdf_produces_no_findings() {
let result = pipeline.run(fixture("benign.pdf")).unwrap(); let result = pipeline.run(fixture("benign.pdf")).unwrap();
assert_eq!(result.scan_report.malicious_count(), 0); assert_eq!(result.scan_report.malicious_count(), 0);
assert!(result.quarantine_path.is_none()); assert!(result.quarantine_path.is_none());
assert!(result.cleansed_path.is_none()); // Clean documents now produce a "clean" output file in the
// cleanse_dir (the original file copied under
// `clean_<timestamp>_<sha>.pdf`). This makes the scanner a
// proper pipeline stage. The cleansed_path field is overloaded:
// it holds either the cleansed derivative (malicious case) or
// the clean-output copy (clean case).
assert!(
result.cleansed_path.is_some(),
"clean PDF should produce a clean-output file in the cleanse_dir"
);
let clean_path = result.cleansed_path.unwrap();
assert!(
clean_path
.file_name()
.and_then(|n| n.to_str())
.map(|n| n.starts_with("clean_") && n.ends_with(".pdf"))
.unwrap_or(false),
"clean-output file should be named clean_<ts>_<sha>.pdf, got: {clean_path:?}"
);
// The clean-output file should have the same contents as the original.
let original = std::fs::read(fixture("benign.pdf")).unwrap();
let clean_bytes = std::fs::read(&clean_path).unwrap();
assert_eq!(
original, clean_bytes,
"clean-output file should be a byte-for-byte copy of the original"
);
}
#[test]
fn benign_realworld_pdf_produces_zero_findings() {
// This is the corpus-based regression test for the false-positive
// redesign. The fixture (`benign_realworld.pdf`) is a multi-page
// PDF that mimics the structure of a technical book like *Linux
// from Scratch* — it contains 13 hyperlinks covering every URL
// pattern that USED TO produce a false positive:
//
// - https://www.linuxfromscratch.org/
// - https://lists.linuxfromscratch.org/listinfo/lfs-support (keyword "support")
// - http://ftp.osuosl.org/pub/lfs/ (http: scheme)
// - https://github.com/LFS-project/build-scripts (github.com host)
// - mailto:lfs-support@linuxfromscratch.org (mailto: + @ in path)
// - https://www.kernel.org/pub/linux/kernel/
// - tel:+1-555-123-4567 (tel: scheme)
// - https://ftp.ru.debian.org/debian/ (.ru TLD)
// - https://en.wikipedia.org/wiki/Microsoft_Windows (brand in path)
// - https://github.com/microsoft/vscode (brand in path)
// - https://www.ietf.org/rfc/rfc2616.txt
// - https://example.com/account/verify (keywords in path)
// - https://example.com/login (keyword in path)
//
// Plus prose containing: wget, exploit, payload, /bin/sh, PowerShell.
//
// The assertion is the theorem: a clean technical PDF produces ZERO
// findings (not "fewer than N", not "0 malicious but maybe some
// suspicious" — literally zero findings in the report). If any
// finding appears, the detector that produced it is wrong by
// construction, not the document.
let tmp = tempdir().unwrap();
let pipeline = Pipeline::with_config(Config::with_workspace(tmp.path()));
let result = pipeline.run(fixture("benign_realworld.pdf")).unwrap();
assert_eq!(
result.scan_report.findings.len(),
0,
"benign real-world PDF must produce ZERO findings, got {}: {:#?}",
result.scan_report.findings.len(),
result.scan_report.findings.iter().map(|f| (
f.classification.clone(),
f.context_notes.clone(),
f.payload_preview.chars().take(80).collect::<String>(),
)).collect::<Vec<_>>(),
);
assert_eq!(result.scan_report.malicious_count(), 0);
assert!(result.quarantine_path.is_none());
// The clean-output file should exist (the pipeline-stage behavior).
assert!(
result.cleansed_path.is_some(),
"clean real-world PDF should produce a clean-output file"
);
let clean_path = result.cleansed_path.unwrap();
assert!(
clean_path
.file_name()
.and_then(|n| n.to_str())
.map(|n| n.starts_with("clean_") && n.ends_with(".pdf"))
.unwrap_or(false),
"clean-output file should be named clean_<ts>_<sha>.pdf, got: {clean_path:?}"
);
// The clean-output file should be a byte-for-byte copy of the original.
let original = std::fs::read(fixture("benign_realworld.pdf")).unwrap();
let clean_bytes = std::fs::read(&clean_path).unwrap();
assert_eq!(
original, clean_bytes,
"clean-output file should be a byte-for-byte copy of the original"
);
} }
#[test] #[test]