corbel/README.md

265 lines
12 KiB
Markdown
Executable File

# CorbelPurge
> Strict Rust document sanitizer & threat neutralizer for PDF, EPUB, Markdown, and DOCX.
**Author:** Jeremy Anderson — [dcos.net](https://dcos.net) — [info@dcos.net](mailto:info@dcos.net)
**Repository:** [https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)
**License:** GPL-3.0-or-later
![Corbel-Purge-ss](./corbel-purge-ss.png)
## Overview
CorbelPurge is a strict, local-only document security scanner. It parses
documents into a unified intermediate representation, runs detectors
that are backed by verifiable properties of the bytes on disk, and
produces cleansed derivatives with all executable content stripped.
Malicious payloads are carved into quarantine tarballs with full
forensic reports.
The scanner is built around a single design invariant: **every
detector must be backed by a verifiable property, either of the
document itself or of an external authority.** No thresholds, no
per-file or per-domain exceptions, no statistical "suspicious" tier.
## What it does
Given a file path, `Pipeline::run()` in `src/core/pipeline.rs`:
1. **Detects format** via `DocumentFormat::from_path()` (extension-based:
`.pdf`, `.epub`, `.md`/`.markdown`, `.docx`).
2. **Parses** via `parsers::Dispatcher` into a `Document` containing
`TextNode` (static text with semantic context) and
`ExecutableVector` (active content like JS streams, embedded files,
script tags, VBA macros) items.
3. **Scans** with a two-pass engine:
- `heuristics::inspect_vector()` classifies every executable vector
against the two-category detector model: Category 1 (verifiable
executable intent — file signatures, shellcode prologues,
executable URI schemes) and Category 2 (verifiable impersonation —
exact-host homographs, credential URLs, mixed-script hosts).
- `context_filter::evaluate()` checks text nodes for structural
signatures (`/JavaScript`, `<script`, etc.) and weaponization
indicators (long hex runs, base64 blobs, multi-shell commands).
Words and function names that appear in legitimate technical
literature (`wget`, `exploit`, `payload`, `eval(`) are not
treated as signatures.
- `cve_tags::match_cve()` annotates findings with known exploit IDs.
4. **Reports** — both clean and malicious scans produce a JSON and
Markdown report in `corbel_quarantine/` named
`report_<timestamp>_<sha_prefix>.{json,md}`. The clean path provides
an audit trail; the malicious path adds a quarantine tarball and
carved payloads.
5. **Quarantines** (when malicious findings exist) — `quarantine::handle()`
carves payloads into `quarantine_<ts>_<sha>.tar.gz` with
`original.<ext>`, `report.json`, `report.md`, and one `.bin` per
payload (plus paired `.hex` and `.info` files).
6. **Cleanses** (when recommended) — produces a sanitized derivative:
- **Markdown mode** (default): `cleanse::sanitizer::sanitize()`
emits a safe Markdown file. Text nodes at malicious locations
are stripped; hyperlinks lose their destinations.
- **PreserveFormat mode** (`--preserve-format`):
`cleanse::repackage::repackage()` rebuilds the original format
with malicious entries removed. EPUB entries are stripped from
the ZIP, DOCX macros/embeddings/external-links are removed, PDF
objects are deleted via lopdf.
7. **Clean-output** (when scan is clean and `emit_clean_output` is on,
the default) — copies the source file to
`corbel_clean/clean_<ts>_<sha>.<ext>` so the scanner acts as a
pipeline stage. Use `--move-clean` for queue-draining semantics.
Supported formats: **PDF**, **EPUB**, **Markdown**, **DOCX**.
## Detection model
Two categories. Nothing else fires.
### Category 1 — Verifiable executable intent
The vector contains a structure whose only purpose is to execute
code or spawn a process. Presence is the threat.
| Detector | Triggers on |
|---|---|
| Active script in PDF | `/JavaScript` or `/JS` action stream |
| Program launch in PDF | `/Launch` action with `/F`, `/Win`, `/Mac`, `/Unix` |
| External program exec in EPUB | `<script>` tag in XHTML |
| VBA macro in DOCX | `word/vbaProject.xml` present in package |
| Executable embedded file | Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0 |
| Executable URI scheme | `javascript:`, `vbscript:`, `data:text/html`, `file:` in a hyperlink or action |
| PDF form with `/AA` | AcroForm dictionary contains the `/AA` (Additional Actions) entry |
| PDF widget with `/AA` | Widget annotation with `/AA` entry |
| Shellcode prologue | Known Metasploit / NOP-sled / syscall-stub byte sequences |
### Category 2 — Verifiable impersonation
The vector lies about identity in a way that is provably wrong.
| Detector | Triggers on |
|---|---|
| Exact-host homograph | URL authority byte-equal (after ASCII case-folding) to a known homograph string (`micros0ft.com`, `paypa1.com`, etc.) |
| Credential URL | RFC 3986 authority component contains `user:pass@` (before the first `/` after `://`) |
| Mixed-script host | URL host mixes Latin with Cyrillic or Greek (a property of the codepoints) |
### What is NOT a detector
The following were removed because they are statistical guesses
about the world, not verifiable properties of the document:
- **Substring keyword matching** — words like `login`, `verify`,
`support`, `account` are not threats. Every legitimate login page
contains them.
- **Phishing TLDs** — `.ru`, `.cn`, `.xyz` are not threats. A Russian
URL is a Russian URL.
- **URL shorteners as a threat signal** — `bit.ly` is not a threat.
If the destination is hostile, the underlying detector catches it.
- **Substring brand matching** — `https://github.com/microsoft/vscode`
is not a threat. The host is `github.com`; the brand substring in
the path is irrelevant.
- **IP-address hosts** — `192.168.1.1` is a valid network address.
RFCs and router manuals reference them.
- **Substring text-node signatures** — words like `wget`, `exploit`,
`payload`, `eval(`, `powershell`, `/bin/sh` in prose are not
threats. Only structural tokens (`/JavaScript`, `<script`,
`shellcode`) and weaponization indicators (hex runs, base64 blobs,
multi-shell commands) fire.
## Pipeline architecture
```text
input path ──► parser ──► Document (UIR)
│
▼
scanner ──► ScanReport
│
┌────────────────────┼────────────────────┐
│ │ │
▼ ▼ ▼
(clean) (malicious) (educational)
│ │ │
▼ ▼ ▼
clean_<ts>_<sha>.<ext> quarantine tarball whitelisted
+ report.json / .md + carved payloads (logged only)
+ cleansed file
```
## Library usage
```rust
use corbel_purge::{Pipeline, Config, CleanseMode};
// Scan a file on disk
let pipeline = Pipeline::with_config(Config::with_workspace("/tmp/output"));
let result = pipeline.run("suspicious.pdf")?;
println!("malicious: {}", result.scan_report.malicious_count());
// Scan in-memory bytes (e.g. from an email gateway)
let bytes = std::fs::read("upload.docx")?;
let result = pipeline.run_on_bytes(
bytes,
corbel_purge::DocumentFormat::Docx,
Some("upload.docx".into()),
)?;
// PreserveFormat mode
let mut config = Config::default();
config.cleanse_mode = CleanseMode::PreserveFormat;
let pipeline = Pipeline::with_config(config);
let result = pipeline.run("book.epub")?;
```
## Source layout
```text
src/
├── main.rs # CLI: scan, scan-dir, study, --version, arg parsing
├── lib.rs # Crate root, CorbelError enum, public re-exports
├── util.rs # sha256_hex(), truncate_with_ellipsis(), read_with_cap()
├── bin/
│ └── gui.rs # iced 0.13 dashboard GUI (gui feature)
├── core/
│ ├── types.rs # Document, TextNode, ExecutableVector, Finding, ScanReport,
│ │ # Location, TextContext, VectorType, ThreatClassification,
│ │ # MaliciousType, Recommendation, DocumentMetadata
│ ├── config.rs # Config, CleanseMode, env overrides, workspace dirs
│ └── pipeline.rs # Pipeline, PipelineResult, DocumentFormat::from_path()
├── parsers/
│ ├── mod.rs # DocumentParser trait, Dispatcher
│ ├── pdf_parser.rs # lopdf-based PDF structure walker
│ ├── epub_parser.rs # ZIP-based EPUB container inspector
│ ├── md_parser.rs # pulldown-cmark text extractor
│ └── docx_parser.rs # OOXML ZIP inspector (regex XML extraction)
├── scanner/
│ ├── mod.rs # scan() entrypoint, CVE tag injection
│ ├── heuristics.rs # classify_vector() — two-category detector model
│ ├── context_filter.rs # evaluate() — structural signatures + weaponization
│ ├── signatures.rs # KNOWN_FILE_SIGNATURES, SHELLCODE_PATTERNS,
│ │ # HOMOGRAPH_HOSTS, URL_SHORTENER_DOMAINS (retained
│ │ # for future reputation work, not consulted)
│ └── cve_tags.rs # CVE_TABLE, match_cve(), cve_tag()
├── quarantine/
│ ├── mod.rs # handle() — tarball writer, write_clean_report()
│ ├── extractor.rs # extract_payloads(), ExtractedPayload, filename sanitization
│ ├── hexdump.rs # hex_dump(), build_payload_info()
│ └── reporter.rs # build_json_report(), build_markdown_report()
├── cleanse/
│ ├── mod.rs # cleanse() dispatcher (Markdown vs PreserveFormat)
│ ├── sanitizer.rs # sanitize() — Markdown re-serializer
│ └── repackage.rs # repackage() — format-preserving (EPUB/DOCX/PDF)
└── ui/
├── mod.rs # GUI module gate (behind `gui` feature)
├── viewer.rs # Stub
└── alert_modal.rs # Stub
```
## Design philosophy
1. **Verifiable properties only.** A detector either proves a finding
by a property of the bytes, or it doesn't fire. No thresholds.
2. **No special-casing.** No per-file or per-domain exceptions. The
clean-document guarantee is a theorem about the detectors, not an
empirical observation about one specific file.
3. **Unix philosophy.** Each module does one thing. The pipeline is a
linear sequence of small steps. Step-down logic (early returns) for
every fork of choices.
4. **`#![forbid(unsafe_code)]`** at the crate root. The scanner never
touches `unsafe` Rust.
5. **Local-only.** No network access during scanning. All analysis is
against the bytes on disk. External threat-intel feeds are loaded
from local JSON files at startup.
## Adding a new format
1. Create `src/parsers/<format>_parser.rs` implementing `DocumentParser`.
2. Add a variant to `DocumentFormat` in `src/core/types.rs` (and its
`Display` impl).
3. Add a match arm in `Dispatcher::parse()` in `src/parsers/mod.rs`.
4. Add an extension check in `DocumentFormat::from_path()` in
`src/core/pipeline.rs`.
5. Add repackage support in `src/cleanse/repackage.rs` (optional).
The scanner, quarantine, and cleanse modules consume only the
`Document` UIR and require no changes for basic format support.
## Non-goals
- No document editing or annotation.
- No execution of embedded scripting (PDF JS, EPUB scripts, VBA macros).
- No network access during scanning — all analysis is local.
- No sandbox hardening of the tool itself.
## Documentation
- **[QUICKSTART.md](QUICKSTART.md)** — Build guide and operator reference
- **[MANIFEST.md](MANIFEST.md)** — Detailed technical specification
- **[TODO.md](TODO.md)** — Development roadmap
- **[FIX-NOTES-false-positive-redesign.md](FIX-NOTES-false-positive-redesign.md)** —
Design notes for the two-category detector model
- **[FIX-NOTES-indoc-links.md](FIX-NOTES-indoc-links.md)** — In-document link handling
- **[FIX-NOTES-glyphfix.md](FIX-NOTES-glyphfix.md)** — Glyph rendering fixes
## License
GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson.
[https://git.dcos.net/dcosnet/corbel](https://git.dcos.net/dcosnet/corbel)