corbel/README.md

12 KiB
Executable File

CorbelPurge

Strict Rust document sanitizer & threat neutralizer for PDF, EPUB, Markdown, and DOCX.

Author: Jeremy Anderson — dcos.net — info@dcos.net Repository: https://git.dcos.net/dcosnet/corbel License: GPL-3.0-or-later

Corbel-Purge-ss

Overview

CorbelPurge is a strict, local-only document security scanner. It parses documents into a unified intermediate representation, runs detectors that are backed by verifiable properties of the bytes on disk, and produces cleansed derivatives with all executable content stripped. Malicious payloads are carved into quarantine tarballs with full forensic reports.

The scanner is built around a single design invariant: every detector must be backed by a verifiable property, either of the document itself or of an external authority. No thresholds, no per-file or per-domain exceptions, no statistical "suspicious" tier.

What it does

Given a file path, Pipeline::run() in src/core/pipeline.rs:

  1. Detects format via DocumentFormat::from_path() (extension-based: .pdf, .epub, .md/.markdown, .docx).
  2. Parses via parsers::Dispatcher into a Document containing TextNode (static text with semantic context) and ExecutableVector (active content like JS streams, embedded files, script tags, VBA macros) items.
  3. Scans with a two-pass engine:
    • heuristics::inspect_vector() classifies every executable vector against the two-category detector model: Category 1 (verifiable executable intent — file signatures, shellcode prologues, executable URI schemes) and Category 2 (verifiable impersonation — exact-host homographs, credential URLs, mixed-script hosts).
    • context_filter::evaluate() checks text nodes for structural signatures (/JavaScript, <script, etc.) and weaponization indicators (long hex runs, base64 blobs, multi-shell commands). Words and function names that appear in legitimate technical literature (wget, exploit, payload, eval() are not treated as signatures.
    • cve_tags::match_cve() annotates findings with known exploit IDs.
  4. Reports — both clean and malicious scans produce a JSON and Markdown report in corbel_quarantine/ named report_<timestamp>_<sha_prefix>.{json,md}. The clean path provides an audit trail; the malicious path adds a quarantine tarball and carved payloads.
  5. Quarantines (when malicious findings exist) — quarantine::handle() carves payloads into quarantine_<ts>_<sha>.tar.gz with original.<ext>, report.json, report.md, and one .bin per payload (plus paired .hex and .info files).
  6. Cleanses (when recommended) — produces a sanitized derivative:
    • Markdown mode (default): cleanse::sanitizer::sanitize() emits a safe Markdown file. Text nodes at malicious locations are stripped; hyperlinks lose their destinations.
    • PreserveFormat mode (--preserve-format): cleanse::repackage::repackage() rebuilds the original format with malicious entries removed. EPUB entries are stripped from the ZIP, DOCX macros/embeddings/external-links are removed, PDF objects are deleted via lopdf.
  7. Clean-output (when scan is clean and emit_clean_output is on, the default) — copies the source file to corbel_clean/clean_<ts>_<sha>.<ext> so the scanner acts as a pipeline stage. Use --move-clean for queue-draining semantics.

Supported formats: PDF, EPUB, Markdown, DOCX.

Detection model

Two categories. Nothing else fires.

Category 1 — Verifiable executable intent

The vector contains a structure whose only purpose is to execute code or spawn a process. Presence is the threat.

Detector Triggers on
Active script in PDF /JavaScript or /JS action stream
Program launch in PDF /Launch action with /F, /Win, /Mac, /Unix
External program exec in EPUB <script> tag in XHTML
VBA macro in DOCX word/vbaProject.xml present in package
Executable embedded file Bytes match a known executable magic (MZ, ELF, Mach-O, OLE2, RTF, SWF, LNK, HTA, VBA stubs) at offset 0
Executable URI scheme javascript:, vbscript:, data:text/html, file: in a hyperlink or action
PDF form with /AA AcroForm dictionary contains the /AA (Additional Actions) entry
PDF widget with /AA Widget annotation with /AA entry
Shellcode prologue Known Metasploit / NOP-sled / syscall-stub byte sequences

Category 2 — Verifiable impersonation

The vector lies about identity in a way that is provably wrong.

Detector Triggers on
Exact-host homograph URL authority byte-equal (after ASCII case-folding) to a known homograph string (micros0ft.com, paypa1.com, etc.)
Credential URL RFC 3986 authority component contains user:pass@ (before the first / after ://)
Mixed-script host URL host mixes Latin with Cyrillic or Greek (a property of the codepoints)

What is NOT a detector

The following were removed because they are statistical guesses about the world, not verifiable properties of the document:

  • Substring keyword matching — words like login, verify, support, account are not threats. Every legitimate login page contains them.
  • Phishing TLDs — .ru, .cn, .xyz are not threats. A Russian URL is a Russian URL.
  • URL shorteners as a threat signal — bit.ly is not a threat. If the destination is hostile, the underlying detector catches it.
  • Substring brand matching — https://github.com/microsoft/vscode is not a threat. The host is github.com; the brand substring in the path is irrelevant.
  • IP-address hosts — 192.168.1.1 is a valid network address. RFCs and router manuals reference them.
  • Substring text-node signatures — words like wget, exploit, payload, eval(, powershell, /bin/sh in prose are not threats. Only structural tokens (/JavaScript, <script, shellcode) and weaponization indicators (hex runs, base64 blobs, multi-shell commands) fire.

Pipeline architecture

input path ──► parser ──► Document (UIR)
                                 │
                                 ▼
                              scanner ──► ScanReport
                                 │
            ┌────────────────────┼────────────────────┐
            │                    │                    │
            ▼                    ▼                    ▼
        (clean)              (malicious)          (educational)
            │                    │                    │
            ▼                    ▼                    ▼
    clean_<ts>_<sha>.<ext>  quarantine tarball   whitelisted
    + report.json / .md     + carved payloads    (logged only)
                            + cleansed file

Library usage

use corbel_purge::{Pipeline, Config, CleanseMode};

// Scan a file on disk
let pipeline = Pipeline::with_config(Config::with_workspace("/tmp/output"));
let result = pipeline.run("suspicious.pdf")?;
println!("malicious: {}", result.scan_report.malicious_count());

// Scan in-memory bytes (e.g. from an email gateway)
let bytes = std::fs::read("upload.docx")?;
let result = pipeline.run_on_bytes(
    bytes,
    corbel_purge::DocumentFormat::Docx,
    Some("upload.docx".into()),
)?;

// PreserveFormat mode
let mut config = Config::default();
config.cleanse_mode = CleanseMode::PreserveFormat;
let pipeline = Pipeline::with_config(config);
let result = pipeline.run("book.epub")?;

Source layout

src/
├── main.rs                  # CLI: scan, scan-dir, study, --version, arg parsing
├── lib.rs                   # Crate root, CorbelError enum, public re-exports
├── util.rs                  # sha256_hex(), truncate_with_ellipsis(), read_with_cap()
├── bin/
│   └── gui.rs               # iced 0.13 dashboard GUI (gui feature)
├── core/
│   ├── types.rs             # Document, TextNode, ExecutableVector, Finding, ScanReport,
│   │                         # Location, TextContext, VectorType, ThreatClassification,
│   │                         # MaliciousType, Recommendation, DocumentMetadata
│   ├── config.rs            # Config, CleanseMode, env overrides, workspace dirs
│   └── pipeline.rs          # Pipeline, PipelineResult, DocumentFormat::from_path()
├── parsers/
│   ├── mod.rs                # DocumentParser trait, Dispatcher
│   ├── pdf_parser.rs         # lopdf-based PDF structure walker
│   ├── epub_parser.rs        # ZIP-based EPUB container inspector
│   ├── md_parser.rs          # pulldown-cmark text extractor
│   └── docx_parser.rs        # OOXML ZIP inspector (regex XML extraction)
├── scanner/
│   ├── mod.rs                # scan() entrypoint, CVE tag injection
│   ├── heuristics.rs         # classify_vector() — two-category detector model
│   ├── context_filter.rs     # evaluate() — structural signatures + weaponization
│   ├── signatures.rs         # KNOWN_FILE_SIGNATURES, SHELLCODE_PATTERNS,
│   │                         # HOMOGRAPH_HOSTS, URL_SHORTENER_DOMAINS (retained
│   │                         # for future reputation work, not consulted)
│   └── cve_tags.rs           # CVE_TABLE, match_cve(), cve_tag()
├── quarantine/
│   ├── mod.rs                # handle() — tarball writer, write_clean_report()
│   ├── extractor.rs          # extract_payloads(), ExtractedPayload, filename sanitization
│   ├── hexdump.rs            # hex_dump(), build_payload_info()
│   └── reporter.rs           # build_json_report(), build_markdown_report()
├── cleanse/
│   ├── mod.rs                # cleanse() dispatcher (Markdown vs PreserveFormat)
│   ├── sanitizer.rs          # sanitize() — Markdown re-serializer
│   └── repackage.rs          # repackage() — format-preserving (EPUB/DOCX/PDF)
└── ui/
    ├── mod.rs                # GUI module gate (behind `gui` feature)
    ├── viewer.rs              # Stub
    └── alert_modal.rs        # Stub

Design philosophy

  1. Verifiable properties only. A detector either proves a finding by a property of the bytes, or it doesn't fire. No thresholds.
  2. No special-casing. No per-file or per-domain exceptions. The clean-document guarantee is a theorem about the detectors, not an empirical observation about one specific file.
  3. Unix philosophy. Each module does one thing. The pipeline is a linear sequence of small steps. Step-down logic (early returns) for every fork of choices.
  4. #![forbid(unsafe_code)] at the crate root. The scanner never touches unsafe Rust.
  5. Local-only. No network access during scanning. All analysis is against the bytes on disk. External threat-intel feeds are loaded from local JSON files at startup.

Adding a new format

  1. Create src/parsers/<format>_parser.rs implementing DocumentParser.
  2. Add a variant to DocumentFormat in src/core/types.rs (and its Display impl).
  3. Add a match arm in Dispatcher::parse() in src/parsers/mod.rs.
  4. Add an extension check in DocumentFormat::from_path() in src/core/pipeline.rs.
  5. Add repackage support in src/cleanse/repackage.rs (optional).

The scanner, quarantine, and cleanse modules consume only the Document UIR and require no changes for basic format support.

Non-goals

  • No document editing or annotation.
  • No execution of embedded scripting (PDF JS, EPUB scripts, VBA macros).
  • No network access during scanning — all analysis is local.
  • No sandbox hardening of the tool itself.

Documentation

License

GPL-3.0-or-later. Copyright (c) 2026 Jeremy Anderson. https://git.dcos.net/dcosnet/corbel