# Growing a Thicket *Building a super-ingest + RAG console that treats a document library like a living ecosystem, not a filing cabinet.* **Jeremy Anderson** — [dcos.net](https://dcos.net) — info@dcos.net September 2026 --- A few weeks ago a friend's Gemini session produced ~270 lines of Python titled "brain ingest": walk a folder, extract text from PDFs and EPUBs, write Markdown notes into an Obsidian vault, chunk, embed, and pour the results into Qdrant. It was a good sketch. It also had four bugs I cared about, an import surface that required every heavy dependency up front, and a LightRAG binding pinned to one release of a fast-moving API. We turned it into **Thicket**: a super-ingest + retrieval console with the same brushed-aluminum, amber-LED MMD3 interface as my transcoder [OpenTranscode](http://git.dcos.net/dcosnet/OpenTranscode) — because the tool you actually use is the tool you actually enjoy opening. The name is the design brief. A thicket is dense, self-connected, and grows on its own. Feed it documents; it grows a thicket. ## The architecture bet: one engine, two drivers The single most important decision was making the pipeline core Qt-free. `IngestPipeline` knows nothing about widgets: it walks a stage table and reports progress through four plain callables. Two drivers share it: - a **QThread worker** that wires those callables to Qt signals, and - a **headless CLI** (`thicket --ingest ... --vault ...`) that wires them to `print`, works over SSH, and fits in a cron line. One behavior change lands in both drivers at once. There is no "GUI logic" to port or keep in sync — that category of bug cannot exist. ## Stage tables instead of branch nests Every per-document decision is data, not control flow. The pipeline stage order is a tuple of `(status, gate, runner)`; file-type extraction is a dict keyed by extension; readiness checks before an ingest run are a table where the first enabled-but-not-ready entry wins. Adding a stage, a format, or a gate means adding an entry — never a new `if` ladder. The QA pass measured this concretely: the extractor went from a five-branch dispatch to a one-line dict lookup. ## Lazy by default, graceful by contract The console starts on a bare system. Every heavy import — FastEmbed, the Qdrant client, EbookLib, LightRAG — happens at point of use, and a background probe reports module and service readiness into the LED status strip. Stages that cannot run are blocked at INGEST with the exact reason and the exact fix; nothing crashes, nothing silently no-ops. Where a choice of paths exists, the code steps down the chain explicitly, best option first: Qdrant search uses `query_points` and falls to `search` on older clients; the LightRAG Ollama binding resolves across the package's API generations; podman networking that cannot create a tap device gets `--network=host`. Unix philosophy in practice: try the clean path, degrade to the simple one, keep going. ## Re-ingestion is an overwrite The subtlest correctness property in the system: **the point count after N re-ingests equals the count after the first.** Deterministic point IDs (MD5-UUID of `title|path|index`) plus a delete-by-filter before every upsert make editing a source and re-ingesting it an exact replacement — a shrunken document leaves no stale tail chunks. We verify this in the live QA run: 4 points, re-ingest, still 4 points, same retrieval rankings. The vault side holds the same invariant: re-ingesting a source refreshes its note in place, and a *different* document with the same title claims a digest-suffixed sibling rather than clobbering it. ## Local embeddings, on purpose FastEmbed runs ONNX on your own cores; nothing leaves the machine. Retrieval quality still respects the model's contract — BGE models want an instruction prefix on the *query* side and bare passages on the *document* side, so Thicket prefixes queries only. It is the kind of detail that silently costs you ten points of relevance when a sketch gets it wrong. ## The QA pass that shaped the code Before calling it production-ready, the codebase went through a five-hat review — senior QA, Linux engineer, architect, admin, and devops PM. What changed: - **Table-driven dispatch everywhere** it beat nested conditionals (stage table, extension table, readiness table, stage-color map). - **Loops reduced to comprehensions and `next()`** where iteration was bookkeeping; explicit loops remain only where iteration *is* the semantics (chunk word windows, queue walks). - **Bounded reads** on files we do not own (frontmatter collision checks read at most 32 lines). - **One decisive failure path per scope** — per-file isolation in the pipeline, a single report-and-disable path in the retrieval worker. - Comments state invariants and contracts. Version-history narration does not survive review; the code reads like decisions, not like an argument with itself. Standards kept in view: PEP 8 throughout, SEI CERT practices (bounded I/O, precise exception scope — the Qdrant step-down catches `AttributeError`, not the world), MISRA-style bounded structured control flow, and POSIX assumptions (paths via `pathlib`, no platform branches, systemd/cron-friendly headless mode). ## Live-fire verification The release gate was not the test suite alone (26 tests, no services required) but a live run: podman Qdrant up, three documents through the full GUI pipeline, two semantic queries returning correctly ranked hits, an idempotent re-ingest, and a dry-run probe reporting every module and service green. Screenshot or it didn't happen — the console looks the part too: knobs for chunk size, overlap, and Top-K; an LED queue table tracking every file's stage; a phosphor log; a retrieval strip at the bottom. ## What's next The graph stage (LightRAG over Ollama) is wired and gating on readiness, but it is deliberately optional — entity extraction is an LLM pass per document, and the vault + vector stages already answer the daily question: *where did I read that?* The thicket grows. Pull it, feed it a shelf of books, and see what surfaces. --- **Thicket** — AGPL-3.0-or-later — [git.dcos.net/dcosnet/Thicket](http://git.dcos.net/dcosnet/Thicket) Jeremy Anderson — info@dcos.net — [dcos.net](https://dcos.net) --- ## Addendum — v1.8: the thicket grows roots *September 2026, after the first full functionality matrix.* The sketch became a workstation tool. What changed since the first essay: **Destinations, not stages.** The original three-stage line (notes → Qdrant → LightRAG) became a destination model: `obsidian` for notes-only, or any of ten open-source vector stores — qdrant, chroma, lancedb, faiss, milvus, weaviate, pgvector, duckdb, sqlite-vec, mariadb — one payload schema, one exact-replacement contract, cosine scores identical to four decimals across all ten. The console's OUT line resolves the real destination live; the status footer follows the target selector keystroke by keystroke. **Sources are never deleted.** The dangerous delete-after-verify checkbox from the transcoder lineage is gone. Verified sources move to an archive directory and bzip2-compress in place — the incoming tree stays clean, the archive always decompresses. **The corpus is technical.** Fenced code blocks became atomic language-tagged chunks; prose became paragraph-aligned; shebangs stopped masquerading as headings; twenty-one source/config extensions ingest as listings; the default embedder became jina-code, and "force browsers to refuse plain http" retrieves the nginx block where HSTS actually lives. **The workstation has a filesystem.** /mnt/AI/corpus/{cold,hot} is now the default corpus flow — cold in, hot brain — with ~/$VAR expansion everywhere and THICKET_AI_ROOT for other hosts. **Two interaction modes.** SEARCH embeds locally; ASK turns natural language into read-only SQL over the Postgres/MariaDB corpora through Vanna 2's agent API and the same Ollama selector. "Which document has the most chunks?" is now a console question. **The matrix.** `scripts/func_test.py` runs eighteen live checks — every destination with idempotency and retrieval assertions, both graph engines, both archives, ask on both SQL targets. Its first run caught six real bugs, including a subtle one: milvus-lite's embedded server keeps the database file lock after client close, so a second process gets `[Errno 11]` — fixed with an explicit `close()` lifecycle that releases the server manager. Production readiness is not a claim; it is a rerunnable script. — Jeremy Anderson · info@dcos.net · dcos.net