8.5 KiB
Growing a Thicket
Building a super-ingest + RAG console that treats a document library like a living ecosystem, not a filing cabinet.
Jeremy Anderson — dcos.net — info@dcos.net September 2026
A few weeks ago a friend's Gemini session produced ~270 lines of Python titled "brain ingest": walk a folder, extract text from PDFs and EPUBs, write Markdown notes into an Obsidian vault, chunk, embed, and pour the results into Qdrant. It was a good sketch. It also had four bugs I cared about, an import surface that required every heavy dependency up front, and a LightRAG binding pinned to one release of a fast-moving API.
We turned it into Thicket: a super-ingest + retrieval console with the same brushed-aluminum, amber-LED MMD3 interface as my transcoder OpenTranscode — because the tool you actually use is the tool you actually enjoy opening.
The name is the design brief. A thicket is dense, self-connected, and grows on its own. Feed it documents; it grows a thicket.
The architecture bet: one engine, two drivers
The single most important decision was making the pipeline core
Qt-free. IngestPipeline knows nothing about widgets: it walks a
stage table and reports progress through four plain callables.
Two drivers share it:
- a QThread worker that wires those callables to Qt signals, and
- a headless CLI (
thicket --ingest ... --vault ...) that wires them toprint, works over SSH, and fits in a cron line.
One behavior change lands in both drivers at once. There is no "GUI logic" to port or keep in sync — that category of bug cannot exist.
Stage tables instead of branch nests
Every per-document decision is data, not control flow. The pipeline
stage order is a tuple of (status, gate, runner); file-type
extraction is a dict keyed by extension; readiness checks before an
ingest run are a table where the first enabled-but-not-ready entry
wins. Adding a stage, a format, or a gate means adding an entry —
never a new if ladder. The QA pass measured this concretely: the
extractor went from a five-branch dispatch to a one-line dict lookup.
Lazy by default, graceful by contract
The console starts on a bare system. Every heavy import — FastEmbed, the Qdrant client, EbookLib, LightRAG — happens at point of use, and a background probe reports module and service readiness into the LED status strip. Stages that cannot run are blocked at INGEST with the exact reason and the exact fix; nothing crashes, nothing silently no-ops.
Where a choice of paths exists, the code steps down the chain
explicitly, best option first: Qdrant search uses query_points and
falls to search on older clients; the LightRAG Ollama binding
resolves across the package's API generations; podman networking that
cannot create a tap device gets --network=host. Unix philosophy in
practice: try the clean path, degrade to the simple one, keep going.
Re-ingestion is an overwrite
The subtlest correctness property in the system: the point count
after N re-ingests equals the count after the first. Deterministic
point IDs (MD5-UUID of title|path|index) plus a delete-by-filter
before every upsert make editing a source and re-ingesting it an
exact replacement — a shrunken document leaves no stale tail chunks.
We verify this in the live QA run: 4 points, re-ingest, still 4
points, same retrieval rankings.
The vault side holds the same invariant: re-ingesting a source refreshes its note in place, and a different document with the same title claims a digest-suffixed sibling rather than clobbering it.
Local embeddings, on purpose
FastEmbed runs ONNX on your own cores; nothing leaves the machine. Retrieval quality still respects the model's contract — BGE models want an instruction prefix on the query side and bare passages on the document side, so Thicket prefixes queries only. It is the kind of detail that silently costs you ten points of relevance when a sketch gets it wrong.
The QA pass that shaped the code
Before calling it production-ready, the codebase went through a five-hat review — senior QA, Linux engineer, architect, admin, and devops PM. What changed:
- Table-driven dispatch everywhere it beat nested conditionals (stage table, extension table, readiness table, stage-color map).
- Loops reduced to comprehensions and
next()where iteration was bookkeeping; explicit loops remain only where iteration is the semantics (chunk word windows, queue walks). - Bounded reads on files we do not own (frontmatter collision checks read at most 32 lines).
- One decisive failure path per scope — per-file isolation in the pipeline, a single report-and-disable path in the retrieval worker.
- Comments state invariants and contracts. Version-history narration does not survive review; the code reads like decisions, not like an argument with itself.
Standards kept in view: PEP 8 throughout, SEI CERT practices (bounded
I/O, precise exception scope — the Qdrant step-down catches
AttributeError, not the world), MISRA-style bounded structured
control flow, and POSIX assumptions (paths via pathlib, no platform
branches, systemd/cron-friendly headless mode).
Live-fire verification
The release gate was not the test suite alone (26 tests, no services required) but a live run: podman Qdrant up, three documents through the full GUI pipeline, two semantic queries returning correctly ranked hits, an idempotent re-ingest, and a dry-run probe reporting every module and service green. Screenshot or it didn't happen — the console looks the part too: knobs for chunk size, overlap, and Top-K; an LED queue table tracking every file's stage; a phosphor log; a retrieval strip at the bottom.
What's next
The graph stage (LightRAG over Ollama) is wired and gating on readiness, but it is deliberately optional — entity extraction is an LLM pass per document, and the vault + vector stages already answer the daily question: where did I read that?
The thicket grows. Pull it, feed it a shelf of books, and see what surfaces.
Thicket — AGPL-3.0-or-later — git.dcos.net/dcosnet/Thicket Jeremy Anderson — info@dcos.net — dcos.net
Addendum — v1.8: the thicket grows roots
September 2026, after the first full functionality matrix.
The sketch became a workstation tool. What changed since the first essay:
Destinations, not stages. The original three-stage line (notes →
Qdrant → LightRAG) became a destination model: obsidian for
notes-only, or any of ten open-source vector stores — qdrant, chroma,
lancedb, faiss, milvus, weaviate, pgvector, duckdb, sqlite-vec,
mariadb — one payload schema, one exact-replacement contract, cosine
scores identical to four decimals across all ten. The console's OUT
line resolves the real destination live; the status footer follows the
target selector keystroke by keystroke.
Sources are never deleted. The dangerous delete-after-verify checkbox from the transcoder lineage is gone. Verified sources move to an archive directory and bzip2-compress in place — the incoming tree stays clean, the archive always decompresses.
The corpus is technical. Fenced code blocks became atomic language-tagged chunks; prose became paragraph-aligned; shebangs stopped masquerading as headings; twenty-one source/config extensions ingest as listings; the default embedder became jina-code, and "force browsers to refuse plain http" retrieves the nginx block where HSTS actually lives.
The workstation has a filesystem. /mnt/AI/corpus/{cold,hot} is now the default corpus flow — cold in, hot brain — with ~/$VAR expansion everywhere and THICKET_AI_ROOT for other hosts.
Two interaction modes. SEARCH embeds locally; ASK turns natural language into read-only SQL over the Postgres/MariaDB corpora through Vanna 2's agent API and the same Ollama selector. "Which document has the most chunks?" is now a console question.
The matrix. scripts/func_test.py runs eighteen live checks —
every destination with idempotency and retrieval assertions, both
graph engines, both archives, ask on both SQL targets. Its first run
caught six real bugs, including a subtle one: milvus-lite's embedded
server keeps the database file lock after client close, so a second
process gets [Errno 11] — fixed with an explicit close() lifecycle
that releases the server manager. Production readiness is not a claim;
it is a rerunnable script.
— Jeremy Anderson · info@dcos.net · dcos.net