Thicket/BLOG.md

8.5 KiB

Growing a Thicket

Building a super-ingest + RAG console that treats a document library like a living ecosystem, not a filing cabinet.

Jeremy Anderson — dcos.net — info@dcos.net September 2026


A few weeks ago a friend's Gemini session produced ~270 lines of Python titled "brain ingest": walk a folder, extract text from PDFs and EPUBs, write Markdown notes into an Obsidian vault, chunk, embed, and pour the results into Qdrant. It was a good sketch. It also had four bugs I cared about, an import surface that required every heavy dependency up front, and a LightRAG binding pinned to one release of a fast-moving API.

We turned it into Thicket: a super-ingest + retrieval console with the same brushed-aluminum, amber-LED MMD3 interface as my transcoder OpenTranscode — because the tool you actually use is the tool you actually enjoy opening.

The name is the design brief. A thicket is dense, self-connected, and grows on its own. Feed it documents; it grows a thicket.

The architecture bet: one engine, two drivers

The single most important decision was making the pipeline core Qt-free. IngestPipeline knows nothing about widgets: it walks a stage table and reports progress through four plain callables.

Two drivers share it:

  • a QThread worker that wires those callables to Qt signals, and
  • a headless CLI (thicket --ingest ... --vault ...) that wires them to print, works over SSH, and fits in a cron line.

One behavior change lands in both drivers at once. There is no "GUI logic" to port or keep in sync — that category of bug cannot exist.

Stage tables instead of branch nests

Every per-document decision is data, not control flow. The pipeline stage order is a tuple of (status, gate, runner); file-type extraction is a dict keyed by extension; readiness checks before an ingest run are a table where the first enabled-but-not-ready entry wins. Adding a stage, a format, or a gate means adding an entry — never a new if ladder. The QA pass measured this concretely: the extractor went from a five-branch dispatch to a one-line dict lookup.

Lazy by default, graceful by contract

The console starts on a bare system. Every heavy import — FastEmbed, the Qdrant client, EbookLib, LightRAG — happens at point of use, and a background probe reports module and service readiness into the LED status strip. Stages that cannot run are blocked at INGEST with the exact reason and the exact fix; nothing crashes, nothing silently no-ops.

Where a choice of paths exists, the code steps down the chain explicitly, best option first: Qdrant search uses query_points and falls to search on older clients; the LightRAG Ollama binding resolves across the package's API generations; podman networking that cannot create a tap device gets --network=host. Unix philosophy in practice: try the clean path, degrade to the simple one, keep going.

Re-ingestion is an overwrite

The subtlest correctness property in the system: the point count after N re-ingests equals the count after the first. Deterministic point IDs (MD5-UUID of title|path|index) plus a delete-by-filter before every upsert make editing a source and re-ingesting it an exact replacement — a shrunken document leaves no stale tail chunks. We verify this in the live QA run: 4 points, re-ingest, still 4 points, same retrieval rankings.

The vault side holds the same invariant: re-ingesting a source refreshes its note in place, and a different document with the same title claims a digest-suffixed sibling rather than clobbering it.

Local embeddings, on purpose

FastEmbed runs ONNX on your own cores; nothing leaves the machine. Retrieval quality still respects the model's contract — BGE models want an instruction prefix on the query side and bare passages on the document side, so Thicket prefixes queries only. It is the kind of detail that silently costs you ten points of relevance when a sketch gets it wrong.

The QA pass that shaped the code

Before calling it production-ready, the codebase went through a five-hat review — senior QA, Linux engineer, architect, admin, and devops PM. What changed:

  • Table-driven dispatch everywhere it beat nested conditionals (stage table, extension table, readiness table, stage-color map).
  • Loops reduced to comprehensions and next() where iteration was bookkeeping; explicit loops remain only where iteration is the semantics (chunk word windows, queue walks).
  • Bounded reads on files we do not own (frontmatter collision checks read at most 32 lines).
  • One decisive failure path per scope — per-file isolation in the pipeline, a single report-and-disable path in the retrieval worker.
  • Comments state invariants and contracts. Version-history narration does not survive review; the code reads like decisions, not like an argument with itself.

Standards kept in view: PEP 8 throughout, SEI CERT practices (bounded I/O, precise exception scope — the Qdrant step-down catches AttributeError, not the world), MISRA-style bounded structured control flow, and POSIX assumptions (paths via pathlib, no platform branches, systemd/cron-friendly headless mode).

Live-fire verification

The release gate was not the test suite alone (26 tests, no services required) but a live run: podman Qdrant up, three documents through the full GUI pipeline, two semantic queries returning correctly ranked hits, an idempotent re-ingest, and a dry-run probe reporting every module and service green. Screenshot or it didn't happen — the console looks the part too: knobs for chunk size, overlap, and Top-K; an LED queue table tracking every file's stage; a phosphor log; a retrieval strip at the bottom.

What's next

The graph stage (LightRAG over Ollama) is wired and gating on readiness, but it is deliberately optional — entity extraction is an LLM pass per document, and the vault + vector stages already answer the daily question: where did I read that?

The thicket grows. Pull it, feed it a shelf of books, and see what surfaces.


Thicket — AGPL-3.0-or-later — git.dcos.net/dcosnet/Thicket Jeremy Anderson — info@dcos.net — dcos.net


Addendum — v1.8: the thicket grows roots

September 2026, after the first full functionality matrix.

The sketch became a workstation tool. What changed since the first essay:

Destinations, not stages. The original three-stage line (notes → Qdrant → LightRAG) became a destination model: obsidian for notes-only, or any of ten open-source vector stores — qdrant, chroma, lancedb, faiss, milvus, weaviate, pgvector, duckdb, sqlite-vec, mariadb — one payload schema, one exact-replacement contract, cosine scores identical to four decimals across all ten. The console's OUT line resolves the real destination live; the status footer follows the target selector keystroke by keystroke.

Sources are never deleted. The dangerous delete-after-verify checkbox from the transcoder lineage is gone. Verified sources move to an archive directory and bzip2-compress in place — the incoming tree stays clean, the archive always decompresses.

The corpus is technical. Fenced code blocks became atomic language-tagged chunks; prose became paragraph-aligned; shebangs stopped masquerading as headings; twenty-one source/config extensions ingest as listings; the default embedder became jina-code, and "force browsers to refuse plain http" retrieves the nginx block where HSTS actually lives.

The workstation has a filesystem. /mnt/AI/corpus/{cold,hot} is now the default corpus flow — cold in, hot brain — with ~/$VAR expansion everywhere and THICKET_AI_ROOT for other hosts.

Two interaction modes. SEARCH embeds locally; ASK turns natural language into read-only SQL over the Postgres/MariaDB corpora through Vanna 2's agent API and the same Ollama selector. "Which document has the most chunks?" is now a console question.

The matrix. scripts/func_test.py runs eighteen live checks — every destination with idempotency and retrieval assertions, both graph engines, both archives, ask on both SQL targets. Its first run caught six real bugs, including a subtle one: milvus-lite's embedded server keeps the database file lock after client close, so a second process gets [Errno 11] — fixed with an explicit close() lifecycle that releases the server manager. Production readiness is not a claim; it is a rerunnable script.

— Jeremy Anderson · info@dcos.net · dcos.net