Thicket/BLOG.md

191 lines
8.5 KiB
Markdown

# Growing a Thicket
*Building a super-ingest + RAG console that treats a document library
like a living ecosystem, not a filing cabinet.*
**Jeremy Anderson** — [dcos.net](https://dcos.net) — info@dcos.net
September 2026
---
A few weeks ago a friend's Gemini session produced ~270 lines of
Python titled "brain ingest": walk a folder, extract text from PDFs
and EPUBs, write Markdown notes into an Obsidian vault, chunk, embed,
and pour the results into Qdrant. It was a good sketch. It also had
four bugs I cared about, an import surface that required every heavy
dependency up front, and a LightRAG binding pinned to one release of
a fast-moving API.
We turned it into **Thicket**: a super-ingest + retrieval console with
the same brushed-aluminum, amber-LED MMD3 interface as my transcoder
[OpenTranscode](http://git.dcos.net/dcosnet/OpenTranscode) — because
the tool you actually use is the tool you actually enjoy opening.
The name is the design brief. A thicket is dense, self-connected, and
grows on its own. Feed it documents; it grows a thicket.
## The architecture bet: one engine, two drivers
The single most important decision was making the pipeline core
Qt-free. `IngestPipeline` knows nothing about widgets: it walks a
stage table and reports progress through four plain callables.
Two drivers share it:
- a **QThread worker** that wires those callables to Qt signals, and
- a **headless CLI** (`thicket --ingest ... --vault ...`) that wires
them to `print`, works over SSH, and fits in a cron line.
One behavior change lands in both drivers at once. There is no "GUI
logic" to port or keep in sync — that category of bug cannot exist.
## Stage tables instead of branch nests
Every per-document decision is data, not control flow. The pipeline
stage order is a tuple of `(status, gate, runner)`; file-type
extraction is a dict keyed by extension; readiness checks before an
ingest run are a table where the first enabled-but-not-ready entry
wins. Adding a stage, a format, or a gate means adding an entry —
never a new `if` ladder. The QA pass measured this concretely: the
extractor went from a five-branch dispatch to a one-line dict lookup.
## Lazy by default, graceful by contract
The console starts on a bare system. Every heavy import — FastEmbed,
the Qdrant client, EbookLib, LightRAG — happens at point of use, and
a background probe reports module and service readiness into the LED
status strip. Stages that cannot run are blocked at INGEST with the
exact reason and the exact fix; nothing crashes, nothing silently
no-ops.
Where a choice of paths exists, the code steps down the chain
explicitly, best option first: Qdrant search uses `query_points` and
falls to `search` on older clients; the LightRAG Ollama binding
resolves across the package's API generations; podman networking that
cannot create a tap device gets `--network=host`. Unix philosophy in
practice: try the clean path, degrade to the simple one, keep going.
## Re-ingestion is an overwrite
The subtlest correctness property in the system: **the point count
after N re-ingests equals the count after the first.** Deterministic
point IDs (MD5-UUID of `title|path|index`) plus a delete-by-filter
before every upsert make editing a source and re-ingesting it an
exact replacement — a shrunken document leaves no stale tail chunks.
We verify this in the live QA run: 4 points, re-ingest, still 4
points, same retrieval rankings.
The vault side holds the same invariant: re-ingesting a source
refreshes its note in place, and a *different* document with the same
title claims a digest-suffixed sibling rather than clobbering it.
## Local embeddings, on purpose
FastEmbed runs ONNX on your own cores; nothing leaves the machine.
Retrieval quality still respects the model's contract — BGE models
want an instruction prefix on the *query* side and bare passages on
the *document* side, so Thicket prefixes queries only. It is the kind
of detail that silently costs you ten points of relevance when a
sketch gets it wrong.
## The QA pass that shaped the code
Before calling it production-ready, the codebase went through a
five-hat review — senior QA, Linux engineer, architect, admin, and
devops PM. What changed:
- **Table-driven dispatch everywhere** it beat nested conditionals
(stage table, extension table, readiness table, stage-color map).
- **Loops reduced to comprehensions and `next()`** where iteration
was bookkeeping; explicit loops remain only where iteration *is*
the semantics (chunk word windows, queue walks).
- **Bounded reads** on files we do not own (frontmatter collision
checks read at most 32 lines).
- **One decisive failure path per scope** — per-file isolation in the
pipeline, a single report-and-disable path in the retrieval worker.
- Comments state invariants and contracts. Version-history narration
does not survive review; the code reads like decisions, not like an
argument with itself.
Standards kept in view: PEP 8 throughout, SEI CERT practices (bounded
I/O, precise exception scope — the Qdrant step-down catches
`AttributeError`, not the world), MISRA-style bounded structured
control flow, and POSIX assumptions (paths via `pathlib`, no platform
branches, systemd/cron-friendly headless mode).
## Live-fire verification
The release gate was not the test suite alone (26 tests, no services
required) but a live run: podman Qdrant up, three documents through
the full GUI pipeline, two semantic queries returning correctly
ranked hits, an idempotent re-ingest, and a dry-run probe reporting
every module and service green. Screenshot or it didn't happen — the
console looks the part too: knobs for chunk size, overlap, and
Top-K; an LED queue table tracking every file's stage; a phosphor
log; a retrieval strip at the bottom.
## What's next
The graph stage (LightRAG over Ollama) is wired and gating on
readiness, but it is deliberately optional — entity extraction is an
LLM pass per document, and the vault + vector stages already answer
the daily question: *where did I read that?*
The thicket grows. Pull it, feed it a shelf of books, and see what
surfaces.
---
**Thicket** — AGPL-3.0-or-later — [git.dcos.net/dcosnet/Thicket](http://git.dcos.net/dcosnet/Thicket)
Jeremy Anderson — info@dcos.net — [dcos.net](https://dcos.net)
---
## Addendum — v1.8: the thicket grows roots
*September 2026, after the first full functionality matrix.*
The sketch became a workstation tool. What changed since the first
essay:
**Destinations, not stages.** The original three-stage line (notes →
Qdrant → LightRAG) became a destination model: `obsidian` for
notes-only, or any of ten open-source vector stores — qdrant, chroma,
lancedb, faiss, milvus, weaviate, pgvector, duckdb, sqlite-vec,
mariadb — one payload schema, one exact-replacement contract, cosine
scores identical to four decimals across all ten. The console's OUT
line resolves the real destination live; the status footer follows the
target selector keystroke by keystroke.
**Sources are never deleted.** The dangerous delete-after-verify
checkbox from the transcoder lineage is gone. Verified sources move to
an archive directory and bzip2-compress in place — the incoming tree
stays clean, the archive always decompresses.
**The corpus is technical.** Fenced code blocks became atomic
language-tagged chunks; prose became paragraph-aligned; shebangs
stopped masquerading as headings; twenty-one source/config extensions
ingest as listings; the default embedder became jina-code, and
"force browsers to refuse plain http" retrieves the nginx block where
HSTS actually lives.
**The workstation has a filesystem.** /mnt/AI/corpus/{cold,hot} is now
the default corpus flow — cold in, hot brain — with ~/$VAR expansion
everywhere and THICKET_AI_ROOT for other hosts.
**Two interaction modes.** SEARCH embeds locally; ASK turns natural
language into read-only SQL over the Postgres/MariaDB corpora through
Vanna 2's agent API and the same Ollama selector. "Which document has
the most chunks?" is now a console question.
**The matrix.** `scripts/func_test.py` runs eighteen live checks —
every destination with idempotency and retrieval assertions, both
graph engines, both archives, ask on both SQL targets. Its first run
caught six real bugs, including a subtle one: milvus-lite's embedded
server keeps the database file lock after client close, so a second
process gets `[Errno 11]` — fixed with an explicit `close()` lifecycle
that releases the server manager. Production readiness is not a claim;
it is a rerunnable script.
— Jeremy Anderson · info@dcos.net · dcos.net