# Thicket **Super-ingest + RAG console** — feed it documents; it grows a thicket: dense, interconnected, searchable. Import PDF / EPUB / Markdown / plain text / source-and-config files into a destination of your choice, wrapped in the same retro-futuristic MMD3 console UI as OpenTranscode (brushed aluminum, amber LEDs, green phosphor log, rotary knobs). ``` ┌─▶ Obsidian vault ──▶ MinIO bucket* ┐ documents ──▶ notes │ (Markdown + (s3:// URIs) ├─▶ ingested-archive/ │ YAML frontmatter) │ (bz2, never deleted) └─▶ vector index ──▶ knowledge graph* ─┘ (10 targets, (LightRAG or local FastEmbed) Graphify, Ollama) TARGET picks the destination: obsidian = notes only, any vector store = notes + that store. * optional stages ``` | Doc | What it is | |---|---| | [QUICKSTART.md](QUICKSTART.md) | clone → services → first ingest → first retrieval | | [BLOG.md](BLOG.md) | design essay + v1.8 addendum | | [LICENSE](LICENSE) | AGPL-3.0-or-later, full text | ## Destinations and stages Vault notes always run — they are the product. Everything else stacks: | Piece | What it does | Needs | |---|---|---| | **Vault notes** *(always on)* | Normalized Markdown notes (YAML frontmatter: title, source file, optional `source_uri`, ingest timestamp, `brain/ingested` + `source/` tags) into the notes dir (default `/Ingested_Brain/`). PDFs get `## Page N` headings; EPUBs are chapter-split; code files render as language-tagged listings. `TARGET=obsidian` = notes only. | `python-slugify` | | **Vector index** *(a TARGET away)* | Structure-preserving contextual chunking, local ONNX embeddings (FastEmbed), exact-replacement indexing — the point count after N re-ingests equals the count after the first. Ten interchangeable targets (below). | `fastembed` + the target's library | | **Knowledge graph** *(optional)* | **LightRAG** (merged entity graph in `/.lightrag`, one Ollama pass per document) or **Graphify** (batch `graphify extract --backend ollama` → `graph.json`, `GRAPH_REPORT.md`, interactive `graph.html` in `/.graphify`). | engine package, Ollama | | **MinIO archive** *(optional)* | Originals uploaded to an S3-compatible bucket (key = input-tree path, idempotent); the note records `s3://bucket/key`. Open-source MinIO has no vector API — it is an archive stage, not a vector target. | `minio`, MinIO service | | **Source archive** *(optional)* | Verified sources *move* out of the incoming tree into `ingested-archive/` and are bzip2-compressed (streaming, level 9). Sources are never deleted; the archive always holds a decompressible original. | nothing | Sources are never deleted — by design, not by flag. Guards (MAX MB, skip-unchanged) skip files entirely; skips never archive. ### Vector targets All ten share one payload schema (`document_title`, `section_header`, `content`, `chunk_kind`, `lang`, …), one retrieval strip, and the same exact-replacement semantics — cosine scores agree to four decimals across targets (verified by the live matrix). | Target | Mode | Library | Service needed | Install extra | |---|---|---|---|---| | `qdrant` *(default)* | service | `qdrant-client` | Qdrant container | `.[ingest]` | | `chroma` | embedded | `chromadb` | none | `.[chroma]` | | `lancedb` | embedded | `lancedb` | none | `.[lancedb]` | | `faiss` | file | `faiss-cpu` | none | `.[faiss]` | | `milvus` | embedded (Lite) | `pymilvus[milvus_lite]` | none | `.[milvus]` | | `weaviate` | service | `weaviate-client` | Weaviate container (8080/50051) | `.[weaviate]` | | `pgvector` | service | `psycopg` + `pgvector` | Postgres + pgvector (libpq env: `PG*`) | `.[pgvector]` | | `duckdb` | embedded | `duckdb` (+vss index, steps down to exact scan) | none | `.[duckdb]` | | `sqlitevec` | embedded | `sqlite-vec` (vec0 tables) | none | `.[sqlitevec]` | | `mariadb` | service | `PyMySQL` | MariaDB 11.7+ (env: `MARIADB_*`) | `.[mariadb]` | Embedded targets keep data under `/.thicket//` — a vault is one portable tree. Service targets read Unix-standard environments: `PGHOST`/`PGUSER`/`PGPASSWORD`/`PGDATABASE` (or `PGDSN`), `MARIADB_HOST`/`MARIADB_USER`/`MARIADB_PASSWORD`/`MARIADB_DATABASE`. `pip install -e ".[targets]"` installs every target extra at once. **Upgrading from 1.1.x:** points indexed before 1.2.0 lack the internal `doc_key`; the first re-ingest of each document cleans them up automatically — or start a fresh collection name. ## AI filesystem layout awareness When the canonical `/mnt/AI` tree exists, Thicket adopts its corpus flow as defaults — zero configuration: | Canonical path | Thicket role | |---|---| | `/mnt/AI/corpus/cold` | IN — incoming raw documents | | `/mnt/AI/corpus/hot` | VAULT — the active brain: notes + `.thicket/` vector data + graphs in one tree | | `/mnt/AI/corpus/books` | standing library (reported by the probe; point IN at it to ingest) | | `/mnt/AI/corpus/archive` | destination of the source-archive (bz2) stage | | `/mnt/AI/backends/thicket.env` | connection profiles (600) — fills `PG*` / `MARIADB_*` / `MINIO_*` gaps; the real environment always wins | Without the layout, home-directory defaults apply — behavior identical. Override the root with `THICKET_AI_ROOT`. Every path field accepts `~` and `$VAR` references, resolved through one shared expansion point in the GUI, CLI, and core. ## Technical corpora (code, configuration, policy) Ingestion preserves what makes technical documents useful: - **Fenced code blocks are atomic** — never split mid-listing, never merged with prose, indentation and blank lines verbatim; the language tag travels in the chunk context (`| code:python`) and the payload (`chunk_kind`, `lang`). Oversized blocks window by lines with overlap. - **Prose chunks are paragraph-aligned** — lists, commands, and tables keep their line structure in stored content. - **ATX headings require `#` + whitespace** — shebangs and `#comments` never masquerade as section headers. - **Source files ingest directly**: `.py .sh .bash .zsh .yaml .yml .toml .ini .conf .cfg .json .sql .rs .go .c .h .cpp .js .ts .tf .nix` are wrapped as language-tagged listings. - **Code-strong default embedder**: `jinaai/jina-embeddings-v2-base-code` (English + code, 768-dim, 8k context); curated alternatives in the EMBED selector. Dimension is fixed per collection — switching models means a new collection name. - **Graph tip**: directories of real code want the `graphify` engine — tree-sitter gives it per-symbol structure no prose pass can match. ## The console - **Paths panel** — IN, VAULT, and a live **OUT** line resolving the actual destination per TARGET (notes dir, `.thicket/` data path, or service/collection URI), re-resolved on every edit - **Pipeline panel** — TARGET is the destination (`obsidian` or one of ten vector stores) with collection/host/port and a per-target connection hint line; stage toggles (graph, MinIO, bz2 archive); EMBED / OLLAMA LLM / GRAPH ENGINE selectors; FILTER, NOTES DIR, MAX MB, Skip-unchanged - **Document queue** — LED matrix table with per-file stage (QUEUED → EXTRACT → MINIO → VAULT → INDEX → GRAPH → ARCHIVE → DONE / SKIP / ERROR), color-coded, with detail column - **Knobs** — chunk size, chunk overlap, retrieval Top-K - **Retrieval strip** — **SEARCH** (semantic across the collection) and **ASK (SQL)** (natural language over SQL-backed targets via Vanna 2 + the Ollama LLM); results land in the log - **Status footer** — live target + service readiness, updated the instant TARGET changes - **Transport** — `> INGEST`, `[] STOP` (cooperative), `~~ SCAN QUEUE` (preview), `? ABOUT` - **Probe** — every dependency and service reported at startup; a stage that cannot run is blocked at INGEST with the exact reason and fix ## Install **One venv, always**: `/mnt/AI/runtime/thicket-venv` is the single environment. The source-anchored bootstrap mirrors every dependency as a git checkout under `/mnt/AI/distfiles/git/`, builds them into that same venv, and drops a launcher in `/mnt/AI/tools/bin`: ```bash scripts/bootstrap_sources.sh # add --force-source for /mnt/AI/tools/bin/thicket --dry-run # native builds from git ``` The venv is created and owned by the bootstrap at `/mnt/AI/runtime/thicket-venv` — nothing is written inside the project checkout (`.gitignore` keeps it archive-clean). Plain pip path: ```bash # venv lives in the AI tree: /mnt/AI/runtime/thicket-venv true # (scripts/bootstrap_sources.sh creates it) /mnt/AI/runtime/thicket-venv/bin/pip install -e . # console only (PySide6) /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # + parsers, FastEmbed, qdrant /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # + every vector target /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph,graphify]" # + both graph engines /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ask,minio]" # + Vanna ask, MinIO archive ``` Services (only for the stages that need them): ```bash podman run -d --name thicket-qdrant -p 6333:6333 \ -v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant ollama pull llama3.1 && ollama pull nomic-embed-text ``` The first ingest downloads the embedding model locally (~160 MB for the default jina-code embedder). ## Usage One command owns every mode — no flags loads the console: ```bash python thicket.py # the console — also: ./thicket.py ``` **Headless** (same core, no Qt — SSH / cron friendly): ```bash python thicket.py --ingest /mnt/AI/corpus/cold --vault /mnt/AI/corpus/hot python thicket.py --ingest ~/Books --vault ~/Vault --target obsidian # notes only python thicket.py --ingest ~/Books --vault ~/Vault --target chroma # embedded, no service python thicket.py --ingest ~/Books --vault ~/Vault --lightrag --graph-engine graphify \ --ollama-llm llama3.1:latest python thicket.py --ingest ~/Books --vault ~/Vault --skip-unchanged --max-mb 200 python thicket.py --ingest ~/Books --vault ~/Vault --archive --minio python thicket.py --ask "which document has the most chunks?" --target pgvector ``` **Probe & report**: `python thicket.py --dry-run` · `--version` Three interchangeable entry points share one environment: the `thicket` command (installed to `~/.local/bin` by the bootstrap — works from any directory), `/mnt/AI/tools/bin/thicket`, and `/mnt/AI/runtime/thicket-venv/bin/python thicket.py`. A bare `python thicket.py` uses the system interpreter, which cannot see the venv — that is Python's rule, not a Thicket setting. ## Architecture ``` thicket/ ├── thicket.py # one-command launcher (GUI default, flags pass through) ├── cli.py # version / dry-run / headless ingest / ask / GUI ├── layout.py # /mnt/AI taxonomy awareness + ~/$VAR expansion ├── extractors.py # PDF / EPUB / MD / TXT / code-config -> (title, text) ├── vault_writer.py # notes: frontmatter, escaping, collisions, source_uri ├── chunker.py # structure-preserving chunking (fences, langs, paragraphs) ├── embedder.py # FastEmbed wrapper (curated catalog, jina-code default) ├── vector_stores.py # registry + 10 targets, close() lifecycle ├── qdrant_store.py # qdrant target ├── graph_store.py # graph engines: LightRAG (async lifecycle) + Graphify ├── minio_archive.py # S3 object archive stage ├── ask_vanna.py # Vanna 2 agent: natural-language SQL over SQL targets ├── pipeline_core.py # the stage table — shared by GUI worker and CLI ├── pipeline_worker.py # QThreads: probe / ingest / search / ask ├── env_probe.py # module + service readiness (incl. liveness tables) ├── ui_theme.py # MMD3 QSS (OpenTranscode visual lineage) ├── ui_window.py # ThicketWindow + launch_gui() └── widgets/radio_knob.py ``` Design invariants: - **Qt-free core.** The QThread worker and the headless CLI drive the same `IngestPipeline`; one behavior change lands in both at once. - **Table-driven dispatch.** Stage order, file-type routing, target registry, readiness gating, UI stage colors — data tables, not branch nests. - **Lazy heavy imports + graceful degradation.** Every heavy dependency loads at point of use; missing pieces report themselves and block only the stage that needs them. - **Idempotent re-ingest.** Deterministic IDs plus delete-by-`doc_key` matching both key generations) make re-ingesting an edited source an exact replacement — proven per-target by the live matrix. - **Per-file isolation.** One broken document logs an ERROR; the queue moves on. Postgres transactions roll back per document. - **Explicit resource lifecycle.** Every store implements `close()`; Milvus Lite's embedded server is released so the next process can open the database. - **Step-down chains.** Qdrant `query_points`→`search`; LightRAG bindings across API generations; DuckDB vss→exact scan; podman rootless→`--network=host`. ## Development ```bash /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[dev,targets,graph,graphify,ask,minio]" .venv/bin/pytest # 62 tests, no services needed QT_QPA_PLATFORM=offscreen \ .venv/bin/python scripts/smoke_gui.py # headless GUI smoke + screenshot .venv/bin/python scripts/func_test.py # LIVE matrix: every destination, # engine, archive, ask — 18 checks /mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run # live readiness report ``` Standards applied: PEP 8; SEI CERT practices (bounded reads, precise exception scope); MISRA-style bounded, structured control flow; POSIX assumptions (`pathlib`, no platform branches, cron-safe headless mode). ## License AGPL-3.0-or-later — Jeremy Anderson · info@dcos.net · [dcos.net](https://dcos.net) · 2026. See [LICENSE](LICENSE) for the full text. Invoked (not bundled) components carry their own licenses: PySide6 (LGPL-3.0), pypdf (BSD), EbookLib (AGPL-3.0), BeautifulSoup (MIT), python-slugify (MIT), FastEmbed (Apache-2.0), Qdrant (Apache-2.0), Chroma (Apache-2.0), LanceDB (Apache-2.0), FAISS (MIT), pymilvus (Apache-2.0), weaviate-client (BSD-3), psycopg (LGPL-3.0), pgvector (PostgreSQL), PyMySQL (MIT), sqlite-vec (MIT), LightRAG (MIT), Graphify (Apache-2.0/MIT), Vanna (MIT), MinIO client (Apache-2.0), Ollama (MIT).