Super-ingest + RAG console — feed it documents; it grows a thicket: dense, interconnected, searchable.
Go to file
Jeremy Anderson 84eac17dc2 Thicket - Super-Injest: ScreenShot 2026-09-28 13:55:22 -04:00
scripts Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
tests Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
thicket Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
.gitignore Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
BLOG.md Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
CHANGELOG.md Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
LICENSE Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
QUICKSTART.md Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
README.md Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
pyproject.toml Thicket - Super-Injest 2026-09-28 13:48:27 -04:00
ss.png Thicket - Super-Injest: ScreenShot 2026-09-28 13:55:22 -04:00
thicket.py Thicket - Super-Injest 2026-09-28 13:48:27 -04:00

README.md

Thicket

Super-ingest + RAG console — feed it documents; it grows a thicket: dense, interconnected, searchable. Import PDF / EPUB / Markdown / plain text / source-and-config files into a destination of your choice, wrapped in the same retro-futuristic MMD3 console UI as OpenTranscode (brushed aluminum, amber LEDs, green phosphor log, rotary knobs).

                    ┌─▶ Obsidian vault ──▶ MinIO bucket*  ┐
documents ──▶ notes │    (Markdown +         (s3:// URIs)  ├─▶ ingested-archive/
                    │    YAML frontmatter)                │     (bz2, never deleted)
                    └─▶ vector index ──▶ knowledge graph* ─┘
                         (10 targets,      (LightRAG or
                          local FastEmbed)  Graphify, Ollama)

        TARGET picks the destination: obsidian = notes only,
        any vector store = notes + that store.      * optional stages
Doc What it is
QUICKSTART.md clone → services → first ingest → first retrieval
BLOG.md design essay + v1.8 addendum
LICENSE AGPL-3.0-or-later, full text

Destinations and stages

Vault notes always run — they are the product. Everything else stacks:

Piece What it does Needs
Vault notes (always on) Normalized Markdown notes (YAML frontmatter: title, source file, optional source_uri, ingest timestamp, brain/ingested + source/<ext> tags) into the notes dir (default <vault>/Ingested_Brain/). PDFs get ## Page N headings; EPUBs are chapter-split; code files render as language-tagged listings. TARGET=obsidian = notes only. python-slugify
Vector index (a TARGET away) Structure-preserving contextual chunking, local ONNX embeddings (FastEmbed), exact-replacement indexing — the point count after N re-ingests equals the count after the first. Ten interchangeable targets (below). fastembed + the target's library
Knowledge graph (optional) LightRAG (merged entity graph in <vault>/.lightrag, one Ollama pass per document) or Graphify (batch graphify extract --backend ollama → graph.json, GRAPH_REPORT.md, interactive graph.html in <vault>/.graphify). engine package, Ollama
MinIO archive (optional) Originals uploaded to an S3-compatible bucket (key = input-tree path, idempotent); the note records s3://bucket/key. Open-source MinIO has no vector API — it is an archive stage, not a vector target. minio, MinIO service
Source archive (optional) Verified sources move out of the incoming tree into ingested-archive/ and are bzip2-compressed (streaming, level 9). Sources are never deleted; the archive always holds a decompressible original. nothing

Sources are never deleted — by design, not by flag. Guards (MAX MB, skip-unchanged) skip files entirely; skips never archive.

Vector targets

All ten share one payload schema (document_title, section_header, content, chunk_kind, lang, …), one retrieval strip, and the same exact-replacement semantics — cosine scores agree to four decimals across targets (verified by the live matrix).

Target Mode Library Service needed Install extra
qdrant (default) service qdrant-client Qdrant container .[ingest]
chroma embedded chromadb none .[chroma]
lancedb embedded lancedb none .[lancedb]
faiss file faiss-cpu none .[faiss]
milvus embedded (Lite) pymilvus[milvus_lite] none .[milvus]
weaviate service weaviate-client Weaviate container (8080/50051) .[weaviate]
pgvector service psycopg + pgvector Postgres + pgvector (libpq env: PG*) .[pgvector]
duckdb embedded duckdb (+vss index, steps down to exact scan) none .[duckdb]
sqlitevec embedded sqlite-vec (vec0 tables) none .[sqlitevec]
mariadb service PyMySQL MariaDB 11.7+ (env: MARIADB_*) .[mariadb]

Embedded targets keep data under <vault>/.thicket/<target>/ — a vault is one portable tree. Service targets read Unix-standard environments: PGHOST/PGUSER/PGPASSWORD/PGDATABASE (or PGDSN), MARIADB_HOST/MARIADB_USER/MARIADB_PASSWORD/MARIADB_DATABASE. pip install -e ".[targets]" installs every target extra at once.

Upgrading from 1.1.x: points indexed before 1.2.0 lack the internal doc_key; the first re-ingest of each document cleans them up automatically — or start a fresh collection name.

AI filesystem layout awareness

When the canonical /mnt/AI tree exists, Thicket adopts its corpus flow as defaults — zero configuration:

Canonical path Thicket role
/mnt/AI/corpus/cold IN — incoming raw documents
/mnt/AI/corpus/hot VAULT — the active brain: notes + .thicket/ vector data + graphs in one tree
/mnt/AI/corpus/books standing library (reported by the probe; point IN at it to ingest)
/mnt/AI/corpus/archive destination of the source-archive (bz2) stage
/mnt/AI/backends/thicket.env connection profiles (600) — fills PG* / MARIADB_* / MINIO_* gaps; the real environment always wins

Without the layout, home-directory defaults apply — behavior identical. Override the root with THICKET_AI_ROOT. Every path field accepts ~ and $VAR references, resolved through one shared expansion point in the GUI, CLI, and core.

Technical corpora (code, configuration, policy)

Ingestion preserves what makes technical documents useful:

  • Fenced code blocks are atomic — never split mid-listing, never merged with prose, indentation and blank lines verbatim; the language tag travels in the chunk context (| code:python) and the payload (chunk_kind, lang). Oversized blocks window by lines with overlap.
  • Prose chunks are paragraph-aligned — lists, commands, and tables keep their line structure in stored content.
  • ATX headings require # + whitespace — shebangs and #comments never masquerade as section headers.
  • Source files ingest directly: .py .sh .bash .zsh .yaml .yml .toml .ini .conf .cfg .json .sql .rs .go .c .h .cpp .js .ts .tf .nix are wrapped as language-tagged listings.
  • Code-strong default embedder: jinaai/jina-embeddings-v2-base-code (English + code, 768-dim, 8k context); curated alternatives in the EMBED selector. Dimension is fixed per collection — switching models means a new collection name.
  • Graph tip: directories of real code want the graphify engine — tree-sitter gives it per-symbol structure no prose pass can match.

The console

  • Paths panel — IN, VAULT, and a live OUT line resolving the actual destination per TARGET (notes dir, .thicket/ data path, or service/collection URI), re-resolved on every edit
  • Pipeline panel — TARGET is the destination (obsidian or one of ten vector stores) with collection/host/port and a per-target connection hint line; stage toggles (graph, MinIO, bz2 archive); EMBED / OLLAMA LLM / GRAPH ENGINE selectors; FILTER, NOTES DIR, MAX MB, Skip-unchanged
  • Document queue — LED matrix table with per-file stage (QUEUED → EXTRACT → MINIO → VAULT → INDEX → GRAPH → ARCHIVE → DONE / SKIP / ERROR), color-coded, with detail column
  • Knobs — chunk size, chunk overlap, retrieval Top-K
  • Retrieval strip — SEARCH (semantic across the collection) and ASK (SQL) (natural language over SQL-backed targets via Vanna 2 + the Ollama LLM); results land in the log
  • Status footer — live target + service readiness, updated the instant TARGET changes
  • Transport — > INGEST, [] STOP (cooperative), ~~ SCAN QUEUE (preview), ? ABOUT
  • Probe — every dependency and service reported at startup; a stage that cannot run is blocked at INGEST with the exact reason and fix

Install

One venv, always: /mnt/AI/runtime/thicket-venv is the single environment. The source-anchored bootstrap mirrors every dependency as a git checkout under /mnt/AI/distfiles/git/, builds them into that same venv, and drops a launcher in /mnt/AI/tools/bin:

scripts/bootstrap_sources.sh                  # add --force-source for
/mnt/AI/tools/bin/thicket --dry-run           # native builds from git

The venv is created and owned by the bootstrap at /mnt/AI/runtime/thicket-venv — nothing is written inside the project checkout (.gitignore keeps it archive-clean). Plain pip path:

# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
true  # (scripts/bootstrap_sources.sh creates it)
/mnt/AI/runtime/thicket-venv/bin/pip install -e .                    # console only (PySide6)
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]"          # + parsers, FastEmbed, qdrant
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]"         # + every vector target
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph,graphify]"  # + both graph engines
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ask,minio]"       # + Vanna ask, MinIO archive

Services (only for the stages that need them):

podman run -d --name thicket-qdrant -p 6333:6333 \
    -v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
ollama pull llama3.1 && ollama pull nomic-embed-text

The first ingest downloads the embedding model locally (~160 MB for the default jina-code embedder).

Usage

One command owns every mode — no flags loads the console:

python thicket.py            # the console — also: ./thicket.py

Headless (same core, no Qt — SSH / cron friendly):

python thicket.py --ingest /mnt/AI/corpus/cold --vault /mnt/AI/corpus/hot
python thicket.py --ingest ~/Books --vault ~/Vault --target obsidian   # notes only
python thicket.py --ingest ~/Books --vault ~/Vault --target chroma     # embedded, no service
python thicket.py --ingest ~/Books --vault ~/Vault --lightrag --graph-engine graphify \
    --ollama-llm llama3.1:latest
python thicket.py --ingest ~/Books --vault ~/Vault --skip-unchanged --max-mb 200
python thicket.py --ingest ~/Books --vault ~/Vault --archive --minio
python thicket.py --ask "which document has the most chunks?" --target pgvector

Probe & report: python thicket.py --dry-run · --version

Three interchangeable entry points share one environment: the thicket command (installed to ~/.local/bin by the bootstrap — works from any directory), /mnt/AI/tools/bin/thicket, and /mnt/AI/runtime/thicket-venv/bin/python thicket.py. A bare python thicket.py uses the system interpreter, which cannot see the venv — that is Python's rule, not a Thicket setting.

Architecture

thicket/
├── thicket.py               # one-command launcher (GUI default, flags pass through)
├── cli.py                   # version / dry-run / headless ingest / ask / GUI
├── layout.py                # /mnt/AI taxonomy awareness + ~/$VAR expansion
├── extractors.py            # PDF / EPUB / MD / TXT / code-config -> (title, text)
├── vault_writer.py          # notes: frontmatter, escaping, collisions, source_uri
├── chunker.py               # structure-preserving chunking (fences, langs, paragraphs)
├── embedder.py              # FastEmbed wrapper (curated catalog, jina-code default)
├── vector_stores.py         # registry + 10 targets, close() lifecycle
├── qdrant_store.py          # qdrant target
├── graph_store.py           # graph engines: LightRAG (async lifecycle) + Graphify
├── minio_archive.py         # S3 object archive stage
├── ask_vanna.py             # Vanna 2 agent: natural-language SQL over SQL targets
├── pipeline_core.py         # the stage table — shared by GUI worker and CLI
├── pipeline_worker.py       # QThreads: probe / ingest / search / ask
├── env_probe.py             # module + service readiness (incl. liveness tables)
├── ui_theme.py              # MMD3 QSS (OpenTranscode visual lineage)
├── ui_window.py             # ThicketWindow + launch_gui()
└── widgets/radio_knob.py

Design invariants:

  • Qt-free core. The QThread worker and the headless CLI drive the same IngestPipeline; one behavior change lands in both at once.
  • Table-driven dispatch. Stage order, file-type routing, target registry, readiness gating, UI stage colors — data tables, not branch nests.
  • Lazy heavy imports + graceful degradation. Every heavy dependency loads at point of use; missing pieces report themselves and block only the stage that needs them.
  • Idempotent re-ingest. Deterministic IDs plus delete-by-doc_key matching both key generations) make re-ingesting an edited source an exact replacement — proven per-target by the live matrix.
  • Per-file isolation. One broken document logs an ERROR; the queue moves on. Postgres transactions roll back per document.
  • Explicit resource lifecycle. Every store implements close(); Milvus Lite's embedded server is released so the next process can open the database.
  • Step-down chains. Qdrant query_points→search; LightRAG bindings across API generations; DuckDB vss→exact scan; podman rootless→--network=host.

Development

/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[dev,targets,graph,graphify,ask,minio]"
.venv/bin/pytest                              # 62 tests, no services needed
QT_QPA_PLATFORM=offscreen \
  .venv/bin/python scripts/smoke_gui.py       # headless GUI smoke + screenshot
.venv/bin/python scripts/func_test.py         # LIVE matrix: every destination,
                                              # engine, archive, ask — 18 checks
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run         # live readiness report

Standards applied: PEP 8; SEI CERT practices (bounded reads, precise exception scope); MISRA-style bounded, structured control flow; POSIX assumptions (pathlib, no platform branches, cron-safe headless mode).

License

AGPL-3.0-or-later — Jeremy Anderson · info@dcos.net · dcos.net · 2026. See LICENSE for the full text.

Invoked (not bundled) components carry their own licenses: PySide6 (LGPL-3.0), pypdf (BSD), EbookLib (AGPL-3.0), BeautifulSoup (MIT), python-slugify (MIT), FastEmbed (Apache-2.0), Qdrant (Apache-2.0), Chroma (Apache-2.0), LanceDB (Apache-2.0), FAISS (MIT), pymilvus (Apache-2.0), weaviate-client (BSD-3), psycopg (LGPL-3.0), pgvector (PostgreSQL), PyMySQL (MIT), sqlite-vec (MIT), LightRAG (MIT), Graphify (Apache-2.0/MIT), Vanna (MIT), MinIO client (Apache-2.0), Ollama (MIT).