|
|
||
|---|---|---|
| scripts | ||
| tests | ||
| thicket | ||
| .gitignore | ||
| BLOG.md | ||
| CHANGELOG.md | ||
| LICENSE | ||
| QUICKSTART.md | ||
| README.md | ||
| pyproject.toml | ||
| thicket.py | ||
README.md
Thicket
Super-ingest + RAG console — feed it documents; it grows a thicket: dense, interconnected, searchable. Import PDF / EPUB / Markdown / plain text / source-and-config files into a destination of your choice, wrapped in the same retro-futuristic MMD3 console UI as OpenTranscode (brushed aluminum, amber LEDs, green phosphor log, rotary knobs).
┌─▶ Obsidian vault ──▶ MinIO bucket* ┐
documents ──▶ notes │ (Markdown + (s3:// URIs) ├─▶ ingested-archive/
│ YAML frontmatter) │ (bz2, never deleted)
└─▶ vector index ──▶ knowledge graph* ─┘
(10 targets, (LightRAG or
local FastEmbed) Graphify, Ollama)
TARGET picks the destination: obsidian = notes only,
any vector store = notes + that store. * optional stages
| Doc | What it is |
|---|---|
| QUICKSTART.md | clone → services → first ingest → first retrieval |
| BLOG.md | design essay + v1.8 addendum |
| LICENSE | AGPL-3.0-or-later, full text |
Destinations and stages
Vault notes always run — they are the product. Everything else stacks:
| Piece | What it does | Needs |
|---|---|---|
| Vault notes (always on) | Normalized Markdown notes (YAML frontmatter: title, source file, optional source_uri, ingest timestamp, brain/ingested + source/<ext> tags) into the notes dir (default <vault>/Ingested_Brain/). PDFs get ## Page N headings; EPUBs are chapter-split; code files render as language-tagged listings. TARGET=obsidian = notes only. |
python-slugify |
| Vector index (a TARGET away) | Structure-preserving contextual chunking, local ONNX embeddings (FastEmbed), exact-replacement indexing — the point count after N re-ingests equals the count after the first. Ten interchangeable targets (below). | fastembed + the target's library |
| Knowledge graph (optional) | LightRAG (merged entity graph in <vault>/.lightrag, one Ollama pass per document) or Graphify (batch graphify extract --backend ollama → graph.json, GRAPH_REPORT.md, interactive graph.html in <vault>/.graphify). |
engine package, Ollama |
| MinIO archive (optional) | Originals uploaded to an S3-compatible bucket (key = input-tree path, idempotent); the note records s3://bucket/key. Open-source MinIO has no vector API — it is an archive stage, not a vector target. |
minio, MinIO service |
| Source archive (optional) | Verified sources move out of the incoming tree into ingested-archive/ and are bzip2-compressed (streaming, level 9). Sources are never deleted; the archive always holds a decompressible original. |
nothing |
Sources are never deleted — by design, not by flag. Guards (MAX MB, skip-unchanged) skip files entirely; skips never archive.
Vector targets
All ten share one payload schema (document_title, section_header,
content, chunk_kind, lang, …), one retrieval strip, and the same
exact-replacement semantics — cosine scores agree to four decimals
across targets (verified by the live matrix).
| Target | Mode | Library | Service needed | Install extra |
|---|---|---|---|---|
qdrant (default) |
service | qdrant-client |
Qdrant container | .[ingest] |
chroma |
embedded | chromadb |
none | .[chroma] |
lancedb |
embedded | lancedb |
none | .[lancedb] |
faiss |
file | faiss-cpu |
none | .[faiss] |
milvus |
embedded (Lite) | pymilvus[milvus_lite] |
none | .[milvus] |
weaviate |
service | weaviate-client |
Weaviate container (8080/50051) | .[weaviate] |
pgvector |
service | psycopg + pgvector |
Postgres + pgvector (libpq env: PG*) |
.[pgvector] |
duckdb |
embedded | duckdb (+vss index, steps down to exact scan) |
none | .[duckdb] |
sqlitevec |
embedded | sqlite-vec (vec0 tables) |
none | .[sqlitevec] |
mariadb |
service | PyMySQL |
MariaDB 11.7+ (env: MARIADB_*) |
.[mariadb] |
Embedded targets keep data under <vault>/.thicket/<target>/ — a vault
is one portable tree. Service targets read Unix-standard environments:
PGHOST/PGUSER/PGPASSWORD/PGDATABASE (or PGDSN),
MARIADB_HOST/MARIADB_USER/MARIADB_PASSWORD/MARIADB_DATABASE.
pip install -e ".[targets]" installs every target extra at once.
Upgrading from 1.1.x: points indexed before 1.2.0 lack the internal
doc_key; the first re-ingest of each document cleans them up
automatically — or start a fresh collection name.
AI filesystem layout awareness
When the canonical /mnt/AI tree exists, Thicket adopts its corpus flow
as defaults — zero configuration:
| Canonical path | Thicket role |
|---|---|
/mnt/AI/corpus/cold |
IN — incoming raw documents |
/mnt/AI/corpus/hot |
VAULT — the active brain: notes + .thicket/ vector data + graphs in one tree |
/mnt/AI/corpus/books |
standing library (reported by the probe; point IN at it to ingest) |
/mnt/AI/corpus/archive |
destination of the source-archive (bz2) stage |
/mnt/AI/backends/thicket.env |
connection profiles (600) — fills PG* / MARIADB_* / MINIO_* gaps; the real environment always wins |
Without the layout, home-directory defaults apply — behavior identical.
Override the root with THICKET_AI_ROOT. Every path field accepts ~
and $VAR references, resolved through one shared expansion point in
the GUI, CLI, and core.
Technical corpora (code, configuration, policy)
Ingestion preserves what makes technical documents useful:
- Fenced code blocks are atomic — never split mid-listing, never
merged with prose, indentation and blank lines verbatim; the language
tag travels in the chunk context (
| code:python) and the payload (chunk_kind,lang). Oversized blocks window by lines with overlap. - Prose chunks are paragraph-aligned — lists, commands, and tables keep their line structure in stored content.
- ATX headings require
#+ whitespace — shebangs and#commentsnever masquerade as section headers. - Source files ingest directly:
.py .sh .bash .zsh .yaml .yml .toml .ini .conf .cfg .json .sql .rs .go .c .h .cpp .js .ts .tf .nixare wrapped as language-tagged listings. - Code-strong default embedder:
jinaai/jina-embeddings-v2-base-code(English + code, 768-dim, 8k context); curated alternatives in the EMBED selector. Dimension is fixed per collection — switching models means a new collection name. - Graph tip: directories of real code want the
graphifyengine — tree-sitter gives it per-symbol structure no prose pass can match.
The console
- Paths panel — IN, VAULT, and a live OUT line resolving the
actual destination per TARGET (notes dir,
.thicket/data path, or service/collection URI), re-resolved on every edit - Pipeline panel — TARGET is the destination (
obsidianor one of ten vector stores) with collection/host/port and a per-target connection hint line; stage toggles (graph, MinIO, bz2 archive); EMBED / OLLAMA LLM / GRAPH ENGINE selectors; FILTER, NOTES DIR, MAX MB, Skip-unchanged - Document queue — LED matrix table with per-file stage (QUEUED → EXTRACT → MINIO → VAULT → INDEX → GRAPH → ARCHIVE → DONE / SKIP / ERROR), color-coded, with detail column
- Knobs — chunk size, chunk overlap, retrieval Top-K
- Retrieval strip — SEARCH (semantic across the collection) and ASK (SQL) (natural language over SQL-backed targets via Vanna 2 + the Ollama LLM); results land in the log
- Status footer — live target + service readiness, updated the instant TARGET changes
- Transport —
> INGEST,[] STOP(cooperative),~~ SCAN QUEUE(preview),? ABOUT - Probe — every dependency and service reported at startup; a stage that cannot run is blocked at INGEST with the exact reason and fix
Install
One venv, always: /mnt/AI/runtime/thicket-venv is the single environment.
The source-anchored bootstrap mirrors every dependency as a git
checkout under /mnt/AI/distfiles/git/, builds them into that same
venv, and drops a launcher in /mnt/AI/tools/bin:
scripts/bootstrap_sources.sh # add --force-source for
/mnt/AI/tools/bin/thicket --dry-run # native builds from git
The venv is created and owned by the bootstrap at
/mnt/AI/runtime/thicket-venv — nothing is written inside the project
checkout (.gitignore keeps it archive-clean). Plain pip path:
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
true # (scripts/bootstrap_sources.sh creates it)
/mnt/AI/runtime/thicket-venv/bin/pip install -e . # console only (PySide6)
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # + parsers, FastEmbed, qdrant
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # + every vector target
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph,graphify]" # + both graph engines
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ask,minio]" # + Vanna ask, MinIO archive
Services (only for the stages that need them):
podman run -d --name thicket-qdrant -p 6333:6333 \
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
ollama pull llama3.1 && ollama pull nomic-embed-text
The first ingest downloads the embedding model locally (~160 MB for the default jina-code embedder).
Usage
One command owns every mode — no flags loads the console:
python thicket.py # the console — also: ./thicket.py
Headless (same core, no Qt — SSH / cron friendly):
python thicket.py --ingest /mnt/AI/corpus/cold --vault /mnt/AI/corpus/hot
python thicket.py --ingest ~/Books --vault ~/Vault --target obsidian # notes only
python thicket.py --ingest ~/Books --vault ~/Vault --target chroma # embedded, no service
python thicket.py --ingest ~/Books --vault ~/Vault --lightrag --graph-engine graphify \
--ollama-llm llama3.1:latest
python thicket.py --ingest ~/Books --vault ~/Vault --skip-unchanged --max-mb 200
python thicket.py --ingest ~/Books --vault ~/Vault --archive --minio
python thicket.py --ask "which document has the most chunks?" --target pgvector
Probe & report: python thicket.py --dry-run · --version
Three interchangeable entry points share one environment: the thicket
command (installed to ~/.local/bin by the bootstrap — works from any
directory), /mnt/AI/tools/bin/thicket, and /mnt/AI/runtime/thicket-venv/bin/python thicket.py.
A bare python thicket.py uses the system interpreter, which cannot see
the venv — that is Python's rule, not a Thicket setting.
Architecture
thicket/
├── thicket.py # one-command launcher (GUI default, flags pass through)
├── cli.py # version / dry-run / headless ingest / ask / GUI
├── layout.py # /mnt/AI taxonomy awareness + ~/$VAR expansion
├── extractors.py # PDF / EPUB / MD / TXT / code-config -> (title, text)
├── vault_writer.py # notes: frontmatter, escaping, collisions, source_uri
├── chunker.py # structure-preserving chunking (fences, langs, paragraphs)
├── embedder.py # FastEmbed wrapper (curated catalog, jina-code default)
├── vector_stores.py # registry + 10 targets, close() lifecycle
├── qdrant_store.py # qdrant target
├── graph_store.py # graph engines: LightRAG (async lifecycle) + Graphify
├── minio_archive.py # S3 object archive stage
├── ask_vanna.py # Vanna 2 agent: natural-language SQL over SQL targets
├── pipeline_core.py # the stage table — shared by GUI worker and CLI
├── pipeline_worker.py # QThreads: probe / ingest / search / ask
├── env_probe.py # module + service readiness (incl. liveness tables)
├── ui_theme.py # MMD3 QSS (OpenTranscode visual lineage)
├── ui_window.py # ThicketWindow + launch_gui()
└── widgets/radio_knob.py
Design invariants:
- Qt-free core. The QThread worker and the headless CLI drive the
same
IngestPipeline; one behavior change lands in both at once. - Table-driven dispatch. Stage order, file-type routing, target registry, readiness gating, UI stage colors — data tables, not branch nests.
- Lazy heavy imports + graceful degradation. Every heavy dependency loads at point of use; missing pieces report themselves and block only the stage that needs them.
- Idempotent re-ingest. Deterministic IDs plus delete-by-
doc_keymatching both key generations) make re-ingesting an edited source an exact replacement — proven per-target by the live matrix. - Per-file isolation. One broken document logs an ERROR; the queue moves on. Postgres transactions roll back per document.
- Explicit resource lifecycle. Every store implements
close(); Milvus Lite's embedded server is released so the next process can open the database. - Step-down chains. Qdrant
query_points→search; LightRAG bindings across API generations; DuckDB vss→exact scan; podman rootless→--network=host.
Development
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[dev,targets,graph,graphify,ask,minio]"
.venv/bin/pytest # 62 tests, no services needed
QT_QPA_PLATFORM=offscreen \
.venv/bin/python scripts/smoke_gui.py # headless GUI smoke + screenshot
.venv/bin/python scripts/func_test.py # LIVE matrix: every destination,
# engine, archive, ask — 18 checks
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run # live readiness report
Standards applied: PEP 8; SEI CERT practices (bounded reads, precise
exception scope); MISRA-style bounded, structured control flow; POSIX
assumptions (pathlib, no platform branches, cron-safe headless mode).
License
AGPL-3.0-or-later — Jeremy Anderson · info@dcos.net · dcos.net · 2026. See LICENSE for the full text.
Invoked (not bundled) components carry their own licenses: PySide6 (LGPL-3.0), pypdf (BSD), EbookLib (AGPL-3.0), BeautifulSoup (MIT), python-slugify (MIT), FastEmbed (Apache-2.0), Qdrant (Apache-2.0), Chroma (Apache-2.0), LanceDB (Apache-2.0), FAISS (MIT), pymilvus (Apache-2.0), weaviate-client (BSD-3), psycopg (LGPL-3.0), pgvector (PostgreSQL), PyMySQL (MIT), sqlite-vec (MIT), LightRAG (MIT), Graphify (Apache-2.0/MIT), Vanna (MIT), MinIO client (Apache-2.0), Ollama (MIT).