Thicket/README.md

277 lines
14 KiB
Markdown

# Thicket
**Super-ingest + RAG console** — feed it documents; it grows a thicket:
dense, interconnected, searchable. Import PDF / EPUB / Markdown / plain
text / source-and-config files into a destination of your choice,
wrapped in the same retro-futuristic MMD3 console UI as OpenTranscode
(brushed aluminum, amber LEDs, green phosphor log, rotary knobs).
```
┌─▶ Obsidian vault ──▶ MinIO bucket* ┐
documents ──▶ notes │ (Markdown + (s3:// URIs) ├─▶ ingested-archive/
│ YAML frontmatter) │ (bz2, never deleted)
└─▶ vector index ──▶ knowledge graph* ─┘
(10 targets, (LightRAG or
local FastEmbed) Graphify, Ollama)
TARGET picks the destination: obsidian = notes only,
any vector store = notes + that store. * optional stages
```
| Doc | What it is |
|---|---|
| [QUICKSTART.md](QUICKSTART.md) | clone → services → first ingest → first retrieval |
| [BLOG.md](BLOG.md) | design essay + v1.8 addendum |
| [LICENSE](LICENSE) | AGPL-3.0-or-later, full text |
## Destinations and stages
Vault notes always run — they are the product. Everything else stacks:
| Piece | What it does | Needs |
|---|---|---|
| **Vault notes** *(always on)* | Normalized Markdown notes (YAML frontmatter: title, source file, optional `source_uri`, ingest timestamp, `brain/ingested` + `source/<ext>` tags) into the notes dir (default `<vault>/Ingested_Brain/`). PDFs get `## Page N` headings; EPUBs are chapter-split; code files render as language-tagged listings. `TARGET=obsidian` = notes only. | `python-slugify` |
| **Vector index** *(a TARGET away)* | Structure-preserving contextual chunking, local ONNX embeddings (FastEmbed), exact-replacement indexing — the point count after N re-ingests equals the count after the first. Ten interchangeable targets (below). | `fastembed` + the target's library |
| **Knowledge graph** *(optional)* | **LightRAG** (merged entity graph in `<vault>/.lightrag`, one Ollama pass per document) or **Graphify** (batch `graphify extract --backend ollama` → `graph.json`, `GRAPH_REPORT.md`, interactive `graph.html` in `<vault>/.graphify`). | engine package, Ollama |
| **MinIO archive** *(optional)* | Originals uploaded to an S3-compatible bucket (key = input-tree path, idempotent); the note records `s3://bucket/key`. Open-source MinIO has no vector API — it is an archive stage, not a vector target. | `minio`, MinIO service |
| **Source archive** *(optional)* | Verified sources *move* out of the incoming tree into `ingested-archive/` and are bzip2-compressed (streaming, level 9). Sources are never deleted; the archive always holds a decompressible original. | nothing |
Sources are never deleted — by design, not by flag. Guards (MAX MB,
skip-unchanged) skip files entirely; skips never archive.
### Vector targets
All ten share one payload schema (`document_title`, `section_header`,
`content`, `chunk_kind`, `lang`, …), one retrieval strip, and the same
exact-replacement semantics — cosine scores agree to four decimals
across targets (verified by the live matrix).
| Target | Mode | Library | Service needed | Install extra |
|---|---|---|---|---|
| `qdrant` *(default)* | service | `qdrant-client` | Qdrant container | `.[ingest]` |
| `chroma` | embedded | `chromadb` | none | `.[chroma]` |
| `lancedb` | embedded | `lancedb` | none | `.[lancedb]` |
| `faiss` | file | `faiss-cpu` | none | `.[faiss]` |
| `milvus` | embedded (Lite) | `pymilvus[milvus_lite]` | none | `.[milvus]` |
| `weaviate` | service | `weaviate-client` | Weaviate container (8080/50051) | `.[weaviate]` |
| `pgvector` | service | `psycopg` + `pgvector` | Postgres + pgvector (libpq env: `PG*`) | `.[pgvector]` |
| `duckdb` | embedded | `duckdb` (+vss index, steps down to exact scan) | none | `.[duckdb]` |
| `sqlitevec` | embedded | `sqlite-vec` (vec0 tables) | none | `.[sqlitevec]` |
| `mariadb` | service | `PyMySQL` | MariaDB 11.7+ (env: `MARIADB_*`) | `.[mariadb]` |
Embedded targets keep data under `<vault>/.thicket/<target>/` — a vault
is one portable tree. Service targets read Unix-standard environments:
`PGHOST`/`PGUSER`/`PGPASSWORD`/`PGDATABASE` (or `PGDSN`),
`MARIADB_HOST`/`MARIADB_USER`/`MARIADB_PASSWORD`/`MARIADB_DATABASE`.
`pip install -e ".[targets]"` installs every target extra at once.
**Upgrading from 1.1.x:** points indexed before 1.2.0 lack the internal
`doc_key`; the first re-ingest of each document cleans them up
automatically — or start a fresh collection name.
## AI filesystem layout awareness
When the canonical `/mnt/AI` tree exists, Thicket adopts its corpus flow
as defaults — zero configuration:
| Canonical path | Thicket role |
|---|---|
| `/mnt/AI/corpus/cold` | IN — incoming raw documents |
| `/mnt/AI/corpus/hot` | VAULT — the active brain: notes + `.thicket/` vector data + graphs in one tree |
| `/mnt/AI/corpus/books` | standing library (reported by the probe; point IN at it to ingest) |
| `/mnt/AI/corpus/archive` | destination of the source-archive (bz2) stage |
| `/mnt/AI/backends/thicket.env` | connection profiles (600) — fills `PG*` / `MARIADB_*` / `MINIO_*` gaps; the real environment always wins |
Without the layout, home-directory defaults apply — behavior identical.
Override the root with `THICKET_AI_ROOT`. Every path field accepts `~`
and `$VAR` references, resolved through one shared expansion point in
the GUI, CLI, and core.
## Technical corpora (code, configuration, policy)
Ingestion preserves what makes technical documents useful:
- **Fenced code blocks are atomic** — never split mid-listing, never
merged with prose, indentation and blank lines verbatim; the language
tag travels in the chunk context (`| code:python`) and the payload
(`chunk_kind`, `lang`). Oversized blocks window by lines with overlap.
- **Prose chunks are paragraph-aligned** — lists, commands, and tables
keep their line structure in stored content.
- **ATX headings require `#` + whitespace** — shebangs and `#comments`
never masquerade as section headers.
- **Source files ingest directly**: `.py .sh .bash .zsh .yaml .yml .toml
.ini .conf .cfg .json .sql .rs .go .c .h .cpp .js .ts .tf .nix` are
wrapped as language-tagged listings.
- **Code-strong default embedder**: `jinaai/jina-embeddings-v2-base-code`
(English + code, 768-dim, 8k context); curated alternatives in the
EMBED selector. Dimension is fixed per collection — switching models
means a new collection name.
- **Graph tip**: directories of real code want the `graphify` engine —
tree-sitter gives it per-symbol structure no prose pass can match.
## The console
- **Paths panel** — IN, VAULT, and a live **OUT** line resolving the
actual destination per TARGET (notes dir, `.thicket/` data path, or
service/collection URI), re-resolved on every edit
- **Pipeline panel** — TARGET is the destination (`obsidian` or one of
ten vector stores) with collection/host/port and a per-target
connection hint line; stage toggles (graph, MinIO, bz2 archive);
EMBED / OLLAMA LLM / GRAPH ENGINE selectors; FILTER, NOTES DIR,
MAX MB, Skip-unchanged
- **Document queue** — LED matrix table with per-file stage (QUEUED →
EXTRACT → MINIO → VAULT → INDEX → GRAPH → ARCHIVE → DONE / SKIP /
ERROR), color-coded, with detail column
- **Knobs** — chunk size, chunk overlap, retrieval Top-K
- **Retrieval strip** — **SEARCH** (semantic across the collection) and
**ASK (SQL)** (natural language over SQL-backed targets via Vanna 2 +
the Ollama LLM); results land in the log
- **Status footer** — live target + service readiness, updated the
instant TARGET changes
- **Transport** — `> INGEST`, `[] STOP` (cooperative), `~~ SCAN QUEUE`
(preview), `? ABOUT`
- **Probe** — every dependency and service reported at startup; a stage
that cannot run is blocked at INGEST with the exact reason and fix
## Install
**One venv, always**: `/mnt/AI/runtime/thicket-venv` is the single environment.
The source-anchored bootstrap mirrors every dependency as a git
checkout under `/mnt/AI/distfiles/git/`, builds them into that same
venv, and drops a launcher in `/mnt/AI/tools/bin`:
```bash
scripts/bootstrap_sources.sh # add --force-source for
/mnt/AI/tools/bin/thicket --dry-run # native builds from git
```
The venv is created and owned by the bootstrap at
`/mnt/AI/runtime/thicket-venv` — nothing is written inside the project
checkout (`.gitignore` keeps it archive-clean). Plain pip path:
```bash
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
true # (scripts/bootstrap_sources.sh creates it)
/mnt/AI/runtime/thicket-venv/bin/pip install -e . # console only (PySide6)
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # + parsers, FastEmbed, qdrant
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # + every vector target
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph,graphify]" # + both graph engines
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ask,minio]" # + Vanna ask, MinIO archive
```
Services (only for the stages that need them):
```bash
podman run -d --name thicket-qdrant -p 6333:6333 \
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
ollama pull llama3.1 && ollama pull nomic-embed-text
```
The first ingest downloads the embedding model locally (~160 MB for the
default jina-code embedder).
## Usage
One command owns every mode — no flags loads the console:
```bash
python thicket.py # the console — also: ./thicket.py
```
**Headless** (same core, no Qt — SSH / cron friendly):
```bash
python thicket.py --ingest /mnt/AI/corpus/cold --vault /mnt/AI/corpus/hot
python thicket.py --ingest ~/Books --vault ~/Vault --target obsidian # notes only
python thicket.py --ingest ~/Books --vault ~/Vault --target chroma # embedded, no service
python thicket.py --ingest ~/Books --vault ~/Vault --lightrag --graph-engine graphify \
--ollama-llm llama3.1:latest
python thicket.py --ingest ~/Books --vault ~/Vault --skip-unchanged --max-mb 200
python thicket.py --ingest ~/Books --vault ~/Vault --archive --minio
python thicket.py --ask "which document has the most chunks?" --target pgvector
```
**Probe & report**: `python thicket.py --dry-run` · `--version`
Three interchangeable entry points share one environment: the `thicket`
command (installed to `~/.local/bin` by the bootstrap — works from any
directory), `/mnt/AI/tools/bin/thicket`, and `/mnt/AI/runtime/thicket-venv/bin/python thicket.py`.
A bare `python thicket.py` uses the system interpreter, which cannot see
the venv — that is Python's rule, not a Thicket setting.
## Architecture
```
thicket/
├── thicket.py # one-command launcher (GUI default, flags pass through)
├── cli.py # version / dry-run / headless ingest / ask / GUI
├── layout.py # /mnt/AI taxonomy awareness + ~/$VAR expansion
├── extractors.py # PDF / EPUB / MD / TXT / code-config -> (title, text)
├── vault_writer.py # notes: frontmatter, escaping, collisions, source_uri
├── chunker.py # structure-preserving chunking (fences, langs, paragraphs)
├── embedder.py # FastEmbed wrapper (curated catalog, jina-code default)
├── vector_stores.py # registry + 10 targets, close() lifecycle
├── qdrant_store.py # qdrant target
├── graph_store.py # graph engines: LightRAG (async lifecycle) + Graphify
├── minio_archive.py # S3 object archive stage
├── ask_vanna.py # Vanna 2 agent: natural-language SQL over SQL targets
├── pipeline_core.py # the stage table — shared by GUI worker and CLI
├── pipeline_worker.py # QThreads: probe / ingest / search / ask
├── env_probe.py # module + service readiness (incl. liveness tables)
├── ui_theme.py # MMD3 QSS (OpenTranscode visual lineage)
├── ui_window.py # ThicketWindow + launch_gui()
└── widgets/radio_knob.py
```
Design invariants:
- **Qt-free core.** The QThread worker and the headless CLI drive the
same `IngestPipeline`; one behavior change lands in both at once.
- **Table-driven dispatch.** Stage order, file-type routing, target
registry, readiness gating, UI stage colors — data tables, not branch
nests.
- **Lazy heavy imports + graceful degradation.** Every heavy dependency
loads at point of use; missing pieces report themselves and block only
the stage that needs them.
- **Idempotent re-ingest.** Deterministic IDs plus delete-by-`doc_key`
matching both key generations) make re-ingesting an edited source an
exact replacement — proven per-target by the live matrix.
- **Per-file isolation.** One broken document logs an ERROR; the queue
moves on. Postgres transactions roll back per document.
- **Explicit resource lifecycle.** Every store implements `close()`;
Milvus Lite's embedded server is released so the next process can
open the database.
- **Step-down chains.** Qdrant `query_points`→`search`; LightRAG
bindings across API generations; DuckDB vss→exact scan; podman
rootless→`--network=host`.
## Development
```bash
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[dev,targets,graph,graphify,ask,minio]"
.venv/bin/pytest # 62 tests, no services needed
QT_QPA_PLATFORM=offscreen \
.venv/bin/python scripts/smoke_gui.py # headless GUI smoke + screenshot
.venv/bin/python scripts/func_test.py # LIVE matrix: every destination,
# engine, archive, ask — 18 checks
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run # live readiness report
```
Standards applied: PEP 8; SEI CERT practices (bounded reads, precise
exception scope); MISRA-style bounded, structured control flow; POSIX
assumptions (`pathlib`, no platform branches, cron-safe headless mode).
## License
AGPL-3.0-or-later — Jeremy Anderson · info@dcos.net · [dcos.net](https://dcos.net) · 2026.
See [LICENSE](LICENSE) for the full text.
Invoked (not bundled) components carry their own licenses: PySide6
(LGPL-3.0), pypdf (BSD), EbookLib (AGPL-3.0), BeautifulSoup (MIT),
python-slugify (MIT), FastEmbed (Apache-2.0), Qdrant (Apache-2.0),
Chroma (Apache-2.0), LanceDB (Apache-2.0), FAISS (MIT), pymilvus
(Apache-2.0), weaviate-client (BSD-3), psycopg (LGPL-3.0), pgvector
(PostgreSQL), PyMySQL (MIT), sqlite-vec (MIT), LightRAG (MIT),
Graphify (Apache-2.0/MIT), Vanna (MIT), MinIO client (Apache-2.0),
Ollama (MIT).