277 lines
14 KiB
Markdown
277 lines
14 KiB
Markdown
# Thicket
|
|
|
|
**Super-ingest + RAG console** — feed it documents; it grows a thicket:
|
|
dense, interconnected, searchable. Import PDF / EPUB / Markdown / plain
|
|
text / source-and-config files into a destination of your choice,
|
|
wrapped in the same retro-futuristic MMD3 console UI as OpenTranscode
|
|
(brushed aluminum, amber LEDs, green phosphor log, rotary knobs).
|
|
|
|
```
|
|
┌─▶ Obsidian vault ──▶ MinIO bucket* ┐
|
|
documents ──▶ notes │ (Markdown + (s3:// URIs) ├─▶ ingested-archive/
|
|
│ YAML frontmatter) │ (bz2, never deleted)
|
|
└─▶ vector index ──▶ knowledge graph* ─┘
|
|
(10 targets, (LightRAG or
|
|
local FastEmbed) Graphify, Ollama)
|
|
|
|
TARGET picks the destination: obsidian = notes only,
|
|
any vector store = notes + that store. * optional stages
|
|
```
|
|
|
|
| Doc | What it is |
|
|
|---|---|
|
|
| [QUICKSTART.md](QUICKSTART.md) | clone → services → first ingest → first retrieval |
|
|
| [BLOG.md](BLOG.md) | design essay + v1.8 addendum |
|
|
| [LICENSE](LICENSE) | AGPL-3.0-or-later, full text |
|
|
|
|
## Destinations and stages
|
|
|
|
Vault notes always run — they are the product. Everything else stacks:
|
|
|
|
| Piece | What it does | Needs |
|
|
|---|---|---|
|
|
| **Vault notes** *(always on)* | Normalized Markdown notes (YAML frontmatter: title, source file, optional `source_uri`, ingest timestamp, `brain/ingested` + `source/<ext>` tags) into the notes dir (default `<vault>/Ingested_Brain/`). PDFs get `## Page N` headings; EPUBs are chapter-split; code files render as language-tagged listings. `TARGET=obsidian` = notes only. | `python-slugify` |
|
|
| **Vector index** *(a TARGET away)* | Structure-preserving contextual chunking, local ONNX embeddings (FastEmbed), exact-replacement indexing — the point count after N re-ingests equals the count after the first. Ten interchangeable targets (below). | `fastembed` + the target's library |
|
|
| **Knowledge graph** *(optional)* | **LightRAG** (merged entity graph in `<vault>/.lightrag`, one Ollama pass per document) or **Graphify** (batch `graphify extract --backend ollama` → `graph.json`, `GRAPH_REPORT.md`, interactive `graph.html` in `<vault>/.graphify`). | engine package, Ollama |
|
|
| **MinIO archive** *(optional)* | Originals uploaded to an S3-compatible bucket (key = input-tree path, idempotent); the note records `s3://bucket/key`. Open-source MinIO has no vector API — it is an archive stage, not a vector target. | `minio`, MinIO service |
|
|
| **Source archive** *(optional)* | Verified sources *move* out of the incoming tree into `ingested-archive/` and are bzip2-compressed (streaming, level 9). Sources are never deleted; the archive always holds a decompressible original. | nothing |
|
|
|
|
Sources are never deleted — by design, not by flag. Guards (MAX MB,
|
|
skip-unchanged) skip files entirely; skips never archive.
|
|
|
|
### Vector targets
|
|
|
|
All ten share one payload schema (`document_title`, `section_header`,
|
|
`content`, `chunk_kind`, `lang`, …), one retrieval strip, and the same
|
|
exact-replacement semantics — cosine scores agree to four decimals
|
|
across targets (verified by the live matrix).
|
|
|
|
| Target | Mode | Library | Service needed | Install extra |
|
|
|---|---|---|---|---|
|
|
| `qdrant` *(default)* | service | `qdrant-client` | Qdrant container | `.[ingest]` |
|
|
| `chroma` | embedded | `chromadb` | none | `.[chroma]` |
|
|
| `lancedb` | embedded | `lancedb` | none | `.[lancedb]` |
|
|
| `faiss` | file | `faiss-cpu` | none | `.[faiss]` |
|
|
| `milvus` | embedded (Lite) | `pymilvus[milvus_lite]` | none | `.[milvus]` |
|
|
| `weaviate` | service | `weaviate-client` | Weaviate container (8080/50051) | `.[weaviate]` |
|
|
| `pgvector` | service | `psycopg` + `pgvector` | Postgres + pgvector (libpq env: `PG*`) | `.[pgvector]` |
|
|
| `duckdb` | embedded | `duckdb` (+vss index, steps down to exact scan) | none | `.[duckdb]` |
|
|
| `sqlitevec` | embedded | `sqlite-vec` (vec0 tables) | none | `.[sqlitevec]` |
|
|
| `mariadb` | service | `PyMySQL` | MariaDB 11.7+ (env: `MARIADB_*`) | `.[mariadb]` |
|
|
|
|
Embedded targets keep data under `<vault>/.thicket/<target>/` — a vault
|
|
is one portable tree. Service targets read Unix-standard environments:
|
|
`PGHOST`/`PGUSER`/`PGPASSWORD`/`PGDATABASE` (or `PGDSN`),
|
|
`MARIADB_HOST`/`MARIADB_USER`/`MARIADB_PASSWORD`/`MARIADB_DATABASE`.
|
|
`pip install -e ".[targets]"` installs every target extra at once.
|
|
|
|
**Upgrading from 1.1.x:** points indexed before 1.2.0 lack the internal
|
|
`doc_key`; the first re-ingest of each document cleans them up
|
|
automatically — or start a fresh collection name.
|
|
|
|
## AI filesystem layout awareness
|
|
|
|
When the canonical `/mnt/AI` tree exists, Thicket adopts its corpus flow
|
|
as defaults — zero configuration:
|
|
|
|
| Canonical path | Thicket role |
|
|
|---|---|
|
|
| `/mnt/AI/corpus/cold` | IN — incoming raw documents |
|
|
| `/mnt/AI/corpus/hot` | VAULT — the active brain: notes + `.thicket/` vector data + graphs in one tree |
|
|
| `/mnt/AI/corpus/books` | standing library (reported by the probe; point IN at it to ingest) |
|
|
| `/mnt/AI/corpus/archive` | destination of the source-archive (bz2) stage |
|
|
| `/mnt/AI/backends/thicket.env` | connection profiles (600) — fills `PG*` / `MARIADB_*` / `MINIO_*` gaps; the real environment always wins |
|
|
|
|
Without the layout, home-directory defaults apply — behavior identical.
|
|
Override the root with `THICKET_AI_ROOT`. Every path field accepts `~`
|
|
and `$VAR` references, resolved through one shared expansion point in
|
|
the GUI, CLI, and core.
|
|
|
|
## Technical corpora (code, configuration, policy)
|
|
|
|
Ingestion preserves what makes technical documents useful:
|
|
|
|
- **Fenced code blocks are atomic** — never split mid-listing, never
|
|
merged with prose, indentation and blank lines verbatim; the language
|
|
tag travels in the chunk context (`| code:python`) and the payload
|
|
(`chunk_kind`, `lang`). Oversized blocks window by lines with overlap.
|
|
- **Prose chunks are paragraph-aligned** — lists, commands, and tables
|
|
keep their line structure in stored content.
|
|
- **ATX headings require `#` + whitespace** — shebangs and `#comments`
|
|
never masquerade as section headers.
|
|
- **Source files ingest directly**: `.py .sh .bash .zsh .yaml .yml .toml
|
|
.ini .conf .cfg .json .sql .rs .go .c .h .cpp .js .ts .tf .nix` are
|
|
wrapped as language-tagged listings.
|
|
- **Code-strong default embedder**: `jinaai/jina-embeddings-v2-base-code`
|
|
(English + code, 768-dim, 8k context); curated alternatives in the
|
|
EMBED selector. Dimension is fixed per collection — switching models
|
|
means a new collection name.
|
|
- **Graph tip**: directories of real code want the `graphify` engine —
|
|
tree-sitter gives it per-symbol structure no prose pass can match.
|
|
|
|
## The console
|
|
|
|
- **Paths panel** — IN, VAULT, and a live **OUT** line resolving the
|
|
actual destination per TARGET (notes dir, `.thicket/` data path, or
|
|
service/collection URI), re-resolved on every edit
|
|
- **Pipeline panel** — TARGET is the destination (`obsidian` or one of
|
|
ten vector stores) with collection/host/port and a per-target
|
|
connection hint line; stage toggles (graph, MinIO, bz2 archive);
|
|
EMBED / OLLAMA LLM / GRAPH ENGINE selectors; FILTER, NOTES DIR,
|
|
MAX MB, Skip-unchanged
|
|
- **Document queue** — LED matrix table with per-file stage (QUEUED →
|
|
EXTRACT → MINIO → VAULT → INDEX → GRAPH → ARCHIVE → DONE / SKIP /
|
|
ERROR), color-coded, with detail column
|
|
- **Knobs** — chunk size, chunk overlap, retrieval Top-K
|
|
- **Retrieval strip** — **SEARCH** (semantic across the collection) and
|
|
**ASK (SQL)** (natural language over SQL-backed targets via Vanna 2 +
|
|
the Ollama LLM); results land in the log
|
|
- **Status footer** — live target + service readiness, updated the
|
|
instant TARGET changes
|
|
- **Transport** — `> INGEST`, `[] STOP` (cooperative), `~~ SCAN QUEUE`
|
|
(preview), `? ABOUT`
|
|
- **Probe** — every dependency and service reported at startup; a stage
|
|
that cannot run is blocked at INGEST with the exact reason and fix
|
|
|
|
## Install
|
|
|
|
**One venv, always**: `/mnt/AI/runtime/thicket-venv` is the single environment.
|
|
The source-anchored bootstrap mirrors every dependency as a git
|
|
checkout under `/mnt/AI/distfiles/git/`, builds them into that same
|
|
venv, and drops a launcher in `/mnt/AI/tools/bin`:
|
|
|
|
```bash
|
|
scripts/bootstrap_sources.sh # add --force-source for
|
|
/mnt/AI/tools/bin/thicket --dry-run # native builds from git
|
|
```
|
|
|
|
The venv is created and owned by the bootstrap at
|
|
`/mnt/AI/runtime/thicket-venv` — nothing is written inside the project
|
|
checkout (`.gitignore` keeps it archive-clean). Plain pip path:
|
|
|
|
```bash
|
|
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
|
|
true # (scripts/bootstrap_sources.sh creates it)
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e . # console only (PySide6)
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # + parsers, FastEmbed, qdrant
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # + every vector target
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph,graphify]" # + both graph engines
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ask,minio]" # + Vanna ask, MinIO archive
|
|
```
|
|
|
|
Services (only for the stages that need them):
|
|
|
|
```bash
|
|
podman run -d --name thicket-qdrant -p 6333:6333 \
|
|
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
|
|
ollama pull llama3.1 && ollama pull nomic-embed-text
|
|
```
|
|
|
|
The first ingest downloads the embedding model locally (~160 MB for the
|
|
default jina-code embedder).
|
|
|
|
## Usage
|
|
|
|
One command owns every mode — no flags loads the console:
|
|
|
|
```bash
|
|
python thicket.py # the console — also: ./thicket.py
|
|
```
|
|
|
|
**Headless** (same core, no Qt — SSH / cron friendly):
|
|
|
|
```bash
|
|
python thicket.py --ingest /mnt/AI/corpus/cold --vault /mnt/AI/corpus/hot
|
|
python thicket.py --ingest ~/Books --vault ~/Vault --target obsidian # notes only
|
|
python thicket.py --ingest ~/Books --vault ~/Vault --target chroma # embedded, no service
|
|
python thicket.py --ingest ~/Books --vault ~/Vault --lightrag --graph-engine graphify \
|
|
--ollama-llm llama3.1:latest
|
|
python thicket.py --ingest ~/Books --vault ~/Vault --skip-unchanged --max-mb 200
|
|
python thicket.py --ingest ~/Books --vault ~/Vault --archive --minio
|
|
python thicket.py --ask "which document has the most chunks?" --target pgvector
|
|
```
|
|
|
|
**Probe & report**: `python thicket.py --dry-run` · `--version`
|
|
|
|
Three interchangeable entry points share one environment: the `thicket`
|
|
command (installed to `~/.local/bin` by the bootstrap — works from any
|
|
directory), `/mnt/AI/tools/bin/thicket`, and `/mnt/AI/runtime/thicket-venv/bin/python thicket.py`.
|
|
A bare `python thicket.py` uses the system interpreter, which cannot see
|
|
the venv — that is Python's rule, not a Thicket setting.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
thicket/
|
|
├── thicket.py # one-command launcher (GUI default, flags pass through)
|
|
├── cli.py # version / dry-run / headless ingest / ask / GUI
|
|
├── layout.py # /mnt/AI taxonomy awareness + ~/$VAR expansion
|
|
├── extractors.py # PDF / EPUB / MD / TXT / code-config -> (title, text)
|
|
├── vault_writer.py # notes: frontmatter, escaping, collisions, source_uri
|
|
├── chunker.py # structure-preserving chunking (fences, langs, paragraphs)
|
|
├── embedder.py # FastEmbed wrapper (curated catalog, jina-code default)
|
|
├── vector_stores.py # registry + 10 targets, close() lifecycle
|
|
├── qdrant_store.py # qdrant target
|
|
├── graph_store.py # graph engines: LightRAG (async lifecycle) + Graphify
|
|
├── minio_archive.py # S3 object archive stage
|
|
├── ask_vanna.py # Vanna 2 agent: natural-language SQL over SQL targets
|
|
├── pipeline_core.py # the stage table — shared by GUI worker and CLI
|
|
├── pipeline_worker.py # QThreads: probe / ingest / search / ask
|
|
├── env_probe.py # module + service readiness (incl. liveness tables)
|
|
├── ui_theme.py # MMD3 QSS (OpenTranscode visual lineage)
|
|
├── ui_window.py # ThicketWindow + launch_gui()
|
|
└── widgets/radio_knob.py
|
|
```
|
|
|
|
Design invariants:
|
|
|
|
- **Qt-free core.** The QThread worker and the headless CLI drive the
|
|
same `IngestPipeline`; one behavior change lands in both at once.
|
|
- **Table-driven dispatch.** Stage order, file-type routing, target
|
|
registry, readiness gating, UI stage colors — data tables, not branch
|
|
nests.
|
|
- **Lazy heavy imports + graceful degradation.** Every heavy dependency
|
|
loads at point of use; missing pieces report themselves and block only
|
|
the stage that needs them.
|
|
- **Idempotent re-ingest.** Deterministic IDs plus delete-by-`doc_key`
|
|
matching both key generations) make re-ingesting an edited source an
|
|
exact replacement — proven per-target by the live matrix.
|
|
- **Per-file isolation.** One broken document logs an ERROR; the queue
|
|
moves on. Postgres transactions roll back per document.
|
|
- **Explicit resource lifecycle.** Every store implements `close()`;
|
|
Milvus Lite's embedded server is released so the next process can
|
|
open the database.
|
|
- **Step-down chains.** Qdrant `query_points`→`search`; LightRAG
|
|
bindings across API generations; DuckDB vss→exact scan; podman
|
|
rootless→`--network=host`.
|
|
|
|
## Development
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[dev,targets,graph,graphify,ask,minio]"
|
|
.venv/bin/pytest # 62 tests, no services needed
|
|
QT_QPA_PLATFORM=offscreen \
|
|
.venv/bin/python scripts/smoke_gui.py # headless GUI smoke + screenshot
|
|
.venv/bin/python scripts/func_test.py # LIVE matrix: every destination,
|
|
# engine, archive, ask — 18 checks
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run # live readiness report
|
|
```
|
|
|
|
Standards applied: PEP 8; SEI CERT practices (bounded reads, precise
|
|
exception scope); MISRA-style bounded, structured control flow; POSIX
|
|
assumptions (`pathlib`, no platform branches, cron-safe headless mode).
|
|
|
|
## License
|
|
|
|
AGPL-3.0-or-later — Jeremy Anderson · info@dcos.net · [dcos.net](https://dcos.net) · 2026.
|
|
See [LICENSE](LICENSE) for the full text.
|
|
|
|
Invoked (not bundled) components carry their own licenses: PySide6
|
|
(LGPL-3.0), pypdf (BSD), EbookLib (AGPL-3.0), BeautifulSoup (MIT),
|
|
python-slugify (MIT), FastEmbed (Apache-2.0), Qdrant (Apache-2.0),
|
|
Chroma (Apache-2.0), LanceDB (Apache-2.0), FAISS (MIT), pymilvus
|
|
(Apache-2.0), weaviate-client (BSD-3), psycopg (LGPL-3.0), pgvector
|
|
(PostgreSQL), PyMySQL (MIT), sqlite-vec (MIT), LightRAG (MIT),
|
|
Graphify (Apache-2.0/MIT), Vanna (MIT), MinIO client (Apache-2.0),
|
|
Ollama (MIT).
|