Thicket/QUICKSTART.md

176 lines
6.4 KiB
Markdown

# Thicket — Quick Start
Five minutes from clone to retrieval. Linux, Python 3.12+.
## 0. Zero-config on an AI workstation
If `/mnt/AI` exists, Thicket detects it: IN defaults to
`/mnt/AI/corpus/cold`, the vault to `/mnt/AI/corpus/hot`, and the
bz2 archive stage targets `/mnt/AI/corpus/archive`. Drop documents in
`corpus/cold`, press INGEST. (Point IN at `corpus/books` to chew
through the standing library.)
## 1. Install
```bash
git clone http://git.dcos.net/dcosnet/Thicket.git # or your local copy
cd Thicket
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
true # (scripts/bootstrap_sources.sh creates it)
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # parsers + FastEmbed + Qdrant client
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph]" # optional: LightRAG stage
```
That's it for setup — ONE venv, always. From here on:
```bash
thicket # anywhere — ~/.local/bin command
/mnt/AI/runtime/thicket-venv/bin/python thicket.py # explicit interpreter
/mnt/AI/tools/bin/thicket # AI-tree launcher
```
(A bare `python thicket.py` uses the system interpreter, which cannot
see the venv — Python's rule. The `thicket` command exists so you
never need to think about it.)
## 2. Services (optional — depends on your vector target)
The vault-note stage needs nothing but the install. The vector stage
targets ten open-source stores; six of them run **embedded** (files
under `<vault>/.thicket/`, zero services):
```bash
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[chroma]" # or lancedb / faiss / milvus /
# duckdb / sqlitevec, or
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # every target extra at once
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault --target chroma
```
Only the default `qdrant` target needs a service:
```bash
# Qdrant (service-backed target)
podman run -d --name thicket-qdrant -p 6333:6333 \
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
# docker works identically; --network=host sidesteps rootless
# networking issues on some kernels.
# Ollama (graph engines + ASK)
ollama pull llama3.1
ollama pull nomic-embed-text
```
## 3. Check readiness
```bash
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run
```
You want `READY: vault + Qdrant stages available.` — every MISSING
module or DOWN service is listed with the exact fix.
## 4. First ingest
**GUI** (the console):
```bash
/mnt/AI/runtime/thicket-venv/bin/python thicket.py
```
1. Set **IN** to a folder of `.pdf` / `.epub` / `.md` / `.txt` files
and **VAULT** to your Obsidian vault root.
2. Press **SCAN QUEUE** — the LED table previews every document found.
3. Press **> INGEST**. Watch the stage column walk each file through
EXTRACT → VAULT → INDEX → … → DONE (graph/archives add their own
stages; the OUT line under IN/VAULT shows exactly where data lands).
4. Type a question in the retrieval strip and press **SEARCH** —
hits land in the log with score, document, and section.
**Headless** (SSH / cron friendly — identical pipeline, no Qt):
```bash
/mnt/AI/runtime/thicket-venv/bin/python thicket.py \
--ingest ~/Downloads/Raw_Books_And_Papers \
--vault ~/Documents/ObsidianVault
```
The first run downloads the embedding model (~160 MB, once — the
default jina-code embedder is tuned for technical corpora); notes
appear in `<vault>/Ingested_Brain/`, vectors in the `second_brain`
collection.
## 4b. Ask questions in natural language (SQL targets)
With the corpus in Postgres or MariaDB, skip SQL entirely:
```bash
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ask "which document has the most chunks?" \
--target pgvector --ollama-llm llama3.1:latest
```
Connection settings come from `/mnt/AI/backends/thicket.env`
(bootstrap writes it, chmod 600); exported `PG*` / `MARIADB_*` /
`MINIO_*` variables override the file, and defaults apply last.
Needs `pip install -e ".[ask]"` (Vanna 2) and Ollama.
## 4c. Archive originals to MinIO (optional)
```bash
MINIO_ENDPOINT=localhost:9000 MINIO_ACCESS_KEY=... MINIO_SECRET_KEY=... \
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault \
--target chroma --minio
```
Notes gain a `source_uri: s3://bucket/key` frontmatter line.
## 5. Verify retrieval
From the GUI retrieval strip, or headless:
```bash
.venv/bin/python - <<'EOF'
from thicket.embedder import DEFAULT_EMBED_MODEL, EmbeddingEngine
from thicket.vector_stores import create_store
engine = EmbeddingEngine(DEFAULT_EMBED_MODEL); engine.load()
store = create_store("qdrant", collection="second_brain", dim=engine.dim)
store.set_embedder(engine); store.ensure_collection()
try:
for hit in store.search(engine.embed_query("your question here"), limit=3):
p = hit["payload"]
print(f"[{hit['score']:.3f}] {p['chunk_kind']:5s} "
f"{p['document_title']} § {p['section_header']}")
finally:
store.close()
EOF
```
## Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| `Qdrant: DOWN` in probe / dry-run | service not running | start the container (step 2) |
| `fastembed MISSING` | extras not installed | `/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]"` |
| `ebooklib MISSING` | EPUBs will fail | same extras install as above |
| INGEST blocked: dimension mismatch | collection built with a different embedding model | pick a new collection name (or delete the old collection) |
| `Ollama: DOWN` | only the graph stage needs it | start Ollama, or leave the graph stage off |
| Empty documents SKIP | scanned PDFs have no text layer | OCR first (e.g. `ocrmypdf`), then ingest |
| Re-ingest count unchanged | that is correct — re-ingest replaces, never duplicates | nothing to fix |
## Day-two operations
```bash
podman stop thicket-qdrant && podman start thicket-qdrant # restart service
podman volume rm thicket_qdrant # wipe vectors (notes stay)
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --target obsidian # notes only
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --skip-unchanged # cheap re-runs
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --archive # bz2 the sources away
.venv/bin/python scripts/func_test.py # full live matrix
```
---
Questions: Jeremy Anderson — info@dcos.net — https://dcos.net