176 lines
6.4 KiB
Markdown
176 lines
6.4 KiB
Markdown
# Thicket — Quick Start
|
|
|
|
Five minutes from clone to retrieval. Linux, Python 3.12+.
|
|
|
|
## 0. Zero-config on an AI workstation
|
|
|
|
If `/mnt/AI` exists, Thicket detects it: IN defaults to
|
|
`/mnt/AI/corpus/cold`, the vault to `/mnt/AI/corpus/hot`, and the
|
|
bz2 archive stage targets `/mnt/AI/corpus/archive`. Drop documents in
|
|
`corpus/cold`, press INGEST. (Point IN at `corpus/books` to chew
|
|
through the standing library.)
|
|
|
|
## 1. Install
|
|
|
|
```bash
|
|
git clone http://git.dcos.net/dcosnet/Thicket.git # or your local copy
|
|
cd Thicket
|
|
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
|
|
true # (scripts/bootstrap_sources.sh creates it)
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # parsers + FastEmbed + Qdrant client
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph]" # optional: LightRAG stage
|
|
```
|
|
|
|
That's it for setup — ONE venv, always. From here on:
|
|
|
|
```bash
|
|
thicket # anywhere — ~/.local/bin command
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py # explicit interpreter
|
|
/mnt/AI/tools/bin/thicket # AI-tree launcher
|
|
```
|
|
|
|
(A bare `python thicket.py` uses the system interpreter, which cannot
|
|
see the venv — Python's rule. The `thicket` command exists so you
|
|
never need to think about it.)
|
|
|
|
## 2. Services (optional — depends on your vector target)
|
|
|
|
The vault-note stage needs nothing but the install. The vector stage
|
|
targets ten open-source stores; six of them run **embedded** (files
|
|
under `<vault>/.thicket/`, zero services):
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[chroma]" # or lancedb / faiss / milvus /
|
|
# duckdb / sqlitevec, or
|
|
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # every target extra at once
|
|
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault --target chroma
|
|
```
|
|
|
|
Only the default `qdrant` target needs a service:
|
|
|
|
```bash
|
|
# Qdrant (service-backed target)
|
|
podman run -d --name thicket-qdrant -p 6333:6333 \
|
|
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
|
|
# docker works identically; --network=host sidesteps rootless
|
|
# networking issues on some kernels.
|
|
|
|
# Ollama (graph engines + ASK)
|
|
ollama pull llama3.1
|
|
ollama pull nomic-embed-text
|
|
```
|
|
|
|
## 3. Check readiness
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run
|
|
```
|
|
|
|
You want `READY: vault + Qdrant stages available.` — every MISSING
|
|
module or DOWN service is listed with the exact fix.
|
|
|
|
## 4. First ingest
|
|
|
|
**GUI** (the console):
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py
|
|
```
|
|
|
|
1. Set **IN** to a folder of `.pdf` / `.epub` / `.md` / `.txt` files
|
|
and **VAULT** to your Obsidian vault root.
|
|
2. Press **SCAN QUEUE** — the LED table previews every document found.
|
|
3. Press **> INGEST**. Watch the stage column walk each file through
|
|
EXTRACT → VAULT → INDEX → … → DONE (graph/archives add their own
|
|
stages; the OUT line under IN/VAULT shows exactly where data lands).
|
|
4. Type a question in the retrieval strip and press **SEARCH** —
|
|
hits land in the log with score, document, and section.
|
|
|
|
**Headless** (SSH / cron friendly — identical pipeline, no Qt):
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py \
|
|
--ingest ~/Downloads/Raw_Books_And_Papers \
|
|
--vault ~/Documents/ObsidianVault
|
|
```
|
|
|
|
The first run downloads the embedding model (~160 MB, once — the
|
|
default jina-code embedder is tuned for technical corpora); notes
|
|
appear in `<vault>/Ingested_Brain/`, vectors in the `second_brain`
|
|
collection.
|
|
|
|
## 4b. Ask questions in natural language (SQL targets)
|
|
|
|
With the corpus in Postgres or MariaDB, skip SQL entirely:
|
|
|
|
```bash
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ask "which document has the most chunks?" \
|
|
--target pgvector --ollama-llm llama3.1:latest
|
|
```
|
|
|
|
Connection settings come from `/mnt/AI/backends/thicket.env`
|
|
(bootstrap writes it, chmod 600); exported `PG*` / `MARIADB_*` /
|
|
`MINIO_*` variables override the file, and defaults apply last.
|
|
|
|
Needs `pip install -e ".[ask]"` (Vanna 2) and Ollama.
|
|
|
|
## 4c. Archive originals to MinIO (optional)
|
|
|
|
```bash
|
|
MINIO_ENDPOINT=localhost:9000 MINIO_ACCESS_KEY=... MINIO_SECRET_KEY=... \
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault \
|
|
--target chroma --minio
|
|
```
|
|
|
|
Notes gain a `source_uri: s3://bucket/key` frontmatter line.
|
|
|
|
## 5. Verify retrieval
|
|
|
|
From the GUI retrieval strip, or headless:
|
|
|
|
```bash
|
|
.venv/bin/python - <<'EOF'
|
|
from thicket.embedder import DEFAULT_EMBED_MODEL, EmbeddingEngine
|
|
from thicket.vector_stores import create_store
|
|
|
|
engine = EmbeddingEngine(DEFAULT_EMBED_MODEL); engine.load()
|
|
store = create_store("qdrant", collection="second_brain", dim=engine.dim)
|
|
store.set_embedder(engine); store.ensure_collection()
|
|
try:
|
|
for hit in store.search(engine.embed_query("your question here"), limit=3):
|
|
p = hit["payload"]
|
|
print(f"[{hit['score']:.3f}] {p['chunk_kind']:5s} "
|
|
f"{p['document_title']} § {p['section_header']}")
|
|
finally:
|
|
store.close()
|
|
EOF
|
|
```
|
|
|
|
## Troubleshooting
|
|
|
|
| Symptom | Cause | Fix |
|
|
|---|---|---|
|
|
| `Qdrant: DOWN` in probe / dry-run | service not running | start the container (step 2) |
|
|
| `fastembed MISSING` | extras not installed | `/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]"` |
|
|
| `ebooklib MISSING` | EPUBs will fail | same extras install as above |
|
|
| INGEST blocked: dimension mismatch | collection built with a different embedding model | pick a new collection name (or delete the old collection) |
|
|
| `Ollama: DOWN` | only the graph stage needs it | start Ollama, or leave the graph stage off |
|
|
| Empty documents SKIP | scanned PDFs have no text layer | OCR first (e.g. `ocrmypdf`), then ingest |
|
|
| Re-ingest count unchanged | that is correct — re-ingest replaces, never duplicates | nothing to fix |
|
|
|
|
## Day-two operations
|
|
|
|
```bash
|
|
podman stop thicket-qdrant && podman start thicket-qdrant # restart service
|
|
podman volume rm thicket_qdrant # wipe vectors (notes stay)
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --target obsidian # notes only
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --skip-unchanged # cheap re-runs
|
|
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --archive # bz2 the sources away
|
|
.venv/bin/python scripts/func_test.py # full live matrix
|
|
```
|
|
|
|
---
|
|
|
|
Questions: Jeremy Anderson — info@dcos.net — https://dcos.net
|