6.4 KiB
Thicket — Quick Start
Five minutes from clone to retrieval. Linux, Python 3.12+.
0. Zero-config on an AI workstation
If /mnt/AI exists, Thicket detects it: IN defaults to
/mnt/AI/corpus/cold, the vault to /mnt/AI/corpus/hot, and the
bz2 archive stage targets /mnt/AI/corpus/archive. Drop documents in
corpus/cold, press INGEST. (Point IN at corpus/books to chew
through the standing library.)
1. Install
git clone http://git.dcos.net/dcosnet/Thicket.git # or your local copy
cd Thicket
# venv lives in the AI tree: /mnt/AI/runtime/thicket-venv
true # (scripts/bootstrap_sources.sh creates it)
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" # parsers + FastEmbed + Qdrant client
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[graph]" # optional: LightRAG stage
That's it for setup — ONE venv, always. From here on:
thicket # anywhere — ~/.local/bin command
/mnt/AI/runtime/thicket-venv/bin/python thicket.py # explicit interpreter
/mnt/AI/tools/bin/thicket # AI-tree launcher
(A bare python thicket.py uses the system interpreter, which cannot
see the venv — Python's rule. The thicket command exists so you
never need to think about it.)
2. Services (optional — depends on your vector target)
The vault-note stage needs nothing but the install. The vector stage
targets ten open-source stores; six of them run embedded (files
under <vault>/.thicket/, zero services):
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[chroma]" # or lancedb / faiss / milvus /
# duckdb / sqlitevec, or
/mnt/AI/runtime/thicket-venv/bin/pip install -e ".[targets]" # every target extra at once
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault --target chroma
Only the default qdrant target needs a service:
# Qdrant (service-backed target)
podman run -d --name thicket-qdrant -p 6333:6333 \
-v thicket_qdrant:/qdrant/storage docker.io/qdrant/qdrant
# docker works identically; --network=host sidesteps rootless
# networking issues on some kernels.
# Ollama (graph engines + ASK)
ollama pull llama3.1
ollama pull nomic-embed-text
3. Check readiness
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --dry-run
You want READY: vault + Qdrant stages available. — every MISSING
module or DOWN service is listed with the exact fix.
4. First ingest
GUI (the console):
/mnt/AI/runtime/thicket-venv/bin/python thicket.py
- Set IN to a folder of
.pdf/.epub/.md/.txtfiles and VAULT to your Obsidian vault root. - Press SCAN QUEUE — the LED table previews every document found.
- Press > INGEST. Watch the stage column walk each file through EXTRACT → VAULT → INDEX → … → DONE (graph/archives add their own stages; the OUT line under IN/VAULT shows exactly where data lands).
- Type a question in the retrieval strip and press SEARCH — hits land in the log with score, document, and section.
Headless (SSH / cron friendly — identical pipeline, no Qt):
/mnt/AI/runtime/thicket-venv/bin/python thicket.py \
--ingest ~/Downloads/Raw_Books_And_Papers \
--vault ~/Documents/ObsidianVault
The first run downloads the embedding model (~160 MB, once — the
default jina-code embedder is tuned for technical corpora); notes
appear in <vault>/Ingested_Brain/, vectors in the second_brain
collection.
4b. Ask questions in natural language (SQL targets)
With the corpus in Postgres or MariaDB, skip SQL entirely:
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ask "which document has the most chunks?" \
--target pgvector --ollama-llm llama3.1:latest
Connection settings come from /mnt/AI/backends/thicket.env
(bootstrap writes it, chmod 600); exported PG* / MARIADB_* /
MINIO_* variables override the file, and defaults apply last.
Needs pip install -e ".[ask]" (Vanna 2) and Ollama.
4c. Archive originals to MinIO (optional)
MINIO_ENDPOINT=localhost:9000 MINIO_ACCESS_KEY=... MINIO_SECRET_KEY=... \
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ~/Books --vault ~/Vault \
--target chroma --minio
Notes gain a source_uri: s3://bucket/key frontmatter line.
5. Verify retrieval
From the GUI retrieval strip, or headless:
.venv/bin/python - <<'EOF'
from thicket.embedder import DEFAULT_EMBED_MODEL, EmbeddingEngine
from thicket.vector_stores import create_store
engine = EmbeddingEngine(DEFAULT_EMBED_MODEL); engine.load()
store = create_store("qdrant", collection="second_brain", dim=engine.dim)
store.set_embedder(engine); store.ensure_collection()
try:
for hit in store.search(engine.embed_query("your question here"), limit=3):
p = hit["payload"]
print(f"[{hit['score']:.3f}] {p['chunk_kind']:5s} "
f"{p['document_title']} § {p['section_header']}")
finally:
store.close()
EOF
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Qdrant: DOWN in probe / dry-run |
service not running | start the container (step 2) |
fastembed MISSING |
extras not installed | /mnt/AI/runtime/thicket-venv/bin/pip install -e ".[ingest]" |
ebooklib MISSING |
EPUBs will fail | same extras install as above |
| INGEST blocked: dimension mismatch | collection built with a different embedding model | pick a new collection name (or delete the old collection) |
Ollama: DOWN |
only the graph stage needs it | start Ollama, or leave the graph stage off |
| Empty documents SKIP | scanned PDFs have no text layer | OCR first (e.g. ocrmypdf), then ingest |
| Re-ingest count unchanged | that is correct — re-ingest replaces, never duplicates | nothing to fix |
Day-two operations
podman stop thicket-qdrant && podman start thicket-qdrant # restart service
podman volume rm thicket_qdrant # wipe vectors (notes stay)
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --target obsidian # notes only
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --skip-unchanged # cheap re-runs
/mnt/AI/runtime/thicket-venv/bin/python thicket.py --ingest ... --archive # bz2 the sources away
.venv/bin/python scripts/func_test.py # full live matrix
Questions: Jeremy Anderson — info@dcos.net — https://dcos.net