Configuration reference
The config loader is a strict JSONC parser (comments + trailing commas
allowed). Every section is optional except embedding (without models
there is nothing to embed or route). dimple init writes the full
annotated schema. ${ENV_VAR} placeholders resolve from the environment
— tokens never need to live in the config file. Validation reports ALL
problems at once (not one error at a time).
{ // ── store ──────────────────────────────────────────────────────────── "storeUrl": "file:./.dimple/dimple.db", // remote (Turso): "libsql://<org>-<db>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}"
// ── durable jobs ───────────────────────────────────────────────────── "jobs": { "dbUrl": "file:./.dimple/jobs.db", // separate DB for the job queue "execution": "background", // background | inline | external "concurrency": 4, "leaseTimeoutMs": 120000, "maxAttempts": 10, // clamped to [1, 100] "pollIntervalMs": 1000, },
// ── embeddings (REQUIRED) ──────────────────────────────────────────── "embedding": { "defaultModel": "minilm", // the ACTIVE model for writes/queries // ONNX Runtime session options EVERY model inherits (per-model blocks // override them field by field). Split cores across parallel workers: // "intraOpNumThreads": 4 // 4 on 8 cores with 2 workers (0 = runtime default = 1/core) // "interOpNumThreads": 1, // only used when executionMode is "parallel" // "executionMode": "sequential", // sequential | parallel // "graphOptimizationLevel": "all", // disabled | basic | extended | all | layout // "enableCpuMemArena": false, // dimple default (bulk ingest does not need the arena) // "enableMemPattern": true, // "gpuMemLimit": 2147483648, // CUDA arena cap in bytes // "arenaExtendStrategy": 0, // 0 = next power of two, 1 = as requested "models": [ // local — any Transformers.js ONNX model (auto-installed on first use) { "id": "minilm", "provider": "transformers", "model": "Xenova/all-MiniLM-L6-v2", "dimensions": 384, "autoInstall": true, // optional peer auto-install (default true) // per-model runtime options (all optional, defaults shown): // "device": "cpu", // auto|gpu|cpu|wasm|webgpu|cuda|dml|coreml|webnn // "dtype": "q8", // fp32|fp16|q8|int8|uint8|q4|bnb4|q4f16|q2|q2f16|q1|q1f16 // "pooling": "mean", // mean|cls|none // "normalize": true, // "cacheDir": "~/.cache/dimple", // "vectorType": "float32", // float32|float16|float8|float1bit // per-model ORT session options — override the global block per field: // "sessionOptions": { "intraOpNumThreads": 2, "enableCpuMemArena": true }, // "wasm": { "numThreads": 4 }, // wasm backend ONLY — intra-op threads do not reach it }, // hosted — any OpenAI-compatible endpoint (default https://api.openai.com/v1) { "id": "openai", "provider": "openai", "model": "text-embedding-3-small", "dimensions": 1536, "apiKey": "${OPENAI_API_KEY}", "baseURL": "https://api.openai.com/v1", }, // native provider families — baseURL baked in, no baseURL needed // { "id": "gemini", "provider": "google", "model": "gemini-embedding-2", "dimensions": 3072, "apiKey": "…" } // { "id": "cohere", "provider": "cohere", "model": "embed-english-v3.0", "dimensions": 1024, "apiKey": "…" } // { "id": "titan", "provider": "bedrock", "model": "amazon.titan-embed-text-v2:0", "dimensions": 1024 } ], },
// ── topic tree shape (all optional — defaults shown) ───────────────── "topics": { "maxLeafItems": 128, // leaf overflow → split "minLeafItems": 32, // leaf underflow → borrow/merge "maxChildren": 16, // fan-out cap per split "minChildren": 4, "maxDepth": 6, "beamWidth": 3, // beam width for the leaf resolver (Store.resolveTopics) "splitDeferredHardLimit": 512, // force-split backstop (4× maxLeafItems) "rebalanceCooldownMs": 60000, // min delay between rebalances "lloydPasses": 2, // k-means refinement passes },
// ── artifact graph (all optional — defaults shown) ─────────────────── "graph": { "enabled": true, "maxDepth": 2, // traversal depth "maxNeighbors": 12, // per-node fan-out "maxTotal": 80, // total neighborhood cap "supersedeDemotion": 0.5, // score penalty for superseded results "conflictBands": { "nearDuplicate": 0.9, "conflictLow": 0.6 }, "clusterK": 8, // dream concept-clustering K "clusterMinSize": 3, // min members for a concept },
// ── dream/repair LLM (optional — enables `dimple dream`) ───────────── "llm": { "baseURL": "https://api.openai.com/v1", // any OpenAI-compatible endpoint "apiKey": "${OPENAI_API_KEY}", "modelId": "gpt-4o-mini", "temperature": 0.2, // 0–2 "maxOutputTokens": 512, // positive integer "reasoningEffort": "low", // low | medium | high },
// ── logging (all optional — console is silent by default) ──────────── "logging": { "enabled": true, "level": "info", // trace | debug | info | warn | error "format": "jsonl", // jsonl | pretty (console rendering) "file": "./.dimple/logs.jsonl", // durable JSONL sink (off unless set) "scopes": { "@dimple/store": "debug" }, // per-scope level overrides (scopes are runtime logger names — not installable packages) "ringSize": 4096, // in-memory live tail },
// ── retrieval/write defaults (all optional) ─────────────────────────── "defaultK": 60, // RRF smoothing constant "defaultLimit": 10, // default result count "defaultTtlSecs": 604800, // default write TTL (7 days; null = never expires) "sweepIntervalMs": 300000, // expiry sweep cadence (5 minutes)}Remote databases (Turso)
Section titled “Remote databases (Turso)”storeUrl and jobs.dbUrl accept remote libSQL URLs (libsql://,
https://, ws:// — Turso or any libsql-compatible host). The same
engine runs everything the local file does — FTS5, vector search, the
topic tree, durable jobs — over the wire (verified end-to-end against
the real Turso service).
Turso tokens are per-database, so each URL carries its own token.
${ENV_VAR} resolution keeps them out of config files:
{ "storeUrl": "libsql://<org>-<memdb>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}", "jobs": { "dbUrl": "libsql://<org>-<memdb>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}" },}Notes:
- One DB serves both memories and the durable queue (separate table
namespaces) — the simplest setup shares one URL + one token
(
TURSO_TOKENworks as the fallback). A dedicated jobs DB stays possible viaTURSO_JOBS_URL/TURSO_JOBS_TOKEN - The jobs DB must also be remote when the store is remote (the durable queue + workflow state are cross-invocation)
file:-only behaviors (parent-dir creation, migration lock) skip for remote URLs- On edge runtimes, use hosted embedding providers (the local ONNX machinery is local/CLI-only)
Embedding models
Section titled “Embedding models”Providers
Section titled “Providers”provider |
Models | baseURL needed? |
|---|---|---|
transformers |
any Transformers.js ONNX model — local, no keys (auto-installed on first use) | ❌ |
openai (default) |
any OpenAI-compatible endpoint — OpenAI, vLLM, ollama (http://localhost:11434/v1), Groq, Mistral, … |
optional (defaults to OpenAI) |
google |
gemini-embedding-001, gemini-embedding-2, … (native provider) |
❌ no |
cohere |
embed-english-v3.0, embed-multilingual-v3.0, … (native provider) |
❌ no |
bedrock |
amazon.titan-embed-text-v2:0, amazon.nova-embed-text-v2:0, … (AWS creds from env) |
❌ no |
dimensions must match the model’s real output — a mismatch fails typed
on first use. Multiple models coexist: each gets its own vector tables
and its own topic tree; defaultModel picks the active one (queries
route per model, backfill via the durable memory.embed job).
Transformers runtime options (per model, all optional)
Section titled “Transformers runtime options (per model, all optional)”| Field | Values | Default | What it does |
|---|---|---|---|
device |
auto | gpu | cpu | wasm | webgpu | cuda | dml | coreml | webnn |
cpu |
execution backend (cpu/wasm tested; CUDA on Linux x64 needs the provider and CUDA 12, see CUDA on Linux x64) |
dtype |
fp32 | fp16 | q8 | int8 | uint8 | q4 | bnb4 | q4f16 | q2 | q2f16 | q1 | q1f16 |
q8 |
quantizes the MODEL WEIGHTS (inference memory/speed) — never the output vectors |
pooling |
mean | cls | none |
mean |
sentence-pooling strategy |
normalize |
boolean | true |
L2-normalize outputs |
cacheDir |
path | HF cache | override the model download cache |
autoInstall |
boolean | true |
auto-install the optional @huggingface/transformers peer on first use (scripts skipped) |
sessionOptions |
see Session options | ORT defaults, except enableCpuMemArena: false |
ONNX Runtime session options — thread counts, execution mode, graph optimization, arena/pattern switches, and the CUDA pair (gpuMemLimit caps the GPU arena in bytes; arenaExtendStrategy picks the growth policy: 0 = kNextPowerOfTwo, 1 = kSameAsRequested). Pair the CUDA keys with device: "cuda"; verified ~3.7× vs CPU on a 4GB card |
wasm |
{ numThreads? } |
ORT env default | the WASM backend’s own thread count — intra-op threads do not reach it |
Session options (ONNX Runtime)
Section titled “Session options (ONNX Runtime)”sessionOptions is one shared struct, accepted at two levels:
embedding.sessionOptions is the default every model inherits, and a model’s
own block overrides it field by field (precedence: built-in defaults →
global → model). A field left out anywhere keeps ORT’s own default.
| Field | Values | Default | What it does |
|---|---|---|---|
intraOpNumThreads |
integer 0–1024 | 0 (runtime) |
threads used to parallelize INSIDE each operator. ORT’s 0 means one thread per physical core per session, which is why N dimple processes on one box oversubscribe |
interOpNumThreads |
integer 0–1024 | 0 (runtime) |
threads used to parallelize BETWEEN operators — only read when executionMode is parallel |
executionMode |
sequential | parallel |
sequential |
run operators one after another, or in parallel |
graphOptimizationLevel |
disabled | basic | extended | all | layout |
all |
how much the graph is fused/rewritten before inference (load-time work) |
enableCpuMemArena |
boolean | false |
dimple default: the CPU arena pre-allocates and holds memory across runs (onnxruntime#11627), which bulk ingest does not need |
enableMemPattern |
boolean | ORT default | trace one allocation for a repeated input shape and reuse it next run — ORT applies the pattern only in sequential execution mode |
gpuMemLimit |
number (bytes) | unbounded | cap the CUDA arena — ORT’s default grows by powers of two, so ≤8GB cards OOM mid-batch and poison later queries |
arenaExtendStrategy |
0 | 1 |
0 |
CUDA arena growth policy: 0 = kNextPowerOfTwo, 1 = kSameAsRequested |
Tuning for parallel processes
Section titled “Tuning for parallel processes”ONNX Runtime gives every session its own intra-op thread pool sized at one thread per physical core. Two dimple processes on an 8-core box therefore run 16 threads and ingest no faster than one — the cores are shared and the oversubscription costs you.
Split them explicitly:
{ "embedding": { // 8 physical cores, 2 ingest workers → 4 threads each "sessionOptions": { "intraOpNumThreads": 4 }, "defaultModel": "minilm", "models": [ { "id": "minilm", "provider": "transformers", "model": "Xenova/all-MiniLM-L6-v2", "dimensions": 384, }, ], },}The recipe is intraOpNumThreads = floor(physicalCores / workers); set it
once globally (each worker process reads the same config), and override per
model when one model deserves a different share. Two caveats:
enableCpuMemArenastaysfalseby default — the arena buys a small amount of single-process speed while holding memory across runs, so set"enableCpuMemArena": true(globally or per model) when you are running one process and want that speed back.- On the wasm backend
intraOpNumThreadsis ignored: ORT documents thread counts as Node-binding-only, and the wasm backend reads its ownenv.wasm.numThreadsinstead — use thewasmblock ("wasm": { "numThreads": 4 }). The standalonedimplebinary pins that knob to1; the embedded default is ORT’s own (0= system-determined).
CUDA on Linux x64
Section titled “CUDA on Linux x64”device: "cuda" runs inference on the GPU through ONNX Runtime’s CUDA
execution provider. The provider is not in the npm tarball: onnxruntime-node
downloads it from NuGet in its own postinstall, and the CUDA runtime libraries
must be visible to the dynamic loader. bun run setup:cuda does both steps,
and bun run setup:cuda --check reports what is present and whether the
provider resolves.
- The script runs onnxruntime-node’s own installer when the provider is missing. That downloads about 196 MB from NuGet, about 300 MB once extracted. The installer verifies only the HTTPS download and the ZIP structure, with no checksum, so treat it as any postinstall binary download.
- It stages the CUDA 12 and cuDNN 9 runtime wheels into
~/.cache/dimple/cuda12withuv pip install --targetand writes anenv.shthere. source ~/.cache/dimple/cuda12/env.shbefore starting dimple, or pass the printedLD_LIBRARY_PATHinline.
Version constraints: onnxruntime-node 1.24.3, the version transformers.js pins, requires CUDA 12.8 or newer and cuDNN 9.x. CUDA 13 builds exist from ONNX Runtime 1.27, but only for PyPI and NuGet; there is no Node binding, so a system CUDA 13 installation does not satisfy it. The setup script pins the CUDA 12 wheels for that reason.
If the provider or a runtime library is missing, model load fails with the
provider error naming the missing library, for example libcublasLt.so.12;
there is no silent CPU fallback. Bun is default-secure and does not run
lifecycle scripts for packages outside its trusted list, so the
onnxruntime-node postinstall does not run by default here; the setup script
does the same work explicitly. CPU-only users need nothing: the default
install stays CPU.
Vector storage quantization (vectorType)
Section titled “Vector storage quantization (vectorType)”The per-memory search vectors are stored in the model’s vec_<dims>_<model>
table and searched with libSQL’s native vector_distance_cos:
vectorType |
Encoding | 384-dim stored | Tested |
|---|---|---|---|
float32 (default) |
float32 | 1536 B | ✅ |
float16 |
bfloat16 | ~790 B | seam exists |
float8 |
int8 | 395 B (4×) | ✅ write/query/health/split |
float1bit |
binary | ~48 B (32×) | ⚠️ l2 unsupported for 1bit |
The topic tree’s centroid math and memories.embedding always stay
float32 — quantization applies to the stored search vectors only.
Validation notes
Section titled “Validation notes”All failures are typed with precise messages: temperature ∈ [0, 2];
maxAttempts and maxOutputTokens are positive integers (≤ 100 for
maxAttempts); intraOpNumThreads, interOpNumThreads and wasm.numThreads
are integers in [0, 1024]; defaultModel must name a configured model (unknown keys
fail at first use); dimensions mismatches fail on first embed; JSONC is
strict (duplicate keys rejected); and multiple problems are reported in
one pass.
Environment variables
Section titled “Environment variables”Beyond the ${ENV_VAR} placeholders inside the file, the loader reads
three variables directly:
| Variable | What it does |
|---|---|
DIMPLE_CONFIG |
Explicit config file path (overrides project discovery) |
DIMPLE_CONFIG_CONTENT |
Inline JSONC config content (no file needed) |
DIMPLE_CWD |
Treat this directory as the project root: config discovery starts there, relative file: URLs resolve against it, and index() relative paths anchor to it. For daemons, systemd units, and tests running from a different cwd |
DIMPLE_CWD does not change the process cwd — it only redirects the
paths dimple resolves. Leave it unset for normal use; relative
file:./.dimple/... defaults keep working against the process cwd.