Skip to content

Configuration reference

The config loader is a strict JSONC parser (comments + trailing commas allowed). Every section is optional except embedding (without models there is nothing to embed or route). dimple init writes the full annotated schema. ${ENV_VAR} placeholders resolve from the environment — tokens never need to live in the config file. Validation reports ALL problems at once (not one error at a time).

{
// ── store ────────────────────────────────────────────────────────────
"storeUrl": "file:./.dimple/dimple.db",
// remote (Turso): "libsql://<org>-<db>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}"
// ── durable jobs ─────────────────────────────────────────────────────
"jobs": {
"dbUrl": "file:./.dimple/jobs.db", // separate DB for the job queue
"execution": "background", // background | inline | external
"concurrency": 4,
"leaseTimeoutMs": 120000,
"maxAttempts": 10, // clamped to [1, 100]
"pollIntervalMs": 1000,
},
// ── embeddings (REQUIRED) ────────────────────────────────────────────
"embedding": {
"defaultModel": "minilm", // the ACTIVE model for writes/queries
// ONNX Runtime session options EVERY model inherits (per-model blocks
// override them field by field). Split cores across parallel workers:
// "intraOpNumThreads": 4 // 4 on 8 cores with 2 workers (0 = runtime default = 1/core)
// "interOpNumThreads": 1, // only used when executionMode is "parallel"
// "executionMode": "sequential", // sequential | parallel
// "graphOptimizationLevel": "all", // disabled | basic | extended | all | layout
// "enableCpuMemArena": false, // dimple default (bulk ingest does not need the arena)
// "enableMemPattern": true,
// "gpuMemLimit": 2147483648, // CUDA arena cap in bytes
// "arenaExtendStrategy": 0, // 0 = next power of two, 1 = as requested
"models": [
// local — any Transformers.js ONNX model (auto-installed on first use)
{
"id": "minilm",
"provider": "transformers",
"model": "Xenova/all-MiniLM-L6-v2",
"dimensions": 384,
"autoInstall": true, // optional peer auto-install (default true)
// per-model runtime options (all optional, defaults shown):
// "device": "cpu", // auto|gpu|cpu|wasm|webgpu|cuda|dml|coreml|webnn
// "dtype": "q8", // fp32|fp16|q8|int8|uint8|q4|bnb4|q4f16|q2|q2f16|q1|q1f16
// "pooling": "mean", // mean|cls|none
// "normalize": true,
// "cacheDir": "~/.cache/dimple",
// "vectorType": "float32", // float32|float16|float8|float1bit
// per-model ORT session options — override the global block per field:
// "sessionOptions": { "intraOpNumThreads": 2, "enableCpuMemArena": true },
// "wasm": { "numThreads": 4 }, // wasm backend ONLY — intra-op threads do not reach it
},
// hosted — any OpenAI-compatible endpoint (default https://api.openai.com/v1)
{
"id": "openai",
"provider": "openai",
"model": "text-embedding-3-small",
"dimensions": 1536,
"apiKey": "${OPENAI_API_KEY}",
"baseURL": "https://api.openai.com/v1",
},
// native provider families — baseURL baked in, no baseURL needed
// { "id": "gemini", "provider": "google", "model": "gemini-embedding-2", "dimensions": 3072, "apiKey": "…" }
// { "id": "cohere", "provider": "cohere", "model": "embed-english-v3.0", "dimensions": 1024, "apiKey": "…" }
// { "id": "titan", "provider": "bedrock", "model": "amazon.titan-embed-text-v2:0", "dimensions": 1024 }
],
},
// ── topic tree shape (all optional — defaults shown) ─────────────────
"topics": {
"maxLeafItems": 128, // leaf overflow → split
"minLeafItems": 32, // leaf underflow → borrow/merge
"maxChildren": 16, // fan-out cap per split
"minChildren": 4,
"maxDepth": 6,
"beamWidth": 3, // beam width for the leaf resolver (Store.resolveTopics)
"splitDeferredHardLimit": 512, // force-split backstop (4× maxLeafItems)
"rebalanceCooldownMs": 60000, // min delay between rebalances
"lloydPasses": 2, // k-means refinement passes
},
// ── artifact graph (all optional — defaults shown) ───────────────────
"graph": {
"enabled": true,
"maxDepth": 2, // traversal depth
"maxNeighbors": 12, // per-node fan-out
"maxTotal": 80, // total neighborhood cap
"supersedeDemotion": 0.5, // score penalty for superseded results
"conflictBands": { "nearDuplicate": 0.9, "conflictLow": 0.6 },
"clusterK": 8, // dream concept-clustering K
"clusterMinSize": 3, // min members for a concept
},
// ── dream/repair LLM (optional — enables `dimple dream`) ─────────────
"llm": {
"baseURL": "https://api.openai.com/v1", // any OpenAI-compatible endpoint
"apiKey": "${OPENAI_API_KEY}",
"modelId": "gpt-4o-mini",
"temperature": 0.2, // 0–2
"maxOutputTokens": 512, // positive integer
"reasoningEffort": "low", // low | medium | high
},
// ── logging (all optional — console is silent by default) ────────────
"logging": {
"enabled": true,
"level": "info", // trace | debug | info | warn | error
"format": "jsonl", // jsonl | pretty (console rendering)
"file": "./.dimple/logs.jsonl", // durable JSONL sink (off unless set)
"scopes": { "@dimple/store": "debug" }, // per-scope level overrides (scopes are runtime logger names — not installable packages)
"ringSize": 4096, // in-memory live tail
},
// ── retrieval/write defaults (all optional) ───────────────────────────
"defaultK": 60, // RRF smoothing constant
"defaultLimit": 10, // default result count
"defaultTtlSecs": 604800, // default write TTL (7 days; null = never expires)
"sweepIntervalMs": 300000, // expiry sweep cadence (5 minutes)
}

storeUrl and jobs.dbUrl accept remote libSQL URLs (libsql://, https://, ws:// — Turso or any libsql-compatible host). The same engine runs everything the local file does — FTS5, vector search, the topic tree, durable jobs — over the wire (verified end-to-end against the real Turso service).

Turso tokens are per-database, so each URL carries its own token. ${ENV_VAR} resolution keeps them out of config files:

{
"storeUrl": "libsql://<org>-<memdb>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}",
"jobs": { "dbUrl": "libsql://<org>-<memdb>.turso.io?authToken=${TURSO_MEMORIES_TOKEN}" },
}

Notes:

  • One DB serves both memories and the durable queue (separate table namespaces) — the simplest setup shares one URL + one token (TURSO_TOKEN works as the fallback). A dedicated jobs DB stays possible via TURSO_JOBS_URL / TURSO_JOBS_TOKEN
  • The jobs DB must also be remote when the store is remote (the durable queue + workflow state are cross-invocation)
  • file:-only behaviors (parent-dir creation, migration lock) skip for remote URLs
  • On edge runtimes, use hosted embedding providers (the local ONNX machinery is local/CLI-only)
provider Models baseURL needed?
transformers any Transformers.js ONNX model — local, no keys (auto-installed on first use) ❌
openai (default) any OpenAI-compatible endpoint — OpenAI, vLLM, ollama (http://localhost:11434/v1), Groq, Mistral, … optional (defaults to OpenAI)
google gemini-embedding-001, gemini-embedding-2, … (native provider) ❌ no
cohere embed-english-v3.0, embed-multilingual-v3.0, … (native provider) ❌ no
bedrock amazon.titan-embed-text-v2:0, amazon.nova-embed-text-v2:0, … (AWS creds from env) ❌ no

dimensions must match the model’s real output — a mismatch fails typed on first use. Multiple models coexist: each gets its own vector tables and its own topic tree; defaultModel picks the active one (queries route per model, backfill via the durable memory.embed job).

Transformers runtime options (per model, all optional)

Section titled “Transformers runtime options (per model, all optional)”
Field Values Default What it does
device auto | gpu | cpu | wasm | webgpu | cuda | dml | coreml | webnn cpu execution backend (cpu/wasm tested; CUDA on Linux x64 needs the provider and CUDA 12, see CUDA on Linux x64)
dtype fp32 | fp16 | q8 | int8 | uint8 | q4 | bnb4 | q4f16 | q2 | q2f16 | q1 | q1f16 q8 quantizes the MODEL WEIGHTS (inference memory/speed) — never the output vectors
pooling mean | cls | none mean sentence-pooling strategy
normalize boolean true L2-normalize outputs
cacheDir path HF cache override the model download cache
autoInstall boolean true auto-install the optional @huggingface/transformers peer on first use (scripts skipped)
sessionOptions see Session options ORT defaults, except enableCpuMemArena: false ONNX Runtime session options — thread counts, execution mode, graph optimization, arena/pattern switches, and the CUDA pair (gpuMemLimit caps the GPU arena in bytes; arenaExtendStrategy picks the growth policy: 0 = kNextPowerOfTwo, 1 = kSameAsRequested). Pair the CUDA keys with device: "cuda"; verified ~3.7× vs CPU on a 4GB card
wasm { numThreads? } ORT env default the WASM backend’s own thread count — intra-op threads do not reach it

sessionOptions is one shared struct, accepted at two levels: embedding.sessionOptions is the default every model inherits, and a model’s own block overrides it field by field (precedence: built-in defaults → global → model). A field left out anywhere keeps ORT’s own default.

Field Values Default What it does
intraOpNumThreads integer 0–1024 0 (runtime) threads used to parallelize INSIDE each operator. ORT’s 0 means one thread per physical core per session, which is why N dimple processes on one box oversubscribe
interOpNumThreads integer 0–1024 0 (runtime) threads used to parallelize BETWEEN operators — only read when executionMode is parallel
executionMode sequential | parallel sequential run operators one after another, or in parallel
graphOptimizationLevel disabled | basic | extended | all | layout all how much the graph is fused/rewritten before inference (load-time work)
enableCpuMemArena boolean false dimple default: the CPU arena pre-allocates and holds memory across runs (onnxruntime#11627), which bulk ingest does not need
enableMemPattern boolean ORT default trace one allocation for a repeated input shape and reuse it next run — ORT applies the pattern only in sequential execution mode
gpuMemLimit number (bytes) unbounded cap the CUDA arena — ORT’s default grows by powers of two, so ≤8GB cards OOM mid-batch and poison later queries
arenaExtendStrategy 0 | 1 0 CUDA arena growth policy: 0 = kNextPowerOfTwo, 1 = kSameAsRequested

ONNX Runtime gives every session its own intra-op thread pool sized at one thread per physical core. Two dimple processes on an 8-core box therefore run 16 threads and ingest no faster than one — the cores are shared and the oversubscription costs you.

Split them explicitly:

{
"embedding": {
// 8 physical cores, 2 ingest workers → 4 threads each
"sessionOptions": { "intraOpNumThreads": 4 },
"defaultModel": "minilm",
"models": [
{
"id": "minilm",
"provider": "transformers",
"model": "Xenova/all-MiniLM-L6-v2",
"dimensions": 384,
},
],
},
}

The recipe is intraOpNumThreads = floor(physicalCores / workers); set it once globally (each worker process reads the same config), and override per model when one model deserves a different share. Two caveats:

  • enableCpuMemArena stays false by default — the arena buys a small amount of single-process speed while holding memory across runs, so set "enableCpuMemArena": true (globally or per model) when you are running one process and want that speed back.
  • On the wasm backend intraOpNumThreads is ignored: ORT documents thread counts as Node-binding-only, and the wasm backend reads its own env.wasm.numThreads instead — use the wasm block ("wasm": { "numThreads": 4 }). The standalone dimple binary pins that knob to 1; the embedded default is ORT’s own (0 = system-determined).

device: "cuda" runs inference on the GPU through ONNX Runtime’s CUDA execution provider. The provider is not in the npm tarball: onnxruntime-node downloads it from NuGet in its own postinstall, and the CUDA runtime libraries must be visible to the dynamic loader. bun run setup:cuda does both steps, and bun run setup:cuda --check reports what is present and whether the provider resolves.

  1. The script runs onnxruntime-node’s own installer when the provider is missing. That downloads about 196 MB from NuGet, about 300 MB once extracted. The installer verifies only the HTTPS download and the ZIP structure, with no checksum, so treat it as any postinstall binary download.
  2. It stages the CUDA 12 and cuDNN 9 runtime wheels into ~/.cache/dimple/cuda12 with uv pip install --target and writes an env.sh there.
  3. source ~/.cache/dimple/cuda12/env.sh before starting dimple, or pass the printed LD_LIBRARY_PATH inline.

Version constraints: onnxruntime-node 1.24.3, the version transformers.js pins, requires CUDA 12.8 or newer and cuDNN 9.x. CUDA 13 builds exist from ONNX Runtime 1.27, but only for PyPI and NuGet; there is no Node binding, so a system CUDA 13 installation does not satisfy it. The setup script pins the CUDA 12 wheels for that reason.

If the provider or a runtime library is missing, model load fails with the provider error naming the missing library, for example libcublasLt.so.12; there is no silent CPU fallback. Bun is default-secure and does not run lifecycle scripts for packages outside its trusted list, so the onnxruntime-node postinstall does not run by default here; the setup script does the same work explicitly. CPU-only users need nothing: the default install stays CPU.

The per-memory search vectors are stored in the model’s vec_<dims>_<model> table and searched with libSQL’s native vector_distance_cos:

vectorType Encoding 384-dim stored Tested
float32 (default) float32 1536 B ✅
float16 bfloat16 ~790 B seam exists
float8 int8 395 B (4×) ✅ write/query/health/split
float1bit binary ~48 B (32×) ⚠️ l2 unsupported for 1bit

The topic tree’s centroid math and memories.embedding always stay float32 — quantization applies to the stored search vectors only.

All failures are typed with precise messages: temperature ∈ [0, 2]; maxAttempts and maxOutputTokens are positive integers (≤ 100 for maxAttempts); intraOpNumThreads, interOpNumThreads and wasm.numThreads are integers in [0, 1024]; defaultModel must name a configured model (unknown keys fail at first use); dimensions mismatches fail on first embed; JSONC is strict (duplicate keys rejected); and multiple problems are reported in one pass.

Beyond the ${ENV_VAR} placeholders inside the file, the loader reads three variables directly:

Variable What it does
DIMPLE_CONFIG Explicit config file path (overrides project discovery)
DIMPLE_CONFIG_CONTENT Inline JSONC config content (no file needed)
DIMPLE_CWD Treat this directory as the project root: config discovery starts there, relative file: URLs resolve against it, and index() relative paths anchor to it. For daemons, systemd units, and tests running from a different cwd

DIMPLE_CWD does not change the process cwd — it only redirects the paths dimple resolves. Leave it unset for normal use; relative file:./.dimple/... defaults keep working against the process cwd.