Skip to content

CLI changelog

  • Session options in dimple.jsonc — embedding.sessionOptions sets ONNX Runtime session options for every local model, and each model’s own sessionOptions overrides it field by field: thread counts, execution mode, graph optimization level, the CPU arena and memory-pattern switches, the CUDA pair, and a per-model wasm.numThreads for the WASM backend. The parallel-ingest recipe is intraOpNumThreads = floor(physicalCores / workers); the configuration reference carries the tuning guide.

  • CPU-only session options no longer request CUDA — a session-options block emits a CUDA execution provider only when gpuMemLimit or arenaExtendStrategy is set, so a CPU-only model is never pinned to a provider the runtime may not have.

  • dimple index --defer-embed — Write units and defer their embeddings plus topic routing to the durable memory.embed.batch jobs. The command returns once the units are stored and lexically searchable, and the summary reports the pending count; a later run or the MCP server’s workers drain them. The default stays fully synchronous.

  • Faster indexing — New-unit writes go through the batched store path, measured on a pinned 235-unit corpus at 9,218 ms → 6,944 ms minimum for index plus 20 queries. ORT session options now reach the runtime (they were silently dropped) and the CPU arena is bounded.

  • MCP SDK 2.3.0 — dimple mcp runs on @modelcontextprotocol/server 2.3.0 and @modelcontextprotocol/client 2.3.0. Tool-schema conversion moved off the per-request path, tasks/get and tasks/cancel are servable on the 2026-07-28 era, and the one-server-per-request rule matches the existing per-request factory.

  • Standalone binaries + launcher — The dimple CLI now ships as prebuilt standalone binaries, with the npm package’s dimple command acting as a launcher. Six platform packages (@ponraaj/dimple-linux-x64, -linux-arm64, -linux-x64-musl, -darwin-arm64, -darwin-x64, -win32-x64) carry a single-file build of the CLI with the wasm ONNX engine embedded, so the local embedder runs offline without an install script; each package declares its os/cpu/libc and ships the binary alone. The launcher resolves the platform package for the current machine, runs it with inherited stdio, and propagates its exit code; DIMPLE_BINARY overrides the path. When no prebuilt package is present (unsupported platform, --omit=optional, or a failed resolution) it falls back to the Node bundle (dist/cli.js) with a message saying so.

    The six platform packages are exact optionalDependencies of @ponraaj/dimple and ship in lockstep with it; no published tarball carries install scripts (npm 12 blocks dependency install scripts by default). This release verifies the cross-target binaries structurally (file type, pinned Bun runtime label, matching @libsql native) but has no macOS/Windows/arm64 runner, so those binaries are not executed by CI yet. Image models are unsupported (sharp is stubbed with a named error); text and code embeddings are the supported path.

  • One MCP answer per call — MCP tool results now ship one answer instead of two copies of it. A successful tool call returns the payload once, minified, in content[0].text, and no longer repeats it in structuredContent; every MCP client already reads that channel because content is the mandatory one. Failures are unchanged — isError: true plus the small structuredContent.error.{code,message} machine envelope (still documented in the MCP guide) — and so is the dormant task-handle path, which keeps both channels because its top-level fields are a wire contract (the Tasks extension is not advertised in this build, pending the upstream era-gate fix).

    Effect on a real call, measured on a pinned fixture of five 720 B memories: memory_search(limit=5) drops from 10,815 B to 5,056 B (−53%), and memory_get from 2,161 B to 1,111 B (−48.6%). Both figures are the returned result object serialized as JSON; the JSON-RPC envelope adds ~35 B on top. Clients that read payloads from structuredContent must read the text channel instead; the payload bytes are identical, only the indentation and the duplicate are gone.

    Also enforced now, with tests that refuse to be weakened: a per-tool inputSchema budget of 4,000 characters (the 16,384-character client drop is a silent failure mode) and a 24,000-character budget for the whole advertised tools/list. Today’s surface is 21 tools, 11,970 B, with memory_search the largest schema at 1,901 B.

    No tool was added, removed, renamed, or reshaped: the 21 names, titles, descriptions, and input schemas are byte-identical. The plan and the measured baseline live in docs/mcp-efficiency-design.md; the rationale is borrowed from the Sentry and PostHog MCP servers.

  • MCP annotations and opt-in narrowing — Every tool in the dimple MCP surface now carries the full annotation quartet (readOnlyHint, destructiveHint, idempotentHint, openWorldHint), so a client can gate on it without guessing: the read-only tools are memory_get, memory_search, memory_embed, graph_candidates, graph_coverage, graph_explain, maintenance_health, docs_search and open_in_dimple; openWorldHint is false everywhere, because dimple is a local, closed world and a constant hint would carry no signal.

    destructiveHint follows the MCP definition (false means the call performs only additive updates), applied through one rule: a tool is destructive when it can irreversibly hide, remove, or repurpose a live entity the caller did not name. Seven tools qualify — memory_forget, memory_index, graph_add_edge, graph_confirm, graph_reject, graph_expire and maintenance_dream. The three that remove or retire the entity the caller named are the obvious set (memory_forget deletes the memory; graph_reject and graph_expire retire the edge); the other four matter because confirming (or inserting as confirmed) a memory_supersedes edge flips the old memory active → superseded, which drops it out of every read path and releases its logical key for reuse — graph_confirm, a confirmed: true graph_add_edge, memory_index’s superseding write plus its reconcile of vanished units, and the dream job’s confirmed supersedes relation all reach it.

    The loopback HTTP endpoint honours two opt-in query parameters for clients that want a narrower surface (stdio is unaffected — it has no URL and always serves everything):

    • ?readonly=1 — only the tools annotated readOnlyHint: true
    • ?tools=memory_get,graph_explain — exactly the named tools
    • both together — the union of the two sets

    Tool names are matched exactly (no hyphen/underscore rewriting); an unknown name is answered 400 with a message listing the valid tools. That body enumerates the tool surface, so it is produced only for a local caller — a request with no Origin, or with a loopback origin (127.0.0.1, localhost, ::1 over http/https). Any other origin goes to the SDK handler untouched, where the daemon’s own localhost host/origin guards decide. With no query parameters the full 21-tool surface is served, exactly as before — nothing is hidden unless the client asks.

  • Updated dependencies

    • @ponraaj/dimple-sdk@0.0.0-alpha.11
  • CUDA session options for local embeddings (see the SDK changelog)
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.10
  • Queries are faster across the board (ANN vector search, lighter result payloads, single-pass graph context) — see the SDK changelog
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.9
  • dimple upgrade — checks the npm registry for a newer CLI release and installs it with the package manager that installed you (auto-detected; --manager overrides). --check prints the available version without installing; no project config needed
  • code.exec auto-install — the sandbox’s secure-exec peer (and its platform sidecar) installs itself on first code-mode run
  • Keyword search matches again — FTS5 treated the whole query as a verbatim phrase, so multi-word queries (e.g. import helper main) returned nothing unless the exact phrase appeared contiguously. Queries are now AND over tokens: every word must appear, anywhere.
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.8
  • MCP server — dimple mcp exposes dimple to any MCP client over the 2026-07-28 protocol: a search tool to discover the API surface, and a sandboxed code.exec tool with the full dimple API available inside the sandbox (dimple.write, dimple.search, dimple.graph.*, …), plus resources, prompts, and change notifications
  • Code-mode execution — run TS/JS in a secure, fresh-per-run sandbox with the whole API as a dimple namespace: one line per operation, native control flow, and Promise.all for parallel calls (deny-by-default permissions; secure-exec is an optional peer)
  • API discovery — the search tool returns matching methods with signatures, when-to-use scenarios, and examples, so agents pick the right operation before writing code
  • Daemon lifecycle — dimple mcp start / stop / reload / status plus --print for client config
  • semantic / keyword search profiles now actually switch the retrieval legs (previously accepted but ignored in ranking)
  • updateEmbedding is idempotent — re-embedding an existing memory replaces it in place instead of failing
  • MCP results exclude embedding vectors (~1.5 KB of JSON noise per result) — smaller agent context, faster round-trips
  • Prompt arguments with options (limit, includePending) work with standard MCP clients (the 2026 spec passes them as strings)
  • TypeScript guests run on plain Node (typescript ^5; newer majors are rejected with a clear pin hint)
  • The packed CLI runs on plain Node (no Bun-only runtime APIs) and installs cleanly in consumer projects
  • dimple init produced an invalid config template (baseUrl vs baseURL) — every command after a fresh init failed; the template now boots as-is with zero env vars
  • First-boot races when the jobs database / log files’ directories did not exist yet
  • Runtime config reload — swap config (models, DB, settings) without restarting; failed reloads leave the current stack untouched
  • Indexing lifecycle — re-indexing a changed file updates its memory in place; deleted files and removed definitions are demoted; renames surface as review candidates; code units never expire
  • Call-graph edges — indexed code links callers to callees, same-file and across imports
  • Query budgets — cap results by units, characters, or time
  • Edge review — confirm / reject / expire candidate edges and explain a memory’s graph neighborhood
  • Optional local models — transformers is auto-installed on first use; hosted embedding providers need nothing installed
  • Remote databases — Turso (libsql://) for store + jobs with env-var token resolution
  • CLI polish — stdout is data-only (logs go to stderr), compact --json, dimple db-schema, fully annotated dimple init
  • Search profiles — semantic / keyword modes plus content-type filtering
  • Worker-ready builds — runs on edge runtimes without a filesystem service
  • Confirming a supersede edge now flips the old memory in the same transaction; get / forget / updateEmbedding accept logical keys
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.5
  • Hosted embedding providers — OpenAI-compatible endpoints (OpenAI, ollama, vLLM, org gateways), Google, Cohere, and Amazon Bedrock
  • Per-model runtime options — device, precision/dtype, pooling, normalization, cache directory
  • Smaller embeddings on disk — quantized vector types (float8 is 4× smaller)
  • Multiple models side by side — per-model trees; switching models backfills embeddings idempotently
  • Tag-based publishing — v<version> tags trigger verified npm publishes
  • Packaging — self-contained type declarations, READMEs, changelog in tarballs
  • --version now reports the real package version
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.4
  • --version now reports the real package version (was hardcoded)
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.3
  • Per-model embedding options (device/dtype/pooling/normalize/cacheDir), full transformers dtype list, DB-level vector quantization (vectorType), package READMEs
  • Updated dependencies
    • @ponraaj/dimple-sdk@0.0.0-alpha.2