Skip to content

SDK changelog

  • Configurable ONNX Runtime session options — embedding.sessionOptions is the new global default and each model’s sessionOptions overrides it field by field (built-in defaults → global → model). The struct covers intraOpNumThreads, interOpNumThreads (integers in [0, 1024], 0 = runtime default), executionMode, graphOptimizationLevel, enableCpuMemArena, enableMemPattern, the CUDA pair gpuMemLimit / arenaExtendStrategy, and a per-model wasm: { numThreads } block for the WASM backend, which reads its thread count from the ORT env rather than from the session options. @dimple/config exports the pure mergeSessionOptions(global, model); @dimple/embedder exports the widened buildSessionOptions and the new applyWasmOptions. The point is parallel ingest: ONNX Runtime sizes every session’s intra-op pool at one thread per physical core, so two dimple processes on an 8-core box run 16 threads and scale nothing. Set intraOpNumThreads = floor(physicalCores / workers); the configuration reference carries the tuning guide.

  • CPU-only session options no longer inject CUDA — any sessionOptions object used to emit executionProviders: [{ name: "cuda", device_id: 0 }], pinning CPU-only models to a CUDA provider. The provider is now emitted only when gpuMemLimit or arenaExtendStrategy is present, and enableCpuMemArena still always lands, defaulting to false.

  • Deferred embedding mode — index({ deferEmbed: true }) writes every unit without a vector and hands the ones that need one to the durable queue, so the call returns once the rows and the lexical index are written. IndexResult.pendingEmbeddings reports the count. A new memory.embed.batch job drains up to 512 ids per execution through one batched store phase — one ensureTopicTree, one ensureVectorTable, one membership query, one insertVectors, one routeInsertMany, one CASE update — and stamps memories.embedding/embedding_model, so skip detection and findStale settle once a drain completes. The CLI exposes it as dimple index --defer-embed; the MCP memory_index tool defaults to it because the server’s workers drain the queue. The synchronous path is unchanged and stays the default.

  • Index throughput — New-unit writes now go through the existing writeMany + updateEmbeddingMany recipe: one transaction, one multi-row vector insert, one routeInsertMany tree pass, and the candidate scan off as in writeMany, instead of one transaction, one DDL pass, one whole-tree routing read, and one candidate scan per unit. Measured on a pinned 235-unit corpus: index plus 20 queries 9,218 ms → 6,944 ms minimum (median 9,659 → 7,717). Updated units keep the per-unit supersede path.

  • ORT session options apply — The embedder passed sessionOptions (camelCase) where transformers.js reads session_options, so every session option was silently dropped, including the CUDA execution-provider settings. The key is fixed, enableCpuMemArena: false bounds the CPU arena (onnxruntime#11627: a 2 MB model can jump from 200 MB to 6 GB during inference), and CUDA options now reach the runtime.

  • Typed LayerNode graph (breaking) — Replace the runtime-only layer-graph helper with a typed port of OpenCode v2’s LayerNode (packages/shared/src/effect/node.ts), and adopt it at both composition sites. Breaking changes, in the packages they occur:

    • @dimple/shared (@dimple/shared/effect): make now requires deps (it used to default to none), accepts service as well as name, and returns a typed Provider<A, E, Tag> instead of the untyped LayerNode class; group and compile are typed as well, and compile infers the compiled layer’s outputs and error channel instead of the caller stating a type. Added tags, unbound, mapLayer, tag-scoped shared memo maps, and the Node, Provider, Replacements, Tag, Output, Error, Tags, and TagConfig types. Fourteen call sites gained an explicit deps: [].
    • @dimple/shared/effect and @dimple/core: LayerNode is now the module’s value namespace (LayerNode.make, LayerNode.compile, LayerNode.tags, and the graph types inside it), matching the reference’s export * as LayerNode from "./layer-node.js". It is no longer a class, so instanceof LayerNode does not apply; the replacement contracts are checked at compile time. compileLayers, groupLayers, and makeLayer keep their names.
    • @dimple/retriever: the hybridRetrievalNode export is removed, and hybridRetrievalConfigured now takes its provider nodes — hybridRetrievalConfigured(options, [storeNode, embedderNode]). The retriever package has no default store or embedder, so there is no self-contained default declaration.
    • @dimple/workflow: workersConfigured now takes its provider nodes — workersConfigured(defs, options, [jobStackNode, registryNode]) — and the workersNode export is removed. The embedded default job stack would have built a second provider and bound the worker to it.
    • @dimple/core: layer()‘s error channel is now the real union of its nodes’ build errors (CoreConfigError | EmbeddingError | HybridError | RegistryError | SqlError | StoreError) instead of the caller-stated CoreConfigError, which erased failures the graph could already produce. The output union is unchanged, and CoreGraphServices is derived from the compiled layer rather than hand-written. Boot failures keep the same runtime contract: the SDK catches the raw failure and remaps it to the closed DimpleError surface with remapError.
    • @dimple/logger: layer’s return type dropped its stale Scope requirement (the acquisition is Layer.effect-owned, so it was never part of the layer’s requirements).

    Graph composition now rejects at compile time a node whose declared deps do not cover its layer’s requirements (naming each missing service), a replacement that drops outputs, widens the error channel, or crosses tags, and a compile result that is claimed to provide services its graph does not expose. Compilation itself is unchanged: declarations stay lazy and Effect still owns acquisition, memoization, scopes, and finalization.

  • Effect 4.0.0 runtime — Upgrade the bundled Effect runtime to the released 4.0.0 (the effect/unstable/* to effect/* module move, the Schema.TaggedErrorClass to Schema.TaggedError rename, and the queue store API changes), replace the vendored job-queue store with the upstream SQL store plus TTL cleanup, and rebuild the composition root on a declared layer graph. No public API changes.

    The rc.118 to 4.0.0 step is additive for this codebase: no source change was needed. Two upstream changes in the window are worth naming even though dimple uses neither: Chunk.separate and Chunk.partition now return their two halves in the opposite order ([values, errors] — a silent swap, not a type error), and Schema.brand / Schema.fromBrand were retightened while internal SchemaAST helpers (getCandidates removed, getContextOwner genericized) changed.

  • Self-contained published declarations — Fix the published declarations importing private @dimple/* packages. The internalizer rewrote only bare specifiers, so a declaration under dist/internal/** kept names like @dimple/shared/errors and @dimple/shared/types/language — subpath exports that the private packages resolve through an exports map a consumer does not have. It also emitted a fixed ./internal/<pkg>/index.js path regardless of the importing file’s depth, so every vendored declaration outside dist/ pointed one or two directories too shallow. A consumer typechecking the package with skipLibCheck: false got 67 TS2307 Cannot find module errors from the SDK’s own .d.ts files, 11 of them naming an unpublished package.

    The rewriter now resolves each specifier through the target package’s exports map — honouring types before import, and subpath and * patterns — so @dimple/shared/effect lands on the declaration that subpath names, and computes the path relative to the importing file, so nesting depth no longer matters. The build then fails loudly, naming the file and the specifier, if any @dimple/ module specifier survives or if any relative specifier in the tree points at no shipped declaration. The guard uses a detector written independently of the rewriter’s matcher, so tightening the rewriter cannot disarm it.

    No public API change: the shipped types are the same declarations, now resolvable without the private workspace packages.

  • CUDA session options — sessionOptions: { gpuMemLimit, arenaExtendStrategy } on embedding models: ORT’s CUDA arena is unbounded by default (grows by powers of two), so GPUs with limited VRAM (≤8GB) OOM mid-batch and leave the process unable to allocate even a cublas handle afterwards. Capping the arena is the verified fix — ~3.7× faster than CPU on a 4GB card (12.3 ms/turn vs 46.0), stable across write + query
  • writeMany batch routing — bulk ingest no longer routes per unit: the topic tree is read once, every insertion is computed in memory (bit-identical to sequential writes), and one transaction + one batched vector upsert apply the batch (per-unit routing re-read the whole tree and committed a transaction per write, ~42 ms/turn). Embeds also run in 64-text sub-batches — large corpora no longer OOM the ONNX pass (results are bit-identical regardless of batch size), and the batch’s embedding updates land as ONE statement
  • ANN vector search — the vector leg now searches a libsql_vector_idx index (vector_top_k with exact re-ranking of the candidates): the same results as the brute-force scan, dramatically faster on large corpora (the scan remains the fallback when the index or extension is unavailable)
  • Lighter query results — search results no longer carry the 384-float embedding vector by default; opt back in with includeEmbedding: true
  • Faster hybrid queries — graph context (conflicts / supports / superseded-by) is fetched once over the fused result set instead of per-leg over 2× the rows
  • One-transaction dream runs — every judgement edge, derived memory, and the watermark advance in a single transaction (was one per write)
  • graph.enabled: false — now actually disables the write-path candidate scan (the flag existed in the schema but was never consulted)
  • Deterministic ranking with tied distances — search results now sort by id on exact distance ties (approximate-neighbor candidate order varied run to run)
  • Tree health after merges — a merge that leaves an internal node with a single child now collapses it up the tree, so the childCount invariant always holds
  • Faster waits — Runtime.wait wakes on an in-process completion signal instead of polling alone (the poll remains the cross-process fallback)
  • writeMany bulk ingest — batch insert + batch embed + topic routing in one pass: ~10–30× faster than per-write write for corpus loads (the per-write candidate scan is O(leaf size) and made bulk ingest quadratic). All-or-nothing: a duplicate logical key rejects the whole batch. Bulk content skips near-duplicate/conflict candidate edges — run dream afterwards if you want them
  • First-use auto-install — the optional @huggingface/transformers embedder peer installs itself on first use (project package manager, scripts skipped, concurrent installs deduped, timeout-safe, win32-safe)
  • DIMPLE_CWD — treat any directory as the project root for config discovery, relative file: database URLs, and index() relative paths. For daemons, systemd units, and tests running from a different working directory
  • Keyword search matches again — FTS5 treated the whole query as a verbatim phrase, so multi-word queries (e.g. import helper main) returned nothing unless the exact phrase appeared contiguously. Queries are now AND over tokens: every word must appear, anywhere.
  • MCP server — dimple mcp exposes dimple to any MCP client over the 2026-07-28 protocol: a search tool to discover the API surface, and a sandboxed code.exec tool with the full dimple API available inside the sandbox (dimple.write, dimple.search, dimple.graph.*, …), plus resources, prompts, and change notifications
  • Code-mode execution — run TS/JS in a secure, fresh-per-run sandbox with the whole API as a dimple namespace: one line per operation, native control flow, and Promise.all for parallel calls (deny-by-default permissions; secure-exec is an optional peer)
  • API discovery — the search tool returns matching methods with signatures, when-to-use scenarios, and examples, so agents pick the right operation before writing code
  • Daemon lifecycle — dimple mcp start / stop / reload / status plus --print for client config
  • semantic / keyword search profiles now actually switch the retrieval legs (previously accepted but ignored in ranking)
  • updateEmbedding is idempotent — re-embedding an existing memory replaces it in place instead of failing
  • MCP results exclude embedding vectors (~1.5 KB of JSON noise per result) — smaller agent context, faster round-trips
  • Prompt arguments with options (limit, includePending) work with standard MCP clients (the 2026 spec passes them as strings)
  • TypeScript guests run on plain Node (typescript ^5; newer majors are rejected with a clear pin hint)
  • The packed CLI runs on plain Node (no Bun-only runtime APIs) and installs cleanly in consumer projects
  • dimple init produced an invalid config template (baseUrl vs baseURL) — every command after a fresh init failed; the template now boots as-is with zero env vars
  • First-boot races when the jobs database / log files’ directories did not exist yet
  • Runtime config reload — swap config (models, DB, settings) without restarting; failed reloads leave the current stack untouched
  • Indexing lifecycle — re-indexing a changed file updates its memory in place; deleted files and removed definitions are demoted; renames surface as review candidates; code units never expire
  • Call-graph edges — indexed code links callers to callees, same-file and across imports
  • Query budgets — cap results by units, characters, or time
  • Edge review — confirm / reject / expire candidate edges and explain a memory’s graph neighborhood
  • Optional local models — transformers is auto-installed on first use; hosted embedding providers need nothing installed
  • Remote databases — Turso (libsql://) for store + jobs with env-var token resolution
  • CLI polish — stdout is data-only (logs go to stderr), compact --json, dimple db-schema, fully annotated dimple init
  • Search profiles — semantic / keyword modes plus content-type filtering
  • Worker-ready builds — runs on edge runtimes without a filesystem service
  • Confirming a supersede edge now flips the old memory in the same transaction; get / forget / updateEmbedding accept logical keys
  • Hosted embedding providers — OpenAI-compatible endpoints (OpenAI, ollama, vLLM, org gateways), Google, Cohere, and Amazon Bedrock
  • Per-model runtime options — device, precision/dtype, pooling, normalization, cache directory
  • Smaller embeddings on disk — quantized vector types (float8 is 4× smaller)
  • Multiple models side by side — per-model trees; switching models backfills embeddings idempotently
  • Tag-based publishing — v<version> tags trigger verified npm publishes
  • Packaging — self-contained type declarations, READMEs, changelog in tarballs
  • --version now reports the real package version
  • --version now reports the real package version (was hardcoded)
  • Per-model embedding options (device/dtype/pooling/normalize/cacheDir), full transformers dtype list, DB-level vector quantization (vectorType), package READMEs