Skip to content

Retrieval

query text
│
▼
embed
│
▼
scoping filters: namespace · topicKeys · contentType
(both legs) · conceptFilter (vector leg only)
│
▼
vector leg: ANN over the model's vector table
(max(16 × fetched, 64) candidates, exact cosine re-rank)
│
▼
FTS5 lexical leg (same filters)
│
▼
RRF fusion (k = 60)
│
▼
recency boost (optional)
│
▼
budgets: maxUnits clamps the per-leg fetch;
maxChars truncates content; maxMs flags truncated
│
▼
ranked results + one audit row

The retriever asks each leg for limit × 2 results. The vector leg runs an approximate nearest-neighbour search over the active model’s vector table, then re-ranks the candidates with exact cosine in SQL (vector_top_k where the libSQL vector index is available, an exact ORDER BY scan otherwise). The ANN stage fetches max(16 × fetched, 64) candidates because the scoping filters apply after the fetch (that is max(32 × limit, 64) at the default per-leg fetch); the exact-scan fallback walks the filtered table.

Scoping is applied per leg, before fusion: namespace, topics.topicKeys, and contentType restrict both legs; profile.conceptFilter scopes the vector leg only. The store also exposes a deterministic beam-descent leaf resolver (Store.resolveTopics): it walks from __root__, keeps the top beamWidth children per level (maxDepth caps the walk, ties break lexicographically), and returns the leaves a query should search. The SDK query path does not call it yet — the resolver is covered by the topics benchmarks and tests, and wiring it into hybrid is a known gap.

The FTS5 leg runs after the vector leg and catches exact names/identifiers the vector leg might miss.

Algorithm — Reciprocal Rank Fusion:

for each leg, rank its hits:
rrf_score(memory) = Σ_legs 1 / (k + rank_leg(memory)) k = 60 (default)
final score = rrf_score × (1 + recencyBoost × recencyFactor)
recencyFactor = 1 / (1 + ageDays) (hyperbolic decay)

Then re-rank (stable — equal scores keep fused order). The fused set is not graph-augmented on this path: hybrid disables the store’s per-leg post-pass (augment: false) and never calls Store.augmentSearchResults, so query results carry no supersededBy / conflicts / supportedBy fields and no confirmed-supersede demotion. The store’s own search legs (searchVector, searchFts5, recency) apply that post-pass by default; wiring it onto the fused set is a known gap.

Profile modes — "keyword" maps to pure-FTS ({vector: 0, fts: 1}), "semantic" to pure-vector ({vector: 1, fts: 0}), "hybrid" to the default fusion. The audited mode reflects the actual legs.

Scope Mechanism Use case
namespace filter on both legs code-vs-message partition
topics.topicKeys membership subquery on both legs “only from the deploy topic”
contentType (string or kinds[]) content_type IN (...) on both legs code vs plain text within one namespace
profile.conceptFilter concept annotation filter (vector leg) query sovereignty over clusters
budget: { maxUnits, maxChars, maxMs } maxUnits clamps the per-leg fetch; maxChars/maxMs post-process (SDK facade) LLM context guardrail — results.truncated

Every query writes one fused retrieval_log row — the per-leg breakdown lives in query_profile (counts, top scores, weights, mode, filters), the contributing ids in sources. The query’s identity is a FNV-1a hash of the raw query (the vector mode hashes the full query vector so the hash is a true query identity, not a truncated prefix).

The row is written off the query path by the durable audit.write job — retrieval never blocks on audit, and a failure is visible (typed), never swallowed. “The answer to ‘why did the agent say this?’ is always reconstructible from logs.” Every query also emits a sync hybrid.search log line (INFO), empty-store queries included.