Retrieval
The pipeline
Section titled “The pipeline” query text │ ▼ embed │ ▼ scoping filters: namespace · topicKeys · contentType (both legs) · conceptFilter (vector leg only) │ ▼ vector leg: ANN over the model's vector table (max(16 × fetched, 64) candidates, exact cosine re-rank) │ ▼ FTS5 lexical leg (same filters) │ ▼ RRF fusion (k = 60) │ ▼ recency boost (optional) │ ▼ budgets: maxUnits clamps the per-leg fetch; maxChars truncates content; maxMs flags truncated │ ▼ ranked results + one audit rowThe vector leg
Section titled “The vector leg”The retriever asks each leg for limit × 2 results. The vector leg runs
an approximate nearest-neighbour search over the active model’s vector
table, then re-ranks the candidates with exact cosine in SQL
(vector_top_k where the libSQL vector index is available, an exact
ORDER BY scan otherwise). The ANN stage fetches max(16 × fetched, 64)
candidates because the scoping filters apply after the fetch (that is
max(32 × limit, 64) at the default per-leg fetch); the exact-scan
fallback walks the filtered table.
Scoping is applied per leg, before fusion: namespace,
topics.topicKeys, and contentType restrict both legs;
profile.conceptFilter scopes the vector leg only. The store also
exposes a deterministic beam-descent leaf resolver
(Store.resolveTopics): it walks from __root__, keeps the top
beamWidth children per level (maxDepth caps the walk, ties break
lexicographically), and returns the leaves a query should search. The
SDK query path does not call it yet — the resolver is covered by the
topics benchmarks and tests, and wiring it into hybrid is a known gap.
The FTS5 leg runs after the vector leg and catches exact names/identifiers the vector leg might miss.
Hybrid fusion
Section titled “Hybrid fusion”Algorithm — Reciprocal Rank Fusion:
for each leg, rank its hits: rrf_score(memory) = Σ_legs 1 / (k + rank_leg(memory)) k = 60 (default)
final score = rrf_score × (1 + recencyBoost × recencyFactor)recencyFactor = 1 / (1 + ageDays) (hyperbolic decay)Then re-rank (stable — equal scores keep fused order). The fused set is
not graph-augmented on this path: hybrid disables the store’s
per-leg post-pass (augment: false) and never calls
Store.augmentSearchResults, so query results carry no supersededBy /
conflicts / supportedBy fields and no confirmed-supersede demotion.
The store’s own search legs (searchVector, searchFts5, recency)
apply that post-pass by default; wiring it onto the fused set is a known
gap.
Profile modes — "keyword" maps to pure-FTS ({vector: 0, fts: 1}),
"semantic" to pure-vector ({vector: 1, fts: 0}), "hybrid" to the
default fusion. The audited mode reflects the actual legs.
Scoping & sovereignty
Section titled “Scoping & sovereignty”| Scope | Mechanism | Use case |
|---|---|---|
namespace |
filter on both legs | code-vs-message partition |
topics.topicKeys |
membership subquery on both legs | “only from the deploy topic” |
contentType (string or kinds[]) |
content_type IN (...) on both legs |
code vs plain text within one namespace |
profile.conceptFilter |
concept annotation filter (vector leg) | query sovereignty over clusters |
budget: { maxUnits, maxChars, maxMs } |
maxUnits clamps the per-leg fetch; maxChars/maxMs post-process (SDK facade) |
LLM context guardrail — results.truncated |
Every query writes one fused retrieval_log row — the per-leg
breakdown lives in query_profile (counts, top scores, weights, mode,
filters), the contributing ids in sources. The query’s identity is a
FNV-1a hash of the raw query (the vector mode hashes the full query
vector so the hash is a true query identity, not a truncated prefix).
The row is written off the query path by the durable audit.write
job — retrieval never blocks on audit, and a failure is visible (typed),
never swallowed. “The answer to ‘why did the agent say this?’ is always
reconstructible from logs.” Every query also emits a sync
hybrid.search log line (INFO), empty-store queries included.