Conversation
BenchmarkDB.recall never passed wrrf_k to pg_recall, so the benchmark path always used the pg_recall default of 60 while the production handler (handlers/recall.py) passes settings.WRRF_K. CORTEX_MEMORY_WRRF_K now moves the benchmark exactly as it moves a live recall. Two tests pin the contract (env override reaches pg_recall; default equals the production setting). Measured scope, 2026-09-24, pgvector/pgvector:pg16, 12 memories, rerank off: on the PostgreSQL path this changes no score. The recall_memories() SQL function fuses max-normalised signal scores with a weighted sum (no rank transform); p_wrrf_k only enters the agent-topic boost term. Top-5 scores were identical to six decimals at k=10, 60 and 1000. Rank-based RRF with k runs only on the SQLite path (sqlite_store_search.py). A k sweep through reproduce.sh therefore cannot measure first-stage RRF sensitivity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
… output BenchmarkDB was limited to PostgreSQL, so the default SQLite backend and its rank-based RRF were never measured through the production recall pipeline. The new mode is opt-in (backend="sqlite" or CORTEX_BENCH_BACKEND=sqlite), uses a throwaway database file removed on close, and keeps PostgreSQL as the default. The LongMemEval harness now reports the backend it ran on. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
The SQLite recall path summed weight/(k+rank) per signal, while PL/pgSQL recall_memories sums weighted scores divided by each signal's pool maximum (vector shifted by +1). Port the PostgreSQL semantics to the SQLite path: cosine from the sqlite-vec L2 distance, BM25 from FTS5, heat_base, and an exp(-0.01 * age_days) recency. The trigram signal has no SQLite counterpart and stays absent. wrrf_k now only scales the agent-topic bonus, as on PostgreSQL. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
The pilot ingests each LongMemEval-S question once, restores the store from an in-memory snapshot before every arm (a recall reinforces the memories it returns), and recalls under the previous rank-based fusion and the score fusion, with and without the reranker. The analysis script runs a paired percentile bootstrap. Moving the fusion into _fused_scores keeps recall_memories under the method-size cap without changing its result. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
paired_n500.jsonl holds one line per question with the top-10 session ids and first-hit rank for four arms (RRF k=60 and score fusion, each with and without the reranker); summary.json holds the paired bootstrap (10 000 resamples, seed 20260924), overall, on the first 50 questions, and per category. The runs used the code of commit 410ca2f; the later refactor reproduces the recorded retrievals on questions 100 to 103. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe Signed-off-by: cdeust <cdeust@icloud.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Draft experiment, not for merge. It ports the PostgreSQL max-normalised score fusion into the SQLite recall path and adds a SQLite mode to the benchmark database plus a paired benchmark, to test whether the fusion explains the gap between the SQLite and PostgreSQL backends. It does not: on LongMemEval-S (500 questions, production pipeline with reranker) the paired MRR difference is -0.009 [-0.022, +0.005] and recall@10 is -0.018 [-0.032, -0.006] against the existing RRF (k=60). The branch is kept for further work on the SQLite backend.
No issue is closed by this draft.
Type of change
Test plan
check_craftsmanship.pypass.test_benchmark_db_still_reachable_lazily_and_fails_loudly_without_psycopg(psycopg is installed on the test machine).Audit notes
Results, per-question data and the analysis script are in
benchmarks/results/sqlite_score_fusion/. Signal differences left as they are: no trigram signal on SQLite, BM25 instead ofts_rank_cd, rawheat_base, nopost_tool_capturefilter. A recall reinforces heat, so the paired driver restores the store from a snapshot before each arm.Reviewer checklist
🤖 Generated with Claude Code
https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe