Skip to content

Experiment: SQLite score fusion and paired LongMemEval benchmark (do not merge) - #646

Draft
cdeust wants to merge 6 commits into
mainfrom
exp/sqlite-score-fusion
Draft

cdeust wants to merge 6 commits into
mainfrom
exp/sqlite-score-fusion

Conversation

@cdeust

@cdeust cdeust commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Summary

Draft experiment, not for merge. It ports the PostgreSQL max-normalised score fusion into the SQLite recall path and adds a SQLite mode to the benchmark database plus a paired benchmark, to test whether the fusion explains the gap between the SQLite and PostgreSQL backends. It does not: on LongMemEval-S (500 questions, production pipeline with reranker) the paired MRR difference is -0.009 [-0.022, +0.005] and recall@10 is -0.018 [-0.032, -0.006] against the existing RRF (k=60). The branch is kept for further work on the SQLite backend.

No issue is closed by this draft.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would change existing behavior)
  • Refactor (no functional change; rules/coding-standards.md compliance)
  • Documentation only
  • Experiment with a negative result, kept as a draft

Test plan

  • 100 SQLite tests pass locally; ruff, format and check_craftsmanship.py pass.
  • One existing test fails before and after this change: test_benchmark_db_still_reachable_lazily_and_fails_loudly_without_psycopg (psycopg is installed on the test machine).
  • The test asserting that k changes SQLite scores is inverted, since k now only affects the agent-topic bonus, as on PostgreSQL.
  • A paired PostgreSQL run on the same 500 questions is still needed to attribute the remaining gap (SQLite RRF MRR 0.867 against 0.905 in the README, unpaired).

Audit notes

Results, per-question data and the analysis script are in benchmarks/results/sqlite_score_fusion/. Signal differences left as they are: no trigram signal on SQLite, BM25 instead of ts_rank_cd, raw heat_base, no post_tool_capture filter. A recall reinforces heat, so the paired driver restores the store from a snapshot before each arm.

Reviewer checklist

  • CHANGELOG.md updated (not applicable to an experiment).
  • No secrets or personal data in the diff.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe

cdeust and others added 6 commits September 24, 2026 19:41
BenchmarkDB.recall never passed wrrf_k to pg_recall, so the benchmark path
always used the pg_recall default of 60 while the production handler
(handlers/recall.py) passes settings.WRRF_K. CORTEX_MEMORY_WRRF_K now moves
the benchmark exactly as it moves a live recall. Two tests pin the contract
(env override reaches pg_recall; default equals the production setting).

Measured scope, 2026-09-24, pgvector/pgvector:pg16, 12 memories, rerank off:
on the PostgreSQL path this changes no score. The recall_memories() SQL
function fuses max-normalised signal scores with a weighted sum (no rank
transform); p_wrrf_k only enters the agent-topic boost term. Top-5 scores
were identical to six decimals at k=10, 60 and 1000. Rank-based RRF with k
runs only on the SQLite path (sqlite_store_search.py). A k sweep through
reproduce.sh therefore cannot measure first-stage RRF sensitivity.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>
… output

BenchmarkDB was limited to PostgreSQL, so the default SQLite backend and its
rank-based RRF were never measured through the production recall pipeline.
The new mode is opt-in (backend="sqlite" or CORTEX_BENCH_BACKEND=sqlite),
uses a throwaway database file removed on close, and keeps PostgreSQL as the
default. The LongMemEval harness now reports the backend it ran on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>
The SQLite recall path summed weight/(k+rank) per signal, while PL/pgSQL
recall_memories sums weighted scores divided by each signal's pool maximum
(vector shifted by +1). Port the PostgreSQL semantics to the SQLite path:
cosine from the sqlite-vec L2 distance, BM25 from FTS5, heat_base, and an
exp(-0.01 * age_days) recency. The trigram signal has no SQLite counterpart
and stays absent. wrrf_k now only scales the agent-topic bonus, as on
PostgreSQL.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>
The pilot ingests each LongMemEval-S question once, restores the store from an
in-memory snapshot before every arm (a recall reinforces the memories it
returns), and recalls under the previous rank-based fusion and the score
fusion, with and without the reranker. The analysis script runs a paired
percentile bootstrap. Moving the fusion into _fused_scores keeps
recall_memories under the method-size cap without changing its result.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>
paired_n500.jsonl holds one line per question with the top-10 session ids and
first-hit rank for four arms (RRF k=60 and score fusion, each with and without
the reranker); summary.json holds the paired bootstrap (10 000 resamples,
seed 20260924), overall, on the first 50 questions, and per category. The runs
used the code of commit 410ca2f; the later refactor reproduces the recorded
retrievals on questions 100 to 103.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mkvy5MtGG89cNLuQSU8xLe
Signed-off-by: cdeust <cdeust@icloud.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant