Repository navigation
Evaluation proposal: reproducible scope-isolation benchmark for DeerMem long-term memory #5198
PeaceMaker-best
started this conversation in
Ideas
Replies: 1 comment
|
One guardrail I’d add is an end-to-end “absence” assertion, not just a rejected write. For each transient canary, run a fresh session through the normal retrieval path and assert that it is not surfaced for the same user, another user, the default agent, or a different custom agent. That distinguishes “the extractor happened not to propose a memory” from “a scoped fact can never become visible later.” I’d report those separately: admission false positives, storage-routing leaks, and retrieval leaks. A small matrix of the same canary across user × agent × fresh-session boundaries would make a regression much easier to localize. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
I would like to contribute a small, reproducible evaluation for DeerMem's long-term-memory admission and identity-isolation boundaries.
The proposed first slice does not change production memory behavior. It turns two already-observed failure classes into a versioned regression benchmark:
user_id x agent_namebucket.If the scope is useful, I am happy to implement it as a standalone benchmark following the repository's existing benchmark layout.
Motivation
DeerFlow now has strong deterministic safeguards after:
scope,durability, andauthoritylabels plus a fail-closed write gate; and__default__bucket, causing personality contamination in ordinary conversations.The remaining evaluation gap is one layer above the current unit tests:
test_memory_scope_gate.pyproves that the deterministic gate behaves correctly when supplied with already-classified model output.Because the extractor is model-driven, production code can remain unchanged while classification quality regresses. For example, a model may label a one-run permission (
you may force-push this branch) or a repository-local constraint (this PR must not change the API) as durable user memory. The deterministic gate only protects the system if the required labels are correct.Proposed first slice
Add a standalone evaluation under a path such as:
It would contain two deliberately small suites.
1. Semantic admission suite
Use synthetic, versioned conversations with unique non-sensitive canary terms and run them through the production DeerMem extraction prompt, normalization, deterministic gate, and temporary storage.
Representative cases:
2. Identity-routing suite
Exercise the production queue/storage boundary with deterministic model output and verify:
__default__or Agent B;This suite is deterministic and should not require a live provider.
Metrics
The report would keep the measurements small and directly tied to the failure modes:
durable_retention_rate: expected durable canaries persisted / expected durable canaries;unsafe_persistence_rate: transient, project, or transactional canaries persisted / unsafe canaries;cross_agent_contamination_rate: facts visible outside their expected agent bucket / scoped facts;cross_user_contamination_rate: facts visible outside their expected user bucket / scoped facts; andatomic_correction_success_rate: valid user-level replacements that both remove the old fact and persist the replacement.Reproducibility and safety boundaries
Non-goals
This first slice would not:
If the evaluation exposes a repeatable production failure, any prompt/gate behavior change would be proposed separately and would use this benchmark as before/after evidence.
Relationship to existing work
Questions for maintainers
backend/scripts/benchmark/be the preferred shape for this evaluation?If this direction is acceptable, I can keep the implementation to one evaluation-only PR with no production behavior changes.
All reactions