Persistent agent memory · Evidence release

Beyond recall
accuracy.

A two-axis evaluation of whether persistent memory preserves authority and whether tool-using agents resume safely after interruption.

Evidence boundary: 700 seeded integrity trials, installed-product crash checks, controlled four-arm replay, and one explicitly non-causal live pilot.
700Integrity trialsSeven attack categories
400Represented passesC, D, F, and L
300Not representableA, H, and I
0Fail or errorAmong represented trials

01 · Research question

Relevant memory is not necessarily authorized memory.

Retrieval benchmarks ask whether an agent can find a fact. This work asks two earlier questions: may the fact carry authority, and what can the agent conclude about an interrupted external action?

AtMem is evaluated as the authority plane. AtFlows is evaluated as an optional observation plane whose delivery cannot authorize, suppress, or repeat a tool call.

02 · Memory integrity

Authority before ranking

Each category contains 100 seeded trials. Unsupported authority transitions remain NOT_REPRESENTABLE; they are not converted into passes.

PASSNOT_REPRESENTABLE
0 / 300Factual contamination
0 / 100Secret retention
100 / 100Taint preservation
0 / 200Trust laundering

03 · Operational continuity

Preserve uncertainty after a crash

The process was killed after the sixth public retail tool committed its state change but before the worker received the response.

Repeat requests after restart

Lower is safer in this fault cell. The destination itself rejected repeated exchanges, so the attributable result is avoiding a repeat request—not preventing a second external effect.

AtMem decisionneeds_confirmation

No declared query or destination-enforced idempotency contract existed.

No-fault replay18 / 18native requests matched in every arm
Tool sequence6identical native tool calls
Final stateEqualsame trajectory and store hash
Provider calls0recorded-response qualification

04 · Negative-result case study

The live pilot is evidence,
not a causal comparison.

One exposed task was run once per arm. The baseline scored 1; AtMem, AtFlows, and both scored 0. The first model responses differed before tools ran, so the result cannot establish a product benefit or penalty.

ConfigurationRewardCallsEst. cost

05 · Claim ledger

What the evidence does—and does not—show

Supported

  • 400 represented integrity trials passed.
  • AtMem made zero repeat requests in one controlled interruption.
  • One recorded no-fault trajectory remained identical across four configurations.
  • Installed document cases ended with one exact output.

Not supported

  • All seven attack categories passed.
  • General retrieval or factual superiority.
  • Distributed exactly-once execution.
  • A production-wide recovery rate or causal live-pilot comparison.

06 · Technical paper

Read the complete method

The complete 12-page paper is rendered below so it remains readable inside the Hugging Face Space.

Technical paper page 1 Technical paper page 2 Technical paper page 3 Technical paper page 4 Technical paper page 5 Technical paper page 6 Technical paper page 7 Technical paper page 8 Technical paper page 9 Technical paper page 10 Technical paper page 11 Technical paper page 12