memgovern 0.5.0: a memgovern audit CLI that shows who your agent's memory trusts
Everybody audits what goes into agent memory. Almost nobody audits who put it there.
pip install memgovern — then:
memgovern --db ~/.local/share/memgovern/memory.db audit
trust report
============
user trust=0.833 writes=41 wins=8 losses=0
tooling trust=0.500 writes=12 wins=0 losses=0
sketchy-feed trust=0.200 writes=9 wins=0 losses=6 <-- LOW: treat this source's writes with suspicion
quarantined writes: 2
#12 [tripwire] key='deploy.flags' source='sketchy-feed'
reason: tripwire: injection-marker:ignore previous instructions
#15 [conflict] key='user.theme' source='sketchy-feed'
reason: tripwire: low-trust-source-vs-high-trust-holder
memories: alive=27 pending=2 superseded=6
expired_alive=0 pending_conflicts=1 audit_events=63
What is memgovern
Everyone builds the store and search side of agent memory. memgovern does
the neglected half: write and delete — when a memory should fade (decay &
forgetting), when it should die (tombstones, never silent erasure), and who
wins when two memories disagree (conflict arbitration). Zero dependencies,
SQLite under the hood.
It deliberately does not compete on memory storage — that's claude-mem's
territory. memgovern is only the governance layer: per-source credibility
scores, quarantine-before-arbitrate, and a structural poisoning defense, with
the core package kept stdlib-only.
Why the trust report exists
Memory poisoning is a measured attack surface. The CodePoisonRAG study
(adversa.ai) implanted 85 poisoning artifacts across three code-generation
agents with 0.80–0.93 success rates — nothing in the write path was asking
who was writing. memgovern has kept a per-source trust ledger since v0.2:
sources earn trust by winning fair arbitrations (Bayesian-smoothed win rate,
prior 0.5, 30-day half-life decay toward neutral), and three tripwires —
injection markers, burst writes, and a low-trust source contradicting a
high-trust holder — quarantine suspicious writes for human review instead of
applying them.
Until now that ledger was reachable only through the Python API
(store.trust_report()) or the memory_trust_report MCP tool. v0.5 adds the
human terminal for it: memgovern audit. A source below 0.35 trust gets the
LOW flag. --json emits machine-readable output; --source NAME filters to
one source.
Honest limits, as always: the defense is structural, not semantic — a patient
attacker can farm trust with benign writes, and the marker list misses
paraphrases. The report makes low-credibility sources visible; it doesn't
pretend poisoning is solved. The code, the ledger math, and the tripwire
rules are all documented in the README: hahahahahahahahah6/memgovern.
Top comments (0)