Skip to the content.

memshelf — Roadmap

Milestones are deliberately thin. M0 validates the pattern with zero new code; every later milestone must justify itself against what M0 already achieves.

M0 — Pattern validation, no code — complete (2026-07-13 → 2026-07-22)

Exit criteria met: recall test 5/5, ledger + recall-cost numbers written down (demo.md), and the annoyance log became the M1 backlog verbatim. Case B closed 2026-07-22 (33 episodes, zero loss; verdict episode on the shelf).

Prove the loop works with docshelf as-is plus conventions. Protocol, kit, and measurement methodology: docs/M0.md; the prompt-only skill and recall-rule snippet live in adapters/claude-code/.

Exit criteria: 5 known-answer recall questions answered correctly from a fresh session via INDEX navigation; ledger numbers written down; the annoyance log filled. That log is the M1 backlog.

M1 — memshelf-mcp thin server

Only what M0 proved annoying, expected:

Exit criteria: dogfooded on two real projects for two weeks; a full shelve→compact→recall cycle survives without manual repair; doctor clean.

M2 — Policy, hygiene & the context advisor

Exit criteria: a shelf with 100+ episodes keeps INDEX under ~10 KB and recall precision doesn’t degrade (re-run the M0 question set); the advisor’s shelve proposals are accepted (not overridden) most of the time in dogfood use.

M3 — Retrieval upgrades, reuse layer & second surface

Exit criteria: search-miss rate measurably better than grep baseline on the dogfood shelves; one non-author user runs the chat-project flow from docs alone; one real “fork from episode” session succeeds end-to-end.

Explicitly deferred