jevmemjevmem

Does jevmem work?

In 72 real Claude Code sessions, Claude followed the project's saved decision in 66 with jevmem (66/72), against 28 with no project memory (28/72) and 67 with the same lines in a hand-written CLAUDE.md (67/72): jevmem does about as well as a hand-written CLAUDE.md, without you writing it.

Last updated: · jevmem 0.6.4 · Install · Source on GitHub

Every test set here was written by the author, and none is an independent benchmark. Each result below says what it was measured on, links its results file in the repository, and says what it does not show. To try it yourself: install.

Does Claude act on the saved line? #

A/B results: followed the project's decision 28 of 72 with no project memory, 66 of 72 with jevmem, 67 of 72 with a hand-written CLAUDE.md.A/B results: followed the project's decision 28 of 72 with no project memory, 66 of 72 with jevmem, 67 of 72 with a hand-written CLAUDE.md.

No project memory jevmem Same lines in CLAUDE.md
Followed the project's decision 28 of 72 66 of 72 67 of 72
Tried a change the project forbids 10 of 18 0 of 18 —
Repeated an approach that had already failed 3 of 15 0 of 15 0 of 15

So jevmem does about as well as a hand-written CLAUDE.md, without you writing it. CLAUDE.md did better on a convention nothing in the prompt points at (3 of 3 against 0 of 3), so rules every task must follow still belong there. Method, dates and builds: What's new.

Method. 24 tasks in three small projects (a TypeScript API, a React front end, a plain-JavaScript CLI), each with a project memory of 34 to 42 lines. Each task's right answer depends on one saved line that the repository does not state. Every session is real Claude Code (claude -p, Claude Code 2.1.281, claude-sonnet-5) in a fresh copy of the project, three runs per task and arm. A check written and committed before any session ran decides whether the line was followed; there is no LLM judge. The no-memory and CLAUDE.md arms ran on 2026-09-28 (results), and the jevmem arm ran on 2026-09-29 on the recall code that ships in 0.6 (results). The full method and every row: docs/benchmark.md.

Limits. Three small projects, one model, sessions of at most 25 turns, three runs per task. CLAUDE.md did better on a convention nothing in the prompt points at (3 of 3 sessions against 0 of 3). The checks test whether the saved line was followed, not whether Claude's answer was better.

Does the guard catch rule breaks? #

On the guard's second held-out set, 274 tool calls in five new projects, the guard caught 66 of 68 rule breaks (66/68) and asked about 3–4 of the 206 calls that break no rule. The set was written after a day's trial of the guard in jevmem's own repository, before the fixes it measures, and was run once on the 0.6.0 build and once on the 0.6.1 build, both on 2026-09-30 (0.6.0, 0.6.1). The one call that differs between the two runs breaks no rule and scored either side of the threshold.

In real Claude Code sessions on the A/B's 6 constraint tasks (18 sessions, rerun on the release build on 2026-09-30 with Claude Code 2.1.284), Claude did not attempt the forbidden change in any session (0/18; 10/18 with no memory in the 2026-09-28 run), and the guard checked all 78 of those sessions' Bash, Edit and Write calls and asked once (results).

Limits. The guard looks at the words a rule and a call share. A script or a make target that already exists and does the forbidden thing is missed, and so is a call that shares a single word with its rule. It checks Bash, Edit and Write calls, not MCP tools or the files a script changes when it runs. More: the guard's limits.

Is deciding what to save fast, cheap and right? #

Median time to decide one message on 66 held-out turns: jevmem 0.28 s, six current LLMs 2.78 to 4.29 s.Median time to decide one message on 66 held-out turns: jevmem 0.28 s, six current LLMs 2.78 to 4.29 s.

Deciding what to save takes 0.28 s and costs $0.00016 per message, in the background: Claude doesn't wait for it. jevmem tied the best LLM on save or skip (98.5%); two LLMs were better at picking the kind of line.

The full benchmark: accuracy, cost and how it was run

66 held-out turns, all seven deciders given the same state (method, regression set, pricing, p95, retries). The six LLM rows are v0.4.2's run of 2026-09-23; jevmem's row is 0.6.0's run of the same set on 2026-09-30 (results; every mode, three builds), where 0.5.9 and v0.4.2 score the same and cost less:

Decider save/skip save+kind contradictions p50 $/decision
GPT-6 Astra 98.5% 98.5% 5/5 3,469 ms $0.007489
GPT-6 Luna 93.9% 93.9% 5/5 2,927 ms $0.000089
Claude Fable 5.1 95.5% 95.5% 5/5 4,290 ms $0.013256
Claude Opus 5.5 97.0% 97.0% 5/5 2,784 ms $0.005186
Gemini 3.8 Flash 92.4% 92.4% 5/5 2,850 ms $0.001174
Grok 4.7 90.9% 90.9% 4/5 3,320 ms $0.004602
jevmem 0.6.0 auto 98.5% 95.5% 5/5 276 ms $0.000157

The 0.28 s is the Jev API decision (p95 527 ms; a saved turn's line costs one more request, $0.000159 per decision with it). Since v0.5.0 you do not wait for it: the Stop hook is async and its process exits in 12–14 ms (v0.5.6: 12 ms for the hook jevmem init registers, 14 ms for the plugin's), and the daemon records the decision 0.26–0.28 s after the hook starts (results, cost and latency).

On 66 held-out turns, jevmem 0.6.0's median decision took 0.28 s, against 2.8–4.3 s for six current LLMs. Its accuracy was within the LLMs' range: 98.5% save/skip (tied with GPT-6 Astra for highest) and 95.5% save+kind, against 90.9–98.5% for the LLMs. GPT-6 Astra (98.5%) and Claude Opus 5.5 (97.0%) were more accurate on save+kind; Claude Fable 5.1 tied; GPT-6 Luna, Gemini 3.8 Flash and Grok 4.7 were less accurate. It found 5/5 contradictions, as did five of the six LLMs. GPT-6 Luna was cheaper ($0.000089 against $0.000157) but less accurate (93.9%) and about 11× slower. Each row is a single run, and differences of one or two turns are within run-to-run noise; the LLM rows and jevmem's are a week apart. If the most accurate decision matters most, GPT-6 Astra or Claude Opus 5.5 are better, at about 33–48× the cost per decision and 10–13× the latency. jevmem is for when you want a fast, cheap decision on every message.

Does the right line come back? #

On the second retrieval held-out set (90 prompts over three new projects of 20, 80 and 250 lines, run once on 2026-09-28), jevmem 0.6's recall found 75/78 of the lines the prompts needed, and 0.5.9 found 55/78. Of those 78 lines, 18 are dead ends, which 0.5.9 cannot read; on the other 60, 0.6 found 57 and 0.5.9 found 55. Of the lines 0.6 put in front of Claude, 96/97 were wanted or fine, and 1/18 unrelated prompts got a line (0.6, 0.5.9).

Limits. Prompts that need two lines got both in 3 of 6. With more than 250 live lines, only the 250 that share the most words with the prompt are asked about (on a 500-line dev file, recall was 37/46). That the right lines reach Claude is tested; whether its answers get better is not.

Dead ends, background subagents and planted lines #

Can a local model do Jev's job? #

Not well enough to offer, in the one test so far. On 2026-10-02, jevmem 0.6.4 was pointed at a local model server, Ollaya 0.9.0, on an Apple M4 with 16 GB, next to a Jev run of the same set the same afternoon. On the 66 held-out turns, in jevmem's default mode, Jev was right on save or skip for 65/66 turns at a median of 0.23 s a turn; winnow:e4b for 60/66 at 28.5 s; laya:typed-decisions for 19/66.

The results files: Jev, winnow:e4b, laya:typed-decisions. The commands as run: scripts/local-model-jev.sh, scripts/local-model-ollaya.sh.

Limits. One run each, on one Mac, with two models. The set was written for Jev's behaviour. Both local models needed the client's timeout raised from 10 s to 180 s. winnow:e4b was measured on these 66 turns only, not on recall or the guard.

What these results do not show #

The limits, in short #

Every limit, with the numbers: the FAQ.