# Does jevmem work?

In 72 real Claude Code sessions, Claude followed the project's saved decision in 66 with jevmem (66/72), against 28 with no project memory (28/72) and 67 with the same lines in a hand-written `CLAUDE.md` (67/72): jevmem does about as well as a hand-written `CLAUDE.md`, without you writing it.

Last updated: 2026-10-04 · jevmem 0.6.4 · [Install](https://avinash-jetwani.github.io/jevmem/install/) · [Source on GitHub](https://github.com/Avinash-jetwani/jevmem)

Every test set here was written by the author, and none is an independent benchmark. Each result below says what it was measured on, links its results file in the repository, and says what it does not show. To try it yourself: [install](https://avinash-jetwani.github.io/jevmem/install/).

## Does Claude act on the saved line?

![A/B results: followed the project's decision 28 of 72 with no project memory, 66 of 72 with jevmem, 67 of 72 with a hand-written CLAUDE.md.](https://avinash-jetwani.github.io/jevmem/img/results-light.svg)

| | No project memory | jevmem | Same lines in `CLAUDE.md` |
|---|---|---|---|
| Followed the project's decision | 28 of 72 | **66 of 72** | 67 of 72 |
| Tried a change the project forbids | 10 of 18 | **0 of 18** | — |
| Repeated an approach that had already failed | 3 of 15 | **0 of 15** | 0 of 15 |

So jevmem does about as well as a hand-written `CLAUDE.md`, without you writing it.
`CLAUDE.md` did better on a convention nothing in the prompt points at (3 of 3 against 0 of 3), so rules every task must follow still belong there.
Method, dates and builds: [What's new](https://avinash-jetwani.github.io/jevmem/whats-new/).

**Method.** 24 tasks in three small projects (a TypeScript API, a React front end, a plain-JavaScript CLI), each with a project memory of 34 to 42 lines. Each task's right answer depends on one saved line that the repository does not state. Every session is real Claude Code (`claude -p`, Claude Code 2.1.281, `claude-sonnet-5`) in a fresh copy of the project, three runs per task and arm. A check written and committed before any session ran decides whether the line was followed; there is no LLM judge. The no-memory and `CLAUDE.md` arms ran on 2026-09-28 ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/ab-2026-09-28.json)), and the jevmem arm ran on 2026-09-29 on the recall code that ships in 0.6 ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/ab-jevmem-2026-09-29-3b.json)). The full method and every row: [docs/benchmark.md](https://github.com/Avinash-jetwani/jevmem/blob/main/docs/benchmark.md#outcome-ab-does-claude-act-on-the-memory).

**Limits.** Three small projects, one model, sessions of at most 25 turns, three runs per task. `CLAUDE.md` did better on a convention nothing in the prompt points at (3 of 3 sessions against 0 of 3). The checks test whether the saved line was followed, not whether Claude's answer was better.

## Does the guard catch rule breaks?

On the guard's second held-out set, 274 tool calls in five new projects, the guard caught 66 of 68 rule breaks (66/68) and asked about 3–4 of the 206 calls that break no rule. The set was written after a day's trial of the guard in jevmem's own repository, before the fixes it measures, and was run once on the 0.6.0 build and once on the 0.6.1 build, both on 2026-09-30 ([0.6.0](https://github.com/Avinash-jetwani/jevmem/blob/main/results/guard-heldout-v2-2026-09-30.json), [0.6.1](https://github.com/Avinash-jetwani/jevmem/blob/main/results/guard-heldout-v2-2026-09-30-v061.json)). The one call that differs between the two runs breaks no rule and scored either side of the threshold.

In real Claude Code sessions on the A/B's 6 constraint tasks (18 sessions, rerun on the release build on 2026-09-30 with Claude Code 2.1.284), Claude did not attempt the forbidden change in any session (0/18; 10/18 with no memory in the 2026-09-28 run), and the guard checked all 78 of those sessions' Bash, Edit and Write calls and asked once ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/ab-guard-2026-09-30.json)).

**Limits.** The guard looks at the words a rule and a call share. A script or a make target that already exists and does the forbidden thing is missed, and so is a call that shares a single word with its rule. It checks Bash, Edit and Write calls, not MCP tools or the files a script changes when it runs. More: [the guard's limits](https://avinash-jetwani.github.io/jevmem/guard/#limits).

## Is deciding what to save fast, cheap and right?

![Median time to decide one message on 66 held-out turns: jevmem 0.28 s, six current LLMs 2.78 to 4.29 s.](https://avinash-jetwani.github.io/jevmem/img/benchmark-light.svg)

Deciding what to save takes 0.28 s and costs $0.00016 per message, in the background: Claude doesn't wait for it.
jevmem tied the best LLM on save or skip (98.5%); two LLMs were better at picking the kind of line.

**The full benchmark: accuracy, cost and how it was run**

66 held-out turns, all seven deciders given the same state ([method, regression set, pricing, p95, retries](https://github.com/Avinash-jetwani/jevmem/blob/main/docs/benchmark.md)). The six LLM rows are v0.4.2's run of 2026-09-23; jevmem's row is 0.6.0's run of the same set on 2026-09-30 ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/eval-heldout-2026-09-30-v060.json); [every mode, three builds](https://github.com/Avinash-jetwani/jevmem/blob/main/docs/benchmark.md#every-mode-three-builds)), where 0.5.9 and v0.4.2 score the same and cost less:

| Decider | save/skip | save+kind | contradictions | p50 | $/decision |
|---|---|---|---|---|---|
| GPT-6 Astra | 98.5% | 98.5% | 5/5 | 3,469 ms | $0.007489 |
| GPT-6 Luna | 93.9% | 93.9% | 5/5 | 2,927 ms | $0.000089 |
| Claude Fable 5.1 | 95.5% | 95.5% | 5/5 | 4,290 ms | $0.013256 |
| Claude Opus 5.5 | 97.0% | 97.0% | 5/5 | 2,784 ms | $0.005186 |
| Gemini 3.8 Flash | 92.4% | 92.4% | 5/5 | 2,850 ms | $0.001174 |
| Grok 4.7 | 90.9% | 90.9% | 4/5 | 3,320 ms | $0.004602 |
| **jevmem 0.6.0 `auto`** | **98.5%** | **95.5%** | **5/5** | **276 ms** | $0.000157 |

The 0.28 s is the Jev API decision (p95 527 ms; a saved turn's line costs one more request, $0.000159 per decision with it). Since v0.5.0 you do not wait for it: the `Stop` hook is async and its process exits in 12–14 ms (v0.5.6: 12 ms for the hook `jevmem init` registers, 14 ms for the plugin's), and the daemon records the decision 0.26–0.28 s after the hook starts ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/ops-2026-09-26-v056.json), [cost and latency](https://github.com/Avinash-jetwani/jevmem/blob/main/docs/cost.md)).

On 66 held-out turns, jevmem 0.6.0's median decision took 0.28 s, against 2.8–4.3 s for six current LLMs.
Its accuracy was within the LLMs' range: 98.5% save/skip (tied with GPT-6 Astra for highest) and 95.5% save+kind, against 90.9–98.5% for the LLMs. GPT-6 Astra (98.5%) and Claude Opus 5.5 (97.0%) were more accurate on save+kind; Claude Fable 5.1 tied; GPT-6 Luna, Gemini 3.8 Flash and Grok 4.7 were less accurate. It found 5/5 contradictions, as did five of the six LLMs.
GPT-6 Luna was cheaper ($0.000089 against $0.000157) but less accurate (93.9%) and about 11× slower.
Each row is a single run, and differences of one or two turns are within run-to-run noise; the LLM rows and jevmem's are a week apart. If the most accurate decision matters most, GPT-6 Astra or Claude Opus 5.5 are better, at about 33–48× the cost per decision and 10–13× the latency. jevmem is for when you want a fast, cheap decision on every message.

## Does the right line come back?

On the second retrieval held-out set (90 prompts over three new projects of 20, 80 and 250 lines, run once on 2026-09-28), jevmem 0.6's recall found 75/78 of the lines the prompts needed, and 0.5.9 found 55/78. Of those 78 lines, 18 are dead ends, which 0.5.9 cannot read; on the other 60, 0.6 found 57 and 0.5.9 found 55. Of the lines 0.6 put in front of Claude, 96/97 were wanted or fine, and 1/18 unrelated prompts got a line ([0.6](https://github.com/Avinash-jetwani/jevmem/blob/main/results/recall-heldout2-2026-09-28-now.json), [0.5.9](https://github.com/Avinash-jetwani/jevmem/blob/main/results/recall-heldout2-2026-09-28-v059.json)).

**Limits.** Prompts that need two lines got both in 3 of 6. With more than 250 live lines, only the 250 that share the most words with the prompt are asked about (on a 500-line dev file, recall was 37/46). That the right lines reach Claude is tested; whether its answers get better is not.

## Dead ends, background subagents and planted lines

- **Dead ends.** On decide's third held-out set (100 turns in five new projects, run on the release build on 2026-09-30), jevmem saved 25 of 25 dead ends, each with its reason ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/dead-ends-heldout-v3-2026-09-30-v060.json)).
- **Background subagents.** On 30 real Claude Code 2.1.281 sessions, 14 of them with a background subagent, replayed through the 0.6.0 build on 2026-09-30, turns were saved or skipped right in 32 of 33 (0.5.9: 27 of 33) ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/stops-heldout-v4-2026-09-30-v060.json)).
- **Planted lines.** On a 44-line test set (2026-09-25), the check on lines jevmem did not write blocked 20 of 22 planted lines, with 0 false blocks on 22 legitimate rules; the 2 it missed were instructions disguised as normal process ([results](https://github.com/Avinash-jetwani/jevmem/blob/main/results/memory-injection-2026-09-25-run1.json)).

## Can a local model do Jev's job?

Not well enough to offer, in the one test so far. On 2026-10-02, jevmem 0.6.4 was pointed at a local model server, Ollaya 0.9.0, on an Apple M4 with 16 GB, next to a Jev run of the same set the same afternoon. On the 66 held-out turns, in jevmem's default mode, Jev was right on save or skip for 65/66 turns at a median of 0.23 s a turn; `winnow:e4b` for 60/66 at 28.5 s; `laya:typed-decisions` for 19/66.

The results files: [Jev](https://github.com/Avinash-jetwani/jevmem/blob/main/results/local-model-2026-10-02-jev.json), [winnow:e4b](https://github.com/Avinash-jetwani/jevmem/blob/main/results/local-model-2026-10-02-winnow-e4b.json), [laya:typed-decisions](https://github.com/Avinash-jetwani/jevmem/blob/main/results/local-model-2026-10-02-laya-typed-decisions.json). The commands as run: [scripts/local-model-jev.sh](https://github.com/Avinash-jetwani/jevmem/blob/main/scripts/local-model-jev.sh), [scripts/local-model-ollaya.sh](https://github.com/Avinash-jetwani/jevmem/blob/main/scripts/local-model-ollaya.sh).

**Limits.** One run each, on one Mac, with two models. The set was written for Jev's behaviour. Both local models needed the client's timeout raised from 10 s to 180 s. `winnow:e4b` was measured on these 66 turns only, not on recall or the guard.

## What these results do not show

- **Early:** 0.6.4; every eval set was written by the author, and none is an independent benchmark.
- **Not the most accurate:** GPT-6 Astra and Claude Opus 5.5 scored higher on save+kind; jevmem's edge is speed and cost.
- **Recall quality is not measured:** that relevant lines are injected is tested; whether answers get better is not.
- **Long-run drift is not measured:** the harness covers five-turn sessions, not weeks of use.
- **Automatic capture is Claude Code only** (and Codex while `jevmem watch` runs); Cursor and Claude Desktop save only when the agent calls `add_memory`.
- **The poisoning gate is a filter, not a guarantee:** it missed 2 of 22 planted lines in our eval (2026-09-25; both worded as ordinary process), it does not apply when an agent opens `JEVMEM.md` as a file, and on a fresh clone its first check costs one noul per line. Review `JEVMEM.md` diffs like code ([SECURITY.md](https://github.com/Avinash-jetwani/jevmem/blob/main/SECURITY.md#memory-poisoning)).
- **Jev outages delay turns, up to a limit; other Jev errors drop them:** each Jev call has a 2 s budget. When it times out, the network fails, or Jev answers 408, 429 or 5xx (529 included), the scrubbed turn waits in `.jevmem/queue.jsonl` and is retried with backoff (15 s, 30 s, then 1, 2 and 5 min, then every 10 min) on the next hook run or by the idle daemon, in order, and saved once. A turn still unsaved after 24 hours, or past 200 queued turns, is dropped. Any other error is not retried and drops the turn at once: a 400 from Jev, for example, or a 401 when the key is wrong, which drops every turn until the key is fixed. Each drop leaves a line in `.jevmem/log.jsonl`, and the retry-queue line of `jevmem stats` counts them.
- **A plugin update can leave an open session without the hooks:** when the plugin synced from claude.ai updates (Claude Code downloads updates in the background each time it starts), Claude Code moves the previous copy aside, and a session that was already open with it loses jevmem's hooks, the guard included, until you run `/reload-plugins` there or start a new session; Claude Code shows a hook error and goes on without them ([anthropics/claude-code#97847](https://github.com/anthropics/claude-code/issues/97847)). `jevmem doctor` shows the version on disk.

## The limits, in short

- jevmem needs a TypeSafe API key ([where to get one, and the install steps](https://avinash-jetwani.github.io/jevmem/install/)).
- Message text is sent to TypeSafe to be scored, with common secrets scrubbed first ([what leaves your machine](https://avinash-jetwani.github.io/jevmem/privacy/)).
- It is automatic in Claude Code, automatic in Codex while `jevmem watch` runs, and in Cursor only when the agent calls it ([what each tool does](https://avinash-jetwani.github.io/jevmem/install/#works-with)).
- Rules every task must follow still belong in `CLAUDE.md` ([jevmem next to CLAUDE.md](https://avinash-jetwani.github.io/jevmem/compare/)).
- The guard is a backstop, not a sandbox ([what it misses](https://avinash-jetwani.github.io/jevmem/guard/#limits)).

Every limit, with the numbers: [the FAQ](https://avinash-jetwani.github.io/jevmem/faq/#what-are-jevmems-limits).
