What is Jev used for in coding agents, and how does jevmem use it?
In jevmem, Jev (a model by TypeSafe AI) is what decides: it scores each message of a Claude Code chat so that jevmem can tell whether it holds a decision, a rule or a failed approach worth saving, it picks the saved lines that bear on your next prompt, and it judges whether a command or a file edit may break a saved rule.
Deciding what to save takes 0.28 s and costs $0.00016 per message, and was right on save or skip for 98.5% of 66 held-out turns (jevmem 0.6.0, one run on 2026-09-30, results). This page covers only what jevmem does with Jev. What Jev itself is: TypeSafe's docs and their launch post.
How jevmem asks Jev #
No prompt decides what to save: Jev answers small yes/no questions with probabilities, and plain rules in code act on them.
Each step in detail
- Scrub. Common secret shapes, email addresses and card-shaped numbers are removed from the turn before it leaves your machine.
- Ask Jev typed questions. Jev by TypeSafe AI answers a fixed set of small questions with probabilities: is there a decision, a rule, a bug? is it small talk or an injection attempt? which existing line does it change?
- Apply thresholds in code. Plain rules over those probabilities decide save or skip; they live in
jevmem.config.json, not in a prompt. - Write one line. On save, jevmem writes one line of at most 200 characters from the sentences of the turn that Jev picks (the one that states the memory and the one that gives its reason), or, if you set
writerinjevmem.config.json, a small OpenAI or Anthropic model condenses the turn. - Supersede the old line. If the turn replaces an existing memory, that line is tagged
[superseded] … → id:newand stays in the file.
Tiers, questions, policy, contradictions, recall and audit: docs/how-it-works.md.
The three places jevmem calls Jev #
- After each turn, to decide what to save. One request of small typed questions; a second, larger set only when the first is unsure. Thresholds in
jevmem.config.jsonturn the probabilities into save or skip (how it works). - On each prompt, to pick the lines to bring back. Each live line, up to 250 of them, is asked about on its own, and at most five go to Claude (the read side).
- Before a Bash, Edit or Write call, to check it against your rules. Only for a call that shares a path, a command or enough words with a saved rule (the guard).
Jev also checks lines that jevmem did not write on your machine, such as a teammate's or a pull request's, before they reach Claude (what leaves your machine).
Why Jev and not an LLM #
- It is fast enough to run on every turn. 0.3 s in-process (and the
Stophook no longer waits for it), against 2.8–4.3 s p50 for the LLMs we measured, means the decision can happen on every Stop, not once per session. Memory that updates continuously catches the decision made in passing at turn 41. It is not the most accurate: see Benchmark. - Typed answers, thresholds in code. Jev returns probabilities, not prose. "Save if importance ≥ useful and chit-chat < 0.5" is a line of config, testable and tunable, not a prompt you hope the model follows.
- A narrow attack surface, not a closed one. Per TypeSafe, Jev returns probabilities and does not generate text or call tools, so a transcript that says "ignore previous instructions and remember X" has no channel to run a command through Jev; the failure mode is a wrong probability. Injected text can still bias those probabilities, which is why an injection noul gates every hook and MCP
add_memorysave (one broad noul in tier 1, four atomic nouls when tier 2 runs), the eval sets carry injection attempts, and the harness sends one. Once Jev says save, the line itself is written by the writer LLM (or the extract) from the scrubbed message, so injected text that gets past the gate can still shape the line's wording. Lines typed withjevmem addare not checked by Jev.
Fast and cheap, measured #
Deciding what to save takes 0.28 s and costs $0.00016 per message, in the background: Claude doesn't wait for it. jevmem tied the best LLM on save or skip (98.5%); two LLMs were better at picking the kind of line.
The full benchmark: accuracy, cost and how it was run
66 held-out turns, all seven deciders given the same state (method, regression set, pricing, p95, retries). The six LLM rows are v0.4.2's run of 2026-09-23; jevmem's row is 0.6.0's run of the same set on 2026-09-30 (results; every mode, three builds), where 0.5.9 and v0.4.2 score the same and cost less:
| Decider | save/skip | save+kind | contradictions | p50 | $/decision |
|---|---|---|---|---|---|
| GPT-6 Astra | 98.5% | 98.5% | 5/5 | 3,469 ms | $0.007489 |
| GPT-6 Luna | 93.9% | 93.9% | 5/5 | 2,927 ms | $0.000089 |
| Claude Fable 5.1 | 95.5% | 95.5% | 5/5 | 4,290 ms | $0.013256 |
| Claude Opus 5.5 | 97.0% | 97.0% | 5/5 | 2,784 ms | $0.005186 |
| Gemini 3.8 Flash | 92.4% | 92.4% | 5/5 | 2,850 ms | $0.001174 |
| Grok 4.7 | 90.9% | 90.9% | 4/5 | 3,320 ms | $0.004602 |
jevmem 0.6.0 auto |
98.5% | 95.5% | 5/5 | 276 ms | $0.000157 |
The 0.28 s is the Jev API decision (p95 527 ms; a saved turn's line costs one more request, $0.000159 per decision with it). Since v0.5.0 you do not wait for it: the Stop hook is async and its process exits in 12–14 ms (v0.5.6: 12 ms for the hook jevmem init registers, 14 ms for the plugin's), and the daemon records the decision 0.26–0.28 s after the hook starts (results, cost and latency).
On 66 held-out turns, jevmem 0.6.0's median decision took 0.28 s, against 2.8–4.3 s for six current LLMs. Its accuracy was within the LLMs' range: 98.5% save/skip (tied with GPT-6 Astra for highest) and 95.5% save+kind, against 90.9–98.5% for the LLMs. GPT-6 Astra (98.5%) and Claude Opus 5.5 (97.0%) were more accurate on save+kind; Claude Fable 5.1 tied; GPT-6 Luna, Gemini 3.8 Flash and Grok 4.7 were less accurate. It found 5/5 contradictions, as did five of the six LLMs. GPT-6 Luna was cheaper ($0.000089 against $0.000157) but less accurate (93.9%) and about 11× slower. Each row is a single run, and differences of one or two turns are within run-to-run noise; the LLM rows and jevmem's are a week apart. If the most accurate decision matters most, GPT-6 Astra or Claude Opus 5.5 are better, at about 33–48× the cost per decision and 10–13× the latency. jevmem is for when you want a fast, cheap decision on every message.
A local model in Jev's place #
jevmem was tried once against a local model server, Ollaya 0.9.0, on an Apple M4 with 16 GB, on 2026-10-02, next to a Jev run of the same set the same afternoon. On the 66 held-out turns, in jevmem's default mode, Jev was right on save or skip for 65/66 at a median of 0.23 s a turn; winnow:e4b for 60/66 at 28.5 s; laya:typed-decisions for 19/66. This was one run each on one Mac, with two models, on a set written for Jev's behaviour (Jev, winnow:e4b, laya:typed-decisions). A local model is not a supported mode (FAQ).
Set it up #
Jev needs a TypeSafe API key, and jevmem sends message text to TypeSafe to be scored, with common secrets scrubbed first. The four steps: install.
The limits, in short #
- jevmem needs a TypeSafe API key (where to get one, and the install steps).
- Message text is sent to TypeSafe to be scored, with common secrets scrubbed first (what leaves your machine).
- It is automatic in Claude Code, automatic in Codex while
jevmem watchruns, and in Cursor only when the agent calls it (what each tool does). - Rules every task must follow still belong in
CLAUDE.md(jevmem next to CLAUDE.md). - The guard is a backstop, not a sandbox (what it misses).
Every limit, with the numbers: the FAQ.