jevmemjevmem

What is Jev used for in coding agents, and how does jevmem use it?

In jevmem, Jev (a model by TypeSafe AI) is what decides: it scores each message of a Claude Code chat so that jevmem can tell whether it holds a decision, a rule or a failed approach worth saving, it picks the saved lines that bear on your next prompt, and it judges whether a command or a file edit may break a saved rule.

Last updated: · jevmem 0.6.4 · Install · Source on GitHub

Deciding what to save takes 0.28 s and costs $0.00016 per message, and was right on save or skip for 98.5% of 66 held-out turns (jevmem 0.6.0, one run on 2026-09-30, results). This page covers only what jevmem does with Jev. What Jev itself is: TypeSafe's docs and their launch post.

How jevmem asks Jev #

How jevmem decides: scrub secrets, ask Jev typed questions, apply thresholds in code, write one line, supersede the old line.How jevmem decides: scrub secrets, ask Jev typed questions, apply thresholds in code, write one line, supersede the old line.

No prompt decides what to save: Jev answers small yes/no questions with probabilities, and plain rules in code act on them.

Each step in detail
  1. Scrub. Common secret shapes, email addresses and card-shaped numbers are removed from the turn before it leaves your machine.
  2. Ask Jev typed questions. Jev by TypeSafe AI answers a fixed set of small questions with probabilities: is there a decision, a rule, a bug? is it small talk or an injection attempt? which existing line does it change?
  3. Apply thresholds in code. Plain rules over those probabilities decide save or skip; they live in jevmem.config.json, not in a prompt.
  4. Write one line. On save, jevmem writes one line of at most 200 characters from the sentences of the turn that Jev picks (the one that states the memory and the one that gives its reason), or, if you set writer in jevmem.config.json, a small OpenAI or Anthropic model condenses the turn.
  5. Supersede the old line. If the turn replaces an existing memory, that line is tagged [superseded] … → id:new and stays in the file.

Tiers, questions, policy, contradictions, recall and audit: docs/how-it-works.md.

The three places jevmem calls Jev #

Jev also checks lines that jevmem did not write on your machine, such as a teammate's or a pull request's, before they reach Claude (what leaves your machine).

Why Jev and not an LLM #

Fast and cheap, measured #

Median time to decide one message on 66 held-out turns: jevmem 0.28 s, six current LLMs 2.78 to 4.29 s.Median time to decide one message on 66 held-out turns: jevmem 0.28 s, six current LLMs 2.78 to 4.29 s.

Deciding what to save takes 0.28 s and costs $0.00016 per message, in the background: Claude doesn't wait for it. jevmem tied the best LLM on save or skip (98.5%); two LLMs were better at picking the kind of line.

The full benchmark: accuracy, cost and how it was run

66 held-out turns, all seven deciders given the same state (method, regression set, pricing, p95, retries). The six LLM rows are v0.4.2's run of 2026-09-23; jevmem's row is 0.6.0's run of the same set on 2026-09-30 (results; every mode, three builds), where 0.5.9 and v0.4.2 score the same and cost less:

Decider save/skip save+kind contradictions p50 $/decision
GPT-6 Astra 98.5% 98.5% 5/5 3,469 ms $0.007489
GPT-6 Luna 93.9% 93.9% 5/5 2,927 ms $0.000089
Claude Fable 5.1 95.5% 95.5% 5/5 4,290 ms $0.013256
Claude Opus 5.5 97.0% 97.0% 5/5 2,784 ms $0.005186
Gemini 3.8 Flash 92.4% 92.4% 5/5 2,850 ms $0.001174
Grok 4.7 90.9% 90.9% 4/5 3,320 ms $0.004602
jevmem 0.6.0 auto 98.5% 95.5% 5/5 276 ms $0.000157

The 0.28 s is the Jev API decision (p95 527 ms; a saved turn's line costs one more request, $0.000159 per decision with it). Since v0.5.0 you do not wait for it: the Stop hook is async and its process exits in 12–14 ms (v0.5.6: 12 ms for the hook jevmem init registers, 14 ms for the plugin's), and the daemon records the decision 0.26–0.28 s after the hook starts (results, cost and latency).

On 66 held-out turns, jevmem 0.6.0's median decision took 0.28 s, against 2.8–4.3 s for six current LLMs. Its accuracy was within the LLMs' range: 98.5% save/skip (tied with GPT-6 Astra for highest) and 95.5% save+kind, against 90.9–98.5% for the LLMs. GPT-6 Astra (98.5%) and Claude Opus 5.5 (97.0%) were more accurate on save+kind; Claude Fable 5.1 tied; GPT-6 Luna, Gemini 3.8 Flash and Grok 4.7 were less accurate. It found 5/5 contradictions, as did five of the six LLMs. GPT-6 Luna was cheaper ($0.000089 against $0.000157) but less accurate (93.9%) and about 11× slower. Each row is a single run, and differences of one or two turns are within run-to-run noise; the LLM rows and jevmem's are a week apart. If the most accurate decision matters most, GPT-6 Astra or Claude Opus 5.5 are better, at about 33–48× the cost per decision and 10–13× the latency. jevmem is for when you want a fast, cheap decision on every message.

A local model in Jev's place #

jevmem was tried once against a local model server, Ollaya 0.9.0, on an Apple M4 with 16 GB, on 2026-10-02, next to a Jev run of the same set the same afternoon. On the 66 held-out turns, in jevmem's default mode, Jev was right on save or skip for 65/66 at a median of 0.23 s a turn; winnow:e4b for 60/66 at 28.5 s; laya:typed-decisions for 19/66. This was one run each on one Mac, with two models, on a set written for Jev's behaviour (Jev, winnow:e4b, laya:typed-decisions). A local model is not a supported mode (FAQ).

Set it up #

Jev needs a TypeSafe API key, and jevmem sends message text to TypeSafe to be scored, with common secrets scrubbed first. The four steps: install.

The limits, in short #

Every limit, with the numbers: the FAQ.