Does jevmem need a key, what does it cost, and what are its limits?
jevmem needs a TypeSafe API key; deciding what to save costs $0.00016 per message on that key; a local model is not a supported mode; your team gets the memory file through git; and the known limits are listed at the end of this page.
The cost is the average over 66 held-out turns with jevmem 0.6.0, one run on 2026-09-30, at TypeSafe's listed price for Jev (results). To set jevmem up: install.
Do I need a key? #
Yes: jevmem needs a TypeSafe API key, because Jev, the model that decides what to save, runs on TypeSafe's API. jevmem key saves it in ~/.jevmem/env, readable only by you. You need no OpenAI or Anthropic key: jevmem writes each line itself unless you set writer in jevmem.config.json.
What does it cost? #
jevmem itself is free and open source (MIT). The requests it makes to Jev are billed by TypeSafe to your key.
- Deciding what to save costs $0.00016 per message (66 held-out turns, jevmem 0.6.0, 2026-09-30, results).
- Bringing lines back costs more as the file grows: 300 prompts a day is about $0.03 with 18 live lines, $0.09 with 74 and $0.28 with 220 (the retrieval held-out set's files, run through the real hook on 2026-09-28, results).
Both are input tokens at $0.042 per million, the price in TypeSafe's launch post; check TypeSafe's own pricing for what your key is charged. Every request is logged in .jevmem/log.jsonl, and jevmem stats adds them up. The whole table: docs/cost.md.
Can jevmem run on a local model? #
Not as a supported mode. It was tried once, on 2026-10-02, with a local model server, Ollaya 0.9.0, on an Apple M4 with 16 GB, next to a Jev run of the same set the same afternoon. On the 66 held-out turns, in jevmem's default mode, Jev was right on save or skip for 65/66 turns at a median of 0.23 s a turn; winnow:e4b for 60/66 at 28.5 s; laya:typed-decisions for 19/66.
The results files: Jev, winnow:e4b, laya:typed-decisions. The limits of that test: one run each, on one Mac, with two models; the set was written for Jev's behaviour; and winnow:e4b was measured on these 66 turns only, not on recall or the guard.
Does my team get the memory through git? #
Yes: JEVMEM.md is a file in your repo, so your team gets the same file through git, and a change to it can be reviewed in a pull request like code. Lines a teammate or a pull request adds are checked before Claude sees them: on a 44-line test set (2026-09-25), that check blocked 20 of 22 planted lines, with 0 false blocks on 22 legitimate rules (results). The scores behind each line stay on the machine that saved it, in .jevmem/, which is not in git.
Which tools does it work with? #
| Saving | Bringing it back | |
|---|---|---|
| Claude Code | Automatic, every turn | Automatic, every prompt |
| Codex | Automatic while jevmem watch runs |
When the agent asks, over MCP |
| Cursor | When the agent calls it, over MCP | When the agent asks, over MCP |
| Claude Desktop | When you ask it to, over MCP | When you ask it to, over MCP |
The MCP server is on the MCP Registry as io.github.Avinash-jetwani/jevmem. Client setup: docs/mcp.md · docs/install.md.
What leaves my machine? #
Message text is sent to TypeSafe to be scored, with common secrets scrubbed first, and only from a project you have turned on. The whole list: What leaves my machine?
What are jevmem's limits? #
Known limits in 0.6.4 #
- A plain statement can fall under the content threshold. A rule said without must, never or prefer ("user-facing copy is British English") can be skipped: 2 of the 12 genuine rules on the genuine-rule held-out set, in 0.5.9 too (docs/benchmark.md).
- No way to mark your own rules as verified. A line you add with
jevmem addor by hand is unverified: the poisoning gate checks it before recall serves it, and the guard asks about it but never denies on it. - A line is at most two sentences, and Jev can pick the wrong one (docs/benchmark.md). Jev picks the sentence that states the memory and the one that gives its reason; a fact told in two sentences keeps one of them, and on the line-text held-out set one rule's line was its consequence without the rule (1 of 63), and one to-do kept "don't start on it today" as its second sentence. The line comes from the text decide chose: a bug whose cause only Claude's reply found gets your description of it. When the pick request fails, the line is one sentence chosen by words, as in 0.5.9.
- A second line one line crowds out. When the first line takes the whole "most relevant" choice, a second line is kept only at relevance 0.97 or more: 3 of 6 two-line prompts on retrieval held-out v2 lost their second line.
- More than 250 live lines. Only the 250 sharing the most words with the prompt are asked about; on a 500-line dev file, 8 of 9 missed lines were never sent. Sending all of them would cost about twice as much per prompt on such a file.
- Rules every task must follow. Recall judges each line against the prompt; a convention nothing in the prompt points at belongs in
CLAUDE.md. - The guard sees the words a rule and a call share (docs/guardrails.md). A rule that names neither the tool nor the host (
psql "$PROD_WAREHOUSE_DSN"under "the production warehouse is never queried from a shell") gets no candidate, and one shared word is not enough (git tag -d v0.5.10under "never delete a tag",npm version 0.6.0under "no version bump"). - Jev reads a script's text as run. A script that runs
jevmem guard test "git push origin directory"(a dry run of the guard) is asked about under "never push to that branch by hand" (Jev 0.57 to 0.71). - An ambiguous rule gets an ambiguous answer. A
Signed-off-byunder the owner's own name scored 0.07 against "commits are underonly, with no Co-Authored-By or other trailers"; write rules as you mean them.
Honest limits #
- Early: 0.6.4; every eval set was written by the author, and none is an independent benchmark.
- Not the most accurate: GPT-6 Astra and Claude Opus 5.5 scored higher on save+kind; jevmem's edge is speed and cost.
- Recall quality is not measured: that relevant lines are injected is tested; whether answers get better is not.
- Long-run drift is not measured: the harness covers five-turn sessions, not weeks of use.
- Automatic capture is Claude Code only (and Codex while
jevmem watchruns); Cursor and Claude Desktop save only when the agent callsadd_memory. - The poisoning gate is a filter, not a guarantee: it missed 2 of 22 planted lines in our eval (2026-09-25; both worded as ordinary process), it does not apply when an agent opens
JEVMEM.mdas a file, and on a fresh clone its first check costs one noul per line. ReviewJEVMEM.mddiffs like code (SECURITY.md). - Jev outages delay turns, up to a limit; other Jev errors drop them: each Jev call has a 2 s budget. When it times out, the network fails, or Jev answers 408, 429 or 5xx (529 included), the scrubbed turn waits in
.jevmem/queue.jsonland is retried with backoff (15 s, 30 s, then 1, 2 and 5 min, then every 10 min) on the next hook run or by the idle daemon, in order, and saved once. A turn still unsaved after 24 hours, or past 200 queued turns, is dropped. Any other error is not retried and drops the turn at once: a 400 from Jev, for example, or a 401 when the key is wrong, which drops every turn until the key is fixed. Each drop leaves a line in.jevmem/log.jsonl, and the retry-queue line ofjevmem statscounts them. - A plugin update can leave an open session without the hooks: when the plugin synced from claude.ai updates (Claude Code downloads updates in the background each time it starts), Claude Code moves the previous copy aside, and a session that was already open with it loses jevmem's hooks, the guard included, until you run
/reload-pluginsthere or start a new session; Claude Code shows a hook error and goes on without them (anthropics/claude-code#97847).jevmem doctorshows the version on disk.
The limits, in short #
- jevmem needs a TypeSafe API key (where to get one, and the install steps).
- Message text is sent to TypeSafe to be scored, with common secrets scrubbed first (what leaves your machine).
- It is automatic in Claude Code, automatic in Codex while
jevmem watchruns, and in Cursor only when the agent calls it (what each tool does). - Rules every task must follow still belong in
CLAUDE.md(jevmem next to CLAUDE.md). - The guard is a backstop, not a sandbox (what it misses).
Every limit, with the numbers: the FAQ.