🦛 Hippo: memory for AI agents that learns what is wrong
Hippo learns what is wrong and ranks it down. Good memory is knowing what to forget: what turned out wrong, what got replaced, what nobody used.
Hippo keeps your coding agents' memories in a SQLite store on your machine, with markdown mirrors you can read and commit. Search is BM25 out of the box, with no model and no network call; embeddings are an optional install. hippo init installs hooks for Claude Code and OpenCode, adds 2 hooks to Codex's hooks.json when Codex is installed (Codex runs them once you trust them in /hooks), and adds instructions to an existing AGENTS.md for Codex, Cursor, OpenClaw and Pi. Any MCP client can connect. Mark a memory wrong and it ranks lower; run hippo supersede and the old fact leaves recall. Zero runtime deps.
Install it, then run hippo init inside one project. Init creates the project's .hippo/ store and adds a block to the CLAUDE.md or AGENTS.md already there. On your machine it adds hooks for the agents it finds, such as Claude Code's in ~/.claude/settings.json, and a daily 6:15am run. What hippo init changes lists all of it and the flags that skip each part.
npm install -g hippo-memory && hippo init
Setting up every git repo under a folder in one go is a second step. The Quick start says what it changes, then gives the command.
Package installation alone does not enable automatic preservation on every agent. Complete the documented setup and required host trust; capture and compaction coverage depend on the integration. See automatic-save coverage.
Having an AI agent install it? Point it at llms-install.md: it installs, wires hippo into the agents it finds, and verifies with hippo doctor.
Works with: Claude Code, Codex, Cursor, OpenClaw, OpenCode, Pi, any MCP client
Imports from: ChatGPT, Claude (CLAUDE.md), Cursor (.cursorrules), Slack, markdown
Storage: SQLite backbone with markdown mirrors. Git-trackable, human-readable.
Dependencies: Zero runtime deps. Node.js 22.16+. Optional embeddings: bring-your-own local Transformers.js (npm i @huggingface/transformers, or legacy @xenova/transformers) or an opt-in API embedder (OpenAI/Voyage/Cohere). Nothing is auto-installed.
Contents: Why · Receipts · Quick start · Agent setup · MCP server · How it works · Features · CLI · Comparison · Benchmarks · FAQ · Contributing
Why this exists
Most "AI memory" systems save everything and search later. That's storage with search on top. A note that turned out wrong ranks the same as one that held up, and an old fact sits beside the one that replaced it.
Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. Memories you keep using get stronger. Those two are the parts we measured helping retrieval, on a synthetic test (mechanism audit, round 2). When a fact changes, run hippo supersede and the old version leaves recall; we have not measured whether that helps. The design borrows from the hippocampus (decay, three layers, sleep consolidation), but that is inspiration. In our tests, decay tied with decay switched off and sleep lowered recall (Benchmarks).
It also fixes the portability problem. Your ChatGPT memories don't travel to Claude. Your .cursorrules don't travel to Codex. Hippo is one store behind every agent. CLAUDE.md, Cursor rules, ChatGPT exports, Slack history, all in one SQLite store, all queryable from any tool that speaks MCP or HTTP.
Receipts
Numbers, not adjectives. Every claim links to the benchmark or the test that proves it.
Every measurement we have ever published is indexed in docs/evals/,
pre-registrations kept next to their results, including the runs that failed and the one
claim we retracted.
- Sequential Learning Benchmark. benchmarks/sequential-learning/. 50 tasks, 10 buried traps. Measures whether agents learn from past mistakes, not just retrieve text. v0.11.0 informal magnitude RETRACTED v1.7.9; mechanism remains shipped. See CHANGELOG.md v1.7.9 entry.
- LongMemEval oracle split, all 500 questions in one pooled store, BM25 only, no embeddings, v0.11: R@5 = 74.0% (benchmarks/README.md, scripts in benchmarks/longmemeval/). A different setup from the per-haystack results under Benchmarks, so the two are not a before and after.
- On a private 300-query developer store, R@1 0.41 to 0.62 with
hippo recall ", the free local cross-encoder against Jev (full eval). The opt-in TypeSafe Jev reranker, off by default, about 0.0004 USD a recall. 2000-draw paired bootstrap; the margin held in 20 of 20 seeds and a permutation null reached it in 0 of 200 runs. Ranking only: three graded tests on one 150-question LongMemEval set did not show a better answer rate than the free local cross-encoder, and that negative result is in the same doc. What it buys today is a shorter context: on that set, 2 memories ranked by Jev answered as well as 5 ranked by the cross-encoder." --reranker jev - Staged Slack corpus, 10 incident scenarios: recall beat transcript replay in 10 of 10 (benchmarks/e1.3/). The answers sit mid-channel by design. In every scenario hippo's top 10 results held all the answer messages, and the channel's last 10 messages held none.
- Slack connector, 1000-event ingestion smoke: 0 outbound HTTP (benchmarks/e1.3/). Proven by a
globalThis.fetchspy that throws on call, not a hardcoded zero. Recall makes no network call by default. One default does:hippo sleepsends memory text to Anthropic for fact extraction whenANTHROPIC_API_KEYis set. Sleep runs at the end of every Claude Code and OpenCode session and in the daily job, so with the key set the call happens without you asking.{"extraction":{"enabled":false}}in.hippo/config.jsonturns that off. Opt-in features such as the Jev reranker above, the LLM reranker and the API embedders also call out. - 3,500+ tests on a real database. No module mocks and no mocked store; only paid network calls are stubbed. Project rule. The one mocks-vs-prod divergence that bit us early is now the constraint that kept the next ten releases honest.
- 3-cluster fixture where BM25 alone cannot discriminate: dlPFC goal-conditioned cluster discrimination passes 3 of 3 queries. Full goal stack with policy weighting and lifespan-windowed outcome propagation, one query per goal; deterministic test in
benchmarks/micro/results/b3-depth.json.
What it does for your agent
- Keeps errors longer. Tag a failure with
--tag errorand it gets twice the half-life of an ordinary memory, so the lesson is still in the store the next time a recall matches it. In Claude Code, a hook stores failed tool calls as error memories for you. - Survives tool switches. Use Claude Code on Monday, Cursor on Tuesday, Codex on Wednesday. They all read the same
.hippo/store, so the memories come with you. - Ingests systems of record. Slack and GitHub today (
POST /v1/connectors/slack/events,POST /v1/connectors/github/events). Jira and Notion next. Webhooks land askind='raw'memories with full provenance and GDPR-correct deletion. - Knows where every memory came from. Every row carries
kind,scope,owner, andartifact_ref. Right-to-be-forgotten is a single API call, not an audit nightmare. - Plays nice with multi-tenant. API keys, scrypt-hashed. Audit log on every mutation. Tenant A literally cannot see tenant B's memories. Proven by negative test.
Quick start
Start in one project. What hippo init changes lists everything the second command writes.
npm install -g hippo-memory
In a project: create its store and wire in the agents it uses
hippo init
Optional: many repos at once. Read what it changes first. hippo init --scan looks for git repos in the folder and up to three levels below it, skipping dot-folders and node_modules. Each repo gets a .hippo/ store, seeded with lessons from the last 365 days of its commits and with its agent memories, and is added to the daily run's list. For the agents it finds, it installs the same user-level hooks as hippo init: 7 Claude Code hook entries in ~/.claude/settings.json and the OpenCode plugin. It also sets up the daily 6:15am run, a crontab line on Linux and macOS or a scheduled task on Windows. It adds no block to any repo's CLAUDE.md or AGENTS.md. --no-hooks, --no-schedule and --no-learn leave out the hooks, the daily run and the history import.
hippo init --scan ~
After setup, hippo sleep runs when a Claude Code or OpenCode session ends, and in the daily 6:15am job for every project. Codex runs it at session end only if you installed its wrapper. It does five things:
- Learns from today's git commits
- Imports what your coding agents remember, about this project and about you (Agent memories)
- Consolidates memories (decay, merge, prune)
- Deduplicates identical memories, keeping the stronger copy
- Shares high-value lessons to a global store so they surface in every project
# Manual usage
hippo remember "FRED cache silently dropped the tips_10y series" --tag error
hippo recall "data pipeline issues" --budget 2000
Full release history: CHANGELOG.md · GitHub Releases
What hippo init changes
Run hippo init inside one project. This is everything it writes, in the project and on your machine:
- The project's store. A
.hippo/folder: SQLite plus markdown mirrors. On the first run in a git repo it learns lessons from the last 30 days of commits. - Instruction files. A block between
andin the project'sCLAUDE.mdorAGENTS.md, only if that file already exists. Codex, Cursor, OpenClaw, OpenCode and Pi readAGENTS.md. - Claude Code, when the project has
CLAUDE.mdor.claude/settings.json: 7 hook entries in~/.claude/settings.json, one each on SessionEnd, UserPromptSubmit, PreCompact, PostCompact and PostToolUseFailure and two on SessionStart. Framework Integrations says what each one runs. - OpenCode, when the project has
.opencode/oropencode.json: a plugin at~/.config/opencode/plugins/hippo.ts. - Codex, when the project has
AGENTS.mdor.codexand Codex is installed ($CODEX_HOME, else~/.codex, exists): 2 hook entries in Codex'shooks.json, one on UserPromptSubmit that sends your pinned memories plus up to 5 that match the prompt with every prompt and one on SessionStart after a compaction. Codex runs them only after you trust them once in/hooks. Init also printshippo hook install codex, the opt-in that wraps the Codex launcher to capture sessions;hippo hook uninstall codexremoves hippo's hooks and the wrapper. - A daily run at 6:15am, one per machine: a crontab line on Linux and macOS, a scheduled task named
hippo-daily-runneron Windows. It runshippo learn --git --days 1and thenhippo sleepin every project listed in~/.hippo/workspaces.json, and init adds this project to that list. - Agent memories. On every run, the notes your coding agents keep about this project go into its store, and the ones about you go into the global store. Agent memories lists what is read.
cd my-project
hippo init
Initialized Hippo at /my-project/.hippo
Directories: buffer/ episodic/ semantic/ conflicts/
Files: hippo.db stats.json
Auto-installed claude-code hook in CLAUDE.md
Auto-installed hippo session-end SessionEnd hook in claude-code settings
(one line per hook entry)
Scheduled machine-level daily runner (6:15am) via crontab
To leave parts out: --no-hooks skips the instruction files and hooks, --no-schedule the daily run, and --no-learn the git history and agent memory import. HIPPO_SKIP_AUTO_INTEGRATIONS=1 skips the same files and hooks that --no-hooks does.
Agent memories
Most coding agents now keep their own notes between sessions. Hippo reads them, whatever the tool, so what one agent learned reaches the others. It reads files only and never writes to another tool's folders.
When: hippo init (every run), init --scan, init --global, hippo setup, every hippo sleep and the daily run. At session end a folder with its own store gets it through sleep; a folder without one sends its project's notes to the global store, marked with the project's name. After a Claude Code compaction, the session's own notes folder is read as well.
What is read, per tool (each tool's own environment variables and settings decide where its home is):
- Claude Code: the project's auto memory notes under
~/.claude/projects/(front matter required,/memory/ MEMORY.mdskipped), and theautoMemoryDirectoryfolder from your user settings. - Codex: the User Profile, preferences and tips in
~/.codex/memories/memory_summary.md. - Gemini CLI: the "Gemini Added Memories" section of
~/.gemini/GEMINI.md, and the project's auto memory folder when that feature is on. - GitHub Copilot Chat in VS Code: the memory tool's user memories and the repository memories of this project's workspace.
- OpenClaw: the workspace's
MEMORY.md. - Qwen Code: the project's auto memory folder and your user memories.
hippo dormant lists it and can restore it). A note shorter than 10 characters, one that looks like it holds a secret (an API key, a password, an auth header or a token), and one whose text you rejected with hippo reject are skipped. Email addresses are stored masked, and notes are cut at 1,500 characters.
Not read: Windsurf (the file format is not documented, and Cascade reached end of life on 1 July 2026); Cursor, Copilot CLI and GitHub's Copilot Memory (the memories live on the vendor's servers); Kiro (the local store is not documented); Cline and Roo memory banks (files in the repository, which hippo import --markdown covers); Amp, Aider, Continue, OpenCode and pi (no memory feature found).
hippo import --agents runs the import by hand; in a folder without a store it does what session end does there. With --dry-run it shows each tool's home, the folders found and what would change, and writes nothing. To choose tools, set "agentMemories": { "tools": ["claude-code", "codex"] } in .hippo/config.json ([] turns the import off), or HIPPO_AGENT_MEMORY_TOOLS=claude-code,codex in the environment (none turns it off), which wins over config.
Cross-Tool Import
Your memories shouldn't be locked inside one tool. Hippo pulls them in from anywhere.
# ChatGPT memory export
hippo import --chatgpt memories.json
Claude's CLAUDE.md (skips existing hippo hook blocks)
hippo import --claude CLAUDE.md
Cursor rules
hippo import --cursor .cursorrules
Any markdown file (headings become tags)
hippo import --markdown MEMORY.md
Any text file
hippo import --file notes.txt
All import commands support --dry-run (preview without writing), --global (write to ~/.hippo/), and --tag (add extra tags). Duplicates are detected and skipped automatically.
Conversation Capture
Extract memories from raw conversation text. No LLM needed: pattern-based heuristics find decisions, rules, errors, and preferences.
# Pipe a conversation in
cat session.log | hippo capture --stdin
Or point at a file
hippo capture --file conversation.md
Preview first
hippo capture --file conversation.md --dry-run
Slack ingestion (E1.3)
Hippo accepts Slack Events API webhooks at POST /v1/connectors/slack/events. Configure SLACK_SIGNING_SECRET (validated on every request) and point Slack at https://. Messages land as kind='raw' memories with slack://team/channel/ts provenance and a slack:public:Cxxx or slack:private:Cxxx scope. Source deletions are honored (GDPR).
Backfill an existing channel: SLACK_BOT_TOKEN=xoxb-... hippo slack backfill --channel C0000. Inspect malformed events: hippo slack dlq list.
Multi-workspace deployments populate slack_workspaces (team_id, tenant_id) to route events per tenant; single-workspace falls back to HIPPO_TENANT.
Active task snapshots
Long-running work needs short-term continuity, not just long-term memory. Hippo can persist the current in-flight task so a later continue has something concrete to recover.
hippo snapshot save \
--task "Ship SQLite backbone" \
--summary "Tests/build/smoke are green, next slice is active-session recovery" \
--next-step "Implement active snapshot retrieval in context output"
hippo snapshot show
hippo context --auto --budget 1500
hippo snapshot clear
hippo context --auto includes the active task snapshot before long-term memories, so agents get both the immediate thread and the deeper lessons.
Session event trails
Manual snapshots are useful, but real work also needs a breadcrumb trail. Hippo can now store short session events and link them to the active snapshot so context output shows the latest steps, not just the last summary.
hippo session log \
--id sess_20260326 \
--task "Ship continuity" \
--type progress \
--content "Schema migration is done, next step is CLI wiring"
hippo snapshot save \
--task "Ship continuity" \
--summary "Structured session events are flowing" \
--next-step "Surface them in framework hooks" \
--session sess_20260326
hippo session show --id sess_20260326
hippo context --auto --budget 1500
Hippo mirrors the latest trail to .hippo/buffer/recent-session.md so you can inspect the short-term thread without opening SQLite.
Session handoffs
When you're done for the day (or switching to another agent), create a handoff so the next session knows exactly where to pick up:
hippo handoff create \
--summary "Finished schema migration, tests green" \
--next "Wire handoff injection into context output" \
--session sess_20260403 \
--artifact src/db.ts
hippo handoff latest # show the most recent handoff
hippo handoff show 3 # show a specific handoff by ID
hippo session resume # re-inject latest handoff as context
Working memory
Working memory is a bounded scratchpad for current-state notes. It's separate from long-term memory. Entries stay until you run hippo wm flush.
hippo wm push --scope repo \
--content "Investigating flaky test in store.test.ts, line 42" \
--importance 0.9
hippo wm read --scope repo # show current working notes
hippo wm clear --scope repo # wipe the scratchpad
hippo wm flush --scope repo # flush on session end
The buffer holds a maximum of 20 entries per scope. When full, the lowest-importance entry is evicted.
Explainable recall
See why a memory was returned:
hippo recall "data pipeline" --why --limit 5
--- mem_a1b2c3 [episodic] [observed] [local] score=0.847
BM25: matched [data, pipeline]; cosine: 0.82
...memory content...
How It Works
Input enters the buffer. Important things get encoded into episodic memory. During "sleep," related episodes are merged into one semantic memory, by word overlap. Weak memories decay and disappear.
The store is SQLite (.hippo/hippo.db). The markdown files are mirrors written after each change. index.json is no longer refreshed by writes, deletes or recalls: it is written only when you call rebuildIndex() from the package, so a copy an older version left on disk goes stale. Read the store through the CLI, the MCP server or the HTTP API.
flowchart TD
I[New information] --> B[Buffer<br/>session-only, no decay]
B -->|encode: tags, strength, half-life| E[Episodic Store<br/>timestamped, decay by default<br/>retrieval strengthens, errors stick]
E -->|hippo sleep<br/>replay + merge| S[Semantic Store<br/>merged memories, stable<br/>schema-aware]
E -.->|decay| X[forgotten]
S -.->|recall| E
classDef bio fill:#fff4dc,stroke:#a8742d,color:#2b1b00
classDef forgotten fill:#f5f5f5,stroke:#999,color:#666,stroke-dasharray:5 5
class B,E,S bio
class X forgotten
Key Features
A memory's life across a typical session, before walking each feature in turn:
sequenceDiagram
autonumber
actor Agent
participant B as Buffer
participant E as Episodic
participant S as Semantic
Agent->>B: hippo remember "cache dropped tips_10y" --error
B->>E: encode (half_life=730d, valence=neg)
Note over E: strength=1.0
Agent->>E: hippo recall "data pipeline"
E-->>Agent: returns memory (rank 1)
Note over E: half_life 730d → 732d, retrieval_count++
Agent->>E: hippo outcome --good
Note over E: reward_factor 1.0 → 1.25
Agent->>S: hippo sleep
S->>E: merge 3 related episodic → 1 semantic
Note over E,S: original episodic decays, pattern survives
Decay by default
Every memory has a half-life: 365 days by default. Until 1.46.0 the default was 7 days. A pre-registered evaluation found 7 days lost the current version of a fact far more often: it was in the top five 29% of the time at 7 days and 75% at 365 (result). 730 days and decay off both tied with 365. So 365 was not tuned: it is the tested value that tied with the others, and on that test decay made no measurable difference to recall. hippo sleep moves memories still on the old 7-day base to the new one, once, and records each move in the audit log. Set defaultHalfLifeDays in .hippo/config.json to choose your own.
hippo remember "always check cache contents after refresh"
stored with half_life: 365d, strength: 1.0
two years later with no retrieval:
hippo inspect mem_a1b2c3
strength: 0.25 (decayed by 2 half-lives)
Retrieval strengthens
Use it or lose it. Each recall boosts the half-life by 2 days.
hippo recall "cache issues"
finds mem_a1b2c3, retrieval_count: 1 -> 2
half_life extended: 365d -> 367d
strength recalculated from retrieval timestamp
hippo recall "cache issues" # again next week
retrieval_count: 2 -> 3
half_life: 367d -> 369d
this memory is learning to survive
Active invalidation
When you migrate from one tool to another, old memories about the replaced tool should die immediately. Hippo detects migration and breaking-change commits during hippo learn --git and actively weakens matching memories.
hippo learn --git
feat: migrate from webpack to vite
Invalidated 3 memories referencing "webpack"
Learned: migrate from webpack to vite
You can also invalidate manually:
hippo invalidate "REST API" --reason "migrated to GraphQL"
Invalidated 5 memories referencing "REST API".
Architectural decisions
One-off decisions don't repeat, so they can't earn their keep through retrieval alone. hippo decide stores them with verified confidence and the store's default half-life, the same as any other memory, and sleep never retires the memory behind a decision. On a store made before 1.52.7, decisions get 90 days until the store's first hippo sleep on 1.52.7 or later moves them to the default.
hippo decide "Use PostgreSQL for all new services" --context "JSONB support"
Decision recorded: #1
memory: mem_a1b2c3
Later, when the decision changes:
hippo decide "Use CockroachDB for global services" \
--context "Need multi-region" \
--supersedes mem_a1b2c3
Decision recorded: #2
memory: mem_d4e5f6
supersedes memory: mem_a1b2c3 (decision #1 superseded)
Error memories stick
Tag a memory as an error and it gets 2x the half-life automatically.
hippo remember "deployment failed: forgot to run migrations" --error
half_life: 730d instead of 365d
emotional_valence: negative
strength formula applies 2.0x multiplier (HIPPO_LOSS_AVERSION_RATIO=0.75 to keep v1.13.4 1.5x)
production incidents don't fade quietly
Confidence tiers
Every memory carries a confidence level: verified, observed, inferred, or stale. This tells agents how much to trust what they're reading.
hippo remember "API rate limit is 100/min" --verified
hippo remember "deploy usually takes ~3 min" --observed
hippo remember "the flaky test might be a race condition" --inferred
When context is generated, confidence is shown inline:
[verified] API rate limit is 100/min per the docs
[observed] Deploy usually takes ~3 min
[inferred] The flaky test might be a race condition
Agents can see at a glance what's established fact vs. a pattern worth questioning.
A memory not recalled for 30 days is shown as aged when it is read, and recalling it clears that. Pinned and verified memories are exempt. Nothing in the store changes.
Conflict tracking
Hippo detects obvious contradictions between overlapping memories and keeps them visible instead of silently letting both masquerade as truth. Shared tags alone do not count; the statements themselves need to overlap in content.
hippo sleep # refreshes open conflicts
hippo conflicts # inspect them
Open conflicts are stored in SQLite, mirrored under .hippo/conflicts/, and linked back into each memory's conflicts_with field.
Observation framing
Memories aren't presented as bare assertions. By default, Hippo frames them as observations with dates, so agents treat them as context rather than commands.
hippo context --framing observe # default
Output: "Previously observed (2026-03-10): deploy takes ~3 min"
hippo context --framing suggest
Output: "Consider: deploy takes ~3 min"
hippo context --framing assert
Output: "Deploy takes ~3 min"
Three modes: observe (default), suggest, assert. Choose based on how directive you want the memory to be.
Sleep consolidation
Run hippo sleep and related episodes merge into one memory.
hippo sleep
Running consolidation...
#
Results:
Active memories: 23
Removed (decayed): 4
Merged episodic: 6
New semantic: 2
Two or more related episodes get merged into a single semantic memory. The originals decay. The pattern survives.
Sleep keeps the store tidy. It has not been shown to improve recall. In round 2 of the mechanism audit, a slept LongMemEval store scored 3.6 points lower at hit@5 than the same store never slept, and no scorer showed sleep helping (PR #232).
Experimental: learned memory-value rescue (opt-in, default off). With
{"memoryValue":{"enabled":true}} in .hippo/config.json, sleep consults a learned
linear memory-value scorer before deleting a decayed memory: a memory that scores in the
top 30% of its tenant by learned value is kept ("rescued") even though its strength fell
below the decay threshold. The scorer can only rescue, never delete: with the flag on,
sleep deletes a strict subset of what it would delete with the flag off. Every rescue is
recorded in the audit log (hippo audit list --op mv_rescue). The weights were learned
on the LongMemEval retention benchmark (held-out retention 0.4897 vs 0.4203 for the best
hand-set baseline); caveat: their usage-feature signs reflect that benchmark's simulated
usage, NOT real usage value, so treat the flag as an experiment, not a recommendation.
Tenants with fewer than 10 non-pinned memories never rescue (rank statistics are noise at
tiny scale).
Faded memories go dormant, not gone (on by default). Sleep moves a memory that faded below the decay threshold into a dormant store instead of deleting it. A dormant memory leaves recall and context exactly like a deleted one and sits out every later sleep, so your agent's context stays as lean as before, but nothing is lost:
hippo dormant # list, newest first (--json, --limit <n>)
hippo dormant "staging hostname" # search: every term must match
hippo dormant restore mem_a1b2c3 # back to active memory, as if just recalled
hippo dormant forget mem_a1b2c3 # delete for good
A restored memory comes back with a fresh recall clock, so it gets a full half-life before
it can fade again, and every restore is logged (hippo audit list --op dormant_restore) as
a "forgot it, then needed it" signal. Two guardrails: a faded memory that the secret
detector flags is deleted, never kept dormant, and a dormant memory nobody restores within
dormant.retentionDays (default 180, 0 keeps them forever) is deleted for good. Rejecting
a value (hippo reject) removes its dormant copies too. To delete faded memories straight
away as before, set {"dormant":{"enabled":false}} in .hippo/config.json. Sleep never
removes pinned memories, raw receipts (Slack, GitHub, vault imports) or the memories a
Claude Code compaction saved either way, and duplicate removal and junk cleanup still
delete other memories. hippo forget still deletes a compaction memory.
See what memory costs in tokens. Every block of memory text hippo hands an agent (the
per-prompt hook, the block hippo compact-resume restores after compaction, hippo context,
hippo recall, the MCP tools, the HTTP API) is recorded in a token ledger: counts, surface
and session, never the text. A block stays in the conversation, so each later model call
reads it again until the host compacts. When a Claude Code session ends, hippo counts those
calls in the session's transcript and records the re-read tokens for the per-prompt hook's
blocks and the compact-resume block, dated by the day of the calls. The other surfaces show sent tokens only: their rows
cannot tell a sub-agent's call from its parent's. hippo tokens shows sent and re-read
totals for the last 30 days (--days, --json). A session that is still open, or that
crashed, shows what was sent only. Re-reads usually bill at the provider's cached-input rate,
a fraction of the full input price. Counts are estimates (characters / 4), the same estimate
every budget uses. Rows older than 90 days are pruned.
See why a memory did or did not reach the agent. Turn on the delivery ledger with
{"deliveryLedger":{"enabled":true}} in .hippo/config.json (off by default). The flag is
read from the store the token ledger writes to: the project's local store when it has one,
else the global store. Each per-prompt hook call then records one event (session, turn
number, whether the block was sent, reused, empty or disabled, counts and token totals) and
one row per candidate memory: emitted, reused or rejected, with the stage and the
reason it was dropped. With prompt recall on, recent memories dropped by the quality filter
are not recorded yet. It holds ids, hashes, counts and reasons only, never prompt or memory
text; the prompt hash is unsalted, so a very short prompt can be guessed. A ledger failure prints one stderr line and never changes what the hook prints. Rows
older than 90 days are pruned; at a heavy 300 prompts a day that is about 190 MB per store.
Outcome feedback
Did the recalled memories actually help? Tell Hippo. It tightens the feedback loop.
hippo recall "why is the gold model broken"
... you read the memories and fix the bug ...
hippo outcome --good
Applied positive outcome to 3 memories
reward factor increases, decay slows
hippo outcome --bad
Applied negative outcome to 3 memories
reward factor decreases, decay accelerates
Outcomes are cumulative. A memory with 5 positive outcomes and 0 negative has a reward factor of ~1.42, making its effective half-life 42% longer. A memory with 0 positive and 3 negative has a factor of 0.625, so it decays 1.6 times as fast; each bad mark past the good ones also halves its strength, up to three times, and recall stops strengthening it. Mixed outcomes converge toward neutral (1.0).
This is the mechanism with the clearest measured win. On the synthetic E1 test, plain BM25 plus the outcome nudge cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0% (mechanism audit, round 2). Every mark in E1 is correct; real marks are noisier, since --bad marks the whole recall batch.
Token budgets
Recall only what fits. No context stuffing.
# fits within Claude's 2K token window for task context
hippo recall "deployment checklist" --budget 2000
need more for a big task
hippo recall "full project history" --budget 8000
machine-readable for programmatic use
hippo recall "api errors" --budget 1000 --json
Results are ranked by relevance strength recency. The highest-signal memories fill the budget first.
The budget counts the whole block as printed: the heading, each memory's label, date and tags,
and any snapshot or hint lines, so the token figure in the heading is the size of what the
model reads. Recall always keeps its first --min-results memories (default 1), even one
larger than the budget; hippo context skips a memory that does not fit and keeps filling.
--json returns the memories the text form would print.
Auto-learn from git
Hippo can scan your commit history and extract lessons from fix/revert/bug commits automatically.
# Learn from the last 7 days of commits
hippo learn --git
Learn from the last 30 days
hippo learn --git --days 30
Scan multiple repos in one pass
hippo learn --git --repos "~/project-a,~/project-b,~/project-c"
The --repos flag accepts comma-separated paths. Hippo scans each repo's git log, extracts fix/revert/bug lessons, deduplicates against existing memories, and stores new ones. Pair with hippo sleep afterwards to consolidate.
Ideal for a weekly cron:
hippo learn --git --repos "~/repo1,~/repo2" --days 7
hippo sleep
Watch mode
Wrap any command with hippo watch to auto-learn from failures:
hippo watch "npm run build"
if it fails, Hippo captures the error automatically
next time an agent asks about build issues, the memory is there
CLI Reference
| Command | What it does |
|---------|-------------|
| hippo init | Create .hippo/, install agent hooks and the daily run (what it changes) |
| hippo init --global | Create global store at ~/.hippo/ |
| hippo init --no-hooks | Create .hippo/ without auto-installing hooks |
| hippo remember " | Store a memory |
| hippo remember " | Store with tag (repeatable) |
| hippo remember " | Store as error (2x half-life) |
| hippo remember " | Store with no decay |
| hippo remember " | Set confidence: verified (default) |
| hippo remember " | Set confidence: observed |
| hippo remember " | Set confidence: inferred |
| hippo remember " | Store in global ~/.hippo/ store |
| hippo recall " | Retrieve relevant memories (local + global) |
| hippo recall " | Recall within token limit (default: 4000) |
| hippo recall " | Cap result count |
| hippo recall " | Show match reasons and source buckets |
| hippo recall " | Also surface memories N hops away in the entity/relation graph (0..3, default off) |
| hippo recall " | Output as JSON |
| hippo context --auto | Smart context injection (auto-detects task from git) |
| hippo context " | Context injection with explicit query (default: 1500) |
| hippo context --limit | Cap memory count in context |
| hippo context --budget 0 | Skip entirely (zero token cost) |
| hippo context --framing | Framing: observe (default), suggest, assert |
| hippo context --format | Output format: markdown (default) or json |
| hippo import --chatgpt | Import from ChatGPT memory export (JSON or txt) |
| hippo import --claude | Import from CLAUDE.md or Claude memory.json |
| hippo import --cursor | Import from .cursorrules or .cursor/rules |
| hippo import --markdown | Import from structured markdown (headings -> tags) |
| hippo import --file | Import from any text file |
| hippo import --dry-run | Preview import without writing |
| hippo import --global | Write imported memories to ~/.hippo/ |
| hippo capture --stdin | Extract memories from piped conversation text |
| hippo capture --file | Extract memories from a file |
| hippo capture --dry-run | Preview extraction without writing |
| hippo sleep | Run consolidation (decay + merge) |
| hippo sleep --dry-run | Preview consolidation without writing |
| hippo status | Memory health: counts, strengths, last sleep |
| hippo outcome --good | Strengthen last recalled memories |
| hippo outcome --bad | Weaken last recalled memories |
| hippo outcome --id | Target a specific memory |
| hippo inspect | Full detail on one memory |
| hippo forget | Force remove a memory |
| hippo dormant [ | List faded memories sleep kept instead of deleting |
| hippo dormant restore | Bring a dormant memory back to active memory |
| hippo dormant forget | Delete a dormant memory permanently |
| hippo doctor [--json] | Check the install: Node, store, schema, sleep, agent hooks; each problem names its fix. It never changes hippo.db, though SQLite may leave empty hippo.db-wal and hippo.db-shm files beside it |
| hippo support-bundle [--out | Write a redacted JSON file for a support ticket: versions, doctor checks, config, store counts and log names, never memory text; --include-logs adds each log's last 200 lines, which can quote it |
| hippo tokens [--days n] | Estimated tokens of memory text handed to agents, per surface, what later model calls re-read of the hook and compact-resume blocks, and what skipping unchanged hook blocks saved |
| hippo failures [--days n] | Failed tool calls the capture-error hook saw, by outcome, and how many errors first happened in another session |
| hippo embed | Embed all memories for semantic search |
| hippo embed --status | Show embedding coverage |
| hippo watch " | Run command, auto-learn from failures |
| hippo learn --git | Scan recent git commits for lessons |
| hippo learn --git --days | Scan N days back (default: 7) |
| hippo learn --git --repos | Scan multiple repos (comma-separated) |
| hippo daily-runner | Sweep registered workspaces and run daily learn+sleep |
| hippo conflicts | List detected open memory conflicts |
| hippo conflicts --json | Output conflicts as JSON |
| hippo resolve | Show both conflicting memories for comparison |
| hippo resolve | Resolve: keep winner, weaken loser |
| hippo resolve | Resolve: keep winner, delete loser |
| hippo promote | Copy a local memory to the global store |
| hippo share | Share with attribution + transfer scoring |
| hippo share | Share even if transfer score is low |
| hippo share --auto | Auto-share all high-scoring memories |
| hippo share --auto --dry-run | Preview what would be shared |
| hippo peers | List projects contributing to global store |
| hippo sync | Pull global memories into local project |
| hippo invalidate " | Actively weaken memories matching an old pattern |
| hippo invalidate " | Include what replaced it |
| hippo decide " | Record architectural decision |
| hippo decide " | Include reasoning |
| hippo decide " | Supersede a previous decision |
| hippo hook list | Show available framework hooks |
| hippo hook install | Install hook (claude-code also adds its 7 settings.json hook entries: session start and end, each prompt, compaction, failed tool calls) |
| hippo hook uninstall | Remove hook |
| hippo handoff create --summary "..." | Create a session handoff |
| hippo handoff latest | Show the most recent handoff |
| hippo handoff show | Show a specific handoff by ID |
| hippo session latest | Show latest task snapshot + events |
| hippo session resume | Re-inject latest handoff as context |
| hippo current show | Compact current state (task + session events) |
| hippo card create --title "..." | Create a work-queue card (--repo, --contract, --budget, repeatable --depends-on ) |
| hippo card show | Show a card, its deps, runs, comments and latest handoff |
| hippo card list [--status | List cards, newest-updated first |
| hippo card claim | Claim a ready or blocked card; prints its run id and the time its 4-hour lease expires |
| hippo card heartbeat | Extend a claimed card's lease |
| hippo card block | Block a running card; the reason is recorded as a comment |
| hippo card review | Move a running card to review |
| hippo card complete | Complete a card in review: success marks it done and promotes children whose parents are all done; failure or partial shelves it |
| hippo card reclaim | Sweep every card whose lease has passed and return it to ready |
| hippo card comment | Add a comment to a card |
| hippo wm push --scope | Push to working memory |
| --content "..."hippo wm read --scope | Read working memory entries |
| hippo wm clear --scope | Clear working memory |
| hippo wm flush --scope | Flush working memory (session end) |
| hippo dashboard | Open web dashboard at localhost:3333 (memory map and card board) |
| hippo dashboard --port | Use custom port |
| hippo mcp | Start MCP server (stdio transport) |
On heartbeat, block, review and complete, a given --run is checked against the card's live run and the command is refused, unchanged, if the two do not match.
Framework Integrations
Auto-install (recommended)
hippo init detects your agent framework and patches the right config file automatically:
| Framework | Detected by | Patches |
|-----------|------------|---------|
| Claude Code | CLAUDE.md or .claude/settings.json | CLAUDE.md + 7 hook entries in ~/.claude/settings.json (listed below) |
| Codex | AGENTS.md or .codex | AGENTS.md + UserPromptSubmit/SessionStart(compact) hooks in Codex's hooks.json when Codex is installed (trust them once in /hooks); session capture is opt-in with hippo hook install codex, which wraps the Codex launcher |
| Cursor | AGENTS.md | AGENTS.md, which Cursor reads from the project root |
| OpenClaw | .openclaw or AGENTS.md | AGENTS.md; the native plugin is a separate install: openclaw plugins install hippo-memory |
| OpenCode | .opencode/ or opencode.json | AGENTS.md + TS plugin at ~/.config/opencode/plugins/hippo.ts (subscribes to session.idle + session.created) |
| Pi | .pi or .pi/agent | AGENTS.md; copy the Pi extension for session hooks |
Init patches an instruction file only if it already exists. It also sets up a daily run and imports your coding agents' own memories; What hippo init changes lists everything.
Manual install
If you prefer explicit control:
hippo hook install claude-code # patches CLAUDE.md + adds the 7 settings.json hook entries listed below
hippo hook install codex # patches AGENTS.md + adds hooks to Codex's hooks.json + wraps the detected Codex launcher
hippo hook install cursor # patches AGENTS.md
hippo hook install openclaw # patches AGENTS.md
hippo hook install opencode # patches AGENTS.md + installs the opencode TS plugin
This adds a ... block that tells the agent to:
- Run
hippo context --auto --budget 1500at session start - Run
hippo remember "the moment it finds out why something failed, never as a closing step" --error - Everywhere but Claude Code, whose own auto memory does this job: run a plain
hippo rememberthe moment it learns something that should outlive the session, leaving out secrets and personal details - Capture a short summary with
hippo capture --stdinwhen the session ends, but only where no hook captures the session: Cursor, OpenClaw, OpenCode, Pi, and Codex without its wrapper
hippo init swaps a block an older hippo wrote for the current one, as long as nobody edited it. It leaves an edited block alone and says so, and never touches text outside the markers.
For Claude Code, it also adds 7 hook entries to ~/.claude/settings.json:
- a
SessionEndhook that runshippo sleepand thenhippo capturewhen the session exits. Capture matches the last 20 user and 10 assistant messages of the transcript against word patterns for decisions, rules, errors and preferences. It uses no model and does not read earlier turns, so record lessons withhippo rememberas you go. - a
SessionStarthook that prints the previous session's consolidation output - a
UserPromptSubmithook that runshippo context --pinned-only --include-recent 5 --format additional-contextevery turn. It re-injects pinned memories (hippo remember) plus up to 5 memories that share words with your prompt. When nothing matches, it adds only the pinned ones. Since 1.55.0 this replaces the 5 newest memories, which cut the median block from 847 to 533 tokens in our eval.--pin {"pinnedInject":{"promptRecall":false}}brings back the 5 newest, so fresh same-session lessons appear on the next prompt whatever you ask. The block is rendered without live strength percentages, so it stays byte-identical while its memories do not change, and it is sent only when it changed since the session's last prompt: an unchanged block is skipped, resent every 10 skips (pinnedInject.refreshTurns,0never resends) and resent after compaction. The prompt-matched memories go in a separate block that is never skipped, so they are sent on every prompt they match.{"pinnedInject":{"skipUnchanged":false}}sends the pinned block every turn as before. Opt out entirely with{"pinnedInject":{"enabled":false}}in.hippo/config.json. - a
PreCompacthook that runshippo pre-compactbefore the transcript gets summarized. It records the compaction in the store, saves a working-state snapshot (task/summary/next step) so mid-session compaction can't drop it, and asks the summariser to end its summary with a "Memories for hippo" list: the lessons, decisions and corrections from the session that should outlive it. TheSessionEndhook still owns extracting durable memories from the transcript. - a second
SessionStarthook (matchercompact) that runshippo compact-resume, printing that snapshot back into context right after compaction, if it is under 15 minutes old. - a
PostCompacthook that runshippo post-compact. It keeps the summary in the store with secrets scrubbed, and saves each item of that list as a memory that sleep never deletes: at most 10 per compaction, skipping an item an earlier compaction already saved and any item that looks like a secret. An item over 500 characters stays in the compaction's record only. It then prints one line, such as "Hippo saved 3 memories from this compaction and restored your task snapshot." If the store is busy, the summary waits in the store'scompactions-spool/folder andhippo sleepfinishes the save;hippo doctornames any compaction left unfinished for over 10 minutes. The session that compacted does not have tho