cass-memory
!Platform !Runtime !Status !License
Procedural memory for AI coding agents. Transforms scattered agent sessions into persistent, cross-agent memory—so every agent learns from every other agent's experience.
One-liner install (Linux/macOS):
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/cass_memory_system/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
Or via package managers:
# macOS/Linux (Homebrew)
brew install dicklesworthstone/tap/cm
Windows (Scoop)
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cm
🤖 Agent Quickstart (JSON)
Always use --json in agent contexts. stdout = data, stderr = diagnostics, exit 0 = success.
# 1) Get task-specific memory before you start
cm context "implement auth rate limiting" --json
2) See the minimum viable workflow
cm quickstart --json
3) Build the playbook (memory onboarding)
cm onboard status --json
cm onboard sample --fill-gaps --json
cm onboard read /path/to/session.jsonl --template --json
cm onboard mark-done /path/to/session.jsonl
Table of Contents
- Why This Exists
- How It Works
- Key Features
- For AI Agents
- Installation
- CLI Reference
- The ACE Pipeline
- Data Models
- Scoring Algorithm
- Configuration
- MCP Server
- Architecture & Engineering
- Deep Dive: Core Algorithms
- Privacy & Security
- Trauma Guard: Safety System
- Performance Characteristics
- Starter Playbooks
- Extensibility
- Troubleshooting
- Design Philosophy
- Comparison with Alternatives
- Roadmap
💡 Why This Exists
The Problem
AI coding agents accumulate valuable knowledge through sessions: debugging strategies, code patterns, user preferences, project-specific insights. But this knowledge is:
- Trapped in sessions — Each session ends, context is lost forever
- Agent-specific — Claude Code doesn't know what Cursor learned yesterday
- Unstructured — Raw conversation logs aren't actionable as guidance
- Subject to collapse — Naive summarization loses critical nuances and details
The Solution
cass-memory implements a three-layer cognitive architecture that transforms raw session logs into actionable, confidence-tracked rules:
| Layer | Role | Implementation |
|-------|------|----------------|
| Episodic Memory | Raw ground truth from all agents | cass search engine |
| Working Memory | Structured session summaries | Diary entries |
| Procedural Memory | Distilled rules with tracking | Playbook bullets |
This mirrors how human expertise develops: raw experiences (episodic) are consolidated into structured memories (working), which eventually become automatic knowledge (procedural).
Who Benefits
- AI Agents: Get relevant rules and historical context before starting any task
- Developers: Build institutional memory that persists across tools and sessions
- Teams: Share patterns discovered by any team member's AI assistant
- Power Users: Create sophisticated workflows that leverage cross-agent learning
🔄 How It Works
┌─────────────────────────────────────────────────────────────────────┐
│ EPISODIC MEMORY (cass) │
│ Raw session logs from all agents — the "ground truth" │
│ Claude Code │ Codex │ Cursor │ Aider │ PI │ Gemini │ ChatGPT │ ... │
└───────────────────────────┬─────────────────────────────────────────┘
│ cass search
▼
┌─────────────────────────────────────────────────────────────────────┐
│ WORKING MEMORY (Diary) │
│ Structured session summaries bridging raw logs to rules │
│ accomplishments │ decisions │ challenges │ outcomes │
└───────────────────────────┬─────────────────────────────────────────┘
│ reflect + curate (automated)
▼
┌─────────────────────────────────────────────────────────────────────┐
│ PROCEDURAL MEMORY (Playbook) │
│ Distilled rules with confidence tracking │
│ Rules │ Anti-patterns │ Feedback │ Decay │
└─────────────────────────────────────────────────────────────────────┘
Every agent's sessions feed the shared memory. A pattern discovered in Cursor automatically helps Claude Code on the next session.
✨ Key Features
Cross-Agent Learning
Sessions from all your AI coding agents feed a unified knowledge base:
Claude Code session → ┐
Cursor session → │→ Unified Playbook → All agents benefit
Codex session → │
Aider session → │
PI session → ┘
A debugging technique discovered in Cursor is immediately available to Claude Code. No manual knowledge transfer required.
Confidence Decay System
Rules aren't immortal. A rule helpful 8 times in January but never validated since loses confidence over time:
- 90-day half-life: Confidence halves every 90 days without revalidation
- 4x harmful multiplier: One mistake counts 4× as much as one success
- Maturity progression:
candidate→established→proven
Anti-Pattern Learning
Bad rules don't just get deleted—they become warnings:
"Cache auth tokens for performance"
↓ (3 harmful marks)
"PITFALL: Don't cache auth tokens without expiry validation"
When a rule is marked harmful multiple times, it's automatically inverted into an anti-pattern that warns future agents away from the same mistake.
Scientific Validation
New rules aren't blindly accepted. Before a rule joins your playbook, it's validated against your cass history:
Proposed rule: "Always check token expiry before auth debugging"
↓
Evidence gate: Search cass for sessions where this applied
↓
Result: 5 sessions found, 4 successful outcomes → ACCEPT
Rules without historical evidence are flagged as candidates until proven.
Graceful Degradation
The system works even when components are missing:
| Condition | Behavior | |-----------|----------| | No cass | Playbook-only scoring, no history snippets | | No playbook | Empty playbook, commands still work | | No LLM | Deterministic reflection, no semantic enhancement | | Offline | Cached playbook + local diary |
You always get value, even in degraded conditions.
🤖 For AI Agents: The Most Important Section
cass-memory is designed specifically for AI coding agents—not just as an afterthought, but as a first-class design goal. When you're an AI agent working on a codebase, cross-agent memory becomes invaluable: rules learned by other agents, context about design decisions, debugging approaches that worked, and institutional memory that would otherwise be lost.
Why Cross-Agent Memory Matters
Imagine you're Claude Code working on a React authentication bug. With cass-memory, you can instantly access:
- Rules extracted from previous Claude sessions about auth patterns
- Debugging strategies discovered by Cursor last week
- Anti-patterns identified when Codex hit the same issue
- Historical context from Aider about the codebase's auth architecture
cass-memory unifies all this collective intelligence into a single, searchable, actionable knowledge base.
The One Command You Need
cm context "<your task>" --json
Run this command before starting any non-trivial task. It returns:
- Relevant rules from the playbook (scored by task relevance)
- Historical context from past sessions (yours and other agents')
- Anti-patterns to avoid (things that have caused problems)
- Suggested searches for deeper investigation
Self-Documenting API
cass-memory teaches agents how to use it—no external documentation required:
# Quick capability check and self-explanation
cm quickstart --json
→ Returns complete explanation of the system and how to use it
System health and available features
cm doctor --json
→ { checks: [...], recommendations: [...], ... }
Get relevant context for a task
cm context "implement user authentication" --json
→ { relevantBullets: [...], antiPatterns: [...], historySnippets: [...] }
Structured Output Format
Every command supports --json for machine-readable output:
cm context "fix the auth timeout bug" --json
{
"success": true,
"task": "fix the auth timeout bug",
"relevantBullets": [
{
"id": "b-8f3a2c",
"content": "Always check token expiry before other auth debugging",
"effectiveScore": 8.5,
"maturity": "proven",
"relevanceScore": 0.92,
"reasoning": "Extracted from 5 successful sessions"
}
],
"antiPatterns": [
{
"id": "b-x7k9p1",
"content": "Don't cache auth tokens without expiry validation",
"effectiveScore": 3.2
}
],
"historySnippets": [
{
"source_path": "~/.claude/sessions/session-001.jsonl",
"agent": "claude",
"origin": { "kind": "local" },
"snippet": "Fixed timeout by increasing token refresh interval...",
"score": 0.87
}
],
"suggestedCassQueries": [
"cass search 'authentication timeout' --robot --days 30"
],
"degraded": null
}
Design principle: stdout contains only parseable JSON data; all diagnostics go to stderr.
Filtering remote vs local history: historySnippets[].origin.kind is "local" or "remote"; remote hits include origin.host.
Error Handling for Agents
Errors are structured and include recovery hints:
{
"success": false,
"code": "PLAYBOOK_NOT_FOUND",
"error": "Playbook file not found at ~/.cass-memory/playbook.yaml",
"hint": "Run 'cm init' to create a new playbook",
"retryable": false
}
Token Budget Management
AI agents have context limits. cass-memory provides controls to manage output size:
| Flag | Effect |
|------|--------|
| --limit N | Cap number of rules returned |
| --min-score N | Only return rules above score threshold |
| --no-history | Skip historical snippets (faster, smaller) |
| --json | Structured output for parsing |
Inline Feedback (During Work)
When a rule helps or hurts during your work, leave inline feedback:
// [cass: helpful b-8f3a2c] - this rule saved me from a rabbit hole
// [cass: harmful b-x7k9p1] - this advice was wrong for our use case
These comments are automatically parsed during reflection and update rule confidence.
Session Outcome Recording
After completing a task, record the outcome:
# Record successful outcome (positional: status, rules)
cm outcome success b-8f3a2c,b-xyz789 --summary "Fixed auth bug"
Record failure
cm outcome failure b-x7k9p1 --summary "Rule led to wrong approach"
Apply recorded outcomes to playbook
cm outcome-apply
What NOT to Do
You do NOT need to:
- Run
cm reflect(automation handles this) - Run
cm markmanually (use inline comments instead) - Manually add rules to the playbook
- Worry about the learning pipeline
Ready-to-Paste Blurb for AGENTS.md / CLAUDE.md
## Memory System: cass-memory
The Cass Memory System (cm) is a tool for giving agents an effective memory based on the ability to quickly search across previous coding agent sessions across an array of different coding agent tools (e.g., Claude Code, Codex, Gemini-CLI, Cursor, etc) and projects (and even across multiple machines, optionally) and then reflect on what they find and learn in new sessions to draw out useful lessons and takeaways; these lessons are then stored and can be queried and retrieved later, much like how human memory works.
The cm onboard command guides you through analyzing historical sessions and extracting valuable rules.
Quick Start
bash
1. Check status and see recommendations
cm onboard status2. Get sessions to analyze (filtered by gaps in your playbook)
cm onboard sample --fill-gaps3. Read a session with rich context
cm onboard read /path/to/session.jsonl --template4. Add extracted rules (one at a time or batch)
cm playbook add "Your rule content" --category "debugging"Or batch add:
cm playbook add --file rules.json5. Mark session as processed
cm onboard mark-done /path/to/session.jsonlBefore starting complex tasks, retrieve relevant context:bash
cm context "This returns:
- relevantBullets: Rules that may help with your task
- antiPatterns: Pitfalls to avoid
- historySnippets: Past sessions that solved similar problems
- suggestedCassQueries: Searches for deeper investigation
Protocol
- START: Run
cm context "<task>" --json before non-trivial work
- WORK: Reference rule IDs when following them (e.g., "Following b-8f3a2c...")
- FEEDBACK: Leave inline comments when rules help/hurt:
- // [cass: helpful b-xyz] - reason
- // [cass: harmful b-xyz] - reason
- END: Just finish your work. Learning happens automatically.
Key Flags
| Flag | Purpose |
|------|---------|
| --json | Machine-readable JSON output (required!) |
| --limit N | Cap number of rules returned |
| --no-history | Skip historical snippets for faster response |
stdout = data only, stderr = diagnostics. Exit 0 = success.
🎓 Agent-Native Onboarding
Building a playbook from scratch can be daunting—but you don't need to spend money on LLM API calls to do it. The agent-native onboarding system leverages the AI coding agent you're already paying for (via Claude Max, ChatGPT Pro, Cursor Pro, etc.) to analyze historical sessions and extract rules.
The Zero-Cost Approach
Traditional approaches to building a playbook might suggest using external LLM APIs to analyze sessions and extract rules. But if you're already running an AI coding agent, that agent can do the analysis work itself—at no additional cost.
Traditional approach:
Sessions → External LLM API → Rules → $ cost per session
Agent-native approach:
Sessions → Your Agent (already paid for) → Rules → $0 additional cost
The cm onboard command guides your agent through analyzing historical sessions and extracting valuable rules.
Quick Start
# 1. Check status and see recommendations
cm onboard status
2. Get sessions to analyze (filtered by gaps in your playbook)
cm onboard sample --fill-gaps
3. Read a session with rich context
cm onboard read /path/to/session.jsonl --template
4. Add extracted rules (one at a time or batch)
cm playbook add "Your rule content" --category "debugging"
Or batch add:
cm playbook add --file rules.json
5. Mark session as processed
cm onboard mark-done /path/to/session.jsonl
Gap Analysis
Not all playbook categories are equally represented. The gap analysis system identifies underrepresented categories so you can prioritize which sessions to analyze.
Categories tracked:
debugging— Error resolution, bug fixing, tracingtesting— Unit tests, mocks, assertions, coveragearchitecture— Design patterns, module structure, abstractionsworkflow— Task management, CI/CD, deploymentdocumentation— Comments, READMEs, API docsintegration— APIs, HTTP, JSON parsing, endpointscollaboration— Code review, PRs, team coordinationgit— Version control, branching, mergingsecurity— Auth, encryption, vulnerability preventionperformance— Optimization, caching, profiling
critical | 0 rules | High |
| underrepresented | 1-2 rules | Medium |
| adequate | 3-10 rules | Low |
| well-covered | 11+ rules | None |
# View detailed gap analysis
cm onboard gaps
Sample sessions prioritized for gap-filling
cm onboard sample --fill-gaps
Progress Tracking
Onboarding progress persists across sessions, so agents can resume where they left off even after context window limits are reached.
State stored at: ~/.cass-memory/onboarding-state.json
{
"version": 1,
"startedAt": "2025-01-15T10:30:00Z",
"lastUpdatedAt": "2025-01-15T14:45:00Z",
"processedSessions": [
{
"path": "/Users/x/.claude/sessions/session-001.jsonl",
"processedAt": "2025-01-15T11:00:00Z",
"rulesExtracted": 3
}
],
"stats": {
"totalSessionsProcessed": 5,
"totalRulesExtracted": 12
}
}
Commands for progress management:
# Check progress
cm onboard status
Mark a session as done (even if no rules extracted)
cm onboard mark-done /path/to/session.jsonl
Reset progress to start fresh
cm onboard reset
Batch Rule Addition
After analyzing a session, add multiple rules at once using the batch add feature:
# Create a JSON file with rules
cat > rules.json << 'EOF'
[
{"content": "Always run tests before committing", "category": "testing"},
{"content": "Check token expiry before auth debugging", "category": "debugging"},
{"content": "AVOID: Mocking entire modules in tests", "category": "testing"}
]
EOF
Add all rules at once
cm playbook add --file rules.json
Or pipe from another command
echo '[{"content": "Rule from stdin", "category": "workflow"}]' | cm playbook add --file -
Track which session the rules came from
cm playbook add --file rules.json --session /path/to/session.jsonl
The --session flag automatically updates onboarding progress, crediting the session with the number of rules extracted.
Template Output for Rich Context
The --template flag provides agents with rich contextual information to guide extraction:
cm onboard read /path/to/session.jsonl --template --json
Template output includes:
{
"metadata": {
"path": "/path/to/session.jsonl",
"workspace": "/Users/x/project",
"messageCount": 127,
"topicHints": ["debugging", "testing", "git"]
},
"context": {
"relatedRules": [
{"id": "b-abc123", "content": "...", "similarity": 0.72}
],
"playbookGaps": {
"critical": ["security", "performance"],
"underrepresented": ["collaboration"]
},
"suggestedFocus": "This session may contain security patterns - you have NO rules in this area!"
},
"extractionFormat": {
"schema": {"content": "string", "category": "string"},
"categories": ["debugging", "testing", "architecture", ...],
"examples": [...]
},
"sessionContent": "..."
}
Template features:
- Topic hints: Categories detected from session content using keyword matching
- Related rules: Existing playbook rules similar to the session content (via semantic search if enabled)
- Playbook gaps: Categories with missing or few rules
- Suggested focus: AI-generated guidance on what to prioritize based on gaps and detected topics
- Examples: Sample rules showing proper format for each category
Gap-Targeted Sampling Algorithm
When using --fill-gaps, the sampling algorithm prioritizes sessions likely to contain patterns for underrepresented categories:
1. Analyze playbook → Identify critical/underrepresented categories
- Generate search queries from category keywords:
- "security" → "security auth token"
- "performance" → "performance optimize cache"
- Search cass with gap-targeted queries
- Score each session against gaps:
- Session matches critical category: +3 points
- Session matches underrepresented category: +2 points
- Session matches adequate category: +1 point
- Sort sessions by gap score (highest first)
- Filter out already-processed sessions
- Return top N sessions
Category keyword detection uses a lightweight approach (no external APIs required):
// Example: detecting categories from text
const CATEGORY_KEYWORDS = {
debugging: ["debug", "error", "fix", "bug", "trace", "stack", ...],
testing: ["test", "mock", "assert", "expect", "jest", "vitest", ...],
security: ["security", "auth", "token", "encrypt", "permission", ...],
// ... more categories
};
// Count keyword matches → category with most matches wins
Targeted Sampling Options
Filter sessions by various criteria to focus on specific areas:
# Filter by workspace/project
cm onboard sample --workspace /Users/x/my-project
Filter by agent
cm onboard sample --agent claude
cm onboard sample --agent cursor
Filter by recency
cm onboard sample --days 30
Combine filters
cm onboard sample --agent claude --days 14 --workspace /Users/x/api-project
Include already-processed sessions (for re-analysis)
cm onboard sample --include-processed
Adjust sample size
cm onboard sample --limit 20
Onboarding Workflow for Agents
Recommended protocol for AI agents doing onboarding:
## Onboarding Protocol
Phase 1: Assessment
- Run
cm onboard status --json to check current progress
- Run
cm onboard gaps --json to see category distribution
- Decide target: ~20 rules across diverse categories for a good initial playbook
Phase 2: Session Analysis Loop
For each session until target reached:
cm onboard sample --fill-gaps --json → get prioritized sessions
cm onboard read <path> --template --json → get session with context
- Analyze session content, identify 2-5 reusable patterns
- Format as rules following extraction guidelines
cm playbook add --file rules.json --session <path> → add rules
- Repeat with next session
Phase 3: Verification
cm onboard status → verify progress
cm onboard gaps → check remaining gaps
cm stats --json → verify playbook health
Why Agent-Native Onboarding?
- Zero additional cost: Uses the AI agent you're already paying for
- Better analysis: Your agent understands your codebase context
- Resumable: Progress persists across sessions
- Gap-aware: Automatically prioritizes underrepresented categories
- Batch-friendly: Add multiple rules efficiently
- Self-documenting: Template output guides extraction
📦 Installation
One-Liner (Recommended)
Recommended: Homebrew (macOS/Linux)
brew install dicklesworthstone/tap/cm
This method provides:
- Automatic updates via
brew upgrade - Dependency management
- Easy uninstall via
brew uninstall
Windows: Scoop
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cm
Alternative: Install Script
Linux/macOS:
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/cass_memory_system/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
Direct Downloads:
- Linux x64
- Linux ARM64 (releases after v0.2.14; e.g. Linux containers on Apple Silicon, Graviton)
- macOS Apple Silicon
- macOS Intel
- Windows x64
Installer Options
# Full options
install.sh [--version vX.Y.Z] [--dest DIR] [--system] [--easy-mode] [--verify]
[--from-source] [--quiet]
Examples
install.sh --version v0.2.2 --verify # Specific version
install.sh --system --verify # Install to /usr/local/bin (requires sudo)
install.sh --from-source # Build from source (requires bun)
| Flag | Effect |
|------|--------|
| --version | Install specific version (default: latest) |
| --dest DIR | Install to custom directory (default: ~/.local/bin) |
| --system | Install to /usr/local/bin (requires sudo) |
| --easy-mode | Non-interactive, auto-configure PATH |
| --verify | Run self-test after install |
| --from-source | Build from source instead of downloading binary |
| --quiet | Suppress output |
Installer Robustness
The installer is designed for reliability:
- No API rate limits: Uses GitHub redirect detection instead of API calls, avoiding the 60 requests/hour unauthenticated limit
- Checksum verification: Downloads
.sha256files and verifies binary integrity before installation - Concurrent install protection: Lock file prevents multiple installers from running simultaneously
- Graceful fallbacks: If binary download fails, automatically falls back to building from source
- Cross-platform checksums: Uses
sha256sumon Linux,shasum -a 256on macOS
From Source
git clone https://github.com/Dicklesworthstone/cass_memory_system.git
cd cass_memory_system
bun install
bun run build
sudo mv ./dist/cass-memory /usr/local/bin/cm
Prerequisites
- cass CLI: The episodic memory layer. Install from cass repo
- LLM API Key (optional): For AI-powered reflection. Set
ANTHROPIC_API_KEY,OPENAI_API_KEY, orGOOGLE_GENERATIVE_AI_API_KEY
Building cass from source requires rustup + nightly. The cass source tree's
.cargo/config.tomlunconditionally passes the nightly-only-Z threadsrustflag,
and its rust-toolchain.toml pins a nightly channel that only a rustup proxy will
honor. A non-rustup Rust (e.g. Homebrew'scargo/rustc) fails before compilation
with `the optionZis only accepted on the nightly compiler. If you need to
build a cass fix from source (e.g. a commit not yet in a stable release), install
rustup and run rustup toolchain install nightly first.
cm doctor reports whether cass is missing entirely or merely has an unavailable
lexical index (fix the latter with cass index --full).
Verify Installation
cm --version
cm doctor --json
Initial Setup
# Initialize (creates global config and playbook)
cm init
Or with a starter playbook for common patterns
cm init --starter typescript # or: react, python, go
cm starters # list available starters
Automating Reflection
The key to the system is automated reflection. Set up a cron job or hook:
# Daily reflection on recent sessions
cm reflect --days 7 --json
Only one project's sessions, only some agents
cm reflect --workspace ~/code/api --agent claude,codex --json
Via cron (runs at 2am daily)
0 2 * /usr/local/bin/cm reflect --days 7 >> ~/.cass-memory/reflect.log 2>&1
For Claude Code, install the auto-reflect hook once and every finished session is reflected in the background:
cm hook install # this project (.claude/settings.json)
cm hook install --global # every project (~/.claude/settings.json)
cm hook status # where it is installed
cm hook uninstall # remove it (other hooks are left alone)
This adds a SessionEnd hook that runs cm hook session-end. It reads the transcript path from the hook payload and starts cm reflect --session detached, so Claude Code never waits on it. Output goes to ~/.cass-memory/hooks.log. It skips transcripts written by cm's own LLM subprocess calls, and sessions that end inside a background reflect, so it cannot loop. Reflection keeps your usual budget limits and processed-session tracking, so re-running a session costs nothing.
For other agents, call the same entry point from a wrapper script or the agent's own end-of-session hook:
cm hook session-end --transcript ~/.codex/sessions/2026/10/06/rollout.jsonl
Long sessions: reflect often. The diary step reads at most diaryMaxInputChars (default 50,000) characters of a transcript: the start, the end, and error/correction lines picked from the middle (see Reflection Settings). A session that goes on for hours can be much longer than that. cm reflect remembers how far into each session it got, and when a session has grown it reflects only the new part. Running it often (from a hook, or cron every 30–60 minutes) therefore keeps each new part small enough to be read in full. Raising diaryMaxInputChars is the other option when your model has the context for it.
Should agents write their own lessons? They don't need to. cm is built so that a separate reflection step pulls lessons out of the transcript, and the working agent can just finish its task. If your agents do end with a list such as Lessons for memory: ..., that is fine input: the diary step reads it like any other text in the transcript, as long as it falls inside the window above (the end of a session always does). Treat such lists as hints for the diary, not as rules: cm still checks candidate rules against history and feedback before trusting them.
🛠️ CLI Reference
Quick Reference Table
| Command | Purpose | Agent Use |
|---------|---------|-----------|
| cm context " | Get rules + history for task | Primary |
| cm quickstart --json | Self-documentation | Setup |
| cm robot-docs [topic] | Machine-readable docs: commands, examples, exit codes, JSON schemas | Integration |
| cm doctor --json | System health check | Diagnostics |
| cm playbook list | Show all rules | Inspection |
| cm similar " | Find similar rules | Search |
| cm stats --json | Playbook metrics | Analytics |
| cm trauma list | Show dangerous patterns | Safety |
| cm guard --status | Check safety hook status | Safety |
Error Output & Exit Codes
When --json (or --format json) is enabled, errors are printed to stdout as a single JSON object:
{
"success": false,
"error": "…",
"code": "…",
"exitCode": 2,
"recovery": ["…", "…"],
"cause": "…",
"docs": "README.md#-troubleshooting",
"hint": "…",
"details": {}
}
Exit codes are categorized for scripting:
- 1
internal (bug / unexpected) - 2
user input / usage - 3
configuration - 4
filesystem - 5
network - 6
cass - 7
LLM/provider
Agent Commands (Primary Workflow)
# Get context before starting a task (THE MAIN COMMAND)
cm context "implement user authentication" --json
Self-documenting explanation of the system
cm quickstart --json
Find rules similar to a query
cm similar "error handling best practices"
Machine-readable CLI docs (JSON): all topics, or one of
guide | commands | examples | exit-codes | schemas
cm robot-docs
cm robot-docs schemas
cm robot-docs always prints one JSON document to stdout (data.schemaVersion, data.version, data.topics). commands is read from the live CLI definition, so it lists exactly the commands, arguments and options this build accepts. schemas gives JSON Schema (draft-07) for the JSON envelope (success and error) and for the data of context, quickstart and onboard status; the schemas are open, so ignore fields you don't know.
Playbook Commands (Inspect & Manage Rules)
# List all active rules
cm playbook list
Get detailed info for a specific rule
cm playbook get b-8f3a2c
Add a new rule manually
cm playbook add "Always run tests before committing"
Remove/deprecate a rule
cm playbook remove b-xyz789 --reason "No longer applicable"
Export playbook for backup or sharing
cm playbook export > playbook-backup.yaml
Import playbook from file. Rules matching an existing one by text (exact,
or near-identical wording) merge their feedback into it instead of becoming
duplicates; contradictions with existing rules are reported.
cm playbook import shared-playbook.yaml
Find rules that contradict each other (with a keep/retire suggestion)
cm playbook conflicts --json
Strip feedback that came from cm's own LLM subprocess transcripts (or any
session path pattern), recount helpful/harmful, re-derive maturity.
Backs up each playbook first; bullets left without genuine support are
reported, and deprecated only with --deprecate-orphans.
cm playbook scrub --cm-subprocess-calls --dry-run --json
cm playbook scrub --from-sessions "/old-project/" --deprecate-orphans
Show top N most effective rules
cm top 10
Find rules without recent feedback (stale)
cm stale --days 60
Show why a rule exists (provenance)
cm why b-8f3a2c
Get playbook health metrics
cm stats --json
Learning Commands (Feedback & Reflection)
# Process recent sessions into rules (usually automated)
cm reflect --days 7 --json
Manual feedback on a rule
cm mark b-8f3a2c --helpful
cm mark b-xyz789 --harmful --reason "Caused regression"
Record session outcome (positional: status, rules)
cm outcome success b-8f3a2c,b-def456
cm outcome failure b-x7k9p1 --summary "Auth approach failed"
Apply recorded outcomes to playbook
cm outcome-apply
Validate a proposed rule against history
cm validate "Always check null before dereferencing"
Deprecate a rule permanently
cm forget b-xyz789 --reason "Superseded by better pattern"
Revert a curation decision
cm undo b-xyz789
Check sessions for rule violations
cm audit --days 30
System Commands (Setup & Diagnostics)
# Initialize configuration and playbook
cm init
cm init --starter typescript # With starter template
cm starters # List available starters
Check system health
cm doctor --json
cm doctor --fix # Auto-fix issues
Check for / install a newer release
cm update --check # Exit code 0 = up to date, 1 = update available
cm update # Re-runs the documented install script (interactive only)
Export rules for AGENTS.md
cm project --format agents.md
Show LLM cost and usage statistics
cm usage
Run MCP HTTP server for agent integration (default: 127.0.0.1:8765)
cm serve
Privacy controls
cm privacy status
cm privacy enable # Enable cross-agent enrichment
cm privacy disable # Disable cross-agent enrichment
Safety Commands (Trauma Guard)
# View active trauma patterns
cm trauma list
Add a dangerous pattern
cm trauma add "DROP TABLE" --description "Database table deletion" --severity critical
Temporarily disable a pattern
cm trauma heal t-abc123 --reason "Intentional migration cleanup"
Permanently remove a pattern
cm trauma remove t-abc123
Scan sessions for potential traumas
cm trauma scan --days 30
Import patterns from file
cm trauma import shared-traumas.yaml
Install safety hooks
cm guard --install # Claude Code pre-tool hook
cm guard --git # Git pre-commit hook
cm guard --install --git # Both
cm guard --status # Check installation status
Command Output Modes
All commands support multiple output formats:
| Flag | Output |
|------|--------|
| (default) | Human-readable text |
| --json | Machine-readable JSON |
| --quiet | Minimal output |
🔬 The ACE Pipeline
cass-memory implements the Agentic Context Engineering (ACE) framework with four stages:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ GENERATOR │ ──▶ │ REFLECTOR │ ──▶ │ VALIDATOR │ ──▶ │ CURATOR │
│ │ │ │ │ │ │ │
│ Pre-task │ │ LLM pattern │ │ Evidence │ │ Deterministic│
│ context │ │ extraction │ │ gate against │ │ delta merge │
│ hydration │ │ from diary │ │ cass history │ │ (NO LLM!) │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘
Stage 1: Generator (cm context)
Hydrates the agent with relevant rules before starting a task.
Implementation (src/commands/context.ts):
async function generateContext(task: string): Promise<ContextResult> {
// 1. Extract keywords from task
const keywords = extractKeywords(task);
// 2. Relevance: BM25 over the playbook (IDF-weighted share of the query a
// bullet covers, 0..10), blended with embedding similarity when enabled
const lexical = scoreLexicalRelevance(bullets, keywords);
const scored = bullets.map(b => {
const relevanceScore = blend(lexical.get(b.id), cosine(taskEmbedding, b.embedding) * 10);
// 3. Track record only reorders within a bounded band: [1 - w, 1 + w],
// damped while a bullet has few marks (feedbackWeight, default 0.25)
return { ...b, relevanceScore, finalScore: relevanceScore * getFeedbackMultiplier(b) };
}).sort(byFinalScore);
// 4. Keep what is relevant and fits: absolute + relative relevance floors,
// maxBulletsInContext, then contextTokenBudget
const { selected, stats } = selectContextBullets(scored, config);
// 5. Search cass for historical context
const historySnippets = await safeCassSearch(task, {
limit: config.maxHistoryInContext,
days: config.sessionLookbackDays
});
return {
task,
relevantBullets: selected.filter(isRule).map(toContextBullet), // compact: no event log
antiPatterns: selected.filter(isAntiPattern).map(toContextBullet),
historySnippets,
retrieval: stats, // candidates / returned / dropped by relevance, limit, budget
suggestedCassQueries: generateSuggestedQueries(task, keywords)
};
}
Bullets in cm context output are a compact projection (id, content, category, kind, maturity, tags, counts, scores, reasoning). The full record, including the feedback-event log and source sessions, is one cm playbook get away.
Stage 2: Reflector (cm reflect)
Processes unprocessed sessions to extract new rules.
How It Works:
- Find unprocessed sessions via cass timeline API
- Generate diary entries - structured summaries of each session:
- Run LLM reflector - multi-iteration pattern extraction
- Output deltas - add/helpful/harmful/deprecate operations
Analyze this session diary and extract actionable rules.
For each pattern you identify:
- Formulate as an imperative rule ("Always...", "Never...", "When X, do Y")
- Categorize (testing, debugging, git, architecture, etc.)
- Assess scope (global, workspace, language, framework)
- Provide reasoning for why this is valuable
Output format: JSON array of proposed rules
Stage 3: Validator
Scientific gate ensuring rules have evidence.
Implementation (src/ace/validator.ts):
async function validateRule(rule: ProposedRule): Promise<ValidationResult> {
// 1. Evidence count gate - fast pre-filter
const evidenceCount = await countEvidenceInCass(rule);
if (evidenceCount < MIN_EVIDENCE_COUNT) {
return { verdict: 'NEEDS_MORE_EVIDENCE', evidenceCount };
}
// 2. LLM validation - deep analysis
const llmVerdict = await llmValidate(rule, evidenceCount);
// 3. Final verdict
return {
verdict: llmVerdict.valid ? 'ACCEPT' : 'REJECT',
reasoning: llmVerdict.reasoning,
evidenceCount
};
}
Stage 4: Curator
Critical: NO LLM in this stage - pure deterministic logic to prevent context collapse.
Why No LLM? LLMs tend to "drift" when iteratively refining their own outputs. By keeping curation deterministic, we:
- Prevent feedback loops that degrade rule quality
- Ensure reproducible behavior
- Avoid expensive redundant LLM calls
function curate(playbook: Playbook, deltas: Delta[]): Playbook {
for (const delta of deltas) {
switch (delta.type) {
case 'add':
// Conflict detection - find contradicting rules
const conflicts = findConflicts(playbook, delta.rule);
if (conflicts.length > 0) {
resolveConflicts(playbook, conflicts, delta);
}
// Duplicate detection - hash-based and semantic
const duplicate = findDuplicate(playbook, delta.rule);
if (duplicate) {
mergeDuplicate(duplicate, delta.rule);
continue;
}
// Add new rule
playbook.bullets.push(delta.rule);
break;
case 'helpful':
// Update feedback
addFeedback(playbook, delta.bulletId, 'helpful', delta.timestamp);
maybePromoteMaturity(playbook, delta.bulletId);
break;
case 'harmful':
// Update feedback with harmful multiplier consideration
addFeedback(playbook, delta.bulletId, 'harmful', delta.timestamp);
maybeConvertToAntiPattern(playbook, delta.bulletId);
break;
case 'deprecate':
deprecateRule(playbook, delta.bulletId, delta.reason);
break;
}
}
// Maturity promotions based on feedback thresholds
promoteQualifiedRules(playbook);
return playbook;
}
📊 Data Models
Playbook Bullet
The core data structure for rules:
interface PlaybookBullet {
// Identity
id: string; // "b-{timestamp}-{random}"
content: string; // The actual rule
category: string; // "testing" | "git" | "debugging" | etc.
// Classification
scope: "global" | "workspace" | "language" | "framework" | "task";
type: "rule" | "anti-pattern";
kind: "project_convention" | "stack_pattern" | "workflow_rule" | "anti_pattern";
isNegative: boolean;
// Lifecycle
state: "draft" | "active" | "retired";
maturity: "candidate" | "established" | "proven" | "deprecated";
// Feedback & Scoring
helpfulCount: number; // Raw count
harmfulCount: number; // Raw count
feedbackEvents: FeedbackEvent[]; // Full history with timestamps
effectiveScore?: number; // Calculated decay-adjusted score
// Provenance
sourceSessions: string[]; // Sessions that generated this rule
sourceAgents: string[]; // Agents involved (claude, codex, cursor, etc.)
reasoning?: string; // Why this rule exists
sourceSession?: string; // Flexible field for reflection metadata
// Metadata
tags: string[];
embedding?: number[]; // Semantic search vector (768 dimensions)
pinned: boolean; // Prevent auto-deprecation
deprecated: boolean;
// Timestamps
createdAt: string; // ISO 8601
updatedAt: string; // ISO 8601
}
Feedback Event
Immutable record of feedback:
interface FeedbackEvent {
id: string; // UUID
type: "helpful" | "harmful";
timestamp: string; // ISO 8601
source: "inline" | "manual" | "outcome" | "audit";
sessionPath?: string; // Session where feedback was given
reason?: string; // Optional explanation
}
Diary Entry
Working memory structure:
interface DiaryEntry {
id: string; // Deterministic hash of content
sessionPath: string; // Original cass session
timestamp: string; // ISO 8601
agent: string; // "claude" | "codex" | "cursor" | etc.
workspace?: string; // Project directory
// Session status
status: "success" | "failure" | "mixed";
// Core content
accomplishments: string[]; // What was completed
decisions: string[]; // Design choices made
challenges: string[]; // Problems encountered
keyLearnings: string[]; // Reusable insights
preferences: string[]; // User style revelations
// Cross-agent enrichment
relatedSessions?: RelatedSession[];
// Search optimization
searchAnchors: string[]; // Keywords for retrieval
tags: string[]; // File/component tags
}
Session Outcome
Explicit outcome recording:
interface SessionOutcome {
id: string; // UUID
timestamp: string; // ISO 8601
status: "success" | "failure" | "mixed";
// Associated rules
rulesUsed: string[]; // Bullet IDs that were followed
rulesViolated?: string[]; // Bullet IDs that were ignored
// Context
task?: string; // What was attempted
summary?: string; // Brief description
sessionPath?: string; // Source session
// Feedback impact
applied: boolean; // Whether outcome-apply has processed this
appliedAt?: string; // When it was applied
}
📈 Scoring Algorithm
Effective Score Calculation
The effective score determines rule ranking, maturity transitions, and anti-pattern conversion:
function getEffectiveScore(bullet: PlaybookBullet): number {
const HARMFUL_MULTIPLIER = 4; // One harmful = 4× one helpful
const HALF_LIFE_DAYS = 90; // 90 days for half decay
// Calculate decayed helpful count
const decayedHelpful = bullet.feedbackEvents
.filter(e => e.type === "helpful")
.reduce((sum, event) => {
const daysAgo = daysSince(event.timestamp);
const decayFactor = Math.pow(0.5, daysAgo / HALF_LIFE_DAYS);
return sum + decayFactor;
}, 0);
// Calculate decayed harmful count
const decayedHarmful = bullet.feedbackEvents
.filter(e => e.type === "harmful")
.reduce((sum, event) => {
const daysAgo = daysSince(event.timestamp);
const decayFactor = Math.pow(0.5, daysAgo / HALF_LIFE_DAYS);
return sum + decayFactor;
}, 0);
// Effective score with harmful multiplier
return decayedHelpful - (HARMFUL_MULTIPLIER * decayedHarmful);
}
Score Decay Visualization
Initial score: 10.0 (10 helpful marks today)
After 90 days (half-life): 5.0
After 180 days: 2.5
After 270 days: 1.25
After 365 days: 0.78
Why Time Decay?
- Codebases evolve - A pattern helpful for React 17 may not apply to React 19
- Tools change - Agent capabilities improve over time
- Prevents stale rules - Forces ongoing validation
- Natural cleanup - Unvalidated rules fade rather than accumulate
Maturity State Machine
`` ┌──────────┐ ┌─────────────┐ ┌────────┐ │ candidate│──────▶│ established │───▶│ proven │ └──────────┘ └─────────────┘ └────────┘ │ │ │ │ │ (harmful >25%) │ │ ▼ │ │ ┌─────────────┐ │
... (README truncated for length)