Profile
Back to NewsBack
GitHub Trending 30 min
Reader Mode
Dicklesworthstone/cass_memory_system: Procedural memory for AI coding agents: transforms scattered session history into persistent, cross-agent memory so every agent learns from every other

Dicklesworthstone/cass_memory_system: Procedural memory for AI coding agents: transforms scattered session history into persistent, cross-agent memory so every agent learns from every other

cass-memory

cass-memory - Procedural memory for AI coding agents

!Platform !Runtime !Status !License

Procedural memory for AI coding agents. Transforms scattered agent sessions into persistent, cross-agent memory—so every agent learns from every other agent's experience.

One-liner install (Linux/macOS):

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/cass_memory_system/main/install.sh?$(date +%s)" \
  | bash -s -- --easy-mode --verify

Or via package managers:

# macOS/Linux (Homebrew)
brew install dicklesworthstone/tap/cm

Windows (Scoop)

scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket scoop install dicklesworthstone/cm


🤖 Agent Quickstart (JSON)

Always use --json in agent contexts. stdout = data, stderr = diagnostics, exit 0 = success.

# 1) Get task-specific memory before you start
cm context "implement auth rate limiting" --json

2) See the minimum viable workflow

cm quickstart --json

3) Build the playbook (memory onboarding)

cm onboard status --json cm onboard sample --fill-gaps --json cm onboard read /path/to/session.jsonl --template --json cm onboard mark-done /path/to/session.jsonl

Table of Contents


💡 Why This Exists

The Problem

AI coding agents accumulate valuable knowledge through sessions: debugging strategies, code patterns, user preferences, project-specific insights. But this knowledge is:

  1. Trapped in sessions — Each session ends, context is lost forever
  2. Agent-specific — Claude Code doesn't know what Cursor learned yesterday
  3. Unstructured — Raw conversation logs aren't actionable as guidance
  4. Subject to collapse — Naive summarization loses critical nuances and details
You've solved authentication bugs three times this month across different agents. Each time, you started from scratch because the knowledge from previous sessions was inaccessible.

The Solution

cass-memory implements a three-layer cognitive architecture that transforms raw session logs into actionable, confidence-tracked rules:

| Layer | Role | Implementation | |-------|------|----------------| | Episodic Memory | Raw ground truth from all agents | cass search engine | | Working Memory | Structured session summaries | Diary entries | | Procedural Memory | Distilled rules with tracking | Playbook bullets |

This mirrors how human expertise develops: raw experiences (episodic) are consolidated into structured memories (working), which eventually become automatic knowledge (procedural).

Who Benefits

  • AI Agents: Get relevant rules and historical context before starting any task
  • Developers: Build institutional memory that persists across tools and sessions
  • Teams: Share patterns discovered by any team member's AI assistant
  • Power Users: Create sophisticated workflows that leverage cross-agent learning

🔄 How It Works

┌─────────────────────────────────────────────────────────────────────┐
│                    EPISODIC MEMORY (cass)                           │
│   Raw session logs from all agents — the "ground truth"             │
│   Claude Code │ Codex │ Cursor │ Aider │ PI │ Gemini │ ChatGPT │ ... │
└───────────────────────────┬─────────────────────────────────────────┘
                            │ cass search
                            ▼
┌─────────────────────────────────────────────────────────────────────┐
│                    WORKING MEMORY (Diary)                           │
│   Structured session summaries bridging raw logs to rules           │
│   accomplishments │ decisions │ challenges │ outcomes               │
└───────────────────────────┬─────────────────────────────────────────┘
                            │ reflect + curate (automated)
                            ▼
┌─────────────────────────────────────────────────────────────────────┐
│                    PROCEDURAL MEMORY (Playbook)                     │
│   Distilled rules with confidence tracking                          │
│   Rules │ Anti-patterns │ Feedback │ Decay                          │
└─────────────────────────────────────────────────────────────────────┘

Every agent's sessions feed the shared memory. A pattern discovered in Cursor automatically helps Claude Code on the next session.


✨ Key Features

Cross-Agent Learning

Sessions from all your AI coding agents feed a unified knowledge base:

Claude Code session    →  ┐
Cursor session         →  │→  Unified Playbook  →  All agents benefit
Codex session          →  │
Aider session          →  │
PI session             →  ┘

A debugging technique discovered in Cursor is immediately available to Claude Code. No manual knowledge transfer required.

Confidence Decay System

Rules aren't immortal. A rule helpful 8 times in January but never validated since loses confidence over time:

  • 90-day half-life: Confidence halves every 90 days without revalidation
  • 4x harmful multiplier: One mistake counts 4× as much as one success
  • Maturity progression: candidate → established → proven
This prevents stale rules from polluting your playbook while rewarding consistently helpful guidance.

Anti-Pattern Learning

Bad rules don't just get deleted—they become warnings:

"Cache auth tokens for performance"
    ↓ (3 harmful marks)
"PITFALL: Don't cache auth tokens without expiry validation"

When a rule is marked harmful multiple times, it's automatically inverted into an anti-pattern that warns future agents away from the same mistake.

Scientific Validation

New rules aren't blindly accepted. Before a rule joins your playbook, it's validated against your cass history:

Proposed rule: "Always check token expiry before auth debugging"
    ↓
Evidence gate: Search cass for sessions where this applied
    ↓
Result: 5 sessions found, 4 successful outcomes → ACCEPT

Rules without historical evidence are flagged as candidates until proven.

Graceful Degradation

The system works even when components are missing:

| Condition | Behavior | |-----------|----------| | No cass | Playbook-only scoring, no history snippets | | No playbook | Empty playbook, commands still work | | No LLM | Deterministic reflection, no semantic enhancement | | Offline | Cached playbook + local diary |

You always get value, even in degraded conditions.


🤖 For AI Agents: The Most Important Section

cass-memory is designed specifically for AI coding agents—not just as an afterthought, but as a first-class design goal. When you're an AI agent working on a codebase, cross-agent memory becomes invaluable: rules learned by other agents, context about design decisions, debugging approaches that worked, and institutional memory that would otherwise be lost.

Why Cross-Agent Memory Matters

Imagine you're Claude Code working on a React authentication bug. With cass-memory, you can instantly access:

  • Rules extracted from previous Claude sessions about auth patterns
  • Debugging strategies discovered by Cursor last week
  • Anti-patterns identified when Codex hit the same issue
  • Historical context from Aider about the codebase's auth architecture
This cross-pollination of knowledge across different AI agents is transformative. Each agent has different strengths, different context windows, and encounters different problems. cass-memory unifies all this collective intelligence into a single, searchable, actionable knowledge base.

The One Command You Need

cm context "<your task>" --json

Run this command before starting any non-trivial task. It returns:

  • Relevant rules from the playbook (scored by task relevance)
  • Historical context from past sessions (yours and other agents')
  • Anti-patterns to avoid (things that have caused problems)
  • Suggested searches for deeper investigation

Self-Documenting API

cass-memory teaches agents how to use it—no external documentation required:

# Quick capability check and self-explanation
cm quickstart --json

→ Returns complete explanation of the system and how to use it

System health and available features

cm doctor --json

→ { checks: [...], recommendations: [...], ... }

Get relevant context for a task

cm context "implement user authentication" --json

→ { relevantBullets: [...], antiPatterns: [...], historySnippets: [...] }

Structured Output Format

Every command supports --json for machine-readable output:

cm context "fix the auth timeout bug" --json
{
  "success": true,
  "task": "fix the auth timeout bug",
  "relevantBullets": [
    {
      "id": "b-8f3a2c",
      "content": "Always check token expiry before other auth debugging",
      "effectiveScore": 8.5,
      "maturity": "proven",
      "relevanceScore": 0.92,
      "reasoning": "Extracted from 5 successful sessions"
    }
  ],
  "antiPatterns": [
    {
      "id": "b-x7k9p1",
      "content": "Don't cache auth tokens without expiry validation",
      "effectiveScore": 3.2
    }
  ],
  "historySnippets": [
    {
      "source_path": "~/.claude/sessions/session-001.jsonl",
      "agent": "claude",
      "origin": { "kind": "local" },
      "snippet": "Fixed timeout by increasing token refresh interval...",
      "score": 0.87
    }
  ],
  "suggestedCassQueries": [
    "cass search 'authentication timeout' --robot --days 30"
  ],
  "degraded": null
}

Design principle: stdout contains only parseable JSON data; all diagnostics go to stderr.

Filtering remote vs local history: historySnippets[].origin.kind is "local" or "remote"; remote hits include origin.host.

Error Handling for Agents

Errors are structured and include recovery hints:

{
  "success": false,
  "code": "PLAYBOOK_NOT_FOUND",
  "error": "Playbook file not found at ~/.cass-memory/playbook.yaml",
  "hint": "Run 'cm init' to create a new playbook",
  "retryable": false
}

Token Budget Management

AI agents have context limits. cass-memory provides controls to manage output size:

| Flag | Effect | |------|--------| | --limit N | Cap number of rules returned | | --min-score N | Only return rules above score threshold | | --no-history | Skip historical snippets (faster, smaller) | | --json | Structured output for parsing |

Inline Feedback (During Work)

When a rule helps or hurts during your work, leave inline feedback:

// [cass: helpful b-8f3a2c] - this rule saved me from a rabbit hole

// [cass: harmful b-x7k9p1] - this advice was wrong for our use case

These comments are automatically parsed during reflection and update rule confidence.

Session Outcome Recording

After completing a task, record the outcome:

# Record successful outcome (positional: status, rules)
cm outcome success b-8f3a2c,b-xyz789 --summary "Fixed auth bug"

Record failure

cm outcome failure b-x7k9p1 --summary "Rule led to wrong approach"

Apply recorded outcomes to playbook

cm outcome-apply

What NOT to Do

You do NOT need to:

  • Run cm reflect (automation handles this)
  • Run cm mark manually (use inline comments instead)
  • Manually add rules to the playbook
  • Worry about the learning pipeline
The system learns from your sessions automatically. Your job is just to query context before working.

Ready-to-Paste Blurb for AGENTS.md / CLAUDE.md

## Memory System: cass-memory

The Cass Memory System (cm) is a tool for giving agents an effective memory based on the ability to quickly search across previous coding agent sessions across an array of different coding agent tools (e.g., Claude Code, Codex, Gemini-CLI, Cursor, etc) and projects (and even across multiple machines, optionally) and then reflect on what they find and learn in new sessions to draw out useful lessons and takeaways; these lessons are then stored and can be queried and retrieved later, much like how human memory works.

The cm onboard command guides you through analyzing historical sessions and extracting valuable rules.

Quick Start

bash

1. Check status and see recommendations

cm onboard status

2. Get sessions to analyze (filtered by gaps in your playbook)

cm onboard sample --fill-gaps

3. Read a session with rich context

cm onboard read /path/to/session.jsonl --template

4. Add extracted rules (one at a time or batch)

cm playbook add "Your rule content" --category "debugging"

Or batch add:

cm playbook add --file rules.json

5. Mark session as processed

cm onboard mark-done /path/to/session.jsonl
Before starting complex tasks, retrieve relevant context:
bash cm context "" --json
This returns:
  • relevantBullets: Rules that may help with your task
  • antiPatterns: Pitfalls to avoid
  • historySnippets: Past sessions that solved similar problems
  • suggestedCassQueries: Searches for deeper investigation

Protocol

  1. START: Run cm context "<task>" --json before non-trivial work
  2. WORK: Reference rule IDs when following them (e.g., "Following b-8f3a2c...")
  3. FEEDBACK: Leave inline comments when rules help/hurt:
- // [cass: helpful b-xyz] - reason - // [cass: harmful b-xyz] - reason
  1. END: Just finish your work. Learning happens automatically.

Key Flags

| Flag | Purpose | |------|---------| | --json | Machine-readable JSON output (required!) | | --limit N | Cap number of rules returned | | --no-history | Skip historical snippets for faster response |

stdout = data only, stderr = diagnostics. Exit 0 = success.


🎓 Agent-Native Onboarding

Building a playbook from scratch can be daunting—but you don't need to spend money on LLM API calls to do it. The agent-native onboarding system leverages the AI coding agent you're already paying for (via Claude Max, ChatGPT Pro, Cursor Pro, etc.) to analyze historical sessions and extract rules.

The Zero-Cost Approach

Traditional approaches to building a playbook might suggest using external LLM APIs to analyze sessions and extract rules. But if you're already running an AI coding agent, that agent can do the analysis work itself—at no additional cost.

Traditional approach:
  Sessions → External LLM API → Rules → $ cost per session

Agent-native approach: Sessions → Your Agent (already paid for) → Rules → $0 additional cost

The cm onboard command guides your agent through analyzing historical sessions and extracting valuable rules.

Quick Start

# 1. Check status and see recommendations
cm onboard status

2. Get sessions to analyze (filtered by gaps in your playbook)

cm onboard sample --fill-gaps

3. Read a session with rich context

cm onboard read /path/to/session.jsonl --template

4. Add extracted rules (one at a time or batch)

cm playbook add "Your rule content" --category "debugging"

Or batch add:

cm playbook add --file rules.json

5. Mark session as processed

cm onboard mark-done /path/to/session.jsonl

Gap Analysis

Not all playbook categories are equally represented. The gap analysis system identifies underrepresented categories so you can prioritize which sessions to analyze.

Categories tracked:

  • debugging — Error resolution, bug fixing, tracing
  • testing — Unit tests, mocks, assertions, coverage
  • architecture — Design patterns, module structure, abstractions
  • workflow — Task management, CI/CD, deployment
  • documentation — Comments, READMEs, API docs
  • integration — APIs, HTTP, JSON parsing, endpoints
  • collaboration — Code review, PRs, team coordination
  • git — Version control, branching, merging
  • security — Auth, encryption, vulnerability prevention
  • performance — Optimization, caching, profiling
Category status thresholds: | Status | Rule Count | Priority | |--------|------------|----------| | critical | 0 rules | High | | underrepresented | 1-2 rules | Medium | | adequate | 3-10 rules | Low | | well-covered | 11+ rules | None |

# View detailed gap analysis
cm onboard gaps

Sample sessions prioritized for gap-filling

cm onboard sample --fill-gaps

Progress Tracking

Onboarding progress persists across sessions, so agents can resume where they left off even after context window limits are reached.

State stored at: ~/.cass-memory/onboarding-state.json

{
  "version": 1,
  "startedAt": "2025-01-15T10:30:00Z",
  "lastUpdatedAt": "2025-01-15T14:45:00Z",
  "processedSessions": [
    {
      "path": "/Users/x/.claude/sessions/session-001.jsonl",
      "processedAt": "2025-01-15T11:00:00Z",
      "rulesExtracted": 3
    }
  ],
  "stats": {
    "totalSessionsProcessed": 5,
    "totalRulesExtracted": 12
  }
}

Commands for progress management:

# Check progress
cm onboard status

Mark a session as done (even if no rules extracted)

cm onboard mark-done /path/to/session.jsonl

Reset progress to start fresh

cm onboard reset

Batch Rule Addition

After analyzing a session, add multiple rules at once using the batch add feature:

# Create a JSON file with rules
cat > rules.json << 'EOF'
[
  {"content": "Always run tests before committing", "category": "testing"},
  {"content": "Check token expiry before auth debugging", "category": "debugging"},
  {"content": "AVOID: Mocking entire modules in tests", "category": "testing"}
]
EOF

Add all rules at once

cm playbook add --file rules.json

Or pipe from another command

echo '[{"content": "Rule from stdin", "category": "workflow"}]' | cm playbook add --file -

Track which session the rules came from

cm playbook add --file rules.json --session /path/to/session.jsonl

The --session flag automatically updates onboarding progress, crediting the session with the number of rules extracted.

Template Output for Rich Context

The --template flag provides agents with rich contextual information to guide extraction:

cm onboard read /path/to/session.jsonl --template --json

Template output includes:

{
  "metadata": {
    "path": "/path/to/session.jsonl",
    "workspace": "/Users/x/project",
    "messageCount": 127,
    "topicHints": ["debugging", "testing", "git"]
  },
  "context": {
    "relatedRules": [
      {"id": "b-abc123", "content": "...", "similarity": 0.72}
    ],
    "playbookGaps": {
      "critical": ["security", "performance"],
      "underrepresented": ["collaboration"]
    },
    "suggestedFocus": "This session may contain security patterns - you have NO rules in this area!"
  },
  "extractionFormat": {
    "schema": {"content": "string", "category": "string"},
    "categories": ["debugging", "testing", "architecture", ...],
    "examples": [...]
  },
  "sessionContent": "..."
}

Template features:

  • Topic hints: Categories detected from session content using keyword matching
  • Related rules: Existing playbook rules similar to the session content (via semantic search if enabled)
  • Playbook gaps: Categories with missing or few rules
  • Suggested focus: AI-generated guidance on what to prioritize based on gaps and detected topics
  • Examples: Sample rules showing proper format for each category

Gap-Targeted Sampling Algorithm

When using --fill-gaps, the sampling algorithm prioritizes sessions likely to contain patterns for underrepresented categories:

1. Analyze playbook → Identify critical/underrepresented categories
  1. Generate search queries from category keywords:
- "security" → "security auth token" - "performance" → "performance optimize cache"
  1. Search cass with gap-targeted queries
  2. Score each session against gaps:
- Session matches critical category: +3 points - Session matches underrepresented category: +2 points - Session matches adequate category: +1 point
  1. Sort sessions by gap score (highest first)
  2. Filter out already-processed sessions
  3. Return top N sessions
Category keyword detection uses a lightweight approach (no external APIs required):
// Example: detecting categories from text
const CATEGORY_KEYWORDS = {
  debugging: ["debug", "error", "fix", "bug", "trace", "stack", ...],
  testing: ["test", "mock", "assert", "expect", "jest", "vitest", ...],
  security: ["security", "auth", "token", "encrypt", "permission", ...],
  // ... more categories
};

// Count keyword matches → category with most matches wins

Targeted Sampling Options

Filter sessions by various criteria to focus on specific areas:

# Filter by workspace/project
cm onboard sample --workspace /Users/x/my-project

Filter by agent

cm onboard sample --agent claude cm onboard sample --agent cursor

Filter by recency

cm onboard sample --days 30

Combine filters

cm onboard sample --agent claude --days 14 --workspace /Users/x/api-project

Include already-processed sessions (for re-analysis)

cm onboard sample --include-processed

Adjust sample size

cm onboard sample --limit 20

Onboarding Workflow for Agents

Recommended protocol for AI agents doing onboarding:

## Onboarding Protocol

Phase 1: Assessment

  1. Run cm onboard status --json to check current progress
  2. Run cm onboard gaps --json to see category distribution
  3. Decide target: ~20 rules across diverse categories for a good initial playbook

Phase 2: Session Analysis Loop

For each session until target reached:
  1. cm onboard sample --fill-gaps --json → get prioritized sessions
  2. cm onboard read <path> --template --json → get session with context
  3. Analyze session content, identify 2-5 reusable patterns
  4. Format as rules following extraction guidelines
  5. cm playbook add --file rules.json --session <path> → add rules
  6. Repeat with next session

Phase 3: Verification

  1. cm onboard status → verify progress
  2. cm onboard gaps → check remaining gaps
  3. cm stats --json → verify playbook health

Why Agent-Native Onboarding?

  1. Zero additional cost: Uses the AI agent you're already paying for
  2. Better analysis: Your agent understands your codebase context
  3. Resumable: Progress persists across sessions
  4. Gap-aware: Automatically prioritizes underrepresented categories
  5. Batch-friendly: Add multiple rules efficiently
  6. Self-documenting: Template output guides extraction

📦 Installation

One-Liner (Recommended)

Recommended: Homebrew (macOS/Linux)

brew install dicklesworthstone/tap/cm

This method provides:

  • Automatic updates via brew upgrade
  • Dependency management
  • Easy uninstall via brew uninstall

Windows: Scoop

scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cm

Alternative: Install Script

Linux/macOS:

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/cass_memory_system/main/install.sh?$(date +%s)" \
  | bash -s -- --easy-mode --verify

Direct Downloads:

Installer Options

# Full options
install.sh [--version vX.Y.Z] [--dest DIR] [--system] [--easy-mode] [--verify]
           [--from-source] [--quiet]

Examples

install.sh --version v0.2.2 --verify # Specific version install.sh --system --verify # Install to /usr/local/bin (requires sudo) install.sh --from-source # Build from source (requires bun)

| Flag | Effect | |------|--------| | --version | Install specific version (default: latest) | | --dest DIR | Install to custom directory (default: ~/.local/bin) | | --system | Install to /usr/local/bin (requires sudo) | | --easy-mode | Non-interactive, auto-configure PATH | | --verify | Run self-test after install | | --from-source | Build from source instead of downloading binary | | --quiet | Suppress output |

Installer Robustness

The installer is designed for reliability:

  • No API rate limits: Uses GitHub redirect detection instead of API calls, avoiding the 60 requests/hour unauthenticated limit
  • Checksum verification: Downloads .sha256 files and verifies binary integrity before installation
  • Concurrent install protection: Lock file prevents multiple installers from running simultaneously
  • Graceful fallbacks: If binary download fails, automatically falls back to building from source
  • Cross-platform checksums: Uses sha256sum on Linux, shasum -a 256 on macOS

From Source

git clone https://github.com/Dicklesworthstone/cass_memory_system.git
cd cass_memory_system
bun install
bun run build
sudo mv ./dist/cass-memory /usr/local/bin/cm

Prerequisites

  • cass CLI: The episodic memory layer. Install from cass repo
  • LLM API Key (optional): For AI-powered reflection. Set ANTHROPIC_API_KEY, OPENAI_API_KEY, or GOOGLE_GENERATIVE_AI_API_KEY
Building cass from source requires rustup + nightly. The cass source tree's
.cargo/config.toml unconditionally passes the nightly-only -Z threads rustflag,
and its rust-toolchain.toml pins a nightly channel that only a rustup proxy will
honor. A non-rustup Rust (e.g. Homebrew's cargo/rustc) fails before compilation
with `the option Z is only accepted on the nightly compiler. If you need to
build a cass fix from source (e.g. a commit not yet in a stable release), install
rustup and run rustup toolchain install nightly first.
cm doctor reports whether cass is missing entirely or merely has an unavailable
lexical index (fix the latter with cass index --full).

Verify Installation

cm --version
cm doctor --json

Initial Setup

# Initialize (creates global config and playbook)
cm init

Or with a starter playbook for common patterns

cm init --starter typescript # or: react, python, go cm starters # list available starters

Automating Reflection

The key to the system is automated reflection. Set up a cron job or hook:

# Daily reflection on recent sessions
cm reflect --days 7 --json

Only one project's sessions, only some agents

cm reflect --workspace ~/code/api --agent claude,codex --json

Via cron (runs at 2am daily)

0 2 * /usr/local/bin/cm reflect --days 7 >> ~/.cass-memory/reflect.log 2>&1

For Claude Code, install the auto-reflect hook once and every finished session is reflected in the background:

cm hook install            # this project (.claude/settings.json)
cm hook install --global   # every project (~/.claude/settings.json)
cm hook status             # where it is installed
cm hook uninstall          # remove it (other hooks are left alone)

This adds a SessionEnd hook that runs cm hook session-end. It reads the transcript path from the hook payload and starts cm reflect --session detached, so Claude Code never waits on it. Output goes to ~/.cass-memory/hooks.log. It skips transcripts written by cm's own LLM subprocess calls, and sessions that end inside a background reflect, so it cannot loop. Reflection keeps your usual budget limits and processed-session tracking, so re-running a session costs nothing.

For other agents, call the same entry point from a wrapper script or the agent's own end-of-session hook:

cm hook session-end --transcript ~/.codex/sessions/2026/10/06/rollout.jsonl

Long sessions: reflect often. The diary step reads at most diaryMaxInputChars (default 50,000) characters of a transcript: the start, the end, and error/correction lines picked from the middle (see Reflection Settings). A session that goes on for hours can be much longer than that. cm reflect remembers how far into each session it got, and when a session has grown it reflects only the new part. Running it often (from a hook, or cron every 30–60 minutes) therefore keeps each new part small enough to be read in full. Raising diaryMaxInputChars is the other option when your model has the context for it.

Should agents write their own lessons? They don't need to. cm is built so that a separate reflection step pulls lessons out of the transcript, and the working agent can just finish its task. If your agents do end with a list such as Lessons for memory: ..., that is fine input: the diary step reads it like any other text in the transcript, as long as it falls inside the window above (the end of a session always does). Treat such lists as hints for the diary, not as rules: cm still checks candidate rules against history and feedback before trusting them.


🛠️ CLI Reference

Quick Reference Table

| Command | Purpose | Agent Use | |---------|---------|-----------| | cm context "" --json | Get rules + history for task | Primary | | cm quickstart --json | Self-documentation | Setup | | cm robot-docs [topic] | Machine-readable docs: commands, examples, exit codes, JSON schemas | Integration | | cm doctor --json | System health check | Diagnostics | | cm playbook list | Show all rules | Inspection | | cm similar "" | Find similar rules | Search | | cm stats --json | Playbook metrics | Analytics | | cm trauma list | Show dangerous patterns | Safety | | cm guard --status | Check safety hook status | Safety |

Error Output & Exit Codes

When --json (or --format json) is enabled, errors are printed to stdout as a single JSON object:

{
  "success": false,
  "error": "…",
  "code": "…",
  "exitCode": 2,
  "recovery": ["…", "…"],
  "cause": "…",
  "docs": "README.md#-troubleshooting",
  "hint": "…",
  "details": {}
}

Exit codes are categorized for scripting:

  • 1 internal (bug / unexpected)
  • 2 user input / usage
  • 3 configuration
  • 4 filesystem
  • 5 network
  • 6 cass
  • 7 LLM/provider

Agent Commands (Primary Workflow)

# Get context before starting a task (THE MAIN COMMAND)
cm context "implement user authentication" --json

Self-documenting explanation of the system

cm quickstart --json

Find rules similar to a query

cm similar "error handling best practices"

Machine-readable CLI docs (JSON): all topics, or one of

guide | commands | examples | exit-codes | schemas

cm robot-docs cm robot-docs schemas

cm robot-docs always prints one JSON document to stdout (data.schemaVersion, data.version, data.topics). commands is read from the live CLI definition, so it lists exactly the commands, arguments and options this build accepts. schemas gives JSON Schema (draft-07) for the JSON envelope (success and error) and for the data of context, quickstart and onboard status; the schemas are open, so ignore fields you don't know.

Playbook Commands (Inspect & Manage Rules)

# List all active rules
cm playbook list

Get detailed info for a specific rule

cm playbook get b-8f3a2c

Add a new rule manually

cm playbook add "Always run tests before committing"

Remove/deprecate a rule

cm playbook remove b-xyz789 --reason "No longer applicable"

Export playbook for backup or sharing

cm playbook export > playbook-backup.yaml

Import playbook from file. Rules matching an existing one by text (exact,

or near-identical wording) merge their feedback into it instead of becoming

duplicates; contradictions with existing rules are reported.

cm playbook import shared-playbook.yaml

Find rules that contradict each other (with a keep/retire suggestion)

cm playbook conflicts --json

Strip feedback that came from cm's own LLM subprocess transcripts (or any

session path pattern), recount helpful/harmful, re-derive maturity.

Backs up each playbook first; bullets left without genuine support are

reported, and deprecated only with --deprecate-orphans.

cm playbook scrub --cm-subprocess-calls --dry-run --json cm playbook scrub --from-sessions "/old-project/" --deprecate-orphans

Show top N most effective rules

cm top 10

Find rules without recent feedback (stale)

cm stale --days 60

Show why a rule exists (provenance)

cm why b-8f3a2c

Get playbook health metrics

cm stats --json

Learning Commands (Feedback & Reflection)

# Process recent sessions into rules (usually automated)
cm reflect --days 7 --json

Manual feedback on a rule

cm mark b-8f3a2c --helpful cm mark b-xyz789 --harmful --reason "Caused regression"

Record session outcome (positional: status, rules)

cm outcome success b-8f3a2c,b-def456 cm outcome failure b-x7k9p1 --summary "Auth approach failed"

Apply recorded outcomes to playbook

cm outcome-apply

Validate a proposed rule against history

cm validate "Always check null before dereferencing"

Deprecate a rule permanently

cm forget b-xyz789 --reason "Superseded by better pattern"

Revert a curation decision

cm undo b-xyz789

Check sessions for rule violations

cm audit --days 30

System Commands (Setup & Diagnostics)

# Initialize configuration and playbook
cm init
cm init --starter typescript  # With starter template
cm starters                   # List available starters

Check system health

cm doctor --json cm doctor --fix # Auto-fix issues

Check for / install a newer release

cm update --check # Exit code 0 = up to date, 1 = update available cm update # Re-runs the documented install script (interactive only)

Export rules for AGENTS.md

cm project --format agents.md

Show LLM cost and usage statistics

cm usage

Run MCP HTTP server for agent integration (default: 127.0.0.1:8765)

cm serve

Privacy controls

cm privacy status cm privacy enable # Enable cross-agent enrichment cm privacy disable # Disable cross-agent enrichment

Safety Commands (Trauma Guard)

# View active trauma patterns
cm trauma list

Add a dangerous pattern

cm trauma add "DROP TABLE" --description "Database table deletion" --severity critical

Temporarily disable a pattern

cm trauma heal t-abc123 --reason "Intentional migration cleanup"

Permanently remove a pattern

cm trauma remove t-abc123

Scan sessions for potential traumas

cm trauma scan --days 30

Import patterns from file

cm trauma import shared-traumas.yaml

Install safety hooks

cm guard --install # Claude Code pre-tool hook cm guard --git # Git pre-commit hook cm guard --install --git # Both cm guard --status # Check installation status

Command Output Modes

All commands support multiple output formats:

| Flag | Output | |------|--------| | (default) | Human-readable text | | --json | Machine-readable JSON | | --quiet | Minimal output |


🔬 The ACE Pipeline

cass-memory implements the Agentic Context Engineering (ACE) framework with four stages:

┌──────────────┐     ┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│  GENERATOR   │ ──▶ │  REFLECTOR   │ ──▶ │  VALIDATOR   │ ──▶ │   CURATOR    │
│              │     │              │     │              │     │              │
│ Pre-task     │     │ LLM pattern  │     │ Evidence     │     │ Deterministic│
│ context      │     │ extraction   │     │ gate against │     │ delta merge  │
│ hydration    │     │ from diary   │     │ cass history │     │ (NO LLM!)    │
└──────────────┘     └──────────────┘     └──────────────┘     └──────────────┘

Stage 1: Generator (cm context)

Hydrates the agent with relevant rules before starting a task.

Implementation (src/commands/context.ts):

async function generateContext(task: string): Promise<ContextResult> {
  // 1. Extract keywords from task
  const keywords = extractKeywords(task);

// 2. Relevance: BM25 over the playbook (IDF-weighted share of the query a // bullet covers, 0..10), blended with embedding similarity when enabled const lexical = scoreLexicalRelevance(bullets, keywords); const scored = bullets.map(b => { const relevanceScore = blend(lexical.get(b.id), cosine(taskEmbedding, b.embedding) * 10); // 3. Track record only reorders within a bounded band: [1 - w, 1 + w], // damped while a bullet has few marks (feedbackWeight, default 0.25) return { ...b, relevanceScore, finalScore: relevanceScore * getFeedbackMultiplier(b) }; }).sort(byFinalScore);

// 4. Keep what is relevant and fits: absolute + relative relevance floors, // maxBulletsInContext, then contextTokenBudget const { selected, stats } = selectContextBullets(scored, config);

// 5. Search cass for historical context const historySnippets = await safeCassSearch(task, { limit: config.maxHistoryInContext, days: config.sessionLookbackDays });

return { task, relevantBullets: selected.filter(isRule).map(toContextBullet), // compact: no event log antiPatterns: selected.filter(isAntiPattern).map(toContextBullet), historySnippets, retrieval: stats, // candidates / returned / dropped by relevance, limit, budget suggestedCassQueries: generateSuggestedQueries(task, keywords) }; }

Bullets in cm context output are a compact projection (id, content, category, kind, maturity, tags, counts, scores, reasoning). The full record, including the feedback-event log and source sessions, is one cm playbook get away.

Stage 2: Reflector (cm reflect)

Processes unprocessed sessions to extract new rules.

How It Works:

  1. Find unprocessed sessions via cass timeline API
  2. Generate diary entries - structured summaries of each session:
- Accomplishments (what was completed) - Decisions (design choices made) - Challenges (problems encountered) - Key learnings (reusable insights)
  1. Run LLM reflector - multi-iteration pattern extraction
  2. Output deltas - add/helpful/harmful/deprecate operations
The Reflector Prompt (simplified):
Analyze this session diary and extract actionable rules.
For each pattern you identify:
  1. Formulate as an imperative rule ("Always...", "Never...", "When X, do Y")
  2. Categorize (testing, debugging, git, architecture, etc.)
  3. Assess scope (global, workspace, language, framework)
  4. Provide reasoning for why this is valuable
Output format: JSON array of proposed rules

Stage 3: Validator

Scientific gate ensuring rules have evidence.

Implementation (src/ace/validator.ts):

async function validateRule(rule: ProposedRule): Promise<ValidationResult> {
  // 1. Evidence count gate - fast pre-filter
  const evidenceCount = await countEvidenceInCass(rule);
  if (evidenceCount < MIN_EVIDENCE_COUNT) {
    return { verdict: 'NEEDS_MORE_EVIDENCE', evidenceCount };
  }

// 2. LLM validation - deep analysis const llmVerdict = await llmValidate(rule, evidenceCount);

// 3. Final verdict return { verdict: llmVerdict.valid ? 'ACCEPT' : 'REJECT', reasoning: llmVerdict.reasoning, evidenceCount }; }

Stage 4: Curator

Critical: NO LLM in this stage - pure deterministic logic to prevent context collapse.

Why No LLM? LLMs tend to "drift" when iteratively refining their own outputs. By keeping curation deterministic, we:

  • Prevent feedback loops that degrade rule quality
  • Ensure reproducible behavior
  • Avoid expensive redundant LLM calls
Curation Operations:

function curate(playbook: Playbook, deltas: Delta[]): Playbook {
  for (const delta of deltas) {
    switch (delta.type) {
      case 'add':
        // Conflict detection - find contradicting rules
        const conflicts = findConflicts(playbook, delta.rule);
        if (conflicts.length > 0) {
          resolveConflicts(playbook, conflicts, delta);
        }

// Duplicate detection - hash-based and semantic const duplicate = findDuplicate(playbook, delta.rule); if (duplicate) { mergeDuplicate(duplicate, delta.rule); continue; }

// Add new rule playbook.bullets.push(delta.rule); break;

case 'helpful': // Update feedback addFeedback(playbook, delta.bulletId, 'helpful', delta.timestamp); maybePromoteMaturity(playbook, delta.bulletId); break;

case 'harmful': // Update feedback with harmful multiplier consideration addFeedback(playbook, delta.bulletId, 'harmful', delta.timestamp); maybeConvertToAntiPattern(playbook, delta.bulletId); break;

case 'deprecate': deprecateRule(playbook, delta.bulletId, delta.reason); break; } }

// Maturity promotions based on feedback thresholds promoteQualifiedRules(playbook);

return playbook; }


📊 Data Models

Playbook Bullet

The core data structure for rules:

interface PlaybookBullet {
  // Identity
  id: string;                    // "b-{timestamp}-{random}"
  content: string;               // The actual rule
  category: string;              // "testing" | "git" | "debugging" | etc.

// Classification scope: "global" | "workspace" | "language" | "framework" | "task"; type: "rule" | "anti-pattern"; kind: "project_convention" | "stack_pattern" | "workflow_rule" | "anti_pattern"; isNegative: boolean;

// Lifecycle state: "draft" | "active" | "retired"; maturity: "candidate" | "established" | "proven" | "deprecated";

// Feedback & Scoring helpfulCount: number; // Raw count harmfulCount: number; // Raw count feedbackEvents: FeedbackEvent[]; // Full history with timestamps effectiveScore?: number; // Calculated decay-adjusted score

// Provenance sourceSessions: string[]; // Sessions that generated this rule sourceAgents: string[]; // Agents involved (claude, codex, cursor, etc.) reasoning?: string; // Why this rule exists sourceSession?: string; // Flexible field for reflection metadata

// Metadata tags: string[]; embedding?: number[]; // Semantic search vector (768 dimensions) pinned: boolean; // Prevent auto-deprecation deprecated: boolean;

// Timestamps createdAt: string; // ISO 8601 updatedAt: string; // ISO 8601 }

Feedback Event

Immutable record of feedback:

interface FeedbackEvent {
  id: string;                    // UUID
  type: "helpful" | "harmful";
  timestamp: string;             // ISO 8601
  source: "inline" | "manual" | "outcome" | "audit";
  sessionPath?: string;          // Session where feedback was given
  reason?: string;               // Optional explanation
}

Diary Entry

Working memory structure:

interface DiaryEntry {
  id: string;                    // Deterministic hash of content
  sessionPath: string;           // Original cass session
  timestamp: string;             // ISO 8601
  agent: string;                 // "claude" | "codex" | "cursor" | etc.
  workspace?: string;            // Project directory

// Session status status: "success" | "failure" | "mixed";

// Core content accomplishments: string[]; // What was completed decisions: string[]; // Design choices made challenges: string[]; // Problems encountered keyLearnings: string[]; // Reusable insights preferences: string[]; // User style revelations

// Cross-agent enrichment relatedSessions?: RelatedSession[];

// Search optimization searchAnchors: string[]; // Keywords for retrieval tags: string[]; // File/component tags }

Session Outcome

Explicit outcome recording:

interface SessionOutcome {
  id: string;                    // UUID
  timestamp: string;             // ISO 8601
  status: "success" | "failure" | "mixed";

// Associated rules rulesUsed: string[]; // Bullet IDs that were followed rulesViolated?: string[]; // Bullet IDs that were ignored

// Context task?: string; // What was attempted summary?: string; // Brief description sessionPath?: string; // Source session

// Feedback impact applied: boolean; // Whether outcome-apply has processed this appliedAt?: string; // When it was applied }


📈 Scoring Algorithm

Effective Score Calculation

The effective score determines rule ranking, maturity transitions, and anti-pattern conversion:

function getEffectiveScore(bullet: PlaybookBullet): number {
  const HARMFUL_MULTIPLIER = 4;    // One harmful = 4× one helpful
  const HALF_LIFE_DAYS = 90;       // 90 days for half decay

// Calculate decayed helpful count const decayedHelpful = bullet.feedbackEvents .filter(e => e.type === "helpful") .reduce((sum, event) => { const daysAgo = daysSince(event.timestamp); const decayFactor = Math.pow(0.5, daysAgo / HALF_LIFE_DAYS); return sum + decayFactor; }, 0);

// Calculate decayed harmful count const decayedHarmful = bullet.feedbackEvents .filter(e => e.type === "harmful") .reduce((sum, event) => { const daysAgo = daysSince(event.timestamp); const decayFactor = Math.pow(0.5, daysAgo / HALF_LIFE_DAYS); return sum + decayFactor; }, 0);

// Effective score with harmful multiplier return decayedHelpful - (HARMFUL_MULTIPLIER * decayedHarmful); }

Score Decay Visualization

Initial score: 10.0 (10 helpful marks today)

After 90 days (half-life): 5.0 After 180 days: 2.5 After 270 days: 1.25 After 365 days: 0.78

Why Time Decay?

  1. Codebases evolve - A pattern helpful for React 17 may not apply to React 19
  2. Tools change - Agent capabilities improve over time
  3. Prevents stale rules - Forces ongoing validation
  4. Natural cleanup - Unvalidated rules fade rather than accumulate

Maturity State Machine

`` ┌──────────┐ ┌─────────────┐ ┌────────┐ │ candidate│──────▶│ established │───▶│ proven │ └──────────┘ └─────────────┘ └────────┘ │ │ │ │ │ (harmful >25%) │ │ ▼ │ │ ┌─────────────┐ │

... (README truncated for length)

Chat with me