Athena
Local-first memory and guardrails for AI agents in your IDE.
Your context lives in plain Markdown on your disk, works across Claude Code, Antigravity, Cursor, Gemini CLI and VS Code — and the rules are enforced by hooks, not just prompts.
!20-second demo: /start recalls last session → work → /end
Quickstart · How It Works · Docs · Why Athena? · Safety
The Problem
You've spent months training ChatGPT to understand you. Then a model update resets the personality. You switch to Claude or Gemini — you start from zero.
Platform memory is unreliable, opaque, and locked to one provider. You don't own it and you can't take it with you.
Athena moves the memory layer to your machine: plain Markdown files that you own, version-control, and point at any model. The model is just whoever's on shift.
Quickstart
git clone https://github.com/winstonkoh87/Athena-Public.git && cd Athena-Public
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[local]" # lightweight — no cloud deps
athena init --ide claude # or: antigravity, cursor, gemini, vscode, kilocode, roocode
athena doctor # expect 0 failures
Then type /start in your IDE's AI chat panel. Work normally. Type /end to save.
Full install (cloud sync + reranking): pip install -e ".[full]"
See Getting Started for Windows, advanced config, and Supabase setup.
How It Works
┌─────────────────────────────────────────────────────┐
│ Your IDE (Claude Code / Antigravity / Cursor / …) │
└───────────────────┬─────────────────────────────────┘
│ hooks (code-enforced, not prompt-based)
▼
┌─────────────────────────────────────────────────────┐
│ Athena SDK │
│ ├── Lifecycle: /start loads ~2K tokens, /end saves │
│ ├── Memory: session logs + canonical facts │
│ ├── Search: hybrid (keyword + optional vectors) │
│ └── Guardrails: ruin check, secret scan, grounding │
└───────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────┐
│ Your Disk (plain Markdown — git-versioned) │
│ └── Optional: Supabase pgvector for cloud sync │
└─────────────────────────────────────────────────────┘
Three modes, one loop:
| Mode | Boot | Token Cost | When |
|:-----|:-----|:-----------|:-----|
| Lightweight | Just chat, then /end | ~500 | Quick questions |
| Standard | /start → work → /end | ~2K–10K | Daily use |
| Deep | /ultrastart → work → /ultraend | ~20K | Complex planning |
What's enforced in code vs. by prompt
Most AI-agent READMEs state every claim in the same confident voice. This one doesn't.
| Claim | Status | Evidence | |:------|:-------|:---------| | Storage & retrieval — memories stored and surfaced when relevant | ✅ Shipped | Hybrid RAG with cross-encoder rerank, hardened through production failures | | Portability — Markdown on your disk, movable across models | ✅ Shipped | Structural — inspect the repo | | Governed autonomy — hooks block destructive commands and secrets | ✅ Shipped (partial) | Ruin check blocks 14/16 destructive commands; 2 bypasses are known and tracked | | Compounding personalization — session 500 recalls session 5 | 🟡 N=1 evidence | 1,900+ sessions by the author; no multi-user study | | Anti-sycophancy — personalization doesn't silently increase agreement | 🟡 Partial mitigation | Code-enforced meta-awareness gate (Claude Code only); see honest limits |
Why publish this table? Because the failure mode of this product category is self-mythologizing — describing aspirations in the present tense. Athena's own convention (Epistemic Status) requires labeling every mechanism ascode-enforced,agent-discretion, oraspirational. This table is that convention applied to the README.
Measured, Not Claimed
# Run the tests yourself
pytest tests/ -v --tb=short
Run the retrieval evaluator yourself (requires Supabase keys)
python examples/scripts/evaluator.py --gold-set .agent/eval/gold_set.json
| Metric | Value | Verification Command |
|:-------|:------|:---------------------|
| Retrieval Hit@5 (Strict) | 0.569 (37 / 65) | python examples/scripts/evaluator.py |
| Retrieval MRR@5 (Strict) | 0.472 | python examples/scripts/evaluator.py |
| Retrieval Hit@5 (Lenient) | 0.892 (deprecated) | Partial substring match (inflated) |
| Unit & Integration Tests | 558 passed (100%) | pytest tests/ |
| Secret Leaks (1,248 commits) | 0 detected | Gitleaks in CI |
| Code Quality & Lints | 0 ruff findings | ruff check src/ |
Anti-Goodhart Invariant: Why did our reported Hit@5 shift from 0.89 to 0.57? Lenient substring matchers count partial word overlaps as "hits," inflating benchmark scores by ~36% without improving retrieval. We killed the lenient matcher because vanity metrics mask regressions. See the full breakdown: Anti-Goodhart Benchmarking in RAG.
Agent Compatibility
| IDE | Config file | Tested |
|:----|:------------|:-------|
| Claude Code | CLAUDE.md | ✅ |
| Antigravity | AGENTS.md | ✅ |
| Cursor | .cursor/rules.md | ✅ |
| Gemini CLI | .gemini/AGENTS.md | ✅ |
| VS Code + Copilot | .vscode/settings.json | ✅ |
| Kilo Code | .kilocode/rules/athena.md | ✅ |
| Roo Code | .roo/rules/athena.md | ✅ |
Documentation
| Doc | What it covers | |:----|:---------------| | Getting Started | Install, configure, first session | | Your First Session | 20-minute guided tutorial | | How It Works | Architecture and search pipeline | | Why Athena? | Philosophy, use cases, cost analysis | | Engineering Depth | Technical deep dives | | CLI Reference | All commands | | FAQ | Common questions | | Safety | Limitations and responsible use | | Changelog | Release history |
Contributing
PRs welcome. Please read CONTRIBUTING.md first.
The codebase uses ruff for linting, pytest for tests, and CI must pass before merge.
MIT License · Contributing · Safety · Security · Code of Conduct
Built and battle-tested solo across 1,900+ sessions by Winston Koh.
Clone it. Boot it. Make it yours.