Aura
DEMO: https://youtu.be/iTyxeugcZtI?si=B91No0Hjz3eKLMwz
Aura is a local AI research project that runs entirely on your Mac. It explores how to build an AI agent that can think continuously, remember things long-term, govern its own behavior with auditable receipts, steer its own emotional tone during text generation, and keep itself running reliably over time.
Aura is not proof of consciousness, life, or personhood. Nothing here settles those questions. Some file and module names sound like they might — they're named after the mechanisms they implement, not achievements they've proven.
The actual claim is narrower and testable: Aura's internal state measurably affects what it generates, what it remembers, what tools it's allowed to use, what tasks it picks up on its own, and how it repairs itself — through code paths that leave audit trails you can inspect.
That's a smaller claim than the vocabulary suggests. It's also one you can verify.
!Python 3.12+
!Platform: macOS Apple Silicon
For the full technical deep dive, read ARCHITECTURE.md. For the same ideas explained without math, read HOW_IT_WORKS.md. For the standard of evidence we hold ourselves to on autonomy and novel output claims, read docs/BEHAVIORAL_PROOF_STANDARD.md.
**The main research program is docs/RECURSIVE_LATENT_CORTEX.md.** It asks: can a frozen 32-billion-parameter model be made to reason deeper without changing any of its stored weights? It builds the machinery to test that, runs a pre-registered experiment, and reports that the hoped-for capability improvement didn't appear. Summary in Recursive Latent Cortex below.
Evidence map: Claims should point to runnable tests, proof bundles, receipts,
or replayable logs. Test counts change with the repo; use pytest --collect-only,
make proof-bundle, and TESTING.md for current numbers rather
than trusting prose.
If you want to see it work, keep reading.
Evidence boundary
This is a working AI research project. It is not proof of consciousness, subjective experience, legal personhood, or moral standing.
The repo enforces that in code, not just in words. A boundary guard treats
loaded terms — "consciousness guarantee," "personhood proof" — as names for
test batteries rather than claims, unless independent evidence says otherwise.
A module named qualia_synthesizer.py is a name. Names are not evidence.
What is actually claimed, and what each claim costs:
- Governance is a design goal, not a guaranteed fact. Important actions
- Autonomous self-improvement is not proven mature. There's scaffolding
- The cognitive layer hasn't been shown to earn its cost. This is the
CLAIMS_MATRIX.md claim 31.
- Production maturity is limited. This is research software being
- Emotional steering is causal, and that's testable. Internal state
- Identity persistence uses retrieval, not just prompt text. Coherence
- φ (phi) is bounded. A bounded integration metric inspired by
- The hardware target is specific. Bryan's Apple Silicon M5-class Mac,
- Resource costs persist and constrain what Aura can do. That's a
For the things deliberately not claimed — including anything physical — read CLAIMS_NOT_SUPPORTED.md. It's the most useful page here if you're skeptical, which you should be.
See it learn (one command)
Don't take the evidence claims on faith. Run it. Apple Silicon, 20–40 minutes, about 5 GB of disk:
make setup # once: venv + requirements
make demo-learning
Here's what happens. A small local model attempts seeded reasoning tasks. Every attempt is graded by an exact checker — the verifier is the reward signal, so there's nothing to game. It then trains a LoRA adapter (a small set of learnable weights) on the verified wins and losses, and the adapter has to pass a sealed held-out test set on fresh inputs it's never seen before its weights get merged and published.
Then it does the whole thing again, building on what it just published.
Every generation gets recorded in a hash-chained ledger
(core/learning/rsi_lineage.py). The verdict is computed from those
receipts, not written by hand afterward. When it refuses to promote, that
prints just as loudly as a gain. Raw responses, evaluation reports, and
cycle receipts all stay on disk.
The same machinery runs autonomously inside the live runtime
(core/learning/compounding_scheduler.py): gated by idle time,
governance-approved, memory-controlled, with promoted weights hot-swapped
into live text generation.
Production Evidence Surface
Everything below has a working implementation, receipts, and validation artifacts. Ideas that don't are left off this list — not softened, not hedged, just excluded until they earn a place.
Release gates generate a proof bundle. You shouldn't have to guess at maturity from how confident the prose sounds.
core/brain/llm/continuous_substrate.pyis a configurable 64-to-512 neuron
get_state_summary() derives valence/arousal/dominance/phi from the live
state vector (grounded in external data via adapt_projections()),
without changing callers.
core/brain/llm/substrate_token_generator.pyis the substrate-first readout:
core/brain/llm/sensorimotor_grounding.pymaps camera/screen/audio
core/consciousness/phi_core.pyimplements real IIT-style integration
core/consciousness/hierarchical_phi.pyimplements the 32-node hierarchical
core/consciousness/affective_steering.pyis a real activation-steering
training/caa_32b_validation.pyvalidates production-model steering
core/consciousness/stdp_external_validation.pyruns the external-usefulness
core/self_modification/fault_pipeline.pyand
core/self_modification/repair_approval.py implement the closed-loop
bug-repair path with deterministic localization, tier-aware approval,
patch tracking, and calibration.
core/architect/implements the Autonomous Architecture Governor: a
docs/AUTONOMOUS_ARCHITECTURE_GOVERNOR.md.
core/runtime/autonomy_conductor.pyandcore/runtime/activation_audit.py
core/runtime/overt_action_loop.pyis the practical "what does she do?"
core/adaptation/online_lora_governor.pyconnects Will-approved
mlx_lm lora process is active, so long training runs are preserved.
core/goals/default_goals.pyseeds durable, tool-attached IN_PROGRESS goals
- The full memory architecture (episodic, semantic, vector, knowledge graph,
core/memory/sqlite_vector_store.py, not as plaintext JSON arrays.
Evidence boundaries on the production parts:
- φ is computed over **cognitive-emotional state nodes and sampled mesh
- Steering credit requires
CAA_32B_RESULTS.json: steered behavior on the
- Plasticity credit requires
STDP_EXTERNAL_VALIDATION.json:
Test attestation: make proof-bundle writes the current evidence bundle:
DECISIVE_RESULTS.json, CAA_32B_RESULTS.json,
STDP_EXTERNAL_VALIDATION.json, GOVERNANCE_COVERAGE.json,
SELF_REPAIR_LINEAGE.json, LONGEVITY_RUN.json,
MUTATION_TEST_REPORT.json, BOOT_HEALTH.json, ACTIVATION_REPORT.json,
SECURITY_SCAN.json, and CANONICAL_PROOF_BUNDLE.json.
Recursive Latent Cortex
Full page: docs/RECURSIVE_LATENT_CORTEX.md.
The question: a frozen 32B model is a fixed-depth pipeline — 64 layers, applied once per token. **Can you make it think longer on a hard problem without changing any stored weights?**
Two approaches have been tried, and they gave different answers.
Approach 1: Frozen Loop (didn't work)
The first approach was a frozen loop: thought slots placed beside the prompt, a window of middle layers run over them repeatedly on a schedule, with the refined slots' key-value states persisted so every generated token can attend to them. Model weights are hash-checked before and after every episode to prove nothing changed; episode-scoped fast weights are provably erased; equal-compute accounting is built in so "more compute helped" can't be mistaken for "the architecture helped."
The pre-registered experiment — seed committed before any task was generated, n=24 per family, with statistical corrections — refuted it. *On a model that wasn't trained for recurrence at this scale, the frozen loop doesn't just fail to help — it hurts.*
The explanation: answer tokens only ever traversed the middle block once, so no extra depth was applied to the actual answer computation. Only the scratchpad was recurring.
Approach 2: Trained Intrinsic Recurrence (promising, bounded)
The second approach, trained intrinsic recurrence (docs/INTRINSIC_RECURRENCE.md), makes the real token stream re-enter the middle block, so a 64-layer model effectively runs 160 layers deep at T=4 with the same weights, and trains it on typed, exactly-checkable program traces instead of answers. This training path is separate from the typed semantic-machine evidence below. CP566 measured a learned semantic machine whose terminal state conditioned the model's answer; it did not isolate a gain from repeated middle-layer passes in the resident model.
The scoped computation layer connects unit-checked equation graphs, logical consequences, candidate constraints and retained recipes. Its calculations remain conditional on supplied premises; autonomous source grounding and general-transfer gains are still unproved.
The frozen RLC baseline records the evidence available on 2026-09-08 and its limitations. The fresh v16 natural-language replication reached 26/96 exact answers, below its pre-registered 48/96 floor; ordinary generation was not run after that futility stop. The v19 repair reached 93/96 on the exposed development set, with coefficient lesion 0/96. That development result still needs fresh replication. Neither record establishes broad reasoning gain or grants serving authority.
Results summary
| What was tested | Verdict |
|---|---|
| Mechanics (KV rewind, stability bounds, slot ablation moving the answer distribution, erasure, invariants) | PROVEN on real MLX weights |
| Live runtime integration on the resident 32B | PROVEN |
| Capability gain, frozen loop, 1.5B | REFUTED — vanilla 21/72 beat every one of 7 latent arms (7–13/72) |
| Capability gain, frozen loop, 32B | CONJECTURE, negative point estimate — latent 0.375 vs vanilla 0.417, overlapping intervals |
| Capability gain, typed semantic machine + state-conditioned decode, 32B | BOUNDED_WOW_SIGNAL — 60/60 against 16/60 for ordinary decode, ablation-dependent, p = 5.7 × 10⁻¹⁴ |
| Cross-generation recovery, trained semantic tissue, 27B | BOUNDED_WOW_SIGNAL — 60/60 against 0/60 ordinary decode on a separate fresh set; wire 6, coefficient lesion 4, wrong-state 0; p = 8.67 × 10⁻¹⁹ |
| Family-blind procedure learning into neural tissue | SUPPORTED, BOUNDED — one depth-2 procedure learned from 16 examples, then 96/96 exact on fresh inputs; coefficient and wrong-input controls failed 96/96, no-procedure solved 1/96, shuffled-output nulls found 0/15 |
| Resident decode of the learned neural procedure | SUPPORTED, BOUNDED — treatment 8/8, ordinary 1/8, wire 1/8, coefficient lesion 1/8, wrong-input 0/8, wrong-state 0/8; seven gains, no regressions, p = 0.0078125 |
| Resident 27B language-to-program transfer | SUPPORTED, BOUNDED — exact execution emitted 134/256 held-out answers from learned model-bound semantics; exact program recovery was 133/256, against hidden-state shuffle 14/256, coefficient lesion 0/256 and label permutation 4/256 |
| Frozen fresh-cohort semantic transfer | SUPPORTED, REPLICATED, BOUNDED — the unchanged transducer emitted 114/256 exact held-out answers on a separately seeded numeric set after a clean worker restart, against hidden-state shuffle 10/256 and coefficient lesion 0/256 |
| Shared variable-geometry semantic programs | SUPPORTED, BOUNDED — one transducer with no family router recovered 258/368 complete programs and exact execution emitted 292/368 answers across arithmetic, sequence and fork/join geometries; hidden-token shuffle 0/368 and coefficient lesion 0/368, paired exact p = 2.16 × 10⁻⁷⁸ |
| Learned programs on the universal floor | SUPPORTED, BOUNDED — all 368/368 accepted frozen test programs had identical outcomes under the existing exact executor and Aura's universal metered floor: 366 matching values and two matching typed refusals across all three families, with 20/20 primitive semantics covered |
| Broad reasoning gain, fusion, frontier performance | NOT CLAIMED |
The current native semantic-selection work uses the resident model's own language pathway and role-relative program registers. Source-selected adaptation reaches 46/50 on an exposed proposal bank against incumbent 41/50 and unfitted 39/50, with one regression. A matched fitting-source erasure control gets the same 46 correct outcomes, so that gain cannot be attributed to newly learned source meaning. Target-blind decoding reached 72/72 on an easier fixed-template development set, but the wider retained canary returns only 10/14 procedure matches and 11/14 provably equivalent answers. None is a promoted candidate or broad/frontier reasoning result.
BOUNDED_WOW_SIGNAL is the adjudicator's own verdict string, and bounded is
the key word: the limitations ship inside the same receipt as the verdict.
The four-domain experiment in detail: On a frozen set of 60 typed tasks — coding, calibration, misleading premise, scientific inference — the semantic-machine treatment answered 60/60 exactly against 16/60 for ordinary decode, with a matched wire base at 7, a coefficient lesion at 5, and a wrong-state control at 0. Forty-four ordinary failures converted, none regressed, paired one-sided exact p = 5.7 × 10⁻¹⁴. The coefficient lesion reduces the result, supporting dependence on those learned coefficients within this experiment. State conditioning, formatting assistance and retries remain part of the recorded treatment contract; equal-compute broad reasoning comparisons remain a separate requirement.
The 27B cross-generation replication: The 2026-08-24 cortex migration
repeated that bounded claim on the fused Qwen3.8-27B resident model. A
separately seeded 60-task, 300-decode campaign returned treatment 60/60,
ordinary decode 0/60, matched wire 6/60, coefficient lesion 4/60, and
wrong-state 0/60, with no regressions and exact one-sided p = 8.67 × 10⁻¹⁹.
Independent verification replayed all 300 journal rows before the frozen
adjudicator returned BOUNDED_WOW_SIGNAL again. This shows the bounded
typed-tissue/executor mechanism works across two model generations; it is
not a head-to-head 27B vs 32B quality benchmark because the test sets and
models differ. The ordinary 27B arm produced no parseable answers under that
campaign's decode contract. Its 0/60 therefore cannot establish lack of
knowledge or performance under a different completion budget.
Family-blind procedure learning: A generic enumerative inducer received
sixteen input-output examples, no family label and no family solver, froze
idiv(add(in0, in1), in2), and passed it through a family-blind SSA lowerer
into the existing learned arithmetic tissue. That path was exact on 96/96
fresh inputs. Coefficient and guaranteed-wrong-input ablations disrupted all
96, the no-procedure control solved 1/96, no depth-one shortcut fit, and
fifteen shuffled-output searches found no program. This establishes bounded
procedure learning and neural execution over a fixed set of primitive
operations.
A second frozen canary carried that same learned program through the fused resident 27B's answer surface. Treatment was 8/8 exact against 1/8 ordinary decode; syntax-only wire and coefficient lesion were also 1/8, while wrong-input and wrong-state controls were 0/8. Seven ordinary failures converted with no regressions, exact paired one-sided p = 0.0078125. Independent replay reconstructed all 48 decodes and the 50-event journal. This does not establish natural-language compilation, open-domain reasoning, unrestricted serving, static fusion or frontier performance.
Language-to-program transfer: A generic linear transducer learned token spans, primitive operations and register arguments from five construction families without expected answers. On four held-out construction combinations, exact execution emitted 134/256 correct answers and recovered 133/256 complete programs. Hidden-token shuffle reached 14/256, coefficient lesion 0/256 and label permutation 4/256. An independent, source-bound replay reloaded all 576 feature records, reproduced the coefficients and report exactly, and recounted all 1,344 task-arm rows. The result is bounded to the synthetic arithmetic grammar; serving and broad-domain transfer remain open. The campaign was re-verified after tooling changes; the first certificate remains the immutable historical record, and the current claim reads the certificate bound to the exact measured commit.
Fresh-cohort replication: The coefficients were frozen and evaluated on a separately seeded set with no task overlap. A clean worker restart changed its process identity but left the complete neural-function basis identical. Without fitting or refitting, exact execution emitted 114/256 held-out answers; hidden-state shuffle reached 10/256 and coefficient lesion 0/256. Independent verification reloaded both 576-record bundles, replayed the frozen result exactly, and recounted 1,728 task-arm rows. This establishes fresh-example reuse inside the same bounded language and primitive system. It does not establish new procedure families, open-domain transfer, or serving authority.
Variable-geometry programs: One shared transducer now receives tokenizer-grounded typed inputs and learns the remaining program structure from resident 27B hidden states. It predicts how many steps, which primitive each step applies, and whether each argument refers to an input or an earlier result. The same coefficient set handles two-step arithmetic, two-step sequence programs, and three-step fork/join arithmetic without a family router. On held-out constructions it recovered 258/368 complete programs, and exact execution emitted 292/368 correct answers. Hidden-token shuffle and coefficient lesion each recovered 0/368 programs; both paired exact tests gave p = 2.16 × 10⁻⁷⁸. This establishes learned variable-geometry programs over the declared typed vocabulary. A new schema still needs support examples, and this result grants neither serving authority nor a broad natural-language claim.
Universal floor execution: Those learned programs no longer stop at a separate Python operation table. Every declared integer and sequence primitive now compiles into Aura's universal metered floor, along with the learned SSA references and typed public inputs. A separate no-refit replay compared both engines on all 368 accepted frozen test programs: 366 produced the same value and two produced the same typed refusal, for 368/368 agreement across arithmetic, sequence and fork/join. All 20 declared primitives have both floor semantics and a type signature; a new primitive without either is refused. The neural front end and endogenous substrate now share execution semantics. Later v14 evidence recovered 79/96 programs and answers on a fresh, fit-withheld synthetic program family using shared primitives. Its endogenous replay also reached 79/96, versus 0/96 under coefficient lesion. Open-domain schema learning remains unproven.
One family — misleading premise — gained nothing, and the reason is worth stating rather than averaging away: ordinary decode was already at ceiling there (15/15), and the controller preserved all fifteen instead of manufacturing a gain by regressing its own baseline. The other three families supplied the 44.
It has dated serving-path qualification. semantic_neural_serving.py refuses
to serve unless a descriptor-bound activation record says active_by_default.
The CP1011 27B package is rlc-27b-recovery-05346acd618d1c925f16; its recorded
runtime verification is 120/120 exact, 120/120 ablation-disrupted, 120/120
through both foreground and service integrations, and unsupported language
refused, at a 9.229 / 38.696 ms median / maximum. These are qualification
measurements, not a live-status probe. Present activation must be checked
against the running process, source revision, model identity and
eligible-request receipt.
It is still not a broad reasoning gain, not static fusion, not frontier performance, and it still cannot answer ordinary chat — admission is decided by an answer-blind parser over the task grammar, and unsupported language never reaches the lane.
Before you read either page: the two negative results from the August reconciliation campaign — a 13-vs-5 and its 9-vs-4 reproduction — were void, because the promotion gate had been wired to the one decode policy that removes the vanilla floor. A win had been structurally impossible there, and those two runs measured a system that was never switched on. The July pre-registration above is untouched by that defect and its verdicts stand. docs/RLC_RECONCILIATION.md has the fourteen defects in dependency order.
Language substrate and generality
The question: can Aura learn a new way of representing knowledge from experience, invent a reusable abstraction, apply it in an unrelated domain, and get better at doing this over time?
Relation induction — core/cognition/relation_language.py learns
structured transformation rules from paired observations. Validated on
held-out transitions; must compress (substitution tables refused, noise
invents nothing). Works across words, colours, records, and grids. Composition
("mirror then rotate") reaches 20 of 120 battery problems unreachable
without it. Transfer across worlds and representations measured with a
cost-of-wrong-prior null.
Endogenous language pathway — 18 modules in core/brain/llm/endogenous_*.py.
A trained, causal path from Aura's 74-dimension cognitive state into the
transformer's output distribution: z_Aura → Δlogits → language. First fit
on 1,629 live turns (9B lane): held-out gain 0.0208 nats, paired recovery
54.0% (p = 0.043). No generation has been biased by the pathway yet.
Ghost substrate — core/ghost/ (6 modules). Hash-chained continuity across
substrate swaps. Four components: causal integration, ghost line, hack guard,
provenance.
Whole-system Φ and Inner Light — honest integration measurement, not a
consciousness claim. The Inner Light test scores Aura's activity on four
neuroscience markers against six negative controls; only the intact system
is 4/4. Enforced in code: PhiEstimate.claim.
Full details: docs/LANGUAGE_SUBSTRATE_AND_GENERALITY.md
Why Aura is Different
Most "AI companion" projects do the same thing: store a mood number, paste it into the system prompt, let the model act it out. The model says it's feeling energetic because it read the words "feeling energetic."
That's a costume. Aura is built the other way around.
When Aura is in an emotional state, that state becomes a direction vector added to the transformer's hidden activations during generation. The model's internal computation changes — not just the text it reads. This is the same family of techniques interpretability researchers use to steer behavior (CAA, activation addition, residual-stream interventions).
Underneath that, a substrate that never stops. Emotions decay and pull on each other. Neurochemicals rise and fall on their own clocks. A global workspace picks which thought wins the tick. A dream cycle consolidates memory while she's idle. And one gate — the Unified Will — signs off on everything that leaves the system.
It's a research project. It's also one you can talk to while it's running.
Table of Contents
- Quick start
- Evidence boundary
- Behavioral proof standard
- Recursive Latent Cortex — the flagship research program
- Language substrate and generality — learning new abstractions from experience
- Tracked vs local workspace
- Architecture overview
- Decisive evidence runner
- Decision authority
- Inference-time steering
- IIT 4.0 computation
- Consciousness modules
- Reality Reach and physical claim honesty
- Memory architecture — the 116-module memory subsystem
- RSI architecture — self-improvement pipeline and safety boundaries
- Reasoning engines — deterministic reasoning, sandboxed execution, and what's out of the model's hands
- Documentation status map — which docs are current, historical, or generated
- Docs index · Changelog · Agent guide
- Benchmarks
- Testing
- Personality training
- Data layer
- What this isn't
- License
Quick start
make setup # .venv + requirements/core.txt + requirements/dev.txt
or, for a fail-closed production install with no fallbacks:
make setup-prod
Full stack + UI
python aura_main.py --desktop
Background cognition only, no UI
python aura_main.py --headless
Reload code changes without restarting
curl -X POST http://localhost:8000/api/system/hot-reload
Other boot modes: --cli (interactive console), --server (API only),
--gui-window (attach a window to a running server), --watchdog,
--philosophy (stream substrate/phi/affect/Will as JSONL), --skeletal
(bypass heavy subsystems), --profile minimal, plus --stop and --reboot.
The installed console script exposes the same entry point as aura, which
also carries the operational subcommands (doctor, conformance,
verify-state, verify-memory, rebuild-index, backup, restore,
migrate, chaos, plugin).
Requirements: Python 3.12+, macOS on Apple Silicon, 64 GB RAM recommended. The
primary model is Aura-Cortex (fused Qwen3.8-27B, migrated from the historical
32B checkpoint). The Brainstem fallback is 9B and loads on demand. First boot
takes 30–60 seconds while Metal compiles shaders.
Hardware honesty: Bryan's target machine is an M5-class Apple Silicon Mac with 64 GB unified memory. The 27B Cortex works there as a primary conversation lane, while heartbeat/background work still belongs to the substrate, Brainstem, or Reflex lanes. On lower-memory machines, the hardware auditor rejects heavy weights for real-time tiers; use the 1.5B or 9B lanes there.
There's also a Dockerfile and docker-compose.yml if you want Redis and Celery
running alongside. The tracked workspace defaults to owner_autonomous posture
for this single-owner machine: autonomy on, outbound/network-enabled skills
available, and self-repair left active. If you want a tighter deployment,
override the AURA_* security settings in your local environment, including
AURA_INTERNAL_ONLY=1 for localhost-only binding.
Tracked vs local workspace
This repository is the baseline, not the whole story on any given machine.
Canonical skills live under core/skills/. The top-level skills/ package
is a compatibility layer for older imports and nothing new should go there.
If you're auditing: a local workspace can hold private modules listed in
.gitignore. They aren't in the tracked review surface, and they can change
the risk profile of that specific machine. Reading this repo tells you about
this repo. If you're auditing a real deployment, read the disk too.
The reproducibility consequence, stated plainly. Because of the above,
plus model weights, plus the local vector stores and the 6.5M-document
corpus, none of which are in git: *the public source is not sufficient to
reproduce a demonstration from this repository.* A third party cannot verify
from the tracked tree alone what ran. That is a real limitation of every result
here — not a caveat on some of them — and it's why every claim in
CLAIMS_MATRIX.md that rests on a local run is classified locally
demonstrated ("passed on this machine, this profile, this project's battery")
rather than independently demonstrated.
Closing this gap needs a frozen release pinning exact model hashes, vector-store
hashes and configuration, plus an independent run on a third-party machine.
Neither exists yet. Until they do, "external validation" stays not proven
(claim 12), and no amount of local evidence changes that, because local
evidence is the thing being questioned.
Architecture overview
The short version:
User input -> HTTP API -> KernelInterface.process()
-> AuraKernel.tick():
Consciousness -> Affect -> Motivation -> Routing -> Response generation
-> State commit (SQLite) -> Response
Every tick is event-sourced. Each phase produces a new immutable state version, the tick holds a lock while the pipeline runs, state commits to SQLite, the lock releases.
Crash in the middle of that and the write-ahead log replays on restart. No half-written thought survives.
Kernel (core/kernel/)
Tick-based cognitive cycle. One tick = one unit of thought. Phases run in order,
state versions, state commits, lock released.
Brain (core/brain/)
Local LLM router with automatic failover:
- Primary (Cortex) —
Aura-Cortex(fused Qwen3.8-27B, migrated from
- Secondary (Solver) — Qwen 2.5 / Qwen 3 72B for deep reasoning, hot-swapped
- Tertiary (Brainstem) — Qwen 3.5 9B 4-bit, lazy-loaded to save memory for
- Reflex — Qwen 2.5 1.5B 4-bit on CPU as an emergency fallback.
- Cloud — Gemini Flash/Pro, with personal info scrubbed and rate-limited.
- Last resort — rule-based static responses that can't fail.
core/config.py: fast_model is the Cortex
(Aura-Cortex / fused Qwen3.8-27B), deep_model defaults to the Cortex
(promoted to Qwen2.5-72B-Instruct-4bit via AURA_DEEP_MODEL),
chat_model the Brainstem (Qwen3.5-9B-4bit), and vision_model is
pinned to the Cortex build so vision and conversation share one identity.
Two non-LLM lanes were replaced in August 2026 based on measurements:
- Speech-to-text is one streaming-native Parakeet TDT pass
core/voice/duplex/streaming_asr.py) serving both duplex stages, replacing
a two-stage Whisper setup (small.en for partials, large-v3-turbo for the
final). Measured on this host over 12.4s of real speech, median of 5 warm
runs: Parakeet 166 ms vs Whisper-small 193 ms vs Whisper-large-v3-turbo
317 ms. One decode is cheaper than the old partial and about half the
old final, so both stages run the same weights on one model-lane lease.
faster_whisper remains as the CPU fallback.
- Embeddings are
Qwen3-Embedding-0.6Bat 384 dimensions
core/memory/embedding_model.py), replacing all-MiniLM-L6-v2. MiniLM
declares max_seq_length: 256 while the ingestion path chunks at 800 words —
1,122 tokens through the model's own tokenizer, so **77% of every full chunk
never reached the encoder**, silently. On four documents whose distinguishing
sentence sits past token 256, MiniLM scored 1/4 on tail retrieval (it ranked
the same document first every time) against Qwen3's 3/4, at 10.7 vs 20.2
ms/query.
What this ladder costs, stated plainly. It's good for availability and bad for attribution. A visible success could mean the Cortex answered; it could also mean the Cortex failed, the cognitive pipeline failed, steering never ran, and rung 4 or rung 6 produced the text you're reading. Those are very different events and they look identical in a transcript. The same is true of post-generation shaping in the chat route — intent classifiers, canonical answer contracts, identity and shape repair, retries — any of which can replace what the machinery actually produced.
So: **a transcript without lane and phase attribution is not evidence about
the architecture**, and no demo in this repository should be read as one.
Where a comparison is being made rather than a story told, the harness in
core/evaluation/matched_budget.py counts fallbacks, retries and human
intervention against the denominator and reports a clean_success_rate
alongside the raw one — because a run that needed rescuing is not a run the
architecture completed.
Decisive evidence runner
For the smallest hostile-review bundle, run:
bash scripts/run_decisive_test.sh
It generates tests/DECISIVE_RESULTS.json and tests/SCALE_SWEEP_RESULTS.json
covering black-box prompt hygiene, rich-prompt steering controls, phi reference
sanity checks, mutual-information permutation baselines, hardware feasibility,
resource-stakes persistence, and a bounded scale-sensitivity sweep. When
mlx_lm is available, the A/B step actually invokes Qwen2.5-1.5B for all four
conditions (black-box / terse text / rich adversarial text / baseline); the
source field in the JSON is live_mlx in that case and synthetic_fallback
otherwise.
Long-run autonomy harness
python tests/long_run_autonomy.py --ticks 1000
Drives adaptive mood coefficients, the resource-stakes ledger, emergent goals,
mesh cognition, the structural mutator, lineage, and self-awareness together
through N ticks with perturbations. No manual resets. Writes
tests/LONG_RUN_AUTONOMY_RESULTS.json with the 8-metric panel (viability,
coherence, calibration, report consistency, planning depth, recovery time,
memory integrity, action diversity) and an audit of which modules were touched
per tick.
The live desktop Cortex uses Aura's Apple Silicon MLX runtime. Circuit breakers, a GPU semaphore, a proactive cortex watchdog, and 429 handling keep the pipeline from cascading when something misbehaves.
Affect (core/affect/)
An 8-emotion model (based on Plutchik's wheel) plus somatic dimensions (energy,
tension, valence, arousal). These values don't just color the prompt. They
adjust sampling parameters (temperature, token budget, repetition penalty) via
the affective circumplex, and they feed the steering engine that injects
activation vectors into the model's hidden states.
Identity (core/identity.py, core/identity/heartstone.py)
An immutable constitutional core plus a mutable persona that drifts with sleep
and dream consolidation. There's active defense against prompt injection — the
dream cycle simulates identity perturbation and tries to repair drift back
toward the anchor.
Agency (core/agency/)
Self-initiated behavior scored along curiosity, continuity, social, and creative
dimensions. Refusal is a real option here — it isn't content filtering, it's a
decision the agent can make. Volition levels 0–3 gate progressively autonomous
behavior up to and including self-modification.
Skills (core/skills/, legacy wrappers in skills/)
115 modules: shell with sandboxing, web search and browse, coding, sleep and
dream consolidation, local media generation, social media (Twitter, Reddit),
screen capture, filesystem, browser automation, network recon, malware
analysis, self-evolution and self-repair, inter-agent messaging, knowledge
base, curiosity-driven exploration. The canonical tracked implementations live
under core/skills/; the top-level skills/ package is retained only as a
legacy compatibility layer for older imports. Every skill call carries a
capability token and has to pass the Will gate.
Orchestrator (core/orchestrator/)
About 3,335 lines in main.py split across 11 mixins (6,300 lines total
across all orchestrator modules): message handling, message pipeline, incoming
logic, response processing, tool execution, autonomy, cognitive background,
context streaming, learning and evolution, personality bridge, output
formatting. Handlers under orchestrator/handlers/ dispatch by message type.
This is the glue between the tick pipeline, the LLM router, and the
consciousness stack.
Somatic cortex (core/somatic/)
A body-schema map of available capabilities, a capability-discovery daemon
that periodically scans for new hardware or software, a motor cortex that runs
a 50 ms reflex loop for pre-approved actions (no LLM in the loop), and an
action-feedback channel that pipes success or failure back into affect.
Autonomy (core/autonomy/)
Self-modification pipeline (propose → sandbox test → simulate → Will authorize →
hot reload), value evolution (drive weights adapt from experience), scar
formation (critical events leave persistent markers), and a boredom accumulator
that nudges the system toward novelty when prediction error stays low too long.
Self-modification engine (core/self_modification/)
A pattern-detection error-intelligence layer, meta-learning, AST-level safety
analysis, shadow-runtime validation, a kernel refiner, a ghost-boot validator
that tests modifications without actually restarting, a shadow AST healer, and
code repair. Nothing modifies itself without Will sign-off.
Resilience (core/resilience/)
63 modules for not crashing: a stability guardian, circuit breakers with
persistent state, a cognitive write-ahead log, graceful degradation that
sheds capability under pressure, a healing swarm, a sovereign watchdog, a
resource arbitrator, a lock watchdog that hunts deadlocks, a memory governor,
an integrity monitor, an antibody system for threat response, and a diagnostic
hub.
Interface (interface/)
FastAPI and WebSocket with streaming. The main UI is vanilla JS
(interface/static/aura.js) with a live neural feed, telemetry, chat, and
substrate visualization. The memory dashboard is React + Vite + Tailwind
(interface/static/memory/). Routes cover chat, inner-state inspection,
memory browsing, system management, and privacy. Parakeet TDT for speech-to-text.
Hot-reload button in the UI for code changes.
Memory (core/memory/)
116 modules across four tiers: working (in-process), episodic
(SQLite with emotional tags and recency ranking), semantic (vector search via
Qwen3-Embedding-0.6B at 384 dimensions), and strategic (long-horizon goal
state). Key subsystems include a hippocampal indexer, reconsolidation engine,
associative entity graph, knowledge graph (39K bytes), black-hole quarantine
vault for poisoned memories, scar formation for catastrophic failures, and a
memory write gateway enforcing privacy, rate-limiting, deduplication, and
retention policies. The full memory architecture is documented in
docs/MEMORY_ARCHITECTURE.md.
Reasoning (core/reasoning/, core/brain/reasoning_amplifier*.py)
Built-in deterministic reasoning engines that handle hard problems without
relying on the model: a System-2 engine (native_system2.py, 80K bytes), a
proof kernel, natural deduction, linear arithmetic, and a proof-answer solver.
The Reasoning Amplifier v2 (core/brain/reasoning_amplifier_v2.py, 86K bytes)
runs five execution modes — FAST (1 candidate), NORMAL (3), DEEP (9 +
adversarial Courtroom Judge), EXTREME (Courtroom + sandbox repair + memory),
and PROOF (refuses to answer unless verifier-clean). An adversarial Courtroom
(core/brain/courtroom.py) runs Solver/Skeptic/Judge roles on complex
assertions. See docs/REASONING_ENGINES.md.
Verification (core/verify/)
37 modules implementing structural validation discipline (modeled after
LLVM's -verify-each). 52 standing runtime invariants enforce service
container acyclicity, lock ordering (with zero inversion errors), OOM
spine immunity, and epistemic consistency. Any check that raises an unhandled
exception is treated as a violation rather than a pass — there is no
"unknown means okay."
Hardware acceleration (rust_extensions/aura_m1_ext)
A Rust PyO3 crate linking macOS pthread_set_qos_class_self_np to schedule
foreground reasoning on Performance cores (QOS_CLASS_USER_INITIATED) and
background sensing/IO on Efficiency cores (QOS_CLASS_UTILITY). Also provides
high-performance AST parsing with rustpython_parser and zero-copy SHA-256
caching.
Decision authority
Anything the system actually does — sending a response, calling a tool, writing
a memory, starting an initiative, changing state — has to pass through one
function: UnifiedWill.decide() in core/governance/will.py (core/will.py is the facade).
Action request
-> UnifiedWill.decide() [core/governance/will.py]
-> SubstrateAuthority [field coherence, somatic veto]
-> CanonicalSelf [identity alignment]
-> Affect valence [emotional weighting]
-> WillDecision (receipt with provenance)
-> Domain-specific checks [AuthorityGateway, CapabilityTokens]
-> Action runs, or is refused/deferred/constrained
Every decision produces a receipt. If an action doesn't carry a valid
WillReceiptEntry, it didn't happen. Receipts are logged with their source,
domain, outcome, reason, constraints, substrate receipt ID, executive intent
ID, and capability token ID. See OWNERSHIP.md for the full
map of who owns what.
Inference-time steering
The steering engine (core/consciousness/affective_steering.py) hooks into
MLX transformer blocks and adds learned direction vectors to the model's
hidden states while tokens are being generated:
# Simplified from affective_steering.py
h = original_forward(args, *kwargs)
composite = hook.compute_composite_vector_mx(dtype=h.dtype)
if composite is not None:
h = h + alpha * composite
return h
This is contrastive activation addition — the technique from Turner et al. 2023, Zou et al. 2023, and Rimsky et al. 2024. The direction vectors come from the current emotional state, and they get injected at configurable layers.
On top of that, the precision sampler
(core/consciousness/precision_sampler.py) adjusts temperature based on
metabolic state (Pneuma arousal/circumplex) and adjusts top_p based on
MHAF topological attractor count, and the affective circumplex
(core/affect/affective_circumplex.py) maps somatic state to generation
parameters.
So there are three places emotion can touch generation:
- Hidden states — activation vectors added to the model's internal
- Sampling — temperature and top-p adjusted by emotional state. Changes
- Context — natural-language emotional cues in the system prompt.
The first is the interesting one. The third is what most "emotional AI" projects stop at.
IIT 4.0 computation
Aura computes Integrated Information (φ) at two scales simultaneously.
16-node cognitive complex — core/consciousness/phi_core.py
- Binarize