AI2Apps
A local-first AI application and Agent platform for personal AI nodes.
AI2Apps turns local and connected model runtimes into durable Apps, Agents, and versioned Services. It provides a server-owned Agent Harness, application shell, tool and capability gateway, package trust system, multi-user identity, and remote access above one or more model backends.
Apple Silicon and the embedded oMLX runtime are the first implementation. AI2Apps also keeps its platform contracts hardware-neutral so that external providers and future NVIDIA/CUDA or AMD/ROCm nodes can expose the same model and Service capabilities.
AI2Apps is not affiliated with, sponsored by, or endorsed by the oMLX
project or its maintainers. The oMLX name identifies the origin of the
embedded runtime. See NOTICE for attribution.
中文说明 · Platform architecture · ACPF capability provisioning · Backend plan · Local Knowledge/RAG · Security baseline · Release gate
Product model
AI2Apps is organized around four product objects and one runtime layer:
User / API client
↓
App interaction, UI, instances, Sessions, files and artifacts
↓
Agent goals, instructions, model policy, tools and durable execution
↓
Service stable, versioned and auditable capabilities
↓
Runtime local oMLX/Fusion, Cloud, external or future federated providers
- Apps own user interaction and persistent application state. Built-in
- Agents run through an authoritative server-side Harness. Runs, steps,
- Services expose models, tools, workspace operations, processes, browser
- AI nodes keep application and Session data local while supporting Cloud
- Model runtimes remain replaceable backends. Cache-MoE and Fusion are
Implemented platform capabilities
The current alpha includes:
- a SQLite-backed App, AppInstance, Session, Message, AgentRun, Step, Event,
- persistent asynchronous Agent execution with Tool calls, user interactions,
- a Service Registry and Tool Gateway for embedded, sandboxed managed-process,
- signed and content-addressed
.ai2service,.ai2agent,.ai2app, and local
.ai2patch package flows with verification, lifecycle, rollback, and Safe
Mode;
- Session-scoped workspaces, ResourceHandles, artifacts, document parsing and
- Chat as a per-user singleton App with independently isolated thread Sessions;
- Coder projects and threads for Codex, OpenCode, and Claude CLIs, including
- local installation identity, multiple member roles, per-user ownership,
- a desktop shell, mobile-ready App contracts, managed remote access, and
- OpenAI-compatible model APIs backed by the existing oMLX runtime.
Local inference and Fusion
The embedded oMLX backend retains model loading, attention, fused MoE kernels, continuous batching, paged KV caching, audio/VLM engines, embeddings, reranking, and MCP integration. AI2Apps adds local inference capabilities for large MoE models such as DeepSeek V4 Flesh:
- configurable flat or hierarchical scope catalogs;
- shared-expert scope probing and per-scope static expert banks;
- device-side Top-K routing with exact and explicitly enabled lossy policies;
- expert-major SSD storage and cache-aware fallback loading;
- Session-safe KV/prefix-cache namespaces and adaptive L1 expert residency;
- reproducible memory, quality, prefill, decode, miss, and I/O release gates.
Repository layout
ai2apps/ Platform, App, Agent, Service and product code
agents/ Durable Agent Runtime and built-in Agents
api/ Versioned platform APIs
apps/ App definitions, access policy and lifecycle
packages/ Signed package trust and Service lifecycle
services/ Service Registry, adapters and Tool Gateway
storage/ SQLite schema, migrations and repositories
workspace/ documents/ Session resources, artifacts and document tools
web/ Desktop/mobile Shell and built-in App UI
apps/omlx-mac/ Native macOS application shell
omlx/ Embedded and modified oMLX model runtime
engine/flesh.py DeepSeek V4 Flesh request orchestration
cache/ KV and routed-expert storage
patches/deepseek_v4/ Scope routing, expert banks and kernels
configs/ Scope catalogs and model profiles
scripts/ Conversion, profiling, packaging and benchmarks
docs/ Architecture, product contracts and experiment records
tests/ Platform, security, API and inference tests
The omlx Python namespace and OMLX_* variables remain compatibility
interfaces for the embedded runtime. New integrations should use the
ai2apps command and AI2APPS_* configuration where available. Runtime data
currently remains under ~/.omlx so existing models and settings survive the
product migration.
Install
The bundled local model backend currently requires an Apple Silicon Mac, Python 3.11–3.13, and Metal-capable macOS.
brew install uv
uv sync --dev
source .venv/bin/activate
ai2apps --version
ai2apps info
Alternatively, create a Python 3.11–3.13 virtual environment and run:
python -m pip install -e '.[dev]'
Start AI2Apps
ai2apps serve --model-dir ~/models --port 8000
- App shell / dashboard:
- Chat:
- OpenAI base URL:
- Platform API root:
- Chat completions:
POST /v1/chat/completions - Model catalog:
GET /v1/models
omlx
executable remains a compatibility alias; new documentation and integrations
should use ai2apps.
DeepSeek V4 Flesh research overrides
Models prepared through AI2Apps use their verified Scope Pack automatically. Manual research environments can override the expert store and profile:
export AI2APPS_DEEPSEEK_V4_EXPERT_STORE=/path/to/expert-store
export AI2APPS_DEEPSEEK_V4_SCOPE_PROFILE=/path/to/scope-profile.json
export AI2APPS_DEEPSEEK_V4_SCOPE_NAME=general
export AI2APPS_DEEPSEEK_V4_SCOPE_PROBE_DEPTH=16
export AI2APPS_DEEPSEEK_V4_SCOPE_LOSSY_MODE=exact
ai2apps serve --model-dir /path/to/models
Lossy modes are opt-in. Use exact for quality-sensitive serving and evaluate
other policies with representative prompts before deployment.
Development and release gates
Before developing an AI2Apps App or System App, read the AI2Apps App development guide, including the shared cross-environment Artifact download UX contract.
Before developing an installable Service or model Package, read the
Service/Package runtime and Sandbox development guide.
For the Model Worker protocol, Adapter API, and checkpoint contract, see the
Model Worker Package manual. A local
Harness or terminal run does not reproduce installed Sandbox permissions; a
real .ai2service installation and activation is required before release.
The active experimental branch is experiment/moe-cache. Preserve existing
oMLX model, attention, router, scheduler, and fused-kernel behavior unless an
AI2Apps feature requires a small, isolated compatibility patch. Platform code
should remain under ai2apps and depend on model runtimes through adapters.
pytest -q
python scripts/bench_scope_once.py --help
python scripts/bench_moe_expert_store.py --help
ai2apps-release-gate --mode preflight --run-tests
Inference comparisons use identical prompts and generated tokens and record source commit, resident/peak memory, cold TPS, and steady TPS. The original static oracle gate requires exact Top-10 parity, zero runtime misses, lower resident memory, and at least 85% of full-resident steady-state TPS.
Platform changes must additionally preserve ownership isolation, capability enforcement, package verification, restart recovery, API compatibility, and bounded resource behavior.
Origin, license, and trademarks
AI2Apps is based on oMLX commit
49ec271
and retains upstream copyright and attribution notices. Modified files and the
repository history identify AI2Apps changes.
This is a multi-license distribution. Its default license is the Apache License 2.0. The AI2Apps Official Cloud Connector is source-available under the Business Source License 1.1: production clients may use it without a separate commercial license when they connect exclusively to an Official AI2Apps Cloud Service. Using it to implement, connect to, or provide an Alternative Cloud Service requires a commercial license. This version changes to Apache-2.0 on 2029-09-07. Earlier versions keep the license terms under which they were distributed.
Copyright 2025 oMLX contributors; Copyright 2026 AI2Apps contributors. No software license grants rights to the AI2Apps name, logo, or other marks; see the AI2Apps trademark policy. Apache-2.0 also does not grant broad rights to upstream trade names or marks. AI2Apps does not use the oMLX name or logo as its product identity and makes no claim of upstream affiliation, sponsorship, certification, or endorsement.