One Switch
Put every LLM channel you own behind one local address. When one goes down, the next one takes over.
English | 简体中文
One Switch runs a proxy on your machine. You register all the channels you have — different providers, different accounts, different models — put them in the order you want them tried, and point every AI client at a single local address. From then on it does the work: identify the protocol, pick a channel, send the request, move on when a channel fails, and record exactly what happened.
Your client only ever sees the attempt that succeeded.
Why it's worth installing
- Configure once, use it everywhere. Every client points at one local address. Swapping providers, accounts or models later means editing One Switch, not hunting through each tool's settings.
- Channels break; your work doesn't. Network hiccups, connection timeouts, rate limits, exhausted quota, rejected keys and upstream 5xx all push the request to the next channel automatically. A response that has already started streaming is never spliced together from a second one — you get a failure instead of a Frankenstein answer.
- Every request tells you the truth. Which provider and model actually served it, which attempt succeeded, how long it took, how fast the first token arrived, tokens per second, how much of the prompt was cached — all of it stored and queryable.
- Provider quirks without writing code. Need to add a
User-Agent, drop a header, or pin a field to0.7? Request rewrite rules do it, and the editor validates the change against a test case on the spot. - Your data stays on your machine. Local listener on
127.0.0.1, keys in the OS-encrypted store, no account, no cloud sync, no relay server. Requests only go to the upstreams you configured. - English and Chinese UI, light / dark / follow-system themes, lives in the system tray, auto-launch at login, built-in updater.
Screenshots
Logical models: drag to set priority, and every model carries its own recent track record.
!Logical models: channel queue with live metrics
Smart Routing: a node graph for "which requests land in which channel group", saved as versions you can roll back.
!Smart Routing: node graph and request matching
Request Logs: one row per request, expandable into the full execution detail and raw usage.
!Request Logs: per-attempt detail and usage
Analytics: success rate, latency, TTFT, TPS, cache hits, model ranking, failure reasons.
!Analytics: metric cards, usage distribution and model ranking
Request Rewrite: stat cards, the rule list, and the templates behind New rule.
!Request Rewrite: rule list and template menu
Up and running in three steps
1. Add a channel
Open Model Management → New provider. Fill in the name, API key, timeout, and the default endpoint for each protocol this provider speaks.
Then add the real model IDs under that provider (gpt-4.1-mini, deepseek-reasoner, claude-sonnet-4, …) and tick the protocols each one supports. A model that speaks several protocols still occupies a single row in the queue.
2. Order the queue
Go to Logical Models and drag the models you just added into the order you want them tried. Switch off whatever should sit out.
Things worth doing while you're here:
- Every row shows when that model last succeeded, how many consecutive failures it has, plus its TPS and TTFT — enough to decide who deserves to go first.
- Each logical model card has a Failover / Manual switch. In Manual mode the request is pinned to one upstream model: that row is marked as selected and the rest go on standby.
- To find out whether a channel actually works, use Model Management → Connection test, tick several channels and protocols and verify them concurrently. Each target receives one minimal real request, which may cost a little and will show up in the request logs.
3. Repoint your client
Access Config is a three-step guide: confirm the service is running, pick your client type, copy the address it asks for. Addresses are built from the current listener, and every one of them has a copy button.
| Your client | Base URL |
| --- | --- |
| OpenAI compatible (Chat Completions / Responses) | http://127.0.0.1:9300/v1 |
| Anthropic | http://127.0.0.1:9300 |
⚠️ The trap everyone falls into: Anthropic clients append/v1/messagesthemselves, so the Base URL stops at the port. Add/v1and the request becomes/v1/v1/messages, which the proxy does not recognise — you get a 404.
The model name is up to you: default, or any non-empty name. One Switch swaps it for the real model ID of whichever channel it selects. If your client insists on an API key, any placeholder will do — the real keys are injected per provider.
Check that the service is alive:
curl http://127.0.0.1:9300/v1/models
The listener host and port live in Settings → Network → Local Listener; save and the proxy moves to the new port.
Failover rules
| What happens upstream | What One Switch does |
| --- | --- |
| Network error, connection timeout, streaming idle timeout | Try the next channel |
| 401, 403 | Try the next channel, and count the failure against that provider |
| 408, 429 | Try the next channel |
| 5xx | Try the next channel |
| Any other 4xx (a malformed request, say) | Returned to you as-is — another channel would not fix it |
| Breaks off after the response has started streaming to you | Aborts the request rather than splicing in another channel's output |
The defaults are 3 consecutive failures before a provider enters cooldown, a 30-second initial cooldown that grows with each failure up to 5 minutes, and a 30-second streaming idle timeout. All three live in Settings → Reliability → Failover.
Supported protocols
| Protocol | Local path | Typical upstreams |
| --- | --- | --- |
| OpenAI Chat Completions | /v1/chat/completions | OpenAI, DeepSeek, Volcengine Ark, OpenRouter, Ollama — anything OpenAI compatible |
| OpenAI Responses | /v1/responses | Services that implement the Responses API |
| Anthropic Messages | /v1/messages | Claude and compatible services |
These paths are recognised with or without the /v1 prefix. GET /v1/models is served locally and returns your model names; it is never forwarded upstream.
About protocol conversion: nothing is converted by default — requests pass through untouched, which is both the safest and the fastest behaviour. If you genuinely need a Claude client to talk to an OpenAI-only channel, turn conversion on for that endpoint binding. Conversion is a best-effort compatibility layer and some parameters may be lost; a single failover will only ever consider channels that match natively or that you have explicitly enabled conversion for.
Smart Routing
This is the landing page after install, and it answers one question: which requests belong to which channel group.
- Four ready-made policies you can drop straight onto the canvas: Logical model hit (use the requested model when it names a logical model, otherwise fall back to the default), Route by user agent (recognise Cursor or Claude CLI and split accordingly, everything else falls back), LLM request complexity (let a model judge difficulty and send the hard ones to the strong channel, the rest to the fast cheap one), and JS script request handling (a sandboxed script scores the request and buckets it).
- Or build your own from nodes: input → protocol discovery → conditions → logical model selection → output. Each node does exactly one thing.
- Test run takes a real request body and shows which branch it takes and what each node produced. Nothing is forwarded upstream.
- Every save leaves a version behind, so a bad policy is one rollback away.
default logical model is the safety net: anything no policy matches ends up there.
Request Rewrite
Maintain rules on the Request Rewrite page to smooth over small differences between providers:
- Two match conditions only: client protocol and upstream protocol. Leave them empty to apply everywhere — no protocol matrix to work out first.
- Actions run in the request or response stage. Headers can be set, appended or removed; JSON bodies can have values set, paths deleted, or strings replaced by literal or regular expression via
$.path. - A rule can be global (applies to every channel) or bound to specific models with an explicit execution order.
- The editor carries a test case, so you see the effect of a change immediately without sending a real request.
OneSwitch/), Drop a request header, and Set a request field.
New rules are enabled; disable one from the list if you change your mind. Every change is a structured add/delete/replace and nothing ever executes a script. Response-stage actions only touch complete non-streaming JSON — the body of a streaming response is never rewritten.
Data and privacy
- The proxy listens on
127.0.0.1by default and does not expose itself to the local network. - API keys are kept in the OS-encrypted store. Exporting providers has an include plaintext API keys switch that is on by default — an export carrying keys is a working credential, so pass it between your own devices and nowhere else.
| File | Contents |
| --- | --- |
| ~/.one-switch/one-switch-config-v1.db | Providers, models, routing and rewrite rules — your configuration, worth backing up |
| ~/.one-switch/one-switch-data-v1.db | Request logs, captured bodies, usage and health state — safe to delete, you only lose history |
The development build uses ~/.one-switch-development instead, so a dev instance never touches your real data.
A few more things worth knowing:
- Body capture is on by default. The full request, response and streamed content are stored locally, including both sides of any protocol conversion. Headers are redacted automatically (
authorization,x-api-key,cookieand friends); bodies are not — which is exactly why it can debug any request, and why the logs may contain sensitive content. Bodies are kept for 7 days by default while the request records themselves are kept forever. Turn capture off, change the retention windows, or clear history from Settings → Data → Request Logs. - To support very long contexts properly, neither the proxy nor body capture caps request size; the body is read into memory in full. Enormous bodies will cost real memory and disk. That is a deliberate trade-off in this version.
- There is no cloud sync, no account and no remote relay. Requests only go to the upstreams you configured.
Install
Download the installer for your platform from GitHub Releases:
- macOS:
.dmg, one for Apple Silicon and one for Intel (plus a.zip, which the updater uses) - Windows:
.exeinstaller, x64 and ARM64 - Linux:
.AppImage, x64 and ARM64
Not supported yet
Better to be clear about the edges than let you find them the hard way:
- No system-wide proxy — you change the Base URL in each AI tool yourself.
- No Gemini
/v1beta/models/*endpoints. - No team collaboration, multi-user permissions, cloud sync or remote access.
- A single request's failover only picks among channels that match the protocol natively or where you explicitly enabled conversion. It never guesses across protocols.
Local development
Node.js 22+ and pnpm 11 are required:
pnpm install
pnpm dev
Useful commands:
pnpm dev # dev session: console dev server + Electron
pnpm dev:preview # console only, for looking at the UI in a browser
pnpm typecheck # TypeScript
pnpm lint # ESLint plus the layering and package-boundary guards
pnpm test # the whole test suite
pnpm build # compile every package (no installers)
pnpm release:mac # build macOS arm64 / x64 installers
pnpm release:win # build Windows arm64 / x64 installers
pnpm release:linux # build Linux arm64 / x64 installers
The repository is a pnpm workspace: packages/{contracts,core,console} are libraries that can be consumed on their own, packages/toolkit holds the cross-package development scripts, apps/app is the desktop host, and Turborepo runs the tasks. The stack is Electron + React + TypeScript + Vite + Drizzle ORM + SQLite.
Design goals, behaviour contracts and acceptance criteria have a single authority in product/; build and packaging details live in packaging.md.
Feedback
Issues and ideas are welcome in Issues. The version number, operating system, protocol and a redacted runtime log go a long way — but please do not paste API keys, full prompts or other sensitive content.