English · 中文 · 🌐 dvalincode.dev
Open security engineering for code written by humans and AI agents.
Every repair carries its own proof.
When an agent fixes a security finding, someone has to decide whether the fix worked. Almost every tool asks the model that wrote it — which is the one question a model cannot answer against its own interest.
Dvalin decides instead, and hands you the proof. It re-scans, runs your project's own tests itself, and reads the exit codes from processes it started. Who wrote the repair — our agent, Claude Code, Codex, Copilot, a person — is recorded and never consulted. The result is a Verified Fix Record: a small JSON file anyone can re-check offline, on a laptop with no network and no Dvalin state.
dvalin verify-fix fix-record.json
Fix record 2c9d71ac03e0 · VERIFIED · scan-and-checks
executor: claude-code (recorded, not consulted)
targets: 1 before · 0 remaining
coverage: complete → complete
introduced: 0 (gate high/new)
outcome: verified
✓ test: npm run test (exit 0)
audit: run verify-36509f42 @ 414644c75af0
That record says something narrow on purpose: *these findings were gone, and these checks were observed to pass.* It is not a claim that your code is safe, and Dvalin will not let it be read as one — every record carries what the scan actually covered, and a repair no check could confirm does not pass. The open profile →
A repair is a change, and a change can add as well as remove. So the record also
carries what the re-scan saw that the first scan did not, and the gate threshold
the verdict was reached under: a fix that removes an eval and introduces an SQL
injection is recorded as regressed and does not verify. Neither does a record
whose issuer never looked — introduced: not determined fails, because a
verifier that skips the question must not score better than one that asks it and
finds something.
Dvalin is the independent security runtime between code generation and merge. Humans, coding agents, and CI call the same versioned contract for discovery, remediation, and verification. It runs independently, or interoperates with specialist systems such as Codex Security through portable SARIF. Its built-in coding capability is a remediation executor — not the trust boundary, and not an attempt to compete with every general-purpose coding agent. See the security-agent strategy.
⏱️ 30 seconds, no install, no API key
npx dvalincode security scan .
After installing the package: dvalin scan .
That is the whole thing. It runs the built-in rules for injection, hardcoded
secrets, XSS, eval, and unsafe shell use against the current directory and
prints what it found. No account, no model, no config, no code leaves your
machine. The default policy runs only Dvalin Built-in, so the first scan always
works. Add optional engines explicitly, or inspect their fixed install commands:
dvalin scanners list
dvalin scanners install semgrep # review the command
dvalin scanners install semgrep --yes # execute it under Dvalin policy
Measure with the scanner your CI gate uses
If Snyk is what blocks your merges, verify against Snyk. A fix judged "clean" by a different engine is a fix the gate can still reject, and that mismatch is what turns one agent repair into several rounds of push, fail, and rebase.
npm install -g snyk && snyk auth # or set SNYK_TOKEN
dvalin scan . --scanners builtin,snyk-code,snyk-oss
snyk-code (SAST) and snyk-oss (dependencies) are engines like any other:
their findings carry the same fingerprints, coverage and gate, and a failed
login is reported as an engine error, which makes the scan partial, never
clean. They run only when named, because Snyk Code uploads source to Snyk.
Name them in dvalin.security.json to make them part of the policy, and
reverify will judge every fix with them.
Ignores Snyk itself reports as accepted are honored and listed as coverage
exclusions. Ignores the change under review adds are not: when a fix is
judged, .snyk, .semgrepignore, .trivyignore and .dvalincodeignore are
restored to the base commit's version and new inline markers (deepcode
ignore, nosemgrep, nosec, NOSONAR, …) are blanked, so a finding that was
silenced rather than fixed still fails the fix. A suppression is a risk
decision a person makes in its own change, never a repair.
The same judgement applies wherever Dvalin says "verified": the fix loop,
dvalin verify, the MCP dvalin_verify_findings tool, single-round
--fix --verify, and CI reverify. Those records are dvalin-fix-record/v3,
which also fails a repair that hid a finding instead of fixing it — a sibling
sink, a deleted vulnerable file, deleted tests or assertions. Outside a git
repository, where the change cannot be determined, records stay v2 and say
that evasion was not evaluated.
Let it loop until the gate would pass
dvalincode dvalin . --scanners builtin,snyk-code,snyk-oss --until-clean --max-rounds 3 --executor codex
--until-clean turns one fix attempt into a bounded loop. The executor edits;
Dvalin re-scans with the same engines and runs the project's checks itself;
the executor is sent only what is still wrong — targets still present,
findings its change introduced, failing checks with their last lines of
output, suppressions it added — and fixes again. Dependency findings go first.
The loop stops, and says which, when every target is gone with nothing
blocking introduced and every check passing (verified), when the round
budget runs out, when a round leaves the same problems open or the open set
stops shrinking (stalled), when the targets need a person — a dependency
with no fixed version (not-auto-fixable) — or when the engines or checks
cannot run (unverifiable). Suppressions the executor adds are undone before
each scan and stay open problems until removed. Every outcome issues a fix
record, and every round is logged under ~/.dvalincode/security/fix-loops/;
dvalin loop-stats summarizes those logs — how often loops reach verified,
the median and p90 rounds to green, which stop rule fired, and what was still
open when a loop stalled (--json, --since, --executor to slice it).
Rounds are kept short without changing what decides them. The engines run side
by side, so a scan costs the slowest engine rather than the sum. Between rounds,
the engines that analyse one file at a time (Dvalin Built-in, Semgrep) rescan
only the files the change touched and the files the targets are in; the others
always scan the whole tree. That narrowed scan only steers the executor: any
round that would stop the loop is re-judged on a full scan of the same tree, and
only the full scan decides the outcome and goes into the record.
--full-rescan scans everything every round, and loop-stats reports scan time
per round and how often a full scan changed the decision.
For code findings the loop first asks for a reproduction: tests only, which
Dvalin runs on the unfixed code and requires to fail — by a failing assertion,
not a missing runner or a not-yet-written import. The tests are hashed; the
fix must then make them pass unchanged, and that run is a check in the
record. A finding no test could demonstrate stops the loop as
not-reproduced before anything is fixed: it may be a false positive, and a
person triages it. The runner is inferred (vitest, jest, mocha, node --test,
pytest, go test) or set with --repro-command 'npx vitest run {files}' or
reproduce in dvalin.security.json — never chosen by the executor.
Each fix round also rejects the evasions an agent reaches for when the goal is
"the scanner went quiet": moving the sink to a sibling (eval →
new Function, exec → spawn with shell: true), deleting the vulnerable
file, deleting tests, removing assertions. In practice the scanner alone is
fooled by the first one — the rule stops matching — and the reproduction test
and the evasion check are what catch it.
Keep it on today's main, and close it in CI
dvalincode dvalin . --scanners builtin,snyk-code,snyk-oss --until-clean \
--rebase-onto origin/main --executor codex --draft-pr --sign-key ci.key
With --rebase-onto the loop fetches and rebases the fix onto upstream before
it starts and after every round, then judges it there. Conflicts go to the
executor — the conflicted files only, with "do not run git" — and Dvalin
continues the rebase itself; markers left in a file are a conflict not
resolved. If they cannot be resolved the rebase is aborted, the tree is left as
it was, and the loop stops as rebase-conflict. After each rebase the baseline
is re-scanned on the new base, so a finding a teammate landed on main is not
blamed on the fix.
--draft-pr commits the record at .dvalin/fix-record.json, and
docs/examples/dvalin-fix-loop.yml
re-executes it on the pull request with reverify: true: base and head
re-scanned, checks re-run, and the reproduction re-run with the fix reverted in
place (it must fail) and restored (it must pass) — then signed with the CI key.
For an incremental “no new high-risk findings” gate, commit the policy and baseline with the repository:
dvalin init
dvalin baseline
dvalin scan
This creates dvalin.security.json and .dvalin/baseline.json. Suppressions
require a reason and may have an owner and expiry date. Scan output is a
versioned envelope with a deterministic gate result and a resumable workflow ID.
Or put it on every pull request — nothing to install at all
# .github/workflows/security.yml
permissions:
contents: read
security-events: write
steps:
- uses: actions/checkout@v5
with:
fetch-depth: 0 # so the scan can reach the base commit
- uses: arthurpanhku/[email protected]
with:
fail-on: high
diff: true # only report on what this PR changed
Findings land inline on the pull request diff and in your Security tab. No API key, no secrets, no model — the scan is deterministic and local to the runner. Full example →
diff: true reports only on lines the pull request changed, so the gate blocks
what this change adds instead of everything the repository already carried.
That is what makes the check adoptable on a codebase that was not clean to
begin with. Drop it to scan the whole repository.
Every comment states what the scan covered — complete, partial, or
unknown — beside the result, because "no findings" from a run where half the
engines were missing is not the same answer as "no findings" from a complete one.
And publish the proof next to the diff
If your pipeline produced a fix record, hand it to the same action:
- uses: arthurpanhku/[email protected]
with:
fix-record: fix-record.json
The runner re-derives the record from the file alone — recomputing its hash and re-deriving its verdict from its own evidence — and posts the result on the pull request. A record that was edited after it was issued fails here, and fails the job.
Re-derivation proves the record is self-consistent. It cannot prove the record was issued by anyone who actually ran anything: every field in it, the exit codes included, is something a forger could write, and a self-consistent forgery re-derives. Two things close that gap — re-execute the claim on the runner, and sign what the runner observed.
🔏 Verified Fix Record
✅ ce504a995395 · VERIFIED · scan-and-checks
- repaired by claude-code — recorded, and not consulted for this verdict
- targets: 1 before → 0 remaining
- coverage: complete → complete
- introduced: none (gate high/new)
- outcome: verified
- ✓ test:
npm run test (exit 0)
- audit chain: verify-eeb1bae7 @ 80881867270d
A repair that regressed says so in the same place, and fails the job with it:
❌ 916e2eeaf065 · NOT VERIFIED · scan-and-checks
- introduced: 1 finding(s) the first scan did not report (gate high/new)
- critical dvalin/sql-injection — src/db.ts:31
- outcome: regressed
Re-execute it on the runner, and sign what you saw
- uses: actions/checkout@v5
with:
fetch-depth: 0 # the base commit has to be reachable
- run: npm ci # the project's checks run on the runner
- uses: arthurpanhku/[email protected]
with:
fix-record: fix-record.json
reverify: true
signing-key: ${{ secrets.DVALIN_SIGNING_KEY }} # optional
With reverify: true the claimed record contributes only which targets it
says it fixed and who the executor was. Everything else is observed again:
- the base commit is checked out and scanned, so each claimed target must
- the head is scanned, so "gone" is observed on the runner;
- the project's checks are run on the runner, with the checks, gate and
dvalin.security.json — a pull
request that relaxes its own gate or deletes its own checks is judged under
the rules it is trying to change;
- a fresh record is issued from those observations, and the job fails unless it
With signing-key the fresh record is signed (Ed25519, over its hash). The key
is taken into memory and removed from the environment before the project's
checks run, so the code under review cannot read it. Downstream — a release
job, another repository, an auditor — can then require that signature:
dvalin keygen --out ci # once; store ci.key as a secret, publish ci.pub
dvalin verify-fix record.json --trusted-key ci.pub
A record that is unsigned, or signed only by a key you did not name, fails
under --trusted-key. A valid signature says *which key vouched for this
record*; whether that key deserves trust is your decision, and the reason to
trust a CI key is that the CI job re-executed the verification rather than
copying it. Locally, dvalin reverify runs the same
re-execution, and dvalin sign-fix signs a record that re-derives.
FVP-1 §4a, §5a →
Or let your agent call it
If an agent is writing the code, something other than that agent has to check it. DvalinCode is an MCP server, so any agent that speaks MCP can:
claude mcp add dvalin -- npx -y dvalincode mcp-serve --workspace .
One command configures the editor you actually use:
npx dvalincode mcp-install cursor # .cursor/mcp.json
npx dvalincode mcp-install vscode # .vscode/mcp.json
npx dvalincode mcp-install claude-code # .mcp.json
The formats differ in a way that fails silently — VS Code keys its servers under
servers, Cursor under mcpServers — so the command writes the right one and
merges into whatever is already there. Editors and MCP →
dvalin_scan accepts diff: "uncommitted", which reports only on what the
agent just wrote rather than everything the repository already carried — the
difference between a usable answer and a wall of pre-existing findings. It never
runs a model, edits the target workspace, or persists Dvalin state, so clients
can allow the preview by default. When a finding will be repaired, the agent
explicitly calls dvalin_begin_verification to record a small local workflow;
it can then retrieve the finding by fingerprint and request an independent
re-scan through dvalin_get_finding and dvalin_verify_findings.
That last one is the point: an agent that has just written a repair can ask for
an independent verdict on it. Dvalin re-scans, runs the project's own checks
itself, and returns a Verified Fix Record — what was targeted, what remains,
what the repair introduced that was not there before, the gate the verdict was
reached under, which commands ran and the exit codes Dvalin observed, and how
much of the codebase was actually covered. Whoever wrote the repair is recorded
and never consulted. dvalin_verify_fix re-derives such a record offline, so the reviewer
receiving it does not have to trust the tool that issued it.
FVP-1 → Responses include MCP structuredContent; scanner
readiness is available through dvalin_list_scanners. The same server exposes
dvalin_run_task as an optional implementation helper, plus session and audit
evidence tools.
Every client below has been driven to a real tool call rather than only a handshake — which client, which version, and on what date is a table rather than a sentence, because hand-written version numbers go stale quietly. Integration support ↓ · Agent integrations →
The repository also contains one dual Codex/Claude plugin payload with both native manifests, the shared security-gate skill, and the local MCP server configuration. Installing it makes the gate discoverable from task context instead of requiring every developer to remember a scan prompt. Codex honors the scan's read-only MCP annotation. Claude Code requires one explicit MCP permission by design; the plugin documents the exact scan-only allow rule instead of asking users to bypass all permissions.
Wherever you already work
One server, reached the way each tool expects:
| Harness | How Dvalin reaches it |
|---|---|
| Claude Code | dual plugin or claude mcp add · standalone skill |
| Codex | dual plugin or codex mcp add · SARIF interop with Codex Security |
| Cursor | dvalincode mcp-install cursor |
| VS Code | dvalincode mcp-install vscode · extension for Problems, coverage/gate status, and offline VFR verification — built, not yet published |
| Windsurf · Zed | stdio MCP through their own settings — server command |
| Any MCP client | MCP registry: io.github.arthurpanhku/dvalincode |
| GitHub Actions | Marketplace action — findings inline on the pull request diff |
| Any CI | dvalin scan . --fail-on high, SARIF out for code scanning |
The MCP config formats are not interchangeable — VS Code keys its servers under
servers, Cursor under mcpServers, and the wrong one fails silently — so
mcp-install writes the right shape and merges into whatever is already there.
Editors and MCP →
Or interoperate with Codex Security
Codex Security can export a completed, sealed scan as SARIF. Import that portable projection without coupling Dvalin to Codex Security's private state directory:
DVALIN_CODEX_SCAN_DIR=/tmp/codex-security-results
npx @openai/codex-security scan . --output-dir "$DVALIN_CODEX_SCAN_DIR"
npx @openai/codex-security export "$DVALIN_CODEX_SCAN_DIR" \
--export-format sarif --source-root "$PWD" --output /tmp/codex-security.sarif
dvalin import /tmp/codex-security.sarif .
dvalin scan . --fail-on high
The import creates stable Dvalin remediation cases; --no-persist validates the
handoff without changing the backlog. Keep Codex Security's original manifest,
findings, and coverage artifacts together—Dvalin imports the SARIF finding
projection but does not rewrite its sealed bundle or reinterpret its coverage.
Integration guide →
Then let it fix what it found
dvalincode dvalin . --fix --verify --draft-pr
This step does use a model — your model, any OpenAI-compatible endpoint. It prepares focused repairs in an isolated worktree, runs your tests, and requires a clean re-scan before anything can proceed to a draft PR. It never auto-merges, and a clean scan is never treated as proof that the code is safe.
This animation is made from the real application, not a mock. The input is an Apache-2.0-licensed example adapted from OWASP NodeGoat, whose contribution route evaluated user-controlled text.
🛡️ What that run actually did
Dvalin turns open-source scanner evidence into a controlled scan → fix → test → re-scan → draft-PR workflow. Here is the run in the animation above, measured:
| Real NodeGoat-derived run | Before | After Dvalin remediation |
|---|---:|---:|
| Security health (triage heuristic) | 22 / 100 · F | 100 / 100 · A |
| Findings | 10 (eval across 3 rules, 2 engines) | 0 |
| Tests | 2 passing | 3 passing, including an injection regression test |
| Scanner fleet | 4 / 4 completed | 4 / 4 completed |
The scanning and hardening control plane uses open-source components:
- Semgrep CE and its community rules for
- Trivy for filesystem vulnerabilities,
- OSV-Scanner with the open
- DvalinCode's MIT-licensed built-in rules, remediation orchestration, test and
The scanners find and rank evidence. The configured model proposes source changes; DvalinCode constrains that work, records the diff, runs project tests, re-scans the changed tree, and keeps PR publication explicit. It does not auto-merge and it does not claim that a clean scan proves the absence of bugs. Choose an open-weight model through Ollama if the repair-proposal step must also stay fully local and open; hosted model licensing depends on the provider.
You can still prove what the agent did after the fact:
dvalincode report verify # re-derive the hash chain of the last run's audit log
🧩 Integration support
Code written with an AI assistant passes through four sets of hands before it merges: the agent that writes it, the editor the developer reads it back in, the pull request that gates it, and the reviewer who has to believe the result. A security answer that exists in only one of those places is not a gate — it is a suggestion the next stage is free to ignore.
Dvalin is one MCP server and one deterministic scan behind all four, so the answer does not change depending on who asks it.
| Stage of the loop | Where you are | How Dvalin gets there | Status |
|---|---|---|---|
| Writing the code | Claude Code | dual plugin · claude mcp add · mcp-install claude-code | ✅ session verified |
| | Codex | dual plugin · codex mcp add · SARIF interop | ✅ session verified · capture below |
| | Cursor | dvalincode mcp-install cursor | ⚙️ config verified |
| | Windsurf · Zed | stdio MCP through their own settings | ⚙️ documented, unverified |
| | Any MCP client | registry io.github.arthurpanhku/dvalincode | ⚙️ published |
| Reading it back | VS Code | mcp-install vscode · extension for Problems, coverage and gate status | ✅ editor verified |
| Gating the merge | GitHub Actions | Marketplace action — findings on the diff, fix records re-derived on the runner | ✅ runs on this repository's own CI |
| | Any CI | dvalin scan . --fail-on high, SARIF out for code scanning | ✅ the exit code is the contract |
| Believing the result | anyone, offline | dvalin verify-fix record.json --trusted-key ci.pub | ✅ no workspace, no network, no Dvalin state |
| | the merge gate | reverify: true — base and head re-scanned, checks re-run on the runner, fresh record signed | ✅ covered by the test suite |
✅ means a real client was driven end to end and the tool call was observed. ⚙️ means the configuration is generated and its shape is tested, but no session has been captured. The difference is not smoothed over here, because a config file that loads is not evidence that a tool was ever called.
What each claim rests on
| Client | Version | Checked | How |
|---|---|---|---|
| Claude Code | CLI 2.1.260 | 2026-09-04 | server connected, dvalin_scan called with only that tool allow-listed, three findings returned — capture below |
| Claude Code | CLI 2.1.251 | 2026-08-31 | weekly harness-interop — a real handshake against the built binary, which needs no credentials |
| Codex | CLI 0.153.2 | 2026-09-18 | real dvalin_scan call in an ephemeral read-only sandbox; the completed MCP event was recorded as JSONL and returned one dvalin/eval finding — capture below |
| Codex | CLI 0.151.0 | 2026-08-31 | weekly harness-interop — the server spec is accepted and stored as stdio. Not a handshake: the tool-call step stays skipped until CODEX_API_KEY is set |
| VS Code | 1.134.0 · extension 0.18.0 | 2026-09-03 | packaged VSIX in a clean profile; finding, gate and coverage rendered in the editor — capture below |
| Cursor | — | 2026-09-04 | mcp-install cursor writes .cursor/mcp.json under mcpServers; no session has been captured |
That weekly workflow exists because an earlier version of this claim named two CLI versions by hand, both went stale within weeks, and nothing said so. It now re-checks against whatever those tools shipped that week, so the dates above either move on their own or stop moving in public.
The same finding, in each place
Claude Code — one read-only tool allow-listed, and the MCP call itself:
!A Claude Code session calling the Dvalin MCP server
Codex — one read-only MCP call, with the completed JSONL tool event shown separately:
!A Codex session calling the Dvalin MCP server
VS Code — the same scanner contract, as squiggles and a status bar:
!A Dvalin finding and its coverage status inside VS Code
🏛️ And it survives a security review
That last command is the part that matters once more than one person depends on this. DvalinCode is a full coding agent — terminal, web GUI, and desktop app — built so that an organization, not the developer, bounds what it may do: a policy file constrains modes, commands, paths, tools, and models; every run is hash-chained into a tamper-evident audit log; nothing reaches a provider that the egress guard did not allow. A repo policy can only ever narrow the machine-level one.
If you are the person who has to approve this class of tool, start at APPROVABILITY-PLAN.md and the Evidence Pack that every release ships of itself.
| 🏠 Home | One place for read-only Ask and approval-gated Collaborate workflows. Switch intent without leaving the project or conversation. |
| ⚡ Code | Focused autonomous coding with full tool access and Ask / Plan / Auto / Bypass permission levels. Security and browser routines no longer compete with the core coding workflow. |
| 🛡️ Dvalin | Dedicated white-box security engineering: orchestrate the built-in scanner plus installed Semgrep CE, Trivy, and OSV-Scanner; triage findings; create isolated fixes; run tests and re-scan; then explicitly publish a reviewable draft PR. Dvalin guide → |
| 🏦 Regulated teams | Designed for finance, healthcare, security-sensitive SaaS, and internal platform teams that need AI coding under policy, audit, data minimization, and supply-chain review — not just developer convenience. |
| 🛡️ Secure remediation | Run a multi-engine scan or import SARIF from CodeQL, GitHub Code Scanning, Semgrep, or compatible scanners, then create an isolated remediation worktree and turn findings into focused repair tasks with source context, verification evidence, and PR-ready reporting. Workflow → |
| 📚 Skills | Upload, download, and inspect local skill bundles. DvalinCode ships built-in secure-code-scan and secure-code-remediation skills, plus agent tools for listing skills, reading skill instructions, scanning, listing cases, and preparing remediation worktrees. Format → |
| 🛡️ Audit trail | Every run emits a tamper-evident, hash-chained JSONL log — every file read/written, every command, every approval. A Run Report renders it as Markdown; dvalincode report verify proves the chain is intact. Threat model → |
🔒 Org policy & trust | A company — not the developer — bounds the agent. A dvalin.policy.json constrains modes, shell commands, file paths, tools, and models; a repo policy can only ever narrow the machine-level one, never widen it. Each run records the governing policy's hash. dvalincode trust prints the install's live security posture — active policy + hashes, audit status, runtime — so a reviewer can verify it directly. Policy reference → · Approvability plan → |
| 🏛️ Governance evidence | OpenSSF Scorecard, CodeQL, Dependabot, pinned GitHub Actions, CODEOWNERS, and ISO/IEC 42001 AIMS alignment docs are maintained as reviewable project evidence, and every release ships an Evidence Pack the binary produced of itself. Scorecard map → · ISO 42001 alignment → · Release evidence → |
| 📐 Open specs | PCP-1 — the provider-boundary contract (egress containment, credential containment, audit, policy binding) written as a vendor-neutral profile with test procedures, so any agent runtime can run it against its own adapters and publish the result. Not a DvalinCode test file; a checklist anyone can hold us to as well. Provider Conformance Profile → |
| 🖥️ First-class GUI | Modern web UI with code highlighting, file @-references, / slash commands, Git branch indicator, live token + cost counter, multi-profile LLM config, and a dark / light / system theme switcher. |
| 🖥️ Terminal or web — one binary | Run it bare for an interactive terminal agent with streaming output, inline approvals, and red/green diffs, or dvalincode serve to host the web GUI for browser/remote use. Both frontends drive the same agent core. |
| 🖥️ Native desktop app | DvalinCode.app — a real dock application (OS-native webview, no Electron) over the same engine. On macOS the one-line installer puts it in /Applications automatically; launch it straight from Launchpad. |
| 🪶 Zero-dependency binary | Single ~25MB executable per platform. No Node, no Python, no Docker. |
| 🔐 Local-first | Sessions, config, profiles, and audit logs live in ~/.dvalincode/. .dvalincodeignore blocks the agent from reading sensitive files. AGENTS.md in your repo becomes persistent project instructions. |
| 💾 Portable & exportable | Export all local data (memory, sessions, config, audit) to one file and import it on another machine — your setup moves with you. Any conversation downloads as a clean Markdown transcript. |
🎯 Core Goal
Make every code-producing human or agent pass the same independent security gate.
DvalinCode is built as an agent-compatible security runtime, not another general coding-agent benchmark entry. The core product is scan evidence, policy, baseline, deterministic verification, and portable interfaces that a human developer, an external agent, or CI can all call. The bundled coding agent stays capable enough to implement and test focused remediations reliably; its model prose never decides whether the security gate passed.
- Any model — every OpenAI-compatible endpoint is a first-class citizen, local models included. Your workflow should never be hostage to one vendor's pricing, rate limits, or quality swings.
- Safe by default — three-tier approvals with diff preview, an undo stack, and sandboxed shell execution. An agent you can trust on full-auto.
- Small enough to audit — one ~25MB binary, a handful of runtime dependencies, a codebase you can read in a weekend. Trust through inspection, not promises. As of v0.5, every agent run is auditable too: a tamper-evident, hash-chained log of every action, verifiable after the fact.
- Open enough to embed — the agent core speaks a clean REST + WebSocket API, ready to be wired into your own product, CI, or internal tools.
- Approvable by any company — governance is built in, not bolted on. An org policy bounds the blast radius (controllable),
dvalincode trustmakes the posture self-verifiable (transparent), and the hash-chained log proves what every run did (auditable). Those three together are exactly what a security review needs to say yes — and what cloud, closed, mutable-log agents structurally struggle to provide. Approvability plan →
✅ Why Teams Pick DvalinCode
DvalinCode is differentiated by approvability. It is built for teams that need AI coding to pass security, compliance, and data-governance review before it can touch production repositories.
- Closed-loop secure remediation — scan locally or import SARIF from
dvalin/remediate/... worktree; then send a focused repair prompt with
source context and verification instructions.
- Skills as governed operating procedures — upload, download, and inspect
- Model freedom without policy drift — use DeepSeek, OpenAI, Claude via
- Security evidence, not just security claims — OpenSSF Scorecard support,
- Local-first by default — sessions, config, profiles, memory, and audit
~/.dvalincode/; .dvalincodeignore and policy controls
bound what the agent can read, write, or execute.
🛡️ Security & Governance
DvalinCode maintains project-level governance evidence for open-source and enterprise review. This is the differentiator for teams where AI coding must pass security approval before it can reach production repositories:
- Threat model — the full attack surface of an agentic coding runtime
AGENTS.md, poisoned MCP servers, prompt-injection escalation,
egress, audit tampering, supply chain, sandbox escape), each mapped to the
control that defends it and the honest residual gap. Threat model →
- OpenSSF Scorecard support — scheduled Scorecard workflow, SARIF upload,
- ISO/IEC 42001 alignment — an AI management system scope, AI policy, role
- AI change impact assessment — a reusable template for changes that affect
- Regulated-use posture — local-first data handling, policy-controlled
- Dvalin security engineering — the dedicated Dvalin workspace combines the
These documents are implementation evidence and operating procedures; they do not claim third-party ISO certification.
⭐ What's New in v0.14.0 — Dvalin security engineering
- Home unifies Chat and Cowork — the GUI now has a single Home workspace
- Code is focused again — the old Security and Routines panels have been
- Dvalin is a first-class workspace — orchestrate the built-in scanner plus
- One flow from evidence to draft PR — selected findings can launch an
- Agent loops converge sooner and cost less — investigation-before-edit and
- Provider and evaluation upgrades — native Anthropic prompt caching and
⭐ What's New in v0.12.4 — finish the task before stopping
- Process narration no longer ends a task — responses such as “let me
- Truncated responses automatically recover — provider finish reasons are
- Normal coding turns get room to finish — the per-turn action limit is now
- Completion is explicit — Code mode is instructed to return a tool-free
⭐ What's New in v0.12.3 — resilient long-running Code mode
- Long coding turns keep going — Code mode now compacts context during an
- Interruptions are resumable — completed tool state is persisted when a
continue
resumes from the actual workspace progress.
- Visible, quieter agent activity — running sessions show a sidebar loading
- GitHub workflows from Code mode — network-aware
gitand GitHub CLI
gh) operations now support pull, push, PR creation, and Actions/repository
commands through the governed shell approval path.
- Safer releases — package and CLI versions are synchronized, and
prepublishOnly runs the build, typecheck, and test suite before publishing.
- Simple tasks stay simple — the Action budget is enforced across the whole
⭐ What's New in v0.12.2 — 🖥️ Desktop app milestone: it just works
- 🖥️ The native desktop app now works out of the box on macOS —
DvalinCode.app
- 📦 The one-line installer installs the app — on macOS,
curl … install.sh | bash now also puts DvalinCode.app (with the
DvalinCode icon) into /Applications, so the desktop window launches
straight from Launchpad after a CLI install. Opt out with
DVALINCODE_NO_APP=1; pin with DVALINCODE_GUI_VERSION.
- ✅ Desktop is no longer "experimental" on macOS — the window and the
v0.9.0 — 🛡️ Secure remediation · Skills · CodeQL hardening
- 🛡️ Secure remediation workflow — run a built-in local scan or import SARIF
- 📚 Skills — upload, download, inspect, and reuse local skill bundles.
- 🔐 CodeQL path hardening — user-controlled workspace, remediation, and
- 🎨 App icons — dark and light theme application icons now ship with the web
v0.8.0 — 🔒 Governance: controllable · transparent · auditable
- 🔒 Org policy — a
dvalin.policy.jsonlets a company, not the developer, bound the agent: which modes, shell commands, file paths, tools, and models are allowed. Two layers (machine~/.dvalincode/policy.json+ repo) resolve by narrowing — a repo policy can only ever make the machine policy stricter, never widen it. With no policy file, behavior is identical to before. Enforced at a single chokepoint; every denial is an inline⛔ Blocked by policyplus apolicy_violationaudit event. Policy reference → - 🔎
dvalincode trust— prints this install's live security posture in one command — active policy + source hashes, audit status, runtime, dependencies — so a reviewer can verify what the agent may and may not do directly, instead of taking claims on trust.--jsonfor tooling. dvalincode policy check— validatesdvalin.policy.jsonagainst the schema, prints the resolved policy + canonical hash (after narrowing with the machine layer), and exits non-zero on failure — for CI and policy authoring. Policy reference →- 🧾 Policy-aware audit — every run records the hash of the governing policy (and which files contributed) in
run_start, so the tamper-evident log proves which rules were in force. - 📐 Approvability plan — the through-line is documented in docs/APPROVABILITY-PLAN.md: make DvalinCode trivially approvable by any company — controllable, transparent, auditable.
v0.7.0 — 🧪 Desktop app (beta)
- 🧠 Portable memory & full data export/import — the upgraded local memory mechanism, plus every session, config, profile, and audit log, can now be bundled into a single file and restored on another machine. Migrate your whole setup in one step:
dvalincode export/dvalincode import, or the Export / Import buttons in the GUI Settings panel. - 📝 Download any AI interaction as Markdown — every conversation can be saved as a clean Markdown transcript (user turns, assistant replies, tool calls + results, decisions — all inline). Use the download icon on any session in the sidebar,
dvalincode session md, orGET /api/sessions/:id/markdown. - 🖥️ Native desktop app — a real application window (not a browser tab) over the same engine:
DvalinCode.appon macOS, plus Windows/Linux builds. Built with webview-bun using the OS-native webview (WKWebView / WebView2 / WebKitGTK) — no Electron, stays a small self-contained binary. - 🧩 A third frontend, one core — the desktop app, terminal UI, and web GUI all drive the same shared turn-runner. The current
dvalincodebinary is now positioned purely as the CLI (terminal +serve). - Status: the desktop binaries are experimental / unverified — grab them from the latest pre-release and please report how the window behaves on your OS.
v0.6.0 — terminal agent ·
serve · shared turn-runner
- 🖥️ Terminal agent — run
dvalincodebare for an interactive terminal coding agent, Claude-Code-style: streaming responses, inline[y/N]write approvals with red/green diffs,/mode·/clear·/git·/plan·/compact·/undo·/help, Ctrl-C to interrupt, and a guided first-run provider setup. Defaults to read-only Chat, switchable live. - 🌐
dvalincode serve— the web GUI now lives behind a command, so the same binary deploys headless on a server:dvalincode serve --host 0.0.0.0 --no-open. - **🧩 One engine, two fronte