QoderAI/better-harness: An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experiments, inspect evidence, and compare outcomes. Turn task evidence into actionable team and organization insights.
Delegate coding to agents. Improve the loop around them.
Better Harness provides open-source insights for the Agent Work Loop. It runs
through your Coding Agent and turns project and session evidence into prioritized
improvements and verifiable next steps. Missing evidence stays explicit.
Choose the host you already use to get its exact installation, verification,
invocation, and report-output steps. Better Harness does not use one universal
entrypoint across every host.
This README shows inline setup for the most common hosts. Additional supported
hosts (Qwen Code, Pi, Kimi Code, WorkBuddy, and Grok) keep their steps and
boundaries in the installation guide and the
public Host Adapter Matrix; see
More adapters. README placement is a display choice, not a
support-level claim.
Better Harness scopes behavior claims to relevant Task Episodes and the
surrounding project mechanisms. Qoder and Cursor produce host-native Canvas
reports; Claude Code, Codex, Qwen Code, GitHub Copilot, and Kimi Code produce
self-contained HTML with paired Markdown. Missing or partial evidence remains
explicit. See the Host Adapter Matrix for current
coverage and output differences.
See it in action
The report keeps missing evidence explicit and turns supported gaps into
prioritized findings with an impact, expected output, scoped repair, and
acceptance checks.
For delivery tracing, the interactive Harness Inspector
follows product intent through agent activity, sessions, files, and commits in
a read-only workspace, keeping evidence strength and limitations visible:
After you have comparable reports over time, the history view shows how the five
Agent Work Loop dimensions move:
The static final frame summarizes historical Harness reports. It shows recorded
trends, not causal proof of improvement. See how the demo was recorded.
Why Better Harness?
AI coding agents change code fast, but the workflow around them is often the
weak point:
🎯 Fuzzy goals — the agent confidently solves the wrong problem.
🧭 Improvised steps — work happens on paths nobody can reproduce.
✅ "It works" without proof — validation is incomplete or missing.
🚢 Speed over safeguards — review and delivery checks get bypassed.
🧠 Lessons lost — the same friction comes back on the next task.
Reviewing only the final diff misses these system-level problems. Better Harness
analyzes the workflow around the diff: it gathers project evidence (and session
evidence where supported), evaluates five connected dimensions, and turns
concrete gaps into prioritized findings — each tied to its evidence, expected
outcome, repair boundary, and validation route, so a team can improve one issue
at a time.
How Better Harness works
Better Harness uses a
feedforward-and-feedback
loop that combines guidance available before work starts with signals available
after the agent acts:
Feedforward guides — AGENTS.md, specs, Skills, and acceptance criteria
Across that loop, it evaluates five parts of delivery — the Agent Work Loop:
| Dimension | The question it answers | Backed by |
| --- | --- | --- |
| Task Understanding | Does the agent know the goal and what "done" means? | Rules, AGENTS.md, specs, DESIGN.md |
| Controlled Execution | Is the work on supported, repeatable paths? | Skills, commands, MCP tools, sandbox boundaries |
| Change Validation | Is there evidence the change actually works? | Tests, lint, Hooks, observable diagnostics |
| Reliable Delivery | Does AI speed bypass quality checks or acceptance? | Human review, approvals, CI/CD, recovery paths |
| Learning Capture | Does the next task benefit from this one? | Loop Discovery, reusable SDLC Skills, Memory |
Running /better-harness establishes a task-bounded baseline and, depending on
the host, produces a visual report, a Markdown report, or both. The report
combines the five-part overview, prioritized findings, detected agent assets,
and an evidence brief. Each finding includes a repair action that drafts a
scoped fix plan for review.
Better Harness is deliberately honest: unobserved behavior stays explicit instead
of becoming an unsupported score or claim. Passing a current check proves that
the intervention was exercised; only a comparable later result can prove that
the loop improved.
What is open
Better Harness opens three connected layers, not only a slash-command prompt:
Engineering practices — evidence and judgment guidance across
The three layers share the same boundary: configured assets can establish that
a mechanism exists, but only linked task evidence can establish that it was used
or improved an outcome.
Architecture
The architecture keeps the three evidence domains independent until unified
analysis by the lead agent. Every result retains a visible evidence source,
owner, and validation route.
Installation
Installation differs by coding agent. Install Better Harness separately for
each host, except that Qoder CLI can use the version bundled with Qoder Desktop.
After installing or updating a plugin, start a new session or task so the host
reloads its plugin inventory.
Claude Code
Register this repository as a Claude Code marketplace:
/plugin marketplace add QoderAI/better-harness
Then install Better Harness:
/plugin install better-harness@better-harness
Verify discovery from the shell:
claude plugin details better-harness@better-harness
The details should include Skills (1) better-harness. Then start a new Claude
session in the repository you want to analyze and run the report prompt:
/better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Claude Code defaults to a self-contained report.html with paired report.md
and findings.json under the repository's .claude/better-harness report root.
Ask for inline or no-files output to keep the result in chat only. Workspace-
matching local Claude sessions are included when available; missing evidence
stays explicit rather than being inferred.
Codex
Codex Desktop
Open Settings > Plugins.
Select + Add > From Marketplace.
Enter the Git repository URL, set its Git ref, and leave Sparse paths
empty for this single-plugin repository.
Select Add marketplace, then install Better Harness from the new
marketplace.
Start a new task in the repository you want to analyze and run the report
prompt:
@better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Use https://github.com/QoderAI/better-harness.git with Git ref main.
codex plugin marketplace add \
'https://github.com/QoderAI/better-harness.git' \
--ref main
Then inspect and install Better Harness:
codex plugin list --marketplace better-harness
codex plugin add better-harness@better-harness
Start a new Codex task in the repository you want to analyze and run the report
prompt:
$better-harness:better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Use the repository URL with marketplace add, not a raw marketplace.json
URL. Current Codex builds use plugin add and --marketplace; examples that
use plugin install or --source target a different CLI contract.
Qoder
Better Harness is built into the Qoder desktop app, so no
Marketplace or local plugin installation is required there. Choose either
entry point:
From a session: Open the repository you want to analyze, start a new
session, and run the report prompt:
/better-harness analyze this project's AI coding workflow and generate an evidence-backed report
From Quest (Qoder 1.18.0+): Open Quest, then select
Better Harness (Beta) from the left sidebar.
Qoder CLI
If Qoder Desktop is installed, Better Harness is already available in Qoder
CLI. No marketplace or plugin installation is required. Start a new Qoder CLI
session in the repository you want to analyze and run the report prompt:
/better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Only when using Qoder CLI without Qoder Desktop, inspect the current manual
installation disposition before following:
Replace .qoder to .qoder-cn in urls for Qoder CN series.
Then start a new Qoder CLI session before using /better-harness.
Cursor
The Cursor plugin is not published to the marketplace. The repository carries
the source-local manifest, but the current local Cursor help does not verify the
historical --plugin-dir contract. Better Harness therefore reports the
installation plan as unavailable instead of emitting that command:
Cursor session evidence is supported through workspace-matched transcripts,
metadata, and audit logs. A session that was loaded through a separately
verified native route can be checked with better-harness plugin verify --host
cursor --surface agent; partial or unavailable coverage remains explicit.
GitHub Copilot
Register this repository as a Copilot plugin marketplace, then install Better
Harness:
Prefer marketplace installs. Direct repository, URL, and local-path installs are
deprecated in Copilot CLI.
Copilot session evidence is supported through workspace-matched Copilot CLI
transcripts under ~/.copilot/session-state/. Copilot records no per-response
token usage, and VS Code Copilot Chat has no supported durable transcript; both
remain explicit evidence boundaries.
Inspect and plan plugin lifecycle changes (Beta)
The standalone CLI can inspect local Better Harness installation evidence for
every host without contacting a registry or changing host configuration:
better-harness plugin status --host all
better-harness doctor --platform all
Build a host-specific install, update, or removal plan before using that host's
native UI or CLI. Plans preserve native steps as typed argv data for deliberate
external execution; the human view does not turn them into shell command
strings, and Better Harness does not execute them:
better-harness plugin plan install --host qwen --surface cli --scope user
better-harness plugin verify --host qwen --surface cli
Host differences remain explicit. Qoder Desktop is bundled, Cursor is
session-only while its native command contract is being reconciled, Pi
lifecycle commands without current native evidence remain manual or
unavailable, and WorkBuddy has no managed Better Harness plugin lifecycle
surface.
More adapters
Beyond the hosts above, Better Harness also supports Qwen Code, Pi, Kimi Code,
WorkBuddy, and Grok. Their exact install, invocation, and
evidence boundaries live in the docs so this README stays focused:
From a source checkout, npm run preview -- --open serves a bundled fixture.
Canvas preview requires an installed Qoder runtime, or an explicit
--sdk-media/--sdk-root path. It listens on 127.0.0.1 by default and is a
local inspection tool, not an authenticated sharing service.
Contribute
You do not need to understand the whole runtime to contribute. Start with the
smallest surface that matches the improvement you want to make:
| What you can contribute | Start here | Example contribution |
| --- | --- | --- |
| Workflow guidance and engineering practices | skills/ or references/ | Add sourced guidance for a language, framework, review pattern, or recurring agent workflow. |
| Evaluation models and executable analysis | models/ or scripts/ | Add an evidence-backed evaluation lens, detector, or agent-friendly analysis command with fixtures and tests. |
| Delivery controls and host support | hooks/ or the new Coding Agent guide | Add a narrow lifecycle check or document and validate evidence support for another Coding Agent host. |
| Reports and visual language | templates/reporting/ or templates/style/ | Add a report mode, reusable reporting contract, or directive-only visual style with validation evidence. |
| Examples and operating models | case-studies/ | Share a redacted, evidence-bounded example of how a team applies Agent Work Loop analysis and delivery practices. |
Add tests, fixtures, or preview evidence when the contribution changes
runtime behavior or rendered output.
Open a focused pull request that explains what changed, why, and how it was
validated.
Not sure where an idea belongs? Open an issue
before building a new top-level surface or changing a public report, schema,
packaging, or compatibility contract.