Profile
Back to NewsBack
GitHub Trending 31 min
Reader Mode
microsoft/waza: CLI / Framework for Agent Skills - create, test, measure and improve skill quality and effectiveness

microsoft/waza: CLI / Framework for Agent Skills - create, test, measure and improve skill quality and effectiveness

8 hours ago

Waza

A Go CLI for evaluating AI agent skills — scaffold eval suites, run benchmarks, and compare results across models.

📖 Getting Started / Docs

Installation

Binary Install (recommended)

Download and install the latest pre-built binary with the Bash install script on macOS, Linux, or Windows Bash environments such as Git Bash, MSYS2, or Cygwin:

curl -fsSL https://raw.githubusercontent.com/microsoft/waza/main/install.sh | bash

The Bash script auto-detects the OS and architecture of the environment where Bash is running (linux/darwin/windows, amd64/arm64), downloads the latest standalone waza CLI release, verifies the checksum, and installs to /usr/local/bin (or ~/bin if not writable).

For native Windows PowerShell:

irm https://raw.githubusercontent.com/microsoft/waza/main/install.ps1 | iex

The PowerShell script downloads the latest standalone native Windows waza binary, verifies the checksum, and installs to an existing waza.exe location or %LOCALAPPDATA%\Microsoft\Waza. On Windows, piping the Bash command from PowerShell may invoke WSL and install the Linux binary inside WSL.

Or browse the GitHub Releases page and choose the standalone waza binary assets for the version you want.

Install from Source

Requires Go 1.26+:

Note: due to the use of LFS artifacts you cannot install waza using go install. To install waza outside of a normal release, clone the repository:

git clone https://github.com/microsoft/waza.git
cd waza

ensure git LFS-based artifacts are available (for embedded copilot binaries)

git lfs install git lfs pull

go build -o waza ./cmd/waza ./waza <waza command line>

Waza bundles the GitHub Copilot CLI used by the copilot-sdk executor and extracts it and its runtime assets to a versioned directory in the local user cache on first use (or under COPILOT_HOME/cache when set). Set COPILOT_CLI_PATH only when you need to force a specific Copilot CLI binary. Installation failures are reported without falling back to a CLI on PATH; set COPILOT_CLI_INSTALL_VERBOSE=1 for installation diagnostics.

Azure Developer CLI (azd) Extension

Waza is also available as an azd extension:

# Add the waza extension registry
azd ext source add -n waza -t url -l https://raw.githubusercontent.com/microsoft/waza/main/registry.json

Install the extension

azd ext install microsoft.azd.waza

Verify it's working

azd waza --help

Once installed, all waza commands are available under azd waza. For example:

azd waza init my-eval --interactive
azd waza run examples/code-explainer/eval.yaml -v

Update Notifications

Waza automatically checks for new versions in the background. If an update is available, a notice appears after command output:

A newer version of waza is available: v0.24.0 → v0.28.0. Run: waza update

Run waza update to download and execute the official OS-specific installer after an explicit confirmation prompt. It uses the Bash installer on macOS/Linux and the PowerShell installer on native Windows. Use waza update --yes to skip the prompt in scripted environments. The check is non-blocking (never slows commands), cached for 24 hours, and can be disabled with --no-update-check or WAZA_NO_UPDATE_CHECK=1.

Quick Start

For New Users: Get Started in 5 Minutes

See Getting Started Guide for a complete walkthrough:

# Initialize a new project
waza init my-project && cd my-project

Create a new skill

waza new skill my-skill

Define the skill in skills/my-skill/SKILL.md

Write evaluation tasks in evals/my-skill/tasks/

Add test fixtures in evals/my-skill/fixtures/

Run evaluations

waza run my-skill

Check skill readiness

waza check my-skill

All Commands

# Build
make build

Initialize a project workspace

waza init [directory]

Update waza to the latest release

waza update

Create a new skill

waza new skill skill-name

Create a new eval scaffold from an existing SKILL.md

waza new eval skill-name

Generate a task YAML by recording a prompt run

waza new task from-prompt "Explain this code and suggest fixes" evals/code-explainer/tasks/recorded-task.yaml

Check if a skill is ready for submission

waza check skills/my-skill

Suggest an eval suite from SKILL.md

waza suggest skills/my-skill --dry-run waza suggest skills/my-skill --apply

Discover shared registry graders and add one to an eval

waza registry search factual --kind grader waza registry add github.com/waza-evals/fact#[email protected] --eval eval.yaml --name factuality

Verify eval coverage against SKILL.md requirements

waza spec verify skills/my-skill evals/my-skill/eval.yaml waza spec verify skills/my-skill evals/my-skill/eval.yaml --fail --format github-actions

Resolve remote grader refs and write waza.lock

waza get evals/my-skill/eval.yaml

Note: 'generate' is available as an alias for 'new' (see below for new command)

Note: Custom agents (.agent.md) are supported — see https://microsoft.github.io/waza/guides/custom-agents/

Run evaluations (works with both skills and custom agents)

waza run examples/code-explainer/eval.yaml --context-dir examples/code-explainer/fixtures -v

Grade output from a previous waza run --output results.json ...

waza grade eval.yaml --results results.json

Compare results across models

waza compare results-gpt4.json results-sonnet.json

Capture snapshots during a run and replay them later for determinism checks

waza run eval.yaml --snapshot ./snapshots/ waza replay ./snapshots/my-task-run1.json

Run offline adversarial / fault-injection packs against a skill

waza adversarial --list-packs waza adversarial --skill ./skills/my-skill --model gpt-4o

Check whether a schema artifact needs migration

waza migrate eval.yaml

Generate eval coverage grid

waza coverage --format markdown

Count tokens in skill files

waza tokens count skills/

Compare skill token budgets vs main

waza tokens compare main --skills --threshold 10

Suggest token optimizations

waza tokens suggest skills/

Agent-assisted eval authoring

Point your coding agent at the Writing Eval Specs guide to create or update evals and guide eval-driven implementation. It provides the canonical workflow, target-specific validation steps, and a reusable delegation prompt.

Commands

waza update

Update waza to the latest release by running the official OS-specific installer after confirmation.

| Flag | Description | |------|-------------| | --yes, -y | Skip the confirmation prompt |

Example:

waza update
waza update --yes

waza init [directory]

Initialize a waza project workspace with separated skills/ and evals/ directories. Idempotent — creates only missing files.

| Flag | Description | |------|-------------| | --no-skill | Skip the first-skill creation prompt |

Creates:

  • skills/ — Skill definitions directory
  • evals/ — Evaluation suites directory
  • .github/workflows/eval.yml — CI/CD pipeline for running evals on PR
  • .gitignore — Waza-specific exclusions
  • README.md — Getting started guide for your project
Example:
waza init my-project

Optionally creates first skill interactively

waza init my-project --no-skill

Skip skill creation prompt

waza new skill

Create a new skill with scaffolded structure and evaluation suite. Detects workspace context and adapts output. In interactive mode, the wizard collects spec-aligned metadata: name, description, trigger phrases, and anti-trigger phrases.

| Flag | Short | Description | |------|-------|-------------| | --template | -t | Template pack (coming soon) |

Modes:

Project mode (detects skills/ directory):

project/
├── skills/{skill-name}/SKILL.md
└── evals/{skill-name}/
    ├── eval.yaml                 # or files.evalFile
    ├── tasks/*.yaml              # or files.taskGlob / files.taskFileSuffix
    └── fixtures/

APM-managed skills are detected from their compiled output without symlinks:

project/
├── skills/{skill-name}/apm.yml
├── skills/{skill-name}/.apm/skills/{skill-name}/SKILL.md
└── skills/{skill-name}/eval.yaml

When both skills/{skill-name}/SKILL.md and the APM compiled .apm/skills/{skill-name}/SKILL.md exist for the same skill, the top-level SKILL.md takes precedence.

Standalone mode (no skills/ detected):

{skill-name}/
├── SKILL.md
├── evals/
│   ├── eval.yaml                 # or files.evalFile
│   ├── tasks/*.yaml              # or files.taskGlob / files.taskFileSuffix
│   └── fixtures/
├── .github/workflows/eval.yml
├── .gitignore
└── README.md

Example:

# In project mode (explained Modes section, above): creates skills/code-explainer/SKILL.md + evals/code-explainer/
waza new skill code-explainer

In standalone mode (explained Modes section, above): creates code-explainer/ self-contained directory

waza new skill code-explainer

waza new eval

Scaffold an eval suite from an existing SKILL.md (reads frontmatter trigger hints from USE FOR and DO NOT USE FOR).

Creates:

  • evals//
  • evals//tasks/positive-trigger-1
  • evals//tasks/positive-trigger-2
  • evals//tasks/negative-trigger-1
| Flag | Description | |------|-------------| | --output | Custom path for the eval file (tasks are generated under sibling tasks/) |

Generated eval and task filenames are configurable in .waza.yaml:

files:
  evalFile: waza-eval.yaml
  taskGlob: tasks/*.waza-task.yaml
  taskFileSuffix: .waza-task.yaml

Example:

# Default output location
waza new eval code-explainer

Custom eval path

waza new eval code-explainer --output evals/custom-code-explainer/eval.yaml

waza new task from-prompt

Run a prompt through Copilot and generate a task YAML with inferred validators based on observed behavior (response text, tool usage, and invoked skills).

| Flag | Description | |------|-------------| | --model | Copilot model to run for recording (default: claude-sonnet-4.5) | | --testname | Test name and ID written into the generated task (default: auto-generated-test) | | --tags | Comma-separated tags to attach to the generated task | | --timeout | Max time for prompt execution (default: 5m) | | --overwrite | Overwrite the output task file if it already exists | | --root

| Root directory used for skill discovery (default: .) |

Example:

# Record a prompt and generate a reusable task YAML
waza new task from-prompt "Refactor this function for readability" evals/code-explainer/tasks/refactor-readability.yaml

Add metadata and overwrite an existing file

waza new task from-prompt "Explain this diff and risks" evals/code-explainer/tasks/diff-analysis.yaml \ --testname diff-analysis \ --tags recorded,regression \ --overwrite

waza run

Run an evaluation benchmark from a spec file.

| Flag | Short | Description | |------|-------|-------------| | --context-dir

| | Fixture directory (default: ./fixtures relative to spec) | | --output | -o | Save results to JSON | | --output-dir | | Directory for structured output; each run creates a UTC-timestamped subdirectory of . Mutually exclusive with --output. | | --verbose | -v | Detailed progress output | | --transcript-dir | | Save per-task transcript JSON files | | --task | | Filter tasks by name/ID pattern (repeatable) | | --parallel | | Run tasks concurrently | | --workers | | Concurrent workers (default: auto, requires --parallel) | | --trials | | Run each task n times to detect flakiness (omit to use config.trials_per_task; if provided, n must be >= 1) | | --interpret | | Print plain-language result interpretation | | --format | | Output format: default or github-comment (default: default) | | --cache | | Enable result caching to speed up repeated runs | | --no-cache | | Explicitly disable result caching | | --cache-dir | | Cache directory (default: .waza-cache) | | --reporter | | Output reporters: json (default), junit: (repeatable) | | --baseline | | A/B testing mode — runs each task twice (without skill = baseline, with skill = normal) and computes improvement scores | | --discover | | Auto skill discovery — walks directory tree for SKILL.md + eval.yaml (root/tests/evals) | | --strict | | Fail if any SKILL.md lacks eval coverage (use with --discover) | | --suggest | | Generate a Copilot suggestion report based on test outcomes (mock engine emits a deterministic fake report) | | --output-dir | | Directory for structured output; each run creates a UTC timestamped subdirectory. Mutually exclusive with --output. | | --tags | | Filter tasks by tags, using glob patterns (repeatable) | | --model | | Override model (repeatable for multi-model comparison) | | --recommend | | Generate heuristic recommendation after multi-model run | | --judge-model | | Model for LLM-as-judge graders (overrides execution model) | | --session-log | | Enable session event logging (NDJSON) | | --session-dir | | Directory for session log files (default: current directory) | | --no-summary | | Skip writing combined summary.json for multi-skill runs | | --update-snapshots | | Update or create diff grader snapshot files to match current output | | --skip-graders | | Skip grading (execution only); grade later with waza grade | | --keep-workspace | | Preserve temp workspaces after execution for debugging | | --auto-file-issue | | Auto-file or update a GitHub issue for failing runs (requires gh and GITHUB_REPOSITORY) | | --otel-exporter | | Export OpenTelemetry traces using otlp, stdout, or file. Off by default. See OpenTelemetry Tracing. | | --otel-endpoint | | OTLP endpoint (host:port or URL); only used with --otel-exporter=otlp | | --otel-headers | | Comma-separated key=value OTLP headers (e.g. for auth) | | --otel-file | | File path for span JSON when --otel-exporter=file | | --otel-include-payloads | | Include prompt/tool-arg/tool-result/completion content in spans (default: redacted to sha256+length) | | --snapshot | | Capture self-contained snapshot.json per task for later waza replay. | | --snapshot-env-allow | | Allow-list of env var name patterns embedded in snapshots (default-deny; supports WAZA_* wildcards). | | --redact | | YAML redaction policy applied to snapshot output (merged with built-in defaults). |

Result Caching

Enable caching with --cache to store test results and skip re-execution on repeated runs:

# First run executes all tests and caches results
waza run eval.yaml --cache

Second run uses cached results (much faster)

waza run eval.yaml --cache

Clear the cache when needed

waza cache clear

Cached results are automatically invalidated when:

  • Spec configuration changes (model, timeout, graders, etc.)
  • Task definitions change
  • Fixture files change
Note: Caching is automatically disabled for evaluations using non-deterministic graders (behavior, prompt).

waza get [eval.yaml | ref]

Resolve remote grader refs and write waza.lock. When passed an eval file, waza get resolves every graders[].ref, downloads module contents into ~/.waza/cache/{host}/{org}/{repo}/{sha}/, and pins each ref to a commit SHA and sha256: content digest. waza run requires a valid lock and cache entry for remote refs; it does not silently resolve unlocked refs during a run.

waza get eval.yaml
waza get github.com/waza-evals/fact#[email protected]

Exit Codes

The run command uses exit codes to enable CI/CD integration:

| Exit Code | Condition | Description | |-----------|-----------|-------------| | 0 | Success | All tests passed | | 1 | Test failure | One or more tests failed validation | | 2 | Configuration error | Invalid spec, missing files, or runtime error |

Example CI usage:

# Fail the build if any tests fail
waza run eval.yaml || exit $?

Capture specific exit codes

waza run eval.yaml EXIT_CODE=$? if [ $EXIT_CODE -eq 1 ]; then echo "Tests failed - check results" elif [ $EXIT_CODE -eq 2 ]; then echo "Configuration error" fi

Post results as PR comment (GitHub Actions)

waza run eval.yaml --format github-comment > comment.md gh pr comment $PR_NUMBER --body-file comment.md

Generate JUnit XML for CI test reporting

waza run eval.yaml --reporter junit:results.xml

Both JSON output and JUnit XML

waza run eval.yaml -o results.json --reporter junit:results.xml

Note: waza generate is an alias for waza new. Both commands support the same functionality with the --output-dir flag for specifying custom output locations.

waza compare [files...]

Compare results from multiple evaluation runs side by side — per-task score deltas, pass rate differences, and aggregate statistics.

| Flag | Short | Description | |------|-------|-------------| | --format | -f | Output format: table or json (default: table) |

waza replay

Replay a task snapshot to verify deterministic reproduction. Snapshots are produced by waza run --snapshot

and capture the prompt, fixture digests, ordered tool events, environment allow-list, and redacted grader outcomes.

# Capture during a run
waza run eval.yaml --snapshot ./snapshots/

Re-check internal consistency (offline, fast)

waza replay ./snapshots/my-task-run1.json

Bisect two snapshots and find first divergent turn

waza replay ./snapshots/a.json --bisect ./snapshots/b.json --json

| Flag | Description | |------|-------------| | --mode | Replay mode: model-replay (default, offline consistency check) or live (planned) | | --bisect | Path to second snapshot to bisect against the primary | | --json | Emit machine-readable JSON instead of human summary | | --strict | Re-check final status and grader outcome consistency (default true) |

Exit codes: 0 match, 1 divergence, 2 load/parse error.

waza adversarial

Run offline adversarial / fault-injection packs against a skill. Two built-in packs ship with the binary: prompt-injection (probes resistance to indirect prompt injection via fixture files) and scope-bypass (probes refusal of out-of-scope actions like sending email, deleting files, or installing packages). Every task is golden: true, so unsafe outcomes also flip waza gate to exit 2.

# List built-in packs
waza adversarial --list-packs

Run every pack against a skill

waza adversarial --skill ./skills/code-review --model gpt-4o

Read pack selection from eval.yaml (schema 1.2 adversarial: block)

waza adversarial --spec eval.yaml --output adversarial.json

Non-blocking CI smoke

waza adversarial --packs prompt-injection --on-unsafe-outcome warn

| Flag | Description | |------|-------------| | --packs | Comma-separated pack names (default: every built-in pack) | | --list-packs | Print the pack catalog and exit | | --spec | Inherit adversarial: block from an eval.yaml | | --on-unsafe-outcome | fail (exit 2, default) or warn (exit 0) | | --engine, --skill, --model | Forwarded to the underlying engine | | --output | Write the full results.json to a file |

Exit codes: 0 all packs PASSED, 2 unsafe outcome with policy=fail (matches waza gate), 3 config error. See the Adversarial harness guide for details.

waza migrate

Check a public schema artifact and migrate it to the current schema version when a future major schema requires it. The current schema is 1.4, so v1 eval.yaml and results.json files are compatible and no file changes are made.

waza migrate eval.yaml
waza migrate results.json

waza coverage [root]

Generate a skill-to-eval coverage grid showing which skills are fully covered, partially covered, or missing evals.

Note: Full coverage requires tasks (via tasks: or tasks_from:) and 2+ grader types. The coverage percentage reflects only fully covered skills.

| Flag | Short | Description | |------|-------|-------------| | --format | -f | Output format: text, markdown, or json (default: text) | | --path

| | Additional directory to scan for skills/evals (repeatable) |

waza spec verify [skill-path] [eval.yaml]

Verify that an eval suite exercises the promises made in SKILL.md. The command deterministically parses the description, USE FOR triggers, DO NOT USE FOR triggers, and parameter blocks into requirement IDs such as req-use-001 and req-dont-001, then maps each requirement to matching task IDs.

| Flag | Description | |------|-------------| | --skill | Path to SKILL.md or a skill directory | | --eval | Path to eval.yaml | | --format | Output format: human, json, or github-actions | | --warn | Report uncovered requirements and exit 0 (default); set false to suppress GitHub Actions warning annotations | | --fail | Exit 1 when uncovered requirements are greater than or equal to --threshold | | --threshold | Uncovered requirement threshold for --fail (default: 1) | | --semantic | Opt in to LLM-assisted semantic matching after deterministic matching | | --judge-model | Judge model for --semantic (defaults to config.judge_model, then config.model) |

Example:

waza spec verify skills/pr-summarizer evals/pr-summarizer/eval.yaml
waza spec verify --skill skills/pr-summarizer --eval evals/pr-summarizer/eval.yaml --format json

waza models

List models available for evaluation via the Copilot SDK. Shows model IDs and metadata that can be used with --model flags in waza run, waza quality, and other commands.

Requires authentication via copilot login. Custom provider configuration only applies when creating or resuming Copilot SDK sessions.

| Flag | Description | |------|-------------| | --json | Output as JSON |

Examples:

# List available models in table format
waza models

Output available models as JSON

waza models --json

waza cache clear

Clear all cached evaluation results to force re-execution on the next run.

| Flag | Description | |------|-------------| | --cache-dir

| Cache directory to clear (default: .waza-cache) |

waza registry search

Search configured registry indexes for reusable graders, eval bundles, and datasets. The default public registry source is https://github.com/waza-evals; project-level .waza.yaml can override sources with a top-level registries: list.

Registry search currently returns bundled sample metadata while live index integration is pending.

| Flag | Description | |------|-------------| | --kind | Filter by grader, eval, or dataset | | --registry | Search only the named registry source | | --format | Output table or json (default: table) |

waza registry add

Append a remote grader preset reference to eval.yaml, resolve it with the same remote grader resolver as waza get, and update waza.lock with the pinned commit SHA and sha256: content digest.

| Flag | Description | |------|-------------| | --eval | Eval file to update (default: eval.yaml) | | --name | Local alias for the grader | | --set key=value | Add a local override, repeatable (for example, --set config.threshold=0.9) | | --allow-exec | Allow remote program graders without interactive confirmation |

waza dev [skill-path]

Iteratively score and improve skill frontmatter in a SKILL.md file.

Use --copilot for a non-interactive, single-pass markdown report that:

  1. Summarizes current skill details and token usage
  2. Loads trigger test prompts as examples (when trigger_tests.yaml exists)
  3. Requests Copilot suggestions for improving skill selection
  4. Prints the report to stdout without applying any changes
When --copilot is set, iterative mode flags (--target, --max-iterations, --auto) are invalid.

| Flag | Description | |------|-------------| | --target | Target adherence level for iterative mode: low, medium, medium-high, high (default: medium-high) | | --max-iterations | Maximum improvement iterations for iterative mode (default: 5) | | --auto | Apply improvements without prompting in iterative mode | | --copilot | Generate a non-interactive markdown report with Copilot suggestions | | --model | Model to use with --copilot |

waza check [skill-path]

Check if a skill is ready for submission with a comprehensive readiness report.

Performs five types of checks:

  1. Compliance scoring — Validates frontmatter adherence (Low/Medium/Medium-High/High)
  2. Token budget — Checks if SKILL.md is within token limits (configurable in .waza.yaml tokens.limits)
  3. Evaluation suite — Checks for the presence of eval.yaml
  4. Spec compliance — Validates the skill against the agentskills.io spec (frontmatter structure, required fields, naming rules, directory match, description length, compatibility, license, and version)
  5. Advisory checks — Detects quality and maintainability issues (reference module count, complexity classification, negative delta risk patterns, procedural content, and over-specificity)
Provides a plain-language summary and actionable next steps to improve the skill.

Example output:

🔍 Skill Readiness Check
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Skill: code-explainer

📋 Compliance Score: High ✅ Excellent! Your skill meets all compliance requirements.

📊 Token Budget: 450 / 500 tokens ✅ Within budget (50 tokens remaining).

🧪 Evaluation Suite: Found ✅ eval.yaml detected. Run 'waza run eval.yaml' to test.

📐 Spec Compliance (agentskills.io) ✅ spec-frontmatter Frontmatter structure valid with required fields ✅ spec-allowed-fields All frontmatter fields are spec-allowed ✅ spec-name Name follows spec naming rules ✅ spec-dir-match Directory name matches skill name ✅ spec-description Description is valid ✅ spec-license License field present ✅ spec-version metadata.version present

🔬 Advisory Checks ✅ module-count Found 2 reference modules (2-3 is optimal) ✅ complexity Complexity: detailed (350 tokens, 2 modules) ✅ negative-delta-risk No negative delta risk patterns detected ✅ procedural-content Description contains procedural language ✅ over-specificity No over-specificity patterns detected

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 📈 Overall Readiness ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

✅ Your skill is ready for submission!

🎯 Next Steps ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

✨ No action needed! Your skill looks great.

Consider: • Running 'waza run eval.yaml' to verify functionality • Sharing your skill with the community

Usage:

# Check current directory
waza check

Check specific skill

waza check skills/my-skill

Suggested workflow

waza check skills/my-skill # Check readiness waza dev skills/my-skill # Improve compliance if needed waza check skills/my-skill # Verify improvements

waza quality

Use an LLM-as-Judge to evaluate skill content quality across five dimensions: clarity, completeness, trigger precision, scope coverage, and anti-patterns.

| Flag | Description | |------|-------------| | --model | Model to use as judge (default: project default model) | | --format table\|json | Output format (default: table) | | --rubric | Path to custom rubric file (reserved for future use) |

Examples:

# Evaluate skill quality (table output)
waza quality skills/code-explainer

JSON output for CI integration

waza quality skills/code-explainer --format json

Use a specific model as judge

waza quality skills/code-explainer --model gpt-4o

waza suggest

Use an LLM to analyze SKILL.md and generate suggested evaluation artifacts.

| Flag | Description | |------|-------------| | --model | Model to use for suggestions (default: project default model) | | --dry-run | Print suggested output to stdout (default) | | --apply | Write files to disk | | --force | Allow --apply to overwrite existing eval/task/fixture files (requires --apply) | | --count | Generate exactly N tasks (default: at least 3 + 1 negative) | | --focus | Steer generation toward one of triggers, negative-triggers, edge-fixtures, do-not-use-for, parameters | | --output-dir

| Output directory (default: /evals) | | --format yaml\|json | Output format (default: yaml) |

Each generated task entry carries a confidence score in [0, 1] and a rationale string pointing to the SKILL.md span it was derived from. Both appear in dry-run output but are kept outside the written task YAML (the task schema rejects unknown fields).

Before writing files, --apply derives a missing or blank string name from the task id (for example, summarize-document becomes Summarize Document). Explicit nonblank names are preserved. Null and other invalid name values still fail task schema validation; dry-run output is unchanged.

--apply is merge-safe: an existing eval.yaml is never overwritten (new task files are picked up by its existing tasks: glob), and existing task files (by path or by task id) cause --apply to fail with a diff unless --force is also passed. Existing fixture files are also preserved unless --force is set.

Examples:

# Preview generated eval/task/fixture files as YAML
waza suggest skills/code-explainer --dry-run

Generate exactly 5 negative-trigger tasks and merge them into the existing suite

waza suggest skills/code-explainer --focus negative-triggers --count 5 --apply

Write generated files to disk

waza suggest skills/code-explainer --apply

Overwrite a previously generated suite

waza suggest skills/code-explainer --apply --force

Print JSON-formatted suggestion payload

waza suggest skills/code-explainer --format json

waza tokens count [paths...]

Count tokens in markdown files. Paths may be files or directories (scanned recursively for .md/.mdx).

| Flag | Description | |------|-------------| | --format | Output format: table or json (default: table) | | --sort | Sort by: tokens, name, or path (default: path) | | --min-tokens | Filter files below n tokens | | --no-total | Hide total row in table output |

waza tokens compare [refs...]

Compare markdown token counts between git refs.

With no arguments, compares HEAD to the working tree. With one ref, compares that ref to the working tree. With two refs, compares the first ref to the second.

| Flag | Description | |------|-------------| | --format | Output format: table or json (default: table) | | --show-unchanged | Include unchanged files in output | | --strict | Exit with code 1 if any file exceeds its absolute token limit | | --skills | Only compare SKILL.md files under configured skill roots | | --threshold | Fail when any existing file increases by more than n percent (0 = disabled) |

Use --skills to restrict comparison to SKILL.md files under configured skill roots (skills/, .github/skills/, APM .apm/skills/ outputs, and paths.skills from .waza.yaml). In skills mode the default base ref is origin/main (falling back to main).

Use --threshold for CI gating — newly added files are exempt from threshold checks (no baseline) but still subject to absolute limit checks with --strict.

# Compare all markdown tokens between HEAD and working tree
waza tokens compare

Skill-aware comparison vs main with CI threshold

waza tokens compare main --skills --threshold 10

JSON output for CI pipelines

waza tokens compare main --skills --threshold 10 --strict --format json

waza tokens profile [skill-name | path]

Structural analysis of SKILL.md files — reports token count, section count, code block count, and workflow step detection with a one-line summary and warnings.

| Flag | Description | |------|-------------| | --format | Output format: text or json (default: text) | | --tokenizer | Tokenizer: bpe or estimate (default: bpe) |

Example output:

📊 my-skill: 1,722 tokens (detailed ✓), 8 sections, 4 code blocks
   ⚠️  no workflow steps detected

waza tokens suggest [paths...]

Suggest ways to reduce token usage in markdown files. Paths may be files or directories (scanned recursively for .md/.mdx).

| Flag | Description | |------|-------------| | --format | Output format: text or json (default: text) | | --min-savings | Minimum estimated token savings for heuristic suggestions | | --copilot | Enable Copilot-powered suggestions | | --model | Model to use with --copilot |

waza serve

Start the waza dashboard server to visualize evaluation results. The HTTP server opens in your browser automatically and scans the specified directory for .json result files.

Optionally, run a JSON-RPC 2.0 server (for IDE integration) instead of the HTTP dashboard using the --tcp flag.

| Flag | Default | Description | |------|---------|-------------| | --port | 3000 | HTTP server port | | --no-browser | false | Don't auto-open the browser | | --results-dir

| . | Directory to scan for result files | | --tcp | (off) | TCP address for JSON-RPC (e.g., :9000); defaults to loopback for security | | --tcp-allow-remote | false | Allow TCP binding to non-loopback addresses (⚠️ no authentication) |

Examples:

Start the HTTP dashboard on port 3000:

waza serve

Start the HTTP dashboard on a custom port and scan a results directory:

waza serve --port 8080 --results-dir ./results

Start the dashboard without auto-opening the browser:

waza serve --no-browser

Start a JSON-RPC server for IDE integration:

waza serve --tcp :9000

Dashboard Views:

The dashboard displays evaluation results with:

  • Task-level pass/fail status
  • Raw resolved task prompts with JSON formatting and copy-to-clipboard
  • Score distributions across trials
  • Model comparisons
  • Aggregated metrics and trends
For detailed documentation on the dashboard and result visualization, see docs/GUIDE.md.

waza results

Manage evaluation results stored in cloud or local storage.

waza results list

List all evaluation runs from configured cloud storage or local results directory.

| Flag | Description | |------|-------------| | --limit | Maximum results to display (default: 20) | | --format | Output format: table or json (default: table) |

# List recent results
waza results list

List with custom limit

waza results list --limit 20

Output as JSON

waza results list --format json

waza results compare

Compare two evaluation runs side by side. Displays per-task score deltas, pass rate differences, and key metrics.

| Flag | Description | |------|-------------| | --format | Output format: table or json (default: table) |

# Compare two runs
waza results compare run-20250226-001 run-20250226-002

Output as JSON for further processing

waza results compare run-20250226-001 run-20250226-002 --format json

waza grade

Run graders against agent output without executing an agent. Designed for standalone grading of previous eval runs.

| Flag | Description | |------|-------------| | --task | Task ID to grade | | --results | Path to waza run output JSON | | --workspace

| Agent workspace directory for file-based graders; must point to the agent's actual workspace (default: .) | | --judge-model | Model for prompt graders | | -o, --output | Write full EvaluationOutcome JSON (compatible with waza compare) | | -v, --verbose | Verbose output |

waza run eval.yaml --output results.json
waza grade eval.yaml --results results.json

waza session list

List session event logs in a directory.

| Flag | Description | |------|-------------| | --dir

| Directory to search for session logs (default: .) |

waza session list
waza session list --dir ./sessions

waza session view

Render a session timeline from an NDJSON event log.

waza session view session-2025-06-15.ndjson

Cloud Storage

Waza can automatically upload evaluation results to Azure Blob Storage for team collaboration and historical tracking.

Waza pins Azure Blob requests to service version 2026-10-06 to avoid the azblob v1.8.1 rollout issue that can cause 400 InvalidHeaderValue with the SDK's newer default.

Configuration

Add a storage: section to your .waza.yaml:

storage:
  provider: azure-blob
  accountName: "myteamwaza"
  containerName: "waza-results"
  enabled: true

| Field | Description | Required | |-------|-------------|----------| | provider | Cloud provider (azure-blob currently supported) | Yes | | accountName | Azure Storage account name | Yes | | containerName | Blob container name (default: waza-results) | No | | enabled | Enable/disable uploads (default: true when configured) | No |

Authentication

Waza uses DefaultAzureCredential — it automatically detects and uses available credentials in this order:

  1. Environment variables (AZURE_CLIENT_ID, AZURE_CLIENT_SECRET, AZURE_TENANT_ID)
  2. Managed Identity (on Azure services)
  3. Azure CLI (az login)
  4. Visual Studio Code (if signed in)
  5. Azure PowerShell (if signed in)
In most cases, running az login is all you need:
az login
waza run eval.yaml  # Results auto-upload to Azure Storage

How It Works

  1. Auto-upload on run: When storage: is configured, waza run automatically uploads results to Azure Blob Storage
  2. Organized by skill: Results are stored as {skill-name}/{run-id}.json
  3. Local copy kept: Results are also saved locally (via -o flag)
  4. List remote results: Use waza results list to browse uploaded runs
  5. Compare runs: Use waza results compare to diff two remote results

Example Workflow

# Configure once (edit .waza.yaml)
cat > .waza.yaml <<EOF
storage:
  provider: azure-blob
  accountName: "myteamwaza"
  containerName: "waza-results"
  enabled: true
EOF

Authenticate

az login

Run evaluations — results auto-upload

waza run evals/my-skill/eval.yaml -v

Browse uploaded results

waza results list

Compare two runs

waza results compare run-id-1 run-id-2

For step-by-step setup and troubleshooting, see Getting Started with Azure Storage guide.

Building

make build          # Compile binary to ./waza
make test           # Run tests with coverage
make lint           # Run golangci-lint
make fmt            # Format code and tidy modules
make install        # Install to GOPATH

Project Structure

cmd/waza/              CLI entrypoint and command definitions
  tokens/              Token counting subcommand
internal/
  config/              Configuration with functional options
  execution/           AgentEngine interface (mock, copilot)
  graders/             Validator registry and built-in graders
  metrics/             Scoring metrics
  models/              Data structures (EvalSpec, TestCase, EvaluationOutcome)
  orchestration/       EvalRunner for coordinating execution
  reporting/           Result formatting and output
  transcript/          Per-task transcript capture
  wizard/              Interactive init wizard
examples/              Example eval suites
skills/                Example skills

Eval Spec Format

name: my-eval
skill: my-skill
schemaVersion: "1.2"
version: "1.0"

config: trials_per_task: 3 max_attempts: 3 # Retry failed graders up to 3 times (default: 1, no retries) timeout_seconds: 300 parallel: false executor: copilot-sdk # required when using reasoning_effort/judge_reasoning_effort model: claude-sonnet-4-20250514 reasoning_effort: high # Optional; copilot-sdk task/responder sessions only judge_model: gpt-5-mini judge_reasoning_effort: low # Optional default for prompt graders group_by: model # Group results by model (or other dimension) instruction_files: - .github/instructions/project.instructions.md

Custom input variables available as {{.Vars.key}} in tasks and hooks

inputs: api_version: v2 environment: production max_retries: 3

hooks: before_run: - command: "echo 'Starting evaluation'" working_directory: "." exit_codes: [0] error_on_fail: false

after_run: - command: "echo 'Evaluation complete'" working_directory: "." exit_codes: [0] error_on_fail: false

before_task: - command: "echo 'Running task: {{.TaskName}}'" working_directory: "." exit_codes: [0] error_on_fail: false

after_task: - command: "echo 'Task {{.TaskName}} completed'" working_directory: "." exit_codes: [0] error_on_fail: false

mcp_mocks: - name: github tools: list_issues: input_schema: type: object required: [owner, repo] responses: - match: owner: microsoft repo: waza return: issues: - number: 363 title: MCP server mocks for hermetic eval - match_regex: repo: "^waza-.*" return: issues: []

graders: - ref: github.com/waza-evals/fact#[email protected] name: factuality_strict weight: 2.0 config: threshold: 0.9

- type: text name: pattern_check config: regex_match: ["\\d+ tests passed"]

- type: behavior name: efficiency config: max_tool_calls: 20 max_duration_ms: 300000

- type: action_sequence name: workflow_check config: matching_mode: in_order_match expected_actions: ["bash", "edit", "report_progress"]

Task definitions: glob patterns or CSV dataset

tasks: - "tasks/*.yaml"

Optional: Generate tasks from CSV dataset

tasks_from: ./test-cases.csv

range: [1, 10] # Only include rows 1-10 (0-indexed, skips header)

Pin reasoning_effort and judge_reasoning_effort to low, medium, high, xhigh, or max when benchmarking model-and-effort combinations. Both settings require executor: copilot-sdk. Omit either setting to preserve the Copilot SDK/model default. A prompt grader can override the judge default with graders[].config.reasoning_effort, including task/checkpoint graders and graders with continue_session: true, which resume the task session with the overridden or default judge effort. Agent effort remains eval-level, not per-task.

With explicit effort, choose a concrete model from waza models. Waza checks the runtime's supported-effort metadata before creating or resuming hosted Copilot sessions; unknown models, unavailable metadata, and unsupported efforts produce actionable errors instead of silently using a different effort. Custom-provider efforts are forwarded directly because the hosted catalog does not describe those models. Result setup metadata and cache keys include both eval-level effort settings; regrading replaces the judge effort, including clearing a previously pinned value when omitted.

schemaVersion uses MAJOR.MINOR format. Missing values are interpreted as the current schema version (currently 1.4). Readers allow same-major minor additions with warnings for unknown fields, but reject different majors with a hint to run waza migrate .

Remote grader refs use Go-module-style paths: //[/path][#export]@. The remote module must provide a waza.registry.yaml manifest and export a grader preset. Config-only grader presets expand to built-in grader types by default; remote program graders require explicit trust with waza registry add --allow-exec or interactive confirmation. Run waza get eval.yaml after manually adding or changing refs so waza.lock records the resolved commit and digest.

results.json is currently emitted at schemaVersion 1.4. Version 1.1 added per-turn checkpoints (runs[].checkpoints[], see #358) and the normalized runs[].tool_events[] array (turn, sequence, tool_call_id, tool_name, args, result, success, error, duration_ms; see #366). Version 1.2 added runs[].snapshot_path for waza run --snapshot artifacts (#367) and the eval-level adversarial: block consumed by waza adversarial --spec (#365). Version 1.3 adds session_digest.tool_policy_mode and tool_policy_denials (#585). Version 1.4 adds sanitized runs[].command_invocations for declarative CLI mocks (#634). See docs/PRD and schema-changes for details.

For custom .agent.md targets, copilot-sdk enforces the selected agent's tools: declaration on initial and resumed turns: omitted means unrestricted, [] denies all tools, and a populated list allows only named tools. Runtime enforcement and the implicit tool_constraint grader share built-in aliases such as fileRead/readFile/view. Denials fail the run and appear in results, --session-log run events, and the dashboard trajectory digest. This is a tool boundary, not host filesystem or network sandboxing. See custom agent policies for MCP names, task overrides, and limitations.

MCP Mock Servers

Use top-level mcp_mocks with schemaVersion: "1.1" for deterministic Copilot SDK evals that need MCP tools without live services. Waza launches each mock as a local stdio MCP server, so CI runs do not need network ports, external credentials, or real service state. Waza exposes every tool declared by each mock to the Copilot CLI automatically; do not add a separate tools allowlist.

schemaVersion: "1.1"
mcp_mocks:
  - name: github
    fixtures: fixtures/mcp/github

Inline responses support exact full-argument matching (match), JSON Schema matching (match_schema), and per-field regex matching (match_regex). Unknown tools and unmatched calls fail loudly with an MCP tool error that names the missing fixture.

Command Mocks

Use top-level command_mocks when the agent must invoke a CLI such as az, gh, or kubectl. Waza creates temporary executable shims for each task and prepends their directory to the Copilot runtime's existing PATH; the host PATH remains available. This requires schemaVersion: "1.3" or newer and executor: copilot-sdk.

schemaVersion: "1.3"
config:
  executor: copilot-sdk
  model: claude-sonnet-4.6

command_mocks: - name: az expect_calls: 3 responses: - args: ["account", "show", "--output", "json"] stdout: subscriptionId: "00000000-0000-0000-0000-000000000000" name: Test Subscription - args_regex: ["group", "show", "--name", ".+"] fixture: fixtures/commands/az-group-show.json - args: ["deployment", "group", "create"] stderr: "deployment failed" exit_code: 1

Each response matches the complete argument vector, excluding the executable. args_regex applies a full-string regular expression to each argument at the same position. Responses are checked in order; the first match wins. stdout strings are emitted as-is and structured values are JSON-encoded. fixture reads a response file relative to eval.yaml. environment can require exact environment-variable values, and workdir can require an exact workspace-relative directory. expect_calls checks the exact invocation count after the task finishes.

Task-level command_mocks replace the eval-level list; use command_mocks: [] on a task to disable inherited mocks. Mocked invocations (command, sanitized arguments, exit code, and response index) are available to graders and appear under runs[].command_invocations in results.json. Verbose output shows the same sanitized records. An unmatched invocation of a declared executable exits with an actionable missing-

... (README truncated for length)

Chat with me