
test-harness
Install vigiles and test a Claude Code harness — hooks, skills, agents, settings, CLAUDE.md — by picking the right tier (unit / deterministic / eval) and writing a test that passes. Use to check that a hook fires or blocks, that a skill triggers, that injected context lands; or to observe a run — which tools it called, whether it stayed inside its declared allowed-tools, what files and side effects it produced, how to intercept a call without executing it.
Related Skills
Install vigiles and test a Claude Code harness — hooks, skills, agents, settings, CLAUDE.md — by picking the right tier (unit / deterministic / eval) and writing a test that passes. Use to check that a hook fires or blocks, that a skill triggers, that injected context lands; or to observe a run — which tools it called, whether it stayed inside its declared allowed-tools, what files and side effects it produced, how to intercept a call without executing it.
Test the Claude Code harness — the hooks, skills, settings, and CLAUDE.md that
steer an agent — as the assembled machine it ships as. vigiles gives three tiers,
cheapest first; this skill picks the right one, writes the test, and runs it.
The guiding rule: start at the cheapest tier that can answer the question, and
climb only when it genuinely can't. Two of the three tiers need no model and no
API key, so they run on every commit for free — reach for the paid real-model
tier only when the question actually requires a real model.
Step 0 — Pick the tier (the judgment call)
Match what you're testing to the cheapest tier that can answer it:
| What you're testing | Tier | Cost | API |
|---|---|---|---|
| "Does this hook block/allow event X?" — pure hook logic, every event type (incl. Edit/Write, PreCompact, SessionEnd, SubagentStop) | Unit | free, milliseconds, no claude |
runHook |
| "Is the hook actually wired into the assembled plugin and does it fire in a real session?" | Deterministic | free, no API key (real claude + scripted mock) |
runHarnessTest + scriptModel |
"Did the injected context (a SessionStart hook, a /command) actually reach the model?" |
Deterministic | free, no API key | runHarnessTest → trace.modelRequests / assertRequestContains |
| "Does this skill's description trigger when it should (recall) and stay quiet when it shouldn't (precision)?" | Eval | paid (real model) | measureTriggerRate (+ irrelevantPrompts) → assertTriggerRate({ min, maxFalsePositive }) |
| "Can I measure triggering on a cheaper model and trust it as a floor?" | Eval | paid (two runs) | compareContainment(weak, strong) → formatContainment |
| "Is this exact skill's output any good?" — absolute quality, no on/off baseline (the default for testing one skill) | Eval | paid (real model) | measure({ checks: [judged(rubric)] }) → assertRates({ min }) |
| "Does this harness change move what the agent does, relative to off?" — A/B lift, regression, signal vs noise | Eval | paid (real model) | runEval (arms) + assertSignificant |
Most harness questions — block/allow, wired-in, context-landed — never need a
model. Only "does the model trigger / behave differently" needs the eval tier.
⚠️ A trigger-rate of 0% on EVERY prompt is a wiring bug until proven otherwise.
It reads like a verdict on the description, and three separate setup mistakes
produce it: a bare id in fired where the namespaced <plugin>:<skill> is
required; pluginDir where a loose .claude/skills needs skillsDir; and a
missing fixture, since a run starts in an empty directory and a prompt about
a file that isn't there is one the model is right to decline. Rule all three out
before reporting it. (A partial rate is a real number — don't second-guess it.)
Don't tune against a cheaper model until you've checked it's actually a floor.
compareContainment(weak, strong) answers that: it reports prompts that fired on
the weak model but NOT the strong one, and each one means the weak model is not a
lower bound but a different router. Prompts that fired only on the strong model
are expected and are not a failure. Measured once (21 skills, 84 prompts, haiku
vs sonnet): 3 weak-only, and one skill higher on haiku — so containment is
not established, which is why the floor stays.
If the unit and deterministic tiers can both answer it, prefer unit: it's
faster and reaches events the deterministic mock can't drive.
Step 0.4 — Observing a run (what it CALLED, WROTE, and TOUCHED)
The table above is keyed on the harness surface under test. Half the real
questions are keyed on the observation instead — "what did this skill
actually do?" — and they have answers already. Reach for these before building
anything; every one of them ships today.
| The question you're actually asking | Use |
|---|---|
| Which tools did it call, and with what arguments? | trace.toolCalls · tool / toolWith checks · parseToolCalls (vigiles) |
| Did it call a tool it must not? | notTool(name) |
| Did it call only tools from a known set? | onlyTools([...]) — the white-list, symmetric to assertWroteOnly |
Did it stay inside the allowed-tools its own frontmatter declares? |
skillContract(dir).surface — builds that check FROM the declaration |
| What files did the run write? | filesWritten · wrote(path) / didNotWrite(path) · r.file(path) |
| Did it write only where it was supposed to? | assertWroteOnly([...]) / assertNoWrite() — needs { sandbox: "auto" } |
| Run a tool call but don't let it execute — capture the args instead | the interceptTools option on measure / runEval (a ToolIntercept[]) |
| Did a subagent do it, and which one? | subagent(name, [...]) · SubagentTrace |
| Was it an MCP tool? | mcp(server, toolName) |
| Assert the whole effect boundary deterministically | assertChecks + the checks above (see examples/harness/effect-boundary.harness.mjs) |
interceptTools is the one worth knowing about, because it is not obvious it
exists: it denies a tool its real execution via an auto-wired PreToolUse
hook while still recording the call and its arguments into the trace. That is
how you test a skill that would otherwise mutate a real external service — a
calendar, an upload — without mocking anything yourself.
Verify a skill against its own declaration with skillContract — it reads
the allowed-tools: the skill already claims and hands back ready checks, so
the claim is verified instead of restated:
import { skillContract, assertChecks } from "vigiles";
const c = skillContract(".claude/skills/my-skill");
assertChecks(trace, [c.activation, ...c.surface]);
Two of its states are findings, not clean bills, and their surface check
fails rather than passing on nothing: undeclared (no allowed-tools: line, so
the skill inherits every tool) and malformed (frontmatter that isn't valid
YAML, so a strict loader reads no contract at all — one unquoted : does it).
⚠️ What is still NOT checked. onlyTools compares tool names, so a narrow
allowlist entry like Bash(node scripts/x.mjs:*) is satisfied by any Bash call
at all. Scope inside a tool is unverified — say so rather than implying the
assertion is total.
Step 0.5 — Set honest expectations (what's testable, and at what cost)
Be explicit with the user about which bucket each surface falls into — never let
"we'll test it" hide whether that's free, sub-priced, or needs a container. Every
surface sorts into one of three buckets:
- A — Free & deterministic (no model, runs in CI on every commit): a hook's
block/allow decision (runHook), a tool-contract / "did NOT call the forbidden
tool" check, structural facts (vigiles audit), and record-replay of any tool
a skill shells out to (record the real result once, replay it via a PATH stub). - B — Model-gated, on your subscription (real model, no metered API): does a
skill's description fire (measureTriggerRate, recall + precision) and
does its guidance actually produce good output (score it directly:
measure({ checks: [judged(rubric)] })+assertRates— the absolute oracle;
use arunEvalA/B on-vs-off only when you need the relative lift). This is
the half a prose / guidance skill lives in —
its worth is behavioral, so only a model can judge it. That is not "uncovered"
and not free: it's fully testable on the sub. State it that way. - C — Needs a real service (a real browser / DB / redis / a11y runtime): vigiles
composes with a container here; it does not fake real semantics. Name the
service and hand off — don't pretend a cheap tier substitutes for it.
So a prose-skill library is roughly ~100% testable (some free, most on your sub),
~0% needs-a-container — not "poorly covered." An accessibility/browser plugin is
the worst case, with a large bucket C. When you report coverage, give two
numbers: "% testable at all (free + sub)" vs "% that needs a container", and say
which surfaces are free vs sub-priced. The model-gated half is the point of the
eval pillar (affordable on the sub), not a gap — and testing a prose skill's
behavior requires a real model for everyone (promptfoo, the SDKs, all of it);
vigiles just does it on your subscription instead of metered API.
Step 1 — Ensure vigiles is installed
Check whether vigiles is a dependency (package.json), and install it as a
dev dependency if not:
npm i -D vigiles # or: pnpm add -D vigiles / yarn add -D vigiles
The deterministic tier additionally needs the claude CLI on PATH (no API key):
npm i -g @anthropic-ai/claude-code. The eval tier needs model auth. If the
claude CLI is missing, you can still write and run unit-tier tests.
Step 2 — Locate the harness surface to test
Find what the project actually ships, in this order:
.claude/settings.json/.claude/settings.local.json— inlinehooks..claude-plugin/plugin.json— a plugin manifest (hooks,skills,agents,mcpServers).hooks/hooks.json— the plugin hooks convention (e.g. obra/superpowers).skills/<name>/SKILL.md,agents/<name>.md,commands/<name>.md.
Pick one concrete thing to pin down — a specific PreToolUse hook, a specific
SessionStart injection, a specific skill.
Step 3 — Write the test for the chosen tier
Unit (runHook) — hand a hook a synthesized event, assert the decision:
import { runHook, assertHookBlocked } from "vigiles";
const r = runHook(hookCommand, {
hook_event_name: "PreToolUse",
tool_name: "Bash",
tool_input: { command: "git commit --no-verify" },
});
assertHookBlocked(r); // exit 2 / decision:"block" / permissionDecision:"deny"
Testing a hook you didn't write (a vendored third-party script)? Mark it
{ trusted: false } and it runs confined under bubblewrap by default (read-only
host, cleared env, no network egress). Add { recordEgress: true } to also
record what it tries to reach — r.egress plus assertNoEgress(r) /
assertEgressOnly(r, [...]) — the supply-chain check for "what does this skill
phone home to / install from?". When the hook's setup needs a real install,
{ egress: { allow: ["registry.npmjs.org"] } } lets it reach only that
allowlist (a packet-layer nft wall, so a raw socket off-list is dropped too) →
r.egress (allowed hosts) + r.egressDropped. Be precise about the boundaries:
see
docs/sandboxing.md (it blocks destruction and
egress, but does NOT isolate reads of host files, and only under bwrap).
Deterministic (runHarnessTest) — load the real plugin, drive a scripted
mock model, assert the hook fired (or the context landed):
import {
runHarnessTest,
assertHookFired,
assertRequestContains,
} from "vigiles";
// `scriptModel` is the Claude-Code TRANSPORT, deliberately not re-exported from
// the harness-agnostic root surface — import it from the harness package:
import { scriptModel } from "vigiles/claude-code";
const r = await runHarnessTest({
pluginDir: "./", // or { settings: { hooks: {...} } }
transcript: true,
model: scriptModel([{ text: "ok" }]),
});
assertHookFired(r, "SessionStart");
assertRequestContains(r, "expected injected text"); // did it actually land?
Eval — absolute (paid_measure + paid_judged) — testing one skill, the usual case:
score its output directly against a rubric. No on/off baseline — this is the
"is it any good?" oracle (what promptfoo/DeepEval lead with), and the right
default when there's nothing to compare against:
An eval file describes its eval — it must never run one at the top level,
because importing such a file spends real money. Write <name>.eval.mjs:
import { defineEval, skill, assertRates } from "vigiles";
import { paid_judged } from "vigiles/eval"; // a Check whose default judge bills
export default defineEval({
measure: {
pluginDir: "./",
task: "…a task the skill should handle…",
checks: [
skill("my-plugin:my-skill"), // it fired
paid_judged("the answer correctly does X and avoids Y"), // …and the output is good
],
trials: 6,
},
assert: (report) => assertRates(report, { min: 0.8 }), // each check ≥ 80% of trials
});
Run it with npx vigiles eval <file> — never node <file>, which refuses.
Eval — relative (paid_runEval + assertSignificant) — when the question is
lift over no-skill (regression, or proving a change isn't noise): A/B the
change on vs off and gate on significance, not eyeballing:
import { defineEval, assertSignificant } from "vigiles";
export default defineEval({
runEval: {
arms: { off: {}, on: { pluginDir: "./" } },
task: "…a task the harness change should affect…",
measure: (ctx) => ({ ok: /* a bare predicate over the trace */ true }),
trials: 6,
cache: "readwrite",
},
assert: (report) =>
assertSignificant(report, { baseline: "off", arm: "on", metric: "ok" }),
});
Never hand-roll the runner — it silently eats stderr
Do not reach for execFileSync / spawnSync to drive the thing under test.
The failure is quiet and repeats: execFileSync returns stdout only on
success, while advisory output — including vigiles's own compiled-hook
notice() — is written to stderr. A hand-rolled runner therefore reports a
perfectly healthy react hook as dead, and an assertion about a warning can
never pass. (Observed three times in one repo, twice after the first fix.)
Every vigiles result already carries both streams, so the bug is
unrepresentable:
| Runner | Result | Carries |
|---|---|---|
runScript |
ScriptRunResult |
exitCode, stdout, stderr, filesWritten? |
runHook |
HookRunResult |
all of the above, plus blocked / decision |
runHarnessTest |
HarnessTestResult |
exitCode, stdout, stderr, cwd + the Trace |
Testing a plain helper script (a bash/node/python program that isn't a hook)?
Use runScript — it runs any command and reports what it did:
import { runScript } from "vigiles";
const r = runScript("bash scripts/check-links.sh", { cwd: repoDir });
assert.equal(r.exitCode, 0);
assert.match(r.stderr, /0 broken links/); // advisory output lives HERE
runHook is exactly runScript plus the hook protocol (event → stdin, exit code
→ allow/deny). Pick by the question you're asking: a hook has a decision, a
script has effects. That's why ScriptRunResult has no decision field —
a field that is always meaningless is worse than no field.
⚠️ Asserting what a script wrote requires confinement. filesWritten is
recorded by diffing the work dir, which only a confined run does — so it is
undefined after a plain run. That is deliberately not the same as []
("recorded, wrote nothing"): assertNoWrite / assertWroteOnly throw on an
unrecorded result rather than pass having inspected nothing. Pass
{ sandbox: "auto" } (Linux + bubblewrap) to actually record writes.
Step 4 — Run it
In a runner (node:test / vitest / jest) the tests are plain async functions. Or
use the zero-setup CLI, which discovers and runs the files:
npx vigiles test # *.harness.{mjs,ts} — unit + deterministic, no API key
npx vigiles eval --trials=6 # *.eval.{mjs,ts} — real model (local / nightly, not CI)
Unit-tier runHook tests need no claude and always run — write and run them
even with no claude installed. A tier that genuinely can't run reports a loud
⊘ SKIPPED (tallied separately, never a fake ✓); a standalone script emits one
via skip(reason) from vigiles. A skip passes by default, but in a CI
job that asserts the capability is present, run vigiles test --no-skip so a
skipped tier fails — a green-with-skips is untested surface. Keep unit +
deterministic tests in CI (free); run evals locally or on a schedule with auth.
After a real-model run: TELL THE USER WHAT IT SPENT
Whenever you run a real-model eval (runEval / measureArms / measureTriggerRate
/ measure), surface the spend to the user in your reply — don't let a paid run
be silent. runEval prints a cost block to stderr and every report carries usage
(report.arms[*].usage: totalCostUsd + token counts). Relay, in plain words:
- tokens spent and the API-equivalent
$(total_cost_usd— what it would
cost at metered API rates); - how it was billed — "on your Claude subscription ($0 metered)" if you're
logged in, or a ⚠ warning ifANTHROPIC_API_KEYis set (that run was billed
per token — tell them to unset it andclaude loginto run free).
We do not show "% of your subscription" — Anthropic doesn't expose a plan's
quota, so any percentage would be invented. Tokens + API-equivalent $ + the
billed-to line is the honest, complete picture. Keep the user's cost visible, always.
CI — don't hand-write the steps
These tiers belong in CI, and there is a published Action for it. Run vigiles init: it
writes .github/workflows/vigiles.yml, wiring the Action (zernie/vigiles@v1) for the jobs
that can use it plus a plain npx vigiles test job for this tier — that one needs
repo-local node_modules, which the Action does not install, so it stays hand-rolled on
purpose.
If the repo already has a workflow, the Action's inputs are documented in
docs/github-action.md. Read them there rather than guessing:
the input list is defined in action.yml, and a copy of it here would be a second source of
truth that goes stale without anything noticing — which is exactly what happened to this
file's own sibling docs and to a consumer's CI comment, both measured on 2026-08-18.
Step 5 — Lock the eval so CI stays honest (you do this automatically)
Real-model evals run on the user's subscription — locally, never in CI. So as
part of writing an eval, you keep its result fresh for them. Do these two things
without being asked:
-
Give every eval a
name. That's what the lock keys on. -
After running it, commit the lock:
vigiles eval --update # records the result → .vigiles/eval-locks/<name>.lock.jsonThen commit that file. CI runs
vigiles eval --check(no model) to verify it
still matches the inputs — so a later edit that forgets to re-eval fails loud
instead of shipping stale numbers.
When you later change a skill's description or prompts, re-run vigiles eval --update and commit the updated lock — the change altered what the eval
measures. (vigiles also nudges you: when a lock exists, a SKILL.md edit triggers
a non-blocking reminder.)
Why it's cheap: --check only hashes inputs (skill text, prompts, model). A
threshold change in the test re-uses the saved numbers (no model); only an
input change needs a fresh --update. Full mechanics:
docs/harness-testing.md.
When the user didn't say what to test
Don't ask them to specify — pick something real and demonstrate. Scan the
harness surface (Step 2), choose the cheapest meaningful test, write it, run it,
and show the result. Good default picks, in order:
- A
PreToolUsehook → unit-test that it blocks the thing it's meant to block (and allows a safe sibling). - A
SessionStarthook that injects context → deterministic test that the text actually reaches the model (assertRequestContains). - A skill → deterministic test that it resolves via
pluginDir, then offer the paidmeasureTriggerRateeval as a follow-up.
Then say which tier you used and why, and offer to climb a tier if the cheaper
test can't fully answer their question.
Reference
The full guide — every tier, testing skills for real, "fired ≠ landed", the
safe-by-default sandbox, the coverage matrix, and how it compares to promptfoo —
is in docs/harness-testing.md.



