feat: create a new probe agent to probe code verifications using usage-pattern testing

This commit is contained in:
2026-08-27 13:59:28 -06:00
parent 35e4c82f27
commit 93e90106a8
8 changed files with 690 additions and 7 deletions
+124
View File
@@ -0,0 +1,124 @@
# Probe
A **black-box usage-pattern verifier**. Where every other reviewer reads *text* — the diff
([`code-reviewer`](../code-reviewer/README.md)), the plan ([`adversary`](../adversary/README.md)),
the attack surface ([`security-reviewer`](../security-reviewer/README.md)) — `probe` asks the one
question none of them can answer without running the thing:
> **"Does the changed consumer-facing surface actually behave as the spec promises when used,
> starting from nothing?"**
It boots the system locally from a clean slate, runs any existing usage suites first (regression
check), derives expected behaviors from the **spec** — never the implementation — and authors
tests for the uncovered usage patterns: cold-start/empty-state calls, idempotent re-calls, invalid
input, auth on new routes, partial-update (patch-vs-replace) semantics, serialization edges,
pagination limits, error-shape consistency. These are exactly the defects invisible to static
review.
## Why it's separate from the other reviewers
| | `code-reviewer` | `adversary` | `security-reviewer` | `probe` |
|---|---|---|---|---|
| Question | Is the code good? | Does it match the plan? | Can it be abused? | Does it *work* when used? |
| Method | Reads the diff | Diff vs. criteria | Source→sink tracing | **Runs the system**, black-box |
| Blind spot it covers | slop, bugs, coupling | skipped criteria, drift | injection, authz gaps | behavioral quirks, regressions, contract surprises |
| Output | severity findings | `CONFORMS`/`DIVERGES` | `PASS`/`FAIL` | `PASS`/`FAIL`/`INCONCLUSIVE` |
The independence is behavioral: expectations are written from the spec/contract **before** reading
handler code, so the implementer's misreadings can't become the probe's assertions — the same
principle that makes `adversary` valuable, applied to runtime behavior.
## Verdict (blocking, three-way)
```
USAGE_PROBE: PASS
Surface: <...>. Existing suites: <N run, all green | none found>. New tests: <M authored at <path>, all green>.
```
```
USAGE_PROBE: FAIL
Behavioral findings:
1. <surface + case> — <spec'd behavior> — <observed behavior> — REPRO: <exact request + response> — <test file>
```
```
USAGE_PROBE: INCONCLUSIVE
Could not establish a clean local environment: <verbatim error>. Missing: <the recipe/fixture that would unblock>.
```
- **`FAIL` blocks completion** — the caller resumes the SAME implementer session with the findings
pasted verbatim, then re-runs `probe` once to confirm.
- **`INCONCLUSIVE` is the honest third state**: the environment, not the code, is the blocker. It
routes the fix to the local-run recipe (often a plan gap the `gatekeeper` should have caught) and
is never disguised as `PASS` or `FAIL`.
Every `FAIL` finding carries an exact reproduction (request/command + response received) and the
test file that proves it.
## How it probes
Driven by the [`usage-pattern-testing`](../../skills/usage-pattern-testing/SKILL.md) skill:
1. **Spec first** — expected behaviors written from acceptance criteria + API contract before any
implementation reads.
2. **Regression first** — discover and run existing usage suites; every failure classified as
BUG / EXPECTED-CHANGE / ENV before anything new is authored.
3. **Delta only** — new tests cover only the usage patterns existing suites miss, written in the
repo's suite conventions so they're adoptable as permanent regression coverage.
4. **Clean, local, isolated** — ephemeral state, mocked externals, full teardown; bounded retries
for startup only, never to mask flakiness.
Toolbox by surface — the repo's existing suite format always comes first, and these are examples,
not requirements: [Hurl](https://hurl.dev) or `curl` scripts for HTTP/REST/JSON (Hurl files double
as committed suites), `grpcurl` for pure gRPC, direct invocation for CLIs.
Unlike the read-only reviewers, `probe` **writes test files** (and only test files) — the tests
are a deliverable alongside the verdict. It never modifies implementation code.
## Usage
Spawned by `sisyphus` (post-coder, when the change touches consumer-facing surface) or `architect`
(Phase E, alongside `adversary`). The spawn prompt IS its entire context — include the change, the
spec, and the local-run recipe:
```sh
agent__spawn --agent probe --prompt "
## TASK
Probe the changed API surface for TASK-NNN from the consumer's perspective. Return PASS/FAIL/INCONCLUSIVE.
## CHANGE
Run get_diff --base <ref>, or: <paste the changed-surface summary>
## SPEC — expected behavior to verify against
<paste acceptance criteria + API contract sections (or contract file paths) VERBATIM>
## LOCAL-RUN RECIPE
<how to boot the stack clean: build, deps/stubs, ports, migrations, teardown — or the doc that has it>
## EXISTING SUITES
<paths + run commands, or 'discover them'>
"
```
Direct invocation for ad-hoc use:
```sh
coyote -a probe --agent-variable project_dir /path/to/repo \
"Probe the /widgets endpoints changed in the last commit against this spec: <paste spec>"
```
### Tools
- `get_diff [--base <ref>]` — staged → unstaged → `HEAD~1` fallback (or an explicit base SHA/branch) to locate the changed surface.
- `get_changed_files [--base <ref>]` — quick changed-file map.
- Plus `fs_*`/`ast_grep` for suite discovery and contract reads, `fs_write`/`fs_patch` for authoring test files, and `execute_command` for booting the stack and running suites.
- Probing tools (`curl`, Hurl, grpcurl, the repo's own harness) are invoked via `execute_command`
(no wrapper tool — probing needs their full CLI surface), and none is a hard requirement: the
[`usage-pattern-testing`](../../skills/usage-pattern-testing/SKILL.md) skill has probe reuse the repo's existing suite tooling first and fall back to what's available.
The optional [`sbx-mixin.yaml`](sbx-mixin.yaml) preinstalls Hurl + grpcurl for sandbox runs.
## Related
- [`usage-pattern-testing`](../../skills/usage-pattern-testing/SKILL.md) — the methodology it runs on.
- [`adversary`](../adversary/README.md) — static plan-conformance counterpart (text), where `probe` is dynamic (behavior).
- [`gatekeeper`](../gatekeeper/README.md) — ensures plans ship the local-run recipe `probe` consumes.
+129
View File
@@ -0,0 +1,129 @@
name: probe
description: Black-box usage-pattern verifier - exercises a change's consumer-facing surface (HTTP APIs, RPCs, CLIs) as a real cold-start consumer against a locally running instance with clean, isolated state. Runs existing usage suites first for regressions (whatever format the repo uses - Hurl files, curl scripts, collections), authors spec-first tests for uncovered patterns in the repo's suite conventions (tools like Hurl and grpcurl are examples, not requirements), and returns a blocking USAGE_PROBE PASS/FAIL/INCONCLUSIVE verdict. Complements code-reviewer (quality), adversary (plan conformance), and security-reviewer (abuse). Designed to be delegated to by sisyphus and architect.
version: 1.0.0
auto_continue: true
max_auto_continues: 25
inject_todo_instructions: true
skills_enabled: true
enabled_skills:
- usage-pattern-testing
variables:
- name: project_dir
description: Project directory containing the change under test - where suites are discovered, the stack is booted, and new tests are written
default: '.'
- name: auto_confirm
description: Auto-confirm command execution
default: '1'
global_tools:
- ast_grep.sh
- fs_read.sh
- fs_cat.sh
- fs_grep.sh
- fs_glob.sh
- fs_ls.sh
- fs_write.sh
- fs_patch.sh
- execute_command.sh
instructions: |
You are the usage-pattern probe. You answer ONE question: **does the changed consumer-facing
surface actually behave as the spec promises when used, starting from a clean slate?** Every
other reviewer reads text — the diff, the plan, the code. You are the only gate that BOOTS the
system locally and exercises it the way a consumer will: cold, black-box, spec-first.
You are NOT the code-quality reviewer (`code-reviewer`), NOT the plan-conformance reviewer
(`adversary`), and NOT the security reviewer (`security-reviewer`). You judge observable
behavior. Your value is behavioral independence: expectations derived from the spec BEFORE
reading the implementation, so the implementer's misreadings cannot become your assertions.
## Step 0: Load the skill
Before anything else, `skill__load` `usage-pattern-testing`. It carries your methodology: the
spec-first independence rule, the regression-first protocol (find and run existing suites before
authoring anything), the usage-pattern checklist (cold start, idempotency, invalid input, auth,
partial-update semantics, serialization edges, pagination, error shapes), the clean-environment
discipline, the failure-classification table (BUG / EXPECTED-CHANGE / ENV), the per-surface
toolbox (the repo's existing suite tooling comes first; Hurl/curl for HTTP, grpcurl for gRPC,
and direct invocation for CLIs are examples, not requirements), and the exact verdict format.
The skill body is your source of truth for HOW to probe; these instructions handle
workflow and I/O.
## Input (the spawn prompt IS your entire context)
You are given:
1. **The change** — a diff pasted inline, a summary of the changed surface, or an instruction to
run `git diff`/`get_diff` (optionally against a base ref) in {{project_dir}}.
2. **The spec** — acceptance criteria, plan section, or API contract (or paths to the contract
files: IDL/schema/OpenAPI/proto). This is what you derive expected behaviors FROM.
3. **A local-run recipe** (strongly preferred) — how to boot the system locally from a clean
state: build command, dependencies to start/stub, ports, migration/seed steps, teardown. If
absent, look for one in the repo's contributor docs and dev scripts before inventing your own.
4. **Pointers to existing usage suites** (optional) — where black-box tests already live and how
to run them. If absent, discover them per the skill.
If the spec is missing, STOP and say so: behavior cannot be judged without a promise to judge
against. Do not infer the spec from the implementation.
## Workflow
1. Load `usage-pattern-testing`.
2. Identify the changed consumer-facing surface from the diff/summary. No consumer-facing surface
→ return PASS with a one-line "no probeable surface" note; do not boot anything.
3. **Spec first:** write down expected behaviors as concrete request→response pairs from the
spec/contract, BEFORE reading handler code (implementation reads are for ports/config/startup
wiring only).
4. Discover existing usage suites; bring up the clean local environment per the recipe; run the
existing suites FIRST and classify every failure (regression vs expected contract change vs
environment).
5. Map existing coverage against your expected behaviors; author tests for the uncovered
patterns only, in the repo's suite location and conventions, walking the skill's
usage-pattern checklist.
6. Run the new tests. Classify every failure. Reproduce non-deterministic results twice and read
the server logs before classifying.
7. Tear the environment down. Emit the verdict in the skill's exact format.
## Output — verdict (MANDATORY, exact format)
End with EXACTLY one of the skill's three sentinels so the caller can route on it:
- `USAGE_PROBE: PASS` — existing suites green (or none), new spec-first tests green. List
surface probed, suites run, and tests authored (with paths, so the caller can adopt them).
- `USAGE_PROBE: FAIL` — behavioral findings, each with the spec'd behavior quoted, the observed
behavior, the EXACT reproduction (request/command + response received), and the test file.
- `USAGE_PROBE: INCONCLUSIVE` — a clean local environment could not be established. State what
failed verbatim and EXACTLY what recipe/fixture/mock would unblock. Include any partial
results. INCONCLUSIVE is honest and routes the fix to the environment recipe — NEVER disguise
it as PASS or FAIL.
## Rules
1. **Never modify implementation code.** Your only writes are new/updated TEST files (in the
repo's suite conventions) and throwaway environment scaffolding you tear down. The
implementer owns all fixes.
2. **Spec-first or nothing.** Expectations written from the spec before implementation reads.
If the spec and the contract files disagree, that is a finding — report it, don't pick one
silently.
3. **Regressions before new coverage.** Existing suites run first; a regression is only
acceptable when the spec explicitly changed that contract (then flag the stale test for
update — never delete or silence it).
4. **Clean, local, isolated.** Fresh ephemeral state, mocked externals, no dependence on
pre-existing data or running services, full teardown. Bounded retries for startup only —
never to mask a flaky assertion.
5. **Classify every failure** as BUG / EXPECTED-CHANGE / ENV per the skill table. The verdict
depends on the classification being honest.
6. **Committed tests are the deliverable** alongside the verdict: write them where the repo's
suites live so the caller can adopt them as permanent regression coverage. Report their paths.
7. Be terse and decisive. Three reproducible behavioral findings beat fifteen speculative ones.
If everything works as spec'd, it PASSes — say so.
## Context
- Project: {{project_dir}}
- CWD: {{__cwd__}}
- Shell: {{__shell__}}
## Available Tools
{{__tools__}}
+76
View File
@@ -0,0 +1,76 @@
schemaVersion: '1'
kind: mixin
name: agent-probe
description: >
Optional convenience for the probe agent: preinstalls Hurl (HTTP
usage-pattern tests) and grpcurl (gRPC probing) — the example tools its
skill reaches for — and allows the GitHub release endpoints the fallback
installers download from. Neither tool is required: probe reuses the repo's
existing suite tooling first and falls back to what's available. Hurl
prefers the distro package: the prebuilt GitHub tarball dynamically links
libxml2.so.2, which newer distros no longer ship (e.g. Ubuntu 26.04 moved
to libxml2.so.16). The services under probe run on localhost, which needs
no network allowance. POSIX-only: sbx runs these commands with /bin/sh (dash).
network:
allowedDomains:
# Latest-release lookup + tarball downloads (GitHub redirects release
# assets to *.githubusercontent.com object hosts)
- 'api.github.com:443'
- 'github.com:443'
- 'objects.githubusercontent.com:443'
- 'release-assets.githubusercontent.com:443'
commands:
install:
- command: |
set -eu
if command -v hurl >/dev/null 2>&1; then
hurl --version
exit 0
fi
if command -v apt-get >/dev/null 2>&1; then
sudo apt-get update
if apt-cache policy hurl 2>/dev/null | grep -q 'Candidate: [0-9]'; then
sudo apt-get install -y --no-install-recommends hurl
hurl --version
exit 0
fi
fi
arch="$(uname -m)"
case "$arch" in
aarch64|arm64) arch="aarch64" ;;
*) arch="x86_64" ;;
esac
curl -fsSL https://api.github.com/repos/Orange-OpenSource/hurl/releases/latest -o /tmp/hurl-release.json
ver="$(sed -n 's/.*"tag_name": *"\([^"]*\)".*/\1/p' /tmp/hurl-release.json | head -1)"
curl -fsSL "https://github.com/Orange-OpenSource/hurl/releases/download/${ver}/hurl-${ver}-${arch}-unknown-linux-gnu.tar.gz" -o /tmp/hurl.tgz
mkdir -p /tmp/hurl-extract
tar -xzf /tmp/hurl.tgz -C /tmp/hurl-extract
bin="$(find /tmp/hurl-extract -type f -name hurl | head -1)"
sudo install -m 0755 "$bin" /usr/local/bin/hurl
rm -rf /tmp/hurl.tgz /tmp/hurl-extract /tmp/hurl-release.json
hurl --version
user: '1000'
description: Install Hurl (distro package preferred, GitHub tarball fallback) for the probe agent's HTTP usage-pattern tests
- command: |
set -eu
if command -v grpcurl >/dev/null 2>&1; then
grpcurl -version
exit 0
fi
arch="$(uname -m)"
case "$arch" in
aarch64|arm64) arch="arm64" ;;
*) arch="x86_64" ;;
esac
curl -fsSL https://api.github.com/repos/fullstorydev/grpcurl/releases/latest -o /tmp/grpcurl-release.json
ver="$(sed -n 's/.*"tag_name": *"v\([^"]*\)".*/\1/p' /tmp/grpcurl-release.json | head -1)"
curl -fsSL "https://github.com/fullstorydev/grpcurl/releases/download/v${ver}/grpcurl_${ver}_linux_${arch}.tar.gz" -o /tmp/grpcurl.tgz
mkdir -p /tmp/grpcurl-extract
tar -xzf /tmp/grpcurl.tgz -C /tmp/grpcurl-extract
sudo install -m 0755 /tmp/grpcurl-extract/grpcurl /usr/local/bin/grpcurl
rm -rf /tmp/grpcurl.tgz /tmp/grpcurl-extract /tmp/grpcurl-release.json
grpcurl -version
user: '1000'
description: Install grpcurl (static GitHub release binary) for the probe agent's gRPC probes
+78
View File
@@ -0,0 +1,78 @@
#!/usr/bin/env bash
set -eo pipefail
# @env LLM_OUTPUT=/dev/stdout
# @env LLM_AGENT_VAR_PROJECT_DIR=.
# @describe Usage-pattern probe tools
_project_dir() {
local dir="${LLM_AGENT_VAR_PROJECT_DIR:-.}"
(cd "${dir}" 2>/dev/null && pwd) || echo "${dir}"
}
# @cmd Get the git diff whose consumer-facing surface is under probe. Returns staged changes, or unstaged if nothing is staged, or the HEAD~1 diff if the working tree is clean.
# @option --base Optional base ref to diff against (e.g., "main", "HEAD~3", a commit SHA, or a task's base SHA)
get_diff() {
local project_dir
project_dir=$(_project_dir)
# shellcheck disable=SC2154
local base="${argc_base:-}"
local diff_output=""
if [[ -n "${base}" ]]; then
diff_output=$(cd "${project_dir}" && git diff "${base}" 2>&1) || true
else
diff_output=$(cd "${project_dir}" && git diff --cached 2>&1) || true
if [[ -z "${diff_output}" ]]; then
diff_output=$(cd "${project_dir}" && git diff 2>&1) || true
fi
if [[ -z "${diff_output}" ]]; then
diff_output=$(cd "${project_dir}" && git diff HEAD~1 2>&1) || true
fi
fi
if [[ -z "${diff_output}" ]]; then
echo "No changes found to probe in ${project_dir}." >> "$LLM_OUTPUT"
return 0
fi
local file_count
file_count=$(echo "${diff_output}" | grep -c '^diff --git' || true)
{
echo "Diff contains changes to ${file_count} file(s):"
echo ""
echo "${diff_output}"
} >> "$LLM_OUTPUT"
}
# @cmd Get the list of changed files with stats (a quick map for locating the changed consumer-facing surface).
# @option --base Optional base ref to diff against
get_changed_files() {
local project_dir
project_dir=$(_project_dir)
local base="${argc_base:-}"
local stat_output=""
if [[ -n "${base}" ]]; then
stat_output=$(cd "${project_dir}" && git diff --stat "${base}" 2>&1) || true
else
stat_output=$(cd "${project_dir}" && git diff --cached --stat 2>&1) || true
if [[ -z "${stat_output}" ]]; then
stat_output=$(cd "${project_dir}" && git diff --stat 2>&1) || true
fi
if [[ -z "${stat_output}" ]]; then
stat_output=$(cd "${project_dir}" && git diff --stat HEAD~1 2>&1) || true
fi
fi
if [[ -z "${stat_output}" ]]; then
echo "No changes found in ${project_dir}." >> "$LLM_OUTPUT"
return 0
fi
{
echo "Changed files:"
echo ""
echo "${stat_output}"
} >> "$LLM_OUTPUT"
}