Tier 1 · Agent authors. Assumes you've completed setup. This page gets you a first rubric-scored agent run end-to-end, with LLM-judge results you can review.
Build a scaffold, write rubric evals, run them, and have an LLM judge score your agent's output against your quality criteria.
Agent effectiveness measures how well your agent performs its job — not whether it triggers, but whether its output meets quality criteria. You define those criteria in a rubric, and an LLM judge scores the agent's output against them. Effectiveness is independent of trigger accuracy: an agent can be delegated to reliably and still produce weak output.
- Create a project scaffold that gives your agent realistic context to work with
- Write your test configuration and rubric using the
/write-agent-eval-rubricskill - Run the eval to produce agent output
- Evaluate the results with the LLM judge
- Import the results into the analytics database
- View the results in the skillwalker-web dashboard
Scaffolds are realistic project directories that your agent analyzes inside the Test Sandbox. They contain source files, configs, and intentionally-planted signals (bugs, missing docs, architectural issues, implementation gaps) that your agent should detect.
Use the /build-agent-eval-scaffold skill to generate one:
/build-agent-eval-scaffold plugin:agent
For example:
/build-agent-eval-scaffold r-and-d:gap-analyzer
The skill interviews you about the technology stack, project shape, and specific signals to plant, then generates a scaffold directory at evals/{agent-name}/scaffolds/{scaffold-name}/.
For details on the scaffold creation workflow, see Building Agent Eval Scaffolds. For how scaffolds work inside the Test Sandbox, see Test Scaffolding.
Use the /write-agent-eval-rubric skill to create a quality rubric and configure test entries:
/write-agent-eval-rubric plugin:agent
For example:
/write-agent-eval-rubric r-and-d:gap-analyzer
The skill reads your scaffold files and interviews you to collect criteria in four categories:
- Presence — things the agent's output must identify (e.g., "The analysis identifies the missing fetch/proxy endpoint")
- Specificity — output must reference concrete details (file names, line numbers, comparison areas)
- Depth — output must be actionable (gap classifications, evidence pairs, remediation suggestions)
- Absence — things the output must not do (hallucinations, incorrect classifications, false gaps)
It generates two things:
evals/{agent-name}/rubrics/{agent-name}-quality.md— the rubric file with categorized criteria- Updates to
evals/{agent-name}/tests.json— agent-prompt test entries withllm-judgeexpectations
Note: The skill can create agent-prompt tests from scratch if none exist yet, or add rubric expectations to existing tests.
For details on the skill's full workflow and criteria categories, see Writing Agent Eval Rubrics. For the complete tests.json field reference, see Evals Reference.
Run all tests in your eval:
./build/skillwalker test-run --eval {agent-name}This runs your agent against the prompt and scaffold inside the Test Sandbox. The LLM judge does not run yet — it evaluates stored output in the next step.
Tip: To run a single test in isolation (useful for debugging):
./build/skillwalker test-run --eval {agent-name} --test "Prompt: some test name"Tip: To see raw Claude output for troubleshooting:
./build/skillwalker test-run --eval {agent-name} --debugFor the full list of CLI flags, see CLI.
Run the evaluation pipeline to have the LLM judge score your agent's output against the rubric:
./build/skillwalker test-evalThe test-run step captures Claude's output; test-eval scores it against your rubric. These are separate commands because the LLM judge consumes tokens — running a second Claude invocation to evaluate each test. Keeping them separate lets you run tests now and evaluate later when you have tokens to spare, or batch multiple test runs before evaluating them all at once.
The judge prints progress per test, showing each criterion as pass or fail with reasoning. Results are written to tests/output/{run-id}/test-results.jsonl.
For details on how the judge constructs its prompt, scores criteria, and handles errors, see LLM Judge Evaluation.
Import your test run and evaluation results into the analytics database:
./build/skillwalker update-analytics-dataThis is idempotent — runs already imported are skipped. For more detail on analytics data and CLI queries, see Analytics.
Launch the skillwalker-web dashboard to inspect your test run and judge results:
./build/skillwalker-webOpen http://localhost:3099 in your browser. Navigate to your test run to see per-criterion pass/fail results from the LLM judge, including the reasoning behind each score. For a full walkthrough of the dashboard, see Viewing Results.
Next: Building Rubric Evals — iterate on rubric criteria and re-score stored output without re-running the agent. Related: Improving Agent Trigger Accuracy — if the agent isn't being delegated to reliably, measure and improve when Claude calls it.