Tier 1 · Skill authors. Assumes you've completed setup. This page gets you a first skill-call eval run end-to-end, with results you can review.
Create skill-call tests, run them, and review whether Claude routes prompts to your skill correctly.
Skill trigger accuracy measures whether Claude correctly routes user prompts to your custom skill. When it's low, your skill either fires on unrelated prompts (false positives) or fails to fire when it should (false negatives). You improve trigger accuracy by tuning your skill's description so Claude has a sharper sense of when to call it.
- Write your test configuration using the
/write-scil-evalsskill - Run the eval to measure current trigger accuracy
- Import the results into the analytics database
- View the results in the skillwalker-web dashboard
Use the /write-scil-evals skill to generate an eval for your skill:
/write-scil-evals plugin:skill
For example, to test the code-review skill in the r-and-d plugin:
/write-scil-evals r-and-d:code-review
The skill interviews you to collect three categories of prompts:
- Positive triggers — prompts that should trigger your skill (3-5 recommended)
- Negative triggers — prompts that should not trigger your skill (3+ recommended)
- Sibling triggers — prompts that should trigger other skills in the same plugin, not yours (3+ if applicable)
It generates two things in evals/{skill-name}/:
tests.json— the test configuration with one entry per promptprompts/skill-call-*.md— individual prompt files
For details on the skill's full workflow and prompt category conventions, see Writing Skill-Call Evals. For the complete tests.json field reference, see Evals Reference.
Run all tests in your eval:
./build/skillwalker test-run --eval {skill-name}Skillwalker executes each prompt inside the Test Sandbox, records whether your skill was triggered, and prints a pass/fail summary.
Tip: To run a single test in isolation (useful for debugging):
./build/skillwalker test-run --eval {skill-name} --test "Skill Call: some test name"Tip: To see raw Claude output for troubleshooting:
./build/skillwalker test-run --eval {skill-name} --debugFor the full list of CLI flags, see CLI.
Import your test run results into the analytics database:
./build/skillwalker update-analytics-dataThis is idempotent — runs already imported are skipped. For more detail on analytics data and CLI queries, see Analytics.
Launch the skillwalker-web dashboard to inspect your test run:
./build/skillwalker-webOpen http://localhost:3099 in your browser. You'll see your test run in the Test Run History page, and can click through to see per-test pass/fail results. For a full walkthrough of the dashboard, see Viewing Results.
Next: Building SCIL Evals — manual test authoring and iteration strategies for the Skill Call Improvement Loop, which refines your skill's description from test failures. For the loop's mechanics and CLI flags, see Skill Call Improvement Loop. Related: Improving Skill Effectiveness — once triggering is reliable, measure how well the skill does its job.