Tier 1 · Anyone querying results. Assumes you've completed setup, run at least one eval, and imported data with
./build/skillwalker update-analytics-data. This page gets you queryable cross-run metrics from the CLI.
Import test run data into a DuckDB database backed by Parquet files, then query it from the CLI or the web dashboard. This page covers importing data, running queries, and finding your way around the analytics output.
After running tests (and optionally evaluating them with test-eval), import the results:
./build/skillwalker update-analytics-dataThis scans tests/output/ for test run directories and converts their JSONL files into Parquet tables under analytics/data/. The conversion is idempotent — runs already imported are skipped, so you can run this command as often as you like.
The importer reads three files from each run directory:
test-config.jsonl— test configuration recordstest-run.jsonl— Claude stream-json eventstest-results.jsonl— evaluation results (present only iftest-evalhas been run)
It also imports SCIL and ACIL loop output when present (scil-iteration.jsonl, scil-summary.json, acil-iteration.jsonl, acil-summary.json).
Skillwalker provides two built-in analytics queries:
Aggregate pass/fail metrics across all runs, grouped by test name:
./build/skillwalker analytics per-testUse this to spot trends — which tests pass reliably, which ones are flaky, and how pass rates change over time.
Drill into a specific run:
./build/skillwalker analytics test-run-details --run-id <run-id>Use this to inspect a single run's results without opening the web dashboard. Run IDs are timestamps in YYYYMMDDTHHmmss format — you can find them in the tests/output/ directory names or in the web dashboard's Test Run History page.
Both queries support JSON and CSV output:
./build/skillwalker analytics per-test --format json
./build/skillwalker analytics per-test --format csvParquet files are stored at analytics/data/. The key tables are:
test-config.parquet— one row per test case per runtest-run.parquet— one row per Claude event per runtest-results.parquet— one row per expectation or criterion evaluatedscil-iteration.parquet— one row per SCIL iterationscil-summary.parquet— one row per SCIL runacil-iteration.parquet— one row per ACIL iterationacil-summary.parquet— one row per ACIL run
For the complete field reference for each table, see Parquet Schema.
For a visual interface to your analytics data, use the skillwalker-web dashboard. See Viewing Results for a full walkthrough, including the Per-Test Analytics page that surfaces cross-run trends, eval breakdowns, and cost analysis.
- Viewing Results — using the skillwalker-web dashboard
- Parquet Schema — field reference for analytics Parquet files
- CLI — full CLI command reference
- Data Package — shared data layer: types, DuckDB queries, and analytics functions
Next: Parquet Schema — the complete field reference for every analytics table you just imported. Related: Viewing Results — explore the same data visually in the skillwalker-web dashboard.