Skip to content

Latest commit

 

History

History
86 lines (54 loc) · 3.59 KB

File metadata and controls

86 lines (54 loc) · 3.59 KB

Analytics

Tier 1 · Anyone querying results. Assumes you've completed setup, run at least one eval, and imported data with ./build/skillwalker update-analytics-data. This page gets you queryable cross-run metrics from the CLI.

Import test run data into a DuckDB database backed by Parquet files, then query it from the CLI or the web dashboard. This page covers importing data, running queries, and finding your way around the analytics output.

Importing data

After running tests (and optionally evaluating them with test-eval), import the results:

./build/skillwalker update-analytics-data

This scans tests/output/ for test run directories and converts their JSONL files into Parquet tables under analytics/data/. The conversion is idempotent — runs already imported are skipped, so you can run this command as often as you like.

The importer reads three files from each run directory:

  • test-config.jsonl — test configuration records
  • test-run.jsonl — Claude stream-json events
  • test-results.jsonl — evaluation results (present only if test-eval has been run)

It also imports SCIL and ACIL loop output when present (scil-iteration.jsonl, scil-summary.json, acil-iteration.jsonl, acil-summary.json).

CLI queries

Skillwalker provides two built-in analytics queries:

Per-test metrics

Aggregate pass/fail metrics across all runs, grouped by test name:

./build/skillwalker analytics per-test

Use this to spot trends — which tests pass reliably, which ones are flaky, and how pass rates change over time.

Test run details

Drill into a specific run:

./build/skillwalker analytics test-run-details --run-id <run-id>

Use this to inspect a single run's results without opening the web dashboard. Run IDs are timestamps in YYYYMMDDTHHmmss format — you can find them in the tests/output/ directory names or in the web dashboard's Test Run History page.

Output formats

Both queries support JSON and CSV output:

./build/skillwalker analytics per-test --format json
./build/skillwalker analytics per-test --format csv

Data location and schema

Parquet files are stored at analytics/data/. The key tables are:

  • test-config.parquet — one row per test case per run
  • test-run.parquet — one row per Claude event per run
  • test-results.parquet — one row per expectation or criterion evaluated
  • scil-iteration.parquet — one row per SCIL iteration
  • scil-summary.parquet — one row per SCIL run
  • acil-iteration.parquet — one row per ACIL iteration
  • acil-summary.parquet — one row per ACIL run

For the complete field reference for each table, see Parquet Schema.

Viewing in the dashboard

For a visual interface to your analytics data, use the skillwalker-web dashboard. See Viewing Results for a full walkthrough, including the Per-Test Analytics page that surfaces cross-run trends, eval breakdowns, and cost analysis.

Related documentation

  • Viewing Results — using the skillwalker-web dashboard
  • Parquet Schema — field reference for analytics Parquet files
  • CLI — full CLI command reference
  • Data Package — shared data layer: types, DuckDB queries, and analytics functions

Next: Parquet Schema — the complete field reference for every analytics table you just imported. Related: Viewing Results — explore the same data visually in the skillwalker-web dashboard.