Tier 5 · Contributor reference. Internal documentation for the
packages/evalspackage. If you're a user looking to write and run quality evals, see Building Rubric Evals.
The evaluation engine. Change this package when you need to touch how test run results are scored against expectations — the deterministic boolean checks, the semantic LLM judge, rubric parsing, the judge prompt builder, or the top-level evaluateTestRun() orchestrator.
After the test-run pipeline executes prompts inside Test Sandboxes and writes JSONL event streams, this package reads those streams and scores each test case's expectations, producing structured pass/fail records for both deterministic boolean checks and semantic LLM judge assessments.
- Last Updated: 2026-05-15
- Package:
@testdouble/skillwalker-evals(packages/evals, workspace package, not published) - Runtime: TypeScript on Bun, ESNext target, strict mode
- Tests: Vitest (co-located
.test.tsfiles)
It serves three consumers:
test-evalcommand -- Evaluates all expectations (boolean + LLM judge) for a completed test run and converts results intoTestResultRecordentries written totest-results.jsonl.- SCIL improvement loop (
step-5-run-eval) -- UsesevaluateSkillCalldirectly to check whether a specific skill was invoked during a prompt execution. - ACIL improvement loop (
step-5-run-eval) -- UsesevaluateAgentCalldirectly to check whether a specific agent was invoked during a prompt execution.
index.ts -- Public barrel export
src/
types.ts -- EvalResult, BooleanEvalResult, LlmJudgeEvalResult, progress events
boolean-evals.ts -- Deterministic expectation evaluators
rubric-parser.ts -- Parses rubric markdown into RubricSection objects (transcript + file sections)
llm-judge-prompt.ts -- Builds the prompt sent to the judge model (including output file content)
llm-judge-eval.ts -- Orchestrates LLM judge evaluation (with auto-fail for missing files)
evaluate.ts -- Top-level evaluateTestRun() orchestrator
boolean-evals.test.ts -- Unit tests for boolean evaluators
rubric-parser.test.ts -- Unit tests for rubric section parsing
llm-judge-eval.test.ts -- Unit tests for LLM judge evaluation
| Dependency | Usage |
|---|---|
@testdouble/skillwalker-data |
JSONL I/O, stream event types, config records, getResultText, getSkillInvocations, getAgentInvocations, parseStreamJsonLines, readJsonlFile, buildTestCaseId |
@testdouble/claude-integration |
runClaude() for invoking the judge model |
Four deterministic evaluators that check test run event streams without any LLM calls:
| Expectation Type | Function | Behavior |
|---|---|---|
result-contains |
evaluateResultContains |
Passes when the final result text includes the expected string (case-sensitive) |
result-does-not-contain |
evaluateResultDoesNotContain |
Passes when the final result text does not include the expected string |
skill-call |
evaluateSkillCall |
Passes when the skill invocation list matches the expected presence/absence of a skill file |
agent-call |
evaluateAgentCall |
Passes when the agent invocation list matches the expected presence/absence of an agent type |
All four return false when no result event exists in the stream (except evaluateSkillCall/evaluateAgentCall with shouldBeCalled: false, which return true for empty events).
evaluateAllExpectations filters out llm-judge expectations and maps the remainder through evaluateExpectation, returning an array of ExpectationResult records.
Semantic evaluation that sends skill output to a second Claude model for rubric-based scoring:
- Reads the rubric markdown file from the eval's
rubrics/directory. - Parses rubric sections using
parseRubricSections— producesRubricSectionobjects withtype: 'transcript'for standard criteria andtype: 'file'with afilePathfor file-scoped criteria. - Loads output files from
output-files.jsonlin the run directory, matching records bybuildTestCaseId(eval, test.name)(only when the rubric contains file sections). - Builds a judge prompt containing scaffold files (if any), a conversation transcript, the final skill output, output file content (for file-scoped sections), and the numbered rubric criteria. File-scoped criteria for missing files are separated as auto-fails.
- Invokes
runClaude()with the specified model (default:opus) and parses the JSON response. If all criteria are auto-fails, the judge is not invoked. - Merges judge results with auto-fail results. Scores each criterion as passed (1.0), partial (0.5), or failed (0.0). Computes an aggregate score as
passedCount / totalCriteria. - Compares the aggregate score against the threshold (default:
1.0) to determine overall pass/fail.
Error handling wraps each judge evaluation in a try/catch. Failures (rubric not found, sandbox timeout, malformed judge response) produce a result with status: 'infrastructure-error' rather than crashing the pipeline.
Parses rubric markdown into structured RubricSection objects. Each section has a type ('transcript' or 'file') and a criteria array of bullet-point strings. File sections also carry a filePath string.
The parser splits on ## File: {path} headers — everything before the first file header is a transcript section, and each file header starts a new file section. Within each section, lines starting with - are extracted as criteria. Headings other than ## File: are ignored.
A backward-compatible parseRubricCriteria function flattens all sections into a single criteria list. Returns an empty array for rubrics with no bullet lines, which triggers an infrastructure-error in the judge evaluator.
Constructs a multi-section prompt for the judge model:
- System framing -- "You are evaluating the output of a Claude Code skill run." (or "agent run" for agent-prompt tests).
- Scaffold files -- If the test uses a scaffold directory, recursively reads all files (up to 5 KB each) and includes them as named sections.
- Transcript -- Formats tool-use events from the stream as
[Tool: name] argswith truncated results (500 chars max), giving the judge visibility into intermediate steps. - Final skill output -- The complete result text.
- Output file content -- For each
## File:section in the rubric whose file exists inoutput-files.jsonl, the file content is injected as a named section (# Output File: {path}). Missing files are excluded and their criteria are auto-failed. - Rubric criteria -- Numbered criteria list with JSON response format instructions. File-scoped criteria are prefixed with
[File: path]. Auto-failed criteria (missing files) are excluded from the judge prompt.
The function returns both the prompt string and an autoFailCriteria array. Auto-failed criteria bypass the judge entirely — they are merged into the final results with passed: false and reasoning "Output file was not produced by the agent."
evaluateTestRun is the main entry point consumed by the test-eval command:
- Reads
test-config.jsonlandtest-run.jsonlfrom the run directory. - Validates compatibility (events must have
test_casefields; older runs without this field are rejected with a descriptive error). - Groups events by test case ID.
- For each test case, evaluates boolean expectations first, then LLM judge expectations.
- Emits progress events via the optional
onProgresscallback for real-time CLI output. - Returns all
EvalResultrecords (bothBooleanEvalResultandLlmJudgeEvalResult).
EvalResult-- Union ofBooleanEvalResult | LlmJudgeEvalResult.BooleanEvalResult-- Containskind: 'boolean', test identifiers, expectation type/value, pass/fail, and status.LlmJudgeEvalResult-- Containskind: 'llm-judge', test identifiers, judge model, aggregate score and threshold, rubric file path, and an array ofLlmJudgeCriterionResultentries.LlmJudgeCriterionResult-- Per-criterion result withpassed, optionalconfidence: 'partial', andreasoning.
EvalProgressEvent-- Union type foreval-start,eval-complete, andeval-errorevents, consumed by CLI progress logging.OnProgress-- Callback type(event: EvalProgressEvent) => void.
All results carry a status field:
'evaluated'-- Normal evaluation completed.'infrastructure-error'-- Evaluation failed due to environment issues (missing rubric, sandbox timeout, parse failure). The result includeserror_messageandpassed: false.
- Skillwalker Architecture — System architecture, package boundaries, and dependency graph
- Execution Package — The
runTestEval()orchestrator that consumesevaluateTestRun(), the SCIL loop that usesevaluateSkillCalldirectly, and the ACIL loop that usesevaluateAgentCalldirectly - CLI Package — Thin Yargs wrapper that delegates to the execution package
- Data Package — Shared data layer providing JSONL I/O, stream event types, config records, and
getResultText/getSkillInvocations - Claude Integration —
runClaude()used to invoke the judge model inside the Test Sandbox - LLM Judge Evaluation — Detailed judge mechanics: prompt construction, scoring, output format, and error handling
- Evals Reference —
tests.jsonfield reference includingllm-judgeandskill-callexpectation formats - Parquet Schema — Analytics schema for evaluation results stored as Parquet
- Building Rubric Evals — Step-by-step guide to writing and running LLM-judge quality evals
- Building SCIL Evals — Step-by-step guide to writing and running trigger accuracy evals
- Agent Call Improvement Loop — ACIL mechanics: agent detection, temp plugin isolation, holdout splits, scoring
- Writing Agent-Call Evals — Skill for generating agent-call evals
Next: LLM Judge Evaluation — detailed judge mechanics: prompt construction, scoring, output format, and error handling.
Related: Execution Package — the runTestEval(), SCIL, and ACIL orchestrators that consume this package's evaluators.