Skip to content

Latest commit

 

History

History
248 lines (193 loc) · 11.6 KB

File metadata and controls

248 lines (193 loc) · 11.6 KB

Method

How the numbers in RESULTS.md were produced, and what they do and do not license you to conclude. The rules are fixed in advance: frozen metrics, stated denominators, controls reported next to results, and no pooling of unlike tasks into one number.


Systems under test

Everything measured here is the C runtime, reached through the ctypes binding in python/blink_train/runtime.py. The PyTorch model is a training artefact; it never produces a reported number. This matters because the thing that ships is the C library and the int8 container, not the fp32 graph.

Two presets:

blink-tiny blink-small
width / blocks / FFN 64 / 2 / 128 192 / 4 / 384
attention rank / heads 64 / 2 192 / 4
pooling stride 4 8
blocks with a position mixer 1 2
hashed bigram buckets 4,096 32,768
state / question / option byte limits 256 / 128 / 48 512 / 192 / 64

Both are the same architecture at two scales and the same .blink container read by the same blink_score.


Corpora

Synthetic (primary; no download, fully deterministic)

Four task families, generated by blink_train.data.synthetic_rows from one seed:

family what it asks what it requires
routing which team handles this ticket associating a cue word with an option description
judgment does the evidence support, contradict, or fail to settle a claim comparing the claim against the state
selection which city does a named person live in binding one entity in the question to one fact in the state
comparison which product scored highest, or lowest reading several values and ranking them, with the direction set by the question

The four were chosen to separate capabilities, not to flatter the model. routing needs only a lexical association; selection and comparison need a question-by-option interaction and are reported separately for that reason.

Two properties matter for the split:

  • Groups, not rows. One scenario produces several questions about one state. Splits are made on the scenario, so two questions about the same state can never straddle a split.
  • Held-out entities. The names, cities and products in the test split are drawn from a pool disjoint from the one used in training. A model cannot score by memorising that Noor goes with a particular answer, because Noor never appears in the test split. The queue labels are shared on purpose: they are the options, and reading options is the task.

One consequence of the entity split is worth stating: the held-out pool is smaller than the training pool, so the families whose option list is built from entities (selection, comparison) offer at most five options in the test split against six in training. Chance therefore differs between the splits, which is why the uniform baseline is computed per split and printed next to every accuracy rather than assumed. Test rows average slightly fewer options, so the test uniform baseline is slightly higher than the training one.

What this corpus does not establish: it has a small closed vocabulary and templated sentence frames. It measures whether the mechanism works. It does not measure open-domain language understanding, and no claim here should be read as measuring that.

WANLI (external)

WANLI (Liu et al., 2022, pinned to revision 61c95318fd71c55b6ba355d76253254615f387ec) supplies the external check. Premise becomes the state, hypothesis becomes the question, and the three NLI relations become three described options whose display order is shuffled per row so the label index carries no signal. Rows are grouped by premise and groups never straddle a split.

WANLI is large enough to train on, which is the point: asking a 390k-parameter byte model to transfer zero-shot onto a task it has never seen is not an interesting experiment, and reporting such a number as if it were an accuracy comparison would be misleading.

External fixtures (transfer, reported as such)

authored144 and perturbations108, two fixtures published with SemIf (MIT), convert one to one. They are reported only as a transfer result for a model trained on another corpus, with the expectation stated in advance that a model this size trained on synthetic routing tasks will be at or near chance on them. They are included because a negative result on a real external fixture is more informative about this model class than another synthetic number.

Note also that some of these rows exceed the byte limits of both presets (up to 432 bytes of state, 223 of question). The harness clips and reports clipped_rows; the runtime itself refuses over-long input rather than truncating silently.


Metrics

Fixed before any model was trained, and computed by eval/evaluate.py.

metric why it is here
top-1, top-3 the headline, full denominator
balanced accuracy immune to a skewed label prior
macro F1 per-class behaviour, not just the majority class
NLL, Brier proper scoring rules: they punish confident mistakes
ECE (15 equal-width bins), MCE calibration, with the reliability table kept in the JSON
risk–coverage accuracy when the model may abstain below a confidence threshold
per-family strata unlike tasks are never pooled into one accuracy

Option counts vary per row, so the chance level is not 1/N for a fixed N. The uniform baseline is mean_rows(1 / options) and the majority baseline is always answering the most common label index. Both are printed next to every accuracy.

Intervals are a percentile bootstrap over groups, not rows, because several rows can share a state; resampling rows would give an interval that is optimistically narrow.

Figures quoted as mean ± standard deviation are over independent training runs. On the MPS backend a seed does not pin a run, so that spread mixes seed variation with backend non-determinism; see REPRODUCE.md. Three runs is enough to see whether a difference is larger than the noise, and not enough to estimate the noise well.


Controls

An accuracy without a control is not evidence. Three are run on every evaluation:

control what a failure looks like
shuffled state each menu is scored against another row's state. Scoring well here means the model is reading the option prior and the question, not the evidence.
shuffled question each state is paired with another row's question. Scoring well here means the question is being ignored.
untrained weights a container of random weights from blink_synth, scored on the same rows. Landing above chance here would mean the harness, not the model, is producing the number.

Neither shuffled control should reach the intact score, and neither is expected to fall all the way to the uniform baseline: with the question intact but the state shuffled, a model can still exclude options that no question of that shape ever selects, and that residual is real signal about the prior rather than about the evidence.


Perturbations

Output-blind, declared before the results were read:

  • Reversed options. The menu order is reversed and the label index moves with it. A model that reads text rather than position should be unmoved. Because Blink computes each option's logit independently, the expected flip rate is zero up to floating point, and a non-zero rate would be a bug rather than a weakness. tests/c/test_inputs.c checks the same property directly.
  • Irrelevant suffix. A fixed, unrelated sentence is appended to the state. Flip rate and mean total variation of the probabilities are reported.

Timing

bench/bench_latency.c, single-threaded, one process, no other load declared. Each configuration reports p50, p90, p99 and the maximum, because a mean hides the tail an embedded caller actually feels. Sampling stops at 2,000 samples or 1.5 seconds of wall clock, whichever comes first, with a floor of 30; the achieved count is reported with every row so a short sample is visible rather than hidden.

Three paths are timed separately, and the distinction is the architecture:

  • decide — state, menu and question encoded on every call.
  • cached_state — state and menu held, only the question changes. This is the path a service loop uses when it asks several questions about one request.
  • new_menu — state held, a new option list on every decision.

The timed region includes everything inside the library: embedding lookups, the encoder, the key/value writes and the option attention. It excludes opening the model and sizing the session, which happen once per process.

The SIMD comparison rebuilds the same sources with -DBLINK_SCALAR_ONLY=1 (make scalar) and reruns the same benchmark. The scalar loop is the normative definition of the kernel, so this is also a check that it still works.

bench/bench_memory.c reports three things that live in three different places: the mapped weight bytes (shared between processes, never copied), the session arena (caller-owned, sized by the declared limits), and the measured growth in resident size — not peak — when 64 sessions are created over one model and each scores a full-length state. Peak RSS is a process-wide high-water mark and cannot show a marginal cost, which is why the sharing measurement runs first, on an otherwise clean process, and reads the current resident size through task_info on Darwin and /proc/self/statm on Linux.


Reproducibility

  • bash scripts/run_all.sh runs every stage and writes results/MANIFEST.json, which records the SHA-256 of every input, artefact and report together with the compiler, its exact command line, the platform, and the git revision.
  • Data generation is seeded; the corpus manifest carries the checksum of each split.
  • Training is seeded, and the run log for every epoch is kept in results/train-*.jsonl.
  • Row-level predictions are written to results/predictions/ so any reported aggregate can be recomputed without rerunning a model.
  • The build pins -ffp-contract=off, so a fused multiply-add cannot change a result between compilers.

Two things are not bit-reproducible and are not claimed to be: PyTorch training on a GPU backend, and wall-clock timings. The exported container and every accuracy number computed from it are reproducible from a given checkpoint.


Interpretation rules

These are the boundaries of every claim in this repository.

  1. The probabilities are conditional on the supplied options. They answer "which of these", not "how likely is this in the world". Recalibrate on your own workload before treating a threshold as operational confidence.
  2. A typed output can still be wrong. Blink cannot produce a schema violation, and that is a guarantee about the shape of the answer only.
  3. Calibration is reported on the test split, and the temperature was fitted on validation. Calibration measured on the split a temperature was fitted on is not evidence.
  4. Per-family numbers are the result. The pooled accuracy over the synthetic corpus is an average over four deliberately unlike tasks and its value depends on the mix. Read the strata.
  5. Nothing here is a comparison with Jev. No Jev endpoint was run. The only Jev numbers quoted anywhere in this repository are published ones, and they are quoted as context for the interface, not as a baseline that Blink was measured against.
  6. The synthetic corpus is templated. It shows that a mechanism works on held-out entities. It does not show open-domain understanding; WANLI is the external check for that, and it is reported separately with its own denominator.