Skip to content

Spec rulings

Rulings spec version: 1.0.0. A ruling is never edited to mean something else; it is superseded, and the spec version moves. limen spec list prints this same registry from the installed package.

LMN-ADP

LMN-ADP-001 — Refuse a poisoned environment before importing

Status: active. bench.tasks reads BENCH_STRIP_ANNOTATIONS at import time into a module constant; a truthy value silently regrades against the unannotated corpus. The adapter checks the environment and refuses BEFORE any foreign import, and never modifies the checkout (imports run with bytecode writing disabled for the process).

LMN-ADP-002 — Split is injected from the filename, always

Status: active. The checkout's regrade functions default a missing 'split' to 'dev', which silently mis-grades test-split records against the dev oracle (tier A) or crashes (tiers B/C). The adapter sets rec['split'] explicitly on every record from the archive filename and never relies on the default.

LMN-ADP-003 — Per-draw verdicts via public singleton regrade, no rounding

Status: active. Per-draw verdicts come from the checkout's public regrade API called on singleton copies (raw_outputs = [one draw]); the aggregate rate of a one-element list is that draw's verdict. Anything non-integral raises; the adapter never rounds a rate into a verdict. refactor x test is refused outright as upstream declares it non-reproducible.

LMN-AUD

LMN-AUD-001 — Per-item instability is min(s, k-s)/k, symmetric, never compounded with correctness

Status: active. u_i is the fraction of a system's k draws disagreeing with the item's own majority verdict: u_i = min(s, k-s)/k. At even k with s = k/2 there is no majority verdict and u_i = 0.5, the maximum. u_i is computed for both systems of a pair symmetrically and is never combined with correctness: an always-fail cell is stable. Retry-free coverage (arXiv 2606.00920) compounds stability with correctness and is reported beside u_i, never instead of it.

LMN-AUD-002 — The v1 threshold is u_i == 0 for both systems, crude by design and versioned

Status: active. The stable-for-both partition uses the declared threshold version u0: u_i == 0 for both systems. It is crude by design; the principled benchmark it must be measured against is an IDR-style threshold (Li, Brown, Huang, Bickel 2011). A threshold change is a new threshold version and a spec supersession, never an edit, and every ruling prints the threshold version it used.

LMN-AUD-003 — The replicate noise band is enumerated, never sampled

Status: active. The band is the p95 (lower interpolation, exact rationals) of |self-gap| over all complementary half-splits of one system's k draws, both systems of the pair pooled, computed on the ruling's item set with each half normalized by its own width; the maximum is printed beside it. Unordered splits are deduplicated by keeping the half containing draw position 0 at even k. No RNG is involved, and the enumeration is never thinned or sampled: above the printed cap of 400000 splits (k of 23 or more) the band and every band-dependent statistic rule UNAVAILABLE with the refusal stated. A zero band is degenerate and bounds nothing.

LMN-AUD-004 — Rulings carry a capped, deterministic decisive-item witness

Status: active. Every SURVIVES ruling carries the greedy removal margin (stable items removed largest signed contribution first, item id ascending on ties, band held fixed, until the ruling changes); every SIGN-INVERTS or FALLS-INTO-NOISE carries the greedy re-inclusion witness, or NO_WITNESS when even full re-inclusion does not rule SURVIVES. At most 25 item ids are printed with the full count. The witness suffices to change the answer and is deterministic; minimality is claimed only on the stable side, where contributions are quantized.

LMN-AUD-005 — Gap-survival rulings never ship without their selection mitigations

Status: active. The same draws classify items unstable and estimate the gap, so naive exclusion erodes the gap even under a null. Every gap_survival section carries the disjoint classify/audit split analysis and a selection null that resamples every cell i.i.d. Bernoulli(p-hat) and re-runs the entire partition-and-gap pipeline, banding gap shrinkage, unstable share, and the ruling frequencies. The noise band is held fixed under the null: the null models per-cell rates, and the band, alone among audit statistics, is sensitive to cross-item draw alignment. No API returns the naive ruling alone.

LMN-AUD-006 — Ruling precedence is fixed

Status: active. In order: a pooled tie or an empty stable-for-both partition rules UNAVAILABLE; SIGN-INVERTS when the stable-gap sign is the negation of the all-items sign, both nonzero, taking precedence over the band with the within-band flag printed; FALLS-INTO-NOISE when the stable gap is a tie or its magnitude is below the band p95; otherwise SURVIVES. All signs are integer pass-count signs; all comparisons are exact rationals.

LMN-AUD-007 — Differentiation from retry-free coverage ships in every audit

Status: active. Every gap_survival section reports both partitions' sizes, the Jaccard overlap of the excluded sets with its integer counts, the count of stable-but-always-wrong items (the exact set on which the two criteria differ, by containment), each system's retry-free coverage, and whether the comparison's verdict differs under the per-system coverage difference versus the u_i partition. A joint-kept RFC gap would be identically zero by construction and is deliberately not used. The section exists to falsify the re-labelling objection, not to decorate it.

LMN-AUD-008 — Stratum rulings have a printed floor

Status: active. A per-stratum ruling is issued only when the stratum holds at least the declared floor of aligned items (default 30, recorded in the envelope options); below it the stratum rules UNAVAILABLE with the floor printed. Unlabelled items are counted per key and never silently pooled into a stratum. The saturation rollup is an association across strata of one archive; no mechanism is claimed (NO_SATURATION_MECHANISM_CLAIM).

LMN-CORE

LMN-CORE-001 — Binary verdicts only; no thresholding

Status: active. A verdict enters limen as a literal 0 or 1. limen never derives a verdict from a score: choosing a threshold is a measurement decision that belongs to the evaluation, not to the instrument that audits it.

LMN-CORE-002 — Table invariants are hard errors

Status: active. Duplicate (model, task, item_id, draw_id), non-binary verdicts, and partially present per-cell optional fields raise; limen never silently repairs its input.

LMN-CORE-003 — Cells below min-k are excluded and counted

Status: active. Cells with fewer than min_k (default 2) draws are excluded from analysis and reported with their count; the build refuses only when no cell remains. One draw cannot be observed to flip.

LMN-CORE-004 — Canonical draw order

Status: active. If every draw_id in the table parses as an integer, draws order numerically, else lexicographically. 'Draw position d' means the d-th draw in this order, everywhere.

LMN-CORE-005 — Long CSV is the canonical on-disk form

Status: active. The generic long-format CSV (model, task, item_id, draw_id, verdict, plus optional score, collected_at, model_version, raw_sha256) is the canonical serialization of the verdict table; the calibration corpus and limen synth write it deterministically.

LMN-CORE-006 — Absent readers are stated, not silent

Status: superseded. Readers deliberately absent from this release (inspect_ai .eval per-epoch, HELM, Parquet) are named in the reader registry and the docs; their absence is a scope decision, not an accident.

LMN-CORE-007 — Absent readers are stated, not silent

Status: active. Readers deliberately absent (HELM, Parquet) are named in the reader registry and the docs; their absence is a scope decision, not an accident. Supersedes LMN-CORE-006 after the inspect_ai .eval per-epoch reader shipped: each epoch of each sample is one draw, read directly from the zip without an inspect_ai dependency, consuming exactly the layer inspect's epoch reducers collapse.

LMN-CORE-008 — Per-item labels are optional, all-or-nothing, and item-consistent

Status: active. Rows may carry label_ columns naming per-item strata. Within a cell labels are all-or-nothing and identical across draws, and identical across every cell of the same (task, item); violations are hard errors. Label names match [a-z0-9_]+ and values are non-empty. Label-free rows keep their exact pre-label digest bytes, so the dataset digest of an unlabeled table is unchanged by label support.

LMN-DRF

LMN-DRF-001 — UNAVAILABLE is never PASS

Status: active. The drift guard's overall state is FAIL if any sub-check fails, else UNAVAILABLE if any sub-check could not run (missing collected_at or model_version), else PASS. A missing field can never launder into PASS.

LMN-DRF-002 — Position-proxy mode can fail but cannot pass

Status: active. When timestamps are absent and the caller declares that within-cell draw order is collection order, the order-based sub-checks run on positions: a found effect is FAIL (real evidence of trouble), a clean result is UNAVAILABLE with the disclaimer attached — collected_at is still missing.

LMN-EMIT

LMN-EMIT-001 — Canonical bytes

Status: active. Ruling documents serialize as json.dumps(sort_keys=True, indent=2, ensure_ascii=True) plus a trailing newline, floats rounded half-even to 6 places with -0.0 normalized, lists in documented sort order, gzip members written with mtime=0. Regeneration from the same input is byte-identical.

LMN-EMIT-002 — No count without its denominator

Status: active. Every rate in a ruling document is emitted as {count, denominator, rate}; a bare rate never appears.

LMN-EMIT-003 — Seeds are derived, never wall-clock

Status: active. Every stochastic procedure derives its seeds as sha256 over (rulings_version, scope, procedure, index). No unseeded RNG and no clock anywhere in a ruling body's production.

LMN-EMIT-004 — Body and provenance are split

Status: active. The ruling body contains no timestamp, package version, hostname, or absolute path — nothing that varies across a faithful regeneration. Provenance lives in a sidecar outside the byte comparison; the body pins its input via dataset_digest.

LMN-EMIT-005 — The scope block ships in every report

Status: active. Every report embeds the fixed does_not_show scope codes (NO_MODEL_QUALITY_CLAIM and companions) so a ruling document cannot circulate without its own limits.

LMN-EMIT-006 — Schema report/v1 has no variance-components section

Status: superseded. The item/model/draw variance decomposition is deliberately deferred: 8 draw levels give a wide interval, and shipping a wide interval as a headline teaches people to trust the wrong thing. Adding the section is a schema version bump, not a field.

LMN-EMIT-007 — Schema report/v2 carries the variance-components section, subordinated

Status: active. report/v1 deliberately had no variance-components section; adding it was a schema bump, exactly as LMN-EMIT-006 prescribed, and this ruling supersedes it. In report/v2 the section exists only under the LMN-VAR guardrails: subordinate placement, mandatory intervals, bucket and never-headline notes, and a gate that never reads it (LMN-VAR-004).

LMN-EMIT-008 — Additive schema evolution

Status: active. A committed golden may be replaced under --write without --confirm-spec-bump only when the regenerated document strictly adds: after stripping every content_hash and exempting the spec_version and limen_schema stamps, every committed leaf is present and byte-equal at the same path, arrays are equal-length and compared positionally, and the rulings spec has moved by at least MINOR. A changed or removed value is a changed recorded meaning and still demands a spec MAJOR and the explicit flag. Adding is cheap by design so that adding is never smuggled in as changing.

LMN-FLK

LMN-FLK-001 — Per-item flakiness is the pairwise-disagreement U-statistic

Status: active. f = s(k-s)/C(k,2), the fraction of draw pairs whose verdict differs; unbiased for 2q(1-q) at every k. Items classify always_pass / always_fail / mixed, and mixed iff f > 0.

LMN-FLK-002 — The constant-verdict fraction is an upper bound on TARa@N and says so

Status: active. The constant-verdict fraction is reported with an explicit note that it upper-bounds TARa@N (Atil et al., arXiv:2408.04667): TARa counts parsed-answer agreement, and two different wrong answers grade to the same 0 verdict, so verdict constancy >= TARa@N — equality is unverifiable from a verdict table. The field is emitted only when k is uniform across cells; at ragged k it is null because the quantity at mixed N is not comparable to the published metric.

LMN-GRD

LMN-GRD-001 — A grader defect is a byte-identical flip

Status: active. An unordered draw pair with identical raw completion bytes and differing verdicts is a grader defect: the model produced the same bytes and the grader decided differently. Identity is byte identity — no strip, no case-fold, no normalization. The count is a share of all discordant draw pairs, and is reported beside raw flakiness, never replacing it.

LMN-GRD-002 — No raw text means UNAVAILABLE, not zero

Status: active. When raw hashes are absent the grader-defect state is UNAVAILABLE with null counts. 'Zero defects found' is a finding; 'no text to check' is not.

LMN-GTE

LMN-GTE-001 — Exit codes preserve the three-way distinction

Status: active. 0: every requested check passed. 1: a requested check measurably failed. 2: a requested check could not be evaluated (UNAVAILABLE section, missing pair, unreadable report). Both 1 and 2 are red in CI, so UNAVAILABLE can never slip through as success; a measured failure outranks a missing section when both occur.

LMN-GTE-002 — The claimed improvement is the report's own delta

Status: active. The gate audits a report against itself: the claimed improvement for a pair is the report's own |delta_pool|. The gate accepts no externally typed effect size; with --pair task:A>B the asserted direction must match the report's pooled sign or the pair fails with claim_contradicts_pooled_data.

LMN-GTE-003 — Every gate failure reprints the quality-claim boundary

Status: active. Every FAIL verdict reprints NO_MODEL_QUALITY_CLAIM: a sign ruling is a verdict on the measurement, never on the models, and a failed pair never means the other model wins.

LMN-GTE-004 — The gate accepts report/v1 and report/v2

Status: active. report/v2 is additive: every field the gate reads is byte-identical in shape and meaning in report/v1, so both schemas are accepted; v2-only checks rule UNEVALUABLE on v1 documents with the regeneration named. An unknown or newer schema exits 2 (unevaluable), never 0. A schema leaves the accepted set only when a change removes or alters a gate-read field, which is a spec MAJOR.

LMN-NSE

LMN-NSE-001 — The MDD is cited, conservative, and carries its assumptions

Status: active. MDD = t_{0.975, k-1} * sqrt((sd_a^2 + sd_b^2)/k), after Kalibera & Jones (ISMM 2013) specialized to one level of repetition, with the conservative df = k-1 and a hardcoded t-table (floor lookup for untabulated df). Every MDD block prints its assumption list verbatim, flags k < 4 as low_k, and flags zero observed spread as degenerate rather than treating it as zero noise.

LMN-RNK

LMN-RNK-001 — Signs are computed on integers

Status: active. Every leaderboard sign is the sign of an integer pass-count difference, never of a float subtraction; a drawn tie must be exactly a tie.

LMN-RNK-002 — Ties are their own count

Status: active. A single-draw tie is neither agreement nor disagreement with the pooled direction; n_agree, n_flip, and n_tie are all printed with denominator k, and the flip rate keeps ties in its denominator (a tie-excluded rate is printed beside it).

LMN-RNK-003 — A pooled tie is SIGN-UNSTABLE

Status: active. When the pooled difference is exactly zero no directional claim is supported: the ruling is SIGN-UNSTABLE with pooled_tie set, per-draw agreement fields are null (agreement with a nonexistent direction is undefined), and the gate fails the pair.

LMN-RNK-004 — Zero flips carry their exact upper bound

Status: active. When no flip is observed in k draws, the ruling prints the one-sided 95% upper bound on the per-draw flip probability, 1 - 0.05**(1/k) (0.312 at k=8), so SIGN-STABLE cannot be read as flip probability ~ 0.

LMN-RNK-005 — Stable-only numbers never ship without their mitigations

Status: active. The same draws that classify an item as flaky also rank the models, so naive exclusion looks tidier even under a null. Every stable-items-only number is emitted together with the disjoint classify/rank split analysis and the selection null; there is no API that returns the naive view alone.

LMN-RNK-006 — The selection null resamples verdicts; permutation is invariant

Status: active. Within-cell draw-label permutation changes nothing here — classification, scores, and tau are functions of per-cell pass counts only — so a permutation 'null' would tautologically pass. The null resamples each cell i.i.d. Bernoulli(p-hat) and re-runs the entire selection pipeline, with seeds derived from the rulings version.

LMN-VAR

LMN-VAR-001 — Variance components are exact EMS method-of-moments, stdlib only

Status: active. The two-facet crossed items x draws decomposition of the binary verdict uses balanced-design ANOVA sums of squares computed exactly on integers (SS_item = sum s_i^2/k - T^2/nk, SS_draw = sum t_d^2/n - T^2/nk, SS_residual by subtraction) and the EMS relations E[MS_item] = s2_res + ks2_item, E[MS_draw] = s2_res + ns2_draw, E[MS_res] = s2_res. A likelihood GLMM is deliberately not used: it is not computable pure-stdlib, not byte-deterministic, and lives on a link scale that does not decompose the observed score. Ragged k is UNAVAILABLE, never approximated. Negative moment estimates truncate to zero and the raw value is printed beside the truncated estimate.

LMN-VAR-002 — Component intervals are seeded item-bootstrap percentiles

Status: active. Cells are resampled with their whole draw vectors, seeds derived per LMN-EMIT-003, bands by lower-interpolation percentiles, with the truncation share printed. The draw-component interval is conditional on the observed k draws and its assumptions list says so. Satterthwaite chi-square intervals are deliberately not used: the effective df is continuous and stdlib has no chi-square inverse, mean squares of Bernoulli verdicts are not scaled chi-squares, and zero-truncation breaks the pivot.

LMN-VAR-003 — The section is subordinate and the draw facet is a bucket

Status: active. Every variance_components section carries the fixed bucket note (NO_FACTOR_ATTRIBUTION) and never-headline note, and emits the fixed low-k warning whenever the draw facet has fewer than 20 levels. The CLI summary never prints a component; report.md renders the section only after the scope block with its warnings above the numbers; no component appears without its interval.

LMN-VAR-004 — The gate never reads variance_components

Status: active. No gate check may consult the section: deleting every variance_components section from a report leaves every gate line, verdict and exit code identical. The section informs planning, never pass/fail.

LMN-VAR-005 — The model facet is descriptive, never a component

Status: active. Models are fixed choices, not a sample from a universe. The task-level model facet is the ddof=1 variance of the pooled scores, labelled descriptive, with no interval and no random-model claim; a three-facet random-model decomposition is deliberately absent because a model variance component would be a model-quality claim (NO_MODEL_QUALITY_CLAIM) estimated on one to three degrees of freedom.

LMN-VAR-006 — Design effect and planning are defined and cited

Status: active. deff = 1 + (k-1)icc_item + (n-1)icc_draw = (ks2_item + ns2_draw + s2_res)/s2_total; n_eff = nk/deff; null when total variance is zero. The draw-facet contribution to the pooled score scales exactly as 1/k, so halving it doubles k; planning numbers cite Kalibera & Jones (ISMM 2013) and issue no sufficiency certificate (NO_K1_CERTIFICATE).