Spec rulings
Rulings-spec version 2.2.0. Generated by tools/render_rulings.py -- do not edit by hand.
PMK-ARC-001: The filename is the producer label
An archive's route identity comes from its filename slug, which the experimenter fixes at collection time. Row content plays no part in that identity. A slug resolves to a route only through the candidate set's own declared slugs (a slug alone cannot say which dash was a '/'). Nothing anywhere in punchmark derives identity from completion text except the detector under test, whose output is scored against this label.
PMK-ARC-002: Error stubs are counted, surfaced, never featurized
A row carrying exactly {sample, profile, language} and no raw_outputs is an API error stub: the collection failed for that item and the stub occupies its slot with zero draws. Stubs are counted per archive and reported in every ruling (n_stub_rows). Policy caps them, and a stub share above the cap forces UNDETERMINED. Stubs never reach feature extraction. There are no partial rows: any other unknown shape is a typed refusal.
PMK-ARC-003: Two row schemas, tier optional
The native corpus ships rows with and without the tier key (the ladder archives carry none). Readers accept both schemas. Tier is metadata for stratified reporting only, and it is never a feature. Ragged draw counts are tolerated and counted. They are never padded.
PMK-ARC-004: Text only, by construction
The row type carries completion strings and identity fields. There is no place to put logprobs, headers, timing or finish reasons, so a detector that degrades when a provider withdraws a side channel cannot be built here by accident. The instrument is defined on the intersection of what archives reliably carry: response text.
PMK-CAL-001: The resampling cluster is the base sample
Items are not independent: a sample's variants, profiles and languages share program content. The cluster unit for every subsample, split and splice is the base sample name, and clusters move whole. Certificates report cluster counts beside item counts, because the effective sample size is bounded by the number of clusters and a row count would overstate it.
PMK-CAL-002: The null is cross-fitted (superseded by PMK-CAL-005) (superseded)
Superseded: this ruling drew the 2-fold cluster split independently PER ARCHIVE. The corpus shares item content across routes by design, so per-archive folds let route A's held-out cluster leak into the training half through route B's near-identical responses to the same items. That leak inflated the null's left tail until true substitutions fell inside it. PMK-CAL-005 keeps the cross-fitting principle with one fold map shared across archives.
PMK-CAL-003: Null material is same-route, same-window
The false-alarm rate is calibrated exclusively on subsets and split-halves of single archives: one route, one collection window. Same-route comparisons across windows (e.g. a dev archive against a test archive of the same route) are a labeled cross-window diagnostic with no verdict semantics. They are never used as null material. One residual exposure remains: a substitution DURING the calibration window would be baked into the null. That exposure cannot be closed retrospectively and is carried verbatim into every write-up.
PMK-CAL-004: Thresholds are conservative empirical quantiles
The threshold at declared false-alarm rate far is the conservative empirical far-quantile of the null draws (flagging at T < t keeps the empirical flag rate <= far). At scoring time the operating point for an archive of n items is the calibrated point with the largest m not exceeding n. A smaller m has a wider null, which can only widen the flag region's distance from the null. An archive below the smallest calibrated m gets UNDETERMINED. No threshold is ever extrapolated for it.
PMK-CAL-005: The null is cross-fitted with one shared fold map
The null distribution of the set statistic T is built only from out-of-fold scores: cluster names are split into two folds by ONE global assignment shared across every training archive. The corpus shares item content across routes by design, so cluster X must sit in the same fold for every route. Per-archive folds leak held-out content through the other routes' responses to the same items and corrupt the null. That failure was observed in practice, and it is why this ruling supersedes PMK-CAL-002. Each fold is scored by the model fitted on the other. An in-sample null is optimistic and silently overshoots the declared false-alarm rate on held-out data. The shipped model is fitted on everything. Its thresholds come from the cross-fitted null and are therefore slightly conservative. That bias is declared, and no correction is applied.
PMK-CAL-006: Thresholds are per declared route
The null that fixes a threshold is built from the DECLARED route's own same-window archives only. Nulls are never pooled across routes. Pooling would let noise from one near-inseparable candidate pair (observed: two routes whose task answers are byte-identical on most items) widen every route's threshold, silently spending power the separated pairs actually have. An operating point is therefore a (task, route, far, m) cell, and a ruling's threshold always came from the null of the route it rules on.
PMK-CAL-007: A quantile needs draws; draws are budgeted over usable archives
An operating point at false-alarm rate far exists only when the cell's null has at least 5 expected tail draws (n_draws * far >= 5). Below that the empirical quantile degenerates to the sample minimum, whose true exceedance probability EXCEEDS the declared far while showing zero empirical flags. Unresolvable (cell, far) pairs are dropped: the later lookup returns None and the ruling comes out UNDETERMINED (refusal-first, PMK-RUL-002). A fit whose whole grid is unresolvable is refused, with the arithmetic in the message. Null draws are budgeted over the cell's USABLE archives, so an archive skipped for the m or min-clusters floors cannot silently shrink the cell below n_null.
PMK-COR-001: The manifest is the corpus identity
The calibration corpus a ruling cites is identified by the content hash of its MANIFEST, meaning the pinned checkout commit plus per-source sha256. Shipped bytes play no part in that identity. A manifest hash is exactly as binding over pins as over copies: if any pinned source changes, the manifest changes, and every ruling id downstream changes with it.
PMK-COR-002: The shipped corpus mode is manifest + local rebuild
The reference calibration corpus ships no completion text. Re-publishing the upstream benchmark's committed private-test completions in a second repository would be a second, separate disclosure decision, and punchmark declines to take it. Instead, the manifest pins the upstream commit and hashes, and 'corpus rebuild --corpus
PMK-CRT-001: A certificate derives from exactly one ruling
The certificate is a one-line text plus a certificate/v1 JSON document, both generated from one verified ruling body. It never aggregates rulings and never carries a number that is not in its ruling or model file. HOLDS maps to exit 0, DOES NOT HOLD to exit 1, UNDETERMINED to exit 2, so an undetermined certificate can never read as success.
PMK-CRT-002: No weights claim, ever
Every certificate carries the sentence: 'This certifies the route label as served within the named candidate set; it is not a statement about model weights.' The sentence is part of the golden-pinned line format. Dropping or weakening it is a spec MAJOR change. A public route name does not denote a fixed configuration, and this instrument does not repair that.
PMK-CRT-003: Verdicts are closed-set
Identification and substitution verdicts are relative to the declared candidate set. A producer outside the set will be mapped to its nearest member. SAME-PRODUCER means 'best match within candidate set C', and it is not an identity proof. The certificate names the candidate set id and count in the same sentence as the verdict.
PMK-DET-001: The detector seam
A detector implements DetectorModel/FittedModel over ResponseSets. It never opens a file and never imports a reader. It emits per-row per-candidate scores with a shared orientation (higher = closer to that candidate) and serializes to plain data that round-trips exactly. Every tunable choice inside a detector is itself a numbered ruling.
PMK-DET-002: The trivial reference detector
The scaffold ships 'trivial': per-(route, task) character-unigram centroids scored by negative L1 distance. It is tuning-free by design, so calibration, power, rulings, certificates and the gate are exercised end-to-end against planted truth before any real detector exists. It is a reference implementation. It is not a recommendation.
PMK-DET-003: The chargram multinomial
The calibrated detector is a per-(route, task) Jeffreys-smoothed multinomial over hashed character n-gram buckets: theta_{r,b} = (c_{r,b} + 0.5) / (C_r + 0.5 * D) with the chargram/v1 spec (orders 3-5, D = 2^18, crc32 hashing). Per-row evidence is the per-gram- normalized log-likelihood, so a row is one evidence unit at any draw count. The fit is closed form with no optimizer. Posteriors are never used: thresholds are empirical null quantiles. Log-probs are rounded to 6 places at fit time so params are compact, byte-stable and round-trip exactly.
PMK-DET-004: The primary view is CANON@1, frozen
Shipped calibrations fit and score on the CANON@1 text view, chosen and frozen before any held-out scoring. A detector may not be fitted on one view and scored on another (the model file records the view and reconstruction refuses a mismatch). RAW@1 and ABL@1 exist so the recorded formatting-ablation study can ask whether the fingerprint is a prompt-template artefact. Results on those views never ship as a calibration.
PMK-EMIT-001: Committed artifacts are byte-stable
Every committed artifact (fitted-model file, ruling line, certificate, manifest, golden) is serialized through canonical.py: floats rounded half-even to 6 places with -0.0 normalized, non-finite floats refused, sorted keys, 2-space indent, ASCII, trailing newline, gzip with mtime=0 and no embedded filename. Regeneration from identical inputs must be byte-identical. The drift gate compares the bytes themselves, so output that means the same thing but differs in bytes still counts as drift.
PMK-EMIT-002: Documents carry their own content address
Ids of model files, rulings, certificates and manifests are content addresses:
PMK-EMIT-003: Every seed is derived, never sampled
No stochastic procedure seeds itself from the clock or from process state. Every seed derives from labelled parts (procedure name, scope, index, the caller's seed) through sha256, so identical invocations replay identical resamples on any machine.
PMK-FEA-001: Features consume raw_outputs strings and nothing else
Feature extraction receives the completion strings of one row and no identity or bookkeeping fields. Stubs, tier presence and schema shape correlate with split and archive in the native corpus, so a feature that could see them would learn the split instead of the producer. Enforced by signature: the featurizer takes a sequence of strings.
PMK-FEA-002: One row is one evidence unit
At temperature 0 most draw sets collapse to one or two distinct strings, so draws are pooled per row (counts per distinct string, scaled by multiplicity) and likelihoods are normalized per gram. A row contributes the same evidence weight whether its k draws are identical or not. k is never treated as k independent observations. Power statements are made in items, qualified by measured effective draws.
PMK-FEA-003: Text views are versioned ids
A text normalization is a versioned view (RAW@1, CANON@1, ABL@1) named in every fitted model and ruling id. Changing a normalization means creating a new view version. An existing view is never edited. The primary view is frozen before any held-out scoring, and fitting on one view while scoring on another is refused.
PMK-GTE-001: Tri-state exit codes
Gate and certify preserve the three-way distinction: 0 pass/holds, 1 measured fail, 2 unevaluable-or-undetermined. Both non-zero: UNAVAILABLE is never PASS, and a measured failure outranks a missing input. argparse usage errors land on exit 2, the 'not a measured outcome' side.
PMK-GTE-002: An empty comparison never passes
A gate that checked nothing has not gated anything. A baseline with no operating points, a model file with none, or a selection that matches nothing is unevaluable (exit 2), and it is never a pass. A typo in a filter must not turn CI green.
PMK-GTE-003: Operating points move only with a declared bump
The gate fails (exit 1) when any calibrated operating point, the calibration corpus hash, or the candidate set differs from the committed baseline while the detector version and the spec MAJOR are unchanged. A silently moved threshold is the failure mode this instrument exists to make visible in others. Punchmark holds itself to the same rule, in-repo (the drift test) and downstream (this gate).
PMK-POW-001: Substitutions are seeded on the score table
The power population is built by splicing: for declared route r0 and substitute r', a fraction rho of a subset's rows (whole clusters, matched by item key) are replaced with the same items' scored rows from r''s archive. The splice happens on cached scores and never touches text, so the evaluation set is guaranteed to contain a substitution of known size and identity.
PMK-POW-002: rho = 0 must reproduce the null
An unspliced (rho = 0) power draw is a null draw. The self-check asserts the flag rate of rho = 0 subsets at the calibrated threshold comes out <= the declared false-alarm rate up to resampling noise. A violation means the splice machinery or the calibration is broken, and the run refuses to report power.
PMK-POW-003: The minimum resolvable separation is a first-class output
For every ordered candidate pair, task and set size, punchmark reports rho: the smallest substituted fraction on the grid detected with the calibrated power at the declared false-alarm rate, or None when even a full swap is unresolvable. rho makes a null result legible as a power limit rather than as reassurance. A SAME-PRODUCER verdict is refused (UNDETERMINED) when the caller's rho_target is below what the archive could have resolved.
PMK-POW-004: Seeded fidelity is a standing bound
Spliced substitutions are a best case: real substitutions arrive with a date offset, hit time-contiguous traffic, and may involve a producer outside the candidate set. The fidelity of the seeded population to real vendor changes cannot be validated from committed data and is printed as a standing bound wherever power numbers appear.
PMK-POW-005: rho* quantifies single-donor substitution only
The splice draws its swapped rows from ONE substitute per power point: the donor loop enumerates candidates one at a time and _swap replaces whole clusters from that single donor. rho therefore bounds substitution by a single candidate at the stated fraction, and a mixture -- two or more candidates each serving some share of the archive -- is outside what rho bounds, whether or not every mixture component is in the candidate set. The score output prints this scope beside the verdict so a ruling's power statement is not read wider than the population it was measured on.
PMK-RUL-001: Three-zone verdicts; UNDETERMINED is first-class
A ruling is SAME-PRODUCER, SUBSTITUTED or UNDETERMINED. SAME-PRODUCER means exactly: a substitution of fraction >= rho_target by any candidate in the set would have been flagged with the calibrated power, and none was. It never means 'the model did not change'. UNDETERMINED is a verdict in its own right: the archive was too small, too clustered, too stubbed, below the calibrated floor, or below the power needed for the caller's rho_target. It is never rounded into either other verdict.
PMK-RUL-002: Refusal conditions are enumerated and mandatory
A ruling must come out UNDETERMINED when: usable items are below the m floor; clusters are below the c floor; the stub share exceeds its cap; no operating point is calibrated at the requested false-alarm rate for the archive's size; the power table cannot support the caller's rho_target. A ruling must be refused outright (no verdict) when: the archive has no window sidecar; the declared route is outside the candidate set; the task was not calibrated. Absence of an alarm is never reported without the power that an alarm would have required.
PMK-RUL-003: The store is append-only; rulings are superseded, never edited
Ruling lines are appended to a JSONL store and never modified. Every line's id is a content hash re-verified on every read, and a supersedes field names an earlier ruling in the same store. An edited line is tamper, detected and refused. Changing one's mind is a new ruling that supersedes the old one, with both remaining readable.
PMK-RUL-004: A ruling id pins the contractual four
The ruling id hashes a body that includes the detector id and version, the candidate set id, the operating point (far and threshold), and the calibration corpus hash. The body also carries the archive hash, window, spec version and verdict. Identical inputs reproduce identical ids, and any change to what a verdict rests on changes the id.
PMK-RUL-005: Cross-split scoring is declared, never inferred
An archive whose filename task is a different split of a calibrated task family (e.g. comprehend_test against the model task comprehend) may be scored under that model task only through an explicit caller declaration (--task-as). The ruling records both names: task stays the filename truth, scored_as records the model task the thresholds and power table came from. Punchmark never infers the mapping from names or content, and a certificate for an aliased ruling says 'scored as' in the same clause as the task.
PMK-SDC-001: The window is declared, never inferred
The collection window is by-construction metadata the caller declares in a window/v1 sidecar. Punchmark never infers time from content, filenames or mtimes. Without a sidecar there is no (route, window) unit, so fit, score and certify refuse. The refusal message prints the exact JSON the caller must write.
PMK-SDC-002: The sidecar must agree with the filename
The sidecar's declared route must slug-match the archive filename and its task must equal the filename task. The filename is the oracle (PMK-ARC-001), so a sidecar may resolve a slug to a route name but never contradict it. Disagreement is a refusal. It is not reported as a warning.
PMK-SDC-003: The sidecar binds to the bytes
A sidecar pins the sha256 of the archive it describes and is refused against any other bytes, so window metadata cannot be quietly re-pointed at different data.