# 088 pre-inference protocol — 2026-10-01 Question: On a fixed small instruct checkpoint, how much does an output contract alter measured sentiment-label correctness and format compliance? Model HuggingFaceTB/SmolLM2-135M-Instruct revision 12fd25f77366fa6b3b4b768ec3050bf629380bac; same tokenizer. Official Transformers 4.49.0, torch2.6.0 CPU float32/eager attention, 4 threads, batch1, eval/no gradients, seed88. No training or checkpoint selection. This is an already instruction-tuned starting point, not a base-pretrained model. Public Stanford SST-2 validation revision 8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb. Sort all872 rows by SHA256('88:'+idx). First64 are development; next128 reserved confirmation, NOT inferred or used for examples; remaining680 unused. Official test not downloaded. No label balancing or output-dependent selection. First2 development rows smoke; same64 formal after resource check, therefore formal is developmental, not untouched confirmation. Training split and smol-smoltalk not downloaded this run; smol-smoltalk remains candidate mixing data for later unit work. Two paired contracts on exactly the same64 sentences: plain lowercase label vs JSON object with only sentiment key. Both use same system message, same official chat template with add_generation_prompt=True, same classification instruction and label order. Only output instruction changes. Zero few-shot. Greedy max_new_tokens32, no sampling, one beam, checkpoint EOS. No input truncation; fail if input>512 rather than silently slice. JSON has more prompt/answer tokens, so paired results estimate total contract change, NOT equal-compute or pure syntax causal effect. Primary metrics, all divided by64: format-valid rate F and joint format-and-correct rate J. Secondary explicit-label correctness C: find case-insensitive whole words positive/negative in response; exactly one distinct label required; absent/both count wrong. This is an automated proxy, not human content truth; explanations/negations may fool it. Plain format is stripped response exactly positive/negative. JSON format is a syntactically valid object with exactly one nonduplicate key sentiment and lowercase allowed value; whitespace permitted, fences/additional keys forbidden. Retain four F×C buckets, conditional C|F with denominator, missing/ambiguous count, token counts, EOS/limit stop, raw output/token IDs and per-class accuracy. Constant positive/negative classifiers form a no-model reference. No accuracy-driven retries or new prompts. Resource rule: smoke2×2 completions; if projected64×2 walltime<900s and RSS<4GiB, continue; otherwise retain incomplete and revise budget explicitly before further runs. Formal timeout900s. No CI or generalization claim. Diagnostic cases: smallest dataset idx in each observed F×C bucket per condition; cases generate hypotheses, not causal labels. All failures retained. Independent audit will recompute metrics from raw outputs and original labels; no human gold claimed. Cache identity includes model/tokenizer revisions + file hashes, prompt/config, data revision/hash, selected IDs, code hashes and decoding. Output directory must be new. No implicit author-machine paths. Training checkpoints absent because no training. Original validation membership and potential pretraining contamination remain limitations.