# 089 pre-run protocol — 2026-10-02 Question: Does response-only supervision have the intended shifted targets and gradients, and how do four tiny updates affect format/content? Reuse 088 checkpoint/tokenizer, system/task/plain/JSON contracts, CPU float32 eager/4 threads/greedy32. Official SST-2 train hash-sort SHA256("89:"+idx), first8; no outcome-based selection. Validate no exact normalized sentence overlap with selected development or reserved confirmation before training (abort if any). Validation hash-sort SHA256("88:"+idx), first8 of prior64 development; next128 after64 remain unqueried. Two arms from same pristine weights: full sequence vs assistant response through EOS inclusive. Trailing template newline excluded from both. JSON targets from original binary labels. All model parameters trainable, SGD lr0.001/no momentum/no decay, clip1, batch2, sequential4steps, one pass8rows. Same inputs/order/steps; supervised token count and denominator necessarily differ, NOT equal supervised-token or equal-update-norm experiment. No packing; right pad EOS with attention0, labels-100 only at padding (and prompt in assistant arm). No pre-shift of labels. Primary checks: manual shifted CE matches official; first response/EOS/pad targets correct; ignored logits direct gradient zero but prompt input embedding gradient nonzero; finite nonzero global gradient, parameter changes. Behavioral diagnostics: F/C/J before/after on8 development×2 contracts. All raw outputs retained. No success threshold or superiority claim, no tuning from outcomes. Independent truncation diagnostics at64/96/128/256 on all8 train sequences; never silently train a truncated answer. Formal training aborts above256. Report lost/all-empty targets. Do not force an artificial long sample. Smoke: first2 train/1step perarm, first2 dev×2contracts, max900s/6GiB; estimate full costs before formal4step run. Save final state_dict locally (not public archive), hashes, step records, per-position masks; raw review/token traces local-only due unknown dataset license. Metrics computed from outputs not teacher-forced loss. No early stopping, search, confirmation evaluation, human annotation, invented baseline, or full SmolLM training reproduction. Single seed89, tiny diagnostic experiment.