# 092 frozen exploratory protocol — 2026-10-05, before inference Question: does the previously selected 091 target_only checkpoint retain its measured gains across instructions and a training-excluded task family? Reuse091 public commit7bd49bd7d6c1cd950d58506cfe7d4e962d161c85 and its predeclared selected target_only arm, NOT select a new arm based on092 scores. Base model means original already instruction-tuned checkpoint, not an untrained model. Six templates in templates.py: original JSON; task-sentence paraphrase with identical JSON contract; swapped task/contract order with every character otherwise preserved; original plain contract; original news; paraphrased news task with unchanged A/B/C/D mapping. Same source text byte-for-byte. No examples, label-order changes, translated inputs or tokenizer chat-template changes. No template sweep or choice of winner. Original JSON vs paraphrase is the PRIMARY comparison, reordered is secondary, plain is output-contract diagnostic. Report all cells and per-row wins/ties/losses for F/C/J. Two model states, same64 SST2 development rows and32 balanced AG News development rows as091. 640generations total, no confirmation inference. 192 reserved rows untouched. News is excluded from target_only incremental training, but already used in091 model selection and may occur in pretraining; thus task-family training holdout, NOT blind unseen-task evaluation. Human semantic/label-preservation review is REQUIRED and PENDING; prepare all 96 rows with six templates without model answers in private review packet, blank human decisions. Automated label/text identity checks do not substitute for a human. All numerical results before human review remain exploratory. No confidence intervals or multi-seed claim. Run CPUfloat32/eager/4threads, greedy32, seed92 for eval; max512/no truncation. Fixed original12fd25f... and tokenizer, pinned data/resources. Use hash-verified local091 target_only checkpoint for full run. Reader may regenerate exactly same64SGD steps on32SST2 rows from pristine weights, seed91/lr.001/batch2/clip1/no momentum/decay/eval-mode autograd/assistant-through-EOS mask/1024supervisedtokens. Smoke first regenerates all64steps (train budget unchanged) then evaluates2SST+4news per model=32generations; verify parameter digest against saved091 weights. Training regeneration needed for portable artifact, not a new experimental arm. Budget1200seconds/6GiB RSS. Extrapolate inference from smoke before full run. Save timing/RSS/shapes/commands/exitcodes/full private outputs and public non-echo outputs, redacting any answer containing >=8consecutive source words. Redacted outputs keep hashes/derived scores and are counted in audit limitations. Raw review/news data and weights not redistributed because data cards license unknown. All self-written modules included, identity contains code/data/model/tokenizer/prompt/decode and checkpoint hashes. Do not tune on failure or consume reserved data.