# 091 preregistered development protocol Saved 2026-10-04 before any 091 model inference/training. Reuse pristine SmolLM2-135M-Instruct, SST2 model/data revisions and 090's32 class-balanced training rows/64 development/128 untouched confirmation. This is small-budget task rehearsal, not proof of catastrophic forgetting or recreation of pretraining. Retention task is independent AG News 4-class classification (A World, B Sports, C Business, D Sci/Tech), pinned fancyzhx/ag_news snapshot eb185aade064a813bc0b7f42de02595523103ca4. Add row-position idx before SHA256('91:'+idx) ordering within each class. Train first8/class; test first8/class becomes32-row DEVELOPMENT, next16/class64 reserved untouched. Never present these test-derived development rows as a final benchmark. No length filtering or truncation; fail if selected train>256 or eval>512. Normalize whitespace/lowercase exact overlap checks; if overlap occurs fail, do not silently replace. No semantic dedup or human labels claimed. Same raw weights, SGD lr.001/momentum0/decay0/clip1, CPUfloat32/eager/4threads/seed91/eval autograd. Single run per arm, no checkpoint selection. Target step is2 opposite-label SST2 JSON examples,8targets/row=16. News step is8 examples (two of each class), each one-letter answer+EOS=2targets, accumulated in4microbatches of2. Sum token loss across microbatches divided by16 BEFORE backward; clip/step once. Check token lengths rather than assume. Row streams cycle in fixed class-interleaved order; no shuffle. Main arms target_only64target steps; replay25 pattern T,T,T,N repeated16; replay50 pattern T,N repeated32. Each64updates/1024supervised tokens, news fractions0/25/50%. Target_half32target steps (512tokens) diagnoses target dilution matched to replay50; intentionally not equal total budget. All target examples match090 order; later news batches never use development. Rehearsal raises input tokens/FLOPs/row exposure; not a compute-matched claim. Primary outcomes: SST2 JSON strict format-and-label joint accuracy and AG News exact single-letter joint accuracy. News content proxy extracts a unique standalone uppercase A/B/C/D; full text kept, not human semantic truth. Existing SST2 plain and separate format/content counts remain diagnostics. Compare paired per-example transitions; noCI/no multi-seed. All160prompts per state (64plain+64JSON+32news), baseline+4arms=800generations, greedy32. Baseline ability is to be measured, not assumed. If baseline news at/below25%, report floor limitation rather than choose another task. Development-only selection rule: among three main arms, require news joint correct>=baseline news correct-1; maximize target JSON joint correct, break ties by smaller news fraction. If no arm qualifies, choose none. Do not use128SST/64news confirmation samples in selection or generation. Provisional choice only; retain every failure/negative result. Smoke uses first2 SST dev/first4 balanced news dev; main arms4steps (one pattern) and target_half2. Training pools unchanged. Measure before formal; conservative extrapolation to1200s and6GiB RSS. No performance-driven budget adjustment. Resources hashed in lock and all execution identities; raw texts/input IDs/checkpoints private, not in public package. ## Pre-inference amendment (2026-10-04) Initial preprocessing-only smoke failed because selected news row112808 is422tokens; all answer targets were correctly2tokens. No model loaded, no training or generated output observed. Preserve exact selected rows; raise NEWS training limit256→512, keep SST2 limit256 and eval512. No truncation/filtering/row substitution. Initial failure retained in initial_preprocess_failed. Re-run resource smoke before formal; this is a disclosed resource amendment, not a performance-driven change.