# 090 protocol frozen before inference — 2026-10-03 Question: at equal assistant-supervised tokens and updates, do more distinct SST-2 training rows change held-development joint JSON accuracy? This is a tiny single-seed exploration, not the official pretraining scaling-law reproduction. Primary contrast repeat8 vs unique32. Fixed SmolLM2-135M-Instruct/tokenizer revision inherited089, same pristine pretrained weights for every arm, official train/validation hashes. Training order: SHA256('89:'+idx) within each original label, select first16 per label, interleave negative then positive. Small nested set is first4 per label; check its IDs equal089's8 without outcome selection. No row screening except this declared label stratification, no length-based selection; abort on max256 overflow or duplicate normalized texts among32 or exact normalized overlap with all64development+128reservedconfirmation. Semantic/phrase overlap not ruled out. Assistant-answer-through-real-EOS supervision only; identical JSON target construction and prompts from089. Assert each response is8tokens in both classes, and each batch one label of each class. Batch2, no packing, dynamic right padding, no truncation, labels-100 on prompt and padding, realEOS retained. 8rows×16epochs versus32rows×4epochs:128exposures,64steps,1024supervised tokens per primary arm. Auxiliary repeat8_short:8×4epochs=32exposures/16steps/256supervised tokens. This auxiliary is intentionally unequal budget to examine added repetition; it cannot isolate a unique-data effect. Training input lengths/padding/attention shape cost differ and must be logged; equality means supervised-token budget, not equal FLOPs or elapsed time. All134515008 parameters, SGD lr.001/no momentum/no decay, clip1, eval-mode autograd/no dropout, CPUfloat32/eager/4threads/seed90. Fixed end checkpoint, no early stopping, no choosing based on development. Order is deterministic repeated cycles, not reshuffled. Untrained-in-this-run initial model is already instruction-tuned. Main score JSON joint F AND C on all64development from088; Fformat and Cexplicit-unique-label proxy reported separately plus plain diagnostics. Same greedy32. Report paired wins/ties/losses and all attempts; no CI, multi-seed, human semantics or benchmark claim.128confirmation stays unqueried. Measure teacher-forced answer NLL on same small8 and expanded32 (overlapping training groups, not independent generalization) before and after each arm; does decreasing NLL lead to greedy validJSON? Log per-target token loss by position at final evaluation. Per-step varying-batch loss is not a held-set learning curve. Smoke first2 development, all declared32training rows for identity but execute2updates per arm (same first4 rows; control of mechanics only), save state hashes and generated outputs. Budget900s/6GiB; estimate full run before launch. Full training must not launch if smoke projection exceeds budget. Save checkpoints/token traces locally, not public; distribution contains indices/hashes/outputs and exact download script. Cache identity includes all executed source/config/resources/model/tokenizer/template/prompt/decode/split settings.