# 093 Preregistered protocol — 2026-10-06, before inference Question: does selected091 target_only behavior persist across actual training-order randomness, and how does this differ from sample uncertainty for a fixed checkpoint? Exploratory DEVELOPMENT only; 192 confirmation rows remain untouched. No selection of best seed or further configuration tuning. Frozen: SmolLM2-135M-Instruct/model+tokenizer12fd25f77366fa6b3b4b768ec3050bf629380bac, same20 resource hashes and091 splits (32SST training,64SST development,32news development). News is retention diagnostic with an8/32 floor, NOT an established general capability. Same134515008parameters, float32 CPU4threads/eager/eval-with-autograd, SGD.001/momentum0/decay0/clip1, batch2,64steps,1024answer-through-EOS targets, greedy32. Masks/labels/prompts/scoring unchanged. All arms start from pristine weights. Negative control: fixed93,fixed94 differ ONLY torch manual seed; original deterministic class-interleaved order, no dropout or sampled decode. Expect possibly identical results; measure state/loss/output equality, do not fabricate stochasticity. Main variation: shuffle93,shuffle94,shuffle95 independently shuffle16negative and16positive rows each epoch with Python random.Random(seed), then zip negative-positive, batch2, four epochs. Changes order/pairing, keeps each batch label-balanced and each row seen4times. This adds a stochastic schedule absent from091. Input and supervised tokens fixed; padding/attention and wall time may change. Python RNG seed plus torch seed recorded; no weight reinitialization, data resampling, dropout or decoding variance tested. Primary outcome: SST JSON joint accuracy after minus pristine baseline for each shuffled seed. Secondary: JSON explicit-label content proxy, plain joint and news joint, failure/format counts. Report all five arms, all predictions/steps/checkpoint hashes and failures. Across3 shuffled seeds report arithmetic mean and sample SD (ddof1) and range, not a population-confidence claim. Two fixed arms are controls, excluded from main3-seed summary. For each fixed model pair and each outcome compute per-row signed binary difference. Bootstrap10000 replicates with replacement of paired row indices, RNG9300, percentile2.5/97.5 using numpy linear quantiles. This is a descriptive conditional iid-row sensitivity interval on reused development rows, NOT correction for adaptive selection or proven sampling coverage. SST source grouping unavailable: no claims of independent movies. News rows resampled within each gold class (fixed balanced design). Reuse identical resample indices across arms; never flatten3x64 predictions as192 independent examples. No combined seed+sample interval or p-value. Same-run baseline is shared, no pseudo-independent baseline replicas. Stop: first5-arm smoke(2steps/arm,2SST+4news) must finish below1200s/6GiB with finite losses and nonzero updates. Estimate full run conservatively from smoke total multiplied by24; budget1200s. Formal uses full64steps/arm/96development. Failure preserves partial outputs and explicit status, no seed replacement. No rerun to find favorable outcomes. Generated long news echoes excluded from public package with explicit hash/count evidence; raw outputs remain local and independently audited first.