# 091 Source map Accessed 2026-10-04. Independent teaching orchestration/mask construction/audits; directly calls unmodified official installed model, tokenizer, loss and optimizer. No copied upstream executable code. Four Transformers files and SGD compared byte-for-byte against actual official pinned raw files; all equal, see upstream_verification.json. Model and data lock has20 exact download resources with SHA256/size. No fabricated main-line numbers. | Formula/module | Official pinned source, exact verified location | Local implementation and differences | |---|---|---| | Chat template / single-turn role boundaries | [apply_chat_template L1527](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1527), [tokenization without duplicated special tokens L1723](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1723); [actual model template](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/tokenizer_config.json) | common.py build() checks exact prefix identity, ends supervision through real assistant EOS, removes trailing template newline. Independent teaching mask, not official SmolLM training recipe or general multi-turn assistant mask. audit.py re-encodes all64 target/news rows and checks prefix/response lengths | | Vocabulary projection / model loss call | [LlamaForCausalLM forward, L859–863](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/models/llama/modeling_llama.py#L859) | run.py calls model(input_ids, attention_mask, labels) directly | | Shifted causal CE with ignored targets | [ForCausalLMLoss L32–48](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/loss/loss_utils.py#L32), [mean/sum denominator L24–29](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/loss/loss_utils.py#L24) | run.py passes unshifted labels. Manual CE(logits[:,:-1],labels[:,1:]) / valid count is equivalent to official pad-label-and-shift formulation. No num_items_in_batch passed; denominator is valid tokens in each batch. collate() sets -100 from true length and response boundary, never by all IDs==EOS | | Greedy output | [generate L1879](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L1879), [argmax branch L3259](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L3259) | common.py and mixture.py evaluate() preserve SST2 contracts, add fixed news A/B/C/D mapping, greedy32, direct API use | | SGD update | [official PyTorch SGD](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/torch/optim/sgd.py) | run.py uses torch.optim.SGD, no momentum/weight decay; gradient clip1 before step. [step L104](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/torch/optim/sgd.py#L104), [single-tensor update L353](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/torch/optim/sgd.py#L353) | | Data / labels / checkpoints | [fixed SST-2 card](https://huggingface.co/datasets/stanfordnlp/sst2/resolve/8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb/README.md), [model card](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/README.md) | prepare.py hashes fixed train/validation, no data bundled; model/tokenizer same12fd25f77366fa6b3b4b768ec3050bf629380bac. Labels0negative/1positive; source indices preserved | | F/C/J, supervised budgets and paired outputs | Custom experimental protocol, not official benchmark | scoring.py unchanged; audit.py independently reparses both tasks and reconstructs the mixed schedule, input/padded/answer budgets. News supervision2 tokens, target8; run.py accumulates weighted microbatch gradients using n_valid/16, clipping only after the full16target update. Budget B=sum of active shifted labels; pure teaching orchestration, no scaling-law fit | Transformers repository https://github.com/huggingface/transformers, commit a22a4378d97d06b7a1d9abad6e0086d30fdea199, v4.49.0, Apache-2.0 (license verbatim retained). PyTorch repository https://github.com/pytorch/pytorch, commit2236df1770800ffea5697b11b0bb0d910b2e59e1 obtained from installed torch.version.git_version, v2.6.0, [BSD-style license](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/LICENSE). Model card Apache-2.0; dataset card license unknown. Original datasets/checkpoints not redistributed; see THIRD_PARTY.md. Self-written prepare.py/scoring.py and extracted common.py helpers come from [089 fixed version](https://github.com/distance539/ai-research-lab/tree/6478572fdaf69d27e0453b851b157fb928d9df4b/cycle06/089-sft-template-label-masks). All needed modules included; no previous folder runtime dependency. run.py implements token-budget task rehearsal; mixture.py contains the full new task logic. Both are independent teaching code. audit.py independently verifies outputs and budgets. No executable upstream code copied. SOURCE_MAP table labels inherited formulas and APIs, not reproduction of official SmolLM training or Muennighoff's pretraining experiments. PyTorch LICENSE retained separately as LICENSE-PYTORCH.txt. ## New modules in091 - `mixture.py/select`: independent deterministic row-position selection; AG News [pinned data card](https://huggingface.co/datasets/fancyzhx/ag_news/blob/eb185aade064a813bc0b7f42de02595523103ca4/README.md), unknown data license. Class order World/Sports/Business/Sci-Tech mapped to A/B/C/D. Upstream dataset120000train/7600test; we use32train and32test-derived DEVELOPMENT,64reserved. No benchmark leaderboard claim. - `mixture.py/schedule`: token fraction rather than row fraction. Each news update8rows vs target2; cursor streams independent and repeated deterministically. All needed self-written modules are included. - `run.py`: official loss mean times n_valid/16 before each backward; four news microbatches yield one16target update. This is derived independent orchestration, not copying an official trainer. Taskwise updates and clipping mean a mixture fraction is not an exact global convex objective trajectory. - `audit.py`: independent parsing and raw-token reconstruction; recomputes budgets and source identity; no upstream executable snippets copied. No original-paper implementation claim for implicit inference. - Prior self-written base: [090 fixed code](https://github.com/distance539/ai-research-lab/tree/11c29a7deefedbf06a6f7e52ca6f7af125176535/cycle06/090-data-size-token-budget), common.py/scoring.py/prepare.py included in full and unmodified. SOURCE_MAP describes original APIs; no hidden dependency on the090 folder.