# 093 SOURCE_MAP — accessed2026-10-06 Direct API use of unmodified official installed code; no upstream executable snippets copied. Five files freshly downloaded anonymously and compared byte-for-byte with installed packages, all equal: upstream_verification.json. Four Transformers4.49.0 files use commit a22a4378d97d06b7a1d9abad6e0086d30fdea199, Apache-2.0. PyTorch2.6.0 SGD uses installed torch.version.git_version2236df1770800ffea5697b11b0bb0d910b2e59e1, BSD-style license. Licenses retained verbatim as LICENSE-APACHE-2.0.txt/LICENSE-PYTORCH.txt. No fabricated main-line references. | Module/formula | Verified official permanent file/function | Local path and adaptation | |---|---|---| | Chat template | [apply_chat_template L1527](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1527), [special-token handling L1723](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1723) | common.py/build checks exact token prefix, includes real EOS, removes final template newline; independently constructed single-turn mask, not general conversation trainer | | Llama forward logits/loss | [LlamaForCausalLM L859](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/models/llama/modeling_llama.py#L859) | run.py calls official model directly; input[B,L], logits[B,L,49152] | | Shifted CE over supervised labels | [ForCausalLMLoss L32](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/loss/loss_utils.py#L32), [valid mean denominator L24](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/loss/loss_utils.py#L24) | common.py/collate independent labels -100 by lengths/prefix; run.py manual CE(logits[:,:-1],labels[:,1:]) checked each step;16supervised tokens/update | | Greedy generation | [generate L1879](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L1879), [argmax L3259](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L3259) | common.py/evaluate,mixture.py/news_evaluate direct API,greedy32,prompts unchanged | | SGD | [step L104](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/torch/optim/sgd.py#L104), [single-tensor update L353](https://github.com/pytorch/pytorch/blob/2236df1770800ffea5697b11b0bb0d910b2e59e1/torch/optim/sgd.py#L353) | run.py official SGDlr.001,no momentum/decay,clip1;eval-with-autograd intentional | | Frozen model/tokenizer | [model card](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/README.md), [template](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/tokenizer_config.json) | prepare.py/resources.lock.json; revision12fd25f77366fa6b3b4b768ec3050bf629380bac; no weights bundled | | Public data | [SST2 card](https://huggingface.co/datasets/stanfordnlp/sst2/blob/8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb/README.md), [AG News card](https://huggingface.co/datasets/fancyzhx/ag_news/blob/eb185aade064a813bc0b7f42de02595523103ca4/README.md) | mixture.py/select unchanged091; actual files/revisions/SHA256 in resources.lock.json. Cards licenseunknown, no raw data redistributed | | Order variation | Independent teaching code; [PyTorch2.6 reproducibility documentation](https://docs.pytorch.org/docs/2.6/notes/randomness.html) explains RNG scope, not our schedule | schedule.py:Python Random(seed), classwise shuffle each epoch then neg/pos pairs; common pretrained start, no random reinitialization | | delta_s=mean_i(after_si-before_i), SD across seeds, paired bootstrap | Independent teaching statistics; methodological context [Reimers & Gurevych2017](https://aclanthology.org/D17-1035/), [Dror et al.2018](https://aclanthology.org/P18-1128/). No paper code copied, no executable upstream commit applicable | analysis.py:d=a-b,vals.std(ddof=1),same paired resample indices for both models;10000replicates,seed9300,linear percentile. News stratifies gold class. No combined seed/sample CI or adaptive-selection correction | Self-written common.py,mixture.py,scoring.py,prepare.py,export_evidence.py copied in full from [091 commit7bd49bd7d6c1cd950d58506cfe7d4e962d161c85](https://github.com/distance539/ai-research-lab/tree/7bd49bd7d6c1cd950d58506cfe7d4e962d161c85/cycle06/091-data-mixture-retention); unmodified helpers. mixture.schedule is unused. run.py is adapted091 orchestration: replaces mixed-task arms with five order runs, removes mixture selection and adds frozen statistics. schedule.py,analysis.py,audit.py are new independent teaching code. All modules included; no dependency on an author-machine path. This is not the official SmolLM training recipe or a reproduction of LSTM paper numbers. Audit independently parses scores and checks row IDs, gold labels, prompt hashes, EOS/masks, budgets and paired arithmetic. Bootstrap audit replays the same frozen implementation and does not claim independent coverage verification. Checkpoint state_hash hashes sorted tensor names/shapes/bytes, avoiding filename-dependent torch archive hashes.