# 088 Source map Accessed 2026-10-01. Local orchestration/scoring/audit are independent teaching implementations; no third-party source is copied into executable modules. Actual inference calls unmodified installed Transformers. Three installed source files below were read and compared byte-for-byte with raw official fixed-commit files; all equal. `resources.lock.json` stores actual URL/hash/size, `prepare.py` downloads and checks those exact files. Full upstream checkout is not bundled. | Module / equation | Official pinned source and verified line | Local implementation / relationship | |---|---|---| | Chat-template rendering and assistant prefix | [Transformers apply_chat_template L1527](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1527), [avoid duplicate special tokens L1723](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/tokenization_utils_base.py#L1723) | run.py message/main directly calls official tokenizer, asserts separately rendered token sequence equality; custom system message explicitly overrides template fallback | | Hidden states and vocabulary projection z=Wh | [LlamaModel L488](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/models/llama/modeling_llama.py#L488), [LlamaForCausalLM L752](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/models/llama/modeling_llama.py#L752), [lm_head L859](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/models/llama/modeling_llama.py#L859) | run.py uses AutoModelForCausalLM, eager CPU float32; no changes to model math; runtime shapes saved | | Greedy y=argmax z | [generate L1879](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L1879), [_sample L3116 and greedy branch L3259](https://github.com/huggingface/transformers/blob/a22a4378d97d06b7a1d9abad6e0086d30fdea199/src/transformers/generation/utils.py#L3259) | run.py explicitly supplies GenerationConfig, no sampling/beam search; first generated token checked against forward argmax | | Model architecture, weights, tokenizer and template | [SmolLM2 config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/config.json), [tokenizer_config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/tokenizer_config.json), [model card](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/12fd25f77366fa6b3b4b768ec3050bf629380bac/README.md) | prepare.py retrieves nine pinned files; no weights/tokenizer redistributed; runtime overrides stored BF16 to float32 | | Dataset schema, label mapping, split | [Stanford SST-2 fixed card](https://huggingface.co/datasets/stanfordnlp/sst2/blob/8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb/README.md), [fixed validation file](https://huggingface.co/datasets/stanfordnlp/sst2/resolve/8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb/data/validation-00000-of-00001.parquet) | run.py hash-sorts source idx, creates disjoint local development/confirmation IDs. Data downloaded only; no original sentences bundled | | F,C,J and F×C buckets | No claim these metrics are official SmolLM2/GLUE protocol | scoring.py independent explicit-label evaluator; audit.py independent reparse from outputs, checks original source labels and prompt hashes. J=mean(F*C), conditional denominator explicitly retained | Verified Transformers Git commit: `a22a4378d97d06b7a1d9abad6e0086d30fdea199` (installed 4.49.0). Repository https://github.com/huggingface/transformers; Apache-2.0. Original license retained verbatim in LICENSE-APACHE-2.0.txt. Source and installed identity evidence in upstream_verification.json. Verified model/tokenizer revision: `12fd25f77366fa6b3b4b768ec3050bf629380bac`; model card Apache-2.0. Verified data revision: `8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb`; dataset card explicitly says license unknown. Do not infer dataset licensing from Transformers or Stanford software licensing. See THIRD_PARTY.md. All line numbers above were checked against downloaded fixed source, not guessed from main. Model and data references point to exact config/files rather than pretending they contain Python functions. No copied/adapted upstream executable code; only direct API use. This is not reproduction of SmolLM2 pretraining/SFT paper benchmark numbers.