# License and attribution Independent teaching code under Apache-2.0, LICENSE-APACHE-2.0.txt. prepare.py/scoring.py and common.py helper functions derive from our publicly published089 commit6478572fdaf69d27e0453b851b157fb928d9df4b; all required modules included. common.py is a literal extracted helper subset, run.py/audit.py newly written orchestration. No upstream executable source copied. Transformers4.49.0 a22a4378d97d06b7a1d9abad6e0086d30fdea199: Apache-2.0, license retained. PyTorch2.6.0 2236df1770800ffea5697b11b0bb0d910b2e59e1: BSD-style, LICENSE-PYTORCH.txt retained. Calls installed official APIs unchanged. Model/tokenizer SmolLM2-135M-Instruct12fd25f77366fa6b3b4b768ec3050bf629380bac Apache-2.0, downloaded only. SST2 stanfordnlp/sst2 revision8d51e7e4887a4caaa95b3fbebbf53c0490b58bbb card license unknown. Raw review text, parquet, prompts and reversible input IDs excluded. Download from original publisher; code license does not confer dataset rights. Public derived records contain IDs/gold labels, hashes, lengths, aggregate/individual loss and short model outputs. Weights/checkpoints/caches/venv excluded. Source links and hashes in SOURCE_MAP.md/resources.lock.json/upstream_verification.json. Muennighoff et al. Scaling Data-Constrained Language Models (2023 arXiv v1) is conceptual motivation only; no paper code copied and no scaling-law reproduction claim.