CoT Distillation
Research
Token-weighted chain-of-thought distillation.
Submitted manuscript
AAAI 2027 · Submitted
From Structure to Preference: Token Weighting for Chain-of-Thought Distillation in Large Language Models
Co-author. Responsible for model training, hyperparameter tuning and vLLM TP=4 evaluation, including LoRA / DPO training and controlled comparisons.
Contributed to a three-stage framework of structure recovery, key-token-weighted supervision and preference optimization. Token importance is estimated by the drop in teacher likelihood of the reference answer after perturbing each rationale token, then used in weighted SFT and an auxiliary loss.
| Student / comparison | GSM8K | SVAMP |
|---|---|---|
| Qwen2.5-7B-Instruct | 94.01% | 94.00% |
| Gain over strongest baseline | +5.51 pp | +10.10 pp |
OpenReview ↗ · Results are reported in the submitted manuscript. Submission does not imply acceptance.