CoT Distillation

Research

Token-weighted chain-of-thought distillation.

Submitted manuscript

AAAI 2027 · Submitted

From Structure to Preference: Token Weighting for Chain-of-Thought Distillation in Large Language Models

Shankui Han, Zhaoyu Li, Weiwen Yuan, Yuyuan Yang, Haoran Feng, Jinke Ren

Co-author. Responsible for model training, hyperparameter tuning and vLLM TP=4 evaluation, including LoRA / DPO training and controlled comparisons.

Contributed to a three-stage framework of structure recovery, key-token-weighted supervision and preference optimization. Token importance is estimated by the drop in teacher likelihood of the reference answer after perturbing each rationale token, then used in weighted SFT and an auxiliary loss.

Student / comparisonGSM8KSVAMP
Qwen2.5-7B-Instruct94.01%94.00%
Gain over strongest baseline+5.51 pp+10.10 pp

OpenReview ↗ · Results are reported in the submitted manuscript. Submission does not imply acceptance.