学习笔记
34
OPD_SNR
Model_Parallel
ZeRO
Connections Between On Policy Distillation And RL
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
Instruction tuning with loss over instructions
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
RETAINING BY DOING: THE ROLE OF ON-POLICY DATA IN MITIGATING FORGETTING
Proximal Gradient and Subgradients
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
More...