Practical post-training for small math reasoning models: teacher distillation, full-parameter SFT, GRPO/RLVR, reproducible evaluation, and sample-level flip analysis.
ReasonTune-Lab studies whether solutions distilled from
Qwen2.5-Math-7B-Instruct provide a useful supervised warm-start for
full-parameter GRPO on Qwen2.5-0.5B-Instruct. The repository implements the
complete experiment loop from GSM8K data construction to EvalScope evaluation
on GSM8K and MATH-500.
- Built a reproducible text-only math post-training pipeline with YAML experiment configs and thin command-line entry points.
- Generated concise teacher solutions through a local OpenAI-compatible vLLM service and filtered them by answer correctness, format, and length.
- Ran full-parameter SFT and accuracy-reward GRPO with
ms-swift. - Preserved per-sample EvalScope reviews for aggregate metrics and pairwise flip analysis.
- Recorded exact commands, config snapshots, git revisions, logs, and SwanLab tracking metadata for server-side runs.
| Component | Choice |
|---|---|
| Student | Qwen2.5-0.5B-Instruct |
| Teacher | Qwen2.5-Math-7B-Instruct |
| Post-training | Full-parameter SFT and GRPO/RLVR |
| Training stack | PyTorch, ms-swift, TRL, DeepSpeed |
| Serving and evaluation | vLLM, EvalScope |
| Tracking | SwanLab |
| Benchmarks | GSM8K test and MATH-500 |
Accuracy from deterministic zero-shot evaluation:
| Model | Checkpoint | GSM8K | MATH-500 |
|---|---|---|---|
| Base | original | 49.66% | 30.80% |
| SFT | final, step 192 | 49.20% | 33.20% |
| GRPO from base | final, step 587 | 52.77% | 33.80% |
| GRPO from SFT-90 | final, step 587 | 53.30% | 33.80% |
| GRPO from SFT-192 | best reward, step 340 | 53.07% | 36.00% |
The complete nine-checkpoint table, relative gains, response lengths, and flip
counts are documented in
the experiment analysis. Machine-readable
tables and generated charts are available under
outputs/analysis/eval_comparison/.
- Direct GRPO produced the clearest improvement. The final direct-GRPO checkpoint improved the base model by 3.11 percentage points on GSM8K and 3.00 points on MATH-500.
- Distilled SFT transferred differently across benchmarks. SFT improved MATH-500 by 2.0-2.4 points but slightly reduced GSM8K accuracy.
- SFT warm-start added small, task-dependent gains over direct GRPO. SFT-90 favored GSM8K, while SFT-192 produced the stronger MATH-500 result.
- Validation reward was not a universal checkpoint selector. Best-reward and final checkpoints ranked differently on GSM8K and MATH-500, so external benchmark evaluation remained necessary.
- The current experiment measures effectiveness, not training stability. Each route has one run; multi-seed experiments are required before claiming lower variance or more stable optimization.
GSM8K train
├── 5,000 examples -> teacher distillation -> filtering -> SFT train/validation
└── remaining examples -> GRPO train/validation
Qwen2.5-0.5B-Instruct
├── base evaluation
├── full-parameter SFT
├── direct full-parameter GRPO
└── SFT warm-start -> full-parameter GRPO
selected checkpoints
-> vLLM OpenAI-compatible serving
-> EvalScope GSM8K + MATH-500 evaluation
-> aggregate metrics + pairwise flip analysis
Shared answer extraction and normalization live in
src/rtlab/math/answers.py, so data filtering,
evaluation, and analysis do not maintain separate answer parsers.
The primary comparison is:
base -> GRPO
vs
base -> distilled SFT -> GRPO
Two SFT checkpoints were retained:
sft_full_90: lowest validation-loss checkpoint.sft_full_192: final checkpoint from a resumed run targeting two total SFT epochs.
The committed SFT YAML remains the one-epoch first-pass template. The
two-epoch checkpoint was produced through the explicit resume workflow
documented in docs/training.md.
For each GRPO initialization, the analysis retains:
- the checkpoint with the highest validation reward;
- the final checkpoint after one GRPO epoch.
All mainline SFT and GRPO runs use full-parameter training. LoRA, multiple teachers, larger students, reward-model training, and mixed-dataset scaling are deliberately outside this first-stage comparison.
src/rtlab/
├── data/ JSONL IO, GSM8K splitting, distillation, filtering
├── math/ answer extraction, normalization, and comparison
├── model/ model download and vLLM command construction
├── train/ reproducible ms-swift training commands
├── eval/ EvalScope config loading and run metadata
└── analysis/ aggregate metrics, relative gains, and flip analysis
scripts/ thin CLI entry points grouped by responsibility
configs/ model, data, training, evaluation, and analysis YAML
docs/ pipeline, training, evaluation, and experiment reports
tests/ unit tests with small synthetic fixtures
Datasets, model weights, raw training outputs, logs, and raw EvalScope predictions remain local or server-side. The repository commits source code, configs, tests, documentation, and compact derived analysis tables.
Requirements:
- Python 3.12
uv- CUDA-capable Linux environment for training and vLLM serving
Install the environment:
uv syncInspect a training command without launching a job:
PYTHONPATH=src uv run python scripts/train/run_swift.py \
--config configs/train/sft_full.yaml \
--dry-runRun the analysis after EvalScope outputs have been copied under
outputs/eval/<experiment_id>/:
PYTHONPATH=src uv run python scripts/analysis/summarize_eval_results.py \
--config configs/analysis/eval_comparison.yaml \
--eval-root outputs/eval \
--output-dir outputs/analysis/eval_comparison \
--write-plotsPYTHONPATH=src uv run python scripts/model/download_model.py \
--config configs/models/qwen25_math_7b_instruct.yaml
PYTHONPATH=src uv run python scripts/model/download_model.py \
--config configs/models/qwen25_05b_instruct.yaml
PYTHONPATH=src uv run python scripts/model/serve_vllm.py \
--config configs/models/qwen25_math_7b_instruct.yamlThe serving command is long-running. Run the health check and distillation in a second shell.
PYTHONPATH=src uv run python scripts/data/split_gsm8k.py \
--config configs/data/teacher_distill.yaml
PYTHONPATH=src uv run python scripts/data/distill_with_teacher.py \
--config configs/data/teacher_distill.yaml
PYTHONPATH=src uv run python scripts/data/filter_distilled_data.py \
--config configs/data/teacher_distill.yaml
PYTHONPATH=src uv run python scripts/data/build_training_data.py \
--config configs/data/teacher_distill.yamlPYTHONPATH=src uv run python scripts/train/run_swift.py \
--config configs/train/sft_full.yaml
PYTHONPATH=src uv run python scripts/train/run_swift.py \
--config configs/train/grpo_base_full.yaml
PYTHONPATH=src uv run python scripts/train/run_swift.py \
--config configs/train/grpo_sft_full.yamlUse --model-path to select the exact SFT checkpoint for a warm-start GRPO run.
Serve one checkpoint through scripts/model/serve_vllm.py, align the model
name and output directory in an evaluation YAML, then run:
PYTHONPATH=src uv run python scripts/eval/run_evalscope.py \
--config configs/eval/base_eval.yaml \
--limit 10
PYTHONPATH=src uv run python scripts/eval/run_evalscope.py \
--config configs/eval/base_eval.yamlThe smoke run checks the endpoint and output structure before the full
evaluation. Preserve the complete reports/, reviews/, and predictions/
directories for later comparison.
- Experiment analysis
- Experiment plan
- Experiment registry
- Data pipeline
- Training
- Evaluation
- Model download and serving
- Archived pre-restart project notes
- Repeat the main routes across multiple random seeds and report variance.
- Analyze training reward, KL, entropy, and completion length curves alongside external benchmark accuracy.
- Inspect beneficial and harmful flips by MATH-500 difficulty level.
- Improve distilled-data quality or vary SFT data size before adding broader teacher, student-size, or dataset comparisons.