Skip to content

Repository files navigation

English | 简体中文

ReasonTune-Lab

Practical post-training for small math reasoning models: teacher distillation, full-parameter SFT, GRPO/RLVR, reproducible evaluation, and sample-level flip analysis.

ReasonTune-Lab studies whether solutions distilled from Qwen2.5-Math-7B-Instruct provide a useful supervised warm-start for full-parameter GRPO on Qwen2.5-0.5B-Instruct. The repository implements the complete experiment loop from GSM8K data construction to EvalScope evaluation on GSM8K and MATH-500.

Highlights

  • Built a reproducible text-only math post-training pipeline with YAML experiment configs and thin command-line entry points.
  • Generated concise teacher solutions through a local OpenAI-compatible vLLM service and filtered them by answer correctness, format, and length.
  • Ran full-parameter SFT and accuracy-reward GRPO with ms-swift.
  • Preserved per-sample EvalScope reviews for aggregate metrics and pairwise flip analysis.
  • Recorded exact commands, config snapshots, git revisions, logs, and SwanLab tracking metadata for server-side runs.
Component Choice
Student Qwen2.5-0.5B-Instruct
Teacher Qwen2.5-Math-7B-Instruct
Post-training Full-parameter SFT and GRPO/RLVR
Training stack PyTorch, ms-swift, TRL, DeepSpeed
Serving and evaluation vLLM, EvalScope
Tracking SwanLab
Benchmarks GSM8K test and MATH-500

Key Results

Accuracy from deterministic zero-shot evaluation:

Model Checkpoint GSM8K MATH-500
Base original 49.66% 30.80%
SFT final, step 192 49.20% 33.20%
GRPO from base final, step 587 52.77% 33.80%
GRPO from SFT-90 final, step 587 53.30% 33.80%
GRPO from SFT-192 best reward, step 340 53.07% 36.00%

The complete nine-checkpoint table, relative gains, response lengths, and flip counts are documented in the experiment analysis. Machine-readable tables and generated charts are available under outputs/analysis/eval_comparison/.

Findings

  1. Direct GRPO produced the clearest improvement. The final direct-GRPO checkpoint improved the base model by 3.11 percentage points on GSM8K and 3.00 points on MATH-500.
  2. Distilled SFT transferred differently across benchmarks. SFT improved MATH-500 by 2.0-2.4 points but slightly reduced GSM8K accuracy.
  3. SFT warm-start added small, task-dependent gains over direct GRPO. SFT-90 favored GSM8K, while SFT-192 produced the stronger MATH-500 result.
  4. Validation reward was not a universal checkpoint selector. Best-reward and final checkpoints ranked differently on GSM8K and MATH-500, so external benchmark evaluation remained necessary.
  5. The current experiment measures effectiveness, not training stability. Each route has one run; multi-seed experiments are required before claiming lower variance or more stable optimization.

Pipeline

GSM8K train
├── 5,000 examples -> teacher distillation -> filtering -> SFT train/validation
└── remaining examples -> GRPO train/validation

Qwen2.5-0.5B-Instruct
├── base evaluation
├── full-parameter SFT
├── direct full-parameter GRPO
└── SFT warm-start -> full-parameter GRPO

selected checkpoints
-> vLLM OpenAI-compatible serving
-> EvalScope GSM8K + MATH-500 evaluation
-> aggregate metrics + pairwise flip analysis

Shared answer extraction and normalization live in src/rtlab/math/answers.py, so data filtering, evaluation, and analysis do not maintain separate answer parsers.

Experiment Design

The primary comparison is:

base -> GRPO
vs
base -> distilled SFT -> GRPO

Two SFT checkpoints were retained:

  • sft_full_90: lowest validation-loss checkpoint.
  • sft_full_192: final checkpoint from a resumed run targeting two total SFT epochs.

The committed SFT YAML remains the one-epoch first-pass template. The two-epoch checkpoint was produced through the explicit resume workflow documented in docs/training.md.

For each GRPO initialization, the analysis retains:

  • the checkpoint with the highest validation reward;
  • the final checkpoint after one GRPO epoch.

All mainline SFT and GRPO runs use full-parameter training. LoRA, multiple teachers, larger students, reward-model training, and mixed-dataset scaling are deliberately outside this first-stage comparison.

Repository Structure

src/rtlab/
├── data/       JSONL IO, GSM8K splitting, distillation, filtering
├── math/       answer extraction, normalization, and comparison
├── model/      model download and vLLM command construction
├── train/      reproducible ms-swift training commands
├── eval/       EvalScope config loading and run metadata
└── analysis/   aggregate metrics, relative gains, and flip analysis

scripts/        thin CLI entry points grouped by responsibility
configs/        model, data, training, evaluation, and analysis YAML
docs/           pipeline, training, evaluation, and experiment reports
tests/          unit tests with small synthetic fixtures

Datasets, model weights, raw training outputs, logs, and raw EvalScope predictions remain local or server-side. The repository commits source code, configs, tests, documentation, and compact derived analysis tables.

Quick Start

Requirements:

  • Python 3.12
  • uv
  • CUDA-capable Linux environment for training and vLLM serving

Install the environment:

uv sync

Inspect a training command without launching a job:

PYTHONPATH=src uv run python scripts/train/run_swift.py \
  --config configs/train/sft_full.yaml \
  --dry-run

Run the analysis after EvalScope outputs have been copied under outputs/eval/<experiment_id>/:

PYTHONPATH=src uv run python scripts/analysis/summarize_eval_results.py \
  --config configs/analysis/eval_comparison.yaml \
  --eval-root outputs/eval \
  --output-dir outputs/analysis/eval_comparison \
  --write-plots

Reproducing the Experiment

1. Download and serve models

PYTHONPATH=src uv run python scripts/model/download_model.py \
  --config configs/models/qwen25_math_7b_instruct.yaml

PYTHONPATH=src uv run python scripts/model/download_model.py \
  --config configs/models/qwen25_05b_instruct.yaml

PYTHONPATH=src uv run python scripts/model/serve_vllm.py \
  --config configs/models/qwen25_math_7b_instruct.yaml

The serving command is long-running. Run the health check and distillation in a second shell.

2. Build distilled SFT and GRPO data

PYTHONPATH=src uv run python scripts/data/split_gsm8k.py \
  --config configs/data/teacher_distill.yaml

PYTHONPATH=src uv run python scripts/data/distill_with_teacher.py \
  --config configs/data/teacher_distill.yaml

PYTHONPATH=src uv run python scripts/data/filter_distilled_data.py \
  --config configs/data/teacher_distill.yaml

PYTHONPATH=src uv run python scripts/data/build_training_data.py \
  --config configs/data/teacher_distill.yaml

3. Run full-parameter SFT and GRPO

PYTHONPATH=src uv run python scripts/train/run_swift.py \
  --config configs/train/sft_full.yaml

PYTHONPATH=src uv run python scripts/train/run_swift.py \
  --config configs/train/grpo_base_full.yaml

PYTHONPATH=src uv run python scripts/train/run_swift.py \
  --config configs/train/grpo_sft_full.yaml

Use --model-path to select the exact SFT checkpoint for a warm-start GRPO run.

4. Evaluate selected checkpoints

Serve one checkpoint through scripts/model/serve_vllm.py, align the model name and output directory in an evaluation YAML, then run:

PYTHONPATH=src uv run python scripts/eval/run_evalscope.py \
  --config configs/eval/base_eval.yaml \
  --limit 10

PYTHONPATH=src uv run python scripts/eval/run_evalscope.py \
  --config configs/eval/base_eval.yaml

The smoke run checks the endpoint and output structure before the full evaluation. Preserve the complete reports/, reviews/, and predictions/ directories for later comparison.

Documentation

Limitations and Next Steps

  • Repeat the main routes across multiple random seeds and report variance.
  • Analyze training reward, KL, entropy, and completion length curves alongside external benchmark accuracy.
  • Inspect beneficial and harmful flips by MATH-500 difficulty level.
  • Improve distilled-data quality or vary SFT data size before adding broader teacher, student-size, or dataset comparisons.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages