Tune default SFT and LoRA hyperparameters

This commit is contained in:
Codex
2026-06-24 23:26:14 +08:00
parent 0b1a05cff5
commit e79c7ea9bf
3 changed files with 92 additions and 8 deletions
+18
View File
@@ -89,11 +89,29 @@ The default experiment uses:
- 1 epoch
- LoRA rank 32 for LoRA runs
- bf16 full fine-tuning for full runs
- full SFT learning rate `1e-5`
- LoRA learning rate `5e-5`
- warmup ratio `0.1`
- explicit cosine LR scheduler via `--lr_scheduler_type cosine`
- `max_length=262144`
- conservative per-device train batch size `1`
- checkpoint save every 1000 steps
- validation every 1000 steps
- TensorBoard logging under `runs/`
Batch size is intentionally conservative because B300/g0049 has ~275GB per GPU but the default context length is 262144 tokens. Increase batch size from the shell only after checking memory:
```bash
export PER_DEVICE_BATCH_SIZE=2
export GRAD_ACCUM_STEPS=2
export LORA_PER_DEVICE_BATCH_SIZE=2
export FULL_PER_DEVICE_BATCH_SIZE=1
export QWEN35_9B_LORA_R32_PER_DEVICE_BATCH_SIZE=2
export QWEN36_27B_FULL_BF16_PER_DEVICE_BATCH_SIZE=1
```
Run-specific variables have the highest precedence, then global `PER_DEVICE_BATCH_SIZE` / `GRAD_ACCUM_STEPS`, then train-type defaults, then the safe default of 1.
Override model IDs or paths with:
```bash