Tune default SFT and LoRA hyperparameters
This commit is contained in:
@@ -89,11 +89,29 @@ The default experiment uses:
|
||||
- 1 epoch
|
||||
- LoRA rank 32 for LoRA runs
|
||||
- bf16 full fine-tuning for full runs
|
||||
- full SFT learning rate `1e-5`
|
||||
- LoRA learning rate `5e-5`
|
||||
- warmup ratio `0.1`
|
||||
- explicit cosine LR scheduler via `--lr_scheduler_type cosine`
|
||||
- `max_length=262144`
|
||||
- conservative per-device train batch size `1`
|
||||
- checkpoint save every 1000 steps
|
||||
- validation every 1000 steps
|
||||
- TensorBoard logging under `runs/`
|
||||
|
||||
Batch size is intentionally conservative because B300/g0049 has ~275GB per GPU but the default context length is 262144 tokens. Increase batch size from the shell only after checking memory:
|
||||
|
||||
```bash
|
||||
export PER_DEVICE_BATCH_SIZE=2
|
||||
export GRAD_ACCUM_STEPS=2
|
||||
export LORA_PER_DEVICE_BATCH_SIZE=2
|
||||
export FULL_PER_DEVICE_BATCH_SIZE=1
|
||||
export QWEN35_9B_LORA_R32_PER_DEVICE_BATCH_SIZE=2
|
||||
export QWEN36_27B_FULL_BF16_PER_DEVICE_BATCH_SIZE=1
|
||||
```
|
||||
|
||||
Run-specific variables have the highest precedence, then global `PER_DEVICE_BATCH_SIZE` / `GRAD_ACCUM_STEPS`, then train-type defaults, then the safe default of 1.
|
||||
|
||||
Override model IDs or paths with:
|
||||
|
||||
```bash
|
||||
|
||||
Reference in New Issue
Block a user