Tpo-Torch – Target Policy Optimization for Stable RLHF Alignment in PyTorch
1 points
5 hours ago
| 1 comment
| github.com
| HN
Griffith-7
5 hours ago
[-]
"Standard PPO in RLHF often suffers from training instability and high sensitivity to hyperparameters. I built Tpo-torch to provide a clean, modular PyTorch implementation of Target Policy Optimization (TPO) for more stable alignment. Would love to hear thoughts from anyone working on LLM post-training!"
reply