Reinforcement Learning from Human Feedback
94 points
9 hours ago
| 3 comments
| rlhfbook.com
| HN
https://arxiv.org/abs/2504.12501
dang
3 hours ago
[-]
Related. Others?

RLHF Book - https://news.ycombinator.com/item?id=42902936 - Feb 2025 (37 comments)

reply
verdverm
7 hours ago
[-]
Last time I saw Nathan say something about the book, he's actively working on the next version and looking for feedback, check his socials
reply
leggerss
5 hours ago
[-]
You could say he's also learning from human feedback
reply
klelatti
8 hours ago
[-]
Web version with links, etc:

https://rlhfbook.com/

reply
dang
3 hours ago
[-]
Thanks! We've switched to that above from https://arxiv.org/abs/2504.12501, and put the latter in the toptext.
reply