title

Reinforcement Learning from Human Feedback

 

LLM alignment and post-training

Nathan Lambert
Foreword by Thomas Wolf

MANNING
SHELTER ISLAND