references
references
Chapter 1
- Paul F Christiano et al. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
- Nisan Stiennon et al. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021, 2020.
- Long Ouyang et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730–27744, 2022.
- Reiichiro Nakano et al. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
- Yuntao Bai et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
- Nathan Lambert et al. Tulu 3: Push-ing frontiers in open language model post-training. arXiv preprint arXiv:, 2024.
- Aaron Grattafiori et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024.