Learning to solve hard problems in RL for LLMs by never giving up
- LLMs 近 90 天出现 19 次
- 上一次:同一天稍早 · Why I'm still bearish on LLMs after Navier-Stokes
发生了什么
This is a blog post for my recent paper on RL post-training of LLMs: introducing the Matthew Effect and proposing to solve it with Never Give Up. It is presented interactively and less formally, more like how I give the talk. For a deeper, more technical dive, check out the paper on arxiv and code on github. What is your eval actually measuring? # Every good RL practitioner has no doubt seen an eval curve go up. Her…
摘要按规则整理自下方来源原文