source avatarMinty

Share

Can AI learn what to optimize from examples instead of just copying them? Supervised fine-tuning (SFT) teaches a model by having it imitate expert answers. A new paper introduces PARED, an alignment method that tries to extract a reusable reward from those same demonstrations. PARED compares the expert answers with the model’s own responses across chosen features such as helpfulness, harmlessness, and topic. It uses those differences to learn an explicit reward, which then guides additional reinforcement learning after SFT. Using 4,000 GPT-5.1-generated demonstrations, responses from PARED-trained models were preferred to those from the original SFT checkpoint in 86.2% and 88.4% of non-tied comparisons under two reward-update approaches, as judged by Gemini 2.5 Flash. Demonstrations could end up being useful for more than imitation. They could also help define an objective that the model continues learning from after SFT.

No.0 picture
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.