-
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Paper • 2508.05629 • Published • 191 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
LIMA: Less Is More for Alignment
Paper • 2305.11206 • Published • 27 -
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
Paper • 2408.16673 • Published
Collections
Discover the best community collections!
Collections including paper arxiv:2510.01171
-
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Paper • 2510.01171 • Published • 19 -
CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation
Viewer • Updated • 7.2k • 122 -
CHATS-Lab/Verbalized-Sampling-Open-Ended-QA
Viewer • Updated • 104k • 270 -
CHATS-Lab/Verbalized-Sampling-Random-Number-Generator
Viewer • Updated • 16.8k • 114
-
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Paper • 2508.05629 • Published • 191 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
LIMA: Less Is More for Alignment
Paper • 2305.11206 • Published • 27 -
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
Paper • 2408.16673 • Published
-
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Paper • 2510.01171 • Published • 19 -
CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation
Viewer • Updated • 7.2k • 122 -
CHATS-Lab/Verbalized-Sampling-Open-Ended-QA
Viewer • Updated • 104k • 270 -
CHATS-Lab/Verbalized-Sampling-Random-Number-Generator
Viewer • Updated • 16.8k • 114