Qwen3-8B teacher-regularized RL method collection for the ToolUse dataset. Includes GRPO-TR, RLSD-TR, SDPO-TR, and SRPO-TR models.
-
SeongryongJung/Qwen3-8B-Tooluse-GRPO-TR
Text Generation • 8B • Updated • 49 -
SeongryongJung/Qwen3-8B-Tooluse-RLSD-TR
Text Generation • 8B • Updated • 221 -
SeongryongJung/Qwen3-8B-ToolUse-SDPO-TR
Reinforcement Learning • 8B • Updated • 17 -
SeongryongJung/Qwen3-8B-ToolUse-SRPO-TR
Reinforcement Learning • 8B • Updated • 23 • 1