-
End-to-End Goal-Driven Web Navigation
Paper • 1602.02261 • Published -
Learning Language Games through Interaction
Paper • 1606.02447 • Published -
Naturalizing a Programming Language via Interactive Learning
Paper • 1704.06956 • Published -
Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Paper • 1802.08802 • Published • 2
Collections
Discover the best community collections!
Collections including paper arxiv:2305.18290
-
Attention Is All You Need
Paper • 1706.03762 • Published • 142 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 45.9M • • 3.37k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 15.3M • 4.16k
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
-
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 150 -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Paper • 2501.12948 • Published • 461
-
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Paper • 2509.22576 • Published • 137 -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 24 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 150
-
High-Resolution Image Synthesis with Latent Diffusion Models
Paper • 2112.10752 • Published • 17 -
Adding Conditional Control to Text-to-Image Diffusion Models
Paper • 2302.05543 • Published • 60 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252 -
The Llama 3 Herd of Models
Paper • 2407.21783 • Published • 119
-
Attention Is All You Need
Paper • 1706.03762 • Published • 142 -
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 66 -
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Paper • 2101.03961 • Published • 13 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12
-
End-to-End Goal-Driven Web Navigation
Paper • 1602.02261 • Published -
Learning Language Games through Interaction
Paper • 1606.02447 • Published -
Naturalizing a Programming Language via Interactive Learning
Paper • 1704.06956 • Published -
Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Paper • 1802.08802 • Published • 2
-
Attention Is All You Need
Paper • 1706.03762 • Published • 142 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 45.9M • • 3.37k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 15.3M • 4.16k
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 150
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 39 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
-
High-Resolution Image Synthesis with Latent Diffusion Models
Paper • 2112.10752 • Published • 17 -
Adding Conditional Control to Text-to-Image Diffusion Models
Paper • 2302.05543 • Published • 60 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 150 -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Paper • 2501.12948 • Published • 461
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252 -
The Llama 3 Herd of Models
Paper • 2407.21783 • Published • 119
-
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Paper • 2509.22576 • Published • 137 -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 24 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
Attention Is All You Need
Paper • 1706.03762 • Published • 142 -
LoRA: Low-Rank Adaptation of Large Language Models
Paper • 2106.09685 • Published • 66 -
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Paper • 2101.03961 • Published • 13 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12