-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 180 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 178 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 791
Collections
Discover the best community collections!
Collections including paper arxiv:2605.12500
-
GLM-5: from Vibe Coding to Agentic Engineering
Paper • 2602.15763 • Published • 220 -
zai-org/GLM-5.2
Text Generation • 753B • Updated • 928k • • 5.13k -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 263 -
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Paper • 2503.11576 • Published • 177
-
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
Paper • 2605.03849 • Published • 128 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 151
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 43 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 145 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 75 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 148
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Paper • 2607.28618 • Published • 304 -
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Paper • 2607.28227 • Published • 156 -
Metis: Memory Foundation Model
Paper • 2607.26760 • Published • 182 -
Kimi K3: Open Frontier Intelligence
Paper • 2607.24653 • Published • 522
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
sensenova/SenseNova-U1-8B-MoT
Any-to-Any • 18B • Updated • 2.77k • 290 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V3
Any-to-Any • 18B • Updated • 8.61k • 63 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V2
Any-to-Any • 18B • Updated • 130 • 29 -
sensenova/SenseNova-U1-8B-MoT-Infographic
Any-to-Any • 18B • Updated • 103 • 56
-
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Paper • 2412.03069 • Published • 34 -
Are Emergent Abilities of Large Language Models a Mirage?
Paper • 2304.15004 • Published • 8 -
Scaling Image Tokenizers with Grouped Spherical Quantization
Paper • 2412.02632 • Published • 10 -
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Paper • 2410.13848 • Published • 37
-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 180 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 178 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 791
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Paper • 2607.28618 • Published • 304 -
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Paper • 2607.28227 • Published • 156 -
Metis: Memory Foundation Model
Paper • 2607.26760 • Published • 182 -
Kimi K3: Open Frontier Intelligence
Paper • 2607.24653 • Published • 522
-
GLM-5: from Vibe Coding to Agentic Engineering
Paper • 2602.15763 • Published • 220 -
zai-org/GLM-5.2
Text Generation • 753B • Updated • 928k • • 5.13k -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 263 -
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Paper • 2503.11576 • Published • 177
-
Code as Agent Harness
Paper • 2605.18747 • Published • 224 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 61
-
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 198 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 73 -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
Paper • 2605.03849 • Published • 128 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 151
-
sensenova/SenseNova-U1-8B-MoT
Any-to-Any • 18B • Updated • 2.77k • 290 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V3
Any-to-Any • 18B • Updated • 8.61k • 63 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V2
Any-to-Any • 18B • Updated • 130 • 29 -
sensenova/SenseNova-U1-8B-MoT-Infographic
Any-to-Any • 18B • Updated • 103 • 56
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 43 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 145 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 75 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 148
-
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Paper • 2412.03069 • Published • 34 -
Are Emergent Abilities of Large Language Models a Mirage?
Paper • 2304.15004 • Published • 8 -
Scaling Image Tokenizers with Grouped Spherical Quantization
Paper • 2412.02632 • Published • 10 -
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Paper • 2410.13848 • Published • 37