-
Adam's Law: Textual Frequency Law on Large Language Models
Paper • 2604.02176 • Published • 109 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 115 -
A Very Big Video Reasoning Suite
Paper • 2602.20159 • Published • 201 -
LightMem: Lightweight and Efficient Memory-Augmented Generation
Paper • 2510.18866 • Published • 117
Collections
Discover the best community collections!
Collections including paper arxiv:2603.16870
-
dLLM: Simple Diffusion Language Modeling
Paper • 2602.22661 • Published • 154 -
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
Paper • 2603.15594 • Published • 150 -
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Paper • 2603.13398 • Published • 155 -
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
Paper • 2603.06569 • Published • 120
-
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 555 -
mHC: Manifold-Constrained Hyper-Connections
Paper • 2512.24880 • Published • 336 -
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
Paper • 2601.00393 • Published • 133 -
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 195
-
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Paper • 2604.05015 • Published • 232 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 115 -
LPM 1.0: Video-based Character Performance Model
Paper • 2604.07823 • Published • 79 -
A Simple Baseline for Streaming Video Understanding
Paper • 2604.02317 • Published • 71
-
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
Paper • 2601.15369 • Published • 22 -
Stable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Model
Paper • 2601.15892 • Published • 56 -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
Paper • 2601.16208 • Published • 55 -
NAACL: Noise-AwAre Verbal Confidence Calibration for LLMs in RAG Systems
Paper • 2601.11004 • Published • 31
-
Wolf: Captioning Everything with a World Summarization Framework
Paper • 2407.18908 • Published • 33 -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Paper • 2407.19985 • Published • 37 -
TPDiff: Temporal Pyramid Video Diffusion Model
Paper • 2503.09566 • Published • 45 -
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
Paper • 2506.07464 • Published • 14
-
Adam's Law: Textual Frequency Law on Large Language Models
Paper • 2604.02176 • Published • 109 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 115 -
A Very Big Video Reasoning Suite
Paper • 2602.20159 • Published • 201 -
LightMem: Lightweight and Efficient Memory-Augmented Generation
Paper • 2510.18866 • Published • 117
-
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Paper • 2604.05015 • Published • 232 -
Demystifing Video Reasoning
Paper • 2603.16870 • Published • 115 -
LPM 1.0: Video-based Character Performance Model
Paper • 2604.07823 • Published • 79 -
A Simple Baseline for Streaming Video Understanding
Paper • 2604.02317 • Published • 71
-
dLLM: Simple Diffusion Language Modeling
Paper • 2602.22661 • Published • 154 -
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
Paper • 2603.15594 • Published • 150 -
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Paper • 2603.13398 • Published • 155 -
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
Paper • 2603.06569 • Published • 120
-
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
Paper • 2601.15369 • Published • 22 -
Stable-DiffCoder: Pushing the Frontier of Code Diffusion Large Language Model
Paper • 2601.15892 • Published • 56 -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
Paper • 2601.16208 • Published • 55 -
NAACL: Noise-AwAre Verbal Confidence Calibration for LLMs in RAG Systems
Paper • 2601.11004 • Published • 31
-
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
Paper • 2509.26507 • Published • 555 -
mHC: Manifold-Constrained Hyper-Connections
Paper • 2512.24880 • Published • 336 -
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
Paper • 2601.00393 • Published • 133 -
LTX-2: Efficient Joint Audio-Visual Foundation Model
Paper • 2601.03233 • Published • 195
-
Wolf: Captioning Everything with a World Summarization Framework
Paper • 2407.18908 • Published • 33 -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Paper • 2407.19985 • Published • 37 -
TPDiff: Temporal Pyramid Video Diffusion Model
Paper • 2503.09566 • Published • 45 -
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
Paper • 2506.07464 • Published • 14