openeta: embodied task agent
AI & ML interests
LLM
Recent Activity
View all activity
Papers
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
An open-source audio understanding model supporting speech recognition, environmental sound analysis, music understanding, time-aware QA, and complex
-
MOSS Audio 8B Thinking
š¢28Generate answers to audio or video prompts
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text ⢠5B ⢠Updated ⢠135k ⢠80 -
OpenMOSS-Team/MOSS-Audio-4B-Thinking
Audio-Text-to-Text ⢠5B ⢠Updated ⢠31.9k ⢠36 -
OpenMOSS-Team/MOSS-Audio-8B-Instruct
Audio-Text-to-Text ⢠9B ⢠Updated ⢠49.4k ⢠48
-
OpenMOSS-Team/MOSS-VL-Instruct-0408
Video-Text-to-Text ⢠11B ⢠Updated ⢠24.6k ⢠103 -
OpenMOSS-Team/MOSS-VL-Base-0408
Video-Text-to-Text ⢠11B ⢠Updated ⢠63 ⢠62 -
OpenMOSS-Team/MOSS-VL-Instruct-0708
Video-Text-to-Text ⢠11B ⢠Updated ⢠632 ⢠28 -
OpenMOSS-Team/MOSS-VL-Base-0708
Video-Text-to-Text ⢠11B ⢠Updated ⢠94 ⢠18
-
AI Can Learn Scientific Taste
Paper ⢠2603.14473 ⢠Published ⢠433 -
OpenMOSS-Team/SciJudgeBench
Viewer ⢠Updated ⢠732k ⢠722 ⢠11 -
OpenMOSS-Team/SciJudge-4B-2605
Text Generation ⢠4B ⢠Updated ⢠185 ⢠6 -
OpenMOSS-Team/SciJudge-30B-2605
Text Generation ⢠31B ⢠Updated ⢠263 ⢠3
True Speech-to-Speech Langugage Model
First Omni-modal Future Forecasting Benchmark
https://github.com/OpenMOSS/FRoM-W1
Proactive Robot Manipulation in Omni-modal Context
Open source weights of Lorsa modules introduced in "Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition".
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Paper ⢠2502.14837 ⢠Published ⢠4 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_16
Text Generation ⢠6B ⢠Updated ⢠13 ⢠1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_32
Text Generation ⢠6B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_64
Text Generation ⢠7B ⢠Updated ⢠16 ⢠1
A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription.
-
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
Paper ⢠2601.01554 ⢠Published ⢠66 -
MOSS Transcribe Diarize
š¢103Transcribe audio/video with speaker diarization
-
OpenMOSS-Team/MOSS-Transcribe-preview-2B
Automatic Speech Recognition ⢠2B ⢠Updated ⢠1.58k ⢠48 -
OpenMOSS-Team/MOSS-Transcribe-Diarize
Audio-Text-to-Text ⢠0.9B ⢠Updated ⢠194k ⢠376
-
OpenMOSS-Team/moss-video-preview-base
Video-Text-to-Text ⢠11B ⢠Updated ⢠15 ⢠15 -
OpenMOSS-Team/moss-video-preview-sft
Video-Text-to-Text ⢠11B ⢠Updated ⢠77 ⢠17 -
OpenMOSS-Team/moss-video-preview-realtime-sft
Video-Text-to-Text ⢠11B ⢠Updated ⢠176 ⢠27 -
OpenMOSS-Team/Realtime-QA-100K
Viewer ⢠Updated ⢠100k ⢠365 ⢠7
-
OpenMOSS-Team/MOSS-TTS
Text-to-Speech ⢠8B ⢠Updated ⢠17.7k ⢠416 -
OpenMOSS-Team/MOSS-TTS-Local-Transformer
Text-to-Speech ⢠3B ⢠Updated ⢠7.91k ⢠30 -
OpenMOSS-Team/MOSS-TTS-Realtime
Text-to-Speech ⢠2B ⢠Updated ⢠54.6k ⢠101 -
OpenMOSS-Team/MOSS-TTS-Nano-100M
Text-to-Speech ⢠Updated ⢠65.6k ⢠236
Opensource Lorsas and Transcoders
-
OpenMOSS-Team/MOSS-TTSD-v1.0
Text-to-Speech ⢠8B ⢠Updated ⢠10.2k ⢠59 -
OpenMOSS-Team/MOSS-TTSD-v0.7
Text-to-Speech ⢠2B ⢠Updated ⢠74 ⢠18 -
OpenMOSS-Team/MOSS-TTSD-v0.5
Text-to-Speech ⢠2B ⢠Updated ⢠493 ⢠54 -
OpenMOSS-Team/MOSS-TTSD-v0
Text-to-Speech ⢠2B ⢠Updated ⢠7 ⢠28
Evaluating Agentic Backend Coding Capabilities in Real-World Development Scenarios
-
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Paper ⢠2601.11077 ⢠Published ⢠67 -
OpenMOSS-Team/ABC-Bench
Viewer ⢠Updated ⢠224 ⢠168 ⢠4 -
OpenMOSS-Team/Qwen3-32B-ABC
Text Generation ⢠33B ⢠Updated ⢠10 ⢠3 -
OpenMOSS-Team/Qwen3-8B-ABC
Text Generation ⢠8B ⢠Updated ⢠8 ⢠3
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
An Efficient Training Framework for Diffusion Language Models
-
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
Paper ⢠2503.10480 ⢠Published ⢠57 -
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Paper ⢠2506.23127 ⢠Published ⢠2 -
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
Paper ⢠2506.21230 ⢠Published ⢠1 -
OpenMOSS-Team/Embodied_R1-ScienceWorld
8B ⢠Updated ⢠7 ⢠1
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_8-refactor
Text Generation ⢠0.1B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_32-refactor
Text Generation ⢠0.1B ⢠Updated ⢠4 ⢠1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_16-refactor
Text Generation ⢠0.1B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/SmolLM-360M-MLA-d_kv_8-refactor
Text Generation ⢠0.3B ⢠Updated ⢠4 ⢠1
-
OpenMOSS-Team/moss-moon-003-sft-plugin
Text Generation ⢠Updated ⢠29 ⢠72 -
OpenMOSS-Team/moss-moon-003-sft
Text Generation ⢠Updated ⢠91 ⢠129 -
OpenMOSS-Team/moss-moon-003-base
Text Generation ⢠Updated ⢠25 ⢠132 -
OpenMOSS-Team/moss-moon-003-sft-int4
Text Generation ⢠Updated ⢠18 ⢠41
openeta: embodied task agent
A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription.
-
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
Paper ⢠2601.01554 ⢠Published ⢠66 -
MOSS Transcribe Diarize
š¢103Transcribe audio/video with speaker diarization
-
OpenMOSS-Team/MOSS-Transcribe-preview-2B
Automatic Speech Recognition ⢠2B ⢠Updated ⢠1.58k ⢠48 -
OpenMOSS-Team/MOSS-Transcribe-Diarize
Audio-Text-to-Text ⢠0.9B ⢠Updated ⢠194k ⢠376
An open-source audio understanding model supporting speech recognition, environmental sound analysis, music understanding, time-aware QA, and complex
-
MOSS Audio 8B Thinking
š¢28Generate answers to audio or video prompts
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text ⢠5B ⢠Updated ⢠135k ⢠80 -
OpenMOSS-Team/MOSS-Audio-4B-Thinking
Audio-Text-to-Text ⢠5B ⢠Updated ⢠31.9k ⢠36 -
OpenMOSS-Team/MOSS-Audio-8B-Instruct
Audio-Text-to-Text ⢠9B ⢠Updated ⢠49.4k ⢠48
-
OpenMOSS-Team/moss-video-preview-base
Video-Text-to-Text ⢠11B ⢠Updated ⢠15 ⢠15 -
OpenMOSS-Team/moss-video-preview-sft
Video-Text-to-Text ⢠11B ⢠Updated ⢠77 ⢠17 -
OpenMOSS-Team/moss-video-preview-realtime-sft
Video-Text-to-Text ⢠11B ⢠Updated ⢠176 ⢠27 -
OpenMOSS-Team/Realtime-QA-100K
Viewer ⢠Updated ⢠100k ⢠365 ⢠7
-
OpenMOSS-Team/MOSS-VL-Instruct-0408
Video-Text-to-Text ⢠11B ⢠Updated ⢠24.6k ⢠103 -
OpenMOSS-Team/MOSS-VL-Base-0408
Video-Text-to-Text ⢠11B ⢠Updated ⢠63 ⢠62 -
OpenMOSS-Team/MOSS-VL-Instruct-0708
Video-Text-to-Text ⢠11B ⢠Updated ⢠632 ⢠28 -
OpenMOSS-Team/MOSS-VL-Base-0708
Video-Text-to-Text ⢠11B ⢠Updated ⢠94 ⢠18
-
OpenMOSS-Team/MOSS-TTS
Text-to-Speech ⢠8B ⢠Updated ⢠17.7k ⢠416 -
OpenMOSS-Team/MOSS-TTS-Local-Transformer
Text-to-Speech ⢠3B ⢠Updated ⢠7.91k ⢠30 -
OpenMOSS-Team/MOSS-TTS-Realtime
Text-to-Speech ⢠2B ⢠Updated ⢠54.6k ⢠101 -
OpenMOSS-Team/MOSS-TTS-Nano-100M
Text-to-Speech ⢠Updated ⢠65.6k ⢠236
-
AI Can Learn Scientific Taste
Paper ⢠2603.14473 ⢠Published ⢠433 -
OpenMOSS-Team/SciJudgeBench
Viewer ⢠Updated ⢠732k ⢠722 ⢠11 -
OpenMOSS-Team/SciJudge-4B-2605
Text Generation ⢠4B ⢠Updated ⢠185 ⢠6 -
OpenMOSS-Team/SciJudge-30B-2605
Text Generation ⢠31B ⢠Updated ⢠263 ⢠3
Opensource Lorsas and Transcoders
-
OpenMOSS-Team/MOSS-TTSD-v1.0
Text-to-Speech ⢠8B ⢠Updated ⢠10.2k ⢠59 -
OpenMOSS-Team/MOSS-TTSD-v0.7
Text-to-Speech ⢠2B ⢠Updated ⢠74 ⢠18 -
OpenMOSS-Team/MOSS-TTSD-v0.5
Text-to-Speech ⢠2B ⢠Updated ⢠493 ⢠54 -
OpenMOSS-Team/MOSS-TTSD-v0
Text-to-Speech ⢠2B ⢠Updated ⢠7 ⢠28
True Speech-to-Speech Langugage Model
Evaluating Agentic Backend Coding Capabilities in Real-World Development Scenarios
-
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Paper ⢠2601.11077 ⢠Published ⢠67 -
OpenMOSS-Team/ABC-Bench
Viewer ⢠Updated ⢠224 ⢠168 ⢠4 -
OpenMOSS-Team/Qwen3-32B-ABC
Text Generation ⢠33B ⢠Updated ⢠10 ⢠3 -
OpenMOSS-Team/Qwen3-8B-ABC
Text Generation ⢠8B ⢠Updated ⢠8 ⢠3
First Omni-modal Future Forecasting Benchmark
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
https://github.com/OpenMOSS/FRoM-W1
An Efficient Training Framework for Diffusion Language Models
Proactive Robot Manipulation in Omni-modal Context
-
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
Paper ⢠2503.10480 ⢠Published ⢠57 -
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Paper ⢠2506.23127 ⢠Published ⢠2 -
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
Paper ⢠2506.21230 ⢠Published ⢠1 -
OpenMOSS-Team/Embodied_R1-ScienceWorld
8B ⢠Updated ⢠7 ⢠1
Open source weights of Lorsa modules introduced in "Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition".
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_8-refactor
Text Generation ⢠0.1B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_32-refactor
Text Generation ⢠0.1B ⢠Updated ⢠4 ⢠1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_16-refactor
Text Generation ⢠0.1B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/SmolLM-360M-MLA-d_kv_8-refactor
Text Generation ⢠0.3B ⢠Updated ⢠4 ⢠1
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Paper ⢠2502.14837 ⢠Published ⢠4 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_16
Text Generation ⢠6B ⢠Updated ⢠13 ⢠1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_32
Text Generation ⢠6B ⢠Updated ⢠7 ⢠1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_64
Text Generation ⢠7B ⢠Updated ⢠16 ⢠1
-
OpenMOSS-Team/moss-moon-003-sft-plugin
Text Generation ⢠Updated ⢠29 ⢠72 -
OpenMOSS-Team/moss-moon-003-sft
Text Generation ⢠Updated ⢠91 ⢠129 -
OpenMOSS-Team/moss-moon-003-base
Text Generation ⢠Updated ⢠25 ⢠132 -
OpenMOSS-Team/moss-moon-003-sft-int4
Text Generation ⢠Updated ⢠18 ⢠41