Spatial-Interactor-Qwen3-VL-4B-GGUF

Spatial-Interactor-Qwen3-VL-4B is a full-parameter BF16 checkpoint from ZJU-OmniAI's Spatial-Interactor project, built on Qwen3-VL-4B-Instruct to learn spatial reasoning through interaction with the observable physical world, as detailed in the accompanying paper. It's trained in two stages: supervised fine-tuning on the LSI-108K dataset's L1-L2 split (local world-state and ego-motion transition modeling) alongside a public spatial QA mixture, followed by On-Policy Distillation (OPD) that integrates verifiable answer rewards with same-prefix privileged self-distillation to guide intermediate reasoning over long-horizon video trajectories — with the visual encoder kept frozen throughout while only the language model and multimodal projector are updated. Critically, the privileged transition trace used during training is discarded at inference time, so the deployed model takes the same image/video-plus-question inputs as its base model with no extra trace, reward model, or teacher branch; for video evaluation, the paper's main results use 32 ordered frames with chronological order preserved. It's one of four checkpoints in the Spatial-Interactor collection, loadable via the standard Transformers interface in place of the base model identifier, and released under Apache-2.0 following the base model's license.

Model Files

File Name Quant Type File Size File Link
Spatial-Interactor-Qwen3-VL-4B.BF16.gguf BF16 8.83 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q3_K_L.gguf Q3_K_L 2.41 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q3_K_M.gguf Q3_K_M 2.24 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q4_K_M.gguf Q4_K_M 2.72 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q4_K_S.gguf Q4_K_S 2.6 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q5_K_M.gguf Q5_K_M 3.16 GB Download
Spatial-Interactor-Qwen3-VL-4B.Q5_K_S.gguf Q5_K_S 3.09 GB Download
Spatial-Interactor-Qwen3-VL-4B.mmproj-bf16.gguf mmproj-bf16 839 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
505
GGUF
Model size
4B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Spatial-Interactor-Qwen3-VL-4B-GGUF

Quantized
(2)
this model

Collection including prithivMLmods/Spatial-Interactor-Qwen3-VL-4B-GGUF