Instructions to use kagakouko/Spatial-Interactor-Qwen3-VL-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kagakouko/Spatial-Interactor-Qwen3-VL-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="kagakouko/Spatial-Interactor-Qwen3-VL-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("kagakouko/Spatial-Interactor-Qwen3-VL-4B") model = AutoModelForMultimodalLM.from_pretrained("kagakouko/Spatial-Interactor-Qwen3-VL-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kagakouko/Spatial-Interactor-Qwen3-VL-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kagakouko/Spatial-Interactor-Qwen3-VL-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kagakouko/Spatial-Interactor-Qwen3-VL-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B
- SGLang
How to use kagakouko/Spatial-Interactor-Qwen3-VL-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kagakouko/Spatial-Interactor-Qwen3-VL-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kagakouko/Spatial-Interactor-Qwen3-VL-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kagakouko/Spatial-Interactor-Qwen3-VL-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kagakouko/Spatial-Interactor-Qwen3-VL-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use kagakouko/Spatial-Interactor-Qwen3-VL-4B with Docker Model Runner:
docker model run hf.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B
Spatial-Interactor Qwen3-VL-4B
Learning Spatial Reasoning through Interaction with the Observable Physical World
This is the full-parameter BF16 Spatial-Interactor checkpoint based on Qwen/Qwen3-VL-4B-Instruct. It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories.
The privileged transition trace is used only during training. At inference, this checkpoint takes the same image/video and question inputs as its base model, with no extra trace, reward model, or teacher branch.
Overview
Presentation
Load
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-4B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto",
)
Use the base model's image/video input format. No privileged trace or additional teacher is needed for inference. Weights, tokenizer, processor, and chat template are included. See the training guide for SFT and OPD.
Citation
For citation, use the project BibTeX.
License
This checkpoint is released under Apache-2.0, following the base model license. Users must also comply with licenses and terms governing input datasets and media.
- Downloads last month
- 128