🎬 InternVideo3-8B-Instruct

InternVideo3 is a multimodal LLM for long-horizon video understanding and agentic reasoning, featuring M²LA (Multimodal Multi-head Latent Attention) for efficient long-context processing. Upload a video (or image) and ask a question, or try text-only conversation.

📄 Paper · 🤗 Model

Examples