How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
# Run inference directly in the terminal:
llama cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
# Run inference directly in the terminal:
llama cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
Use Docker
docker model run hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:
Quick Links
Liquid AI
Try LFM β€’ Docs β€’ LEAP β€’ Discord

LFM2.5-1.2B-Instruct

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Find more details in the original model card: https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct

πŸƒ How to run LFM2.5

Example usage with llama.cpp:

llama-cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF --conversation \
    --temp 0.1 --top-k 50 --repeat-penalty 1.05

QAD Q4_0 GGUF

The Quantization-Aware Distillation (QAD) checkpoint is available as LFM2.5-1.2B-Instruct-QAD-Q4_0.gguf.

This is distinct from the post-training-quantized LFM2.5-1.2B-Instruct-Q4_0.gguf; both use the GGUF Q4_0 format.

Example usage with llama.cpp:

llama-cli -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF \
  --hf-file LFM2.5-1.2B-Instruct-QAD-Q4_0.gguf \
  -p "What is C. elegans?"

QAD source weights (safetensors)

The original FP32 QAD source checkpoint is available in qad/, with its model config, tokenizer, generation defaults, and the same chat template as the released QAD GGUF. It can be loaded in Transformers by passing subfolder="qad":

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "LiquidAI/LFM2.5-1.2B-Instruct-GGUF"
tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder="qad")
model = AutoModelForCausalLM.from_pretrained(
    repo_id, subfolder="qad", dtype="auto", device_map="auto"
)

inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "What is 2 + 2?"}],
    tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

These weights are intended for fine-tuning and experimentation. Published QAD results apply to the Q4_0 GGUF; direct FP32/BF16 inference and other quantization formats may behave differently. See the source checkpoint documentation for validation details and the license.

πŸ“¬ Contact

Downloads last month
350,989
GGUF
Model size
1B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for LiquidAI/LFM2.5-1.2B-Instruct-GGUF

Quantized
(114)
this model
Quantizations
1 model

Spaces using LiquidAI/LFM2.5-1.2B-Instruct-GGUF 7

Collection including LiquidAI/LFM2.5-1.2B-Instruct-GGUF

Article mentioning LiquidAI/LFM2.5-1.2B-Instruct-GGUF