Liquid AI
Try LFM β€’ Docs β€’ LEAP β€’ Discord

LFM2.5-350M-GGUF

LFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency.

Find more details in the original model card: https://huggingface.co/LiquidAI/LFM2.5-350M

πŸƒ How to run LFM2.5

Example usage with llama.cpp:

llama-cli -hf LiquidAI/LFM2.5-350M-GGUF --conversation \
    --temp 0.1 --top-k 50 --repeat-penalty 1.05

QAD Q4_0 GGUF

The Quantization-Aware Distillation (QAD) checkpoint is available as LFM2.5-350M-QAD-Q4_0.gguf.

This is distinct from the post-training-quantized LFM2.5-350M-Q4_0.gguf; both use the GGUF Q4_0 format.

Example usage with llama.cpp:

llama-cli -hf LiquidAI/LFM2.5-350M-GGUF \
  --hf-file LFM2.5-350M-QAD-Q4_0.gguf \
  -p "What is C. elegans?"

QAD source weights (safetensors)

The original FP32 QAD source checkpoint is available in qad/, with its model config, tokenizer, generation defaults, and the same chat template as the released QAD GGUF. It can be loaded in Transformers by passing subfolder="qad":

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "LiquidAI/LFM2.5-350M-GGUF"
tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder="qad")
model = AutoModelForCausalLM.from_pretrained(
    repo_id, subfolder="qad", dtype="auto", device_map="auto"
)

inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "What is 2 + 2?"}],
    tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

These weights are intended for fine-tuning and experimentation. Published QAD results apply to the Q4_0 GGUF; direct FP32/BF16 inference and other quantization formats may behave differently. See the source checkpoint documentation for validation details and the license.

Downloads last month
72,782
GGUF
Model size
0.4B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for LiquidAI/LFM2.5-350M-GGUF

Quantized
(78)
this model

Spaces using LiquidAI/LFM2.5-350M-GGUF 3

Collection including LiquidAI/LFM2.5-350M-GGUF

Article mentioning LiquidAI/LFM2.5-350M-GGUF