How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf antirez/Laguna-S-2.1-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf antirez/Laguna-S-2.1-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf antirez/Laguna-S-2.1-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf antirez/Laguna-S-2.1-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf antirez/Laguna-S-2.1-GGUF:Q8_0
# Run inference directly in the terminal:
./llama-cli -hf antirez/Laguna-S-2.1-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf antirez/Laguna-S-2.1-GGUF:Q8_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf antirez/Laguna-S-2.1-GGUF:Q8_0
Use Docker
docker model run hf.co/antirez/Laguna-S-2.1-GGUF:Q8_0
Quick Links

Laguna S 2.1 GGUF

This repository contains a reduced-memory Laguna S 2.1 quantization for DwarfStar.

Mixed Q2_K/Q3_K variant

laguna-s-2.1-RoutedQ2_K-Last27Q3_K.gguf keeps every non-routed tensor byte-identical to Poolside's laguna-s-2.1-Q4_K_M.gguf at revision 706fa69799926b6afde1af9e24ca2a4923f110a1. Only routed expert tensors were requantized using the source importance matrix:

  • routed layers 1 through 20: Q2_K gate, up, and down
  • routed layers 21 through 47: Q3_K gate, up, and down
  • all other tensors: unchanged from the official Q4_K_M GGUF

The file is 48,260,803,968 bytes (44.946 GiB), intended for full-residency inference on 64 GiB systems. Runtime memory also depends on context size and KV-cache allocation.

SHA-256:

61fc66596597985cb9408a8530de6322d9e0d5b1d2ad4ed6503938018e0ce903

DwarfStar

Q3_K routed Laguna inference is supported starting with DwarfStar commit 938227a2.

./download_model.sh laguna-q2-q3
./ds4 -m gguf/laguna-s-2.1-RoutedQ2_K-Last27Q3_K.gguf -p "Hello"

On an Apple M5 Max, the tested model reached approximately 514 tokens/second for a 4096-token prefill and 63 tokens/second steady-state generation.

Against 100 official continuation vectors, the mixed model obtained average NLL 0.2583, 87/100 first-token matches, and average matching-prefix length 9.50 tokens. The corresponding full Q4_K_M measurements were 0.2352, 92/100, and 10.86.

Downloads last month
1,601
GGUF
Model size
118B params
Architecture
laguna
Hardware compatibility
Log In to add your hardware

2-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for antirez/Laguna-S-2.1-GGUF

Quantized
(94)
this model