Animated model card and benchmark chart
Browse files- README.md +4 -24
- assets/benchmarks.svg +336 -0
- assets/model.svg +37 -0
README.md
CHANGED
|
@@ -29,35 +29,15 @@ Needle does three jobs, all of them on the device:
|
|
| 29 |
|
| 30 |
## Model
|
| 31 |
|
| 32 |
-
|
| 33 |
-
| --- | --- |
|
| 34 |
-
| Inputs | Text prompts, plus tool definitions or an extraction schema |
|
| 35 |
-
| Outputs | Tool calls or structured extraction |
|
| 36 |
-
| Model | 29-121M Laddered Simple Attention Networks, CQ2 quantisation |
|
| 37 |
-
| Training | 360B tokens of proprietary structured dataset |
|
| 38 |
-
| Speed | 400-4k tokens/s decode and 1-10k tokens/s prefill on a Raspberry Pi 5 |
|
| 39 |
-
| Customisation | Fine-tuning lifts every subnetwork 18 to 36 points on DroidCall; from 4 layers (29M) up the tuned subnetwork passes DeepSeek V4 Flash |
|
| 40 |
|
| 41 |
Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the [release page](https://cactuscompute.com/needle). The repo holds the 20-layer `needle3.cact`, the `needle3.safetensors` checkpoint to fine-tune, and an engine per platform.
|
| 42 |
|
| 43 |
## Benchmarks
|
| 44 |
|
| 45 |
-
Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 49 |
-
| DeepSeek V4 Flash (cloud) | - | 88.4 | 60.5 | 77.2 | 80.0 | 69.4 | 66.7 |
|
| 50 |
-
| **Needle3-20L-121M** | 121M | 86.0 | 47.0 | 50.2 | 40.7 | 30.2 | 24.7 |
|
| 51 |
-
| LFM2.5 1.2B | 1.2B | 82.4 | 35.5 | 62.0 | 48.0 | 43.0 | 38.0 |
|
| 52 |
-
| **Needle3-16L-98M** | 98M | 80.7 | 40.0 | 41.3 | 28.5 | 23.5 | 19.2 |
|
| 53 |
-
| Qwen3.5 0.8B | 800M | 76.0 | 28.0 | 56.8 | 49.0 | 35.0 | 34.0 |
|
| 54 |
-
| LFM2.5 350M | 350M | 72.8 | 32.5 | 59.1 | 20.0 | 34.0 | 29.0 |
|
| 55 |
-
| LFM2.5 230M | 230M | 69.3 | 11.5 | 46.3 | 53.0 | 27.0 | 22.0 |
|
| 56 |
-
| FunctionGemma 270M | 270M | 65.1 | 16.5 | 46.6 | 27.0 | 29.0 | 14.0 |
|
| 57 |
-
| Needle 2 | 45M | 63.5 | 17.0 | - | - | - | - |
|
| 58 |
-
| Apple FM | 3.0B | 57.6 | - | - | - | - | - |
|
| 59 |
-
| **Needle3-8L-52M** | 52M | 36.8 | 36.5 | 28.2 | 15.3 | 16.6 | 10.1 |
|
| 60 |
-
| **Needle3-4L-29M** | 29M | 11.7 | 21.0 | 19.5 | 6.9 | 7.7 | 4.3 |
|
| 61 |
|
| 62 |
The interactive frontier plot, the architecture and the fine-tuning results are at [cactuscompute.com/needle](https://cactuscompute.com/needle).
|
| 63 |
|
|
|
|
| 29 |
|
| 30 |
## Model
|
| 31 |
|
| 32 |
+

|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the [release page](https://cactuscompute.com/needle). The repo holds the 20-layer `needle3.cact`, the `needle3.safetensors` checkpoint to fine-tune, and an engine per platform.
|
| 35 |
|
| 36 |
## Benchmarks
|
| 37 |
|
| 38 |
+
Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.
|
| 39 |
+
|
| 40 |
+

|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
|
| 42 |
The interactive frontier plot, the architecture and the fine-tuning results are at [cactuscompute.com/needle](https://cactuscompute.com/needle).
|
| 43 |
|
assets/benchmarks.svg
ADDED
|
|
assets/model.svg
ADDED
|
|