henryNdubuaku commited on
Commit
55db00b
·
verified ·
1 Parent(s): bcb0317

Animated model card and benchmark chart

Browse files
Files changed (3) hide show
  1. README.md +4 -24
  2. assets/benchmarks.svg +336 -0
  3. assets/model.svg +37 -0
README.md CHANGED
@@ -29,35 +29,15 @@ Needle does three jobs, all of them on the device:
29
 
30
  ## Model
31
 
32
- | | |
33
- | --- | --- |
34
- | Inputs | Text prompts, plus tool definitions or an extraction schema |
35
- | Outputs | Tool calls or structured extraction |
36
- | Model | 29-121M Laddered Simple Attention Networks, CQ2 quantisation |
37
- | Training | 360B tokens of proprietary structured dataset |
38
- | Speed | 400-4k tokens/s decode and 1-10k tokens/s prefill on a Raspberry Pi 5 |
39
- | Customisation | Fine-tuning lifts every subnetwork 18 to 36 points on DroidCall; from 4 layers (29M) up the tuned subnetwork passes DeepSeek V4 Flash |
40
 
41
  Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the [release page](https://cactuscompute.com/needle). The repo holds the 20-layer `needle3.cact`, the `needle3.safetensors` checkpoint to fine-tune, and an engine per platform.
42
 
43
  ## Benchmarks
44
 
45
- Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits; Needle 3 runs through the shipped CQ2-bit binary with the confidence gate on, baselines at f16 under vLLM, DeepSeek V4 Flash through its cloud API.
46
-
47
- | Model | Params | Mobile Actions | DroidCall | BFCL v4 | DSTC8 F1 | SNIPS gold F1 | SNIPS 7-way F1 |
48
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
49
- | DeepSeek V4 Flash (cloud) | - | 88.4 | 60.5 | 77.2 | 80.0 | 69.4 | 66.7 |
50
- | **Needle3-20L-121M** | 121M | 86.0 | 47.0 | 50.2 | 40.7 | 30.2 | 24.7 |
51
- | LFM2.5 1.2B | 1.2B | 82.4 | 35.5 | 62.0 | 48.0 | 43.0 | 38.0 |
52
- | **Needle3-16L-98M** | 98M | 80.7 | 40.0 | 41.3 | 28.5 | 23.5 | 19.2 |
53
- | Qwen3.5 0.8B | 800M | 76.0 | 28.0 | 56.8 | 49.0 | 35.0 | 34.0 |
54
- | LFM2.5 350M | 350M | 72.8 | 32.5 | 59.1 | 20.0 | 34.0 | 29.0 |
55
- | LFM2.5 230M | 230M | 69.3 | 11.5 | 46.3 | 53.0 | 27.0 | 22.0 |
56
- | FunctionGemma 270M | 270M | 65.1 | 16.5 | 46.6 | 27.0 | 29.0 | 14.0 |
57
- | Needle 2 | 45M | 63.5 | 17.0 | - | - | - | - |
58
- | Apple FM | 3.0B | 57.6 | - | - | - | - | - |
59
- | **Needle3-8L-52M** | 52M | 36.8 | 36.5 | 28.2 | 15.3 | 16.6 | 10.1 |
60
- | **Needle3-4L-29M** | 29M | 11.7 | 21.0 | 19.5 | 6.9 | 7.7 | 4.3 |
61
 
62
  The interactive frontier plot, the architecture and the fine-tuning results are at [cactuscompute.com/needle](https://cactuscompute.com/needle).
63
 
 
29
 
30
  ## Model
31
 
32
+ ![Needle 3 at a glance](assets/model.svg)
 
 
 
 
 
 
 
33
 
34
  Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the [release page](https://cactuscompute.com/needle). The repo holds the 20-layer `needle3.cact`, the `needle3.safetensors` checkpoint to fine-tune, and an engine per platform.
35
 
36
  ## Benchmarks
37
 
38
+ Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.
39
+
40
+ ![Needle 3 against baselines on six benchmarks](assets/benchmarks.svg)
 
 
 
 
 
 
 
 
 
 
 
 
 
41
 
42
  The interactive frontier plot, the architecture and the fine-tuning results are at [cactuscompute.com/needle](https://cactuscompute.com/needle).
43
 
assets/benchmarks.svg ADDED
assets/model.svg ADDED