Spaces:
Running
Running
docs: pase visual con emojis al paquete de difusion v0.12
Browse filesCo-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- hf-post-v012.md +26 -21
- hf-replies-v012.md +4 -4
hf-post-v012.md
CHANGED
|
@@ -1,40 +1,45 @@
|
|
| 1 |
-
# Show & Tell: TAF Agent v0.12 — "will it fit?" answered before you download, and the whole tool now speaks your language
|
| 2 |
|
| 3 |
-
> **TL;DR** — Free, no-signup, browser-only LLM diagnostics.
|
| 4 |
>
|
| 5 |
-
> **
|
| 6 |
-
> **
|
|
|
|
|
|
|
|
|
|
| 7 |
|
| 8 |
---
|
| 9 |
|
| 10 |
-
## The problem Fit Check solves
|
| 11 |
|
| 12 |
The recurring surprise, straight from this forum: *"Llama-3.1-8B is ~16 GB in fp16 — why is my 24 GB GPU full?"* Because **weights are only half the story**. The KV cache grows linearly with context, and at 128K it can weigh as much as the model:
|
| 13 |
|
| 14 |
```
|
| 15 |
Llama-3.1-8B · fp16 · RTX 4090 (24 GB) · context 131072
|
| 16 |
-
weights 15.0 GB
|
| 17 |
-
KV cache 16.0 GB ← 128 KB per token, 50% of the total
|
| 18 |
-
scratch 0.9 GB
|
| 19 |
-
─────────────────
|
| 20 |
-
total 31.9 GB → 🚨 DOES NOT FIT — and it's the KV cache, not the model
|
| 21 |
-
max context that DOES fit: ~66,000 tokens
|
| 22 |
-
cheapest rescues, in order: q8_0 KV cache → Q6_K weights → partial offload
|
| 23 |
```
|
| 24 |
|
| 25 |
-
That's the whole mode: model id (geometry fetched from `config.json`) + precision (fp16/bf16/int8/nf4 or any GGUF quant) + GPU + target context → full budget, a verdict that names **which side is the problem** (weights-bound vs KV-bound), the **max context that fits**, and the cheapest fix. Same budget math as the Launch-Flag Generator, so the two modes can never disagree.
|
| 26 |
|
| 27 |
-
## Your language, end to end
|
| 28 |
|
| 29 |
Until v0.11 the menus were translated but demos and recipe results came out in English. v0.12 fixes it at the root: the Python recipe engine emits message codes that the UI localizes (EN/ES/FR/ZH, English always kept as fallback), and the guided demos follow your browser language automatically. Guarded by a real-browser regression test (Spanish locale, clean storage, full demo + full Profile, zero English residue) that now runs in CI on every push, along with a sweep of every mode in every language.
|
| 30 |
|
| 31 |
-
## What else is in the box (29 modes, all browser-only)
|
|
|
|
|
|
|
| 32 |
|
| 33 |
-
|
| 34 |
|
| 35 |
-
## Honesty section (as always)
|
| 36 |
|
| 37 |
-
- Everything from `config.json` is a **prediction**, not a measurement. When it matters, measure: the Diagnose CLI generates the command, and Prediction-vs-Reality compares.
|
| 38 |
-
- The quant γ-shift constants are hand-set, mostly uncalibrated (flagged in-tool).
|
| 39 |
-
- Verdict pill labels ("GO", "MEMORY-LIMITED"…) are still English-only; next batch.
|
| 40 |
-
- Nothing you type leaves your browser. No server, no telemetry, no signup.
|
|
|
|
| 1 |
+
# 💾 Show & Tell: TAF Agent v0.12 — "will it fit?" answered before you download, and the whole tool now speaks your language 🌍
|
| 2 |
|
| 3 |
+
> **⚡ TL;DR** — Free, no-signup, browser-only LLM diagnostics.
|
| 4 |
>
|
| 5 |
+
> 🆕 **💾 Fit Check** answers the most-asked question on this forum — *"will this model fit on my GPU at my context length?"* — including the part every VRAM calculator forgets (**the KV cache**), before you download a single byte.
|
| 6 |
+
> 🆕 **🌍 Four languages end-to-end** (EN/ES/FR/ZH): not just menus — guided demos, recipe verdicts and explanations follow your browser language.
|
| 7 |
+
>
|
| 8 |
+
> 🎬 **Try it in 10 seconds** (the demo runs itself): https://karlesmarin.github.io/tafagent/?demo=fitcheck
|
| 9 |
+
> 🚀 **Live**: https://huggingface.co/spaces/karlexmarin/taf-agent · 📦 **Source**: https://github.com/karlesmarin/tafagent
|
| 10 |
|
| 11 |
---
|
| 12 |
|
| 13 |
+
## 🧨 The problem Fit Check solves
|
| 14 |
|
| 15 |
The recurring surprise, straight from this forum: *"Llama-3.1-8B is ~16 GB in fp16 — why is my 24 GB GPU full?"* Because **weights are only half the story**. The KV cache grows linearly with context, and at 128K it can weigh as much as the model:
|
| 16 |
|
| 17 |
```
|
| 18 |
Llama-3.1-8B · fp16 · RTX 4090 (24 GB) · context 131072
|
| 19 |
+
⚖️ weights 15.0 GB
|
| 20 |
+
🧊 KV cache 16.0 GB ← 128 KB per token, 50% of the total!
|
| 21 |
+
🔧 scratch 0.9 GB
|
| 22 |
+
─────────────────────
|
| 23 |
+
📦 total 31.9 GB → 🚨 DOES NOT FIT — and it's the KV cache, not the model
|
| 24 |
+
✂️ max context that DOES fit: ~66,000 tokens
|
| 25 |
+
💡 cheapest rescues, in order: q8_0 KV cache → Q6_K weights → partial offload
|
| 26 |
```
|
| 27 |
|
| 28 |
+
That's the whole mode: model id (geometry fetched from `config.json`) + precision (fp16/bf16/int8/nf4 or any GGUF quant) + GPU + target context → full budget, a verdict that names **which side is the problem** (⚖️ weights-bound vs 🧊 KV-bound), the **max context that fits**, and the cheapest fix. Same budget math as the Launch-Flag Generator, so the two modes can never disagree.
|
| 29 |
|
| 30 |
+
## 🌍 Your language, end to end
|
| 31 |
|
| 32 |
Until v0.11 the menus were translated but demos and recipe results came out in English. v0.12 fixes it at the root: the Python recipe engine emits message codes that the UI localizes (EN/ES/FR/ZH, English always kept as fallback), and the guided demos follow your browser language automatically. Guarded by a real-browser regression test (Spanish locale, clean storage, full demo + full Profile, zero English residue) that now runs in CI on every push, along with a sweep of every mode in every language.
|
| 33 |
|
| 34 |
+
## 🧰 What else is in the box (29 modes, all browser-only)
|
| 35 |
+
|
| 36 |
+
📇 Profile (5-recipe TAF Card from a model id) · 🪟 Context Unmasker (is `max_position_embeddings` honest?) · 📜 Chat-template Sniffer (the lm-eval #1841 silent-halving fix) · ⚖️ Quant-regime · 🔍 NIAH→Reason · 🎯 LongScore (RULER+HELMET) · 🎯 Arena-Elo CI reconstructor · 🧪 Contamination prior · 🧵 YaRN planner · 🧊 GGUF Bridge (reads headers via HTTP Range — no download) · 🚀 Launch Flags · 🔁 Cache Diff · 🔬 Spec-Decode compatibility · 🌍 Token Tax · and more.
|
| 37 |
|
| 38 |
+
Every mode has a **🎬 Demo** button that walks you through it, step by step, in your language.
|
| 39 |
|
| 40 |
+
## ⚖️ Honesty section (as always)
|
| 41 |
|
| 42 |
+
- 📐 Everything from `config.json` is a **prediction**, not a measurement. When it matters, measure: the Diagnose CLI generates the command, and Prediction-vs-Reality compares.
|
| 43 |
+
- ⚠️ The quant γ-shift constants are hand-set, mostly uncalibrated (flagged in-tool).
|
| 44 |
+
- 🚧 Verdict pill labels ("GO", "MEMORY-LIMITED"…) are still English-only; next batch.
|
| 45 |
+
- 🔒 Nothing you type leaves your browser. No server, no telemetry, no signup.
|
hf-replies-v012.md
CHANGED
|
@@ -10,9 +10,9 @@ There's a concrete way to answer this for YOUR gpu + YOUR context length, becaus
|
|
| 10 |
|
| 11 |
Rough rules that fall out of the arithmetic:
|
| 12 |
|
| 13 |
-
- **Short context (≤8K)**: the KV cache is small, so the question is almost purely "quant cliff vs parameter count". A bigger model at Q4_K_M usually beats a smaller one at fp16 — Q4_K_M sits before the quality cliff for most ≥7B models, and parameters buy more than precision there.
|
| 14 |
-
- **Long context (≥32K)**: the KV cache starts to dominate the budget, and it does NOT shrink with weight quantization (it depends on layers × kv-heads × head_dim × context). Here a smaller model — or a q8_0-quantized KV cache — often wins, because the big-model-Q3 option leaves no room for the cache.
|
| 15 |
-
- **The cliff**: below Q4 (Q3_K_M, Q2_K) quality degradation is model-dependent and can be steep; below-4-bit "big model" setups often lose to a clean 8B.
|
| 16 |
|
| 17 |
If you want the numbers for your exact case without downloading anything: I built a free browser tool that does this budget (weights + KV + scratch) from the model's `config.json`, tells you which side is the problem if it doesn't fit, the max context that DOES fit, and scans the quant ladder for the first one that works. Demo that runs itself: https://karlesmarin.github.io/tafagent/?demo=fitcheck — no signup, nothing leaves your browser.
|
| 18 |
|
|
@@ -28,7 +28,7 @@ KV bytes = 2 (K+V) × layers × kv_heads × head_dim × context × bytes/elem
|
|
| 28 |
Llama-3.1-8B: 2 × 32 × 8 × 128 × 131072 × 2 (fp16) = 16 GB (128 KB per token!)
|
| 29 |
```
|
| 30 |
|
| 31 |
-
So at the advertised 128K context: 15 GB of weights + 16 GB of cache + scratch ≈ 32 GB — it was never going to fit in 24 GB, and it's not the model's "fault": it's the context. Three levers, cheapest first: quantize the cache (`q8_0` halves it), cap the context (~66K fits in 24 GB with fp16 weights), or drop weight precision.
|
| 32 |
|
| 33 |
I packaged this arithmetic (plus the "which side is the problem" verdict and the max-context solver) into a free browser tool — geometry fetched from the model card, nothing downloaded: https://karlesmarin.github.io/tafagent/?demo=fitcheck
|
| 34 |
|
|
|
|
| 10 |
|
| 11 |
Rough rules that fall out of the arithmetic:
|
| 12 |
|
| 13 |
+
- 📏 **Short context (≤8K)**: the KV cache is small, so the question is almost purely "quant cliff vs parameter count". A bigger model at Q4_K_M usually beats a smaller one at fp16 — Q4_K_M sits before the quality cliff for most ≥7B models, and parameters buy more than precision there.
|
| 14 |
+
- 🧊 **Long context (≥32K)**: the KV cache starts to dominate the budget, and it does NOT shrink with weight quantization (it depends on layers × kv-heads × head_dim × context). Here a smaller model — or a q8_0-quantized KV cache — often wins, because the big-model-Q3 option leaves no room for the cache.
|
| 15 |
+
- ⛔ **The cliff**: below Q4 (Q3_K_M, Q2_K) quality degradation is model-dependent and can be steep; below-4-bit "big model" setups often lose to a clean 8B.
|
| 16 |
|
| 17 |
If you want the numbers for your exact case without downloading anything: I built a free browser tool that does this budget (weights + KV + scratch) from the model's `config.json`, tells you which side is the problem if it doesn't fit, the max context that DOES fit, and scans the quant ladder for the first one that works. Demo that runs itself: https://karlesmarin.github.io/tafagent/?demo=fitcheck — no signup, nothing leaves your browser.
|
| 18 |
|
|
|
|
| 28 |
Llama-3.1-8B: 2 × 32 × 8 × 128 × 131072 × 2 (fp16) = 16 GB (128 KB per token!)
|
| 29 |
```
|
| 30 |
|
| 31 |
+
So at the advertised 128K context: ⚖️ 15 GB of weights + 🧊 16 GB of cache + scratch ≈ 32 GB — it was never going to fit in 24 GB, and it's not the model's "fault": it's the context. Three levers, cheapest first: 💡 quantize the cache (`q8_0` halves it), ✂️ cap the context (~66K fits in 24 GB with fp16 weights), or ⚖️ drop weight precision.
|
| 32 |
|
| 33 |
I packaged this arithmetic (plus the "which side is the problem" verdict and the max-context solver) into a free browser tool — geometry fetched from the model card, nothing downloaded: https://karlesmarin.github.io/tafagent/?demo=fitcheck
|
| 34 |
|