aaxaxax Claude Fable 5 commited on
Commit
5b950c2
·
1 Parent(s): 4563ab7

docs: pase visual con emojis al paquete de difusion v0.12

Browse files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Files changed (2) hide show
  1. hf-post-v012.md +26 -21
  2. hf-replies-v012.md +4 -4
hf-post-v012.md CHANGED
@@ -1,40 +1,45 @@
1
- # Show & Tell: TAF Agent v0.12 — "will it fit?" answered before you download, and the whole tool now speaks your language
2
 
3
- > **TL;DR** — Free, no-signup, browser-only LLM diagnostics. New in v0.12: **💾 Fit Check** answers the most-asked question on this forum — *"will this model fit on my GPU at my context length?"* — including the part every VRAM calculator forgets (the KV cache), **before you download a single byte**. And the tool now works **end-to-end in EN/ES/FR/ZH**: not just menus — guided demos, recipe verdicts and explanations follow your browser language.
4
  >
5
- > **Try it in 10 seconds** (the demo runs itself): https://karlesmarin.github.io/tafagent/?demo=fitcheck
6
- > **Live**: https://huggingface.co/spaces/karlexmarin/taf-agent · **Source**: https://github.com/karlesmarin/tafagent
 
 
 
7
 
8
  ---
9
 
10
- ## The problem Fit Check solves
11
 
12
  The recurring surprise, straight from this forum: *"Llama-3.1-8B is ~16 GB in fp16 — why is my 24 GB GPU full?"* Because **weights are only half the story**. The KV cache grows linearly with context, and at 128K it can weigh as much as the model:
13
 
14
  ```
15
  Llama-3.1-8B · fp16 · RTX 4090 (24 GB) · context 131072
16
- weights 15.0 GB
17
- KV cache 16.0 GB ← 128 KB per token, 50% of the total
18
- scratch 0.9 GB
19
- ─────────────────
20
- total 31.9 GB → 🚨 DOES NOT FIT — and it's the KV cache, not the model
21
- max context that DOES fit: ~66,000 tokens
22
- cheapest rescues, in order: q8_0 KV cache → Q6_K weights → partial offload
23
  ```
24
 
25
- That's the whole mode: model id (geometry fetched from `config.json`) + precision (fp16/bf16/int8/nf4 or any GGUF quant) + GPU + target context → full budget, a verdict that names **which side is the problem** (weights-bound vs KV-bound), the **max context that fits**, and the cheapest fix. Same budget math as the Launch-Flag Generator, so the two modes can never disagree.
26
 
27
- ## Your language, end to end
28
 
29
  Until v0.11 the menus were translated but demos and recipe results came out in English. v0.12 fixes it at the root: the Python recipe engine emits message codes that the UI localizes (EN/ES/FR/ZH, English always kept as fallback), and the guided demos follow your browser language automatically. Guarded by a real-browser regression test (Spanish locale, clean storage, full demo + full Profile, zero English residue) that now runs in CI on every push, along with a sweep of every mode in every language.
30
 
31
- ## What else is in the box (29 modes, all browser-only)
 
 
32
 
33
- Profile (5-recipe TAF Card from a model id), Context Unmasker (is `max_position_embeddings` honest?), Chat-template Sniffer (the lm-eval #1841 silent-halving fix), Quant-regime, NIAH→Reason, LongScore (RULER+HELMET), Arena-Elo CI reconstructor, Contamination prior, YaRN planner, GGUF Bridge (reads headers via HTTP Range — no download), Launch Flags, Cache Diff, Spec-Decode compatibility, Token Tax, and more. Every mode has a **🎬 Demo** button that walks you through it.
34
 
35
- ## Honesty section (as always)
36
 
37
- - Everything from `config.json` is a **prediction**, not a measurement. When it matters, measure: the Diagnose CLI generates the command, and Prediction-vs-Reality compares.
38
- - The quant γ-shift constants are hand-set, mostly uncalibrated (flagged in-tool).
39
- - Verdict pill labels ("GO", "MEMORY-LIMITED"…) are still English-only; next batch.
40
- - Nothing you type leaves your browser. No server, no telemetry, no signup.
 
1
+ # 💾 Show & Tell: TAF Agent v0.12 — "will it fit?" answered before you download, and the whole tool now speaks your language 🌍
2
 
3
+ > **⚡ TL;DR** — Free, no-signup, browser-only LLM diagnostics.
4
  >
5
+ > 🆕 **💾 Fit Check** answers the most-asked question on this forum — *"will this model fit on my GPU at my context length?"* — including the part every VRAM calculator forgets (**the KV cache**), before you download a single byte.
6
+ > 🆕 **🌍 Four languages end-to-end** (EN/ES/FR/ZH): not just menus — guided demos, recipe verdicts and explanations follow your browser language.
7
+ >
8
+ > 🎬 **Try it in 10 seconds** (the demo runs itself): https://karlesmarin.github.io/tafagent/?demo=fitcheck
9
+ > 🚀 **Live**: https://huggingface.co/spaces/karlexmarin/taf-agent · 📦 **Source**: https://github.com/karlesmarin/tafagent
10
 
11
  ---
12
 
13
+ ## 🧨 The problem Fit Check solves
14
 
15
  The recurring surprise, straight from this forum: *"Llama-3.1-8B is ~16 GB in fp16 — why is my 24 GB GPU full?"* Because **weights are only half the story**. The KV cache grows linearly with context, and at 128K it can weigh as much as the model:
16
 
17
  ```
18
  Llama-3.1-8B · fp16 · RTX 4090 (24 GB) · context 131072
19
+ ⚖️ weights 15.0 GB
20
+ 🧊 KV cache 16.0 GB ← 128 KB per token, 50% of the total!
21
+ 🔧 scratch 0.9 GB
22
+ ─────────────────────
23
+ 📦 total 31.9 GB → 🚨 DOES NOT FIT — and it's the KV cache, not the model
24
+ ✂️ max context that DOES fit: ~66,000 tokens
25
+ 💡 cheapest rescues, in order: q8_0 KV cache → Q6_K weights → partial offload
26
  ```
27
 
28
+ That's the whole mode: model id (geometry fetched from `config.json`) + precision (fp16/bf16/int8/nf4 or any GGUF quant) + GPU + target context → full budget, a verdict that names **which side is the problem** (⚖️ weights-bound vs 🧊 KV-bound), the **max context that fits**, and the cheapest fix. Same budget math as the Launch-Flag Generator, so the two modes can never disagree.
29
 
30
+ ## 🌍 Your language, end to end
31
 
32
  Until v0.11 the menus were translated but demos and recipe results came out in English. v0.12 fixes it at the root: the Python recipe engine emits message codes that the UI localizes (EN/ES/FR/ZH, English always kept as fallback), and the guided demos follow your browser language automatically. Guarded by a real-browser regression test (Spanish locale, clean storage, full demo + full Profile, zero English residue) that now runs in CI on every push, along with a sweep of every mode in every language.
33
 
34
+ ## 🧰 What else is in the box (29 modes, all browser-only)
35
+
36
+ 📇 Profile (5-recipe TAF Card from a model id) · 🪟 Context Unmasker (is `max_position_embeddings` honest?) · 📜 Chat-template Sniffer (the lm-eval #1841 silent-halving fix) · ⚖️ Quant-regime · 🔍 NIAH→Reason · 🎯 LongScore (RULER+HELMET) · 🎯 Arena-Elo CI reconstructor · 🧪 Contamination prior · 🧵 YaRN planner · 🧊 GGUF Bridge (reads headers via HTTP Range — no download) · 🚀 Launch Flags · 🔁 Cache Diff · 🔬 Spec-Decode compatibility · 🌍 Token Tax · and more.
37
 
38
+ Every mode has a **🎬 Demo** button that walks you through it, step by step, in your language.
39
 
40
+ ## ⚖️ Honesty section (as always)
41
 
42
+ - 📐 Everything from `config.json` is a **prediction**, not a measurement. When it matters, measure: the Diagnose CLI generates the command, and Prediction-vs-Reality compares.
43
+ - ⚠️ The quant γ-shift constants are hand-set, mostly uncalibrated (flagged in-tool).
44
+ - 🚧 Verdict pill labels ("GO", "MEMORY-LIMITED"…) are still English-only; next batch.
45
+ - 🔒 Nothing you type leaves your browser. No server, no telemetry, no signup.
hf-replies-v012.md CHANGED
@@ -10,9 +10,9 @@ There's a concrete way to answer this for YOUR gpu + YOUR context length, becaus
10
 
11
  Rough rules that fall out of the arithmetic:
12
 
13
- - **Short context (≤8K)**: the KV cache is small, so the question is almost purely "quant cliff vs parameter count". A bigger model at Q4_K_M usually beats a smaller one at fp16 — Q4_K_M sits before the quality cliff for most ≥7B models, and parameters buy more than precision there.
14
- - **Long context (≥32K)**: the KV cache starts to dominate the budget, and it does NOT shrink with weight quantization (it depends on layers × kv-heads × head_dim × context). Here a smaller model — or a q8_0-quantized KV cache — often wins, because the big-model-Q3 option leaves no room for the cache.
15
- - **The cliff**: below Q4 (Q3_K_M, Q2_K) quality degradation is model-dependent and can be steep; below-4-bit "big model" setups often lose to a clean 8B.
16
 
17
  If you want the numbers for your exact case without downloading anything: I built a free browser tool that does this budget (weights + KV + scratch) from the model's `config.json`, tells you which side is the problem if it doesn't fit, the max context that DOES fit, and scans the quant ladder for the first one that works. Demo that runs itself: https://karlesmarin.github.io/tafagent/?demo=fitcheck — no signup, nothing leaves your browser.
18
 
@@ -28,7 +28,7 @@ KV bytes = 2 (K+V) × layers × kv_heads × head_dim × context × bytes/elem
28
  Llama-3.1-8B: 2 × 32 × 8 × 128 × 131072 × 2 (fp16) = 16 GB (128 KB per token!)
29
  ```
30
 
31
- So at the advertised 128K context: 15 GB of weights + 16 GB of cache + scratch ≈ 32 GB — it was never going to fit in 24 GB, and it's not the model's "fault": it's the context. Three levers, cheapest first: quantize the cache (`q8_0` halves it), cap the context (~66K fits in 24 GB with fp16 weights), or drop weight precision.
32
 
33
  I packaged this arithmetic (plus the "which side is the problem" verdict and the max-context solver) into a free browser tool — geometry fetched from the model card, nothing downloaded: https://karlesmarin.github.io/tafagent/?demo=fitcheck
34
 
 
10
 
11
  Rough rules that fall out of the arithmetic:
12
 
13
+ - 📏 **Short context (≤8K)**: the KV cache is small, so the question is almost purely "quant cliff vs parameter count". A bigger model at Q4_K_M usually beats a smaller one at fp16 — Q4_K_M sits before the quality cliff for most ≥7B models, and parameters buy more than precision there.
14
+ - 🧊 **Long context (≥32K)**: the KV cache starts to dominate the budget, and it does NOT shrink with weight quantization (it depends on layers × kv-heads × head_dim × context). Here a smaller model — or a q8_0-quantized KV cache — often wins, because the big-model-Q3 option leaves no room for the cache.
15
+ - ⛔ **The cliff**: below Q4 (Q3_K_M, Q2_K) quality degradation is model-dependent and can be steep; below-4-bit "big model" setups often lose to a clean 8B.
16
 
17
  If you want the numbers for your exact case without downloading anything: I built a free browser tool that does this budget (weights + KV + scratch) from the model's `config.json`, tells you which side is the problem if it doesn't fit, the max context that DOES fit, and scans the quant ladder for the first one that works. Demo that runs itself: https://karlesmarin.github.io/tafagent/?demo=fitcheck — no signup, nothing leaves your browser.
18
 
 
28
  Llama-3.1-8B: 2 × 32 × 8 × 128 × 131072 × 2 (fp16) = 16 GB (128 KB per token!)
29
  ```
30
 
31
+ So at the advertised 128K context: ⚖️ 15 GB of weights + 🧊 16 GB of cache + scratch ≈ 32 GB — it was never going to fit in 24 GB, and it's not the model's "fault": it's the context. Three levers, cheapest first: 💡 quantize the cache (`q8_0` halves it), ✂️ cap the context (~66K fits in 24 GB with fp16 weights), or ⚖️ drop weight precision.
32
 
33
  I packaged this arithmetic (plus the "which side is the problem" verdict and the max-context solver) into a free browser tool — geometry fetched from the model card, nothing downloaded: https://karlesmarin.github.io/tafagent/?demo=fitcheck
34