mbovingfred commited on
Commit
90f53f1
·
verified ·
1 Parent(s): 9f2f141

Upload folder using huggingface_hub

Browse files
Files changed (9) hide show
  1. .gitignore +9 -0
  2. AGENTS.md +7 -0
  3. README.md +46 -6
  4. app.py +400 -0
  5. docs/FEATURES.md +26 -0
  6. requirements.txt +3 -0
  7. studio.py +95 -0
  8. tests/test_space_contract.py +50 -0
  9. tests/test_studio.py +77 -0
.gitignore ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ __pycache__/
2
+ *.py[cod]
3
+ .venv/
4
+ .pytest_cache/
5
+ .gradio/
6
+ *.nemo
7
+ *.wav
8
+ *.mp3
9
+ .DS_Store
AGENTS.md ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ # Project instructions
2
+
3
+ - Keep this project deployable as a Hugging Face Gradio Space on ZeroGPU.
4
+ - Preserve the import order in `app.py`: set CUDA-related environment variables, import `spaces`, then import libraries that may touch CUDA.
5
+ - Never commit Hugging Face tokens, API keys, generated audio, model checkpoints, or caches.
6
+ - Run `python -m unittest discover -s tests -v` and `python -m py_compile app.py studio.py` after changes.
7
+ - Update `docs/FEATURES.md` whenever features are added, edited, removed, renamed, made reachable, made unreachable, or newly covered by automated tests. Every feature must state its coverage and link to the exact test file when covered.
README.md CHANGED
@@ -1,13 +1,53 @@
1
  ---
2
- title: Magpie Tts Studio
3
- emoji: ⚡
4
- colorFrom: indigo
5
- colorTo: red
6
  sdk: gradio
7
  sdk_version: 6.24.0
8
- python_version: '3.12'
9
  app_file: app.py
 
 
10
  pinned: false
 
 
 
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: Magpie Voice Studio
3
+ emoji: 🎙️
4
+ colorFrom: yellow
5
+ colorTo: green
6
  sdk: gradio
7
  sdk_version: 6.24.0
 
8
  app_file: app.py
9
+ python_version: "3.12"
10
+ startup_duration_timeout: 1h
11
  pinned: false
12
+ license: other
13
+ license_name: nvidia-open-model-license
14
+ license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license
15
+ short_description: Multilingual speech in five voices and 12 languages
16
  ---
17
 
18
+ # Magpie Voice Studio
19
+
20
+ A polished ZeroGPU demo for NVIDIA's [MagpieTTS Multilingual 357M](https://huggingface.co/nvidia/magpie_tts_multilingual_357m) model. Write a short script, choose one of 12 languages and five fixed voices, and render 22.05 kHz speech in the browser.
21
+
22
+ The app exposes a stable `/synthesize` Gradio API and MCP tool in addition to its visual interface. It does not support voice cloning.
23
+
24
+ ## Run locally
25
+
26
+ Use Python 3.12 with an NVIDIA CUDA GPU:
27
+
28
+ ```bash
29
+ python -m venv .venv
30
+ source .venv/bin/activate
31
+ pip install -r requirements.txt
32
+ python app.py
33
+ ```
34
+
35
+ The model checkpoint is downloaded from Hugging Face on first launch. If model access requires authentication, set `HF_TOKEN` to a read token with access to the repository.
36
+
37
+ ## Test
38
+
39
+ The fast test suite covers configuration, validation, ZeroGPU contracts, and Space metadata without loading the 5.5 GB model:
40
+
41
+ ```bash
42
+ python -m unittest discover -s tests -v
43
+ ```
44
+
45
+ See the exhaustive [feature inventory and automated test coverage matrix](docs/FEATURES.md).
46
+
47
+ ## Model and license
48
+
49
+ - [MagpieTTS model card](https://huggingface.co/nvidia/magpie_tts_multilingual_357m)
50
+ - [NanoCodec decoder](https://huggingface.co/nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps)
51
+ - [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license)
52
+
53
+ Model outputs may contain errors or accents. Review generated audio before production use.
app.py ADDED
@@ -0,0 +1,400 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Magpie Voice Studio — a multilingual TTS demo for Hugging Face Spaces."""
2
+
3
+ import os
4
+ import time
5
+ import traceback
6
+
7
+ os.environ.setdefault("NUMBA_DISABLE_CUDA", "1")
8
+ os.environ.setdefault("OMP_NUM_THREADS", "2")
9
+ os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
10
+
11
+ import spaces
12
+ import gradio as gr
13
+ from huggingface_hub import hf_hub_download
14
+ from nemo.collections.tts.modules.magpietts_inference.utils import (
15
+ ModelLoadConfig,
16
+ load_magpie_model,
17
+ )
18
+
19
+ from studio import (
20
+ LANGUAGE_MAP,
21
+ SAMPLE_RATE,
22
+ SPEAKER_MAP,
23
+ estimate_gpu_duration,
24
+ result_summary,
25
+ validate_request,
26
+ )
27
+
28
+
29
+ MODEL_ID = "nvidia/magpie_tts_multilingual_357m"
30
+ MODEL_FILENAME = "magpie_tts_multilingual_357m.nemo"
31
+ CODEC_MODEL_ID = "nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps"
32
+
33
+ checkpoint_path = hf_hub_download(
34
+ repo_id=MODEL_ID,
35
+ filename=MODEL_FILENAME,
36
+ token=os.environ.get("HF_TOKEN"),
37
+ )
38
+ model_config = ModelLoadConfig(
39
+ nemo_file=checkpoint_path,
40
+ codecmodel_path=CODEC_MODEL_ID,
41
+ legacy_codebooks=False,
42
+ legacy_text_conditioning=False,
43
+ hparams_from_wandb=None,
44
+ )
45
+
46
+ print(f"Loading {MODEL_ID} from {checkpoint_path}")
47
+ model, _ = load_magpie_model(model_config)
48
+ model.eval().to("cuda")
49
+ print("MagpieTTS loaded and ready.")
50
+
51
+
52
+ @spaces.GPU(duration=estimate_gpu_duration)
53
+ def synthesize(
54
+ text: str,
55
+ language: str = "English",
56
+ speaker: str = "Sofia",
57
+ apply_text_normalization: bool = True,
58
+ ) -> tuple[tuple[int, object], str]:
59
+ """Synthesize mono WAV-quality speech from text with NVIDIA MagpieTTS.
60
+
61
+ Args:
62
+ text: UTF-8 text to speak, up to 600 characters.
63
+ language: One of the 12 language labels shown in the app.
64
+ speaker: One of Aria, Jason, John, Leo, or Sofia.
65
+ apply_text_normalization: Expand numbers and abbreviations before synthesis.
66
+
67
+ Returns:
68
+ An audio sample tuple and a short generation summary.
69
+ """
70
+
71
+ try:
72
+ request = validate_request(
73
+ text, language, speaker, apply_text_normalization
74
+ )
75
+ except ValueError as exc:
76
+ raise gr.Error(str(exc)) from exc
77
+
78
+ started_at = time.perf_counter()
79
+ try:
80
+ audio, audio_length = model.do_tts(
81
+ request.text,
82
+ language=request.language_code,
83
+ apply_TN=request.apply_text_normalization,
84
+ speaker_index=request.speaker_index,
85
+ )
86
+ audio_samples = audio[0, : int(audio_length[0])].detach().cpu().numpy()
87
+ except Exception as exc:
88
+ traceback.print_exc()
89
+ raise gr.Error(f"Speech synthesis failed: {exc}") from exc
90
+
91
+ elapsed = time.perf_counter() - started_at
92
+ print(
93
+ f"Synthesized {len(request.text)} chars in {elapsed:.2f}s "
94
+ f"({request.language_code}, speaker={request.speaker_index})"
95
+ )
96
+ return (
97
+ (SAMPLE_RATE, audio_samples),
98
+ result_summary(language, speaker, len(request.text), elapsed),
99
+ )
100
+
101
+
102
+ CSS = """
103
+ @import url('https://fonts.googleapis.com/css2?family=DM+Mono:wght@400;500&family=Instrument+Serif:ital@0;1&display=swap');
104
+
105
+ :root {
106
+ --ink: #171714;
107
+ --paper: #f3f0e8;
108
+ --paper-2: #e7e2d6;
109
+ --acid: #c8ff38;
110
+ --orange: #ff6838;
111
+ --line: rgba(23, 23, 20, 0.2);
112
+ }
113
+
114
+ body, .gradio-container {
115
+ background:
116
+ radial-gradient(circle at 85% 4%, rgba(255, 104, 56, .16), transparent 24rem),
117
+ linear-gradient(rgba(23,23,20,.035) 1px, transparent 1px),
118
+ linear-gradient(90deg, rgba(23,23,20,.035) 1px, transparent 1px),
119
+ var(--paper) !important;
120
+ background-size: auto, 32px 32px, 32px 32px, auto !important;
121
+ color: var(--ink) !important;
122
+ font-family: 'DM Mono', monospace !important;
123
+ }
124
+
125
+ .dark .gradio-container { color: var(--body-text-color); }
126
+ .gradio-container { max-width: 1240px !important; margin: 0 auto !important; }
127
+ .contain, main { max-width: 1240px !important; margin-inline: auto !important; }
128
+ .wrap { gap: 0 !important; }
129
+
130
+ .hero {
131
+ position: relative;
132
+ overflow: hidden;
133
+ min-height: 350px;
134
+ padding: 38px 42px 34px;
135
+ border: 1.5px solid var(--ink);
136
+ border-radius: 4px 4px 0 0;
137
+ background: var(--ink);
138
+ color: var(--paper);
139
+ }
140
+ .hero::after {
141
+ content: '';
142
+ position: absolute;
143
+ width: 380px;
144
+ height: 380px;
145
+ right: -75px;
146
+ top: -135px;
147
+ border: 1px solid rgba(243,240,232,.34);
148
+ border-radius: 50%;
149
+ box-shadow: 0 0 0 54px rgba(243,240,232,.055), 0 0 0 108px rgba(243,240,232,.035);
150
+ }
151
+ .eyebrow, .panel-index {
152
+ color: var(--acid);
153
+ font: 500 12px/1.2 'DM Mono', monospace;
154
+ letter-spacing: .16em;
155
+ text-transform: uppercase;
156
+ }
157
+ .hero h1 {
158
+ max-width: 860px;
159
+ margin: 52px 0 20px;
160
+ font: 400 clamp(55px, 8vw, 102px)/.82 'Instrument Serif', serif;
161
+ letter-spacing: -.055em;
162
+ }
163
+ .hero h1 em { color: var(--acid); font-weight: 400; }
164
+ .hero-copy {
165
+ max-width: 670px;
166
+ margin: 0;
167
+ color: rgba(243,240,232,.72);
168
+ font-size: 14px;
169
+ line-height: 1.65;
170
+ }
171
+ .signal {
172
+ position: absolute;
173
+ right: 44px;
174
+ bottom: 34px;
175
+ display: flex;
176
+ align-items: end;
177
+ gap: 5px;
178
+ height: 54px;
179
+ }
180
+ .signal i { width: 5px; background: var(--acid); animation: pulse 1.3s ease-in-out infinite alternate; }
181
+ .signal i:nth-child(1), .signal i:nth-child(7) { height: 18px; }
182
+ .signal i:nth-child(2), .signal i:nth-child(6) { height: 33px; animation-delay: -.4s; }
183
+ .signal i:nth-child(3), .signal i:nth-child(5) { height: 47px; animation-delay: -.8s; }
184
+ .signal i:nth-child(4) { height: 25px; animation-delay: -.2s; }
185
+ @keyframes pulse { to { transform: scaleY(.45); opacity: .55; } }
186
+
187
+ .stats {
188
+ display: grid;
189
+ grid-template-columns: repeat(4, 1fr);
190
+ border: 1.5px solid var(--ink);
191
+ border-top: 0;
192
+ background: var(--acid);
193
+ }
194
+ .stat { padding: 15px 20px; border-right: 1px solid var(--ink); }
195
+ .stat:last-child { border-right: 0; }
196
+ .stat b { display:block; font-size: 18px; }
197
+ .stat span { font-size: 10px; text-transform: uppercase; letter-spacing: .1em; opacity: .65; }
198
+
199
+ #studio-shell {
200
+ gap: 0 !important;
201
+ border: 1.5px solid var(--ink);
202
+ border-top: 0;
203
+ background: rgba(243,240,232,.86);
204
+ }
205
+ #controls-panel, #output-panel { padding: 30px !important; }
206
+ #controls-panel { border-right: 1.5px solid var(--ink); }
207
+ .panel-head { margin: 0 0 22px; }
208
+ .panel-head h2 { margin: 8px 0 0; font: 400 37px/1 'Instrument Serif', serif; }
209
+ .panel-index { color: rgba(23,23,20,.52); }
210
+
211
+ .gradio-container label span, .gradio-container .label-wrap span {
212
+ font-family: 'DM Mono', monospace !important;
213
+ text-transform: uppercase;
214
+ letter-spacing: .08em;
215
+ font-size: 10px !important;
216
+ }
217
+ .gradio-container textarea, .gradio-container input {
218
+ font-family: 'DM Mono', monospace !important;
219
+ }
220
+ .gradio-container .block {
221
+ border-color: var(--line) !important;
222
+ border-radius: 2px !important;
223
+ box-shadow: none !important;
224
+ }
225
+ #script-box textarea { min-height: 190px !important; font-size: 16px !important; line-height: 1.6 !important; }
226
+ #generate-btn {
227
+ min-height: 58px;
228
+ margin-top: 8px;
229
+ border: 1.5px solid var(--ink) !important;
230
+ border-radius: 2px !important;
231
+ background: var(--ink) !important;
232
+ color: var(--acid) !important;
233
+ font: 500 13px 'DM Mono', monospace !important;
234
+ letter-spacing: .12em;
235
+ text-transform: uppercase;
236
+ transition: transform .18s ease, box-shadow .18s ease !important;
237
+ }
238
+ #generate-btn:hover { transform: translate(-3px,-3px); box-shadow: 6px 6px 0 var(--orange) !important; }
239
+
240
+ #audio-output { margin-top: 6px; min-height: 245px; background: var(--paper-2) !important; }
241
+ #result-meta { min-height: 42px; border: 0 !important; background: transparent !important; }
242
+ #result-meta p { font-size: 11px; letter-spacing: .03em; }
243
+ .output-note { margin-top: 20px; padding: 17px; border-left: 4px solid var(--orange); background: rgba(255,104,56,.09); font-size: 11px; line-height: 1.6; }
244
+
245
+ #examples-wrap { border: 1.5px solid var(--ink); border-top: 0; padding: 28px; }
246
+ #examples-wrap h3 { margin: 0 0 4px; font: 400 29px 'Instrument Serif', serif; }
247
+ #examples-wrap p { margin-top: 0; opacity: .65; font-size: 11px; }
248
+
249
+ .footer {
250
+ display: flex;
251
+ justify-content: space-between;
252
+ gap: 24px;
253
+ padding: 21px 2px;
254
+ color: rgba(23,23,20,.66);
255
+ font-size: 10px;
256
+ line-height: 1.6;
257
+ text-transform: uppercase;
258
+ letter-spacing: .07em;
259
+ }
260
+ .footer a { color: var(--ink) !important; text-decoration: underline; text-underline-offset: 3px; }
261
+
262
+ @media (max-width: 780px) {
263
+ .hero { min-height: 405px; padding: 28px 24px; }
264
+ .hero h1 { margin-top: 46px; font-size: 57px; }
265
+ .signal { right: 25px; bottom: 24px; transform: scale(.75); transform-origin: bottom right; }
266
+ .stats { grid-template-columns: repeat(2, 1fr); }
267
+ .stat:nth-child(2) { border-right: 0; }
268
+ .stat:nth-child(-n+2) { border-bottom: 1px solid var(--ink); }
269
+ #controls-panel { border-right: 0; border-bottom: 1.5px solid var(--ink); }
270
+ #controls-panel, #output-panel { padding: 24px !important; }
271
+ .footer { flex-direction: column; padding-inline: 5px; }
272
+ }
273
+
274
+ @media (prefers-reduced-motion: reduce) {
275
+ .signal i { animation: none; }
276
+ #generate-btn { transition: none !important; }
277
+ }
278
+ """
279
+
280
+ HERO = """
281
+ <header class="hero">
282
+ <div class="eyebrow">Magpie / Voice Lab &nbsp;—&nbsp; NVIDIA NeMo</div>
283
+ <h1>Give text a<br><em>global voice.</em></h1>
284
+ <p class="hero-copy">A multilingual speech studio powered by MagpieTTS 357M. Write once, then hear one of five consistent voices speak across twelve languages.</p>
285
+ <div class="signal" aria-hidden="true"><i></i><i></i><i></i><i></i><i></i><i></i><i></i></div>
286
+ </header>
287
+ <div class="stats">
288
+ <div class="stat"><b>12</b><span>languages</span></div>
289
+ <div class="stat"><b>05</b><span>fixed voices</span></div>
290
+ <div class="stat"><b>22.05</b><span>kHz output</span></div>
291
+ <div class="stat"><b>357M</b><span>parameters</span></div>
292
+ </div>
293
+ """
294
+
295
+ FOOTER = """
296
+ <footer class="footer">
297
+ <span>Built with MagpieTTS · No voice cloning · ZeroGPU</span>
298
+ <span><a href="https://huggingface.co/nvidia/magpie_tts_multilingual_357m" target="_blank">Model card ↗</a> &nbsp; <a href="https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license" target="_blank">License ↗</a></span>
299
+ </footer>
300
+ """
301
+
302
+ EXAMPLES = [
303
+ ["Welcome to Magpie Voice Studio, where one voice can travel the world.", "English"],
304
+ ["Bienvenue dans le studio vocal multilingue de Magpie.", "French · Français"],
305
+ ["La tecnología puede acercar las voces y las culturas.", "Spanish · Español"],
306
+ ["音声技術で、ことばの世界をもっと身近に。", "Japanese · 日本語"],
307
+ ["언어의 경계를 넘어, 자연스러운 목소리를 만나보세요.", "Korean · 한국어"],
308
+ ["A tecnologia de voz conecta pessoas em todo o mundo.", "Portuguese · Português"],
309
+ ]
310
+
311
+
312
+ with gr.Blocks(theme=gr.themes.Base(), css=CSS, title="Magpie Voice Studio") as demo:
313
+ gr.HTML(HERO)
314
+
315
+ with gr.Row(elem_id="studio-shell"):
316
+ with gr.Column(scale=6, elem_id="controls-panel"):
317
+ gr.HTML(
318
+ '<div class="panel-head"><span class="panel-index">01 / Script</span>'
319
+ '<h2>Compose your take</h2></div>'
320
+ )
321
+ script = gr.Textbox(
322
+ label="Text to synthesize",
323
+ placeholder="Type a sentence, a product line, or a short narration…",
324
+ lines=7,
325
+ max_lines=12,
326
+ max_length=600,
327
+ elem_id="script-box",
328
+ )
329
+ with gr.Row():
330
+ language = gr.Dropdown(
331
+ choices=list(LANGUAGE_MAP),
332
+ value="English",
333
+ label="Language",
334
+ scale=3,
335
+ )
336
+ speaker = gr.Dropdown(
337
+ choices=list(SPEAKER_MAP),
338
+ value="Sofia",
339
+ label="Voice",
340
+ scale=2,
341
+ )
342
+ normalize = gr.Checkbox(
343
+ value=True,
344
+ label="Normalize numbers and abbreviations",
345
+ info="Recommended. Turn off when using custom IPA between | pipes |.",
346
+ )
347
+ generate = gr.Button("Generate voice →", variant="primary", elem_id="generate-btn")
348
+
349
+ with gr.Column(scale=5, elem_id="output-panel"):
350
+ gr.HTML(
351
+ '<div class="panel-head"><span class="panel-index">02 / Playback</span>'
352
+ '<h2>Listen to the take</h2></div>'
353
+ )
354
+ audio = gr.Audio(
355
+ label="Generated speech",
356
+ type="numpy",
357
+ interactive=False,
358
+ elem_id="audio-output",
359
+ )
360
+ metadata = gr.Markdown(
361
+ "**STANDING BY** &nbsp;·&nbsp; Your rendered take will appear here.",
362
+ elem_id="result-meta",
363
+ )
364
+ gr.HTML(
365
+ '<div class="output-note"><b>Studio note.</b> The five speakers are fixed model voices—not cloned voices. Non-English speech may retain an English accent.</div>'
366
+ )
367
+
368
+ with gr.Column(elem_id="examples-wrap"):
369
+ gr.HTML(
370
+ "<h3>Try a phrase</h3><p>Pick an example to load its text and language, then generate your take.</p>"
371
+ )
372
+ gr.Examples(
373
+ examples=EXAMPLES,
374
+ inputs=[script, language],
375
+ outputs=[audio, metadata],
376
+ fn=synthesize,
377
+ cache_examples=True,
378
+ cache_mode="lazy",
379
+ )
380
+
381
+ gr.HTML(FOOTER)
382
+
383
+ generate.click(
384
+ fn=synthesize,
385
+ inputs=[script, language, speaker, normalize],
386
+ outputs=[audio, metadata],
387
+ api_name="synthesize",
388
+ concurrency_limit=1,
389
+ )
390
+ script.submit(
391
+ fn=synthesize,
392
+ inputs=[script, language, speaker, normalize],
393
+ outputs=[audio, metadata],
394
+ api_name=False,
395
+ concurrency_limit=1,
396
+ )
397
+
398
+
399
+ if __name__ == "__main__":
400
+ demo.launch(mcp_server=True)
docs/FEATURES.md ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Feature inventory and automated test coverage
2
+
3
+ Status reflects the current `main` branch. “Covered” means at least one automated test exercises the feature or its deploy-time contract; “Untested” is explicit by design.
4
+
5
+ | Area | Feature | Reachability | Automated coverage |
6
+ |---|---|---|---|
7
+ | Synthesis | Generate speech from UTF-8 text | UI, Enter key, `/synthesize` API, MCP | Untested end-to-end; requires live ZeroGPU inference |
8
+ | Synthesis | Twelve supported language mappings | UI, API, MCP | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
9
+ | Synthesis | Five fixed speaker voices | UI, API, MCP | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
10
+ | Synthesis | Optional text normalization | UI, API, MCP | Covered request mapping: [`tests/test_studio.py`](../tests/test_studio.py); model behavior untested |
11
+ | Synthesis | Automatic terminal punctuation | UI, API, MCP | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
12
+ | Synthesis | Empty and whitespace-only input rejection | UI, API, MCP | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
13
+ | Synthesis | 600-character input limit | UI, API, MCP | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
14
+ | Synthesis | 22.05 kHz mono audio output | UI, API, MCP | Constant covered: [`tests/test_studio.py`](../tests/test_studio.py); waveform format untested locally |
15
+ | Runtime | NVIDIA MagpieTTS 357M and NanoCodec loading | Space startup | Untested locally; verified from live Space logs after deployment |
16
+ | Runtime | Dynamic, bounded ZeroGPU reservation | All synthesis routes | Covered: [`tests/test_studio.py`](../tests/test_studio.py) |
17
+ | Runtime | CUDA-safe import order and no forbidden Space runtime pins | Space startup | Covered: [`tests/test_space_contract.py`](../tests/test_space_contract.py) |
18
+ | Runtime | Serialized inference to protect shared model state | UI and API queue | Covered by source contract: [`tests/test_space_contract.py`](../tests/test_space_contract.py) |
19
+ | Interface | Responsive editorial voice-studio layout | Browser UI | Untested (visual) |
20
+ | Interface | Language, voice, and normalization controls | Browser UI | Untested (browser interaction) |
21
+ | Interface | Audio playback and downloadable Gradio audio | Browser UI | Untested (browser interaction) |
22
+ | Interface | Six multilingual example prompts with lazy caching | Browser UI | Covered by source contract: [`tests/test_space_contract.py`](../tests/test_space_contract.py); clicks untested |
23
+ | Interface | Reduced-motion accessibility behavior | Browser UI | Covered by source contract: [`tests/test_space_contract.py`](../tests/test_space_contract.py) |
24
+ | Interface | Model limitations, license, and source links | Browser UI, README | Untested (content review) |
25
+ | Integration | Stable `/synthesize` Gradio endpoint | API | Covered by source contract: [`tests/test_space_contract.py`](../tests/test_space_contract.py); live call required for e2e |
26
+ | Integration | MCP server exposing documented synthesis tool | MCP | Covered by source contract: [`tests/test_space_contract.py`](../tests/test_space_contract.py); MCP client call untested |
requirements.txt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ nemo_toolkit[tts] @ git+https://github.com/NVIDIA-NeMo/Speech.git@24c4f58a764350a62248094c70e487892e9d1a00
2
+ git+https://github.com/NVIDIA/NeMo-text-processing.git@1f1263579fe57ba7ed783cad3dddee710fcc5064
3
+ kaldialign
studio.py ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Pure configuration and validation helpers for Magpie Voice Studio."""
2
+
3
+ from __future__ import annotations
4
+
5
+ from dataclasses import dataclass
6
+
7
+
8
+ MAX_TEXT_LENGTH = 600
9
+ SAMPLE_RATE = 22_050
10
+
11
+ SPEAKER_MAP = {
12
+ "Aria": 0,
13
+ "Jason": 1,
14
+ "John": 2,
15
+ "Leo": 3,
16
+ "Sofia": 4,
17
+ }
18
+
19
+ LANGUAGE_MAP = {
20
+ "Arabic · العربية": "ar-MSA",
21
+ "Chinese · 中文": "zh",
22
+ "English": "en",
23
+ "French · Français": "fr",
24
+ "German · Deutsch": "de",
25
+ "Hindi · हिन्दी": "hi",
26
+ "Italian · Italiano": "it",
27
+ "Japanese · 日本語": "ja",
28
+ "Korean · 한국어": "ko",
29
+ "Portuguese · Português": "pt-BR",
30
+ "Spanish · Español": "es",
31
+ "Vietnamese · Tiếng Việt": "vi",
32
+ }
33
+
34
+
35
+ @dataclass(frozen=True)
36
+ class SynthesisRequest:
37
+ """A validated speech synthesis request."""
38
+
39
+ text: str
40
+ language_code: str
41
+ speaker_index: int
42
+ apply_text_normalization: bool
43
+
44
+
45
+ def validate_request(
46
+ text: str | None,
47
+ language: str,
48
+ speaker: str,
49
+ apply_text_normalization: bool,
50
+ ) -> SynthesisRequest:
51
+ """Validate and normalize values received from the Gradio API."""
52
+
53
+ clean_text = (text or "").strip()
54
+ if not clean_text:
55
+ raise ValueError("Enter some text to synthesize.")
56
+ if len(clean_text) > MAX_TEXT_LENGTH:
57
+ raise ValueError(
58
+ f"Keep the script under {MAX_TEXT_LENGTH} characters for this demo."
59
+ )
60
+ if language not in LANGUAGE_MAP:
61
+ raise ValueError("Choose one of the supported languages.")
62
+ if speaker not in SPEAKER_MAP:
63
+ raise ValueError("Choose one of the available Magpie voices.")
64
+
65
+ if clean_text[-1] not in ".?!。!?":
66
+ clean_text += "."
67
+
68
+ return SynthesisRequest(
69
+ text=clean_text,
70
+ language_code=LANGUAGE_MAP[language],
71
+ speaker_index=SPEAKER_MAP[speaker],
72
+ apply_text_normalization=bool(apply_text_normalization),
73
+ )
74
+
75
+
76
+ def estimate_gpu_duration(text: str | None, *_args: object, **_kwargs: object) -> int:
77
+ """Estimate a bounded ZeroGPU reservation from the requested text length."""
78
+
79
+ character_count = len((text or "").strip())
80
+ return min(180, max(75, 55 + character_count // 3))
81
+
82
+
83
+ def result_summary(
84
+ language: str,
85
+ speaker: str,
86
+ character_count: int,
87
+ elapsed_seconds: float,
88
+ ) -> str:
89
+ """Create the compact result metadata displayed under the audio player."""
90
+
91
+ return (
92
+ f"**READY** &nbsp;·&nbsp; {language} &nbsp;·&nbsp; {speaker} &nbsp;·&nbsp; "
93
+ f"{character_count} chars &nbsp;·&nbsp; {SAMPLE_RATE / 1000:g} kHz &nbsp;·&nbsp; "
94
+ f"{elapsed_seconds:.1f}s inference"
95
+ )
tests/test_space_contract.py ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from pathlib import Path
2
+ import unittest
3
+
4
+
5
+ ROOT = Path(__file__).resolve().parents[1]
6
+
7
+
8
+ class SpaceContractTests(unittest.TestCase):
9
+ @classmethod
10
+ def setUpClass(cls):
11
+ cls.app = (ROOT / "app.py").read_text()
12
+ cls.readme = (ROOT / "README.md").read_text()
13
+ cls.requirements = (ROOT / "requirements.txt").read_text()
14
+
15
+ def test_spaces_import_precedes_cuda_touching_libraries(self):
16
+ self.assertLess(self.app.index("import spaces"), self.app.index("import gradio"))
17
+ self.assertLess(self.app.index("import spaces"), self.app.index("from nemo"))
18
+
19
+ def test_gpu_function_is_bound_to_stable_api(self):
20
+ self.assertIn("@spaces.GPU(duration=estimate_gpu_duration)", self.app)
21
+ self.assertIn('api_name="synthesize"', self.app)
22
+ self.assertIn("concurrency_limit=1", self.app)
23
+
24
+ def test_examples_use_lazy_cache(self):
25
+ self.assertIn('cache_mode="lazy"', self.app)
26
+ self.assertIn("cache_examples=True", self.app)
27
+
28
+ def test_mcp_server_and_accessibility_contracts_are_enabled(self):
29
+ self.assertIn("demo.launch(mcp_server=True)", self.app)
30
+ self.assertIn("prefers-reduced-motion", self.app)
31
+
32
+ def test_runtime_managed_packages_are_not_pinned(self):
33
+ package_names = {
34
+ line.split("=", 1)[0].strip().lower()
35
+ for line in self.requirements.splitlines()
36
+ if line.strip() and not line.startswith("#")
37
+ }
38
+ self.assertNotIn("gradio", package_names)
39
+ self.assertNotIn("spaces", package_names)
40
+ self.assertNotIn("huggingface_hub", package_names)
41
+
42
+ def test_space_metadata_and_feature_inventory_link_exist(self):
43
+ self.assertIn("sdk: gradio", self.readme)
44
+ self.assertIn('python_version: "3.12"', self.readme)
45
+ self.assertIn("docs/FEATURES.md", self.readme)
46
+ self.assertTrue((ROOT / "docs" / "FEATURES.md").is_file())
47
+
48
+
49
+ if __name__ == "__main__":
50
+ unittest.main()
tests/test_studio.py ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import unittest
2
+
3
+ from studio import (
4
+ LANGUAGE_MAP,
5
+ MAX_TEXT_LENGTH,
6
+ SAMPLE_RATE,
7
+ SPEAKER_MAP,
8
+ estimate_gpu_duration,
9
+ result_summary,
10
+ validate_request,
11
+ )
12
+
13
+
14
+ class StudioConfigurationTests(unittest.TestCase):
15
+ def test_all_twelve_languages_are_exposed(self):
16
+ self.assertEqual(len(LANGUAGE_MAP), 12)
17
+ self.assertEqual(
18
+ set(LANGUAGE_MAP.values()),
19
+ {"ar-MSA", "de", "en", "es", "fr", "hi", "it", "ja", "ko", "pt-BR", "vi", "zh"},
20
+ )
21
+
22
+ def test_all_five_voices_map_to_unique_indices(self):
23
+ self.assertEqual(set(SPEAKER_MAP), {"Aria", "Jason", "John", "Leo", "Sofia"})
24
+ self.assertEqual(set(SPEAKER_MAP.values()), set(range(5)))
25
+
26
+ def test_sample_rate_is_22050_hz(self):
27
+ self.assertEqual(SAMPLE_RATE, 22_050)
28
+
29
+
30
+ class RequestValidationTests(unittest.TestCase):
31
+ def test_valid_request_is_stripped_and_punctuated(self):
32
+ request = validate_request(" Hello world ", "English", "Sofia", True)
33
+ self.assertEqual(request.text, "Hello world.")
34
+ self.assertEqual(request.language_code, "en")
35
+ self.assertEqual(request.speaker_index, 4)
36
+ self.assertTrue(request.apply_text_normalization)
37
+
38
+ def test_supported_non_latin_punctuation_is_preserved(self):
39
+ request = validate_request("こんにちは。", "Japanese · 日本語", "Aria", False)
40
+ self.assertEqual(request.text, "こんにちは。")
41
+ self.assertFalse(request.apply_text_normalization)
42
+
43
+ def test_empty_input_is_rejected(self):
44
+ for value in (None, "", " "):
45
+ with self.subTest(value=value), self.assertRaisesRegex(ValueError, "Enter some text"):
46
+ validate_request(value, "English", "Sofia", True)
47
+
48
+ def test_too_long_input_is_rejected(self):
49
+ with self.assertRaisesRegex(ValueError, str(MAX_TEXT_LENGTH)):
50
+ validate_request("x" * (MAX_TEXT_LENGTH + 1), "English", "Sofia", True)
51
+
52
+ def test_unknown_language_is_rejected(self):
53
+ with self.assertRaisesRegex(ValueError, "supported languages"):
54
+ validate_request("Hello", "Klingon", "Sofia", True)
55
+
56
+ def test_unknown_speaker_is_rejected(self):
57
+ with self.assertRaisesRegex(ValueError, "available Magpie voices"):
58
+ validate_request("Hello", "English", "Unknown", True)
59
+
60
+
61
+ class RuntimeHelperTests(unittest.TestCase):
62
+ def test_duration_estimate_is_bounded_and_scales(self):
63
+ self.assertEqual(estimate_gpu_duration("short"), 75)
64
+ self.assertGreater(estimate_gpu_duration("x" * 300), 75)
65
+ self.assertEqual(estimate_gpu_duration("x" * 10_000), 180)
66
+
67
+ def test_result_summary_contains_generation_metadata(self):
68
+ summary = result_summary("English", "Sofia", 42, 3.26)
69
+ self.assertIn("English", summary)
70
+ self.assertIn("Sofia", summary)
71
+ self.assertIn("42 chars", summary)
72
+ self.assertIn("22.05 kHz", summary)
73
+ self.assertIn("3.3s inference", summary)
74
+
75
+
76
+ if __name__ == "__main__":
77
+ unittest.main()