[FEEDBACK] Inference Playground ๐Ÿ•น๏ธ

#12
by enzostvs - opened
Hugging Face org

Inference Playground - Hugging Face

playground

This discussion is dedicated to providing feedback on the Inference Playground and Serverless Inference API.
๐Ÿ‘‰ Visit the playground: huggingface.co/playground

About the Inference Playground:

The Inference Playground is a user interface designed to simplify testing our serverless inference API with chat models. It lists available models for you to try, allowing you to experiment with each model's settings, test available models via a UI, and copy code snippets.

To view all available settings, refer to the Serverless Inference for Chat Completion documentation.

๐Ÿ‘‰ Browse available chat models

If you need more usage, you can subscribe to PRO โœจ.

The drop-down list only provides the ability to add a model for testing. How can I remove a model from testing?

Thanks,

Bon

Hugging Face org

Hi @bontruong , while hovering a model you can decide to remove it (see screenshot below)
just one thing: once the prompt has been run from a model you can't delete it anymore, but can still remove it for the future responses
Screenshot 2026-03-09 at 2.55.59โ€ฏPM

This keeps occuring even on re-login {"error":"Failed to perform inference: Invalid username or password."}

deleted
This comment has been hidden
Hugging Face org

Yes its on purpose @inventivecornflower, it to kept the main flow. You can fully reset it or add another flow with a right click if needed

้€้กน่งฃๆžไธ€ไธ‹ๅ•†ไธšๆจกๅž‹๏ผŒ็œ‹็œ‹ไฝ ๅฏนไธญๅ›ฝๅธ‚ๅœบไบ†่งฃๆœ‰ๅคšๆทฑ
๏ผˆ1๏ผ‰1ๅ•ๅ›žๆœฌ2ๅ•ๆ็Žฐ3ๅ•ๅˆ†็บขๆจกๅผ๏ผˆ้ซ˜ๆณขๅˆฉ่ฃ‚ๅ˜๏ผŒ็”จๆˆท่บบๅนณๆ”ถ็›Š๏ผ‰
๏ผˆ2๏ผ‰279ๆจกๅผไธƒไบบๆ‹ผๅ›ขๆจกๅผ๏ผˆโ€œ2ไบบๅ›žๆœฌใ€7ไบบๆˆๅ›ขใ€9๏ผˆไน…๏ผ‰ไบซๅˆ†็บขโ€๏ผ‰
๏ผˆ3๏ผ‰139ไน…ๅˆ†็บขๆจกๅผ๏ผˆ1 ๆฌก่ดญไนฐ่ฟ”ๆœฌ,3 ๅ•ๅˆ†ไฝฃ,9 ๆฑ ๅˆ†็บข๏ผ‰
๏ผˆ4๏ผ‰ๅ€่Žฑ้ฒœ็พŠๅฅถๆจกๅผ
๏ผˆ5๏ผ‰ๅ•†ๅฎถ่ฎฉๅˆฉ็บขๅŒ…่กฅ่ดดๆจกๅผ๏ผˆ5ๅ€ๅˆ†็บขๆ”ถ็›Šใ€ๅ•†ๅฎถ็บขๅŒ…ใ€ๆŽ’้˜Ÿๅ…ๅ•๏ผ‰
๏ผˆ6๏ผ‰ๅ…ƒๆฐ”ๅ…ˆ้”‹ๆฐดๆœบๆŽ’้˜Ÿ่ฟ”ๆจกๅผ๏ผˆๆฐดๆœบ+้—จๅบ—ใ€่ฟ›ไบŒๅ‡บไธ€๏ผ‰๏ผˆ7๏ผ‰้“พๅŠจ3+1๏ผˆๆป‘่ฝๆœบๅˆถ๏ผ‰321ๆจกๅผไธ‰ไธ‰ๅคๅˆถ๏ผŒ้“พๅŠจ2+1๏ผˆ่ฃ‚ๅ˜ๅฟซ ๅฟซ้€ŸไธŠ็บฟ ๅ‡็บงๆ–ฐๅŠŸ่ƒฝ๏ผ‰
๏ผˆ8๏ผ‰ๆŽจไธ‰่ฟ”ไธ€ไธƒๆ˜Ÿๅ›ข้˜Ÿๆœบๅˆถ๏ผˆไฝŽๆณขๆฏ”๏ผŒๅฟซ้€Ÿ่ฃ‚ๅ˜๏ผ‰
๏ผˆ9๏ผ‰ไธ€ไบฉ็”ฐๅˆไผ™ไบบๅ…ฑๅŒๅฏŒ่ฃ•็ณป็ปŸ๏ผˆๅ•†ไธšๆจกๅผๅฏผๅธˆ-็Ž‹ๅ†ฒๅ›ข้˜Ÿ็ญ–ๅˆ’ๆจกๅผ๏ผ‰
๏ผˆ10๏ผ‰ๅพช็Žฏ่ดญๆจกๅผ๏ผˆๆถˆ่ดน่ฟ”ๅˆฉใ€่ดก็Œฎๅ€ผไธŽๅˆ†็บข๏ผ‰
๏ผˆ11๏ผ‰็ปฟ่‰ฒ็งฏๅˆ†ๅขžๅ€ผๆจกๅผ๏ผˆๅ•่พนไธŠๆ‰ฌ๏ผŒๅชๆถจไธ่ทŒ๏ผ‰
๏ผˆ12๏ผ‰ไปฃ็†ๅˆ†้”€็ณป็ปŸ๏ผˆๅ›ข้˜Ÿ็บงๅทฎ ๅŒบๅŸŸไปฃ็† ๅ›ข้˜Ÿๅˆ†ๆถฆ ่‡ชๅŠจๅˆ†่ดฆ๏ผ‰
๏ผˆ13๏ผ‰ๆ–ฐ้›ถๅ”ฎๅˆ†้”€ๅ•†ๅŸŽ๏ผˆๆ‹ผๅ›ข๏ผŒ็งฏๅˆ†ๅ•†ๅŸŽ๏ผŒ็ง’ๆ€๏ผŒ็ ไปท๏ผŒ็งฏๅˆ†ๅ…‘ๆข๏ผŒ้—จๅบ—่‡ชๆ็ญ‰๏ผ‰
ๆœ€ๅŽๅˆคๅˆซๅ“ชไธ€็งๆจกๅž‹ๆœ€ๆœ‰ๆ•ˆ่ฃ‚ๅ˜ๆœ€ๅฟซ
ๅ“ชไธ€็งๆจกๅž‹ๆ—ข่ƒฝๅฟซ้€Ÿ่ฃ‚ๅ˜ๅˆๅ…ทๆœ‰ๅฏๆŒ็ปญๆ€ง
ๆœ‰ๆฒกๆœ‰ๅฏ่ƒฝ้€š่ฟ‡ๆ•ดๅˆไผ˜ๅŒ–ไปฅไธŠๆจกๅž‹ๅˆ›ๆ–ฐไธ€็งๆ›ดๅŠ ้€‚ๅˆๅฝ“ๅ‰ๅธ‚ๅœบ็š„ๆจกๅผ๏ผŸ

Keep getting inference failures:
{"error":"Failed to perform inference: JWT has expired: "exp" claim timestamp check failed"}

Makes it pretty un-usable as a tool. Above was in the middle of a chat with Kimi-2.6 though Novita, tried other models & providers, all with the same error in the same chat.

Also Playground needs a way to stop an inference in progress if you see that the thinking is wrong so you can correct your query other than just waste tokens on something that you know will not meet your requirements.

Need a way to increase response limits and timeouts as have in many cases been cut-off in output generation which wastes time and money making the endeavour to use playground also useless.

Hugging Face org

Hi @SenecaInExile sorry for the delay, didnt see your message.
1 - stop button: I'll add one to make it stoppable
2 - regarding your first answer it was a known issue, do you still have the issue ?

Hi @SenecaInExile sorry for the delay, didnt see your message.
1 - stop button: I'll add one to make it stoppable
Thanks, would help.

2 - regarding your first answer it was a known issue, do you still have the issue ?
Yes, seems to be more prevalent if you are repositioning in the 'desktop' while one of the models is working. (i.e. moving around, trying to re-position, etc). don't know if it's a direct correlation but if there is any movement while one or more models is processing the preponderance of errors (truncated responses, timestamp errors, etc) are much more frequent.

Some other items I'm noticing:
3) should have a means to copy the entire entry or response prompt block for that turn. You can grab the entire flow in json format, but would really like the means similar to say lm studio and other tools where you can copy the entire block /turn (including embedded code, etc in that step/block) to memory or to save it in markdown format. Currently have to basically highlight & scroll to copy the block manually and then paste that in another document. Not too user friendly.

  1. should be able to collapse/expand blocks. If you paste in say a handover document or other description document that is long you are scrolling down in playground for page and pages to get to the top/bottom. The quick navigation item on the bottom helps but it's way too small when you have say 500 or even 1000 lines of a document to paste in for review. (turns it into a vertical line). To collapse old steps would be great.

  2. should be able to save or default to what models to use when starting a new chat. (perhaps as a profile setting or just restore last used models) opposed to defaulting to the whatever is 'popular' at the moment. Just an annoyance that I have to remove and then add in the models that I want each time I start a new chat.

Hugging Face org

Thanks for the idea @SenecaInExile
I'll look into it this week

I keep getting this in playground: ("error:" Failed to perform inference. JWT has expired:"exp" claim timestamp check failed")

You've got to extend the timeout in playground to allow the user to answer clarifying questions in a chat session. Currently if output is provided the session terminates way too fast (less than a minute). And there is no way to continue the session without starting over from scratch which just wastes thousands of tokens for no reason. Either lengthen the time to allow the user to respond meaningfully, OR allow a continued session by sending current turn history context with a new connection (sub optimal as it still wastes tokens). But the current state of affairs is the worst option.

Sign up or log in to comment