Spaces:
Running on CPU Upgrade
[FEEDBACK] Inference Playground ๐น๏ธ
Inference Playground - Hugging Face
This discussion is dedicated to providing feedback on the Inference Playground and Serverless Inference API.
๐ Visit the playground: huggingface.co/playground
About the Inference Playground:
The Inference Playground is a user interface designed to simplify testing our serverless inference API with chat models. It lists available models for you to try, allowing you to experiment with each model's settings, test available models via a UI, and copy code snippets.
To view all available settings, refer to the Serverless Inference for Chat Completion documentation.
๐ Browse available chat models
If you need more usage, you can subscribe to PRO โจ.
The drop-down list only provides the ability to add a model for testing. How can I remove a model from testing?
Thanks,
Bon
Hi @bontruong , while hovering a model you can decide to remove it (see screenshot below)
just one thing: once the prompt has been run from a model you can't delete it anymore, but can still remove it for the future responses
This keeps occuring even on re-login {"error":"Failed to perform inference: Invalid username or password."}
Yes its on purpose @inventivecornflower, it to kept the main flow. You can fully reset it or add another flow with a right click if needed
้้กน่งฃๆไธไธๅไธๆจกๅ๏ผ็็ไฝ ๅฏนไธญๅฝๅธๅบไบ่งฃๆๅคๆทฑ
๏ผ1๏ผ1ๅๅๆฌ2ๅๆ็ฐ3ๅๅ็บขๆจกๅผ๏ผ้ซๆณขๅฉ่ฃๅ๏ผ็จๆท่บบๅนณๆถ็๏ผ
๏ผ2๏ผ279ๆจกๅผไธไบบๆผๅขๆจกๅผ๏ผโ2ไบบๅๆฌใ7ไบบๆๅขใ9๏ผไน
๏ผไบซๅ็บขโ๏ผ
๏ผ3๏ผ139ไน
ๅ็บขๆจกๅผ๏ผ1 ๆฌก่ดญไนฐ่ฟๆฌ,3 ๅๅไฝฃ,9 ๆฑ ๅ็บข๏ผ
๏ผ4๏ผๅ่ฑ้ฒ็พๅฅถๆจกๅผ
๏ผ5๏ผๅๅฎถ่ฎฉๅฉ็บขๅ
่กฅ่ดดๆจกๅผ๏ผ5ๅๅ็บขๆถ็ใๅๅฎถ็บขๅ
ใๆ้ๅ
ๅ๏ผ
๏ผ6๏ผๅ
ๆฐๅ
้ๆฐดๆบๆ้่ฟๆจกๅผ๏ผๆฐดๆบ+้จๅบใ่ฟไบๅบไธ๏ผ๏ผ7๏ผ้พๅจ3+1๏ผๆป่ฝๆบๅถ๏ผ321ๆจกๅผไธไธๅคๅถ๏ผ้พๅจ2+1๏ผ่ฃๅๅฟซ ๅฟซ้ไธ็บฟ ๅ็บงๆฐๅ่ฝ๏ผ
๏ผ8๏ผๆจไธ่ฟไธไธๆๅข้ๆบๅถ๏ผไฝๆณขๆฏ๏ผๅฟซ้่ฃๅ๏ผ
๏ผ9๏ผไธไบฉ็ฐๅไผไบบๅ
ฑๅๅฏ่ฃ็ณป็ป๏ผๅไธๆจกๅผๅฏผๅธ-็ๅฒๅข้็ญๅๆจกๅผ๏ผ
๏ผ10๏ผๅพช็ฏ่ดญๆจกๅผ๏ผๆถ่ดน่ฟๅฉใ่ดก็ฎๅผไธๅ็บข๏ผ
๏ผ11๏ผ็ปฟ่ฒ็งฏๅๅขๅผๆจกๅผ๏ผๅ่พนไธๆฌ๏ผๅชๆถจไธ่ท๏ผ
๏ผ12๏ผไปฃ็ๅ้็ณป็ป๏ผๅข้็บงๅทฎ ๅบๅไปฃ็ ๅข้ๅๆถฆ ่ชๅจๅ่ดฆ๏ผ
๏ผ13๏ผๆฐ้ถๅฎๅ้ๅๅ๏ผๆผๅข๏ผ็งฏๅๅๅ๏ผ็งๆ๏ผ็ ไปท๏ผ็งฏๅๅ
ๆข๏ผ้จๅบ่ชๆ็ญ๏ผ
ๆๅๅคๅซๅชไธ็งๆจกๅๆๆๆ่ฃๅๆๅฟซ
ๅชไธ็งๆจกๅๆข่ฝๅฟซ้่ฃๅๅๅ
ทๆๅฏๆ็ปญๆง
ๆๆฒกๆๅฏ่ฝ้่ฟๆดๅไผๅไปฅไธๆจกๅๅๆฐไธ็งๆดๅ ้ๅๅฝๅๅธๅบ็ๆจกๅผ๏ผ
Keep getting inference failures:
{"error":"Failed to perform inference: JWT has expired: "exp" claim timestamp check failed"}
Makes it pretty un-usable as a tool. Above was in the middle of a chat with Kimi-2.6 though Novita, tried other models & providers, all with the same error in the same chat.
Also Playground needs a way to stop an inference in progress if you see that the thinking is wrong so you can correct your query other than just waste tokens on something that you know will not meet your requirements.
Need a way to increase response limits and timeouts as have in many cases been cut-off in output generation which wastes time and money making the endeavour to use playground also useless.
Hi @SenecaInExile sorry for the delay, didnt see your message.
1 - stop button: I'll add one to make it stoppable
2 - regarding your first answer it was a known issue, do you still have the issue ?
Hi @SenecaInExile sorry for the delay, didnt see your message.
1 - stop button: I'll add one to make it stoppable
Thanks, would help.
2 - regarding your first answer it was a known issue, do you still have the issue ?
Yes, seems to be more prevalent if you are repositioning in the 'desktop' while one of the models is working. (i.e. moving around, trying to re-position, etc). don't know if it's a direct correlation but if there is any movement while one or more models is processing the preponderance of errors (truncated responses, timestamp errors, etc) are much more frequent.
Some other items I'm noticing:
3) should have a means to copy the entire entry or response prompt block for that turn. You can grab the entire flow in json format, but would really like the means similar to say lm studio and other tools where you can copy the entire block /turn (including embedded code, etc in that step/block) to memory or to save it in markdown format. Currently have to basically highlight & scroll to copy the block manually and then paste that in another document. Not too user friendly.
should be able to collapse/expand blocks. If you paste in say a handover document or other description document that is long you are scrolling down in playground for page and pages to get to the top/bottom. The quick navigation item on the bottom helps but it's way too small when you have say 500 or even 1000 lines of a document to paste in for review. (turns it into a vertical line). To collapse old steps would be great.
should be able to save or default to what models to use when starting a new chat. (perhaps as a profile setting or just restore last used models) opposed to defaulting to the whatever is 'popular' at the moment. Just an annoyance that I have to remove and then add in the models that I want each time I start a new chat.
I keep getting this in playground: ("error:" Failed to perform inference. JWT has expired:"exp" claim timestamp check failed")
You've got to extend the timeout in playground to allow the user to answer clarifying questions in a chat session. Currently if output is provided the session terminates way too fast (less than a minute). And there is no way to continue the session without starting over from scratch which just wastes thousands of tokens for no reason. Either lengthen the time to allow the user to respond meaningfully, OR allow a continued session by sending current turn history context with a new connection (sub optimal as it still wastes tokens). But the current state of affairs is the worst option.
