Back to Catalog
Create Endpoint
thinkingmachines

Inkling-Small-NVFP4

Catalog model officially supported by Inference Endpoints.

This model is from our Model Catalog, and comes with an optimized configuration. Deployment has been verified by Hugging Face.

/
$5.50 / h
per running replica

Contact us if you'd like to request a custom solution or instance type.

Nvidia RTX PRO 6000 Blackwell
2x GPUs · 192 GB 46x vCPUs · 512 GB
$5.5 / h
available
Hardware should be compatible with the selected model.
  • Only you can access your endpoint, using a Hugging Face Token generated from your personal account.
Number of replicas
Automatically scale the number of replicas within Min and Max based on compute usage. Min is always 0 if Scale-To-Zero is active.
More options
Autoscaling Strategy
Control what type of trigger will cause your Endpoint to scale up.
Default Env
Environment variables that will be provided to your container during deployment.
Secret Env
Same as Default, but people with access to this endpoint will not be able to read these values after creation.
VPC Config
Check to activate and configure AWS PrivateLink