Deploy and run open-source models directly from the Hugging Face Hub using leading serverless cloud inference providers.
### Key Features
– **Unified Provider Access**: Seamlessly run Hub models using premier serverless and dedicated compute backends directly from the Hugging Face interface.
– **OpenAI-Compatible APIs**: Switch between inference backends with minimal code changes using standardized schema outputs.
– **Optimized Latency**: Leverage optimized inference engines configured by specialized hardware partners for instant scaling.
### Use Cases
– Production deployment of highly-demanding open-source architectures, such as massive Mixture of Experts (MoEs), without managing physical GPU clusters.
### Developer Pros & Cons
– **Pro:** Instant scaling and benchmarking of different hardware backends with zero infrastructure setup.
– **Con:** Dependency on Hugging Face’s Hub routing layer for initial setup and credential mapping.