Alphabet Inc Class CGoogle Gemini is delivered on premises through Google Distributed Cloud and operated by Cirrascale, extending Gemini's enterprise reach.
Cirrascale Cloud Services announced the production release of the Cirrascale Inference Platform, a complete software stack for enterprise-grade AI inference, at the AI Infra Summit in Santa Clara. The platform lets enterprises run open-source models, their own private models, and closed model ecosystems from a single serverless platform, including Google Gemini delivered on premises through Google Distributed Cloud and operated by Cirrascale. Its model and hardware selection layer automatically routes each request to the right model and runs it on the best available accelerator across NVIDIA, AMD, Tenstorrent, and Qualcomm with no code changes required to switch, and teams can fine-tune models on their own private data without that data leaving their environment. The platform includes a turnkey private chat experience connected to a company knowledge base, built-in controls to manage AI spend across teams, and governance guardrails for agentic workloads, aligned with HIPAA, SOC 2, and FedRAMP requirements where required. The Cirrascale Inference Platform is available now across Cirrascale's U.S. and international regions.
Alphabet Inc Class CGoogle Gemini is delivered on premises through Google Distributed Cloud and operated by Cirrascale, extending Gemini's enterprise reach.
Advanced Micro Devices IncCirrascale's inference platform routes workloads across AMD accelerators, expanding demand for AMD AI hardware.
NVIDIA CorporationThe platform runs inference on NVIDIA accelerators as one of its supported hardware options, supporting NVIDIA AI demand.
Qualcomm IncorporatedQualcomm accelerators are among the hardware options the platform can route inference workloads to.
CTO Realty Growth IncCirrascale launched the production release of its enterprise AI inference platform, a new product offering spanning multiple accelerators and model ecosystems.
Tenstorrent accelerators are included in the platform's hardware selection layer for running inference.