started · updated
Gimlet Labs and Cerebras partner for ultrafast AI inference
Gimlet Labs and Cerebras Systems have announced a collaboration to provide high-speed AI inference through the Gimlet Cloud. By integrating Cerebras’ wafer-scale compute with Gimlet’s disaggregated inference cloud, the partnership aims to deliver speeds of up to 3,000 tokens per second.
The technology is designed for demanding real-time and agentic applications, such as voice and video AI, where low latency is critical for user experience. The Gimlet Cloud will combine the Cerebras Wafer Scale Engine with GPUs, using advanced orchestration to run different phases of inference on the most suitable silicon. The first Cerebras-powered Gimlet Cloud datacenter is expected to be operational later this year.