started · updated
Ninefold Intelligence enhances AI inference with Alaya-DSpark and GLM-5.3 integration
Ninefold Intelligence has expanded its AI infrastructure capabilities through new optimizations and model integrations. The company released Alaya-DSpark, an open-source speculative decoding acceleration tool specifically optimized for the GLM-5.2 series. By utilizing a “draft-first, main model verification” approach, Alaya-DSpark has demonstrated an increase in inference throughput from approximately 121 tokens per second to 352 tokens per second, representing a 2.9x speedup while maintaining output consistency.
Additionally, Ninefold Intelligence’s Alaya Token platform has completed deep adaptation for Zhipu AI’s next-generation GLM-5.3 model. This integration allows for standardized, scalable delivery of the flagship model’s capabilities, including improved coding and agentic task performance. The platform provides optimized support for long-context features and high-concurrency inference.
Separately, industry observations suggest Zhipu AI may be testing new models under the name Omen Alpha on the OpenCode platform. Clues from the service indicate Omen Alpha supports a 500,000-token context window and multimodal inputs, potentially representing a new branch of the GLM model series.
Entities
Alaya Token · Alaya-DSpark · GLM-5.3 · Ninefold Intelligence · Zhipu AI