Saturday, August 1, 2026

Engy.ai Kimi K3 Inference Release on Bittensor SN53

Futuristic AI core connected to a network of glowing nodes on a subtle grid, symbolizing decentralized inference on SN53.

Engy.ai Kimi K3 Inference Release on Bittensor SN53

Engy.ai is preparing to add Kimi K3 and GLM-5.2 inference to Bittensor Subnet 53, expanding the network’s catalogue of large open-weight AI models. Mining access had not yet opened when Engy.ai announced the integration, with activation expected to follow shortly. Bittensor identifies SN53 as a verified-inference subnet where miners serve committed models and validators check their outputs.

The rollout would allow miners to process requests for both models and compete for subnet rewards. The update remains an infrastructure launch in progress rather than a completed production deployment, and Engy.ai has not published a fixed activation time or capacity target.

Verified Inference Sits at the Center of SN53

Engy.ai is designed to prove that the requested model, rather than a cheaper substitute, produced each response. Its repository says model weights and quantization settings are pinned through published Merkle roots, while outputs carry activation fingerprints used during verification. The system aims to connect distributed GPU providers with auditable model execution, although parts of its broader audit mechanism are still being rolled out.

The team said it ran the full Kimi K3 model on 80 RTX 5090 GPUs linked through standard Ethernet, initially reaching 20 tokens per second for a single stream. It also claimed that GLM-5.2 throughput rose from 30 to 110 tokens per second on the same fleet after optimization. Those figures are project-reported engineering results and have not been independently benchmarked.

Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with 104 billion activated parameters and a context window of up to 1 million tokens. Its Stable LatentMoE design activates 16 of 896 routed experts per token, while Kimi Delta Attention and Attention Residuals support long-sequence processing. This differs from descriptions attributing the model to Gated Multi-Head Latent Attention.

Open Weights Expand the Subnet’s Model Options

Moonshot AI has released Kimi K3’s full weights, and vLLM has added deployment support for the architecture. SN53 is therefore preparing to serve an open-weight model rather than merely forwarding requests to Moonshot’s hosted API, although operating it still requires substantial hardware and distributed-inference engineering.

GLM-5.2 is also an open-weight model built for long-horizon coding and agent tasks, with Z.ai documenting a 1-million-token context window. NIST’s Center for AI Standards and Innovation has separately evaluated both models in preliminary cyber-capability testing. That independent assessment confirms their technical availability, not Engy.ai’s projected pricing or performance.

Bittensor-focused investor Mark Jeffrey suggested that lower-cost Kimi K3 access could arrive within roughly one day, but the estimate was not an official deadline. Claims that SN53 will materially undercut centralized providers remain prospective until Engy.ai publishes live pricing, throughput and reliability data.

Once mining opens, miners will need to meet Engy.ai’s model and proof requirements, while validators assess execution and set weights on-chain. The practical test will be consistent output quality and uptime across distributed hardware, not simply whether the subnet can load the models. For now, support for Kimi K3 and GLM-5.2 remains announced but pending final miner activation.

The status of the launch remains pending final activation for miners. Once live, the integration will allow validators and miners on SN53 to facilitate inference requests for these specific models, testing the subnet’s ability to maintain performance parity with centralized providers while reducing overall operational expenses.

Shatoshi Pick
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.