Skip to content
context Other

TileRT: a persistent kernel accelerates low-latency LLM inference on NVIDIA GPUs

only one source so far

SemiAnalysis describes TileRT – an engine that compiles the decoding graph into a single persistent kernel on NVIDIA B200 GPUs to achieve 500 tokens/s/user, about 3× more than traditional engines on GB300 NVL72. Already deployed in production at Xiaomi and ZAI.

In its analysis, SemiAnalysis describes TileRT, an inference engine that statically compiles the entire decoding graph into a single so-called persistent kernel on NVIDIA GPUs, eliminating the overhead associated with repeatedly launching and synchronizing many individual kernels. According to the article, this overhead is the main reason why actual inference speed on GPUs falls far short of the theoretical limit determined by HBM memory bandwidth – for the GLM-5 model in NVFP4, the theoretical roofline on a B200 server would allow up to 3 047 tokens/s/user, but in practice GPUs come nowhere near this figure.

On the InferenceX GLM5 FP8 744B benchmark, TileRT achieved a verified 500 tokens/s/user on a single B200 decode server, according to SemiAnalysis, which is about 3× faster than GB300 NVL72 with traditional inference engines. At the same cost per output token (iso-cost), TileRT is said to deliver up to 2× greater interactivity than traditional engines. TileRT was created by the same developer community behind the DSL tool TileLang.

The engine integrates into the existing ecosystem through PD disaggregation: TileRT handles latency-sensitive decoding, while throughput-focused engines such as vLLM and SGLang continue to handle the prefill phase. According to the article, the TileRT decode engine is already deployed in production at Xiaomi for the MiMo V2.5 Pro UltraSpeed model and at ZAI for GLM 5.1 HighSpeed.

The rest of the article, which according to the introduction was intended to further explore the comparison of TileRT with specialized low-latency chips such as NVIDIA Groq LPU, Cerebras or SambaNova, is not available. You can find details in the source article.

What changed

Why it matters

According to SemiAnalysis, users are willing to pay more for lower latency and faster token generation, which increases gross margins for providers. TileRT shows that acceleration similar to what specialized chips (Groq, Cerebras, SambaNova) have offered so far can also be achieved through software optimization on standard NVIDIA GPUs, with production deployments already in place at Xiaomi and ZAI.

Relevant practical impact

What this means

01

For a business

According to SemiAnalysis, companies running LLM inference can use TileRT to reduce response latency on NVIDIA GPUs without having to invest in specialized low-latency chips, which could support premium "fast mode" offerings with higher margins; deployments at Xiaomi and ZAI demonstrate real production use.

Development
What to decide If a company runs LLM inference with a focus on latency, consider evaluating TileRT as an alternative to purchasing specialized chips (Groq, Cerebras).
More business impacts →
GPU inference Latency NVIDIA TileRT

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
SemiAnalysis (newsletter feed — hardware, chips, AI economics) independent context · first detected Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX