Startup Kog is developing software to accelerate LLM inference on standard GPUs
French startup Kog is developing Kog Inference Engine, software intended to accelerate LLM inference on standard GPUs without new hardware. A demo with the small model Laneformer 2B achieved 3000 tokens/s; the target is a 30-fold speedup for larger models.
French startup Kog is developing Kog Inference Engine (KIE), software intended to accelerate large language model inference through software optimization on standard data center GPUs, such as AMD MI300X and NVIDIA H200, without the need for new specialized hardware. In May 2026, the company released a technical preview that it says showed that “extremely fast decoding of individual requests is possible on standard data center GPUs that companies already own”. According to CEO Gaël Delalleau, this earned the company 200 concrete business leads.
The demo from Kog achieved 3000 tokens per second for a single request, but used a small, purpose-built model with roughly 2 billion parameters called Laneformer 2B, which the company subsequently open-sourced. However, the company aims for a 30-fold inference speedup for larger models — according to Delalleau, since launch the company has therefore focused entirely on accelerating development for larger models, because customers are unwilling to fine-tune small models for their own needs. Delalleau expects the first use case to be software engineering, where users of tools such as Claude Code often wait many minutes or even hours for results; according to the article, Anthropic charges for speed through a surcharge for Fast Mode in the Claude model. Kog also has partners focused on generating games and applications from prompts, where faster results would mean higher revenue.
According to Delalleau, the company builds on an in-depth, hands-on engineering approach to each GPU type — it devotes weeks to months of low-level research to each new chip, which, with a team of 11 people, limits the number of chips it can support. The company compares its approach to the Hazy Research lab at Stanford University and distinguishes it from that of French company ZML, which bypasses CUDA from Nvidia in a different way. Kog is backed by French institutions Bpifrance and the French Tech 2030 program and works with Scaleway; according to Delalleau, the company plans to achieve a 10-fold speedup on its first large model in September 2026, which should serve as a basis for a Series A funding round.
Why it matters
If the software approach from Kog also proves effective for large models, it could reduce costs and waiting times in AI workflows without the need to invest in specialized chips such as Cerebras — which is relevant mainly for companies running or making intensive use of LLM inference, for example in software development or generating applications from prompts. For now, however, this is an unconfirmed target (30x speedup) based on a demo with a small 2-billion-parameter model, not on real-world deployment with large models.
Relevant practical impact
What this means
For a business
If the target of a 30x speedup is also confirmed for larger models, it could reduce compute costs and waiting times in AI workflows for companies running inference on standard GPUs, without the need to buy specialized hardware.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.