AIToday
Large Language ModelsAI Business & IndustryTechCrunch AIPublished: Aug 15, 2026, 01:00 JST5 min read

French startup Kog claims 30x faster AI inference on standard GPUs

French startup Kog claims 30x faster AI inference on standard GPUs

Key takeaway

  • Kog, a French startup, is demonstrating that standard datacenter GPUs can deliver drastically faster AI inference through specialized software optimization rather than new hardware.

  • The company's demo showed 3,000 tokens per second on a small model and claims its method can achieve 30x faster inference on larger language models by unlocking GPU capabilities that existing tools don't fully utilize.

  • With 200 early business leads and a target of proving the approach on a major model by September, Kog is positioning itself to address a critical pain point—inference speed—for enterprises already investing in conventional chips.

3 Key Points

  1. What happened

    Kog, a French startup founded by Gaël Delalleau, is using deep GPU-level software optimization to accelerate AI inference on conventional datacenter GPUs like the AMD MI300X and Nvidia H200. The company's tech preview in May demonstrated 3,000 tokens per second on a small 2 billion parameter model (Laneformer 2B), and claims its approach can achieve 30x faster LLM inference by unlocking existing hardware capabilities that vendors don't fully expose.

  2. Why it matters

    Inference speed has become a critical bottleneck for AI deployments, and Kog's promise to unlock performance gains through software alone—without requiring new chips—appeals to enterprises already owning standard GPUs. The startup has accumulated 200 business leads and is targeting software engineering workflows, where tools like Claude Code sometimes force users to wait hours for results; Anthropic itself charges a price premium for faster processing, signaling the market values speed. For developers and businesses relying on AI for professional tasks, faster inference could reduce both costs and operational friction.

  3. What to watch

    Delalleau stated Kog aims to demonstrate its approach on a "first major model at 10x speed" by September, a milestone he flagged as essential to securing Series A funding. The company's team of 11 is limited by the manual, months-long effort required to optimize each new GPU; it plans to develop agent-based automation in the longer term to support more chips. The startup is backed by Scaleway, Bpifrance, and France's Tech 2030 program, positioning it within Europe's sovereign AI capability efforts.

In Depth

Read the full story

Kog, founded by Gaël Delalleau, is tackling one of AI infrastructure's most pressing challenges: the speed of inference—the computational step where an AI model produces an answer to a user prompt. While competitors like Cerebras have pursued custom silicon to accelerate inference, Kog is betting that conventional GPUs already in enterprise datacenters contain untapped performance waiting to be unlocked through software alone. In May, the startup published a tech preview on Hacker News demonstrating extremely fast single-request decoding on standard datacenter GPUs, specifically the AMD MI300X and Nvidia H200, achieving 3,000 tokens per second on a small open-sourced model called Laneformer 2B with 2 billion parameters. The company claims its approach can deliver 30x faster LLM inference overall, though proving this on larger models remains pending. The early response exceeded expectations: Kog accumulated 200 tangible business leads from the preview. Delalleau indicated that software engineering is the most immediate use case, as tools like Claude Code sometimes force users to wait hours for results. Anthropic itself has monetized speed by charging a price multiple for Claude's Fast Mode, validating that customers will pay for faster inference. Kog is also in talks with design partners enabling users to generate games and apps from prompts, where reduced latency would translate directly into more revenue per session. However, Kog faces a technical leap: its demo worked on a purpose-built small model, whereas enterprise demand is skewed toward large language models. Delalleau is confident the same optimization techniques will apply to larger systems, arguing that the misconception—that GPUs are poorly suited for decoding—stems from a failure to exploit newer GPUs' growing memory bandwidth. Delalleau himself studied solid-state physics at France's École Polytechnique and spent years in offensive cybersecurity (white-hat hacking), participating in DEFCON's CTF tournament four times. This background shaped his team's methodology: on the science side, understanding GPU physics to maximize efficiency; on the hacking side, reverse-engineering hardware down to assembly language and binary to repurpose it for unintended goals. The downside is labor-intensive: Kog dedicates several weeks or even months of GPU engineering research for each new chip it supports. With only 11 employees, this limits the chips the startup can target in the near term, though Delalleau hopes to automate the process via agent-based pipelines over time. For now, Kog is focused on delivering its first major model optimization at 10x speed by September, a milestone Delalleau flagged as necessary to demonstrate customer traction and secure Series A funding. The startup is backed by Scaleway, Bpifrance, and France's Tech 2030 program, positioning it within Europe's broader push for sovereign AI capability—a potential tailwind as the continent seeks independence from U.S. chip and software vendors.

Context & Analysis

Kog's entry into the GPU inference optimization market reflects a broader shift in AI infrastructure priorities: with chips like Nvidia's GPUs already deployed across enterprises, the next frontier for performance gains lies in software. The startup's claim of 30x speedup is ambitious, but grounded in a specific observation—that modern GPUs ship with memory bandwidth capabilities that existing software (including Nvidia's CUDA framework) does not fully exploit. Delalleau's background in solid-state physics and reverse-engineering (via competitive hacking) positions him to pursue this "low-level" optimization work, drawing parallels to Hazy Research at Stanford. However, the approach has a clear bottleneck: the startup's team of 11 requires months per GPU to conduct the detailed engineering work needed to unlock new performance, limiting scalability. Delalleau acknowledges that the market for fine-tuned small models is not yet mature; instead, Kog is targeting the use cases where inference latency directly hits revenue or user experience—software engineering workflows, game and app generation, and professional AI tools where customers currently accept multi-hour wait times. The 200 tangible business leads accumulated since the May launch suggest real demand, though the move from a 2 billion parameter demo to proven gains on large language models remains the critical test.

FAQ

Which GPUs does Kog's software work with?
Kog's tech preview demonstrated its approach on the AMD MI300X and Nvidia H200 datacenter GPUs. The company notes it does not currently extend support to GPUs in laptops.
What is Kog's main claim about speed improvement?
Kog claims its approach can achieve 30x faster LLM inference. Its demo showed 3,000 tokens per second using a small 2 billion parameter model called Laneformer 2B.
When does Kog expect to prove its approach works on large language models?
CEO Gaël Delalleau said he expects to implement the approach on a "first major model at 10x speed" by September, after which the company plans to demonstrate customer traction and raise a Series A.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAWS shows how to blend SageMaker and Bedrock models in one agent workflow

The AI news that matters, in one minute each morning.

Sign up free