While Cerebras builds custom AI chips, French startup Kog squeezes performance out of the GPUs enterprises already own. By rewriting low-level software for existing datacenter hardware, Kog eliminates the latency bottlenecks that slow AI workflows.
Challenging the Necessity of Custom AI Silicon
The AI industry has pivoted to purpose-built chips for inference. Kog's CEO, Gaël Delalleau, says the belief that standard GPUs can’t handle decoding is wrong. Modern GPUs—Nvidia H200, AMD MI300X—offer memory bandwidth that current software leaves idle.
Kog’s Kog Inference Engine (KIE) targets "extremely fast single-request decoding," a must-have for real-time apps. In a recent demo the company hit 3,000 tokens per second (TPS) with Laneformer 2B, an open-source 2-billion-parameter model tuned for their stack. The demo uses a small model, but Kog plans to extend those gains to much larger LLMs.
A "Hacker" Approach to GPU Engineering
Unlike hardware-agnostic providers such as ZML, Kog digs deep. The team reverse-engineers GPUs at the assembly and binary levels, treating the silicon as a set of physical laws to master.
That low-level focus extracts every ounce of efficiency, but it costs time. With a lean crew of 11 engineers, Kog spends weeks or months dissecting each new GPU architecture before adding support. The method delivers top-tier performance but limits how quickly Kog can cover new hardware.
Targeting High-Value AI Use Cases
Kog focuses on sectors where latency equals lost revenue:
- Software engineering: Tools like Claude Code stall for minutes. Kog aims to turn "waiting hours" into near-instant results.
- Generative app/game design: Prompt-to-app pipelines need rapid iteration to keep users engaged and revenue flowing.
The company’s next milestone is a 10× speedup on its first major large-scale model. Backed by Scaleway, Bpifrance and the French Tech 2030 program, Kog positions itself as a key player in Europe’s AI sovereignty drive.
Key Takeaways
- Optimization over hardware: Kog pushes the memory bandwidth of Nvidia H200 and AMD MI300X GPUs to the limit with extreme low-level software engineering.
- Breaking the latency barrier: The startup targets a 30× boost in LLM inference speed for real-time professional workflows such as automated coding and generative design.
- Deep-level engineering: By adopting a "hacker" mindset, Kog reverse-engineers GPU assembly and binary code to achieve performance that standard stacks cannot reach.
Kog, a French startup, says its Kog Inference Engine aims to run large-language-model (LLM) decoding up to 30 times faster on the same Nvidia H200 or AMD MI300X GPUs that data-center operators already own.
Why the push for speed matters now
LLM inference now throttles products like code-completion assistants, on-the-fly content generators, and interactive game-design tools. In those settings, the gap between a user’s prompt and the model’s reply directly impacts productivity and revenue. High-level APIs such as CUDA leave a large slice of GPU memory bandwidth idle, especially during token-by-token decoding, which dominates real-time use cases.
Kog’s low-level answer
Kog skips conventional software layers and talks straight to the silicon. Its engineers reverse-engineer the GPU at the assembly level, treating the hardware as a set of constraints to obey rather than a black box to abstract. The result is a custom execution path that keeps data flowing through the GPU’s memory pipes at near-full capacity.
In a recent demo the company ran Laneformer 2B—a 2-billion-parameter open-source model tuned for its stack—and reported a high token throughput. The model is modest compared with larger commercial systems, but the speed gain shows what a purpose-built software stack can extract from the same hardware.
The engineering trade-off
The upside comes with a steep cost curve. Kog’s 11-engineer team spends weeks or months dissecting each new GPU architecture before it can be supported. That depth yields high performance on the targeted cards, but it also means Kog cannot instantly add support for every new GPU that hits the market. Customers benefit only if their fleets include the Nvidia H200 or AMD MI300X GPUs Kog has already optimized.
Who stands to gain
- மென்பொருள் பொறியியல் கருவிகள் (Software-engineering tools) – Claude Code போன்ற குறியீடுகளை உருவாக்கும் தயாரிப்புகள், மாடல் ஒரு கோரிக்கையைச் செயலாக்கும் போது பெரும்பாலும் பல நிமிடங்கள் முடங்கிப் போகின்றன. அந்த காத்திருப்பு நேரத்தைக் குறைத்து சில வினாடிகளாக மாற்றினால், ஊடாடும் மேம்பாடு (interactive development) மிகவும் சீராக இருக்கும்.
- உருவாக்கத் திறன் கொண்ட வடிவமைப்பு வழிமுறைகள் (Generative design pipelines) – உரைத் தூண்டுதல்களை (textual prompts) செயலி முன்மாதிரிகளாகவோ (app prototypes) அல்லது விளையாட்டுச் சொத்துக்களாகவோ (game assets) மாற்றும் ஸ்டுடியோக்களுக்கு, படைப்பாளர்கள் தொடர்ந்து ஈடுபாட்டுடன் இருக்க விரைவான மறுசெயல்முறை (rapid iteration) தேவைப்படுகிறது. வேகமான டீகோடிங் (Faster decoding), வடிவமைப்புச் சுழற்சிகளைக் குறைப்பதோடு மாற்ற விகிதங்களையும் (conversion rates) அதிகரிக்கிறது.
இந்த இரண்டு துறைகளிலும், தாமதம் (latency) என்பது நேரடியாகப் பணிக்கான நேர இழப்பு அல்லது வாடிக்கையாளர் வெளியேற்றமாக (churn) மாறுகிறது, எனவே 30 மடங்கு வேக அதிகரிப்பு ஒரு தீர்மானிக்கும் போட்டித்திறனாக அமையும்.
விரிவான பார்வை
தனிப்பயனாக்கப்பட்ட AI சிலிக்கான்களை (custom AI silicon) உருவாக்கும் தொழில்துறைப் போக்குக்கு மாறாக Kog-ன் உத்தி உள்ளது. Cerebras போன்ற நிறுவனங்கள், வாட் (watt) ஒன்றுக்கு அதிகத் திறனை (higher throughput) வழங்கும் சிப்களுக்காகப் பில்லியன் கணக்கான டாலர்களைச் செலவிடுகின்றன. இன்றைய GPU-க்களில் டீகோடிங்கிற்குத் தேவையான போதுமான மெமரி பேண்ட்வித் (memory bandwidth) ஏற்கனவே உள்ளது; ஆனால் அதைச் சரியாகப் பயன்படுத்தக்கூடிய மென்பொருளே தற்போது இல்லை என்று Kog வாதிடுகிறது. இந்தக் கூற்று பெரிய அளவில் உண்மையாக இருக்கும்பட்சத்தில், டெவலப்பர்கள் விலையுயர்ந்த வன்பொருள் மேம்படுத்தல்களைத் தள்ளிப்போட்டுவிட்டு, நிகழ்நேரத்திற்கு நெருக்கமான அனுமானத்தை (near-real-time inference) அடைய மென்பொருள் மேம்படுத்தலையே நம்பியிருக்கலாம்.
சாத்தியமான சவால்கள்
- பராமரிப்புச் சுமை (Maintenance burden) – ஒவ்வொரு புதிய GPU தலைப்பிற்கும் புதிய ரிவர்ஸ்-இன்ஜினியரிங் (reverse-engineering) தேவைப்படும். சந்தை விரிவடையும் போது, சிறிய குழுவினர் அதிக பணிச்சுமையால் திணறக்கூடும்.
- பெரிய மாடல்களுக்கான அளவிடுதல் திறன் (Scalability to larger models) – இந்த டெமோவில் 2 பில்லியன் அளவுருக்கள் கொண்ட மாடல் (2B-parameter model) பயன்படுத்தப்பட்டது. இதே நுட்பங்களை மிகப் பெரிய அமைப்புகளுக்கு விரிவுபடுத்துவது மெமரித் திறன் வரம்புகளை எட்டக்கூடும் அல்லது இன்னும் வெளிப்படுத்தப்படாத கூடுதல் பொறியியல் நுணுக்கங்களைக் கோரக்கூடும்.
- மாற்றுப் பாதைகள் (Alternative paths) – கிளவுட் வழங்குநர்கள் ஏற்கனவே சிறப்பு சிப்கள் மற்றும் மென்பொருள்களை உள்ளடக்கிய இன்ஃபரன்ஸ்-மேம்படுத்தப்பட்ட இன்ஸ்டன்ஸ்களை (inference-optimized instances) வழங்குகிறார்கள். சில பணிச்சுவைகளுக்கு, Kog-ன் தொழில்நுட்பத்திலிருந்து கிடைக்கும் சிறிய லாபம், ஒரு மேலாண்மை செய்யப்பட்ட சேவையின் (managed service) வசதியை விட அதிகமாக இருக்காது.
கவனிக்க வேண்டியவை
- முதல் பெரிய மாடல் வெளியீடு (First large-model rollout) – Kog தனது அடுத்த மைல்கல் ஒரு "முக்கியமான பெரிய அளவிலான மாடலில்" 10 மடங்கு வேகத்தை அதிகரிப்பதாகும் என்று கூறுகிறது. 2B டெமோவிற்கும் ஒரு நிறுவனத் தர மாடலுக்கும் (enterprise-grade model) இடையிலான இடைவெளியே இந்த அணுகுமுறையின் தெளிவான சோதனையாக இருக்கும்.
- வன்பொருள் ஆதரவு விரிவாக்கம் (Hardware support expansion) – புதிய GPU-கள் அல்லது பிற முடுக்கிகளுக்கான (accelerators) ஆதரவைச் சேர்ப்பது, இந்த குறைந்த நிலை மாடல் (low-level model) விரைவான வன்பொருள் மாற்றங்களுடன் (hardware cadence) இணைந்து செயல்படுகிறதா என்பதைத் தெரிவிக்கும்.
முடிவுரை
மென்பொருள் கட்டமைப்பை (software stack) ஆரம்பத்திலிருந்து மீண்டும் எழுதும்போது, தற்போதுள்ள GPU-க்களிலிருந்து அதிகப்படியான வேலையைச் சாத்தியமாக்க முடியும் என்பதை Kog காட்டுகிறது. 30 மடங்கு வேகமான LLM டீகோடிங் என்பது டெவலப்பர்கள் AI சேவைகளுக்கான விலையை நிர்ணயிப்பதையும் கட்டமைப்பதையும் மாற்றியமைக்கக்கூடும்—ஒவ்வொரு புதிய சிலிக்கான் துண்டிற்கும் தேவைப்படும் தீவிர பொறியியல் முயற்சியை நிறுவனத்தால் தொடர்ந்து செய்ய முடிந்தால் மட்டுமே இது சாத்தியம்.
