← Back to list

Is Kompact AI-IIT Madras’s LLMs in CPU Breakthrough Overstated?

Chidhambararajan R in TheSeriousProgrammer · 2025-04-13 21:23 · 124 claps · 7.2 min read
#ai #llm #deep-learning
Open on Medium ↗
Wiki topics: LLM · Large Language Models ML · Machine Learning AI · AI · General EDU · Education & Learning

Is Kompact AI-IIT Madras’s LLMs in CPU Breakthrough Overstated?

LLMs can run in CPUs, but does it make sense to run them there?

I recently read multiple articles from multiple news outlets which all pointed to a claim which stated that there has been a break through from kompact.ai (a collaboration between IIT Madras and ziroh.com) which allows LLMs run in CPUs instead of GPUs. I was pretty skeptical of it, but it was just filled with nationalism PR, so I decided to get to the bottom of it.

Spoiler: I managed to run LLMs with similar performance in CPUs with just 3hrs of effort

First of all most of those articles are misleading, LLMs can indeed run in CPUs, its just that cost/token doesnt make sense in CPUs when compared to GPUs (more on that later)

Busting their Demo :

The hardware they used was Intel Xeon Silver 4510 CPU (24 cores) (46GB RAM. Hence to see if there was any innovation, I had to run LLMs in a similar hardware and compare the results

Screenshot from the youtube video

Screenshot from the youtube video

In the above screenshot of their demo you can see that, they have achieved an impressive 43 tokens per second without any quantization and on a CPU!!

I didnt have a cpu with 24 cores at home (any normal consumer wont have that kind of hardware at home, it might not even be in the standard deviation (range) of distribution of power of compute of normal people)

Here is what I did:

I procured a VM with Intel Xeon processor with 12 cores in it ( I went for this because this costed only 1.5USD an hour, the 24 core variant costed 3USD an hour, but I can easily compare my results to their results by doubling the figures)

For inference software I went for https://github.com/intel/intel-extension-for-pytorch.git the official LLM inference stack from intel for intel hardware (had to barely make any changes to their example codes)

And I ran the same llm with bfloat16 precision ( Yeah I hear the tech wizards complaining that bfloat16 is not full precision, but in ML the full precision is barely utilized by the model, often switching to half precision doesnt impact the model in any way at all. Hell, even models arent trained in fp32 these days)

The results (drum rolls please!!) :

I got a speed of 22.5 tokens/second but my hardware has only half the cores as the IIT demo , so for a fair comparison, I double my result which is ~44 tokens/second, which is extremely close to the IIT demo of 43 tokens per second (Its essentially equal for a rough comparison)

And how much time did i invest in this entire flow?

A bit of R&D for inference framework: 45 m Time To create AWS Instance: 04m 12.3s First Successful LLM inference run: 37m 1.8s Time to build interface with benchmark code: 1hr

It just took me 3 hours !!!! Just 3 hours to recreate the same performance, so did they actually do any innovation on their own?

The other demos where run in an extremely expensive hardware to for me to replicate as a passion project ( USD-INR expenses for hobby projects are too much)

Here are the different possible scenarios of what might have happened in their side:

Scenario 1:

Its very likely that they built their system around intel’s inference software and called it a day!!

If this is the case, it is very concerning!! Let us say I declare one day that I created a revolutionary compression algorithm and created a service for the same, but upon inspection you realize that its just WinRAR but wrapped up, every one in the tech community will just beat me up (The problem here is claiming that I created something new, but just using what someone has already made in the market behind the screens. Nothing wrong , but its not innovation)

In this case, I might have let it pass if it had been made by students, but this is made in collaboration with a 2 year old AI company ziroh.com and its also endorsed by the chairs of IIT Madras (the chairs of IIT Madras dont easily endorse something without validation, there might be more to the story here!!)

So many media outlets and influencers are milking this as an indian innovation (I hate people exploiting nationalism) which is clearly not an innovation in this scenario.

Scenario 2(:

Its possible that they developed their own inference framework (but what’s the point here? there is already one in market which possibly out performs you)

Even in this scenario, we cant let it pass, especially with the program chairs heavily endorsing the same!! (It’s puzzling that the project’s novelty wasn’t compared to existing frameworks like Intel’s, which are widely known.”!!). IIT Madras is very well known for its high quality research

So why am I pissed ?

In both the above scenarios the claimed breakthrough is not a break through at all!! Many influencers and media outlets are blindly milking the same without any R&D

Eventually companies who attempt to partner with IIT Madras to use this tech, will get to know the same. This risks undermining confidence in Indian DeepTech if the breakthrough doesn’t hold up. Companies and student groups which perform legit deep tech research will also be impacted because of the same

Remember, they didnt announce a stack (which incorporates existing frameworks and makes the job of companies easier), they announced a fundamental breakthrough, so I am not here to talk about the possible of ease of use by Kai VM, just criticizing their claimed breakthrough

What can IIT Madras do?

There is nothing wrong in announcing breakthroughs, but it will be good if you release a solid technical report on the details analysis which showcase that their system is superior (Given IIT Madras’s research legacy, clearer documentation would be expected)

So what ? LLMs run on CPUs!! Its still a win!!

Nope, the basic hardware which they used was a intel xeon silver 4515 which costs 630USD in intel website, and it gave 43 tokens/second for a 1.5b llm. RTX 2060 GPU can the run the same in similar speeds and it costs only 200 USD. We run LLMs and AI Models in GPUs because, GPUs are the most cost effective way to run them and that’s because GPUs are funamentally designed to do parallel computations unlike CPUs and LLMs are just filled with parallel computation. Not because we cant run them on other types of hardware.

Using CPUs for LLMs instead of GPUS in PROD

Using CPUs for LLMs instead of GPUS in PROD

Some people might claim that CPUs are more accessible in rural areas, but the perfomance of non industrial ones are not good enough to be reliable. Consumer grade CPUs can only run highly quantized variants of the LLMs which often times hallucinates a lot

Conclusion

Kompact AI’s claims of a CPU-based LLM breakthrough sound groundbreaking, but my 3-hour experiment suggests otherwise. Matching their 43 tokens/second with existing tools raises doubts about their “new AI engine.” While niche optimizations are possible, their focus on costly CPUs over GPUs feels impractical. To silence skeptics like me, Kompact AI should open-source their “Kai VM” or share a whitepaper. Until then, this “innovation” risks overselling India’s DeepTech potential, which deserves better.

What I can do

If anyone from IIT Madras reads this, you have the reputation of the one of the best institutes in the world, kindly release a report!! I will certainly with full heart do a revision and make sure to put it in the top of the article, if any of my claims are proved to be invalid by IIT Madras or Ziroh.com

Statements from IIT Madras Heads in the video:

00:09:56,959 → 00:10:03,040 (Prof. Kamakoti): …this is a product that we are conceived and designed by zero labs and fine-tuned and validated by IIT Madras…

00:27:20,720 → 00:28:06,640 (Prof. Kamakoti): …this is extremely deep tech to understand this… we have to know the instruction set architecture… micro architecture… cache… operating system… threading… scheduling… compilation… effective database… this is very very deep tech very very involved and that’s going that’s the real big contribution here…

00:38:17,160 → 00:38:27,599 (Prof. Sadagopan): …watching at close quarters zero labs and rishies the creators of compact AI from day minus one…

01:05:19,039 → 01:05:28,440 (Mr. Aj Goyel): …these brilliant engineers including Rishi Kesh Indian scientists and engineers who have invented real cutting edge technologies…

01:06:49,559 → 01:07:08,400 (Mr. Aj Goyel): …the second big invention which has happened from Zerolab’s team is this new AI computing engine. This is truly a deep tech inventions… acknowledge the kind of inventions it has happened and largely done by Indian engineers.

01:07:33,720 → 01:08:12,599 (Mr. Aj Goyel): compact AI… it’s a completely brand new computing AI engine… built grounds up with a complete new mathematical approach… new technology innovations in the computing science… truly a cuttingedge work done to make CPUs perform at a highest level…

01:17:45,120 → 01:17:47,840 (Mr. Rushikesh): …this is one of the thing that compact AI has solved.

01:33:59,520 → 01:34:15,320 (Prof. Madhusudanan): …what caught my attention was their algorithmic approach… the novelty of their approach which is unique. Unlike conventional inferencing pipelines the solution aimed to bridge gap between the hardware and software…

01:48:40,080 → 01:49:05,670 (Mr. Rushikesh): …one of the core invention in today’s compact AI is a new VM for AI it’s a Kai VM… it’s it’s a very new way of uh doing things… it’s a very very big inventions in that space as well as a new mathematics underlying underneath…


메타데이터
post_id
60027c13ea53
slug
is-kompact-ai-iit-madrass-llms-in-cpu-breakthrough-overstated-60027c13ea53
url
https://blogs.chidha.dev/is-kompact-ai-iit-madrass-llms-in-cpu-breakthrough-overstated-60027c13ea53
canonical_url
https://blogs.chidha.dev/is-kompact-ai-iit-madrass-llms-in-cpu-breakthrough-overstated-60027c13ea53
author_url
https://medium.com/@chidhambararajan
status
ok
fetched_at
2026-06-14 16:15:44