← Back to list

Your AI Has Only One Sense. Your Customers Have Five. Missing the leverage?

Think about the last time you really understood something important about a person.

Technology Innovation Hub, IIT Mandi · 2026-08-10 10:35 · 5 claps · 5.4 min read
#multimodal-ai #finance #incubation-program #ai #collaboration
Open on Medium ↗
Wiki topics: RAG · RAG & Retrieval MM · Multimodal & Generative Media AI · AI · General ECO · Economy · General 📊 · Economic Policy

Your AI Has Only One Sense. Your Customers Have Five. Missing the leverage?

Think about the last time you really understood something important about a person.

You didn’t do it from their words alone. You read their tone. You watched their face. You noticed the pause before they answered, the way they held themselves, the things they left unsaid but you still “heard”. You took in several signals at once, and the meaning lived in how those signals fit together.

Now think about how most AI works today. It reads one thing at a time. A chatbot reads text. A fraud model reads transaction data. A verification system reads a document. Each one is clever in its own lane, but each is working with one sense switched on and the rest switched off.

That gap is where the next wave of AI is being built. It is called Multimodal AI. And if you are a founder building anything that touches financial services, it may already describe what you are doing, even if you have never used the term.

What “multimodal” actually means

Multimodal AI is simply AI that takes in and reasons across more than one kind of signal at the same time: vision, voice, text, gesture, and even physiological signals, the way a human does.

The important word is together. It is not about running a vision model and a voice model side by side and stapling the outputs. It is the fusion. The system understands that the slight tremor in a voice, the hesitation in an on-screen action, and the mismatch in a document all point to the same conclusion, when none of them would have on their own.

A few examples make it concrete.

Fraud you cannot see in the numbers. A transaction can look perfectly ordinary on paper. But fuse it with device behaviour, typing rhythm, and the way a person moves through a screen, and a synthetic or coerced session starts to stand out. The signal is not in any one stream. It is in the combination.

Understanding a borrower you have thin data on. For millions of Indians, traditional credit history barely exists. But a short video interview, a spoken conversation, and a set of documents, read together, can build a far richer and fairer picture than a credit score alone.

Banking that meets people where they are. A customer sends a photo of a cheque, asks a question by voice in their own language, and expects one coherent answer. That is three modalities in a single interaction, and handling it gracefully is a multimodal problem.

Reading the market, not just the numbers. Price and volume tell one story. Add the tone of an earnings call, the body language on a results-day video, and the swing of sentiment across news and social feeds, and a prediction model starts to see what the spreadsheet alone would miss.

Advice that actually fits the person. A good personal advisor listens to what a client says, watches how they say it, and notices the hesitation before a big decision. An assistant that senses tone, expression, and text together can offer guidance that feels personal rather than generic.

Spotting bias you would otherwise miss. When you can look at decisions across several signals at once, patterns of unfairness that hide inside a single data stream become visible, which matters for anyone building lending or advisory systems that have to be fair as well as accurate.

If any of that sounds close to what you are building, keep reading, because there is now a place built for exactly this, ready to help you build it.

The hard part, and where India just built an answer

Here is the catch. Multimodal AI is powerful precisely because it needs rich, real-world data across several signals captured at the very same instant and perfectly synchronised. And that data, especially for Indian faces, voices, accents, and everyday situations, is genuinely rare and extraordinarily hard to create. It is the single biggest thing standing between a promising idea and a working product.

This is the problem the MI-RA Lab (Multimodal Intelligence for Real-World Applications) was built to solve. Set up by IIT Mandi iHub and HCi Foundation, it is India’s first and only dedicated multimodal AI research facility, and there is genuinely nothing else like it in the country.

Step inside and it looks the part. At its heart is an anechoic chamber (a room engineered for total acoustic isolation, so no stray echo contaminates the recording). Inside it sits an egg-shaped volumetric capture rig (a steel chassis that surrounds the subject on all sides), ringed with around 16 to 20 Blackmagic cinema cameras, precision microphones, and lighting that can recreate anything from bright daylight to near darkness.

The rig uses a spherical equidistant geometry (every camera, from the ones on the face to the ones on the legs, sits at exactly the same distance from the subject), which keeps image resolution uniform across the whole body and makes calibration far simpler. It is built in corner-collapse quadrants (four detachable sections that fold away to the room’s edges) so that ceiling-mounted OptiTrack motion capture (a wide-area system that tracks full-body movement and gesture) gets a clear line of sight.

On and around the subject sit EMG and EDA biosensors (EMG reads muscle activity; EDA reads the tiny changes in skin conductance that track emotional arousal). And here is the genuinely hard part of the engineering: every one of these streams, video, spatial audio, physiology, and motion, is synchronised to sub-millisecond accuracy. So, the tremor in a voice, the flicker of a micro-expression, and the shift in a heartbeat are all captured as the single, connected moment they actually were.

The ambition behind it is deliberately large. The ambition is to build something like an “ImageNet for multimodality”, a foundational, India-first dataset of synchronised human signals that a whole generation of AI products can be built on. For a startup, that is not lab tourism. It is access to the one ingredient that is otherwise almost impossible to get.

The programme built around it

That lab now sits at the centre of the Multimodal AI Innovation Program, a joint initiative of ICICI Bank and IIT Mandi iHub and HCi Foundation, with the Reserve Bank Innovation Hub (RBIH) as ecosystem partner. Much of its focus sits squarely in behavioural finance, where how people actually decide, hesitate, and react matters as much as the numbers, which is exactly where multimodal signals earn their keep.

For a selected cohort of startups, it offers a combination that is genuinely hard to assemble on your own:

• Access to the MI-RA Lab, its datasets, its capture infrastructure, and a research-grade environment to develop, test, and validate multimodal solutions.

• Real problems and pilot pathways with ICICI Bank and other financial companies, not simulations.

• Mentorship from senior bankers, founders, and investors.

• A structured programme across technology, product, go-to-market, and fundraising, ending in a Demo Day in front of curated investors.

Participation is free. Selection is on merit.

You don’t need to have it all figured out

Here is the part worth saying plainly, because our early applications suggest some founders are talking themselves out of it.

You do not need to already be a “multimodal AI company”. If you are working with one modality today and can see how a second or third would make your product sharper, that is exactly the kind of startup this programme exists to help. The one thing we look for is that you are past the idea stage, with a working prototype or MVP, and ready to build.

If you have read this far and something in it felt familiar, that is the signal.

Applications close on 28 August

The window is open now, and it closes on 28 August. A cohort of 12 to 15 startups will be selected.

If you are building AI for financial services, or you can see how adding another sense would change what your product is capable of, this is your invitation to apply.

Apply here: https://www.ihubiitmandi.in/mmai-icici/


메타데이터
post_id
8803be4da397
slug
your-ai-has-only-one-sense-your-customers-have-five-missing-the-leverage-8803be4da397
url
https://medium.com/@tihiitmandi/your-ai-has-only-one-sense-your-customers-have-five-missing-the-leverage-8803be4da397
canonical_url
https://medium.com/@tihiitmandi/your-ai-has-only-one-sense-your-customers-have-five-missing-the-leverage-8803be4da397
author_url
https://medium.com/@tihiitmandi
status
ok
fetched_at
2026-08-16 01:40:05