What is Multimodal AI? Images, audio and text in one model

What is Multi Modal AI? — watch on YouTube
5 min

In 30 seconds

  • Multimodal means one model handles all four modalities, text, images, audio and video, on the way in and, for some models, on the way out.
  • Hand a model a photo of a paper invoice and it reads the picture and the words in the same pass, then gives you the total and the due date.
  • Image recognition tells you which bucket a picture belongs in. A multimodal model spotted the Q3 dip in a revenue chart on its own, which older models could not do.

Multimodal AI (also written multi-modal) means one model that handles more than text. The same model takes images, audio and video as input, and some of them produce those formats as output. Photograph a paper invoice, hand it over, and it reads the total and the due date back to you.

What does multimodal AI mean?

Older models did one thing: text in, text out. Then image models, video models and audio models arrived, each on its own. Multimodal means a single model covering all four modalities, text, images, audio and video, on the input side and on the output side.

Is ChatGPT multimodal? What about Claude and Gemini?

ChatGPT, Claude and Gemini are all multimodal. They are not multimodal equally. Gemini is strong on output, and Nano Banana, one of its submodels, is where that shows: very good images, very good video. Claude has no output side here at all and will not generate an image for you. Send Claude a screenshot of your phone and it works out what is happening on the screen.

ChatGPT Claude Gemini
Multimodal yes yes yes
Reads an image you send it yes yes yes
Generates images or video not covered in the video no yes, through Nano Banana

That split matters when you pick a tool. Reading a photo of a job sheet and generating an image are two separate capabilities, and a vendor can be excellent at one and absent on the other.

How does AI read an image?

Upload a paper invoice and the model reads the picture and the words in the same pass.

It then reasons with itself about what it is looking at: the price is this, the due date is that. Reasoning gets its own lesson. For now, treat it as the model talking to itself about the image. The third step is the answer you wanted, the total and the date it falls due.

How is multimodal AI different from image recognition?

We handed a model a bar chart of revenue running from Q1 to Q4 and it picked out the dip in Q3 on its own. Older models could not do that at all. Train one specifically to flag a quarter where revenue fell and it might manage that one question, and nothing else you ask about the chart. Give the multimodal model hundreds of charts and it will work through them.

Image recognition tells you which bucket the picture belongs in. A multimodal model tells you what is in it.

Image recognition is old. Machine learning has been doing it for more than a decade, and it works. Point it at a field and it will tell you which vegetables are going bad, because it was trained on a large amount of data about vegetables going bad, and a category or a true or a false is the whole of what comes back. A multimodal model is trained on the whole internet, so it reasons about what is inside the image instead of matching a pattern it was drilled on.

The same gap shows up against OCR. OCR pulls the characters off the page. A multimodal model reads the characters and the layout together, which is how it works out which of the numbers on an invoice is the total.

Image recognition Multimodal model
Trained on one narrow dataset the whole internet
Gives you back a category, or true or false reasoning about what the image shows
Reading a paper invoice OCR returns the characters on the page the total and the due date, because it reads the layout too
Spot bad vegetables in a field yes, that is what it was built for yes
Spot the Q3 dip in a revenue chart only if trained to look for falling quarters notices it on its own

What can multimodal AI do with your business documents?

Most of the information in a field-services business sits in images. Whiteboards. Paper. Slide decks. Photos taken on site.

We work on this directly. After a whiteboard session someone photographs the board and feeds it in. Calls get recorded and transcribed, and the transcript goes in alongside the photo, so the model has the discussion as well as the diagram.

The job is documenting and categorising the material you already hold, so the model has enough context to be useful. Start by listing where your operation keeps information that has never been typed: the paper invoices, the handwritten service tickets, the whiteboard in the dispatch office. Then work down that list and decide what gets photographed, what gets transcribed, and who is doing it each week.

Full transcript

Expand

Hey, my name is Mahan. I'm the CMO of Naro. My name is Luka and I'm the CEO of Nairon. Luka, over the past couple episodes, we've gone deep into how AI works, breaking it down to different pieces. But what we haven't covered in depth yet is how AI has become multimodal. So, in this video, let's discuss what that means so the audience has a better understanding of AI's capabilities in 2026. Essentially multimodal means multifunctional across different types of inputs. Though previously models have been able to only receive text and give off text. In the last few years we've seen image models, we've seen video models, audio models both on the input side and on the output side. So I have an example over here of a old model which was able to just take text and an example of a multimodal uh model which is able to take text images audio and essentially output all of that as well.

Now the models that we use today, JBT, Claude, Gemini, all of them are multimodal. Okay. However, they're not multimodal equally. Gemini is really good at output in multimodal. So really good with nano banana. Those are names of different sort of submodals. And it's able to produce very good videos and very good images. Claude, for example, doesn't have an output multimodal system where it's able to generate images for you.

Uh, but it's able to input. So, it's if you give it a screenshot, for example, of your phone, it's able to sort of understand what's happening on your phone. So, not every model builds their multimodal models uh the same.

And so, just to give a quick example to the audience, what does it look like in slow motion when you give uh an NLM say an image to analyze? Yeah. And so particularly here we have an image of an invoice. So if you were to upload a paper invoice um to to an LLM, what the LLM is going to do is actually understand the image and the text together. So it's not like a separate function. It really runs together. And so the LLM might actually reason with itself and sort of analyze the image and think to itself and say, "Oh, I see the price is that and this was the due date." And sort of extract these information. And by the way, in the next video, we're going to cover reasoning and how that works with an LLM, but it's essentially the LLMs way of sort of working with itself to figure out what's happening. And so it does that with images quite a lot. And then you can see over here in the in the third step, it actually produces the final outcome um which is, you know, telling you what the total is and then what due date it is, etc.

And this has been very revolutionary over the last couple of months because as we both know, most business information is not just text format. Yeah. And so what we see with a lot of businesses is that most of their information is in images. Um it's in whiteboards perhaps on paper as well. Um and so a lot of the stuff that we do is we try to now start documenting and sort of categorizing all of that information so that we can give the LLM enough context to do all the cool stuff that LM are able to do, what agents are able to do, which we'll get to in other videos as well.

But it can really be a a plethora of different things. It can be a slide deck. Yeah, exactly. And you hit a point there. We use whiteboards all the time. And now all we have to do is just take an image of what we've gone through in the session and feed it into the LM and there you go. That's stored there for Evernote.

Well, not only that, we also record our calls, transcribe them and actually use the transcript of the meeting to get even more context of what we use on the whiteboards. I mean, image recognition is not a new concept. um a lot of the old machine learning techniques were and and have been used for even more than a decade on image recognition. However, there's a difference between image recognition and being a multimodal AI that's able to understand images. And that's the that's the piece that I'm getting to understand the image. Um, image recognition was essentially a machine learning technique to identify patterns in images and basically give you either a categorization or a true and false about some image. Right? For example, if you're um looking at a field and you want to spot what what vegetables might be going going bad, an image recognition uh machine model was able to tell you if the vegetables were going bad or not because it had a plethora of data on on that specific thing.

The multimodal LLMs that we use, they're trained on the whole internet, right? there's so much more context and so it's able to reason about what's inside the actual image. And so we gave it a bar chart uh of revenues all the way from Q1 to Q4 and it was actually able to reason and actually notice that there was a dip in Q3 something that the old models weren't able to do at all. I mean maybe if you asked it and trained it on identify if any quarters revenue fell then perhaps it was able to tell you but this was able to reason a lot more and you can give it hundreds of graphs like this and have it do a lot of data analysis.

Essentially that understanding has really come from the fact that AI is much better now at reasoning than it was before which is something that we'll get into more in the next video. So guys if you enjoyed this video subscribe to the channel you can subscribe to our newsletter links down in the description below. You can also find us on LinkedIn where we're very active on a daily basis. And uh until next time, we'll see you in the next video.

Next in AI FundamentalsWhat is AI Reasoning?
Ready to put AI employees to work?
Book a call
256-bit SSL Secured
© 2026 Nairon, Inc. All rights reserved.
PrivacyTerms & ConditionsCookie PolicyAcceptable Use