>Foundations

What is Multimodal AI?

Multimodal AI is AI that can work with more than one type of input or output — for example text, images, audio and video together.

Multimodal models can describe an image, read a chart, or turn a sketch into code, opening up new everyday use cases.

kenai shows teams practical multimodal workflows — like turning screenshots, documents and slides into useful output.

Related terms
Generative AI (GenAI)
Generative AI is AI that creates new content — text, images, code, audio or video — rather than only analysing existing data.
Large Language Model (LLM)
A Large Language Model (LLM) is an AI system trained on vast amounts of text to understand and generate human language, powering tools like Claude and ChatGPT.
Artificial Intelligence (AI)
Artificial Intelligence (AI) is software that performs tasks normally requiring human intelligence — understanding language, recognising patterns, reasoning and making decisions.

Learn to use this — hands-on.

kenai workshops and bootcamps turn AI theory into team capability.

> Book a workshopTry the AI Readiness tool