>Foundations
What is Multimodal AI?
Multimodal AI is AI that can work with more than one type of input or output — for example text, images, audio and video together.
Multimodal models can describe an image, read a chart, or turn a sketch into code, opening up new everyday use cases.
kenai shows teams practical multimodal workflows — like turning screenshots, documents and slides into useful output.
Related terms
Generative AI (GenAI) →
Generative AI is AI that creates new content — text, images, code, audio or video — rather than only analysing existing data.
Large Language Model (LLM) →
A Large Language Model (LLM) is an AI system trained on vast amounts of text to understand and generate human language, powering tools like Claude and ChatGPT.
Artificial Intelligence (AI) →
Artificial Intelligence (AI) is software that performs tasks normally requiring human intelligence — understanding language, recognising patterns, reasoning and making decisions.