Fourth episode of the series. Today we tackle a word that has become central in 2026 AI presentations and which, for how it gets used, sounds much more technical than it actually is: multimodal.
Let's start with the word itself. Modality in AI language means "type of data": text, images, audio, video. A unimodal model works with one type. A multimodal model works with two or more types together.
That's it. It's the basic concept.
What makes the thing interesting is not the definition. It's what becomes possible when an AI can look, listen, read, and understand — all at the same moment.
The metaphor
Think of a specialist you can consult by email.
If your specialist can only read your words, you have to tell them everything: the problem, the context, the details, what the object looks like, what sound it makes, where the symptom is. That's what a text-only AI does (base ChatGPT, text Claude). It works, but much of the time is spent describing instead of solving.
If your specialist can also see the photo you send them, hear the audio recording, read the floor plan you forward — then a big part of describing disappears. You say "this thing makes this noise when I turn it on" and send a 5-second video. The specialist knows exactly what you're talking about.
This is a multimodal AI. They understand you with less effort on your part. And often they give you more precise answers, because they have more context to start from.
The four business cases where everything changes
Three years ago, these applications were advanced research or demos. Today, in 2026, they are real products you can buy or build.
1. Customer service: the user sends the photo
Real example. A customer of an appliance company writes to customer care: "the washing machine is leaking water, what do I do?". Classic customer service: ten back-and-forth messages to figure out where it leaks from, how much, in what conditions, model, year of purchase.
Multimodal customer service: the customer sends the photo of the washing machine. The AI recognizes the model, visually identifies the leak (under the gasket, from the filter, from the drain hose), and in two exchanges provides the right procedure.
What's needed for it to work: a multimodal model (today easily accessible via API), a knowledge base of the product (see the RAG hardwords piece), an integration with the ticketing system.
Typical ROI: average ticket resolution time drops 40-60%. For high-volume companies, it's huge.
2. E-commerce: image search
Example. A user sees a sofa in a photo on Instagram, would like to buy it but doesn't know brand/model. Opens your site, uploads the photo. The multimodal AI recognizes visual characteristics — shape, color, structure, fabric, legs — and finds in your products the 5 most similar.
It works very well for furniture, fashion, accessories. It works less well for categories where the difference is functional and not visual (one mouse from another mouse).
Typical ROI: conversion rate +20-35% on products where image search is activated. Higher on marketplaces, lower on single-product brands.
3. Manufacturing: visual quality control
Classic Italian example. A production line that assembles components. Rare defects, but expensive. Quality control traditionally done by expert operators looking "by eye".
Multimodal AI: a camera over the line, a model trained to recognize specific defects. The model flags every suspicious piece to a human operator who decides. The model doesn't replace the operator — it amplifies them.
What was random quality control becomes systematic quality control. Defects that previously escaped get caught. False positives get educated.
Typical ROI: depends heavily on the sector, but for Italian precision engineering manufacturing companies — think your Loccioni — it's one of the most mature and best-quantifiable cases.
4. Compliance and documentation: "reading complex documents"
Example for professional studios, consultancies, healthcare. An accountant has to check a balance sheet: 80-page PDF with tables, charts, marginal notes. The text-only model struggles with complex PDFs (tables "break" in the text). The multimodal model looks at the image of the page and understands the table the way a human would.
Same goes for medical records, contracts with highlighted clauses, floor plans with annotations, technical sheets.
Typical ROI: time spent "deciphering" a complex document drops 70-80%. The part where AI can err (a wrong interpretation) is managed with a final human review, which remains indispensable.
How to know if you need it
A quick checklist. If you answer yes to one of these, multimodal AI is a serious lead for you.
- Do your customers send you photos, videos, audio in daily communications?
- Do your products have a visually recognizable physical component?
- Does your work include "reading" documents with many images, tables, schemas?
- Do you have specialized operators doing visual or auditory checks?
If the answer to all is no, text AI (more mature, cheaper) probably suffices. If the answer to one is yes, a multimodal pilot is worth evaluating.
When you don't need it
Three red flags.
They propose "multimodal AI" for problems that are text-only. Resist. Adding modalities you don't need is just complexity and cost.
They propose a "custom" multimodal model without having done a pilot. Commercial models (GPT-5, Claude 4, Gemini 3) are natively multimodal and work very well on almost all common use cases. Custom is needed only for really specific cases (specialized medicine, precision industrial vision). For the rest, start with the commercial model.
They propose a system that does "everything" — text, images, audio, video, advanced reasoning. It's probably a package sale, not a targeted solution. Start from a use case, scale when the first works.
Summary in three lines
A multimodal AI is a model that understands more types of data together — text, images, audio, video.
It changes everything in business cases where communication with the customer or the process isn't only textual.
It's worth starting from a pilot on a specific case, not from an "enterprise multimodal AI strategy".
Monday, August 10 we return to the branding pillar with a piece on seasonal brands: who knows how to use the emotional calendar (Mulino Bianco, Lavazza) and who fakes it. Three Italian SMEs that made the calendar an asset, not a cost.
Are you wondering if multimodal AI can really help you? Under artificial intelligence for business we do half-day AI use-case audits. We tell you if it makes sense, and how much it costs to do well. Let's talk.
