Multimodal content: images, video, audio
Alt text, captions, transcripts: how to make your images, videos and podcasts genuinely usable by generative AIs and search engines alike.
Multimodal content refers to everything on a site that isn't plain text: images, videos, podcasts, infographics. For a generative AI to make use of this content, it needs to come with a sufficiently precise textual description — without one, it remains largely invisible to an engine.
Why non-textual content remains largely invisible to AI
The language models powering generative AIs are, at their core, trained on text. Some recent models include image understanding or automatic description capabilities, but this capability remains uneven across models and contexts, and it doesn't replace a clear, explicit accompanying text. For a traditional search engine as much as a generative one, an image or video without a text description remains, to a large extent, content it can neither index precisely nor cite reliably.
This creates a frequent blind spot: a company might invest in a polished product demo video, or an infographic that perfectly summarizes its expertise, without that content ever being picked up in a generated answer — simply because no accompanying text lets an engine understand what it contains.
Alt text: the foundation for images
Alt text is the textual description attached to an image in a page's code. Originally designed for accessibility (screen readers for visually impaired users), it has also become an essential signal for search engines and, by extension, for GEO.
A few simple principles:
- Describe what the image actually shows, factually, without keyword stuffing.
- Adjust the level of detail to the image's importance: a decorative photo can stay minimal, a technical diagram or product shot deserves a fuller description.
- Avoid generic text ("image1.jpg", "product photo") that provides no usable information.
- For purely decorative images, an empty alt text is preferable to an artificial one that adds nothing.
Well-written alt text lets an engine understand that an image illustrates a specific point on the page, and therefore consider it part of coherent content — a criterion tied to the principles of GEO-ready content.
Captions: valuable additional context
Unlike alt text, invisible to a human visitor, a caption is displayed below or next to the image. It provides additional context — a photo's source, a figure tied to a chart, an explanation of a diagram — which helps both visitors and engines understand visual content without needing to interpret the image itself.
Captions are especially useful for charts and infographics containing data: restating the key figures shown visually in a caption, or ideally in the surrounding text, lets an engine pick them up correctly, whereas data present only within an image remains very hard to extract reliably.
Transcriptions for video and audio
For video and audio (podcasts, webinars, interviews), a full text transcript remains the most reliable way to make this content usable by an engine, generative or otherwise. A well-structured transcript also makes it possible to locate and cite a specific passage, which ties into the chunking logic used by many search systems (see the glossary).
A few best practices:
- Publish the full transcript on the same page as the video or audio, rather than on a separate, hard-to-link page.
- Structure the transcript with subheadings or timestamps for longer content, to make it easier to locate a specific passage.
- For subtitles embedded in the video, check their accuracy: an unreviewed automatic transcript often contains errors that can distort the meaning of a cited passage.
- Add a summary at the start of the transcript for longer content, useful both for time-pressed visitors and to give quick access to the main topic.
Structured data applied to multimodal content
Beyond accompanying text, specific structured data types exist for images and videos ("ImageObject" and "VideoObject" in schema.org), which let you specify metadata such as a video's duration, publication date, or description. This technical markup doesn't replace good text description, but reinforces it by making it readable in a structured way by the systems that analyze it automatically.
A practical method to catch up
Few sites handle all their multimodal content correctly from the start. Here's a step-by-step method to close existing gaps:
- Inventory your most strategic non-textual content: product demo videos, infographics with key data, high-value podcasts or webinars.
- Prioritize checking the alt text of images on your most visited or commercially important pages.
- Add transcripts to videos and audio content that don't have them yet, starting with those containing the most useful information.
- Restate in text form the key figures and information that currently exist only in visual form.
- Track whether this content starts getting picked up in generated answers, notably through analyzing your citation sources.
Multimodal content remains an often-overlooked axis of GEO, even though it frequently represents a site's richest and most differentiated content. Making it readable for engines, beyond its visual or audio quality, is a step not to skip.