Much of the conversation around artificial intelligence (AI) focuses on its impact on productivity and automation. Less attention has been given to how it transforms [H(1] Content is created, distributed, and consumed despite its potential. As AI adoption expands and places new demands on digital infrastructure, multimedia is becoming an increasingly important part of that broader change.
Imagine you are watching a movie at home. The footage is adjusted scene by scene to maintain detail and lighting. Or imagine watching a live soccer match and seeing the moments, angles, and players you care about most as the game unfolds.
These are not far-fetched stories, but early signs of broader changes across multimedia, with AI at the center. AI is enabling more personalized, interactive experiences and increasingly rich and immersive environments across creation, editing, management, and analysis.
However, realizing this potential at scale is not easy. Maximizing the benefits of AI in multimedia requires greater standardization. Currently, the integration of complex machine learning (ML) models varies by platform, hardware, and implementation, resulting in a fragmented landscape. Without a more consistent approach, we risk making it difficult for AI-driven tools to work together seamlessly.
Enabling interoperable AI innovation across multimedia
In recent years, thousands of innovations and iterations have been incorporated throughout multimedia systems, each one pushing the industry one step further.
For example, video streaming platforms are experimenting with AI-driven compression, and creation and editing tools are using generative models to transform users’ ideas into rich media. More personalized and adaptive experiences are emerging across audio and visual services.
As AI becomes more widely integrated into multimedia, it will become increasingly important to ensure these technologies work together seamlessly. Innovation is happening rapidly across platforms and ecosystems, creating new opportunities for interoperability and collaboration.
The two video platforms may use different AI-based compression methods, so your content will require additional processing to optimize playback across services and devices. A common approach can help alleviate these inefficiencies and support smoother interoperability and more predictable performance across multimedia systems.
Initiatives such as MPEG-AI play a key role here, helping to establish a shared technology foundation that supports compatibility between systems while enabling continued innovation across the industry. In this sense, standards are not only technical enablers, but also part of the broader infrastructure needed to support the next wave of AI innovation.
Vision of MPEG-AI as a bridge between AI and multimedia
Effectively serving as a standardization umbrella, MPEG-AI is being developed as a framework to increase consistency in AI and multimedia interactions. At its core, MPEG-AI aims to consider two complementary aspects: how AI can be used to enhance multimedia systems and how multimedia can be structured to better serve AI-driven processes.
These two aspects are often described as “AI as a multimedia coding tool” and “multimedia consumed by AI systems.” Together, these provide a foundation for considering how to more effectively integrate AI across the multimedia industry.
MPEG-AI’s vision is realized through its applications. First, we look at AI as a powerful multimedia coding tool. Through neural networks, AI can enhance multimedia coding by applying deep learning models to improve video compression and quality. MPEG-AI aims to standardize the use of AI-driven compression methods and enable interoperability between platforms.
At the same time, multimedia data such as videos and images can be structured in a way that makes it more accessible to AI systems to support tasks such as object recognition and object tracking. MPEG-AI is expected to play a key role in defining how multimedia can be encoded and structured in an AI-friendly manner. Facilitating feature coding for machine use has the potential to transform input feature maps obtained from neural networks into decodable bitstreams for any machine task.
Looking further ahead, the MPEG-AI family of standards will continue to provide improved digital experiences for consumers for years to come. People will benefit from AI-enhanced applications across devices, including virtual reality, mixed reality, and more immersive gaming experiences with a supercharged metaverse.
Outside the home, multimedia AI has the potential to power smart cities and improve agent AI use cases.
Innovations that have been talked about for a long time can become reality in months rather than years, and will be rolled out to the masses thanks to a comprehensive family of standards: MPEG-AI.
What’s next for MPEG-AI?
While much of the focus today is on AI-enhanced content experiences, there is also a focus on the infrastructure needed to support AI-driven multimedia systems at scale. One area that is receiving increasing attention is tensor data, which includes intermediate feature data generated within neural networks.
This has important implications for new applications such as split inference, where AI workloads are shared between edge devices and cloud infrastructure. In these scenarios, you can run parts of the neural network locally on your device and send the compressed feature data to the cloud for further processing. Improving how this data is processed and distributed has the potential to enable more scalable and efficient AI-powered multimedia services.
Progress is already being made in this area. Recent developments in neural network coding and machine feature coding are exploring how tensor compression standards can support a broader range of use cases beyond just model parameters. This includes applications such as machine-driven feature coding, federated learning, and new immersive media formats such as Gaussian splats and 3D scene representations.
In addition to efficiency and scalability, future standard developments are also expected to support improved reliability and trustworthiness of AI systems, such as mechanisms to verify the authenticity of neural network data and updates.
The industry is currently at a critical inflection point for AI as a multimedia coding tool and as the multimedia that AI consumes. Establishing a common framework such as MPEG-AI is a critical step to supporting a more consistent and scalable future of AI-driven multimedia.
