Metrics can deceive, but eyes can’t: This AI technique proposes perceptual quality metrics for video frame interpolation

AI Video & Visuals


Source: https://www.ecva.net/papers/eccv_2022/papers_ECCV/html/1542_ECCV_2022_paper.php

Advances in display technology have made our viewing experience more intense and comfortable. Watching something at 4K 60FPS is much more satisfying than 1080P 30FPS. The first immerses you in the content as if you were witnessing it. However, this content isn’t easy to distribute, so it’s not for everyone. One minute of his 4K 60FPS video costs about six times the data cost of 1080P 30 FPS, making it inaccessible to many users.

However, it is possible to address this issue by increasing the resolution and frame rate of the delivered video. Super-resolution techniques work on increasing the resolution of the video, while video interpolation techniques focus on increasing the number of frames in the video.

Video frame interpolation is used to add new frames to a video sequence by estimating the motion between existing frames. This technique is widely used in various applications such as slow motion video, frame rate conversion and video compression. The resulting video usually looks nicer.

🚀 Check out 100’s of AI Tools at the AI ​​Tools Club

In recent years, great progress has been made in research on video frame interpolation. It produces intermediate frames very accurately and provides a comfortable viewing experience.

However, measuring the quality of interpolation results has been a difficult task for many years. Existing methods mostly use off-the-shelf metrics to measure the quality of interpolation results. Interpolated video frames often have inherent artifacts, so existing quality metrics may not match human perception when measuring interpolated results.

Some methods conduct subjective tests to get a more accurate measurement, but that takes time, with the exception of some methods that use user research. So how do you accurately measure the quality of your video interpolation method? Time to answer that question.

A group of researchers has published a dedicated perceptual quality metric for measuring video frame interpolation results. They designed a new neural network architecture for video perceptual quality assessment based on Swin Transformers.

The network receives as input a set of frames, one from the original video sequence and one interpolated frame. Outputs a score representing the perceptual similarity between two frames. The first step in realizing this kind of network was to prepare the dataset, and that’s where we started. They constructed a large video frame-interpolated perceptual similarity dataset. This dataset contains pairs of frames from different videos and human judgments of their perceptual similarity. This dataset is used to train the network using a combination of L1 and SSIM goal metrics.

L1 loss measures the absolute difference between prediction scores and ground truth scores, while SSIM loss measures structural similarity between two images. Combining these two losses trains the network to predict scores that are accurate and consistent with human perception. The main advantage of the proposed method is that it is frame-of-reference agnostic. Therefore, it can run on client devices where that information is not normally available.


Please check paper.don’t forget to join 20,000+ ML SubReddits, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more. If you have any questions regarding the article above or missed something, feel free to email me. Asif@marktechpost.com

🚀 Check out 100’s of AI Tools at the AI ​​Tools Club

Ekrem Cetinkaya graduated with a Bachelor of Science degree. He completed his master’s degree in 2018. He graduated in 2019 from Ožegin University, Turkiye, Istanbul. he wrote his master’s degree. A paper on image denoising using deep convolutional networks. He got his Ph.D. He completed his doctoral dissertation in 2023 at the University of Klagenfurt, Austria, titled “HTTP Adaptive Using Machine Learning to Enhance Video Coding for His Streaming”. His research interests include deep learning, his vision of computers, video encoding, and multimedia networking.

🚀 Transform Selfies into AI Generated Headshots: Try the #1 AI Headshot Generator Today



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *