Abstract
Previous methods of continuous spatial video super-resolution (C-STVSR) employ implicit neural representations (INRs) for continuous encoding, but they often struggle to capture the complexity of video data, relying on simple coordinate concatenation and pre-trained optical flow networks for motion representations. Interestingly, contrary to general observation, we find that adding encodings does not improve and does not degrade performance. This problem is especially pronounced when combined with pre-trained optical flow networks that may limit the flexibility of the model. To address these issues, we propose the BF-STVSR, a C-STVSR framework with two important modules tailored to better represent the spatial and temporal properties of the video. Our approach achieves cutting edge with a variety of metrics, including PSNR and SSIM, demonstrating spatial detail and natural temporal consistency.
A research team led by Professor Jaejun Yoo of Unist's Graduate School of Artificial Intelligence has announced the development of an Advanced Artificial Intelligence (AI) model, “BF-STVSR (Instant Video Super Solution Based on Two-way Flows.”).
Resolution and frame rate are key factors that determine the quality of your video. Higher resolutions will result in sharper images with more detailed visuals, but higher frame rates ensure smoother movements without sudden jumps.
Traditional AI-based video repair techniques typically rely heavily on pre-trained optical flow prediction networks for motion estimation, which handle resolution and frame rate improvements individually. Light flow calculates the direction and velocity of the object's movement to generate an intermediate frame. However, this approach involves complex calculations, is prone to accumulated errors, and limits both the speed and quality of video repair.
In contrast, “BF-STVSR” introduces a signal processing method tailored to video characteristics, allowing models to independently learn bidirectional movement between frames, without relying on external optical flow networks. By collaboratively inferring the contour and motion flow of objects, the model effectively enhances both resolution and frame rate at the same time, resulting in more natural and coherent video reconstruction.

Applying this AI model to low-resolution low frame-rate video demonstrated superior performance compared to existing models, as evidenced by peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) scores. The rise in PSNR and SSIM values indicates that even videos with critical motions retain clear, undistorted human figures and details, producing more realistic results.
Professor Jaejoong Yoo said, “This technology has a wide range of applications, from restoring security camera footage and black box recordings captured on low-end devices to rapid enhancing compressed streaming videos for high-quality media content.
The study was led by the first author, Eunjin Kim, and co-authored by Heeonjin Kim. Their findings have been accepted in their presentation at the 2025 Conference on Computer Vision and Pattern Recognition (CVPR), one of the most prestigious conferences in the field of computer vision. The CVPR, held in Nashville, USA from June 11th to 15th, received 13,008 submissions and only 22.1% (2,878 papers).
The project was supported by the Ministry of Science (MSIT), the National Research Foundation of Korea (NRF), the Institute for Information and Communications Technology Planning and Evaluation (IITP), and the UNIST Supercomputing Centre.
Journal Reference
Eunjin Kim, Heeonjin Kim, Kyong Hwan Jin, Jaejun Yoo, “BF-SPLINE: B-SPLINES and FORIER-BEST FRENDS OF FORIER-BEST FRENDS FOR High Fidelity Spatial Video Super Resolution,” CVPR 2025, (2025).
/Public release. This material of the Organization of Origin/Author is a point-in-time nature and may be edited for clarity, style and length. Mirage.news does not take any institutional position or aspect, and all views, positions and conclusions expressed here are the views of the authors alone.
