Open source platforms evaluate how computer vision tasks and downstream inference compression methods

Machine Learning


The growing demand for computer vision applications that rely heavily on image and video data requires efficient compression techniques that are specifically tailored to these tasks. Interdigital's Hiomin Choi, Heaji Han of Hanbat National University, Chris Roseworn of Canon, and Interdigital's Fabian Rasape are tackling this challenge by introducing Compressai-Vision, a new open-source software platform designed to rigorously evaluate how computer vision is compressed. The platform provides a standardized environment for testing how accurate different compression tools maintain vision tasks, taking into account both local and remote processing scenarios. By providing a common basis for comparison, Compressai-vision accelerates the development of optimized compression technologies and has already gained recognition through adoption by the Maving Pictures Experts Group.

SION tasks, related neural network models, and datasets require a concatenated platform as a common basis for implementing and evaluating compression methods optimized for downstream visual acuity tasks. Compressai-Vision is being introduced as a comprehensive evaluation platform where new coding tools are competing to maintain task accuracy in the context of two different inference scenarios, while efficiently compressing the inputs of the visual network. This evaluation platform has it.

Video compression evaluation focused on machine learning

This study introduces Compressai-vision, an open source framework designed to evaluate video compression techniques dedicated to machine learning applications, often referred to as machine video coding. Addresses the critical needs of efficiently compressing video data while storing information essential for artificial intelligence tasks such as object detection, pose estimation, and tracking. Traditional video compression metrics do not accurately reflect the performance of these AI models and require a new evaluation approach. Key contributions and features include: * Focus on machine learning performance: By directly measuring the effect of compression on the accuracy of AI models, Compressai-Vision moves beyond traditional metrics.

Integrates with popular AI frameworks such as Detectron2 and MMope to evaluate performance after compression and decompression. *Comprehensive Dataset Support: The framework supports a variety of datasets commonly used in machine learning, including open images, FLIR thermal datasets, SFU-HW objects, Tencent video datasets (TVDs), and event human. * Integration with AI frameworks: Seamlessly integrates with popular AI frameworks such as detectron2, mmophe, and yolo, allowing end-to-end evaluation of compressed video data. *Open Source and Extensible: Compressai-Vision, which is open source, facilitates community contributions and makes it easy to customize and expand to support new datasets, AI models, and compression technologies.

  • Support for the latest video codecs: using framework, H. 264, H. 265, H. You can evaluate a variety of video codecs, such as 266. *General Test Conditions: This study establishes general test conditions for video coding of machines and ensures fair and reproducible evaluation results.

This method involves compressing the video data with a selected codec, decompressing it, and feeding it to an AI model to measure performance. This framework is important for developing and evaluating video compression techniques optimized for machine learning applications. It bridges the gap between traditional video compression and AI requirements, enabling more efficient and effective video analysis in areas such as autonomous driving, robotics, and surveillance. Compressai-Vision is a valuable tool for researchers and developers, providing a standardized, comprehensive platform for evaluating and comparing a variety of compression technologies.

Compressai-Vision evaluates video coding for computer vision

Scientists have developed Compressai-vision, a comprehensive evaluation platform designed to evaluate video compression methods specialized for computer vision tasks. This work introduces the platform's capabilities through extensive testing using standard codecs and various datasets. The experiments show significant compression gain using FCTM V6.

One codec for VCM-RS V0. 12 codecs spanning multiple datasets. For the SFU-HW-OBJ dataset, FCTM achieved bitrate savings of 79.35% and 69.02% in class C and D, respectively, while maintaining comparable task accuracy.

On average, FCTM reduced the bitrate by -58. 33%, -41. 43%, and -72. 70% under random access, low latency, and all-intra configurations, respectively, when compared to VCM-RS results. Conversely, when assessing VCM-RS under FCM CTTC, the team discovered it, and FCTM significantly outperforms other methods on TVD datasets, reaching nearly unlost accuracy at higher bitrates.

Further analysis revealed the use of VTM-23. 3 As an inner codec for the FCTM V6. 1 provided superior performance compared to using the JM-19. 1 or HM-18. 0. These results demonstrate the ability of Compressai-Vision to consistently evaluate coding performance across a variety of internal codec configurations and heterogeneous inference pipelines. The platform is poised to support cutting-edge vision transformer architectures and multitasking networks, allowing you to investigate the impact of compressive noise on embedded spaces and optimize how to code a variety of tasks.

Compression assessment of computer vision tasks

Compressai-Vision represents a significant advance in the evaluation of video compression technology, particularly for computer vision applications. The researchers have developed a comprehensive platform that allows for comparative analysis of coding tools while maintaining the accuracy of downstream visual acuity tasks evaluated in both remote and split inference scenarios. The platform facilitates a detailed investigation of bitrate and task accuracy across different datasets, providing valuable insight into the trade-offs between compression efficiency and performance. The open source nature of Compressai-Vision ensures scalability, promotes contribution from the wider research community, and encourages continuous development and innovation. The authors acknowledge that the platform is currently focusing on convolutional neural networks, plans to expand support for visual transformer architectures, allowing investigation into the effects of compressive noise on embedded spaces. Future work will explore multitasking networks to optimize coding methods for parallel processing of various machine vision tasks.

👉Details
🗞 Compressai-vision: Open source software for evaluating how computer vision tasks are compressed
🧠arxiv: https://arxiv.org/abs/2509.20777



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *