Повышение эффективности передачи видео в компьютерных сетях с помощью нейросетевого кодирования / Improving the Efficiency of Video Transmission in Computer Networks Using Neural Network Coding тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Ибрагим Мурудж Халид Ибрагим

  • Ибрагим Мурудж Халид Ибрагим
  • кандидат науккандидат наук
  • 2025, «Московский физико-технический институт (национальный исследовательский университет)»
  • Специальность ВАК РФ00.00.00
  • Количество страниц 122
Ибрагим Мурудж Халид Ибрагим. Повышение эффективности передачи видео в компьютерных сетях с помощью нейросетевого кодирования / Improving the Efficiency of Video Transmission in Computer Networks Using Neural Network Coding: дис. кандидат наук: 00.00.00 - Другие cпециальности. «Московский физико-технический институт (национальный исследовательский университет)». 2025. 122 с.

Оглавление диссертации кандидат наук Ибрагим Мурудж Халид Ибрагим

Contents

Contents----------------------------------------------------------------------------------------------------------------3

Introduction-----------------------------------------------------------------------------------------------------------8

Thesis Topic-----------------------------------------------------------------------------------------------------------8

Motivation-------------------------------------------------------------------------------------------------------------9

Background------------------------------------------------------------------------------------------------------------9

Main Goal-------------------------------------------------------------------------------------------------------------11

Thesis Contribution-------------------------------------------------------------------------------------------------12

Thesis Structure-----------------------------------------------------------------------------------------------------15

Chapter

LITERATURE REVIEW-----------------------------------------------------------------------------------------17

1.1 Introduction-------------------------------------------------------------------------------------------------17

1.2 Video Transmission----------------------------------------------------------------------------------------17

1.2.1 Adaptive Bitrate Streaming--------------------------------------------------------------------------17

1.2.2 Compression----------------------------------------------------------------------------------------------18

1.2.3 Content Delivery Networks-----------------------------------------------------------------------------18

1.2.4 QoS Improvements---------------------------------------------------------------------------------------18

1.2.5 Error-Correcting Systems-------------------------------------------------------------------------------18

1.2.6 P2P Video Streaming------------------------------------------------------------------------------------19

1.2.7 Network Protocol-----------------------------------------------------------------------------------------19

1.3 Video Compression Techniques-------------------------------------------------------------------------19

1.3.1 Spatial Compression (Intra-frame Compression)----------------------------------------------------19

1.3.2 Temporal Compression (Inter-frame Compression)-------------------------------------------------19

1.4 Video Coding Standards----------------------------------------------------------------------------------20

1.4.1 Early video coding standards---------------------------------------------------------------------------20

1.4.2 Recent Video Coding Standards-----------------------------------------------------------------------21

1.5 Improving Video Compression Using Deep Learning---------------------------------------------26

1.5.1 Neural Networks and Deep learning-------------------------------------------------------------------26

1.5.2 Neural Network Coding (NNC)------------------------------------------------------------------------27

1.5.3 Adaptive Compression--------------------------------------------------------------------------------27

1.6 Related Work-----------------------------------------------------------------------------------------------28

1.6.1 Pre-processing

1.6.2 Coding------------------------------------------------------------------------------------------------------30

1.6.3 Post-processing-------------------------------------------------------------------------------------------31

1.7 Conclusion---------------------------------------------------------------------------------------------------34

Chapter

VVC OPTIMIZATION: POST-PROCESSING ENHANCEMENT VIA NEURAL NETWORK INTEGRATION-----------------------------------------------------------------------------------------------------35

2.1 Introduction-------------------------------------------------------------------------------------------------35

2.2 Using Neural Networks for Post-Processing Optimization----------------------------------------36

2.2.1 Proposed Res-DCNN model----------------------------------------------------------------------------36

2.2.2 Integration with VVC------------------------------------------------------------------------------------38

2.2.3 Training using TensorFlow--------------------------------------------------------------------------39

2.2.4 Testing with VVC test model (VTM 22.2)-----------------------------------------------------------39

2.3 Implementation and Results-----------------------------------------------------------------------------39

2.3.1 Data Collection-------------------------------------------------------------------------------------------39

2.3.2 Configuration

2.3.3 Evaluation Metric----------------------------------------------------------------------------------------40

2.3.4 Results------------------------------------------------------------------------------------------------------41

2.4 Conclusion---------------------------------------------------------------------------------------------------48

Chapter

VVC OPTIMIZATION: IN-LOOP FILTERING ENHANCEMENT VIA NEURAL NETWORK INTEGRATION

3.1 Introduction

3.2 In-Loop Filtering (ILF) in H.266/VVC----------------------------------------------------------------50

3.3 Using Neural Networks for In-Loop Filtering Optimization-------------------------------------51

3.3.1 RDCNN Architecture------------------------------------------------------------------------------------52

3.3.2 Rate Distortion Optimization (RDO)------------------------------------------------------------------52

3.4 Implementation and Results-----------------------------------------------------------------------------52

3.4.1 Data Collection-------------------------------------------------------------------------------------------53

3.4.2 Preprocessing------------------------------------------------------------------------------------------53

3.4.3 Training using TensorFlow-----------------------------------------------------------------------------53

3.4.4 Model Integration----------------------------------------------------------------------------------------54

3.4.5 Evaluation

3.4.6 Results------------------------------------------------------------------------------------------------------58

3.5 Conclusion

Chapter

VVC OPTIMIZATION: COMBINING IN-LOOP FILTERING ENHANCEMENT AND POSTPROCESSING ENHANCEMENT VIA NEURAL NETWORK INTEGRATION------------------63

4.1 Introduction-------------------------------------------------------------------------------------------------63

4.2 Motivation

4.3 Problem Statement

4.4 Contributions

4.5 Proposed Framework

4.6 Configuration

4.6.1 Training Environment and Setup-----------------------------------------------------------------------66

4.6.2 Dual-Mode Operation------------------------------------------------------------------------------------66

4.6.3 Feature Extraction----------------------------------------------------------------------------------------67

4.6.4 Rate Distortion Optimization (RDO) Strategy-------------------------------------------------------68

4.6.5 URDCNN Architecture----------------------------------------------------------------------------------68

4.7 Experimental Setup----------------------------------------------------------------------------------------69

4.7.1 Datasets-------------------------------------------------------------------------------------------------69

4.7.2 Preprocessing Steps

4.7.3 Encoder and Decoder Integration

4.7.4 Settings Of Quantization Parameter (QP)-------------------------------------------------------------70

4.7.5 Training Hyperparameters

4.7.6 Evaluation Metrics

4.8 Results--------------------------------------------------------------------------------------------------------71

4.8.1 PSNR Improvements-------------------------------------------------------------------------------------71

4.8.2 BD-Rate Reduction

4.8.3 Comparative Analysis

4.8.4 Discussion

4.9 Conclusion

Chapter

OPTIMIZING H.266/VVC ENCODING USING GENETIC ALGORITHMS-----------------------76

5.1 Introduction

5.2 Video Intra Coding

5.3 Intra Coding in H.266/VVC

5.3.1 H.266/VVC Advancements-----------------------------------------------------------------------------77

5.3.2 Partition Structure in VVC------------------------------------------------------------------------------78

5.3.3 Balance between video quality and coding efficiency----------------------------------------------79

5.4 Genetic Algorithms

5.4.1 Workflow-----------------------------------------------------------------------------------------------80

5.4.2 Optimizing Intra Coding Using Genetic Algorithms---------------------------------------------81

5.5 Fitness Evaluation Function

5.5.1 Peak Signal-to-Noise Ratio (PSNR)-------------------------------------------------------------------82

5.5.2 Structural Similarity Index (SSIM)--------------------------------------------------------------------82

5.5.3 Video Multimethod Assessment Fusion (VMAF)---------------------------------------------------83

5.5.4 Compression Ratio (CR)-----------------------------------------------------------------------------83

5.5.5 Bitrate (BR)-----------------------------------------------------------------------------------------------84

5.6 First Algorithm: Genetic Algorithm for Optimizing VVC Coding Tools

5.6.1 Main Steps-------------------------------------------------------------------------------------------------84

5.6.2 Experiments

5.6.3 Results--------------------------------------------------------------------------------------------------85

5.7 Second Algorithm: Genetic Algorithm for Optimizing VVC Encoding Efficiency----------89

5.7.1 Main Steps-------------------------------------------------------------------------------------------------89

5.7.2 Experiments--------------------------------------------------------------------------------------------91

5.7.3 Testing-----------------------------------------------------------------------------------------------------92

5.7.4 Results------------------------------------------------------------------------------------------------------99

5.8 Limitation--------------------------------------------------------------------------------------------------

5.9 Conclusion-------------------------------------------------------------------------------------------------100

Chapter

CONCLUSION AND FUTURE RECOMMENDATION------------------------------------------------102

6.1 Conclusion-------------------------------------------------------------------------------------------------102

6.1.1 Literature Review-----------------------------------------------------------------------------------102

6.1.2 Post-Compression Enhancement---------------------------------------------------------------------102

6.1.3 In-loop Filtering Enhancement--------------------------------------------------------------------102

6.1.4 Combining In-Loop Filtering and Post-Processing Enhancement------------------------------102

6.1.5 Genetic Algorithms Enhancement----------------------------------------------------------------103

6.2 Future Recommendation-------------------------------------------------------------------------------103

Notation

References

List of Figures------------------------------------------------------------------------------------------------------115

List of Tables-------------------------------------------------------------------------------------------------------117

Appendix

118

Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Введение диссертации (часть автореферата) на тему «Повышение эффективности передачи видео в компьютерных сетях с помощью нейросетевого кодирования / Improving the Efficiency of Video Transmission in Computer Networks Using Neural Network Coding»

Introduction

Thesis Topic

The speedy proliferation of multimedia content, particularly video, has had a profound impact on the way in which data is shared. Nevertheless, the increasing request for exquisite and high-quality video items has posed demanding situations in accomplishing efficient video transmission through computer networks [1]. Optimal video transmission is critical for the powerful delivery of advanced virtual video content material through the net. The growing number of people getting access to video content material at the net has led to a full-size rise within the need for seamless streaming experiences, that are defined via minimal buffering, high-definition quality, and low latency. Therefore, it's far crucial to creating and applying technology and procedures that could transmit video material with extra efficiency.

Due to media technology and mobile devices, global video consumption is rising swiftly, making traditional video transmission techniques unsatisfying for current user. Video transmission, especially HD and UHD facts, needs numerous bandwidths, and actual-time video application like live streaming, video conferencing, and online gaming need minimal latency [2]. Delays among transmission and reception due to high latency leads to lower live interactions and customers satisfaction.

In order to fulfill these necessities, professionals inside the business utilize number of technologies and methods, inclusive of state-of-the-art compression algorithms and adaptive bitrate streaming, as well as the implementation of resilient content material delivery networks [3]. The goal isn't always handiest to restriction data size and decrease prices associated with bandwidth, but also to enhance the overall user experience by means of ensuring that video load rapidly, play seamlessly, and show in the highest first-class viable on numerous devices and below various network situations.

There is a pressing call for novel answers that can enhance the effectiveness, dependability, and excellence of video transmission. This thesis suggests utilizing neural network coding (NNC) as an innovative approach to tackle these troubles [4]. The objective of NNC is to utilize machine learning techniques to enhance video coding and packet routing on the way to improve video quality, lower latency, increase scalability, lessen packet loss, and decrease the resource needs of video transmission systems. This approach not best complements the consumer enjoy but also promotes more sustainable practices in digital media transmission.

This thesis explores the utilization of sophisticated methods to enhance the efficiency of video transmission in computer networks. The project goals to enhance the performance and speed of video compression and transmission, by tackling considerable limitations together with big data demands and making sure premier first-class in networked settings. This thesis specializes in a comprehensive technical investigation of Versatile Video Coding (VVC, also known as ITU-T Recommendation H.266). It employs innovative techniques to extend the bounds of video exceptional and compression efficiency.

Motivation

Video content material is the number one supply of internet traffic inside the present-day generation of generation, as clients have an insatiable desire for streaming offerings, video conferencing, and online gaming. The increase in video consumption gives awesome problems for modern-day network infrastructures, particularly in upholding advanced pleasant and making sure a smooth user enjoys. The growing demand for extra powerful video transmission mechanisms is emphasized by main three principal factors: the aim of improving consumer enjoy, the restrictions imposed via existing network assets, and the chronic requirement for high-quality video content material.

Consumers count on uninterrupted and seamless video streaming, unfastened from buffering or delays, regardless of the decision of the material or their geographical location. With improvements in technology, video resolutions are enhancing, transitioning from HD to 4K and beyond [5]. Consequently, the bandwidth needed to effectively transmit those codecs is likewise increasing. If there's any incapability to offer a streaming experience of superb high-quality, it would result in consumer disappointment.

Video content sometimes generates considerable degrees of facts traffic that pressure gift network infrastructures. Several places, in particular rural and underserved ones, continue to battle with insufficient bandwidth that cannot meet the desires of cutting-edge video utilization [6]. Furthermore, the modern-day networks be afflicted by each capacity obstacles and continual troubles inclusive of packet loss, latency problems, and jitter, all of which bring about a deterioration in video quality. The presence of these limits emphasizes the instantaneous necessity for novel methodologies and technology that may improve the usage of bandwidth with the modifying network circumstances whilst retaining video high-quality.

With the development of digital media, humans aren't most effective consuming greater video but also growing a discerning preference for pinnacle-notch visuals. The increasing demand for excessive-definition video material places extra pressure on networks to facilitate the widespread transmission of huge-sized documents with high bitrates. This phenomenon is made extra complicated by way of the rise of high dynamic range (HDR) material and 360-degree movies [7], which provide advanced viewing stories but necessitate extra network resources for efficient distribution.

Content distributors and network carriers face a complicated undertaking environment due to growing expectations for a better watching revel in, network infrastructure restrictions, and growing call for remarkable video content. To overcome those issues, greater advanced video transmission strategies need to be developed and applied to satisfy modern internet intake needs. Advanced methods like neural network coding and other adaptive streaming technologies can assist stakeholders make certain that the network's capacity to deliver video efficiently evolves with patron expectations and technological advances, ensuring a robust digital destiny for video content material delivery.

Background

The size of the uncooked information linked to video files is normally attributed to the demanding frame rates and resolutions wished-for present-day video content material. Without compression, the transmission or storage of this information necessitates large-size bandwidth and storage capability, rendering the big distribution of massive quantities unfeasible and pricey. Video compression technology purpose to mitigate this trouble through decreasing the dimensions of video documents, therefore enabling quicker switch speeds and decreasing garage demands. Video compression encompasses multiple

fundamental technologies that collaborate to diminish the size of video files while endeavoring to maintain video quality. The technologies encompassed are:

■ Intra-frame Compression: This method compresses video by minimizing redundancy inside an individual frame, akin to the compression of photographs. It commonly employs the Discrete Cosine Transform (DCT) to transfer data from the spatial domain to the frequency domain [8]. DCT is utilized to partition the image into segments that vary in significance in relation to the visual quality of the image. The modification enhances compression efficiency by prioritizing the components that have less influence on visual quality when data is eliminated.

■ Inter-frame Compression: Temporal compression, also referred to as redundancy reduction, is a technique that decreases redundancy across multiple frames. It is based on the idea that multiple consecutive frames in a video exhibit a high degree of similarity [9]. Methods like temporal prediction are employed to forecast sections of a frame based on prior frames, encoding solely the disparities rather than the complete frame.

■ Motion Compensation: This technique is employed in inter-frame compression to estimate the motion of objects between frames and encode solely the change, in the form of motion vectors, coupled with a residual difference frame [10].

Traditional video compression algorithms are efficient in reducing the data size of video files. However, they encounter many difficulties, particularly in situations involving dynamic networks, fluctuating bandwidth, and unreliable network connections. While these techniques can decrease the size of the video, they do not adaptively modify the bitrate based on real-time variations in network circumstances. This lack of adjustment might result in buffering and a decrease in the quality of service when the network is congested. Packet loss can appreciably affect compressed video, whilst delays in decoding and encoding can increase total latency, which in flip affects real-time video application. Contemporary video coding standards have significantly enhanced the effectiveness and excellence of video compression, by introducing H.264/AVC (Advanced Video Coding) which ensures broad interoperability across various devices and systems. H.265/HEVC, also known as High Efficiency Video Coding, offers a compression improvement of almost 50% compared to H.264. However, due to its increased processing demands [11], it leads to longer encoding and decoding durations. AV1 (AOMedia Video 1) is characterized by its exceptional compression efficiency, however its decoding speed is slower due to its high complexity. VVC (Versatile Video Coding) has emerged as a video coding technology that surpasses H.265 and AV1 in terms of compression effectiveness. However, it is important to note that there are certain implementation cost constraints and it consumes a large amount of power. The standard is frequently determined by the particular requirements and limitations of the application, striking a balance between sophisticated functionalities and practical limitations.

These troubles spotlight the need for ingenious techniques that enhance video compression performance and also adjust to changing network situations and computational obstacles. The ongoing quest for more desirable strategies is important due to the growing call for super, actual-time video offerings.

Given the rising need for transmitting high-definition videos over computer networks, multiple scientists and research groups contributed towards evolving video compression technologies. Alan C. Bovik advanced video quality assessment when he created the Video Multimethod Assessment Fusion (VMAF) metric. His work optimizes compression by adjusting quality and bitrate tradeoffs, and is one of many algorithms utilized by Netflix. Thomas Wiegand and Gary J. Sullivan are known for creating the H.264/AVC and later the H.265/HEVC standards in video compression standards, which vastly enhanced

video compression efficiency. Their work sustains real time video platforms around the world, allowing for data rate savings up to 50% without substantial quality loss.

In AI assisted video compression, David Taubman's work on scalable and wavelet-based techniques such as JPEG2000 has greatly impacted adaptive and multi-resolution streaming. Other researchers like Rui Zhang and Huifang Sun enhanced inter-frame compression by employing neural networks for frame prediction and residual estimation. Soulef Bouaafia created learning-based approaches that utilize deep learning networks to substitute classical video coding parts, supporting dynamic bitrate modification in real time. The employment of generative adversarial networks (GANs) for video compression is becoming more common, with Anil Kokaram presenting models for detail recovery of high-frequency components in compressed videos. These techniques permit more aggressive compression ratios while maintaining visual fidelity.

Industry and academic collaborations, particularly those from Google and Facebook AI Research (FAIR), have showcased end-to-end learned video codecs, which perform better than classical approaches on restricted datasets. These projects use transformer models and attention mechanisms to improve the modeling of temporal dependencies in video data. All in all, the realm of video compression is experiencing rapid development due to scientists like A. Bovik, T. Wiegand, G. J. Sullivan, D. Taubman, R. Zhang, H. Sun, C. Lin, and L. Ma. Their advancements are commanding the development of the next generation of coding standards with heavy emphasis on AI. These advancements markedly enhance compression, visual quality, and network adaptability. Consequently, AI solutions are increasingly essential for timely, high-quality, and seamless videos. This shift enables myriad use cases, including video streaming, conferencing, surveillance, and virtual environments.

Main Goal

The transmission of top-quality video material via computer networks has grown to be an essential element of communications, permitting various offerings along with streaming systems and video conferencing. Nevertheless, the growing need for enhanced resolution and advanced quality presents great obstacles on the subject of bandwidth intake and data effectiveness. The thesis tackles those difficulties and attempts to greatly improve the efficiency of video transmission through integrating and optimizing new video compression algorithms. The foremost objective is to create novel strategies that enhance the effectiveness and excellence of video transmission in computer networks by means of using the latest improvements in neural network coding. This entails using a pronged method: first off, incorporating superior deep learning algorithms to improve video quality post-compression, improve the filtration process and secondly, optimizing the video encoding manner to correctly strike a balance among speed and quality.

Thesis Objectives

This thesis explores cutting-edge video compression processes, leveraging cutting-edge improvements in synthetic intelligence and computational methodologies to noticeably enhance the efficiency and quality of video transmission in computer networks. Primary objectives include:

■ Analyze video transmission technologies: Present an intensive assessment of each traditional and cutting-edge methodologies for video compression, emphasizing their blessings and constraints. This overview affords a foundation for comprehending the progress achieved on this thesis.

■ Improve the quality of video using neural networks: Create and contain a deep residual convolutional neural network into the VVC framework. The goal of this integration is to significantly improve the visual fidelity of video frames after compression, by tackling the principle exceptional challenges encountered in video streaming.

■ Improving the filtering step of encoding VCC unit using neural networks.

■ Develop a unified framework that integrates both in-loop filtering and post-processing optimization to achieve more efficient compression.

■ Develop a genetic algorithm to enhance the intra-coding process in the H.266/VVC standard. This involves the careful selection of the most excellent coding tools and multi-type tree partitions to obtain a balance between encoding time and video quality.

■ The improvements will be validated through effective analysis and objective measurement. A comprehensive evaluation will be conducted to assess the impact of the proposed improvements on video quality, using metrics such as PSNR. The results will be compared with those of existing methods to highlight performance improvements and identify any limitations.

■ Make significant improvements to the field of video transmission by developing insights and technologies that offer more efficient and higher-quality video streaming solutions.

Thesis Contribution

The thesis affords a substantial addition to the field of digital communications, mainly inside the place of video transmission. The primary contributions are as follows:

A. Integration of deep learning into video compression: It improves the clarity and satisfactory of video frames following compression. Integrating sophisticated neural network systems into the video compression process aids in recuperating the loss that takes place throughout the encoding phase. This now not best enhances the visual fidelity of the video but additionally guarantees effective facts transmission across networks.

B. Video coding optimization: Optimization of video coding is a multidimensional struggle concerning the selection of coding tools' combinations with respect to encoding speed and video quality. It helps provide a better system performance using a fitness function which combines perception quality metrics with coding measurement efficiency. The optimization focuses on compression efficiency and visual quality defined by the PSNR metric. Rate-Distortion Optimization (RDO) and other perceptual refinement techniques are used to mitigate artifact and detail loss during motion preservation and smoothing. The goal is to enhance the video quality while reducing the bit rate and processing power needed.

C. Further enhance the video quality: Better compression across video files achieves through replacing the in-loop filters with a neural network result in improving the balance between compression efficiency and video quality.

D. Global video compression standard advancement: This thesis lends its help to the continuous development of the H.266/VVC standard by incorporating deep learning and innovative coding methods. Because it lays the basis for addition examine and development of video coding structures.

E. Practical applications: These developments have realistic application in streaming offerings, broadcasting, and any virtual platform that calls for efficient and fantastic video transmission. Through the process of improve compression the video even as preserving high quality, these improvements have the ability to decrease costs related to bandwidth and enhance client satisfaction on many platforms.

The thesis makes full-size advances in improving the performance of video transmission with the aid of efficiently integrating neural network coding and genetic algorithms, thereby advancing the modern cutting-edge. This not handiest improves the theoretical basis of video compression tactics, however additionally gives realistic and scalable solutions that can be implemented in numerous technological settings, as a result boosting person experience and resource management in networked structures.

Scientific novelty: The scientific novelty of this dissertation lies in the development and integration of advanced neural network models and genetic algorithms to enhance the efficiency of H.266/VVC video compression. This work is different from previous works in:

■ A novel Residual Deep Convolutional Neural Network was introduced to improve both in-loop filtering and post-processing, significantly enhancing the visual quality of decoded frames while reducing compression artifacts.

■ Relying on genetic algorithms to optimize encoder selection and intra-code partitioning can effectively reduce encoding time while maintaining video quality.

■ It explores strategies for optimizing the ratio of dynamic distortion, which enhances adaptability to diverse video content.

■ It proves that neural network-enhanced coding frameworks can effectively balance bitrate reduction, encoding speed, and quality, paving the way for next-generation compression systems.

Theoretical value of the work: The purpose of the work is to develop and integrate evolutionary deep learning algorithms into H.266/VVC video coding framework, which is the core contribution of research. The work also develops novel methods of feature extraction and adaptive filtering through neural networks which establishes a new theoretical basis for video optimization. Moreover, it formulates a framework for optimizing decision making in coding using genetic algorithms while focusing on speed and quality balance.

Practical value of the work: The work builds supporting models using deep learning techniques enabling relevant speed and bitrate reduction, compared to state-of-the-art techniques, maintaining quality of the images captured. These techniques can be applicable in high-performance multimedia systems such as video conferencing, surveillance, and streaming. The proposed models enable scalable and adaptive as well as multi-user resource-friendly video transmission, which is vital in active networks sensitive to latency.

Publications: The scientific publications were produced based on the research presented in this dissertation are:

1. Ibraheem, M. K. I., Dvorkovich, A. V., Al-khafaji, I. M. A. Improving the Efficiency of Video Transmission in Computer Networks // International Journal on Recent and Innovation Trends in Computing and Communication (IJRITCC), 2023, v. 11, № 9, pp. 2500-2513. DOI: 10.17762/ijritcc.v11i9.9319. [14]

2. Ibraheem I.K., Abdalameer A.I., Hatif Naji A.Z. A Genetic Approach-Based Intra Coding Algorithm for H.266/VVC // Informatics and Automation (Информатика и автоматизация), 2024, v. 23, № 3, pp. 801-830. DOI 10.15622/ia.23.3.6. (Scopus, RSCI) [130]

3. Ibraheem, M. K. I., Dvorkovich, A. V., Al-khafaji, I. M. A. A Comprehensive Literature Review on Image and Video Compression: Trends, Algorithms, and Techniques // Ingénierie des Systèmes d'Information, 2024, v. 29, № 3, pp. 863-876. DOI: 10.18280/isi.290307. (Scopus) [31]

4. Ibraheem, M. K. I., Dvorkovich, A. V. Optimizing H.266/VVC Intra Coding with a Genetic Algorithm: Balancing Speed and Quality // Fusion: Practice and Applications (FPA), 2024, v. 15, № 2, pp. 8-16. DOI: 10.54216/FPA.150201. (Scopus) [131]

5. Ibraheem, M. K. I., Dvorkovich, A. V. Enhancing Versatile Video Coding Efficiency via PostProcessing of Decoded Frames Using Residual Network Integration in Deep Convolutional Neural Networks // 2024 26th International Conference on Digital Signal Processing and its Applications (DSPA), IEEE, 2024, pp. 1-9. DOI: 10.1109/DSPA60853.2024.10510065. (Scopus) [102]

6. Ibraheem MKI, Dvorkovich AV, Al-Temimi AMS. Innovative Integration of Residual Networks for Enhanced In-loop Filtering in VVC Using Deep Convolutional Neural Networks. Computer Optics 2025; 49 (4): 692-701. DOI: 10.18287/2412-6179-CO-1572 (Scopus, Web of Science, RSCI) [95]

These publications include:

- Five articles indexed in Scopus (Articles No. 2, 3, 4, 5, 6).

- Two articles indexed in RSCI (Articles No. 2, 6).

- One article indexed in Web of Science (WoS) (Article No. 6).

The research results cover a wide range of topics related to neural network-based video compression, optimization of H.266/VVC encoding using genetic algorithms, and the integration of deep learning models such as ResNet and RDCNN into video coding pipelines.

Validation of the research results: The main results of the dissertation were presented, discussed, and validated at multiple national and international conferences and journals, confirming their scientific and practical significance.

Approbation: The materials of the dissertation were presented and discussed on the following scientific conferences and workshops:

- 10th International Conference «Engineering & Telecommunication — En&T-2023», Moscow, Russian Federation, November 22-23, 2023.

- 26th International Conference on Digital Signal Processing and its Applications (DSPA-2024), Moscow, Russian Federation, March 27-29, 2024.

- Russian-STW: AI Innovation Summit 2024 / Huawei Workshop, Saint Petersburg, Russian Federation, September 19 - 21, 2024.

Personal contribution of the author: The author played a key role in planning the work, applying the model, and analyzing the results in all research papers. The author contributed to the entire research cycle, from design to publication, ensuring that all technical aspects and results were carefully considered.

Compliance with the specialty passport: The research conducted by the author is included in the research area 2.3.5.:

- (Intelligent systems of machine learning, database and knowledge management, tools for developing digital products) p.4

- (Models, methods, architectures, algorithms, formats, protocols and software for human-machine interfaces, computer graphics, visualization, image and video processing, virtual reality systems, multimodal interaction insocio-cyberphysical systems) p. 7.

Provisions submitted for defense:

■ The developed Residual Deep Convolutional Neural Network (Res-DCNN) model integrated into the H.266/VVC framework allows to achieve an average YUV-PSNR improvement of 1.97 dB and an overall BD-rate reduction of -2.05% (Y), -6.61% (U), and -9.78% (V) across various video sequences, enhancing compression efficiency without increasing bitrate.

■ The proposed RDCNN architecture used as a replacement for conventional loop filters in VVC allows to achieve higher average YUV PSNR values.

■ Additionally, the model demonstrates an overall BD-rate reduction of -2.43% (Y), -6.96% (U), and -9.43% (V) across various video sequences.

■ The developed dual-branch Unified Residual DCNN framework allows to achieve notable BD-Rate reductions of -3.00% (Y), -9.56% (U), and -10.50% (V), outperforming both individual enhancements. It consistently improves compression efficiency and visual quality, reaching an average YUV PSNR of 41.4 dB across various test sequences.

■ The developed genetic algorithm-driven methods for intra coding optimization in the VVC standard allow to achieve up to 20% encoding time reduction, preserving high visual quality and coding efficiency.

Thesis Structure

This thesis seeks to address the difficulties and deficiencies inherent in cutting-edge video transmission structures. It investigates novel techniques to enhance video compression and transmission by way of emphasizing advanced coding techniques just like the Versatile Video Coding (VVC) standard and neural network. The methods make use of modern-day technologies like deep learning and genetic algorithms to improve video quality and coding efficiency, crucial for meeting the rigorous requirements of cutting-edge video streaming services. The thesis comprises the following chapters:

Chapter 1- Literature Review: To provide a general explanation of current video transmission era, with a focal point on traditional and current video compression techniques. Examine previous research to offer a basis for knowledge the contributions given on this thesis.

Chapter 2- Optimizing VVC: Post-Compression Enhancement Via Neural Network Integration:

This chapter explores the new approach of utilizing a deep residual convolutional neural network to enhance the quality of video frames following compression, in particular inside the context of the VVC popular. By employing state-of-the-art neural network algorithms within the VVC framework, this method ensures tremendous enhancements in video excellent submit-compression, thereby addressing the essential necessities of video streaming services.

Chapter 3 - Optimizing VVC: In-Loop Filtering Enhancement Via Neural Network Integration:

This chapter focuses on improving the in-loop filtering step instead of post-processing by replacing traditional VVC loop filtering modules with a deep convolutional neural network applied during the coding process, improving the balance between video quality and bitrate.

Chapter 4 - VVC Optimization: Combining In-Loop Filtering Enhancement and Post-Processing Enhancement Via Neural Network Integration: In this chapter, a unified framework will be developed that integrates both in-loop filtering and post-processing enhancements using a residual dual-branch deep convolutional neural network (URDCNN).

Chapter 5 - Optimizing H.266/VVC Encoding using Genetic Algorithms: Performance optimization algorithms based on genetic algorithms will be presented to improve the efficiency of video in-stream coding (VVC), while balancing coding speed and video quality.

Chapter 6 - Conclusion and Future Recommendation: This chapter summarizes the main contributions of this thesis and highlights the most important areas for future research to focus on.

Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Заключение диссертации по теме «Другие cпециальности», Ибрагим Мурудж Халид Ибрагим

5.9 Conclusion

This chapter introduces two distinct genetic algorithm-driven methods to optimize the H.266/VVC video compression framework, each targeting different aspects of intra coding efficiency. The first algorithm focuses on optimizing the selection of coding tools and MTT partitions, aiming to reduce encoding time while maintaining video quality. By intelligently exploring the vast parameter space of coding options, the algorithm identifies the most efficient combinations for intra coding. The highlight is to achieve a compromise in quality and speed where in the first algorithm the speed and efficiency of the visual outcome is maintained, while in the second algorithm is more detailed focusing on quality and time by optimizing all encoding parameters. In this case, genetic algorithms enhance the PSNR and VMAF metric performance regarding the video quality and the time spent encoding it. The increase in bitrate is acceptable as the case presented offers better total encoding speed and video quality. In brief, the first method dominantly saves time while the quality of the output is minimally impacted and the second method focuses on optimizing the two metrics alongside video quality and time of encoding. The first

method is ideal for applications where speed is crucial, while the second method offers a more balanced solution, improving both efficiency and compression quality for high-performance video encoding systems.

Список литературы диссертационного исследования кандидат наук Ибрагим Мурудж Халид Ибрагим, 2025 год

References

[1] Agustsson, E., Mentzer, F., Tschannen, M., Cavigelli, L., Timofte, R., Benini, L., & Gool, L. V. (2017). Soft-to-hard vector quantization for end-to-end learning compressible representations. Advances in neural information processing systems, 30.

[2] Sidaty, N., Hamidouche, W., Déforges, O., Philippe, P., & Fournier, J. (2019, November). Compression performance of the versatile video coding: HD and UHD visual quality monitoring. In 2019 Picture Coding Symposium (PCS) (pp. 1-5). IEEE.

[3] Pinol, P., Martinez-Rach, M., Garrido, P., Lopez-Granado, O., & Malumbres, M. P. (2018). Error resilient coding techniques for video delivery over vehicular networks. Sensors, 18(10), 3495.

[4] Yang, R., Santamaria, M., Cricri, F., Zhang, H., Lainema, J., Youvalari, R. G., & Hannuksela, M. M. (2022, December). Low-precision post-filtering in video coding. In 2022 IEEE International Symposium on Multimedia (ISM) (pp. 137-140). IEEE.

[5] Nallappan, K., Guerboukha, H., Nerguizian, C., & Skorobogatiy, M. (2018). Live streaming of uncompressed HD and 4K videos using terahertz wireless links. IEEE Access, 6, 58030-58042.

[6] Khalek, A. A., Caramanis, C., & Heath, R. W. (2014). Delay-constrained video transmission: Quality-driven resource allocation and scheduling. IEEE Journal of Selected Topics in Signal Processing, 9(1), 6075.

[7] Shafi, R., Shuai, W., & Younus, M. U. (2020). 360-degree video streaming: A survey of the state of the art. Symmetry, 12(9), 1491.

[8] Jridi, M., Kumar Meher, P., & Alfalou, A. (2013). Zero-quantised discrete cosine transform coefficients prediction technique for intra-frame video encoding. IET image processing, 7(2), 165-173.

[9] Belyaev, E. (2023). An efficient compressive sensed video codec with inter-frame decoding and low-complexity intra-frame encoding. Sensors, 23(3), 1368.

[10] Shaikh, M. A., & Badnerkar, S. S. (2014). Video compression algorithm using motion compensation technique. International Journal of Advanced Research in Electronics and Communication Engineering, 3(6), 625-628.

[11] Deep, V., & Elarabi, T. (2017, April). HEVC/H. 265 vs. VP9 state-of-the-art video coding comparison for HD and UHD applications. In 2017 IEEE 30th Canadian Conference on Electrical and Computer Engineering (CCECE) (pp. 1-4). IEEE.

[12] Nishat, K., Gnawali, O., & Abdelhadi, A. (2020). Adaptive bitrate video streaming for wireless nodes: A survey. arXiv preprint arXiv:2008.00087.

[13] Spiteri, K., Urgaonkar, R., & Sitaraman, R. K. (2020). BOLA: Near-optimal bitrate adaptation for online videos. IEEE/ACM transactions on networking, 28(4), 1698-1711.

[14] Ibraheem, M. K. I., Dvorkovich, A. V., & Al-khafaji, I. M. A. (n.d.). Improving the Efficiency of Video Transmission in Computer Networks. International Journal of Recent Technology and Engineering (IJRTE). Advance online publication. doi:10.17762/ijritcc.v11i9.9319

[15] Djelouah, A., Campos, J., Schaub-Meyer, S., & Schroers, C. (2019). Neural inter-frame compression for video coding. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 64216429).

[16] Agustsson, E., Minnen, D., Johnston, N., Balle, J., Hwang, S. J., & Toderici, G. (2020). Scale-space flow for end-to-end optimized video compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8503-8512).

[17] Ladune, T., Philippe, P., Hamidouche, W., Zhang, L., & Déforges, O. (2020, September). Optical flow and mode selection for learning-based video coding. In 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP) (pp. 1-6). IEEE.

[18] Ballé, J., Minnen, D., Singh, S., Hwang, S. J., & Johnston, N. (2018). Variational image compression with a scale hyperprior. arXiv preprint arXiv:1802.01436.

[19] Minnen, D., Ballé, J., & Toderici, G. D. (2018). Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems, 31.

[20] Sigger, N., Al-Jawed, N., & Nguyen, T. (2022, May). Spatial-temporal autoencoder with attention network for video compression. In International Conference on Image Analysis and Processing (pp. 290300). Cham: Springer International Publishing.

[21] Habibian, A., Rozendaal, T. V., Tomczak, J. M., & Cohen, T. S. (2019). Video compression with rate-distortion autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 7033-7042).

[22] Pandey, C., Kumar, S., & Tiwari, R. (2012). An Innovative Approach towards the Video Compression Methodology of the H. 264 Codec: Using SPIHT Algorithms. International Journal of Soft Computing and Engineering (IJSCE).

[23] Halbach, T. (2003, October). The H. 264 video compression standard. In Proceedings of 6th Nordic Signal Processing Symposium. Cite-seer.

[24] Hua, K. L., Zhang, R., Comer, M., & Pollak, I. (2012). Inter frame video compression with large dictionaries of tilings: algorithms for tiling selection and entropy coding. IEEE transactions on circuits and systems for video technology, 22(8), 1136-1149.

[25] Bross, B., Chen, J., Ohm, J. R., Sullivan, G. J., & Wang, Y. K. (2021). Developments in international video coding standardization after AVC, with an overview of versatile video coding (VVC). Proceedings of the IEEE, 109(9), 1463-1493.

[26] Johnston, N., Vincent, D., Minnen, D., Covell, M., Singh, S., Chinen, T., ... & Toderici, G. (2018). Improved lossy image compression with priming and spatially adaptive bit rates for recurrent networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4385-4393).

[27] Lu, G., Cai, C., Zhang, X., Chen, L., Ouyang, W., Xu, D., & Gao, Z. (2020). Content adaptive and error propagation aware deep video compression. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part II 16 (pp. 456-472). Springer International Publishing.

[28] Hanzo, L., Cherriman, P., & Streit, J. (2007). Video compression and communications: from basics to H. 261, H. 263, H. 264, MPEG4 for DVB and HSDPA-style adaptive turbo-transceivers. John Wiley & Sons.

[29] Ebrahimi, T., & Horne, C. (2000). MPEG-4 natural video coding-An overview. Signal Processing: Image Communication, 15(4-5), 365-385.

[30] Dissanayake, M. B., & Abeyrathna, D. L. (2015). Performance comparison of HEVC and H. 264/AVC standards in broadcasting environments. Journal of Information Processing Systems, 11(3), 483494.

[31] Ibraheem, M. K. I., Dvorkovich, A. V., & Al-khafaji, I. M. A. (2024). A comprehensive literature review on image and video compression: Trends, algorithms, and techniques. Ingénierie des Systèmes d'Information, 29(3). doi: 10.18280/isi.290307.

[32] Lu, G., Zhang, X., Ouyang, W., Chen, L., Gao, Z., & Xu, D. (2020). An end-to-end learning framework for video compression. IEEE transactions on pattern analysis and machine intelligence, 43(10), 3292-3308.

[33] Djelouah, A., Campos, J., Schaub-Meyer, S., & Schroers, C. (2019). Neural inter-frame compression for video coding. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 64216429).

[34] Samuelsson, J., Choi, K., Chen, J., & Rusanovskyy, D. (2019, October). Mpeg-5 evc. In SMPTE 2019 (pp. 1-11). SMPTE.

[35] Battista, S., Meardi, G., Ferrara, S., Ciccarelli, L., Maurer, F., Conti, M., & Orcioni, S. (2022). Overview of the low complexity enhancement video coding (LCEVC) standard. IEEE Transactions on Circuits and Systems for Video Technology, 32(11), 7983-7995.

[36] Soltoggio, A., Stanley, K. O., & Risi, S. (2018). Born to learn: the inspiration, progress, and future of evolved plastic artificial neural networks. Neural Networks, 108, 48-67.

[37] Smys, S., Chen, J. I. Z., & Shakya, S. (2020). Survey on neural network architectures with deep learning. Journal of Soft Computing Paradigm (JSCP), 2(03), 186-194.

[38] Al-Azzawi Z.H.N., Nazarov, A. N., & Ibraheem M.K.I. (2023). Description of the optimization method and challenge of deep learning. Paper presented at the International Conference "Engineering & Telecommunication EN&T-2023", November 22-23, 2023. Retrieved from https://www.elibrary.ru/item.asp?id=65640795&pff=1.

[39] Agarap, A. F. (2018). Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375.

[40] Shewalkar, A., Nyavanandi, D., & Ludwig, S. A. (2019). Performance evaluation of deep neural networks applied to speech recognition: RNN, LSTM and GRU. Journal of Artificial Intelligence and Soft Computing Research, 9(4), 235-245.

[41] Laitinen, J., Mercat, A., Vanne, J., Tavakoli, H. R., Cricri, F., Aksu, E., & Hannuksela, M. (2022, July). Efficient Topology Coding and Payload Partitioning Techniques for Neural Network Compression (NNC) Standard. In 2022 IEEE International Conference on Multimedia and Expo Workshops (ICMEW) (pp. 1-4). IEEE.

[42] Kim, D. H., Jeong, J. Y., Lee, G., & Kim, J. G. (2024, May). Compression method of NeRF model using NNC and VVC. In International Workshop on Advanced Imaging Technology (IWAIT) 2024 (Vol. 13164, pp. 585-590). SPIE.

[43] Baroffio, L., Cesana, M., Redondi, A., Tagliasacchi, M., & Tubaro, S. (2014). Coding visual features extracted from video sequences. IEEE transactions on Image Processing, 23(5), 2262-2276.

[44] Wang, H., Gan, W., Hu, S., Lin, J. Y., Jin, L., Song, L., ... & Kuo, C. C. J. (2016, September). MCL-JCV: a JND-based H. 264/AVC video quality assessment dataset. In 2016 IEEE international conference on image processing (ICIP) (pp. 1509-1513). IEEE.

[45] Xue, T., Chen, B., Wu, J., Wei, D., & Freeman, W. T. (2019). Video enhancement with task-oriented flow. International Journal of Computer Vision, 127, 1106-1125.

[46] Xu, D., Lu, G., Yang, R., & Timofte, R. (2020, December). Learned image and video compression with deep neural networks. In 2020 IEEE International Conference on Visual Communications and Image Processing (VCIP) (pp. 1-3). IEEE.

[47] Han, J., Lombardo, S., Schroers, C., & Mandt, S. (2018). Deep probabilistic video compression.

[48] Liu, M. Y., Huang, X., Yu, J., Wang, T. C., & Mallya, A. (2021). Generative adversarial networks for image and video synthesis: Algorithms and applications. Proceedings of the IEEE, 109(5), 839-862.

[49] Ma, S., Zhang, X., Jia, C., Zhao, Z., Wang, S., & Wang, S. (2019). Image and video compression with neural networks: A review. IEEE Transactions on Circuits and Systems for Video Technology, 30(6), 1683-1698.

[50] Ding, D., Ma, Z., Chen, D., Chen, Q., Liu, Z., & Zhu, F. (2021). Advances in video compression system using deep neural network: A review and case studies. Proceedings of the IEEE, 109(9), 14941520.

[51] Gupta, R., Khanna, M. T., & Chaudhury, S. (2013). Visual saliency guided video compression algorithm. Signal Processing: Image Communication, 28(9), 1006-1022.

[52] Sullivan, G. J., & Wiegand, T. (2005). Video compression-from concepts to the H. 264/AVC standard. Proceedings of the IEEE, 93(1), 18-31.

[53] Tian, C., Xu, Y., Fei, L., & Yan, K. (2019). Deep learning for image denoising: A survey. In Genetic and Evolutionary Computing: Proceedings of the Twelfth International Conference on Genetic and Evolutionary Computing, December 14-17, Changzhou, Jiangsu, China 12 (pp. 563-572). Springer Singapore.

[54] Xu, L., Ren, J., Yan, Q., Liao, R., & Jia, J. (2015, June). Deep edge-aware filters. In International conference on machine learning (pp. 1669-1678). PMLR.

[55] Bazzani, L., Larochelle, H., & Torresani, L. (2016). Recurrent mixture density network for spatiotemporal visual attention. arXiv preprint arXiv:1603.08199.

[56] Jiang, L., Xu, M., Liu, T., Qiao, M., & Wang, Z. (2018). Deepvs: A deep learning based video saliency prediction approach. In Proceedings of the european conference on computer vision (eccv) (pp. 602-617).

[57] Zhang, S., Wei, K., Jia, H., Xie, X., & Gao, W. (2012, November). An efficient foreground-based surveillance video coding scheme in low bit-rate compression. In 2012 Visual Communications and Image Processing (pp. 1-6). IEEE.

[58] Sun, X., Yang, X., Wang, S., & Liu, M. (2020). Content-aware rate control scheme for HEVC based on static and dynamic saliency detection. Neurocomputing, 411, 393-405.

[59] Liu, D., Li, Y., Lin, J., Li, H., & Wu, F. (2020). Deep learning-based video coding: A review and a case study. ACM Computing Surveys (CSUR), 53(1), 1-35.

[60] Cui, W., Zhang, T., Zhang, S., Jiang, F., Zuo, W., & Zhao, D. (2018). Convolutional neural networks based intra prediction for HEVC. arXiv preprint arXiv:1808.05734.

[61] Jin, Z., An, P., & Shen, L. (2020). Video intra prediction using convolutional encoder decoder network. Neurocomputing, 394, 168-177.

[62] Chen, T., Liu, H., Shen, Q., Yue, T., Cao, X., & Ma, Z. (2017, December). Deepcoder: A deep neural network based video compression. In 2017 IEEE Visual Communications and Image Processing (VCIP) (pp. 1-4). IEEE.

[63] Chen, T., Liu, H., Ma, Z., Shen, Q., Cao, X., & Wang, Y. (2021). End-to-end learnt image compression via non-local attention optimization and improved context modeling. IEEE Transactions on Image Processing, 30, 3179-3191.

[64] Mentzer, F., Agustsson, E., Tschannen, M., Timofte, R., & Van Gool, L. (2018). Conditional probability models for deep image compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4394-4402).

[65] Ballé, J., Minnen, D., Singh, S., Hwang, S. J., & Johnston, N. (2018). Variational image compression with a scale hyperprior. arXiv preprint arXiv:1802.01436.

[66] Jia, C., Wang, S., Zhang, X., Wang, S., & Ma, S. (2017, December). Spatial-temporal residue network based in-loop filter for video coding. In 2017 IEEE Visual Communications and Image Processing (VCIP) (pp. 1-4). IEEE.

[67] Meng, X., Chen, C., Zhu, S., & Zeng, B. (2018, March). A new HEVC in-loop filter based on multichannel long-short-term dependency residual networks. In 2018 Data Compression Conference (pp. 187196). IEEE.

[68] Li, D., & Yu, L. (2019, May). An in-loop filter based on low-complexity CNN using residuals in intra video coding. In 2019 IEEE International Symposium on Circuits and Systems (ISCAS) (pp. 1-5). IEEE.

[69] Ding, D., Kong, L., Chen, G., Liu, Z., & Fang, Y. (2019). A switchable deep learning approach for in-loop filtering in video coding. IEEE Transactions on Circuits and Systems for Video Technology, 30(7), 1871-1887.

[70] Ding, D., Chen, G., Mukherjee, D., Joshi, U., & Chen, Y. (2019, November). A CNN-based in-loop filtering approach for AV1 video codec. In 2019 Picture Coding Symposium (PCS) (pp. 1-5). IEEE.

[71] Chen, G., Ding, D., Mukherjee, D., Joshi, U., & Chen, Y. (2019, September). AV1 in-loop filtering using a wide-activation structured residual network. In 2019 IEEE International Conference on Image Processing (ICIP) (pp. 1725-1729). IEEE.

[72] Li, T., Xu, M., Zhu, C., Yang, R., Wang, Z., & Guan, Z. (2019). A deep learning approach for multiframe in-loop filter of HEVC. IEEE Transactions on Image Processing, 28(11), 5663-5678.

[73] He, X., Hu, Q., Zhang, X., Zhang, C., Lin, W., & Han, X. (2018, October). Enhancing HEVC compressed videos with a partition-masked convolutional neural network. In 2018 25th IEEE International Conference on Image Processing (ICIP) (pp. 216-220). IEEE.

[74] Bao, W., Lai, W. S., Zhang, X., Gao, Z., & Yang, M. H. (2019). Memc-net: Motion estimation and motion compensation driven neural network for video interpolation and enhancement. IEEE transactions on pattern analysis and machine intelligence, 43(3), 933-948.

[75] Wang, X., Chan, K. C., Yu, K., Dong, C., & Change Loy, C. (2019). Edvr: Video restoration with enhanced deformable convolutional networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (pp. 0-0).

[76] Yang, R., Xu, M., Wang, Z., & Li, T. (2018). Multi-frame quality enhancement for compressed video. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 66646673).

[77] Tong, J., Wu, X., Ding, D., Zhu, Z., & Liu, Z. (2019, September). Learning-based multi-frame video quality enhancement. In 2019 IEEE International Conference on Image Processing (ICIP) (pp. 929-933). IEEE.

[78] Lu, M., Cheng, M., Xu, Y., Pu, S., Shen, Q., & Ma, Z. (2019, September). Learned quality enhancement via multi-frame priors for HEVC compliant low-delay applications. In 2019 IEEE International Conference on Image Processing (ICIP) (pp. 934-938). IEEE.

[79] Viitanen, M., Sainio, J., Mercat, A., Lemmetti, A., & Vanne, J. (2022). From HEVC to VVC: the first development steps of a practical intra video encoder. IEEE Transactions on Consumer Electronics, 68(2), 139-148.

[80] Bonnineau, C., Hamidouche, W., Fournier, J., Sidaty, N., Travers, J. F., & Déforges, O. (2022). Perceptual quality assessment of HEVC and VVC standards for 8K video. IEEE Transactions on Broadcasting, 68(1), 246-253.

[81] Zhao, Y., Lin, K., Wang, S., & Ma, S. (2022, May). Joint luma and chroma multi-scale CNN in-loop filter for Versatile Video Coding. In 2022 IEEE International Symposium on Circuits and Systems (ISCAS) (pp. 3205-3209). IEEE.

[82] Zhang, F., Ma, D., Feng, C., & Bull, D. R. (2021). Video compression with CNN-based postprocessing. IEEE MultiMedia, 28(4), 74-83.

[83] Ma, D., Zhang, F., & Bull, D. R. (2020). MFRNet: a new CNN architecture for post-processing and in-loop filtering. IEEE Journal of Selected Topics in Signal Processing, 15(2), 378-387.

[84] Zhang, H., Jung, C., Zou, D., & Li, M. (2023). WCDANN: A Lightweight CNN Post-Processing Filter for VVC-Based Video Compression. IEEE Access.

[85] Jin, K. H., McCann, M. T., Froustey, E., & Unser, M. (2017). Deep convolutional neural network for inverse problems in imaging. IEEE transactions on image processing, 26(9), 4509-4522.

[86] Zhou, D. X. (2020). Theory of deep convolutional neural networks: Downsampling. Neural Networks, 124, 319-327.

[87] Zhang, Y., Tian, Y., Kong, Y., Zhong, B., & Fu, Y. (2018). Residual dense network for image superresolution. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 24722481).

[88] Goldsborough, P. (2016). A tour of tensorflow. arXiv preprint arXiv:1610.01178.

[89] Mercat, A., Mäkinen, A., Sainio, J., Lemmetti, A., Viitanen, M., & Vanne, J. (2021). Comparative rate-distortion-complexity analysis of VVC and HEVC video codecs. IEEE Access, 9, 67813-67828.

[90] Li, Y., Zhang, Y., Timofte, R., Van Gool, L., Yu, L., Li, Y., ... & Wang, X. (2023). NTIRE 2023 challenge on efficient super-resolution: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1921-1959).

[91] Deshpande, R. G., Ragha, L. L., & Sharma, S. K. (2018). Video quality assessment through PSNR estimation for different compression standards. Indonesian Journal of Electrical Engineering and Computer Science, 11(3), 918-924.

[92] Lin, J., Akbari, M., Fu, H., Zhang, Q., Wang, S., Liang, J., ... & Tu, C. (2020). Learned variable-rate multi-frequency image compression using modulated generalized octave convolution. arXiv preprint arXiv:2009.13074.

[93] Li, J., Li, B., & Lu, Y. (2021). Deep contextual video compression. Advances in Neural Information Processing Systems, 34, 18114-18125.

[94] Wang, H., Ren, G., Ouyang, T., Zhang, J., Han, W., Liu, Z., & Chen, Z. (2022). Perceptual in-loop filter for image and video compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1770-1773).

[95] Ibraheem MKI, Dvorkovich AV, Al-Temimi AMS. Innovative Integration of Residual Networks for Enhanced In-loop Filtering in VVC Using Deep Convolutional Neural Networks. Computer Optics 2025; 49 (4): 692-701. DOI: 10.18287/2412-6179-CO-1572.

[96] Tsai, C. Y., Chen, C. Y., Yamakage, T., Chong, I. S., Huang, Y. W., Fu, C. M., ... & Lei, S. M. (2013). Adaptive loop filtering for video coding. IEEE Journal of Selected Topics in Signal Processing, 7(6), 934-945.

[97] Karczewicz, M., Hu, N., Taquet, J., Chen, C. Y., Misra, K., Andersson, K., ... & Chen, J. (2021). VVC in-loop filters. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 39073925.

[98] Ozcan, E., Adibelli, Y., & Hamzaoglu, I. (2013). A high performance deblocking filter hardware for high efficiency video coding. IEEE Transactions on Consumer Electronics, 59(3), 714-720.

[99] Pfaff, J., Filippov, A., Liu, S., Zhao, X., Chen, J., De-Luxan-Hernandez, S., ... & Van der Auwera, G. (2021). Intra prediction and mode coding in VVC. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 3834-3847.

[100] Ma, D., Zhang, F., & Bull, D. R. (2021). BVI-DVC: A training database for deep video compression. IEEE Transactions on Multimedia, 24, 3847-3858.

[101] Bouaafia, S., Messaoud, S., Khemiri, R., & Sayadi, F. E. (2021). VVC in-loop filtering based on deep convolutional neural network. Computational Intelligence and Neuroscience, 2021.

[102] Ibraheem, M. K. I., & Dvorkovich, A. V. (2024, March). Enhancing versatile video coding efficiency via post-processing of decoded frames using residual network integration in deep convolutional neural networks. In 2024 26th International Conference on Digital Signal Processing and its Applications (DSPA) (pp. 1-9). IEEE. doi:10.1109/DSPA.2024.10510065.

[103] Chen, S., Chen, Z., Wang, Y., & Liu, S. (2020, August). In-loop filter with dense residual convolutional neural network for VVC. In 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR) (pp. 149-152). IEEE.

[104] Kawamura, K., Y. Kidani, and S. Naito. "CE13-2.6/CE13-2.7: evaluation results of cnn based inloop filtering." Document JVET N 710 (2019): 19-27.

[105] Bouaafia, S., Messaoud, S., Khemiri, R., & Sayadi, F. E. (2021). VVC In-Loop Filtering Based on Deep Convolutional Neural Network. Computational Intelligence and Neuroscience, 2021(1), 9912839.

[106] Zou, N., Zhang, H., Cricri, F., Tavakoli, H. R., Lainema, J., Aksu, E., ... & Rahtu, E. (2021). Learned video compression with intra-guided enhancement and implicit motion information. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1870-1874).

[107] Ibraheem M. K. I., Al-Khafaji I. M., & Al-Azzawi Z. H. (2023). A genetic algorithm-based intra coding algorithm for H.266/VVC. Paper presented at the International Conference "Engineering & Telecommunication EN&T-2023", November 22-23, 2023. Retrieved from https://www.elibrary.ru/item.asp?id=65640761&pff=1.

[108] Dong, T., Kim, K., & Jang, E. S. (2021). Performance evaluation of the codec agnostic approach in MPEG-I video-based point cloud compression. IEEE Access, 9, 167990-168003.

[109] Pfaff, J., Filippov, A., Liu, S., Zhao, X., Chen, J., De-Luxân-Hernândez, S., ... & Van der Auwera, G. (2021). Intra prediction and mode coding in VVC. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 3834-3847.

[110] Saldanha, M., Sanchez, G., Marcon, C., & Agostini, L. (2021, December). Analysis of VVC intra prediction block partitioning structure. In 2021 International Conference on Visual Communications and Image Processing (VCIP) (pp. 1-5). IEEE.

[111] Zhao, H., Zhao, S., Shang, X., & Wang, G. (2023). A Fast Algorithm for VVC Intra Coding Based on the Most Probable Partition Pattern List. Applied Sciences, 13(18), 10381.

[112] Alam, T., Qamar, S., Dixit, A., & Benaida, M. (2020). Genetic algorithm: Reviews, implementations, and applications. arXiv preprint arXiv:2007.12673.

[113] "PyGAD: Genetic Algorithm in Python," GeneticAlgorithmPython, 2023. https://github.com/ahmedfgad/GeneticAlgorithmPython.

[114] HEVC-Projects/CPIH. https://github.com/HEVC-Projects/CPIH.

[115] "UVG dataset," 2023. https://www.kaggle.com/datasets/minhngt02/uvg-yuv.

[116] Choi, H., Hosseini, E., Alvar, S. R., Cohen, R. A., & Bajic, I. V. (2021). A dataset of labelled objects on raw video sequences. Data in Brief, 34, 106701.

[117] Soomro, K. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv: 1212.0402.

[118] Pont-Tuset, J., Perazzi, F., Caelles, S., Arbelâez, P., Sorkine-Hornung, A., & Van Gool, L. (2017). The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675.

[119] Alghamdi, N., Maddock, S., Marxer, R., Barker, J., & Brown, G. J. (2018). A corpus of audio-visual Lombard speech with frontal and profile views. The Journal of the Acoustical Society of America, 143(6), EL523-EL529.

[120] Blender Foundation, "Blender Foundation Open Movies," 2024. [Online]. Available: https://cloud.blender.org/p/gallery.

[121] Montgomery, C., & Lars, H. (1994). Xiph. org video test media (derfs collection). Online, https://media. xiph. org/video/derf, 6.

[122] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).

[123] Tsai, Y. H., Lu, C. R., Chen, M. J., Hsieh, M. C., Yang, C. M., & Yeh, C. H. (2023). Visual Perception Based Intra Coding Algorithm for H. 266/VVC. Electronics, 12(9), 2079.

[124] Chen, J. J., & Su, J. A. (2023). Fast H. 266/VVC intra-coding by mode inheritance. Multimedia Tools and Applications, 82(23), 36041-36065.

[125] Wang, Y., Huang, Q., Tang, B., Sun, H., & Li, X. (2023). Multiscale Motion-Aware and Spatial-Temporal-Channel Contextual Coding Network for Learned Video Compression. arXiv preprint arXiv:2310.12733.

[126] X. Zhang, Z. Dong, J. Hong and P. Cao, "Fast Algorithm for CU Split in H.266/VVC Intra Based on Texture Information," 2023 3rd International Conference on Intelligent Communications and Computing (ICC), Nanchang, China, 2023, pp. 115-118, doi: 10.1109/ICC59986.2023.10421268.

[127] Y. Wang, S. Feng, W. Zhang, K. Li and F. Yang, "Fast H.266/VVC Intra Coding by Early Skipping Joint Coding of Chroma Residuals," in IEEE Signal Processing Letters, vol. 31, pp. 2465-2469, 2024, doi: 10.1109/LSP.2024.3456631.

[128] Xiang, J. (2023). Subjective and objective image and video quality assessment methodologies and metrics (Doctoral dissertation, UNIVERSITY OF BRITISH COLUMBIA (Vancouver).

[129] Netflix, VMAF - video multi-method assessment fusion, https://github.com/Netflix/vmaf.

[130] Ibraheem, M. K. I., Abdalameer, A. K. I. M., & Naji, A. A. Z. H. (2024). A genetic approach-based intra coding algorithm for H. 266/VVC. Информатика и автоматизация, 23(3), 801-830. doi:10.15622/ia.23.3.6.

[131] Ibraheem, M. K., & Dvorkovich, A. V. (2024). Optimizing H. 266/VVC intra coding with a genetic algorithm: Balancing speed and quality. Fusion: Practice & Applications, 15(2). doi:10.54216/FPA.150201.

Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.