Методы самообучения для зависимых последовательных данных / Self-supervised learning for dependent sequential data тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Марусов Александр Эдуардович
- Специальность ВАК РФ00.00.00
- Количество страниц 111
Оглавление диссертации кандидат наук Марусов Александр Эдуардович
Contents
Page
Introduction
Chapter 1. Related works
1.1 Motivation of self-supervised learning
1.2 Generative self-supervised learning
1.2.1 Overview of generative SSL approaches
1.2.2 Temporal SSL generative approaches
1.2.3 Limitations of generative SSL approaches
1.3 Discriminative self-supervised learning
1.3.1 Overview of discriminative SSL approaches
1.3.2 Temporal discriminative SSL approaches
1.3.3 Limitations of discriminative SSL approaches
1.4 Research gaps
Chapter 2. Discriminative self-supervised approaches
2.1 Representation learning for sequential data
2.1.1 Problem statement
2.1.2 Data preprocessing
2.1.3 Methods
2.1.4 Downstream problems
2.1.5 Results
2.2 Discussion
Chapter 3. Contrastive self-supervised learning theory for
continuous dependent sequential data
3.1 A theoretical self-supervised learning approach for continuous dependent sequential data
3.1.1 Hard dependency
3.1.2 Soft dependency
3.2 Estimated similarity matrices
3.2.1 Semantic independence
Page
3.2.2 Continuous dependent sequential data
3.3 DepTS2Vec
3.4 Discussion and future work
Chapter 4. Applications for temporal and spatio-temporal data
4.1 Temporal data
4.1.1 UCR & UEA time series benchmarks
4.2 Spatio-temporal data
4.2.1 Drought prediction (supervised approach)
4.2.2 Weather prediction (self-supervised approach)
4.3 Discussion and future work
Conclusion
List of symbols and abbreviations
Bibliography
List of Figures
List of Tables
Appendix
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Методы обучения представлений для оптимальных процедур детектирования разладок / Representation learning methods for optimal change point detection procedures2025 год, кандидат наук Романенкова Евгения Дмитриевна
Доменная адаптация глубоких сверточных нейросетей для обработки медицинских изображений / Domain Adaptation of Deep Convolutional Neural Networks in Medical Imaging2026 год, кандидат наук Широких Борис Николаевич
Псевдобулевский полиномиальный подход к решению задач компьютерного зрения / Pseudo- Boolean Polynomial approach to solving Computer Vision tasks2025 год, кандидат наук Чикаке Тендай Мапунгвана
Исследование вариантов трансформера для различных задач обработки длинных документов/ Investigation of transformer options for various long documents processing tasks2024 год, кандидат наук Аль Адел Ариж
Оценка неопределенности в задачах обработки естественного языка2026 год, кандидат наук Важенцев Артем Андреевич
Введение диссертации (часть автореферата) на тему «Методы самообучения для зависимых последовательных данных / Self-supervised learning for dependent sequential data»
Introduction
Relevance of the work. Representation learning refers to a class of machine learning methods designed to automatically extract informative and structured representations from raw data. The primary objective is to transform high-dimensional inputs into compressed feature vectors, known as representations or embeddings, that capture underlying patterns and dependencies, thereby facilitating more effective and efficient learning in subsequent tasks.
By extracting the main structure of the data, high-quality representations form the basis for solving many downstream tasks, such as classification and prediction. To be broadly effective, such embeddings are expected to exhibit several essential diverse properties [1]:
- Generalization. Representations should support multiple downstream tasks without requiring extensive re-engineering. This adaptability reduces reliance on large labeled datasets.
- Invariance. Representations remain stable under transformations of the input (so-called augmentations) that are irrelevant to the semantic information for an object (e.g., noise, scaling, or rotation). This ensures robustness and emphasizes task-relevant information.
- Hierarchical organization. Representations capture information at multiple levels of abstraction, enabling both fine-grained and high-level data interpretation. This supports diverse prospective applications ranging from recognition to complex decision-making.
- Disentanglement. Different dimensions of the representation correspond to distinct factors of variation in the data. This facilitates interpretability and improves transferability across tasks.
- Compactness. Essential information is encoded in a compressed yet expressive form, which improves computational efficiency while preserving predictive performance.
Representation learning is especially important for time series because these data often contain complex patterns spread across different scales. High-quality representations enable models to capture both short-term fluctuations and long-term dependencies, which are fundamental for achieving accurate forecasting and classification. Furthermore, robust representations enhance the resilience of models to
noise, distributional shifts, and distortions that are frequently encountered in real-world temporal data.
To ensure effectiveness across a wide range of time-series downstream tasks, embeddings should exhibit several universal properties specific to the sequential dependent data [2].
— Generalization. High-quality embeddings for time series should generalize by transferring temporal patterns and dependencies across diverse downstream tasks—such as classification, forecasting, and anomaly detection—thereby minimizing the reliance on large amounts of labeled data.
— Invariance. Representations of time series should remain stable under temporal transformations that do not alter the underlying semantics, such as small temporal shifts or scaling.
— Hierarchical organization High-quality embeddings for time series should capture information across multiple temporal scales, effectively representing both fine-grained local dynamics and long-range global dependencies that characterize sequential data.
— Disentanglement Effective embeddings for time series should disentangle different sources of variation—such as trend, seasonality, and irregular fluctuations—into distinct latent components, reflecting the structured nature of sequential temporal data.
— Compactness In spatio-temporal cases, compact embeddings reduce size while preserving key information.
Those representations can be learned in the supervised regime. A typical neural network in a supervised task has two main parts: an encoder that creates feature representations, and a classifier that predicts a label based on those features. When trained on a supervised task, the network learns an encoder that automatically generates useful representations of the input signals with transfer capabilities. Nevertheless, for time series data, comprehensive labeled datasets spanning multiple downstream tasks and diverse temporal patterns are rarely available, making supervised learning approaches challenging.
Due to the limited availability of labeled data, the self-supervised learning (SSL) paradigm leverages large amounts of unlabeled data to pretrain the encoder by solving pretext tasks [2]. Once pretrained, the encoder is then used to support downstream supervised tasks with representations. Because labeled data are scarce,
using a pretrained encoder typically leads to better performance on the target task compared to training an encoder from scratch using only labelled data [3].
SSL algorithms are broadly categorized into two main groups: generative and discriminative approaches. We focus on discriminative SSL methods, which aim to produce representations that are particularly well-suited for downstream classification tasks. This is especially evident in contrastive self-supervised learning, a central subtype within this group, which implicitly addresses the classification problem by constructing pairs of instances: positive pairs, representing semantically similar samples, and negative pairs, corresponding to samples from different semantic classes [4]. Then, the contrastive learning aims to draw the representations of positive pairs closer together in the latent space, while simultaneously pushing the representations of negative pairs further apart during encoder training.
SSL techniques were originally developed within the computer vision domain. A central assumption in many of these methods is semantic independence, which presumes that each image belongs to a distinct semantic class. Under this assumption, positive pairs are typically constructed by applying a series of data augmentation techniques, such as random cropping, color jittering, or Gaussian noise, to a single input instance. These augmentations produce different views of the same object that are treated as semantically similar. All other images in the dataset are treated as negative examples for a current signal, regardless of any latent similarities.
However, in domains such as time series or other structured, dependent data, the semantic independence assumption no longer holds. Samples in these domains may exhibit temporal or contextual dependencies, implying that the construction of positive and negative pairs must take such relationships into consideration. Thus SSL methods need to be adapted accordingly to reflect the underlying data structure. Approaches based on semantic independence treat all samples as dissimilar, with each sample forming a positive pair only with itself. Meanwhile the methods that assign a smooth similarity measure between objects define contrasting procedure in a soft way. Existing techniques either ignore semantic information altogether or incorporate it in a heuristic manner without a solid theoretical foundation.
Therefore, self-supervised learning for dependent data requires a principled and interpretable approach that explicitly accounts for data dependencies. Indeed, by explicitly taking data dependencies into account, this approach allows the learned representations to capture more meaningful patterns than traditional methods that treat samples as independent. This leads to better performance in downstream
tasks such as forecasting, classification, and anomaly detection. In addition, the interpretable design of the approach helps to understand how temporal or contextual relationships shape the embeddings, supporting clearer insights and more informed model design. Overall, combining a solid theoretical basis with interpretability makes the representations both reliable and useful across different applications.
Dissertation goals and problems. The main goal of the research is to develop a theoretically grounded SSL approach suitable for dependent data. Specifically, we want to obtain theoretically-grounded loss functions for different dependency types, that are encountered in time series data. We aim to reach the goal through the following problems:
1. We aim to select self-supervised learning approaches that are well-suited for handling dependent sequential data, including both time series and spatio-temporal data.
2. To develop a theoretically grounded self-supervised learning approach tailored for dependent data, which is capable of modeling and handling diverse forms of inter-sample correlations.
3. To integrate the developed grounded objectives in the previous step into a self-supervised learning method tailored for temporal and spatio-temporal tasks.
4. To enhance prediction accuracy in comparison with existing approaches for diverse industrial downstream tasks by developing a dependency-aware method.
Scientific novelty. This dissertation provides several distinct contributions to knowledge, the primary of which are:
1. Although many self-supervised learning methods for dependent data have been proposed [2; 3; 5—9], there is still no universal solution for handling such data. In this thesis, I introduce a fundamentally new approach by developing a unified theoretical approach that accounts for various types of inter-correlations between samples.
2. At present, only a few studies focus on the theory of self-supervised learning, [10—13], and these are primarily limited to the case of semantically independent data. In this thesis, I formulate and solve an optimization problem specifically designed for dependent data. This leads to theoretically grounded loss functions that are tailored to different types of dependencies.
3. Previous approaches for dependent data either use loss functions suitable for the semantic independence case or objective functions that are heuristic-based. For the first time, a novel DepTS2Vec approach, theoretically based and interpretable, was proposed.
4. The introduced DepTS2Vec approach surpasses current state-of-the-art methods across a wide range of temporal and spatio-temporal benchmarks.
Theoretical and practical significance. From a theoretical perspective, this thesis introduces, for the first time, a unified theoretical approach for self-supervised learning on dependent data. It is the first approach to establish a general foundation for handling various types of dependencies. Within this approach, we developed theoretically grounded, dependency-aware loss functions tailored to different dependency structures.
Regarding the practical significance, the proposed drought prediction module was integrated into the Skoltech AI platform to address key ESG challenges across several regions of the Russian Federation. The results have proven valuable in sectors such as petroleum and agriculture. Notably, for drought predictions one and twelve months ahead, our models closed the performance gap by 54% and 16%, respectively, relative to classic approaches. The importance of these practical outcomes is highlighted by their mention in TASS, Vedomosti, and Skoltech's 2024 Annual Report.
Methodology and research methods. This research explores machine learning and deep learning methods to achieve the defined objectives. Specifically, we examine both supervised and self-supervised approaches for learning meaningful representations from dependent data. To ensure accurate and efficient performance, the proposed methods are grounded in optimization theory. The use of precise mathematical language enables a clear and concise expression of the results and their theoretical properties.
Propositions submitted for the defense. This section presents the core statements of this dissertation, formally submitted for defense. They encapsulate the original scientific contributions substantiated throughout the research. 1. The proposed theoretical framework enables modeling a range of dependencies in time series data — both hard and soft — and facilitates the construction of corresponding ground truth similarity matrices. These matrices capture how similar or close different samples are to one another [14].
2. The formulated optimization problem, tailored to the introduced dependency types, yields a closed-form estimated similarity matrix for continuous dependent data [14; 15].
3. The DepTS2Vec method — a theoretically grounded approach — relies on a dependency-aware family of loss functions. These objectives are defined in terms of the proposed estimated and ground truth similarity matrices [15].
4. The designed DepTS2Vec approach consistently outperforms state-of-the-art methods across a diverse set of temporal and spatio-temporal tasks.
Compliance with Passport of Speciality 1.2.1. Artificial intelligence and machine learning. Below is the correspondence between the propositions for the defense and items from the specialty passport.
1. Items 1 and 2 correspond to the point 15: "Mathematical research in statistics, logic, algebra, topology, functional analysis, and other domains, oriented toward addressing problems in artificial intelligence and machine learning".
2. Item 3 corresponds to the point 4: "Development of methods, algorithms, and creation of artificial intelligence and machine learning systems for processing and analyzing natural language texts, images, speech, biomedical data, and other specialized types of data".
3. Item 4 corresponds to the point 17: "Research in the field of multilayer algorithmic structures, including multilayer neural networks."
Personal contribution. The content of this dissertation and the main claims submitted for defense represent the author's original contribution to the published research. The results were obtained either independently by the author or with his direct involvement. The author's role included formulating research objectives, conducting a comprehensive review of relevant literature, designing and performing experiments, and interpreting the results. The problem statements were proposed by the scientific supervisor, A. Zaytsev, and the research findings were discussed in collaboration with co-authors.
Approbation of the results. The results of the thesis have been presented at the following conferences and seminars:
1. Scientific Seminar of Laboratory of applied research Skoltech-Sberbank (LARSS) (2023, 2024, 2025), Moscow, Russia.
2. Scientific Seminar at the Institute of Control Sciences, RAS (2025), Moscow, Russia.
3. 67th All-Russian Scientific Conference of MIPT, (2025), Moscow, Russia.
4. 32nd International Conference Of The Students, Postgraduates And Youngs Scientists "Lomonosov", (2025), Moscow, Russia.
Structure of the dissertation. The dissertation consists of an introduction, four chapters, a conclusion, a bibliography, a list of symbols and abbreviations, a list of figures, and supplementary material.
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Оценка и прогнозирование методами машинного обучения гниения плодовых растений на раннем этапе после сбора урожая (Early Postharvest Decay Assessment and Prediction in Fruit Plants by Machine Learning Methods)2026 год, кандидат наук Стасенко Никита Андреевич
Эффективные методы мультиязычного текстового переноса стиля/Efficient Multilingual Text Style Transfer Methods2026 год, кандидат наук Московский Даниил Алексеевич
Analysis of the vibration effects on combined cycle power plants' mechanical parts /Оценка воздействия вибрации на механическое оборудование электростанций комбинированного цикла2025 год, кандидат наук Аль-Текрити Ватбан Халид Фахми
"The role of executive functions in emotion regulation"2022 год, кандидат наук Мохаммед Абдул-Рахеем
Гарантии обучения и эффективный вывод в задачах структурного предсказания2024 год, кандидат наук Струминский Кирилл Алексеевич
Заключение диссертации по теме «Другие cпециальности», Марусов Александр Эдуардович
Conclusion
This thesis presents a theoretically grounded self-supervised learning methodology for sequential data characterized by inherent dependencies. Through an extensive empirical evaluation across diverse datasets—spanning multiple modalities and dependency structures, including temporal and spatio-temporal—we demonstrate the practical efficacy of the proposed approach in comparison to contemporary self-supervised learning techniques.
In Chapter 1, we establish the foundational concepts of self-supervised learning and provide a comprehensive review of the relevant literature, encompassing both generative and discriminative approaches.
Chapter 2 presents a comparative analysis of multiple contrastive and non-contrastive self-supervised learning approaches applied to a representation learning problem within oil and gas data. The evaluation focused on the quality of the learned representations for both clustering and a suite of classification downstream tasks. The results indicate that both methodological families achieve approximately comparable performance on these benchmarks. However, contrastive methods possess a more direct and natural theoretical interpretation compared to their non-contrastive counterparts. Consequently, the contrastive paradigm was selected as the basis for the subsequent research detailed in this thesis.
Chapter 3 describes the core theoretical contribution of this work. Specifically, we formulate a theoretical framework designed to model dependencies within sequential data in a theoretically optimal manner. We formally consider both hard and soft dependency structures, deriving theoretically optimal similarity matrices for each case. These matrices are subsequently used to construct corresponding objective functions. The resulting loss functions form the foundation of a novel self-supervised learning method, designated DepTS2Vec.
Chapter 4 provides a comprehensive experimental validation of the proposed DepTS2Vec method across a diverse suite of temporal and spatio-temporal downstream tasks. These tasks span multiple applied domains, including human activity recognition, ECG and EEG analysis, and traffic prediction. Furthermore, the evaluation incorporates several weather-related spatio-temporal forecasting problems. The experimental results demonstrate that our approach consistently outperforms
established baseline and state-of-the-art methods across this broad spectrum of benchmarks.
The main contributions of this dissertation can be summarized as follows:
- We propose a theoretical framework for modeling a wide range of dependencies in time series data, distinguishing between two categories: hard and soft dependencies. Each type captures a different notion of closeness within the data. This formulation provides a unified way to represent the semantic relationships inherent in sequential signals. To the best of our knowledge, this is the first work to introduce such dependency types in the context of dependent sequential data.
- For each defined dependency type, we designed a theoretically grounded loss function. This was achieved by deriving an estimated similarity matrix through the solution of an appropriate optimization problem. As a result, we obtained a family of dependency-aware objective functions. Building on this foundation, we introduced our novel DepTS2Vec self-supervised learning framework for dependent sequential data.
- We evaluated the proposed DepTS2Vec method across a range of temporal and spatio-temporal downstream tasks. For the temporal setting, we conducted experiments on the widely used UCR [140] and UEA [141] time series benchmarks, where our approach surpassed TS2Vec [3] by 2.08% and 4.17% in accuracy, respectively. In the drought classification task, where spatio-temporal continuity is a key factor, DepTS2Vec delivered a 7% increase in the ROC-AUC score.
The primary limitation of our approach lies in its reliance on the data continuity assumption. Nonetheless, this condition is met in numerous real-world scenarios, including climate-related applications, human activity recognition, and other domains characterized by smoothly evolving signals.
The proposed method offers a combination of theoretical rigor and practical effectiveness. From the theoretical perspective, the set of objective functions we propose is solidly based on a clear and well-defined dependency framework. On the practical side, the method demonstrates consistently superior performance across a variety of downstream tasks, encompassing both temporal and spatio-temporal settings.
Список литературы диссертационного исследования кандидат наук Марусов Александр Эдуардович, 2026 год
Bibliography
1. Bengio, Y. Representation learning: A review and new perspectives / Y. Ben-gio, A. Courville, P. Vincent // IEEE Transactions on Pattern Analysis and Machine Intelligence. — 2013. — Vol. 35, no. 8. — P. 1798—1828.
2. Woo, G. CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting / G. Woo, C. Liu, et al. // arXiv preprint arXiv:2202.01575. — 2022. — https://arxiv.org/abs/2202.01575 (accessed 18.09.2025).
3. Yue, Z. TS2Vec: Towards universal representation of time series / Z. Yue, Y. Wang, et al. // Proceedings of the AAAI conference on artificial intelligence. Vol. 36. — 2022. — P. 8980—8987.
4. Joshi, S. Data-efficient contrastive self-supervised learning: Most beneficial examples for supervised learning contribute the least / S. Joshi, B. Mirza-soleiman // International Conference on Machine Learning. — PMLR. 2023. — P. 15356—15370.
5. Oord, A. v. d. Representation learning with contrastive predictive coding / A. v. d. Oord, Y. Li, O. Vinyals // arXiv preprint arXiv:1807.03748. — 2018. — https://arxiv.org/abs/1807.03748 (date accessed 18.09.2025).
6. Zhang, K. Self-supervised learning for time series analysis: Taxonomy, progress, and prospects / K. Zhang, Q. Wen, et al. // IEEE Transactions on Pattern Analysis and Machine Intelligence. — 2024. — Vol. 46, no. 10. — P. 6775—6794.
7. Shen, L. Timeseries anomaly detection using temporal hierarchical one-class network / L. Shen, Z. Li, J. Kwok // Advances in Neural Information Processing Systems. — 2020. — Vol. 33. — P. 13016—13026.
8. Xi, L. Semi-supervised time series classification model with self-supervised learning / L. Xi, Z. Yun, et al. // Engineering Applications of Artificial Intelligence. — 2022. — Vol. 116. — P. 105331.
9. Lee, S. Soft contrastive learning for time series / S. Lee, T. Park, K. Lee // 12th International Conference on Learning Representations, ICLR 2024. — 2024.
10. Balestriero, R. Contrastive and non-contrastive self-supervised learning recover global and local spectral embedding methods / R. Balestriero, Y. Le-Cun //. Vol. 35. — 2022. — P. 26671—26685.
11. Jing, L. Understanding dimensional collapse in contrastive self-supervised learning / L. Jing, P. Vincent, et al. //. — 2021. — https://arxiv.org/abs/2110.09348 (date accessed 18.09.2025).
12. Tian, Y. Understanding self-supervised learning dynamics without contrastive pairs / Y. Tian, X. Chen, S. Ganguli // International Conference on Machine Learning. — PMLR. 2021. — P. 10268—10278.
13. Ji, W. The power of contrast for feature learning: A theoretical analysis / W. Ji, Z. Deng, et al. // Journal of Machine Learning Research. — 2023. — Vol. 24, no. 330. — P. 1—78.
14. Marusov, A. Theoretically Justified Contrastive Self-Supervised Methods for Continuous Dependent Data / A. Marusov, A. Zaytsev // Doklady Mathematics. — 2025. — Vol. 527. — P. 192—205.
15. Marusov, A. A theoretical framework for self-supervised contrastive learning for continuous dependent data / A. Marusov, A. Yuhay, A. Zaytsev // ICDM. — 2025. — P. 1—10.
16. Jaiswal, A. A survey on contrastive self-supervised learning / A. Jaiswal, A. R. Babu, et al. // Technologies. — 2020. — Vol. 9, no. 1. — P. 2.
17. Krishnan, R. Self-supervised learning in medicine and healthcare / R. Krish-nan, P. Rajpurkar, E. J. Topol // Nature Biomedical Engineering. — 2022. — Vol. 6, no. 12. — P. 1346—1352.
18. Liu, X. Self-supervised learning: Generative or contrastive / X. Liu, F. Zhang, et al. // IEEE transactions on knowledge and data engineering. — 2021. — Vol. 35, no. 1. — P. 857—876.
19. Gui, J. A survey on self-supervised learning: Algorithms, applications, and future trends / J. Gui, T. Chen, et al. // IEEE Transactions on Pattern Analysis and Machine Intelligence. — 2024. — Vol. 46, no. 12. — P. 9052—9071.
20. Zhai, X. S4l: Self-supervised semi-supervised learning / X. Zhai, A. Oliver, et al. // Proceedings of the IEEE/CVF international conference on computer vision. — 2019. — P. 1476—1485.
21. Liu, Y. Graph self-supervised learning: A survey / Y. Liu, M. Jin, et al. // IEEE transactions on knowledge and data engineering. — 2022. — Vol. 35, no. 6. — P. 5879—5900.
22. Cao, H. Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark / H. Cao // arXiv preprint arXiv:2406.01607. — 2024. — https://arxiv.org/abs/2406.01607 (date accessed 18.09.2025).
23. Qiang, W. On the Universality of Self-Supervised Representation Learning / W. Qiang, J. Wang, et al. // arXiv preprint arXiv:2405.01053. — 2024. — https://arxiv.org/abs/2405.01053 (date accessed 18.09.2025).
24. Feng, Y. Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation / Y. Feng, Y. Sun, et al. // arXiv preprint arXiv:2505.24754. — 2025. — https://arxiv.org/abs/2505.24754 (date accessed 18.09.2025).
25. Das, L. There is more to graphs than meets the eye: Learning universal features with self-supervision / L. Das, S. Munikoti, et al. // arXiv preprint arXiv:2305.19871. — 2023. — https://arxiv.org/abs/2305.19871 (date accessed 18.09.2025).
26. Limkonchotiwat, P. An efficient self-supervised cross-view training for sentence embedding / P. Limkonchotiwat, W. Ponwitayarat, et al. // Transactions of the Association for Computational Linguistics. — 2023. — Vol. 11. — P. 1572—1587.
27. Lugo, L. Efficiency-oriented approaches for self-supervised speech representation learning / L. Lugo, V. Vielzeuf // International Journal of Speech Technology. — 2024. — Vol. 27, no. 3. — P. 765—779.
28. Garcia, K. Efficient Hierarchical Contrastive Self-supervising Learning for Time Series Classification via Importance-aware Resolution Selection / K. Garcia, J. M. Perez, Y. Gao // 2024 IEEE International Conference on Big Data (BigData). — IEEE. 2024. — P. 880—889.
29. Wang, B. Efficient sentence embedding via semantic subspace analysis / B. Wang, F. Chen, et al. // 2020 25th International Conference on Pattern Recognition (ICPR). — IEEE. 2021. — P. 119—125.
30. Pepino, L. EnCodecMAE: Leveraging neural codecs for universal audio representation learning / L. Pepino, P. Riera, L. Ferrer // arXiv preprint arXiv:2309.07391. — 2023. — https://arxiv.org/pdf/2309.07391 (date accessed 18.09.2025).
31. LeCun, Y. Deep learning / Y. LeCun, Y. Bengio, G. Hinton // Nature. — 2015. — Vol. 521, no. 7553. — P. 436—444.
32. Schmidhuber, J. Deep learning in neural networks: An overview / J. Schmid-huber // Neural networks. — 2015. — Vol. 61. — P. 85—117.
33. Zeiler, M. D. Visualizing and understanding convolutional networks / M. D. Zeiler, R. Fergus // European conference on computer vision. — Springer. 2014. — P. 818—833.
34. Sun, C. Revisiting unreasonable effectiveness of data in deep learning era / C. Sun, A. Shrivastava, et al. // Proceedings of the IEEE international conference on computer vision. — 2017. — P. 843—852.
35. Misra, I. Self-supervised learning of pretext-invariant representations / I. Misra, L. v. d. Maaten // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. — 2020. — P. 6707—6717.
36. Hendrycks, D. Using self-supervised learning can improve model robustness and uncertainty / D. Hendrycks, M. Mazeika, et al. // Advances in Neural Information Processing Systems. — 2019. — Vol. 32. — P. 1—12.
37. Rani, V. Self-supervised learning: A succinct review / V. Rani, S. T. Nabi, et al. // Archives of Computational Methods in Engineering. — 2023. — Vol. 30, no. 4. — P. 2761—2775.
38. Wang, Y. Self-supervised learning in remote sensing: A review / Y. Wang, C. M. Albrecht, et al. // IEEE Geoscience and Remote Sensing Magazine. — 2022. — Vol. 10, no. 4. — P. 213—247.
39. Schiappa, M. C. Self-supervised learning for videos: A survey / M. C. Schi-appa, Y. S. Rawat, M. Shah // ACM Computing Surveys. — 2023. — Vol. 55, 13s. — P. 1—37.
40. Baevski, A. Data2vec: A general framework for self-supervised learning in speech, vision and language / A. Baevski, W.-N. Hsu, et al. // International Conference on Machine Learning. Vol. 162. — PMLR. 2022. — P. 1298—1312.
41. Yu, W. Dict-bert: Enhancing language model pre-training with dictionary / W. Yu, C. Zhu, et al. // arXiv preprint arXiv:2110.06490. — 2021. — https://arxiv.org/pdf/2110.06490 (date accessed 18.09.2025).
42. Kumar, M. Dual patchnorm / M. Kumar, M. Dehghani, N. Houlsby // arXiv preprint arXiv:2302.01327. — 2023. — https://arxiv.org/pdf/2302.01327 (date accessed 18.09.2025).
43. Chen, X. DsEE: Dually sparsity-embedded efficient tuning of pre-trained language models / X. Chen, T. Chen, et al. // arXiv preprint arXiv:2111.00160. — 2021. — https://arxiv.org/pdf/2111.00160 (date accessed 18.09.2025).
44. Park, S. M. Trak: Attributing model behavior at scale / S. M. Park, K. Georgiev, et al. // arXiv preprint arXiv:2303.14186. — 2023. — https://arxiv.org/abs/2303.14186 (date accessed 18.09.2025).
45. Dong, C. Efficientbert: Progressively searching multilayer perceptron via warm-up knowledge distillation / C. Dong, G. Wang, et al. // arXiv preprint arXiv:2109.07222. — 2021. — https://arxiv.org/pdf/2109.07222 (date accessed 18.09.2025).
46. Jing, L. Self-supervised spatiotemporal feature learning via video rotation prediction / L. Jing, X. Yang, et al. // arXiv preprint arXiv:1811.11387. — 2018. — https://arxiv.org/abs/1811.11387, (date accessed 18.09.2025).
47. Tao, L. Self-supervised video representation using pretext-contrastive learning / L. Tao, X. Wang, T. Yamasaki // arXiv preprint arXiv:2010.15464. — 2020. — Vol. 2. — P. 2. — https://arxiv.org/pdf/2010.15464 (date accessed 18.09.2025).
48. Wang, J. Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics / J. Wang, J. Jiao, et al. // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. — 2019. — P. 4006—4015.
49. Ahsan, U. Video jigsaw: Unsupervised learning of spatiotemporal context for video action recognition / U. Ahsan, R. Madhok, I. Essa // 2019 IEEE Winter Conference on Applications of Computer Vision (WACV). — IEEE. 2019. — P. 179—189.
50. Huo, Y. Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw / Y. Huo, M. Ding, et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). — 2021. — P. 1693—1702.
51. Kim, D. Self-supervised video representation learning with space-time cubic puzzles / D. Kim, D. Cho, I. S. Kweon // Proceedings of the AAAI conference on artificial intelligence. Vol. 33. — 2019. — P. 8545—8552.
52. Park, D. Self-supervised contextual data augmentation for natural language processing / D. Park, C. W. Ahn // Symmetry. — 2019. — Vol. 11, no. 11. — P. 1393.
53. Yao, B. Self-supervised pre-trained neural network for quantum natural language processing / B. Yao, P. Tiwari, Q. Li // Neural Networks. — 2025. — Vol. 184. — P. 107004.
54. Lewis, M. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension / M. Lewis, Y. Liu, et al. // arXiv preprint arXiv:1910.13461. — 2019. — https://arxiv.org/abs/1910.13461 (date accessed 18.09.2025).
55. Yang, Z. Xlnet: Generalized autoregressive pretraining for language understanding / Z. Yang, Z. Dai, et al. // Advances in Neural Information Processing Systems. — 2019. — Vol. 32. — P. 1—11.
56. Baevski, A. wav2vec 2.0: A framework for self-supervised learning of speech representations / A. Baevski, Y. Zhou, et al. // Advances in Neural Information Processing Systems. — 2020. — Vol. 33. — P. 12449—12460.
57. Hsu, W.-N. Hubert: Self-supervised speech representation learning by masked prediction of hidden units / W.-N. Hsu, B. Bolte, et al. // IEEE/ACM transactions on audio, speech, and language processing. — 2021. — Vol. 29. — P. 3451—3460.
58. Chen, S. Wavlm: Large-scale self-supervised pre-training for full stack speech processing / S. Chen, C. Wang, et al. // IEEE Journal of Selected Topics in Signal Processing. — 2022. — Vol. 16, no. 6. — P. 1505—1518.
59. Chung, Y.-A. An unsupervised autoregressive model for speech representation learning / Y.-A. Chung, W.-N. Hsu, et al. // arXiv preprint arXiv:1904.03240. — 2019. — https://arxiv.org/abs/1904.03240 (date accessed 18.09.2025).
60. Mohamed, A. Self-supervised speech representation learning: A review / A. Mohamed, H.-y. Lee, et al. // IEEE Journal of Selected Topics in Signal Processing. — 2022. — Vol. 16, no. 6. — P. 1179—1210.
61. Tian, Y. What makes for good views for contrastive learning? / Y. Tian, C. Sun, et al. // Advances in Neural Information Processing Systems. —
2020. — Vol. 33. — P. 6827—6839.
62. Suarez, P. J. O. A monolingual approach to contextualized word embeddings for mid-resource languages / P. J. O. Suarez, L. Romary, B. Sagot // arXiv preprint arXiv:2006.06202. — 2020. — https://arxiv.org/abs/2006.06202 (date accessed 18.09.2025).
63. Provable guarantees for self-supervised deep learning with spectral contrastive loss / J. Z. HaoChen [et al.] // Advances in Neural Information Processing Systems. — 2021. — Vol. 34. — P. 5000—5011.
64. Xu, L. Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation / L. Xu, Z. Song, et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision. — 2023. — P. 2228—2238.
65. Kim, J. Collective relevance labeling for passage retrieval / J. Kim, M. Kim, S.-w. Hwang // arXiv preprint arXiv:2205.03273. — 2022. — https://arxiv.org/pdf/2205.03273 (https://arxiv.org/pdf/2205.03273).
66. Yang, J. Focal self-attention for local-global interactions in vision transformers. arXiv 2021 / J. Yang, C. Li, et al. // arXiv preprint arXiv:2107.00641. —
2021. — https://arxiv.org/abs/2107.00641 (date accessed 18.09.2025).
67. Wang, T. Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix / T. Wang, W. Jiang, et al. // International Conference on Machine Learning. Vol. 202. — PMLR. 2022. — P. 22680—22690.
68. Gidaris, S. Unsupervised representation learning by predicting image rotations / S. Gidaris, P. Singh, N. Komodakis // arXiv preprint arXiv:1803.07728. — 2018. — https://arxiv.org/abs/1803.07728 (date accessed 18.09.2025).
69. Bhat, P. Distill on the go: Online knowledge distillation in self-supervised learning / P. Bhat, E. Arani, B. Zonooz // Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition. — 2021. — P. 2678—2687.
70. Hu, Z. An investigation of potential function designs for neural CRF / Z. Hu, Y. Jiang, et al. // arXiv preprint arXiv:2011.05604. — 2020. — https://arxiv.org/pdf/2011.05604 (date accessed 18.09.2025).
71. Henaff, O. Data-efficient image recognition with contrastive predictive coding / O. Henaff // International Conference on Machine Learning. — PMLR. 2020. — P. 4182—4192.
72. Hénaff, O. J. Efficient visual pretraining with contrastive detection / O. J. Henaff, S. Koppula, et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision. — 2021. — P. 10086—10096.
73. Sariyildiz, M. B. Learning visual representations with caption annotations / M. B. Sariyildiz, J. Perez, D. Larlus // European Conference on Computer Vision. Vol. 13684. — Springer. 2020. — P. 153—170.
74. Popov, V. Grad-tts: A diffusion probabilistic model for text-to-speech / V. Popov, I. Vovk, et al. // International Conference on Machine Learning. Vol. 139. — PMLR. 2021. — P. 8599—8608.
75. Feng, S. Y. A survey of data augmentation approaches for NLP / S. Y. Feng, V. Gangal, et al. // arXiv preprint arXiv:2105.03075. — 2021. — https://arxiv.org/abs/2105.03075 (date accessed 18.09.2025).
76. Ye, Q. Teaching machine comprehension with compositional explanations / Q. Ye, X. Huang, et al. // arXiv preprint arXiv:2005.00806. — 2020. — https://arxiv.org/pdf/2005.00806 (date accessed 18.09.2025).
77. Zhang, Y. An unsupervised sentence embedding method by mutual information maximization / Y. Zhang, R. He, et al. // arXiv preprint arXiv:2009.12061. — 2020. — https://arxiv.org/abs/2009.12061 (date accessed 18.09.2025).
78. Sowrirajan, H. Moco pretraining improves representation and transferability of chest x-ray models / H. Sowrirajan, J. Yang, et al. // Medical Imaging with Deep Learning. — PMLR. 2021. — P. 728—744.
79. Chen, L. Self-supervised learning for medical image analysis using image context restoration / L. Chen, P. Bentley, et al. // Medical image analysis. — 2019. — Vol. 58. — P. 101539.
80. Chen, K. Multisiam: Self-supervised multi-instance siamese representation learning for autonomous driving / K. Chen, L. Hong, et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision. — 2021. — P. 7546—7554.
81. Xue, F. Toward hierarchical self-supervised monocular absolute depth estimation for autonomous driving applications / F. Xue, G. Zhuo, et al. // 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). — IEEE. 2020. — P. 2330—2337.
82. Bhattacharyya, P. Ssl-lanes: Self-supervised learning for motion forecasting in autonomous driving / P. Bhattacharyya, C. Huang, K. Czarnecki // Conference on Robot Learning. — PMLR. 2023. — P. 1793—1805.
83. He, K. Momentum contrast for unsupervised visual representation learning / K. He, H. Fan, et al. // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. — 2020. — P. 9729—9738.
84. Liu, S. Query2label: A simple transformer way to multi-label classification / S. Liu, L. Zhang, et al. // arXiv preprint arXiv:2107.10834. — 2021. — https://arxiv.org/abs/2107.10834 (date accessed 18.09.2025).
85. Clark, K. ELECTRA: Pre-training text encoders as discriminators rather than generators / K. Clark, M.-T. Luong, et al. // arXiv preprint arXiv:2003.10555. — 2020. — https://arxiv.org/abs/2003.10555 (date accessed 18.09.2025).
86. Bender, E. M. Climbing towards NLU: On meaning, form, and understanding in the age of data / E. M. Bender, A. Koller // Proceedings of the 58th annual meeting of the association for computational linguistics. — 2020. — P. 5185—5198.
87. Ding, S. Evaluating saliency methods for neural language models / S. Ding, P. Koehn // arXiv preprint arXiv:2104.05824. — 2021. — https://arxiv.org/pdf/2104.05824 (date accessed 18.09.2025).
88. Ghesu, F. C. Contrastive self-supervised learning from 100 million medical images with optional supervision / F. C. Ghesu, B. Georgescu, et al. // Journal of Medical Imaging. — 2022. — Vol. 9, no. 6. — P. 064503—064503.
89. Grill, J.-B. Bootstrap your own latent-a new approach to self-supervised learning / J.-B. Grill, F. Strub, et al. // Advances in Neural Information Processing Systems. — 2020. — Vol. 33. — P. 21271—21284.
90. Oquab, M. DINOv2: Learning robust visual features without supervision / M. Oquab, T. Darcet, et al. // arXiv preprint arXiv:2304.07193. — 2023. — https://arxiv.org/abs/2304.07193 (date accessed 18.09.2025).
91. Deng, J. Imagenet: A large-scale hierarchical image database / J. Deng, W. Dong, et al. // 2009 IEEE conference on computer vision and pattern recognition. — IEEE. 2009. — P. 248—255.
92. Devlin, J. Bert: Pre-training of deep bidirectional transformers for language understanding / J. Devlin, M.-W. Chang, et al. // Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics. Vol. 1. — 2019. — P. 4171—4186.
93. Liu, Y. RoBERTa: A robustly optimized bert pretraining approach / Y. Liu, M. Ott, et al. // arXiv preprint arXiv:1907.11692. — 2019. — https://arxiv.org/abs/1907.11692 (date accessed 18.09.2025).
94. Vaswani, A. Attention is all you need / A. Vaswani, N. Shazeer, et al. // Advances in Neural Information Processing Systems. — 2017. — Vol. 30. — P. 5998—6008.
95. Zhang, Y. DialoGPT: Large-scale generative pre-training for conversational response generation / Y. Zhang, S. Sun, et al. // arXiv preprint arXiv:1911.00536. — 2019. — https://arxiv.org/abs/1911.00536 (date accessed 18.09.2025).
96. Radford, A. Language models are unsupervised multitask learners / A. Radford, J. Wu, et al. // OpenAI blog. — 2019. — Vol. 1, no. 8. — P. 9.
97. Goodfellow, I. J. Generative adversarial nets / I. J. Goodfellow, J. Pouget-Abadie, et al. // Advances in Neural Information Processing Systems. — 2014. — Vol. 27. — P. 2672—2680.
98. Kingma, D. P. Auto-encoding variational bayes / D. P. Kingma, M. Welling // arXiv preprint arXiv:1312.6114. — 2013. — https://arxiv.org/abs/1312.6114 (date accessed 18.09.2025).
99. Sohl-Dickstein, J. Deep unsupervised learning using nonequilibrium thermodynamics / J. Sohl-Dickstein, E. Weiss, et al. // International conference on machine learning. Vol. 37. — PMLR. 2015. — P. 2256—2265.
100. Ho, J. Denoising diffusion probabilistic models / J. Ho, A. Jain, P. Abbeel // Advances in Neural Information Processing Systems. — 2020. — Vol. 33. — P. 6840—6851.
101. Rezende, D. Variational inference with normalizing flows / D. Rezende, S. Mohamed // International Conference on Machine Learning. — PMLR. 2015. — P. 1530—1538.
102. Papamakarios, G. Normalizing flows for probabilistic modeling and inference / G. Papamakarios, E. Nalisnick, et al. // Journal of Machine Learning Research. — 2021. — Vol. 22, no. 57. — P. 1—64.
103. Dinh, L. Density estimation using Real NVP / L. Dinh, J. Sohl-Dick-stein, S. Bengio // arXiv preprint arXiv:1605.08803. — 2016. — https://arxiv.org/abs/1605.08803 (date accessed 18.09.2025).
104. Ferchichi, A. Spatio-temporal modeling of climate change impacts on drought forecast using Generative Adversarial Network: A case study in Africa / A. Ferchichi, M. Chihaoui, A. Ferchichi // Expert Systems with Applications. — 2024. — Vol. 238. — P. 122211.
105. Mirza, M. Conditional generative adversarial nets / M. Mirza, S. Osindero // arXiv preprint arXiv:1411.1784. — 2014. — https://arxiv.org/abs/1411.1784 (date accessed 18.09.2025).
106. Foroumandi, E. Generative adversarial network for real-time flash drought monitoring: A deep learning study / E. Foroumandi, K. Gavahi, H. Morad-khani // Water Resources Research. — 2024. — Vol. 60, no. 5. — e2023WR035600.
107. Wang, Z. Learning latent seasonal-trend representations for time series forecasting / Z. Wang, X. Xu, et al. // Advances in Neural Information Processing Systems. — 2022. — Vol. 35. — P. 38775—38787.
108. Salinas, D. DeepAR: Probabilistic forecasting with autoregressive recurrent networks / D. Salinas, V. Flunkert, et al. // International journal of forecasting. — 2020. — Vol. 36, no. 3. — P. 1181—1191.
109. Rangapuram, S. S. Deep state space models for time series forecasting / S. S. Rangapuram, M. W. Seeger, et al. // Advances in Neural Information Processing Systems. — 2018. — Vol. 31. — P. 7785—7794.
110. Wu, X. KVAE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting / X. Wu, X. Qiu, et al. // arXiv preprint arXiv:2505.23017. — 2025. — https://arxiv.org/html/2505.23017v1 (date accessed 18.09.2025).
111. Koopman, B. O. Hamiltonian systems and transformation in Hilbert space / B. O. Koopman // Proceedings of the National Academy of Sciences. — 1931. — Vol. 17, no. 5. — P. 315—318.
112. Bishop, G. An introduction to the kalman filter / G. Bishop, G. Welch, et al. // Proc of SIGGRAPH, Course. — 2001. — Vol. 8, no. 27599—23175. — P. 41.
113. Tashiro, Y. CSDI: Conditional score-based diffusion models for probabilistic time series imputation / Y. Tashiro, J. Song, et al. // Advances in Neural Information Processing Systems. — 2021. — Vol. 34. — P. 24804—24816.
114. Li, Y. Generative time series forecasting with diffusion, denoise, and disentanglement / Y. Li, X. Lu, et al. // Advances in Neural Information Processing Systems. — 2022. — Vol. 35. — P. 23009—23022.
115. Kollovieh, M. Predict, refine, synthesize: Self-guiding diffusion models for probabilistic time series forecasting / M. Kollovieh, A. F. Ansari, et al. // Advances in Neural Information Processing Systems. — 2023. — Vol. 36. — P. 28341—28364.
116. Medina-Salgado, B. Urban traffic flow prediction techniques: A review / B. Medina-Salgado, E. Sanchez-DelaCruz, et al. // Sustainable Computing: Informatics and Systems. — 2022. — Vol. 35. — P. 100739.
117. Rasul, K. Multivariate probabilistic time series forecasting via conditioned normalizing flows / K. Rasul, A.-S. Sheikh, et al. // arXiv preprint arXiv:2002.06103. — 2020. — https://arxiv.org/abs/2002.06103 (date accessed 18.09.2025).
118. Fan, W. In-flow: Instance normalization flow for non-stationary time series forecasting / W. Fan, S. Zheng, et al. // Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. — 2025. — P. 295—306.
119. Islam, A. A broad study on the transferability of visual representations with contrastive learning / A. Islam, C.-F. R. Chen, et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision. — 2021. — P. 8845—8855.
120. Caron, M. Emerging properties in self-supervised vision transformers / M. Caron, H. Touvron, et al. // Proceedings of the IEEE/CVF international conference on computer vision. — 2021. — P. 9650—9660.
121. Zbontar, J. Barlow twins: Self-supervised learning via redundancy reduction / J. Zbontar, L. Jing, et al. // International Conference on Machine Learning. — PMLR. 2021. — P. 12310—12320.
122. Chen, T. A simple framework for contrastive learning of visual representations / T. Chen, S. Kornblith, et al. // International Conference on Machine Learning. — PMLR. 2020. — P. 1597—1607.
123. Chen, X. Improved baselines with momentum contrastive learning / X. Chen, H. Fan, et al. // arXiv preprint arXiv:2003.04297. — 2020. — https://arxiv.org/abs/2003.04297 (date accessed 18.09.2025).
124. Chen, X. An empirical study of training self-supervised vision transformers / X. Chen, S. Xie, K. He // Proceedings of the IEEE/CVF international conference on computer vision. — 2021. — P. 9640—9649.
125. Tenenbaum, J. B. A global geometric framework for nonlinear dimensionality reduction / J. B. Tenenbaum, V. d. Silva, J. C. Langford // science. — 2000. — Vol. 290, no. 5500. — P. 2319—2323.
126. IBM research. The New Zealand petroleum & minerals online exploration database. — 2015. — https://www.nzpam.govt.nz, (date accessed 18.09.2025).
127. Romanenkova, E. Similarity learning for wells based on logging data /
E. Romanenkova, A. Rogulina, et al. // Journal of Petroleum Science and Engineering. — 2022. — Vol. 215. — P. 110690.
128. Acock, A. C. Working with missing values / A. C. Acock // Journal of Marriage and family. — 2005. — Vol. 67, no. 4. — P. 1012—1028.
129. Sola, J. Importance of input data normalization for the application of neural networks to complex industrial problems / J. Sola, J. Sevilla // IEEE Transactions on nuclear science. — 1997. — Vol. 44, no. 3. — P. 1464—1468.
130. Schroff, F. Facenet: A unified embedding for face recognition and clustering /
F. Schroff, D. Kalenichenko, J. Philbin // Proceedings of the IEEE conference on computer vision and pattern recognition. — 2015. — P. 815—823.
131. Hochreiter, S. Long short-term memory / S. Hochreiter, J. Schmidhuber // Neural computation. — 1997. — Vol. 9, no. 8. — P. 1735—1780.
132. Rumelhart, D. E. Learning internal representations by error propagation : tech. rep. / D. E. Rumelhart, G. E. Hinton, R. J. Williams. — 1987. — https://ieeexplore.ieee.org/document/6302929, (date accessed 18.09.2025).
133. Hewamalage, H. Recurrent neural networks for time series forecasting: Current status and future directions / H. Hewamalage, C. Bergmeir, K. Ban-dara // International Journal of Forecasting. — 2021. — Vol. 37, no. 1. — P. 388—427.
134. Iwana, B. K. An empirical survey of data augmentation for time series classification with neural networks / B. K. Iwana, S. Uchida // Plos one. — 2021. — Vol. 16, no. 7. — e0254841.
135. Rand, W. M. Objective criteria for the evaluation of clustering methods / W. M. Rand // Journal of the American Statistical association. — 1971. — Vol. 66, no. 336. — P. 846—850.
136. Zafarani, R. Social media mining: an introduction / R. Zafarani, M. A. Ab-basi, H. Liu. — Cambridge University Press, 2014. — P. 320.
137. Bazarova, A. Normalizing self-supervised learning for provably reliable Change Point Detection / A. Bazarova, E. Romanenkova, A. Zaytsev // 2024 IEEE International Conference on Data Mining (ICDM). — IEEE. 2024. — P. 21—30.
138. Wang, L. Self-supervised classification of weather systems based on spatiotemporal contrastive learning / L. Wang, Q. Li, Q. Lv // Geophysical Research Letters. — 2022. — Vol. 49, no. 15. — e2022GL099131.
139. Shi, X. Convolutional LSTM network: A machine learning approach for precipitation nowcasting / X. Shi, Z. Chen, et al. // Advances in Neural Information Processing Systems. — 2015. — Vol. 28. — P. 1—9.
140. Dau, H. A. The UCR time series archive / H. A. Dau, A. Bagnall, et al. // IEEE/CAA Journal of Automatica Sinica. — 2019. — Vol. 6, no. 6. — P. 1293—1305.
141. Bagnall, A. The UEA multivariate time series classification archive, 2018 / A. Bagnall, H. A. Dau, et al. // arXiv preprint arXiv:1811.00075. — 2018. — https://arxiv.org/abs/1811.00075 (date accessed 18.09.2025).
142. Ansuini, A. Intrinsic dimension of data representations in deep neural networks / A. Ansuini, A. Laio, et al. // Advances in Neural Information Processing Systems. — 2019. — Vol. 32. — P. 1—14.
143. Change, I. P. O. C. Climate change 2007: The physical science basis / I. P. O. C. Change, et al. // Agenda. — 2007. — Vol. 6, no. 07. — P. 333.
144. Hao, Z. Global integrated drought monitoring and prediction system / Z. Hao, A. AghaKouchak, et al. // Scientific data. — 2014. — Vol. 1, no. 1. — P. 1—10.
145. Wilhite, D. A. Managing drought risk in a changing climate: The role of national drought policy / D. A. Wilhite, M. V. Sivakumar, R. Pulwarty // Weather and climate extremes. — 2014. — Vol. 3. — P. 4—13.
146. Rhee, J. Monitoring agricultural drought for arid and humid regions using multi-sensor remote sensing data / J. Rhee, J. Im, G. J. Carbone // Remote Sensing of environment. — 2010. — Vol. 114, no. 12. — P. 2875—2887.
147. Cihlar, J. Relation between the normalized difference vegetation index and ecological variables / J. Cihlar, L. S. Laurent, J. Dyer // Remote sensing of Environment. — 1991. — Vol. 35, no. 2/3. — P. 279—298.
148. Zhang, Y. Impact-based evaluation of multivariate drought indicators for drought monitoring in China / Y. Zhang, Z. Hao, et al. // Global and Planetary Change. — 2023. — Vol. 228. — P. 104219.
149. McKee, T. The relation of drought frequency and duration to time scales / T. McKee, N. Doesken, J. Kleist // Proceedings of the Eighth Conference on Applied Climatology. — 1993. — Vol. 17. — P. 179—184.
150. Vicente-Serrano, S. A Multiscalar Drought Index Sensitive to Global Warming: The Standardized Precipitation Evapotranspiration Index / S. Vicente-Serrano, B. Santiago, L. Juan // Journal of climate. — 2010. — Vol. 23. — P. 1696—1718.
151. Alley, W. M. The Palmer drought severity index: limitations and assumptions / W. M. Alley // Journal of Applied Meteorology and Climatology. — 1984. — Vol. 23, no. 7. — P. 1100—1109.
152. Dai, A. The climate data guide: Palmer drought severity index (PDSI) / A. Dai, N. C. for Atmospheric Research Staff. — 2017. — https://climatedataguide.ucar.edu (date accessed 18.09.2025).
153. McPherson, R. A. A place-based approach to drought forecasting in south-central Oklahoma / R. A. McPherson, I. L. Corporal-Lodangco, M. B. Rich-man // Earth and Space Science. — 2022. — Vol. 9, no. 10. — e2022EA002315.
154. Liu, Y. An insight into the Palmer drought mechanism based indices: comprehensive comparison of their strengths and limitations / Y. Liu, L. Ren, et al. // Stochastic Environmental Research and Risk Assessment. — 2016. — Vol. 30, no. 1. — P. 119—136.
155. Gorelick, N. Google Earth Engine: Planetary-scale geospatial analysis for everyone / N. Gorelick, M. Hancher, et al. // Remote sensing of Environment. — 2017. — Vol. 202. — P. 18—27.
156. Gao, Z. EarthFormer: Exploring space-time transformers for earth system forecasting / Z. Gao, X. Shi, et al. // Advances in Neural Information Processing Systems. — 2022. — Vol. 35. — P. 25390—25403.
157. Pathak, J. FourCastNet: A global data-driven high-resolution weather model using adaptive fourier neural operators / J. Pathak, S. Sub-ramanian, et al. // arXiv preprint arXiv:2202.11214. — 2022. — https://arxiv.org/abs/2202.11214 (date accessed 18.09.2025).
158. Zeng, A. Are transformers effective for time series forecasting? / A. Zeng, M. Chen, et al. // Proceedings of the AAAI conference on artificial intelligence. Vol. 37. — 2023. — P. 11121—11128.
159. Rasp, S. WeatherBench: a benchmark data set for data-driven weather forecasting / S. Rasp, P. D. Dueben, et al. // Journal of Advances in Modeling Earth Systems. — 2020. — Vol. 12. — e2020MS002203.
160. Tan, C. OpenSTL: A comprehensive benchmark of spatio-temporal predictive learning / C. Tan, S. Li, et al. // Advances in Neural Information Processing Systems. Vol. 36. — 2023. — P. 69819—69831.
161. Fort, S. Deep ensembles: A loss landscape perspective / S. Fort, H. Hu,
B. Lakshminarayanan // arXiv preprint arXiv:1912.02757. — 2019. — https://arxiv.org/abs/1912.02757 (date accessed 18.09.2025).
162. Jain, S. Maximizing overall diversity for improved uncertainty estimates in deep ensembles / S. Jain, G. Liu, et al. // Proceedings of the AAAI conference on artificial intelligence. Vol. 34. — 2020. — P. 4264—4271.
163. Bishop, C. M. Pattern recognition and machine learning. Vol. 4 /
C. M. Bishop, N. M. Nasrabadi. — Springer, 2006. — P. 738.
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.