Оценка неопределенности в задачах обработки естественного языка тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Важенцев Артем Андреевич

  • Важенцев Артем Андреевич
  • кандидат науккандидат наук
  • 2026, «Национальный исследовательский университет «Высшая школа экономики»
  • Специальность ВАК РФ00.00.00
  • Количество страниц 276
Важенцев Артем Андреевич. Оценка неопределенности в задачах обработки естественного языка: дис. кандидат наук: 00.00.00 - Другие cпециальности. «Национальный исследовательский университет «Высшая школа экономики». 2026. 276 с.

Оглавление диссертации кандидат наук Важенцев Артем Андреевич

Contents

Page

Introduction

Chapter 1. Background on Uncertainty Quantification (UQ)

1.1 Notion of Uncertainty

1.2 Uncertainty in Deep Learning

1.2.1 Bayesian Modeling

1.2.2 Deterministic Approximations

1.3 Applications of Uncertainty

1.3.1 Out-of-Distribution Detection

1.3.2 Active Learning

1.3.3 Selective Prediction

Chapter 2. UQ for Text Classification Models

2.1 Related Work

2.2 Background

2.2.1 Baseline Methods

2.2.2 Monte Carlo Dropout

2.2.3 Density-based Methods

2.2.4 Metrics

2.3 Integrating Epistemic and Aleatoric Uncertainty

2.3.1 Out-of-Distribution and In-Distribution Instances

2.3.2 Quantifying Epistemic Uncertainty

2.3.3 Quantifying Aleatoric Uncertainty

2.3.4 Hybrid Uncertainty Quantification (HUQ)

2.3.5 Experimental Setup

2.3.6 Results

2.3.7 Analyses

2.4 Medical Diagnostics Application

2.4.1 Motivation

2.4.2 Related Work

Page

2.4.3 Modification of the Hybrid Method

2.4.4 Experimental Setup

2.4.5 Results

Chapter 3. UQ for Sequence-to-Sequence Models

3.1 Related Work

3.2 Background

3.2.1 Information-based Methods

3.2.2 Ensembling

3.2.3 Density-based Methods

3.3 Experiments

3.3.1 Experimental Setup

3.3.2 Results

Chapter 4. UQ for Large Language Models

4.1 Motivation

4.2 Related Work

4.3 Background

4.3.1 Information-based Methods

4.3.2 Sampling-based Methods

4.3.3 Density-based Methods

4.3.4 Supervised Methods

4.3.5 Black-box Methods

4.3.6 Metrics

4.4 Density-based Uncertainty Quantification

4.4.1 Token-Level Mahalanobis Distance

4.4.2 Experiments

4.5 Attention-based Uncertainty Quantification

4.5.1 Motivation

4.5.2 Trainable Attention-Based Conditional Dependency

4.5.3 Experiments

4.6 Application for Medical Assistants

Conclusions

Page

Bibliography

Supplementary Materials

Russian Translation of the Dissertation

Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Введение диссертации (часть автореферата) на тему «Оценка неопределенности в задачах обработки естественного языка»

Introduction

Background. Uncertainty is an intrinsic part of human existence, appearing in every aspect of our daily lives. We are constantly faced with decisions - whether trivial or important, that are made under conditions of incomplete information. This uncertainty arises because we rarely have full knowledge of the conditions or the potential consequences of our choices. For instance, when deciding whether to buy a ready-made meal, we may not know the quality of the ingredients, the skill level of the cook, or the hygiene standards of the kitchen. Despite these unknowns, we rely on prior knowledge, intuition, and experience to make informed decisions. This reliance on incomplete information is a characteristic of human decision-making. When choosing a meal, for example, we may base our decision on past experiences with similar food items, reviews from other customers, or even the reputation of the restaurant. However, these expectations are never guaranteed. There is always a degree of uncertainty. Moreover, this uncertainty is not just limited to food choices but extends to every decision we make. The complexity of human decisions lies in the fact that each choice we make can have different effects on future decisions. These decisions create a sequence of uncertainty, where the outcome of one decision can change future possibilities.

Neural networks (NNs) are computational models inspired by the structure and function of the human brain, and they exhibit similar patterns of uncertainty in their decision-making processes. Just as humans rely on incomplete information to make decisions, NNs often work in environments where data is noisy, incomplete, or ambiguous. This similarity is not accidental, since NNs are designed to follow the way the human brain processes information, making them susceptible to the same kinds of uncertainties.

Given the parallels between human decision-making and neural networks, it is crucial for AI systems to account for uncertainty in their predictions. Overlooking uncertainty can lead to overconfidence in crucial predictions, potentially resulting in harmful consequences, especially in safety-critical domains like healthcare, finance, or autonomous driving.

Relevance of the work. In recent years, the use of neural networks has significantly increased in many fields. One of the most prominent areas of application

is natural language processing (NLP). The introduction of the Transformer architecture [1] has led to substantial improvements in the quality of neural network predictions in NLP. Despite these advancements, modern classification and generation models still face challenges and may produce numerous incorrect predictions, especially in many sensitive domains, such as medical applications for diagnosis prediction or analysis of social networks for toxicity detection. In such domains, the cost of error is very high and can have serious consequences.

On the contrary, we can abstain from using neural networks and rely solely on human experts. In this case, the number of errors will be lower, however, the cost will be enormous. In medical applications, highly qualified experts would be needed to verify all the data. The analysis of social networks required thousands of experts to annotate all the messages to detect toxic ones. Overall, balancing the need for high-quality performance with the practical limitations of human resources makes it practically nearly impossible to rely only on either experts or neural networks.

To address this dilemma, we need to make a human-machine hybrid system, which may use neural networks as a basic part of the system and leverage the human expert for annotation of only the most complicated and ambiguous cases. This hybrid approach can optimize performance while reducing the overall costs. This integration of human work and neural networks ensures a more effective and reliable system overall.

In order to identify the most complicated and ambiguous instances from the entire test dataset we need an automated system. One of the most prominent and effective approaches for this task is uncertainty quantification (UQ). UQ is a subfield of ML that seeks to model the degree to which model predictions can be trusted. These techniques enable the detection of instances for which the model is prone to error. Ideally, the instances with the most uncertain predictions should correspond to errors.

Many natural language processing tasks are subjective and contain inherent ambiguity. For example, the notion of toxicity is inherently subjective [2] and can be defined in a number of ways that may conflict with one another and differ according to the demographic that the methods are applied to [3]. For many datasets, implicit or ambiguous toxicity can comprise more than 90% of the labeled toxic content [4]. Such ambiguity introduces a high risk of classification mistakes for machine learning (ML) models. Classification mistakes for toxicity detection can result in the removal of legitimate non-toxic content on one hand, and the lack of sanction for toxic

content, on the other. A common method for addressing this concern for content moderation is to abstain from predictions on ambiguous instances and process them with the help of human workers [5].

Large language models (LLMs), which are another type of Transformer-based model, have become very popular in recent times. These models found widespread application and achieved impressive results over various tasks, such as machine translation, question answering, text summarization, etc [6—8]. However, even the most modern LLMs still face similar to those encountered by classification models - they are prone to errors during text generation. Generation errors in LLMs consist of a variety types of errors: hallucinations, non-factual claims, misleading generation, incorrect text, etc [9; 10]. It poses significant obstacles when deploying LLMs in practical applications. This issue becomes more and more critical, particularly as users rely on answers generated by LLMs in various chatbot applications such as ChatGPT, Gemini, and Claude [6—8].

To address this challenge effectively, for LLMs, it is also important to develop an approach to detect potential mistakes using UQ techniques. While many UQ methods already have demonstrated efficiency for classification models, they are not applicable to Transformer-based generation models. Furthermore, the definition of mistakes and uncertainty may vary depending on model size, quality, or the specific generation task. The UQ method will address such problems of LLMs as a hallucination in the generated output, especially for models that work in a zero-shot setting. By developing robust and efficient uncertainty quantification techniques for LLMs, we can enhance the reliability and accuracy of these models in real-world applications.

Dissertation objectives. The primary objective of this dissertation is to enhance the reliability and trustworthiness of decision-making in natural language processing (NLP) tasks through the application of uncertainty quantification techniques. Specifically, this work will address the following key objectives:

1. Conduct an in-depth analysis of UQ methods in a diverse range of NLP tasks and model architectures. This includes examining encoder-only models for text classification, encoder-decoder models for sequence-to-sequence tasks such as machine translation, and decoder-only models for text generation tasks.

2. Investigate state-of-the-art UQ methods for ambiguous text classification tasks such as toxicity detection. Propose a new UQ method for more reliable se-

lective classification, especially in ambiguous tasks, by combining different types of uncertainty.

3. Explore the real-world applicability and impact of the proposed UQ approaches for selective classification in the safe-critical healthcare domain.

4. Develop UQ methods for selective generation that demonstrate the effectiveness and generalization of the proposed approach for free-form generation using LLMs in a wide range of downstream tasks.

Scientific novelty. The scientific novelty of the study consists of the following key propositions, which serve as the points to be defended:

1. A new uncertainty quantification method, HUQ, for selective text classification is proposed. This approach combines epistemic and aleatoric UQ techniques through a new hybrid approach, thereby improving the quality of selective classification in text classification tasks and outperforming the state-of-the-art approaches.

2. For the first time, a comprehensive study of state-of-the-art uncertainty quantification methods for selective classification in medical tasks is conducted. This investigation has revealed that accurate modeling of aleatoric and epistemic uncertainties with state-of-the-art methods and combining them helps to improve the quality of selective prediction in healthcare applications.

3. A large evaluation of density-based uncertainty quantification methods for the out-of-distribution detection for sequence-to-sequence models is performed. The findings of this evaluation demonstrated that the density-based uncertainty quantification approaches are both more effective and computationally efficient than previously explored state-of-the-art ensemble-based for out-of-distribution detection in sequence-to-sequence tasks.

4. A new computationally efficient supervised method for uncertainty quantification of LLMs that uses layer-wise density-based scores as features is proposed. The vast empirical investigation demonstrated the effectiveness of the proposed method for sequence-level selective generation and claim-level fact-checking.

5. The new data-driven approach to uncertainty quantification that models the conditional dependency between the individual token predictions of an LLM is proposed. An empirical demonstration that the proposed method outperforms previous approaches is provided.

Theoretical and practical significance. The dissertation offers significant practical contributions to uncertainty quantification in natural language processing. By tackling critical challenges in the field, such as reliable prediction and hallucination detection, the dissertation provides valuable insights and introduces several state-of-the-art UQ techniques for various types of NLP tasks. Furthermore, the proposed methods are computationally efficient, making them highly applicable to real-world scenarios. The dissertation also demonstrates the effectiveness of these methods in safety-critical domains such as healthcare, underscoring their potential impact in real-life applications.

Research methodology. The methodology included methods of classical machine learning and modern deep learning, aspects of linear algebra, probability theory, and mathematical statistics.

Validation of the research results, reliability. The validity and reliability of the research results and conclusions are robustly demonstrated through a comprehensive empirical investigation. This is achieved by leveraging a diverse range of tasks, datasets, and pre-trained models, ensuring broad applicability and generalization. We rigorously evaluate the proposed methods against numerous state-of-the-art approaches, demonstrating their superiority in various contexts. Additionally, we validate the effectiveness of the proposed techniques in real-world applications, such as the medical domain, where accurate predictions are critical.

Approbation of the work and publications. During the preparation of the dissertation, a total of 12 publications have been prepared. The main results of the dissertation materials are presented in 6 papers, including one published in the Q1 journal and three published in the Rank A*/A conference proceedings. The materials of the works were presented at the following conferences in the form of oral and poster presentations: 61st and 60th Annual Meeting of the Association for Computational Linguistics (2023 and 2022), Conference on Empirical Methods in Natural Language Processing (2024), Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (2025). A detailed list of the main publications of the author included in this thesis is provided below:

1. Artem Vazhentsev, Gleb Kuzmin, Akim Tsvigun, Alexander Panchenko, Maxim Panov, Mikhail Burtsev, and Artem Shelmanov. 2023. Hybrid Uncertainty

Quantification for Selective Text Classification in Ambiguous Tasks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11659-11681, Toronto, Canada. Association for Computational Linguistics. [CORE A*]

2. Artem Vazhentsev*, Akim Tsvigun*, Roman Vashurin*, Sergey Pe-trakov*, Daniil Vasilev, Maxim Panov, Alexander Panchenko, and Artem Shel-manov. 2023. Efficient Out-of-Domain Detection for Sequence to Sequence Models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 1430-1454, Toronto, Canada. Association for Computational Linguistics. [CORE A*]

3. Roman Vashurin*, Ekaterina Fadeeva*, Artem Vazhentsev*, Lyudmila Rvanova, Akim Tsvigun, Daniil Vasilev, Rui Xing, Abdelrahman Boda Sadallah, Kir-ill Grishchenkov, Sergey Petrakov, Alexander Panchenko, Timothy Baldwin, Preslav Nakov, Maxim Panov, and Artem Shelmanov. 2025. Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph. Transactions of the Association for Computational Linguistics. [Q1]

4. Artem Vazhentsev, Lyudmila Rvanova, Ivan Lazichny, Alexander Panchenko, Maxim Panov, Timothy Baldwin, and Artem Shelmanov. 2025. Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models. In Proceedings of the Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics, Albuquerque, New Mexico, USA. Association for Computational Linguistics. [CORE A]

The following is a list of additional papers by the author that are included in this thesis but have not been published:

1. Artem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Alexander Panchenko, Preslav Nakov, Timothy Baldwin, Maxim Panov, and Artem Shelmanov. 2024. Unconditional Truthfulness: Learning Conditional Dependency for Uncertainty Quantification of Large Language Models. CoRR abs/2408.10692 (2024)

2. Artem Vazhentsev, Ivan Sviridov, Alvard Barseghyan, Gleb Kuzmin, Alexander Panchenko, Aleksandr Nesterov, Artem Shelmanov and Maxim Panov. 2025. Uncertainty-aware Abstention in Medical Diagnosis Based on Medical Texts. CoRR abs/2502.18050 (2025)

A list of the additional publications of the author published during the PhD studies and relevant to the primary topic of the thesis but are not included in the thesis text is provided below:

1. Artem Vazhentsev*, Gleb Kuzmin*, Artem Shelmanov*, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gu-sev, Mikhail Burtsev, Manvel Avetisian, and Leonid Zhukov. 2022. Uncertainty Estimation of Transformer Predictions for Misclassification Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8237-8252, Dublin, Ireland. Association for Computational Linguistics. [CORE A*]

2. Ekaterina Fadeeva*, Roman Vashurin*, Akim Tsvigun*, Artem Vazhentsev*, Sergey Petrakov*, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, Timothy Baldwin, and Artem Shelmanov. 2023. LM-Polygraph: Uncertainty Estimation for Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 446-461, Singapore. Association for Computational Linguistics. [CORE A*]

3. Gleb Kuzmin, Artem Vazhentsev, Artem Shelmanov, Xudong Han, Simon Suster, Maxim Panov, Alexander Panchenko, and Timothy Baldwin. (2023). Uncertainty Estimation for Debiased Models: Does Fairness Hurt Reliability? Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 744-770. [CORE B]

4. Nikita Kotelevskii, Aleksandr Artemenkov, Kirill Fedyanin, Fedor Noskov, Alexander Fishkov, Artem Shelmanov, Artem Vazhentsev, Aleksandr Petiushko, and Maxim Panov (2022): Nonparametric Uncertainty Quantification for Single Deterministic Neural Network. In Proceedings of NeurIPS 2022, New Orleans, United States of America. [CORE A*]

5. Akim Tsvigun, Leonid Sanochkin, Daniil Larionov, Gleb Kuzmin, Artem Vazhentsev, Ivan Lazichny, Nikita Khromov, Danil Kireev, Aleksandr Ruba-shevskii, Olga Shahmatova, Dmitry V. Dylov, Igor Galitskiy, and Artem Shelmanov. 2022. ALToolbox: A Set of Tools for Active Learning Annotation of Natural Language Texts. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 406-434, Abu Dhabi, UAE. Association for Computational Linguistics. [CORE A*]

6. Irina Nikishina, Maria Tikhonova, Viktoriia Chekalina, Alexey Zaytsev, Artem Vazhentsev, and Alexander Panchenko. 2024. Industry vs Academia: Running a Course on Transformers in Two Setups. In Proceedings of the Sixth Workshop on Teaching NLP, pages 7-22, Bangkok, Thailand. Association for Computational Linguistics.

Author's contribution. The personal contribution of the author includes conducting a comprehensive literature review related to the research topic, formulating the research aims and objectives under the guidance of research supervisors, and actively participating in the development of the research ideas and experimental methods for each chapter together with the scientific supervisor and co-supervisor. The author personally conducted all the experiments and obtained the results presented in the dissertation.

Specifically, the personal contribution of the author to the published works with co-authors is detailed below:

1. "Hybrid Uncertainty Quantification for Selective Text Classification in Ambiguous Tasks": The author conducted all experimental work, processed and systematized the obtained results, and prepared the paper. Additionally, the applicant developed the proposed HUQ method and demonstrated its effectiveness.

2. "Efficient Out-of-Domain Detection for Sequence to Sequence Models": The author conducted experiments on machine translation and summarization tasks, processed and systematized the results, and prepared the part of the paper. The applicant also proposed adapting density-based methods for sequence-to-sequence models and validated their efficiency. Experiments related to QA tasks, not performed by the author, are not included in the thesis.

3. "Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph": The author contributed significantly to the experimental work, benchmark development, and paper preparation. The thesis includes only the methods description from this work.

4. "Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models": The author proposed the core idea, developed the method, designed the experimental framework, conducted all experiments, processed the results, and prepared the paper.

5. "Unconditional Truthfulness: Learning Conditional Dependency for Uncertainty Quantification of Large Language Models": The author developed the

proposed method, conducted the main experimental work, processed and systematized the results, and prepared the paper.

6. "Uncertainty-aware Abstention in Medical Diagnosis Based on Medical Texts": The author performed the majority of the experimental work, processed and systematized the results, and contributed to part of the paper preparation. Experimental parts not conducted by the author are not included in the thesis.

The structure and volume of the dissertation. The dissertation consists of an introduction, four chapters, and a conclusion. The full volume of the dissertation is 148 pages with 22 figures and 52 tables. The list of references contains 160 entries.

Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Заключение диссертации по теме «Другие cпециальности», Важенцев Артем Андреевич

Conclusion

This thesis has explored the critical role of uncertainty quantification in enhancing the reliability and trustworthiness of decision-making in natural language processing. Moving beyond a theoretical exploration, this thesis has introduced a suite of new methods that directly address the critical need for reliability in NLP systems across a variate of applications, including text classification, text generation, and safety-critical domains such as healthcare. The primary objectives were to conduct a comprehensive analysis of existing state-of-the-art techniques and to propose new advanced methods for a diverse range of NLP tasks and model architectures.

The key findings of this research demonstrate the significant potential of UQ methods to enhance the model's predictions in NLP. First, the proposed HUQ method combines epistemic and aleatoric uncertainties and substantially improves the quality of selective classification, particularly in ambiguous tasks such as toxicity detection. Second, the comprehensive evaluation of state-of-the-art UQ methods in medical applications highlights the importance of accurately modeling both types of uncertainty by using the proposed HUQ method. Third, the thesis demonstrates that computationally efficient density-based UQ methods for out-of-distribution detection in sequence-to-sequence tasks outperform ensemble-based approaches, offering effectiveness and computational efficiency. Fourth, the thesis introduces a computationally efficient supervised UQ method for large language models, leveraging layer-wise density-based scores. Finally, a new supervised UQ method for LLMs has been developed, addressing key challenges such as conditional dependencies between generated tokens during text generation.

The findings of this thesis provide also practical solutions for real-world applications. By improving the reliability of NLP systems, the proposed UQ techniques can significantly reduce the risk of errors in sensitive domains like healthcare diagnostics and content moderation. For instance, the integration of UQ into human-machine hybrid systems enables the more efficient allocation of human resources. Furthermore, these UQ techniques offer a cost-effective solution that makes them highly applicable to large-scale industrial applications, reinforcing their practical utility.

In conclusion, this thesis has not only identified key challenges in UQ for NLP but also made significant contributions to the field. By introducing robust and efficient UQ methods, this work has enhanced the reliability of neural networks

and demonstrated their potential for real-world systems. By bridging the gap between theoretical advancements and practical applications, the proposed methods contribute to the development of more trustworthy and reliable AI systems. Future research could explore the applicability of UQ techniques in multi-modal contexts and to further refine their integration into real-time, safety-critical systems.

Список литературы диссертационного исследования кандидат наук Важенцев Артем Андреевич, 2026 год

Bibliography

1. Attention is All you Need / A. Vaswani [et al.] // Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA / ed. by I. Guyon [et al.]. — 2017. — P. 5998—6008. — URL: https://proceedings. neurips.cc / paper_files / paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.

2. Waseem, Z. Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on Twitter / Z. Waseem // Proceedings of the First Workshop on NLP and Computational Social Science. — Austin, Texas : Association for Computational Linguistics, 11/2016. — P. 138—142. — URL: https://aclanthology.org/W16-5618.

3. Thylstrup, N. Detecting 'Dirt' and 'Toxicity': Rethinking Content Moderation as Pollution Behaviour / N. Thylstrup, Z. Waseem // SSRN Electronic Journal. — 2020. — Jan.

4. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection / T. Hartvigsen [et al.] // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). — Dublin, Ireland : Association for Computational Linguistics, 05/2022. — P. 3309—3326. — URL: https://aclanthology.org/2022.acl-long.234.

5. Roberts, S. T. Behind the Screen: Content Moderation in the Shadows of Social Media / S. T. Roberts. — Yale University Press, 2019. — URL: http: //www.jstor.org/stable/j.ctvhrcz0v (visited on 01/19/2023).

6. GPT-4 Technical Report / OpenAI [et al.]. — 2024. — arXiv: 2303.08774 [cs.CL]. — URL: https://arxiv.org/abs/2303.08774.

7. The Llama 3 Herd of Models / A. Dubey [et al.] // CoRR. — 2024. — Vol. abs/2407.21783. — arXiv: 2407.21783. — URL: https://doi.org/10. 48550/arXiv.2407.21783.

8. Gemma 2: Improving Open Language Models at a Practical Size / M. Riviere [et al.] // CoRR. — 2024. — Vol. abs/2408.00118. — arXiv: 2408.00118. — URL: https://doi.org/10.48550/arXiv.2408.00118.

9. Xiao, Y. On Hallucination and Predictive Uncertainty in Conditional Language Generation / Y. Xiao, W. Y. Wang // Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. — Online : Association for Computational Linguistics, 04/2021. — P. 2734—2744. — URL: https://aclanthology.org/2021.eacl-main.236.

10. On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models? / N. Dziri [et al.] // Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. — Seattle, United States : Association for Computational Linguistics, 07/2022. — P. 5271—5285. — URL: https: //aclanthology.org/2022.naacl-main.387.

11. Der Kiureghian, A. Aleatory or epistemic? Does it matter? / A. Der Ki-ureghian, O. Ditlevsen // Structural safety. — 2009. — Vol. 31, no. 2. — P. 105—112.

12. Hendrycks, D. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks / D. Hendrycks, K. Gimpel // 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. — OpenReview.net, 2017. — URL: https://openreview.net/forum?id=Hkg4TI9xl.

13. Settles, B. Active Learning Literature Survey : Computer Sciences Technical Report / B. Settles ; University of Wisconsin-Madison. — 2009. — No. 1648.

14. Gal, Y. Uncertainty in Deep Learning : PhD thesis / Gal Yarin. — University of Cambridge, 2016. — URL: https://api.semanticscholar.org/CorpusID: 86522127.

15. Weight Uncertainty in Neural Network / C. Blundell [et al.] // Proceedings of the 32nd International Conference on Machine Learning. Vol. 37 / ed. by F. Bach, D. Blei. — PMLR. Lille, France : PMLR, 07/2015. — P. 1613—1622. — (Proceedings of Machine Learning Research). — URL: https://proceedings. mlr.press/v37/blundell15.html.

16. Decomposition of Uncertainty in Bayesian Deep Learning for Efficient and Risk-sensitive Learning / S. Depeweg [et al.] // Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018. Vol. 80 / ed. by J. G. Dy, A. Krause. — PMLR, 2018. — P. 1192—1201. — (Proceedings of Machine Learning Research). — URL: http://proceedings.mlr.press/v80/depeweg18a.html.

17. Gal, Y. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning / Y. Gal, Z. Ghahramani // Proceedings of The 33rd International Conference on Machine Learning. Vol. 48 / ed. by M. F. Bal-can, K. Q. Weinberger. — New York, New York, USA : PMLR, 06/2016. — P. 1050—1059. — (Proceedings of Machine Learning Research). — URL: https: //proceedings .mlr.press/v48/gal16.html.

18. Lakshminarayanan, B. Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles / B. Lakshminarayanan, A. Pritzel, C. Blundell // Proceedings of the 31st International Conference on Neural Information Processing Systems. — Long Beach, California, USA : Curran Associates Inc., 2017. — P. 6405—6416. — (NeurIPS 2017). — URL: https://proceedings. neurips. cc / paper /2017/ hash / 9ef2ed4b7fd2c810847ffa5fa85bce38- Abstract. html.

19. A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks / K. Lee [et al.] // Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montreal, Canada. Vol. 31 / ed. by S. Bengio [et al.]. — 2018. — P. 7167—7177. — URL: https : / / proceedings . neurips . cc / paper / 2018 / hash/abdeb6f575ac5c6676b747bca8d09cc2-Abstract.html.

20. Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection / A. Podolskiy [et al.] // Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021. — AAAI Press, 2021. — P. 13675—13682. — URL: https://doi.org/10. 1609/aaai.v35i15.17612.

21. Zhou, W. Contrastive Out-of-Distribution Detection for Pretrained Transformers / W. Zhou, F. Liu, M. Chen // Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing / ed. by M.-F. Moens [et al.]. — Online, Punta Cana, Dominican Republic : Association for Computational Linguistics, 11/2021. — P. 1100—1111. — URL: https://aclanthology. org/2021.emnlp-main.84/.

22. Malinin, A. Uncertainty Estimation in Autoregressive Structured Prediction / A. Malinin, M. J. F. Gales // 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. — Open-Review.net, 2021. — URL: https://openreview.net/forum?id=jN5y-zb5Q7m.

23. Shifts 2.0: Extending The Dataset of Real Distributional Shifts / A. Malinin [et al.]. — 2022. — URL: https://arxiv.org/abs/2206.15407.

24. Active Learning for Sequence Tagging with Deep Pre-trained Models and Bayesian Uncertainty Estimates / A. Shelmanov [et al.] // Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume / ed. by P. Merlo, J. Tiedemann, R. Tsarfaty. — Online : Association for Computational Linguistics, 04/2021. — P. 1698—1712. — URL: https://aclanthology.org/2021.eacl-main.145.

25. ALToolbox: A Set of Tools for Active Learning Annotation of Natural Language Texts / A. Tsvigun [et al.] // Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations / ed. by W. Che, E. Shutova. — Abu Dhabi, UAE : Association for Computational Linguistics, 12/2022. — P. 406—434. — URL: https:// aclanthology.org/2022.emnlp-demos.41/.

26. El-Yaniv, R. On the Foundations of Noise-free Selective Classification / R. El-Yaniv, Y. Wiener // Journal of Machine Learning Research. — 2010. — Vol. 11, no. 53. — P. 1605—1641. — URL: http://portal.acm.org/citation. cfm?id=1859904.

27. Geifman, Y. Selective Classification for Deep Neural Networks / Y. Geifman, R. El-Yaniv // Proceedings of the 31st International Conference on Neural Information Processing Systems. — Long Beach, California, USA : Curran Associates Inc., 2017. — P. 4885—4894. — (NeurIPS 2017). — URL: https://

proceedings.neurips.cc/paper/2017/file/4a8423d5e91fda00bb7e46540e2b0cf1-Paper.pdf.

28. The Art of Abstention: Selective Prediction and Error Regularization for Natural Language Processing / J. Xin [et al.] // Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). — Online : Association for Computational Linguistics, 08/2021. — P. 1040—1051. — URL: https://aclanthology.org/2021.acl-long.84.

29. Deep learning uncertainty quantification for clinical text classification / A. Peluso [et al.] // Journal of Biomedical Informatics. — 2024. — Vol. 149. — P. 104576. — URL: https : / / www. sciencedirect. com / science / article / pii / S1532046423002976.

30. Out-of-Distribution Detection and Selective Generation for Conditional Language Models / J. Ren [et al.] // The Eleventh International Conference on Learning Representations. — 2023. — URL: https://openreview.net/forum? id=kJUS5nD0vPB.

31. Uncertainty Estimation Using a Single Deep Deterministic Neural Network / J. van Amersfoort [et al.] // Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event. Vol. 119. — PMLR, 2020. — P. 9690—9700. — (Proceedings of Machine Learning Research). — URL: http://proceedings.mlr.press/v119/van-amersfoort20a.html.

32. Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness / J. Liu [et al.] // Advances in Neural Information Processing Systems. Vol. 33 / ed. by H. Larochelle [et al.]. — Curran Associates, Inc., 2020. — P. 7498—7512. — URL: https://proceedings.neurips.cc/ paper/2020/file/543e83748234f7cbab21aa0ade66565f-Paper.pdf.

33. Nonparametric Uncertainty Quantification for Single Deterministic Neural Network / N. Kotelevskii [et al.] // Advances in Neural Information Processing Systems. Vol. 35 / ed. by S. Koyejo [et al.]. — Curran Associates, Inc., 2022. — P. 36308—36323. — URL: https://proceedings.neurips.cc/paper_files/paper/ 2022/file/eb7389b039655fc5c53b11d4a6fa11bc-Paper-Conference.pdf.

34. Geifman, Y. SelectiveNet: A Deep Neural Network with an Integrated Reject Option / Y. Geifman, R. El-Yaniv // Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA. Vol. 97 / ed. by K. Chaudhuri, R. Salakhutdinov. — PMLR,

2019. — P. 2151—2159. — (Proceedings of Machine Learning Research). — URL: http://proceedings.mlr.press/v97/geifman19a.html.

35. Deep Deterministic Uncertainty: A New Simple Baseline / J. Mukhoti [et al.] // IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. — IEEE, 2023. — P. 24384—24394. — URL: https://doi.org/10.1109/CVPR52729.2023.02336.

36. Mitigating Uncertainty in Document Classification / X. Zhang [et al.] // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). — Minneapolis, Minnesota : Association for Computational Linguistics, 06/2019. — P. 3126—3136. — URL: https: //aclanthology.org/N19-1316.

37. Towards More Accurate Uncertainty Estimation In Text Classification / J. He [et al.] // Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 / ed. by B. Webber [et al.]. — Association for Computational Linguistics,

2020. — P. 8362—8372. — URL: https://doi.org/10.18653/v1/2020.emnlp-main.671.

38. On Mixup Training: Improved Calibration and Predictive Uncertainty for Deep Neural Networks / S. Thulasidasan [et al.] // Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada / ed. by H. M. Wallach [et al.]. — 2019. — P. 13888—13899. — URL: https://proceedings.neurips.cc/paper/2019/ hash/36ad8b5f42db492827016448975cc22d-Abstract.html.

39. How Certain is Your Transformer? / A. Shelmanov [et al.] // Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. — Online : Association for Computational Linguistics, 04/2021. — P. 1833—1840. — URL: https://www.aclweb.org/ anthology/2021.eacl-main.157.

40. Uncertainty Estimation of Transformer Predictions for Misclassification Detection / A. Vazhentsev [et al.] // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). — Dublin, Ireland : Association for Computational Linguistics, 05/2022. — P. 8237—8252. — URL: https://aclanthology.org/2022.acl-long.566.

41. Active Learning to Recognize Multiple Types of Plankton / T. Luo [et al.] // J. Mach. Learn. Res. — 2005. — Vol. 6. — P. 589—613. — URL: https://jmlr. org/papers/v6/luo05a.html.

42. Gal, Y. Deep Bayesian Active Learning with Image Data / Y. Gal, R. Islam, Z. Ghahramani // Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017. Vol. 70 / ed. by D. Precup, Y. W. Teh. — PMLR, 2017. — P. 1183—1192. — (Proceedings of Machine Learning Research). — URL: http://proceedings. mlr.press/v70/gal17a.html.

43. Kampffmeyer, M. Semantic segmentation of small objects and modeling of uncertainty in urban remote sensing images using deep convolutional neural networks / M. Kampffmeyer, A.-B. Salberg, R. Jenssen // Proceedings of the IEEE conference on computer vision and pattern recognition workshops. — 2016. — P. 1—9.

44. Bayesian Active Learning for Classification and Preference Learning / N. Houlsby [et al.] // CoRR. — 2011. — Vol. abs/1112.5745. — arXiv: 1112. 5745. — URL: http://arxiv.org/abs/1112.5745.

45. Detection of Adversarial Examples in Text Classification: Benchmark and Baseline via Robust Density Estimation / K. Yoo [et al.] // Findings of the Association for Computational Linguistics: ACL 2022. — Dublin, Ireland : Association for Computational Linguistics, 04/2022. — P. 3656—3672. — URL: https://aclanthology.org/2022.findings-acl.289.

46. Rousseeuw, P. J. Least median of squares regression / P. J. Rousseeuw // Journal of the American statistical association. — 1984. — Vol. 79, no. 388. — P. 871—880. — URL: https://doi.org/10.1080/01621459.1984.10477105.

47. Malinin, A. Predictive Uncertainty Estimation via Prior Networks / A. Ma-linin, M. J. F. Gales // Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018,

NeurIPS 2018, December 3-8, 2018, Montreal, Canada / ed. by S. Bengio [et al.]. — 2018. — P. 7047—7058. — URL: https://proceedings.neurips.cc/ paper/2018/hash/3ea2db50e62ceefceaf70a9d9a56a6f4-Abstract.html.

48. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators / K. Clark [et al.] // 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. — OpenReview.net, 2020. — URL: https : / / openreview . net / forum ? id = r1xMH1BtvB.

49. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding / J. Devlin [et al.] // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). — Stroudsburg, PA, USA : Association for Computational Linguistics, 2019. — P. 4171—4186. — arXiv: 1810.04805v1. — URL: http://aclweb.org/anthology/ N19-1423.

50. ParaDetox: Detoxification with Parallel Data / V. Logacheva [et al.] // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). — Dublin, Ireland : Association for Computational Linguistics, 05/2022. — P. 6804—6818. — URL: https:// aclanthology.org/2022.acl-long.469.

51. Automated Hate Speech Detection and the Problem of Offensive Language / T. Davidson [et al.] // Proceedings of the International AAAI Conference on Web and Social Media. — 2017. — May. — Vol. 11, no. 1. — P. 512—515. — URL: https://ojs.aaai.org/index.php/ICWSM/article/view/14955.

52. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech / M. ElSherief [et al.] // Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. — Online, Punta Cana, Dominican Republic : Association for Computational Linguistics, 11/2021. — P. 345—363. — URL: https://aclanthology.org/2021.emnlp-main.29.

53. Lang, K. NewsWeeder: Learning to Filter Netnews / K. Lang // Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995 / ed. by A. Prieditis,

S. Russell. — Morgan Kaufmann, 1995. — P. 331—339. — URL: https://doi. org/10.1016/b978-1-55860-377-6.50048-7.

54. Recursive Deep Models for Semantic Compositionality Over a Sentiment Tree-bank / R. Socher [et al.] // Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. — Seattle, Washington, USA : Association for Computational Linguistics, 10/2013. — P. 1631—1642. — URL: https://aclanthology.org/D13-1170.

55. McAuley, J. Hidden Factors and Hidden Topics: Understanding Rating Dimensions with Review Text / J. McAuley, J. Leskovec // Proceedings of the 7th ACM Conference on Recommender Systems. — Hong Kong, China : Association for Computing Machinery, 2013. — P. 165—172. — (RecSys '13). — URL: https://doi.org/10.1145/2507157.2507163.

56. Automatic multilabel detection of ICD10 codes in Dutch cardiology discharge letters using neural networks / A. Sammani [et al.] // npj Digital Medicine. —

2021. — Feb. — Vol. 4, no. 1. — P. 37. — URL: https://doi.org/10.1038/s41746-021-00404-9.

57. Juhn, Y. Artificial intelligence approaches using natural language processing to advance EHR-based clinical research / Y. Juhn, H. Liu // Journal of Allergy and Clinical Immunology. — 2020. — Vol. 145, no. 2. — P. 463—469. — URL: https://www.sciencedirect.com/science/article/pii/ S0091674919326041.

58. Lu, H. A comparative study on deep learning models for text classification of unstructured medical notes with various levels of class imbalance / H. Lu, L. Ehwerhemuepha, C. Rakovski // BMC Medical Research Methodology. —

2022. — July. — Vol. 22, no. 1. — P. 181. — URL: https://doi.org/10.1186/ s12874-022-01665-y.

59. Pharmabot: a pediatric generic medicine consultant chatbot / B. E. V. Comendador [et al.] // Journal of Automation and Control Engineering. — 2015. — Vol. 3, no. 2.

60. Mandy: Towards a smart primary care chatbot application / L. Ni [et al.] // International symposium on knowledge and systems sciences. — Springer. 2017. — P. 38—52.

61. Prevalence of harmful diagnostic errors in hospitalised adults: a systematic review and meta-analysis / C. G. Gunderson [et al.] // BMJ Qual Saf. — 2020. — Dec. — Vol. 29, no. 12. — P. 1008—1018.

62. Diagnostic error increases mortality and length of hospital stay in patients presenting through the emergency room / W. E. Hautz [et al.] // Scandinavian journal of trauma, resuscitation and emergency medicine. — 2019. — Apr. — Vol. 27, no. 1. — P. 54. — URL: https://europepmc.org/articles/PMC6505221.

63. An overview of clinical decision support systems: benefits, risks, and strategies for success / R. T. Sutton [et al.] // NPJ digital medicine. — 2020. — Vol. 3, no. 1. — P. 17. — URL: https://www.nature.com/articles/s41746-020-0221-y.

64. Alyahya, M. S. Health care professionals' knowledge and awareness of the ICD-10 coding system for assigning the cause of perinatal deaths in Jordanian hospitals / M. S. Alyahya, Y. S. Khader //J. Multidiscip. Healthc. — 2019. — Feb. — Vol. 12. — P. 149—157.

65. Patient-centered radiology reports with generative artificial intelligence: adding value to radiology reporting / J. Park [et al.] // Scientific Reports. — 2024. — Vol. 14, no. 1. — P. 13218. — URL: https://www.nature.com/ articles/s41598-024-63824-z.

66. Automated ICD coding using extreme multi-label long text transformer-based models / L. Liu [et al.] // Artificial Intelligence in Medicine. — 2023. — Vol. 144. — P. 102662. — URL: https://www.sciencedirect.com/science/ article/pii/S0933365723001768.

67. Clinical text classification research trends: Systematic literature review and open issues / G. Mujtaba [et al.] // Expert Systems with Applications. — 2019. — Vol. 116. — P. 494—520. — URL: https://www.sciencedirect.com/ science/article/pii/S0957417418306110.

68. Eloranta, S. Predictive models for clinical decision making: Deep dives in practical machine learning / S. Eloranta, M. Boman //J Intern Med. — 2022. — Aug. — Vol. 292, no. 2. — P. 278—295.

69. MINIMAR (MINimum Information for Medical AI Reporting): Developing reporting standards for artificial intelligence in health care / T. Hernan-dez-Boussard [et al.] // Journal of the American Medical Informatics Association. — 2020. — June. — Vol. 27, no. 12. — P. 2011—2015. — eprint:

https : / / academic. oup. com / jamia / article - pdf /27/12/ 2011 / 34838637 / ocaa088.pdf. — URL: https://doi.org/10.1093/jamia/ocaa088.

70. A survey of uncertainty in deep neural networks / J. Gawlikowski [et al.] // Artif. Intell. Rev. — USA, 2023. — July. — Vol. 56, Suppl 1. — P. 1513—1589. — URL: https://doi.org/10.1007/s10462-023-10562-9.

71. Kompa, B. Second opinion needed: communicating uncertainty in medical machine learning / B. Kompa, J. Snoek, A. L. Beam // npj Digital Medicine. — 2021. — Jan. — Vol. 4, no. 1. — P. 4. — URL: https://doi.org/10.1038/s41746-020-00367-3.

72. Ambiguity in medical concept normalization: An analysis of types and coverage in electronic health record datasets / D. Newman-Griffis [et al.] // Journal of the American Medical Informatics Association. — 2020. — Dec. — Vol. 28, no. 3. — P. 516—532. — eprint: https://academic.oup.com/jamia/article-pdf/28/3/516/36428822/ocaa269.pdf. — URL: https://doi.org/10.1093/ jamia/ocaa269.

73. Rare Codes Count: Mining Inter-code Relations for Long-tail Clinical Text Classification / J. Chen [et al.] // Proceedings of the 5th Clinical Natural Language Processing Workshop / ed. by T. Naumann [et al.]. — Toronto, Canada : Association for Computational Linguistics, 07/2023. — P. 403—413. — URL: https://aclanthology.org/2023.clinicalnlp-1.43.

74. Out-of-Distribution Detection for Medical Applications: Guidelines for Practical Evaluation / K. Zadorozhny [et al.] // Multimodal AI in Healthcare: A Paradigm Shift in Health Intelligence. — Cham : Springer International Publishing, 2023. — P. 137—153. — URL: https://doi.org/10.1007/978-3-031-14771-5_10.

75. The top 10 causes of death. — 2024. — URL: https://www.who.int/news-room/fact-sheets/detail/the-top- 10-causes-of-death.

76. Publisher Correction: Deep Bayesian Gaussian processes for uncertainty estimation in electronic health records / Y. Li [et al.] // Scientific Reports. — 2021. — Nov. — Vol. 11, no. 1. — P. 22254. — URL: https://doi.org/10.1038/ s41598-021-01680-x.

77. The Role of Uncertainty Quantification for Trustworthy AI / J. Deuschel [et al.] // Unlocking Artificial Intelligence: From Theory to Applications. — Springer, 2024. — P. 95—115.

78. LM-Polygraph: Uncertainty Estimation for Language Models / E. Fadeeva [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations / ed. by Y. Feng, E. Lefever. — Singapore : Association for Computational Linguistics, 12/2023. — P. 446—461. — URL: https://aclanthology.org/2023.emnlp-demo.41.

79. Hybrid Uncertainty Quantification for Selective Text Classification in Ambiguous Tasks / A. Vazhentsev [et al.] // Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). — Toronto, Canada : Association for Computational Linguistics, 07/2023. — P. 11659—11681. — URL: https://aclanthology.org/2023.acl-long.652.

80. DEED: DEep Evidential Doctor / A. Ashfaq [et al.] // Artificial Intelligence. — 2023. — Vol. 325. — P. 104019. — URL: https://www.sciencedirect. com/science/article/pii/S0004370223001650.

81. Analyzing the Role of Model Uncertainty for Electronic Health Records / M. W. Dusenberry [et al.] // Proceedings of the ACM Conference on Health, Inference, and Learning. — Toronto, Ontario, Canada : Association for Computing Machinery, 2020. — P. 204—213. — (CHIL '20). — URL: https://doi. org/10.1145/3368555.3384457.

82. Fortunato, M. Bayesian Recurrent Neural Networks / M. Fortunato, C. Blun-dell, O. Vinyals // CoRR. — 2017. — Vol. abs/1704.02798. — arXiv: 1704. 02798. — URL: http://arxiv.org/abs/1704.02798.

83. Uncertainty-Aware Attention for Reliable Interpretation and Prediction / J. Heo [et al.] // Advances in Neural Information Processing Systems. — 2018. — Vol. 2018—December. — P. 909—918. — Publisher Copyright: © 2018 Curran Associates Inc..All rights reserved.; 32nd Conference on Neural Information Processing Systems, NeurIPS 2018 ; Conference date: 02-12-2018 Through 08-12-2018.

84. Modeling the Uncertainty in Electronic Health Records: a Bayesian Deep Learning Approach / R. Qiu [et al.] // CoRR. — 2019. — Vol. abs/1907.06162. — arXiv: 1907.06162. — URL: http://arxiv.org/ abs/1907.06162.

85. Deep Kernel Learning / A. G. Wilson [et al.] // Proceedings of the 19th International Conference on Artificial Intelligence and Statistics. Vol. 51 / ed. by A. Gretton, C. C. Robert. — Cadiz, Spain : PMLR, 05/2016. — P. 370—378. — (Proceedings of Machine Learning Research). — URL: https://proceedings. mlr.press/v51/wilson16.html.

86. Sensoy, M. Evidential deep learning to quantify classification uncertainty / M. Sensoy, L. Kaplan, M. Kandemir // Advances in neural information processing systems. — 2018. — Vol. 31.

87. MIMIC-III, a freely accessible critical care database / A. E. W. Johnson [et al.] // Sci. Data. — 2016. — May. — Vol. 3, no. 1. — P. 160035.

88. Wen, Z. MeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining / Z. Wen, X. H. Lu, S. Reddy // Proceedings of the 3rd Clinical Natural Language Processing Workshop / ed. by A. Rumshisky [et al.]. — Online : Association for Computational Linguistics, 11/2020. — P. 130—135. — URL: https://aclanthology.org/2020.clinicalnlp-1.15.

89. Mimic-IV-ICD: A new benchmark for eXtreme MultiLabel Classification / T. Nguyen [et al.] // CoRR. — 2023. — Vol. abs/2304.13998. — arXiv: 2304. 13998. — URL: https://doi.org/10.48550/arXiv.2304.13998.

90. A large language model for electronic health records / X. Yang [et al.] // npj Digital Medicine. — 2022. — Vol. 5, no. 1. — P. 194.

91. A comparative study of pretrained language models for long clinical text / Y. Li [et al.] //J. Am. Medical Informatics Assoc. — 2023. — Vol. 30, no. 2. — P. 340—347. — URL: https://doi.org/10.1093/jamia/ocac225.

92. Beltagy, I. Longformer: The Long-Document Transformer / I. Beltagy, M. E. Peters, A. Cohan // CoRR. — 2020. — Vol. abs/2004.05150. — arXiv: 2004.05150. — URL: https://arxiv.org/abs/2004.05150.

93. MASS: Masked Sequence to Sequence Pre-training for Language Generation / K. Song [et al.] // International Conference on Machine Learning. — PMLR.

2019. — P. 5926—5936.

94. Incorporating BERT into Neural Machine Translation / J. Zhu [et al.] // International Conference on Learning Representations. — 2020.

95. Multilingual Denoising Pre-training for Neural Machine Translation / Y. Liu [et al.] // Transactions of the Association for Computational Linguistics. —

2020. — Vol. 8. — P. 726—742. — URL: https://doi.org/10.1162/tacl%5C_ a%5C_00343.

96. PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization / J. Zhang [et al.] // Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event. Vol. 119. — PMLR, 2020. — P. 11328—11339. — (Proceedings of Machine Learning Research). — URL: http://proceedings.mlr.press/v119/zhang20ae. html.

97. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension / M. Lewis [et al.] // Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 / ed. by D. Jurafsky [et al.]. — Online : Association for Computational Linguistics, 07/2020. — P. 7871—7880. — URL: https://aclanthology.org/2020.acl-main.703.

98. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer / C. Raffel [et al.] //J. Mach. Learn. Res. — 2020. — Vol. 21. — 140:1—140:67. — URL: http://jmlr.org/papers/v21/20-074.html.

99. Deep Anomaly Detection with Outlier Exposure / D. Hendrycks [et al.] // International Conference on Learning Representations. — 2019.

100. Liang, S. Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks / S. Liang, Y. Li, R. Srikant // International Conference on Learning Representations. — 2018.

101. Lukovnikov, D. Detecting Compositionally Out-of-Distribution Examples in Semantic Parsing / D. Lukovnikov, S. Däubener, A. Fischer // Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021 / ed. by M. Moens

[et al.]. — Association for Computational Linguistics, 2021. — P. 591—598. — URL: https://doi.org/10.18653/v1/2021.findings-emnlp.54.

102. Xiao, T. Z. Wat zei je? detecting out-of-distribution translations with variational transformers / T. Z. Xiao, A. N. Gomez, Y. Gal // arXiv preprint arXiv:2006.08344. — 2020. — Vol. abs/2006.08344. — arXiv: 2006.08344. — URL: https://arxiv.org/abs/2006.08344.

103. Gidiotis, A. Should We Trust This Summary? Bayesian Abstractive Summarization to The Rescue / A. Gidiotis, G. Tsoumakas // Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022 / ed. by S. Muresan, P. Nakov, A. Villavicencio. — Association for Computational Linguistics, 2022. — P. 4119—4131. — URL: https: //aclanthology.org/2022.findings-acl.325.

104. Ueffing, N. Word-Level Confidence Estimation for Machine Translation / N. Ueffing, H. Ney // Comput. Linguistics. — 2007. — Vol. 33, no. 1. — P. 9—40. — URL: https://doi.org/10.1162/coli.2007.33.L9.

105. Findings of the 2017 Conference on Machine Translation (WMT17) / O. Bojar [et al.] // Proceedings of the Second Conference on Machine Translation, Volume 2: Shared Task Papers. — Copenhagen, Denmark : Association for Computational Linguistics, 09/2017. — P. 169—214. — URL: http://www. aclweb.org/anthology/W17-4717.

106. Findings of the 2020 Conference on Machine Translation (WMT20) / L. Bar-rault [et al.] // Proceedings of the Fifth Conference on Machine Translation. — Online : Association for Computational Linguistics, 11/2020. — P. 1—55. — URL: https://aclanthology.org/2020.wmt-1.1.

107. Findings of the 2014 Workshop on Statistical Machine Translation / O. Bojar [et al.] // Proceedings of the Ninth Workshop on Statistical Machine Translation. — Baltimore, Maryland, USA : Association for Computational Linguistics, 06/2014. — P. 12—58. — URL: http:/ /www.aclweb.org/ anthology/W/W14/W14-3302.

108. Librispeech: an ASR corpus based on public domain audio books / V. Panay-otov [et al.] // Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. — IEEE. 2015. — P. 5206—5210.

109. Narayan, S. Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization / S. Narayan, S. B. Cohen, M. Lapata // Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 -November 4, 2018 / ed. by E. Riloff [et al.]. — Association for Computational Linguistics, 2018. — P. 1797—1807. — URL: https://doi.org/10.18653/v1/ d18-1206.

110. Zhang, R. This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation / R. Zhang, J. R. Tetreault // Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers / ed. by A. Korhonen, D. R. Traum, L. Marquez. — Association for Computational Linguistics, 2019. — P. 446—456. — URL: https://doi.org/10.18653/v1/p19-1043.

111. Wang, L. Neural Network-Based Abstract Generation for Opinions and Arguments / L. Wang, W. Ling // NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016 / ed. by K. Knight, A. Nenkova, O. Rambow. — The Association for Computational Linguistics, 2016. — P. 47—57. — URL: https://doi.org/10. 18653/v1/n16-1007.

112. Manakul, P. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models / P. Manakul, A. Liusie, M. Gales // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. — 2023. — P. 9004—9017. — URL: https://doi. org/10.48550/arXiv.2303.08896.

113. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation / S. Min [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. — 2023. — P. 12076—12100. — URL: https://aclanthology.org/2023.emnlp-main.741.

114. Hallucination detection: Robustly discerning reliable answers in large language models / Y. Chen [et al.] // Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. — 2023. — P. 245—255. — URL: https://arxiv.org/abs/2407.04121.

115. Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration / S. Feng [et al.] // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) / ed. by L.-W. Ku, A. Martins, V. Srikumar. — Bangkok, Thailand : Association for Computational Linguistics, 08/2024. — P. 14664—14690. — URL: https: //aclanthology.org/2024.acl-long.786.

116. Uncertainty in natural language generation: From theory to applications / J. Baan [et al.] // arXiv preprint arXiv:2307.15703. — 2023. — URL: https: //arxiv.org/abs/2307.15703.

117. A Survey of Confidence Estimation and Calibration in Large Language Models / J. Geng [et al.] // Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) / ed. by K. Duh, H. Gomez, S. Bethard. — Mexico City, Mexico : Association for Computational Linguistics, 06/2024. — P. 6577—6595. — URL: https://aclanthology.org/2024.naacl-long.366.

118. Unsupervised Quality Estimation for Neural Machine Translation / M. Fomicheva [et al.] // Transactions of the Association for Computational Linguistics. — Cambridge, MA, 2020. — Vol. 8. — P. 539—555. — URL: https://aclanthology.org/2020.tacl-1.35.

119. Lin, Z. Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models / Z. Lin, S. Trivedi, J. Sun // CoRR. — 2023. — Vol. abs/2305.19187. — arXiv: 2305.19187. — URL: https://doi.org/10. 48550/arXiv.2305.19187.

120. Kuhn, L. Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation / L. Kuhn, Y. Gal, S. Farquhar // The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. — OpenReview.net, 2023. — URL: https://openreview.net/pdf ?id=VD-AYtP0dve.

121. Detecting hallucinations in large language models using semantic entropy / S. Farquhar [et al.] // Nature. — 2024. — Vol. 630, no. 8017. — P. 625—630.

122. Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models / J. Duan [et al.] // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) / ed. by L.-W. Ku, A. Martins, V. Srikumar. — Bangkok, Thailand : Association for Computational Linguistics, 08/2024. — P. 5050—5063. — URL: https://aclanthology.org/2024.acl-long.276.

123. Uncertainty Estimation and Reduction of Pre-trained Models for Text Regression / Y. Wang [et al.] // Transactions of the Association for Computational Linguistics / ed. by B. Roark, A. Nenkova. — Cambridge, MA, 2022. — Vol. 10. — P. 680—696. — URL: https://aclanthology.org/2022.tacl-1.39.

124. Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission / J. He [et al.] // Findings of the Association for Computational Linguistics: NAACL 2024 / ed. by K. Duh, H. Gomez, S. Bethard. — Mexico City, Mexico : Association for Computational Linguistics, 06/2024. — P. 2823—2835. — URL: https://aclanthology.org/2024.findings-naacl.180.

125. Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph / R. Vashurin [et al.] // Transactions of the Association for Computational Linguistics. — 2025. — Mar. — Vol. 13. — P. 220—248. — eprint: https : / / direct. mit. edu / tacl / article - pdf / doi / 10. 1162 / tacl \ _a\ _00737 / 2511955 / tacl \ _a\ _00737. pdf. — URL: https: //doi.org/10.1162/tacl%5C_a%5C_00737.

126. Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus / T. Zhang [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. — 2023. — P. 915—932. — URL: https:// doi.org/10.48550/arXiv.2311.13230.

127. Takayama, J. Relevant and Informative Response Generation using Pointwise Mutual Information / J. Takayama, Y. Arase // Proceedings of the First Workshop on NLP for Conversational AI. — Florence, Italy : Association for Computational Linguistics, 08/2019. — P. 133—138. — URL: https:// aclanthology.org/W19-4115.

128. Poel, L. van der. Mutual Information Alleviates Hallucinations in Abstractive Summarization / L. van der Poel, R. Cotterell, C. Meister // Proceedings of

the 2022 Conference on Empirical Methods in Natural Language Processing. — Abu Dhabi, United Arab Emirates : Association for Computational Linguistics, 12/2022. — P. 5956—5965. — URL: https://aclanthology.org/ 2022.emnlp-main.399.

129. Darrin, M. RainProof: An Umbrella to Shield Text Generator from Out--Of-Distribution Data / M. Darrin, P. Piantanida, P. Colombo // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing / ed. by H. Bouamor, J. Pino, K. Bali. — Singapore : Association for Computational Linguistics, 12/2023. — P. 5831—5857. — URL: https: //aclanthology.org/2023.emnlp-main.357.

130. Efficient Out-of-Domain Detection for Sequence to Sequence Models / A. Vazhentsev [et al.] // Findings of the Association for Computational Linguistics: ACL 2023 / ed. by A. Rogers, J. Boyd-Graber, N. Okazaki. — Toronto, Canada : Association for Computational Linguistics, 07/2023. — P. 1430—1454. — URL: https://aclanthology.org/2023.findings-acl.93.

131. Azaria, A. The Internal State of an LLM Knows When It's Lying / A. Azaria, T. Mitchell // Findings of the Association for Computational Linguistics: EMNLP 2023 / ed. by H. Bouamor, J. Pino, K. Bali. — Singapore : Association for Computational Linguistics, 12/2023. — P. 967—976. — URL: https:// aclanthology.org/2023.findings-emnlp.68.

132. INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection / C. Chen [et al.] // The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. — OpenReview.net, 2024. — URL: https://openreview.net/forum?id=Zj12nzlQbz.

133. LLM Factoscope: Uncovering LLMs' Factual Discernment through Measuring Inner States / J. He [et al.] // Findings of the Association for Computational Linguistics ACL 2024 / ed. by L.-W. Ku, A. Martins, V. Srikumar. — Bangkok, Thailand, virtual meeting : Association for Computational Linguistics, 08/2024. — P. 10218—10230. — URL: https://aclanthology.org/2024. findings-acl.608.

134. Do Androids Know They're Only Dreaming of Electric Sheep? / S. CH-Wang [et al.] // Findings of the Association for Computational Linguistics ACL 2024 / ed. by L.-W. Ku, A. Martins, V. Srikumar. — Bangkok, Thailand,

virtual meeting : Association for Computational Linguistics, 08/2024. — P. 4401—4420. — URL: https://aclanthology.org/2024.findings-acl.260.

135. Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps / Y.-S. Chuang [et al.] // Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing / ed. by Y. Al-Onaizan, M. Bansal, Y.-N. Chen. — Miami, Florida, USA : Association for Computational Linguistics, 11/2024. — P. 1419—1436. — URL: https://aclanthology.org/2024.emnlp-main.84/.

136. Language Models (Mostly) Know What They Know / S. Kadavath [et al.] // CoRR. — 2022. — Vol. abs/2207.05221. — arXiv: 2207.05221. — URL: https: //doi.org/10.48550/arXiv.2207.05221.

137. Lin, C.-Y. ROUGE: A Package for Automatic Evaluation of Summaries / C.-Y. Lin // Text Summarization Branches Out. — Barcelona, Spain : Association for Computational Linguistics, 07/2004. — P. 74—81. — URL: https: //www.aclweb.org/anthology/W04-1013.

138. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback / K. Tian [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing / ed. by H. Bouamor, J. Pino, K. Bali. — Singapore : Association for Computational Linguistics, 12/2023. — P. 5433—5442. — URL: https://aclanthology.org/2023.emnlp-main.330.

139. Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models / W. Su [et al.] // Findings of the Association for Computational Linguistics: ACL 2024 / ed. by L.-W. Ku, A. Martins, V. Srikumar. — Bangkok, Thailand : Association for Computational Linguistics, 08/2024. — P. 14379—14391. — URL: https://aclanthology.org/2024. findings-acl.854/.

140. DeBERTa: Decoding-Enhanced BERT with Disentangled Attention / P. He [et al.] // 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. — OpenReview.net, 2021. — URL: https://openreview.net/forum?id=XPZIaotutsD.

141. LUQ: Long-text Uncertainty Quantification for LLMs / C. Zhang [et al.] // Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing / ed. by Y. Al-Onaizan, M. Bansal, Y.-N. Chen. — Miami, Florida, USA : Association for Computational Linguistics, 11/2024. — P. 5244—5262. — URL: https://aclanthology.org/2024.emnlp-main.299/.

142. Incorporating Uncertainty into Deep Learning for Spoken Language Assessment / A. Malinin [et al.] // Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). — Vancouver, Canada : Association for Computational Linguistics, 07/2017. — P. 45—50. — URL: https://aclanthology.org/P17-2008.

143. AlignScore: Evaluating Factual Consistency with A Unified Alignment Function / Y. Zha [et al.] // Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) / ed. by A. Rogers, J. Boyd-Graber, N. Okazaki. — Toronto, Canada : Association for Computational Linguistics, 07/2023. — P. 11328—11348. — URL: https: //aclanthology.org/2023.acl-long.634.

144. COMET: A Neural Framework for MT Evaluation / R. Rei [et al.] // Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) / ed. by B. Webber [et al.]. — Online : Association for Computational Linguistics, 11/2020. — P. 2685—2702. — URL: https: //aclanthology.org/2020.emnlp-main.213.

145. Shrestha, N. Detecting Multicollinearity in Regression Analysis / N. Shrestha // American Journal of Applied Mathematics and Statistics. — 2020. — June. — Vol. 8. — P. 39—42.

146. Mitigating the Multicollinearity Problem and Its Machine Learning Approach: A Review / J. Chan [et al.] // Mathematics. — 2022. — Apr. — Vol. 10. — P. 1283.

147. Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification / E. Fadeeva [et al.] // Findings of the Association for Computational Linguistics: ACL 2024. — Association for Computational Linguistics, 2024. — URL: https://doi.org/10.48550/arXiv.2403.04696.

148. Mistral 7B / A. Q. Jiang [et al.] // CoRR. — 2023. — Vol. abs/2310.06825. — arXiv: 2310.06825. — URL: https://doi.org/10.48550/arXiv.2310.06825.

149. SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization / B. Gliwa [et al.] // Proceedings of the 2nd Workshop on New Frontiers in Summarization. — Hong Kong, China : Association for Computational Linguistics, 11/2019. — P. 70—79. — URL: https://www.aclweb. org/anthology/D19-5409.

150. See, A. Get To The Point: Summarization with Pointer-Generator Networks / A. See, P. J. Liu, C. D. Manning // Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). — Vancouver, Canada : Association for Computational Linguistics, 07/2017. — P. 1073—1083. — URL: https://www.aclweb.org/anthology/P17-1099.

151. PubMedQA: A Dataset for Biomedical Research Question Answering / Q. Jin [et al.] // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) / ed. by K. Inui [et al.]. — Hong Kong, China : Association for Computational Linguistics, 11/2019. — P. 2567—2577. — URL: https://aclanthology.org/D19-1259.

152. Abacha, A. B. A question-entailment approach to question answering / A. B. Abacha, D. Demner-Fushman // BMC Bioinform. — 2019. — Vol. 20, no. 1. — 511:1—511:23. — URL: https://doi.org/10.1186/s12859-019-3119-4.

153. Lin, S. TruthfulQA: Measuring How Models Mimic Human Falsehoods / S. Lin, J. Hilton, O. Evans // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) / ed. by S. Muresan, P. Nakov, A. Villavicencio. — Dublin, Ireland : Association for Computational Linguistics, 05/2022. — P. 3214—3252. — URL: https: //aclanthology.org/2022.acl-long.229.

154. Training Verifiers to Solve Math Word Problems / K. Cobbe [et al.] // CoRR. — 2021. — Vol. abs/2110.14168. — arXiv: 2110.14168. — URL: https: //arxiv.org/abs/2110.14168.

155. Reddy, S. CoQA: A Conversational Question Answering Challenge / S. Reddy, D. Chen, C. D. Manning // Transactions of the Association for Computational Linguistics. — Cambridge, MA, 2019. — Vol. 7. — P. 249—266. — URL: https: //aclanthology.org/Q19-1016.

156. Welbl, J. Crowdsourcing Multiple Choice Science Questions / J. Welbl, N. F. Liu, M. Gardner // Proceedings of the 3rd Workshop on Noisy Usergenerated Text / ed. by L. Derczynski [et al.]. — Copenhagen, Denmark : Association for Computational Linguistics, 09/2017. — P. 94—106. — URL: https://aclanthology.org/W17-4413.

157. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension / M. Joshi [et al.] // Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) / ed. by R. Barzilay, M.-Y. Kan. — Vancouver, Canada : Association for Computational Linguistics, 07/2017. — P. 1601—1611. — URL: https: //aclanthology.org/P17-1147.

158. Measuring Massive Multitask Language Understanding / D. Hendrycks [et al.] // 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. — OpenReview.net, 2021. — URL: https://openreview.net/forum?id=d7KBjmI3GmQ.

159. Findings of the 2019 Conference on Machine Translation (WMT19) / L. Bar-rault [et al.] // Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1) / ed. by O. Bojar [et al.]. — Florence, Italy : Association for Computational Linguistics, 08/2019. — P. 1—61. — URL: https://aclanthology.org/W19-5301.

160. Qwen2.5 Technical Report / A. Yang [et al.] // arXiv preprint arXiv:2412.15115. — 2024.

Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.