Эффективные методы мультиязычного текстового переноса стиля/Efficient Multilingual Text Style Transfer Methods тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Московский Даниил Алексеевич

  • Московский Даниил Алексеевич
  • кандидат науккандидат наук
  • 2026, «Сколковский институт науки и технологий»
  • Специальность ВАК РФ00.00.00
  • Количество страниц 173
Московский Даниил Алексеевич. Эффективные методы мультиязычного текстового переноса стиля/Efficient Multilingual Text Style Transfer Methods: дис. кандидат наук: 00.00.00 - Другие cпециальности. «Сколковский институт науки и технологий». 2026. 173 с.

Оглавление диссертации кандидат наук Московский Даниил Алексеевич

Contents

Page

Introduction

Chapter 1. Background: Text Style Transfer

1.1 Text Style Transfer

1.2 Prior Works

1.2.1 Disentanglement Methods

1.2.2 Latent Space Manipulation

1.2.3 Lexical Editing Methods

1.2.4 Prompting Approaches

1.3 Supervised Sequence-to-Sequence Approaches

1.4 Discussion

Chapter 2. Cross-lingual Text Style Transfer Methods

2.1 The Challenge of Cross-lingual Style Transfer

2.2 Datasets and Models

2.2.1 Parallel Detoxification Corpora

2.3 Cross-lingual Detoxification Transfer

2.3.1 Backtranslation Pipeline

2.3.2 Training Data Translation

2.3.3 Multitask Learning

2.3.4 Parameter-Efficient Knowledge Transfer with Adapters

2.4 Simultaneous Translation and Detoxification

2.5 Evaluation Setup

2.5.1 Automatic Evaluation

2.5.2 Manual Evaluation

2.6 Results

2.6.1 Monolingual Baselines and Multilingual Model Performance

2.6.2 Cross-lingual Detoxification Transfer Performance

2.6.3 Simultaneous Translation and Detoxification Results

2.6.4 Human Evaluation Results

2.7 Discussion

Chapter 3. Fisher Weighted Tensor Train Matrix Decomposition for

Efficient Low Rank Approximation of Transformer Models

3.1 Low-Rank Approximation Problem

3.1.1 Parametric Efficiency

3.1.2 Limitations of Standard SVD for Task-Specific Compression

3.1.3 Fisher Information

3.1.4 FWSVD: Fisher-Weighted SVD

3.2 FWTTM: Fisher-Weighted Tensor Train Matrix Decomposition

3.2.1 Algorithm Description

3.2.2 Implementation Details

3.2.3 Experimental Setup

3.2.3.1 Baselines

3.2.3.2 Experimental details

3.2.4 Results

3.2.4.1 BERT Compression Experiments

3.2.4.2 BART Compression Experiments

3.2.5 FWTTM Discussion

3.3 Generalized Fisher-Weighted SVD

3.3.1 Fisher Information as a Kronecker Product

3.3.2 The GFWSVD Compression Algorithm

3.3.3 Scalable Computation of Kronecker Factors

3.3.4 Relationship to FWSVD

3.3.5 Experimental Setup

3.3.5.1 BERT Compression

3.3.5.2 Llama Compression

3.3.6 Results

3.3.6.1 BERT Compression Results

3.3.6.2 Llama Compression Results

3.3.7 GFWSVD Discussion

3.4 Discussion

Chapter 4. Scaling Multilingual Text Detoxification

4.1 Multilingual Text Detoxification Task at PAN

4.1.1 Task Design and Organization

4.1.2 Dataset Preparation

4.1.3 Dataset Analysis

4.1.3.1 Toxicity Semantics Analysis

4.1.3.2 Toxic and Detoxified Sentences Lengths Comparison

4.1.4 Evaluation Methodology

4.1.4.1 Automatic Evaluation Setup

4.1.4.2 Manual Evaluation Setup

4.1.5 Results

4.2 An Explainable Analysis of Multilingual Toxic Patterns

4.2.1 Cross-lingual Toxicity Features

4.2.2 Language-specific Patterns

4.2.3 Common Toxic Language Patterns

4.3 Chain-of-Thought Methods for Detoxification

4.3.1 Reasoning Approaches in Text Generation

4.3.2 Clustering of Descriptive Attributes

4.3.3 CoT-enhanced Prompting

4.3.4 Experimental Evaluation

4.4 Comparative Analysis Across Languages

4.4.1 Language Family Effects

4.4.2 Comparative Analysis Across Languages: Insights from a Multilingual Shared Task

4.4.3 Performance Patterns

4.5 Discussion

Chapter 5. LLM-based Methods for Synthetic Parallel Data Generation

5.1 Limitations of Crowdsourced Data Collection

5.1.1 Cost and Time Constraints

5.1.2 Quality Control Challenges

5.1.3 Coverage Limitations

5.2 PseudoParaDetox

5.2.1 Activation Patching Technique

5.2.2 Experimental Setup

5.2.2.1 Source Toxic Data

5.2.2.2 UsedLLMs

5.2.2.3 Detoxification Prompt

5.2.2.4 Activation Patching

5.2.2.5 Fine-tuning Setup

5.2.2.6 Automatic Evaluation Setup

5.2.2.7 Side-by-Side Evaluation Setup

5.2.2.8 Manual Evaluation Setup

5.2.3 Experimental Results

5.2.3.1 Automatic Evaluation Results

5.2.3.2 Side-by-side Evaluation Results

5.2.3.3 Manual Evaluation Results

5.3 SynthDetoxM: A Synthetic Multilingual Detoxification Dataset

5.3.1 Methodology

5.3.1.1 Source Toxic Data Collection

5.3.1.2 Parallel Data Generation Pipeline

5.3.1.3 Few-shot Example Mining

5.3.2 Experimental Setup

5.3.3 Toxicity and Similarity of Synthetic Texts

5.3.4 Automatic Evaluation Setup

5.3.5 Side-by-Side Evaluation Setup

5.3.6 Dataset Statistics and Analysis

5.3.7 Linguistic Analysis of SynthDetoxM

5.3.8 Automatic Evaluation Results

5.3.9 Side-by-Side Evaluation Results

5.4 Discussion

Conclusions

List of symbols and abbreviations

Appendix A

Appendix B

Appendix C

List of Figures

List of Tables

Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Введение диссертации (часть автореферата) на тему «Эффективные методы мультиязычного текстового переноса стиля/Efficient Multilingual Text Style Transfer Methods»

Introduction

Natural Language Processing (NLP) has undergone revolutionary developments in recent years. This development is driven primarily by large transformers-based pretrained language models [1]. These models have already surpassed humans in understanding, interpreting, and generating natural human language, and thus have a wide range of different applications. One such application is Text Style Transfer (TST) - the task of modifying certain stylistic attributes of text, such as toxicity, formality, sentiment, or complexity, without changing its original semantic meaning and fluency [2]. TST has the potential to improve user experience, make text more readable, support creative writing, and enable more sophisticated communication in various contexts.

While other subtasks of TST remain important, text detoxification is particularly crucial and socially relevant application of TST. This work is especially focused on converting text with offensive or toxic content into a non-toxic, polite, or neutral equivalent [2]. In today's age of increasing online activity, the dissemination of toxic content poses numerous challenges for productive and safe digital environments. Automatic detoxification is a preventive action that goes beyond simplistic content filtering or excision, and so can allow for worthwhile information or opinion within ill-phrased messages to be retained while reducing harm. This is a requirement for content moderation within social media, protecting users (especially vulnerable groups) from cyber harassment, and even preventing AI systems themselves from generating inappropriate output [129].

Historically, significant progress in achieving high-quality TST, particularly for detoxification, coincided with the adoption of supervised learning paradigms [2]. The availability of parallel corpora - datasets containing paired examples of text in source and target styles (e.g., toxic-neutral sentence pairs) - proved crucial. Fine-tuning sequence-to-sequence models like BART [3] and T5 [4] on datasets such as ParaDetox for English [2], its Russian counterpart [5], or APPDIA [6], allowed these models to learn the complex mapping between styles effectively, consistently surpassing the performance of earlier unsupervised or rule-based methods and establishing a high benchmark for monolingual task quality [2]. However, the applicability of such methods is still limited by the scarcity of parallel datasets, especially for many low-resource languages and niche styles.

Alongside the emergence of parallel detoxification data, the field has witnessed the rise of large language models (LLMs), which are a further development of earlier decoder-based language models [7]. Trained on trillions of tokens and comprising anywhere from several billion to over 670 billion parameters, LLMs exhibit strong zero-shot and few-shot capabilities across a wide range of tasks [1]. Of course, these generative abilities can also be used to produce highly fluent, coherent text in any desired style and perform TST directly [8].

LLMs can be used to generate large-scale synthetic parallel datasets from style transfer, which we explore later in this manuscript (see Chapter 5 for details) for the case of parallel text detoxification data [130, 9].

The practical applications of text detoxification are diverse and critical for modern digital communication. Firstly, rather than deleting offensive comments or banning users, which may be perceived as censorship, text detoxification enables platforms to rewrite toxic messages automatically, preserving the user's original intent while ensuring a safe online environment. As LLMs become more prevalent, text detoxification acts as an essential safety feature, preventing models from generating harmful content, even when prompted with aggressive language. Finally, text detoxification algorithms can power writing assistants that help users to rephrase emotionally charged emails or messages, providing constructive, professional feedback before they are sent.

Relevance of the work. This work addresses critical NLP challenges with real-world impact. First, online toxicity demands effective detoxification systems to combat harassment while preserving content. Second, multilingual solutions are needed as most languages lack parallel detoxification data. Third, we tackle computational barriers through model compression. Our methods for cross-lingual transfer and synthetic data generation advance text style transfer while enabling practical applications in content moderation and safer AI. The released resources support future multilingual detoxification research.

Scientific novelty. In this manuscript we describe several novel approaches and ideas. We conduct the first comprehensive empirical comparisons of cross-lingual detoxification transfer strategies, demonstrating effective transfer is achievable [10]. Next, we introduce FWTTM, a novel method for task-aware low-rank tensor decomposition specifically applied to compressing Transformer models, showing competitive results particularly at higher compression rates [131, 11]. Furthermore, we propose how modern LLMs, combined with few-shot prompting and activation

patching, can generate human-quality parallel detoxification data across multiple languages and show, that fine-tuning on this synthetic data can yield models that match or surpass limited human data in downstream task performance [130, 9]. We also propose a novel human-annotated parallel detoxification dataset and a set of evaluation metrics [10].

Theoretical and practical significance. This manuscript and the research work it summarizes provides significant theoretical and practical contributions to both the theory and real-world use of multilingual TST. From the theoretical point of view, we describe how NLP models can learn and transfer TST knowledge across languages, how these models could be low-rank approximated and how well LLMs with moderate adjustments can simulate humans and produce parallel data of the same quality. From the practical side, this research makes multilingual TST more accessible especially for the low-resource languages as we release all the collected synthetic parallel data for the public use. Moreover, we also release the code for our low-rank approximation experiments.

Research methodology. The methodology of this dissertation is a combination of cross-lingual model evaluation, model compression, and LLM-based synthetic data generation. We evaluate supervised and zero-shot methods for multilingual text detoxification, explore low-rank adaptation of transformers using Fisher-weighted decompositions, and use modern LLMs for the generation of parallel training data. We do our experiments using both publicly available and custom-constructed datasets, which we further release. We use both automatic and human evaluation, releasing the detailed annotation instructions, to assess the performance of our models across multiple languages and tasks.

Main results submitted for the defense. The following key results of the dissertation are submitted for the defense:

1. Effective zero-shot cross-lingual text detoxification transfer between languages like English and Russian is possible and can be achieved using methods such as adapter-based fine-tuning and backtranslation pipelines, contrary to earlier findings suggesting inherent difficulty.

2. Simultaneous detoxification and translation can be performed effectively in a single end-to-end model trained on synthetically generated cross-lingual parallel data, achieving performance competitive with multi-step pipeline approaches while offering computational efficiency.

3. Incorporating task-specific parameter importance via Fisher information into low-rank decomposition methods, specifically through the proposed FWTTM technique, yields compressed Transformer models that maintain competitive performance on downstream tasks (NLU and text detoxification) compared to standard SVD or TTM, especially at higher compression rates.

4. Modern LLMs, when combined with appropriate prompting strategies (few-shot examples) and techniques to mitigate alignment constraints (activation patching), can serve as effective annotators for generating large-scale, high-quality synthetic parallel text detoxification data across multiple languages.

5. Sequence-to-sequence models trained exclusively on large-scale, high-quality synthetic parallel detoxification data generated by LLMs can achieve performance comparable to or exceeding that of models trained on smaller, human-annotated parallel datasets.

Approbation of the work and publications. The main statements and experimental results of the dissertation were peer-reviewed and published in the proceedings of eight international conferences. The key venues where the work was presented in the form of official publications, poster presentations, and/or video reports include top-tier conferences such as ACL, EMNLP (CORE A*), NAACL (CORE A), and COLING (CORE B). All of these publications are indexed by Web of Science and/or Scopus, attesting to the high quality and international recognition of the research.

1. 60th Annual Meeting of the Association for Computational Linguistics (ACL), Dublin, Ireland, May 2022.

2. 11th International Conference on Analysis of Images, Social Networks, and Texts (AIST), Yerevan, Armenia, September 2023.

3. 13th International Joint Conference on Natural Language Processing and 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL), Nusa Dua, Bali, Indonesia, November 2023.

4. 2024 Conference and Labs of the Evaluation Forum (CLEF), Grenoble, France, September 2024.

5. 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), Miami, Florida, USA, November 2024.

6. 31st International Conference on Computational Linguistics (COLING), Abu Dhabi, United Arab Emirates, January 2025.

7. 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), Albuquerque, New Mexico, USA, April 2025.

Personal contribution. The experimental results presented in this manuscript were obtained by the author or with his direct participation. The author's contribution encompasses establishing research objectives, examining pertinent literature, designing and executing experiments, and analyzing the results obtained.

Specifically, in Chapter 2, the author was responsible for the experimental design and execution in the relevant studies [10, 132]. For Chapter 3, the author conducted the downstream compression experiments for the FWTTM approach [131] and the model compression experiments for the GFWSVD method [12]. Regarding Chapter 4, the author was responsible for the preparation and execution of the text detoxification shared task [133] (including the competition page, backend, and evaluation pipelines) and the subsequent analysis [129]. Finally, for Chapter 5, the author conceived the initial idea and led the project development throughout all critical stages of the study [130].

Structure of the dissertation. The dissertation is organized into five chapters, plus introduction, conclusion, bibliography and supplementary materials. Following this introduction, Chapter 1 lays the foundational concepts regarding TST architectures, multilingual models, parallel data challenges, and evaluation principles. Chapter 2 presents a detailed investigation into methods for cross-lingual text detoxification transfer. In Chapter 3 we introduce and evaluate novel Fisher-weighted low-rank approximation technique FWTTM for compressing transformer text detoxification model. Next, in Chapter 4 we discuss the extrapolation of the parallel data to nine diverse languages and how we have conducted a multilingual text detoxification shared task. In Chapter 5 we explore the use of LLMs for generating synthetic parallel detoxification data firstly for English with the help of activation patching and then for the case of multiple languages and expert models. Finally, the Conclusion summarizes the main contributions, discusses limitations, and outlines directions for future research. Supplementary materials provide additional details on datasets, implementation, and results.

Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Заключение диссертации по теме «Другие cпециальности», Московский Даниил Алексеевич

Conclusions

This dissertation tackled two major problems preventing wider use of text detoxification systems: the lack of training data for many languages and the high computational costs of powerful models. As online communication grows globally, we urgently need better tools to reduce toxic content across different languages. My research developed new methods to solve both the data shortage and efficiency problems, making multilingual text detoxification more practical.

The main contributions of this work fall into three areas. First, we showed that detoxification skills can successfully transfer between languages. Through extensive experiments, we proved that methods like adapter-based fine-tuning and backtranslation can effectively move detoxification abilities from English to other languages like Russian. This work also created better ways to measure performance in cross-lingual settings.

Second, we developed a new model compression technique called Fisher-Weighted Tensor Train Matrix Decomposition (FWTTM). This method smartly reduces model size by focusing on keeping the most important parameters for the detoxification task. Tests showed it works better than standard compression methods, especially when making models much smaller while keeping good performance.

Third, we demonstrated how large language models (LLMs) can create synthetic training data at scale. My methods for prompting LLMs and handling toxic content produced large datasets that trained models nearly as good as those using human-made data. This solves a major problem in the field by providing cheap, scalable data for many languages.

Together, these advances help build better detoxification systems that can work across languages while using fewer resources. The findings improve our theoretical understanding of cross-lingual style transfer, model compression, and synthetic data generation. Practically, they provide concrete tools for developing more efficient and widely applicable content moderation systems.

However, some limitations remain. The studies focused mainly on English and Russian, and on obvious toxicity rather than subtle harmful language. The compression methods, while effective, still involve trade-offs between size and performance. Synthetic data, though useful, may miss some nuances of human language. Evaluation metrics, though improved, still can't perfectly capture all aspects of good style transfer.

Future work should expand to more languages, especially those with few resources. Researchers could develop better ways to combine different compression techniques and improve synthetic data quality. We also need better evaluation methods and more study of ethical issues in this field.

In summary, this thesis makes important progress toward practical multilingual text detoxification. By developing new methods for cross-lingual transfer, model compression, and data generation, it provides foundations for safer online communication worldwide. While challenges remain, this work helps move us closer to effective, efficient systems that can benefit many languages and cultures.

Список литературы диссертационного исследования кандидат наук Московский Даниил Алексеевич, 2026 год

Bibliography

1. Brown, T. B. Language models are few-shot learners / T. B. Brown, B. Mann, N. Ryder, [et al.] // Proc. of the 34th International Conference on Neural Information Processing Systems. Vol. 33.-2020.-P. 1877-1901.

2. Logacheva, V. ParaDetox: Detoxification with Parallel Data / V. Logacheva, D. Dementieva, S. Ustyantsev, [et al.] // Proc. of the 60th Annual Meeting of the Association for Computational Linguistics / Long Papers. Vol. 1. -— 2022. -— P. 6804-6818.

3. Lewis, M. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension / M. Lewis, Y. Liu, N. Goyal, [et al.] // Proc. of the 58th Annual Meeting of the Association for Computational Linguistics.-2020.-P. 7871-7880.

4. Raffel, C. Exploring the limits of transfer learning with a unified text-to-text transformer / C. Raffel, N. Shazeer, A. Roberts, [et al.] // Journal of Machine Learning Research.-2020.-Vol. 21, no. 1.-P. 1532-4435.

5. Dementieva, D. RUSSE-2022: Findings of the First Russian Detoxification Shared Task Based on Parallel Corpora / D. Dementieva, V. Logacheva, I. Nikishina, [et al.] // Computational Linguistics and Intellectual Technologies.-2022.-P. 114-131.

6. Atwell, K. APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media Conversations / K. Atwell, S. Hassan, M. Alikhani // Proc. of the 29th International Conference on Computational Linguistics.-2022.-P. 6063-6074.

7. Radford, A. Improving Language Understanding by Generative Pre-Training /

tech. rep. / A. Radford, K. Narasimhan, T. Salimans, [et al.] / OpenAI. -

2018. -URL: https : / / cdn . openai. com / research - covers / language -

unsupervisedAanguage_understanding_paper.pdf; Visited on 01/05/2025.

8. Reif, E. A Recipe for Arbitrary Text Style Transfer with Large Language Models / E. Reif, D. Ippolito, A. Yuan, [et al.] // Proc. of the 60th Annual Meeting of the

Association for Computational Linguistics / Short Papers. Vol. 2.-2022.-

P. 837-848.

9. Moskovskiy, D. SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators / D. Moskovskiy, N. Sushko, S. Pletenev, [et al.] // Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies / Long Papers. Vol. 1.-2025.-P. 5714-5733.

10. Dementieva, D. Exploring Methods for Cross-lingual Text Style Transfer: The Case of Text Detoxification / D. Dementieva, D. Moskovskiy, D. Dale, [et al.] // Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the

Association for Computational Linguistics / Long Papers. Vol. 1.-2023.-

P. 1083-1101.

11. Pletenev, S. A Computational Study of Matrix Decomposition Methods for Compression of Pre-trained Transformers / S. Pletenev, V. Chekalina, D. Moskovskiy, [et al.] // Proc. of the 37th Pacific Asia Conference on Language, Information and Computation.-2023.-P. 723-742.

12. Chekalina, V. A. Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models / V. A. Chekalina, D. Moskovskiy, M. Kurkin, [et al.] //NeurIPS 2025 Conference

Submission. - 2025. - URL: https : / / openreview . net / forum ? id =

TW2OgVtqQR; Under review.

13. Jin, D. Deep Learning for Text Style Transfer: A Survey / D. Jin, Z. Jin, Z. Hu, [et al.] // Computational Linguistics.-2022.-Vol. 48, no. 1.-P. 155-205.

14. Hovy, E. H. Interpretation in Generation / E. H. Hovy // Proc. of the 6th National Conference on Artificial Intelligence.- 1987.-P. 545-549.

15. Huang, C. Automatic Dialogue Generation with Expressed Emotions / C. Huang, O. Zaiane, A. Trabelsi, [et al.] // Proc. of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies / Short Papers. Vol. 2.-2018.-P. 49-54.

16. Thapliyal, A. V. Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage / A. V. Thapliyal, R. Soricut // Proc. of the 58th

Annual Meeting of the Association for Computational Linguistics.-2020.-

P. 160-170.

17. Cao, Y. Expertise Style Transfer: A New Task Towards Better Communication between Experts and Laymen / Y. Cao, R. Shui, L. Pan, [et al.] // Proc. of the 58th

Annual Meeting of the Association for Computational Linguistics.-2020.-

P. 1061-1071.

18. Shen, T. Style transfer from non-parallel text by cross-alignment / T. Shen, T. Lei, R. Barzilay, [et al.] // Proc. of the 31st International Conference on Neural Information Processing Systems. Vol. 31.-2017.-P. 6833-6844.

19. Fu, Z. Style Transfer in Text: Exploration and Evaluation / Z. Fu, X. Tan, N. Peng,

[et al.] // Proc. of the AAAI Conference on Artificial Intelligence.-2018.-

Vol. 32, no. 1.-P. 663-670.

20. John, V. Disentangled Representation Learning for Non-Parallel Text Style Transfer / V. John, L. Mou, H. Bahuleyan, [et al.] // Proc. of the 57th Annual

Meeting of the Association for Computational Linguistics. - 2019. -

P. 424-434.

21. Xu, H. VAE based Text Style Transfer with Pivot Words Enhancement Learning / H. Xu, S. Lu, Z. Sun, [et al.] // Proc. of the 18th International Conference on Natural Language Processing (ICON).-2021.-P. 162-172.

22. Lample, G. Multiple-Attribute Text Rewriting / G. Lample, S. Subramanian, E. M. Smith, [et al.] // Proc. of the 7th International Conference on Learning Representations.-2019.-P. 1-20.-Visited on 12/12/2023.

23. Li, X. Text Style Transfer: Leveraging a Style Classifier on Entangled Latent Representations / X. Li, S. Sun, Y. Wang // Proc. of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021).-2021.-P. 72-82.

24. Narasimhan, S. Towards Robust and Semantically Organised Latent Representations for Unsupervised Text Style Transfer / S. Narasimhan, S. Dey, M. Desarkar // Proc. of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.-2022.-P. 456-474.

25. Li, J. Delete, Retrieve, Generate: a Simple Approach to Sentiment and Style Transfer / J. Li, R. Jia, H. He, [et al.] // Proc. of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies / Long Papers. Vol. 1.-2018.-P. 1865-1874.

26. Dale, D. Text Detoxification using Large Pre-trained Neural Models / D. Dale, A. Voronov, D. Dementieva, [et al.] // Proc. of the 2021 Conference on Empirical Methods in Natural Language Processing.-2021.-P. 7979-7996.

27. Suzgun, M. Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models / M. Suzgun, L. Melas-Kyriazi, D. Jurafsky // Proc. of the 2022 Conference on Empirical

Methods in Natural Language Processing. - 2022. -P. 2195-2222. -

https://aclanthology.org/2022.emnlp-main.141.pdf.

28. Riley, P. TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling / P. Riley, N. Constant, M. Guo, [et al.] // Proc. of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing / Long Papers. Vol. 1.-2021.-P. 3786-3800.

29. Rao, S. Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer / S. Rao, J. Tetreault // Proc. of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies / Long Papers. Vol. 1.-2018.-P. 129-140.

30. Liu, Y. Multilingual Denoising Pre-training for Neural Machine Translation / Y. Liu, J. Gu, N. Goyal, [et al.] // Transactions of the Association for Computational Linguistics.-2020.-Vol. 8.-P. 726-742.

31. Xue, L. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer / L. Xue, N. Constant, A. Roberts, [et al.] // Proc. of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.-2021.-P. 483-498.

32. Zmitrovich, D. A Family of Pretrained Transformer Language Models for Russian / D. Zmitrovich, A. Abramov, A. Kalmykov, [et al.] // Proc. of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation.-2024.-P. 507-524.

33. Fan, A. Beyond English-Centric Multilingual Machine Translation / A. Fan,

S. Bhosale, H. Schwenk, [et al.] // Journal of Machine Learning Research.-

2021.-Vol. 22, no. 1.-P. 1532-4435.

34. Ng, N. Facebook FAIR's WMT19 News Translation Task Submission / N. Ng, K. Yee, A. Baevski, [et al.] // Proc. of the Fourth Conference on Machine Translation / Shared Task Papers. Vol. 2.-2019.-P. 314-319.

35. Tiedemann, J. OPUS-MT - Building open translation services for the World / J. Tiedemann, S. Thottingal // Proc. of the 22nd Annual Conference of the European Association for Machine Translation.-2020.-P. 479-480.

36. Lison, P. 0penSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles / P. Lison, J. Tiedemann // Proc. of the Tenth International Conference on Language Resources and Evaluation.-2016.-P. 923-929.

37. Tiedemann, J. The Tatoeba Translation Challenge - Realistic Data Sets for Low Resource and Multilingual MT / J. Tiedemann // Proc. of the Fifth Conference on Machine Translation.-2020.-P. 1174-1182.

38. Tiedemann, J. Parallel Data, Tools and Interfaces in OPUS / J. Tiedemann // Proc. of the Eighth International Conference on Language Resources and Evaluation.-2012.-P. 2214-2218.

39. Creutz, M. Open Subtitles Paraphrase Corpus for Six Languages / M. Creutz // Proc. of the Eleventh International Conference on Language Resources and Evaluation.-2018.-P. 1364-1369.

40. Houlsby, N. Parameter-Efficient Transfer Learning for NLP / N. Houlsby, A. Giurgiu, S. Jastrzebski, [et al.] // Proc. of the 36th International Conference on Machine Learning. Vol. 97.-2019.-P. 2790-2799.

41. Liu, Y. RoBERTa: A Robustly Optimized BERT Pretraining Approach / Y. Liu,

M. Ott, N. Goyal, [et al.].-2019.-arXiv: 1907.11692.-URL: http:

//arxiv.org/abs/1907.11692; Visited on 15/11/2022.

42. Adams, C. J. Jigsaw Toxic Comment Classification Challenge / C. J. Adams,

J. Sorensen, J. Elliott, [et al.]. -2017. -URL: https : //kaggle . com/

competitions / jigsaw - toxic - comment - classification - challenge; Visited on 03/10/2022.

43. Kuratov, Y. Adaptation of Deep Bidirectional Multilingual Transformers for

Russian Language / Y. Kuratov, M. Y. Arkhipov.-2019.-arXiv: 1905.

07213.-URL: https://doi.org/10.48550/arXiv.1905.07213; Visited on

9/11/2022.

44. Gusev, I. Russian Texts Detoxification with Levenshtein Editing /1. Gusev.-

2022.-arXiv: 2204.13638.-URL: https://doi.org/10.48550/arXiv.2204.

13638; Visited on 18/11/2022.

45. Sellam, T. BLEURT: Learning Robust Metrics for Text Generation / T. Sellam, D. Das, A. Parikh // Proc. of the 58th Annual Meeting of the Association for Computational Linguistics.-2020.-P. 7881-7892.

46. Babakov, N. A large-scale computational study of content preservation measures for text style transfer and paraphrase generation / N. Babakov, D. Dale, V. Logacheva, [et al.] // Proc. of the 60th Annual Meeting of the Association

for Computational Linguistics: Student Research Workshop. - 2022. -

P. 300-321.

47. Gudkov, V. Automatically Ranked Russian Paraphrase Corpus for Text Generation / V. Gudkov, O. Mitrofanova, E. Filippskikh // Proc. of the Fourth Workshop on Neural Generation and Translation.-2020.-P. 54-59.

48. Martynov, N. RuPAWS: A Russian Adversarial Dataset for Paraphrase Identification / N. Martynov, I. Krotova, V. Logacheva, [et al.] // Proc. of

the Thirteenth Language Resources and Evaluation Conference.-2022.-

P. 5683-5691.

49. Warstadt, A. Neural Network Acceptability Judgments / A. Warstadt, A. Singh, S. R. Bowman // Transactions of the Association for Computational Linguistics.-2019.-Vol. 7.-P. 625-641.

50. Mikhailov, V. RuCoLA: Russian Corpus of Linguistic Acceptability /

V. Mikhailov, T. Shamardina, M. Ryabinin, [et al.]. - 2022. - arXiv:

2210.12814.-URL: https://doi.org/10.48550/arXiv.2210.12814; Visited on

1/12/2022.

51. Logacheva, V. A Study on Manual and Automatic Evaluation for Text Style Transfer: The Case of Detoxification / V. Logacheva, D. Dementieva, I. Krotova, [et al.] // Proc. of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval).-2022.-P. 90-101.

52. Vaswani, A. Attention is all you need / A. Vaswani, N. Shazeer, N. Parmar, [et al.] // Proc. of the 31st International Conference on Neural Information Processing Systems.-2017.-P. 6000-6010.

53. Devlin, J. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding / J. Devlin, M.-W. Chang, K. Lee, [et al.] // Proc. of the 2019 Conference of the North American Chapter of the Association for Computational

Linguistics: Human Language Technologies / Long and Short Papers. Vol. 1.-

2019.-P. 4171-4186.

54. Strubell, E. Energy and Policy Considerations for Deep Learning in NLP / E. Strubell, A. Ganesh, A. McCallum // Proc. of the 57th Annual Meeting of the Association for Computational Linguistics.-2019.-P. 3645-3650.

55. Denil, M. Predicting parameters in deep learning / M. Denil, B. Shakibi, L. Dinh, [et al.] // Proc. of the 27th International Conference on Neural Information Processing Systems / Long Papers. Vol. 2.-2013.-P. 2148-2156.

56. Eckart, C. The Approximation of One Matrix by Another of Lower Rank

/ C. Eckart, G. Young // Psychometrika. - 1936. - Vol. 1, no. 3. -

P. 211-218.

57. Hsu, Y. Language Model Compression with Weighted Low-Rank Factorization / Y. Hsu, T. Hua, S. Chang, [et al.] // Proc. of the 10th International Conference on Learning Representations.-2022.-P. 1-15.

58. Sanh, V. Movement pruning: adaptive sparsity by fine-tuning / V. Sanh, T. Wolf, A. M. Rush // Proc. of the 34th International Conference on Neural Information Processing Systems. Vol. 33.-2020.-P. 20378-20389.

59. Le Cun, Y. Optimal brain damage / Y. Le Cun, J. S. Denker, S. A. Solla // Proc. of

the 3rd International Conference on Neural Information Processing Systems.-

1989.-P. 598-605.

60. Casella, G. C. Theory of Point Estimation / G. C. Casella.-2001.

61. Amari, S. Natural Gradient Works Efficiently in Learning / S. Amari // Neural Comput.- 1998.-Vol. 10, no. 2.-P. 251-276.

62. Kunstner, F. Limitations of the empirical Fisher approximation for natural gradient descent / F. Kunstner, P. Hennig, L. Balles // Proc. of the 32nd

International Conference on Neural Information Processing Systems. -

2019.-P. 4158-4169.

63. Srebro, N. Weighted low-rank approximations / N. Srebro, T. Jaakkola // Proc. of the Twentieth International Conference on International Conference on Machine Learning.-2003.-P. 720-727.

64. Paszke, A. PyTorch: an Imperative Style, High-Performance Deep Learning Library / A. Paszke, S. Gross, F. Massa, [et al.] // Proc. of the 33rd International

Conference on Neural Information Processing Systems. Vol. 32.-2019.-

P. 721-732.

65. Guo, D. Parameter-Efficient Transfer Learning with Diff Pruning / D. Guo, A. Rush, Y. Kim // Proc. of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing / Long Papers. Vol. 1.-2021.-P. 4884-4896.

66. Wang, A. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding / A. Wang, A. Singh, J. Michael, [et al.] // Proc. of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP.-2018.-P. 353-355.

67. Socher, R. Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank / R. Socher, A. Perelygin, J. Wu, [et al.] // Proc. of the

2013 Conference on Empirical Methods in Natural Language Processing.-

2013.-P. 1631-1642.

68. Dolan, W. B. Automatically Constructing a Corpus of Sentential Paraphrases / W. B. Dolan, C. Brockett // Proc. of the Third International Workshop on Paraphrasing.-2005.-P. 9-16.

69. Cer, D. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation / D. Cer, M. Diab, E. Agirre, [et al.] // Proc. of

the 11th International Workshop on Semantic Evaluation (SemEval-2017).-

2017.-P. 1-14.

70. Rahman, A. Resolving Complex Cases of Definite Pronouns: The Winograd Schema Challenge / A. Rahman, V. Ng // Proc. of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning.-2012.-P. 777-789.

71. Williams, A. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference / A. Williams, N. Nangia, S. Bowman // Proc. of the 2018 Conference of the North American Chapter of the Association for Computational

Linguistics: Human Language Technologies / Long Papers. Vol. 1.-2018.-

P. 1112-1122.

72. Oseledets, I. V. Tensor-Train Decomposition /1. V. Oseledets // SIAM J. Sci. Comput.-2011.-Vol. 33, no. 5.-P. 2295-2317.

73. Sanh, V. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and

lighter / V. Sanh, L. Debut, J. Chaumond, [et al.].-2019.-arXiv: 1910.

01108.-URL: https://doi.org/10.48550/arXiv.1910.01108; Visited on

22/11/2022.

74. Lagunas, F. Block Pruning For Faster Transformers / F. Lagunas, E. Charlaix, V. Sanh, [et al.] // Proc. of the 2021 Conference on Empirical Methods in Natural Language Processing.-2021.-P. 10619-10629.

75. Loshchilov, I. Decoupled Weight Decay Regularization / I. Loshchilov, F. Hutter // Proc. of the 7th International Conference on Learning Representations.-2019.-P. 1-8.

76. Van Loan, C. F. Approximation with Kronecker Products / C. F. Van Loan,

N. Pitsianis // Linear Algebra for Large Scale and Real-Time Applications.-

1993.-P. 293-314.

77. Yuan, Z. ASVD: Activation-aware Singular Value Decomposition for

Compressing Large Language Models / Z. Yuan, Y. Shang, Y. Song, [et al.].-

2023.-arXiv: 2312.05821.-URL: https://doi.org/10.48550/arXiv.2312.

05821; Visited on 12/02/2024.

78. Wang, X. SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression / X. Wang, Y. Zheng, Z. Wan, [et al.] //

Proc. of the 13th International Conference on Learning Representations.-

2025.-P. 1-21.

79. Merity, S. Pointer Sentinel Mixture Models / S. Merity, C. Xiong, J. Bradbury, [et

al.] // Proc. of the 5th International Conference on Learning Representations.-

2017.-P. 1-15.

80. Marcus, M. P. Building a Large Annotated Corpus of English: The Penn Treebank / M. P. Marcus, B. Santorini, M. A. Marcinkiewicz // Comput. Linguistics.- 1993.-Vol. 19, no. 2.-P. 313-330.

81. Hendrycks, D. Measuring Massive Multitask Language Understanding / D. Hendrycks, C. Burns, S. Basart, [et al.] // Proc. of the 9th International Conference on Learning Representations.-2021.-P. 1-27.

82. Penedo, G. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale / G. Penedo, H. Kydlícek, L. B. Allal, [et al.] // Proc. of the 38th International Conference on Neural Information Processing Systems. Vol. 37.-2025.-P. 30811-30849.

83. Pavao, A. CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges / A. Pavao, I. Guyon, A.-C. Letournel, [et al.] // Journal of Machine Learning Research.-2023.-Vol. 24, no. 198.-P. 1-6.

84. Dementieva, D. MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages / D. Dementieva, N. Babakov, A. Panchenko // Proc. of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies / Short Papers. Vol. 2.-2024.-P. 124-140.

85. Belchikov, A. Russian Language Toxic Comments Dataset / A. Belchikov.-

2019. -URL: https : / / www. kaggle . com / datasets / blackmoon / russian -

language-toxic-comments; Visited on: 14/10/2024.

86. Kivlichan, I. Jigsaw Multilingual Toxic Comment Classification Challenge /

I. Kivlichan, J. Sorensen, J. Elliott, [et al.].-2020.-URL: https://kaggle.

com/ competitions/jigsaw - multilingual - toxic - comment - classification; Visited on 03/10/2022.

87. Bobrovnyk, K. Automated building and analysis of Ukrainian Twitter corpus for toxic text detection / K. Bobrovnyk // Proc. of the 3rd International Conference

on Computational Linguistics and Intelligent Systems / Workshop. Vol. 2.-

2019.-URL: https://ena.lpnu.ua:8443/server/api/core/bitstreams/c4c645c1-

f465-4895-98dd-765f862cf186/content.

88. Pereira-Kohatsu, J. C. Detecting and Monitoring Hate Speech in Twitter /

J. C. Pereira-Kohatsu, L. Q. Sánchez, F. Liberatore, [et al.] // Sensors. -

2019.-Vol. 19, no. 21.-P. 4654-4691.

89. Taulé, M. NewsCom-TOX: a corpus of comments on news articles annotated for toxicity in Spanish / M. Taulé, M. Nofre, V. Bargiela, [et al.] // Language Resources and Evaluation.-2024.-Vol. 58, no. 4.-P. 1115-1155.

90. Pérez, J. M. RoBERTuito: a pre-trained language model for social media text in Spanish / J. M. Pérez, D. A. Furman, L. Alonso Alemany, [et al.] // Proc. of

the Thirteenth Language Resources and Evaluation Conference.-2022.-

P. 7235-7243.

91. Wiegand, M. Overview of the GermEval 2018 Shared Task on the Identification of Offensive Language / M. Wiegand, M. Siegel, J. Ruppenhofer // Proc.

of GermEval 2018, 14th Conference on Natural Language Processing. -

2018.-P. 1-10.

92. Risch, J. Overview of the GermEval 2021 Shared Task on the Identification of Toxic, Engaging, and Fact-Claiming Comments / J. Risch, A. Stoll, L. Wilms, [et al.] // Proc. of the GermEval 2021 Shared Task on the Identification of Toxic, Engaging, and Fact-Claiming Comments.-2021.-P. 1-12.

93. Ross, B. Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis / B. Ross, M. Rist, G. Carbonell, [et al.] // Proc. of NLP4CMC III: 3rd Workshop on Natural Language Processing for Computer-Mediated Communication. Vol. 17.-2016.-P. 6-9.

94. Mandl, T. Overview of the HASOC track at FIRE 2019 / T. Mandl, S. Modha, P. Majumder, [et al.] // Proc. of the 11th Forum for Information Retrieval Evaluation.-2019.-P. 14-17.

95. Ayele, A. A. Exploring Amharic Hate Speech Data Collection and Classification Approaches / A. A. Ayele, S. M. Yimam, T. D. Belay, [et al.] // Proc. of the 14th International Conference on Recent Advances in Natural Language Processing.-2023.-P. 49-59.

96. Ayele, A. A. The 5Js in Ethiopia: Amharic Hate Speech Data Annotation Using Toloka Crowdsourcing Platform / A. A. Ayele, S. Dinter, T. D. Belay, [et al.] // 2022 International Conference on Information and Communication Technology for Development for Africa (ICT4DA).-2022.-P. 114-120.

97. Mulki, H. L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language / H. Mulki, H. Haddad, C. Bechikh Ali, [et al.] // Proc. of the Third Workshop on Abusive Language Online.-2019.-P. 111-118.

98. Haddad, H. T-HSAB: A Tunisian Hate Speech and Abusive Dataset / H. Haddad, H. Mulki, A. Oueslati // Arabic Language Processing: From Theory to Practice.-2019.-P. 251-263.

99. Mubarak, H. Overview of OSACT4 Arabic Offensive Language Detection Shared Task / H. Mubarak, K. Darwish, W. Magdy, [et al.] // Proc. of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection.-2020.-P. 48-52.

100. Lu, J. Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks / J. Lu, B. Xu, X. Zhang, [et al.] // Proc. of the 61st Annual Meeting of the Association for Computational Linguistics.-2023.-P. 16235-16250.

101. Feng, F. Language-agnostic BERT Sentence Embedding / F. Feng, Y. Yang, D. Cer, [et al.] // Proc. of the 60th Annual Meeting of the Association for Computational Linguistics / Long Papers. Vol. 1.-2022.-P. 878-891.

102. Popovic, M. chrF: character n-gram F-score for automatic MT evaluation / M. Popovic // Proc. of the Tenth Workshop on Statistical Machine Translation.-2015.-P. 392-395.

103. Wei, J. Chain-of-thought prompting elicits reasoning in large language models / J. Wei, X. Wang, D. Schuurmans, [et al.] // Proc. of the 36th International

Conference on Neural Information Processing Systems. Vol. 36.-2022.-

P. 24824-24837.

104. Conneau, A. Cross-lingual language model pretraining / A. Conneau, G. Lample // Proc. of the 33rd International Conference on Neural Information Processing Systems. Vol. 32.-2019.-P. 634-645.

105. Muennighoff, N. Crosslingual Generalization through Multitask Finetuning / N. Muennighoff, T. Wang, L. Sutawika, [et al.] // Proc. of the 61st Annual Meeting of the Association for Computational Linguistics / Long Papers. Vol. 1.-2023.-P. 15991-16111.

106. Arditi, A. Refusal in language models is mediated by a single direction / A. Arditi, O. Obeso, A. Syed, [et al.] // Proc. of the 38th International

Conference on Neural Information Processing Systems. Vol. 37.-2025.-

P. 136037-136083.

107. Dubey, A. The Llama 3 Herd of Models / A. Dubey, A. Jauhri, A. Pandey, [et

al.].-2024.-arXiv: 2407.21783.-URL: https://doi.org/10.48550/

arXiv.2407.21783; Visited on 1/08/2024.

108. Khondaker, M. T. I. DetoxLLM: A Framework for Detoxification with Explanations / M. T. I. Khondaker, M. Abdul-Mageed, L. V. S. Lakshmanan // Proc. of the 2024 Conference on Empirical Methods in Natural Language Processing.-2024.-P. 19112-19139.

109. Sachdeva, P. S. The Measuring Hate Speech Corpus: Leveraging Rasch Measurement Theory for Data Perspectivism / P. S. Sachdeva, R. Barreto, G. Bacon, [et al.] // Proc. of the 1st Workshop on Perspectivist Approaches to NLPerspectives.-2022.-P. 83-94.

110. Hartford, E. Cognitive Computations: The Dolphin Dataset / E. Hartford.-

2023. -URL: https : / / huggingface . co / datasets / cognitivecomputations /

dolphin; Visited on 15/06/2024.

111. Lees, A. A New Generation of Perspective API: Efficient Multilingual Character-level Transformers / A. Lees, V. Q. Tran, Y. Tay, [et al.] // Proc. of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining.-2022.-P. 3197-3207.

112. Semiletov, A. Toxic Russian Comments Dataset / A. Semiletov.-2020.-

URL: https : // www. kaggle. com/ datasets /alexandersemiletov/toxic - russian-comments; Visited on 03/09/2022.

113. Assenmacher, D. RP-Mod & RP-Crowd: Moderator- and Crowd-Annotated German News Comment Datasets / D. Assenmacher, M. Niemann, K. Müller, [et al.] // Proc. of the Neural Information Processing Systems

Track on Datasets and Benchmarks. Vol. 1. - 2021. - URL: https :

/ / datasets - benchmarks - proceedings . neurips . cc / paper / 2021 / hash / c9e1074f5b3f9fc8ea15d152add07294-Abstract-round2.html.

114. Capozzi, A. Clandestino or Rifugiato? Anti-immigration Facebook Ad Targeting in Italy / A. Capozzi, G. De Francisci Morales, Y. Mejova, [et al.] // Proc. of the 2021 CHI Conference on Human Factors in Computing Systems.-2021.

115. Ousidhoum, N. Multilingual and Multi-Aspect Hate Speech Analysis / N. Ousidhoum, Z. Lin, H. Zhang, [et al.] // Proc. of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International

Joint Conference on Natural Language Processing (EMNLP-IJCNLP). -

2019.-P. 4675-4684.

116. Yang, A. Qwen2 Technical Report / A. Yang, B. Yang, B. Hui, [et al.]. -

2024.-arXiv: 2407.10671.-URL: https://doi.org/10.48550/arXiv.

2407.10671; Visited on 15/06/2024.

117. Yang, A. Qwen2.5 Technical Report / A. Yang, B. Yang, B. Zhang, [et al.].-

2024.-arXiv: 2412.15115.-URL: https://doi.org/10.48550/arXiv.2412.

15115; Visited on 22/12/2024.

118. Gomes, A. Updates to the Command R Series / A. Gomes.-2024.-URL:

https://cohere.com/blog/command-series-0824; Visited on 15/10/2024.

119. Rivière, M. Gemma 2: Improving Open Language Models at a Practical Size /

M. Rivière, S. Pathak,P. G. Sessa, [etal.].-2024.-arXiv: 2408.00118.-

URL: https://doi.org/10.48550/arXiv.2408.00118; Visited on 10/09/2024.

120. Dang, J. Aya Expanse: Combining Research Breakthroughs for a New

Multilingual Frontier / J. Dang, S. Singh, D. D'souza, [et al.].- 2024.-

arXiv: 2412.04261.-URL: https://doi.org/10.48550/arXiv.2412.04261;

Visited on 15/12/2024.

121. Conover, M. Free Dolly: Introducing the World's First Truly Open

Instruction-Tuned LLM / M. Conover, M. Hayes, A. Mathur, [et al.]. -

2023.-URL: https://www.databricks.com/blog/2023/04/12/dolly-first-

open-commercially-viable-instruction-tuned-llm; Visited on 02/07/2023.

122. AI, M. AI in Abundance / M. AI.-2024.-URL: https://mistral.ai/news/

september-24-release/; Visited on 15/10/2024.

123. AI, M. Mistral NeMo / M. AI.- 2024.-URL: https://mistral.ai/news/

mistral-nemo/; Visited on 16/10/2024.

124. Mondshine, I. Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs / I. Mondshine,

T. Paz-Argaman, R. Tsarfaty. - 2025. -arXiv: 2502. 09331. -URL:

https://doi.org/10.48550/arXiv.2502.09331; Visited on 05/03/2025.

125. Conneau, A. Unsupervised Cross-lingual Representation Learning at Scale / A. Conneau, K. Khandelwal, N. Goyal, [et al.] // Proc. of the 58th Annual

Meeting of the Association for Computational Linguistics. - 2020. -

P. 8440-8451.

126. Zhang, Z. MELA: Multilingual Evaluation of Linguistic Acceptability / Z. Zhang, Y. Liu, W. Huang, [et al.] // Proc. of the 62nd Annual Meeting

of the Association for Computational Linguistics / Long Papers. Vol. 1. -

2024.-P. 2658-2674.

127. Xu, B. S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation Extraction / B. Xu, Q. Wang, Y. Lyu, [et al.] // Proc. of the 61st Annual Meeting of the Association for Computational Linguistics / Long Papers. Vol. 1.-2023.-P. 8186-8207.

128. Hurst, A. GPT-4o System Card / A. Hurst, A. Lerer, A. P. Goucher, [et al.].-

2024.-arXiv: 2410.21276.-URL: https://doi.org/10.48550/arXiv.2410.

21276; Visited on 30/10/2024.

Author's publications on the dissertation subject

129. Dementieva, D. Multilingual and Explainable Text Detoxification with Parallel Corpora / D. Dementieva, N. Babakov, A. Ronen, A. A. Ayele, N. Rizwan, F. Schneider, X. Wang, S. M. Yimam, D. Moskovskiy, E. Stakovskii, E. Kaufman, A. Elnagar, A. Mukherjee, A. Panchenko // Proceedings - International Conference on Computational Linguistics, COLING. Vol. 206484-1.-2025.-P. 7998-8025.

130. Moskovskiy, D. LLMs to Replace Crowdsourcing For Parallel Data Creation? The Case of Text Detoxification / D. Moskovskiy, S. Pletenev, A. Panchenko // EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2024.-2024.-P. 14361-14373.

131. Pletenev, S. Transformers Compression: A Study of Matrix Decomposition Methods Using Fisher Information / S. Pletenev, D. Moskovskiy, V. Chekalina, M. Seleznyov, S. Zagoruyko, A. Panchenko // Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). Vol. 14486.-2024.-P. 36-48.

132. Moskovskiy, D. Exploring Cross-lingual Text Detoxification with Large Multilingual Language Models.D. Moskovskiy, D. Dementieva, A. Panchenko //

Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop.-2022.-P. 346-354.

133. Dementieva, D. Overview of the Multilingual Text Detoxification Task at PAN 2024 / D. Dementieva, D. Moskovskiy, N. Babakov, A. A. Ayele, N. Rizwan, F. Schneider, X. Wang, S. M. Yimam, D. Ustalov, E. Stakovskii, A. Smirnova, A. Elnagar, A. Mukherjee, A. Panchenko // CEUR Workshop Proceedings. Vol. 3740.-2024.-P. 2432-2461.

Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.