Надежные вопросно-ответные системы с выравниванием знаний по графу знаний (Trustworthy Question Answering system with Knowledge Graphs alignment) тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Сальников Михаил Дмитриевич
- Специальность ВАК РФ00.00.00
- Количество страниц 132
Оглавление диссертации кандидат наук Сальников Михаил Дмитриевич
Contents
Page
Introduction
Chapter 1. Background and Related Work
1.1 Knowledge Graph
1.2 Large Language Models: Role in Question Answering
1.3 Traditional Knowledge Graph Question Answering Methods
1.3.1 Semantic Parsing-based KGQA
1.3.2 Information Retrieval-based KGQA
1.4 LLM-KG Fusion Strategies
1.4.1 Retrieval-Augmented Generation (RAG) with KGs for
Context Augmentation
1.4.2 Integrating Graph Structure and Knowledge into LLM Architectures
1.4.3 Leveraging KGs for Refining LLM-Generated Answer Candidates
1.5 Specialized Aspects and Applications of KGQA
1.5.1 Factoid Question Answering
1.5.2 Multilingual KGQA
1.6 KGQA Datasets and the Rationale for ShortPathQA
Chapter 2. Answer Candidate Type Selection
2.1 Methodology
2.2 Initial Answer Candidate Generation
2.2.1 Answer Candidate Typing
2.2.2 Entity Linking
2.2.3 Candidates Scorer
2.3 Experimental Design and Baselines
2.3.1 Datasets
2.3.2 Evaluation
2.3.3 Ablation Study and Error Analysis
Page
Chapter 3. Controllable Fusion Large Language Models and
Knowledge Graphs
3.1 Answer Candidate Generation
3.2 Subgraph Extraction
3.3 Features based on Extracted Subgraphs
3.3.1 Graph Features
3.3.2 Text Features
3.3.3 Graph2Text Sequence Features
3.4 Rank LLM answer candidates using subgraphs
3.5 Experiments
3.5.1 Dataset
3.5.2 Question Entities
3.5.3 Text Embeddings
3.5.4 Graph2Text with Highlight & Context
3.5.5 Experimental Pipeline
3.5.6 Evaluation
3.6 Results & Discussion
3.6.1 Features Importance
3.6.2 Hits@N Evaluation Results
3.7 ShortPathQA: Dataset for Controllable Fusion Knowledge Graphs
and Large Language Models
3.7.1 Collection of Questions
3.7.2 Dataset Statistics
3.7.3 Baseline Evaluation
Chapter 4. System Demonstrations and Implementation
4.1 System Architecture
4.1.1 Baseline Seq2Seq Pipeline
4.1.2 M3M Pipeline
4.1.3 Answer Candidate Type (ACT) Selection Pipeline
4.1.4 Subgraph Visualization Tool
4.2 Implementation Details
4.2.1 M3M Pipeline Implementation
4.2.2 ACT Selection Pipeline Implementation
Page
4.2.3 Baseline Seq2Seq Pipeline Implementation
4.2.4 Subgraph Visualization Tool Implementation
Conclusion
List of symbols and abbreviations
Bibliography
List of Figures
List of Tables
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Эффективные методы мультиязычного текстового переноса стиля/Efficient Multilingual Text Style Transfer Methods2026 год, кандидат наук Московский Даниил Алексеевич
Исследование вариантов трансформера для различных задач обработки длинных документов/ Investigation of transformer options for various long documents processing tasks2024 год, кандидат наук Аль Адел Ариж
Доменная адаптация глубоких сверточных нейросетей для обработки медицинских изображений / Domain Adaptation of Deep Convolutional Neural Networks in Medical Imaging2026 год, кандидат наук Широких Борис Николаевич
Высокопроизводительный вычислительный поиск новых материалов для твердотельных аккумуляторов (High-throughput computational search for new solid-state battery materials)2025 год, кандидат наук Дембицкий Артем Дмитриевич
Модели и методы автоматической обработки неструктурированных данных в биомедицинской области2023 год, доктор наук Тутубалина Елена Викторовна
Введение диссертации (часть автореферата) на тему «Надежные вопросно-ответные системы с выравниванием знаний по графу знаний (Trustworthy Question Answering system with Knowledge Graphs alignment)»
Introduction
Background and Relevance of the Work. For a long time, a key goal in artificial intelligence has been to make machines that can understand and answer human questions using large collections of structured knowledge. Knowledge Graphs (KGs) are good for this because they can show how facts are connected, which helps give exact answers. The arrival of Large Language Models (LLMs) has changed this area a lot. LLMs are trained on huge amounts of text and are very good at understanding and creating language. They also store a lot of factual knowledge inside themselves. However, when we use LLMs for Knowledge Graph Question Answering (KGQA), they have some problems. A big problem is that they can create answers that sound good but are wrong. This is often called hallucination [1—4]. Also, it is hard to understand how LLMs come up with their answers, so it's difficult to check if the answers are based on facts.
Now that LLMs are used so much, KGs are even more important as a source of true, organized knowledge. LLMs are good at creating natural text and understanding general meanings. But they are not so good at handling facts that change often. They also have trouble adding new information without messing up the knowledge they already have [5]. For example, LLMs alone often cannot answer comparative questions correctly if they do not have access to structured information, as shown in our work on the CAM 2.0 system [81]. The LLM might become worse at general question answering. This happens especially if the new knowledge is mostly about certain things. The LLM might then often give those overrepresented answers and be less likely to say when it's not sure. These results show that we really need good ways to combine the text-making skills of LLMs with the facts from KGs. We should not just try to update the LLM's own knowledge. This dissertation works on this problem. It offers and tests new methods for mixing LLMs and KGs better, to make KGQA systems more trustworthy.
Dissertation Objectives. Goals and problems addressed in this dissertation formulated as a research question: how to create and test new ways to combine Large Language Models with Knowledge Graphs to make Question Answering systems that are more accurate, dependable, and controllable. The main objectives are:
- To make better answer candidates in KGQA. We can do this by using the LLMs' ability to understand meaning, together with specific type rules from
KGs. This includes creating a method called Answer Candidate Type (ACT) Selection. This method predicts the likely semantic type of an answer. It then uses this type to filter and improve candidate answers.
- To facilitate controllable and standardized community research on the fusion of Knowledge Graphs (KGs) and Large Language Models (LLMs) for KGQA, through the development and public release of the ShortPathQA dataset. This resource provides pre-computed KG subgraphs, allowing researchers to concentrate on fusion methodologies rather than complex KG processing tasks.
— To make LLM-generated answers more factually correct and controllable. We can do this by creating a system to re-rank candidate answers. This system uses structural proof from KG subgraphs. This includes finding important subgraphs that connect question entities to LLM-generated candidates. It also includes getting different features from these subgraphs and using various rankers to choose candidates that have stronger KG-based proof.
Scientific Novelty. The new scientific ideas in this dissertation are in creating and testing new ways to mix LLMs and KGs. These ways help solve important problems in KGQA:
1. New Answer Candidate Typing and Selection: The Answer Candidate Type (ACT) Selection method (Chapter 2) is a new way to improve KGQA. It is special because it combines an LLM's skill to guess the semantic type of an answer with the clear type information from a KG (for example, Wikidata's P31 'instance of' property). This type-based checking of LLM-generated candidates works like a strong filter and scoring tool. It greatly improves the quality of candidate answers, even if the LLM first gives the wrong answer but gets its type right.
2. Controllable Fusion via Subgraph Reranking: This work proposes a novel framework for improving the factuality of LLM outputs by reranking answer candidates based on structural evidence from KG subgraphs (Chapter 3). The novelty lies in:
— The systematic extraction of relevant KG subgraphs connecting question entities to multiple LLM-generated answer candidates.
— The comprehensive feature engineering from these subgraphs, encompassing graph-theoretic metrics (e.g., PageRank,
Katz centrality, density), textual features (question-answer concatenation), and various Graph2Text (G2T) sequence representations (Deterministic, T5-based, and GAP-based).
— The application and comparative analysis of diverse ranking models (from simple regression to neural rankers like MPNet) to effectively utilize these multi-modal subgraph features for promoting factually grounded answers.
3. Creation of a Specialized Evaluation Resource (ShortPathQA):
The development and publication of the ShortPathQA dataset (Chapter 3.7) is a significant novel contribution. This dataset directly supports the research on Controllable Fusion via Subgraph Reranking presented in this dissertation and aims to foster further community exploration of such methods. As the first QA resource offering pre-computed KG subgraphs, ShortPathQA allows researchers to focus on the core tasks of subgraph reasoning and reranking. This isolates these tasks from the complexities of upstream entity linking and large-scale KG processing, thereby facilitating more targeted evaluation and development of controllable fusion techniques.
Theoretical and Practical Significance. The research presented in this dissertation holds both theoretical and practical significance for the fields of Natural Language Processing and Artificial Intelligence, particularly in the domain of Knowledge Graph Question Answering.
Theoretically, this work contributes to a deeper understanding of how the complementary strengths of LLMs (which are good at semantic understanding and fluency) and KGs (which provide structured factual knowledge and verifiability) can be synergistically combined. It explores specific mechanisms for this fusion, moving beyond simple retrieval augmentation towards more nuanced methods of candidate refinement and evidence-based reranking. The investigation into type-based filtering (ACT Selection) sheds light on the implicit semantic knowledge captured by LLMs and how it can be explicitly leveraged in conjunction with KG schema for improved reasoning. The study of subgraph-based reranking contributes to theories of evidence-based reasoning in QA, demonstrating how structural path information in KGs can serve as a robust signal for factuality and control in LLM outputs. Furthermore, the creation of the ShortPathQA dataset facilitates more
focused theoretical explorations into subgraph reasoning by providing a standardized benchmark.
Practically, the methodologies developed, particularly ACT Selection and subgraph-based reranking, offer practical pathways to improve the factual accuracy and reliability of QA systems. This is crucial for real-world applications where incorrect information can have significant consequences. By making LLM outputs more controllable and grounded in verifiable KG facts, this research contributes to building more trustworthy AI systems. The ability to trace an answer back to KG evidence, as facilitated by subgraph analysis, enhances interpretability. The ACT Selection method offers a resource-efficient strategy to boost the performance of even smaller LLMs, making high-quality KGQA more accessible when computational resources are constrained. The ShortPathQA dataset provides a valuable, ready-to-use resource for other researchers and practitioners, lowering the barrier to entry for developing and evaluating LLM-KG fusion techniques without the need for extensive KG infrastructure setup. Finally, the system demonstrations (Chapter 4), including API endpoints and visualization tools, provide tangible proofs-of-concept and reusable components for building advanced KGQA applications.
Research Methodology. The research in this dissertation employs a multifaceted methodology combining theoretical development with empirical experimentation and system implementation.
1. Literature Review and Problem Formulation: The work begins with a comprehensive review of existing research in KGQA, LLMs, and LLM-KG fusion strategies (Chapter 1) to identify gaps and formulate precise research questions.
2. Method Development - ACT Selection:
- Observation of LLM behavior regarding answer type prediction, even in cases of factual error.
- Design of a pipeline that:
- Generates initial answer candidates using LLMs (e.g., T5 with Diverse Beam Search).
- Employs multilingual entity linking (e.g., mGENRE) to connect question and candidate entities to Wikidata.
- Infers the expected answer type using the LLM and aggregates type information from KG 'instance-of' relations.
— Ranks candidates using a weighted scoring mechanism incorporating type scores, KG neighborhood scores, LLM generation scores, and question-property similarity.
3. Method Development - Controllable Fusion via Subgraph Reranking:
— Hypothesis that KG path-based evidence can improve LLM answer factuality.
— Design of a reranking pipeline that:
— Generates multiple answer candidates from LLMs using Diverse Beam Search.
— Extracts KG subgraphs (shortest paths) connecting question entities to each candidate answer using Wikidata.
— Engineers diverse features from these subgraphs:
— Graph features: number of nodes/edges, cycles, bridges, average shortest path, density, Katz centrality, PageRank.
— Text features: concatenation of question and answer, encoded with MPNet.
— Graph2Text (G2T) features: using Deterministic linearization, T5-based G2T, and GAP-based G2T models, often with question context and answer highlighting.
— Employs and compares various reranking models: semantic (MPNet cosine similarity), regression (Linear, Logistic), gradient boosting (CatBoost), and neural (MPNet with regression head).
4. Dataset Creation:
— Addressing the need for a focused benchmark for subgraph-based reasoning.
— Collection of questions from Mintaka (filtered for entity answers) and manual curation of new complex questions.
— Generation of answer candidates using LLMs.
— Unified shortest path subgraph extraction from Wikidata for all question-candidate pairs.
— Publication of the dataset with questions, candidates, and pre-computed subgraphs.
5. Experimental Evaluation:
— Utilization of established KGQA datasets (SQWD, RuBQ, Mintaka) and the newly created ShortPathQA.
— Fine-tuning of LLMs (T5, spaCy NER models) on relevant training splits.
— Evaluation metrics appropriate for the tasks: Hits@1 for factoid accuracy in ACT Selection and ShortPathQA classification; Hits@N (N=1,2,3) and F1-score for reranking performance in controllable fusion and ShortPathQA.
— Ablation studies to assess the contribution of individual components and feature sets.
— Error analysis to understand model behavior and the corrective capabilities of the proposed methods.
— Comparison against strong baselines, including standalone LLMs (e.g., ChatGPT) and existing KGQA systems.
6. System Implementation and Demonstration:
— Development of practical system pipelines (baseline Seq2Seq, M3M, ACT Selection) using Python, FastAPI, PyTorch, HuggingFace Transformers, spaCy, etc.
— Creation of a subgraph visualization tool using web technologies (HTML, JavaScript, D3.js) to support research and provide interactive exploration.
Main Results Submitted for the Defense. This dissertation presents the main results of the dissertation.
1. Answer Candidate Typing and Selection (ACT Selection). A novel method that combines semantic types of LLM generated predictions with explicit type information from knowledge graphs. This approach filters and scores answer candidates based on type constraints, consistently improving Hits@1 across datasets while maintaining controllable outputs. This represents the first stable integration of types of LLM predictions with KG schema for candidate filtering and ranking.
2. Controllable Fusion Framework. A subgraph-based reranking system that extracts shortest-path knowledge graph subgraphs for each question-candidate pair and transforms them into multi-modal features including graph measures, semantic text features, and graph-to-text
representations. Multiple rankers are evaluated over these features to promote answers supported by KG structure.
3. ShortPathQA Dataset. A specialized question answering resource containing pre-computed KG subgraphs for every question-candidate pair. This dataset is the first QA resource that standardizes subgraph-based reasoning evaluation and enables reproducible comparison of reranking methods.
4. Complete end-to-end pipelines including baseline sequence-to-sequence models, ACT Selection implementation, public API endpoints, and a web-based subgraph visualization tool. These systems demonstrate technological readiness and facilitate community adoption.
5. Two empirical findings verified across multiple datasets and models: LLMs reliably predict answer of correct types even when factual answers are incorrect, providing stable control signals for KG question answering; and subgraph-based evidence effectively reduces hallucinations while maintaining transparency in language model outputs.
Validation of Research Results and Reliability. The research findings presented in this dissertation are validated through rigorous empirical evaluation on multiple standard benchmark datasets and newly created resources. The reliability of the results is ensured by:
— Standardized Datasets and Metrics: For the ACT Selection method (Chapter 2), evaluations were performed on three Wikidata-based one-hop KGQA datasets: SimpleQuestions-Wikidata (SQWD), RuBQ (English translations), and a subset of Mintaka (one-hop English questions). The primary metric was Hits@1, which directly measures the accuracy of the top-ranked answer.
— For the controllable fusion methods (Chapter 3), experiments were conducted on the Mintaka dataset (excluding count and yes/no questions) and the newly introduced ShortPathQA dataset. Evaluation metrics included Hits@N (N=1, 2, 3) to assess the reranker's ability to position the correct answer within the top N candidates, and F1-score for binary classification tasks on ShortPathQA.
— Comparison with Baselines: The performance of the proposed methods was consistently compared against various baselines.
— In ACT Selection, comparisons included the base Text-to-Text models (T5 variants, both zero-shot and fine-tuned) without ACT, specialized KGQA systems like QAnswer and KEQA, and large conversational models like ChatGPT.
— In controllable fusion, baselines included initial LLM predictions (T5-Large-SSM, T5-XL-SSM, Mistral, Mixtral), random ranking, semantic reranking, and simpler regression-based models before evaluating more complex neural rankers. For ShortPathQA, comparisons included zero-shot LLMs (GPT-4o, LLaMA3-8b-Instruct) and supervised models (MPNet, fine-tuned LLaMA3-8b-Instruct).
— Ablation Studies: Comprehensive ablation studies were conducted to assess the individual contributions of different components of the proposed frameworks (e.g., different scoring mechanisms, feature types). This helps to isolate the impact of novel elements.
— Error Analysis: Qualitative error analysis was performed (e.g., in Chapter 2 and for G2T methods in Chapter 3) to understand specific instances where the proposed methods succeed or fail, providing deeper insights beyond aggregate metrics. For example, ACT Selection demonstrated an ability to predict correct answer types in 94% of instances for T5-Large-SSM on SQWD, a significant improvement over the 61% baseline. The G2T analysis highlighted T5's superior factual accuracy over GAP in preserving entity information.
— Cross-Model and Cross-Dataset Evaluation: The proposed methods were evaluated across different LLM architectures (T5, Mistral, Mixtral) and sizes, and across multiple datasets with varying characteristics, demonstrating the robustness and generalizability of the findings. For example, ACT Selection consistently improved performance across all tested base models and datasets. The subgraph reranking methods also showed consistent Hits@1 improvements across different LLMs.
— Statistical Significance and Reproducibility: While not always explicitly stated as statistical significance tests, the consistent and considerable margins of improvement over baselines across multiple setups suggest meaningful results. The publication of code (e.g., M3M pipeline, ShortPathQA dataset and code) and detailed experimental setups (hyperparameters, model versions) supports reproducibility. For
the ShortPathQA dataset, a manual test set was curated with multiple annotators, ensuring data quality. These measures collectively ensure that the conclusions drawn are well-supported by empirical evidence and that the proposed methods offer reliable improvements in KGQA.
Approbation of the Work and Publications. The research presented in this dissertation has been disseminated through six peer-reviewed publications in international conferences and journals, and through participation in shared tasks and system demonstrations including top-tier conferences such as ACL and LREC/COLING. All this publications are indexed in Scopus. The key contributions have been presented to and validated by the scientific community on the following conferences:
1. 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, Canada, July 2023.
2. The 18th Conference on Natural Language Processing (KONVENS), Ingolstadt, Germany, September 2023.
3. The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING), Torino, Italy, May 2024.
4. The 62nd Annual Meeting of the Association for Computational Linguistics (ACL), TextGraphs-17, Bangkok, Thailand, August 2024.
5. The 30th International Conference on Natural Language & Information Systems (NLDB), Kanazawa, Japan, July 2025.
6. The 31st International Conference on Computational Linguistics (COLING), GenAIK, Abu Dhabi, United Arab Emirates, January 2025.
Author's Personal Contribution.
Author's personal contribution to the research presented in this dissertation is as follows:
— ACT Selection Method Development (Chapter 2): The author was fully responsible for the conceptualization, implementation, and experimental validation of the Answer Candidate Type Selection pipeline, including the design of the four-component scoring mechanism. This chapter based on publication called "Answer Candidate Type Selection: Text-To-Text Language Model for Closed Book Question Answering Meets Knowledge Graphs" [82].
- Subgraph Reranking Framework (Chapter 3): The author led the full development of the controllable fusion methodology and the design of the experiments. The implementation and experimental design for the Graph-To-Text (G2T) models were developed by the author's advisor with significant contributions from Hai Le, Dmitrii Iarosh, and Ivan Lazichny. Olga Tsymboi and Egor Cheremiskin implemented experiments with the Mistral and Mixtral models. This contribution includes analyzing Graph-To-Text approaches limitations [83] (Author responsible for experiments with Graph-To-Text AlignScore analysis and managment of this work).
- ShortPathQA Dataset Creation (Chapter 3.7): The author had full responsibility for the dataset design and the subgraph generation pipeline. The annotation framework and question curation for the manual portion of the dataset were handled by Andrey Sakhovskiy and Irina Nikishina. Andrey Sakhovskiy conducted experiments with encoder-only baselines.
- System Demonstration Creation (Chapter 4): The author had full responsibility for creating the system demonstrations. Dmitrii Iarosh assisted with the development of the subgraph visualization tool and related DevOps tasks.
Structure and volume of the dissertation. This dissertation is structured to systematically present these contributions across four chapters, with a logical progression from foundational concepts through individual methodologies to integrated system implementations and system demonstrations. Introduction and Chapter 1 provides a detailed overview of the background and relevance of the work. Chapter 2 introduces the Answer Candidate Type Selection method. Chapter 3 introduces the Controllable Fusion via Subgraph Reranking method. Chapter 4 introduces the system demonstrations for the ACT Selection and tools for supporting Controllable Fusion methods.
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Оптимизация технологий нефтесервисного обслуживания на основе математического моделирования и анализа данных / Optimization of Oilfield Services Technologies Based on Mathematical Modeling and Data Analysis2026 год, кандидат наук Морозов Антон Дмитриевич
Оценка неопределенности в задачах обработки естественного языка2026 год, кандидат наук Важенцев Артем Андреевич
Псевдобулевский полиномиальный подход к решению задач компьютерного зрения / Pseudo- Boolean Polynomial approach to solving Computer Vision tasks2025 год, кандидат наук Чикаке Тендай Мапунгвана
Модели и методы искусственного интеллекта в обработке медицинских изображений высокого качества / Models and Methods of Artificial Intelligence of High-Quality Medical Images Processing2025 год, кандидат наук Ал-Аззави Зобеда Хатиф Наджи
Извлечение читаемых моделей из логов событий2026 год, кандидат наук Бегичева Антонина Константиновна
Заключение диссертации по теме «Другие cпециальности», Сальников Михаил Дмитриевич
Conclusion
This dissertation addressed the central challenge formulated in the introduction: how to fuse Large Language Models (LLMs) with Knowledge Graphs (KGs) to build question answering systems that are more accurate, reliable, and controllable. The work proposed two complementary methodologies and validated them through extensive experiments, ablations, and system demonstrations. In doing so, it focused not on repeating known results, but on highlighting what is new and practically useful: type-aware control of LLM generations and evidence-based reranking over KG subgraphs.
The main results of the dissertation includes several key contributions. Answer Candidate Type (ACT) Selection represents a new method that was proposed and validated, using the LLM's ability to infer the expected semantic type of an answer and combining it with explicit type information from the KG (e.g., Wikidata P31). This type signal acts as a stable control to filter and rerank candidates. Hits@1 substantially improved across datasets and model sizes. The ablation study confirmed that the full combination of scores (type, one-hop neighbors, Text-to-Text rank, question-property similarity) yields the best performance, while type alone is not sufficient. Importantly, ACT provides a resource-efficient path to close the gap between smaller and larger LLMs by adding controllability rather than parameters.
Controllable Fusion via subgraph reranking introduced a general reranking framework that verifies LLM candidates against shortest-path subgraphs between question entities and answers in the KG. Multi-modal features were engineered (graph statistics such as PageRank, density, Katz centrality; text features; and Graph2Text sequences obtained by Deterministic linearization, T5-based, and GAP-based models). Across candidate sources (T5-Large-SSM, T5-XL-SSM, Mistral, Mixtral), reranking consistently improved quality over initial LLM rankings. Feature-importance analyses demonstrated that PageRank is a strong and stable graph signal, while text-based features dominate overall. Graph2Text T5 produced more faithful verbalizations than GAP in our setting, which explains its stronger downstream reranking performance.
To enable standardized study of controllable fusion, the dissertation introduced ShortPathQA: a dataset for subgraph-based evaluation—the first QA
resource that releases pre-computed shortest-path subgraphs linking question entities and candidate answers. The dataset contains 12,526 questions and 143,061 question-candidate pairs with subgraphs. A manually curated test set complements an automatic split and targets more complex reasoning. Baselines show that models benefit more from graphs on the manual set, indicating that richer subgraphs are particularly helpful for multi-step questions.
The work also delivered system demonstration in the form of production-style components: a baseline Seq2Seq pipeline, ACT Selection, and a web tool for subgraph visualization. All components are exposed as REST APIs (FastAPI) and demonstrate technological readiness. They make the proposed methods reproducible and usable for the community.
Finally, the research established empirical principles for controllable KG-LLM fusion. Two consistent observations guide future designs: (i) LLMs often infer the correct type even when the top-1 answer is wrong, making type a robust control signal for KGQA; (ii) explicit path evidence from KGs reduces hallucinations and improves interpretability when used to rerank candidates. Together, these principles explain why ACT and subgraph reranking are effective and complementary.
Limitations and future extensions. While the proposed methods improved reliability and control, several limitations remain. ACT relies on the availability and quality of type schema and entity linking, and subgraph reranking currently focuses on shortest paths and one-hop settings in candidate sourcing. Zero-shot prompting of LLMs with subgraphs was not always effective; better interfaces for models to consume structured evidence are needed. Finally, scaling to larger multilingual settings and domain-specific KGs requires additional engineering.
The most promising directions for further development of the research include several key areas. From shortest paths to richer subgraphs and multi-hop reasoning represents an important extension beyond single shortest paths to bounded-depth induced subgraphs and relation patterns; early results on the manual set suggest larger gains on compositional questions. Stronger model interfaces for KG evidence involves developing structured prompting, tool-use, or lightweight adapters so LLMs can reliably consume subgraphs without performance drops, complementing the demonstrated gains from supervised finetuning. Dataset growth and analysis tooling encompasses expanding ShortPathQA with more complex questions, richer annotations, and standardized Graph2Text outputs to facilitate detailed error analysis and benchmarking.
In conclusion, the dissertation provides a clear recipe for making LLM-based QA more factual and controllable: use type to constrain and use subgraph evidence to verify. Together with public systems and a new dataset, these contributions form a practical foundation for future research and applications.
Список литературы диссертационного исследования кандидат наук Сальников Михаил Дмитриевич, 2026 год
Bibliography
1. Lin, S., TruthfulQA: Measuring How Models Mimic Human Falsehoods / S. Lin, J. Hilton, O. Evans // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Vol. 1. — Association for Computational Linguistics, 2022. — P. 3214--3252.
2. Roberts, A., How Much Knowledge Can You Pack into the Parameters of a Language Model? / A. Roberts, C. Raffel, N. M. Shazeer // Conference on Empirical Methods in Natural Language Processing. — 2020. — P. 5418—5426.
3. Ji, Z., Survey of Hallucination in Natural Language Generation / Z. Ji, N. Lee, R. Frieske, [et al.] // ACM Computing Surveys. — 2022. — Vol. 55. — P. 1—38.
4. Tonmoy, S. M. T. I., A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models / S. M. T. I. Tonmoy, S. M. M. Zaman, V. Jain, [et al.] // ArXiv. — 2024. — URL: https://doi.org/10.48550/arXiv. 2401.01313; visited on 07/07/2025.
5. Pletenev, S., How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM? / S. Pletenev, M. Marina, D. Moskovskiy, [et al.] // Findings of the Association for Computational Linguistics: NAACL 2025. — Association for Computational Linguistics, 2025. — P. 4309—4322.
6. Singhal, A., Introducing the Knowledge Graph: things, not strings / A. Singhal. — 2012. — URL: https : / / blog . google / products / search / introducing -knowledge-graph-things-not/; visited on 07/07/2025.
7. Bizer, C., DBpedia - A crystallization point for the Web of Data / C. Bizer, J. Lehmann, G. Kobilarov, [et al.] //J. Web Semant. — 2009. — Vol. 7, 3. — P. 154—165.
8. Auer, S., DBpedia: A Nucleus for a Web of Open Data / S. Auer, C. Bizer, G. Kobilarov, [et al.] // The Semantic Web, 6th International Semantic Web Conference, 2nd Asian Semantic Web Conference, ISWC 2007 + ASWC 2007, Busan, Korea, November 11-15, 2007. Vol. 4825. — Springer, 2007. — P. 722—735.
9. Vrandecic, D., Wikidata: a free collaborative knowledgebase / D. Vrandecic, M. Krotzsch // Commun. ACM. — 2014. — Vol. 57, 10. — P. 78—85.
10. Gu, Y., Knowledge Base Question Answering: A Semantic Parsing Perspective / Y. Gu, V. Pahuja, G. Cheng, [et al.] // 4th Conference on Automated Knowledge Base Construction, AKBC 2022, London, UK, November 3-5, 2022. — 2022.
11. Vaswani, A., Attention is All you Need / A. Vaswani, N. Shazeer, N. Parmar, [et al.] // Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems / ed. by I. Guyon, U. von Luxburg, S. Bengio, [et al.]. — 2017. — P. 5998—6008.
12. Chung, H. W., Scaling Instruction-Finetuned Language Models / H. W. Chung, L. Hou, S. Longpre, [et al.] // ArXiv. — 2022. — URL: https://doi.org/10. 48550/arXiv.2210.11416; visited on 25/11/2024.
13. Petroni, F., Language Models as Knowledge Bases? / F. Petroni, T. Rocktäschel, S. Riedel, [et al.] // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP. — Association for Computational Linguistics, 2019. — P. 2463—2473.
14. Raffel, C., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer / C. Raffel, N. Shazeer, A. Roberts, [et al.] // J. Mach. Learn. Res. — 2020. — Vol. 21. — P. 140:1—140:67.
15. Raffel, C., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer / C. Raffel, N. Shazeer, A. Roberts, [et al.] // CoRR. — 2019. — URL: http : / / arxiv . org / abs / 1910 . 10683; visited on 17/12/2024.
16. Brown, T., Language Models are Few-Shot Learners / T. Brown, B. Mann, N. Ryder, [et al.] // Advances in Neural Information Processing Systems. Vol. 33. — Curran Associates, Inc., 2020. — P. 1877—1901.
17. Lewis, M., BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension / M. Lewis, Y. Liu, N. Goyal, [et al.] // Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL. — Association for Computational Linguistics, 2020. — P. 7871—7880.
18. Izacard, G., Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering / G. Izacard, E. Grave // Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL. — Association for Computational Linguistics, 2021. — P. 874—880.
19. Cao, S., KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base / S. Cao, J. Shi, L. Pan, [et al.] // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ACL. — Association for Computational Linguistics,
2022. — P. 6101—6119.
20. Tan, Y., Evaluation of ChatGPT as a Question Answering System for Answering Complex Questions / Y. Tan, D. Min, Y. Li, [et al.] // CoRR.
— 2023. — URL: https://doi.org/10.48550/arXiv.2303.07992; visited on 19/11/2024.
21. Mallen, A., When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories / A. Mallen, A. Asai, V. Zhong, [et al.] // Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics ACL. — Association for Computational Linguistics,
2023. — P. 9802—9822.
22. Cohen, R., LM vs LM: Detecting Factual Errors via Cross Examination / R. Cohen, M. Hamri, M. Geva, [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP. — Association for Computational Linguistics, 2023. — P. 12621—12640.
23. Min, S., FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation / S. Min, K. Krishna, X. Lyu, [et al.] // Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. — Association for Computational Linguistics, 2023. — P. 12076—12100.
24. Ouyang, L., Training language models to follow instructions with human feedback / L. Ouyang, J. Wu, X. Jiang, [et al.] // Advances in Neural Information Processing Systems. Vol. 35. — Curran Associates, Inc., 2022.
— P. 27730—27744.
25. Zhang, C., A review of deep learning in question answering over knowledge bases / C. Zhang, Y. Lai, Y. Feng, [et al.] //AI Open. — 2021. — Vol. 2. — P. 205—215.
26. Yih, W., Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge Base / W. Yih, M. Chang, X. He, [et al.] // Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing ACL. — The Association for Computer Linguistics, 2015. — P. 1321—1331.
27. Balog, K., Entity-Oriented Search. Vol. 39 / K. Balog. — Springer, 2018. — P. 1—351.
28. Bordes, A., Question Answering with Subgraph Embeddings / A. Bordes, S. Chopra, J. Weston // Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing EMNLP. — ACL, 2014. — P. 615—620.
29. Dai, Z., CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge Bases / Z. Dai, L. Li, W. Xu // Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. Vol. 1 / ed. by K. Erk, N. A. Smith. — 2016. — P. 800—810.
30. Lukovnikov, D., Neural Network-based Question Answering over Knowledge Graphs on Word and Character Level / D. Lukovnikov, A. Fischer, J. Lehmann, [et al.] // Proceedings of the 26th International Conference on World Wide Web. — ACM, 2017. — P. 1211—1220.
31. Huang, X., Knowledge Graph Embedding Based Question Answering / X. Huang, J. Zhang, D. Li, [et al.] // Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining. — ACM, 2019. — P. 105—113.
32. Wang, Q., Knowledge Graph Embedding: A Survey of Approaches and Applications / Q. Wang, Z. Mao, B. Wang, [et al.] // IEEE Trans. Knowl. Data Eng. — 2017. — Vol. 29, 12. — P. 2724--2743.
33. Devlin, J., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding / J. Devlin, M. Chang, K. Lee, [et al.] // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Vol. 1. — Association for Computational Linguistics, 2019. — P. 4171—4186.
34. Chakraborty, N., Introduction to neural network-based question answering over knowledge graphs / N. Chakraborty, D. Lukovnikov, G. Maheshwari, [et al.] // WIREs Data Mining Knowl. Discov. — 2021. — Vol. 11, 3. — P. 1389—1390.
35. Lan, Y., A Survey on Complex Knowledge Base Question Answering: Methods, Challenges and Solutions / Y. Lan, G. He, J. Jiang, [et al.] // Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence IJCAI.
— ijcai.org, 2021. — P. 4483—4491.
36. Roy, R. S., Question Answering for the Curated Web: Tasks and Methods in QA over Knowledge Bases and Text Collections / R. S. Roy, A. Anand. — Morgan & Claypool Publishers, 2021. — P. 7—172.
37. Pan, S., Unifying Large Language Models and Knowledge Graphs: A Roadmap / S. Pan, L. Luo, Y. Wang, [et al.] // IEEE Trans. Knowl. Data Eng. — 2024.
— Vol. 36, 7. — P. 3580—3599.
38. Li, Y., A Survey of Graph Meets Large Language Model: Progress and Future Directions / Y. Li, Z. Li, P. Wang, [et al.] // Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence IJCAI. — ijcai.org, 2024. — P. 8123—8131.
39. Lewis, P., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks / P. Lewis, E. Perez, A. Piktus, [et al.] // Advances in Neural Information Processing Systems. Vol. 33. — Curran Associates, Inc., 2020.
— P. 9459—9474.
40. Baek, J., Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answering / J. Baek, A. F. Aji, A. Saffari // CoRR. — 2023. — URL: https://doi.org/10.48550/arXiv.2306.04136; visited on 12/07/2024.
41. He, X., G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering / X. He, Y. Tian, Y. Sun, [et al.] // Advances in Neural Information Processing Systems. Vol. 37. — Curran Associates, Inc., 2024. — P. 132876—132907.
42. Li, M., Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation / M. Li, S. Miao, P. Li // ArXiv. — 2024. — URL: https://arxiv.org/abs/2410.20724; visited on 07/07/2025.
43. Liu, H., Knowledge Graph-Enhanced Large Language Models via Path Selection / H. Liu, S. Wang, Y. Zhu, [et al.] // Findings of the Association for Computational Linguistics, ACL. — Association for Computational Linguistics, 2024. — P. 6311—6321.
44. Wen, Y., MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models / Y. Wen, Z. Wang, J. Sun // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Vol. 1. — Association for Computational Linguistics, 2024. — P. 10370—10388.
45. Wang, R., K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters / R. Wang, D. Tang, N. Duan, [et al.] // CoRR. — 2020. — URL: https://arxiv.org/abs/2002.01808; visited on 09/04/2024.
46. Yasunaga, M., QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering / M. Yasunaga, H. Ren, A. Bosselut, [et al.] // Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT. — 2021. — P. 535--546.
47. Salnikov, M., Large Language Models Meet Knowledge Graphs to Answer Factoid Questions / M. Salnikov, H. Le, P. Rajput, [et al.] // Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation, PACLIC. — 2023. — P. 635—644.
48. Vijayakumar, A. K., Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models / A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, [et al.] // CoRR. — 2016. — URL: http://arxiv.org/abs/1610.02424; visited on 13/10/2024.
49. Page, L., The PageRank Citation Ranking: Bringing Order to the Web / tech. rep. / L. Page, S. Brin, R. Motwani, [et al.] / Stanford Digital Library Technologies Project, Stanford University. — 1998. — URL: http://ilpubs. stanford.edu:8090/422/; isited on 13/10/2024.
50. Katz, L., A new status index derived from sociometric analysis / L. Katz // Psychometrika. — 1953. — Vol. 18, 1. — P. 39—43.
51. Prokhorenkova, L. O., CatBoost: unbiased boosting with categorical features / L. O. Prokhorenkova, G. Gusev, A. Vorobev, [et al.] // Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS. — 2018. — P. 6639—6649.
52. Song, K., MPNet: Masked and Permuted Pre-training for Language Understanding / K. Song, X. Tan, T. Qin, [et al.] // Advances in Neural Information Processing Systems. Vol. 33. — Curran Associates, Inc., 2020. — P. 16857—16867.
53. Wang, Y., KC-GenRe: A Knowledge-constrained Generative Re-ranking Method Based on Large Language Models for Knowledge Graph Completion / Y. Wang, M. Hu, Z. Huang, [et al.] // Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING. — 2024. — P. 9668—9680.
54. Wei, Y., KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph Completion / Y. Wei, Q. Huang, J. T. Kwok, [et al.] // CoRR. — 2024. — URL: https://doi.org/10.48550/arXiv.2402.02389; visited on 28/05/2024.
55. Gupta, V., Retrieve and Re-rank: A Simple and Effective IR Approach to Simple Question Answering over Knowledge Graphs / V. Gupta, M. K. Chinnakotla, M. Shrivastava // Proceedings of the First Workshop on Fact Extraction and VERification, FEVER@EMNLP. — 2018. — P. 22—27.
56. Conneau, A., Unsupervised Cross-lingual Representation Learning at Scale / A. Conneau, K. Khandelwal, N. Goyal, [et al.] // Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL. — 2020. — P. 8440—8451.
57. Sen, P., Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering / P. Sen, A. F. Aji, A. Saffari // Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022. — International Committee on Computational Linguistics, 2022. — P. 1604—1619.
58. Korablinov, V., RuBQ: A Russian Dataset for Question Answering over Wikidata / V. Korablinov, P. Braslavski // The Semantic Web - ISWC 2020
- 19th International Semantic Web Conference. Vol. 12507. — 2020. — P. 97—110.
59. Cao, N. D., Multilingual Autoregressive Entity Linking / N. D. Cao, L. Wu, K. Popat, [et al.] // Trans. Assoc. Comput. Linguistics. — 2022. — Vol. 10.
— P. 274—290.
60. Bordes, A., Large-scale Simple Question Answering with Memory Networks / A. Bordes, N. Usunier, S. Chopra, [et al.] // CoRR. — 2015. — URL: http: //arxiv.org/abs/1506.02075; visited on 27/01/2024.
61. Diefenbach, D., Question Answering Benchmarks for Wikidata / D. Diefenbach, T. P. Tanon, K. D. Singh, [et al.] // Proceedings of the ISWC 2017 Posters & Demonstrations and Industry Tracks co-located with 16th International Semantic Web Conference (ISWC 2017). Vol. 1963. — CEUR-WS.org, 2017. — P. 1—5.
62. Talmor, A., The Web as a Knowledge-Base for Answering Complex Questions / A. Talmor, J. Berant // Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT. Vol. 1. — 2018. — P. 641—651.
63. Rybin, I., RuBQ 2.0: An Innovated Russian Question Answering Dataset / I. Rybin, V. Korablinov, P. Efimov, [et al.] // The Semantic Web - 18th International Conference, ESWC. Vol. 12731. — 2021. — P. 532--547.
64. Reimers, N., Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks / N. Reimers, I. Gurevych // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP. — 2019. — P. 3980—3990.
65. Biemann, C., Text: now in 2D! A framework for lexical expansion with contextual similarity / C. Biemann, M. Riedl //J. Lang. Model. — 2013. --Vol. 1, 1. — P. 55—95.
66. Vajjala, S., What do we really know about State of the Art NER? / S. Vajjala, R. Balasubramaniam // Proceedings of the Thirteenth Language Resources and Evaluation Conference, LREC. -- 2022. -- P. 5983--5993.
67. Bollacker, K. D., Freebase: a collaboratively created graph database for structuring human knowledge / K. D. Bollacker, C. Evans, P. K. Paritosh, [et al.] // Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD. -- 2008. — P. 1247--1250.
68. Diefenbach, D., Towards a question answering system over the Semantic Web / D. Diefenbach, A. Both, K. Singh, [et al.] // Semantic Web. — 2020. -Vol. 11, 3. -- P. 421—439.
69. Lerer, A., Pytorch-BigGraph: A Large Scale Graph Embedding System / A. Lerer, L. Wu, J. Shen, [et al.] // Proceedings of the Second Conference on Machine Learning and Systems, SysML. Vol. 1. — 2019. — P. 120—131.
70. Lysyuk, M., Konstruktor: A Strong Baseline for Simple Knowledge Graph Question Answering / M. Lysyuk, M. Salnikov, P. Braslavski, [et al.] // Natural Language Processing and Information Systems - 29th International Conference on Applications of Natural Language to Information Systems, NLDB. Vol. 14763. -- 2024. — P. 107—118.
71. Wu, Y., Retrieve-Rewrite-Answer: A KG-to-Text Enhanced LLMs Framework for Knowledge Graph Question Answering / Y. Wu, N. Hu, S. Bi, [et al.] // CoRR. — 2023. — URL: https://doi.org/10.48550/arXiv.2309.11206; visited on 03/12/2024.
72. Shimorina, A., Handling Rare Items in Data-to-Text Generation / A. Shimorina, C. Gardent // Proceedings of the 11th International Conference on Natural Language Generation. — 2018. — P. 360—370.
73. Ribeiro, L. F. R., Investigating Pretrained Language Models for Graph-to-Text Generation / L. F. R. Ribeiro, M. Schmitt, H. Schütze, [et al.] // CoRR. — 2020. — Vol. abs/2007.08426. — URL: https://arxiv.org/abs/2007.08426; visited on 31/05/2024.
74. Colas, A. M., GAP: A Graph-aware Language Model Framework for Knowledge Graph-to-Text Generation / A. M. Colas, M. Alvandipour, D. Z. Wang // Proceedings of the 29th International Conference on Computational Linguistics, COLING. — 2022. — P. 5755—5769.
75. Zha, Y., AlignScore: Evaluating Factual Consistency with A Unified Alignment Function / Y. Zha, Y. Yang, R. Li, [et al.] // Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Vol. 1. — Association for Computational Linguistics, 2023. — P. 11328—11348.
76. Figueroa, J. H., Measuring semantic similarity of documents with weighted cosine and fuzzy logic / J. H. Figueroa, F. Perez-Tellez, D. Pinto //J. Intell. Fuzzy Syst. — 2020. — Vol. 39, 2. — P. 2263--2278.
77. Henderson, M. L., Efficient Natural Language Response Suggestion for Smart Reply / M. L. Henderson, R. Al-Rfou, B. Strope, [et al.] // CoRR. — 2017. — URL: http://arxiv.org/abs/1705.00652; visited on 06/06/2025.
78. Loshchilov, I., Decoupled Weight Decay Regularization / I. Loshchilov, F. Hutter // 7th International Conference on Learning Representations, ICLR. — OpenReview.net, 2019. — URL: https : / / openreview . net / forum ? id = Bkg6RiCqY7; visited on 07/07/2025.
79. Dubey, A., The Llama 3 Herd of Models / A. Dubey, A. Jauhri, A. Pandey, [et al.] // CoRR. — 2024. — URL: https://doi.org/10.48550/arXiv.2407.21783; visited on 06/06/2025.
80. Kingma, D. P., Adam: A Method for Stochastic Optimization / D. P. Kingma, J. Ba // 3rd International Conference on Learning Representations, ICLR. — 2015. — URL: http://arxiv.org/abs/1412.6980; visited on 06/06/2025.
Author's publications on the dissertation subject
81. Shallouf, A., CAM 2.0: End-to-End Open Domain Comparative Question Answering System / A. Shallouf, H. Herasimchyk, M. Salnikov, R. A. G. Veliz, N. Mestvirishvili, A. Panchenko, C. Biemann, I. Nikishina // 2024 Joint International Conference on Computational Linguistics, Language Resources
and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. — 2024. — P. 2657—2672.
82. Salnikov, M., Answer Candidate Type Selection: Text-To-Text Language Model for Closed Book Question Answering Meets Knowledge Graphs / M. Salnikov, M. Lysyuk, P. Braslavski, A. Razzhigaev, V. Malykh, A. Panchenko // 19th Conference on Natural Language Processing, KONVENS 2023 - Proceedings of the Conference. — 2023. — P. 155—164.
83. Iarosh, D., On Reducing Factual Hallucinations in Graph-to-Text Generation Using Large Language Models / D. Iarosh, A. Panchenko, M. Salnikov // Proceedings - International Conference on Computational Linguistics, COLING. — 2025. — P. 43—53.
84. Razzhigaev, A., A System for Answering Simple Questions in Multiple Languages / A. Razzhigaev, M. Salnikov, V. Malykh, P. Braslavski, A. Panchenko // Proceedings of the Annual Meeting of the Association for Computational Linguistics. — 2023. — P. 524—537.
85. Sakhovskiy, A., TextGraphs 2024 Shared Task on Text-Graph Representations for Knowledge Graph Question Answering / A. Sakhovskiy, M. Salnikov, I. Nikishina, A. Usmanova, A. Kraft, C. Moller, D. Banerjee, J. Huang, L. Jiang, R. Abdullah, X. Yan, D. Ustalov, E. Tutubalina, R. Usbeck, A. Panchenko // TextGraphs at ACL 2024 - Proceedings of TextGraphs-17: Graph-Based Methods for Natural Language Processing, 62nd Annual Meeting of the Association of Computational Linguistics. — 2024. — P. 116—125.
86. Salnikov, M., ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs / M. Salnikov, A. Sakhovskiy, I. Nikishina, A. Usmanova, A. Kraft, C. Moller, D. Banerjee, J. Huang, L. Jiang, R. Abdullah, X. Yan, E. Tutubalina, R. Usbeck, A. Panchenko // Lecture Notes in Computer Science. Vol. 15836. — 2025. — P. 95—110.
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.