Нейро-символьный метод рассуждений с использованием абстракций для вопросно-ответных систем тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Радюш Даниил Валентинович
- Специальность ВАК РФ00.00.00
- Количество страниц 336
Оглавление диссертации кандидат наук Радюш Даниил Валентинович
Реферат
Synopsis
Введение
Глава 1. Общая характеристика нейро-символьных систем
1.1 Предпосылки создания нейро-символьных систем
1.2 Классификация нейро-символьных систем
1.3 Применение нейро-символьных систем
Выводы по главе
Глава 2. Методы для построения нейро-символьных систем
2.1 Символьные методы
2.2 Языковые модели
2.3 Обучение с подкреплением
2.4 Мультиагентные системы
Выводы по главе
Глава 3. Нейро-символьные вопросно-ответные системы
3.1 Общая характеристика вопросно-ответных систем
3.2 Подходы к построению нейро-символьных вопросно-ответных систем
Выводы по главе
Глава 4. Нейро-символьный метод рассуждений
4.1 Постановка проблемы
4.2 Архитектура мультиагентной нейро-символьной вопросно-ответной системы
4.3 Анализ тестового бенчмарка
4.4 Характеристика имплементации и результаты экспериментов
Выводы по главе
Заключение
Список сокращений и условных обозначений
Словарь терминов
Список литературы
Приложение A
Приложение B
Тексты публикаций
Реферат
Общая характеристика диссертации
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Метод генерации ответно-вопросных реплик с учетом особенностей письменной речи пользователя2025 год, кандидат наук Матвеева Анастасия Андреевна
Нейросетевые методы работы с базами знаний для ответа на вопросы, ведения диалога и обработки текста2023 год, кандидат наук Евсеев Дмитрий Андреевич
Семантический анализатор русскоязычного текста для вопросно-ответной системы2017 год, кандидат наук Мочалова Анастасия Викторовна
Сетевые модели оперативного управления процессом принятия решений в САПР2001 год, кандидат технических наук Семенов, Владимир Геннадьевич
Методы разноуровневого анализа текстов на естественном языке и их приложения в системах информационного поиска и психолингвистических исследованиях2024 год, доктор наук Смирнов Иван Валентинович
Введение диссертации (часть автореферата) на тему «Нейро-символьный метод рассуждений с использованием абстракций для вопросно-ответных систем»
Актуальность темы.
В настоящее время большинство приложений, связанных с понятием искусственного интеллекта, так или иначе используют в своей работе большие языковые модели, которые успели зарекомендовать себя в качестве универсального подхода к решению повседневных, профессиональных и даже творческих задач. Такой подход действительно имеет множество преимуществ, предлагая унифицированный формат взаимодействия систем искусственного интеллекта с человеком посредством текста или устной речи.
В то же время существуют классы задач, при решении которых становится явным отсутствие понимания большими языковыми моделями сути выполняемых действий, что может контрастировать с их восприятием человеком. В связи с этим возникают обоснованные сомнения относительно того, может ли текущая парадигма глубокого обучения, основанная на нейронных сетях, привести к созданию систем, по-настоящему способных адаптироваться к новым задачам.
При этом проблема ограниченности отдельных подходов в области искусственного интеллекта актуализировалась еще в конце XX века, до появления идеи глубокого обучения, что привело к зарождению концепции гибридных систем искусственного интеллекта. Данная концепция основывается на идее сочетать при создании систем искусственного интеллекта несколько принципиально различных подходов как способа преодолеть их ограничения, объединив преимущества и минимизировав недостатки.
В качестве одного из возможных подходов для такой интеграции может рассматриваться нейро-символьная концепция, ориентированная на объединении формализма символьного подхода и универсальных вычислительных способностей нейронных сетей с возможным использованием и других подходов из области искусственного интеллекта, таких как обучение с подкреплением. Помимо этого, в качестве дополняющей идеи может служить рассмотрение систем искусственного интеллекта как сочетания отдельных гетерогенных агентов, что
также может позволить эффективно распределять задачи между ними, в полной мере используя их сильные стороны в зависимости от контекста.
С практической точки зрения разработка нейро-символьных систем искусственного интеллекта ориентирована на получение решений, менее зависимых от количественных и качественных характеристик выборки, лучше адаптирующихся к новым условиям, чья работа может быть хотя бы отчасти логически обоснована и иметь понятную человеку интерпретацию. В первую очередь вышеприведенные свойства могут быть крайне полезны для применения в сферах человеческой жизнедеятельности, где высока цена единичной ошибки и нередко требуется формально строгое обоснование принятым решениям, таких как создание роботизированных систем, медицина и финансы.
Одним из наиболее распространенных типов задач, для которых используются нейро-символьные системы, является построение вопросно-ответных систем. Это связано с тем, что, с одной стороны, вопросно-ответные системы являются широко распространенным типом приложений из области искусственного интеллекта, тесно связанным по сути своей работы с множеством других типов, таких как экспертные системы и чат-боты. С другой стороны, на примере вопросно-ответных систем достаточно явным образом можно показать преимущество нейро-символьного подхода через оценку способности адаптироваться к новым типам запросов, а также предоставлять обоснование полученным ответам. В то же время, несмотря на текущие темпы развития решений, основанных на обработке естественного языка, в этой области все еще остаются актуальными проблемы при решении задач, требующих использование неявных знаний, строгого логического формализма и работы с абстракциями.
Степень разработанности темы.
С активным развитием приложений на основе нейронных сетей в 90-е годы XX века, во многом ставшим возможным благодаря разработке алгоритма обратного распространения ошибки Дэвидом Румельхартом, Джеффри Хинтоном и Рональдом Уильямсом, начало формироваться направление, изучающее совмещение преимуществ нейронных сетей и символьных методов, чья
необходимость при построении систем искусственного интеллекта на теоретическом уровне обосновывается в рамках гипотезы Ньюэлла-Саймона и работ Джерри Фодора. Значительную роль в формировании концепции нейро-символьного подхода сыграли многочисленные исследования Артура Гарцеса, Паскаля Хицлера, Луиса Ламба и Лучиано Серафини, классификация нейро-символьных систем Генри Каутца, а также идеи Даниэля Канемана, в соответствии с которыми символьная и субсимвольная составляющие могут соотноситься с двумя системами принятия решений человека. Интерес к нейро-символьной концепции проявляется и в русскоязычных источниках, например, в работах Александра Смирнова, Николая Шилова, Андрея Пономарева и Александра Демидовского, где уделяется внимание ее применению в контексте систем поддержки принятия решений.
На волне успеха системы IBM Watson, победившей в телевикторине «Jeopardy!», в качестве одного из наиболее распространенных приложений для нейро-символьных методов стали вопросно-ответные системы. В частности, одно из направлений сосредоточилось на аспекте использования структурированных знаний для улучшения качества ответов, вариации чего рассматриваются в статьях Мичихиро Ясунаги и Антуана Босслю. В то же время реализации Neuro-Symbolic Concept Learner (NS-CL) и Neuro-Symbolic Question Answering (NSQA) показали возможность организовывать в системе механизм формального вывода, что также обеспечивает полноценную интерпретируемость работы. Важную роль нейро-символьные вопросно-ответные системы сыграли в контексте бенчмарков, требующих рассуждений, таких как CommonsenseQA (Commonsense Question Answering) и CLEVR (Visual Question Answering), однако с развитием возможностей больших языковых моделей исследовательский интерес стали представлять в первую очередь ситуации, осложненные необходимостью устанавливать основанные на абстракциях закономерности без явной опоры на знания из уже существующей базы. Этой проблеме в значительной степени посвящена научная работа Франсуа Шолле, одним из результатов которой стала
разработка бенчмарка ARC-AGI, актуализировавшего исследования, связанные организацией процесса решения абстрактных задач.
Целью диссертационной работы является увеличение точности вопросно-ответных систем на задачах, требующих оперирования абстракциями, за счет применения нейро-символьного подхода.
Для достижения данной цели в рамках диссертации были поставлены и решены следующие задачи:
Задача 1: Провести обзор и систематизировать существующие на данный момент нейро-символьные подходы к построению систем искусственного интеллекта;
Задача 2: Определить актуальную область для приложения, позволяющую акцентировать преимущества нейро-символьного подхода;
Задача 3: Предложить архитектуру нейро-символьной вопросно-ответной системы, учитывающей особенности выбранной проблемы для решения;
Задача 4: Выполнить разработку отдельных модулей системы и реализовать механизмы их взаимодействия;
Задача 5: Осуществить тестирование системы на практике путем проведения вычислительных экспериментов и проанализировать полученные результаты.
Методы исследования.
Теоретическая часть работы преимущественно опирается на методы анализа, синтеза, классификации и обобщения.
Практическая часть включает эксперименты в форме моделирования с преобладанием методов глубокого обучения, обработки естественного языка и обучения с подкреплением. Ее реализация осуществлена с помощью библиотек и фреймворков языка программирования Python.
Основные положения, выносимые на защиту:
1. Архитектура мультиагентной вопросно-ответной системы, основанной на нейро-символьных рассуждениях, использующих символьное, субсимвольное и интерактивное представления задач, обеспечивающей более высокую точность решений, требующих оперирования абстракциями;
2. Нейро-символьный метод рассуждений, обеспечивающий решение задач посредством формирования знаний с учетом взаимодействия агентных модулей, специализированных по типам представления задач.
Научная новизна диссертации отражена в следующих пунктах:
Научная новизна 1 — разработана архитектура вопросно-ответной системы, обеспечивающей более высокую точность решения требующих оперирования абстракциями задач за счет мультиагентного нейро-символьного подхода, отличающаяся возможностью видоизменять состав, структуру и конфигурацию агентов с учетом рассмотрения символьного, субсимвольного и интерактивного представлений задачи.
Научная новизна 2 — предложен нейро-символьный метод рассуждений в вопросно-ответной системе, обеспечивающий решение задач посредством уменьшения пространства поиска, отличающийся способностью формировать знания о задаче через анализ её представлений и взаимодействия агентных модулей без использования априорных закономерностей о её специфике.
Научно-техническая задача, решаемая в диссертации, заключается в создании программных средств, обеспечивающих расширение типов запросов к вопросно-ответным системам для решения задач, требующих оперирования абстракциями, посредством рассуждений.
Объектом исследования являются мультиагентные нейро-символьные вопросно-ответные системы.
Предметом исследования являются методы нейро-символьных рассуждений для задач, требующих оперирования абстракциями.
Теоретическая значимость результатов диссертационной работы состоит в создании методологии интеграции мультиагентного и нейро-символьного подходов, которая дополняет и расширяет теорию построения гибридных систем искусственного интеллекта, а также методов моделирования рассуждений, акцентируя необходимость их дальнейшего исследования.
Практическая значимость результатов диссертационной работы определяется возможностью дальнейшего совершенствования и адаптации
разработанных подходов для задач с использованием абстракций при малом числе демонстраций, актуальных для робототехники, самоуправляемых автомобилей и медицины.
Достоверность полученных результатов обуславливается на методологическом уровне опорой на известные в научном сообществе и апробированные на конкретных задачах способы построения мультиагетнных и нейро-символьных систем, а также соответствием проведенных в рамках работы экспериментов принципам воспроизводимости и верифицируемости. Для тестирования предложенного подхода использовался общедоступный бенчмарк ARC-AGI, что обеспечивает сопоставимость результатов относительно уже существующих решений.
Апробация результатов работы. Основные результаты работы докладывались и обсуждались на следующих конференциях:
1. LI научная и учебно-методическая конференция Университета ИТМО (2 февраля - 5 февраля 2022 года)
2. XI Конгресс молодых ученых (4 апреля - 8 апреля 2022 года)
3. LII научная и учебно-методическая конференция Университета ИТМО (31 января - 3 февраля 2023 года)
4. XII Конгресс молодых ученых (3 апреля - 6 апреля 2023 года)
5. 19th International Conference on Semantic Systems (SEMANTiCS 2023) (20 сентября - 22 сентября 2023 года)
6. 22nd International Semantic Web Conference (ISWC 2023) (6 ноября - 10 ноября 2023)
7. LIII научная и учебно-методическая конференция Университета ИТМО (31 января - 3 февраля 2024 года)
8. XII Конгресс молодых ученых (3 апреля - 6 апреля 2024 года)
9. LIV научная и учебно-методическая конференция Университета ИТМО (27 января - 31 января 2025 года)
Личный вклад автора.
Все этапы получения изложенных в диссертации результатов были выполнены лично автором: постановка задач исследования, проведение обзорно-аналитической работы [150-151], разработка архитектуры и отдельных модулей нейро-символьной мультиагентной вопросно-ответной системы, а также реализация экспериментов на ее основе, анализ полученных результатов и формулировка возможных направлений для дальнейших исследований.
В совместных публикациях вклад автора заключался в разработке методов извлечения релевантных задаче структурированных знаний из графов знаний и интеграции их в работу систем искусственного интеллекта, основанных на больших языковых моделях [141, 187] и обучении с подкреплением [188], а также в создании и использовании принципиально новых вопросно-ответных бенчмарков для анализа ограничений актуальных реализаций [141-142].
Соавторы совместных публикаций внесли следующий вклад. Ясер М. — участие в разработке бенчмарка, проведение экспериментов и подготовка текста публикации [142]; Каррас О. — участие в разработке бенчмарка, подготовка текста публикации [142]; Цалапати Э. — классификация запросов, анализ результатов и подготовка текста публикации [142]; Барц К. — участие в разработке бенчмарка [142]; Кортес Э. — участие в разработке бенчмарка [142]; Шилин И. — анализ результатов [142]; Ауэр З. — общее руководство и вычитка публикации [142]; Бароне Д. — общее руководство и вычитка публикации [142]; Стокер М. — общее руководство и вычитка публикации [142]; Кубаракис М. — общее руководство и вычитка публикации [142]; Варденга Р. — разработка и имплементация системы, проведение экспериментов, подготовка текста публикации [188]; Смоляков И. — проведение экспериментов [188]; Письмеров А. — проведение экспериментов [188]; Сюэ Ю. — подготовка текста публикации [188]; Мюллер Х. — подготовка текста публикации [188]; Куденко Д. — общее руководство и вычитка публикации [188]; Тойхер Р. — разработка и имплементация системы, проведение экспериментов, подготовка текста публикации [141]; Ковригина Л. — разработка и имплементация систем [141, 188], проведение экспериментов [141], общее руководство и подготовка текста публикаций [141, 187, 188]; Плюхин Д. —
имплементация систем [187-188], проведение экспериментов и подготовка текста публикации [187], участие в разработке бенчмарка и анализ результатов [142]; Муромцев Д. — общее руководство, анализ результатов и вычитка публикаций [141-142, 187-188].
Структура и объем диссертации.
Диссертация состоит из введения, четырех глав, заключения, списка сокращений и условных обозначений, словаря терминов, списка литературы и двух приложений. Полный объем диссертации составляет 332 страницы. Список литературы содержит 212 наименований.
Во введении обосновывается актуальность исследований по теме диссертационной работы, определяется цель и соответствующие задачи как способ её достижения, оценивается научная новизна, теоретическая и практическая значимость представленной диссертации, а также формулируются положения, выносимые на защиту.
Первая глава диссертации вводит в рассмотрение концепцию нейро-символьных систем в области искусственного интеллекта. Для этого, в частности, показывается, что возникновение данной концепции непосредственно вытекает из эволюции исследовательских взглядов на проблему построения универсальных систем искусственного интеллекта в качестве способа преодоления фундаментальных недостатков отдельных подходов. В рамках более детального рассмотрения разбираются акцентируемые при разработке нейро-символьных систем свойства, приводится с соответствующими примерами одна из наиболее распространенных классификаций подобных систем, а также в общем виде анализируется их структура. В заключении приводятся авторские рассуждения, основанные на текущем положении вещей, относительно практической целесообразности применения нейро-символьных систем и конкретных сфер жизнедеятельности, в рамках которых они могут оказаться наиболее актуальными.
В разделе 1.1 анализируется смена доминирующих парадигм построения систем искусственного интеллекта с течением времени, обусловленная невозможностью с их помощью на практике создать достаточно универсальные
приложения, обладающие рядом желательных свойств. В первые десятилетия развития научной области искусственного интеллекта, зародившейся в середине XX века, преобладали логико-ориентированные символьные подходы, позволявшие за счет тщательной проработки базы знаний и правил вывода получать примечательные для того времени результаты при весьма ограниченных вычислительных ресурсах, примерами чему могут служить системы Logic Theorist [3] и SHRDLU [5]. С практической точки зрения создание приложений в то время в значительной степени опиралось на языки программирования, такие как IPL [2] и Lisp [4], что в том числе поспособствовало формированию нового направления — логического программирования — и распространению ряда подходов, являющихся актуальными по сей день. В то же время были установлены и фундаментальные ограничения для основанных исключительно на символьном подходе систем: их работа в существенной степени зависит от уже зафиксированных знаний и правил вывода, тогда как их расширение на постоянной основе может требовать существенных ресурсов. В результате, акцент применения символьных методов сместился с создания универсальных систем искусственного интеллекта на приложения, требующие работы со специализированными и структурированными знаниями, такие как экспертные системы. Одновременно с этим благодаря росту вычислительных возможностей и таким важным теоретическим достижениям, как метод обратного распространения ошибки [14], на первый план стали выходить субсимвольные подходы — статистические модели и нейронные сети. Последующее техническое и алгоритмическое развитие практики использования нейронных сетей позволило в более полной мере использовать их разносторонние возможности, обоснованные на теоретическом уровне теоремой об универсальной аппроксимации [13]. В частности, рост количества слоев и весов у нейронных сетей в рамках так называемого глубокого обучения обеспечил появление моделей, успешно и унифицированно решающих сразу ряд задач в определенной модальности, а впоследствии и в нескольких модальностях одновременно. Тем не менее ряд проблем, связанных с использованием нейронных сетей, остается актуальным, так как обусловлен самой их природой. В первую очередь, будучи
моделью типа black box, они не способны обеспечить логическую обоснованность и интерпретируемость вывода, которые могут быть крайне важны для ряда приложений [16], а их способность к обобщению в существенной мере зависит от характеристик обучающей выборки и особенностей процесса обучения. В этом контексте в качестве альтернативы может рассматриваться использование нейро-символьных систем, призванных объединить логическую обоснованность и интерпретируемость символьных методов с универсальными вычислительными возможностями нейронных сетей.
В разделе 1.2 анализируются конкретные свойства [19], на которые нацелена разработка нейро-символьных систем, такие как: обобщаемость на новые данные, независимость от области знаний, меньшие требования к величине обучающей выборки, логическая обоснованность и интерпретируемость. Таким образом, представляется возможным, что успешно реализованная нейро-символьная система будет эффективнее адаптироваться к новым задачам, требуя меньше обучающих примеров и избегая в явном виде характерного для глубоких нейронных сетей эффекта переобучения, а результаты ее работы будут хотя бы в определенной степени логически обоснованными и интерпретируемыми. Традиционно в нейро-символьных системах выделяют символьную и субсимвольные составляющие, и в зависимости от характера их сочетания принято выделять 5 классов [22]:
1. Symbolic-Neuro-Symbolic;
2. Symbolic[Neuro];
3. Neuro и compile[Symbolic];
4. Neuro ^ Symbolic;
5. Neuro[Symbolic.
В контексте первого класса нейронные сети осуществляют преобразование символов в символы (языковые модели), во втором классе они решают отдельные задачи в рамках символьной системы (AlphaGo [23]), в третьем классе символьная и субсимвольная составляющие на равных участвуют в выводе ответа через разделение задач (Neuro-symbolic Concept Learner [24]), в четвертом классе
числовые признаки используются для улучшения символьного вывода ([25]), а в пятом классе логические операции в системе организованы посредством нейронных сетей (Logic Tensor Networks [26] и Neural Tensor Networks [27]). С научной точки зрения, использование двух данных компонент также принято рассматривать с позиции теории двух систем Канемана [20]: первая система ориентирована на относительно простые и повседневные задачи и не подразумевает осознанного контроля (субсимвольная подсистема — например, предобученная глубокая нейронная сеть), тогда как вторая отвечает уже за требующие логической обоснованности комплексные задачи (символьная подсистема — например, совокупность правил вывода). Если рассматривать нейро-символьные системы с более общей позиции гибридных систем, то в дополнение могут использоваться и другие подходы из области искусственного интеллекта, такие как обучение с подкреплением, которое может позволить упростить поиск оптимального решения через интерактивное представление задачи, а также расширить механизм взаимодействия символьной и субсимвольной компонент. Кроме того, в качестве комплементарной идеи может рассматриваться и концепция мультиагентных систем, чье функционирование также основано на идее разделения задачи на подзадачи, способные решаться посредством специально разработанных для этого моделей.
Раздел 1.3 посвящен практической стороне использования нейро-символьных систем. Так как такого рода системы в существенной степени опираются на многие из существующих достижений в области искусственного интеллекта, то и сфера их применения существенно не отличается. Однако в то же время нейро-символьные системы могут и превосходить аналоги по метрикам за счет более эффективного использования структурированных знаний и механизма логического вывода при сопоставимых или даже меньших затратах на работу [33]. Не менее важными с точки зрения приложений являются логическая обоснованность и интерпретируемость функционирования подобных систем, что позволяет снизить вероятность успешности злонамеренного влияния [32] и получения бессмысленных или даже заведомо ложных результатов (галлюцинации
[87]), а также непосредственно предоставлять пользователю сопутствующую информацию, на основании которой был получен соответствующий результат. Исходя из рассмотренных качеств, можно определить, что наиболее актуальны нейро-символьные системы могут быть в сферах, где особенно высока цена ошибки, и нужно принимать строго обоснованные решения с учетом всех рисков: медицина, автономная робототехника и финансы.
Цель второй главы работы заключается в формировании методологической основы для всей работы путем более детального рассмотрения отдельных концепций построения систем искусственного интеллекта с акцентом на тех идеях, которые актуальны с позиции разработки нейро-символьных систем. Для достижения данной цели в рамках соответствующих разделов анализируются символьные методы, языковые модели как наиболее распространенное к данному моменту воплощение субсимвольной составляющей, концепция обучения с подкреплением, а также мультиагентные системы в качестве потенциально дополняющего нейро-символьную концепцию подхода.
В разделе 2.1 формулируются основные положения символьного подхода в области искусственного интеллекта. Для этого в первую очередь рассматриваются основные формы представления знаний: базы знаний, онтологии, фреймы, семантические сети и графы знаний. Благодаря данным формам представляется возможным обеспечить системы искусственного интеллекта априорными знаниями в подходящем для эффективной обработки виде и облегчить логический вывод на их основе. В частности, в последнее десятилетие значительное распространение на практике получили графы знаний, чье использование обеспечивается существованием специальных форматов данных (стандарт RDF [42], форматы Turtle [43], JSON-LD [44] и RDF/XML [42]) и языка запросов SPARQL, позволяющего унифицировать различного рода взаимодействия с графами. Помимо этого, в параграфе рассматриваются основные способы интеграции символьных признаков в работу систем искусственного интеллекта. В связи с этим затрагивается вопрос получения векторных представлений графов знаний, для чего созданы различные типы моделей (RESCAL [51], DistMult [52],
ComplEx [53], TransE [54], RotatE [55], ConvE [56]), отличающиеся с точки зрения используемых предпосылок, скоринговых функций и тех отношений в графе, которые теоретически могут учитывать (например, симметричность, инверсия, композиция). Обучение такого рода моделей на практике подразумевает использование основанных на рангах предсказаний метрик (mean rank, mean reciprocal rank, Hits@K), а также специального вида функции ошибки:
L = £„.,ДУ+fh r о - f(h'> r 01, (1)
где у — отступ, G — граф, h и t — сущности из графа, r — отношения из графа, h и t' — измененные сущности в триплете, f (h,r,t) — функция, оценивающая вероятность существования триплета в графе.
Полученное векторное представление графа может использоваться в работе системы как посредством классических архитектур нейронных сетей, так и с помощью лучше учитывающих структурированность информации графовых нейронных сетей, в рамках которых эмбеддинг отдельно взятой вершины зависит от эмбеддингов соседних вершин, что реализуется через механизм передачи сообщений:
hu =Ф(хи, ®¥(xu,xv)) , (2)
где hu — обновленное представление вершины и, у и ф — дифференцируемые функции message и update, Ф — инвариантный к перестановкам непараметрический оператор агрегирования, x и x — значения признаков для вершин и и v.
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Разработка алгоритмов оценивания характеристик диалоговой системы на основе применения нечеткого вывода с нейросетевой настройкой2023 год, кандидат наук Игитян Елена Владимировна
Методы сквозного обучения моделей семантического поиска и генерации текста в интеллектуальных диалоговых системах с доступом к неструктурированной базе знаний2025 год, кандидат наук Маслюхин Сергей Михайлович
Методы принятия решений и управления в неструктурированных задачах на основе самоорганизующихся мультиагентных рекурсивных когнитивных архитектур2014 год, кандидат наук Нагоев, Залимхан Вячеславович
Метод информированного исследования на основе графов знаний и механизма рассуждений в обучении с подкреплением2025 год, кандидат наук Письмеров Алексей Максимович
Список литературы диссертационного исследования кандидат наук Радюш Даниил Валентинович, 2025 год
Литература
1. Hamilton К., Nayak A., Bozic В. amd Longo L. Is Neuro-Symbolic A1 Meeting its Promise in Natural Language Processing? A Structured Review// arXiv preprint arXiv:2202.12205vJ - 2022., https://doi.org/10.48550/arXiv.2202.I2205.
2. Weber L., Minervini P., Munchmeyer J., Leser U., Rocktaschel Т., NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language // arXiv preprint arXiv: 1906.06187. 2019.
3. Manhaeve R., Dumancic S., Kimmig A., et al. DeepProbLog: neural probabilistic logic programming // Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS). 2018. P. 3753-3763.
4. Dong H., Mao J., Lin Т., Wang C., Li L., Zhou D. Neural logic machines // arXiv preprint arXiv: 1904.11694. 2019.
5. Riegel R., Gray A., Luus F., et al. Logical Neural Networks // arXiv preprint arXiv:2006.13155. 2020.
6. Johnson J., Fei-Fci Li, Hariharan В., Zitnick C. L., et al. CLEVR: A Diagnostic Dataset for Compositional Language and FJementary Visual Reasoning // arXiv preprint arXiv: 1612.06890. 2016.
7. Mao J., Can C., KohLi P., et al. The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision // Proceedings of the 7th International Conference on Learning Representations (ICLR). 2019.
8. Kapanipathi P., AbdeLaziz I., Ravishankar S., et al. Leveraging Abstract Meaning Representation for Knowledge Base Question Answering // Findings of the Association for Computational Linguistics: ACL-IJCNLP. 2021. P. 3884-3894.
scientific reports
H) Check for updates
open The SciQA Scientific Question
Answering Benchmark for Scholarly Knowledge
Soren Auer1'2, Dante A C. Barone3, Cassiano Bartz3, Eduardo G. Cortes3,
Mohamad YaserJaradeh1'2, Oliver Karras1LJ, Manolis Koubarakis4, Dmitry Mouromtsev5,
Dmitrii Pliukhin5, Daniil Radyushs, IvanShilin5, Markus Stocker1'2 & Eleni Tsalapati4
Knowledge graphs have gained increasing popularity in the last decade in science and technology. However, knowledge graphs are currently relatively simple to moderate semantic structures that are mainly a collection of factual statements. Question answering (QA) benchmarks and systems were so far mainly geared towards encyclopedic knowledge graphs such as DBpedia and Wikidata. We present SciQA a scientific QA benchmark for scholarly knowledge. The benchmark leverages the Open Research Knowledge Graph (ORKG) which includes almost 170,000 resources describing research contributions of almost 15,000 scholarly articles from 709 research fields. Following a bottom-up methodology, we first manually developed a set of 100 complex questions that can be answered using this knowledge graph. Furthermore, we devised eight question templates with which we automatically generated further 2465 questions, that can also be answered with the ORKG. The questions cover a range of research fields and question types and are translated into corresponding SPARQL queries over the ORKG. Based on two preliminary evaluations, we show that the resulting SciQA benchmark represents a challenging task for next-generation QA systems. This task is part of the open competitions at the 22nd International Semantic Web Conference 2023 as the Scholarly Question Answering over Linked Data (OALD) Challenge.
Knowledge graphs have gained increasing popularity in the last decade in science and technology. They enable a versatile and evolving semantic representation of knowledge at the crossroads of various
(a) levels of information structuring: unstructured, semi-structured, structured;
(b) levels of abstraction: conceptual vs. operational;
(c) knowledge representation formalisms: graphs, facts, entity-relationship, logic; and
(d) technology ecosystems.
However, most publicly available knowledge graphs, such as DBpedia or Wikidata, are relatively simple to moderate semantic structures', Although they vary in content, size, coverage, and overlap, they all primarily represent a collection of factual statements arranged in entity descriptions, possibly enriched by class hierarchies and corresponding property definitions. Question answering (QA) benchmarks and systems were so far geared primarily towards encyclopedic knowledge graphs such as DBpedia and Wikidata2,3. Currently, a new type of knowledge graph, called research knowledge graphs, is emerging wliuse contents are bibliographic metadata and scientific elements, such as ideas, theories, approaches, and claims as they are conveyed in scholarly contributions1,1 or OMTCS data structures for personalized medicine*. These novel research knowledge graphs increasingly intertwine three previously largely isolated aspects: semantic representations (semantic intelligence), machine learning (machine intelligence), and crowd and expert sourcing (human intelligence). In particular, scholarly communication is a more challenging application domain for QA due to:
1. The heterogeneity of knowledge representat ion;
JTIB—Leibniz Information Centre for Science and Technology, Hannover, Germany. 2L3S Research Center, Leibniz University Hannover, Hannover, Germany. 'Institute of Informatics, Federal University of Rio Grande do Sul, Porto Alegre, Brazil. ^Department of Informatics and Telecommunications, National and Kapodistrian University of Athens, Athens, Greece. ^Laboratory of Information Science and Semantic Technologies, ITMO University, St. Petersburg, Russia.—email: oliver.karras@tib.eu
is only able to retrieve correct answers for a .subset of handcrafted questions due to the more diverse data and question types 111 SciQA compared to the data JarvisQA was built on10. Second, we present initial insights into using ihe large language model (LLM) ChalGPT11 for answering the handcrafted questions. This evaluation aims to understand how well one of the current most famous 1.1 .Ms is able to answer complex queries on scholarly knowledge (with superlatives, comparisons, etc.). In this evaluation, we also focused on the 100 hand-crafted questions to compare the results of JarvisQA with those ofChalGPT, In both preliminary evaluations, we found that the systems perform rather low. In the best-performing configuration (hwnxisi)' the proof-of-concept implementation of JarvisQA was able to answer 52 questions with 12 correct answers. ChatGPT provided answers for 63 questions, of which only 14 answers were correct. These low numbers substantiate that answering questions about scholarly knowledge is a challenge tor current QA systems and LLMs1 \ for this reason, we conclude that the SciQA benchmark represents a challenging task for next-generation QA systems, as QA systems must now also deal with scientific knowledge in addition to encyclopedic knowledge.
Related work
Hie problem of answering questions expressed in natural language has received a lot of attention recently. Depending on the kind of system queries, e.g., text documents, knowledge graphs, relational databases, or image archives, benchmarks have been developed for evaluating the respective QA systems. Since the focus of this paper is on QA over scholarly knowledge graphs, we concentrate on QA benchmarks over knowledge graphs and linked data. An overview of relevant benchmarks is presented in Table 1.
One of the first datasets is WebQuestions14, which contains 5810 factoid question-answer pairs and is targeted at Freebase. It was created using the Google Suggest API to obtain questions that begin with a "wh"-word. 100K randomly selected questions were submitted to Amazon Mechanical Turk, asking workers to annotate the ones that can be answered by Freebase. In terms of structural complexity, WfbQufstion.s is simple, as many questions only contain one class, one property, and one instance. In 2016, WebQuestions was extended to WtuQuesi jonsSP15, by providing SPARQL queries for the 4737 questions that the annotators could fully process to find the answers.
The dataset SimpleQuestions1'1 also targets Freebase. It was created manually by English-speaking annotators and consists only of factoid questions. It is much larger than WubQuestiuNS, containing 108,442 simple questions paired with their corresponding answers and explanations. Diefenbach et al.'' created the benchmark SimfleQuestionsWjkidata by converting S[mpleQuestions to target Wikidata.
The LC-QuAD dataset" differs from the previous ones, as it includes not only simple, factoid questions but also complex ones, i.e., the respective SPARQL queries contain multiple triple patterns. Tire dataset contains 5000 pairs of question s-SPARQL queries targeting DBpedia. I he questions were generated semi-automatically by extracting sub-graphs containing triples within a 2-hop distance from a seed entity. The generation of the SPARQL queries and questions was facilitated automatically, using templates, and, then, refined manually. After the development of I.C-QuAD, its developers proceeded with the development of LC-QuAD 2.0 which contains 30,000 questions, their paraphrases, and their corresponding SPARQI. queries. LC-QuAD 2.0 targets both Wikidata and DBpedia 2018, and it was created similarly to LC-QuAD. LC-QuAD 2.0 also contains questions of higher complexity: Noil-factoid questions, questions with qualifiers, aggregates, temporal aspects (as qualifiers), and superlatives. The benchmark ComplexWebQuestions20 (34,689 questions) has similar complexity: it contains composition questions, superlatives and comparatives. It was generated from WEBQUESTlONSSPby sampling question-query pairs and automatically creating more complex SPARQL queries. From these queries, a set of questions was generated automatically by using 687 templates, and, then, reformulated by Amazon Mechanical Turk workers.
The benchmarks generated for the QALD challenges (http://qald.aksw.org/) are also of high complexity. 'Ihe QALD- 10 benchmark, generated for testing in the latest challenge (NLIWoD, KSWC2022), contains 394 manually created Wikidata-based questions of varied complexity and each is annotated with a manually specified SPARQL query and its oulput. Each question may contain counts, superlatives, comparatives, and temporal aggregators. The questions are available in 4 different languages, i.a., English, German, Chinese, and Russian.
QA Hcucli murks SciQA LC-QuAD 2.0 QAI.D-9 Web- QuesHons-SP Simple-Questions-Wikidala C om pi ex- Web-Questioos
Questions 2565 30,000 408 4737 21,957 34,689
Domain Scholarly communication Encyclopedic
Languages English English Multi- language English English Lnglish
Question types Complex, factoid, non-factoid Complex, factoid, non-factoid Complex, non-factoid Simple, factoid Simple Complex, factoid
Knowledge hase ORKG Wikidata DBpedia DBpedia Freebase Wikidata l-recbase
Formal language SPARQL SPARQL SPARQL SPARQL SPARQL SPARQL
Answers •i X s / / /
Paraphrases / ✓ X X X x
Table 1. QA benchmark comparison (full comparison is available in the ORKG13),
domain experts participating in the ORKG curation grants25,26 to create the relevant, realistic and useful questions.
(2) Workflow for the 2465 autogenerated questions
We performed this second workflow to enrich the SciQA benchmark as, even though the 100 handcrafted questions were purposefully created, this number of questions is rather small for a benchmark. For this purpose, we expanded the SciQA benchmark through the integration with a set of automatically generated questions, which have been created using a structured approach that involves a combination of handcrafted questions and queries, and the utilization of a LLM (in this case GPT-32 ). Hie objective of the autogenerated questions is to target specific parts of the ORKG by creating queries with placeholders that can be populated with various entities, thereby facilitating the generation of numerous natural language questions.
To create the autogenerated questions, wc followed a structured process. When creating the handcrafted questions, we observed that the data in the ORKG is very heterogeneous, which complicates the automatic generation of questions and queries. For this reason, we decided to set certain restrictions on the generation of the questions and queries. First, we decided to focus on a specific dataset from papers-with-code28 that is available in ORKG. Although this dataset belongs to only one research field (Computer Science), it is extensive, with 2236 papers (around 15% of the total number of papers in the ORKG) that are homogeneously described. 'Ihis homogeneity is important as it facilitates the automatic generation of questions and queries. Second, wc decided to focus on the questions and queries with the shape of a tree, the class which-what, and the type factoid to further narrow down the scope of the automatic generation. This shape, class, and type are the most common in the hand crafted questions and also match the nature of the selected papers-with-code data.
Initially, we crafted a set of eight queries and 32 questions. For each query, we created one question manually and three variations using GPT-327 with careful manual validation. Next, we collected all possible entities for the placeholders in the queries from the ORKG. We then filled the queries' placeholders with all possible entities, selecting one question randomly for each query. Finally, wc collected the results of the created queries and extracted metadata for the final set of questions.
The addition of the autogenerated questions expands the SciQA dataset to a total of 2565 questions and queries, providing a larger corpus for training machine-learning-based question-answering systems. This approach can be particularly useful compared to relying solely on handcrafted questions, which are often limited in number and may not capture the full scope of the underlying data. By contrast, the use of machine-generated questions provides a more diverse and extensive set of questions that can help improve the accuracy and robustness of machine-learning models in answering questions on large knowledge graphs.
SciQA benchmark
In this section, we provide an overview of the SciQA questions and their corresponding SPARQL queries. We first explain how we classified the questions to extract the metadata, before presenting some examples of handcrafted and autogenerated questions in more detail.
Question classifications. An appropriate question typology helps to satisfy two main goals of QA benchmark development. Namely, (1) an extensive coverage of the different topics in various subject areas that appear in the knowledge graph, and (2) the validation of the patterns used for writing questions and queries to ensure a better and more balanced distribution of questions across the possible different types of information requested.
There are many existing approaches to defining taxonomies of question types. Wendy Lehnert29 proposed a conceptual taxonomy with 13 conceptual classes, e.g., causal antecedent, goal orientation, enablement, etc. Li and Roth'" developed a two-layered taxonomy based on the answer type semantics: six coarse classes (abbreviation, entity, description, human, location, numeric value) and 50 fine classes (subclasses of different coarse classes do not overlap). Singhal et al.M designed a small set of simple answer types corresponding to question classes, words, and expected answer types: Person, Location, Organization, Date, Quantity, Duration, Linear Measure. For example, if a question starts with who or whom, its type will be Person. The system Quarc32 defines a question categorization based on the use of certain interrogative pronouns, e.g., who, what, when, where, or why. A similar approach was used by the system AskBill33, where eleven question types were defined with question patterns such as type "QTemporalAge" identified with pattern "How old/At (Which/What) age".
Research data and their descriptions have a very complex structure and semantics. When developing questions to search for information within this data, it is useful to define the types of expected answers and the focus of the questions. The definition of necessary types of expected answers is based on the results of evaluation campaigns ofQAI.D34 and analysis of characteristic problems associated with the task of mapping natural language to formal queries presented in Cimiano and Minock'1. These problems include:
• Lexical ambiguities arise when one word can be interpreted in different ways, i.e., it can refer to different entities or concepts.
• Light expressions such as the verbs "to be" and "to have", and the prepositions "of" and "with" either refer to an ontological property in a highly underspecified way or do not correspond to any property at all.
• Lexical gap between the user's vocabulary and that of the ontology.
• Complex questions that can only be expressed using queries involving aggregation functions, comparisons, superlatives, and temporal reasoning.
www.nature.com/scientificreports/
The definition of the focus of a question makes the search for an answer more specific. Moldovan et al.3a defined the question focus as a word or sequence of words that indicate what information is being asked about in the question. Ferret et al." defined the question focus as "a noun phrase that is likely to be present in the answer" consisting of a head noun and a list of its modifiers. For example, the question " What types of nanocarriers do have therapeutic effect?" has the focus on "types of nanocarriers". According to Mikhailian et aL5s there are two types of question foci:
1. Asking Point (AP), which is denoted explicitly, e.g., words "research problems* in the question "What are the research problems Vernier Iffect is related to?".
2. ExpectedAnswerType (EAT), is an implicit answer that can be inferred from the information provided by the question, e.g., answer type "person' is the EAT for the question' Who are the authors of the SOSA ontology!"
For our methodology, we modified the approach of Moldovan et al.56 by combining the question types, e.g., WFLAT, WHO, WHICH, etc., corresponding to classes from the ORKG schema, e.g., Paper, Problem, etc., and the question patterns that define the expected answer (BOOLEAN, WHAT-WH0. WHAT-WHEN, WHICH -WHERE, WHICH-WHAT, and WHO-WHAT). For instance, the question "Who is the author of the most recent paper about insects?" has the pattern WHO-WHAT. We also classified the questions according to the following dimensions:
• OKKG-content I his classification is based on the structure of the ORKG schema.
- Paper-based: Questions on the contcnt of a single or multiple rcscarch papers, e.g., " Which papers use DBLP as a dataset?"
- Comparison-based: Questions on the content of a comparison, i.e., on the properties that the contributions participating in a comparison share, e.g., "What is the most common knowledge representation method in Semantic Representations of Scholarly Communication?".
• Question content Following the approach of M ikhail ian et al.5s, we classify tile questions into either factoid, i.e., AP, or non-factoid, i.e., FAT. Factoid questions assume an explicit AP mapping to the entities of the ORKG ontology. If the answer to a question requires inference of a sequence of facts, counting, or filtering, we consider such questions non-factoid. We furLher classify them according to superlatives, e.g., "What is the most common lead compound in Anuran Antimicrobial Peptides Activity and Mechanism Against Different Biological Membranes?", negation questions, e.g., "What percentage of comparisons lacks a class link?", questions with counts, e.g.,"What is the total number of species examined in Invasion Biology-Enemy release hypothesis?" ranking questions, i.c„ asking for a min/max value, e.g., "What is the maximum female percentage in Brief Psychotherapy for Depression Studies?", temporal questions, e.g., "How many studies are published after 2019?" or a combination of various types of content, e.g., "Which was the most popular approach for summarization until 2002?".
Finally, we characterised the questions based on important properties of their respective SPARQI, queries:
• Nu uiber of triple patterns In contrast to simple questions, the SPARQL query of complex questions consists of more than a single triple pattern18. As presented in Tables 3 and 4, the dataset contains both simple and complex questions with up to 14 triple patterns.
• Query shape: We identified the shape (single edge, chain, star, cycle, tree, etc.) of the queries according to Bonifati et al.3'. Note that the classification based on the number of triple patterns is incorporated in this classification, as simple questions can be classified as single-edge queries.
Characteristic Number
Research fields SPARQI, queries 48 too
Query shapes Tree Chain Star Forest hd (;r Cycle
47 39 7 5 ] 1
Query classes Which What Boolean What When What Who Who What Which Where
S4 5 4 3 3 1
Query types Factoid Non-factoid Superlative Temporal Count
61 39 26 3 7
Query components Min; 1, Med: S, Mai; 14
Triple patterns Min: 1, Med: 4, Mai: 14
Table 3. Overview of the SciQA handcrafted queries.
SELECT ?range ?srcLabel AVG(?val) AS ?avgVal
WHERE {
r:R153881 p:compareContribution ?contrib .
?paper p:hasContribution ?contrib ;
p:hasPublicationYear ?year .
BINDCxsd:int(?year) AS ?y) .
VALUES('range ?min ?max) {
("2001-2885" 2881 2085)
("2006-2810" 2806 2010)
("2011-2815" 2811 2015)
("2016-2820" 2816 2020)
} FILTER(?min <= ?y && ?y <= ?max) .
?contrib p:hasEnergySources ?energySrc .
?energySrc rdfs¡label ?srcLabel ;
p:hasGeneration ?energyGen .
?energyGen p:hasValue ?genVal .
BIND(xsd:float(?genVal) AS ?val)
} ORDER BY ASC(?range)
2. Handcrafted Question What is the most common knowledge representation method in Semantic Representations of Scholarly Communication?
The second question (ID 3 in SciQA-Handcrafted) belongs to the research field Databases/Information Systems from the domain Computer Science. This non-factoid question is based on the ORKG comparison Semantic Representations of Scholarly Communication12. This comparison provides an overview of publications on semantic representations of scholarly communication by focusing on scholarly communication as a whole and not specific data such as citations. The question is a typical one about the most frequent occurrence of information. Specifically, it is about the most commonly used knowledge representation data model for scientific communication, which in this case is the Resource Description Framework (RI)1, cf. https:// www.w3.org/RDF/). The SPARQL query includes three triple patterns, uses seven query components, and is shaped as a chain.
PREFIX r: <http://orkg.org/orkg/resource/>
PREFIX c: <http://orkg.org/orkg/class/>
PREFIX p: <http://orkg.org/orkg/predicate/>
PREFIX rdfs: <http://www.w3.org/2888/81/rdf-schema\#>
PREFIX xsd: <http://www.w3.org/208l/XMLSchema\#>
SELECT COUNT(?repr) AS ?cnt ?repr WHERE {
r:R8364 p:compareContribution ?cont .
?cont p:seraantic_representation ?sys . ?sys p:knowledge_representation ?repr .
} GROUP BY ?repr ORDER BY DESC(?cnt) LIMIT 1
3. Handcrafted Question Where did the study with maximal geographic scale take place in Genetic Variability (COI Variation) in Studies Large Sampled (>1000 Sequences)?
The third question (ID 78 in SciQA-Handcrafted) belongs to the research field Ecology and Biodiversity of Animals and Ecosystcms, Organismic Interactions from the domain of Zoology. This non factoid question is based on the comparison Genetic Variability (COI Variation) in Studies Large Sampled (>1000 Sequencesfl which compares the genetic variability in studies containing more than 1000 cytochrome c oxidase I (COI) harcoding sequences. The question aims to identify where the study with the maximum geographic scope took place, which in this case is a study conducted in the United States of America, Mexico, and Canada. The SPARQL query has six triple patterns, uses six query components, and is shaped like a tree.
SELECT ?locatiûn, ?location_label WHERE {
{ SELECT (MAX(?gco_scale_num) AS ?max_geo_5calc) WHERE {
orkgr:R149849 orkgp:compareContribution ?contrib . ?contrib orkgp:geographicScale ?geo_scale .
BINDCxsd: integer(?geo_scale) AS 7geo_scale_num) .
Î
Î
orkgr:R149849 orkgp:compareContribution 7contrib . ?contrib orkgp:geographicScale ?geo_scale ;
orkgp:studyLocation ?location .
?location rdfs:label ?location_label .
BIND(xsd: integer(?geo_scale) AS ?geo_scale_num)
>
GROUP BY(?location_label) HAVING(?geo scale nuiti = ?max geo scale)
Autopcncratcd Question Can you provide the highest benchmark result, including the metric and score, for the SequentialMNIST dataset?
The fourth question (TD 1355 in SciQA-Autogenerated) belongs to the research field Computer Science. This non-facloid question is based on the content of the ORKG imported from papers-with-code18. The question is about fetching the top (or the best) evaluation score recorded in the ORKG, the results should be fetched for each distinct evaluation metric used in the evaluation. The related SPARQL query includes ten triple patterns, uses nine query components, and has a shape of a tree.
SELECT DISTINCT WHERE {
necric_lbl (MAX(?value) AS ?score)
SELECT ?metric ?metric_lbl ?value WHERE {
?dataset a orkgc:Dataset ;
rdfs:label ?dataset_lbl . FILTER (SIR(?dataset_lbl) = "Sequential^WNIST") ?benchmark orkgp:HAS_DATASET ?dataset ;
orkgp:HAS_EVALUATION ?eval . ?eva1 orkgp:HAS_VALUE ?value .
OPTIONAL f'eval orkgp:HAS_HETRIC ?metric .
?metric rdfs:label ?mctric_lbl .}
?cont orkgp:HAS_BENCHHARK ?benchmark . OPTIONAL {?cont orkgp:HAS_MODEL ?model .
?model rdfs:label ?model_lbl .}
}
ORDER BY DESC(?value)
}
} GROUP BY ?metric ?metric_lbl
Autogenerated Question List the title and ID of research papers that contain a benchmark over the SST-2 Binary classification dataset?
The fifth queslion (ID 524 in SciQA-Autogcnerated) belongs to the research field Computer Science. This factoid question is based on the content of the ORKG imported from papers-with-code28, which describes the evaluation results of machine learning models benchmarked oil commonly used datasets in the natural language processing and machine learning communities. The question requests the IDs and titles of papers that have models that benchmarked a particular dataset, in this case, the SST-2 Binary classification dataset. The related SPARQL query includes six tri pie patterns, uses four query components, and has the shape of a tree,
www.nature.com/5cientificreports/
SELECT DISTINCT ?paper ?paper_lbl
WHERE {
?dataset a orkgc:Dataset ; rdfs:label ?dataset_lbl .
FILTER (STRC?dataset_lbl) = "SST-2uBinary^classification"J
îbenchmark orkgp: HAS DATASET ?dataset .
?cont orkgp:HAS_BENCHMARK ?benchmark .
?paper orkgp:P31 ?cont ;
} rdfs:label ?paper_lbl
Applicability and feasibility evaluation
In this section, we present two preliminary evaluations using the handcrafted part of the SciQA benchmark. First, we show the results of a proof-of-concept implementation of a QA system based on the JarvisQA system10. Second, we show initial insights into using ChatGPT11 tor answering the handcrafted questions.
Proof-of-concept based on JarvisQA. In a preliminary analysis, we aim to understand how SciQA can be used by a QA system that is focused on scholarly knowledge. For this putpose, we investigate the performance of a proof-of-concept implementation based on JarvisQAlu.
Experimental Setup. farvisQA is fundamentally designed to answer questions about scholarly knowledge. The system is based on BERT41 but works only on tables and tabular views of scholarly knowledge graphs, such as ORKG comparisons. SciQA does not rely only on tables and tabular views (comparisons) but has a broader spectrum of question/answer types. For this reason, we can answer 52 of the handcrafted questions (52%) with JarvisQA as they correspond to its input form. We configured our proof-of-concept implementation of JarvisQA to run on the compatible questions of SciQA, and we use seven distinct experimental setups that JarvisQA provides. I hie to the limited coverage of questions that the system can answer, we limit the results to two categories of questions. The evaluation is conductcd in terms of precision@k, recall@k, and_fl@& metrics.
Ri'yuhy Table 5 shows the evaluation results of these experiments for two main categories of questions: normal and overall. While the category normal refers to single-answer questions, the category overall aggregates singleanswer questions and all other question types that JarvisQA can answer, such as listing and boolean questions. We note that the performance decreases across all the setups for the overall category because of the complex nature of the SciQA benchmark and the answers it expects, unlike what JarvisQA was trained with and thus can answer10.
ChatGPT arid SciQA. Besides the use of SciQA with a QA system focused on scholarly knowledge, we performed an additional preliminary evaluation based on all handcrafted questions using ChatGPT11. Numerous M.Ms adept at solving common natural language tasks have now been released such as ChatGPT11, Galacticn44, LaMDA45, Codex46, or Sparrow47. Some of them show better performance on technical knowledge tasks, e.g., Galactica1'1, some in the medical domain, e.g., PubMedQA18 and MedMCQA4J, etc. In this experiment, we do not aim to test all these LLMs on SciQA, but to estimate the baseline performance of LLMs for scholarly questions about various topics. We chose ChatGPT for our experiment as it is one of the most prominent LLMs at the moment, and it is not domain-specific. ChatGPT should be able to answer the questions from SciQA, as the
JarvisQA setup Normal Overall
Precision Recall Fl Precision Recall Fl
et «¡10 #10 @1 S>10 91 ISIO 910 #1 MO
JarvisQAeus 0.190 0.271 0.191 0.271 0.190 0.271 0.13« 0.190 0.136 0,190 0.136 0.190
JarvisQAirs 0.193 0.2 S4 0.194 0.254 0.194 0.254 0.138 0.179 0.133 0.179 0.138 0.179
JarvisQAecsi 0.134 0.187 0.134 0.188 0.134 0.187 0,098 0.135 0.098 0.135 0.098 0.135
larvisQAiusi 0.169 0.288 0.169 0.288 0.169 0.288 0.122 0.202 0.122 0,202 0.122 0.202
JarvisQAuu-BOS 0.134 0.246 0.134 0.246 0.134 0.246 0,098 0.174 0,098 0.174 0.09S 0.174
larvisQAxm 0.172 0.339 0.172 0.339 0.172 0.339 0.125 0.233 0.125 0.235 0.125 0.235
JarvisQAxxia 0.169 0-246 0.169 0.246 0.169 0.246 0.122 0.171 0.122 0.174 0.122 0.174
Table 5. Evaluation results of running farvisQA against the handcrafted questions of the SciQA benchmark. JarvisQA setups follow similar notations as introduced in"1. Top performing setup is indicated in bold, second best is underlined.
source texts of the papers and the topics of the questions that were used to develop the dataset are mostly openly available on the internet. For this reason, we assume that questions from the SciQA can he potentially processed and answered by LLMs such as ChatGPT. In this way, this evaluation aims to gain initial insights into how well one of the current most famous LLMs, which is not trained specifically on ORKG data, is able to answer complex queries on scholarly knowledge (with superlatives, comparisons, etc.).
Experimental setup- The underlying model of ChatGPT is designed to generate detailed answers to a user's questions. For this reason, wc have added the additional prompt "short:" to each of the 100 handcrafted questions to get shorter answers similar to the answers in SciQA. Although ChatGPT's responses were shorter, they were still very detailed. In (he fulure, a more refined individual tuning of the prompt is necessary to obtain answers in a more similar format to the answers of our dataset. However, we decided that this way of retrieving the answers is sufficient for our preliminary evaluation of SciQA. After collecting the 100 answers to all ques tions, the assessment of its correctness was performed with expert opinion.
Four experts compared the ChatGPT's answers with the answer from the SciQA dataset. If correct facts were mentioned in the text returned by ChatGPT, that answer was assessed as "Correct". Otherwise, ihe result was assessed as "Incorrect". Also, if the system returned a response that it could not answer the question, it was assessed as "No answer" Alter all the answers were independently assessed by the experts, the experts compared their results and discussed any disagreements in a meeting. During this discussion, situations arose in which the experts could not agree whether the answer was derived from the paper or data source mentioned in the question or whether the answer had been generated from common sense or general knowledge. As a result, the experts assessed these answer as "Uncertain." In Table 7, we show four examples of each of the four assessment types "Correct", "Incorrect", "Uncertain", "No answer". These examples include the question from SciQA with the SciQA answer, the answer by ChaLGPT, and the experts' assessment of the answer by ChaLGPT with an explanaliun.
Resu/is. In Table 6, we provide an overview of the results of the experts' assessment. We found from this analysis that ChatGPT was able to generate answers for 63 of the 100 handcrafted questions. Fourteen of these 63 answers are correct, 40 answers arc incorrect, and nine answers arc uncertain. Although these results arc slightly better compared to the results of the best performing configuration of the proof-of-concept implementation of JarvisQA (/«rvifxLsi: 12 correct answers), the performance of ChatGPT in answering questions about scientific knowledge is still low with only 14 correct answers. This preliminary evaluation shows the limited applicability and low accuracy of even the current cutting-edge LLM ChatGPT in answering specific questions about scholarly knowledge.
Discussion
In this section, we discuss three key aspects of SciQA: (1) the higher complexity of scholarly knowledge; (2) file insufficiency of I.I.Ms for dealing with scholarly knowledge, and (3) the advancement of knowledge graphs towards cognitive knowledge graphs for facilitating the integration of semantic and machine learning approaches.
We argue, that scholarly knowledge is more complex than common sense knowledge, e.g., encyclopedic knowledge. We illustrated this complexity with some examples, e.g., Fig. 1. In these examples, the semantic units of knowledge are not simple entity descriptions, where an entity, such as an organization, place, or person is described with a set of RDF triple slatements having the entity identifier as a common subject, In contrast, for scholarly knowledge graphs such as the ORKG a scholarly contribution description comprises numerous interwoven entity descriptions, which include, for example, descriptions of processes and process steps, material characteristics, components, measurements, simulations as well as bibliographic metadata. A single semantic unit in the ORKG commonly consists of more than a dozen tightly interlinked entity descriptions, which only logcthcr reasonably convey information, compared to encyclopedic or common sense knowledge, where often already a single entity description contains sufficient information, This inherently more complex structure of scholarly knowledge graphs makes question answering on them significantly more challenging, which we have demonstrated with our applicability and feasibility evaluation. In particular, the results of the two preliminary evaluations show how challenging it is for a QA system and a LLM to answer questions about scholarly knowledge. Neither a QA system that is intended for scholarly knowledge and even trained on ORKG data nor the LLM ChatGPT was able to perform well in answering the 100 handcrafted questions from SciQA, In both cases, the systems were unable to adequately solve the challenge of scholarly QA (Table 7).
In our preliminary evaluations, we observed that even QA systems specifically designed for academic data, including one of the most advanced LLMs available at the time of this article, struggle to excel on the SciQA benchmark. The SciQA dataset encompasses various question types stemming from distinct entities within the knowledge graph, rather than solely relying on tabular views, which is the primary format larvisQA operates on. Moreover, to accurately respond to SciQA queries, a QA system must comprehend Ihe context of each question,
Answer Nu answer
Correct 1 ritt irrtet Uncertain
* Answers 14 40 ! 37
www.nature.com/scientificreports/
SciQA question SciQA answer ChatGPT answer Assessment and explanation
What is the most common Knowledge representation method in "Semantic representations of scholarly communica tion"? RDF 'Ihe most common knowledge representation method in "Semantic representations of scholarly communication" is the use of semantic web technologies, such as RDF (Resource Description Framework) and OWI (Web Ontology Language), to represent and integrate scholarly communication data. The answer is correct. The comparison mentions several semantic representation methods, each of which is associated with one or more knowledge representation methods.
Does paper "Disambiguating authors in citations on the web and authorship correlations" in Semi-supervised author name disambiguation employe Cosine similarity? Yes No The answer is not corrcct. The paper does indeed mention the metric.
How many studies are published after 2019 tor '"Ihe Use ot Silicon Carbide Thin Films in Different Device Types"? 3 As an AI language model, I cannot access up-to-date information on the number of studies published after 2019 for "The Use of Silicon Carbide Thin Films in Different Device Types". However, a search on Google Scholar or other academic databases can provide the most up-to-date information on this topic. In this case, ChatGPT provided no answer, as there are 3 of 5 papers in the comparison that were published in 2020 or later.
What risk factors for OSA are discovered in the paper with 100% OSA frequency in the "Comparison of the frequency of obstructive sleep apnea (OSA) in patients with mucopolysaccharidosis (MPS)"? Increase in age In the paper "Comparison of the frequency of obstructive sleep apnea (OSA) in patients with mucopolysaccharidosis (MPS)", the authors found that all patients with mucopolysaccharidosis (MPS) in their study had obstructive sleep apnea (OSA). The paper identifies several risk factors for OSA in patients with MPS, including:Upper airway abnormalities: MPS can cause structural abnormalities in the upper airway, such as enlarged tonsils and adenoids, a small or recessed jaw, and a thick and large tongue, which can increase the risk of OSA. Obesity: F.xcess weight can put additional pressure on the airway, increasing the likelihood of airway obstruction during sleep. Age: Ihe risk of OSA increases with age, and patients with MPS may be at higher risk due to the early unset and progressive nature of the disease. Type of MPS: The study found thai patients with MPS 1 and MPS II were at higher risk of OSA than those with MPS III, although all patients with MPS in the study had OSA. Severity of disease: Patients with more severe forms of MPS, as measured by a higher number of enzyme deficiencies and greater dinical manifestations, were more likely to have OSA. Overall, the study highlights the high prevalence of OSA in patients with MPS and the importance of early diagnosis and treatment to prevent potential complications. 'Ihe answer is uncertain as the response is too extensive, and it is impossible to unambiguously determine whether it contains a correct answer or not
Table 7. Sample of ChatGPT answers and our evaluation of the results as part of the qualitative assessment of the handcrafted SciQA benchmark.
which is embedded in the ORKG graph structure. One contributing factor to this challenge is the deficiency in NT.P components tailored to academic data, such as entity linkers and query builders"". Another significant limitation is that LLMs like ChatGPT and BERT do not possess the contextual understanding specific to a knowledge graph, such as the ORKG, which further hinders their performance on the SciQA benchmark. Taking into account all the factors mentioned above, it becomes increasingly clear that there is a pressing need for the research community to rally behind the SciQa benchmark. By collaborating to develop systems that perform well on SciQA, researchers can contribute to improve and expand this QA dataset, as well as make progress in this field of QA tor scholarly knowledge. With this goal in mind, we launched the Scholarly Question Answering over Linked Data (QALD) Challenge with a task using SciQA as one of the open competitions at the 22nd International Semantic Web Conference 202351 With this challenge, we hope to generate more baselines and inspire the community to crcatc an array of scholarly-oriented tools and QA systems. Ultimately, this collaborative effort will foster significant advancements in the field, benefiting academia as a whole.
A reason for this challenge even for LLMs lies in the fact, that these models are very good at recreating common sense knowledge, which can be found in varying forms in several different sources. However, due to their nature of employing probability distribution over sequences of words, they are not good at dealing with knowledge to be found only in a single or very few sources. This issue was also shown recently with the failed LLM
Galacl ica trained un scientific literature, which had to lit taken offline after Ihree days when it became clear Lhal the model's ratio between hallucination and reasonable answers is too unfortunate to be of any use'3. We deem, that this is an inherent characteristic of LLMs, which can also not be addressed with further improvements of the mudels themselves. However, a combination of I.I.Ms with symbulic knowledge representation approaches (such as the ORKG and SciQA) can be a promising avenue for leveraging the potential of A1 and also for domains with more unique knowledge production such as science.
Scholarly knowledge graphs such as ORKG demonstrate the advance of the knowledge graph concept towards more cognitive knowledge graphs, which enable the trustworthy integration of artificial and human intelligence. In cognitive knowledge graphs, the constituents will be more complex elements, such as ideas, theories, approaches, and claims as they are conveyed, for example, in scholarly contributions, but also in other areas such as industrial product models54, common vulnerabilities and exposure descriptions ill developer secui ity " or ONI ICS data for personalized medicine'1. We see these hase constituents of cognitive knowledge graphs to be complex fabrics of entity descriptions arranged according to certain patterns, such as graphlets. In network analysis and graph theory, the notions graphlet5i and motif1' were introduced tu provide a structuring element between whole graphs and individual nodes and edges. Hence, in order to be able to effectively represent and manage more complex knowledge artifacts, the notion of graphlets can be applied to knowledge graphs (as we did in SciQA with research contributions). Cognitive knowledge graphs can be of particular importance to support the step from correlation to causality—while correlation arises from the detection of statistical relationships and patterns in the data, we plan to use rich contextual knowledge from knowledge graphs as additional signals for causality testing. Such integration of symbolic and sub-symbolic intelligence as HybridAI (cf. Breit et al.58) for a recent survey of approaches) can help us to systematically anchor transparency, traceability, explainability, trustworthiness, and reliability in data sciencc and AI methods.
Conclusions and future work
In this section, we draw some conclusions and point out directions for future work. We address the problem of missing QA benchmarks foT scholarly knowledge. So far, QA systems and corresponding benchmarks were mainly geared towards encyclopedic knowledge composed of relatively simple to moderate semantic structures'. In contrast, the consideration of scientific knowledge combined with knowledge graphs is rather new and challenging due to heterogeneous representations, concept drifts and evolution over time, different levels of granularity, and novel semantic structures.
l;or these reasons, we developed the SciQA benchmark for scholarly knowledge as a newchallengingta.sk for the next-generation QA systems with 13 different researchers using a defined bottom-up methodology. SciQA cunLains 100 handcrafted natural language questions willi paraphrases, corresponding human- and machine-readable SPARQI, queries with their results. These questions and queries are analyzed according to several classifications and cover 48 different fine-grained research fields such as Computer Science, Engineering, Chemistry, Geology, Immunology, and Economics (see Table 3). In addition to the handcrafted question-answer pairs, we semi-automatically created a set of 2465 questions derived from eight question templates. This approach is currently limited to the computer scicnce domain, where wc have a large set of homogeneously structured and described data. However, once the ORKG comprises more such homogeneously structured contribution descriptions, the SciQA approach can easily be expanded to further research fields.
The initial results of the evaluation of SciQA using JarvisQA and ChatGPT demonstrate the difficulties of scholarly knowledge in general for a system that is designed to answer questions about scholarly knowledge, or a large language model capable of advanced reasoning and language unders landing. Based on Lhese insights, we conclude that the SciQA benchmark represents a challenging task for QA systems, hut its implementation is realistic and feasible.
This work is the foundation for a longer research and technology development agenda. We envision advancing the concept of knowledge graphs from rather simple, atomic entity descriptions towards richer structured knowledge graphs, comprising fabrics of complex knowledge structures such as knowledge graph cclls ''. We plan to update SciQA annually as the ORKG evolves to include more content for more questions, queries, and answers. We also currently launch the Scholarly Question Answering over Linked Data (QALD) Challenge with a task using SciQA as one of the open competitions at the 22nd International Semantic Web Conference 2l)23j1-52. An extension of this work is to perform QA on federated scholarly knowledge graphs that link ORKG content to metadata about articles, dafasets, people, organizations, etc, published by other scholarly infrastructures60. Given the advanced standardization of the persistent identification, description, interlinking, and exchange of metadata about these entities as well as the provision of (programmatic) access to metadata through systems such as the GraphQL-based PID Graph, the federated integration of ORKG content with metadata about contextual entities is straightforward. This will enable QA on scholarly knowledge understood broadly to include both the scientific knowledge published in articles interlinked with contextual knowledge about its production and consumption.
Data availability
The full SciQA dataset and a snapshot of ORKG data is available from Zenodo (https://doi.org/10.528f/zenodo. 5845197)a and Hugging Face (https://huggingface.co/datasets/orkg/SciQA)''1.
Code availability
The source code of larvisQA10 is available from GitHub (cf. https://github.com/YaserJaradeh/JarvisQA).
Received: 2H Septemher 2022: Accepted: 15 April 2023 Published online: 04 May 2023
References
1. Heist, N., Hertling, S., Ringler, D. 8c Paulheim, H. Knowledge graphs on the weh • An overview. Knowledge Graphs for explainable Artificial Intelligence. 3-22 (2020).
2. Chakraborty, N. et al. Introduction to neural network-based question answering over knowledge graphs. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 11 (2021).
3. Diefenbach, D., López, V., Singh, K. D. & Maret, P. Core techniques of question answering systems over knowledge bases: A survey. Knowl. Inf. Syst. 55, 529 569 (2018).
4. Jaradeh, M. Y. et al. Open research knowledge graph: next generation infrastructure for semantic scholarly knowledge. K CAP, 243-246(2019).
5. Stocker, M. et al. SKG4EOSC—Scholarly knowledge graphs tor EOSC: Establishing a backbone of knowledge graphs for FAIR scholarly information in EOSC. Res. Ideas Outcomes 8. e83789 (2022).
6. Kim, D. et al. Knowledge boosting: A graph-based integration approach with multi-omics data and genomic knowledge for cancer clinical outcome prediction. /. Am. Med. Inform Assoc. 22, 109-120 (2015).
7. Stocker, M. et al. FAIR scientific information with the open research knowledge graph. FAIR CoHNecfhttps://doi.org/10.3233/ FC-221513 (2023).
8. Budde, L. et al. Investigation of the material combination 20mncr5 and x45crsi9-3 in the tailored forming of shafts with bearing seats. Product. Eng. 16,661-671 (2022).
9. Karras, O. Investigation of the material combination 20mncr5 and x45crsi9 3 in the tailored forming of shafts with bearing seats. https://doi.org/10.48366/R288295 (20231.
10. laradeh, M. Y.. Stocker, M. 8c Auer. S. Question answering on scholarly knowledge graphs. TPDL. 19-32 (2020).
11. Lciter, C. et al. Chatgpt; A meta-analysis after 2.5 months. https://doi.org/10.48550/ARXIV.2302.13795 (2023).
12. Saikli, T., Ghusal, T., Mittal, A., Ekbal, A. 8c Bhattacharyya, P. Scienceqa: A novel resource for question answering on scholarly articles, Int. I Digital Libraries 23,289-301. https://doLoii/10.1007/s00799-022-00329-y (2022).
13. Cortes, E. & Karras. O. Question answering over linked data benchmark comparison. https://doi.org/10.48366/R161787 (2022).
14. Berant, J., Chou, A., Frostig, R. & Liang, P. Semantic parsing on freebase from question-answer pairs. EMNLP. 1533-1544 (2013).
15. Yih, W.-T., Richardson, M., Meek, C., Chang, M.-W. 8c Suh, J. The value of semantic parse labeling for knowledge base question answering. ACL. https://doi.org/10.18653/vl/P16 2033 (2016).
16. Bordes, A., Lsunier, N., Chopra, S. & Weston, J. Large scale simple question answering with memory networks. CoRR. abs/1506.02075 (2015).
17. Dicfenbach, D.. Tanon, T. P., Singh. K. D. 8c Maret, P. Question answering benchmarks for Wikidata. ISWC Posters Demos. (2017).
18. Trivcdi, P., Mahcshwari, G., Dubcy, M. 8c Lehmann, J. Lc-quad: A corpus for complex question answering over knowledge graphs. ISWC. 210-218(2017).
19. Dubey, M., Banerjee, D.. Abdelkawi, A. & Lehmann, J. IO-QuAD 2.0: A large dataset for complex question answering over Wikidata and DBpedia. ISWC. 69-78 (2019).
20. Talmor, A. 8c Berant, I. The web as a knowledge-base for answering complex questions. NAACL. 641-651 (2018).
21. Karras, O., Groen, E. C., Khan, J. A. 8c Auer, S. Researcher or crowd member? Why not both! The open research knowledge graph for applying and communicating CrowdRE research, in 2021 IEEE 29th International Requirements Engineering Confcrcnce Work shops (REW). https://doi.org,'10.1109/REW53955.2021.00056 (2021).
22. Ocien, A. Semantic representations of scholarly communication. https://doi.org/10.48366'R8364 (2022).
23. Auer, S. et al. Sciqa benchmark: Dataset and rdf dump, https://doi.org/10.5281/zenodo.7729047 (2023).
24. Oelen, A., Jaradeh, M. Y., Stocker, M. 8c Auer, S. Generate FAIR literature surveys with scholarly knowledge graphs, in ACM/IF.F.F. Joint Conference im Digital Libraries. (2020).
25. lstorkg curation grant program, https://orkg.org/page/lst-curation-grant-program (2021). (Accessed on 03/13/2023).
26. 2nd orkg curation grant program, https://orkg.org/pagc/2nd-curation-grant-program (2021 ). (Accessed on 03/13/2023).
27. Brown, T. B. et al. I-anguage models are few-shot learners. https://doi.org/10.48550/ARXIV.2005.14165 (2020).
28. Papers with code, https://paperswithcode.com/about (2020). (Accessed on 03/13/2023).
29. Lehnert, W. A conceptual theory of question answering, in Readings in Natural Ixmguagc Processing (Morgan Kaufmann, 1986).
30. Li, X. & Roth, L). I .earning question classifiers. ACL. (2002).
31. Singhal, A. et al. AT 8cT at TREC-8. TREC 8,317-330 (1999).
32. Riloff, E. 8c Thclen, M. A rule-based question answering system for reading comprehension tests, in ANLP/NAACL Workshop on Reading comprehension tests as Evaluation for Computer-based Language Understanding Systems (2000).
33. Leidner, J. I.. Question answering over unstructured data without domain restrictions. arXivpreprint cs/0207058 (2002).
34. Lopez, V., Unger, C., Cimiano, P. 8c Motta, E. Evaluating question answering over linked data. Web Semantics. 21,3-13 (2013).
35. Cimiano, P. 8c Minock, M. Natural language interfaces: What is the problem? A data-driven quantitative analysis, in Int. Conf. on Appt. of Natural Ixing. to Inf, Systems (Springer, 2009).
36. Moldovan, D. et al. The structure and performance of an open domain question answering system. ACL 563 570 (2000).
37. Ferret, O. et al. Finding an answer based on the récognition of the question focus. TREC. (2001 ).
38. Mikhailian. A., Dalmas, T. 8c Pinchuk, R. Learning foci tor question answering over topic maps. ACL-IICNLP 325-328 (2009).
39. Bonifati. A.. Martens, W. 8c Timm. T. An analytical study of large SPARQL query logs. VLDB /. 29.655-679 (2020).
40. Kullmann. F. et al. Comparison of Studies on Germany's Energy Supply in 2050 (Tech. Rep Technookonomische Systemanaly.se, 2021).
41. Kullmann, F. et ul. Comparison of studies on Germany's energy supply in 2050. https://doi.org/10.48366/R153801 (2021).
42. Marín, M. A. Genetic variability (COI variation) in studies large sampled (>1000 sequences). https://doi.org/10.48366/R149849 (2022).
43. Devlin, J., Chang, M.-W., Lee, K. 8c Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. https://doi.org/10.18550/ARX1V. 1810.04805 (2018).
44. Taylor, R. et al. Calactica: A large language model for science, https://doi.org/10.48550/ARXI V.2211.09085 (2022).
45. Ihoppilan, R. et al. Lamda: Language models for dialog applications. https://doi.org/10.48550/ARXlV.2201.08239 (2022).
46. Chen, M. et al. Evaluating large language models trained on code. https://doi.org/10.48550/ARXIV.2107.03374 (2021).
47. Glaese, A. et al. Improving alignment of dialogue agents via targeted human judgements. https://doi.org/10.48550/ARXIV.2209. 14375 (2022).
48. Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W. 8c Lu, X. Pubmedqa: A dataset for biomedical research question answering. https://doi. org/10.48550/ARXIV. 1909.06146 (2019).
49. Pal, A., Umapathi, L. K. 8c Sankarasubbu, M. Medmcqa : A large-scale multi-subject multi-choice dataset for medical domain question answering. https://doi.Org/10.48550/ARXlV.2203.l 1371 (2022).
50. laradeh, M. Y., Singh, K., Stocker, M., Both, A. 8c Auer, S. Information extraction pipelines for knowledge graphs. Knowl. Inform. Svs/.https://doi.org/10.1007'sl0115-022-01826-x (2023).
51. Scholarly qald challenge. https://kgqa.github.io/scholarly-QALD-challenge/2023/ (2023). (Accessed on 03/13/2023).
52. Github repository: Scholarly qald challenge. https://github.com/KGQA/scholarly-QALD-challenge (2023). (Acccsscd on 03/13/2023).
53. Why metas latest large language model only survived three days online | mit technology review. https://www.technokigyreview. com/2022/1 l/l8/1063487/meta large-language-model-ai-anly-survived-three-days-gpt-3-science/. (Accessed on 03/13/2023).
54. Graiigel-Gon/alc-v, I. et ul. An rdf-based approach tor implementing industry 4.0 components with administration shells. In 21st IEEE International Conference on Emerging Technologies and Factory Automation, ETFA 2016. Berlin, Germany, September 6-9, 2016,1-8. https://doi.org/10.1109/ETFA.2016.7733503 (IEEE, 2016).
55. Fischer, F. etal. Stack Overflow Considered Harmful? The Impact of Copy &Paste on Android Application Security (2017).
56. Prxzulj, N.. Corneil, D. G, Sr furisica, 1. Modeling interactome; Scale-free or geometric?. Biomformatics 20, 3508-3515, https:// Joi.org/t0.1093/bioinfbrmatics/bth436 (2004).
57. Milo, R. etal. Network motifs: Simple building blocks of complex networks. Science 298. 824-827. https://doi.org/10.] 126/scien ce.298.5594.P24 (2002).
58. Breit, A. el ul. Combining machine learning and semantic web: A systematic mapping study. AC\l Commit. Surv.hUps://(]oi.org/ 10.1145/3586163 (2023).
59. Vogt, L., D'Souza, J„ Stocker, M. St Auer, S. Toward representing research contributions in scholarly knowledge graphs using knowledge graph cells. /CDthttps ;//doi.org/10.1145/3383583.3398530 (2020),
60. Haris, M„ l-arfar, K, K„ stockcr, M. 8: Aucr, 5. federating scholarly infrastructures with (JraphqL. (L'.4 ilLhttps://doi.org/lo, 1 (107/ 978-3-030-91669-SJ24 ¡2021).
61. Hugging lace—orkg/sciqa. https://huggingface.co/datasefs/orkg/SciQA (2023). (Accessed on 03/13/2023).
Acknowledgements
'I his work was co-funded by the European Research Council for the project ScienceGRA PH (Grant agreement ID: 819536) and by the German Federal Ministry of Education and Research (BMBF) under the project Leib nizKILabur (Grant no. 01DD20003), German Research Foundation DFG for KFDI4Ing (No. 442146713) and NFDI4DataScience (No. 460234259). It has, also, received funding from the European Union's Horizon 2020 research and innovation programme under the Marie Sklodowska Curie Grant agreement No. 101032307. It is, also, financed in part by the Coordena^ao de Aperfei^oamento de Pessoal de Nivel Superior-Brasil (CAPES)-T'inance Code 001.
Author contributions
S.A. and D.M. conceived and designed the analysis, S.A., M.Y.J., O.K., D.P., D.R., and F,.T. collected the data, D.A.C.B., C.B., E.G.C., M.Y.J., O.K., D.P., D.R., I.S., and E.T. contributed data or analysis tools, S.A., M.Y.J., O.K., D.M., D.P., D.R., I.S. performed the analysis, S.A., E.G.C., M.Y.J., O.K., M.K., D.M., E.T, wrote the paper, S.A., D.A.C.B., E.G.C., M.Y.J., O.K., M.K., DM., D.R., I.S., M.S. reviewed the manuscript.
Funding
Open Access funding enabled and organized by Projekt DEAL.
Competing interests
The authors declare no competing interests.
Additional information
Correspondence and requests for materials should be addressed to O.K. Reprints and permissions information is available at www.nature.com/reprints.
Publisher's note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
[(ec) © I t)pen Access This article is licensed under a Creative Commons Attribution 4.0 International ifca^BL^« License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. 1 he images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the materia]. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licencc, visit http://creativecommons.Org/licenscs/by/4.0/,
@ The Author(s) 2023
SPARQLGEN: One-Shot Prompt-based Approach for SPARQL Query Generation
Liubov Kovriguina1'2'*, Roman Teucher2, Daniil Radyush' and Dmitry Mouromtsev4
!metaphacts GmbH, Daimlerstrafie 36, 69190, Walldorf, Germany
*Fraunhofer IAIS Dresden, Schloss Birlinghoven 1, 53757, Sankt Augustin, Germany
3ITMO University, Kronverksky Pr. 49, bldg. A, St. Petersburg, 197101, Russia
4UB - Leibnlz-InformationszenLrum Technik und Naturwissenschuften und Universitutsbibliuthek, Welfengarlen IB, 30167 Hannover, Germany
Abstract
In this work, we present a one-shot generative approach (further referred to as SPARQLGEN) for generating SPARQL queries by augmenting Large Language Models (LLMs) with the relevant context within a single prompt. The prompt includes heterogeneous data sources: a question itself, an RDF subgraph required lo answer Ihe question, and an example of a correct SPARQL query for a diflerenl question. In the experiments, GPT-3, a popular pre-trained language model from OpenAL was leveraged, but it. is passible Lo extend the approach to any other generative LLM. We evaluate, how different types of conlext in ihe prompt influence the query generation performance on QALD-9, QALD-10 and Besliary dataset (BESTIARY), which was created to test LLM performance on unseen data, and provide a detailed error analysis. One of the findings is that providing the model with the underlying KG and a random correct query improve the generation results. The approach shows strong results on QALD-9 dataset, but doesn't generalize on QALD-10 and BESTIARY, which can be caused by memorization problem.
Keywords
Knowledge Graphs Question Answering, SPARQL query generation, Augmented Large Language Models, Prompt Template Design
1. Introduction
In the current paper, we propose a one-shot approach for generating SPARQL queries with prompting LLMs, further referred as SPARQLGEN. Our approach lies in augmenting LLMs [l] with a knowledge graph fragment, required to construct the query, and a question-subgraph-query example, randomly sampled from the training set. Assembling all the context, required to generate a query, in a single prompt, is performed via loosely coupled heterogeneous structured information snippets, further referred as prompt elements. The prompt element is represented as a structure, having description and source and a set of pre-processing methods (i.e. for sampling, serializing to string, ranking, linearizing), that are specific to the prompt element and allow to combine heterogeneous data sources within a single prompt (Fig. 1). This allows to quickly and flexibly build custom prompt templates with an arbitrary order and number of elements.
SEMANTICS 2023 EU: 19th International Conference on Semantic Systems, September 20-22, 2023, Leipzig, Germany 'Corresponding author.
lOl lk@melaphacts.com (L. Kovriguina); roman.teucher@iais.fraunhofer.de (R. Teucher); daniil.radyush@gmail.com (D. Radyush); d.muromtsev@gmail.com (D. Mouromtsev)
(tQ © © 2023 Copyright for this pape !i by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)
r—CEIJR Workshop Proceedings (CEUR-WS.org)
The experiments, implemented by now, pursue two main goals: (i) leverage LLiM for SPARQL query generation in one-shot / zero-shot setup without fine-tuning, (ii) estimate the impact of different prompt elements on the query generation task and how the model attends to them:
1.e., does the model focus on reasoning over the provided contextual information, rather than retrieving what it already knows, at the inference step. Since memorization is a known problem of LLMs [2], we created an extra dataset on a fantasy bestiary topic, following the QALD format, to evaluate the method on the data, which could not be seen during LLM training1.
Main contributions of the paper are the following:
• SPARQLGEN, a one-shot method for SPARQL query generation with prompting LLMs, that can be quickly adapted to other datasets and knowledge graphs,
• unseen BESTIARY dataset, that can be particularly useful for evaluating reasoning capabilities of LLMs,
• detailed error analysis of LLM generated queries across QALD-9, QALD-10 and BESTIARY datasets.
2. Related Work
There are several recent approaches for generating SPARQL queries from natural language queries based on neural architectures. Soru et al. [3] designed a sequence-to-sequence system that utilizes bi-directional LSTM for generating SPARQL templates. Using a rule-based approach, the SPARQL query is then created from the generated templates. However, this translational approach cannot handle oul-of-vocabulary tokens. Rony ct al. [4] propose SGPT, an approach using a stack of Transformer encoders to embed linguistic features from natural language questions, as well as entity and relation information, to the GPT-2 model. While entities and relations representations are fed to the model in SGPT, providing their connections in the underlying KG is missing. Thus, generating correct triple sequences in the final SPARQL queries is error prone due to unknown graph structures. Another strong approach was introduced in [5], where authors improve on the state of the art KGQA2 by train the T5 model to generate skeleton SPARQL gueries and truncated KG embeddings, that are used to fetch candidate entities for the skeleton query. However, all these approaches assume training or fine-tuning an existing model, whereas SPARQLGEN doesn't require any training.
3. Datasets
QALD-9 [6] is a small yet challenging multilingual question answering dataset based on DBpedia. The dataset contains 150 questions in 3 to 8 different languages, for our experiments the test data in English were used. QALD-10 is a multilingual question answering datasel based on Wikidata. It contains 394 samples with questions in English, Chinese, Russian and German. The dataset is more complex than QALD-9, for instance, by incorporating property paths instead
The augmented data used for prompting, the BESTIARY dataset, as well as supplementary material, are uploaded to the repository:https://github.com/danrd/sparqlgen
2https://github.com/KGQA/leaderboard
of only using single relations. BESTIARY dataset consists of 100 manually created queries related to a custom Bestiary knowledge graph. This graph contains diverse information about creatures from the Dungeons & Dragons fantasy role-playing game. The graph and dataset description are presented in the supplementary material
4. Architecture Description
SPARQLGEN is implemented as a modular architecture with the following business logic: (1) retrieving context to populate prompt elements, (2) composing and executing the prompt, (3) removing hallucinations and validating the query. At the preprocessing step, each datapoint in the QALD-formatted dataset was augmented with the knowledge, represented as the minimal subgraph, required to execute the query. To combine heterogeneous data sources in the prompt, we designed an abstract structure, called prompt element, that is instantiated during the experiment. Each prompt element has fields description and source, as well as methods for pre-processing the source data (see prompt elements Example, Instruction, Question and Knowt-edgeGraph in Fig. 1). A prompt in SPARQLGEN is a serialized sequence of prompt elements. The implemented structure allows to configure experiments with minimal changes in the code structure and quickly design custom prompt templates.
For one-shot prompting we created a set of guiding examples, which includes 20 question-snbgraph-query samples, selected from QALD-9 training set and representing different query types (ASK or SELECT), patterns (number of hops), combinations of modifiers (FILTER, ORDER BY, ctc.). Guiding examples repeat the structure of prompt, but already provide the corrcct answers (one-shot prompting).
The architecture of the SPARQLGEN pipeline is shown in Fig. 1, Firstly, for each datapoint in QALD format, already augmented with a subgraph, a guiding example is randomly selected (prompt element Example). Then the prompt builder constructs a prompt from the provided datapoint and a guiding example with the order of the prompt elements, defined in the experiment config, as shown in Fig. 1), and the serialized prompt is sent to the GPT-3 Completions endpoint. The resulting query is validated and evaluated. Validation includes removing hallucinated symbols (i.e. generated text prior the query, like System:, Query:, randomly inserted newlines, etc.)
To investigate whether adding subgraph information to the model improves the performance, we enriched each sample in the test sets of QALD-9, QALD-10 and BESTIARY with a subgraph of the source RDF graph of that dataset. These subgraphs contain all the triples that are required to answer the question correctly, but no irrelevant triples. First, we strip the ground truth SPARQL queries from any modifiers leaving a simple SELECT * query. By doing that, we get all possible bindings for the variables in the ground truth query. Second, we extract the individual triples from the query. They contain entities, relations but also still the variables. Third, we take ihe result bindings from the SELECT * query and replace the variables in the extracted triples. The result is a set of triples, which is based on the original ground truth query, representing the subgraph that is sufficient to answer the given query. Further, this subgraph is considerably smaller than taking all triples from all the given entities and relations of the ground truth query (see process diagram in supplementary material).
Table 1
SPARQL query generation accuracy with different prompt configurations for on QALD-93
Prompt Configuration Total Executed Samples not Accuracy
aug- correctly processed
mented due to limited
samples input
Instruction + Question 113 14 0 0.1239
Instruction + Question + Subgraph 113 28 26 0.3218
Instruction + Question + Subgraph + 113 38 30 0.4578
Guiding example
Table 1 shows, that providing the subgraph to the model increases performance. This suggests that the model can make use of ihe information and structure of ihe graph at inference step. One-shot prompting increases the performance as well.
6. Evaluation and Results
We used F1 -macro to evaluate the whole query generation as suggested in GERBIL benchmarking system [7]. Results are reported in Table 2. On QALD-9 SPARQI.GEN reaches 67.07, that matches the performance of pre-trained SGPT system (67.82), which is currently at the top of the leaderboard [8]. However, the approach doesn't generalize well on the recent QALD-10 and unseen BESTIARY.
Table 2
SPARQLGEN performance across datasets
Dataset Description
Fl-macro Knowledge graph Prompt configuration
QALD-9 67.07 DBPedia Example + Instruction + Question + Subgraph
QALD-10 28.75 Wikidata Example + Instruction + Question + Subgraph
BESTIARY 15.01 Bestiary graph Example + Instruction + Question + Subgraph
Objective evaluation with QALD Fl-macro [7] has shown quite diverse results across the 3 datasets. Therefore, we tried to categorize the errors in a small set of categories, that cover all types of errors witnessed during LLM inference. Error classification and distribution is presented in the supplementary materials.
The experimental results show that the model struggles to deal with an unknown knowledge graph. It is safe to assume that GPT-3 has encountered DBpedia as well as Wikidata information
in column 4); 3) for experiments in table 1 we didn't fix hallucinations. When running the experiments on all 3 datasets (Table 2), we addressed these errors: 1) for QALD-9, from 37 samples with non-executable queries we managed to fix 20 queries; 2) also, we added sampling of triples from the subgraph, so that the resulting prompt could never exceed the token limit. Following the evaluation guidelines, suggested in GF.RRIL [7], if both reference and generated query return empty set, QALD-F1 is set to 1,
in pre-training. This is visible in the overall better performance on the corresponding datasets, as well as in the higher number of knowledge related errors in BESTIARY. Namespace errors, incomplete triples and ignoring KG structure errors occur more frequently for BESTIARY. In general, it seems beneficial to introduce the model to the KG in pre-training already.
7. Conclusion and Future Work
The performance of the one-shot SPARQLGEN approach on QALD-9 is only slightly below the SGPT approach, that requires additional training of the stack of Transformer-encoders to leverage the pre-trained GPT-2 model, and significantly outperforms the rest of the QALD-9 leaderboard. Assuming that SPARQLGEN can be adapted to a new dataset and KG faster than approaches, requiring fine-tuning, continuing experimenting with prompt-based approaches in KGQA can be definitely a winning strategy. However, this approach doesn't generalize perfectly, given the evidence from QAI.D-10 and BESTIARY.
Memorization can be one explanation: QALD-10 is a recent dataset, which OpenAI models might not seen, and BESTIARY dataset have never been used for training any model. BESTIARY queries are based on the knowledge graph, that was designed specially for structuring the Dungeons & Dragons domain, and this graph has been never made available to the GPT-3. The availability of the unstructured data about Dungeons & Dragons on the Web doesn't imply, that the LLM can synthesize the knowledge graph, following by generating the adequate query. Given that, we assume that a niche domain (for knowledge engineering) combined with an unseen knowledge graph makes the model rely only on its reasoning skills, when querying BESTIARY. The existence of these skills is questionable at its own: parallel research lines prove that LLMs can and can not reason at full scale. However, during query generation against the BESTIARY KG, GPT-3 copies entities and relations from the provided BESTIARY subgraph, without large hallucinations from DBPedia and Wikidata, indicating that providing a subgraph of domain-specific KGs can improve SPARQI. query generation.
Future work is manifold. First of all, we would like to proceed with fine-tuning open source LLMs on available train sets of QALD / LC-QuAD series and evaluate the generalization of the fine-tuned model 011 more benchmarks. Another direction can be embedding the current SPARQLGEN approach in a reinforcement learning pipeline.
References
[1] G. Mialon, R. Dessi, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Roziere, T. Schick, J. Dwivcdi-Yu, A. Celikyilmaz, et al., Augmented language models: a survey, arXiv preprint arXiv:2302.07842 (2023).
[2] N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramcr, C. Zhang, Quantifying memorization across neural language models, 2023. arXiv: 2202. 07646.
[3] T. Soru, E. Marx, A. Valdestilhas, D. Esteves, D. Moussallem, G. Publio, Neural machine translation for query construction and composition, 2018, arxiv: 1806 .10478,
[4] M. R. A. H. Rony, U. Kumar, R. Teucher, L. Kovriguina, J. Lehmann, Sgpt: A generative
approach for sparql query generation from natural language questions, IEEE Access 10 (2022) 70712-70723. doi:10. 1109/ACCESS. 2022 .3188714.
[5] D. Banerjee, P. A. Nair, J. N. Kaur, R. Usbeck, C. Biemann, Modern baselines for sparql semantic parsing, in: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2022, pp. 2260-2265.
[6] N. Ngomo, 9th challenge on question answering over linked data (qald-9), language 7 (2018) 58-64.
[7] R. Usbeck, M. Roder, M, Hoffmann, F. Conrads, J, Huthmann, A.-C. Ngonga-Ngomo, C. Demmler, C. Unger, Benchmarking question answering systems, Semantic Web 10 (2019) 293-304.
[8] A. Perevalov, X. Yan, L. Kovriguina, L.Jiang, A. Both, R. Usbeck, Knowledge graph question answering leadcrboard: A community resource to prevent a replication crisis, in: Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2022, pp. 2998-3007.
Improving Subgraph Extraction Algorithms for One-Shot SPARQL Query Generation with Large Language Models
Dmitrii Pliukhin''*, Daniil Radyush', Liubov Kovriguina2'* and Dmitry Mouromtsev3
'ITMO University, Kronverksky Pr. 49, hldg. A, St. Petersburg, 197)01, Russia 2Independent Reseurcher, Dresden, Germany
'IIB - Leibniz-Informationszentrum Technik und Naturwissenschaften und Universitätsbibliothek, Weifengarten 1B, 30167 Hannover, Germany
Abstract
Question answering over scholarly knowledge graphs involves many challenges: complex graph patterns, long-tail distributed data, revision and evolution of the scholarly ontologies, and knowledge graphs incompleteness due to constant research dynamics. In this work, we present an LLM-based approach for SPARQL query generation over Open Research Knowledge Graph (ORKG) for the ISWC SciQA Challenge. Our approach proposes a couple of improvements to the recently published SPARQLGEN approach, that performs one-shot SPARQL query generation by augmenting Large Language Models (LLMs) with the relevant context within a single prompt. Similar to SPARQLGEN, we include heterogeneous data sources in the SPARQI, generation prompt: a question itself, an RDF subgraph required to answer the question, and an example of a correct SPARQL query. In the current work, we focused on designing subgraph extraction algorithms, that are close to real-life scenarios of generative KGQA, and replaced the random choice of example question-query pair with similarity scoring.
Keywords
Scholarly Knowledge Graphs, Knowledge Graphs Question Answering, SPARQL query generation, Augmented Large Language Models, Subgraph Extraction
1. Introduction
Scholarly knowledge graphs has become a recent trend and inspired the adaptation of KGQA systems to new complex domains, that are constantly evolving, bringing new facts and concepts to the knowledge graph. Currently, there is a number of scholarly knowledge graphs, that differ in metadata and coverage, being built on top of various ontologies and data sources [1, 2, 3]. For knowledge graph question answering (KGQA) systems, such landscape creates a lot of challenges due to variative graph patterns, ambiguity, and complex user questions. Previous KGQA benchmarks (i.e. QALD series and LC-QuAD datasets) were build upon DBpedia and Wikidata and allowed cumulative improvements of KGQA systems, especially for template-based approaches. With the diversity of scholarly KGs, template-based and pre-traincd approaches may not work as good as before due to adaptation costs. We suggest to employ generative approaches
ISWC 2023: Scholarly QALD Challenge, November 6-10, 2023, Athens, Greece
(O) zeionara@gmail.com (D. Pliukhin); daniil.radyush@gmail.com (D. Radyush); lkovriguina@gmail.com (L. Kovriguina); d.muromtscv@gmail.com (D. Mouromtsev)
^^^^wHJ © .HL' Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)
ry^HH CEUR Workshop Proceedings (CEUR-WS.org)
to KGQA, namely within the augmented large language models paradigm, when the LLM gets all the context required to generate the SPARQL query, in a prompt. LLMs augmenting approaches try to mitigate the fundamental defect of LLMs: pure statistical language modeling over limited context size, by providing LLMs information from relevant data sources. Several strategies are known to do the augmentation: (i) retrieval-augmented language models [4], knowledge injection [5], reasoning [6], etc. According to the classification, proposed in [7], providing LLMs with extra context via prompting belongs to augmenting with eliciting reasoning.
In the current paper, we propose further improvements to SPARQLGEN, a recent one-shot approach for generating SPARQL queries with prompting LLMs [8]. SPARQLGEN approach lies in augmenting LLMs with a knowledge graph fragment, required to construct the query, and a question-subgraph-query example within a single prompt (see Sec.2). This knowledge graph fragment, further referred as subgraph, contains all the triples that are required to build the correct SPARQL query (that is, to answer the question correctly), but no irrelevant triples, and extracting this subgraph is quite challenging. Since the motivation behind SPARQLGEN was to evaluate, whether an LLM can make use of the provided subgraph and infer graph patterns during SPARQL query generation, the algorithm of subgraph extraction is based on target SPARQL queries of QALD-9 and QALD-10 datasets1. The disadvantage of this approach is that it will not work on inference, since it requires the ground truth data. For the SciQA challenge, we have designed a subgraph extraction algorithm, combining similarity search with Hierarchical Navigable Small World (HNSW) method, that can be used as a replacement of the subgraph extraction algorithm in SPARQLGEN and which matches better the real-life KGQA scenarios (see Sec.5). The random selection of the guiding example for one-shot prompting is replaced with selection based on the Levenstein distance.
Main contributions of the paper are the following:
• improvement of SPARQLGEN, a one-shot method for SPARQL query generation with prompting LLMs, that can be quickly adapted to other datasets and knowledge graphs,
• subgraph extraction algorithm, that can be used in LLM-augmented SPARQL query generation scenarios,
• evaluation of the improved SPARQLGEN method on the ORKG benchmark.
2. Related Work
Fine-tuning LLMs to generate SPARQL queries has already a proven track record with top-leaderboard results on QALD-9 with SGPT [9] and LC-QuAD with GETT-QA [10]. Besides that, there are also no-SPARQL KGQA approaches, allowing to load the knowledge graph into the question answering LLM pipeline in a retrieval fashion, i.e. see2, but we excluded them for now despite considering promising.
Rony et al. [9| propose SGPT, an approach using a stack of Transformer encoders to embed linguistic features from natural language questions, as well as entity and relation information, to the GPT-2 model. While entities and relations representations are fed to the model in SGPT,
'See algorithm description Ln the repository: https://github.com/danrd/sparqlgen 3https://github.Com/mommi84/rdf-qa
providing their connections in the underlying KG is missing. Thus, generating correct triple sequences in the final SPARQL queries is error prone due to unknown graph structures. Another strong approach was introduced in [10], where authors improve on the state of the art KGQA1 by training the T5 model to generate skeleton SPARQL queries and truncated KG embeddings, that are used to fetch candidate entities for the skeleton query. However, all these approaches assume training or fine-tuning an existing model.
The recent approach, which we are extending upon, is SPARQLGEN [8], that doesn't require any training and instructs LLMs to generate SPARQL queries by providing them the underlying knowledge graph, guiding examples and information about SPARQL grammar within a single prompt. Assembling all the context, required to generate a query, in a single prompt, is performed via loosely coupled heterogeneous structured information snippets, further referred as prompt elements. The prompt element is represented as a structure, having description and source and a set of pre-processing methods (i.e. for sampling, serializing to string, ranking, linearizing), that are specific to the prompt element and allow to combine heterogeneous data sources within a single prompt. The description field depicts the source data, i.e. "The RDF knowledge graph", and the source contains the data itself, i.e. triples. This altogether allows to quickly and flexibly build custom prompL templates with an arbitrary order and number of elements.
The findings of SPARQLGEN approach show that the model struggles to deal with an unknown knowledge graph. Namespace errors, incomplete triples and ignoring KG structure errors occur more frequently for the unseen dataset and it seems beneficial to introduce the model to the KG in pre-training already (see supplementary material in in the repository: htlps: //github.com/danrd/sparqlgen).
3. SciQA Challenge and Datasets
This section presents an overview of the SciQA Challenge and the datasets used in this study. 3.1. SciQA Challenge
The primary objective of this research endeavor was to participate in the Scholarly Question-Answering over Linked Data (Scholarly QAT.D) challenge. This challenge is centered around Knowledge Graph Question Answering utilizing the ORKG Scholarly Knowledge Graph. The challenge was structured into two distinct stages.
The first stage aimed to develop a model and conduct validation using a held-out subset with known true labels, thereby establishing an initial leaderboard. The second stage focused on assessing the model's quality and generating final results for comparison among the various solutions developed. The central challenge task involved transforming a user's natural language question into a SPARQL query and executing this query on the provided knowledge graph to furnish a response. The efficacy of the proposed solutions was evaluated through a set of 200 questions, employing the Fl-score metric in both stages.
3https://github.com/KGQA/Ieaderboard
3.2. Datasets
To train and evaluate the developed models, the SciQA dataset was designated as the target data source. This dataset is publicly available in the repository: https://zenodo.org/record/7744048 . The SciQA dataset encompasses a rich assortment of records. Each record comprises a question expressed in na tural language, the corresponding SPARQL query, and a list of knowledge graph triples that represent the answer to that question.
Notably, the SciQA dataset is characterized by a significant diversity of questions, spanning various domains, response types, and sizes. The questions also exhibit structural and computational complexity and a degree of ambiguity. In addition to the SciQA dataset, the repository contains a link to the dump of the knowledge graph designed for question answering. This knowledge graph is primarily structured around scholarly data, including papers, contributions, and connections between them, accompanied by pertinent metadata.
The knowledge graph contains a total of 1,133,217 triples and 21,243 contributions, with an overall dump size of 152 megabytes. This knowledge graph is inherently heterogeneous, encompassing diverse data pertaining to papers from an array of research fields. Its structure is notably intricate, rendering the challenge particularly demanding. Specifically, the graph incorporates both discrete and continuous data entries, encompassing a wide range of primitive types such as numbers in various formats, dates, strings, and binary labels.
4. Architecture Description
Our approach is implemented as a modular architecture with the following components: (1) subgraph retrieval (2) example selection context to populate prompt elements, (3) prompt building, (4) prompt execution, (5) removing hallucinations and validating the query, (6) query execution.
During steps (1) and (2), each data point in the SciQA test set was augmented with the context, represented as (i) the subgraph, required to execute the query (see Sec. 5), and (ii) a question-query pair, similar to the question (see below). Then the prompt builder (3) constructs a prompt from the augmented SciQA datapoint and a guiding example with the order of the prompt elements, defined in the experiment config, as shown in Fig, 1), and the serialized prompt is sent to the GPT-3.5 completions endpoint. The resulting query is validated (5) and executed (6). Validation includes removing hallucinated symbols (i.e. generated text prior the query, like System:, Query:, randomly inserted newlines, etc.) The pipeline architecture is shown in Fig. 1. Subgraph extraction is described in more detail in Sec. 5.
Following the SPARQLGEN approach, to combine heterogeneous data sources in the prompt, we used an abstract structure, called prompt element, that is instantiated during the experiment. Each prompt element has fields description and source, as well as methods for pre-processing the source data (see prompt elements Example, Instruction, Question and Knowledge Graph in Fig. 1). A prompt in SPARQLGEN is a serialized sequence of prompt elements. The implemented structure allows to configure experiments with minimal changes in the code structure and quickly design custom prompt templates.
Example Selection For one-shot prompting, the example was sampled from the train subset of the proposed dataset by computing Levenstein distance between input question and every
№ Hyperparameters Steps Queries patterns
1 n_prop - list of n properties n_paper - list of n papers n_cont - list of n contributions 1) For each question's bigram retrieve relevant n_prop, n_paper and n_cont by indexes; 2) Form n_prop, ri_paperand n_cont for the question keeping most relevant objects; 3) Extract subgraphs for n_paper and n_cont from ORKG; 4) Retain only triples with properties from n_prop. {<paper> ?x ?y ?y <property> <label>} {<contribution> <property> <label>}
2 n_prop - list of n properties n_res - list of n resources 1) For each question's bigram retrieve relevant n_prop and n_res by indexes; 2) Form n_prop for the question keeping most relevant properties; 3) Discard resources with cosine similarity to the question less than 0.6; 4) Extract subgraphs for n_res from ORKG: 5) Retain only triples with properties from n_prop. {<resource> ?x ?y ?y <property> <label>} {?x ?y <resource> <resource> <property> <label>}
Figure 2: Subgraphs construction approaches.
object labels leveraging initially constructed with HNSW indexes. (3) subgraphs retrieval: deriving 2-hop paths containing given similar objects. (4) subgraphs merging and converting into string with postprocessing for prompting.
In our experiments we implemented two approaches for involving derived triples in subgraph construction process as part of step (3). The first approach considers papers and contributions as a main source of structured information for LLMs. Therefore, it relies on retrieving a number of papers and contributions with related titles from ORKG, excluding triples with irrelevant predicates. However, in some cases titles are not specific enough and do not contain needed keywords. To address this issue, the second approach is to directly extract triples containing resources and properties with high similarity to the question, though in some eases it requires particular triple patterns. More detailed description of the implemented approaches is provided with Figure 2, whereas the whole process of subgraph retrieval in general is presented in Fig. 3:
6. Experiments and Results
During the experiments, we evaluated, how the subgraph extraction algorithm influences the SPARQL generation quality. In the baseline version, no subgraph was provided, only an example. The results are summarized in Table 1.
According to the experimental results, even without involving subgraph extraction algorithms the proposed architecture is capable to obtain relatively decent result with 0.922 F1 score. Moreover, the table shows that the second approach to subgraph extraction leads to minor F1 score decreasing. This probably means that structured information, extracted from ORKG, predominantly contributes additional noise to the prompts. Consequently, this approach needs further revision or adaptation to the given KG. On the other hand, the first approach fosters slight F1 metric increase demonstrating its usefulness in general that requires investigation in future works.
Table 2
Ablation study
Subgraph approach Hyperparameters Fl score
№1 n_prop=100, n_paper=25, n_cont=25 0.935
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.