Разработка методов автоматического определения тональности предложений на русском языке с использованием семантических правил и синтаксической структуры тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Полетаев Анатолий Юрьевич
- Специальность ВАК РФ00.00.00
- Количество страниц 434
Оглавление диссертации кандидат наук Полетаев Анатолий Юрьевич
Реферат
Synopsis
Введение
ГЛАВА 1. Обзор существующих работ
1.1 Определение общей тональности предложения
1.1.1 Основные работы для предложений на английском языке
1.1.2 Основные работы для предложений на русском языке
1.1.3 Метрики качества, достижимые с помощью методов машинного обучения
1.2 Определение тональности по отношению к именованной сущности
1.3 Определение тональности по отношению к аспекту социально-экономического развития
1.4 Построение деревьев синтаксических единиц предложений
на русском языке
1.5 Выводы
ГЛАВА 2. Использованные корпуса
2.1 Корпус OpenSentimentCorpus
2.2 Корпус отзывов на отели
2.3 Корпус CABSAR
2.4 Корпус предвыборной агитации
2.5 Выводы
ГЛАВА 3. Адаптированный метод определения общей
тональности предложения
3.1 Исходный метод для английского языка
3.2 Адаптация исходного метода к русскому языку
3.2.1 и И10: правила для обработки отрицаний
3.2.2 ИЗ, И6 и И7: правила объединения
3.2.3 К11-И15: правила для обработки сложных предложений
3.3 Эксперименты с адаптированным методом
3.3.1 Результаты на четверти корпуса отзывов на отели
3.3.2 Анализ ошибок
3.3.3 Доработка метода
3.3.4 Итоговые результаты
3.4 Выводы
ГЛАВА 4. Алгоритм построения дерева синтаксических единиц предложения на русском языке по дереву
синтаксических связей
4.1 Описание алгоритма
4.2 Эксперименты
4.2.1 Корпус
4.2.2 Метрики
4.2.3 Результаты
4.3 Выводы
ГЛАВА 5. Новый метод определения общей тональности
предложения
5.1 Общая схема метода
5.2 Словарь тональных слов
5.3 Семантические правила
5.3.1 Правила для фраз, тональность которых является результатом соединения тональности их частей
5.3.2 Правило для фраз с отрицаниями
5.3.3 Правила для фраз, содержащих специальные языковые средства выражения тональности
5.4 Эксперименты
5.4.1 Метрики качества
5.4.2 Анализ ошибок
5.4.3 Обсуждение результатов
5.5 Выводы
ГЛАВА 6. Метод определения тональности по отношению
к именованной сущности
6.1 Эксперименты с нейросетевыми классификаторами
6.1.1 Методика постановки экспериментов
6.1.2 Оценка качества существующих классификаторов
6.1.3 Эксперимент с аугментацией данных
6.1.4 Эксперимент с IAN, использующим контекст упоминания именованной сущности
6.2 Метод определения тональности по отношению к именованной сущности, основанный на семантических правилах
6.2.1 Связь между тональностью по отношению
к именованной сущности и общей тональностью
предложения
6.2.2 Общая схема метода
6.2.3 Семантические правила
6.2.4 Оценка качества
6.3 Ансамбль BERT-SPC и метода, основанного на семантических правилах
6.4 Выводы
ГЛАВА 7. Метод определения тональности по отношению
к аспекту социально-экономического развития
1.1 Нейросетевые классификаторы
1.1.1 Классификатор, обучаемый только на корпусе предвыборной агитации
1.1.2 Классификатор, дообучаемый на корпусе предвыборной агитации
1.1.3 Дополнительные эксперименты
7.2 Метод, основанный на семантических правилах
7.2.1 Общая схема метода
7.2.2 Метод поиска аспектных терминов на основе перечня терминов
7.2.3 Метод поиска аспектных терминов на основе семантической схожести
7.2.4 Гибридный метод поиска аспектных терминов
7.2.5 Дополнительные семантические правила
7.3 Оценка качества метода, основанного на семантических правилах
7.3.1 Метрики качества
7.3.2 Анализ ошибок
7.4 Ансамблевый классификатор
7.5 Выводы
Заключение
Список литературы
Список рисунков
Список таблиц
Приложение А. Выдержка из руководства для разметчиков,
использовавшегося при создании OpenSentimenrCorpus
Приложение Б. Формальная грамматика синтаксических единиц предложений на русском языке и критерии
для синтаксических единиц
7.6 Типы синтаксических единиц
7.6.1 Общая структура предложения
7.6.2 Группа подлежащего
7.6.3 Группа сказуемого
7.6.4 Группа дополнения
7.6.5 Группа определения
7.6.6 Группа обстоятельства
7.6.7 Критерии для элементов предложения в целом и его грамматической основы
7.7 Критерии для элементов групп подлежащего, сказуемого, определения, дополнения и обстоятельства
Приложение В. Список изменений, внесённых в «РуСентиЛекс-
2017»
Приложение Г. Семантические правила, используемые методами
определения тональности
7.8 Описание семантических правил, используемых
для определения общей тональности
7.8.1 Правило для фраз с отрицанием
7.8.2 Правила для фраз, содержащих специальные языковые средства выражения тональности
7.9 Описание семантических правил, используемых
для определения тональности по отношению к именованной сущности
7.10 Описание семантических правил, используемых
при определении тональности по отношению к аспекту социально-экономического развития
7.11 Псевдокод реализации некоторых семантических правил
Приложение Д. Описания групп ошибок, допущенных
созданными методами
7.12 Группы ошибок, допущенных методом определения общей тональности предложения
7.13 Группы ошибок, допущенных методом определения общей тональности предложения
Приложение Е. Справка о внедрении
Приложение Ж. Документы, подтверждающие результаты
интеллектуальной деятельности
Приложение З. Тексты публикаций автора по теме диссертации
11
Реферат
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Методы извлечения и резюмирования критических отзывов пользователей о продукции2016 год, кандидат наук Тутубалина Елена Викторовна
Модели, методы и программные средства извлечения оценочных отношений на основе фреймовой базы знаний2022 год, кандидат наук Русначенко Николай Леонидович
Разработка гибридного алгоритма распознавания именованных сущностей в узбекском языке2025 год, кандидат наук Менглиев Давлатёр Бахтиярович
Методы и алгоритмы распознавания и связывания сущностей для построения систем автоматического извлечения информации из научных текстов2022 год, кандидат наук Бручес Елена Павловна
Методы и алгоритмы аспектного анализа тональности на основе гибридной семантико-статистической модели естественного языка2022 год, кандидат наук Корней Алена Олеговна
Введение диссертации (часть автореферата) на тему «Разработка методов автоматического определения тональности предложений на русском языке с использованием семантических правил и синтаксической структуры»
Общая характеристика работы
Актуальность темы. Автоматическое определение тональности — одна из важных и популярных задач компьютерной лингвистики, заключающаяся в определении авторского отношения [105]. Определяться может как тональность по отношению к какому-либо объекту или аспекту объекта, так и общая тональность, т. е. отношение автора к теме текста [43; 45; 52; 54; 93; 3].
Данная работа посвящена автоматическому определению тональности отдельных предложений с использованием семантических правил и синтаксической структуры. В ходе автоматического определения тональности конкретному предложению сопоставляется один из классов тональности, т. е. решается задача классификации. В данном исследовании используются три класса тональности: положительный, отрицательный и нейтральный. В исследовании рассматриваются методы решения трёх задач, выделяемых в рамках задачи автоматического определения тональности: задачи определения общей тональности предложения, задачи определения тональности предложения по отношению к именованной сущности и задачи определения тональности предложения по отношению к аспекту социально-экономического развития.
Подавляющее большинство методов определения тональности предложений используют либо машинное обучение, либо семантические правила, определяющие, какую тональность имеет предложение или его часть в зависимости от их синтаксической структуры, семантики входящих в них слов, а также другой информации, которая может влиять на тональность [66; 100; 2]. Методы, основанные на семантических правилах, немногочисленны, и большая их часть как для английского, так и для русского языка была создана достаточно давно, в период до 2018 года [45; 46; 61]. Снижение интереса к методам, основанным на семантических правилах, связано с тем, что правила составляются вручную на основе знаний экспертов-лингвистов о средствах выражения тональности, используемых в языке, что приводит к увеличению требуемого на создание
метода времени. В то же время, активно развивавшиеся в 2010-х годах нейросе-тевые классификаторы, например, использующие архитектуры LSTM и BERT, не требуют затрат времени на работу с экспертами-лингвистами — для их обучения необходимы только размеченные по тональности корпуса предложений.
Тем не менее у нейросетевых классификаторов есть существенные недостатки. Один из них — серьёзные требования к объёму корпуса, на котором они обучаются: для достижения высокого качества могут требоваться тысячи или даже десятки тысяч предложений, что серьёзно повышает стоимость создания таких классификаторов [35]. Другой недостаток — зависимость качества определения тональности предложения от его схожести с примерами, на которых обучался классификатор, что также снижает экономическую эффективность нейросетевых классификаторов [72]. Кроме того, поскольку нейросетевые классификаторы представляют собой «чёрный ящик», для них серьёзно затруднён анализ ошибок.
Несмотря на относительно высокую наукоёмкость создания методов, основанных на семантических правилах, они лишены упомянутых недостатков: их не требуется обучать (следовательно, не требуются большие размеченные корпуса), а результаты их работы являются интерпретируемыми, что упрощает анализ ошибок. Кроме того, поскольку семантические правила отражают реально существующие в языке закономерности, единожды составленный набор правил можно использовать для определения тональности предложений из различных предметных областей. За счёт этого на некоторых наборах данных даже достаточно простые методы, основанные на семантических правилах, определяют тональность более качественно, чем нейросетевые классификаторы [85].
Нужно отметить, что методы, основанные на семантических правилах, могут использоваться не только самостоятельно, но и в составе более сложных методов, объединяющих семантические правила и методы машинного обучения [38]. Такие классификаторы могут показывать более высокое качество, чем отдельные методы, использованные для их создания.
Необходимо отметить, что в настоящий момент большая часть методов определения тональности разработана для предложений, извлечённых из записей в социальных сетях или интернет-отзывов на товары и услуги, относящихся преимущественно к разговорному стилю речи [6; 7]. Представляется важным уделить внимание определению тональности предложений из публицистических текстов, которые могут быть полезны для создания различных автоматизированных систем, например, систем мониторинга общественного мнения.
Поскольку предложения публицистического стиля, как правило, являются сложными [34], для их обработки может быть полезным учитывать синтаксическую структуру. В работе [104] показано, что для длинных предложений на английском языке представление синтаксической структуры в виде дерева синтаксических единиц позволяет определять тональность более точно, чем представление в виде дерева синтаксических связей. Разумно предположить, что для предложений на русском языке лучшее качество также может быть получено с использованием деревьев синтаксических единиц.
Несмотря на то, что для русского языка в настоящий момент недоступны открытые построители деревьев синтаксических единиц, доступно несколько построителей деревьев синтаксических связей, что приводит к появлению в исследовании дополнительной задачи — разработке алгоритма построения дерева синтаксических единиц предложения по его дереву синтаксических связей.
Таким образом, поскольку с помощью семантических правил можно преодолеть некоторые недостатки классификаторов, основанных на машинном обучении, а использование синтаксической структуры позволяет создавать более точные семантические правила, можно предполагать, что методы, использующие семантические правила и синтаксическую структуру, позволят повысить качество определения тональности предложений на русском языке.
Целью диссертационной работы является повышение качества определения тональности предложений публицистического стиля на русском языке.
Как основной показатель качества определения тональности в рамках данной работы используется макро-Б-мера.
Для достижения поставленной цели необходимо решить следующие задачи:
1. Создание размеченных корпусов предложений из публицистических текстов для оценки качества методов определения тональности (при отсутствии требуемых корпусов в открытом доступе).
2. Разработка алгоритма построения дерева синтаксических единиц предложения на русском языке по его дереву синтаксических связей.
3. Разработка метода определения общей тональности предложения на русском языке, использующего семантические правила над деревом синтаксических единиц.
4. Разработка метода определения тональности предложения на русском языке по отношению к именованной сущности, использующего семантические правила над деревом синтаксических единиц.
5. Разработка метода определения тональности предложения на русском языке по отношению к аспекту социально-экономического развития, использующего семантические правила над деревом синтаксических единиц.
На защиту выносятся положения, обладающие научной новизной:
1. Алгоритм построения дерева синтаксических единиц предложения на русском языке, отличающийся использованием дерева синтаксических связей, а также специально разработанной на основе справочника Д. Э. Розенталя формальной грамматики русского языка, обеспечивающий высокую точность построения деревьев синтаксических единиц предложений различной структуры, в том числе сложных и осложнённых: Б-мера построения деревьев синтаксических единиц составила 0.85.
2. Метод определения общей тональности предложения на русском языке, отличающийся использованием деревьев синтаксических единиц для алгоритмизации семантических правил, и тонального словаря, обеспечивающий более точное определение общей тональности предложений
из публицистических текстов, чем ранее существовавшие методы, использующие семантические правила. Макро-Б-мера разработанного метода составила 0.76, что на 0.16-0.18 выше, чем у ранее существовавших методов.
3. Метод определения тональности предложения на русском языке по отношению к именованной сущности, отличающийся применением классификатора, использующего семантические правила и деревья синтаксических единиц, объединяемого в ансамбль с нейросетевым классификатором BERT-SPC, обеспечивающий более точное определение тональности, чем классификатор BERT-SPC, использующийся не в составе ансамбля. Макро-Б-мера разработанного метода составила 0.81, что на 0.10 выше, чем у BERT-SPC, и на 0.11 — чем у классификатора, использующего семантические правила и деревья синтаксических единиц.
4. Метод определения тональности предложения на русском языке по отношению к аспекту социально-экономического развития, отличающийся применением классификатора, использующего семантические правила и деревья синтаксических единиц, объединяемого в ансамбль с нейросете-вым классификатором BERT-SPC, обеспечивающий более точное определение тональности, чем классификатор BERT-SPC, использующийся не в составе ансамбля. Макро-Б-мера разработанного метода составила 0.79, что на 0.05 выше, чем у BERT-SPC, и на 0.16 —чем у классификатора, использующего семантические правила и деревья синтаксических единиц.
В ходе исследования были использованы методы обработки естественного языка (NLP), машинного обучения и математической статистики. Программы для экспериментального исследования были разработаны на языке Python с использованием библиотек Stanza, Spacy, Natasha и PyMorphy (для обработки естественного языка), scikit-learn (для расчёта метрик качества) и PyTorch (для работы с нейросетевыми классификаторами).
Теоретическая значимость диссертационного исследования заключается в разработке новых методов определения тональности предложений на рус-
ском языке, использующих семантические правила и синтаксическую структуру. Исследование показало перспективность использования деревьев синтаксических единиц для создания методов определения тональности, использующих семантические правила. Было выявлено, что семантические правила, использующиеся для решения одной из задач определения тональности, могут быть переиспользованы для решения смежных задач, например, правила, созданные для метода определения общей тональности, использовались также методом определения тональности по отношению к именованной сущности.
Кроме того, было определено, что при объединении методов, использующих семантические правила, в ансамбль с наилучшими нейросетевыми классификаторами качество определения тональности может быть существенно повышено по сравнению с качеством нейросетевого классификатора, использующегося не в составе ансамбля.
Практическая значимость диссертационного исследования заключается в реализации разработанных методов в созданном комплексе программ для ЭВМ, который для заданного предложения позволяет определять:
- общую тональность;
- тональность по отношению к упоминаемой в предложении именованной сущности;
- тональность по отношению к заданному аспекту социально-экономического развития.
Предложения могут обрабатываться как в пакетном, так и в интерактивном режимах. Если у пользователя возникает необходимость более подробно проанализировать конкретное предложение, с помощью комплекса можно визуализировать фразовую структуру предложения, тональность каждой фразы и семантические правила, на основе которых эта тональность была определена.
Комплекс программ используется в деятельности ООО «ЛИНГВАХАБ» при обучении русскому языку как иностранному. Он позволяет автоматически проверять, что тональность предложений в текстах, написанных обучающимися при выполнении учебных заданий, соответствует требуемой, например, что в деловом письме нет предложений с отрицательной тональностью по от-
ношению к адресату или компании, в которой он работает. Также комплекс используется при проведении исследований в области прикладной лингвистики для автоматизации поиска в текстах предложений с положительной или отрицательной тональностью (как общей, так и по отношению к именованной сущности или к аспекту социально-экономического развития), по которым эксперты-лингвисты изучают особенности применения в языке различных средств выражения тональности. Качество определения тональности, достижимое с помощью разработанных методов — 0.76 и выше по макро-Б-мере — является достаточным для практического применения комплекса при автоматизации рутинных задач, выполняемых преподавателями и исследователями-лингвистами.
Кроме того, разработанный в ходе выполнения диссертационного исследования алгоритм построения дерева синтаксических единиц по дереву синтаксических связей может использоваться для создания разнообразных методов анализа предложений на русском языке, не обязательно связанных с задачами определения тональности. Формальная грамматика синтаксических единиц предложений на русском языке, описанная в работе, может послужить основой для создания нового построителя деревьев синтаксических единиц, не требующего для своей работы деревьев синтаксических связей.
Достоверность научных достижений подтверждается корректным использованием методов, обоснованием постановки задач, экспериментальными исследованиями на реальных массивах данных, покрывающими разработанные методы.
Соответствие паспорту специальности. Диссертация соответствует пятому пункту паспорта специальности 2.3.8 «Информатика и информационные процессы (технические науки)»: «Лингвистическое обеспечение информационных систем и процессов. Методы и средства проектирования словарей данных, словарей индексирования и поиска информации, тезаурусов и иных лексических комплексов. Методы семантического, синтаксического и прагматического анализа текстовой информации для представления в базах данных и организации интерфейсов информационных систем с пользователями». Были
разработаны, обоснованы и протестированы методы определения тональности, обеспечивающие информационным системам возможность более точного семантического анализа естественного языка, в том числе при взаимодействии с пользователями. Был также разработан алгоритм построения дерева синтаксических единиц предложения на русском языке по дереву синтаксических связей, обеспечивающий возможность более точного синтаксического анализа естественного языка.
Внедрение результатов работы. Разработанный программный комплекс был внедрён в ООО «ЛИНГВАХАБ», где он используется для проведения научных исследований в области прикладной компьютерной лингвистики, а также в составе программного продукта для обучения студентов, изучающих русский язык как иностранный.
Апробация работы. Основные результаты работы докладывались на международных научных конференциях:
1. The 30th Conference of Open Innovations Association FRUCT. Оулу, Финляндия, 2021 (онлайн). Секция: Natural Language Processing and Speech Technologies. Доклад: «Adaptation of Semantic Rule-Based Sentiment Analysis Approach for Russian Language».
2. Международный научный конгресс Университетского консорциума исследователей больших данных. Москва, Россия, 2023. Доклад: «Автоматическое моделирование текста при помощи лингвистических характеристик».
3. The 36th Conference of Open Innovations Association FRUCT. Хельсинки, Финляндия, 2024 (онлайн). Секция: Natural Language Processing. Доклад: «Automatic Detection of Sentiment Towards Explicit Aspect in Russian Publicism Sentences Using Syntactic Structure»
Поддержка работы. Диссертационная работа была выполнена при поддержке гранта РНФ №23-21-00495 «Разработка методов анализа тональности русскоязычных публицистических текстов с использованием синтаксической структуры предложений», в рамках проекта гражданской науки ЯрГУ № CS-02/2022 «Разметка корпусов текстовых данных на русском языке для при-
кладных задач компьютерной лингвистики» и проекта №П2-ГМ5-2021 по программе развития ЯрГУ в рамках программы стратегического академического лидерства «Приоритет-2030» «Внедрение технологий автоматической обработки текста на основе когнитивного анализа для междисциплинарного проектного обучения».
Личный вклад соискателя. В диссертационной работе использованы результаты, в которых соискателю принадлежит определяющая роль. Идея определения тональности предложения с помощью рекурсивного применения семантических правил к узлам синтаксического дерева, а также идея построения дерева синтаксических единиц предложения по его дереву синтаксических связей предложены соискателем самостоятельно. Задача по созданию ансамблевых классификаторов, объединяющий метод, основанный на семантических правилах, и нейросетевой классификатор ВЕИТ-БРС, была поставлена совместно с научным руководителем Парамоновым И. В. Руководства по разметке корпусов, созданных в ходе исследования, созданы соискателем также совместно с Парамоновым И. В. Формализация семантических правил, используемых созданными в ходе исследования методами, выполнена соискателем совместно с экспертом-филологом Бойчук Е. И. Все экспериментальные исследования выполнены автором самостоятельно. В публикациях, выполненных в соавторстве, вклад соавторов следующий.
- Полетаев А. Ю.: идея определения тональности с помощью рекурсивного применения семантических правил и идея использования дерева синтаксических единиц (публикация 3); идея метода определения тональности по отношению к именованной сущности, основанного на семантических правилах (публикация 6); идея метода определения тональности по отношению к аспекту социально-экономического развития, основанного на семантических правилах (публикация 7); формализация семантических правил (публикации 3, 5-7); подготовка данных, анализ источников, программная реализация методов, проведение экспериментов (все публикации); описание экспериментов и анализ их результатов (публикации 2, 3, 5-7); анализ ошибок (публикации 1-3, 5, 7); участие в создании ан-
самблевого классификатора (публикации 6, 7); формирование комплектов для разметчиков, участие в формировании руководств для разметчиков, статистическая обработка результатов разметки (публикация 4).
- Парамонов И. В.: рекомендации по постановке экспериментов; рекомендации по описанию методов (все публикации); описание результатов экспериментов и рекомендации по программной реализации метода, (публикация 1); рекомендации по описанию результатов экспериментов (публикации (публикации 2-3, 5-7); рекомендации по подготовке данных (публикация 2); участие в создании ансамблевого классификатора (публикация 6); анализ источников, организация процесса разметки, участие в формировании руководств для разметчиков (публикация 4).
- Бойчук Е. И.: формирование набора семантических правил (публикации 5-6); консультирование по разработке алгоритма и описании результатов экспериментов (публикация 2); консультирование по постановке экспериментов и описанию их результатов (публикация 7).
Структура и объем диссертации. Диссертационная работа состоит из введения, семи глав, заключения, списка литературы, включающего 131 наименование, и восьми приложений. Работа изложена на 430 страницах машинописного текста, содержит 8 рисунков и 51 таблицу.
Основное содержание работы
Во введении обосновывается актуальность исследований, проводимых в рамках диссертационной работы, формулируется цель, ставятся задачи работы, излагаются научная новизна и практическая значимость представляемой работы, приводятся новые научные результаты, выносимые на защиту.
В первой главе приведён обзор основных работ, посвящённых применению семантических правил для определения тональности предложений: как общей, так и по отношению к именованной сущности и к аспекту социально-экономического развития. Поскольку для определения тональности часто применяются методы машинного обучения, в главе также приводятся типичные метрики качества, достижимые с их помощью. Главу завершает обзор работ, посвящённых методам построения деревьев синтаксических единиц.
В работе [104] для представления предложений использовались синтаксические деревья; семантические правила были реализованы как алгоритмы над деревьями. Эксперименты показали, что при помощи правил, использующих и деревья синтаксических единиц, и деревья синтаксических связей, можно определять тональность длинных (10 и более слов) предложений точнее, чем при использовании только деревьев синтаксических связей: F-мера составила 0.59 и 0.55 соответственно, из чего авторы сделали вывод о том, что для представления сложных предложений лучше подходят деревья синтаксических единиц. Метод определения тональности, описанный в работе [90], заявлен авторами как применимый для различных языков за счёт адаптации используемых семантических правил (первоначально созданных для английского языка). Авторами была проведена адаптация правил для немецкого, китайского и корейского языков. Для английского языка F-мера составила 0.76 на предложениях из Facebook и 0.70 на предложениях из Twitter. Улучшенный вариант этого метода, позволивший получить долю правильных ответов 0.88 на записях из Twitter и 0.76 на отзывах на фильмы, описан в [37]. Существенный недостаток метода состоит в том, что предложение представляется в виде списка слов, а правила реализованы как шаблоны над таким списком: для языков без строгого порядка слов это приводит к снижению качества.
Работ, описывающих методы определения тональности предложений на русском языке, опубликовано значительно меньше, чем работ для английского языка. В работе [95] приведены метрики качества определения общей тональности предложений публицистического стиля из «Живого журнала» с помощью словарного метода: F-мера на различных наборах предложений составила 0.55-0.60.
Среди методов определения общей тональности, использующих машинное обучение, наиболее популярны классификаторы, основанные на нейронных сетях LSTM и BERT. С их помощью в большинстве задач достижима F-мера 0.70-0.75; при использовании комплексных методов обучения качество может быть существенно повышено с увеличением F-меры как минимум до 0.85-0.90. Однако для некоторых предметных областей (например, новостей или финан-
совой аналитики) легко достижимое качество существенно ниже — большинство исследований сообщают о достижении F-меры 0.70-0.80 [69; 98; 109].
Создание методов определения тональности по отношению к именованной сущности для предложений на русском языке долгое время было затруднено из-за недостатка открытых корпусов. В настоящий момент для русского языка отсутствуют актуальные сведения о методах решения этой задачи, использующих семантические правила. Тем не менее, опубликовано несколько работ, описывающих использование нейросетевых классификаторов. В [93] используется классификатор на основе архитектуры IAN и векторов эмбеддингов ELMo. На тестовой выборке корпуса CABSAR макро-Б-мера составила 0.70. В работах [73; 75; 112] используются различные варианты классификатора BERT-SPC; макро-Б-мера на тестовой выборке корпуса RuSentNE-2023 составила от 0.75 до 0.78.
Основная проблема при создании методов определения тональности по отношению к аспекту социально-экономического развития состоит в том, что интересующий аспект не обязательно упоминается в предложении явно. Для решения этой проблемы часто используются аспектные термины — слова предложения, тесно связанные с целевым аспектом, например, «школа» или «студенты» для аспекта «образование»: тональность по отношению к аспекту определяется на основе тональности по отношению к аспектным терминам [43].
Создание автоматического построителя деревьев синтаксических единиц требует формальной грамматики синтаксических единиц. Первая такая грамматика для русского языка была создана ещё в 1960-х годах [33], однако из-за ограничений вычислительной техники и сложности самой грамматики автоматических построителей деревьев на её основе создано не было. В работах [122; 13] описывается метод построения деревьев синтаксических единиц на основе морфологической информации и анализа синтаксических связей, однако эксперименты проводились только для простых предложений. Тем не менее, опыт проекта «Диалинг» [23] показывает, что в настоящий момент качество инструментов автоматического анализа отдельных слов и связей между ними
уже достаточно высоко, чтобы результаты их работы можно было применять для анализа сложных предложений [126; 24].
Вторая глава посвящена корпусам, на которых оценивалось качество методов определения тональности.
Для оценки качества методов определения общей тональности предложений публицистического стиля был создан корпус предложений из состава OpenCorpora (открытого корпуса публицистических текстов, относящихся к различным предметным областям), получивший название OpenSentimentCorpus. Для его создания использовались только предложения из семи и более слов. Разметка корпуса производилась разметчиками-волонтёрами, перед которыми ставилась задача оценить тональность предложений из индивидуального набора. Наборы автоматически формировались так, чтобы каждое предложение оказались бы в наборах хотя бы трёх разметчиков. Разметчики могли оценить тональность предложения как положительную, отрицательную, нейтральную, смешанную или неопределимую. Для оценки качества методов определения общей тональности использовались только предложения, получившие согласованные оценки, среди которых не было ни одной оценки предложения как имеющего неопределимую тональность. Также не использовались предложения со смешанной тональностью. Характеристики полученного корпуса приведены в таблице 1. OpenSentimentCorpus доступен онлайн и может быть загружен по адресу https://github.com/yarfruct/open-sentiment-corpus.
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Методы и средства морфологической сегментации для систем автоматической обработки текстов2022 год, кандидат наук Сапин Александр Сергеевич
Методы и средства морфологической сегментации для систем автоматической обработки текстов2023 год, кандидат наук Сапин Александр Сергеевич
Нейросетевое моделирование и машинное обучение на основе экспериментальных и наблюдательных данных2021 год, доктор наук Сбоев Александр Георгиевич
Методы сравнения и построения устойчивых к шуму программных систем в задачах обработки текстов2019 год, кандидат наук Малых Валентин Андреевич
Список литературы диссертационного исследования кандидат наук Полетаев Анатолий Юрьевич, 2025 год
Литература
1. Jurafsky D., Martin J.H. Speech and Language Processing. 2nd Edition. USA: Prentice-Hall, Inc., 2009. 1024 p.
2. Батура Т.В., Чаринцева М.В. Основы обработки текстовой информации: Учебное пособие. Новосибирск: Институт систем информатики им. А.П. Ершова СО РАН, 2016. 45 с.
3. Андреева С.В. Типология конструктивно-синтаксических единиц в русской речи // Вопросы языкознания. 2004. № 5. С. 32-45.
4. Онипенко Н.К. Об основаниях классификации синтаксических единиц // Труды института русского языка им. В.В. Виноградова. 2019. Т. 20. С. 189-201.
5. Percival W.K. On the historical source of immediate constituent analysis // Notes from the linguistics underground. 1976. pp. 229-242.
6. Waziri Z.Y., Safana M.I. Contrastive analysis of English and Hausa sentence structures and its pedagogical implications // Voices: A Journal of English Studies. 2021. vol. 5. pp. 15-27.
7. Dewi N.M.P., Putra I.G.W.N., Winarta I.B.G.N. Imperative Sentence in «The Guidance iPhone Support Website» // Elysian Journal: English Literature, Linguistics and Translation Studies. 2021. vol. 1. pp. 81-92.
8. Nguyen H.V., Tan N., Quan N.H., Huong T.T., Phat N.H. Building a Chatbot System to Analyze Opinions of English Comments // Informatics and Automation. 2023. vol. 22. no. 2. pp. 289-315.
9. Matchin W., Hickok G. The cortical organization of syntax // Cerebral Cortex. 2020. vol. 30. no. 3. pp. 1481-1498.
10. Ениколопов С.Н., Кузнецова Ю.М., Осипов С.Г., Смирнов И.В., Чудова Н.В. Метод реляционно-ситуационного анализа текста в психологических исследованиях // Психология. Журнал Высшей школы экономики. 2021. Т. 18. № 4. С. 748-769.
11. Zhang Y., Zhang Y. Tree communication models for sentiment analysis // Proceedings of the 57th annual meeting of the association for computational linguistics. 2019. pp. 3518-3527. DOI: 10.18653/v1/P19-1342.
12. Marcus M., Santorini B., Marcinkewicz M.A. Building a large annotated corpus of English: The Penn Treebank // Computational Linguistics. 1993. vol. 19 no. 2. pp.313-330.
13. Розенталь Д.Э., Голуб И.Б., Теленкова М.А. Современный русский язык. 16-e изд. М.: АЙРИС-пресс, 2018. 448 с.
14. Chomsky N. On certain formal properties of grammars // Information and control. 1959. vol. 2. no. 2. pp. 137-167.
15. Chomsky N. Some Puzzling Foundational Issues: the Reading Program // Catalan journal of linguistics. 2019. pp. 263-285. DOI: 10.5565/rev/catjl.287.
16. Muller S. Grammatical theory: From transformational grammar to constraint-based approaches. Fifth revised and extended edition. Berlin: Language Science Press, 2023. 889 p. DOI: 10.17169/langsci.b25.167.
17. Taylor A., Marcus M., Santorini B. The Penn Treebank: an overview // Treebanks: Building and using parsed corpora. Dordrecht: Springer Netherlands, 2003. 407 p. DOI: 10.1007/978-94-010-0201-1.
18. Zhou J., Zhao H. Head-Driven Phrase Structure Grammar Parsing on Penn Treebank // Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. pp. 2396-2408.
19. Gaddy D., Stern M., Klein D. What's Going On in Neural Constituency Parsers? An Analysis // Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2018. vol. 1. pp. 999-1010.
20. Zhang M.S. A survey of syntactic-semantic parsing based on constituent and dependency structures // Science China Technological Sciences. 2020. vol. 63. no. 10. pp. 1898-1920.
21. Yang S., Cui L., Ning R., Wu D., Zhang Y. Challenges to open-domain constituency parsing // Findings of the Association for Computational Linguistics: ACL 2022. 2022. pp. 112-127.
22. Гладкий А.В., Мельчук И.А. Элементы математической лингвистики. М.: Наука, 1969. 192 с.
23. Гладкий А.В. Синтаксические структуры естественного языка. Изд. 2-е. М.: УРСС, 2007. 146 с.
24. Коротаев Н.А. Синтаксические группы А.В Гладкого: анализ конструкций с сочинением // Вестник РГГУ. Серия: Литературоведение. Языкознание. Культурология. 2013. № 8(109). С. 16-36.
25. Кагиров И.А., Леонтьева А.Б. Модуль синтаксического анализа для литературного русского языка// Труды СПИИРАН. 2008. Т. 6. С. 171-183.
26. Leontyeva A., Kagirov I. The module of morphological and syntactic analysis SMART // Text, Speech and Dialogue: 11th International Conference, TSD 2008. 2008. pp. 373-380.
27. Леонтьева Н.Н., Ермаков М.В., Крылов С.А., Семенова С.Ю., Соколова Е.Г. Прикладной семантический словарь РУСЛАН: основная концепция и обновленный подход // Компьютерная лингвистика и интеллектуальные технологии: По материалам ежегодной международной конференции «Диалог». 2020. С. 1049-1064.
28. Москвина А.Д., Орлова Д., Паничева П.В., Митрофанова О.А. Разработка ядра синтаксического анализатора для русского языка на основе библиотек NLTK // Сборник научных статей. Труды XIX Международной объединённой научной
конференции «Интернет и современное общество». Санкт-Петербург: Санкт-Петербургский национальный исследовательский университет информационных технологий, механики и оптики. 2016. C. 44-54.
29. Shelmanov A., Pisarevskaya D., Chistova E., Toldova S., Kobozeva M., Smirnov I. Towards the data-driven system for rhetorical parsing of Russian texts // Proceedings of the Workshop on Discourse Relation Parsing and Treebanking. 2019. pp. 82-87.
30. Гаврилов Д.А Сопоставительное изучение пунктуации в сетевом газетном заголовке: к постановке проблемы // Вестник Чувашского государственного педагогического университета им. И.Я. Яковлева. 2021. № 3(112). С. 3-8.
31. De Marneffe M.C, Manning C.D., Nivre J., Zeman D. Universal Dependencies // Computational Linguistics. 2021. vol. 47. no. 2. pp. 255-308.
32. Lyashevskaya O., Bocharov V., Sorokin A., Shavrina T., Granovsky D., Alexeeva S. Text collections for evaluation of Russian morphological taggers // Journal of Linguistics / Jazykovedny Casopis. 2017. vol. 68. no. 2. pp. 258-267.
33. Kirillovich A., Loukachevitch N., Kulaev M., Bolshina A., Ilvovsky D. Sense-Annotated Corpus for Russian // Proceedings of the 5th International Conference on Computational Linguistics in Bulgaria (CLIB 2022). 2022. pp. 130-136.
34. Volkova L., Bocharov V. An approach to inter-annotation agreement evaluation for the named entities annotation task at OpenCorpora // Communications in Computer and Information Science. 2019. vol. 1119. pp. 33-44.
35. Lagutina K. Topical Text Classification of Russian News: a comparison of BERT and Standard Models //31st Conference of Open Innovations Association FRUCT. 2022. pp.160-166.
36. Yang S., Tu K. Bottom-up constituency parsing and nested named entity recognition with pointer networks // Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022. vol. 1. pp. 2403-2416.
Полетаев Анатолий Юрьевич - ассистент, кафедра компьютерных сетей, факультет информатики и вычислительной техники, Ярославский государственный университет им. П.Г. Демидова. Область научных интересов: анализ и моделирование естественного языка, математическая статистика. Число научных публикаций — 14. anatoliy-poletaev@mail.ru; улица Советская, 14, 150003, Ярославль, Россия; р.т.: +7(910)819-8325.
Парамонов Илья Вячеславович - канд. физ.-мат. наук, доцент, кафедра компьютерных сетей, факультет информатики и вычислительной техники, Ярославский государственный университет им. П.Г. Демидова; руководитель лаборатории, ярославская лаборатория Ассоциации открытых инноваций FRUCT. Область научных интересов: компьютерная лингвистика, искусственные нейронные сети, методология разработки программного обеспечения. Число научных публикаций — 55. ilya.paramonov@fruct.org; улица Советская, 14, 150003, Ярославль, Россия; р.т.: +7(905)633-3993.
Бойчук Елена Игоревна - д-р филол. наук, доцент, старший научный сотрудник, отдел управления наукой и инновациями, Ярославский государственный университет им. П.Г. Демидова. Область научных интересов: компьютерная лингвистика, филология, литературоведение, лингвистика текста, функциональная грамматика, теория языка и сравнительно-сопоставительные исследования. Число научных публикаций — 108. elena-boychouk@rambler.ru; улица Советская, 14,150003, Ярославль, Россия; р.т.: +7(903)824-1797.
Поддержка исследований. Исследование выполнено за счет гранта Российского научного фонда № 23-21-00495 (https://rscf.ru/project/23-21-00495/).
DOI 10.15622/ia.22.6.3
A. POLETAEV, I. PARAMONOV, E. BOYCHUK ALGORITHM OF CONSTITUENCY TREE FROM DEPENDENCY TREE CONSTRUCTION FOR A RUSSIAN-LANGUAGE SENTENCE
Poletaev A., Paramonov I., Boychuk E.. Algorithm of Constituency Tree from Dependency Tree Construction for a Russian-Language Sentence.
Abstract. Automatic syntactic analysis of a sentence is an important computational linguistics task. At present, there are no syntactic structure parsers for Russian that are publicly available and suitable for practical applications. Ground-up creation of such parsers requires building of a treebank annotated according to a given formal grammar, which is quite a cumbersome task. However, since there are several syntactic dependency parsers for Russian, it seems reasonable to employ dependency parsing results for syntactic structure analysis. The article introduces an algorithm that allows to construct the constituency tree of a Russian sentence by a syntactic dependency tree. The formal grammar used by the algorithm is based on the D.E. Rosenthal's classic reference. The algorithm was evaluated on 300 Russian-language sentences. 200 of them were selected from the aforementioned reference, and 100 from OpenCorpora, an open corpus of sentences extracted from Russian news and periodicals. During the evaluation, the sentences were passed to syntactic dependency parsers from Stanza, SpaCy, and Natasha packages, then the resulted dependency trees were processed by the proposed algorithm. The obtained constituency trees were compared with the trees manually annotated by experts in linguistics. The best performance was achieved using the Stanza parser: the constituency parsing F\ -score was 0.85, and the sentence parts tagging accuracy was 0.93, that would be sufficient for many practical applications, such as event extraction, information retrieval and sentiment analysis.
Keywords: computational linguistics, natural language processing, syntactic parsing, constituency tree, dependency tree, formal grammar.
References
1. Jurafsky D., Martin J.H. Speech and Language Processing. 2nd Edition. USA: Prentice-Hall, Inc., 2009. 1024 p.
2. Batura T V., Charinceva M.V. Osnovy obrabotki tekstovoj informacii: Uchebnoe posobie. [Basics of textual information processing: Study Guide]. Novosibirsk: Institut sistem informatiki im. A.P. Ershova SO RAN, 2016. 45 p. (in Russ.).
3. Andrejeva S.V. [Typology of constructive-syntactic units in Russian speech]. Voprosy yazykoznaniya - Problems of linguistics. 2004. no. 5. pp. 32-45. (in Russ.).
4. Onipenko N.K. [About the grounds for the classification of syntactic units]. Trudy Instituta Russkogo Iazyka im. V.V. Vinogradova - Proceedings of the V.V. Vinogradov Russian Language Institute. 2019. vol. 20. pp. 189-201. (in Russ.).
5. Percival W.K. On the historical source of immediate constituent analysis. Notes from the linguistics underground. 1976. pp. 229-242.
6. Waziri Z.Y., Safana M.I. Contrastive analysis of English and Hausa sentence structures and its pedagogical implications. Voices: A Journal of English Studies. 2021. vol. 5. pp. 15-27.
7. Dewi N.M.P., Putra I.G.W.N., Winarta I.B.G.N. Imperative Sentence in «The Guidance iPhone Support Website». Elysian Journal: English Literature, Linguistics and Translation Studies. 2021. vol. 1. pp. 81-92.
8. Nguyen H.V., Tan N., Quan N.H., Huong T.T., Phat N.H. Building a Chatbot System to Analyze Opinions of English Comments. Informatics and Automation. 2023. vol. 22. no. 2. pp. 289-315.
9. Matchin W., Hickok G. The cortical organization of syntax. Cerebral Cortex. 2020. vol. 30. no. 3. pp. 1481-1498.
10. Enikolopov S.N., Kuznetsova Y.M., Osipov G.S., SmirnovI.V., ChudovaN.V. [The Method of Relational-Situational Analysis of Text in Psychological Research]. Psihologiya. Zhurnal vysshej shkoly ekonomiki - Psychology. Journal of the Higher School of Economics]. 2021. vol. 18. no. 4. pp. 748-769. (inRuss.).
11. Zhang Y., Zhang Y. Tree communication models for sentiment analysis. Proceedings of the 57th annual meeting of the association for computational linguistics. 2019. pp. 3518-3527. DOI: 10.18653/v1/P19-1342.
12. Marcus M., Santorini B., Marcinkewicz M.A. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics. 1993. vol. 19 no. 2. pp. 313-330.
13. Rozental D.E., Golub I.B., Telenkova M.A. Sovremennyj russkij jazyk [Modern Russian language (16th Edition)]. Moscow: AJRIS-press, 2018. 448 p. (in Russ.).
14. Chomsky N. On certain formal properties of grammars. Information and control. 1959. vol. 2. no. 2. pp. 137-167.
15. Chomsky N. Some Puzzling Foundational Issues: the Reading Program. Catalan journal of linguistics. 2019. pp. 263-285. DOI: 10.5565/rev/catjl.287.
16. Muller S. Grammatical theory: From transformational grammar to constraint-based approaches. Fifth revised and extended edition. Berlin: Language Science Press, 2023. 889 p. DOI: 10.17169/langsci.b25.167.
17. Taylor A., Marcus M., Santorini B. The Penn Treebank: an overview. Treebanks: Building and using parsed corpora. Dordrecht: Springer Netherlands, 2003. 407 p. DOI: 10.1007/978-94-010-0201-1.
18. Zhou J., Zhao H. Head-Driven Phrase Structure Grammar Parsing on Penn Treebank. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. pp. 2396-2408.
19. Gaddy D., Stern M., Klein D. What's Going On in Neural Constituency Parsers? An Analysis. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2018. vol. 1. pp. 999-1010.
20. Zhang M.S. A survey of syntactic-semantic parsing based on constituent and dependency structures. Science China Technological Sciences. 2020. vol. 63. no. 10. pp. 1898-1920.
21. Yang S., Cui L., Ning R., Wu D., Zhang Y. Challenges to open-domain constituency parsing. Findings of the Association for Computational Linguistics: ACL 2022. 2022. pp. 112-127.
22. Gladkij A.V., Melchuk I.A. Jelementy matematicheskoj lingvistiki [Elements of mathematical linguistics]. Moscow: Nauka, 1969. 192 p. (inRuss.).
23. Gladkij A.V. Sintaksicheskie struktury estestvennogo jazyka [Syntactic structures of natural language (2nd Edition)]. Moscow: URSS, 2007. 146 p. (in Russ.).
24. Korotaev N.A. [A.V. Gladkij syntactic groups: analysis of compound constructions]. Vestnik RGGU. Seriya: Literaturovedenie. Yazykoznanie. Kulturologiya - RSUH Bulletin. «Literary theory. Linguistics. Cultural Studies.» Series.]. 2013. no. 8(109). pp. 16-36. (in Russ.).
25. Kagirov I.A., Leontyeva A.B. [Module for syntax parsing of the literary Russian language]. Trudy SPIIRAN - SPIIRAS Proceedings. 2008. vol. 6. pp. 171-183. (in Russ.).
26. Leontyeva A., Kagirov I. The module of morphological and syntactic analysis SMART. Text, Speech and Dialogue: 11th International Conference, TSD 2008. 2008. pp. 373-380.
27. Leontyeva N.N., Ermakov M.V., Krylov S.A., Semenova S.Yu., Sokolova E.G. [On traditional conception and upgrading of one applied semantic dictionary]. Kompyuternaya lingvistika i intellektualnye tekhnologii: Po materialam ezhegodnoj mezhdunarodnoj konferencii «Dialog» - Computational linguistics and intellectual technologies: Papers from the annual international conference «Dialogue». 2020. pp. 1049-1064. (in Russ.).
28. Moskvina A.D., Orlova D. Panicheva P.V., Mitrofanova O.A. Razrabotka yadra sintaksicheskogo analizatora dlya russkogo yazyka na osnove bibliotek NLTK [Development of the Core for Syntactic Parser for Russian based on NLTK libraries] Kompyuternaya lingvistika i vychislitelnye ontologii: Sbornik nauchnyh statej. Trudy XIX Mezhdunarodnoj ob'edinyonnoj nauchnoj konferencii [Computational linguistics and computational ontologies: Collection of scientific articles. Proceedings of the XlXth International joint scientific conference]. St. Petersburg: Sankt-Peterburgskij nacional'nyj issledovatel'skij universitet informacionnyh tekhnologij, mekhaniki i optiki, 2016. pp. 44-54. (in Russ.).
29. Shelmanov A., Pisarevskaya D., Chistova E., Toldova S., Kobozeva M., Smirnov I. Towards the data-driven system for rhetorical parsing of Russian texts. Proceedings of the Workshop on Discourse Relation Parsing and Treebanking. 2019. pp. 82-87.
30. Gavrilov D.A. [Comparative study of punctuation in an online newspaper headline: statement of the problem]. Vestnik Chuvashskogo Gosudarstvennogo Pedagogicheskogo Universiteta im. I. Y. Yakovleva - I. Yakovlev Chuvash State Pedagogical University Bulletin. 2021. no. 3(112). pp. 3-8. (in Russ.).
31. De Marneffe M.C, Manning C.D., Nivre J., Zeman D. Universal Dependencies. Computational Linguistics. 2021. vol. 47. no. 2. pp. 255-308.
32. Lyashevskaya O., Bocharov V., Sorokin A., Shavrina T., Granovsky D., Alexeeva S. Text collections for evaluation of Russian morphological taggers. Journal of Linguistics / Jazykovedny Casopis. 2017. vol. 68. no. 2. pp. 258-267.
33. Kirillovich A., Loukachevitch N., Kulaev M., Bolshina A., Ilvovsky D. Sense-Annotated Corpus for Russian. Proceedings of the 5th International Conference on Computational Linguistics in Bulgaria (CLIB 2022). 2022. pp. 130-136.
34. Volkova L., Bocharov V. An approach to inter-annotation agreement evaluation for the named entities annotation task at OpenCorpora. Communications in Computer and Information Science. 2019. vol. 1119. pp. 33-44.
35. Lagutina K. Topical Text Classification of Russian News: a comparison of BERT and Standard Models. 31st Conference of Open Innovations Association FRUCT. 2022. pp. 160-166.
36. Yang S., Tu K. Bottom-up constituency parsing and nested named entity recognition with pointer networks. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 2022. vol. 1. pp. 2403-2416.
Poletaev Anatoliy - Assistant, Chair of computer networks, faculty of computer science, P.G. Demidov Yaroslavl State University. Research interests: natural language analysis and modeling, mathematical statistics. The number of publications — 14. anatoliy-poletaev@mail.ru; 14, Sovetskaya St., 150003, Yaroslavl, Russia; office phone: +7(910)819-8325.
Paramonov Ilya - Ph.D., Associate professor, Chair of computer networks, faculty of computer science, P.G. Demidov Yaroslavl State University; Head of the laboratory, Yaroslavl Laboratory of Open Innovations Association FRUCT. Research interests: computational linguistics, artificial neural networks, software development methods. The number of publications — 55. ilya.paramonov@fruct.org; 14, Sovetskaya St., 150003, Yaroslavl, Russia; office phone: +7(905)633-3993.
Boychuk Elena - Ph.D., Dr.Sci., Associate Professor, Senior researcher, Research and development department, P.G. Demidov Yaroslavl State University. Research interests: computational linguistics, philology, literary criticism, text linguistics, functional grammar, language theory and comparative studies. The number of publications — 108. elena-boychouk@rambler.ru; 14, Sovetskaya St., 150003, Yaroslavl, Russia; office phone: +7(903)824-1797.
Acknowledgements. The reported study was funded by the grant of Russian Science Foundation No. 23-21-00495 (https://rscf.ru/en/project/23-21-00495/).
ISSN0146-4116, Automatic Control and Computer Sciences, 2023, Vol. 57, No. 7, pp. 740-749. ©Allerton Press, Inc., 2023. Russian Text © The Author(s), 2022, published in Modelirovanie i Analiz Informatsionnykh Sistem, 2022, Vol. 29, No. 2, pp. 134-147.
Recursive Sentiment Detection Algorithm for Russian Sentences
A. Y. Poletaev"' * (orcid: 0000-0003-0116-4739) and I. V. Paramonov"' ** (ORCID: 0000-0003-3984-8423)
a Demidov Yaroslavl State University, Yaroslavl, 150003 Russia *e-mail: anatoliy-poletaev@mail.ru **e-mail: ilya.paramonov@fruct.org Received April 30, 2022; revised May 22, 2022; accepted May 25, 2022
Abstract— The article is devoted to the task of sentiment detection of Russian sentences. The sentiment is conceived as the author's attitude to the topic of a sentence. This assay considers positive, neutral, and negative sentiment classes, i.e., the task of three-classes classification is solved. The article introduces a rule-based sentiment detection algorithm for Russian sentences. The algorithm is based on the assumption that the sentiment of a phrase can be determined by the sentiments of its parts by the recursive application of appropriate semantic rules to the sentiments of its parts organized as a constituency parse tree. The utilized set of semantic rules was constructed based on a discussion with experts in linguistics. The experiments showed that the proposed recursive algorithm performs slightly worse on the hotel reviews corpus than the adapted rule-based approach: weighted Fl-measures are 0.75 and 0.78, respectively. To measure the algorithm efficiency on complex sentences, we created OpenSentiment-Corpus based on OpenCorpora, an open corpus of sentences extracted from Russian news and periodicals. On OpenSentimentCorpus the recursive algorithm performs be.er than the adapted approach does: Fl-measures are 0.70 and 0.63, respectively. This indicates that the proposed algorithm has an advantage in case of more complex sentences with more subtle ways of expressing the sentiment.
Keywords: sentiment analysis, sentiment detection, semantic rules, sentiment corpus DOI: 10.3103/S0146411623070118
1. INTRODUCTION
This article introduces and evaluates a sentiment detection algorithm for Russian sentences based on the development of our previous research [1]. We consider the task of sentence-level sentiment detection as identifying the attitude of the sentence's author to the topic of the sentence. A sentence is considered as positive if it contains a positive fact, opinion, or emotion expressed by the author and there are no negative facts, opinions, or emotions; or these negative expressions are overlapped by positive ones. In the opposite situation (when negative facts, opinions, or emotions prevail) the sentence is negative. If the sentence is neither positive nor negative we consider it to be neutral [2, 3].
The algorithm we introduce is based on semantic rules. Rule-based sentiment analysis approaches are less common than neural network-based ones, however, they can in some cases surpass the drawbacks of neural networks: semantic rules do not need large corpora to be trained on and their results are rather easy to interpret. In our previous work [1] we adapted a rule-based approach, which was originally proposed for English by Xie et al., [4], for the Russian language. The original approach utilizes semantic rules implemented as patterns applied to a list of words representing a sentence. Such a way is difficult to apply in Russian because it, unlike English, has no strict word order. To resolve the issue, we reconstructed the rules as algorithms over the dependency parse tree, which represents syntactic dependencies between the words of a sentence. The adapted approach was evaluated on a hotel reviews corpus; the F1 -measure of 0.73 was achieved. The state-of-the-art BERT neural network showed slightly better (by 5%) results.
The main goal of this work is to propose and evaluate a novel sentiment detection algorithm that analyzes the clausal structure of a sentence better than the approach used in our previous work. We use the constituency parse tree as a source of information on sentence phrase structure and assume that the sentiment of a phrase is determined by the sentiments of its constituents represented as its children in the constituency parse tree. On this basis, we develop the recursive sentiment detection algorithm over the constituency tree.
We also assume that the sentiment of a sentence highly depends on the sentiment of its predicative core, which is the root of the constituency tree of a sentence. Thus, the proposed algorithm starts its work from the root of the constituency tree of a given sentence and recursively detects the sentiment of every node (the phrase in a sentence) based on the sentiments of its children (constituent parts of a phrase) according to a set of semantic rules.
Most of research dedicated to sentiment analysis of Russian sentences utilizes corpora with rather short and simple sentences (e.g., tweets and reviews) [5]. In most of such sentences sentiment is expressed clearly and directly. Sentences with more complicated structure are usually not considered. To fill the gap, we created new OpenSentimentCorpus based on sentences extracted from news and periodicals from the OpenCorpora corpus (http://opencorpora.org). The algorithm is evaluated and compared with the approach described in our previous work on both the hotel reviews corpus and OpenSentimentCorpus.
The rest of this article is structured as follows. Section 2 describes the related work. In Section 3 we give a general description of the proposed recursive algorithm. Section 4 describes the newly marked corpus we use for evaluation. Section 5 contains the results of evaluation of the introduced algorithm and comparison with the previously adapted rule-based approach. Section 6 describes the errors of the introduced algorithm on the OpenSentimentCorpus. In conclusion we summarize the results and propose ideas for future work.
2. RELATED WORK
The usage of semantic rules for sentence-level sentiment analysis was originally proposed by M. Shaikh et al. [6] and improved with the use of types of words dependencies by L. Tan et al. [3]. Y. Xie et al. [4] introduced an advanced rule-based approach that uses semantic rules handling different ways of how sentiment can be expressed to take different kinds of clauses into consideration. Xie's approach achived /¡-measures of 0.76 and 0.68 on Facebook comments and Twitter tweets datasets. O. Appel et al., [7] improved the approach by addition of new rules and semi-automatic sentiment dictionary enrichment using machine learning. The resulted approach was evaluated by its authors on three datasets. On the first and the second ones consisting of Twitter comments it achieves the accuracy of approximately 0.88. On the third dataset consisting of movie reviews the accuracy is about 0.76.
Our previous work [¡] was devoted to the adaptation of the Appel's rule-based sentiment analysis approach to the Russian language. The main adaptation challenge we faced was that semantic rules implemented as patterns over a list of words representing a sentence rely on the word order in a sentence too much to be used for Russian, which has no strict word order. To overcome the issue we recreated semantic rules as algorithms over the dependency parse tree, which represents syntactic dependencies between the words of a sentence. The adapted approach evaluation on the hotel reviews corpus showed the possibility of reaching the results close to the state-of-the-art for Russian using a semantic rules-based approach instead of commonly used neural networks.
Parse tree (also called syntax tree) is a representation of the hierarchical structure of the words of a sentence [8]. There are two approaches to the analysis of the sentence structure. The first is dependency-based, which considers a sentence as a hierarchy of words ordered according to the government relations. The nodes of the dependency tree are words and the edges represent the government relation. Another approach is constituency-based, which considers a sentence as a hierarchy of phrases. The leafs of the constituency parse tree are words, the branches are phrases, and the edges represent the "part-whole" relation.
Constituency trees and semantic rules are successfully used for sense analysis in a variety of applications, e.g., recognition of questions and instructions [9, ¡0] and inference detection [¡¡]. Turning to the sentiment analysis, there is Stanford Sentiment Treebank (SST) dataset [¡2] that includes sentiment labels for 2^54 phrases in the constituency trees of ¡¡855 sentences. Most of the sentiment analysis approaches using constituency trees utilize neural networks and use SST for training and evaluation. There are two popular approaches. The first approach is Tree-LSTM [¡3, ¡4], the architecture of neural networks that structures LSTM units according to the parse tree, each unit corresponds to a node of the tree. The second approach is aimed at capturing compositional sentiment semantic by predicting the sentiment of a phrase by the sentiments of its constituent parts. One of the recently developed model is SentiBERT that predicts sentiment of a phrase by the transformer representations of its parts [¡5] using machine learning.
The adaptation of aforementioned neural network based approaches for Russian seems to be a difficult task at the moment because of requirement of a large corpus similar to SST, i.e., with a sentiment mark for each node of the constituency tree. As a consequence, the development of rule-based approaches that do not require large corpora to be trained on looks like a promising direction of sentiment analysis for Russian.
3. ALGORITHM DESCRIPTION
The phrase structure of a sentence describes how the words of a sentence are grouped to form utterances [8]. To process the phrase structure we use the constituency trees that represent the aforementioned word grouping. The leafs of the constituency tree are words, the branches are phrases, and the edges represent the "part-whole" relation. The root node of the constituency tree is the predicative core of the sentence.
We choose constituency trees rather than dependency trees (which represent direct hierarchical relations between the words) because the constituency tree better fits the phrase structure of a sentence. As a consequence, it is easier to propose semantic rules using constituency trees than using dependency trees. Since there is no constituency tree parser for Russian, only dependency tree parsers, we constructed the supporting algorithm that builds the constituency tree on the basis of the given dependency tree. That supporting algorithm is based on grouping words that are dependants of the one head word into phrases according to their part-of-speech tags and dependency types.
The sentiment detection algorithm we propose is based on an assumption that the sentiment of a phrase can be determined by sentiments of its parts. We constructed a set of semantic rules that describes the way to determine the sentiment of a given phrase by sentiments of its parts, which are determined using recursive calls ofthe algorithm and rules application. Sentiments of single words (to which a recursive call cannot be applied) are taken from the RuSentiLex-2017 sentiment dictionary [16] (with changes made in our previous paper [1]). Despite the low quality of that dictionary [1], there is still no better sentiment dictionary for Russian.
More formally, let N be a node of the constituency tree, C(N) is a set of children of N, and S(N) is the sentiment of N. The algorithm starts from the root node of the constituency tree Nr that represents the predicative core of the given sentence; hence, S(Nr) is the sentiment of the sentence.
The algorithm calculates S(N) for a given node N as follows:
♦ if C(N) = 0, i.e., the node is a single word, S(N) is the sentiment for N taken from the dictionary;
♦ otherwise, recursively calculate S(Nc) for all Nc e C(N); choose an appropriate semantic rule depending on types of C (N) (see below) and apply it to calculate S(N).
To construct the semantic rules, we separated cases of how sentiment of a phrase depends on sentiments of its parts into three groups.
The first group includes phrases consisting of homogeneous sentence parts, e.g., the positive phrase #ум, честь и совесть# (mind, honor, and conscience), the neutral phrase и талантливый, и ленивый (talented and lazy), the negative phrase разработанный коллективно и обречённый на неудачу (jointly developed and doomed to fail), etc. For these cases we constructed the following rule to detect the sentiment of a phrase that consists of homogeneous parts: the phrase is positive if at least one of its homogeneous parts is positive and there are no negative homogeneous parts; the phrase is negative if at least one of its homogeneous parts is negative and there are no positive homogeneous parts; otherwise the phrase is neutral.
In the second group of the rules, one part of a phrase modifies the sentiment of another part of the phrase. For example, consider the phrases объём производства увеличивается (production value is growing up) and дефицит бюджета увеличивается (budgetgap is increasing). In the first one the word увеличивается modifies the neutral sentiment of the phrase объём производства and makes the whole phrase positive. In the second phrase the same word увеличивается lefts the negative sentiment of дефицит бюджета unchanged; the whole phrase is negative. 26 most common Russian modifier words and phrases we use in this research after a discussion with experts in linguistics are presented in Table 1.
In the third group, the sentiment of a phrase is a result of conjunction of the sentiments of its parts. For example, the neutral noun законопроект (bill) in conjunction with the positive adjective своевременный (well-timed) makes the phrase своевременный законопроект positive. The result of conjunction can vary for different phrases depending on their structure. For example, the negative phrase отчислить лучшего студента (to expel the best student) contains the negative predicate отчислить and the positive object лучшего студента. On the other side, the neutral phrase заслуженно отчислили (expelled reasonably) contains the negative predicate отчислили and the positive adverbial modifier заслуженно.
After a discussion with experts in linguistics, we constructed the rules that represent typical results of conjunction for all possible combinations of Russian sentence parts (see Table 2).
The algorithm choose a rule to apply for a given phrase as follows:
♦ if the phrase consists of homogeneous parts, the rule for homogeneous parts is applied;
♦ if any of the modifier words and phrases is present in the sentence, the rule for the modifier phrase is applied;
♦ otherwise, the rule for combination of sentence parts is applied according to the types of the parts.
Table 1. Modifier words and phrases
Modification Words and phrases Translation
Make neutral and positive phrases negative; make negative phrases positive Не, нет, прекратить, разрушить, перестать, избегать, противоречить, уменьшать, утратить, ни один No, not, cease, destroy, stop, avoid, contradict, lessen, forfeit, neither
Make negative phrases neutral; left other phrases untouched Понимать, принимать Understand, accept
Make neutral phrases positive; left other phrases untouched Способность, обеспечить, позволить, выполнить, помочь, увеличить, достичь, доказать, надеяться, очень, также Ability, ensure, afford, fulfil, assist, increase, achieve prove, hope, very, too
Makes positive phrases neutral; left other phrases untouched Крайне Extremely
Make phrases neutral Вроде бы, несмотря на Apparently, despite
It should be mentioned that the model we use in the proposed algorithm has some limitations. The first one is related to situations when the sentiment of the phrase cannot be detected only by sentiments of their parts. For example, consider the phrase после проведения акции сайт организаторов исчез (the organizers' website disappeared after the promotion campaign). Despite the fact that its parts сайт организаторов and исчез после проведения акции are neutral, the phrase as a whole is negative. The second limitation is that the rules we propose are rather simple and cannot handle all the cases of detection of the phrase sentiment by sentiments of its parts. For example, the phrase мастерски обманывает contains a negative predicate and a positive adverbial modifier. According to the rule, the phrase should be neutral, but it is negative.
Nevertheless, despite all these limitations, it is useful to run experiments using such a simple model to reveal its potential for the sentiment analysis and gather error analysis results for its further improvement.
4. CORPORA USED IN EXPERIMENTS
To measure the proposed algorithm's ability to detect the sentiments of rather simple sentences we evaluated it on the hotel reviews corpus, used in our previous research [1]. In most of the sentences of the corpus sentiment is expressed clearly and directly, and the sentiment detection task is rather simple.
To assess the applicability and usefulness of the proposed algorithm for more complex cases we evaluated it on a multi-domain corpus with significant share of complex sentences.
There are several Russian sentiment datasets available: SentiRuEval-2015 Subtask 2, SentiRuEval2016, RuTweetCorp, LINIS Crowd, Kaggle Russian News Dataset, and RuReviews [2]. Unfortunately, none of them meet our requirements: some corpora consist of texts but not sentences (Kaggle Russian News Dataset, RuReviews), other ones are littered with markup errors (RuTweetCorp, LINIS Crowd) or contain too small amount of complex sentences (SentiRuEval datasets). That is why we created OpenSentimentCor-pus based on OpenCorpora, an open corpus of sentences extracted from Russian news and periodicals from different domains.
We started out by excluding sentences shorter than 7 words from OpenCorpora to increase the share of complex sentences and dump sentences too simple to be relevant for evaluation. The remaining sentences were marked by at least three evaluators.
Each sentence was marked as positive, neutral, or negative according to the sentiment it expresses; mixed if it contained both positive and negative sentiment, e.g., Пелевин пишет не в пример хуже Акунина, но у Пелевина есть что сказать читателю (Pelevin writes far worse than Akunin, but Pelevin has something to say to a reader); or doubtful for sentences with a questionable sentiment, unclear author's position, or any other reason of the evaluator's uncertainty.
The results of markup were processed as follows:
♦ The sentences that had at least one doubtful mark or both positive and negative marks were excluded to get rid of sentences with vague or unclear sentiment.
Table 2. Sentiment of combination of sentence parts
First part Second part Sentiment of the first Sentiment of the second Sentiment
of the phrase of the phrase part of the phrase part of the phrase of the phrase
Subject Predicate Positive Positive Positive
Neutral Positive
Negative Negative
Neutral Positive Positive
Neutral Neutral
Negative Negative
Negative Positive Neutral
Neutral Negative
Negative Negative
Attribute Positive Positive Positive
Neutral Positive
Negative Negative
Neutral Positive Positive
Neutral Neutral
Negative Negative
Negative Positive Negative
Neutral Negative
Negative Negative
Object Positive Positive Positive
Neutral Positive
Negative Neutral
Neutral Positive Neutral
Neutral Neutral
Negative Negative
Negative Positive Neutral
Neutral Negative
Negative Negative
Predicate Adv. Modifier Positive Positive Positive
Neutral Positive
Negative Negative
Neutral Positive Positive
Neutral Neutral
Negative Negative
Negative Positive Neutral
Neutral Negative
Negative Negative
Object Positive Positive Positive
Neutral Positive
Negative Negative
Neutral Positive Neutral
Neutral Neutral
Negative Negative
Negative Positive Negative
Neutral Negative
Negative Negative
Object Attribute Positive Positive Positive
Neutral Neutral
Negative Negative
Neutral Positive Positive
Neutral Neutral
Negative Negative
Negative Positive Neutral
Neutral Negative
Negative Negative
Table 3. Classes of sentences in hotel reviews corpus and OpenSentimentCorpus
Corpus Hotel reviews OpenSentimentCorpus
Positive 639 536
Neutral 232 2441
Negative 333 1510
Total 1204 4487
♦ The final sentiment of each sentence was assigned according to the result of agreement between all the evaluators. If there was no such an agreement (e.g., 2 of 3 marks were positive, or 1 mark was neutral), the sentence was excluded.
The resulted OpenSentimentCorpus is a multi-domain corpus, which fits for evaluation of two-, three, and four-classes classification algorithms. As the algorithm under evaluation distinguishes only positive, neutral, and negative sentiments, we used only positive, negative and neutral sentences in this research. The corpus is large enough to provide a reliable evaluation result.
The distributions of sentences among the classes for both hotel reviews corpus and OpenSentiment-Corpus are shown in Table 3.
Sentences in both corpora are distributed between the classes nonuniformly. About half of the sentences in the hotel reviews corpus is positive, a quarter is negative, and only one fifth is neutral. The main reason for this is the specificity of the reviews domain — most of the authors express their satisfaction or dissatisfaction and do not state facts. On the other hand, the major part of OpenSentimentCorpus contains neutral statements, which do not express their authors' opinion, about a third of a sentences is negative, and only approximately one tenth is positive. Despite this, there are still more than five hundred positive sentences, which is sufficient to evaluate the algorithm.
OpenSentimentCorpus is available online at https://github.com/yarfruct/open-sentiment-corpus.
5. EXPERIMENTS
We evaluated the proposed recursive algorithm on the hotel reviews corpus and OpenSentimentCorpus and compared its performance with the performance of the rule-based approach adapted for Russian and evaluated in [1].
As the performance metrics we used both the simple average and the weighted average of precision, recall, and F-score. The simple average (also called arithmetic mean) is just the sum of the performance metrics for all the classes divided by the number of classes. The weighted average is the sum of the performance metrics for classes multiplied by the number of sentences in every class and divided by total number of sentences in the corpus. The reason of using the weighted average is the imbalance of classes in corpora that may otherwise lead to incorrect assessment of the performance.
The performance metrics and confusion matrices on the hotel reviews corpus are shown in Tables 4 and 5. The proposed recursive algorithm detects positive and negative sentiment in hotel reviews quite precise, but slightly worse than the rule-based approach does. The decreasion is primarily attributable to the increased share of positive sentences incorrectly classified neutral. The main drawback of both the recursive algorithm and the rule-based approach is the negative and neutral sentiments distinction, however the recursive algorithm distinguishes negative and neutral sentences better than the adapted approach does.
The classification performance metrics and confusion matrices on OpenSentimentCorpus are shown in Tables 6 and 7. The average performance of the recursive algorithm is lower on OpenSentimentCorpus than on the hotel reviews corpus, reduction is approximately 5—6%. The most significant flaw is poor quality of positive sentiment detection: only a half of positive sentences is classified correctly, and significant amount of positive sentences is classified negative. On the other hand, the algorithm detects the negative sentiment on OpenSentimentCorpus better than on the hotel reviews corpus, and distinguishes negative and neutral sentiments more accurate. The recursive algorithm performs better than the rule-based approach by approximately 7%. The only significant advantage of the adapted rule-based approach is that it incorrectly classifies smaller amount of positive sentences.
Table 4. Sentiment classification performances of the proposed recursive algorithm and the algorithm adapted in [¡] on the hotel reviews corpus
Algorithm Proposed recursive algorithm Adapted rule-based approach No. of sentences
Class Precision Recall F-score Precision Recall F-score
Positive 0.89 0.77 0.82 0.88 0.88 0.88 639
Neutral 0.43 0.75 0.55 0.48 0.75 0.58 232
Negative 0.86 0.64 0.73 0.94 0.57 0.71 333
Average 0.73 0.72 0.70 0.77 0.73 0.73 1204
Weighted average 0.79 0.73 0.75 0.82 0.77 0.78 1204
Accuracy of the proposed recursive algorithm = 0.73 Accuracy of the adapted rule-based approach = 0.77
Table 5. Sentiment classification confusion matrices of the proposed recursive algorithm and the rule-based approach adapted in [¡] on the hotel reviews corpus
Algorithm Proposed recursive algorithm Adapted rule-based approach Total
Actual Predicted Positive Neutral Negative Positive Neutral Negative
Positive 491 137 11 546 72 3 639
Neutral 33 175 24 50 173 9 232
Negative 29 92 212 24 118 191 333
Table 6. Sentiment classification performances of the proposed recursive algorithm and the adapted rule-based approach on OpenSentimentCorpus
Algorithm Proposed recursive algorithm Adapted rule-based approach No. of sentences
Class Precision Recall F-score Precision Recall F-score
Positive 0.45 0.49 0.47 0.30 0.71 0.42 536
Neutral 0.76 0.74 0.75 0.72 0.65 0.68 2441
Negative 0.70 0.70 0.70 0.77 0.52 0.62 1510
Average 0.63 0.64 0.64 0.60 0.62 0.57 4487
Weighted average 0.70 0.70 0.70 0.68 0.61 0.63 4487
Accuracy of the proposed recursive algorithm = 0.70 Accuracy of the adapted rule-based approach = 0.61
Table 7. Sentiment classification confusion matrices of the proposed recursive algorithm and the adapted rule-based approach on OpenSentimentCorpus
Algorithm Proposed recursive algorithm Adapted rule-based approach Total
Actual Predicted Positive Neutral Negative Positive Neutral Negative
Positive 261 209 66 378 131 27 536
Neutral 244 1800 397 640 1596 205 2441
Negative 81 371 1058 228 504 778 1510
6. ERROR ANALYSIS
To investigate the reasons of the performance decrease on OpenSentimentCorpus, we collected information on the reasons of incorrect classification of ¡50 sentences (50 sentences of each class). The sentences were subdivided into four groups (Table 8) based on the reason of incorrect classification:
♦ incorrect syntax tree parsing;
♦ incorrect detection of sentiment of a single word;
♦ imperfection of the rules;
♦ the sentiment is born by a high-level sentence structure.
Table 8. Error groups of the proposed recursive algorithm on OpenCorpora
Error % of positive sentences % of neutral sentences % of negative sentences
Incorrect syntax tree parsing 12 4 24
Incorrect single word sentiment 24 12 34
detection
Imperfection of the rules 26 22 2
Sentiment is beared by a high-level 38 62 40
sentence structure
Incorrect syntax tree parsing includes errors of part-of-speech tagging, constituency tree parsing, lem-matization, and any other errors done by a syntactic parser. For positive and neutral sentences it is the most infrequent group of errors, but almost a quarter of errors for negative sentences are caused by imperfect work of the syntactic parser. Presumably, the possible reason of this is more complex structure of negative sentences in comparison with positive and neutral ones.
Incorrect single word sentiment detection group includes errors occurred because the sentiment of a single word taken from the sentiment dictionary does not reflect its real sentiment. The first and most important cause of such errors is the imperfection of the sentiment dictionary. There are a lot of words having strong sentiment, such as больно (painful), истерика (hysteria), ломать (to break), взаимовыгодный (reciprocal), позабавить (to amuse), благодаря (thanks to), that are not present in the RuSentiLex dictionary. Words that changes their sentiments depending on context and domain are another important cause of incorrect single word sentiment detection. For example, the word исторический (historic) is generally neutral, but in context of importance of an event it is positive, e.g., исторический момент (historic moment). It should be mentioned that RuSentiLex-2017 provides some information on homonymy, but a separate study is needed to utilize this information for algorithm refinement. Errors caused by incorrect detection of the single word sentiment impacts mostly positive and negative sentences; only on tenth of neutral sentences classification errors belong to this group. The possible reason of such imbalance is that there is less sentiment words in neutral sentences than in positive and negative ones.
The last two groups of errors are connected to limitations of the algorithm. As we stated before, the semantic rules that the algorithm uses are rather simple and cannot detect the sentiment of every phrase correctly. Such imperfection of the rules leads to incorrect classification of sentences that principally can be correctly classified if the rules were perfect. For example, the neutral sentence Я, правда, не видела нового шестого фильма (I really have not seen the new sixth film yet) is incorrectly classified as negative because the applied rule processing не detects the negative sentiment of the part не видела нового шестого фильма, and the entire sentence is classified as negative. A small amount of negative sentences are classified incorrectly due to imperfection of the rules; for positive and neutral sentences these amounts are significantly greater. It seems that the proposed set of semantic rules is biased towards negative sentiment: it detects the negative sentiment of almost every phrase that is really negative, but also detects negative phrases that are not really negative. Analysis of the accuracy of the modifier words and phrases processing rules (Table 9) shows that the chosen modifier words are rather adequate.
The last errors group includes sentences that are classified incorrectly because the sentiment of some phrases cannot be detected only by the sentiments of its parts as the sentiment is born by a high-level sentence structure. For example, consider the sentence С возрастом детям становится всё более интересен и интернет: его регулярно посещают 15% младших школьников и почти 60% подростков (Growing children become more interested in the Internet: 15 percent of kids and 60 percent of teenagers use it regularly). It contains positive information on the growth of the Internet usage, but the general context and speech style indicate that the author does not expresses her opinion, only stating the fact, hence, the sentence is neutral. Such errors more often affects neutral sentences; possibly because significant part of neutral sentences are news that contain positive or negative facts, but general context of such sentences are neutral.
Summing up error analysis, we can state the following. At first, the groups of errors are distributed along the sentiment classes nonuniformly: errors caused by incorrect parsing lead to incorrect classification of negative sentences more often than positive or neutral sentences; correct single words sentiment detection is important for positive and negative sentences; imperfection of the rules affects the quality of positive and neutral sentences classification much stronger than the quality of negative sentences classification; a lack of high-level sentence structure analysis impacts all classes of sentences, but neutral sentences in particular.
Table 9. Accuracy of modifier words and phrases on OpenSentimentCorpus
Modification Words and phrases Accuracy, %
Make neutral and positive phrases negative; make negative phrases positive Не, нет, прекратить, разрушить, перестать, избегать, противоречить, уменьшать, утратить, ни один 56
Make negative phrases neutral; left other phrases untouched Понимать, принимать 76
Make neutral phrases positive; left other phrases untouched Способность, обеспечить, позволить, выполнить, помочь, увеличить, достичь, доказать, надеяться, очень, также 77
Makes positive phrases neutral; left other phrases untouched Крайне 100
Make phrases neutral Вроде бы, несмотря на 80
7. CONCLUSION
In this article we proposed and evaluated a recursive rule-based sentiment analysis algorithm for Russian. To improve the analysis of the clausal structure of a sentence we used the constituency parse tree, and proposed a set of semantic rules that determines the sentiment of a phrase by the sentiments of its parts.
The proposed algorithm was evaluated and compared with the rule-based approach, which was adapted for Russian in our previous work, on two corpora: the hotel reviews corpus and the multi-domain OpenSentimentCorpus that contains rather complex sentences from media. The experiments showed that the proposed recursive algorithm performs on the hotel reviews corpus slightly worse than the adapted rule-based approach: weighted ^¡-measures are 0.75 and 0.78 respectively. On OpenSentimentCorpus the performance of both ones is significantly worse than on the hotel reviews corpus (which is expected due to the complexity of its sentences and more subtle ways of expressing the sentiments), but the proposed algorithm performs better than the adapted approach does: ^¡-measures are 0.70 and 0.63 respectively. This indicates that in complex cases the proposed algorithm can perform significantly better than the adapted approach.
Despite the fact that the average performance on OpenSentimentCorpus is not very high, it should be taken into account that the set of semantic rules used in the experiments is rather simple, and the corpus characteristics are uncommon among corpora usually used in Russian sentiment analysis research. It also should be mentioned that the negative sentiment is detected with rather high quality, which may be a sign of some accordance of the semantic rules with the nature of how the negative sentiment is expressed.
Future research directions will be related to the refinement of the algorithm to achieve higher quality of the positive and neutral sentiments detection. The results of the error analysis suggests that this could be achieved by improvement of a single word sentiment detection (including the sentiment dictionary extension) and addition of new and more precise rules for various ways of sentiment expression.
FUNDING
This work was supported by ongoing institutional funding. No additional grants to carry out or direct this particular research were obtained.
CONFLICT OF INTEREST The authors of this work declare that they have no conflicts of interest.
REFERENCES
¡. Paramonov, I. and Poletaev, A., Adaptation of semantic rule-based sentiment analysis approach for russian language, 202130th Conf. of Open Innovations Association FRUCT, Oulu, Finland, 202!, IEEE, 202!, pp. ¡55—¡64. https://doi.org/¡0.239¡9/fruct53335.202¡.9599992
2. Wilson, T., Wiebe, J., and Hoffmann, P., Recognizing contextual polarity in phrase-level sentiment analysis, Proc. Conf. on Human Language Technology and Empirical Methods in Natural Language Processing, Vancouver, 2005, Stroudsburg, Pa.: Association for Computational Linguistics, 2005, pp. 347—354. https://doi.org/10.3115/1220575.1220619
3. Kien-Weng Tan, L., Na, J.-Ch., Theng, Yi.-L., and Chang, K., Sentence-level sentiment polarity classification using a linguistic approach, Digital Libraries: For Cultural Heritage, Knowledge Dissemination, and Future Creation. ICADL 2011, Lecture Notes in Computer Sciences, vol. 7008, Berlin: Springer, 2011, pp. 77—87. https://doi.org/10.1007/978-3-642-24826-9_13
4. Xie, Yu., Chen, Z., Zhang, K., Cheng, Yu., Honbo, D., Agrawal, A., and Choudhary, A., MuSES: Multilingual sentiment elicitation system for social media data, IEEEIntell. Syst., 2013, vol. 29, no. 4, pp. 34—42. https://doi.org/10.1109/mis.2013.52
5. Smetanin, S. and Komarov, M., Deep transfer learning baselines for sentiment analysis in Russian, Inf. Process. Manage., 2021, vol. 58, no. 3, p. 102484.
https://doi.org/10.1016Xj.ipm.2020.102484
6. Shaikh, M.A.M., Prendinger, H., and Ishizuka, M., Sentiment assessment of text by analyzing linguistic features and contextual valence assignment, Appl. Artif. Intell., 2008, vol. 22, no. 6, pp. 558—601. https://doi.org/10.1080/08839510802226801
7. Appel, O., Chiclana, F., Carter, J., and Fujita, H., A hybrid approach to the sentiment analysis problem at the sentence level, Knowl.-BasedSyst., 2016, vol. 108, pp. 110—124. https://doi.org/10.1016Zj.knosys.2016.05.040
8. Kahane, S. and Mazziotta, N., Syntactic polygraphs. A formalism extending both constituency and dependency, Proc. 14th Meeting on the Mathematics of Language (MoL 2015), Kuhlmann, M., Kanazawa, M., and Kobele, G.M., Eds., Chicago: Association for Computational Linguistics, 2015, pp. 152—164. https://doi.org/10.3115/v1/w15-2313
9. Gao, Y., Lou, J.-G., and Zhang, D., A hybrid semantic parsing approach for tabular data analysis,, 2019. https://doi.org/10.48550/arXiv.1910.10363
10. Li, J., Tan, H., and Bansal, M., Improving cross-modal alignment in vision language navigation via syntactic information, Proc. 2021 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Toutanova, K., Rumshisky, A., Zettlemoyer, L., et al., Eds., Association for Computational Linguistics, 2021, pp. 1041—1050. https://doi.org/10.18653/v1/2021.naacl-main.82
11. Marji, Z., Nighojkar, A., and Licato, J., Probing the natural language inference task with automated reasoning tools, The Thirty-Third Int. Flairs Conf., 2020.
https://doi.org/10.48550/arXiv.2005.02573
12. Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, Ch.D., Ng, A., and Potts, Ch., Recursive deep models for semantic compositionality over a sentiment treebank, Proc. 2013 Conf. on Empirical Methods in Natural Language Processing, Yarowsky, D., Baldwin, T., Korhonen, A., Livescu, K., and Bethard, S., Eds., Seattle, Wash.: Association for Computational Linguistics, 2013, pp. 1631—1642.
13. Tai, K., Socher, R., and Manning, C., Improved semantic representations from tree-structured long short-term memory networks, Proc. 53rd Annu. Meeting of the Association for Computational Linguistics and the 7th Int. Joint Conf. on Natural Language Processing, Zong, Ch. and Strube, M., Eds., Beijing: Association for Computational Linguistics, 2015, vol. 1, pp. 1556—1566.
https://doi.org/10.3115/v1/p15-1150
14. Zhang, Yu. and Zhang, Yu., Tree communication models for sentiment analysis, Proc. 57th Annu. Meeting of the Association for Computational Linguistics, Korhonen, A., Traum, D., and Marquez, L., Eds., Florence: Association for Computational Linguistics, 2019, pp. 3518—3527. https://doi.org/10.18653/v1/p19-1342
15. Yin, D., Meng, T., and Chang, K.-W., SentiBERT: A transferable transformer-based architecture for compositional sentiment semantics, Proc. 58th Annu. Meeting of the Association for Computational Linguistics, Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J., Eds., Association for Computational Linguistics, 2020, pp. 3695—3706. https://doi.org/10.18653/v1/2020.acl-main.341
16. Loukachevitch, N.V. and Levchick, A.V., Creating a general Russian sentiment lexicon, Proc. Tenth Int. Conf. on Language Resources and Evaluation (LREC'16), Calzolari, N., Choukri, Kh., Declerck, Th., , Eds., Portoroz, Slovenia: European Language Resource Association, 2016, pp. 1171—1176. https://aclanthology.org/L16-1186.
Publisher's Note. Allerton Press remains neutral with regard to jurisdictional claims in published
maps and institutional affiliations.
ISSN 0146-4116, Automatic Control and Computer Sciences, 2024, Vol. 58, No. 7, pp. 53-63. © Allerton Press, Inc., 2024. Russian Text © The Author(s), 2023, published in Modelirovanie i Analiz Informatsionnykh Sistem, 2023, Vol. 30, No. 1, pp. 86-100.
Annotation of Text Corpora by Sentiment and Irony in a Project
of Citizen Science
I. V. Paramonov", * (ORCID: 0000-0003-3984-8423) and A. Y. Poletaev", ** (orcid: 0000-0003-0116-4739)
a Demidov Yaroslavl State University, Yaroslavl, 150003 Russia *e-mail: ilya.paramonov@fruct.org **e-mail: anatoliy-poletaev@mail.ru Received February 3, 2023; revised February 24, 2023; accepted February 27, 2023
Abstract—This paper studies the construction of a corpus of sentences annotated by general sentiment into four classes (positive, negative, neutral, and mixed), a corpus of phrasemes annotated by sentiment into three classes (positive, negative, and neutral), and a corpus of sentences annotated by the presence or absence of irony. The annotation is conducted by volunteers within the project Preparing Texts for Algorithms on the People of Science website. Based on the available knowledge of the subject area for each of the problems, guidelines for the annotators are compiled. A methodology for the statistical processing of the annotation results is also developed based on analyzing the distributions and agreement measures of the annotations of different annotators. For annotating sentences by irony and phrasemes by sentiment, the agreement measures are quite high (the full agreement rate is 0.60-0.99), while for annotating sentences by general sentiment, the agreement is low (the full agreement rate is 0.40), apparently due to the higher complexity of the problem. It is also shown that the performance of automatic algorithms for sentence sentiment analysis improves by 12-13% when using a corpus on whose sentences all annotators (3-5 people) agree compared with a corpus annotated by only one volunteer.
Keywords: sentiment analysis, text corpus, statistical analysis, agreement measures, citizen science DOI: 10.3103/S0146411624700263
INTRODUCTION
The construction of annotated text corpora is an urgent problem in computational linguistics, because it provides input data for almost any research in this field. The solution of this problem is extremely laborintensive, because the construction of high-quality corpora requires a great deal of manual work. Since the annotation criteria are often difficult to formalize (e.g., in the case of text annotation by sentiment), it is desirable to annotate by experts in the subject area, which further increases the cost of this type of work.
Due to the high labor intensity, researchers often resort to automatic or semiautomatic text annotation (e.g., annotation of text reviews of products based on their ratings given by users on marketplaces), as well as hiring unskilled volunteer annotators. In both cases, the quality of the corpora will be lower than in the case of manual annotation by experts, but the corpora constructed in this way can still be useful in various studies after aggregating the collected data and at least selectively verifying them [1].
There is no general methodology for compiling guidelines for annotating text corpora for volunteers as well as for processing the results of the annotation. The compilation of guidelines is a rather difficult problem, because, on the one hand, they should cover enough topics to be useful for unskilled annotators in the vast majority of cases, and, on the other hand, they should not be too complex and voluminous to avoid difficulties with their understanding. The applicability of the obtained corpus for solving problems of computational linguistics directly depends on the quality of the methodology for processing the results of the annotation [2].
This paper studies the construction of a corpus of sentences annotated by general sentiment into four classes (positive, negative, neutral, and mixed), a corpus of phrasemes annotated by sentiment into three classes (positive, negative, and neutral), and a corpus of sentences annotated by the presence or absence of irony. The annotation was done by volunteers recruited as part of the project Preparing Texts to Algorithms on the People of Science website1.
1 https://citizen-science.ru/projects/gotovim-teksty-algoritmam.html
1. REVIEW OF RELATED WORKS
There are very few studies in the current scientific literature that focus on the annotation of text corpora and the aggregation of the results of this annotation. In most of the works that deal with the construction of annotated text corpora, the authors describe what they have done superficially or do not describe it at all. We can distinguish two main aspects related to this problem: the development of annotation guidelines and the methodology of annotation and aggregation of its results.
It should be noted that there are several general techniques for aggregating expert annotations, e.g., the Delphi method, but due to their rather high labor intensity, they are poorly applicable for aggregating annotations of a large number of sentences made by a large number of annotators.
The text [3] is one of the most significant works in the field of guideline construction for annotating texts by sentiment. It points out the potential complexity of the problem due to the vague definition of the notion of sentiment itself, provides a classification of sentences that are difficult to annotate, and proposes two approaches to the compilation of guidelines for annotators presumably making it possible to improve the accuracy of results in difficult cases. The first one uses a questionnaire containing statements characterizing certain classes of sentence sentiment and examples of sentence types belonging to the given class (admiration, support, sympathy, etc.). In the second case, annotators are asked to answer questions regarding the emotional state of the author of the sentence, the object in relation to which the sentiment is expressed, and the nature of the effect of the sentence on most people. It is noted that there is no best approach among these two approaches, and the effectiveness of each approach can be determined by the given problem and the skill of the annotators.
Another work by the same author [4] considers the problem of annotating tweets according to the author's attitude toward specific personalities or phenomena (for, against, and other) and whether an opinion is expressed about the object directly. Each tweet was considered by at least eight annotators, which is a very high value among all the existing works in the field. A total of 5412 tweets were collected; 5% of the tweets were originally annotated by their authors, and if an annotator made a mistake while annotating such a tweet, they were notified about it, and if they made a mistake on more than 30% of such tweets in aggregate, they were excluded from the work as dishonest. In aggregating the results, only ratings with which at least 60% of the annotators who worked with them agreed were taken. The degree of agreement was approximately 81.85% for the problem of determining the author's attitude toward a persona or phenomenon and 68.9% for the problem of determining whether an opinion about an object was expressed directly. The degree of agreement was defined as the average of the measures of agreement for each tweet, which in turn was computed as the ratio of the number of annotators who gave the most common response for a given tweet to the total number of annotators of that tweet.
In [5], the approach from [3] with some modifications was used. In particular, only 5- to 15-word sentences were considered. In addition, the idea of checking the annotator's adherence to guidelines, which was solved in the original article by using questions with known answers that the annotator did not know in advance, was replaced by giving participants a set of sentences annotated by the experts as an example before starting the work. Each sentence was annotated by a minimum of two annotators. If they agreed, the third one was not involved; if they disagreed, the third one was introduced. If even the third one did not resolve the problem, two more were introduced. The data of some annotators were excluded from consideration because they made too many mistakes at the first stage. A total of 15 744 sentences were annotated. Krippendorffs alpha [6] was used to evaluate agreement, which shows how close the agreement between the annotators is to a full one. Its value turned out to be 0.66, which indicates significant agreement. Four examples of difficult-to-analyze sentences similar to those described in [3] were also given.
A similar methodology was used in [7]: the annotating was performed by two experts; if the annotations did not converge, the decision was made by a third annotator. A total of 17572 Chinese sentences were annotated. Cohen's kappa [8] was used to evaluate agreement. Its value was found to be 0.71—0.73, which was considered satisfactory.
In another study, also for Chinese [9], a corpus of restaurant reviews was annotated for different aspects (food quality, service, price, etc.). For each aspect, annotators were asked to answer only one question: the attitude of the reviewer toward the aspect (positive, negative, or neutral). Each sentence was annotated by two independent annotators who underwent a brief training; their annotation results were then validated by a third annotator (if the annotations agreed) or by an expert (otherwise). Disputed cases were also handled by an expert. In the end, 46 730 sentences were annotated. The agreement characteristics were not computed.
In [10], a corpus of tweets related to specific brands was annotated. Each tweet was annotated by three annotators who had to indicate which emotions from the given set (trust, happiness, sadness, etc.) were
present in the given tweet. In order to evaluate the agreement of all annotators across categories, Fleiss' kappa was used [11], and Cohen's kappa was used to evaluate pairwise agreement between annotators. Both of these metrics characterize the proportion of matched annotators' ratings that cannot be explained by a random coincidence of opinions. The values of the metrics were 0.372 and 0.354, respectively, indicating moderate agreement among the annotators. The reasons for the disagreement of the annotators were not analyzed.
The article [12] considered the problem of annotating by sentiment a corpus of posts from the social network VKontakte on sociopolitical topics with response options: positive, negative, not sentimental, doubtful, and speech act (congratulations, technical messages, etc.). Only posts containing between 10 and 800 characters were annotated. Six experts in the field of linguistics were involved in the annotation and were given the corresponding guidelines. In the end, 31 185 posts were annotated, with each post being annotated by three annotators. Fleiss' kappa, whose value was 0.58, was used to assess agreement.
From the reviewed articles, we can see that in each case, corpora were annotated differently, and different metrics were used to assess agreement. There are some common features: in almost all the works each text fragment was annotated by at least three annotators (in some cases, only two annotations were used if the annotators agreed with each other); some guidelines for the annotators were used, which were quite similar in different works; actions were taken to clean the annotated corpus from questionable annotating results up to removing the results of the annotators who seemed to be dishonest. In most of the listed works, agreement metrics were considered the primary indicator of annotation quality. This can be because other analysis methods, such as verification by experts, are too labor intensive. However, the degree of agreement between the annotators is quite low, which is apparently due to the vagueness of the notion of sentiment and subjective perception of it by different people.
2. CORPUS ANNOTATION METHODOLOGY
OpenCorpora2, an open corpus of sentences from Russian news and journalistic texts in various subject areas, was used as the source corpus for annotating sentences by sentiment and irony. For the problem of annotating sentences by sentiment, only sentences of seven or more words were used. The list of phrase-mes and their meanings for annotating by sentiment was taken from Wiktionary3.
Annotators were tasked with evaluating sentences or phrasemes from an individual set for sentiment or presence of irony. Sets were automatically generated in such a way that, first, each annotator evaluated each sentence no more than once, and, second, each sentence or phraseme would end up in the sets of at least three annotators. The annotating sets also included the dictionary meanings of the phrasemes for the convenience of the annotators.
A briefing session was conducted for the annotators during which they were familiarized with the subject area and the annotating guidelines.
The guidelines for the problem of annotating sentences by sentiment stated that the sentiment of a sentence shows how the author treats the topic of the sentence. Four classes of sentence sentiments and a class of sentences with indefinable (doubtful) sentiment were identified and their distinguishing features were given, along with examples of sentences from each class. An excerpt from the guidelines in terms of describing the different classes is given below.
• Positive sentences in which the writer makes positive judgments or expresses positive emotions. A positive sentence can mention negative assessments or emotions, but the author must explicitly state that they do not share them. Examples of positive sentences:
Before no one could think that List'ev can act so relaxed on camera.
In his opinion, the management that organized the campaign for achieving the goal deserves the highest praise.
I really enjoyed the movie, and anyone who says it is boring just did not understand anything.
• Negative sentences in which the author makes negative judgments or expresses negative emotions. There is no information positively characterizing the topic, or it is irrelevant. Examples of negative sentences:
Overall, it turned out that everyone lost.
It turned out that the drain hole in the fancy shower stall was clogged and therefore it was completely impossible to take a shower.
2 https://opencorpora.org
3 https://ru.wiktionary.org
• Sentences with mixed sentiment, in which the author makes both positive and negative judgements and expresses emotions corresponding to the author's position. This class has been allocated specifically for sentences that are sentimental (expressing the author's emotions and containing the author's judgements) but are not unambiguously categorized as positive or negative. Examples of such sentences are as follows:
Akunin portrayed the heroes of the Seagull, weak and not particularly remarkable but still decent people, as vicious and criminal.
The deal between Google and RM received prominent media coverage in the summer but did not become a sensation.
Victor Pelevin writes much worse than Akunin, but Pelevin has something to say to the reader.
• Neutral sentences in which the author does not make their own judgments and does not express agreement with the judgments of others. Neutral sentences are unemotional, they simply convey dry facts, and the author's attitude to the topic is not shown. Examples of neutral sentences are as follows:
The information that the movie failed at the box office was not confirmed.
Traditionally, the host country of the world championship is exempted from the qualifying tournament.
The newspaper Prazsky Express claims that the actors behaved correctly and did not intend to disrupt the performance; they only demanded their fees to be paid.
• Sentences with an indefinable (doubtful) sentiment, in which the author's thought is unclear, e.g., because the sentence is incomplete and makes no sense without context. Examples of such sentences are as follows:
The author of the question is sure..., or history is repeating itself.
After all, the Russians could have withdrawn as far as the Urals.
Because this reading is too music from the beginning.
The guidelines also contained a number of general recommendations:
• Detect the author's positive or negative attitude out of context and the associations that arise.
• Classify as mixed sentiment sentences that clearly express the author's attitude but it is not clear whether it is positive or negative.
• Classify as neutral complete sentences containing a complete thought but not expressing the author's attitude; if the thought or the sentence is not complete, then classify the sentence sentiment as indefinable.
• When in doubt, focus on the emotional component of sentences.
• When in doubt, do not discuss any questions with other annotators in order to avoid influencing their perception.
The guidelines for the problem of annotating phrasemes by sentiment were made similarly; the main difference was that only three sentiment classes were retained, i.e., phrasemes with positive sentiment, phrasemes with negative sentiment, and neutral phrasemes. The mixed sentiment class was not used because phrasemes are much shorter than sentences and cannot exhibit both positive and negative attitudes at the same time. For a similar reason, the indefinable sentiment class was not used; i.e., because the phrasemes are quite short and the author's attitude is quite explicit, all cases where it is unclear whether the author's attitude is positive or negative belong to the neutral class.
The guidelines for the problem of annotating sentences for the presence of irony contained examples of sentences with irony. In addition, before starting the work, the annotators were required to study the definition of irony from the dictionary of linguistic terms edited by T.V. Zherebilo [13].
The characteristics of the obtained data are summarized in Table 1. A significant difference in the number of annotations of each annotator in the problems of annotating phrasemes by sentiment and sentences by irony is due to the fact that some annotators did not complete the tasks given to them.
3. METHODOLOGY OF STATISTICAL DATA PROCESSING
The following methodology was used to investigate the annotations made by the annotators.
1. The distribution of annotations of different annotators across classes was studied. Since the sentences and phrasemes were distributed randomly, and the volume of individual annotating sets was quite large, it can be expected that the shares of annotations of each class for all annotators will be approximately the same given that all annotators understood the task equally. Consequently, if the shares of annotations of any class differ significantly among the annotators, it gives grounds to believe that the annotators solved the problem in different ways. The distribution of annotations by classes was studied by the graphical method based on the histograms of the distribution of the share of annotations of this class and box plots
Table 1. Characteristics of corpora and their annotations
Annotation type Corpus volume Number of annotators Number of classes Number of annotations of each element Number of annotations of each annotator
Sentence/sentiment 11608 18 5 3-5 2214-2218
Phraseme/sentiment 1253 12 3 3-5 140-354
Sentence/irony 35013 14 2 1-3 1922-4000
constructed for each class. The chi-squared test was also used to check whether the distribution of annotations by class for each of the annotators coincided with the distribution by class of all annotations of all annotators. Nonmatching distributions also suggest that the annotator's way of performing the task was significantly different from that of the other annotators.
2. The agreement of annotations given by different annotators was evaluated. If the opinions of different annotators regarding the annotation of one sentence or phraseme often differ, this can indicate both a different understanding of the task by the annotators and difficulties in uniformly following guidelines, e.g., due to the excessively high role of individual perceptual features in the process of determining sentiment. Agreement was assessed on the basis of two measures: the proportion of sentences for which the opinions of all annotators who evaluated them agreed and Krippendorffs alpha. Krippendorffs alpha indicates how close the agreement between the annotators is to a full one and is calculated using the formula
a = Ao—Ae, 1 - Ae
where Ao is the number of annotators' ratings that matched and Ae is the number of annotations that would have matched if the annotators had annotated randomly [14]. In [15], the following empirical bounds for different degrees of agreement were distinguished:
— a < 0.2, slight agreement;
— 0.2 < a < 0.4, fair agreement;
— 0.4 < a < 0.6, moderate agreement;
— 0.6< a < 0.8, substantial agreement;
— a >0.8, near-perfect agreement.
It should be noted that Krippendorffs alpha takes into account differences in the annotations of several annotators, not only their full agreement. For example, it will be higher if among three annotators only two annotators' ratings matched than in the case where all three annotators gave different annotations.
3. For annotating sentences by sentiment, it was assessed how much agreement would improve if we discarded sentences whose sentiment was annotated as indefinable by at least one annotator and sentences whose sentiment was annotated as positive by at least one annotator and negative by at least one annotator. A significant (more than 0.05) improvement in the agreement measures would suggest that a significant part of the discrepancies in the annotators' ratings relate to such sentences. Since speech should not contain too many sentences whose sentiment cannot be determined and sentences that appear positive to one person and negative to another, if the annotation turns out to contain many such cases, it is possible that the annotation guidelines contain an incorrect or poorly formulated notion of sentiment.
4. RESULTS OF STATISTICAL PROCESSING
4.1. Results of Processing the Annotation of Sentences by Sentiment
According to the histograms of the distribution of the annotation shares of each annotation class and the box plot (Fig. 1), there are no obvious outliers that would indicate that one of the annotators solved the problem of identifying sentences of a certain class significantly differently from the others. Across all annotators, the proportions of positive, negative, and mixed sentiment class annotations are quite close to each other, with medians of 0.15 for positive, 0.30 for negative, and 0.10 for mixed sentiment. The proportion of neutral annotations varies quite significantly, from 0.30 to 0.80. The proportion of sentences whose sentiment could not be determined varies even more: for half of the annotators, it does not exceed 0.01, but for the other half, it is almost evenly distributed between 0.01 and 0.07; i.e., some annotators were
o ca CD
s c
ca
o
so
-Q
E £
0L
PARAMONOV, POLETAEV 8
0 0.2 0.4 0.6 0.8 1.0 Proportion of positive annotations
10 r
0 ca CD
n n a f o r e
1 £
a
0.2 0.4 0.6 0.8 1.0 Proportion of mixed annotations
8 r
to6
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.