Методы повышения производительности трансформеров на основе приближённого двудольного соответствия и сингулярного разложения тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Али Аммар

  • Али Аммар
  • кандидат науккандидат наук
  • 2025, «Национальный исследовательский университет ИТМО»
  • Специальность ВАК РФ00.00.00
  • Количество страниц 343
Али Аммар. Методы повышения производительности трансформеров на основе приближённого двудольного соответствия и сингулярного разложения: дис. кандидат наук: 00.00.00 - Другие cпециальности. «Национальный исследовательский университет ИТМО». 2025. 343 с.

Оглавление диссертации кандидат наук Али Аммар

РЕФЕРАТ

Synopsis

Introduction

CHAPTER 1. The Evolution of Machine Learning and Deep

Learning

1.1 Introduction to Machine Learning

1.1.1 Data Preparation and Feature Engineering

1.1.2 Basic ML Methods and Algorithms

1.1.3 Training Methods

1.2 Introduction to Deep Learning

1.2.1 Neural Networks and the Need of Deep Learning

1.2.2 Convolution Neural Networks

1.2.3 Transformers

1.3 Machine Learning for Computer Vision

1.3.1 Basic Computer Vision Problems

1.3.2 Loss Functions

1.4 Neural Networks for Graph Prediction

1.4.1 DETR for Object Detection

1.4.2 Object Detection and Graph Prediction Using CNNs

1.4.3 Graph Prediction Using Transformers

1.5 Neural Networks for Object Detection

1.5.1 2D Object Detection

1.5.2 3D Object Detection

1.5.3 Synthetic Data for Object Detection

1.6 Neural Networks for Action Recognition

1.6.1 2D Based Architectures

1.6.2 3D Based Architectures

1.6.3 Sampling Techniques for Action Recognition

1.7 Machine learning for Multi-View 3D Reconstruction

1.7.1 Feature Extraction and Matching

1.7.2 Non-Trainable Mathematical Modeling

1.7.3 E2E Trainable Methods

1.8 Machine Learning for Generative AI

1.8.1 Large Language Models

1.8.2 Vision Language Models

1.8.3 Merging Methods

1.8.4 Knowledge Transfer

1.9 Pruning and Scalability Techniques for Machine Learning

1.9.1 Pruning Methods for Large Language Models

1.9.2 Computational Issues and Training Costs

1.10 Advanced Driver Assistant and Monitoring Systems

1.10.1 AI Systems for Monitoring

1.10.2 Driver Assistant and Monitoring Systems

1.11 Principles of ML Based Driver Assistant and Monitoring Systems Development

1.11.1 Object Detection for Driver Monitoring

1.11.2 Head Pose Estimation for Driver Monitoring

1.11.3 Monocular Depth Estimation for Driver Monitoring

1.11.4 3D Object Detection for Driver Monitoring

1.11.5 Safe Speed Estimation for Driving Assistant Systems

1.11.6 Offline Localization

1.11.7 Sign Language Recognition

CHAPTER 2. Transformer-based architecture for E2E object

detection and graph prediction

2.1 Graph Prediction Method

2.1.1 The Generalization of Object Detection

2.1.2 Mathematical Modeling of the Graph Prediction Problem

2.1.3 Proposed Method for Graph Prediction

2.1.4 Approximated Bipartite Matching Loss

2.1.5 Results and Evaluation

2.1.6 Multi-Functional Neural Network Architecture

2.1.7 Results and Evaluation

2.2 Improved 2D Object Detection Pipelines

2.2.1 Cascade Object Detection

2.2.2 Results and Evaluation

2.2.3 Effective Ensemble Method for Object Detection

2.2.4 Results and Evaluation

2.2.5 Trusted AI Systems for Object Detection Using Synthetic Data162

2.3 3D Object Detection and Segmentation Architecture Improvement

2.3.1 CenterNet Baseline for 3D Object Detection

2.3.2 Improved CenterNet Architecture

2.3.3 Results and Evaluation

2.3.4 Two stages Method for 3D Segmentation

2.3.5 Results and Evaluation

CHAPTER 3. Improving transformer-based models efficiency

3.1 Training-Free Method for LLMs and VLMs

3.1.1 LLM Architecture

3.1.2 VLM Architecture

3.1.3 Delta Model and SVD Approach

3.1.4 Whitening Transformation with SVD

3.1.5 Results and Evaluation

3.2 Pruning and Compression Technique for Tranformers-Based Architecture

3.2.1 Baseline PruneMe Method

3.2.2 Compressor Method

3.2.3 Evaluation and Results

3.2.4 Shared Attention

3.2.5 Evaluation and Results

3.2.6 Transformer Simplification via Linear Transformations

3.2.7 Experiments and Results

3.2.8 Analysis

CHAPTER 4. Driver assistant system

4.1 Trusted AI Systems Development for Action Recognition

4.1.1 Multi-Stage Training Approach with Adaptive Sampling

4.1.2 Vision-SOLAR Method

4.1.3 Results and Evaluation

4.2 Multi-View 3D Reconstruction Method

4.2.1 Proposed Clustering Algorithm DBScan

4.2.2 Proposed Frames-Sorting Method

4.2.3 3D Filtering Using Fundamental Matrix Estimation

4.2.4 Results and Evaluation

4.3 Monocular Depth Estimation

4.3.1 Proposed Architecture

4.3.2 Data Preparation and Pseudo Labeling

4.3.3 Combined Loss Function Construction

4.3.4 Results and Evaluation

4.4 Safe Speed Estimation

4.4.1 Vehicles Around

4.4.2 Road Curvature and Width

4.4.3 Weather and Day/Night Classification

4.4.4 Speed Limit

4.4.5 Estimating Safe Speed

4.4.6 Results

4.5 Offline Localization

4.5.1 General Description

4.5.2 Data

4.5.3 Image Segmentation

4.5.4 Image Matching

4.5.5 Geometric Model Estimation

4.5.6 Image Retrieval

4.5.7 Results and Evaluation

Conclusion

Bibliography

Publications

10

РЕФЕРАТ

Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Введение диссертации (часть автореферата) на тему «Методы повышения производительности трансформеров на основе приближённого двудольного соответствия и сингулярного разложения»

Общая характеристика диссертации Актуальность темы

Архитектура трансформеров (Transformers) зарекомендовала себя как эффективный инструмент в различных областях искусственного интеллекта. Она продемонстрировала высокую эффективность в широком спектре задач: от обработки естественного языка до компьютерного зрения и легла в основу нового поколения систем ИИ, включая большие языковые модели (LLM) и модели генеративного ИИ. Современные прорывы, связанные с достижением state-of-the-art результатов, зачастую основаны на этой архитектуре или её гибридных модификациях.

Несмотря на достигнутый прогресс, трансформеры обладают недостаточно раскрытым потенциалом, реализация которого может привести к дальнейшему повышению эффективности и точности. Ключевым ограничением архитектуры остаётся механизм внимания, вычислительная сложность которого квадратично зависит от длины входной последовательности.

Данное исследование детально анализирует архитектуру трансформеров, уделяя особое внимание DETR — первому полностью End-to-End методу для детекции объектов. На основе архитектуры DETR была разработана модель PairDETR, впервые реализующая полностью end-to-end подход к предсказанию графов, модифицируя архитектуру для предсказания наборов узлов на графе. Введение концепции адаптивных относительных точек и разработка приближённой функции потерь на основе паросочетаний позволили получить полиномиальные решения для задач, традиционно относимых к классу NP-труд-ных..

В работе также исследованы стратегии ускорения трансформеров, направленные на снижение их ресурсоёмкости. Особое внимание уделено оптимизации как процесса обучения, так и этапа вывода с помощью передовых методов, не требующих дообучения.

Глубокий анализ архитектуры трансформеров выявил два ключевых наблюдения:

1) активации модели характеризуются высокой разрежённостью, что позволяет представлять веса в низкоранговом пространстве с помощью усечённого сингулярного разложения; 2) трансформеры демонстрируют скрытую линейность, что открывает возможность аппроксимации многих блоков линейными преобразованиями.

На основе первого наблюдения предлагается использовать усечённое сингулярное разложение (SVD) для представления дельта-весов между моделями, что позволяет осуществлять передачу знаний, аналогичную объединению моделей без дообучения. Опираясь на оба вывода, предложен новый метод сжатия трансформеров, в котором критерии важности блоков используются для выбора между их заменой на линейные преобразования или на низкоранговые представления. Также демонстрируется возможность совместного использования карт внимания между блоками и предлагается подход «shared attention» для частичного снижения квадратичной вычислительной сложности.

Методы протестированы на публичных наборах данных и бенчмарках в различных областях, включая большие языковые модели (LLM), визуально-языковые модели (VLM) и CLIP-подобные архитектуры. Практическая применимость продемонстрирована на примере системы помощи водителю, где предложенные подходы позволяют улучшить существующие компоненты и внедрить новые вспомогательные функции, такие как предсказание безопасной скорости на основе динамики окружающей транспортное среды.

Основная ценность работы заключается в совершенствовании одной из наиболее мощных и широко используемых архитектур машинного обучения — трансформеров. Повышая точность и производительность моделей при одно-

временном снижении аппаратных требований, данное исследование способно оказать значительное влияние на развитие современных систем ИИ.

Цель

Основная цель исследования — снижение вычислительной сложности и повышение производительности моделей трансформеров за счёт применения методов обрезки, основанных на сингулярном разложении матриц, для трансфера знаний без дообучения и оптимизации моделей, а также за счёт введения приближённых функций потерь на основе двудольного сопоставления для сквозного предсказания графов. Исследование охватывает задачи компьютерного зрения и генеративного ИИ, включая обнаружение объектов, предсказание графов и улучшение мультимодальных систем, с акцентом на большие языковые модели (LLM), визуально-языковые модели (VLM) и архитектуры, подобные CLIP.

Задачи

Для достижения поставленных целей исследования система была разделена на несколько ключевых компонентов и применена адаптированная исследовательская методология, включающая следующие этапы:

Задача 1 - Анализ современных исследований

Был проведён всесторонний анализ научных работ, опубликованных в ведущих международных журналах и конференциях. Этот анализ позволил сформировать теоретическую базу для диссертационного исследования и определить направления для дальнейшего развития существующих методов и алгоритмов.

Задача 2 Анализ и усовершенствование современных методов

В рамках данной задачи были изучены ключевые достижения и тенденции в развитии передовых методов и алгоритмов за последние годы. Сформулированы гипотезы, направленные на улучшение функциональных и вычислительных характеристик архитектур трансформеров. Гипотезы подтверждены с помощью математического анализа и экспериментальной проверки

Задача 3 Проектирование и проведение экспериментов

Разработанные методы и алгоритмы были протестированы на различных уровнях сложности, с использованием как публичных, так и специализированных наборов данных, включая данные, относящиеся к системам мониторинга и помощи водителю. Для оценки применялись общепринятые метрики и бенчмарки, соответствующие каждой задаче.

Задача 4 - Разработка, улучшение и внедрение алгоритмов в системы помощи и мониторинга водителя

Были проанализированы и адаптированы различные алгоритмы с целью повышения их точности и снижения вычислительной сложности. Предложены и реализованы усовершенствования, ориентированные на применение в системах помощи и мониторинга водителя. Разработанные решения протестированы как в рамках целевой системы, так и на стандартных задачах с использованием публичных наборов данных, что позволило оценить их обобщённость и устойчивость.

Задача 5 Валидация алгоритмов на публичных и собственных данных

На этапе разработки каждый алгоритм оценивался на общедоступных наборах данных для сравнения с современными подходами и подтверждения его эффективности. Для финальной настройки и тестирования в реальных условиях использовались собственные данные, накопленные в ходе многолетних наблюдений. Дополнительно проведено тестирование на внешних публичных данных

с целью обеспечения надёжности, устойчивости и полноты разработанных решений.

Методы исследования

В работе применяются методы искусственного интеллекта, алгоритмы анализа данных, математические методы представления информации, теория графов, статистическая теория, математический анализ и линейная алгебра.

Для каждой решаемой задачи проведён анализ современных подходов к аналогичным проблемам. На его основе предложены модификации на уровне данных, архитектуры нейронных сетей и стратегий обучения, направленные на повышение точности, устойчивости и снижение вычислительной сложности моделей.

Основные положения, выносимые на защиту

Согласно паспорту специальности, основные положения, представленные для защиты, можно суммировать следующим образом:

1. Метод определения объектов на изображениях и связей между ними, основанный на трансформерах, который расширяет архитектуру DETR за счёт использования адаптивных относительных точек для улучшения производительности и новой функции потерь для приближённого двудольного сопоставления, которая позволяет свести NP полную задачу к полиномиальной сложности.

2. Методы для моделей на базе архитектур трансформеров для объединения знаний нескольких моделей без дообучения за счет использования усечённого сингулярного разложения дельта-весов между двумя моделями, а также эффективному упрощению модели трансформеров за счет пред-

ложения критериев для замены блоков трансформера либо усечённым сингулярным разложением исходных весов, либо линейным преобразованием.

Научная новизна

Научная новизна исследования включает в себя 2 направления:

1. В области компьютерного зрения разработан новый метод, основанный на архитектуре трансформеров и подходу к предсказанию графов для решения задач детектирования объектов на изображениях и поиска ассоциаций. В рамках метода предложена новая функция потерь на основе приближённого двудольногосопоставления и улучшена нейросетевая архитектура ВБХК на основе использования адаптивных относительных точек.

2. Предложен метод для упрощения и объединения весов нейросетевых моделей с архитектурой трансформеров без переобучения, основанный на усечённом сингулярном разложении. В контексте упрощения моделей метод включает селективный критерий, который балансирует низкоранговое представление весов модели и линейные преобразования с использованием метрики важности. Кроме того, в области объединения весов метод применяется к дельта-весам последовательности моделей, что и позволяет осуществлять объединение весов без переобучения моделей.

Научно-техническая задача

Научная цель исследования заключается в том, чтобы проанализировать фундаментальные компоненты архитектуры, основанной на трансформерах, и методы с математической точки зрения и сформулировать гипотезы для их рас-

ширения и улучшения, проверив эти гипотезы посредством экспериментальных и математических методов. Техническая цель состоит в том, чтобы разработать эти системы и непосредственно протестировать их на системе мониторинга и помощи водителю, предложенной в этом исследовании.

Объект исследования

Объектом исследования являются изображения и видео, используя общедоступные наборы данных для решения каждой исследовательской задачи. В рамках архитектуры трансформеров, был сделан акцент на весах и активациях.

Предмет исследования

Предметом исследования являются методы и алгоритмы искусственного интеллекта для моделей на базе трансформеров,включая модели на основе DETR для обнаружения объектов и предсказания графов, мультимодальные модели (на основе CLIP), языковые модели, а также архитектуры, интегрирующие визуальную и языковую модальности, и методы сжатия моделей.

Теоретическая значимость

Заключается в глубоком анализе методов решения поставленных задач и их развитии на нескольких уровнях: от анализа данных и алгоритмов до модификации архитектур трансформеров и математических моделей, направленных на ускорение и повышение эффективности моделей трансформеров. Разрабо-

танные ключевые методы:

- Метод графовых предсказаний для архитектур трансформеров (PairDETR), в рамках которого происходит не только обнаружение объектов, но и определение взаимосвязей между ними.

- Метод объединения знаний моделей машинного обучения на основе транс-формеров посредствам математического моделирования дельта-модели, что значительно сокращает вычислительные затраты на обучение за счет работы с разницей весов (delta weights).

- Метод упрощения моделей на основе архитектуры трансформеров, основанный на замене слоев линейными преобразованиями или их представлении через усеченное сингулярное разложение весов, который превосходит существующие методы упрощения для архитектур трансформеров.

Практическая значимость

Для каждого предложенного метода были разработаны прототипы приложений, которые тестировались локально или участвовали в международных соревнованиях по машинному обучению, занимая первые места. На основе предложенных методов была разработана практическая система мониторинга и помощи водителю, испытания которой проводились в реальных условиях. Эксперименты подтвердили эффективность предложенных решений: система не только стала работать точнее, но и обрела новые функции, такие как оценка безопасной скорости и другие возможности. Данные в рамках системы мониторинга и помощи водителям собирались несколько лет в различных условиях. Разметка данных выполнялась как вручную, так и автоматически (с использованием псевдоразметки), что делает их ценным ресурсом для будущих исследований. Мультимодальная большая языковая модель выступает в роли основного интерфейса взаимодействия с водителем: с помощью функци-

ональных вызовов она активирует нужные сервисы системы в зависимости от запросов пользователя. Теоретические разработки, применённые в этой системе, включают:

- Архитектура Vision-SOLAR для дистилляции и масштабирования У1Х-трансформеров, что значительно сокращает затраты на обучение новых моделей (результаты проверены на открытых наборах данных).

- Комплексная гибридная функция потерь для моноокулярного оценивания глубины, позволяющая компактным моделям показывать точность, сопоставимую с крупными аналогами той же архитектуры.

- Модификация архитектуры CenteгNet, которая улучшает оценку поз и ЭЭ-реконструкцию, а также легко адаптируется для ЭЭ-сегментации с использованием базы ЭЭ-моделей.

- Улучшения традиционных методов ЭЭ-реконструкции, превосходящие современные End-to-End подходы (например, VGGSFM) и применимые к другим SOTA-методам.

- Адаптивная семплизация видео для классификации и распознавания действий, демонстрирующая высокую устойчивость в различных сценариях при использовании в качестве аугментации.

Эти результаты не только подтверждают эффективность предложенных методов, но и расширяют возможности их применения в реальных задачах, от автономного транспорта до промышленных решений.

Определение новых терминов и понятий

В разделе введены термины, которые будут использоваться на протяжении всего исследования, большинство из которых являются названиями новых предложенных методов и алгоритмов. Названия, используемые в исследовании,

согласованы с теми, которые используются в соответствующих научных работах.

- PairDETR: это новая модификация метода DETR, предназначенная для расширения до предсказания пары объектов в E2E формате, чтобы ихобнару-живать и определять связи между ними, в том случае, если она существует.

- SOLAR-vision — метод масштабирования крупных моделей, основанный на подходе SOLAR. Он расширяет модель по глубине, добавляя новые трансфор-мер-блоки в середину сети, что позволяет увеличить её размер и точность без обучения с нуля. Этот подход эффективно объединяет существующие знания с новыми слоями для улучшения производительности.

- Compressor: новый метод оптимизации весов модели (pruning).

- Healing: процесс донастройки (finetuning) модели на небольшом наборе данных после оптимизации для восстановления её производительности.

- Оценка 3D фундаментальной матрицы: оценка фундаментальной матрицы, применяемая к облаку точек после его двойной проекции в два различных 2D пространства, в результате чего получаются две фундаментальные матрицы.

- Shared attention: процесс совместного использования карты внимания между блоками трансформера для минимизации вычислений.

Достоверность

Достоверность положений подтверждается с помощью обоснованного использования методов, обоснования выбора задач и экспериментальных исследований, охватывающих передовые техники и алгоритмы. Полученные результаты признаны научным сообществом: они были опубликованы в высокорейтинговых научных изданиях и представлены на международных конференциях.

Внедрение результатов работы

Результаты исследований были использованы в следующих НИОКР проектах: российский научный фонд (РНФ 18-71-10065) и проект НИРМА Университета ИТМО №620176. Кроме того, многие из предложенных методов использовались для побед в международных соревнованиях, проводимых такими организациями, как NASA, Google, NOAA, SberAI, SberDevices, AIRI, EVRAZ, Чешский технический университет и РЖД.

Апробация результатов работы

Ключевые результаты исследований были представлены и обсуждены на следующих конференциях:

1. Конференция по компьютерному зрению и распознаванию образов (CVPR) 2024 - докладчик.

2. Конференция по компьютерному зрению и распознаванию образов (CVPR) 2024 - воркшоп.

3. AlJourney 2024 - докладчик/победитель.

4. AI Journey 2023 - докладчик/победитель.

5. Международная Российская Конференция по Умной Промышленности 2023 (SmartIndustryCon), 2023.

6. Международная конференция по обучению представлениям ICLR 2023.

7. 33-й Конференция Ассоциации Открытых Инноваций FRUCT, 2023.

8. 15-я Международная конференция "Интеллектуальные системы"(INTELS'22), 2023.

9. XI Конгресс молодых ученых 2022 года, 4-8 апреля 2022, Санкт-Петербург, Россия.

10. Научная и учебно-методическая конференция Университета ИТМО, 2-5 февраля 2022, Санкт-Петербург, Россия.

11. XXVIII международная конференция «28-я Конференция Ассоциации Открытых Инноваций FRUCT», 27-29 января, Москва, Россия.

12. AI Journey 2022 - докладчик/победитель.

13. 2-я Конференция Ассоциации Открытых Инноваций FRUCT, 2022.

14. AI Journey 2021 - докладчик/победитель.

15. XXVI международная конференция «26-я Конференция Ассоциации Открытых Инноваций FRUCT», 23-24 апреля, 2020, Ярославль, Россия.

Личный вклад автора

Разработка идей и методов, реализация и внедрение алгоритмов, обучение и развертывание нейронных сетей, а также написание обзоров литературы являются личным вкладом автора.

Структура и объем диссертации

Диссертация состоит из четырёх глав, как показано на рисунке 0.1. Первая глава посвящена обзору современных исследований и научных публикаций по тематике работы. В ней даётся введение в машинное обучение, охватывающее традиционные методы и основные инструменты, используемые при построении моделей. Далее рассматриваются основы глубокого обучения с акцентом на ключевые архитектуры нейронных сетей, в особенности на трансформеры, лежащие в основе современных state-of-the-art (SOTA) моделей в различных областях искусственного интеллекта.

Затем анализируются приложения компьютерного зрения, включая задачу распознавания объектов. Подчёркивается важность развития этого направления для решения более сложной задачи предсказания графовых структур, отражающих объекты и их взаимосвязи.

ш Теоретическое направление исследований

ш Практическое направление исследований

ш Конкретное направление исследований

Text Теоретический вклад автора

Тех! Практический вклад автора

Рисунок 0.1 — Структура диссератационной работы

Кроме того, в главе рассматриваются современные достижения в области больших языковых моделей (ЬЬМ), методы сжатия трансформеров (обрезка, квантование), а также подходы к тонкой настройке и объединению моделей. Ввиду практической направленности исследования на системы помощи водителю, приводятся основные компоненты, необходимые для их интеграции: распознавание объектов, ЭЭ-реконструкция, распознавание действий и оценка глубины по монокулярным изображениям.

Вторая глава посвящена применению архитектуры трансформеров для распознавания объектов и предсказания графов в задачах компьютерного зрения. Глава начинается с постановки задачи предсказания взаимосвязей между объектами на изображениях или видео. Представлена новая аппроксимация функции потерь на основе бикликового сопоставления, позволяющая свести КР-трудную задачу предсказания графов к полиномиальной сложности. Также предложены модификации архитектуры деформируемого внимания с использованием адаптивных относительных точек.

Далее представлены два метода ансамблирования для улучшения конвейеров распознавания объектов: один ориентирован на максимизацию точности,

второй — на снижение вычислительной сложности. Оба подхода разработаны как архитектурно-агностические. Далее рассматривается устойчивость систем распознавания объектов, анализируется концепция устойчивого распознавания и роль синтетических данных в повышении надёжности и качества моделей.

В главе также описано улучшение архитектуры CenterNet для SD-распозна-вания объектов: введена параллельная масштабируемая ветвь, улучшающая представление признаков глубины.

Третья глава посвящена оптимизации моделей на основе трансформеров с целью снижения времени вывода и уменьшения размера модели при сохранении высокой производительности, что позволяет развертывать их на устройствах с ограниченными вычислительными ресурсами.

В главе вводится применение усечённого сингулярного разложения (SVD) для низкорангового представления весов. Предложен гибридный подход, объединяющий SVD с селективными критериями важности блоков, позволяющий заменять компоненты трансформера либо на низкоранговые представления, либо на линейные преобразования. Это обеспечивает эффективное сжатие моделей без необходимости переобучения. Далее развивается расширение SVD, учитывающее информацию о трюкации (truncation-aware SVD), для анализа дельта-весов между обученными моделями, что предлагается использовать как метод передачи знаний без дообучения.

Кроме того, вводится механизм общего внимания (shared attention) между блоками трансформера, снижающий вычислительную сложность операций внимания при сохранении качества модели. Эта архитектурная модификация значительно уменьшает объём вычислений, связанных с механизмом внимания.

Четвёртая глава посвящена практическому применению разработанных решений в системе мониторинга и помощи водителю. В ней описываются ключевые инновации, реализованные в рамках практической части исследования, включая:

— модели распознавания жестового языка на русском языке;

— библиотеку SD-реконструкции;

— модели оценки глубины по монокулярным изображениям.

— Представлена новая система предсказания безопасной максимальной скорости, основанная на динамических параметрах окружающей среды — количестве транспортных средств, ширине дороги, дистанции между автомобилями и погодных условиях — с интеграцией установленных государственными органами статических ограничений скорости.

Дополнительно описана система помощи при оффлайн-локализации, предназначенная для работы при отсутствии доступа к интернету. Система функционирует за счёт сопоставления текущих изображений с визуальной базой данных городских ориентиров, что позволяет определять местоположение на основе визуальных признаков.

Содержание работы

Ключевые аспекты

На рисунке 0.2 представлены основные теоретические результаты данного исследования в сравнении с существующими работами. В главе II рассматривается расширение архитектур трансформеров, изначально предназначенных для обнаружения объектов, на задачу предсказания графов, дополненное углублённым теоретическим анализом путей повышения эффективности и обоб-щаемости подобных вычислительных конвейеров.

В главе III представлены методы сжатия и прореживания моделей, не требующие переобучения, а также новый подход к передаче знаний без обучения, основанный на анализе сингулярных значений с учётом данных и линейных преобразований. Эти методы разработаны и детально исследованы в третьей главе.

Четвёртая глава посвящена практической реализации системы, в которой предложенные методы интегрируются для построения комплексной системы помощи водителю. В рамках этой системы реализованы дополнительные

Обнаружение объектов и построение ассоциативных графов по изображениям

Трансформер или CNN Трансформер или CNN

Сжатие трансформеров

Методы SOTA

Матрицы весов блоков трансформера

Замена каждой матрицы

представлением низкого ранга -—

I—1из^^рангов^1е веса для блока трансформера с учетом данных

Требуется обучение

Предлагаемый Метод

Матрицы весов блоков трансформера

Замените каждый полный блок или набор блоков на отдельное линейное преобразование

>| ReplaceME j-

Линейные преобразования блоков весов

Не требует обучения Более высокая точность на более чем 20 открытых наборов данных

Баланс между глубиной и шириной сжатия

Предлагаемый Метод

Матрицы весов блоков трансформера

Compressor

Линейные преобразования блоков весов

—изкоранговые веса для блока трансформера с учетом данных

Не требует обучения Учитывает баланс между глубиной и шириной сжатия

Передача знаний трансформаторам

Методы SOTA

Предлагаемый Метод

Выполняется точная настройка моделей в наборах данных, чтобы получить унифицированную модель

Базовая LLM

Модель LLM, натренированная на написание программного кода Метод LORA J-

Мсдель LLM, натренированная на )

Унифицированная LLM

Требуется обучение

математические задачи

Объединение моделей с использованием низкоранговых представлений дельта весов с учетом данных

Базовая LLM

Модель LLM, натренированная на написание программного кода

Дельта-веса

>

SVD с поддержкой данных

Унифицированная LLM

Без обучения

Модель LLM, натренированная на математические задачи

Дельта-веса

Рисунок 0.2 — Основные теоретические результаты настоящей работы в сравнении с современными передовыми методами

функциональные возможности, что подтверждает применимость и масштабируемость предложенных решений в реальных условиях.

Содержание работы в деталях

Во введении выделены ключевые компоненты, рассматриваемые в рамках диссертационного исследования, а также определены основные научные вклады, вносимые в области компьютерного зрения и генеративного искусственного интеллекта. Подчёркивается значимость совершенствования каждого из этих компонентов, поскольку их развитие способствует повышению эффективности

Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК

Список литературы диссертационного исследования кандидат наук Али Аммар, 2025 год

Публикации

Основные результаты по теме диссертации представлены в 8 публикациях, индексированных в базе данных цитирования Scopus.

1. Ali A., Gaikov G., Rybalchenko D., Chigorin A., Laptev I., Zagoruyko S. PairDETR : Joint Detection and Association of Human Bodies and Faces. 2024 IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2024.

2. Kashevnik A., Ali A. 3D Vehicle Detection and Segmentation Based on EfficientNetB3 and Residual Blocks, //Sensors, 2022, Vol. 22, No. 20, pp. 7990 Q1.

3. Kashevnik A., Ali A. Light-Weight 2D Map Construction of Vehicle Environment Using Semi-Supervised Depth Estimation Approach, DSAA,

2022, 7, 15th International Conference "Intelligent Systems" (INTELS'22),

2023, pp.

4. Othman W., Kashevnik A., Ali A., Shilov N. Driver-MVT: In-Cabin Dataset for Driver Monitoring including Video and Vehicle Telemetry Information //Data, 2022, Vol. 7, No. 5, pp. 62.

5. Kashevnik A., Ali A., Lashkov I., Zubok D. Human Head Angle Detection Based on Image Analysis //Advances in Intelligent Systems and Computing, 2021, Vol. 1288, pp. 233-242.

6. Kashevnik A., Ali A., Lashkov I., Shilov N. Seat Belt Fastness Detection Based on Image Analysis from Vehicle In-Cabin Camera //Proceedings of the 26th Conference of Open Innovations Association FRUCT, 2020, pp. 143-150.

7. Kashevnik A., Ali A. Vehicle Offline Localization Based on Computer Vision: an Approach Based on Image Matching Retrieval Algorithms and Implementation //Proceedings of the 33nd Conference of Open Innovations Association FRUCT, 2023, pp. 125-131.

8. Kashevnik A., Ali A. Driving Safe Speed Estimation Based on Outside Environment Vision Analysis //Proceedings of the 32nd Conference of Open Innovations Association FRUCT, 2022, pp. 121-127.

47

Synopsis

General thesis summary

Relevance

The Transformer architecture has proven to be a highly effective tool across various fields of artificial intelligence, demonstrating remarkable performance in a wide range of tasks—from natural language processing to computer vision. It has become the foundation of the new generation of AI systems, including large language models (LLMs) and generative AI models. Modern breakthroughs achieving state-of-the-art (SOTA) results are often based on this architecture or its hybrid modifications.

Despite the significant progress achieved, the potential of Transformers remains underexplored, and unlocking it could lead to further improvements in efficiency and accuracy. A key limitation of the architecture lies in the attention mechanism, whose computational complexity scales quadratically with the length of the input sequence.

This research provides a detailed analysis of the Transformer architecture, with particular focus on DET R—the first fully end-to-end method for object detection. Building upon the DETR architecture, we developed PairDETR, which implements, a fully end-to-end approach to graph prediction by modifying the architecture to predict sets of nodes in a graph. The introduction of adaptive relative points and the development of an approximate bipartite matching loss function enabled polynomial-time solutions to the problem that traditionally classified as NP-hard.

The study also investigates strategies for accelerating Transformers, aimed at reducing their computational and resource demands. Special focus is given to optimizing both the training and inference phases using advanced, training-free techniques.

A deep analysis of the Transformer architecture reveals two key observations: 1) Model activations exhibit high sparsity, allowing weights to be represented in a low-rank space through truncated singular value decomposition (SVD); 2) Transformers demonstrate hidden linearity, opening the possibility of approximating many of their blocks with linear transformations.

Based on the first observation, we propose using truncated SVD to represent delta weights between models (residuals), enabling knowledge transfer analogous to model merging without retraining. Drawing on both insights, we introduce a novel Transformer compression method, in which block importance criteria are used to decide whether to replace specific blocks with linear transformations or low-rank representations which balance depth and width pruning. Furthermore, we demonstrate the feasibility of sharing attention maps across blocks and propose a "shared attention" approach to partially alleviate the quadratic computational complexity.

The proposed methods are evaluated on public datasets and benchmarks across multiple domains, including large language models (LLMs), vision-language models (VLMs), and CLIP-like architectures. Practical applicability of this research is demonstrated in a driver assistance system, where the proposed techniques improve existing components and enable new auxiliary functions, such as predicting safe speed based on the dynamics of the surrounding traffic environment.

The primary value of this work lies in advancing one of the most powerful and widely used architectures in machine learning—Transformers. By enhancing model accuracy and performance while simultaneously reducing hardware requirements, this research has the potential to impact the development of modern AI systems.

The goal of the study

The primary aim of this research is to enhance the efficiency and performance of transformer-based models using truncation-aware singular value decomposition for training-free knowledge transfer and models compression and introduce

approximated bipartite matching loss for end-to-end graph prediction. This encompasses tasks such as object detection, graph prediction, and advancements in generative AI, specifically focusing on large language models (LLMs), vision-language models (VLMs) multi-modal models such as CLIP.

Objectives

To achieve the research objectives, the system was divided into several key components, and an adapted research methodology was applied, consisting of the following steps:

Objective 1: Analysis of Current Research A comprehensive analysis was conducted of scientific publications from leading international journals and conferences. This analysis established the theoretical foundation for the dissertation research and identified directions for further development of existing methods and algorithms.

Objective 2: Analysis and Enhancement of State-of-the-Art Methods Key advancements and trends in the development of cutting-edge methods and algorithms over recent years were examined. Hypotheses aimed at improving the functional and computational characteristics of Transformer architectures were formulated and validated through mathematical analysis and experimental evaluation.

Objective 3: Design and Conduct of Experiments

The developed methods and algorithms were tested at varying complexity levels using both public and specialized datasets, including data related to driver monitoring and assistance systems. Established metrics and benchmarks, appropriate to each task, were employed for evaluation.

Objective 4: Development, Improvement, and Integration of Algorithms into Driver Assistance and Monitoring Systems Various algorithms were analyzed and adapted to enhance their accuracy and reduce computational complexity. Proposed improvements were specifically oriented toward

deployment in driver assistance and monitoring systems. The developed solutions were evaluated both within the target system and on standard benchmark tasks using public datasets, enabling assessment of their generalization and robustness. Objective 5: Validation of Algorithms on Public and Proprietary Data

During development, each algorithm was evaluated on publicly available datasets to enable comparison with state-of-the-art approaches and confirm its effectiveness. For final tuning and real-world testing, proprietary data collected over multiple years of observation were utilized. Additionally, external public datasets were used to ensure the reliability, robustness, and comprehensiveness of the developed solutions.

Research methods

In this research, various methods were used to achieve its objectives and to develop the proposed methods and algorithms. These include essentially artificial intelligence theory, data analysis algorithms, mathematical methods for representing data, graph theory, statistical theory, calculus, and linear algebra. For each problem, we explored the methods used to solve similar ones. Through an analysis of these methods, as followed in the research, we proposed modifications at different levels: data level, neural network architecture level, and training methods level, aimed at improving the approach used to solve this problem.

Assertions that are presented for defense

According to the specialty passport, we can summarize the main assertions that are presented for defense as follows:

1. Transformer-based graph prediction method that extends DETR architecture by proposing adaptive relative points to improve the performance and

introducing an approximated bipartite matching loss to map the loss calculation from NP hard to polynomial time.

2. Training-free methods for transformer-based models using truncation-aware singular value decomposition of delta weights between two models for resource effective knowledge transfer (fusion) as well as proposal of selective criteria to replace transformer blocks either with truncation-aware singular value decomposition of original weights or a linear transform resulting in effective model compression.

The novelty of research

The Scientific Novelty of the research is classified into 2 main points:

1. In the domain of computer vision, an End-to-End Transformer-based graph prediction method has been developed to address object detection and associations. This method incorporates an approximated bipartite matching loss and utilizes adaptive relative points within the deformable DETR architecture.

2. Training-free method based on a truncation-aware singular value decomposition tailored for transformer-based architectures. In the context of pruning, the method integrates a selective criterion that balances the low-rank representation of Transformer model weights and linear transformations using an importance metric. Furthermore, within the domain of knowledge transfer, this method is applied to the delta weights across a sequence of models, enabling a training-free transfer of knowledge.

The scientific and technical objective

The scientific aim of the research is to analyze fundamental components of the transformer-based architecture, and methods mathematically and to construct hypotheses for extending and improving them, verifying these hypotheses through experimental and mathematical methods. The technical goal is to develop these systems and directly test them on the driver monitoring and assistance system proposed in this research.

The research object

In this study, we focus on analyzing images and videos, utilizing publicly available datasets to address each research problem. And on the concept of transformers architecture we focus on the weights and activations as objects.

The research subject

The essential artificial intelligence methods and algorithms for transformer based models. Including DETR based architecture for object detection and Graph prediction, multi modality (CLIP based models), Large language models, Vision language models, and compression methods.

The theoretical significance

The theoretical significance of the research lies in the in-depth study of the methods used to solve the problems and their development on multiple levels including data analysis, algorithms, modified transformer architectures and mathematical representations to accelerate and enhance performance of transformer based models as following:

- The proposed transformer-based graph prediction method (PairDETR) is universal and it is a step forward into modern detection systems where the system is able to detect the objects and define the relation between these objects.

- Knowledge transfer (Fusion) with a training-free approach based on mathematical modeling of the delta model is proposed and it minimizes the training requirements.

- The proposed depth compression method based on representing cutting-layer as a linear transformation or replacing it with a truncation-aware singular value decomposition of the weights is new and outperforms all previous depth compression methods for transformers.

The practical significance

First, mini applications were developed for each proposed algorithm and method, and they were evaluated locally or were developed during ML international competitions and acquired top-1 place. Then, a practical driver monitoring and assistance system was developed using the proposed methods, and experiments were conducted in a real-world environment. The experiments proved the effectiveness of the methods in improving the performance of the system, providing it with

new functionalities such as safe speed estimation system and other new features. Secondly, data for driver monitoring systems was gathered over several years under different conditions, and the ground-truth of the data was prepared manually for some tasks and automatically (using pseudo-labeling techniques), making it available for future research. A multi-modal large language model acts as a primary interface to interact with the driver, utilizing function calls to initiate each system service according to user requests. Some of the theoretical work was developed for the practiclar system in particular including:

- Distilling or expanding ViT transformers using vision-SOLAR method shows promising results to minimize the training required for new different transformers and verified on open datasets.

- Complex-hybrid loss function for Monocular depth estimation can improve superlight models to perform almost equivalent to large models following the same architecture.

- The proposed modification on CenterNet show much better improvement for pose estimation improving the 3D reconstruction, and has adaptability to support 3D segmentation easily using a 3D models database.

- The modifications proposed for traditional 3D reconstruction methods outperforms state-of-the-art end-to-end methods such as VGGSFM and adaptable to be applied on SOTA methods. making it a general improvement.

- Adaptive sampling technique for video classification and action recognition shows much robust results on different scenarios when used as an augmentation through training process.

- The proposed 2D object detection pipeline improvement is generalizable and tested on different systems, a new research on the use of synthetic data for object detection is conducted by Rice University on top of our work.

Definitions of new terms

In this section, we explain the meaning of some of the terms that will be used throughout the research, most of which are names for new proposed methods and algorithms. The names used in the research are consistent with those used in the related scientific papers.

- PairDETR: is a modification over DETR based methods, to extend it to E2E pair prediction, so that it will detect pairs and define the relation in between if exists.

- SOLAR-vision: following SOLAR method used to expand large language models we define SOLAR-vision method that extends vision models such as ViT towards larger and more accurate models by stacking multiple transformer blocks in between old ones to expand the model without the need of scratch training.

- Compressor: a new pruning method

- Healing: a process of finetuning a model on a small set of data after pruning to recover performance.

- 3D Fundamental Matrix Estimation: is a fundamental matrix estimation applied to a point cloud after projecting it twice into two different 2D spaces resulting in two fundamental matrices.

- Shared attention: the process of sharing the attention map in between transformer blocks to minimize the calculations.

Reliability of scientific achievements

Achievements are confirmed through the correct use of methods, the justification of task selection, and experimental studies that cover advanced techniques and algorithms. The results obtained are recognized by the scientific community: they are published in articles and presented at conferences.

Implementation of research results

The research results were used in the following R&D projects: Russian Science Foundation, RSF 18-71-10065 and the NIRMA 620176 project. Also many of the proposed methods are used in winning international competitions hosted by NASA, Google, NOAA, SberAI, SberDevices, AIRI, EVRAZ, Czech Technical University and RZD.

Approbation of research results

Key research results were presented and discussed at the following conferences:

1. Conference on Computer Vision and Pattern Recognition (CVPR) 2024 Speaker

2. Conference on Computer Vision and Pattern Recognition (CVPR) 2024 workshop

3. AIJourney 2024 speaker/winner

4. AI Journey 2023 speaker/winner

5. 2023 International Russian Smart Industry Conference (SmartIndustryCon), 2023

6. 2023 International Conference on Learning Representations ICLR 2023

7. 33rd Conference of Open Innovations Association FRUCT, 2023

8. 15th International Conference "Intelligent Systems" (INTELS'22), 2023

9. XI Конгресс молодых ученых 2022 года, 4-8 Апреля 2022, Санкт-Петербург, Россия.

10. Научная и учебно-методическая конференция Университета ИТМО, 2-5 февраля 2022, Санкт-Петербург, Россия.

11. XXVIII международная конференция «28th Conference of the Open Innovations Association FRUCT», 27-29 Января, Moscow, Russia.

12. AI Journey 2022 speaker/winner

13. 2nd Conference of Open Innovations Association FRUCT, 2022

14. AI Journey 2021 speaker/winner

15. XXVI международная конференция «26th Conference of the Open Innovations Association FRUCT», 23-24 Апреля, 2020, Yaroslavl, Russia.

Personal contribution of the author

Construction of ideas and methods, the development and implementation of algorithms and methods, training and deploying the neural networks, and the writing of literature reviews are personal contribution of the author.

Thesis structure

The thesis is primarily divided into four main chapters, ash shown in Figure 0.13. The first chapter summarizes research and scientific papers addressing the fundamental components under study. It begins with an introduction to machine learning covering traditional methods and the primary tools used for building machine learning models. Next it explores deep learning focusing on key architectures for constructing deep neural networks, with particular emphasis on transformer-based architectures the cornerstone of modern state-of-the-art (SOTA) models across various domains. The discussion then shifts to applications in computer vision including direct uses in object detection. It highlights the importance of advancing this field to address the broader challenge of defining graphical structures that identify objects and their interrelationships.

Additionally the chapter reviews the latest research on large language models, transformer pruning techniques, and contemporary methods for fine-tuning and model merging. Since the practical work of this research involves driver assistance systems, the chapter also examines core components to be integrated into such

Рисунок 0.13 — Structure of the thesis

systems, including object detection, 3D reconstruction, action recognition, and monocular depth estimation. The second chapter focuses on transformer-based object detection and graph prediction in computer vision. It begins by introducing the challenge of predicting relationships (building graphs) between objects in images or video, then presents an approximated bipartite matching loss that simplifies graph prediction from NP-Hard to a P class problem, along with modifications to the transformer deformable attention architecture using an adaptive relative points approach. Next it proposes two ensemble methods to enhance object detection pipelines one optimized for accuracy and another for runtime efficiency. Both designed to work regardless of the model architecture. The discussion then shifts to robustness in object detection exploring the concept of robust detection and the role of synthet ic data in improving model performance. Additionally, the chapter introduces an improvement to the CenterNet architecture for 3D object detection. The modification adds a parallel scaling branch to enhance depth feature representation. The third chapter focuses on optimizing transformer-based models by accelerating their inference speed and reducing model size, while maintaining near original performance with lower hardware and memory requirements. The

chapter first introduces Truncation-aware Singular Value Decomposition (SVD), then presents an integrated approach combining this method with a selective criteria mechanism, this hybrid approach replaces transformer blocks with either a low-rank representations of the weights or linear transformations, which results in an effective training-free compression method, the study further extends truncation-aware SVD to analyze delta weights between trained models proposing its application as a training-free knowledge transfer technique. Additionally, the research proposes a shared attention mechanism across transformer blocks reducing computational overhead while maintaining model efficiency. This architectural modification significantly decreases the computational requirements of the attention operations. The fourth chapter primarily discusses the practical application of the proposed driver monitoring and assistance system. It begins by highlighting the key innovations developed in the practical part of the research including: Russian sign language recognition models, 3D reconstruction library and monocular depth estimation models at the component level. At the system level, a system for predicting safe maximum speed values based on dynamic changes in the surrounding environment (e.g., the number of vehicles, road width, inter-vehicle distance, and weather conditions), while also integrating static speed limits set by traffic authorities. Additionally, the chapter introduces an assistance system for offline localization in case of internet access loss. This system works by linking images together and recognizing specific city landmarks using a visual database of these areas.

The content of the work

Key Aspects

Figure 0.14 presents the main theoretical contributions of this research. Chapter II discusses the extension of transformer architectures—originally designed for object

Object Detection and Graph Prediction in Images

Proposed Method (End-to-End)

Transformer or CNN

Transformer or CNN

SOTA Methods

Transformers Pruning

SOTA Methods

[Weights of the transformer]

Replace each matrix with a low-rank approximation

>| SVD-LLM j-

JData-aware low-rank representaion of the | \ weights of the transformer block J

Needs finetuning

Proposed Method

Weights of the transformer blocks

Replace each transformer block or a set of consecutive blocks by a linear transform

>| ReplaceME j-

Linear transform of the transformer blocks

Training-free Better performance in more than 20 open benchmarks

blocks

Proposed Method

Weights of the transformer blocks

Balnace between depth and width pruning -w Compressor \-

Linear transform of the transformer blocks

Data-aware low-rank representaion of the weights of the transformer block

Training-free Balance Between depth and width prungin

M OR

Knowledge-transfer for transformers

SOTA Methods

Fine-tuning of models in datasets is performed to obtain a unified model.

LLM model that is trained for programming related tasks LORA J-

LLM model that is trained for mathimatical

]->C

Universal LLM

Needs finetuning

Prposed Method

Merging models based on a data-aware SVD on the residuals _(delta weights)_

LLM model that is trained for programming related tasks

Delta-weights

>

Data-aware SVD

Merge

Universal LLM

LLM model that is trained for mathimatical related tasks

Delta-weights

Training-free

Base LLM

related tasks

Base LLM

Рисунок 0.14 — Main theoretical results of the present work.

detection—to the task of graph prediction, complemented by an in-depth theoretical analysis of pathways to enhance the efficiency and generalizability of such computational pipelines.

Chapter III introduces training-free methods for model compression and sparsification, as well as a novel approach to training-free knowledge transfer based on singular value analysis with data adaptation and linear transformations. These methods are developed and thoroughly investigated in the third chapter.

Chapter IV focuses on the practical implementation of a system in which the proposed methods are integrated to build a comprehensive driver assistance system. Within this framework, additional functional capabilities are implemented, demonstrating the applicability and scalability of the proposed solutions under real-world conditions.

The detailed content of the work

In the introduction of this research, we outline the key advancements and improvements to be developed, including: the generalization of object detection to graph prediction using transformer-based architectures and the introduction of training-free pruning and knowledge transfer techniques for these type of neural networks. Additionally, we highlight the most significant contributions of this research in each studied area, covering both theoretical and practical aspects.

The first chapter of this research focuses on providing a summary of the previous research for each component within this study including both theoritical contributions and practical models used for the driver assistant system. This chapter starts with an overview of machine learning in general and the traditional ML methods followed in Section 1.1 on how to process data, along with the algorithms used to build a machine learning system to solve simple problems and the training methods used in such systems. In Section 1.2, we move on to discuss deep learning, with an introduction to the general meaning of deep learning, as well as a discussion of the most commonly used network architectures that will be employed in the subsequent experiments in this research, such as convolutional networks and transformers in which we mainly explain the attention mechanism and its advantages and bottleneck.

In Section 1.3, we proceed to discuss the role of machine learning in the field of computer vision by defining the most prevalent computer vision tasks and the loss functions used to train networks to solve such tasks.

Following the foundational introduction to artificial intelligence and machine learning, Section 1.4 dives into the core tasks that will be discussed in this research, starting with the issue of graph prediction in images and the relations and associations between objects. We first explore the use of transformers for object detection, along with convolutional networks and transformers that have been developed to tackle graph prediction problem. In section 1.5, we discuss research related to neural networks used for solving the problem of object detection in both

2D and 3D, as well as how data generated using modern artificial intelligence tools (synthetic data) can affect the accuracy of object detection systems if used during training as additional.

Then, we continue in the chapter to discuss research related to video analysis for action recognition tasks and which neural network architectures used to solve this type of problems, as well as the proposed methods for analyzing video data and selecting frames efficiently in section 1.6 (sampling techniques).

3D reconstruction is an important feature for our practical system, in section 1.7 we discuss the methods used to solve this problem in multi-view scenario where we go through the methods used for feature extraction and matching, non-trainable and trainable mathematical modeling of non-differential operations.

The use of large language models and vision language models is of great importance due to their ability to perform a variety of tasks effectively and operate as valuable assistants. In section 1.8, we provide a summary of how these networks operate, as well as methods for integrating and developing them, in addition to training methods and transferring knowledge between them.

However, one of the biggest challenges facing this type of neural network is their large amount of parameters. In section 1.9, we discuss how to compress and reduce their size, as well as the high costs associated with their training and operation. We end this chapter in sections 1.10 and 1.11, by discussing driver assistance and monitoring systems in general, starting with how artificial intelligence is used in many monitoring systems, followed by showcase of the role for each discussed components in the development of driver monitoring and assistance systems.

In the second chapter of the research, the focus is on transformer-based object detection and graph prediction in computer vision. It begins by introducing the challenge of predicting relationships (building graphs) between objects in images or video, then presents an approximated bipartite matching loss that simplifies graph prediction from NP-Hard to a P class problem, along with modifications to the transformer deformable attention architecture using an adaptive relative points approach. Next it proposes two ensemble methods to enhance object detection pipelines one optimized for accuracy and another for runtime efficiency. Both

designed to work regardless of the model architecture. The discussion then shifts to robustness in object detection exploring the concept of robust detection and the role of synthetic data in improving model performance. Additionally, the chapter introduces an improvement to the CenterNet architecture for 3D object detection. The modification adds a parallel scaling branch to enhance depth feature representation.

Section 2.1 of this chapter begins with a discussion about the problem of graph detection in images, specifically the pairwise association between objects. We analyze this problem and discuss how to develop a transformer architecture based on the DETR method to solve it. It is explained that the mathematical representation of the problem poses a particular challenge in training, as it is classified as a NP-hard known as maximum flow with fixed costs. We present an approximate solution using dummy nodes, and we propose adaptive relative points to enhance the accuracy of these systems. The results of the research show that the proposed method significantly contributes to improving the field of graph and pair prediction. The proposed system was tested on face and body detection and their association

Ground Truth

Рисунок 0.15 — Proposed DETR adjustment to do object detection and

association.

with each other, primarily because these are among the most common issues in this field, and they also present a significant challenge in highly crowded areas. In the Fig. 0.15, we illustrate the process of how the system trained. The transformer model predicts pairs of boxes and links them automatically. We then apply the proposed approximate bipartite matching between the pairs and the ground truth during training to minimize the matching cost, then calculating the loss function taking into account the dummy nodes when the ground-truth doesn't include pairs but a single

object. The proposed system was tested on different publicly available datasets, and it was found that it significantly improves the association metric (mMR-2). Furthermore, in the study, we present a detailed analysis of the impact of each modification individually. We continue in this section on how to develop this system and add other tasks that can be solved using the same neural network (building a multi-functional model). After training the neural network for object detection and association, we froze the neural network and added additional decoder to solve the problem of head pose estimation. this additional decoder was trained on our internal dataset as illustrated in the Fig.0.16.

The resulting model is capable of doing graph prediction and head pose estimation

set of Ponts predictions

Рисунок 0.16 — Proposed DETR adjustment for pairdetection and head pose

estimation

with competitive results with the baseline, introducing one unified model for multitasking that solves (detection, association and head pose estimation).

In section 2.2, we focus on object detection part, mainly on improving the overall inference structure of these neural networks. Our proposed improvement is applicable to any type of object detection models. However, in our experiments, YOLO as a CNN architecture and Rt-DeTr as transformer based were primarily used.

The first proposed inference improvement relies on a cascade object detection pipeline. At the first stage, small, quantized neural networks are used since the accuracy of object detection is not of critical importance. The focus on this stage is on recognizing the object area in which f-score is critical. This is followed by a

second stage where the object area is cropped from the image and sent to another network that performs refinement to precisely determine the object's location, as illustrated in the Fig. 0.17. As demonstrated by the experiments, the proposed

( TTA with Hfilp )

Рисунок 0.17 — Cascade object detection method

system significantly improved the accuracy of object detection, especially in cases involving multi-scale objects. The primary goal of the initial stage is to locate the objects area, and after the process of cropping and approximation, most objects become similar in scale, which facilitates precise localization for the refiner.

The capability to prune and compress the networks in the first stage makes the system considerably faster than a single network that has similar performance, as it allows training the refiner on relatively small image resolution while maintaining high accuracy. In Table 10, we show the advantage of our proposed method that

Таблица 10 — Ablation study for pipeline development, inference time given for the full testing-set

Model Refiner Inference time in minutes JI public JI private

Yolov8n None 54 0.8905 0.8775

Yolov8s None 96 0.8947 0.8826

Yolov8m None TLE - -

Yolov8n/s/m-int8 Yolov8m 105 0.9159 0.9057

Yolov8n/s/m-int8 Rt-Detr 106 0.9147 0.9063

was mainly developed on Posebowl competition by NASA, and we were able to acquire first place. As shown the latency of the method is almost as using a single

small YOLO version while having much better performance. We continue in the section by proposing an effective method for merging the outputs of neural networks (ensemble) for object detection. We propose an efficient ensemble algorithm to merge the outputs which relies on two main algorithms for merging the detected boxes based on the image scale used in inference: Non-Max Suppression and Weighted Box Fusion. We discuss that by using both methods together, we can achieve the best results and propose an appropriate approach to decide when each method should be used. In table 11, we show the results of our proposed ensemble pipeline

Таблица 11 — Comparison of the proposed methods against other strategies

Models multi-resolution Ensemble mAP

Yolov5m (ours) ✓ NMS 0.33

Yolov5m (ours) ✓ WPF 0.34

Yolov8m (ours) ✓ NMS + WPF 0.37

Yolov8L + EfficientDet ✓ NMS 0.35

Yolov8X ✓ NMS 0.34

that were constructed for AITrain competition for autonomous trains. the proposed pipeline won first place on the competition thanks to its robustness it was able to get almost same detection performance on both private and public testing while other methods failed. This section ends with a discussion on the importance of using synthetic data to enhance the accuracy of object detection systems. It also examines some experiments conducted to improve a particular system using images generated by the Kandinsky network, alongside results that highlight the significance of this step in achieving more precise and comprehensive object detection systems that can operate effectively in the production. In this table 12, we went back to the Posebowl

Таблица 12 — The advantage of using synthetic data in training the refiner only

Model Refiner JI public JI private syn data

Yolov8n/s/m-int8 Yolov8m 0.9159 0.9057 X

Yolov8n/s/m-int8 Yolov8m 0.9323 0.9261 ✓

Yolov8n/s/m-int8 Rt-DeTr 0.9315 0.9267 ✓

competition in which we used Kandinsky to enrich the training set and generalize the

pipeline capabilities and it shows that using this auto-generated data can improve the results winning the first place with the most robust and accurate system. In section 2.3, we discuss 3D object detection and focus on the CenterNet network, where we propose some modifications as illustrated in the fig. 0.18.

Рисунок 0.18 — Improved CenterNet proposed architecture

The modifications applied to the network show a noticeable improvement in the results on the studied data by minimizing the 3DoF of transition error from 1.23 to 0.9 and on rotation error from 0.18 to 0.135. The significance of using the CenterNet network in solving this issue facilitates its integration with the database of various 3D vehicle models, enabling us to perform 3D segmentation and build a comprehensive 3D system for all existing vehicles. A vehicle recognition system was developed and directly linked to the CenterNet network to perform this task, as shown in the fig. 0.19.

In the third chapter, the focus is on optimizing transformer-based models by accelerating their inference speed and reducing model size, while maintaining near original performance with lower hardware and memory requirements. The chapter first introduces Truncation-aware Singular Value Decomposition (SVD), then presents an integrated approach combining this method with a selective criteria mechanism, this hybrid approach replaces transformer blocks with either

в

EfficientNet ВС ns

The Classification model classifies 75 model types, trained with image size 224x224 with images similar to the above. Extracted from the Apollo 3D cars dataset.

Рисунок 0.19 — Integrated vehicle classification to perform 3D segmentation

a low-rank representations of the weights or linear transformations, which results in an effective training-free compression method, the study further extends truncation-aware SVD to analyze delta, weights between trained models proposing its application as a training-free knowledge transfer technique. Additionally, the research proposes a shared attention mechanism across transformer blocks reducing computational overhead while maintaining model efficiency. This architectural modification significantly decreases the computational requirements of the attention operations.

Section 3.1 We explore LLMs and VLMs architecture, then we introduce a new method to transfer knowledge between two models by defining a delta, model that is the subtraction of weights in-between these models. We assume that the delta, model weights are sparse enough to be represented effectively with a, low-rank approximation to produce something similar to LoRA in which we can use to transfer the knowledge between models without any training LoRA weights are estimated

using a truncation-aware SVD algorithm.

models = model f 1 — model^ase, (16)

modeltarget = SVDr (models) + model f 2, (17)

where SVDr stands for the singular value decomposition truncated at rank r of each weights matrix Ws. for any Ws in models

Wb = UZVT and SVDr(Ws) = Ur£rV? (18)

Then we introduce a whitening transformation using Cholesky factor to improve the estimation of the low-rank representation

Ws.x = Ws.L.L—i.x (19)

so that we can now represent SVDr as follows:

SVDr (Ws.L) = Ur £r V? (20)

and at inference we rewrite it this way so that L—1 could be merged with the second part

Ws.^ = .VrT .L—1).x (21)

In table 13, we illustrate the improvements we gain using our knowledge transfer method without any additional training, some other tricks such as RoPE scaling was used to extend the model context size and support longer videos.

Section 3.2 we propose a pruning and compressing technique based on a linear transformation matrix estimation or a low-rank representation based on importance

Model VideoMME NextQA

Video Mistral 50 66

Video Qwen 56 72

Video Mistral (ours) 64 75

Video Qwen (ours) 73 81

Таблица 13 — Evaluation results of our VLMs against baseline on well-know benchmarks VideoMME and NextQA

1. Input: Ordered layers L\,L2,... ,7Èn

2. For each layer Li in ordered layers do

a) Compute the initial cosine distance: ^initial 6) Replace Li with a linear layer

b) Compute the new cosine distance: dlinear

r) Replace Li with a low-rank representation of rank r

g) Compute the new cosine distance: rflow-rank

e) If ^linear < ^low-rank then

1) Choose the linear layer replacement Else

1) Choose the low-rank replacement

h) If achieved target compression ratio then 1) break

Algorithm 2: Minimize Cosine Distance in Decoder Layers

analysis criteria using a set of calibration data.

In algorithm 2, we present the proposed method (compressor) used to prune the models effectively by selectively replacing layers with their low rank representation based on singular value decomposition or by replacing them with linear transformations, we call layers after sorting them based on the initial cosine distance in between activations (ordered layers). In figure 0.20, we show that our method outperforms PruneMe (baseline) by large margin on some benchmarks including (reasoning, common sense and math). Another contribution of this work was proposing the shared attention mechanism, according to LLM and VLM analysis, we noticed that the attention scores obtained from multiplying the query and keys are very similar for later consecutive layers. therefore we share these attention scores instead of recomputing which will give much faster inference and lower memory requirements especially for long-context tasks. In table 14, we

Model WinoGrande TruthfulQA HellaSwag ARC Average

Mistral-7B 76.8 66.9 84.56 63.3 72.89

SA 8 layers 73.9 59.08 79.94 54.5 66.86

Таблица 14 — Mistral 7B Performance Metrics base and shared attention models (SA stands for shared attention).

compare a shared attention model (sharing 8 attentions) with the original mistral

Рисунок 0.20 — Comparison between Mistral 7B and two pruned version one using PruneMe and another using our compressor method

race

Model Performance Comparison boolq

lambada_openai

model and we show that we actually we can keep 92% of the original model performance without the need of any finetuning on different benchmarks. Following that we propose ReplaceMe method a systematic method for pruning transformer blocks through three key components: (1) an optimized block selection criterion that identifies pruning candidates by minimizing the cosine distance between transformer block outputs inspired by other methods, (2) a linear transformation estimation process that compensates for removed blocks via either closed-form least squares solution or numerical optimization of cosine similarity objectives, and (3) a regularization framework incorporating Li/Lo constraints to balance performance and model sparsity. The approach supports pruning of multiple non-contiguous blocks through independent linear transformations while maintaining the original model architecture by merging transformations into existing MLP layers, achieving

efficient compression without structural modifications or additional parameters. As shown on Tab.15, we outperform all other depth pruning methods without an

Method Train-Free C3 CMNLI CHID (test) WSC Hella Swag PIQA Race-M Race-H MMLU CMMLU AVG RP

Llama 2 7B (baseline) 43.8 33.0 41.6 37.5 71.3 78.1 33.1 35.5 46.8 31.8 45.3 100.0%

LLM-Streamline* X 43.3 33.0 24.1 36.5 61.1 71.5 34.8 37.0 45.5 29.4 41.6 92.0%

LLMPruner* X 29.7 33.4 28.4 40.4 54.6 72.0 22.9 22.0 25.3 25.0 35.4 78.2%

SliceGPT* X 31.5 31.6 18.5 43.3 47.5 68.3 27.0 29.4 28.8 24.8 35.1 77.5%

LaCo* X 39.7 34.4 36.1 40.4 55.7 69.8 23.6 22.6 26.5 25.2 37.4 82.7%

UIDL* X 40.2 34.4 21.5 40.4 59.7 69.0 35.2 34.7 44.6 28.9 40.9 90.3%

Ours (Cosine) ✓ 42.5 33.0 25.2 38.5 59.4 71.1 35.4 36.7 46.4 30.4 41.9 92.5%

Ours (LS) ✓ 39.4 33.0 18.9 38.5 58.5 70.5 37.1 36.5 45.2 29.2 40.7 89.9%

Таблица 15 — Comparing other pruning methods after healing and our training free approach ReplaceMe, * indicates that the numbers are taken from streamline paper [1]. After compressing Llama 2 7B with 25% compression ratio. и signifies that the model was trained following pruning, whereas и indicates that the model is training-free.

healing or finetuning while other methods requires substantially amount of training, We also evaluated our method on vision transformers to ensure its generalization.

In the fourth chapter, we discuss the practical application of the proposed driver monitoring and assistance system. It begins by highlighting the key innovations developed in the practical part of the research including: Russian sign language recognition models, 3D reconstruction library and monocular depth estimation models at the component level. At the system level, a system for predicting safe maximum speed values based on dynamic changes in the surrounding environment (e.g., the number of vehicles, road width, inter-vehicle distance, and weather conditions), while also integrating static speed limits set by traffic authorities. Additionally, the chapter introduces an assistance system for offline localization in case of internet access loss. This system works by linking images together and recognizing specific city landmarks using a visual database of these areas. In section 4.1, based on the previous work we presented on section 1.6, we propose two main contributions to improve mViT2 performance for sign language recognition in particular: 1) how to extend state-of-the-art transformer architecture for action recognition in which we propose vision-SOLAR method 0.21 which is basically stacking transformer blocks to get larger and more accurate models. Our

experiments show that SOLAR method works well for vision models as well and we present a new mvit xlarge model.

Рисунок 0.21 — Vision-SOLAR MViT proposed architecture

2) we present a method for adaptive sampling during training towards a more robust and generalized solution, which contributes with huge improvements in comparison with the baseline.

Таблица 16 — Comparison between our model and MViT models

Model Accuracy t Adp. sampling

MViTv2-small-32-2 (baseline) [2] 64.09 X

MViTv2-small-16-6 [2] 80.1 X

MViTv2-small (2nd place) [2] 81.0 X

MViT-tiny-16-6 (ours) 83.1 ✓

MViT-small-16-6 (ours) 84.4 ✓

MViT-base-16-6 (ours) 85.36 ✓

MViT-large-16-6 (ours) 86.2 ✓

SOLAR MViT xlarge-16-6 (ours) 89.1 ✓

In table 16, we illustrate that using our modifications improves the results from 64% which was the baseline provide to 89% on Slovo dataset for Russian sign language recognition winning the first place on the internation ML competition EqualAI.

We end this chapter by revisiting 3D reconstruction method based on multi-scene view input In section 4.2. we discuss the standard pipelines used to solve this task and the end-to-end ones in section 1.7, to improve these pipelines we propose 3 different algorithms. First one is a clustering algorithm based on DBScan in which we cluster the keypoints extracted in a coarse matching phase to identify the object of interest on the multi input images to crop the images accordingly. The proposed algorithm is straight forward, we do keypoints extraction followed by coarse matching to identify the clusters that appear on multiple images and then crop the images that includes these cluster. Second algorithm aimed to sort the multi-scene images for indoor objects, in which we define a distance metric based on the Structural Similarity Index Measure (SSIM). we assume each image as a node and build a complete graph with edge costs calculated using this metric. Then we use TSP solvers to sort the images, this method shows a huge improvement on the 3D reconstruction for indoor objects especially symmetrical ones. Third algorithm is a filtering method for keypoints using 3D fundamental matrix estimation. The idea of this this method is simple we call a 3D fundamental matrix Exzy which satisfy the following:

Yxzy Exzy Yxzy = 0, (22)

where Yxzy is the set of keypoints in the first point cloud and Yxzy their correspondences on the second point cloud. Then we filter these keypoints using the following criteria

projxy(Yxyz),projxy(Yxyz) where YxyExyYxy = 0 and YxzExzY'xz = 0 (23)

where proj is the projection of the 3D keypoints on a particular space. The proposed algorithms improved the mean average accuracy (MAA) metric from 0.15 to 0.27. In table 17, we illustrate the effectiveness of our proposed adjustments, using these

Таблица 17 — Results on both public and private LB for IMC competition against other methods. Aliked-LG* is our best submission which includes some hacks for specific cases.

Model MAA 3D Filtering DBScan TSP

Vggsfm 0.201/0.224 X ✓ X

Aliked-LG 0.153/0.155 X X X

Aliked-LG 0.17/0.175 X ✓ X

Aliked-LG 0.22/0.24 X ✓ ✓

Aliked-LG 0.232/0.252 ✓ ✓ ✓

Aliked-LG* 0.267/0.28 ✓ ✓ ✓

methods we were able to win the first place on google image matching challenge 2024 and presented this algorithms on the CVPR workshop.

Starting with section Section 4.3 of this chapter, we discuss monocular depth estimation in computer vision, we build on top of Adabins method as baseline where we propose to replace the encoder with a much lighter version and get rid of the adaptive bins with the transformer decoder on the head. To reduce the gap in performance caused by this change, we build a complex objective that is a weighted average of: 1) Sobel loss

GT —

+1 0 -1

* img^

Gy —

+1 0

1

^12 1 * img^

(24)

(25)

lossdx — loq(\(Gx , — GXt

^^ U \ \ \ ^prea ^gt

)|+1) * ^ ■ (26)

lossdy — log(\(Gyrnd - Gyt,)| +1) * —

(27)

2) cosine similarity loss 28

pred.gt 1

similarity =--¡-.-^—r.—r.--, lossGOS = |1 — similarity | * — (28)

max(llpredll2.llgtll2, eps) N

3) Virtual Normal Loss (VNL) 29

1 M

lossvn = — * fo (29)

i=0

where f is the distance between the triangle centers of prediction and ground truth as shown in Fig0.22.

Рисунок 0.22 — Virtual normal loss - triangulation and center distance

We conduct experiments on publicly available dataset (NYU dataset) and our internal drivesafely data. To prepare the ground truth of our dataset we used a pseudo-labeling technique be ensembling the top-5 state-of-the-art methods for monocular depth estimation. After training our proposed method shows almost the same performance as Adabins with 30x faster inference on CPU (Intel Core 9) and 15x faster inference on GPU (Nvidia RTX 3090). In table 18, we compare our method

Method Encoder Decoder RMSE ABS_REL log_10 Delta1 Delta2 Delta3

AdaBins Ours (large) Ours (small) EfficientNet-b5 EfficientNet-b0 MobileNetV3 UNet + miniViT UNet++ UNet++ 0.5136 0.5960 0.6088 0.1176 0.1173 0.1197 0.0477 0.04885 0.8638 0.8559 0.8494 0.9658 0.9597 0.9564 0.9884 0.9841 0.9821

Таблица 18 — Evaluation metrics for each method over our dataset

against Adabins on our internal dataset, and it shows competitive performance with much smaller and faster models. we trained mainly two versions one with efficient net as encoder and another using MobileNet v3.

Section 4.4 we propose a system named safe speed estimation. This system is responsible of analyzing the surrounding environment to capture specific parameters, then we use these parameters to estimate the safe speed limit instead of relying on the static speed limits as shown in Fig. 0.23. The system consists of multi components,

Road Width (RW)

Type: Integer Range: 1 - 7

Source: neural network and image processing Unit: number of car lanes

/-N

Number of Other Vehicles (NV)

Type: Integer Range: 0 - 50 Source: Object Detection Unit: vehicle

Closest vehicle (D)

Type: Integer Range: 1 -100 Source: Depth estimation Unit: meters

Safe Speed Estimation

safe_speed = SL*D*RW*cos(theta)/(MD*NV)

4-l

Road Curvature (theta)

Type: Integer Range: -90:90 Source: road segmentation Unit: degrees

Weather (W)

Type: state

values: (rainy -snowy .. etc) Source: Image processing Unit: None

Illumination (I)

Type: state

Range: (day time , night ..etc) Source: Histogram analysis Unit: None

Рисунок 0.23 — Proposed Approach to Safe Speed Limit Estimation

each components is responsible to capture certain features such as weather state, number of vehicles, road width and curvature .. etc. It also pulls the statics speed limit information from yandex maps, then using the following proposed formula 30:

SL*D*RW*cos(d)

safe speed =--—7-f^-Ж/, 30)

J ~ 10 * max(log(NV + 1),0.4) v y

we calculate the dynamic safe speed , it is also bounded to be always smaller or equal to the static speed limit to be consistent with the traffic law. In Section 4.5, we revisit image matching arid propose a modification for LoFTR baseline to improve the quality of the matches using a segmentation model and fundamental matrix filtering. Our proposed adjustments improved the mean average accuracy from 0.533 to 0.721 the extracted matches are used to build a fast and robust basic SLAM system in which we do triangulation and bundle adjustments every few steps. This system is not accurate for offline localizations, and it might diverge. So to provide it with something similar to loop closer we add a retrieval model for the most known landmarks in saint-Petersburg as shown on Fig.

0.24. Using the landmarks recognition model, we correct the estimated location.

Рисунок 0.24 — Diagram shows the process of adding the retrieval NN for auto-correction. T is a threshold that should be adapted according to the used

device

Unfortunately this system relies on the existence of landmarks for a specific region which makes it hard to generalize for any region or on traveling roads.

Publications

The main results on the topic of the dissertation are presented in 8 publications indexed in the Scopus citation database

1. Ali A., Gaikov G., Rybalchenko D., Chigorin A., Laptev I., Zagoruyko S. PairDETR : Joint Detection and Association of Human Bodies and Faces. 2024 IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2024.

2. Kashevnik A., Ali A. 3D Vehicle Detection and Segmentation Based on EfficientNetB3 and Residual Blocks, //Sensors, 2022, Vol. 22, No. 20, pp. 7990 Q1.

3. Kashevnik A., Ali A. Light-Weight 2D Map Construction of Vehicle Environment Using Semi-Supervised Depth Estimation Approach, DSAA,

2022, 7, 15th International Conference "Intelligent Systems" (INTELS'22),

2023, pp.

4. Othman W., Kashevnik A., Ali A., Shilov N. Driver-MVT: In-Cabin Dataset for Driver Monitoring including Video and Vehicle Telemetry Information //Data, 2022, Vol. 7, No. 5, pp. 62.

5. Kashevnik A., Ali A., Lashkov I., Zubok D. Human Head Angle Detection Based on Image Analysis //Advances in Intelligent Systems and Computing, 2021, Vol. 1288, pp. 233-242.

6. Kashevnik A., Ali A., Lashkov I., Shilov N. Seat Belt Fastness Detection Based on Image Analysis from Vehicle In-Cabin Camera //Proceedings of the 26th Conference of Open Innovations Association FRUCT, 2020, pp. 143-150.

7. Kashevnik A., Ali A. Vehicle Offline Localization Based on Computer Vision: an Approach Based on Image Matching Retrieval Algorithms and Implementation //Proceedings of the 33nd Conference of Open Innovations Association FRUCT, 2023, pp. 125-131.

8. Kashevnik A., Ali A. Driving Safe Speed Estimation Based on Outside Environment Vision Analysis //Proceedings of the 32nd Conference of Open Innovations Association FRUCT, 2022, pp. 121-127.

80

Introduction

Relevance of the topic. The Transformer architecture stands out as a powerhouse within AI various fields, showcasing its efficiency in areas ranging from natural language processing to computer vision. It serves as the backbone for the next generation of AI tools particularly in Large Language Models and generative AI. Modern breakthroughs in achieving SOTA results often hinge on this architecture, or its hybrid adaptations.

Even with these advancements, transformers have untapped potential that, if harnessed, could lead to even greater efficiency and accuracy. A key limitation lies in the architecture's dependency on an attention mechanism characterized by quadratic complexity relative to input length.

This research primarily dissects transformer architecture, particularly through a deep dive into the DETR architecture, renowned as the first end-to-end method for object detection. Building upon DETR, we introduce PairDETR, pioneering the first end-to-end graph prediction approach. By tweaking the architecture to predict node sets, introducing the concept of adaptive relative points, and developing an approximate bipartite matching loss, we achieve polynomial-time solutions to problems traditionally deemed NP-hard.

Our exploration extends to strategies aiming to accelerate transformer models, tackling their substantial resource demands. We delve into optimizing both training and inference processes using innovative, training-free techniques.

A thorough examination of transformer architecture discloses two core insights: 1) model activations prove to be sparse, allowing weights to be mapped to a low-rank representation through truncation-aware singular value decomposition; 2) Transformers have an underlying linearity, suggesting that many deep transformer blocks could be approximated with linear transformations (LTs).

Leveraging the first insight, we propose using truncation-aware singular value decomposition to represent the delta weights between models, enabling knowledge transfer akin to model merging without retraining. Inspired by both insights,

we unveil a new compression method for transformers, employing a selective criteria based on block importance analysis to decide whether to replace specific transformer blocks with linear transforms or low-rank representations. Additionally, we demonstrate the potential to share attention maps across certain blocks, proposing a "shared attention"approach to partially mitigate the input length's quadratic complexity.

We put our methods to the test on publicly available datasets and benchmarks across various domains, including LLMs, VLMs, and CLIP-based models, particularly for pruning and knowledge transfer. Practical application is exemplified through a driver assistant system, where we enhance components using our techniques and introduce novel auxiliary systems, like a safe speed prediction feature based on environmental dynamics.

The core significance of this work lies in its focus on refining one of machine learning's most potent and renowned architectures. By elevating model accuracy, performance, and reducing hardware requirements through advanced methodologies, this research stands to make a profound impact on numerous models. Goal of research. The primary aim of this research is to enhance the efficiency and performance of transformer-based models using truncation-aware singular value decomposition for training-free knowledge transfer and models compression and introduce approximated bipartite matching loss for end-to-end graph prediction. This encompasses tasks such as object detection, graph prediction, and advancements in generative AI, specifically focusing on large language models (LLMs), vision-language models (VLMs) multi-modal models such as CLIP. Research objectives.We can summarize the main objectives of this research as following:

1. Proposing a novel method using approximated bipartite matching loss towards end-to-end transformer-based graph prediction models;

2. Transformers training-free pruning and knowledge transfer methods for hardware requirements efficiency;

To achieve the primary objectives of the research, we divided the system into several core components and applied the adopted research methodology to achieve the main goals through:

1. Analysis and review

For each method included on this research, we studied and analyzed work published in international journals or conferences to build a foundational base from which to launch development and add new elements to achieve a better methods and algorithms.

2. Theoretical or Practical Development

In this study, we examined the progression and enhancements of state-of-the-art (SOTA) methods and algorithms over recent years. We formulated hypotheses aimed at specific improvement directions and validated these hypotheses using both mathematical analysis and experimental evidence.

3. Experimental Design

For each topic under this study, we tested modifications on the method or algorithm at various levels and different data. The testing process were mainly applied on different publicly available datasets, and then tested on data dedicated to the driver monitoring and assistant systems as it is the main practical system under study. Widely used test criteria for each method were employed using proper metrics and benchmarks.

Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.