Методы видеомониторинга физиологических характеристик и умственной работоспособности операторов транспортных средств тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Осман Валаа

  • Осман Валаа
  • кандидат науккандидат наук
  • 2025, «Национальный исследовательский университет ИТМО»
  • Специальность ВАК РФ00.00.00
  • Количество страниц 336
Осман Валаа. Методы видеомониторинга физиологических характеристик и умственной работоспособности операторов транспортных средств: дис. кандидат наук: 00.00.00 - Другие cпециальности. «Национальный исследовательский университет ИТМО». 2025. 336 с.

Оглавление диссертации кандидат наук Осман Валаа

Реферат

Synopsis

Introduction

CHAPTER 1 Literature review

1.1 Overview of traditional methods for monitoring vital signs

1.2 Conceptual foundations and traditional Methods for assessing

mental performance

1.3 Theoretical background on deep learning and neural networks

1.4 Technologies and techniques for video-based human state monitoring

1.5 Challenges in existing approaches and Summary

CHAPTER 2 Methods design and development

2.1 Video-based robust vital sign estimation technology for assessing

heart and respiratory rates

2.1.1 Proposed respiratory rate estimation approaches

2.1.2 Proposed heart rate estimation methods

2.2 Agentic Large Language model for vital signs estimation

2.3 Two-stage method for operator mental performance estimation

2.3.1 Oxygen Saturation Assessment Model and Movement Analysis

2.3.2 Mental performance assessment model based on determined

vital signs

2.4 Methodology of collecting a dataset for experimental research

2.5 Summary

CHAPTER 3 Results and Discussion

3.1 Experimental platform

3.2 Experimental setup and comparative analysis of the developed machine learning models for estimating the respiratory rate

3.2.1 Used datasets

3.2.2 Implementation of the skeletonization-based method for estimating the respiratory rate

3.2.3 Implementation of the 3D classification-based Method for estimating the respiratory rate

3.3 Implementation and comparative analysis of the developed models

for monitoring the heart rate

3.3.1 Used Datasets

3.3.2 Implementation details of the 3D-classification based heart

rate estimation

3.3.3 Implementation details of the ViT-BiLSTM approach and the multi-skip connections decoder based approach for estimating the heart rate

3.3.4 Experiments and evaluation

3.4 Implementation and evaluation of the developed models for monitoring the oxygen saturation

3.4.1 Used dataset

3.4.2 Conducted experiments for choosing the best ROI

3.4.3 Implementation of the multi-scale vision transformer based approach

3.5 Implementation of the vital signs assessment based on LLM

3.6 Implementation of the mental performance estimation model

3.6.1 Dataset description

3.6.2 Implementation details and experiments

3.7 Discussion of findings

3.8 Summary

Conclusion

References

163

Publications

Реферат

Общая характеристика диссертации

Введение диссертации (часть автореферата) на тему «Методы видеомониторинга физиологических характеристик и умственной работоспособности операторов транспортных средств»

Актуальность темы

В современных отраслях промышленности, включая производство, энергетику и транспорт, операционная среда становится все более сложной, что требует повышенного внимания и мониторинга состояния работников. Операторы играют ключевую роль в этих условиях и все чаще подвергаются утомлению, стрессу и физиологическим нагрузкам, которые снижают умственную работоспособность, что снижает ситуационную осведомленность оператора и влияет на безопасность. Традиционные методы оценки состояния оператора, такие как электрокардиография (ЭКГ), носимые биосенсоры или когнитивные тесты, часто неприменимы, когда работа требует активного участия оператора и прерывание рабочего процесса недопустимо. Кроме того, носимые устройства могут быть неприемлемы или запрещены в определенных контекстах из-за требований безопасности, что ограничивает возможность масштабируемого непрерывного мониторинга в реальном времени. Хотя умные часы могут измерять некоторые физиологические характеристики, такие как частота сердечных сокращений и насыщение кислородом, оценка на основе интеллектуального анализа видеоизображений операторов предоставляет ценную альтернативу, обеспечивая бесконтактный, пассивный мониторинг без необходимости ношения оператором какого-либо устройства. Это поддерживает непрерывное наблюдение без перерывов, что делает этот подход особенно выгодным, когда носимые устройства являются непрактичными, некомфортными или навязчивыми.

В работе представлены бесконтактные методы, основанные на компьютерном зрении, для оценки физиологических характеристик опреатора (таких как частота сердечных сокращений, частота дыхания, насыщение крови кислородом и артериальное давление) и показателя умственной работоспособности. Ис-

пользуя современные достижения в области компьютерного зрения, обработки сигналов и глубокого обучения, предложенные методы эффективно извлекают физиологические и поведенческие сигналы непосредственно из видеопотока камеры, направленной на оператора без специализированного оборудования или физического контакта. Это обеспечивает пассивный и непрерывный мониторинг, что делает его хорошо подходящим для развертывания в динамичных, критически важных для безопасности рабочих местах.

Отслеживание умственной работоспособности имеет важное значение для обеспечения безопасности в транспортных ролях, таких как вождение автомобиля и пилотирование воздушных судов, где небольшие ошибки могут привести к серьезным последствиям. Когда люди испытывают умственную усталость, скорость их реакции снижается и они совершают больше ошибок, что может привести к авариям. Своевременное выявление сотсояния утомления человека повышает вероятность предотвращения вышеупомянутых происшествий и позволяет создать более безопасную рабочую среду.

Традиционные методы определения утомления человека, предполагающие необходимость совержения некоторого действия (например, нажатия кнопки) в условиях ограниченного времени, продемонстрировали высокую эффективность в оценке состояния человека. Однако, основным недостатком этих методов является необходимость активного участия оператора в процессе определения его состояния, что прерывает рабочий процесс и само по себе может вызывать дополнительную усталость. Кроме того, эти методы предоставляют только дискретные измерения в определенные моменты времени, а не непрерывный мониторинг. Напротив, предложенный в рамках работы метод оценки умственной работоспособности обеспечивает непрерывный пассивный мониторинг, который не нарушает рабочий процесс оператора.

в рамках данной работы совершенствуются методы интеллектуальной обработки видео для оценки физиологических показателей человека. Сочетание глубокого обучения и обработки видео справляется с такими технологическими задачами, такими как тряска во время движения транспортного средства, переменное освещение и необходимость в анализе данных в реальном времени.

Эти достижения расширяют потенциал бесконтактного мониторинга здоровья, прокладывая путь для более умного взаимодействия человека с машинами.

Разработанные методы применимы в различных секторах. В здравоохранении видеомониторинг может улучшить уход за пациентами и дистанционные консультации. В транспортной отрасли своевременное обнаружение усталости у пилотов или водителей может предотвратить аварии.

Таким образом, работа направлена на разработку надежных бесконтактных методов мониторинга, способных оценивать физиологические параметры, такие как частота сердечных сокращений, артериальное давление, частота дыхания и насыщение кислородом, по стандартным ИСБ-видеозаписям лиц операторов в условиях реальных операционных ограничений. Опираясь на эти физиологические измерения, предложено оценивать умственную работоспособность оператора, на основе использования метрику Ли (тест Ландольта) в качестве эталонного показателя. Методы решают задачу поддержания точности измерений при изменяющихся условиях освещения, которые часто подрывают существующие подходы к мониторингу физиологических характеристик на основе камер, с конечной целью обеспечения непрерывной оценки состояния оператора транспортных средств. Объединяя технологический прогресс, приоритеты безопасности и практический дизайн, эта работа предлагает дорожную карту для более умных и здоровых рабочих мест.

Цель

Повышение эффективности мониторинга операторов транспортных средств на основе разработки методов цифровой обработки аудиовизуальной информации. Основное внимание уделяется разработке методов и моделей для оценки физиологических характеристик оператора, таких как частота сердечных сокращений, артериальное давление, насыщеность крови кислородом и частота дыхания, а также к оценка его умственной работоспособности.

Задачи

В данном разделе представлены поставленные и решенные задачи, необходимые для достижения цели в рамках диссертации.

Задача 1: Комплексный анализ предметной области. Систематическое сравнение результатов исследований в данной области.

Задача 2: Оценка состояния оператора. Разработка системы оценки умственной работоспособности и физиологических показателей оператора. Система будет отслеживать физиологические характеристики с поведенческими сигналами оператора дляобеспечения комплексного понимания его когнитивной нагрузки.

Задача 3: Оценка частоты дыхания с использованием технологии скелетизации. Разработка метода оценки частоты дыхания и анализа позы оператора с использованием технологии скелетонизации и оптического потока. Данный метод устойчив к изменениям освещения, что обеспечивает его надежную работу в динамичных условиях и в условиях слабого освещения.

Задача 4: Повышение точности оценки физиологических показателей. Разрабтка методов оценки частоты сердечных сокращений и насыщенности крови кислородом опреаторов на основе областей интереса (ROI).

Задача 5: Изучение эффективности больших языковых моделей для оценки физиологических показателей операторов. Дообучение большой языковой модели (например, Mistral) с использованием технологий ImageBind и LoRA для оценки физиологических показателей человека по последовательности изображений с применением методов оптимизации (таких как: tensor-параллельные вычисления, flash внимание и квантование) для повышения вычислительной эффективности как на этапе обучения, при использовании.

Методы исследования

А рамках исследования использовался комплекс методов и подходов для достижения поставленных целей: методы искусственного интеллекта и машинного обучения, алгоритмы анализа данных, методы математического представления данных, теория статистики, математический анализ и линейная алгебра. Для каждой задачи были проанализированы существующие на сегодняшний день методы и подходы. Этот анализ способствовал разработке новых методов для повышения эффективности и применимости данных подходов к оценке физиологических показателей и когнитивных функций по видеозаписям.

Основные положения, выносимые на защиту

Исследование было проведено в соответствии с требованиями, изложенными в рамке научной специализации 2.3.8 Высшей аттестационной комиссии Российской Федерации. В соответствии с регламентирующими критериями специализации для научной защиты предлагаются следующие основные вклады:

1. Двухэтапный метод обнаружения умственной работоспособности оператора, основанный на оценке физиологических показателей (соответствует пункту 4 и пункту 13 паспорта специальности).

2. Метод оценки физиологических показателей на основе анализа последовательности изображений с использованием крупных языковых моделей через интеграцию мультимодального эмбеддинга и техник работы с пром-том (соответствует пункту 13 паспорта специальности).

3. Технология оценки физиологических показателей на основе анализа видео, объединяющая измерение частоты дыхания, устойчивое к изменениям освещения, и методы детекции сердечного ритма, основанные на анализе области интересов на лице человека (соответствует пункту 4 паспорта специальности)

Научная новизна

Научная новизна диссертации выделяется следующим образом:

1. Двухэтапный метод оценки умственной работоспособности оператора, отличающийся предварительным извлечением важных физиологических показателей с помощью видео и использованием их в качестве признаков для нейронной сети-классификатора, что позволяет отслеживать снижение работоспособности человека через мульти-модальную интеграцию физиологических и поведенческих данных.

2. Метод оценки важных физиологических показателей, основанный на больших языковых моделях, отличающийся использованием техник работы с промтом и мульти-модального эмбеддинга, что позволяет успешно обрабатывать шумные и низкокачественные входные данные.

3. Технология обнаружения физиологических показателей опреатора на основе анализа видео, включающая методы оценки частоты дыхания и сердечного ритма, отличающаяся использованием скелетизации изображения человека и оптического потока, а также применением multi-skip connections для нейронной сети на основе архитектуры трансформера, что обеспечивает устойчивость к изменениям окружения и улучшает производительность.

Объект исследования

Данное исследование фокусируется на анализе изображений и видеозаписей, полученных с помощью RGB-камер из общедоступных и специально собранных наборов данных, для решения поставленных научных задач. Работа сконцентрирована на изучении поведения операторов/водителей в салонах транспортных средств и на рабочих местах за компьютерными станциями.

Предмет исследования

Модели и методы для оценки состояния оператора, включая нейронные сети, машинное обучение, крупные языковые модели и обработку сигналов.

Теоретическая значимость

Теоретическая значимость диссертации заключается в разработке и исследовании методов анализа визуальных данных для оценки физиологических показателей и умственной работоспособности оператора транспортных средств.

Практическая значимость

Практическая значимость диссертации заключается во внедрении разработанных бесконтактных методов мониторинга для оценки физиологических показателей и умственной работоспособности оператора транспортных средств. Данные методы могут быть развернуты на стандартных ЯОБ-камерах без необходимости специализированного оборудования или носимых устройств, что обеспечивает ненавязчивое и непрерывное наблюдение, что отличает разработанные методы от традиционных подходов. Эти методы применимы в транспортном секторе, где их можно интегрировать в кабины транспортных средств для мониторинга физиологического состояния и умственной работоспособности водителей. Путем непрерывной оценки частоты сердечных сокращений, частоты дыхания и умственной работоспособности система может обнаруживать ранние признаки усталости или стресса, активируя предупреждения для предотвращения аварий. В обрабатывающей и энергетической отраслях методы используются для мониторинга операторов в диспетчерских

и на производственных площадках. Анализируя видеопотоки с камер, расположенных перед операторами, система оценивает физиологические показатели, чтобы осуществлять мониторинг умственной работоспособности для сохранение бдительности во время выполнения критических задач. Данное применение было апробировано в крупной энергетической компании, где оно позволило сократить количество инцидентов, связанных с человеческим фактором, за счет раннего вмешательства при появлении у операторов признаков когнитивной перегрузки. Практическая реализация подтверждена тремя зарегистрированными программными системами.

Внедрение результатов работы

Результаты исследования использованы в следующих проектах:

- Проект РНФ 24-21-00300, Методы и модели для определения усталости оператора на основе анализа физиологических показателей, полученных с использованием систем компьютерного зрения, 2024-2025. - Проект РНФ 22-21-00790, Методы и модели для систем поддержки принятия решений в области проектирования сложных систем, 2021-2023. - Проект РНФ 18-71-10065, Модели и методы для интеллектуальной поддержки водителей на основе мониторинга ситуации в салоне автомобиля, 2018-2023. - НИОКР проект ИТМО 620176, Мобильное приложение для автоматизированной оценки функционального состояния человека на основе видеопотока, 2020-2021.

Апробация результатов работы

Ключевые результаты исследования были представлены и обсуждены на следующих конференциях:

1. The 37th Conference of Open Innovations Association FRUCT 14-16 Мая 2025.

2. 9-я Международная Научно-Практическая Конференция "Технологическая перспектива 2023 29-30 ноября 2023.

3. The 33th Conference of Open Innovations Association FRUCT, 24-26 Мая 2023.

4. 52-я Научная и учебно-методическая конференция Университета ИТМО, 31 января - 3 февраля 2023.

5. Intellisys 2022, 1-2 сентября 2022, Амстердам, Нидерланды.

6. XI Конгресс молодых ученых 2022 года, 4-8 апреля 2022, Санкт-Петербург, Россия.

7. Научная и учебно-методическая конференция Университета ИТМО, 2-5 февраля 2022, Санкт-Петербург, Россия.

8. 28-я Конференция Ассоциации открытых инноваций FRUCT, 27-29 января 2021, Москва, Россия.

9. Научная и учебно-методическая конференция Университета ИТМО, 1-4 февраля 2021, Санкт-Петербург, Россия.

10. X Конгресс молодых ученых (онлайн), 14-17 апреля 2021, Санкт-Петербург, Россия.

Личный вклад автора

Личный вклад аспиранта состоит в определении целей и задач исследования, поиске источников информации, концепции и дизайна идей и методологий, создание и применение методов и моделей, обучении и развертывании нейронных сетей, проведения экспериментов и составлении обзоров литературы. Важно отметить, что соавторы внесли различные вклады в зависимости от конкретной статьи. Некоторые соавторы обеспечивали руководство и помощь в организации идей, а также в подготовке статей к публикации, в то время как другие внесли конкретные вклады, основываясь на своей экспертизе. Кашевник А. выступил

в роли научного руководителя. Хамуд Б. внес вклад в разработку и реализацию нейронной сети для оценки артериального давления на основе видео [1] и в разработку метода оценки насыщения кислородом [2], а также подготовку набора данных для оценки умственного утомления [3]. Шилов Н. обеспечил надзор и валидацию методов, в частности методологии оценки умственной производительности [3], так как он внес вклад в установление порога утомления на основе результатов теста с кольцами Ландольта. Рябчиков И. внес вклад в реализацию оценки движения человеческого тела [4; 5]. Али А. внес вклад в разработку и Рюмин Д. в реализацию метода оценки частоты сердечных сокращений [6].

Структура и количество страниц диссертации

Структура диссертации включает введение, три главы, заключение и список литературы. Объем работы составляет 181 страницу, содержащую 41 рисунок и 23 таблиц. Библиографический раздел насчитывает сто сорок восемь (148) источника.

Содержание диссертации

Введение включает актуальность темы, цель и задачи исследования, новизну работы и положения, выносимые на защиту.

Первая глава представляет комплексный анализ физиологических показателей уделяя особое внимание бесконтактным методам измерения ключевых параметров (частота сердечных сокращений, дыхание). Эти технологии минимизируют физический контакт и повышают комфорт оператора. Рассматриваются такие показатели как: сердечный ритм, паттерны дыхания, артериальное

давление и уровень насыщения крови кислородом, изученные с точки зрения их практической значимости для диагностики нарушений здоровья.

В главе детально анализируются технологии видео-мониторинга физиологических показателей и ментального состояния. Анализ литературы подтверждает, что частота сердечных сокращений, кровеное давление, насыщенность крови кислородом и частота дыхания могут быть оценены по видеозаписи лица на основе физиологических принципов. Частота сердечных сокращений измеряется с помощью удаленной фотоплетизмографии, обнаруживающей subtle color variations, поскольку гемоглобин поглощает зеленый свет during cardiac cycles. Артериальное давление оценивается с использованием расчетов времени распространения пульсовой волны, которые требуют сигналов rPPG в combination with a reference signal. Частота дыхания определяется путем отслеживания вызванных дыханием микродвижений посредством алгоритмов оптического потока, которые обнаруживают расширение ноздрей и движение челюсти. Насыщенность крови кислородом рассчитывается путем анализа различного поглощения света на specific wavelengths между оксигенированным и деоксигенированным гемоглобином в лицевых кровеносных сосудах. Каждый параметр использует различные оптические и физиологические свойства, регистрируемые стандартными камерами. Основные подходы включают математические модели, методы обработки сигналов и техники глубокого обучения для оценки частоты сердечных сокращений. Также обсуждаются бесконтактный мониторинг частоты дыхания, предсказание артериального давления, а также дополнительно исследуются методологии для оценки уровня насыщенности крови кислородом с использованием видео лиц операторов.

Ряд исследований [7-10] подтвердили взаимосвязь между умственной работоспособностью и физиологическими показателями, выявив корреляцию изменений частоты сердечных сокращений, артериального давления, сатурации кислорода и частоты дыхания с изменениями показателей памяти, когнитивных функций и уровня стресса. Исследования показали, что частота сердечных сокращений увеличивается при росте умственной нагрузки от низкого до среднего уровня. Повышенное артериальное давление ассоциировано со снижением

когнитивной производительности. Кроме того, частота дыхания возрастает при увеличении стресса и рабочей нагрузки. Также была выявлена значительная положительная корреляция между сатурацией кислорода и показателями памяти. Несмотря на достижения в области интеллектуального анализа видео, в области мониторинга операторов остается ряд проблем. Одна из главных проблем связана с динамическими условиями, такими как переменная освещенность и разнообразные оттенки кожи. Ограниченное разнообразие участников и зависимость от контролируемых внутренних условий вызывают обеспокоенность в отношении обобщаемости модели. Артефакты движения и необходимость в улучшенных методах обработки сигналов усложняют реализацию.

Быстро развивающаяся область бесконтактного измерения физиологических характеристик объединяет математическое моделирование, обработку сигналов и глубокое обучение для повышения точности и полезности в транспортных средствах. Выявленные проблемы подчёркивают возможности для методологического усовершенствования, что обеспечивает необходимость преодоления существующих ограничений для улучшения надёжности, адаптируемости и развертывания технологий видищомониторинга операторов.

Обзор литературы подтверждает актуальность исследования в данной области. Основное внимание уделено разработке методов и алгоритмов, способных последовательно оценивать физиологические показатели человека при различных условиях освещенности.

Вторая глава описывает разработанные методы оценки физиологических показателей и анализа умственной работоспособности. Предложенный двух-этапный метод, показанный на рисунке 8 оценивает ключевые физиологические параметры, такие как частота сердечных сокращений, частота дыхания, артериальное давление и насыщенность крови кислородом, с целью последующей оценки умственной работоспособности. Двухэтапный метод превосходит одно-этапный метод и базируется на идее того, что физиологические изменения предшествуют снижению умственной работоспособности, что позволяет более точно осуществлять мониторинг состояния человека. Входное видео обрабаты-

вается для извлечения лицевых характеристик, строится виртуальный скелет, определяются ключесвые точки и их движение, которые затем подаются в разработанные предварительно обученные модели для оценки физиологических показателей человека. Система использует комбинацию нейронных сетей, таких как OpenPose для оценки позы, Selflow для анализа оптического потока и передовых архитектур, таких как 3D EfficientNet и Vision Transformers, для задач классификации. Выходные данные включают измерения физиологических показателей и умственной работоспособности, которые предоставляются пользователю через REST API.

Рисунок 8 демонстрирует предложенный метод оценки ментальной работоспособности.

Figure 1 — Предложенный двухэтапный метод оценки умственной

работоспособности

Метод разделен на несколько компонентов. Первый этап заключается в предварительной обработке и включает в себя сегментацию лица и скелетонизацию для определения областей интереса (ROI). После этого видеоданные обрабатываются с использованием четырех различных методов для оценки ЧД, ЧСС, АД и SPO2. Эти физиологические характеристики на втором этапе подаются в предложенную модель табличного преобразователя для оценки умственной работоспособности. Архитектура предложенной модели проиллюстрирована на рисунке 9.

5 Physiological features (vital signs) extracted from the first stage

Embedding + Positional

* Z-score normalization Encoding

Transformer la7eг 2

Transformer layer

Transformaf layer 3 Transformer layer 4

Attention-Based pooling

Linear¡ayer

Mental performance

Figure 2 — Модель оценки умственной работоспособности

В начале происходит обработка стандартизированных входных признаков через слой встраивания, который отображает эти признаки в 64-мерное скрытое пространство. Этот процесс встраивания преобразует физиологические признаки в плотные представления, позволяя уловить нелинейные зависимости, присущие данным. Для сохранения позиционной информации признаков были интегрированы обучаемые параметры позиционного кодирования, которые добавляются к встроенным признакам, чтобы модель могла понимать относительную важность различных физиологических измерений. Основой архитектуры является энкодер-трансформер, состоящий из четырех последовательных слоев. Каждый слой включает механизм самовнимания с несколькими головами, используя четыре головы внимания на слой. Слои трансформера обрабатывают встроенные признаки как последовательный вход, что позволяет моделировать взаимосвязи между признаками через вычисленные оценки внимания. Выход затем подается в слой пулинга на основе внимания, за которым следует линейный слой, выводящий оценку ментальной работоспособности.

В главе также подробно описана технология видео-оценки физиологических характеристик, которая объединяет метод измерения частоты дыхания на основе скелетизации, устойчивый к изменениям освещения, и метод обнаружения частоты сердечных сокращений, основанный на анализе областей интереса. Для оценки частоты дыхания был предложен, основанный на обнаружении движения точки грудной клетки, вызванного процессом дыхания, с использованием

обычной ЯСБ-камеры, такой как камера смартфона. Метод показан на рисунке 10.

Figure 3 — Предложенный алгоритм оценки частоты дыхания

Как показано на рисунке, подход состоит из четырех основных этапов:

1. Обнаружение позиции точки грудной клетки: было произведено сравнение двух методов на основе нейронных сетей для оценки позы человека: OpenPose и Integral Human Pose Estimation. Метод OpenPose был выбран благодаря его более высокой точности и способности обнаруживать ключевые точки даже при отсутствии определенных частей тела.

2. Обнаружение движения грудной клетки, представленного проекцией смещения точек грудной клетки между кадрами на ось Y с использованием нейронной сети на основе оптического потока под названием Selflow. Проекция гарантирует, что учитывается только движение, вызванное процессом дыхания.

3. Фильтрация и подавление шумов полученного сигнала смещения.

4. Определение частоты дыхания путем подсчета количества пиков в сигнале.

Для оценки ЧСС система использует подход ЭЭ-классификации с использованием модели EfficientNet-B1 и архитектуры Vision Transformer-Bidirectional Long Short-Term Memory (ViT-BiLSTM). Область интереса извлекается с использованием лицевых отметок, обнаруженных при помощи технологии 3DDFA_V2, и, после этого, разработанная модель обрабатывает последовательность кадров для предсказания частоты сердечных сокращений. Введён декодер с множественными пропусками, показанный на рисунке 4, чтобы улучшить извлечение признаков и точность предсказания.

Список литературы диссертационного исследования кандидат наук Осман Валаа, 2025 год

chin 1.

Forehead 1.16% Forehead+cheeks 1.11%

using MAE criteria is the face region, It's noteworthy that the face is the optimal choice, as in some instances, parts of the face may be partially covered, such as one cheek, chin, or forehead. This partial coverage can result in failures in detecting the Region of Interest (ROI) in case the ROI is the covered part.

3.4.3 Implementation of the multi-scale vision transformer based

approach

As a preprocessing step, we extracted the face ROI using the keypoints extracted by 3DDFA_v2. Subsequently, we resized the image to 224x224 pixels. Since the frame rate for the videos in VIPL-HR is 30 FPS, we sampled 20 frames with one frame skipped between each pair, resulting in a sampling interval of approximately 1.3 seconds. The MVit from the mmaction library was used as the main model with 101 classes corresponding to all possible values for oxygen saturation (0-100%). The dataset was split 20% for testing and 80% for training. Results in table.3.9 showed a better performance of our method with MAE of 0.56% compared to previously published work.

Table 3.9 — Comparison of different methods for blood oxygen saturation

estimation. _

Method MAE (%)

Deep Learning with STMap [142] 1.274

Multi-model fusion module [143] 1.000

CNN+XGBregressor [2] 1.170

Casalino et al [144] 3.334

Past Analytic (Ratio of Ratios) [145] 1.838

mViT based approach (ours) 0.56

3.5 Implementation of the vital signs assessment based on LLM

In this section, we explore the experiments conducted with the Mistral LLM as detailed in Section 2.2. As previously mentioned, we utilized the pretrained models with their original weights and exclusively trained the projection linear layer. The input consisted of 32 frames. The choice of the frames number was determined by analyzing the of how vital signs reveal themselves over time in video recordings.Through repeated testing, we noted that shorter segments under 20 frames usually missed a full heartbeat cycles or breathing patterns. Thirty-two frames struck a practical balance: at standard video speeds, this captured over one second which sufficient to observe full physiological rhythms like the rise and fall of a single breath or the pulsing of blood through facial capillaries, yet compact enough to process efficiently even on modest hardware.

Three experiments were conducted: the first involved using the mean embeddings of all frames, the second utilized the differences between the maximum and minimum embedding matrices and the final experiment employed the embedding of all frames.

The experiments utilized both the LGI-PPGI dataset and the V4V vital signs dataset which incorporates annotations for respiratory rate, heart rate and blood pressure. The objective was to determine the most suitable embedding for use. The assessment was performed using the V4V evaluation subset for Respiratory Rate (RR) and Heart Rate (HR), while for systolic and diastolic blood pressure, 30% of the

V4V training subset was utilized. Table 3.10 and Table 3.11 provides a comparison of the results obtained with the three experiments using MAE and RMSE respectively.

Table 3.10 — Comparison of the used input embedding for vital signs estimation using LLM (MAE values presented in the table)

Vital sign Mean embedding Max-Min embedding All embeddings

Heart rate 11.8 10.5 11.1

respiratory rate 4.6 5.0 4.2

Systolic blood pressure 8.8 14.7 8.8

Diastolic blood pressure 6.2 10.5 5.8

Table 3.11 — Comparison of the used input embedding for vital signs estimation using LLM (RMSE values presented in the table)

Vital sign Mean embedding Max-Min embedding All embeddings

Heart rate 14.7 13.7 14.5

respiratory rate 5.6 6.5 5.58

Systolic blood pressure 17.2 21.9 16.5

Diastolic blood pressure 9.4 14.1 8.79

from the Tables we can notice that the max-min embedding method consistently delivered weaker results across all vital sign measurements, suggesting that extreme value differences fail to capture the critical time based variations needed for accurate physiological monitoring. When we examined diastolic blood pressure specifically, the all-embeddings approach showed a clear advantage by reducing the MAE compared to the mean embedding technique. We believe this happens because the model can now track subtle changes between consecutive frames, capturing blood vessel behavior that gets averaged out in other methods. Thus utilizing the embedding of all images as input for the Large Language Model (LLM) yields superior performance in estimating vital signs compared to using only the mean embedding or the difference between the maximum and minimum embeddings.

Therefore, inputting the embedding of all the frames was selected for further experimentation. Keeping all video frames in the analysis matches current research showing that time-based patterns help detect subtle physical signals from video. This method works especially well for breathing rate because it allows capturing the breathing cycles. When we feed every frame into the model instead of averaging them, the system better recognizes the patterns in the data.

For additional experimentation, we initially trained all attention layers within ImageBind while keeping the remainder of the architecture frozen. Subsequently, we froze the architecture and exclusively trained the projection layer between the ImageBind and the LLM. Table 3.12 shows the results obtained for the V4V evaluation dataset.

Table 3.12 — Vital signs estimation results after training the ImageBind and the pro jection layer

Vital sign MAE RMSE

Heart rate 8.5 10.3

respiratory rate 4.1 6.8

Systolic blood pressure 11.18 14.54

Diastolic blood pressure 9.80 12.72

Enhancing the heart rate MAE by ImageBind finetuning shows how important it is to adapt the visual encoder to the physiological features. This improvement overcome the enhancements achieved through optimization the projection layer alone which suggests that the low level feature extraction requires specialized tuning for health applications. The model maintained strong performance on respiratory rate by an MAE of 4.1 despite the challenges of estimating this parameter from visual data where motion artifacts often dominate the signal. As shown in Table 3.12, the results were significantly enhanced and it achieved the SOTA on the 3 vital signs.

3.6 Implementation of the mental performance estimation model

3.6.1 Dataset description

The Human Fatigue Assessment Video Dataset (HFAVD) [3], which is inhereteed from the OperatorEYEVP dataset introduced in [146], was used for the training of the mental performance model. This dataset comprises video recordings of ten individuals participating in diverse activities across three distinct time intervals (morning, afternoon and evening) over a period of eight to ten days.

Each day experimental protocol commenced with a sleep quality assessment administered before the morning session. Subsequently, participants completed the Visual Analog Scale for Fatigue (VAS-F), a Choice Reaction Time (CRT) task, engaged in reading a scientific text, performed the "Landolt rings"visual correction test, played the "Tetris"game and concluded with a second CRT task. On average, each session lasted approximately one hour. To quantify fatigue, we analyzed the outcomes of the Landolt test which serves as a robust measure of mental performance, capturing cognitive and attentional deficits associated with fatigue.

3.6.2 Implementation details and experiments

The model employed in this study is detailed in Section 2.3.2. This implementation builds upon established transformer architectures while adapting specifically for physiological signal processing constraints. During hyperparameter tuning, we evaluated embedding dimensions from 32 to 128 and found 64 provided optimal balance between representational capacity and computational efficiency for CPU deployment—larger dimensions showed negligible accuracy gains while increasing the inference time by 37%. The four transformer layers were selected after observing performance saturation in validation tests; adding fifth layer reduce

the mae by only 0.2% but double the processing latency. The single head attention configuration proved insufficient for capturing cross-feature dependencies while eight heads introduced overfitting on our limited dataset. The 150-epoch duration was determined through early stopping analysis where validation loss after epoch 142±5 (mean±SD across five runs).

Optimization was performed using the Adam optimizer, with a learning rate set to 0.001 and the model was trained over 150 epochs with early stopping implemented to prevent overfitting. The mean squared error (MSE) loss function was utilized during training. The architecture is designed for CPU-based computation and takes advantage of PyTorch's automatic differentiation capabilities to efficiently compute gradients. Key hyper-parameters include an embedding dimension of 64, which was determined experimentally, along with 4 transformer layers and 4 attention heads. The distribution of the mental performance is shown in figure 3.7 The dataset was

Figure 3.7 — Mental performance distributions: training dataset (left) versus test

dataset (right).

partitioned into training and testing subsets with an 80:20 ratio. The performance of the model, evaluated using the five features BP Systolic, BP Diastolic, Heart Rate, Oxygen Saturation and respiratory rate showed an MAE of 0.5451. Based on the dataset containing respiratory rate, blood pressure, heart rate, oxygen saturation, and mental performance (Au) metrics, the correlation analysis reveals critical relationships between physiological indicators and mental performance calculated based on Landolt test.

-0.2 -

Rft SE3P DBP HK Sp02

Figure 3.8 — Correlation Coefficients between vital signs and mental performance.

As shown in figure 3.8,The analysis shows that respiratory rate exhibits the strongest positive correlation with mental performance (r = 0.22), indicating that higher breathing rates (15-22 breaths/min) correspond to better cognitive function (Au > 1.5), while lower rates (6-10 breaths/min) signal significant performance degradation (Au < 0.5). Blood pressure parameters show moderate negative correlations (systolic r = -0.18, diastolic r = -0.11). Heart rate shows a moderate correlation (r = -0.15), where performance declines as heart rate exceeds 86 BPM. Oxygen saturation shows the weakest but still significant correlation (r = 0.10), with more pronounced effects at lower saturation levels (<97%). The most reliable predictor of mental performance degradation occurs when multiple vital signs simultaneously deviate from optimal ranges, particularly when respiratory rate falls below 12 breaths/min combined with heart rate exceeding 83 BPM, which predicts critical performance decline (Au < 1.0) with 87% accuracy.

To validate the need of two-stages method. we ran an experiment to extract the mental performance from the video directly. To do that we implemented the Video

Transformer model as it was previously used for estimating the heart rate in the first stage. The model was trained end-to-end using the mental performance from Landolt test as groung truth. The L1loss was used as the loss function with Adam optimizer and learning rate of 1e-3. The Table 3.13 shows a comparison between the performance of two-stages and the one-stage method consists of vision transformer model following by LSTM and linear layer.

Table 3.13 — Comparison of two-stage vs. single-stage approach for mental performance assessment_

Approach MAE RMSE Advantages

Two-stage 0.5451 0.6586 Better performance, modular design

Single-stage 0.6893 0.7921 Simpler implementation

This results is moderate which suggest that vital signs again may fails to detect the mental performance, so we made more experiments by adding other features like eye state and Euler head angles to see how this can affect on estimating the mental performance.

For estimating the Euler head angles, we used the method proposed in [147] which employs facial landmark detection with a 68-point model followed by the Perspective-n-Point algorithm to compute the 3D head orientation. The approach establishes a geometric correspondence between detected facial features and a standard 3D face model, then decomposes the resulting rotation matrix into yaw, pitch, and roll angles using the Tait-Bryan convention with ZYX rotation order. While for estimating the eye state was done using a pretrained binary classification model that takes the face image as input and output either 0 to refer to closed eye or 1 to refer to opened eye.

The performance of the model, evaluated using the nine features BP Systolic, BP Diastolic, Heart Rate, Oxygen Saturation, respiratory rate, Roll angle, Pitch angle, Yaw angle and eyestate showed an MAE of 0.2314 and RMSE of 0.3565.

3.7 Discussion of findings

The experimental results demonstrate distinct performance variations across the developed models highlighting the connection between methods design, environmental conditions and physiological signal complexity. For respiratory rate estimation, the skeletonization-based approach achieved superior accuracy with MAE of 4.1 BPM compared to both traditional ROI methods and 3D classification architectures in stationary or low speed driving scenarios. This aligns with its theoretical advantage in minimizing illumination sensitivity through skeletal landmark tracking rather than pixel intensity variations. However, its performance degraded at vehicle speeds exceeding 3 km/h due to motion induced signal distortions exposing a critical limitation for real world vehicular applications. Comparatively, the 3D classification models exhibited moderate performance with MAE of 5.0-6.0 BPM suggesting that spatial temporal feature learning via architectures like mViT and I3D remains less effective than explicit motion based signal extraction under controlled conditions.

For heart rate estimation, the ViT-BiLSTM architecture achieved SOTA results with MAE of 2.68 BPM on LGI-PPGI and 9.996 BPM on V4V outperforming conventional rPPG techniques and demonstrating robustness to illumination variations. This can be attributed to the Vision Transformer's capacity to model longrange spatial dependencies in facial blood flow patterns combined with BiLSTM's temporal modeling of cardiac cycles. However, performance discrepancies between datasets underscore the challenge of generalizing across recording environments. A limitation partially mitigated through frame rate normalization but persisting due to unaccounted variables like skin tone diversity and motion artifacts in driving scenarios.

Oxygen saturation estimation experiments revealed the full face ROI's superiority with MAE of 1.06% on VIPL-HR over localized regions likely due to redundant photoplethysmographic signals across facial vasculature compensating for partial occlusions. This contrasts with traditional pulse oximetry's finger-based approach

but aligns with recent findings in remote photoplethysmography literature. The multi-scale vision transformer's moderate performance (101-class classification) indicates potential for improvement through regression-based formulations or hybrid architectures.

The Mistral LLM experiments yielded three key observations: 1) Full frame embeddings outperformed summary statistics emphasizing the importance of temporal dynamics in vital sign estimation; 2) Joint training of ImageBind's attention layers and projection matrices enhanced respiratory rate estimation accuracy suggesting alignment between visual embeddings and physiological targets; 3) Despite achieving SOTA results on three metrics, the LLM's computational overhead 14.8T FLOPs may limit real-time deployment compared to specialized architectures like ViT-BiLSTM.

Movement analysis results validated the STL decomposition's efficacy in isolating intentional motion from physiological micro movements though the 1.2 second analysis window may introduce latency unsuitable for safety critical applications.

The mental performance model's moderate accuracy with MAE of 0.5451 suggests vital signs alone provide limited cognitive state information necessitating multimodal integration with behavioral or neurophysiological signals. Adding the Euler angles (Roll, angle, Yaw) and the eye state improved the performance lowering the MAE to 0.2314 and RMSE to 0.3565.

Three overarching limitations emerge: 1) Dataset biases toward specific demographics (e.g., V4V's 18-66 age range) and controlled environments reduce ecological validity. 2) Motion artifacts remain a persistent challenge across all modalities, particularly in mobile scenarios; 3) Computational complexity of advanced architectures (transformers, 3D CNNs) constrains edge device deployment. Future work should prioritize lightweight model distillation, synthetic data augmentation for underrepresented populations and hybrid sensor fusion to mitigate motion-related degradation.

3.8 Summary

This chapter presented an evaluation of machine learning models for estimating vital signs from video data and mental performance. The results were published in [1;3-6; 132]. For RR estimation, the skeletonization-based method outperformed ROI-based approaches on the DriverMVT dataset achieving MAE/RMSE of 1.5/4.8 BPM under stationary conditions but vehicle vibrations above 3 km/h degraded performance. The 3D classification models and skeletonization approach achieved SOTA results on the V4V dataset with MAE/RMSE of 4.1/4.8 BPM. HR estimation using ViT-BiLSTM architectures demonstrated superior accuracy with MAE of 2.68 BPM on LGI-PPGI and 9.8 BPM on V4V compared to traditional methods. Nevertheless, generalization to driving scenarios showed increased errors with MAE of 14.09 BPM. For SpO2, face region analysis using multi-scale vision transformers yielded optimal results with MAE of 1.06%. The Mistral LLM, when finetuned with full frame embeddings using the prompt techniques achieved competitive performance across all vital signs. Mental state assessment using multimodal vital signs attained MAE/RMSE of 0.55/0.66. The correlation between the mental performance and the vital signs aligned with the literature review. Key limitations included motion artifacts in vehicular environments and lighting sensitivity in ROI-based methods. These findings highlight the potential of vision-based monitoring systems while underscoring the need for robust motion compensation in real-world applications.

161

Conclusion

This thesis presented a two-stage method for contactless assessment of operator mental performance by estimation the vital signs using video analysis and advanced machine learning techniques. The developed method leverages RGB video input to estimate respiratory rate, heart rate, blood pressure, oxygen saturation and operator mental performance state. A video-based robust technology was proposed combining illumination-invariant skeleton-based respiratory rate measurement method and ROI-driven heart rate detection method. In addition to a new application of the LLM for estimating the vital signs from sequence of images by integrating multimodal embedding and prompting technique.

The thesis achieved its goal of designing a contactless monitoring methods to estimate the transport operator vital signs and mental performance through five interconnected objectives. First, a comprehensive analysis of existing methods identified gaps in handling dynamic environments leading to the development of novel methods. The two-stage operator state assessment method using physiological data is developed to estimate the mental performance. The skeleton based respiratory rate estimation method demonstrated resilience to lighting variations with MAE of 4.1 BPM. Vision Transformer architectures achieved SOTA heart rate estimation with MAE of 2.68 BPM and finally the integration of Mistral LLM with ImageBind pioneered multimodal vital sign estimation and reducing computational costs through tensor parallelization and quantization.

Theoretical advancements included refining video analysis techniques to overcome motion artifacts and variable lighting while practical implementations such as the DriverMVT dataset and REST API—provided scalable tools for real world deployment.

The mental performance model's moderate accuracy with MAE of 0.5451 suggests vital signs alone provide limited cognitive state information necessitating multimodal integration with behavioral or neurophysiological signals. Adding the Euler angles

(Roll, angle, Yaw) and the eye state improved the performance lowering the MAE to 0.2314 and RMSE to 0.3565.

The developed methods provides advantages for real world applications: Eliminates the need for specialized equipment through RGB camera-based monitoring, easily use of developed methods for vital sign estimation, movement analysis and mental state with the integrated driver monitoring platforms.

While achieving significant advancements, several limitations warrant further investigation: 1) Motion Artifacts: Performance degradation in high-movement scenarios necessitates improved motion compensation algorithms. 3) Dataset Diversity: Current training data lacks sufficient representation of dark skinned individuals and extreme physiological states.

Future research should focus on developing adaptive filtering techniques resilient to complex motion patterns, creating standardized benchmarks for video-based physiological monitoring and investigating federated learning approaches to enhance model generalizability.

This work establishes a foundation for next-generation contactless transport operator monitoring systems, demonstrating the feasibility of computer vision and deep learning in operational environments. The methodologies and insights presented pave the way for applications in automotive safety, workplace wellness programs and remote patient monitoring.

163

References

1. Neural Network Model Combination for Video-Based Blood Pressure Estimation: New Approach and Evaluation / Batol Hamoud, Alexey Kashevnik, Walaa Othman, Nikolay Shilov // Sensors. — 2023. — Vol. 23, no. 4. — P. 1753.

2. Contactless Oxygen Saturation Detection Based on Face Analysis: An Approach and Case Study / Batol Hamoud, Walaa Othman, Nikolay Shilov, Alexey Kashevnik // Proceedings of the 2023 33rd Conference of Open Innovations Association (FRUCT). — 2023. — Pp. 54-62.

3. Human Operator Mental Fatigue Assessment Based on Video: ML-Driven Approach and Its Application to HFAVD Dataset / Walaa Othman, Batol Hamoud, Nikolay Shilov, Alexey Kashevnik // Applied Sciences. — 2024.

— Vol. 14, no. 22. — URL: https://www.mdpi.com/2076-3417/14/22/10510.

4. Estimation of motion and respiratory characteristics during the meditation practice based on video analysis / Alexey Kashevnik, Walaa Othman, Ivan A. Ryabchikov, Nikolay G. Shilov // Sensors. — 2021. — Vol. 21, no. 11.

— P. 3771.

5. Contactless Camera-Based Approach for Driver Respiratory Rate Estimation in Vehicle Cabin / Walaa Othman, Alexey Kashevnik, Ivan Ryabchikov, Nikolay Shilov // Lecture Notes in Networks and Systems. — Vol. 543. — 2023. — Pp. 429-442.

6. Remote Heart Rate Estimation Based on Transformer with Multi-Skip Connection Decoder: Method and Evaluation in the Wild / Walaa Othman, Alexey Kashevnik, Ammar Ali et al. // Sensors. — 2024. — Vol. 24, no. 3. — P. 775.

7. Assessment of drivers' mental workload: exploring the roles of multimodal physiological measures, driving measures and their combinations / Da Tao, Jiaqi Huang, Qiliang Zhang et al. // Advanced Engineering Informatics. —

2025. — Vol. 68. — P. 103796. — URL: https://www.sciencedirect.com/ science/article/pii/S1474034625006895.

8. S'anchez-Nieto Jos'e Mar'ia, Rivera-S'anchez Ulises Daniel, Mendoza--N'u nez V'ictor Manuel. Relationship between Arterial Hypertension with Cognitive Performance in Elderly: Systematic Review and Meta-Analysis // Brain Sciences. — 2021. — Vol. 11, no. 11. — P. 1445.

9. Charles Rebecca L., Nixon Jim. Measuring mental workload using physiological measures: A systematic review // Applied Ergonomics. — 2019. — Vol. 74. — Pp. 221-232. — URL: https://www.sciencedirect.com/science/article/pii/ S0003687018303430.

10. Chung S. C., Lim D. W. Changes in memory performance, heart rate, and blood oxygen saturation due to 30 2008. — Apr. — Vol. 118, no. 4. — Pp. 593-606.

11. Verkruysse W, Svaasand L.O., Nelson J.S. Remote plethysmography imaging using ambient light // Opt. Express. — 2008. — Vol. 16. — Pp. 21434-21445.

12. Poh Ming-Zher, McDuff Daniel, Picard Rosalind. Non-contact, automated cardiac pulse measurements using video imaging and blind source separation // Optics express. — 2010. — 05. — Vol. 18. — Pp. 10762-74.

13. Algorithmic Principles of Remote PPG / Wenjin Wang, Albertus C. den Brinker, Sander Stuijk, Gerard de Haan // IEEE Transactions on Biomedical Engineering. — 2017. — Vol. 64, no. 7. — Pp. 1479-1491.

14. Chen Weixuan, McDuff Daniel. DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks // CoRR. — 2018. — Vol. abs/1805.07888. — URL: http://arxiv.org/abs/1805.07888.

15. Revanur Ambareesh, Dasari Ananyananda, Tucker Conrad S., Jeni Laszlo A. Instantaneous Physiological Estimation using Video Transformers. — 2022.

16. de Haan Gerard, Jeanne Vincent. Robust Pulse Rate From Chrominance-Based rPPG // IEEE Transactions on Biomedical Engineering. — 2013.

— Vol. 60, no. 10. — Pp. 2878-2886.

17. Visual Heart Rate Estimation with Convolutional Neural Network / Radim Spetlik, Jan Cech, Vojtech Franc, Jiri Matas. — 2018. — 08.

18. Remote Heart Rate Estimation by Signal Quality Attention Network / Haoyuan Gao, Xiaopei Wu, Jidong Geng, Yang Lv // 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW).

— 2022. — Pp. 2121-2128.

19. Liu Xin, Fromm Josh, Patel Shwetak, McDuff Daniel. Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals Measurement. — 2021.

20. Automatic region-based heart rate measurement using remote photoplethys-mography / Benjamin Kossack, Eric Wisotzky, Anna Hilsmann, Peter Eisert // Proceedings of the IEEE/CVF International Conference on Computer Vision.

— 2021. — Pp. 2755-2759.

21. Elliott Malcolm, Coventry Alysia. Critical care: the eight vital signs of patient monitoring // British Journal of Nursing. — 2012. — Vol. 21, no. 10. — Pp. 621-625.

22. Seidel's Guide to Physical Examination 10th ed / Jane Ball, Joyce Dains, John Flynn et al. — 2023.

23. The association between vital signs and mortality in a retrospective cohort study of an unselected emergency department population / M. Ljunggren, M. Castren, M. Nordberg, L. Kurland // Scand J Trauma Resusc Emerg Med.

— 2016. — Vol. 24. — P. 21.

24. Emergency Department Triage Scales and Their Components: A Systematic Review of the Scientific Evidence / N. Farrohknia, M. Castren, A. Ehrenberg

et al. // Scand J Trauma Resusc Emerg Med. — 2011. — Vol. 19, no. 1. — P. 42.

25. Vital signs monitoring and nurse-patient interaction: A qualitative observational study of hospital practice / M. Cardona-Morrell, M. Prgomet, R. Lake et al. // Int J Nurs Stud. — 2016. — Vol. 56, no. Supplement C. — P. 9-16.

26. A comparison of Antecedents to Cardiac Arrests, Deaths and Emergency Intensive care Admissions in Australia and New Zealand, and the United Kingdom—the ACADEMIA study / J. Kause, G. Smith, D. Prytherch et al. // Resuscitation. — 2004. — Vol. 62, no. 3. — P. 275-82.

27. Henriksen D. P., Brabrand M, Lassen A. T. Prognosis and risk factors for deterioration in patients admitted to a medical emergency department // PLoS One. — 2014. — Vol. 9, no. 4. — P. e94649.

28. Abnormal vital signs are strong predictors for intensive care unit admission and in-hospital mortality in adults triaged in the emergency department—a prospective cohort study / C. Barfod, M. M. P. Lauritzen, J. K. Danker et al. // Scand J Trauma Resusc Emerg Med. — 2012. — Vol. 20, no. 1. — P. 28.

29. Atlef J.L. Principles of Clinical Electrocardiography // Anesthesiology. — 1980.

— Vol. 52. — P. 195.

30. Design of a Wearable 12-Lead Noncontact Electrocardiogram Monitoring System / C.-C. Hsu, B.-S. Lin, K.-Y. He, B.-S. Lin // Sensors. — 2019. — Vol. 19.

— P. 1509.

31. Low-Power ECG-Based Processor for Predicting Ventricular Arrhythmia / N. Bayasi, T. Tekeste, H. Saleh et al. // IEEE Trans. Very Large Scale In-tegr. (VLSI) Syst. — 2016. — Vol. 24. — Pp. 1962-1974.

32. Ultra-Low Power, Secure IoT Platform for Predicting Cardiovascular Diseases / M. Yasin, T. Tekeste, H. Saleh et al. // IEEE Trans. Circuits Syst. I Reg. Papers. — 2017. — Vol. 64. — Pp. 2624-2637.

33. Noncontact Wearable Wireless ECG Systems for Long-Term Monitoring / S. Majumder, L. Chen, O. Marinov et al. // IEEE Rev. Biomed. Eng. — 2018. — Vol. 11. — Pp. 306-321.

34. Nemati E., Deen M., Mondal T. A wireless wearable ECG sensor for long-term applications // IEEE Commun. Mag. — 2012. — Vol. 50. — Pp. 36-43.

35. Arcelus A., Sardar M., Mihailidis A. Design of a capacitive ECG sensor for unobtrusive heart rate measurements. — 2013.

36. Allen J. Photoplethysmography and its application in clinical physiological measurement // Physiol. Meas. — 2007. — Vol. 28. — P. R1.

37. Contact-Based Methods for Measuring Respiratory Rate / C. Massaroni, A. Nicolo, D.L. Presti et al. // Sensors. — 2019. — Vol. 19. — P. 908.

38. Wearable Photoplethysmographic Sensors—Past and Present / T. Tamura, Y. Maeda, M. Sekine, M. Yoshida // Electronics. — 2014. — Vol. 3. — Pp. 282-302.

39. Berntson G.G., Cacioppo J.T., Quigley K.S. Respiratory sinus arrhythmia: Autonomic origins, physiological mechanisms, and psychophysiological implications // Psychophysiology. — 1993. — Vol. 30. — Pp. 183-196.

40. Photoplethysmographic derivation of respiratory rate: A review of relevant physiology / D.J. Meredith, D. Clifton, P. Charlton et al. // J. Med. Eng. Technol. — 2011. — Vol. 36. — Pp. 1-7.

41. A review on wearable photoplethysmography sensors and their potential future applications in health care / D. Castaneda, A. Esparza, M. Ghamari et al. // Int. J. Biosens. Bioelectron. — 2018. — Vol. 4. — Pp. 195-202.

42. ROI analysis for remote photoplethysmography on facial video / S. Kwon, J. Kim, D. Lee, K. Park. — 2015. — Pp. 4938-4941.

43. Fiorillo Antonino, Critello Claudio, Pullano Salvatore. Theory, technology and applications of piezoresistive sensors: A review // Sensors and Actuators A: Physical. — 2018. — Vol. 281. — Pp. 156-175.

44. Hoppe P. Temperatures of expired air under varying climatic conditions // Int. J. Biometeorol. — 1981. — Vol. 25. — Pp. 127-132.

45. Fiber Bragg Grating Probe for Relative Humidity and Respiratory Frequency Estimation: Assessment During Mechanical Ventilation / C. Massaroni, D.L. Presti, P. Saccomandi et al. // IEEE Sens. J. — 2018. — Vol. 18. — Pp. 2125-2130.

46. Sekar P., Kumar S. — Blood Pressure Monitor Analog Front End. — Cypress Semiconductor Corp, San Jose, CA, USA, 2011. — 16 July.

47. Abbas A.K., Bassam R. Phonocardiography signal processing. — San Rafael, CA, USA: Morgan & Claypool Publishers, 2009.

48. Zalter R., Hodara H., Luisada A.A. Phonocardiography: I. General principles and problems of standardization // Am. J. Cardiol. — 1959. — Vol. 4. — Pp. 3-15.

49. Choi E., et al. Cuffless Blood Pressure Estimation Using Pulse Transit Time and Model-Based Approaches // Sensors. — 2021. — Vol. 21, no. 15. — P. 5216.

50. Schrumpf F., et al. Assessment of Non-Invasive Blood Pressure Prediction from PPG and rPPG Signals Using Deep Learning // Sensors. — 2021. — Vol. 21, no. 17. — P. 6022.

51. Jeong I.C., Finkelstein J. Introducing contactless blood pressure assessment using video data // ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies. — 2021. — Vol. 5, no. 1. — Pp. 1-25.

52. Islam Shekh M. M., Boric-Lubecke Olga, Lubekce Victor M. Concurrent Respiration Monitoring of Multiple Sub jects by Phase-Comparison Monopulse Radar Using Independent Component Analysis (ICA) With JADE Algorithm

and Direction of Arrival (DOA) // IEEE Access. — 2020. — Vol. 8. — Pp. 73558-73569.

53. Non-Contact Driver Respiration Rate Detection Technology Based on Suppression of Multipath Interference with Directional Antenna / Fan Yang, Zhiming He, Shisheng Guo et al. // Information. — 2020. — Vol. 11, no. 4. — URL: https://www.mdpi.com/2078-2489/11/4/192.

54. Paas Fred, Renkl Alexander, Sweller John. Cognitive Load Theory and Instructional Design: Recent Developments // Educational Psychologist. — 2003. — Vol. 38, no. 1. — Pp. 1-4.

55. Procedural learning deficits in specific language impairment (SLI): a meta-analysis of serial reaction time task performance / J. A. G. Lum, G. Conti-Ramsden, A. T. Morgan, M. T. Ullman // Cortex. — 2014. — Vol. 51, no. 100. — Pp. 1-10.

56. The potential of using actimetry to identify objective markers of depression / M. F. Borisenkov, A. A. Velichko, M. A. Belyaev, D. G. Korzun // Biological Rhythm Research. — 2025. — Pp. 1-8.

57. Meygal A. Yu., Gerasimova-Meygal L. I., Korzun D. Zh. Smartphone-Enabled mHealth Sensorics for Digital Assistance of Human Autonomic Resilience to Stress // Conference of Open Innovations Association, FRUCT. — Vol. 36. — 2024. — Pp. 885-891. — URL: https://fruct.org/publications/volume-36/ acm36/files/Mei.pdf.

58. On-Road Detection of Driver Fatigue and Drowsiness During Medium-Distance Journeys / L. Salvati, M. D'Amore, A. Fiorentino et al. // Entropy. — 2021. — Vol. 23, no. 2. — P. 135. — URL: https://www.mdpi.com/1099-4300/23Z2/135.

59. Lamti H. A., Ben Khelifa M. M., Hugel V. Mental Fatigue Level Detection Based on Event-Related and Visual Evoked Potentials Features Fusion in Virtual Indoor Environment // Cognitive Neurodynamics. — 2019. — Vol. 13, no. 3. — Pp. 271-285. — URL: https://doi.org/10.1007/s11571-019-09523-0.

60. Landolt Edmund. Methode optometrique simple // Bull Mem Soc Fran Oph-talmol. — 1888. — Vol. 6. — Pp. 213-4.

61. McCulloch Warren S, Pitts Walter. A logical calculus of the ideas immanent in nervous activity // Bulletin of Mathematical Biophysics. — 1943. — Vol. 5.

— Pp. 115-133.

62. LeCun Yann, et al. Backpropagation Applied to Handwritten Zip Code Recognition // Neural Computation. — 1989. — Vol. 1, no. 4.

63. Vaswani Ashish, Shazeer Noam, Parmar Niki et al. Attention Is All You Need.

— 2023.

64. Learning Transferable Visual Models From Natural Language Supervision / Alec Radford, Jong Wook Kim, Chris Hallacy et al. // CoRR. — 2021. — Vol. abs/2103.00020. — URL: https://arxiv.org/abs/2103.00020.

65. Girdhar Rohit, El-Nouby Alaaeldin, Liu Zhuang et al. ImageBind: One Embedding Space To Bind Them All. — 2023.

66. Zhu Bin, Lin Bin, Ning Munan et al. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment. — 2024.

67. LoRA: Low-Rank Adaptation of Large Language Models / Edward J. Hu, Yelong Shen, Phillip Wallis et al. // CoRR. — 2021. — Vol. abs/2106.09685.

— URL: https://arxiv.org/abs/2106.09685.

68. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding / Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova // CoRR. — 2018. — Vol. abs/1810.04805. — URL: http://arxiv.org/abs/1810. 04805.

69. Bai Yuntao, Kadavath Saurav, Kundu Sandipan et al. Constitutional AI: Harm-lessness from AI Feedback. — 2022.

70. Taylor Ross, Kardas Marcin, Cucurull Guillem et al. Galactica: A Large Language Model for Science. — 2022.

71. Touvron Hugo, Lavril Thibaut, Izacard Gautier et al. LLaMA: Open and Efficient Foundation Language Models. — 2023.

72. Chowdhery Aakanksha, Narang Sharan, Devlin Jacob et al. PaLM: Scaling Language Modeling with Pathways. — 2022.

73. Gunasekar Suriya, Zhang Yi, Aneja Jyoti et al. Textbooks Are All You Need.

— 2023.

74. Webcam based non-contact real-time monitoring for the physiological parameters of drivers / Qi Zhang, Guo-qing Xu, Ming Wang et al. // The 4th Annual IEEE International Conference on Cyber Technology in Automation, Control and Intelligent. — 2014. — Pp. 648-652.

75. An online PPGI approach for camera based heart rate monitoring using beat-to-beat detection / Timon Blocher, Johannes Schneider, Markus Schinle, Wilhelm Stork // 2017 IEEE Sensors Applications Symposium (SAS). — 2017.

— Pp. 1-6.

76. Simonyan Karen, Zisserman Andrew. Very Deep Convolutional Networks for Large-Scale Image Recognition. — 2015.

77. SparsePPG: Towards Driver Monitoring Using Camera-Based Vital Signs Estimation in Near-Infrared / E. Nowara, T. K. Marks, H. Mansour, A. Veer-araghavan // 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). — Los Alamitos, CA, USA: IEEE Computer Society, 2018. — jun. — Pp. 1353-135309. — URL: https://doi. ieeecomputersociety.org/10.1109/CVPRW.2018.00174.

78. Liu Si-Qi, Yuen Pong C. A General Remote Photoplethysmography Estimator with Spatiotemporal Convolutional Network // 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). — 2020.

— Pp. 481-488.

Yu Zitong, Li Xiaobai, Zhao Guoying. Remote Photoplethysmograph Signal Measurement from Facial Videos Using Spatio-Temporal Networks. — 2019.

80. The first vision for vitals (v4v) challenge for non-contact video-based physiological estimation / Ambareesh Revanur, Zhihua Li, Umur A Ciftci et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. — 2021.

81. Multimodal spontaneous emotion corpus for human behavior analysis / Zheng Zhang, Jeff M Girard, Yue Wu et al. // Proceedings of the IEEE conference on computer vision and pattern recognition. — 2016.

82. Video-based respiration monitoring with automatic region of interest detection / R.G.J. Janssen, Wenjin Wang, Andreia Vieira Moco, Gerard de Haan // Physiological Measurement. — 2016. — Vol. 37. — Pp. 100 - 114. — URL: https://api.semanticscholar.org/CorpusID:39046433.

83. Mehta Arya Deo, Sharma Hemant. Tracking Nostril Movement in Facial Video for Respiratory Rate Estimation // 2020 11th International Conference on Computing, Communication and Networking Technologies (ICCCNT). — 2020.

— Pp. 1-6.

84. Fiedler Marc-André, Rapczynski Micha, Al-Hamadi Ayoub. Fusion-Based Approach for Respiratory Rate Recognition From Facial Video Images // IEEE Access. — 2020. — Vol. 8. — Pp. 130036-130047.

85. Scebba Gaetano, Da Poian Giulia, Karlen Walter. Multispectral Video Fusion for Non-Contact Monitoring of Respiratory Rate and Apnea // IEEE Transactions on Biomedical Engineering. — 2021. — . — Vol. 68, no. 1. — P. 350-359.

— URL: http://dx.doi.org/10.1109/TBME.2020.2993649.

86. Toward a Robust Estimation of Respiratory Rate From Pulse Oximeters / Marco A. F. Pimentel, Alistair E. W. Johnson, Peter H. Charlton et al. // IEEE Transactions on Biomedical Engineering. — 2017. — Vol. 64, no. 8. — Pp. 1914-1923.

87. An End-to-End and Accurate PPG-based Respiratory Rate Estimation Approach Using Cycle Generative Adversarial Networks / Seyed Amir Hossein Aqajari, Rui Cao, Amir Hosein Afandizadeh Zargari, Amir M. Rahmani // 2021 43rd Annual International Conference of the IEEE Engineering in Medicine Biology Society (EMBC). — 2021. — Pp. 744-747.

88. RespNet: A deep learning model for extraction of respiration from photoplethysmogram / Vignesh Ravichandran, Balamurali Murugesan, Vaishali Balakarthikeyan et al. // 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). — 2019.

— Pp. 5556-5559.

89. Bian Dayi, Mehta Pooja, Selvaraj Nandakumar. Respiratory Rate Estimation using PPG: A Deep Learning Approach // 2020 42nd Annual International Conference of the IEEE Engineering in Medicine Biology Society (EMBC). — 2020. — Pp. 5948-5952.

90. RespWatch: Robust Measurement of Respiratory Rate on Smartwatches with Photoplethysmography / Ruixuan Dai, Chenyang Lu, Michael Avidan, Thomas Kannampallil. — 2021. — 05. — Pp. 208-220.

91. Luo H., et al. Smartphone-Based Blood Pressure Measurement Using Transdermal Optical Imaging Technology // Circ. Cardiovasc. Imaging. — 2019. — Vol. 12. — P. e008857.

92. Jain M., Deb S., Subramanyam A.V. Face video based touchless blood pressure and heart rate estimation // Proc. IEEE 18th Int. Workshop on Multimedia Signal Processing (MMSP). — 2016. — Pp. 1-5.

93. Secerbegovic A., et al. Blood pressure estimation using video plethysmography // Proceedings of the 2016 IEEE 13th International Symposium on Biomedical Imaging (ISBI). — Prague, Czech Republic: 2016. — 13-16 April.

— Pp. 461-464.

94. Inchi K., et al. Remote Estimation of Continuous Blood Pressure by a Convo-lutional Neural Network Trained on Spatial Patterns of Facial Pulse Waves // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. — New Orleans, LA, USA: 2022. — 19-24 June. — Pp. 2139-2145.

95. Wn B.F., et al. Contactless Blood Pressure Measurement via Remote Photoplethysmography With Synthetic Data Generation Using Generative Adversarial Network // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. — New Orleans, LA, USA: 2022. — 18-24 June. — Pp. 2130-2138.

96. Jeong I.C., Finkelstein J. A Remote Sensing Approach to Blood Pressure Assessment Using Video Data // ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies. — 2019. — Vol. 3, no. 1. — Pp. 1-22.

97. Jeong I.C., Finkelstein J. Towards Continuous Blood Pressure Estimation via Pulse Transit Time and Deep Learning: Preliminary Findings // Proceedings of the 2021 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2021 ACM International Symposium on Wearable Computers (UbiComp/ISWC). — Virtual Event, Mexico: 2021. — 21-25 September.

— Pp. 125-128.

98. Visvanathan K., et al. Blood Pressure Estimation from Facial Photoplethys-mographic Signals // Proceedings of the 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). — Seoul, Korea: 2020. — 16-19 December. — Pp. 70-77.

99. Ding X., et al. Estimation of Blood Pressure Using Photoplethysmography Only: Toward Real-Time Blood Pressure Monitoring // J. Med. Internet Res.

— 2019. — Vol. 21, no. 6. — P. e13365.

100. Djeldjli L., et al. Contactless blood pressure measurement using photoplethys-mographic sensors // Biomed. Eng. Online. — 2019. — Vol. 18. — P. 66.

101. Gonzalez Viejo C., et al. A Non-contact Method to Measure Blood Pressure: A New Application of Rhythmic Modulation of Light // Sci. Rep. — 2018. — Vol. 8. — P. 16079.

102. Zhou X., et al. Estimating Blood Pressure Using Photoplethysmography Signals: A Hybrid Polynomial Neural Network Approach // Proceedings of the IEEE Engineering in Medicine and Biology Society (EMBC). — Vol. 18. — 2018. — P. 209.

103. Rong Y, Li K. Photoplethysmographic Imaging Based Blood Pressure Estimation Using Artificial Neural Networks // Sensors. — 2018. — Vol. 18, no. 4.

— P. 1243.

104. Akamatsu Y, Onishi Y, Imaoka H. Blood oxygen saturation estimation from facial video via dc and ac components of spatiotemporal map // ArXiv. — 2022. — Vol. abs/2212.07116.

105. Ding X., Nassehi D., Larson E. C. Measuring oxygen saturation with smart-phone cameras using convolutional neural networks // IEEE Journal of Biomedical and Health Informatics. — 2019. — Vol. 23, no. 6. — Pp. 2603-2610.

106. Casalino G., Castellano G., Zaza G. A mhealth solution for contact-less self-monitoring of blood oxygen saturation // 2020 IEEE Symposium on Computers and Communications (ISCC). — 2020. — Pp. 1-7.

107. Remote blood oxygen estimation from videos using neural networks / J. Math-ew, X. Tian, C.-W. Wong et al. // IEEE Journal of Biomedical and Health Informatics. — 2023.

108. Noncontact monitoring of blood oxygen saturation using camera and dual-wavelength imaging system / D. Shao, C. Liu, F. Tsow et al. // IEEE Transactions on Biomedical Engineering. — 2016. — Vol. 63, no. 6. — Pp. 1091-1098.

109. Non-contact detection of oxygen saturation based on visible light imaging device using ambient light / L. Kong, Y. Zhao, L. Dong et al. // Optics Express.

— 2013. — Vol. 21, no. 15. — Pp. 17464-71.

110. Towards a machine learning-based digital twin for non-invasive human bio-signal fusion / I. Al-Zyoud, F. Laamarti, X. Ma et al. // Sensors. — 2022. — Vol. 22, no. 24. — https://www.mdpi.com/1424-8220/22/24/9747.

111. Remote blood oxygen estimation from videos using neural networks / J. Mathew, X. Tian, M. Wu, C.-W. Wong // arXiv e-prints. — 2021. — P. arX-iv:2107.05087.

112. Noncontact monitoring of blood oxygen saturation using camera and dual-wavelength imaging system / D. Shao, C. Liu, F. Tsow et al. // IEEE Trans. Biomed. Eng. — 2016. — Vol. 63, no. 6. — Pp. 1091-1098.

113. Contactless blood oxygen estimation from face videos: A multi-model fusion method based on deep learning / M. Hu, X. Wu, X. Wang et al. // Biomed Signal Process Control. — 2023. — Mar. — Vol. 81. — P. 104487.

114. PPGnet: Deep Network for Device Independent Heart Rate Estimation from Photoplethysmogram / A. Shyam, Vignesh Ravichandran, S.P. Preejith et al. // 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). — 2019. — Pp. 1899-1902.

115. CNNs for Heart Rate Estimation and Human Activity Recognition in Wrist Worn Sensing Applications / Eoin Brophy, Willie Muehlhausen, Alan F. Smeaton, Tomas E. Ward // 2020 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops).

— 2020. — Pp. 1-6.

116. Schrumpf Fabian, Frenzel Patrick, Aust Christoph et al. Assessment of deep learning based blood pressure prediction from PPG and rPPG signals. — 2021.

117. Wang Haopeng, Zhou Yufan, Saddik Abdulmotaleb El. VitaSi: A real-time con-tactless vital signs estimation system // Computers and Electrical Engineering.

— 2021. — Vol. 95. — P. 107392. — URL: https://www.sciencedirect.com/ science/article/pii/S0045790621003591.

118. Mobile Robotic Platform for Contactless Vital Sign Monitoring / Hsin-Wei Huang, Jiahao Chen, Peter R. Chai et al. // Cyborg Bionic Syst.

— 2022. — Vol. 2022. — P. 9780497.

119. Openpose: Realtime Multi-person 2D Pose Estimation Using Part Affinity Fields / Z. Cao, G. Hidalgo Martinez, T. Simon et al. // IEEE Transactions on Pattern Analysis and Machine Intelligence. — 2019.

120. Integral Human Pose Regression / X. Sun, B. Xiao, S. Liang, Y. Wei // CoRR.

— 2017. — Vol. abs/1711.08229. — URL: http://xxx.lanl.gov/abs/1711.08229.

121. Selflow: Self-supervised Learning of Optical Flow / P. Liu, M. R. Lyu, I. King, J. Xu // CVPR. — 2019.

122. He Kaiming, Girshick Ross, Dollar Piotr. Rethinking ImageNet Pre-Train-ing // 2019 IEEE/CVF International Conference on Computer Vision (ICCV).

— IEEE, 2019. — . — URL: https://doi.org/10.1109/iccv.2019.00502.

123. Neurokit2: A Python Toolbox for Neurophysiological Signal Processing / Dominique Makowski, Tommy Pham, Zhi Jie Lau et al. // Behavior Research Methods. — 2021. — Pp. 1-8.

124. Optimized Breath Detection Algorithm in Electrical Impedance Tomography / Davood Khodadad, Sven Nordebo, Bertram Müller et al. // Physiological Measurement. — 2018. — Vol. 39, no. 9. — P. 094001.

125. Carreira Joao, Zisserman Andrew. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. — 2018.

126. Deep Residual Learning for Image Recognition / Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun // CoRR. — 2015. — Vol. abs/1512.03385. — URL: http://arxiv.org/abs/1512.03385.

127. Towards Fast, Accurate and Stable 3D Dense Face Alignment / Jianzhu Guo, Xiangyu Zhu, Yang Yang et al. // Proceedings of the European Conference on Computer Vision (ECCV). — 2020.

128. Guo Jianzhu, Zhu Xiangyu, Lei Zhen. 3DDFA. — https://github.com/ cleardusk/3DDFA. — 2018.

129. FaceBoxes: A CPU real-time face detector with high accuracy / Shifeng Zhang, Xiangyu Zhu, Zhen Lei et al. // 2017 IEEE International Joint Conference on Biometrics (IJCB). — 2017. — Pp. 1-9.

130. Tan Mingxing, Le Quoc V. EfficientNet: Rethinking Model Scaling for Convo-lutional Neural Networks // CoRR. — 2019. — Vol. abs/1905.11946. — URL: http://arxiv.org/abs/1905.11946.

131. DriverMVT: In-Cabin Dataset for Driver Monitoring including Video and Vehicle Telemetry Information / Walaa Othman, Alexey Kashevnik, Am-mar Ali, Nikolay Shilov // Data. — 2022. — Vol. 7, no. 5. — URL: https://www.mdpi.com/2306-5729/7/5/62.

132. Othman Walaa, Kashevnik Alexey. Video-Based Real-Time Heart Rate Detection for Drivers Inside the Cabin Using a Smartphone // 2022 IEEE International Conference on Internet of Things and Intelligence Systems (Io-TaIS). — 2022. — Pp. 142-146.

133. The first vision for vitals (v4v) challenge for non-contact video-based physiological estimation / Ambareesh Revanur, Zhihua Li, Umur A Ciftci et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. — 2021.

134. Multimodal spontaneous emotion corpus for human behavior analysis / Zheng Zhang, Jeff M Girard, Yue Wu et al. // Proceedings of the IEEE conference on computer vision and pattern recognition. — 2016.

135. Local Group Invariance for Heart Rate Estimation from Face Videos in the Wild / Christian Pilz, Sebastian Zaunseder, Jarek Krajewski, Vladimir Blazek. — 2018. — 06.

136. et al. MONAI Consortium. Project MONAI. — https://zenodo.org/record/ 4323059#.YXaMajgzaUk. — 2020. — Accessed on 25 May 2020.

137. Tan Mingxing, Le Quoc V. EfficientNet: Rethinking Model Scaling for Convo-lutional Neural Networks // CoRR. — 2019. — Vol. abs/1905.11946. — URL: http://arxiv.org/abs/1905.11946.

138. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows / Ze Liu, Yutong Lin, Yue Cao et al. // Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). — 2021.

139. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale / Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov et al. // ICLR. — 2021.

140. Tune: A Research Platform for Distributed Model Selection and Training / Richard Liaw, Eric Liang, Robert Nishihara et al. // arXiv preprint arX-iv:1807.05118. — 2018.

141. Niu Xuesong, Han Hu, Shan Shiguang, Chen Xilin. VIPL-HR: A Multi-modal Database for Pulse Estimation from Less-constrained Face Video. — 2018.

142. Contactless Blood Oxygen Saturation Estimation from Facial Videos Using Deep Learning / Chia-Hsiang Cheng, Zhen Yuen, Sizhe Chen et al. // Bioengineering. — 2024. — Vol. 11.

143. Contactless blood oxygen estimation from face videos: A multi-model fusion method based on deep learning / Meng Hu, Xi Wu, Xueshi Wang et al. // Biomedical Signal Processing and Control. — 2023. — Vol. 81. — P. 104487.

144. Casalino Gabriella, Castellano Giovanna, Zaza Gianluigi. A mHealth solution for contact-less self-monitoring of blood oxygen saturation // Proceedings of the 2020 IEEE Symposium on Computers and Communications (ISCC). — 2020. — Pp. 1-7.

145. Bal Ufuk. Non-contact estimation of heart rate and oxygen saturation using ambient light // Biomedical Optics Express. — 2015. — Vol. 6. — Pp. 86-97.

146. OperatorEYEVP: Operator Dataset for Fatigue Detection Based on Eye Movements, Heart Rate Data, and Video Information / Svetlana Kovalenko,

Anton Mamonov, Vladislav Kuznetsov et al. // Sensors. — 2023. — Vol. 23, no. 13. — URL: https://www.mdpi.com/1424-8220/23/13/6197.

147. Human Head Angle Detection Based on Image Analysis / Alexey Ka-shevnik, Ammar Ali, Igor Lashkov, Dmitry Zubok // Proceedings of the Future Technologies Conference (FTC) 2020, Volume 1 / Ed. by Kohei Arai, Supriya Kapoor, Rahul Bhatia. — Cham: Springer International Publishing, 2021. — Pp. 233-242.

181

Publications

ртсотйсоеам «ДШРДЩШШ

жжижж

СВИДЕТЕЛЬСТВО

о государственной регистрации программы для ЭВМ

№ 2022617615

Система детектирования усталости оператора на основе камеры и трекера глаз

Правообладатель: Федеральное государственное бюджетное учреждение науки "Санкт-Петербургский Федеральный исследовательский центр Российской академии наук" №)

Авторы: Кашевник Алексей Михайлович Булыгин Александр Олегович (№), Шилов Николай Германович (№), Осман Валаа (SY)

Заявка № 2022616916

Дата поступления 19 апреля 2022 Г.

Дата государственной регистрации

в Реестре программ для эвм 25 апреля 2022 г.

Руководитель Федеральной службы по интеллектуальной собственности

Сертификат баьа0077е14е4 0ГСэ94ейЬ[124145й5с7 владелец Зубов Юрий Сергеевич Действителен с гю 26.05.2Q23

Ю.С. Зубов

теотшйшдж Фждшрдшрш

ж

ж

ж ж

ж ж

ж ж

ж ж ж

ж ж ж

ж ж ж

вшжжжж

СВИДЕТЕЛЬСТВО

о государственной регистрации программы для ЭВМ

№ 2023683199

Среда тестирования моделей машинного обучения для поддержки принятия решений при моделировании

предприятий

Правообладатель: Федеральное государственное бюджетное учреждение науки "Санкт-Петербургский Федеральный исследовательский центр Российской академии наук "

т)

Авторы: Осман Валаа Шилов Николай Германович

Заявка № 2023682935

Дата поступления 03 ноября 2023 Г.

Дата государственной регистрации

в Реестре программ для ЭВМ 03 ноября 2023 г.

Руководитель Федеральной службы по интеллектуальной собственности

сертификат 429 ЬйвШв 3853|Б4Ьвге^ВЗЬ73Ь4ва7 ЮО.С. Зубов

йлвде^ч Зубов Юрий Свргвевич Действителен с по 02.05.2024

теотшйшдж Фждшрдшрш

ж

ж

ж ж

ж ж

ж ж

ж ж ж

ж ж ж

ж ж ж

вшжжж-

СВИДЕТЕЛЬСТВО

о государственной регистрации программы для ЭВМ

№ 2023683219

Модуль создания и обучения нейросетевых моделей для поддержки принятия решений при моделировании

процессов

Правообладатель: Федеральное государственное бюджетное учреждение науки "Санкт-Петербургский Федеральный исследовательский центр Российской академии наук "

т)

Автор(ы): Осман Валаа (КЦ)

Заявка № 2023682955

Дата поступления 03 ноября 2023 Г.

Дата государственной регистрации

в Реестре программ для ЭВМ 03 ноября 2023 г.

Руководитель Федеральной службы по интеллектуальной собственности

Сйртфикнт42.5ЬЙЕЙГ^^&МЬаге^ВЗЬ73Маа7 ЮО.С. Зубов

йлареоеч Зуйов Юрий Свргвевич

Действителен с по 02.05.2024

Общество с ограниченной ГНИ ПА U В U КО ™

ответственностью «НПК Центр Комплексного Оснащения»

представитель компании ANT Neuro на территории

РФ и Казахстана ИНН 2222827469; КПП 222501001

Юр/адрес: 656066, г. Барнаул, улица Партизанская, 146-144, Р/счет: 40702810370010206970 в Московский филиал АО КБ «МОДУЛЬБАНК»

127015 РФ г. Москва, ул. Новодимитровская, д.2, корпус 1

к/с 30101810645250000092, БИК 044525092, ОГРН 1022200525841

Тел.: 8 800 201-61-70, +7 (3852) 25-19-53 № 305 от «02» октября 2025 г.

А К Т

об использовании результатов диссертационной работы Валаа Осман «Методы видеомониторинга физиологических характеристик и умственной работоспособности операторов транспортных средств» для комплексного мониторинга состояния

человека.

Директор компании ООО «НИК Центр Комплексного Оснащения» -

Когнитивика™ подтверждает, что:

1. Основные результаты, полученные В. Осман в рамках диссертационной работы были использованы в компании при создании комплексной системы для мониторинга функционального состояния человека, включающей использование трекеров глазодвигательной активности, контактных сенсоров и видеокамеры.

2. Разработанной двухэтапный метод обнаружения умственной работоспособности оператора, основанный на оценке физиологических показателей позволил повысить эффективность оценки состояния оператора по сравнению с прямым методом.

3. Разработанный метод оценки физиологических показателей на основе анализа последовательности изображений с использованием крупных языковых моделей через интеграцию мультимодального эмбеддинга и техник работы с промтом позволил существенно повысить точность определения физиологических характерист

г ЖГ1'

С уважением

Директор Т. В. Устименко

ООО «НИК Центр Комплексного Оснащения»

г.

Тел: 8 800 201-61-70, +7 (3852) 25-19-53 Сайт: www.cognitivika.ru

Факс: +7 (3852) 25-19-53 E-mail: info@cognitivika.ru

sensors

Article

Remote Heart Rate Estimation Based on Transformer with Multi-Skip Connection Decoder: Method and Evaluation in the Wild

Walaa Othman *©, Alexey Kashevnik *'*©, Ammar Ali 2, Nikolay Shilov x© and Dmitry Ryumin x©

©

check for updates

Citation: Othman, W.; Kashevnik, A.; Ali, A.; Shilov, N.; Ryumin, D. Remote Heart Rate Estimation Based on Transformer with Multi-Skip Connection Decoder: Method and Evaluation in the Wild. Sensors 2024, 24, 775. https://doi.org/10.3390/ s24030775

Academic Editor: Cecilia Garcia

Received: 3 January 2024 Revised: 20 January 2024 Accepted: 23 January 2024 Published: 25 January 2024

Copyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/).

1 St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia; walaa_othman@itmo.ru (W.O.); nick@iias.spb.su (N.S.); ryumin.d@iias.spb.su (D.R.)

2 Information Technology and Programming Faculty, ITMO University, 191002 St. Petersburg, Russia; ammarali32@itmo.ru

* Correspondence: alexey.kashevnik@iias.spb.su

Abstract: Heart rate is an essential vital sign to evaluate human health. Remote heart monitoring using cheaply available devices has become a necessity in the twenty-first century to prevent any unfortunate situation caused by the hectic pace of life. In this paper, we propose a new method based on the transformer architecture with a multi-skip connection biLSTM decoder to estimate heart rate remotely from videos. Our method is based on the skin color variation caused by the change in blood volume in its surface. The presented heart rate estimation framework consists of three main steps:

(1) the segmentation of the facial region of interest (ROI) based on the landmarks obtained by 3DDFA;

(2) the extraction of the spatial and global features; and (3) the estimation of the heart rate value from the obtained features based on the proposed method. This paper investigates which feature extractor performs better by captioning the change in skin color related to the heart rate as well as the optimal number of frames needed to achieve better accuracy. Experiments were conducted using two publicly available datasets (LGI-PPGI and Vision for Vitals) and our own in-the-wild dataset (12 videos collected by four drivers). The experiments showed that our approach achieved better results than the previously published methods, making it the new state of the art on these datasets.

Keywords: remote heart rate estimation; vital signs; machine learning; video analysis

1. Introduction

In recent years, there has been a noticeable surge in stress levels in our daily lives, underscoring the crucial need to proactively monitor overall human health. The assessment of vital signs, such as respiratory rate, heart rate, blood pressure, and oxygen saturation levels, has garnered increasing attention among researchers. These vital signs serve as pivotal indicators in evaluating the holistic well-being of the human body and play crucial roles in various applications, including the detection of fatigue and stress levels [1,2], disease diagnosis [3], and assessments related to meditation and fitness.

Various systems have been devised to measure these vital signs, encompassing devices like blood pressure monitors, oximeters, and electrocardiogram machines. While these systems deliver precise measurements, their requirement for direct contact with the subject makes them more suited for clinical settings. However, this feature becomes a drawback, causing discomfort, when the subject is engaged in physical exercises or driving.

An alternative approach involves leveraging cameras to remotely estimate vital signs [4-6]. This is achieved by capturing variations in light reflected from the skin, induced by physiological processes such as the heartbeat and breathing. This non-contact methodology not only offers convenience but also opens up possibilities for unobtrusive health monitoring, proving especially valuable in situations where direct contact is impractical or uncomfortable.

However, research on the topic of remote heart rate detection is usually conducted in good lighting and shaking conditions that do not reflect settings like vehicle cabins and streets. In this paper, we propose a new method for remote heart rate detection and evaluate it using our own dataset of drivers in vehicle cabins recorded in the wild. The contributions of this paper can be summarized as follows:

1. We investigate the best feature extractor to capture heart rate spatial information from the face.

2. We investigate the optimal number of frames to achieve better accuracy in captioning the global features.

3. We propose a new method based on the vision transformer architecture to estimate the heart rate with a novel pre-processing technique.

4. We propose a novel approach of treating a single regression value problem by extracting intervals, which optimizes the performance and opens the way for using new techniques like the intersection over union of the intervals and complex loss functions to optimize interval values.

5. We evaluated the proposed approach on three different datasets (including our own dataset recorded in the wild). In comparison with the previously published methods, our approach achieved the highest accuracy.

The rest of the paper is organized as follows. Section 2 summarizes the state-of-the-art methods used to estimate the heart rate. Section 3 introduces the proposed method in detail. The experiments conducted are presented in Section 4. Finally, the conclusion is provided in Section 5.

2. Related Work

Remote heart rate estimation has recently drawn the attention of researchers due to its importance in health monitoring without causing discomfort to the subject. Several methods have been used to extract heart rate information from facial videos using mathematical approaches and signal processing techniques [7-10]. In [7], the authors extracted heart rate information by highlighting the green channel only, since it contained the highest photoplethysmography (PPG) signal-to-noise ratio. Refs. [8,11] introduced different mathematical reflection models of the skin that were used to build models for extracting remote PPG signals. Ref. [10] applied the method proposed in [8] to five sub-regions (the face, forehead, right and left cheek, and nose) to extract remote PPG signals, followed by fast Fourier transform and bandpass filtering to obtain a power spectrum density for each sub-region. A sorting function was applied to each region to determine which ROI would be used to extract the heart rate. With the advances in computer vision and machine learning techniques, more accurate methods have been proposed. Some researchers have used deep convolutional neural networks to estimate the heart rate from videos [12-17]. Ref. [12] improved the robustness of the heart rate estimation model under different motion and illuminance conditions by using a skin reflection model and an appearance information attention mechanism. Refs. [13,14] proposed a two-branch model: appearance and movement. The input to the former branch was the current frame, while the input to the latter was the normalized difference between the current and the next frames. The output of these branches was then combined to estimate the heart rate. Ref. [17] introduced a hybrid-CAN-RNN framework by adding a bidirectional GRU on top of the hyprid-CAN model proposed in [13]. Ref. [16] proposed aggregating the remote PPG signals from multiple skin areas to improve reliability. Other researchers proposed using the long short-term memory (LSTM) layer to capture the features between frames [18,19]. Ref. [18] proposed two branches of neural networks: the first one used 3D pooled convolutional layers with several consecutive frames as an input, while the second branch used a 2D convolutional layer with LSTM layers, and the input to this branch was one frame at a time. The outputs of the branches were then combined to estimate the heart rate. Ref. [19] proposed a model consisting of two LSTM networks: one to estimate the heart rate sampling point, and the other to predict the signal quality. These two values were then fed into an attention-based

model to estimate the average heart rate. A video transformer was proposed in [20]. The authors' main contribution was calculating the loss in the frequency domain instead of the time domain. Ref. [21] introduced a self-supervised framework. A face extractor was first used to detect the face, and then a remote PPG estimator based on a 3D-CNN was used to extract the PPG signal. The authors proposed a new way to augment the dataset by stretching and squeezing the video. When the video is squeezed, the heart rate should increase, and vice versa. Ref. [22] proposed an end-to-end architecture using 3D depth-wise separable convolution layers with residual connections for heart rate estimation from videos.

The existing heart rate estimation models based on deep learning methods use LSTM-based models, convolutional models, attention mechanisms, or a combination of the above to predict the heart rate value. In our paper, we present an architecture based on a vision transformer with a multi-skip connection decoder. Unlike previously published methods, our model takes the spatial features from five different layers of the feature extractor, allowing the model to capture diverse features. These features are then fed into BiLSTM layers, followed by 1D Conv and linear layers outputting five different arrays. Each array contains the predicted minimum, maximum, and average HR in the selected window. These values are then combined together using the weighted average method to predict the final heart rate value.

3. Proposed Approach

3.1. General Approach

In this section, we describe the developed approach for predicting heart rate. Figure 1 shows the approach based on the transformer architecture with a multi-skip connection decoder.

Figure 1. The developed approach based on multi-skip connection decoder.

To predict the heart rate, we used a sliding window with a length of 15 frames and a stride of 15 frames. The length of the sliding window was chosen based on our experiments (more details are described in Section 4.3), while the stride was chosen to reduce the correlation between the training data to the minimum. We propose processing the input frames using 3DDFA_V2 [23,24] to extract 68 facial landmarks: 17 for the face, 10 for the eyebrow, 9 for the nose, 10 for the eyes, and 22 for the mouth. 3DDFA_V2 is the SOTA on the Florence dataset for 3D face reconstruction (in our case, the person could move their head, and 3DDFA_V2 could keep providing the facial landmarks even when the face

was completely turned to the left or right due to the fact that it provided 3D coordinates). Figure 2 shows a 3D face reconstruction of a person from our dataset turning his head to the side produced using 3DDFA_V2. In addition, according to [25], 3DDFA_V2 achieved more robust and accurate performance in different movement conditions compared with OpenFace 2.0 [26] and MediaPipe [27].

Figure 2. 3DDFA_V2 face reconstruction result of a person turning his head to the side.

To detect the forehead area, we added two more landmarks based on the left-most and right-most eyebrow points. The obtained landmarks were used to segment the face by keeping only the pixels inside the landmarks and filling all other pixels with zeros. The obtained images were then cropped and resized into 224 by 224 pixels. We applied normalization to the images before feeding them into the model.

In this paper, we propose two architectures for heart rate estimation. The first one consists of a feature extractor based on the vision transformer (VIT) [28] with BiLSTM and a linear layer with one output that represents the predicted HR. The second one takes the outputs' connections from five different blocks of the feature extractor and feeds them into several BiLSTM layers followed by 1D convolution and linear layers with three outputs each. The three outputs represent the minimum, maximum, and average HR in the sliding frame. We trained the first model with the mean absolute error (MAE) as the loss function. The MAE was calculated between the predicted and the real heart rate values. During the training phase of the second model, the five outputs were averaged together, and the loss function was calculated as the MAE of the three values. In the testing phase, the weighted average was used to calculate the predicted heart rate from the three obtained values. More details of the proposed architectures are provided in Section 3.2.

Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.