Методы мониторинга безопасности и обнаружения аномалий в среде Kubernetes с использованием агентов и методов машинного обучения тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Дарвиш Гадир
- Специальность ВАК РФ00.00.00
- Количество страниц 290
Оглавление диссертации кандидат наук Дарвиш Гадир
Реферат
Synopsis
Introduction
CHAPTER 1. Literature review
1.1 Overview of core concepts of Kubernetes security
1.1.1 Virtualization and Containers
1.1.2 Background on Kubernetes and its importance
1.1.3 Security challenges in Kubernetes environments
1.1.4 Kubernetes security practices
1.1.5 The need for devsecops in Kubernetes security research
1.1.6 Machine learning approaches for anomaly detection
1.1.7 The role of machine learning in enhancing security
1.1.8 DoS attacks and their impact on Kubernetes security
1.2 Existing monitoring solutions in Kubernetes
1.3 Related work
1.3.1. Kubernetes security
1.3.2. DoS attack detection in Kubernetes
1.3.3. Security in Kubernetes using Machine Learning
1.3.4. DoS attack detection using machine learning
1.4 Research Gaps and problem Statement
1.5 Chapter Synthesis
CHAPTER 2. Method for real-time security monitoring in Kubernetes using a lightweight telemetry agent
2.1 Overview of method for real-time security monitoring in Kubernetes using a
lightweight telemetry agent
2.1.1 Architecture of the monitoring agent
2.1.2 Collection and export of the metrics
2.2 Performance evaluation of the proposed method
2.2.1 Implementation
2.2.2 Integration with Kubernetes in the cloud
2.2.3 Evaluation setup
2.2.4 Evaluation results
2.3 Chapter Synthesis
CHAPTER 3. Method for anomaly detection across multiple Kubernetes frameworks using ensemble machine learning
3.1 Processing data and feature engeneering
3.2 Selecting and preliminary evaluation of machine learning algorithms for anomaly detection in Kubernetes
3.2.1 Evaluation of ML-models for anomaly detection in Kubernetes
3.2.2 Evaluation setup
3.2.3 Evaluation results
3.4 Development of method for anomaly detection across multiple Kubernetes frameworks using ensemble machine learning
3.4.1 Evaluation setup
3.4.2 Evaluation results
3.5 Chapter Synthesis
CHAPTER 4. Method for adaptive anomaly detection using a dual-agent architecture with feedback loop
4.1 Approach utilizing two agents for data collection and analysis
4.2 Overview of method for adaptive anomaly detection using a dual-agent architecture with feedback loop
4.2.1 Integration with existing systems
4.2.1 Scalability and performance
4.2.3 Adaptive learning mechanism for continuous improvement
4.2.4 Implementation and deployment in Kubernetes
4.3 Real-World evaluation of the developed method and case studies
4.3.1 Case study 1: Resource exhaustion attack
4.3.2 Case Study 2: Network Traffic Anomaly
4.3.3 Case Study 3: Unauthorized Access Attempt
4.3.4 Case Study 4: Deployment Anomaly
4.4 Chapter Synthesis
CHAPTER 5. Discussion
5.1 Summary of findings
5.1.1. Performance of the monitoring agent
5.1.2. Accuracy and effectiveness of machine learning models
5.1.3. Real-world deployment and case studies
5.1.4. Scalability and robustness
5.1.5. Overall impact and benefits
5.2 Significance of findings
5.3 Challenges and limitations
5.4 Comparison with existing solutions
5.5 Implications for Kubernetes security
Conclusion
List of Abbreviations
References
APPLICATIONS
Appendix A: Publications
Реферат
Общая характеристика диссертации
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Обнаружение вторжений на основе многозадачного глубокого обучения для сетей Интернета вещей2025 год, кандидат наук Дун Хуэйяо
Методы повышения производительности трансформеров на основе приближённого двудольного соответствия и сингулярного разложения2025 год, кандидат наук Али Аммар
Синтез изображений лиц на основе генеративных методов машинного обучения с применением к распознаванию лиц2022 год, кандидат наук Зено Бассель
Методы видеомониторинга физиологических характеристик и умственной работоспособности операторов транспортных средств2025 год, кандидат наук Осман Валаа
Система управления базами знаний для управления процессами интеллектуального анализа данных2021 год, кандидат наук Мань Тяньсин
Введение диссертации (часть автореферата) на тему «Методы мониторинга безопасности и обнаружения аномалий в среде Kubernetes с использованием агентов и методов машинного обучения»
Актуальность темы исследования
Широкое распространение Kubernetes как основной платформы для оркестровки контейнеров коренным образом изменило подходы к проектированию, развертыванию и сопровождению современных программных систем. Переход организаций на микросервисную архитектуру и контейнеризированные приложения для обеспечения масштабируемости, надежности и скорости разработки обуславливает потребность в возможностях автоматизации и управления, которые предоставляет Kubernetes для оркестрации жизненных циклов контейнеров в распределенных вычислительных средах. Однако эта трансформация одновременно породила сложные проблемы безопасности, в частности, в области обеспечения доступности и производительности сервисов в условиях злонамеренных воздействий.
Одной из наиболее критических угроз в этой области являются атаки типа «отказ в обслуживании» (DoS), направленные на истощение ресурсов системы, что приводит к деградации сервисов, увеличению задержек или полной недоступности. В системах на базе Kubernetes последствия DoS-атак усугубляются в силу динамической природы платформы, ее распределенной архитектуры и активного использования API-коммуникаций и автоматизированного планирования. Злоумышленники могут использовать ошибки конфигурации, уязвимые компоненты, незащищенные сервисы или механизмы оркестрации для организации крупномасштабных сбоев, которые сложно обнаружить традиционными методами обеспечения безопасности.
Традиционные меры противодействия, такие как статические правила межсетевых экранов, инструменты периодического сканирования и системы обнаружения вторжений на основе правил, уже недостаточны в облачных средах (cloud-native). Эти устаревшие инструменты не способны распознавать адаптивный и многовекторный характер современных атак. Кроме того, их недостаточная
оперативность и ограниченная контекстная осведомленность часто приводят к высокому проценту ложных срабатываний, пропуску реальных угроз и неэффективному реагированию на инциденты. Неспособность современных систем учитывать поведенческий контекст и оперативно адаптироваться к новым угрозам ограничивает их применимость в таких средах, как Kubernetes, где рабочие нагрузки, шаблоны трафика и конфигурации изменяются быстро и часто.
В условиях роста сложности киберугроз потребность в интеллектуальных, контекстно-зависимых и адаптивных системах защиты становится все более актуальной. Распределенные DoS-атаки влияют не только на производительность отдельных сервисов, но и могут распространяться через взаимосвязанные контейнеры и микросервисы, потенциально нарушая работу целых экосистем. Эта проблема требует создания нового поколения решений для мониторинга безопасности и обнаружения угроз, которые сочетают наблюдение (observability) в реальном времени, масштабируемую архитектуру и методы машинного обучения (МО) для интеллектуального выявления аномалий.
Существующие научные работы демонстрируют применение методов ML для обнаружения аномалий и сетевых атак. В то время как традиционные системы эффективны в статических, однородных средах, кластеры Kubernetes характеризуются неоднородностью и динамическим планированием, что делает многие существующие решения неэффективными. Исследователи, такие как Садик, Као и Али, предложили методы обнаружения аномалий с использованием МО; однако этим методам часто не хватает работы в реальном времени, способности к обобщению для различных фреймворков или проверки в промышленных условиях. К ограничениям современных подходов относятся узкий охват собираемых данных (только на системном уровне), неспособность к обобщению для разнородных рабочих нагрузок и отсутствие механизмов адаптивного обучения.
Цель диссертации
Целью данной диссертации является обеспечение высокой точности обнаружения, низкого уровня ложных срабатываний и минимального влияния на производительность за счет разработки агент-ориентированных методов, основанных на машинном обучении, для обнаружения DoS-атак в средах Kubemetes в реальном времени.
Задачи исследования
Для достижения поставленной цели решались следующие задачи:
1. Анализ существующих механизмов обнаружения DoS-атак и выявление их ограничений в контексте Kubemetes.
2. Разработка и реализация метода мониторинга безопасности в реальном времени в средах Kubemetes на основе легковесного агента, способного собирать метрики как на системном, так и на прикладном уровне в реальном времени.
3. Разработка метода обнаружения аномалий для множества прикладных фреймворков в Kubemetes с использованием предварительно обученных моделей машинного обучения, применяемых к интегрированной телеметрии от разнородных рабочих нагрузок.
4. Разработка и валидация метода адаптивного обнаружения аномалий на основе стратегии с двумя агентами, которая сочетает вывод моделей в реальном времени и непрерывное переобучение на основе обратной связи.
5. Разработка метода уточнения моделей на основе обратной связи, который включает пользовательский ввод (истинно/ложно положительные/отрицательные срабатывания) для динамического улучшения производительности обнаружения с течением времени.
6. Экспериментальная оценка предложенных методов на реальном кластере Kubemetes с использованием различных рабочих нагрузок и смоделированных DoS-атак для оценки точности обнаружения, производительности и нагрузки.
7. Сравнение предложенных методов с существующими решениями с использованием метрик, таких как точность обнаружения, задержка, адаптивность и ресурсоэффективность.
Объект и предмет исследования
Объектом исследования являются методы мониторинга в реальном времени и обнаружения аномалий в распределенных вычислительных системах на базе Kubemetes. Предметом исследования являются методы интеграции сбора телеметрии на основе агентов и техник обнаружения на основе машинного обучения для выявления DoS-атак, и аномалий в контейнеризированных, микросервис-ориентированных средах.
Положения, выносимые на защиту
1. Метод мониторинга безопасности в реальном времени в средах КиЬегпе1еБ, отличающийся использованием легковесного агента сбора телеметрии, который собирает метрики системного и прикладного уровня с минимальными ресурсными затратами. Данный метод обеспечивает непрерывную наблюдаемость распределенных рабочих нагрузок и поддерживает обнаружение в реальном времени.
2. Метод обнаружения аномалий для множества прикладных программных фреймворков в КиЬегпе1еБ, характеризующийся интеграцией ансамблевых моделей машинного обучения, обученных на данных телеметрии из множества источников. Данный метод хорошо обобщается для различных фреймворков и обеспечивает высокую точность обнаружения.
3. Метод адаптивного обнаружения аномалий в КиЬегпе1еБ, основанный на двухагентной архитектуре, которая разделяет сбор метрик и вывод модели. Он включает в себя цикл обратной связи, который обновляет модели обнаружения в режиме реального времени, позволяя адаптироваться к изменяющейся рабочей нагрузке и новым моделям атак.
Научная новизна
• Разработан метод интеграции метрик двух уровней (системного и прикладного) для мониторинга безопасности Kubernetes, обеспечивающий повышенную контекстную осведомленность об аномалиях за счет совместного анализа использования ресурсов узлом и специфического для приложений поведения.
• Предложен метод обнаружения аномалий для различных фреймворков, основанный на моделях машинного обучения с учителем, обученных на унифицированном наборе телеметрии, охватывающем разнородные среды (Flask, FastAPI, Django, Node.js, Golang), и проверенный в ходе экспериментов с рабочими нагрузками в реальном времени.
• Создан метод адаптивного обнаружения аномалий с интеграцией обратной связи, который включает цикл онлайн-обучения, динамически обновляющий модели обнаружения на основе операционных данных (истинно/ложно положительные/отрицательные срабатывания), что позволяет адаптироваться к новым угрозам.
• Все методы были оценены с использованием количественных метрик производительности (точность, F1 -мера, задержка обнаружения) и валидированы в практических развертываниях Kubernetes, что подтвердило высокую точность обнаружения (до 99,5%).
Теоретическая значимость
• Предложен унифицированный метод интеграции агент-ориентированной телеметрии и адаптивных моделей МО в реальном времени, который развивает исследования в области безопасности Kubernetes, объединяя статические и динамические подходы к обнаружению аномалий в cloud-native системах.
• Продемонстрирован метод, обеспечивающий способность к обобщению и масштабируемость моделей обнаружения, проверенный для множества прикладных
фреймворков и сценариев развертывания, что вносит теоретический вклад в развитие адаптивных систем кибербезопасности для распределенных систем.
Практическая значимость
• Разработан метод, пригодный для развертывания в реальном времени в промышленных кластерах с минимальным влиянием на производительность (показано использование ЦП <5% на узел и задержка телеметрии <150 мс), что делает его suitable для непрерывного наблюдения во время выполнения и обнаружения угроз.
• Создано комплексное решение для автоматизированного обнаружения аномалий и реагирования, интегрирующее дообучение моделей, генерацию оповещений и сбор пользовательской обратной связи, что позволяет сократить ручные операции безопасности и повысить эффективность управления инцидентами.
Методология и методы исследования
Для решения задач, поставленных в исследовании, использовались методы теории информационной безопасности, системного анализа, машинного обучения, статистической оценки и экспериментальной проверки. Апробация результатов
Результаты были представлены на 7 международных и отечественных конференциях по облачным вычислениям и кибербезопасности:
1. LI научная и учебно-методическая конференция университета ИТМО,
02.02.2022-05.02.2022
2. КМУ ИТМО, 04.04.2022-08.04.2022
3. LII научная и учебно-методическая конференция университета ИТМО,
31.01.2023-03.02.2023
4. XII конгресс молодых ученых ИТМО (КМУ 2023), 03.04.2023-06.04.2023
5. LIII научная и учебно-методическая конференция университета ИТМО, 29.01.2024
6. XIII конгресс молодых ученых ИТМО (КМУ 2024), 08.04.2024-11.04.2024
7. XIV конгресс молодых ученых ИТМО (КМУ 2025), 07.04.2025-11.04.2025
Автор принял участие в 4 технологических и научных конкурсах:
1. Kubernetes monitoring with Prometheus for security purposes - 15.06.2022
2. Overcoming challenges in dos attack detection for cloud-native Kubernetes environments - 27.02.2023
3. Kubernetes security essentials: navigating threats and protecting the cluster -15.01.2024
4. Best practices and security analysis for Kubernetes environment - 09.09.202409.12.2024
Публикации и научные достижения
1. По результатам диссертационного исследования автором опубликовано 6 научных работах, из них статей в журналах, рекомендованных ВАК - 3, индексируемых в базе данных «Scopus» - 3.
Внедрение результатов
Разработанные метода были внедрены в 5-узловом кластере Azure Kubernetes с разнородными рабочими нагрузками. Они были интегрирована с Prometheus, Grafana и пользовательской панелью управления. Смоделированные DoS-атаки и реальные рабочие нагрузки продемонстрировали эффективность и адаптивность методов.
Личный вклад автора
Автор получил результаты диссертации самостоятельно. Совместно с научным руководителем осуществлялась постановка цели и задач исследования, обсуждение плана исследования и полученных результатов, авторство и подача всех результатов исследования. Автором самостоятельно разработан агент мониторинга и конвейер подготовки данных, обучены и оценены все модели МО, спроектирована и протестирована двух-агентная архитектура.
Структура и объем диссертации
Диссертация состоит из введения, пяти глав, заключения, списка использованных источников. Общий объем работы 223 страницы, 25 рисунков, 16 таблиц. Список литературы содержит 147 наименований.
Введение обосновывает актуальность темы диссертационного исследования, проблему исследования, цели и важность исследования, focusing on обеспечение безопасности Kubernetes и решение проблем, связанных с DoS-атаками. В нем представлены цели исследования и ключевые задачи.
Глава 1 представляет собой обзор литературы и представляет комплексный анализ подходов к обеспечению безопасности Kubernetes. Данный раздел раскрывает меры безопасности в Kubernetes и то, как методы МО используются для обнаружения аномалий. Выделены существующие пробелы в исследованиях и даны пояснения, как данное исследование отвечает на существующие вызовы.
В последние годы Kubernetes стал доминирующей платформой для оркестрации контейнеризированных рабочих нагрузок, благодаря своей автоматизации, масштабируемости и гибкости. Однако, Kubernetes является уязвимым для развивающихся угроз безопасности, включая атаки типа DoS. Его распределенная архитектура, виртуальные компоненты и высокая интенсивность внутренней коммуникации создают широкую поверхность атаки, которую традиционные механизмы безопасности не могут адекватно покрыть.
Многие существующие решения для обнаружения аномалий посвящены исследованию использования МО для выявления инцидентов безопасности в облачных или сетевых средах. Тем не менее, большинство этих усилий не полностью учитывают уникальное поведение и динамику рабочих нагрузок Kubernetes. Значительное число этих подходов построены на статических наборах данных, используют офлайн-обучение и разработаны для традиционных, централизованных
инфраструктур. Их применение в быстро меняющихся, муультифреймворковых средах Kubemetes остается ограниченным.
Детальный анализ литературы выявил ряд ограничений, которые препятствуют эффективному применению рассмотренных решений при применении к реальным кластерам Kubemetes. К ним относятся отсутствие обработки в реальном времени, плохая адаптивность, узкий охват метрик и ограниченное обобщение для гетерогенных рабочих нагрузок. Таблица 1 представляет результаты сравнения существующих подходов, их ограничения и недостатки.
Таблица 1 - Ключевые ограничения существующих подходов для обнаружения DoS на основе МО в Kubemetes
Существующее решение_
Ограничения
Примеры
Пакетная обработка или периодическое сканирование
Зависимость от последующего (post-hoc) анализа предотвращает немедленное реагирование на угрозы; не может адаптироваться к динамическим паттернам рабочих нагрузок
[123] отмечают зависимость от статических наборов данных
Статические модели без механизмов обратной связи
Отсутствие возможностей непрерывного обучения; требуют ручного переобучения для новых шаблонов угроз_
[123, 124] показывают, что модели деградируют без онлайн-обучения
Ограничение конкретными типами метрик
Узкая область мониторинга создает слепые зоны обнаружения; пропускают кросс-уровневые корреляции_
[124] выделяют недостаточный охват телеметрии
Единый ML-подход на решение
Ограниченное методологическое разнообразие снижает robustness и адаптивность обнаружения
Недостаточная интеграция множественных парадигм обучения
Статические
политики
безопасности
Плохое обобщение между средами; требуют ручной настройки для новых развертываний_
[124] показывают требования к специфичной для среды настройке_
Реализации лабораторного масштаба Вычислительные накладные расходы и архитектурные узкие места в продукционных средах [73] сообщают о ресурсоемких операциях
Высокий уровень ошибок Усталость от оповещений и пропущенные обнаружения due to плохого баланса чувствительности-специфичности [123] отмечают подавляющие уровни ложных тревог
Эти ограничения указывают на необходимость разработки нового поколения решений, которые сочетают обработку телеметрии с низкой задержкой, адаптивное обучение и масштабируемое развертывание. Текущие исследования часто не учитывают многоуровневую сложность Kubernetes - где значимые аномалии могут проявляться только при анализе корреляций в использовании ЦП, поведении памяти, производительности приложений и сетевого трафика.
При наличии исследований по внедрению МО в конвейер безопасности Kubernetes, большинство из них не оценивало модели МО для различных фреймворков или рабочих нагрузок. В результате их производительность часто значительно снижается за пределами узких экспериментальных сред. Эта проблема особенно остро проявляется в системах обнаружения аномалий, которые не поддаются обобщению и требуют ручной перенастройки или повторной маркировки при развертывании в новых кластерах.
Данная диссертация addresses эти пробелы, предлагая комплексный methods framework обнаружения аномалий в реальном времени, который интегрирует адаптивные ML-модели с агентом сбора метрик, работающим с множеством фреймворков и уровней. В отличие от предыдущих работ, он подчеркивает обобщаемость, уточнение моделей на основе обратной связи и готовность к промышленному использованию (production-ready scalability) — bridging разрыв между исследовательскими прототипами и развертываемыми решениями безопасности в экосистемах Kubernetes.
В данной работе устраняются эти ограничения и недостатки, предлагая комплексные методы обнаружения аномалий в режиме реального времени, которые объединяют адаптивные модели МО с мультифреймворковым, многоуровневым средством сбора метрик. В отличие от предыдущих работ, в нем особое внимание уделяется обобщаемости, уточнению модели на основе обратной связи и масштабируемости, готовой к производству, что позволяет преодолеть разрыв между исследовательскими прототипами и развертываемыми решениями безопасности в экосистемах Kubernetes.
Глава 2. В данной главе представлен метод мониторинга безопасности в реальном времени в средах Kubernetes. Этот метод включает развертывание легковесного агента мониторинга на каждом узле Kubernetes. Агент активно собирает телеметрию в реальном времени как с системного уровня (например, использование ЦП, потребление памяти, сетевой трафик), так и с прикладного уровня (например, HTTP-запросы, открытые дескрипторы файлов). Собранные метрики затем предварительно обрабатываются, нормализуются и безопасно передаются для обнаружения. Агент разработан для работы с минимальными накладными расходами и легкой интеграции с Kubernetes API, обеспечивая непрерывную видимость контейнеризированных рабочих нагрузок.
На рисунке 1 показана пошаговая последовательность операций предлагаемого метода.
Метод включает этапы инициализации, загрузки безопасной конфигурации, сбора метрик в реальном времени с системного и прикладного уровней, предварительной обработки и нормализации данных, а также их безопасной передачи в модуль обнаружения. Процесс также включает проверки во время выполнения для обеспечения минимального влияния на производительность и целостности операций. Механизмы обработки ошибок и обратной связи позволяют агенту самостоятельно настраиваться и сохранять устойчивость в непредвиденных условиях.
Данный структурированный подход гарантирует, что агент остается легковесным, безопасным и адаптивным к динамическим средам Kubernetes.
В этой главе представлены архитектура и функциональность легковесного агента мониторинга, разработанного как ключевой компонент предложенного метода обнаружения аномалий. Агент отвечает за сбор детальной телеметрии как с системного, так и с прикладного уровней в средах Kubernetes. Учитывая распределенную и динамическую природу кластеров Kubernetes, создание эффективного, масштабируемого компонента мониторинга с низкими накладными расходами было необходимо для обеспечения обнаружения аномалий в реальном времени без ущерба для производительности.
Агент мониторинга работает непрерывно на каждом узле Kubernetes и тесно интегрируется с Kubernetes API для получения критически важных метрик времени выполнения. К ним относятся загрузка ЦП, использование памяти, операции ввода-вывода на диск, сетевой трафик и статистика HTTP-запросов для подов и контейнеров. Такая двухуровневая наблюдаемость - сбор данных как об инфраструктуре, так и о поведении приложений - является ключевым преимуществом предлагаемого подхода. Собранные данные служат основным входом для моделей машинного обучения, обеспечивая идентификацию потенциальных угроз на основе реального операционного поведения, а не заранее заданных сигнатур или статических правил.
Рисунок 1 - Блок-схема предлагаемого метода мониторинга безопасности в реальном времени с использованием легковесного агента в Kubernetes
Внутренняя структура агента является модульной и проиллюстрирована на Рисунке 2. Агент состоит из пяти основных компонентов: сборщик метрик (metrics collector), модуль предварительной обработки (preprocessor), модуль передачи (secure transmitter), контроллер (controller) и слой безопасности (security layer). Сборщик извлекает телеметрию в реальном времени, в то время как модуль предварительной обработки выполняет фильтрацию, нормализацию и агрегацию для обеспечения чистоты и готовности данных к анализу. Модуль передачи безопасно пересылает обработанные метрики в центральную систему обнаружения. Тем временем контроллер управляет политиками планирования и дискретизации агента, позволяя динамически настраивать свое поведение на основе текущих уровней рабочей нагрузки. Включение выделенного слоя безопасности обеспечивает конфиденциальность и целостность всех передаваемых данных, что критически важно в средах, чувствительных к безопасности.
® Data Transmitter ^ Data Preprocessor
: Encryption Compression Batching Normalization i Filtering Aggregation
Metrics Collector
System metrics Application Metrics
Configuration Manager ® Security Module
Transmission Protocols Collection Intervals Data Retention Authentication Data Integrity Compliance
В
Agent Controller
Performance Monitoring Initialization and Shutdown Error Handling
Рисунок 2 - Архитектура легковесного агента мониторинга Kubemetes Для оценки практичности и эффективности агента была проведена серия тестов производительности в смоделированной среде Kubemetes. Результаты обобщены в Таблице 2. Агент показал высокую производительность при различных условиях нагрузки, со средним использованием ЦП всего 1,5% и пиковым использованием памяти ниже 150 МБ. Задержка сбора данных оставалась в допустимых пределах, в среднем 150 миллисекунд. Что наиболее важно, операции агента оказали незначительное влияние на отзывчивость приложений, с вариацией времени ответа
сервиса менее 1%. Эти результаты подтверждают, что агент может работать в продукционных кластерах без внесения заметного снижения производительности или конкуренции за ресурсы.
Таблица 2 - Результаты оценки производительности разработанного метода
Метрика Значение
Средняя задержка сбора данных 150 мс
Пиковая задержка сбора данных 250 мс
Среднее использование ЦП 1,5% ^ЦП
Пиковое использование ЦП 3,0% ^ЦП
Среднее использование памяти 100 МБ
Пиковое использование памяти 150 МБ
Целостность данных 99,8% точность
Полнота данных 99,6%
Влияние на время ответа приложения <1% вариация
Масштабируемость Линейная
Робастность Высокая
Кроме того, конструкция агента подчеркивает расширяемость и переносимость, что делает его подходящим для развертывания в различных средах Kubernetes, включая облачные, локальные (on-premise) и периферийные (edge) среды. Его способность обрабатывать телеметрию в реальном времени обеспечивает своевременный ввод для конвейера обнаружения аномалий, а его модульная конструкция поддерживает будущие улучшения, такие как добавление новых типов метрик или интеграция с внешними системами мониторинга. В целом, агент мониторинга не только служит основой метода обнаружения, но и предлагает самостоятельный вклад в безопасные, наблюдаемые и адаптивные операции Kubernetes.
Глава демонстрирует, что предложенный агент достигает оптимального баланса между полнотой наблюдения и эффективностью. Его конструкция, ориентированная на работу в реальном времени с низкими накладными расходами, делает его практичным инструментом для реализации интеллектуальных систем безопасности на основе машинного обучения в экосистемах Kubernetes.
Глава 3. В данной главе представлен метод обнаружения аномалий для множества прикладных фреймворков в Kubernetes. Был разработан метод машинного обучения с учителем для обнаружения аномалий в Kubernetes, который обладает способностью к обобщению для различных прикладных фреймворков (Flask, Django, FastAPI, Node.js, Golang). Метод включает обучение ансамблевых классификаторов (Random Forest, XGBoost, LightGBM) на интегрированных метриках системного и прикладного уровня, собранных с реальных рабочих нагрузок и смоделированных DoS-атак. Производительность моделей оценивалась с использованием Fl-меры, точности, ROC-AUC и матриц ошибок, что подтвердило их эффективность в обнаружении аномалий для различных фреймворков.
Глава 3 посвящена применению моделей машинного обучения для обнаружения аномалий в средах Kubernetes, с фокусом на выявлении шаблонов поведения, связанных с атаками типа DoS. Модели обучались на данных телеметрии в реальном времени, собранных специально разработанным агентом мониторинга, представленным в предыдущей главе. Этот набор данных включал метрики как системного, так и прикладного уровня, обеспечивая комплексное представление об использовании ресурсов в нормальных условиях и в условиях атаки. Данные для обучения были явно размечены для состояний "атака" и "не атака" для проведения обучения с учителем.
Процесс отбора признаков был сфокусирован на метриках, наиболее релевантных к истощению ресурсов и аномалиям рабочих нагрузок. Как подробно описано в Таблице 3, итоговый набор признаков включал использование ЦП,
Похожие диссертационные работы по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Применение методов машинного обучения к данным геномики, транскриптомики и биомедицинской визуализации для решения медицинских задач /Application of Machine Learning to Genomic, Transcriptomic, and Imaging Data in Medical Problems2026 год, кандидат наук Сарачаков Александр Евгеньевич
Повышение эффективности детерминированных алгоритмов управления с использованием нейронных сетей /Enhancement of Deterministic Control Algorithms Using Neural Networks2025 год, кандидат наук Кафа Висам
Методы машинного обучения для сквозных систем автоматического распознавания речи2023 год, кандидат наук Лаптев Александр Алексеевич
Многозначная классификация и распознавание именованных сущностей на основе переноса обучения по зашумленным меткам для малоресурсных языков2023 год, кандидат наук Шахин Зейн
Многомодальная мягкая биометрия в условиях частичного перекрытия лица2024 год, кандидат наук Маркитантов Максим Викторович
Список литературы диссертационного исследования кандидат наук Дарвиш Гадир, 2025 год
Литература
1. Sultan S., Ahmad I., Dimitriou T. Container security: Issues, challenges, and the road ahead // IEEE Access. 2019. V. 7. P. 5297652996. https://doi.org/10.1109/ACCESS.2019.2911732
2. Shamim Md.S.I., Bhuiyan F.A., Rahman A. XI Commandments of kubernetes security: A systematization of knowledge related to kubernetes security practices // Proc. of the 2020 IEEE Secure Development (SecDev). 2020. P. 58-64. https://doi.org/10.1109/ SecDev45635.2020.00025
3. Darwesh G., Hammoud J., Vorobeva A.A. Security in kubernetes: best practices and security analysis // Вестник УРФО. Безопасность в информационной сфере. 2022. Т. 22. № 2. С. 63-69. https://doi. org/10.14529/SECUR220209
4. Mondal S.K., Pan R., Kabir H.M.D., Tian T., Dai H.N. Kubernetes in IT administration and serverless computing: An empirical study and research challenges // Journal of Supercomputing. 2022. V. 78. N 2. P. 2937-2987. https://doi.org/10.1007/s11227-021-03982-3
5. Shamim S.I. Mitigating security attacks in kubernetes manifests for security best practices violation // Proc. of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/ FSE). 2021. P. 1689-1690. https://doi.org/10.1145/3468264.3473495
6. Yu D., Jin Y., Zhang Y., Zheng X. A survey on security issues in services communication of Microservices-enabled fog applications // Concurrency and Computation: Practice and Experience. 2019. V. 31. N 22. P. e4436. https://doi.org/10.1002/CPE.4436
7. Lou J.-G., Fu Q., Yang S., Xu Y., Li J. Mining invariants from console logs for system problem detection // Proc. of the USENIX Annual Technical Conference. 2010. P. 1-14.
8. Lin C.H., Tien C.W., Pao H.K. Efficient and effective NIDS for cloud virtualization environment // Proc. of the 4th IEEE International Conference on Cloud Computing Technology and Science Proceedings. 2012. P. 249-254. https://doi.org/10.1109/ cloudcom.2012.6427583
9. Gomez M.E. Full Packet Capture Infrastructure Based on Docker Containers: Tech. rep. SANS Institute InfoSec Reading Room, 2016.
10. Tien C.-W., Huang T.-Y., Tien C.-W., Huang T.-C., Kuo S.-Y. KubAnomaly: Anomaly detection for the Docker orchestration platform with neural network approaches // Engineering Reports. 2019. V. 1. N 5. P. e12080. https://doi.org/10.1002/eng2.12080
11. Chang C.-C., Yang S.-R., Yeh E.-H., Lin P., Jeng J.-Y. A Kubernetes-based monitoring platform for dynamic cloud resource provisioning // Proc. of the GLOBECOM 2017 — 2017 IEEE Global Communications Conference. 2017. P. 1-6. https://doi.org/10.1109/ GLOCOM.2017.8254046
12. Shah J., Dubaria D. Building modern clouds: Using Docker, Kubernetes & Google Cloud Platform // Proc. of the 2019 IEEE 9th Annual Computing and Communication Workshop and Conference (CCWC). 2019. P. 0184-0189. https://doi.org/10.1109/ CCWC.2019.8666479
13. Song M., Zhang C., Haihong E. An auto scaling system for API Gateway based on Kubernetes // Proc. of the 2018 IEEE 9th International Conference on Software Engineering and Service Science (ICSESS). 2018. P. 109-112. https://doi.org/10.1109/ ICSESS.2018.8663784
14. Burns B., Grant B., Oppenheimer D., Brewer E., Wilkes J. Borg, Omega, and Kubernetes // Queue. 2016. V. 14. N 1. P. 70-93. https:// doi.org/10.1145/2898442.2898444
Authors
Ghadeer Darwesh — PhD Student, ITMO University, Saint Petersburg, 197101, Russian Federation, https://orcid.org/0000-0003-1116-9410, ghadeerdarwesh32@gmail.com
Jaafar Hammoud — PhD Student, ITMO University, Saint Petersburg, 197101, Russian Federation, sc 57222044000, https://orcid.org/0000-0002-2033-0838, hammoudgj@gmail.com
Alisa A. Vorobeva — PhD, Associate Professor, Saint Petersburg, 197101, Russian Federation, sc 57191359167, https://orcid.org/0000-0001-6691-6167, Alice_w@mail.ru
Авторы
Дарвиш Гадир — аспирант, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, https://orcid.org/0000-0003-1116-9410, ghadeerdarwesh32@gmail.com
Хаммуд Жаафар — аспирант, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, sc 57222044000, https://orcid.org/0000-0002-2033-0838, hammoudgj@gmail.com
Воробьева Алиса Андреевна — кандидат технических наук, доцент, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, sc 57191359167, https://orcid.org/0000-0001-6691-6167, Alice_w@mail.ru
Received 25.11.2022
Approved after reviewing 20.02.2023
Accepted 16.05.2023
Статья поступила в редакцию 25.11.2022 Одобрена после рецензирования 20.02.2023 Принята к печати 16.05.2023
Работа доступна по лицензии Creative Commons «Attribution-NonCommercial»
НАУЧНО-ТЕХНИЧЕСКИЙ ВЕСТНИК ИНФОРМАЦИОННЫХ ТЕХНОЛОГИЙ, МЕХАНИКИ И ОПТИКИ _
•__ноябрь-декабрь2024 Том 24 №6 http://ntv.ifmo.ru/ научно-технический вестник
I/ITMO SCIENTIFIC AND TECHNICAL JOURNAL OF INFORMATION TECHNOLOGIES, MECHANICS AND OPTICS ИНФОРМАЦИОННЫХ ТЕХНОЛОГИЙ. МЕХАНИКИ И ОПТИКИ
November-December 2024 Vol. 24 No 6 http://ntv.ifmo.ru/en/
ISSN 2226-1494 (print) ISSN 2500-0373 (online)
doi: 10.17586/2226-1494-2024-24-6-
Enhancing Kubernetes security with machine learning: а proactive approach to anomaly detection
Ghadeer Darwesh1, Jaafar Hammoud2, Alisa A. Vorobeva3®
1,2,3 ITMO University, Saint Petersburg, 197101, Russian Federation
1 ghadeerdarwesh32@gmail.com, https://orcid.org/0000-0003-1116-9410
2 hammoudgj@gmail.com, https://orcid.org/0000-0002-2033-0838
3 vorobeva@itmo.ru®, https://orcid.org/0000-0001-6691-6167
Abstract
Kubernetes has become a cornerstone of modern software development enabling scalable and efficient deployment of microservices. However, this scalability comes with significant security challenges, particularly in detecting specific attack types within dynamic and ephemeral environments. This study presents a focused application of Machine Learning (ML) techniques to enhance security in Kubernetes by detecting Denial of Service (DoS) attacks and differentiating between DoS attacks, resource overload caused by attacks, and natural resource overloads. We developed a custom monitoring agent that collects telemetry data from various sources, including real-world workloads, actual attack scenarios, simulated hacking attempts, and induced overloading on containers and pods, ensuring comprehensive coverage. The dataset comprising these diverse sources was meticulously labeled and preprocessed, including normalization and temporal analysis. We employed and evaluated various ML classifiers, with Random Forest and AdaBoost emerging as the top performers, achieving F1 macro scores of 0.9990 ± 0.0006 and 0.9990 ± 0.0003, respectively. The novelty of our approach lies in its ability to accurately distinguish between different types of resource overloads and provide robust detection of DoS attacks within Kubernetes environments. These models demonstrated a high degree of accuracy in detecting security incidents, significantly reducing false positives and false negatives. Our findings highlight the potential of ML models to provide a targeted, proactive security framework for Kubernetes, offering robust protection against specific attack vectors while maintaining system reliability. Keywords
Kubernetes security, microservices, machine learning, anomaly detection, containerization, cybersecurity, telemetry data, real-time threat detection
For citation: Darwesh G., Hammoud J., Vorobeva A.A. Enhancing Kubernetes security with machine learning: а proactive approach to anomaly detection. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2024, vol. 24, no. 6, pp. (in Russian). doi: 10.17586/2226-1494-2024-24-6-
УДК 004.056
Повышение безопасности Kubernetes с использованием машинного обучения: проактивный подход к обнаружению аномалий
Гадир Дарвиш1, Жаафар Хаммуд2, Алиса Андреевна Воробьева3®
1'2'3 Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация
1 ghadeerdarwesh32@gmail.com, https://orcid.org/0000-0003-1116-9410
2 hammoudgj@gmail.com, https://orcid.org/0000-0002-2033-0838
3 vorobeva@itmo.ru®, https://orcid.org/0000-0001-6691-6167
Аннотация
Введение. Kubernetes — ключевая платформа для масштабируемого и эффективного развертывания микросервисов. С увеличением масштабируемости возрастает сложность выявления и своевременного обнаружения специфических типов атак в динамичных средах Kubernetes. Метод. В работе предложен подход для повышения безопасности Kubernetes, позволяющий детектировать атаки типа «отказ в обслуживании» (Denial
© Darwesh G., Hammoud J., Vorobeva A.A., 2024
of Service, DoS), основанный на использовании методов машинного обучения. Подход базируется на данных, полученных от пользовательского агента мониторинга, осуществляющего сбор телеметрической информации из различных источников, включая реальные рабочие нагрузки, сценарии атак, имитацию взлома и перегрузку ресурсов в контейнерах и подах. Полученные данные размечаются и обрабатываются, включая нормализацию и временной анализ для создания полноценного набора данных. Основные результаты. В ходе экспериментов протестированы различные классификаторы машинного обучения. Наиболее высокие показатели качества получены с использованием алгоритмов Random Forest и AdaBoost, дающие макро F1-оценки 0,9990 ± 0,0006 и 0,9990 ± 0,0003 соответственно. Разработанный подход позволяет эффективно отличать перегрузки ресурсов, вызванные атаками от естественных перегрузок, и обеспечивает точное выявление DoS-атак. Предложенная модель машинного обучения демонстрирует высокую точность в обнаружении инцидентов безопасности, существенно снижая количество ложных срабатываний. Обсуждение. Полученные результаты показывают, что модели машинного обучения могут стать основой для создания проактивной системы безопасности Kubernetes, которая обеспечит надежную защиту от специфических векторов атак, сохраняя при этом стабильность системы. Полученные результаты могут быть полезны исследователям и специалистам в области кибербезопасности приложения Kubernetes. Ключевые слова
безопасность Kubernetes, микросервисы, машинное обучение, обнаружение аномалий, контейнеризация, кибербезопасность, телеметрические данные, обнаружение угроз в реальном времени
Ссылка для цитирования: Дарвиш Г., Хаммуд Ж., Воробьева А.А. Повышение безопасности Kubernetes с использованием машинного обучения: проактивный подход к обнаружению аномалий // Научно-технический вестник информационных технологий, механики и оптики. 2024. Т. 24, № 6. С. (на англ. яз.). doi: 10.17586/22261494-2024-24-6-
Introduction
The microservices architecture has emerged as a transformative approach in software development, enabling the decomposition of monolithic applications into smaller, independently deployable services. This architectural style promotes scalability, flexibility, and rapid deployment, addressing the demands of modern software applications [1]. By leveraging containerization technologies such as Docker and Kubernetes, microservices can be orchestrated efficiently, providing a robust framework for developing and maintaining complex, large-scale applications [2]. However, the distributed and dynamic nature of microservices introduces significant security challenges. Traditional security mechanisms often fail to adequately protect microservices due to their inability to adapt to the continuous integration and deployment cycles inherent in these environments.
Security in Kubernetes extends beyond traditional perimeter defenses. In a containerized environment, each pod represents a potential entry point for attackers. Vulnerabilities in container images, misconfigurations in Kubernetes manifests, or even compromised nodes can lead to catastrophic breaches if left unchecked [3, 4]. In this context, Machine Learning (ML) models can be trained to recognize patterns and anomalies in data, making them highly effective at detecting attacks that may otherwise go unnoticed by traditional Intrusion Detection Systems (IDS). The ability of ML models to learn and adapt makes them particularly suited for the dynamic and complex nature of cybersecurity. Both supervised and unsupervised ML algorithms are used in this domain. Supervised learning can classify whether network traffic is normal or potentially harmful, while unsupervised learning can identify previously unseen attack patterns [5-7].
However, the application of ML in cybersecurity is not without challenges. One of the main challenges is the quality and quantity of data required to train effective
ML models. Cybersecurity datasets need to be large and diverse to encompass the wide range of potential attacks, and they also need to be labeled accurately to train supervised learning algorithms. This often requires significant resources and expertise [8]. Another challenge is the interpretability of ML models. Many effective ML models, such as neural networks, are often described as "black boxes" because their internal workings are not easily interpretable by humans. This can make it difficult to understand why a particular prediction was made, which is often important in cybersecurity contexts [9, 10].
In this article, we explore the concept of using machine learning to enhance security in Kubernetes environments by detecting DoS attacks. Our approach not only identifies DoS attacks but also differentiates between resource overloads caused by attacks and natural fluctuations in resource usage, a critical distinction for reducing false positives. We collect telemetry data from multiple sources, including real-world workloads, actual attack scenarios, simulated hacking attempts, and induced overloads on containers and pods. By leveraging machine learning algorithms to analyze this diverse data, we aim to detect anomalous behaviors indicative of potential attacks. This proactive approach empowers organizations to identify and mitigate attacks in Kubernetes environments before they escalate into full-blown breaches.
Related Works
Recent advancements in container technology have spurred significant research into securing these environments. This section reviews notable works that have contributed to enhancing the security of containerized applications, particularly in detecting attacks such as DoS and differentiating between attack-induced resource overloads and natural system overloads.
Researchers in [11] introduced a real-time Host-based IDS for Linux containers. Their approach monitors system calls from the host kernel to detect anomalies in container
behavior. The system achieved a high detection rate of 100 % with a low false positive rate of around 2 %.
In [12], researchers developed an online anomaly detection system for Docker containers using an optimized isolation forest algorithm. By assigning weights to resource metrics and using weighted feature selection, their system improves detection accuracy. This approach effectively detects anomalies in both simulated and real cloud environments with minimal performance overhead, which is crucial for distinguishing between attack-related overloads and natural variations in resource usage.
Researchers in [13] proposed a probabilistic real-time IDS for Docker containers. Their IDS uses n-grams of system calls and probabilistic models like Maximum Likelihood Estimator and Simple Good Turing to detect malicious applications. The system achieved accuracy ranging from 87 % to 97 % on various datasets, demonstrating its potential in detecting a wide range of attacks, including DoS.
Researchers in [14] evaluated the performance of anomaly-based IDS at the container level for multi-tenant applications. They used the Bag of System Calls technique and a sliding window with eight machine learning algorithms. Decision Tree and Random Forest algorithms provided the best results, achieving an F-Measure of 99.8 %. The study also found that Decision Tree is faster and consumes less Central Processing Unit (CPU) and memory compared to Random Forest, which is advantageous for real-time attack detection in resource-constrained environments.
In [15], researchers presented an approach to evaluate intrusion detection effectiveness in container-based systems using attack injection. They used a TPC-C workload with a database engine running as a container and monitored its system calls. The approach was effective in different scenarios, consistently detecting most attacks with precision values showing more variance, which could be beneficial in scenarios involving varied types of attacks like DoS.
A study in [16] focused on container vulnerability exploit detection, evaluating static and dynamic detection schemes using 28 real-world vulnerability exploits. Static scanning detected only 3 out of 28 vulnerabilities, while dynamic anomaly detection schemes detected 22 exploits. This highlights the importance of dynamic detection methods in identifying attack-induced anomalies that static methods might miss.
Researchers in [17] introduced Compiler Description Language (CDL), a classified distributed learning framework for detecting security attacks in containerized applications. CDL integrates online application classification with anomaly detection to address the challenge of insufficient training data for dynamic, short-lived containers. CDL improved detection rates and reduced false positive rates significantly compared to traditional methods, making it a promising approach for distinguishing between different causes of resource overloads in Kubernetes environments.
In [18], researchers analyzed security attacks and detection techniques for Docker containers, highlighting the need for effective security measures due to the high efficiency and widespread use of Docker in development
and deployment. Their study proposed a detailed analysis of existing security mechanisms and attacks, presenting a detection framework that proved effective in experimental evaluations, particularly in identifying DoS attacks.
Lastly, in [19], researchers conducted a comprehensive analysis of Docker container attack and defense mechanisms. They identified significant gaps in existing defenses, particularly in handling dynamic attack landscapes. Their evaluation framework, using an extensive dataset of 51 real-world vulnerabilities, demonstrated that static scanning tools and dynamic anomaly detection approaches both have limitations, with high false positive rates and inadequate training data being major issues. This underscores the need for more robust detection mechanisms that can accurately differentiate between attack-induced and natural resource overloads.
Problem Statement: Challenges in Detecting Attacks in Kubernetes and Microservices Environments
Adopting Kubernetes for its agility and scalability also introduces significant security challenges. This section examines the key challenges in detecting attacks in these environments and why traditional security measures may be insufficient.
1. Complexity and Dynamism: Kubernetes environments are highly dynamic, with containers continuously being created, scaled, and terminated in response to varying workloads. In a microservices architecture, where applications consist of numerous independently deployable services each running in its own container, the complexity is further amplified [12, 19]. This dynamism increases the difficulty in distinguishing between normal operational behaviors and potential attacks, such as DoS attacks or resource overloads caused by malicious activities. Differentiating these from natural, benign overloads in the system is a critical challenge for effective attack detection.
2. Attack Surface Expansion: The proliferation of containers and microservices significantly expands the attack surface in Kubernetes environments. Attackers have numerous entry points to exploit, ranging from vulnerabilities in container images and misconfigurations in Kubernetes manifests to compromised nodes and insecure Application Programming Interfaces. Traditional security tools often focus on perimeter defenses and lack the visibility needed into containerized workloads, leaving Kubernetes environments vulnerable to sophisticated attacks, including DoS attacks that can be easily mistaken for legitimate resource spikes [18, 19].
3. Ephemeral Nature of Containers: Containers are ephemeral by design, meaning they are short-lived and can be easily replaced or terminated. Traditional security measures, which rely on static configurations and long-lived assets, struggle to adapt to the transient nature of containers. Consequently, security teams may find it difficult to maintain visibility into the security posture of Kubernetes workloads and respond effectively to attacks, particularly those involving intentional resource exhaustion that mimics natural system behavior [19].
4. Lack of Contextual Awareness: In Kubernetes environments, security events and telemetry data are generated at a high volume and velocity. Without proper context distinguishing between normal operational behavior and attack-induced anomalies becomes a daunting task. For example, an unexpected surge in resource usage might either be a result of a legitimate operational demand or a DoS attack. Traditional security measures often lack the intelligence to analyze telemetry data in real-time and identify subtle deviations from normal behavior that indicate potential attacks [9, 10]. This lack of contextual awareness increases the likelihood of false positives and false negatives, undermining the effectiveness of security measures in Kubernetes environments.
5. Scalability and Performance Impact: As Kubernetes clusters scale to accommodate growing workloads, traditional security measures may introduce scalability and performance overhead. Agent-based security solutions commonly used in traditional environments may struggle to keep up with the dynamic nature of Kubernetes deployments leading to increased resource consumption and performance degradation [6]. This challenge is particularly pronounced when trying to detect DoS attacks and differentiate them from legitimate high-load scenarios where the security system must operate efficiently without hindering overall system performance.
These challenges underscore the need for advanced, adaptive security solutions that can effectively detect and differentiate between various types of attacks, including DoS attacks, in Kubernetes and microservices environments. By leveraging machine learning techniques, this study aims to address these challenges, providing a robust framework for proactive attack detection and mitigation in these complex, dynamic systems.
Dataset Collection with Custom Monitoring Agent
In this section, we detail the dataset used for training and evaluating our machine learning models for attack detection in Kubernetes environments. The dataset is comprehensive, incorporating telemetry data collected from a variety of sources to ensure the models can effectively differentiate between DoS attacks, resource overloads caused by attacks, and natural system overloads.
Custom Monitoring Agent. To gather comprehensive telemetry data from Kubernetes nodes and applications, we developed a custom monitoring agent tailored to our specific requirements. This agent is designed to collect a diverse range of system-level and application-level metrics, providing valuable insights into resource utilization, network activity, and application behavior. The custom monitoring agent is deployed across all nodes in the Kubernetes cluster, ensuring thorough data collection and monitoring coverage.
Data Collection Sources. The dataset comprises telemetry data collected from multiple sources, including: — Real-world Workloads. Metrics were gathered from production Kubernetes clusters running under typical operational conditions. This real-world data provides
a baseline for normal system behavior and natural resource usage patterns.
— Actual Attack Scenarios. We intentionally induced DoS attacks and other resource-exhausting activities in controlled environments. This data helps in training the models to recognize the specific signatures and patterns associated with such attacks.
— Simulated Hacking Attempts. To further enrich the dataset, we simulated various attack vectors that could potentially target Kubernetes environments. These simulations were designed to mimic real-world hacking activities, including attempts to overload resources or exploit vulnerabilities.
— Induced Overloading on Containers and Pods. We also conducted experiments to deliberately overload containers and pods in a non-malicious manner. This data is crucial for teaching the models to distinguish between malicious overloads caused by attacks and benign overloads resulting from legitimate high-demand scenarios.
Data Collection from Kubernetes Nodes. Our
monitoring agent collects an extensive array of system-level metrics from each Kubernetes node at regular intervals. These metrics include CPU usage, disk I/O, network traffic, memory utilization, process activity, and TCP/UDP socket statistics. By capturing these metrics, we obtain a holistic view of the cluster health and performance, enabling the identification of anomalous behavior that may indicate potential attacks.
Data Collection from Applications. In addition to monitoring Kubernetes nodes, our custom agent gathers metrics directly from the target applications. These application-level metrics encompass CPU and memory usage, open file descriptors, and detailed HyperText Transfer Protocol (HTTP) request statistics. Monitoring application behavior in real-time allows us to detect deviations from normal operation such as unusual spikes in resource consumption. Specifically, we identify and analyze patterns in HTTP request handling by categorizing them into predefined threat models, which helps in distinguishing between benign anomalies and potential attacks. This approach enhances the precision of our detection mechanisms, ensuring that the identified anomalies are meaningful and actionable.
Timestamped Data Collection. All metrics collected by our monitoring agent are timestamped to facilitate temporal analysis of system behavior. Timestamped data enables the correlation of events across different components of the Kubernetes environment, aiding in the identification of patterns indicative of security incidents, including those that may evolve gradually or occur in bursts such as DoS attacks.
Dataset Integrity and Quality. Ensuring the integrity and quality of the collected dataset is paramount for the effectiveness of our machine learning models. We employ rigorous data validation techniques to address missing values, outliers, and data quality issues, ensuring that our models are trained on clean and reliable data. High-quality datasets enhance the accuracy and robustness of the machine learning models leading to more reliable detection of attacks and reducing the chances of false positives and negatives.
Experiments and Model Evaluation
The success of machine learning models in detecting attacks within Kubernetes environments relies heavily on the quality of the dataset, the preprocessing steps, and the evaluation methods used. In this section, we outline the experiments conducted to train and evaluate our models, focusing on their ability to detect DoS attacks, differentiate between attack-induced resource overloads, and natural resource spikes.
Data Preprocessing. The initial phase of our machine learning project involves comprehensive data processing. This crucial step includes loading the dataset, cleansing unnecessary columns, transforming time-related data into discrete components, and partitioning the dataset into training and testing subsets. Utilizing the pandas library, we leveraged its robust DataFrame structure to efficiently manipulate our structured data.
Firstly, we removed the 'id' column, which does not contribute to our analysis. Subsequently, the 'time' column was converted to a datetime format, from which 'hour', 'minute', and 'second' components were extracted using pandas date-time functionality. This step ensures that temporal data can be effectively utilized in model training, which is crucial for detecting patterns related to DoS attacks that may occur at specific times or under certain conditions.
Post data preprocessing, the dataset was split into features (X) and the target variable (y). We applied normalization to the features using the StandardScaler from scikit-learn, ensuring that each feature contributes equally to the model training. The dataset was then divided into training and testing sets to facilitate model evaluation on unseen data, enhancing the model generalizability [20].
Model Selection and Training. We employed a diverse array of classifiers from the scikit-learn library, including Logistic Regression, Decision Tree Classifier, Random Forest Classifier, Support Vector Classifier, K-Nearest Neighbors Classifier, Gradient Boosting Classifier, AdaBoost Classifier, and Extra Trees Classifier. Each classifier was trained using default parameters on the training set, learning the underlying patterns in the data to make accurate predictions.
Given the focus on distinguishing between DoS attacks, attack-induced resource overloads, and natural overloads, special attention was given to models that could capture complex patterns and interactions within the data. For instance, ensemble methods like Random Forest and AdaBoost were particularly effective in this regard due to their ability to handle a variety of data distributions and interactions.
Model Evaluation. Upon training, the models were evaluated on the testing set comparing the predictions against the actual values to calculate accuracy. However, to ensure a more robust performance estimate, we implemented cross-validation. StratifiedKFold cross-validation with five splits was employed, preserving the percentage of samples for each class in each fold. This method provides a comprehensive performance measure by averaging the results across multiple rounds of training and testing.
In addition to accuracy, we calculated the F1 macro scores for each model using cross-validation. The F1 macro scores accounts for both precision and recall, providing a balanced measure of model performance across all classes, including the detection of DoS attacks and differentiation between resource overloads (Fig. 1). Fig. 1 shows the F1 macro scores obtained for each machine learning classifier over multiple cross-validation folds. Each subplot demonstrates how a specific classifier performs consistently across different folds, providing insights into the robustness and generalizability of the models. The high consistency observed across folds indicates minimal variance in the performance, underlining the reliability of the classification models in detecting DoS attacks and distinguishing between resource overloads. The models were also evaluated based on their ability to minimize false positives and false negatives, which is crucial for maintaining the reliability of security systems in Kubernetes environments.
Novel Contributions. One of the novel aspects of our work is the ability of our machine learning models to distinguish between attack-induced and natural resource overloads. This is particularly important in Kubernetes environments where resource usage can fluctuate naturally due to varying workloads. By accurately identifying these differences, our models reduce the risk of false positives, thereby enhancing the operational efficiency of the security system. This ability to differentiate adds a layer of precision that is often lacking in traditional security approaches, making our solution particularly suited for dynamic, cloud-native environments.
The results from the cross-validation show that the models generally achieved high F1 macro scores, with the Random Forest Classifier and AdaBoost Classifier achieving top performance (0.9990 ± 0.0006 and 0.9990 ± ± 0.0003, respectively). These results indicate strong model performance across various classifiers with minimal variability across folds, highlighting the robustness of our approach.
Confusion Matrices. To further elucidate model performance, we generated confusion matrices for each classifier (Fig. 2). Fig. 2 presents the confusion matrices derived from the cross-validation predictions of each classifier. The rows represent the actual class labels ('attack' and 'no attack'), while the columns show the predicted class labels. Diagonal elements indicate the number of correctly classified instances for each class, and the off-diagonal elements represent misclassifications. These matrices provide a visual assessment of the types of errors made by each classifier, helping to identify whether a model is prone to mistaking natural resource spikes for DoS attacks or vice versa. This analysis allows for targeted refinement of the models to reduce both false positives and false negatives. These matrices offer detailed insights into how well each model discriminates between the classes, highlighting areas where the model excels and where it requires improvement. For example, the confusion matrices help identify whether a model is particularly prone to mistaking natural overloads for DoS attacks or vice versa, allowing us to refine the models further.
9JOOS OIOBJ^-Jj
<u
•B
л
M
■ S3
I
1)
■ S3 л
о л
E
=i
О
Ö о О
£
pq^i этих
pqiq этих
Conclusion
The dynamic and distributed nature of Kubernetes and microservices architecture necessitates advanced security measures that extend beyond traditional approaches. Our study has demonstrated that machine learning can significantly enhance security in Kubernetes environments by providing robust, real-time detection of attacks, including Denial of Service attacks. Our models are particularly effective in distinguishing between attack-induced resource overloads and natural fluctuations in resource usage, a critical capability for reducing false positives and maintaining operational efficiency.
By leveraging telemetry data collected from multiple sources, including real-world workloads, simulated hacking attempts, and induced overloading scenarios, we have developed a comprehensive dataset that enables our Machine Learning (ML) models to accurately detect and differentiate between various types of attacks. The integration of these models into production environments
offers a proactive and adaptive approach to security, well-suited to the continuous integration and deployment cycles of modern software applications.
This research highlights the importance of integrating ML-based security solutions to maintain the integrity and reliability of Kubernetes deployments. The rigorous data processing and robust model evaluation techniques, such as cross-validation, confusion matrices, and F1 macro scores, have ensured the development of highly accurate and reliable machine learning models.
Looking forward, future work should focus on enhancing the interpretability of ML models, making it easier for security teams to understand and act upon the predictions made. Additionally, expanding the dataset to include even more diverse and complex scenarios will further improve the models ability to detect and respond to sophisticated attacks. Through these efforts, the security of Kubernetes environments can be continually improved, ensuring that they remain resilient in the face of evolving cyber threats.
References
1. Nobre J., Pires E.J., Reis A. Anomaly detection in microservice-based systems. Applied Sciences, 2023, vol. 13, no. 13, pp. 7891. https:// doi.org/10.3390/app13137891
2. De Lauretis L. From monolithic architecture to microservices architecture. Proc. of the 2019 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW), 2019, pp. 9396. https://doi.org/10.1109/issrew.2019.00050
3. Darwesh G., Hammoud J., Vorobeva A.A. A novel approach to feature collection for anomaly detection in Kubernetes environment and agent for metrics collection from Kubernetes nodes. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2023, vol. 23, no. 3, pp. 538-546. https://doi.org/10.17586/2226-1494-2023-23-3-538-546
4. Ghadeer D., Jaafar H., Vorobeva A.A. Security in kubernetes: best practices and security analysis. Vestnik UrFO. Security in the Information Sphere, 2022, no. 2(44), pp. 63-69.
5. Jacob S., Qiao Y., Ye Y., Lee B. Anomalous distributed traffic: Detecting cyber security attacks amongst microservices using graph convolutional networks. Computers & Security, 2022, vol. 118, pp. 102728. https://doi.org/10.1016/j.cose.2022.102728
6. Peralta-Garcia E., Quevedo-Monsalbe J., Tuesta-Monteza V., Arcila-Diaz J. Detecting structured query language injections in web microservices using machine learning. Informatics, 2024, vol. 11, no. 2, pp. 15. https://doi.org/10.3390/informatics11020015
7. Vinayakumar R., Alazab M., Soman K.P., Poornachandran P., Al-Nemrat A., Venkatraman S. Deep learning approach for intelligent intrusion detection system. IEEE Access, 2019, vol. 7, pp. 4152541550. https://doi.org/10.1109/ACCESS.2019.2895334
8. Zhang L., Cushing R., de Laat C., Grosso P. A real-time intrusion detection system based on OC-SVM for containerized applications. Proc. of the 2021 IEEE 24th International Conference on Computational Science and Engineering (CSE), 2021, pp. 138-145. https://doi.org/10.1109/cse53436.2021.00029
9. Raj P., Vanga S., Chaudhary A. Cloud-Native Computing: How to Design, Develop, and Secure Microservices and Event-Driven Applications. John Wiley & Sons, 2022, 352 p.
10. Torkura K.A., Sukmana M.I.H., Meinel C. Integrating continuous security assessments in microservices and cloud native applications. Proc. of the 10th International Conference on Utility and Cloud Computing, (UCC'17), 2017, pp. 171-180. https://doi. org/10.1145/3147213.3147229
11. Abed A.S., Clancy C., Levy D.S. Intrusion detection system for applications using linux containers. Lecture Notes in Computer Science, 2015, vol. 9331, pp. 123-135. https://doi.org/10.1007/978-3-319-24858-5_8
12. Zou Z., Xie Y., Huang K., Xu G., Feng D., Long D. A docker container anomaly monitoring system based on optimized isolation
Литература
1. Nobre J., Pires E.J., Reis A. Anomaly detection in microservice-based systems // Applied Sciences. 2023. V. 13. N 13. P. 7891. https://doi. org/10.3390/app13137891
2. De Lauretis L. From monolithic architecture to microservices architecture // Proc. of the 2019 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW). 2019. P. 9396. https://doi.org/10.1109/issrew.2019.00050
3. Darwesh G., Hammoud J., Vorobeva A.A. A novel approach to feature collection for anomaly detection in Kubernetes environment and agent for metrics collection from Kubernetes nodes // Научно-технический вестник информационных технологий механики и оптики. 2023. Т. 23. № 3. С. 538-546. https://doi.org/10.17586/2226-1494-2023-23-3-538-546
4. Ghadeer D., Jaafar H., Vorobeva A.A. Security in kubernetes: best practices and security analysis // Вестник УрФО. Безопасность в информационной сфере. 2022. № 2(44). С. 63-69.
5. Jacob S., Qiao Y., Ye Y., Lee B. Anomalous distributed traffic: Detecting cyber security attacks amongst microservices using graph convolutional networks // Computers & Security. 2022. V. 118. P. 102728. https://doi.org/10.1016/j.cose.2022.102728
6. Peralta-Garcia E., Quevedo-Monsalbe J., Tuesta-Monteza V., Arcila-Diaz J. Detecting structured query language injections in web microservices using machine learning // Informatics. 2024. V. 11. N 2. P. 15. https://doi.org/10.3390/informatics11020015
7. Vinayakumar R., Alazab M., Soman K.P., Poornachandran P., Al-Nemrat A., Venkatraman S. Deep learning approach for intelligent intrusion detection system // IEEE Access. 2019. V. 7. P. 4152541550. https://doi.org/10.1109/ACCESS.2019.2895334
8. Zhang L., Cushing R., de Laat C., Grosso P. A real-time intrusion detection system based on OC-SVM for containerized applications // Proc. of the 2021 IEEE 24th International Conference on Computational Science and Engineering (CSE). 2021. P. 138-145. https://doi.org/10.1109/cse53436.2021.00029
9. Raj P., Vanga S., Chaudhary A. Cloud-Native Computing: How to Design, Develop, and Secure Microservices and Event-Driven Applications. John Wiley & Sons, 2022. 352 p.
10. Torkura K.A., Sukmana M.I.H., Meinel C. Integrating continuous security assessments in microservices and cloud native applications // Proc. of the 10th International Conference on Utility and Cloud Computing, (UCC'17). 2017. P. 171-180. https://doi. org/10.1145/3147213.3147229
11. Abed A.S., Clancy C., Levy D.S. Intrusion detection system for applications using linux containers // Lecture Notes in Computer Science. 2015. V. 9331. P. 123-135. https://doi.org/10.1007/978-3-319-24858-5_8
12. Zou Z., Xie Y., Huang K., Xu G., Feng D., Long D. A docker container anomaly monitoring system based on optimized isolation
forest. IEEE Transactions on Cloud Computing, 2022, vol. 10, no. 1, pp. 134-145. https://doi.org/10.1109/tcc.2019.2935724
13. Srinivasan S., Kumar A., Mahajan M., Sitaram D., Gupta S. Probabilistic real-time intrusion detection system for docker containers. Communications in Computer and Information Science, 2019, vol. 969, pp. 336-347. https://doi.org/10.1007/978-981-13-5826-5_26
14. Cavalcanti M., Inacio P., Freire M. Performance evaluation of container-level anomaly-based intrusion detection systems for multi-tenant applications using machine learning algorithms. Proc. of the 16th International Conference on Availability, Reliability and Security (ARES'21), 2021, pp. 1-9. https://doi.org/10.1145/3465481.3470066
15. Flora J., Gonçalves P., Antunes N. Using attack injection to evaluate intrusion detection effectiveness in container-based systems. Proc. of the IEEE 25th Pacific Rim International Symposium on Dependable Computing (PRDC), 2020, pp. 60-69. https://doi.org/10.1109/ prdc50213.2020.00017
16. Tunde-Onadele O., He J., Dai T., Gu X. A study on container vulnerability exploit detection. Proc. of the IEEE International Conference on Cloud Engineering (IC2E), 2019, pp. 121-127. https:// doi.org/10.1109/ic2e.2019.00026
17. Lin Y., Tunde-Onadele O., Gu X. CDL: Classified distributed learning for detecting security attacks in containerized applications. Proc. of the 36th Annual Computer Security Applications Conference (ACSAC ' 20), 2020, pp. 179-1 88. https://doi. org/10.1145/3427228.3427236
18. Huang L., Ma D., Li S., Zhang X., Wang H. Text level graph neural network for text classification. Proc. of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3444-3450. https://doi.org/10.18653/ v1/d19-1345
19. Haq M.S., Nguyen T.D., Tosun A.S., Vollmer F., Korkmaz T., Sadeghi A.-R. SoK: A comprehensive analysis and evaluation of docker container attack and defense mechanisms. Proc. of the IEEE Symposium on Security and Privacy (SP), 2024, pp. 4573-4590. https://doi.org/10.1109/sp54263.2024.00268
20. Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., Vanderplas J., Passos A., Cournapeau D., Brucher M., Perrot M., Duchesnay É. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 2011, vol. 12, pp. 2825-2830.
forest // IEEE Transactions on Cloud Computing. 2022. V. 10. N 1. P. 134-145. https://doi.org/10.1109/tcc.2019.2935724
13. Srinivasan S., Kumar A., Mahajan M., Sitaram D., Gupta S. Probabilistic real-time intrusion detection system for docker containers // Communications in Computer and Information Science. 2019. V. 969. P. 336-347. https://doi.org/10.1007/978-981-13-5826-5_26
14. Cavalcanti M., Inacio P., Freire M. Performance evaluation of container-level anomaly-based intrusion detection systems for multi-tenant applications using machine learning algorithms // Proc. of the 16th International Conference on Availability, Reliability and Security (ARES'21). 2021. P. 1-9. https://doi.org/10.1145/3465481.3470066
15. Flora J., Gonjalves P., Antunes N. Using attack injection to evaluate intrusion detection effectiveness in container-based systems // Proc. of the IEEE 25th Pacific Rim International Symposium on Dependable Computing (PRDC). 2020. P. 60-69. https://doi.org/10.1109/ prdc50213.2020.00017
16. Tunde-Onadele O., He J., Dai T., Gu X. A study on container vulnerability exploit detection // Proc. of the IEEE International Conference on Cloud Engineering (IC2E). 2019. P. 121-127. https:// doi.org/10.1109/ic2e.2019.00026
17. Lin Y., Tunde-Onadele O., Gu X. CDL: Classified distributed learning for detecting security attacks in containerized applications // Proc. of the 36th Annual Computer Security Applications Conference (ACS AC'20). 2020. P. 1 79-1 88. https://doi. org/10.1145/3427228.3427236
18. Huang L., Ma D., Li S., Zhang X., Wang H. Text level graph neural network for text classification // Proc. of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. P. 3444-3450. https://doi.org/10.18653/v1/ d19-1345
19. Haq M.S., Nguyen T.D., Tosun A.S., Vollmer F., Korkmaz T., Sadeghi A.-R. SoK: A comprehensive analysis and evaluation of docker container attack and defense mechanisms // Proc. of the IEEE Symposium on Security and Privacy (SP). 2024. P. 4573-4590. https://doi.org/10.1109/sp54263.2024.00268
20. Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., Blondel M., Prettenhofer P., Weiss R., Dubourg V., Vanderplas J., Passos A., Cournapeau D., Brucher M., Perrot M., Duchesnay É. Scikit-learn: Machine learning in Python // Journal of Machine Learning Research. 2011. V. 12. P. 2825-2830.
Authors
Ghadeer Darwesh — PhD Student, ITMO University, Saint Petersburg, 197101, Russian Federation, sc 57226287648, https://orcid.org/0000-0003-1116-9410, ghadeerdarwesh32@gmail.com
Jaafar Hammond — PhD Student, ITMO University, Saint Petersburg, 197101, Russian Federation, sc 57222044000, https://orcid.org/0000-0002-2033-0838, hammoudgj@gmail.com
Alisa A. Vorobeva — PhD, Associate Professor, ITMO University, Saint Petersburg, 197101, Russian Federation, sc 57191359167, https://orcid. org/0000-0001-6691-6167, vorobeva@itmo.ru
Авторы
Дарвиш Гадир — аспирант, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, sc 57226287648, https://orcid.org/0000-0003-1116-9410, ghadeerdarwesh32@gmail.com
Хаммуд Жаафар — аспирант, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, sc 57222044000, https://orcid.org/0000-0002-2033-0838, hammoudgj@gmail.com
Воробьева Алиса Андреевна — кандидат технических наук, доцент, Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация, sc 57191359167, https://orcid.org/0000-0001-6691-6167, vorobeva@itmo.ru
Received 20.06.2024
Approved after reviewing 08.10.2024
Accepted 15.11.2024
Статья поступила в редакцию 20.06.2024 Одобрена после рецензирования 08.10.2024 Принята к печати 15.11.2024
Работа доступна по лицензии Creative Commons «Attribution-NonCommercial»
НАУЧНО-ТЕХНИЧЕСКИЙ ВЕСТНИК ИНФОРМАЦИОННЫХ ТЕХНОЛОГИЙ, МЕХАНИКИ И ОПТИКИ _
•__сентябрь-октябрь2025 Том 25 №5 http://ntv.ifmo.ru/ научно-технический вестник
l/ITMO SCIENTIFIC AND TECHNICAL JOURNAL OF INFORMATION TECHNOLOGIES, MECHANICS AND OPTICS ИНФОРМАЦИОННЫХ ТЕХНОЛОГИЙ. МЕХАНИКИ И ОПТИКИ
September-October 2025 Vol. 25 No 5 http://ntv.itmo.ru/en/
ISSN 2226-1494 (print) ISSN 2500-0373 (online)
doi: 10.17586/2226-1494-2025-25-5-910-922
Enhanced detection of denial-of-service attacks in Kubernetes: a multi-framework machine learning approach integrating node and application metrics
Ghadeer Darwesh1, Jaafar Hammoud2, Alisa A. Vorobeva3®
1,2,3 ITMO University, Saint Petersburg, 197101, Russian Federation
1 ghadeerdarwesh32@gmail.com, https://orcid.org/0000-0003-1116-9410
2 hammoudgj@gmail.com, https://orcid.org/0000-0002-2033-0838
3 vorobeva@itmo.ru®, https://orcid.org/0000-0001-6691-6167
Abstract
The widespread adoption of Kubernetes as a platform for orchestrating containerized applications has heightened the need for effective security mechanisms, particularly to counter Denial-of-Service (DoS) attacks. This article proposes an approach to DoS attack detection based on two key components the use of comprehensive metrics and the application of ensemble Machine Learning models. The approach involves the collection and analysis of comprehensive metrics from node-level (CPU, memory) and application-level (network activity, file descriptors) data from containers running on various frameworks (Flask, Django, FastAPI, Node.js, Golang). To implement this approach, a dataset containing 49,990 instances of network activity, characterized by 28 features (comprehensive metrics), was created. Statistical analysis (Student's t-test, Pearson correlation) identified the metrics most relevant for attack detection, including total CPU time (cpu_sec_total) and resident memory usage (resident_memory_total). A comparison of nine Machine Learning models for attack detection was conducted, including ensemble methods (Random Forest, XGBoost, LightGBM) which demonstrated the highest effectiveness, achieving 100 % accuracy (F1-score equals 1.0) and perfect class separation (AUC equals 1.0). The XGBoost model also eliminated false positives (precision equals 1.0). Feature importance analysis revealed the most significant metrics for classification: CPU usage (cpu_sec_total, cpu_sec_idle), network packet transmission (transmit_packets), system load average, and memory usage (virtual_memory_total, resident_ memory_total). The work emphasizes the importance of integrating multi-level metrics for building resilient anomaly detection systems. The proposed approach is scalable and independent of specific frameworks, making it applicable for protecting containerized environments. The research results serve as a foundation for developing proactive Kubernetes security systems capable of countering sophisticated attack vectors. Keywords
Kubernetes, DoS attack detection, machine learning, node-level metrics, application-level metrics, anomaly detection, ensemble models
For citation: Darwesh G., Hammoud J., Vorobeva A.A. Enhanced detection of denial-of-service attacks in Kubernetes: a multi-framework machine learning approach integrating node and application metric. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2025, vol. 25, no. 5, pp. 910-922. doi: 10.17586/2226-1494-2025-25-5-910-922
УДК 004.056
Повышение эффективности обнаружения DoS-атак в Kubernetes: подход на основе машинного обучения с интеграцией метрик уровня узлов и приложений для мультифреймворковых сред
Гадир Дарвиш1, Жаафар Хаммуд2, Алиса Андреевна Воробьева3®
1'2'3 Университет ИТМО, Санкт-Петербург, 197101, Российская Федерация
1 ghadeerdarwesh32@gmail.com, https://orcid.org/0000-0003-1116-9410
2 hammoudgj@gmail.com, https://orcid.org/0000-0002-2033-0838
3 vorobeva@itmo.ru®, https://orcid.org/0000-0001-6691-6167
© Darwesh G., Hammoud J., Vorobeva A.A., 2025
Аннотация
Широкое распространение Kubemetes как платформы для оркестрации контейнеризированных приложений атакам типа «Отказ в обслуживании» (Denial-of-Service, DoS). В работе предложен подход к обнаружению DoS-атак, основанный на двух ключевых компонентах: использование комплексных метрик и применение ансамблевых моделей машинного обучения. Подход предполагает сбор и анализ комплексных метрик: уровня узлов (Central Processing Unit (CPU), память) и уровня приложений (сетевая активность, файловые дескрипторы) из контейнеров, работающих на различных фреймворках (Flask, Django, FastAPI, Node.js, Golang). Для реализации подхода создан набор данных, содержащий 49 990 экземпляров сетевой активности, охарактеризованных 28 признаками (комплексными метриками). Статистический анализ (t-критерий Стьюдента, корреляция Пирсона) выявил наиболее релевантные для детектирования атак метрики, включая общее время использования CPU (cpu_sec_total) и объем задействованной оперативной памяти (resident_memory_total). Сравнение девяти моделей машинного обучения для детектирования атак, включая ансамблевые методы (Random Forest, XGBoost, LightGBM), показало наивысшую эффективность (Fl-мера равна 1,0) и полное разделение классов (AUC равна 1,0). Применение модели XGBoost позволило исключить ложноположительные срабатывания (precision равна 1.0). Анализ важности признаков выявил наиболее значимые для классификации метрики, связанные с использованием CPU (cpu_sec_total, cpu_sec_idle), передачей сетевых пакетов (transmit_ packets), средней загрузкой системы и использованием памяти (virtual_memory_total, resident_memory_total). Проведенное исследование показало важность интеграции разноуровневых метрик для создания устойчивых систем обнаружения аномалий. Предложенный подход является масштабируемым и независимым от конкретных фреймворков, что делает его применимым для защиты контейнеризированных сред. Результаты исследования служат основой для разработки проактивных систем безопасности Kubernetes, способных противостоять сложным векторам атак. Ключевые слова
Kubernetes, обнаружение DoS-атак, машинное обучение, метрики уровня узлов, метрики уровня приложений, обнаружение аномалий, ансамблевые модели
Ссылка для цитирования: Дарвиш Г., Хаммуд Ж., Воробьева А.А. Повышение эффективности обнаружения DoS-атак в Kubernetes: подход на основе машинного обучения с интеграцией метрик уровня узлов и приложений для мультифреймворковых сред // Научно-технический вестник информационных технологий, механики и оптики. 2025. Т. 25, № 5. С. 910-922 (на англ. яз.). doi: 10.17586/2226-1494-2025-25-5-910-922
Introduction
Kubernetes, the de facto standard for container orchestration, has revolutionized cloud-native architectures by automating the deployment, scaling, and management of containerized applications. However, the complexity and dynamic nature of Kubernetes clusters expose them to a wide range of security threats, with Denial-of-Service (DoS) attacks standing out as a prominent risk. These attacks exploit Kubernetes resource scaling mechanisms to inundate cluster resources, potentially leading to service disruptions and substantial financial repercussions [1, 2]. The containerized workloads managed by Kubernetes pose unique challenges in detecting and mitigating DoS attacks. Containers often demonstrate dynamic, ephemeral, and unpredictable resource utilization patterns which blur the line between legitimate traffic bursts and malicious activity. Conventional DoS detection methods, such as static thresholds and signature-based techniques, fall short in such settings due to their inability to adapt to Kubernetes highly dynamic operational environments. These limitations often result in high false-positive rates, rendering traditional approaches unsuitable for production environments [1, 2]. Machine Learning (ML) has emerged as a promising paradigm for tackling these challenges. By leveraging anomaly detection techniques, ML-based systems can dynamically identify deviations in traffic and resource utilization patterns without relying on predefined rules or static thresholds. However, many existing studies focus on specific application frameworks, such as Flask, or narrow workloads, thereby limiting their generalizability across diverse operational contexts [1, 2]. Recent advancements in Kubernetes security emphasize
the importance of hybrid approaches that combine ML with runtime monitoring and rule-based mechanisms. Tools such as extended Berkeley Packet Filter (eBPF) and Express Data Path (XDP) offer lightweight, high-performance anomaly detection at the kernel level, enabling real-time security insights. Despite these innovations, significant gaps remain in evaluating the effectiveness of such approaches across diverse frameworks and workloads, leaving room for improvement in generalizability and robustness [3, 4]. Building upon our prior work [5], where we developed an ML-based DoS detection framework tailored to the Flask framework, this study extends the scope to encompass multiple application frameworks, including Django, FastAPI, Flask, Golang, and Node.js. This broader scope addresses the generalizability concerns raised in our earlier work and ensures applicability to a wider range of Kubernetes environments.
This study makes three key contributions: it provides a comparative analysis of ML-based DoS detection across multiple frameworks to improve generalizability and robustness; it integrates lightweight runtime monitoring tools to enhance detection efficiency; and it evaluates advanced classifiers for distinguishing between natural workload variations and attack-induced anomalies in Kubernetes environments [5-7].
Literature Review and Previous Work
DoS attacks are among the most prevalent security threats in containerized environments. Kubernetes, with its dynamic orchestration and auto-scaling capabilities, provides an efficient platform for managing modern workloads but also presents a large attack surface for
adversaries to exploit. DoS attacks in Kubernetes often target resource management mechanisms, overwhelming cluster components like Central Processing Unit (CPU), memory, and network bandwidth to render applications unresponsive [1, 2]. Studies emphasize the difficulty in distinguishing between natural workload variations and attack-induced overloads, as both may manifest as anomalies in resource usage [1, 4].
Conventional DoS detection techniques rely on predefined thresholds or static rules to identify anomalous behaviors. While computationally inexpensive, these methods are rigid and fail to adapt to dynamic environments like Kubernetes [1, 2, 4].
Recent advancements have introduced more sophisticated approaches, such as Host-Based Intrusion Detection Systems (HIDS). Researchers in [8] developed a real-time HIDS for Linux containers that monitors system calls from the host kernel to detect anomalies. Their method achieved a 100 % detection rate with a false positive rate of just 2 %.
Several studies focus on Docker containers, the most widely used container runtime. For example: Researchers in [9] proposed an online anomaly detection system using an optimized isolation forest algorithm. By assigning weights to resource metrics and incorporating weighted feature selection, this approach improved accuracy while maintaining minimal performance overhead, crucial for differentiating between attack-induced and natural resource overloads. In [10], a probabilistic real-time IDS was proposed using n-grams of system calls and probabilistic models like Maximum Likelihood Estimator. This system achieved detection accuracy between 87 % and 97 % across datasets. Dynamic approaches using anomaly-based methods have demonstrated significant improvements over static techniques. In [11], researchers evaluated dynamic schemes on 28 real-world container vulnerability exploits, with results indicating that dynamic methods detected 22 out of 28 exploits compared to only three detected by static methods.
ML-based anomaly detection systems offer significant advantages in identifying DoS attacks. Techniques like Random Forest, Gradient Boosting, and Neural Networks have been widely adopted for detecting anomalies in Kubernetes. The introduction of eBPF and XDP has enabled lightweight, kernel-level anomaly detection for real-time insights [1, 2]. Studies such as [12] have demonstrated the potential of ML classifiers like Decision Trees and Random Forests in distinguishing between legitimate and malicious activities at the container level. These approaches achieved F-measures of 99.8 % for attack detection while maintaining low resource overheads.
Most existing research evaluates DoS detection methods on specific frameworks, such as Flask or FastAPI, without accounting for their generalizability to other workloads [1, 2, 4]. Researchers in [13] highlighted the need for cross-framework evaluations by analyzing security mechanisms in Docker containers across multiple deployment scenarios. Their results emphasized that framework-agnostic detection systems are critical for robust Kubernetes security.
Studies are often constrained to single frameworks, making their findings less applicable to diverse Kubernetes workloads [1, 2, 4].
Despite these advancements, the current research landscape reveals several persistent and interconnected limitations that hinder the development of robust, practical detection systems. A primary constraint is the lack of generalizability across technological stacks. The majority of studies evaluate proposed methods within the context of a single application framework, neglecting validation across diverse environments [1, 2, 4, 13]. This significantly limits the applicability of such solutions in real-world Kubernetes clusters which are inherently heterogeneous and host applications built with different languages and frameworks.
Furthermore, a significant efficiency-effectiveness trade-off remains unresolved. While static methods are computationally efficient yet inflexible, the more effective dynamic and ML approaches typically demand large volumes of training data and introduce substantial computational overhead, making them costly to deploy in resource-sensitive environments [10, 14].
Finally, the critical challenge of integration is still largely unaddressed. Only a limited number of studies have explored the combination of real-time monitoring techniques (e.g., eBPF) with ML systems to achieve the dual objectives of high accuracy and low-latency detection in Kubernetes production settings [1].
Thus, the identified gaps — limited generalizability, an unoptimized accuracy-overhead trade-off, and a lack of integrated real-time solutions — form the core research challenge addressed by this work.
This study directly addresses these limitations by proposing a comprehensive detection framework validated across multiple application frameworks and programming languages, including Flask, Django, FastAPI, Golang, and Node.js. Our approach leverages lightweight, eBPF-based runtime monitoring to minimize performance impact and detection latency. We conduct an extensive comparative evaluation of ML classifiers to identify optimal strategies for DoS detection in Kubernetes. Through these contributions, this research provides a foundation for developing scalable, framework-agnostic security solutions capable of protecting complex Kubernetes deployments against evolving DoS threats.
Data and Statistical Study
These contributions aim to provide a framework-agnostic, efficient, and accurate solution to securing Kubernetes environments against evolving threats. The dataset1 consists of 49,990 instances with 28 features, encompassing both node-level and app-level metrics collected from multiple frameworks deployed in Kubernetes by using a collector developed in [15]. The target variable, attack, is binary, indicating the presence (1) or absence (0) of DoS attacks. Table 1 presents the node-level metrics gathered from the frameworks, while Table 2 details the app-level metrics.
We conducted a comprehensive statistical analysis to demonstrate the robustness and reliability of the dataset. This analysis provides insights into the distribution, central
1 Available at: https://github.com/ghadeerda/Kubernetes-model-agent (accessed: 29.08.2025).
Table 1. This table lists and describes the metrics collected at the node level, including CPU, memory, disk, and network-related
features
Column Name Description
id Unique identifier for the observation
time Timestamp of the data record
cpu_sec_idle Percentage of CPU idle time during the interval
disk_av_per Available disk space as a percentage
disk_read Amount of data read from the disk (in bytes)
disk_write Amount of data written to the disk (in bytes)
net_receive Network data received (in bytes)
mem_pressure Memory pressure indicator (value reflects memory load)
mem_av_per Available memory as a percentage
forks_total Total number of process forks
intr Number of interrupts handled by the CPU
loadl, load5, load15 CPU load average over 1, 5, and 15 minutes respectively
receive_drop Number of network packets dropped during reception
receive_errs Number of network reception errors
transmit_packets Number of packets transmitted over the network
ipv4_sock_inuse Number of IPv4 sockets currently in use
est_conn Number of established connections
lis_conn Number of listening connections
open_fds Number of open file descriptors
attack Indicator for attack presence (1 for attack, 0 for no attack)
Table 2. This table provides details of the application-level metrics, including resource utilization features specific to the application
frameworks
Column Name Description
id Unique identifier for the observation
time Timestamp of the data record
cpu_sec_total Total CPU usage in seconds
virtual_memory_total Total virtual memory usage (in bytes)
resident_memory_total Total resident memory usage (in bytes)
open_fds Number of open file descriptors
attack Indicator for attack presence (1 for attack, 0 for no attack)
tendencies, and variability of both node-level and app-level metrics, highlighting the dataset ability to capture diverse system behaviors under different conditions. By examining the relationships between features and their correlations with the target variable (attack), we confirmed that the dataset effectively encapsulates the characteristics required for detecting DoS attacks. This statistical study not only validates the dataset integrity but also underscores its potential for developing and benchmarking robust anomaly detection models in containerized environments.
To thoroughly understand the dataset characteristics and assess its suitability for detecting DoS attacks, we conducted an in-depth statistical analysis. This analysis aimed to explore key metrics at both node-level and app-level, evaluate their relationships, and identify patterns that distinguish between attack and non-attack states. By
combining descriptive statistics, hypothesis testing, and correlation analysis, we established a solid foundation for understanding the dataset structure and its potential for training robust ML models. The following sections detail the findings from these analyses.
Descriptive Statistics
During attacks, CPU usage increases significantly while idle time decreases, indicating elevated system load. Virtual and resident memory also spike, reflecting stress on memory resources. Open file descriptors rise, suggesting heavier application activity. Disk and network I/O metrics show moderate changes but still contribute to overall anomaly detection.
Hypothesis Testing
T-tests revealed statistically significant differences between attack and non-attack states. Features such as
cpu_sec_total, resident_memory_total, and open_fds were highly significant, while disk_read, disk_write, and virtual_memory_total showed moderate significance. These findings confirm that attacks consistently alter resource usage patterns.
Correlation Analysis
Correlation analysis showed that cpu_sec_total, resident_memory_total, and open_fds are positively correlated with attacks, while cpu_sec_idle and disk_ av_per are negatively correlated. Some memory-related features showed redundancy, while others like forks_total and disk_av_per contributed unique information.
Key Insights
The most predictive indicators of DoS attacks include increased CPU and memory usage, along with higher counts of open file descriptors. These resource usage
patterns correspond to the system stress and resource exhaustion typically induced by attacks. Fig. 1 presents histograms comparing the distributions of key. Fig. 2 shows boxplots of metrics. These visualizations were generated by the authors based on the labeled dataset of 49,990 instances, which includes node- and application-level metrics collected from real Kubernetes workloads using five frameworks (Flask, Django, FastAPI, Node.js, and Golang). The patterns observed in these figures are derived from the statistical analysis discussed earlier (t-tests and correlation), and form the empirical basis for the ML models used in this study.
Sample Representativeness and Realism of Workload Simulation
To ensure a realistic and representative dataset, we designed the data collection process to emulate real
8000
6000
^ 4000
2000
Attack —No Attack Attack
24 cpu_sec_total
6000
cy4000
2000
0.5 1.0
cpu_sec_idle
b
a
0
0
0
d
c
e
open_fds
Fig. 1. Histograms of Frequency Distribution of CPU and memory usage metrics (cpu_sec_total, resident_memory_total) under attack and non-attack conditions, based on the collected dataset: cpu_usage_total (a); cpu_sec_idle (b); resident_memory_total (c);
virtual_memory_total (d); open_fds (e)
O l O l
Attack (O: No Attack, l: Attack) Attack (O: No Attack, l: Attack)
e
5O
1П
Ol Attack (O: No Attack, l: Attack)
Fig. 2. Distribution shifts for key metrics. Feature Boxplots comparing CPU idle time (cpu_sec_idle) and open file descriptors (open_fds) between normal and attack states in Kubernetes workloads: cpu_usage_total (a); cpu_sec_idle (b); resident_memory_total (c); virtual_memory_total (d); open_fds (e)
Kubernetes environments. The dataset includes both node-and application-level metrics from applications built with five frameworks — Flask, Django, FastAPI, Golang, and Node.js — covering diverse architectures and performance patterns. Benign traffic was generated using synthetic tools and real user-like behavior, simulating varying request rates, CPU/memory loads, and I/O operations to reflect real-world workload dynamics such as traffic spikes and batch processing. Attack data was produced using standard DoS tools to stress both network and application layers, mimicking real adversarial scenarios like volumetric floods and resource exhaustion. This blend of framework diversity, metric richness, and controlled anomaly injection results in a dataset that reflects real operational conditions and provides a solid foundation for training effective, generalizable ML models.
ML Models and Detection Approach
We evaluated multiple supervised learning algorithms for detecting DoS attacks, focusing on their ability to generalize across different frameworks and languages (Flask, Django, FastAPI, Nodejs, Golang). The classifiers include: Logistic Regression: A simple linear model for baseline comparisons [16]. Random Forest: A tree-based ensemble model known for its robustness to overfitting [17]. Gradient Boosting: An iterative boosting model suitable for handling imbalanced datasets [18]. Support Vector Machine (SVM): Effective for high-dimensional feature spaces [19]. Decision Tree: A fast and interpretable model [20]. Naive Bayes: Suitable for datasets where feature independence can be assumed [21]. K-Nearest Neighbors: A distance-based approach for capturing non-
linear patterns [22]. XGBoost: A high-performance gradient boosting model [23]. LightGBM: A scalable and efficient tree-based model optimized for large datasets [24].
All features were standardized using StandardScaler to ensure uniform scaling across models. Any missing values were imputed based on the mean or median of the respective feature.
A stratified 5-fold cross-validation strategy was employed to ensure robust evaluation across diverse data splits. For tree-based models, feature importance scores were computed to interpret the contribution of individual metrics.
Accuracy, precision, recall, and F1-score were computed for each classifier. Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) scores were used to assess model performance. In addition to the Confusion Matrix which provided a detailed breakdown of true positives, true negatives, false positives, and false negatives for each classifier.
The detection pipeline was implemented using Python, leveraging libraries, such as Scikit-learn, XGBoost, and LightGBM. The training and evaluation processes were automated to facilitate reproducibility and ensure consistent results across multiple frameworks.
Results and Discussion
This section presents a detailed analysis of the results derived from the machine learning classifiers applied to the combined dataset, integrating both application and node-level metrics from multiple frameworks. The discussion includes model performance metrics, feature importance analysis, and insights into classifier behavior for detecting anomalies.
The cross-validation accuracy of the classifiers was evaluated to assess their ability to generalize across different data subsets. The accuracy boxplot shows that ensemble models like Random Forest, Gradient Boosting, XGBoost, and LightGBM consistently outperformed
0.96-
other models, achieving nearly perfect performance with minimal variance. Among all classifiers: Random Forest and Decision Tree achieved perfect classification accuracy with no false predictions. XGBoost achieved an accuracy of 100 % with robust predictive power. Linear classifiers like Logistic Regression and simpler models like Naive Bayes showed relatively lower but acceptable accuracies. The classifier cross-validation accuracy plot highlights the stability of ensemble models compared to others (Fig. 3).
The confusion matrices provide detailed insights into the classification behavior of each ML model. Fig. 4 illustrates these matrices for all evaluated classifiers, showing the distribution of true positives, true negatives, false positives, and false negatives. Notably, XGBoost and Random Forest achieved perfect classification, with zero misclassifications of either attack or non-attack instances. In contrast, models like Naive Bayes and SVM showed some weaknesses. For example, Naive Bayes falsely identified 4,967 normal cases as attacks, reflecting its limitation in environments with overlapping feature distributions. SVM exhibited a slightly elevated false positive rate due to its sensitivity to kernel parameterization. These confusion matrices were generated by the authors using the labeled dataset of 49,990 samples, collected from a realistic Kubernetes cluster and detailed in the dataset description. The results confirm the relative strengths of ensemble models and provide a comparative evaluation of precision and recall across all classifiers.
Feature Importance
Feature importance was analyzed using: tree-based native importance for ensemble models and permutation importance for other classifiers where native importance is unavailable.
The following trends were observed: For tree-based models like Random Forest, LightGBM, and XGBoost, the most critical features were: CPU utilization metrics (cpu_sec_total, cpu_sec_idle), Packet transmission metrics (transmit_packets), System load averages (load1, load5, load15), and Memory metrics (virtual_memory_total, resident_memory_total). XGBoost and Gradient Boosting
Logistic Random Gradient Support Decision Navie K-Nearest XGBoost LightGBM Regression Forest Boosting Vector Tree Bayes Neighbors
Machine
Fig. 3. Boxplot comparing cross-validation accuracy across all evaluated machine learning models
0.92-
a
È3
О
о
<
0.880.84-
о о <3
о"
<N
О
о <3
о"
о о <3
о"
<N
о о о
о о <3
о"
о о о
о о <3
о"
(N
о о о
и
л
s
1102 24,468
m
0
2
2
й _ С U
.а
о -о
4967 24,006
8 9
7
^
in
oo" 2
й _ С U
.а
о -о й ü £ с
о
0
4 4
in in"
2
in
,9 0
2, 12
2
о Рч
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.