Все выпуски
- 2026 Том 18
- 2025 Том 17
- 2024 Том 16
- 2023 Том 15
- 2022 Том 14
- 2021 Том 13
- 2020 Том 12
- 2019 Том 11
- 2018 Том 10
- 2017 Том 9
- 2016 Том 8
- 2015 Том 7
- 2014 Том 6
- 2013 Том 5
- 2012 Том 4
- 2011 Том 3
- 2010 Том 2
- 2009 Том 1
-
Lidar and camera data fusion in self-driving cars
Компьютерные исследования и моделирование, 2022, т. 14, № 6, с. 1239-1253Sensor fusion is one of the important solutions for the perception problem in self-driving cars, where the main aim is to enhance the perception of the system without losing real-time performance. Therefore, it is a trade-off problem and its often observed that most models that have a high environment perception cannot perform in a real-time manner. Our article is concerned with camera and Lidar data fusion for better environment perception in self-driving cars, considering 3 main classes which are cars, cyclists and pedestrians. We fuse output from the 3D detector model that takes its input from Lidar as well as the output from the 2D detector that take its input from the camera, to give better perception output than any of them separately, ensuring that it is able to work in real-time. We addressed our problem using a 3D detector model (Complex-Yolov3) and a 2D detector model (Yolo-v3), wherein we applied the image-based fusion method that could make a fusion between Lidar and camera information with a fast and efficient late fusion technique that is discussed in detail in this article. We used the mean average precision (mAP) metric in order to evaluate our object detection model and to compare the proposed approach with them as well. At the end, we showed the results on the KITTI dataset as well as our real hardware setup, which consists of Lidar velodyne 16 and Leopard USB cameras. We used Python to develop our algorithm and then validated it on the KITTI dataset. We used ros2 along with C++ to verify the algorithm on our dataset obtained from our hardware configurations which proved that our proposed approach could give good results and work efficiently in practical situations in a real-time manner.
Ключевые слова: autonomous vehicles, self-driving cars, sensors fusion, Lidar, camera, late fusion, point cloud, images, KITTI dataset, hardware verification.
Lidar and camera data fusion in self-driving cars
Computer Research and Modeling, 2022, v. 14, no. 6, pp. 1239-1253Sensor fusion is one of the important solutions for the perception problem in self-driving cars, where the main aim is to enhance the perception of the system without losing real-time performance. Therefore, it is a trade-off problem and its often observed that most models that have a high environment perception cannot perform in a real-time manner. Our article is concerned with camera and Lidar data fusion for better environment perception in self-driving cars, considering 3 main classes which are cars, cyclists and pedestrians. We fuse output from the 3D detector model that takes its input from Lidar as well as the output from the 2D detector that take its input from the camera, to give better perception output than any of them separately, ensuring that it is able to work in real-time. We addressed our problem using a 3D detector model (Complex-Yolov3) and a 2D detector model (Yolo-v3), wherein we applied the image-based fusion method that could make a fusion between Lidar and camera information with a fast and efficient late fusion technique that is discussed in detail in this article. We used the mean average precision (mAP) metric in order to evaluate our object detection model and to compare the proposed approach with them as well. At the end, we showed the results on the KITTI dataset as well as our real hardware setup, which consists of Lidar velodyne 16 and Leopard USB cameras. We used Python to develop our algorithm and then validated it on the KITTI dataset. We used ros2 along with C++ to verify the algorithm on our dataset obtained from our hardware configurations which proved that our proposed approach could give good results and work efficiently in practical situations in a real-time manner.
-
Нейросетевой анализ транспортных потоков городских агломераций на основе данных публичных камер видеообзора
Компьютерные исследования и моделирование, 2021, т. 13, № 2, с. 305-318Адекватное моделирование сложной динамики городских транспортных потоков требует сбора больших объемов данных для определения характера соответствующих моделей и их калибровки. Вместе с тем оборудование специализированных постов наблюдения является весьма затратным мероприятием и не всегда технически возможно. Совокупность этих факторов приводит к недостаточному фактографическому обеспечению как систем оперативного управления транспортными потоками, так и специалистов по транспортному планированию с очевидными последствиями для качества принимаемых решений. В качестве способа обеспечить массовый сбор данных хотя бы для качественного анализа ситуаций достаточно давно применяется обзорные видеокамеры, транслирующие изображения в определенные ситуационные центры, где соответствующие операторы осуществляют контроль и управление процессами. Достаточно много таких обзорных камер предоставляют данные своих наблюдений в общий доступ, что делает их ценным ресурсом для транспортных исследований. Вместе с тем получение количественных данных с таких камер сталкивается с существенными проблемами, относящимися к теории и практике обработки видеоизображений, чему и посвящена данная работа. В работе исследуется практическое применение некоторых мейнстримовских нейросетевых технологий для определения основных характеристик реальных транспортных потоков, наблюдаемых камерами общего доступа, классифицируются возникающие при этом проблемы и предлагаются их решения. Для отслеживания объектов дорожного движения применяются варианты сверточных нейронных сетей, исследуются способы их применения для определения базовых характеристик транспортных потоков. Простые варианты нейронной сети используются для автоматизации при получении обучающих примеров для более глубокой нейронной сети YOLOv4. Сеть YOLOv4 использована для оценки характеристик движения (скорость, плотность потока) для различных направлений с записей камер видеонаблюдения.
Ключевые слова: искусственные нейронные сети, машинное зрение, машинное обучение, сопровождение объекта, сверточные нейронные сети.
Neural network analysis of transportation flows of urban aglomeration using the data from public video cameras
Computer Research and Modeling, 2021, v. 13, no. 2, pp. 305-318Correct modeling of complex dynamics of urban transportation flows requires the collection of large volumes of empirical data to specify types of the modes and their identification. At the same time, setting a large number of observation posts is expensive and technically not always feasible. All this results in insufficient factographic support for the traffic control systems as well as for urban planners with the obvious consequences for the quality of their decisions. As one of the means to provide large-scale data collection at least for the qualitative situation analysis, the wide-area video cameras are used in different situation centers. There they are analyzed by human operators who are responsible for observation and control. Some video cameras provided their videos for common access, which makes them a valuable resource for transportation studies. However, there are significant problems with getting qualitative data from such cameras, which relate to the theory and practice of image processing. This study is devoted to the practical application of certain mainstream neuro-networking technologies for the estimation of essential characteristics of actual transportation flows. The problems arising in processing these data are analyzed, and their solutions are suggested. The convolution neural networks are used for tracking, and the methods for obtaining basic parameters of transportation flows from these observations are studied. The simplified neural networks are used for the preparation of training sets for the deep learning neural network YOLOv4 which is later used for the estimation of speed and density of automobile flows.
-
Сверточные нейронные сети семейства YOLO для мобильных систем компьютерного зрения
Компьютерные исследования и моделирование, 2024, т. 16, № 3, с. 615-631Работа посвящена анализу известных классов моделей сверточных нейронных сетей и исследованию выбранных из них перспективных моделей для детектирования летающих объектов на изображениях. Под детектированием объектов (англ. — Object Detection) здесь понимаются обнаружение, локализация в пространстве и классификация летающих объектов. Комплексное исследование выбранных перспективных моделей сверточных нейронных сетей проводится с целью выявления наиболее эффективных из них для создания мобильных систем компьютерного зрения реального времени. Показано, что наиболее приемлемыми для детектирования летающих объектов на изображениях с учетом сформулированных требований к мобильным системам компьютерного зрения реального времени и, соответственно, к лежащим в их основе моделям сверточных нейронных сетей являются модели семейства YOLO, причем наиболее перспективными следует считать пять моделей из этого семейства: YOLOv4, YOLOv4-Tiny, YOLOv4-CSP, YOLOv7 и YOLOv7-Tiny. Для обучения, валидации и комплексного исследования этих моделей разработан соответствующий набор данных. Каждое размеченное изображение из набора данных включает от одного до нескольких летающих объектов четырех классов: «птица», «беспилотный летательный аппарат самолетного типа», «беспилотный летательный аппарат вертолетного типа» и «неизвестный объект» (объекты в воздушном пространстве, не входящие в первые три класса). Исследования показали, что все модели сверточных нейронных сетей по скорости детектирования объектов на изображении (по скорости вычисления модели) значительно превышают заданное пороговое значение, однако только модели YOLOv4-CSP и YOLOv7, причем только частично, удовлетворяют требованию по точности детектирования (классификации) летающих объектов. Наиболее сложным для детектирования классом объектов является класс «птица». При этом выявлено, что наиболее эффективной по точности классификации является модель YOLOv7, модель YOLOv4-CSP на втором месте. Обе модели рекомендованы к использованию в составе мобильной системы компьютерного зрения реального времени при условии увеличения в созданном наборе данных числа изображений с объектами класса «птица» и дообучения этих моделей с тем, чтобы они удовлетворяли требованию по точности детектирования летающих объектов каждого из четырех классов.
Ключевые слова: детектирование летающих объектов на изображениях, сверточная нейронная сеть, YOLO, мобильная система компьютерного зрения.
Convolutional neural networks of YOLO family for mobile computer vision systems
Computer Research and Modeling, 2024, v. 16, no. 3, pp. 615-631The work analyzes known classes of convolutional neural network models and studies selected from them promising models for detecting flying objects in images. Object detection here refers to the detection, localization in space and classification of flying objects. The work conducts a comprehensive study of selected promising convolutional neural network models in order to identify the most effective ones from them for creating mobile real-time computer vision systems. It is shown that the most suitable models for detecting flying objects in images, taking into account the formulated requirements for mobile real-time computer vision systems, are models of the YOLO family, and five models from this family should be considered: YOLOv4, YOLOv4-Tiny, YOLOv4-CSP, YOLOv7 and YOLOv7-Tiny. An appropriate dataset has been developed for training, validation and comprehensive research of these models. Each labeled image of the dataset includes from one to several flying objects of four classes: “bird”, “aircraft-type unmanned aerial vehicle”, “helicopter-type unmanned aerial vehicle”, and “unknown object” (objects in airspace not included in the first three classes). Research has shown that all convolutional neural network models exceed the specified threshold value by the speed of detecting objects in the image, however, only the YOLOv4-CSP and YOLOv7 models partially satisfy the requirements of the accuracy of detection of flying objects. It was shown that most difficult object class to detect is the “bird” class. At the same time, it was revealed that the most effective model is YOLOv7, the YOLOv4-CSP model is in second place. Both models are recommended for use as part of a mobile real-time computer vision system with condition of additional training of these models on increased number of images with objects of the “bird” class so that they satisfy the requirement for the accuracy of detecting flying objects of each four classes.
-
Обнаружение малоразмерных объектов на аэрофотоснимках с использованием сверточных нейронных сетей
Компьютерные исследования и моделирование, 2026, т. 18, № 4, с. 855-870В работе решается задача обнаружения малоразмерных объектов на аэрофотоснимках в видимом спектре. Высокая вариативность фона и слабая выраженность признаков делают детекцию малоразмерных объектов нетривиальной задачей. В таких условиях низкая эффективность классических методов компьютерного зрения на основе дескрипторов стимулирует переход к архитектурам на основе глубоких нейронных сетей, которые демонстрируют большую обобщающую способность и устойчивость к ложным срабатываниям. В качестве метода поиска объектов на изображении был выбран одноэтапный подход на базе нейросетевой предиктивной модели. Проведен анализ нейросетевых архитектур, относящихся к данному классу, приведены их достоинства и недостатки. В качестве базовой архитектуры детектора была взята YOLO версии 11 с добавлением блока многомасштабной агрегации признаков (используется для объединения информации, полученной из слоев искусственной нейронной сети с разными масштабами) и блока селективной интеграции с учетом размерности (блок автоматически определяет, по какому измерению (канал, высота, ширина) обрабатывать данные, и выборочно объединяет их). Данные модификации направлены как на повышение вычислительной эффективности и возможности развертывания нейросетевых моделей на борту беспилотных летательных аппаратов для анализа изображений в режиме реального времени, так и для улучшения точности детекции объектов малых размеров. Учитывая сложность аннотации изображений, для обучения и тестирования (в соотношении 90/10% соответственно) использовалась открытая синтетическая база изображений, содержащая порядка 4000 изображений. Обучающая выборка была расширена путем различных случайных аугментаций. Для оценки качества полученных предиктивных моделей использовалось среднее значение точности по всем классам как с порогом 50% пересечения с экспертной разметкой, так и с варьируемым порогом 50–95%. Переобучение контролировалось при помощи анализа кривых потерь. Предложенные модификации архитектуры YOLO позволили уменьшить время обработки изображений в два раза при сохранении точности распознавания.
Ключевые слова: распознавание малоразмерных объектов, компьютерное зрение, сверточные нейронные сети, аэрофотоснимки.
Small object detection in aerial images using convolutional neural networks
Computer Research and Modeling, 2026, v. 18, no. 4, pp. 855-870This paper addresses the problem of detecting small objects in visible-spectrum aerial imagery. High background variability and weak feature saliency make small object detection a non-trivial task. Under such conditions, classical computer vision algorithms based on hand-crafted descriptors exhibit low efficiency, prompting a shift towards deep neural network architectures, which demonstrate superior generalization capability and robustness to false positives. We selected a one-stage approach based on a neural network predictive model as the primary object detection method. Also, several neural network architectures belonging to this class were analyzed, outlining their advantages and disadvantages. As the baseline detector, we adopted YOLO version 11 and incorporated a multi-scale feature aggregation module (which combines information from neural network layers operating at different scales) and a dimension-aware selective integration module (which automatically determines the dimension (channel, height, or width) along which to process features and fuses them selectively). These modifications aim to both enhance computational efficiency, enabling deployment of neural network models onboard aerial vehicles for real-time image analysis, and improve small object detection accuracy. Given the complexity of image annotation, we used an open synthetic image database containing approximately 4 000 images for training and testing (with a 90/10 split, respectively). We extend the training set using various random augmentation techniques. To evaluate the performance of the resulting predictive models, we employed mean Average Precision across all classes, using both a fixed 50% intersectionover- union threshold and a varying threshold from 50% to 95%. Overfitting was monitored by analyzing loss curves during training process. The proposed modifications to the YOLO architecture reduced image processing time by a factor of two while maintaining detection accuracy.
-
Классификатор на базе оптимизированной архитектуры YOLO11 с применением инструментов CBAM и Grad-CAM для надежной диагностики больных растений
Компьютерные исследования и моделирование, 2026, т. 18, № 4, с. 871-889Надежная диагностика болезней растений компьютерным зрением требует не только высокой точности классификации, но и стабильной обобщающей способности и интерпретируемого принятия решений. При этом механизмы внимания безусловно полезны для моделей глубокого обучения, но эффективность решения задачи зависит от места и способа интеграции этих механизмов в архитектуру нейронной сети. В данном исследовании представлен классификатор на базе оптимизированной архитектуры YOLOv11m, а также результаты анализа результатов диагностики болезней растений в зависимости от сверточных блоков CBAM (Convolutional Block Attention Module), внедренных на разных уровнях архитектуры. Выполнен сравнительный анализ трех архитектурных вариантов модели: базовой архитектуры YOLO11m-Cls без механизма внимания, варианта с добавлением CBAM только в магистральную часть backbone и гибридной архитектуры, сочетающей сокращенное внимание в магистральной части backbone с дополнительным блоком CBAM на уровне классификационного слоя head. Все модели были обучены и протестированы в одинаковых экспериментальных условиях с использованием обширного набора данных, содержащего около 90 тысяч изображений 38 типов болезней растений. Экспериментальные результаты ясно показывают, что равномерное внедрение CBAM в магистральную часть backbone снижает устойчивость модели и ухудшает ее способность к обобщению, что проявляется в повышении валидационной функции потерь и понижении точности (Top-1 ≈ 90,5%) по сравнению с базовой архитектурой. Напротив, гибридная архитектура обеспечивает баланс между устойчивостью и способностью детектировать с повышением параметров точности Top-1 до 99,71% и Top-5 до 99,99% на этапе валидации, близким по значениям к параметрам базовой модели. Анализ интерпретируемости на основе метода Grad-CAM также положительно характеризует гибридную модель, которая формирует улучшенные и биологически обоснованные карты активации, четко фокусирующиеся именно на признаках заболевания и практически исключающие отвлечение внимания на фоновые элементы изображения. Важно отметить, что целью данной работы является не достижение наилучшей производительности среди всех классификаторов, а предоставление доказательств на уровне архитектурного проектирования относительно того, каким образом размещение механизма внимания влияет на устойчивость и интерпретируемость в классификационных моделях на базе YOLO. Полученные результаты предоставляют практические рекомендации по проектированию надежных и интерпретируемых систем искусственного интеллекта для агропромышленного сектора.
Ключевые слова: диагностика болезней растений, компьютерное зрение, сверточная нейронная сеть, YOLO11, Convolutional Block Attention Module (CBAM), Explainable Artificial Intelligence (XAI), Grad-CAM.
Design-optimized YOLO11 classification via strategic CBAM attention injection and Grad-CAM explainability for reliable plant disease diagnosis
Computer Research and Modeling, 2026, v. 18, no. 4, pp. 871-889Reliable plant disease diagnosis requires not only high classification accuracy, but also stable generalization and interpretable decision-making. Although attention mechanisms are useful to these deep learning models, the performance also depends on where and how they are integrated into the network architecture. This study presents a design-optimized YOLO11m-based classification framework that systematically investigates the impact of Convolutional Block Attention Module (CBAM) injection at different architectural levels for plant disease diagnosis. We perform a comparative modelcontrolled analysis of three model architectures: (i) the baseline YOLO11m-Cls architecture lacking attention, (ii) only adding the backbone block along with CBAM and (iii) a hybrid architecture that includes reduced backbone attention combined with CBAM added at classification head. All models are trained and tested under the same experimental settings using a largescale dataset with around 90 000 images of 38 types of plant diseases. Experimental results clearly show that the uniform injection of CBAM into the backbone reduces stability but causes generalization to worsen with a higher validation loss and significantly lower Top-1 accuracy (≈ 90.5%), while the hybrid attention design balances stability and discrimination, with Top-1 accuracy up to 99.71%, Top-5 accuracy up to 99.99% and near-baseline validation behavior respectively Grad-CAMbased interpretability analysis also demonstrates that the hybrid model generates enhanced and biologically interpretable activation maps, which are more disease-specific with less distraction from background. Notably, the aim of our work is not for achieving superior performance over all classifiers but instead only to provide design-level evidence on how placing attention modulates rigidity and interpretability in YOLO-based classification models. The results provide practical architectural considerations for building dependable and interpretable AI systems in agriculture.
Журнал индексируется в Scopus
Полнотекстовая версия журнала доступна также на сайте научной электронной библиотеки eLIBRARY.RU
Журнал входит в систему Российского индекса научного цитирования.
Журнал включен в базу данных Russian Science Citation Index (RSCI) на платформе Web of Science
Международная Междисциплинарная Конференция "Математика. Компьютер. Образование"





