Текущий выпуск Номер 4, 2026 Том 18

Все выпуски

Результаты поиска по 'the nearest pattern':
Найдено статей: 2
  1. Воронина М.Ю., Орлов Ю.Н.
    Определение автора текста методом сегментации
    Компьютерные исследования и моделирование, 2022, т. 14, № 5, с. 1199-1210

    В работе описывается метод распознавания авторов литературных текстов по близости фрагментов, на которые разделен отдельный текст, к эталону автора. Эталоном является эмпирическое распределение частот буквосочетаний, построенное по обучающей выборке, куда вошли экспертно отобранные достоверно известные произведения данного автора. Совокупность эталонов разных авторов образует библиотеку, внутри которой и решается задача об идентификации автора неизвестного текста. Близость между текстами понимается в смысле нормы в L1 для вектора частот буквосочетаний, который строится для каждого фрагмента и для текста в целом. Автором неизвестного текста назначается тот, эталон которого чаще всего выбирается в качестве ближайшего для набора фрагментов, на которые разделен текст. Длина фрагмента оптимизируется исходя из принципа максимального различия расстояний от фрагментов до эталонов в задаче распознавания «свой–чужой». Тестирование метода проведено на корпусе отечественных и зарубежных (в переводе) авторов. Были собраны 1783 текста 100 авторов суммарным объемом примерно 700 млн знаков. Чтобы исключить тенденциозность отбора авторов, рассматривались авторы, фамилии которых начинались на одну и ту же букву (в данном случае Л). Ошибка идентификации по биграммам составила 12%. Наряду с достаточно высокой точностью данный метод обладает еще одним важным свойством: он позволяет оценить вероятность того, что эталон автора рассматриваемого текста в библиотеке отсутствует. Эта вероятность может быть оценена по результатам статистики ближайших эталонов для малых фрагментов текста. В работе исследуются также статистические цифровые портреты писателей: это совместные эмпирические распределения вероятности того, что некоторая доля текста идентифицируется на заданном уровне доверия. Практическая важность этих статистик в том, что носители соответствующих распределений практически не пересекаются для своих и чужих эталонов, что позволяет распознать эталонное распределение буквосочетаний на высоком уровне доверия.

    Voronina M.Y., Orlov Y.N.
    Identification of the author of the text by segmentation method
    Computer Research and Modeling, 2022, v. 14, no. 5, pp. 1199-1210

    The paper describes a method for recognizing authors of literary texts by the proximity of fragments into which a separate text is divided to the standard of the author. The standard is the empirical frequency distribution of letter combinations, built on a training sample, which included expertly selected reliably known works of this author. A set of standards of different authors forms a library, within which the problem of identifying the author of an unknown text is solved. The proximity between texts is understood in the sense of the norm in L1 for the frequency vector of letter combinations, which is constructed for each fragment and for the text as a whole. The author of an unknown text is assigned the one whose standard is most often chosen as the closest for the set of fragments into which the text is divided. The length of the fragment is optimized based on the principle of the maximum difference in distances from fragments to standards in the problem of recognition of «friend–foe». The method was tested on the corpus of domestic and foreign (translated) authors. 1783 texts of 100 authors with a total volume of about 700 million characters were collected. In order to exclude the bias in the selection of authors, authors whose surnames began with the same letter were considered. In particular, for the letter L, the identification error was 12%. Along with a fairly high accuracy, this method has another important property: it allows you to estimate the probability that the standard of the author of the text in question is missing in the library. This probability can be estimated based on the results of the statistics of the nearest standards for small fragments of text. The paper also examines statistical digital portraits of writers: these are joint empirical distributions of the probability that a certain proportion of the text is identified at a given level of trust. The practical importance of these statistics is that the carriers of the corresponding distributions practically do not overlap for their own and other people’s standards, which makes it possible to recognize the reference distribution of letter combinations at a high level of confidence.

  2. Щербань И.В., Лысенко Л.В., Щербань О.Г., Калитин К.Ю.
    Статистический анализ и моделирование паттернов активации обонятельной луковицы на основе немаркированных пространственных точечных процессов
    Компьютерные исследования и моделирование, 2026, т. 18, № 4, с. 1005-1019

    В нейробиологии исследование механизмов кодирования запахов требует анализа пространственных паттернов активации обонятельных структур (гломерул), реконструируемых по данным мультифотонной микроскопии. Однако отсутствие формального статистического аппарата для анализа обобщенных карт, полученных на группе животных, ограничивает воспроизводимость результатов и затрудняет синтез прогностических моделей. Для преодоления этих ограничений разработана методология анализа карт ольфакторной активности, суммарно зарегистрированных на нескольких животных, основанная на теории немаркированных случайных точечных процессов. Методология включает процедуру предобработки данных и алгоритм численного анализа, реализованный в среде R с использованием пакета spatstat. Предложенный подход позволяет: (1) перейти от исходных карт гломерулярной активности к точечным паттернам с сохранением информации о размерах гломерул; (2) выполнить анализ точечных паттернов и рассчитать пространственно-морфологические характеристики специфичных для каждого одоранта областей (доменов) устойчивой активации гломерул; (3) выполнить статистическую проверку гипотез о пространственной случайности точечных паттернов с использованием $K$-функции Рипли, $G$-функции ближайших соседей и метода Монте-Карло; (4) синтезировать параметрическую модель парных взаимодействий Штрауса, параметры которой $(r_{PI}, \gamma)$ имеют ясную биологическую интерпретацию — масштаб пространственного взаимодействия гломерул и силу комодуляции ответов соответственно.

    Валидация методологии выполнена на экспериментальных данных, полученных на 24 лабораторных крысах (10 особей стимулированы камфорой, 14 — метилбензоатом). Подобранные модели Штрауса продемонстрировали близкие, но специфичные для каждого одоранта параметры: радиус взаимодействия $r_{PI}$ составил 150 мкм для камфоры и 120 мкм для метилбензоата, коэффициент взаимодействия $\gamma$ — 0,95 и 0,89 соответственно. Валидация с использованием $Q$-$Q$-графиков сглаженных остатков подтвердила адекватность моделей, а рассчитанные на этапе (2) суммарная площадь доменов (0,60 мм2 и 0,73 мм2) и плотность реакций (38 и 43 точки/мм2) полностью согласуются с параметрическими сигнатурами $(r_{PI}, \gamma)$ моделей.

    Предложенный подход обеспечивает воспроизводимую количественную оценку гломерулярных доменов в единой стереотаксической системе координат и может быть распространен на другие одоранты и биологические виды. Все выводы получены на наркотизированных животных; экстраполяция на активные обонятельные стратегии бодрствующих животных требует дополнительных исследований.

    Shcherban I.V., Lysenko L.V., Shcherban O.G., Kalitin K.Y.
    Statistical analysis and modeling of olfactory bulb activation patterns using unmarked spatial point processes
    Computer Research and Modeling, 2026, v. 18, no. 4, pp. 1005-1019

    In neuroscience, the study of odor coding mechanisms requires the analysis of spatial activation patterns of olfactory structures (glomeruli) reconstructed from multiphoton microscopy data. However, the lack of a formal statistical framework for analyzing population-level summary maps limits result reproducibility and hinders the development of predictive models. To address these limitations, we developed a novel methodology for the analysis of olfactory activity maps aggregated across multiple animals, based on the theory of unmarked spatial point processes. The methodology includes a data preprocessing procedure and a numerical analysis algorithm implemented in the R environment using the spatstat package.

    The proposed approach enables: (1) transformation of raw glomerular activity maps into point patterns while preserving information about glomerular sizes (replacing size information with local point density is a methodological compromise reflecting the “functional weight” of glomerular input); (2) analysis of point patterns based on spatial morphometric characteristics of domains — regions of stable glomerular activation in the olfactory bulb, each approximated by an ellipse, with ellipse parameters (center coordinates in stereotaxic space, major and minor axis lengths, orientation angles), areas, and intra-ellipse point densities reflecting odorant-specific response signatures; (3) statistical hypothesis testing for spatial randomness (Complete Spatial Randomness) using Ripley’s $K$-function, the nearest-neighbor G-function, and Monte Carlo simulations; (4) synthesis of a parametric pairwise interaction model (Strauss process), whose parameters $(r_{PI}, \gamma)$ have a clear biological interpretation — the spatial interaction scale of glomeruli and the strength of response comodulation, respectively.

    The methodology was validated using experimental data obtained from 24 laboratory rats: 10 animals stimulated with camphor and 14 with methyl benzoate. The fitted Strauss models yielded close but odorant-specific parameters: interaction radii $r_{PI}$ of 150 $\mu$m (camphor) and 120 μm (methyl benzoate); interaction parameters $gamma$ of 0.95 and 0.89, respectively. The total domain areas (0.60 mm2 and 0.73 mm2) and point densities (38 and 43 points/mm2) calculated at the first stage of analysis are fully consistent with the parametric signatures $(r_{PI}, \gamma)$ of the Strauss model. Model validation using $Q$-$Q$ plots of smoothed residuals confirmed their adequacy.

    Our results are consistent with data previously obtained using genetic labeling and functional mapping techniques, demonstrating the correctness of the proposed methodology and the effectiveness of multiphoton laser scanning microscopy for such applications. The proposed framework provides reproducible quantitative assessment of glomerular domains within a unified stereotaxic coordinate system and can be extended to other odorants and biological species. All findings were obtained under anesthesia; extrapolation to active olfactory strategies in awake animals requires further investigation.

Журнал индексируется в Scopus

Полнотекстовая версия журнала доступна также на сайте научной электронной библиотеки eLIBRARY.RU

Журнал включен в базу данных Russian Science Citation Index (RSCI) на платформе Web of Science

Международная Междисциплинарная Конференция "Математика. Компьютер. Образование"

Международная Междисциплинарная Конференция МАТЕМАТИКА. КОМПЬЮТЕР. ОБРАЗОВАНИЕ.