<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">sapi</journal-id><journal-title-group><journal-title xml:lang="ru">Системный анализ и прикладная информатика</journal-title><trans-title-group xml:lang="en"><trans-title>«System analysis and applied information science»</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2309-4923</issn><issn pub-type="epub">2414-0481</issn><publisher><publisher-name>Belarusian National Technical University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.21122/2309-4923-2022-4-60-64</article-id><article-id custom-type="elpub" pub-id-type="custom">sapi-595</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>Обработка информации и принятие решений</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>Data processing and decision–making</subject></subj-group></article-categories><title-group><article-title>Optimizing the performance of a server-based classification for a large business document flow</article-title><trans-title-group xml:lang="en"><trans-title>Optimizing the performance of a server-based classification for a large business document flow</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Cлавин</surname><given-names>О. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Slavin</surname><given-names>O. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Главный научный сотрудник, доктор технических наук</p></bio><email xlink:type="simple">oslavin@isa.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Федеральный исследовательский центр “Информатика и управление” Российской академии наук</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Federal Research Center “Informatics and Management”  of the Russian Academy of Sciences; &#13;
Smart Engines Service LLC</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2022</year></pub-date><pub-date pub-type="epub"><day>24</day><month>02</month><year>2023</year></pub-date><volume>0</volume><issue>4</issue><fpage>60</fpage><lpage>64</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Cлавин О.А., 2023</copyright-statement><copyright-year>2023</copyright-year><copyright-holder xml:lang="ru">Cлавин О.А.</copyright-holder><copyright-holder xml:lang="en">Slavin O.A.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://sapi.bntu.by/jour/article/view/595">https://sapi.bntu.by/jour/article/view/595</self-uri><abstract><p>.</p></abstract><trans-abstract xml:lang="en"><p>The document categorization problem in the case of a large business document flow is considered. Textual and visual embeddings were employed for classification. Textual embeddings were extracted via OCR Tesseract. The Viola and Jones method was applied to generate visual embeddings. This paper describes the performance optimization technology for the implemented classification algorithm. Servers with Intel CPUs were used for the algorithm execution. For single-threaded implementation, high-level and low-level optimizations were performed. High-level optimization was based on the parametrization of the recognition algorithms and the employment of intermediate data. Low-level optimization was carried out via compiler tools allowing for an extended set of SIMD instructions. The implementation of parallelization with several multithreaded applications on multiple servers was also described. The proposed solution was tested using own test data sets of business documents. The proposed method can be applied in modern information systems to analyze the content of a large flow of digital document images.</p></trans-abstract><kwd-group xml:lang="en"><kwd>text analysis</kwd><kwd>document recognition</kwd><kwd>document classification</kwd><kwd>speedup</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Башкатова, A. Цифровая экономика плодит все больше бумаг: Россияне не скоро перестанут носить в организации справки // Независимая Газета. – 2019 – 14 ноя. [Электронный ресурс] – Режим доступа: https://www.ng.ru/ economics/2019-11-14/4_7727_paper.html, – Загл. с экрана – Яз. рус. Дата доступа – 08.11.2022.</mixed-citation><mixed-citation xml:lang="en">Bashkatova, A. Cifrovaya ekonomika plodit vse bol’she bumag: Rossiyane ne skoro perestanut nosit’ v organizacii spravki // Nezavisimaya Gazeta – 2019 – 14 nov. . [Электронный ресурс] – Режим доступа: https://www.ng.ru/economics/2019-11-14/4_7727_ paper.html, – Загл. с экрана – Яз. рус. Дата доступа – 08.11.2022.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Liu, L., Wang, Z., Qiu, T., Chen, Q., Lu, Y., Suen, C.Y. Document image classification: Progress over two decades, Neurocomputing 2021, 453: 223-240.</mixed-citation><mixed-citation xml:lang="en">Liu, L., Wang, Z., Qiu, T., Chen, Q., Lu, Y., Suen, C.Y. Document image classification: Progress over two decades, Neurocomputing 2021, 453: 223-240.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Byun, Y., Lee, Y. Form classification using DP matching. ACM Symposium on Applied Computing 2000; 1: 1–4.</mixed-citation><mixed-citation xml:lang="en">Byun, Y., Lee, Y. Form classification using DP matching. ACM Symposium on Applied Computing 2000; 1: 1–4.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. 2019. [Электронный ресурс] – Режим доступа: https://arxiv.org/abs/1810.04805/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation><mixed-citation xml:lang="en">Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. 2019. [Электронный ресурс] – Режим доступа: https://arxiv.org/abs/1810.04805/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Rubin, T.N., Chambers, A., Smyth, P., Steyvers, M. Statistical topic models for multi-label document classification. Machine Learning – 2011, Vol. 88, № 1, 157–208. https://doi.org/10.1007/s10994-011-5272-5.</mixed-citation><mixed-citation xml:lang="en">Rubin, T.N., Chambers, A., Smyth, P., Steyvers, M. Statistical topic models for multi-label document classification. Machine Learning – 2011, Vol. 88, № 1, 157–208.  https://doi.org/10.1007/s10994-011-5272-5.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Vorontsov, K.V., Potapenko, A.A. Tutorial on probabilistic topic modeling: Additive regularization for stochastic matrix factorization. Communications in Computer and Information Science – 2014, Vol. 436, pp. 29-46. https://doi.org/10.1007/978-3-31912580-0_3.</mixed-citation><mixed-citation xml:lang="en">Vorontsov,  K.V., Potapenko, A.A. Tutorial on probabilistic topic modeling: Additive regularization for stochastic matrix factorization. Communications in Computer and Information Science  –  2014,  Vol. 436, pp. 29-46. https://doi.org/10.1007/978-3-31912580-0_3.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">NIST Special Database 2 [Электронный ресурс] – Режим доступа: https://www.nist.gov/srd/nist-special-database-2/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation><mixed-citation xml:lang="en">NIST Special Database 2 [Электронный ресурс] – Режим доступа: https://www.nist.gov/srd/nist-special-database-2/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Tobacco-3482 [Электронный ресурс] – Режим доступа: https://www.kaggle.com/patrickaudriaz/tobacco3482jpg/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation><mixed-citation xml:lang="en">Tobacco-3482 [Электронный ресурс] – Режим доступа: https://www.kaggle.com/patrickaudriaz/tobacco3482jpg/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">OCR Tesseract [Электронный ресурс] – Режим доступа: https://github.com/tesseract-ocr/tesseract/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation><mixed-citation xml:lang="en">OCR Tesseract [Электронный ресурс] – Режим доступа: https://github.com/tesseract-ocr/tesseract/, – Загл. с экрана – Яз. англ. Дата доступа – 08.11.2022.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Tereshin, A.A., Usilin, S.A., Arlazarov, V.V. Performance Improvement of Multi-class Detection Using Greedy Algorithm for Viola-Jones Cascade Selection. Proceedings Volume 10696, Tenth International Conference on Machine Vision (ICMV 2017); 106960D (2018). https://doi.org/10.1117/12.2310101</mixed-citation><mixed-citation xml:lang="en">Tereshin, A.A., Usilin, S.A., Arlazarov, V.V. Performance Improvement of Multi-class Detection Using Greedy Algorithm for Viola-Jones Cascade Selection. Proceedings Volume 10696, Tenth International Conference on Machine Vision (ICMV 2017); 106960D (2018). https://doi.org/10.1117/12.2310101</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Slavin, O.A., Farsobina, V., Myshev, A.V. Analyzing the content of business documents recognized with a large number of errors using modified Levenshtein distance. Cyber-Physical Systems: Intelligent Models and Algorithms. – 2022, Springer Nature Switzerland AG., Vol. 417, pp. 267 – 279. https://doi.org/10.1007/978-3-030-95116-0</mixed-citation><mixed-citation xml:lang="en">Slavin, O.A., Farsobina, V., Myshev, A.V. Analyzing the content of business documents recognized with a large number of errors using modified Levenshtein distance. Cyber-Physical Systems: Intelligent Models and Algorithms. –  2022, Springer Nature Switzerland AG., Vol. 417, pp. 267 – 279. https://doi.org/10.1007/978-3-030-95116-0</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Slavin, O.A. Using Special Text Points in the Recognition of Documents. Studies in Systems, Decision and Control. – 2020, Springer Nature Switzerland AG., Vol 259. pp. 43–53. https://doi.org/10.1007/978-3-030-32579-4_4</mixed-citation><mixed-citation xml:lang="en">Slavin, O.A. Using Special Text Points in the Recognition of Documents. Studies in Systems, Decision and Control. – 2020, Springer Nature Switzerland AG., Vol 259. pp. 43–53. https://doi.org/10.1007/978-3-030-32579-4_4</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Konaka, F., Miura, T. Semantic similarity for sequenced shingles, – 2015 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing (PACRIM), pp. 12-17. https://doi.org/10.1109/PACRIM.2015.7334801.</mixed-citation><mixed-citation xml:lang="en">Konaka, F., Miura, T. Semantic similarity for sequenced shingles, – 2015 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing (PACRIM), pp. 12-17. https://doi.org/10.1109/PACRIM.2015.7334801.</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Acar, U.A., Blelloch, G.E., Harper, R. Selective memorization. ACM SIGPLAN Notices, – 2003, Vol. 38, Issue 1, pp 14–25. https://doi.org/10.1145/640128.604133</mixed-citation><mixed-citation xml:lang="en">Acar, U.A., Blelloch, G.E., Harper, R. Selective memorization. ACM SIGPLAN Notices, –  2003, Vol. 38, Issue 1, pp 14–25. https://doi.org/10.1145/640128.604133</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Tatarowicz, A.L., Curino, C., Jones, E. P. C. and Madden, S. Lookup Tables: Fine-Grained Partitioning for Distributed Databases. – 2012 IEEE 28th International Conference on Data Engineering, pp. 102-113. https://doi.org/10.1109/ICDE.2012.26.</mixed-citation><mixed-citation xml:lang="en">Tatarowicz, A.L., Curino, C., Jones, E. P. C. and Madden, S. Lookup Tables: Fine-Grained Partitioning for Distributed Databases. – 2012 IEEE 28th International Conference on Data Engineering, pp. 102-113. https://doi.org/10.1109/ICDE.2012.26</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
