<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">sapi</journal-id><journal-title-group><journal-title xml:lang="ru">Системный анализ и прикладная информатика</journal-title><trans-title-group xml:lang="en"><trans-title>«System analysis and applied information science»</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2309-4923</issn><issn pub-type="epub">2414-0481</issn><publisher><publisher-name>Belarusian National Technical University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.21122/2309-4923-2026-2-78-81</article-id><article-id custom-type="elpub" pub-id-type="custom">sapi-817</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>Информационные технологии</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>Information technologies</subject></subj-group></article-categories><title-group><article-title>Эволюция архитектур трансформеров: от самовнимания к эффективным Sparse-моделям</article-title><trans-title-group xml:lang="en"><trans-title>Evolution of transformer architectures: from self-attention to efficient sparse models</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Гуляев</surname><given-names>В. О.</given-names></name><name name-style="western" xml:lang="en"><surname>Gulyaev</surname><given-names>V. O.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Гуляев Владислав Олегович - Магистрант.г. СамараE-mail: vladislavgulaev03@gmail.com</p></bio><bio xml:lang="en"><p>Vladislav O. Gulyaev - Master's student.SamaraE-mail: vladislavgulaev03@gmail.com</p></bio><email xlink:type="simple">vladislavgulaev03@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Захарова</surname><given-names>О. И.</given-names></name><name name-style="western" xml:lang="en"><surname>Zakharova</surname><given-names>O. I.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Захарова Оксана Игоревнаг. СамараE-mail: o.zaharova@psuti.ru</p></bio><bio xml:lang="en"><p>Oksna I. Zaharova SamaraE-mail: o.zaharova@psuti.ru</p></bio><email xlink:type="simple">o.zaharova@psuti.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Поволжский государственный университет телекоммуникаций и информатики</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Volga State University of Telecommunications and Informatics</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>17</day><month>07</month><year>2026</year></pub-date><volume>0</volume><issue>2</issue><fpage>78</fpage><lpage>81</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Гуляев В.О., Захарова О.И., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Гуляев В.О., Захарова О.И.</copyright-holder><copyright-holder xml:lang="en">Gulyaev V.O., Zakharova O.I.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://sapi.bntu.by/jour/article/view/817">https://sapi.bntu.by/jour/article/view/817</self-uri><abstract><p>Статья посвящена исследованию новых архитектур искусственных нейронных сетей, направленных на улучшение классических трансформеров путем уменьшения вычислительных затрат и увеличения эффективности. Рассматриваются основные этапы эволюции трансформеров, начиная с введения механизма самовнимания и заканчивая современными Sparse-моделями. Особое внимание уделяется методам оптимизации, таким как параметрически эффективный fine-tuning (PEFT), адаптивные слои (adapters), а также использование внешних хранилищ знаний и mem-векторов для расширения долговременной памяти моделей. Статья завершается обсуждением текущих достижений, ограничений и перспективных направлений будущих исследований в данной области.</p></abstract><trans-abstract xml:lang="en"><p>The article is devoted to the study of new architectures of artificial neural networks aimed at improving classical transformers by reducing computational costs and increasing efficiency. The main stages of the evolution of transformers are considered, starting with the introduction of the mechanism of self-attention and ending with modern sparse models. Particular attention is paid to optimization techniques such as parametrically efficient fine-tuning (PEFT), adaptive layers (adapters), as well as the use of external knowledge repositories and memory vectors to expand the long-term memory of models. The article concludes with a discussion of current achievements, limitations, and promising areas for future research in this area.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>трансформеры</kwd><kwd>механизм самовнимания</kwd><kwd>sparse-модели</kwd><kwd>PEFT</kwd><kwd>mem-векторы</kwd></kwd-group><kwd-group xml:lang="en"><kwd>transformers</kwd><kwd>self-attention mechanism</kwd><kwd>sparse models</kwd><kwd>PEFT</kwd><kwd>memory vectors</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. arXiv [Preprint]. 2017. Available from: https://arxiv.org/abs/1706.03762 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. arXiv [Preprint]. 2017. Available from: https://arxiv.org/abs/1706.03762 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Child R., Gray S., Radford A., Sutskever I. Generating long sequences with sparse transformers. arXiv [Preprint]. 2019. Available from: https://arxiv.org/abs/1904.10509 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Child R., Gray S., Radford A., Sutskever I. Generating long sequences with sparse transformers. arXiv [Preprint]. 2019. Available from: https://arxiv.org/abs/1904.10509 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Hu E.J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., et al. LoRA: Low-rank adaptation of large language models. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2106.09685 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Hu E.J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., et al. LoRA: Low-rank adaptation of large language models. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2106.09685 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Bulatov A., Kuratov Y., Burtsev M.S. Recurrent memory transformer. arXiv [Preprint]. 2022. Available from: https://arxiv.org/abs/2207.06881 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Bulatov A., Kuratov Y., Burtsev M.S. Recurrent memory transformer. arXiv [Preprint]. 2022. Available from: https://arxiv.org/abs/2207.06881 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Ramesh A., Pavlov M., Goh G., Scott G., Voss C., Radford A., et al. Zero-shot text-to-image generation. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2102.12092 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Ramesh A., Pavlov M., Goh G., Scott G., Voss C., Radford A., et al. Zero-shot text-to-image generation. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2102.12092 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Dhariwal P., Jun H., Payne C., Kim J.W., Radford A., Sutskever I. Jukebox: A Generative model for music. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.00341 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Dhariwal P., Jun H., Payne C., Kim J.W., Radford A., Sutskever I. Jukebox: A Generative model for music. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.00341 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., et al. Language models are few-shot learners. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.14165 (accessed 12 Jun 2026).</mixed-citation><mixed-citation xml:lang="en">Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., et al. Language models are few-shot learners. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.14165 (accessed 12 Jun 2026).</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
