Evolution of transformer architectures: from self-attention to efficient sparse models
https://doi.org/10.21122/2309-4923-2026-2-78-81
Abstract
The article is devoted to the study of new architectures of artificial neural networks aimed at improving classical transformers by reducing computational costs and increasing efficiency. The main stages of the evolution of transformers are considered, starting with the introduction of the mechanism of self-attention and ending with modern sparse models. Particular attention is paid to optimization techniques such as parametrically efficient fine-tuning (PEFT), adaptive layers (adapters), as well as the use of external knowledge repositories and memory vectors to expand the long-term memory of models. The article concludes with a discussion of current achievements, limitations, and promising areas for future research in this area.
About the Authors
V. O. GulyaevRussian Federation
Vladislav O. Gulyaev - Master's student.
Samara
E-mail: vladislavgulaev03@gmail.com
O. I. Zakharova
Russian Federation
Oksna I. Zaharova
Samara
E-mail: o.zaharova@psuti.ru
References
1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. arXiv [Preprint]. 2017. Available from: https://arxiv.org/abs/1706.03762 (accessed 12 Jun 2026).
2. Child R., Gray S., Radford A., Sutskever I. Generating long sequences with sparse transformers. arXiv [Preprint]. 2019. Available from: https://arxiv.org/abs/1904.10509 (accessed 12 Jun 2026).
3. Hu E.J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., et al. LoRA: Low-rank adaptation of large language models. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2106.09685 (accessed 12 Jun 2026).
4. Bulatov A., Kuratov Y., Burtsev M.S. Recurrent memory transformer. arXiv [Preprint]. 2022. Available from: https://arxiv.org/abs/2207.06881 (accessed 12 Jun 2026).
5. Ramesh A., Pavlov M., Goh G., Scott G., Voss C., Radford A., et al. Zero-shot text-to-image generation. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2102.12092 (accessed 12 Jun 2026).
6. Dhariwal P., Jun H., Payne C., Kim J.W., Radford A., Sutskever I. Jukebox: A Generative model for music. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.00341 (accessed 12 Jun 2026).
7. Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., et al. Language models are few-shot learners. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.14165 (accessed 12 Jun 2026).
Review
For citations:
Gulyaev V.O., Zakharova O.I. Evolution of transformer architectures: from self-attention to efficient sparse models. «System analysis and applied information science». 2026;(2):78-81. (In Russ.) https://doi.org/10.21122/2309-4923-2026-2-78-81
JATS XML





















