Preview

«System analysis and applied information science»

Advanced search

Evolution of transformer architectures: from self-attention to efficient sparse models

https://doi.org/10.21122/2309-4923-2026-2-78-81

Abstract

The article is devoted to the study of new architectures of artificial neural networks aimed at improving classical transformers by reducing computational costs and increasing efficiency. The main stages of the evolution of transformers are considered, starting with the introduction of the mechanism of self-attention and ending with modern sparse models. Particular attention is paid to optimization techniques such as parametrically efficient fine-tuning (PEFT), adaptive layers (adapters), as well as the use of external knowledge repositories and memory vectors to expand the long-term memory of models. The article concludes with a discussion of current achievements, limitations, and promising areas for future research in this area.

About the Authors

V. O. Gulyaev
Volga State University of Telecommunications and Informatics
Russian Federation

Vladislav O. Gulyaev - Master's student.
Samara
E-mail: vladislavgulaev03@gmail.com





O. I. Zakharova
Volga State University of Telecommunications and Informatics
Russian Federation

Oksna I. Zaharova 
Samara
E-mail: o.zaharova@psuti.ru





References

1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. arXiv [Preprint]. 2017. Available from: https://arxiv.org/abs/1706.03762 (accessed 12 Jun 2026).

2. Child R., Gray S., Radford A., Sutskever I. Generating long sequences with sparse transformers. arXiv [Preprint]. 2019. Available from: https://arxiv.org/abs/1904.10509 (accessed 12 Jun 2026).

3. Hu E.J., Shen Y., Wallis P., Allen-Zhu Z., Li Y., Wang S., et al. LoRA: Low-rank adaptation of large language models. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2106.09685 (accessed 12 Jun 2026).

4. Bulatov A., Kuratov Y., Burtsev M.S. Recurrent memory transformer. arXiv [Preprint]. 2022. Available from: https://arxiv.org/abs/2207.06881 (accessed 12 Jun 2026).

5. Ramesh A., Pavlov M., Goh G., Scott G., Voss C., Radford A., et al. Zero-shot text-to-image generation. arXiv [Preprint]. 2021. Available from: https://arxiv.org/abs/2102.12092 (accessed 12 Jun 2026).

6. Dhariwal P., Jun H., Payne C., Kim J.W., Radford A., Sutskever I. Jukebox: A Generative model for music. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.00341 (accessed 12 Jun 2026).

7. Brown T.B., Mann B., Ryder N., Subbiah M., Kaplan J., Dhariwal P., et al. Language models are few-shot learners. arXiv [Preprint]. 2020. Available from: https://arxiv.org/abs/2005.14165 (accessed 12 Jun 2026).


Review

For citations:


Gulyaev V.O., Zakharova O.I. Evolution of transformer architectures: from self-attention to efficient sparse models. «System analysis and applied information science». 2026;(2):78-81. (In Russ.) https://doi.org/10.21122/2309-4923-2026-2-78-81

Views: 119

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2309-4923 (Print)
ISSN 2414-0481 (Online)