Preview

«System analysis and applied information science»

Advanced search

Parallelism strategies as a key factor for deploying Large Language Models on consumer gpus

https://doi.org/10.21122/2309-4923-2026-1-54-59

Abstract

The exponential growth in the size of Large Language Models (LLMs) creates significant barriers to their local deployment, primarily due to Video RAM (VRAM) shortages on single devices. The aim of this work is to identify and substantiate the most effective parallelism strategy for LLM inference on consumer Graphics Processing Unit (GPU) clusters connected via a slow PCIe bus. Research methods included a series of experiments comparing a monolithic architecture (NVIDIA RTX A6000) and a distributed system (2x NVIDIA RTX 3090) using the vLLM framework. The impact of Tensor Parallelism (TP) and Pipeline Parallelism (PP) on key metrics – throughput, latency (TTFT, TPOT), and power consumption stability – was analyzed while running the DeepSeek-R1-Distill-Llama-14B model. The results unequivocally indicate the unsuitability of Tensor Parallelism for systems without NVLink due to critical synchronization delays. It is proven that Pipeline Parallelism is the only viable strategy for PCIe clusters, ensuring high throughput despite the presence of idle periods («bubbles») and a less stable power consumption profile compared to the monolithic solution. In conclusion, recommendations for using multi-GPU configurations are formulated: they represent the optimal economic choice for memory-critical tasks, such as Retrieval-Augmented Generation (RAG), allowing VRAM scaling at a significantly lower cost than professional analogs.

About the Authors

K. S. Kurochka
Sukhoi State Technical University of Gomel
Belarus

Konstantin S. Kurochka – PhD in Engineering, Associate Professor.

Gomel
kurochka@gstu.by



Yu. S. Basharymau
Sukhoi State Technical University of Gomel
Belarus

Yury S. Basharymau – Assistant.

Gomel
basharymauyury@gmail.com



Yu. D. Youzhanka
Sukhoi State Technical University of Gomel
Belarus

Yury D. Youzhanka – Master’s Student.

Gomel
yuevzhenko@gmail.com



References

1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30. https://doi.org/10.48550/arXiv.1706.03762.

2. Kurochka K.S., Youzhanka Y.D. The technology of translating natural language merchandising rules into digital planograms. Informatics. 2025;22(4):55−64 (in Russian). http://dx.doi.org/10.37661/1816-0301-2025-22-4-55-64.

3. Kaplan J., McCandlish S., Henighan T., Brown T.B., Chess B., Child R., et al. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. 2020. https://doi.org/10.48550/arXiv.2001.08361.

4. Minaee S., Mikolov T., Nikzad N., Chenaghlu M., Socher R., Amatriain X., et al. Large Language Models: A survey. arXiv preprint arXiv:2402.06196. 2024. https://doi.org/10.48550/arXiv.2402.06196.

5. Kurochka K.S., Basharymau Y.S. Neural network model for automated test generation for students in the Moodle system based on the analysis of methodological materials. Digital Transformation. 2025;31(3):66–75 (in Russian). http://dx.doi.org/10.35596/1729-7648-2025-31-3-66-75.

6. Shoeybi M., Patwary M., Puri R., Le Gresley P., Casper J., Catanzaro B. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053. 2019. https://doi.org/10.48550/arXiv.1909.08053.

7. Pope R., Douglas S., Chowdhery A., Devlin J., Bradbury J., Heek J., et al. Efficiently scaling transformer inference. Proceedings of Machine Learning and Systems (MLSys 2023). 2023;(5):606–624. Available at: https://proceedings.mlsys.org/paper_files/paper/2023/hash/c4be71ab8d24cdfb45e3d06dbfca2780-Abstract-mlsys2023.html (accessed 10 February 2026).

8. Song Y., Mi Z., Xie H., Chen H. Powerinfer: Fast Large Language Model serving with a consumer-grade GPU. Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP’24). 2024:590–606. https://doi.org/10.1145/3694715.3695964.

9. Huang Y., Cheng Y., Bapna A., Firat O., Chen D., Chen M., et al. GPipe: Efficient training of giant neural networks using pipeline parallelism. Advances in Neural Information Processing Systems (NeurIPS 2019). 2019;32. Available at: https://proceedings.neurips.cc/papers/search?q=GPipe%3A+Efficient+training+of+giant+neural+networks+using+pipeline+parallelism+ (accessed 10 February 2026).

10. Kwon W., Li Z., Zhuang S., Sheng Y., Zheng L., Yu C.H., et al. Efficient memory management for Large Language Model serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles (SOSP’23). 2023:611–626. https://doi.org/10.1145/3600006.3613165.

11. Liu A., Feng B., Wang B., Wang B., Liu B., Zhao C., et al. DeepSeek-V2: A strong, economical, and efficient Mixture-of-Experts language model. arXiv preprint arXiv:2405.04434. 2024. https://doi.org/10.48550/arXiv.2405.04434.

12. Lewis P., Perez E., PiktusA., Petroni F., KarpukhinV., Goyal N., et al. Retrieval-augmented generation for knowledgeintensive NLP tasks. Advances in Neural Information Processing Systems (NeurIPS 2020). 2020;33. Available at:https://proceedings.neurips.cc/papers/search?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks+ (accessed 10 February 2026).


Review

For citations:


Kurochka K.S., Basharymau Yu.S., Youzhanka Yu.D. Parallelism strategies as a key factor for deploying Large Language Models on consumer gpus. «System analysis and applied information science». 2026;(1):54-59. (In Russ.) https://doi.org/10.21122/2309-4923-2026-1-54-59

Views: 464

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2309-4923 (Print)
ISSN 2414-0481 (Online)