Parallelism strategies as a key factor for deploying Large Language Models on consumer gpus
https://doi.org/10.21122/2309-4923-2026-1-54-59
Abstract
The exponential growth in the size of Large Language Models (LLMs) creates significant barriers to their local deployment, primarily due to Video RAM (VRAM) shortages on single devices. The aim of this work is to identify and substantiate the most effective parallelism strategy for LLM inference on consumer Graphics Processing Unit (GPU) clusters connected via a slow PCIe bus. Research methods included a series of experiments comparing a monolithic architecture (NVIDIA RTX A6000) and a distributed system (2x NVIDIA RTX 3090) using the vLLM framework. The impact of Tensor Parallelism (TP) and Pipeline Parallelism (PP) on key metrics – throughput, latency (TTFT, TPOT), and power consumption stability – was analyzed while running the DeepSeek-R1-Distill-Llama-14B model. The results unequivocally indicate the unsuitability of Tensor Parallelism for systems without NVLink due to critical synchronization delays. It is proven that Pipeline Parallelism is the only viable strategy for PCIe clusters, ensuring high throughput despite the presence of idle periods («bubbles») and a less stable power consumption profile compared to the monolithic solution. In conclusion, recommendations for using multi-GPU configurations are formulated: they represent the optimal economic choice for memory-critical tasks, such as Retrieval-Augmented Generation (RAG), allowing VRAM scaling at a significantly lower cost than professional analogs.
Keywords
About the Authors
K. S. KurochkaBelarus
Konstantin S. Kurochka – PhD in Engineering, Associate Professor.
Gomel
kurochka@gstu.by
Yu. S. Basharymau
Belarus
Yury S. Basharymau – Assistant.
Gomel
basharymauyury@gmail.com
Yu. D. Youzhanka
Belarus
Yury D. Youzhanka – Master’s Student.
Gomel
yuevzhenko@gmail.com
References
1. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30. https://doi.org/10.48550/arXiv.1706.03762.
2. Kurochka K.S., Youzhanka Y.D. The technology of translating natural language merchandising rules into digital planograms. Informatics. 2025;22(4):55−64 (in Russian). http://dx.doi.org/10.37661/1816-0301-2025-22-4-55-64.
3. Kaplan J., McCandlish S., Henighan T., Brown T.B., Chess B., Child R., et al. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. 2020. https://doi.org/10.48550/arXiv.2001.08361.
4. Minaee S., Mikolov T., Nikzad N., Chenaghlu M., Socher R., Amatriain X., et al. Large Language Models: A survey. arXiv preprint arXiv:2402.06196. 2024. https://doi.org/10.48550/arXiv.2402.06196.
5. Kurochka K.S., Basharymau Y.S. Neural network model for automated test generation for students in the Moodle system based on the analysis of methodological materials. Digital Transformation. 2025;31(3):66–75 (in Russian). http://dx.doi.org/10.35596/1729-7648-2025-31-3-66-75.
6. Shoeybi M., Patwary M., Puri R., Le Gresley P., Casper J., Catanzaro B. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053. 2019. https://doi.org/10.48550/arXiv.1909.08053.
7. Pope R., Douglas S., Chowdhery A., Devlin J., Bradbury J., Heek J., et al. Efficiently scaling transformer inference. Proceedings of Machine Learning and Systems (MLSys 2023). 2023;(5):606–624. Available at: https://proceedings.mlsys.org/paper_files/paper/2023/hash/c4be71ab8d24cdfb45e3d06dbfca2780-Abstract-mlsys2023.html (accessed 10 February 2026).
8. Song Y., Mi Z., Xie H., Chen H. Powerinfer: Fast Large Language Model serving with a consumer-grade GPU. Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP’24). 2024:590–606. https://doi.org/10.1145/3694715.3695964.
9. Huang Y., Cheng Y., Bapna A., Firat O., Chen D., Chen M., et al. GPipe: Efficient training of giant neural networks using pipeline parallelism. Advances in Neural Information Processing Systems (NeurIPS 2019). 2019;32. Available at: https://proceedings.neurips.cc/papers/search?q=GPipe%3A+Efficient+training+of+giant+neural+networks+using+pipeline+parallelism+ (accessed 10 February 2026).
10. Kwon W., Li Z., Zhuang S., Sheng Y., Zheng L., Yu C.H., et al. Efficient memory management for Large Language Model serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles (SOSP’23). 2023:611–626. https://doi.org/10.1145/3600006.3613165.
11. Liu A., Feng B., Wang B., Wang B., Liu B., Zhao C., et al. DeepSeek-V2: A strong, economical, and efficient Mixture-of-Experts language model. arXiv preprint arXiv:2405.04434. 2024. https://doi.org/10.48550/arXiv.2405.04434.
12. Lewis P., Perez E., PiktusA., Petroni F., KarpukhinV., Goyal N., et al. Retrieval-augmented generation for knowledgeintensive NLP tasks. Advances in Neural Information Processing Systems (NeurIPS 2020). 2020;33. Available at:https://proceedings.neurips.cc/papers/search?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks+ (accessed 10 February 2026).
Review
For citations:
Kurochka K.S., Basharymau Yu.S., Youzhanka Yu.D. Parallelism strategies as a key factor for deploying Large Language Models on consumer gpus. «System analysis and applied information science». 2026;(1):54-59. (In Russ.) https://doi.org/10.21122/2309-4923-2026-1-54-59
JATS XML





















