<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">sapi</journal-id><journal-title-group><journal-title xml:lang="ru">Системный анализ и прикладная информатика</journal-title><trans-title-group xml:lang="en"><trans-title>«System analysis and applied information science»</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">2309-4923</issn><issn pub-type="epub">2414-0481</issn><publisher><publisher-name>Belarusian National Technical University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.21122/2309-4923-2024-4-13-20</article-id><article-id custom-type="elpub" pub-id-type="custom">sapi-701</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>Системный анализ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>System analysis</subject></subj-group></article-categories><title-group><article-title>Усложнение посредством постепенного вовлечения и предоставления вознаграждения в глубоком обучении с подкреплением</article-title><trans-title-group xml:lang="en"><trans-title>Complexification through gradual involvement and reward Providing in deep reinforcement learning</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Рулько</surname><given-names>Е. B.</given-names></name><name name-style="western" xml:lang="en"><surname>Rulko,</surname><given-names>E. V.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Рулько Евгений Викторович, кандидат технических наук, доцент, начальник научно-исследовательской лаборатории моделирования военных действий г. Минск</p></bio><bio xml:lang="en"><p>Eugene Rulko, РhD, associate professor in computer science. The head of the research laboratory of military operation simulation&#13;
Minsk</p></bio><email xlink:type="simple">eugeni1533@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Военная академия Республики Беларусь</institution><country>Беларусь</country></aff><aff xml:lang="en"><institution>Military academy of the Republic of Belarus</institution><country>Belarus</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2024</year></pub-date><pub-date pub-type="epub"><day>26</day><month>12</month><year>2024</year></pub-date><volume>0</volume><issue>4</issue><fpage>13</fpage><lpage>20</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Рулько Е.B., 2024</copyright-statement><copyright-year>2024</copyright-year><copyright-holder xml:lang="ru">Рулько Е.B.</copyright-holder><copyright-holder xml:lang="en">Rulko, E.V.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://sapi.bntu.by/jour/article/view/701">https://sapi.bntu.by/jour/article/view/701</self-uri><abstract><p>Тренировка нейронной сети, в рамках задач обучения с подкреплением, имеющей достаточную вычислительную емкость для решения сложных задач достаточно проблематична. В реальной жизни процесс решения задач требует системы знаний, где процесс изучения более сложных навыков основывается на использовании уже имеющихся. Аналогично, в ходе биологической эволюции, новые формы жизни базируются на достигнутом на предыдущем этапе уровне структурной сложности. Используя данные идеи, в настоящей работе предложены способы увеличения сложности архитектуры нейронных сетей, в частности способ тренировки сети с меньшем рецептивным полем и использованием натренированных весов в качестве отправной точки для более сложных сетей через постепенное вовлечение некоторых частей, а также способ предполагающий использование более простой сети с целью предоставления вознаграждения для более сложной. Это позволяет получить лучшую производительность в конкретном описанном примере, использующем Q-обучение, по сравнению со сценариями, когда сеть пытается использовать больший вектор входной информации с нуля.</p></abstract><trans-abstract xml:lang="en"><p>Training a relatively big neural network within the framework of deep reinforcement learning that has enough capacity for complex tasks is challenging. In real life the process of task solving requires system of knowledge, where more complex skills are built upon previously learned ones. The same way biological evolution builds new forms of life based on a previously achieved level of complexity. Inspired by that, this work proposes ways of increasing complexity, especially a way of training neural networks with smaller receptive fields and using their weights as prior knowledge for more complex successors through gradual involvement of some parts, and a way where a smaller network works as a source of reward for a more complicated one. That allows better performance in a particular case of deep Q-learning in comparison with a situation when the model tries to use a complex receptive field from scratch.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>глубокое обучение с подкреплением</kwd><kwd>Q-обучение</kwd><kwd>обучение по куррикулумому</kwd><kwd>дистилляционная модель</kwd><kwd>формирование вознаграждения в обучение с подкреплением</kwd></kwd-group><kwd-group xml:lang="en"><kwd>deep reinforcement learning</kwd><kwd>Q-learning</kwd><kwd>curriculum learning</kwd><kwd>distillation model</kwd><kwd>reward</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Zhuangdi Zhu et al. Transfer Learning in Deep Reinforcement Learning: A Survey. 2023. arXiv: 2009.07888.</mixed-citation><mixed-citation xml:lang="en">Zhuangdi Zhu et al. Transfer Learning in Deep Reinforcement Learning: A Survey. 2023. arXiv: 2009.07888.</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Petru Soviany et al. Curriculum Learning: A Survey. 2022. arXiv: 2101.10382.</mixed-citation><mixed-citation xml:lang="en">Petru Soviany et al. Curriculum Learning: A Survey. 2022. arXiv: 2101.10382.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Vassil Atanassov et al. Curriculum-Based Rein-forcement Learning for Quadrupedal Jumping: A Reference-free Design. 2024. arXiv: 2401.16337.</mixed-citation><mixed-citation xml:lang="en">Vassil Atanassov et al. Curriculum-Based Rein-forcement Learning for Quadrupedal Jumping: A Reference-free Design. 2024. arXiv: 2401.16337.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Yash J. Patel et al. Curriculum reinforcement learning for quantum architecture search under hardware errors. 2024. arXiv: 2402.03500.</mixed-citation><mixed-citation xml:lang="en">Yash J. Patel et al. Curriculum reinforcement learning for quantum architecture search under hardware errors. 2024. arXiv: 2402.03500.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">David Hoeller et al. ANYmal Parkour: Learning Agile Navigation for Quadrupedal Robots. 2023. arXiv: 2306.14874.</mixed-citation><mixed-citation xml:lang="en">David Hoeller et al. ANYmal Parkour: Learning Agile Navigation for Quadrupedal Robots. 2023. arXiv: 2306.14874.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Ken Caluwaerts et al. Barkour: Benchmarking Animal-level Agility with Quadruped Robots. 2023. arXiv: 2305.14654.</mixed-citation><mixed-citation xml:lang="en">Ken Caluwaerts et al. Barkour: Benchmarking Animal-level Agility with Quadruped Robots. 2023. arXiv: 2305.14654.</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Andrei A. Rusu et al. Progressive Neural Networks. 2022. arXiv: 1606.04671.</mixed-citation><mixed-citation xml:lang="en">Andrei A. Rusu et al. Progressive Neural Networks. 2022. arXiv: 1606.04671.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Enric Boix-Adsera. Towards a theory of model distillation. 2024. arXiv: 2403.09053.</mixed-citation><mixed-citation xml:lang="en">Enric Boix-Adsera. Towards a theory of model distillation. 2024. arXiv: 2403.09053.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Timo Kaufmann et al. A Survey of Reinforcement Learning from Human Feedback. 2024. arXiv: 2312. 14925 [cs.LG]. URL: https://arxiv.org/abs/2312.14925..</mixed-citation><mixed-citation xml:lang="en">Timo Kaufmann et al. A Survey of Reinforcement Learning from Human Feedback. 2024. arXiv: 2312. 14925 [cs.LG]. URL: https://arxiv.org/abs/2312.14925.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">E. Rulko. Complexification Through Gradual Involvement in Deep Reinforcement Learning. https://github.com/Eugene1533/snake-aipytorch-complexification. 2024.</mixed-citation><mixed-citation xml:lang="en">E. Rulko. Complexification Through Gradual Involvement in Deep Reinforcement Learning. https://github.com/Eugene1533/snake-aipytorch-complexification. 2024.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">P. Loeber. Reinforcement Learning With PyTorch and Pygame. https : / / github . com / patrickloeber/snake-aipytorch.2021.</mixed-citation><mixed-citation xml:lang="en">P. Loeber. Reinforcement Learning With PyTorch and Pygame. https : / / github . com / patrickloeber/snake-aipytorch.2021.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
