Home // International Journal On Advances in Software, volume 18, numbers 3 and 4, 2025 // View article


Integer Sequence Refinement Using Language Models and Reinforcement Learning

Authors:
Carolina Carvalho
Paulo Quaresma

Keywords: Heuristic Optimization; Reinforcement Learning; Language Model; Task Semantic Segmentation; Artificial Neural Network; Neural Architecture Search; Unordered Markov Decision Processes; Bellman Operator.

Abstract:
This article extends the applicability domain of language models to problems where candidate solutions can be expressed as an encoded integer sequence. Considering this sequence, language models can operate in the neural machine translation setting and leverage their optimization power for heuristic search techniques. Reinforcement Learning (RL) is ap- plied to Language Models (LM), regardless of whether character-level or word-level models are used as a basis. To stabilize the learning, several approaches are explored, including functional and architectural decoupling. The framework is then applied to two combinatorial problems, namely the Traveling Salesman Problem benchmark and Neural Architecture Search, which is used to generate a hierarchical (tree-based) text classifier where the blocks are inspired by the InceptionV1 architecture. The decoupling results are the main contribution of this paper, easing the RL and LM stabilization requirements while expanding the resolution domain beyond Markov Decision Processes to non-causal normative heuristic problems, such as Neural Architecture Search (NAS).

Pages: 215 to 229

Copyright: Copyright (c) to authors, 2025. Used with permission.

Publication date: December 30, 2025

Published in: journal

ISSN: 1942-2628