Reinforcement learning for an intelligent and autonomous production control of complex job-shops under time constraints

Обучение с подкреплением для интеллектуального и автономного управления производством на сложных мелкосерийных производствах с учетом временных ограничений
Thomas Altenmüller, Tillmann Stüker, Bernd Waschneck, Andreas Kuhnle, Gisela Lanza
2020-06-01

Q-learningdiscrete-event simulationorder dispatchingproduction controlreinforcement learning
Abstract Reinforcement learning (RL) offers promising opportunities to handle the ever-increasing complexity in managing modern production systems. We apply a Q-learning algorithm in combination with a process-based discrete-event simulation in order to train a self-learning, intelligent, and autonomous agent for the decision problem of order dispatching in a complex job shop with strict time constraints. For the first time, we combine RL in production control with strict time constraints. The simulation represents the characteristics of complex job shops typically found in semiconductor manufacturing. A real-world use case from a wafer fab is addressed with a developed and implemented framework. The performance of an RL approach and benchmark heuristics are compared. It is shown that RL can be successfully applied to manage order dispatching in a complex environment including time constraints. An RL-agent with a gain function rewarding the selection of the least critical order with respect to time-constraints beats heuristic rules strictly by picking the most critical lot first. Hence, this work demonstrates that a self-learning agent can successfully manage time constraints with the agent performing better than the traditional benchmark, a time-constraint heuristic combining due date deviations and a classical first-in-first-out approach.
1
A Q-learning agent combined with process-based discrete-event simulation was trained for autonomous order dispatching in complex job shops.
2
An RL agent rewarding selection of the least time-critical order outperformed the heuristic rule that prioritizes the most critical lot first.
3
Reinforcement learning successfully managed order dispatching in the complex, time-constrained production environment.
4
The RL approach outperformed a benchmark heuristic combining due-date deviations with classical first-in-first-out dispatching.
5
The framework addresses strict time constraints in production control, demonstrated using semiconductor-manufacturing job-shop characteristics and a real wafer-fab use case.

complex job-shop production systems in semiconductor wafer fabrication with strict time constraints

autonomous reinforcement-learning-based order dispatching performance and decision-making under strict time constraints

Publication Details
Publication Date
2020-06-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Thomas Altenmüller
Tillmann Stüker
Bernd Waschneck
Andreas Kuhnle
Gisela Lanza
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%