Diese Arbeit simuliert eine Flugbahnverfolgung mit mehreren Reinforcement Learning Methoden. Das zu verfolgendes Objekt sendet ein Signal, das von der Antenne empfangen wird. Anhand der empfangenen Signalleistung wird beurteilt, wie gut die Antenne dem Objekt folgt. Die Arbeit beschreibt die mathematische Modellierung der Umwelt sowie die Bewegung von Objekt und Antenne. Im Weiteren werden Explorationsstrategien, die Abschätzung der Signalleistung zu einem nächsten Zeitpunkt sowie das Neuronalen Netzes erläutert. Diese Komponenten bilden die Grundlage der betrachteten Verfahren. Die Algorithmen erhalten zu jedem Zeitpunkt einen Zustand und eine Belohnung, aus denen eine optimale Aktion zur Antennensteuerung abgeleitet wird. Die Methoden SARSA, Experience Replay und Deep Deterministic Policy Gradient werden implementiert, die Bearbeitung der Kommunikation mit der Umgebung optimiert und hinsichtlich ihrer Leistungsfähigkeit zur Flugbahnverfolgung verglichen. Von allen untersuchten zeigt die Experience Replay Methode die erfolgreichste Verfolgung hinsichtlich der Differenz von den Azimut- und Elevationswinkel zwischen der Antenne und dem Objekt.
This work simulates trajectory tracking using several reinforcement learning methods. The object to be tracked emits a signal that is received by an antenna. Based on the received signal power, it is assessed how well the antenna follows the object. The work describes the mathematical modeling of the environment as well as the motion of the object and the antenna. Furthermore, exploration strategies, the estimation of the signal power at a future time step, and the neural network are explained. These components form the basis of the considered approaches. At each time step, the algorithms receive a state and a reward, from which an optimal action for antenna control is derived. The methods SARSA, Experience Replay, and Deep Deterministic Policy Gradient are implemented, the handling of communication with the environment is optimized, and their performance in trajectory tracking is compared. Amongst the evaluated approaches, the Experience Replay method demonstrates the most successful tracking in terms of the difference between the azimuth and elevation angles of the antenna and the object.
Flugbahnverfolgung mit mechanisch ausrichtbarer Antenne
Beschreibung
Diese Arbeit simuliert eine Flugbahnverfolgung mit mehreren Reinforcement Learning Methoden. Das zu verfolgendes Objekt sendet ein Signal, das von der Antenne empfangen wird. Anhand der empfangenen Signalleistung wird beurteilt, wie gut die Antenne dem Objekt folgt. Die Arbeit beschreibt die mathematische Modellierung der Umwelt sowie die Bewegung von Objekt und Antenne. Im Weiteren werden Explorationsstrategien, die Abschätzung der Signalleistung zu einem nächsten Zeitpunkt sowie das Neuronalen Netzes erläutert. Diese Komponenten bilden die Grundlage der betrachteten Verfahren. Die Algorithmen erhalten zu jedem Zeitpunkt einen Zustand und eine Belohnung, aus denen eine optimale Aktion zur Antennensteuerung abgeleitet wird. Die Methoden SARSA, Experience Replay und Deep Deterministic Policy Gradient werden implementiert, die Bearbeitung der Kommunikation mit der Umgebung optimiert und hinsichtlich ihrer Leistungsfähigkeit zur Flugbahnverfolgung verglichen. Von allen untersuchten zeigt die Experience Replay Methode die erfolgreichste Verfolgung hinsichtlich der Differenz von den Azimut- und Elevationswinkel zwischen der Antenne und dem Objekt.
This work simulates trajectory tracking using several reinforcement learning methods. The object to be tracked emits a signal that is received by an antenna. Based on the received signal power, it is assessed how well the antenna follows the object. The work describes the mathematical modeling of the environment as well as the motion of the object and the antenna. Furthermore, exploration strategies, the estimation of the signal power at a future time step, and the neural network are explained. These components form the basis of the considered approaches. At each time step, the algorithms receive a state and a reward, from which an optimal action for antenna control is derived. The methods SARSA, Experience Replay, and Deep Deterministic Policy Gradient are implemented, the handling of communication with the environment is optimized, and their performance in trajectory tracking is compared. Amongst the evaluated approaches, the Experience Replay method demonstrates the most successful tracking in terms of the difference between the azimuth and elevation angles of the antenna and the object.