Diese Bachelorarbeit befasst sich mit der Optimierung autonomer Wendemanöver für Agrarroboter
(Arawex) in Steillagen. Da 2D-Planer hier durch Kipprisiko versagen, evaluiert die Arbeit zwei Ansätze
in einer Digital-Twin-Architektur: einen analytischen Reeds-Shepp-Planer mit 3D-Kostenfunktion und
einen Deep-Reinforcement-Learning-Agenten (PPO) in NVIDIA Isaac Sim. Ein quantitativer Vergleich
entfiel, da der analytische Planer nicht vollendet wurde und die RL-Agenten wegen strenger Zeitlimits in
Evaluationen scheiterten. Die Untersuchung liefert dennoch einen zentralen Befund: Beide RL-Modelle
scheiterten architekturübergreifend an der simultanen Ausrichtung von Position und Orientierung (Heading)
am Ziel. Dichte Belohnungssignale führten zur reinen Annäherung, das spärlich belohnte End-
Heading schlug jedoch fehl – zusätzlich erschwert durch die Ackermann-Kinematik. Empfohlen wird
abschliessend eine hybride Architektur aus KI-Adaptivität und analytischen Sicherheitsgarantien.
This bachelor’s thesis examines the optimisation of autonomous turning manoeuvres for agricultural
robots (Arawex) on steep slopes. As 2D planners fail in this context due to the risk of tipping, the thesis
evaluates two approaches within a digital twin architecture: an analytical Reeds-Shepp planner with a
3D cost function and a deep reinforcement learning agent (PPO) in NVIDIA Isaac Sim. A quantitative
comparison was not carried out, as the analytical planner was not completed and the RL agents failed
in evaluations due to strict time limits. Nevertheless, the study yields a key finding: both RL models
failed across architectures when attempting to simultaneously align position and orientation (heading)
with the target. Dense reward signals led to mere approximation, but the sparsely rewarded end-heading
failed – a challenge further compounded by Ackermann kinematics. The conclusion recommends a hybrid
architecture combining AI adaptivity with analytical safety guarantees.
Diese Bachelorarbeit befasst sich mit der Optimierung autonomer Wendemanöver für Agrarroboter
(Arawex) in Steillagen. Da 2D-Planer hier durch Kipprisiko versagen, evaluiert die Arbeit zwei Ansätze
in einer Digital-Twin-Architektur: einen analytischen Reeds-Shepp-Planer mit 3D-Kostenfunktion und
einen Deep-Reinforcement-Learning-Agenten (PPO) in NVIDIA Isaac Sim. Ein quantitativer Vergleich
entfiel, da der analytische Planer nicht vollendet wurde und die RL-Agenten wegen strenger Zeitlimits in
Evaluationen scheiterten. Die Untersuchung liefert dennoch einen zentralen Befund: Beide RL-Modelle
scheiterten architekturübergreifend an der simultanen Ausrichtung von Position und Orientierung (Heading)
am Ziel. Dichte Belohnungssignale führten zur reinen Annäherung, das spärlich belohnte End-
Heading schlug jedoch fehl – zusätzlich erschwert durch die Ackermann-Kinematik. Empfohlen wird
abschliessend eine hybride Architektur aus KI-Adaptivität und analytischen Sicherheitsgarantien.
This bachelor’s thesis examines the optimisation of autonomous turning manoeuvres for agricultural
robots (Arawex) on steep slopes. As 2D planners fail in this context due to the risk of tipping, the thesis
evaluates two approaches within a digital twin architecture: an analytical Reeds-Shepp planner with a
3D cost function and a deep reinforcement learning agent (PPO) in NVIDIA Isaac Sim. A quantitative
comparison was not carried out, as the analytical planner was not completed and the RL agents failed
in evaluations due to strict time limits. Nevertheless, the study yields a key finding: both RL models
failed across architectures when attempting to simultaneously align position and orientation (heading)
with the target. Dense reward signals led to mere approximation, but the sparsely rewarded end-heading
failed – a challenge further compounded by Ackermann kinematics. The conclusion recommends a hybrid
architecture combining AI adaptivity with analytical safety guarantees.