I’m exploring how an agent learns to choose targets in the MicroRTS strategy game. The project uses PPO, custom observation and action wrappers, and SHAP to inspect its decisions. Supervised by Dr Piotr Duch.
The question
How can an agent learn which target to choose as a game changes? I’m investigating this in MicroRTS, a real-time strategy game, with Dr Piotr Duch. The project is about decisions in a simulated game environment.
Training and interpretation
The implementation uses PPO with custom observation and action wrappers. Reward shaping guides training, while game results provide a separate view of performance. I’m also using SHAP to inspect the agent’s decisions. The research is ongoing and the repository is private.
Following one decision
The game state first passes through observation wrappers so the policy receives the representation used for training. PPO produces a decision, and the action wrapper translates it into a targeting action in MicroRTS. The next state and reward return from the game, closing the training loop.
What I’m still investigating
Shaped rewards can help an agent learn, but I also need to look at the actual game outcome. I keep those signals separate when interpreting progress. SHAP gives another view of which inputs influence a decision. This work is still ongoing, so I’m sharing the approach without presenting it as a finished benchmark result.