RL-weapon-targeting

Eldar Mukhtarov

Learning decisions in MicroRTS · Ongoing research

I’m exploring how an agent learns to choose targets in the MicroRTS strategy game. The project uses PPO, custom observation and action wrappers, and SHAP to inspect its decisions. Supervised by Dr Piotr Duch.

The agent acts in MicroRTS, receives feedback and updates its policy during training.
The agent acts in MicroRTS, receives feedback and updates its policy during training.
Open the diagram in a new tab

The question

How can an agent learn which target to choose as a game changes? I’m investigating this in MicroRTS, a real-time strategy game, with Dr Piotr Duch. The project is about decisions in a simulated game environment.

Training and interpretation

The implementation uses PPO with custom observation and action wrappers. Reward shaping guides training, while game results provide a separate view of performance. I’m also using SHAP to inspect the agent’s decisions. The research is ongoing and the repository is private.

Following one decision

The game state first passes through observation wrappers so the policy receives the representation used for training. PPO produces a decision, and the action wrapper translates it into a targeting action in MicroRTS. The next state and reward return from the game, closing the training loop.

What I’m still investigating

Shaped rewards can help an agent learn, but I also need to look at the actual game outcome. I keep those signals separate when interpreting progress. SHAP gives another view of which inputs influence a decision. This work is still ongoing, so I’m sharing the approach without presenting it as a finished benchmark result.

← All projects