AI Plays Pokémon
Creator
Training reinforcement learning agents for Pokémon Red and Crystal, with progress measured through verified gameplay milestones.

Stack
- Python
- PyBoy
- Recurrent PPO
- Reinforcement Learning
The Work
AI Plays Pokémon trains separate recurrent PPO agents for Pokémon Red and Crystal. Each agent uses emulator screenshots and structured game state to choose actions, with feedback from verified story milestones, navigation, battles, catches, and team survival under Nuzlocke rules.
I train across multiple emulator workers and all three starters, preserving model and optimizer checkpoints between runs. Frozen evaluations across every starter on fixed development seeds track verified progress, blackouts, and stalls; promising checkpoints face held-out tests against a fixed reference, with controller assistance reported separately.
What I Did
- Train separate recurrent PPO agents for Red and Crystal using screenshots and structured game state.
- Run multiple PyBoy workers across all three starters and preserve model and optimizer checkpoints.
- Reward verified story progress, navigation, battles, catches, and Nuzlocke team survival.
- Evaluate frozen checkpoints on fixed and held-out seeds, tracking gameplay results separately from training volume and controller assistance.