Reinforcement learning for 2048
The world's highest-scoring algorithm and the fastest simulation engine for the 2048 puzzle game.
From the CV:
- The agent constructs a 2048-tile within 1.2 milliseconds in 99.98% of games, and even the 32768-tile in 75% of games.
- Trained a custom 200M-parameter RAMnet via self-play reinforcement learning on 150M games.
- Combined expectimax search with locally brute-forced critical phases to maximise the probability of survival.
- C++, Cambridge dissertation, September 2023 – May 2024.
[ write-up — to be written ]