Temporal-Difference Learning for 2048

Temporal-Difference Learning for 2048 cover

A temporal-difference learning agent developed during an NCTU CGI Lab internship.

During a research internship at National Chiao Tung University’s Computer Games and Intelligence Lab, I trained a temporal-difference learning agent to play 2048. The agent learns a value function over board states and selects moves greedily with respect to that estimate.

The agent first reached the 2048 tile after 2,000 training rounds, and hit a 93% win rate after 20,000. I also investigated how changing the probability of spawning a 4-tile in the environment affects the final score and the learning speed.

Full C++ implementation on GitHub.

Some results

  • Temporal-Difference Learning for 2048: Some results 1
  • Temporal-Difference Learning for 2048: Some results 2