OBJECT RECORD / AIM-DOC-2016-001

Mastering the game of Go with deep neural networks and tree search

A Nature paper records a Go system that combines learned policy and value networks with tree search.

CURATORIAL READING

Why it is here

Silver and colleagues describe training policy networks first from expert games and then through self-play, learning a value network from self-play positions, and using both to guide Monte Carlo tree search.

Evidence boundary

Keep the claim inside this paper and the game of Go. Its reported five-game result against Fan Hui does not establish general intelligence, and this record does not cover the later Lee Sedol match or reproduce the paper’s figures.

10 / LEARN AND SEARCH

Position in the guided path

From written rules to learned representations

Policy and value networks guide tree search inside the game of Go.

Back to the object room →

PROVENANCE

Source trail