« Back

How accurately can machine learning algorithms predict a person’s future?

Author(s): Emily Cantrell, Pranay Anchuri, Hanzhang Ren, Matthew Salganik

Wednesday 14  |   15:00-15:20

Room: TP43

Session: Behind the scenes of digital public services

Applications of artificial intelligence (AI) are rapidly increasing in public policy. One such application is the use of machine learning algorithms in public services to make predictions about individual people’s futures. For example, several U.S. child protective services agencies use algorithms to predict the likelihood that a child will be removed from the home due to maltreatment. Predictive algorithms are also being used or piloted in criminal justice, homelessness prevention, medical services, school dropout prevention, and more. To understand the appropriateness of using such algorithms, we need to understand how accurately a person’s future can be predicted.

We investigate the predictability of hundreds of life outcomes for youth and their families, and identify patterns in what types of outcomes are most (and least) predictable. Specifically, we use in-depth survey and observational data collected from a child’s birth to age 9 in a U.S. longitudinal study to predict approximately 500 outcomes at the child’s age 15. This dataset includes thousands of variables about children and their families in a wide range of life domains. We test dozens of machine learning approaches for every outcome to ensure a reasonable estimate of the best possible predictive accuracy that can be achieved with these data. Our main finding is that most outcomes in our set are not very predictable. For approximately half of the outcomes, our algorithms do not make predictions any better than a null model that predicts everyone’s outcome to be equal to the mean of the training data. However, some outcomes are much more predictable than others, and we identify patterns in what types of outcomes are most predictable.

These findings call for caution about predictive accuracy when implementing predictive algorithms in public services. However, they also raise questions about how sample size and other characteristics of data affect predictive accuracy. In future work, we plan to explore questions about the capabilities and limitations of predictive algorithms with Dutch administrative register data, which has similar characteristics to the administrative register data of Nordic countries.

Original file: 1112.pdf