Skip to main content

Yusuke Narita Publications

Discussion Paper
Abstract

In many real recommender systems, novel items are added frequently over time. The importance of sufficiently presenting novel actions has widely been acknowledged for improving long-term user engagement. A recent work builds on Off-Policy Learning (OPL), which trains a policy from only logged data, however, the existing methods can be unsafe in the presence of novel actions. Our goal is to develop a framework to enforce exploration of novel actions with a guarantee for safety. To this end, we first develop Safe Off-Policy Policy Gradient (Safe OPG), which is a model-free safe OPL method based on a high confidence off-policy evaluation. In our first experiment, we observe that Safe OPG almost always satisfies a safety requirement, even when existing methods violate it greatly. However, the result also reveals that Safe OPG tends to be too conservative, suggesting a difficult tradeoff between guaranteeing safety and exploring novel actions. To overcome this tradeoff, we also propose a novel framework called Deployment-Efficient Policy Learning for Safe User Exploration, which leverages safety margin and gradually relaxes safety regularization during multiple (not many) deployments. Our framework thus enables exploration of novel actions while guaranteeing safe implementation of recommender systems.

Discussion Paper
Abstract

Automated decision-making algorithms drive applications in domains such as recommendation systems and search engines. These algorithms often rely on off-policy contextual bandits or off-policy learning (OPL). Conventionally, OPL selects actions that maximize the expected reward within an existing action set. However, in many real-world scenarios, actions—such as news articles or video content—change continuously, and the action space evolves over time compared to when the logged data was collected. We define actions introduced after deploying the logging policy as new actions and focus on the problem of OPL with new actions. Existing OPL methods cannot learn and select new actions because no relevant data are logged. To address this limitation, we propose a new OPL method that leverages action features. In particular, we first introduce the Local Combination PseudoInverse (LCPI) estimator for the policy gradient, generalizing the PseudoInverse estimator initially proposed for off-policy evaluation of slate bandits. LCPI controls the trade-off between reward-modeling condition and the condition for data collection regarding the action features, capturing the interaction effects among different dimensions of action features. Furthermore, we propose a generalized algorithm called Policy Optimization for Effective New Actions (PONA), which integrates LCPI, a component specialized for new action selection, with Doubly Robust (DR), which excels at learning within existing actions. We define PONA as a weighted sum of the LCPI and DR estimators, optimizing both the selection of existing and new actions, and allowing the proportion of new action selections to be adjusted by controlling the weight parameter.

Discussion Paper
Abstract

Algorithms make a growing portion of policy and business decisions. We develop a treatment-effect estimator using algorithmic decisions as instruments for a class of stochastic and deterministic algorithms. Our estimator is consistent and asymptotically normal for well-defined causal effects. A special case of our setup is multidimensional regression discontinuity designs with complex boundaries. We apply our estimator to evaluate the Coronavirus Aid, Relief, and Economic Security Act, which allocated many billions of dollars worth of relief funding to hospitals via an algorithmic rule. The funding is shown to have little effect on COVID-19-related hospital activities. Naive estimates exhibit selection bias.

Discussion Paper
Abstract

What happens if selective colleges change their admission policies? We study this question by analyzing the world’s first implementation of nationally centralized meritocratic admissions in the early twentieth century. We find a persistent meritocracy-equity tradeoff. Compared to the decentralized system, the centralized system admitted more high-achievers and produced more occupational elites (such as top income earners) decades later in the labor market. This gain came at a distributional cost, however. Meritocratic centralization also increased the number of urban-born elites relative to rural-born ones, undermining equal access to higher education and career advancement.

Discussion Paper
Abstract

Algorithms produce a growing portion of decisions and recommendations both in policy and business. Such algorithmic decisions are natural experiments (conditionally quasirandomly assigned instruments) since the algorithms make decisions based only on observable input variables. We use this observation to develop a treatment-effect estimator for a class of stochastic and deterministic algorithms. Our estimator is shown to be consistent and asymptotically normal for well-defined causal effects. A key special case of our estimator is a high-dimensional regression discontinuity design. The proofs use tools from differential geometry and geometric measure theory, which may be of independent interest. 

 

The practical performance of our method is first demonstrated in a high-dimensional simulation resembling decision-making by machine learning algorithms. Our estimator has smaller mean squared errors compared to alternative estimators. We finally apply our estimator to evaluate the effect of Coronavirus Aid, Relief, and Economic Security (CARES) Act, where more than $10 billion worth of relief funding is allocated to hospitals via an algorithmic rule. The estimates suggest that the relief funding has little effect on COVID- 19-related hospital activity levels. Naive OLS and IV estimates exhibit substantial selection bias.

Discussion Paper
Abstract

Countries with more democratic political regimes experienced greater GDP loss and more deaths from COVID-19 in 2020. Using five diffferent instrumental variable strategies, we find that democracy is a major cause of the wealth and health losses. This impact is global and is not driven by China and the US alone. A key channel for democracy’s negative impact is weaker and narrower containment policies at the beginning of the outbreak, not the speed of introducing policies.

Discussion Paper
Abstract

Democracy is widely believed to contribute to economic growth and public health. However, we find that this conventional wisdom is no longer true and even reversed; democracy has persistent negative impacts on GDP growth since the beginning of this century. This finding emerges from five different instrumental variable strategies. Our analysis suggests that democracies cause slower growth through less investment, less trade, and slower value-added growth in manufacturing and services. For 2020, democracy is also found to cause more deaths from Covid-19.

Discussion Paper
Abstract

Centralized school assignment algorithms must distinguish between applicants with the same preferences and priorities. This is done with randomly assigned lottery numbers, nonlottery tie-breakers like test scores, or both. The New York City public high school match illustrates the latter, using test scores, grades, and interviews to rank applicants to screened schools, combined with lottery tie-breaking at unscreened schools. We show how to identify causal effects of school attendance in such settings. Our approach generalizes regression discontinuity designs to allow for multiple treatments and multiple running variables, some of which are randomly assigned. Lotteries generate assignment risk at screened as well as unscreened schools. Centralized assignment also identifies screened school effects away from screened school cutoffs. These features of centralized assignment are used to assess the predictive value of New York City’s school report cards. Grade A schools improve SAT math scores and increase the likelihood of graduating, though by less than OLS estimates suggest. Selection bias in OLS estimates is egregious for Grade A screened schools.

Discussion Paper
Abstract

What is the most statistically efficient way to do off-policy optimization with batch data from bandit feedback? For log data generated by contextual bandit algorithms, we consider offline estimators for the expected reward from a counterfactual policy. Our estimators are shown to have lowest variance in a wide class of estimators, achieving variance reduction relative to standard estimators. We then apply our estimators to improve advertisement design by a major advertisement company. Consistent with the theoretical result, our estimators allow us to improve on the existing bandit algorithm with more statistical confidence compared to a state-of-theart benchmark.

Abstract

Many centralized school admissions systems use lotteries to ration limited seats at oversubscribed schools. The resulting random assignment is used by empirical researchers to identify the effect of entering a school on outcomes like test scores. I first find that the two most popular empirical research designs may not successfully extract a random assignment of applicants to schools. When do the research designs overcome this problem? I show the following main results for a class of data-generating mechanisms containing those used in practice: One research design extracts a random assignment under a mechanism if and practically only if the mechanism is strategy-proof for schools. In contrast, the other research design does not necessarily extract a random assignment under any mechanism.

Abstract

In centralized school admissions systems, rationing at oversubscribed schools often uses lotteries in addition to preferences. This partly random assignment is used by empirical researchers to identify the effect of entering a school on outcomes like test scores. This paper formally studies if the two most popular empirical research designs successfully extract a random assignment. For a class of data-generating mechanisms containing those used in practice, I show: One research design extracts a random assignment under a mechanism if and almost only if the mechanism is strategy-proof for schools. In contrast, the other research design does not necessarily extract a random assignment under any mechanism.