# Can AI Really Predict Stock Returns? An original review of *Empirical Asset Pricing via Machine Learning* by Shihao Gu, Bryan Kelly, and Dacheng Xiu. This episode is educational commentary, not investment advice. It does not reproduce the paper's prose, tables, charts, or figures. Exact source and license links appear on the episode page. ## 1. The question Can artificial intelligence predict stock returns? That question sounds like an invitation to hype. It can summon images of a machine seeing tomorrow's prices, beating every human trader, and turning uncertainty into a smooth upward line. The paper in today's episode asks a narrower and much better question. If researchers take decades of U.S. stock data, give several machine-learning models the information that investors could have observed at the time, and then test the models on later years, do the flexible models forecast returns better than conventional statistical methods? The paper is *Empirical Asset Pricing via Machine Learning*, by Shihao Gu, Bryan Kelly, and Dacheng Xiu. It appeared in *The Review of Financial Studies* in 2020. It has become influential because it brings modern prediction methods into one of finance's most difficult settings: estimating the expected return of an individual stock. The answer is interesting because it is neither “AI solves the market” nor “machine learning makes no difference.” The authors report a real historical improvement. But the statistical improvement is small, the portfolio implications are much larger, and the journey from those two facts to a commercial investment product is full of friction. So our real question is not simply whether the model predicts. It is: what exactly was predicted, how honestly was it tested, and what survives after we ask the questions a portfolio manager, risk officer, regulator, or paying customer would ask? ## 2. Why it matters now The commercial language of AI in finance often runs ahead of the evidence. Investment products are described as intelligent, adaptive, or powered by machine learning. A label can be attached to an old strategy without changing much beneath it. This paper gives us a more useful standard. First, define the target. Second, separate training from validation and testing. Third, compare the new model with serious baselines. Fourth, report not only a statistical score but also what happens in portfolios. Finally, expose the uncomfortable details: turnover, drawdowns, small-stock dependence, and sensitivity to research choices. This matters because stock returns are extraordinarily noisy. A company can execute well and still fall because interest rates change. A weak company can rally because expectations were worse. News, liquidity, positioning, taxes, and risk appetite collide in the same monthly return. In that setting, an out-of-sample R-squared of less than one percent can sound trivial. It means the forecast removes only a sliver of squared prediction error compared with a simple benchmark. Yet when thousands of small forecasts are ranked and combined, a sliver may change portfolio construction. That creates a dangerous communication gap. A researcher may correctly say, “The model improved prediction by four-tenths of one percent.” A marketer may turn the associated portfolio result into “AI finds alpha.” Both sentences can refer to the same experiment while giving a listener radically different expectations. For education, the paper is valuable because it teaches us to keep three layers separate: predictive accuracy, historical portfolio performance, and a live commercial product. They are connected, but they are not interchangeable. ## 3. What came before Traditional empirical asset pricing often begins with a regression. Researchers choose a small set of characteristics—perhaps company size, valuation, or recent price momentum—and estimate how those variables relate to future returns. Linear models are attractive because they are transparent and disciplined. But they impose a shape on the world. They may assume that a one-unit change has a similar effect across companies, or that two variables matter separately when their interaction is the real signal. Machine learning offers a larger toolbox. Penalized regressions can shrink noisy coefficients. Principal-components and partial-least-squares methods can compress many variables. Random forests and boosted trees can capture thresholds and interactions. Neural networks can build nonlinear combinations through layers of learned features. Flexibility is both the promise and the hazard. A model with enough freedom can fit patterns that happened once and will never repeat. Finance makes that hazard worse because researchers can try many variables, time periods, portfolio rules, and hyperparameters. A beautiful backtest may be a map of the researcher's search process rather than a durable feature of markets. The authors therefore put unusual emphasis on prediction outside the data used for fitting. They are not the first to apply machine learning to markets, and their paper does not test every possible method. Its importance comes from the breadth of the comparison, the long sample, the common evaluation framework, and the attempt to connect a small forecasting gain to economically interpretable portfolios. ## 4. What the researchers did The data cover nearly thirty thousand U.S. stocks between 1957 and 2016, with more than six thousand two hundred firms in an average month. The sample draws on CRSP and includes firms listed on the New York Stock Exchange, AMEX, and Nasdaq. For each stock, the researchers assemble ninety-four characteristics. These describe ideas such as momentum, liquidity, volatility, valuation, profitability, investment, and trading activity. They combine those firm-level variables with eight macroeconomic predictors and seventy-four industry indicators. Once interactions are represented, the model has nine hundred twenty covariates. The time split is the heart of the design. The first training period runs from 1957 through 1974. Validation uses 1975 through 1986 to choose model settings. The final historical test covers 1987 through 2016. As time advances, the training window expands, and the models are refitted annually. That is stronger than fitting the entire sixty-year sample and reporting how well the model explains the same history. The later test years were held outside the initial estimation and model-selection periods. Still, “out of sample” needs careful language. This was not a robot trading live from 1987 while the researchers watched. It was a historical experiment designed to respect the flow of time. Its credibility depends on using only information that would have been available, constructing the dataset consistently, and not allowing later research choices to leak knowledge backward. The model lineup includes ordinary least squares using all predictors; a simple three-variable regression; penalized linear models; principal components; partial least squares; random forests; gradient-boosted trees; and neural networks with one through five hidden layers. The authors judge the models in several ways. They measure forecast error for individual stocks. They aggregate stock forecasts into a market-level prediction. They sort stocks into portfolios based on predicted returns. And they inspect which families of characteristics contribute most. This is not one magical neural network facing a weak straw man. It is a structured tournament among model families, with a common historical test. ## 5. What they found Start with the least cinematic result: monthly returns of individual stocks remain very hard to predict. The full ordinary-least-squares model has a negative out-of-sample R-squared of minus 3.46 percent in the reported comparison. That means it performs worse than the benchmark. A small three-variable regression reaches positive 0.16 percent. Principal components and partial least squares reach 0.26 and 0.27 percent. Random forests and boosted trees reach 0.33 and 0.34 percent. The best result in that table is the neural network with three hidden layers: 0.40 percent. Not forty percent. Four-tenths of one percent. That small number is the statistically honest headline. The flexible models improve the forecast, but most variation in next month's individual-stock return remains unexplained. When the stock forecasts are aggregated into a bottom-up forecast of the S&P 500, the three-layer neural network reaches a reported monthly out-of-sample R-squared of 1.80 percent. A market-timing exercise based on the forecast has a historical Sharpe ratio of 0.77, compared with 0.51 for buy and hold in the paper's setup. The portfolio sorts are more dramatic. Each month, the researchers rank stocks by predicted return and compare high-prediction with low-prediction groups. For a value-weighted neural-network long-short portfolio, one reported Sharpe ratio is 1.35. Equal weighting produces 2.45. Because equal weighting gives more influence to small firms, the authors repeat the analysis after excluding stocks below the twentieth percentile of New York Stock Exchange market capitalization. The equal-weighted neural result falls, but remains high at a reported 1.69. These are gross historical statistics. They are not a promise, and they are not automatically the return experience of a fund customer. ## 6. What was genuinely new One contribution is scale: many stock characteristics, macro variables, industry information, model families, and decades are evaluated in one framework. A second contribution is the finding that nonlinear interactions matter. Momentum and short-term reversal variables are important. So are liquidity, volatility, and valuation. But the model's value is not merely rediscovering one famous factor. It can allow the relevance of one characteristic to change with another characteristic or with the broader economic state. A third lesson is delightfully anti-hype: deeper is not automatically better. The neural network with three hidden layers performs best in the main stock-level comparison. Adding fourth and fifth layers does not improve the headline result. In a noisy, structured dataset with limited effective history, greater architectural complexity can add estimation error instead of insight. That lesson has commercial value beyond finance. The best deployed model is not necessarily the largest model. It is the model whose added complexity earns its cost under a test that resembles the real decision. The paper also helps change the culture of asset-pricing research. It treats validation, hyperparameter selection, model comparison, and out-of-sample evaluation as central design choices. Those practices are routine in machine learning, but applying them carefully to financial panels forces economists and data scientists to speak a more common language. What it does not deliver is an economic theory of why the predictors work. A neural network can map conditions to forecasts without telling us whether the return premium compensates investors for risk, reflects behavioral mistakes, captures market frictions, or has already begun to disappear. Prediction is evidence. It is not mechanism. ## 7. Commercial meaning Suppose an asset manager wants to turn this research into a product. The first challenge is turnover. The paper reports monthly turnover around one hundred ten to one hundred thirty percent for neural-network portfolios. Roughly speaking, the portfolio is replacing a very large share of its positions each month. Trading costs, bid-ask spreads, market impact, financing, short availability, and the cost of borrowing hard-to-short stocks can consume a signal that looks powerful before implementation. The second challenge is capacity. A result can survive excluding the tiniest stocks and still rely on securities that cannot absorb billions of dollars without prices moving. A small research portfolio and a scalable fund are different engineering problems. The third challenge is risk. The value-weighted four-layer neural portfolio reports a maximum drawdown of 51.78 percent and a worst one-month loss of 33.03 percent. A high historical Sharpe ratio does not mean a smooth ride. Customers may redeem, lenders may tighten terms, or a risk committee may shut down a strategy before a long-run average has time to recover. The fourth challenge is decay. Once a signal becomes known and capital pursues it, the return may shrink. Data definitions change. Market structure changes. The population of firms changes. A model must be monitored for drift, not merely retrained on schedule. The fifth challenge is governance. Which data were licensed? Can every feature be reproduced at the decision time? Who approves a model change? How is an unusual exposure explained? What happens when the prediction engine conflicts with portfolio constraints? So the commercial opportunity is broader than selling a black box. There is value in data quality, feature timing, cost-aware optimization, execution, monitoring, explainability, and risk controls. The research signal may be only one component of the product. The most credible marketing claim would therefore be modest: machine learning can improve the organization of complex historical information and may support better forecasts. The least credible claim would be that AI has discovered a reliable machine for effortless excess returns. ## 8. Limits and a later challenge The first limitation is external validity. The evidence comes from one large historical U.S. dataset. Other countries, later periods, different market regimes, and live execution can behave differently. The second is researcher freedom. Even with a clean time split, decisions about missing data, normalization, model tuning, portfolio construction, and evaluation can influence the conclusion. A later paper by Matias Cattaneo, Yingjie Feng, and William Underwood examines the role of those research-design choices in machine-learning asset pricing. Its abstract reports that design choices can generate large variation, with nonstandard errors as much as five times conventional standard errors. It also reports that about one-third of strategies remain statistically significant after transaction costs. That does not erase the original paper. It strengthens the right interpretation. The original study is an important demonstration and benchmark, not the final answer to whether a specific strategy will work after costs. The third limitation is that portfolio statistics amplify choices. Equal weighting, value weighting, the long and short legs, rebalancing frequency, and the treatment of microcaps can move the result. Whenever a paper travels from prediction score to investment return, we should inspect every bridge between them. The fourth is explanation. The authors themselves are clear that successful prediction does not identify the underlying mechanism or equilibrium. A model can reveal structure without telling us why markets permit it. The fifth is disclosure. The manuscript lists Bryan Kelly's affiliation with AQR Capital Management, and the NBER disclosure record reports consulting income from AQR above five thousand dollars. That context should be visible. It is not evidence that the findings are wrong; it is information a reader deserves when evaluating commercially relevant research. ## 9. A skeptical reading checklist When you encounter the next claim that AI predicts markets, ask ten questions. What was the exact prediction target? Was the test genuinely later in time than training and model selection? What simple baseline did the model beat? Was the improvement statistical, economic, or both? Are the portfolio results gross or net of realistic costs? How much turnover and market impact are implied? Does the result depend on tiny or difficult-to-trade securities? What is the worst drawdown, not only the average Sharpe ratio? How sensitive is the finding to data construction and researcher choices? And what evidence would cause the team to stop or change the model after launch? This checklist does not reject machine learning. It gives machine learning the respect of a serious test. The paper's most durable message may be methodological. In a domain flooded with stories, separate the training period from the test. Compare against disciplined baselines. Report the small predictive number before the large portfolio number. Then price the entire system required to turn a forecast into a product. ## 10. Source card The paper is *Empirical Asset Pricing via Machine Learning*, by Shihao Gu, Bryan Kelly, and Dacheng Xiu, published in *The Review of Financial Studies*, volume 33, issue 5, in 2020. Its digital object identifier is 10.1093/rfs/hhaa009. The episode page links to the final Oxford University Press record, an author-hosted manuscript, the NBER working-paper record, and the recorded license. It also links the later research-design paper used for the robustness perspective. This episode is an independent, original explanation and critique. It does not reproduce the article's prose, charts, tables, or figures. The source links are there so you can inspect the evidence, the version, and the disclosures for yourself. The balanced conclusion is simple. Machine learning produced a small, credible historical improvement in stock-return prediction and striking gross portfolio results. That makes the paper important. Turnover, drawdowns, costs, capacity, model decay, and research-design sensitivity make it commercially demanding. AI may help predict returns. It does not abolish uncertainty, implementation, or judgment.