Minkey Chang

Data Scientist

All posts

Can we really get alpha from market data?

The academic answer has been no for decades: prices already contain public information, so predictable patterns should not survive. Competitions and funds keep trying anyway. This post puts the two views side by side, the efficient market hypothesis and the micro alphas argument, and then asks what it takes to turn a weak signal, if you have one, into a position.

EDA of market data: missing values, correlations, volatility, target distribution.
Exploratory analysis of daily market and macro features: missing values, correlations, volatility over time, and the distribution of the return target.

The efficient market view

If prices reflect all available information, any edge gets arbitraged away, and the efficient market hypothesis says exactly that. The empirical record on return prediction mostly agrees. Goyal and Welch (2008) took the popular predictors of the equity premium, dividend yield, earnings yield, the term spread and the rest, and showed that their out-of-sample forecasts were no better than the historical mean. What looked like predictive power in-sample was overfitting and look-ahead. The default position in academic finance is therefore that alpha from public market data is very hard, and that most of what looks like it is a mistake in the backtest.


Micro alphas

The counter-argument is not that one magic variable works. Hull and co-authors (2024) argue that many weak signals, none of them statistically convincing on its own, can add up to usable predictability if they are combined carefully, with feature selection and a regularised model such as an elastic net. The authors report running the approach in a live fund. The Kaggle Hull Tactical Market Prediction competition is built on this framework: a large set of daily features covering market, macro, volatility and sentiment, a target of S&P 500 excess return, and a position between 0 and 2 that has to respect a volatility constraint. The forecast is only useful once it has been turned into a position, which is the part the efficient market literature rarely tests.


What a weak signal looks like

It helps to have a number in mind. In a forecasting competition run by a hedge fund in 2026, with anonymised features, strict sequential rules and a host who warned that the signal-to-noise ratio was low, my validation score was a skill measure relative to predicting zero. Unwinding it, the model removed about 3% of the squared error of predicting nothing. That is what a micro alpha looks like from the inside: real, repeatable across seeds, and tiny. The public leaderboard, meanwhile, carried scores several times higher, and the host posted a notice that those almost certainly used forward-looking information. Leakage looks like alpha until someone checks, which is the efficient market hypothesis restated as a warning.


From signal to position

A forecast that explains 3% of variance cannot be traded on its point value alone. What matters is how confident the model is on a given day, because sizing has to shrink when uncertainty is high. That is the case for probabilistic forecasting. The Temporal Fusion Transformer (Lim et al., 2021), for example, outputs quantiles rather than a single number, so the spread between the 5th and 95th percentile can drive conservative sizing when it is wide and more aggressive sizing when it is narrow, inside a leverage and volatility budget. In a Hull Tactical style setup the pipeline is then: forecast the excess return with a distribution, read regime information off the same inputs, and map both into a position that stays within the risk limits. Getting a good signal is the smaller half of the problem.


References

  1. Equity premium prediction — Goyal, A., & Welch, I. (2008). A Comprehensive Look at the Empirical Performance of Equity Premium Prediction. Review of Financial Studies, 21(4), 1455–1508.

  2. Micro Alphas — Hull, B., Bakosova, P., Cocquemas, F., Sinclair, E., & Fast, P. (2024). Micro Alphas. Working Paper, Hull Tactical Asset Allocation. SSRN 5035294.

  3. Temporal Fusion Transformer — Lim, B., Arik, S. Ö., Loeff, N., & Pfister, T. (2021). Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting. International Journal of Forecasting, 37(4), 1748–1764.

  4. Hull Tactical Market Prediction — Kaggle competition: hull-tactical-market-prediction.

  5. Hedge fund time series forecasting — Kaggle community competition: ts-forecasting, 2026.

Other posts