Can we make a more risk-aware portfolio agent from utility theory?
Portfolio reinforcement learning usually means a scalar reward each step, a discounted sum of those rewards as the value, and PPO or A2C to optimise it. That is easy to build, but it folds two different attitudes into one discount factor: how much the agent dislikes risk, and how willing it is to trade payoff today for payoff later. Asset pricing separates the two. Recursive utility in the Epstein–Zin form has one parameter for risk aversion and another for intertemporal substitution, and this post is about what happens when that objective replaces the discounted sum inside actor–critic training. The short version, from a paper I wrote for the ICLR 2026 FinAI workshop: on Korean ETF data it raised the Sharpe ratio and cut the drawdown against a discounted baseline under identical code, with two caveats that carry as much weight as the result.
What changes in the Bellman equation
With a discounted sum, the value of a state is a linear expectation of next-period value, and everything downstream, critic targets, advantages, GAE, relies on that linearity. Recursive utility replaces the expectation with a certainty equivalent, a power mean that penalises the bad tail of next-period value, and then aggregates it with current consumption:
Here is risk aversion, not the RL discount; is time preference, is the elasticity of intertemporal substitution, and is current consumption. With the certainty equivalent sits below the plain expectation whenever next-period value is dispersed, so the agent is pushed away from allocations with a fat left tail. When the two parameters collapse into one and you are back to a time-additive objective.
Two practical problems follow. The certainty equivalent has no closed form under real return distributions, and a trading agent does not consume, so needs a meaning.
Making it trainable
State and wealth. The state is : log wealth and last period's portfolio weights. Weights live on the simplex, long only and summing to one. The policy outputs increments that are added to the previous weights and projected back, which trains more stably than predicting weights from scratch. Wealth follows the self-financing rule
Consumption as an accounting variable. I set with a scalar fraction of wealth, learned or fixed. The agent does not spend it. The term exists so that the trade-off between current and future value, which is where the risk-aware behaviour comes from, stays in the objective for a machine that never literally consumes.
Monte Carlo certainty equivalent. Sample candidate next states from the return distribution, evaluate the critic on each, and take the empirical power mean:
It is consistent as grows, and it is noisy at any you can afford.
Target, critic, advantage. Plug the estimate into the aggregator to get a one-step target, train the critic by squared error to it, and use the Bellman residual as the advantage:
Standard GAE assumes a linear recursion and does not apply here. The paper derives a multi-step version with state-dependent weights, and the runs below use it with . Because appears inside the certainty equivalent and in the residual, only critic-based algorithms can use this objective; REINFORCE cannot. The actor is initialised from Campbell–Viceira approximate rules rather than from a random corner of the simplex.
What happened on Korean ETFs
Daily closing prices of 110 Korean ETFs, chosen for listing history and low missingness, split into ten chronological train/test pairs with the training share rising from 50% to 90%. Each split is one trial, and results are the mean and standard deviation over the ten, following the FinRL evaluation protocol. Three objectives ran in the same environment with the same code: a naive discounted return, a per-step Markowitz mean–variance reward, and the recursive target above with and .
| PPO objective | Sharpe | Max drawdown | Cumulative return |
|---|---|---|---|
| Naive | 1.22 ± 1.07 | 12.3% | −6.5% |
| Markowitz | 1.43 ± 1.58 | 13.7% | −3.7% |
| Recursive | 2.07 ± 1.04 | 10.4% | 8.2% |
Volatility was similar across the three, so the Sharpe gain came from the shape of returns and the drawdown, which is what a left-tail-averse objective should buy. Two caveats carry equal weight. The spread across splits is wide and there was one trial per split, so the ranking is suggestive rather than established. And an equal-weight portfolio reached a Sharpe of 2.3 to 2.7 on the same splits, above every PPO run. The claim that survives is recursive versus naive under identical RL code, not RL beating a simple baseline. Ablations showed the results move with and with the training window.
What I would keep
Recursive utility is a legitimate way to put finance preferences into an RL objective, and on these splits it behaved the way the theory says it should. It also costs extra critic evaluations per step, a custom backup, and a restriction to critic-based methods. If the goal is a portfolio policy, start with a simple reward and a strong simple baseline. If the goal is to separate risk aversion from patience in the objective itself, this is a workable way to do it, and the ETF results are one data point for it rather than a verdict.
References
-
This paper — Chang, M. (2026). Portfolio Optimization under Recursive Utility via Reinforcement Learning. ICLR 2026 FinAI Workshop, short paper. arXiv:2603.22880.
-
Recursive utility — Epstein, L. G., & Zin, S. E. (1989). Substitution, Risk Aversion, and the Temporal Behavior of Consumption and Asset Returns: A Theoretical Framework. Econometrica, 57(4), 937–969.
-
Strategic asset allocation — Campbell, J. Y., & Viceira, L. M. (2002). Strategic Asset Allocation: Portfolio Choice for Long-Term Investors. Oxford University Press.
-
RL fundamentals — Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
-
Actor–critic — Mnih, V., et al. (2016). Asynchronous Methods for Deep Reinforcement Learning. ICML.
-
PPO — Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. Preprint.
-
FinRL — Liu, X.-Y., et al. (2020). FinRL: A Deep Reinforcement Learning Library for Automated Stock Trading in Quantitative Finance. ACM International Conference on AI in Finance (ICAIF).
Other posts
- Driving through underground rocks
Steering a drill bit through rock you cannot see: why a good particle filter still drifts, and how simulation fixed the drift.
- Can we really get alpha from market data?
The efficient market view, the micro alpha counter-argument, and why a weak signal only becomes a position once you know its uncertainty.
- What works for forecasting macro economic series with deep learning?
Korean output and investment nowcasting with seven deep models: what the data allows, which families worked, and why it depends on the target.
- Could multivariate time series have their own representations?
Why forecast embeddings are not factors, and how identifiable innovations with diagonal dynamics recover them without losing forecast quality.
- Classifying bird sounds in the field
Sound as a picture: what a spectrogram is, why the mel scale matches how we hear, and how a network finds a bird call in the image.
- Creating and Evaluating Synthetic Tabular Data
Sequential synthesis for credit bureau data, and three checks: pMSE distinguishability, confidence interval overlap, and attribute disclosure.