Minkey Chang

Data Scientist

All posts

Can we make a more risk-aware portfolio agent from utility theory?

Portfolio reinforcement learning usually means a scalar reward each step, a discounted sum of those rewards as the value, and PPO or A2C to optimise it. That is easy to build, but it folds two different attitudes into one discount factor: how much the agent dislikes risk, and how willing it is to trade payoff today for payoff later. Asset pricing separates the two. Recursive utility in the Epstein–Zin form has one parameter for risk aversion and another for intertemporal substitution, and this post is about what happens when that objective replaces the discounted sum inside actor–critic training. The short version, from a paper I wrote for the ICLR 2026 FinAI workshop: on Korean ETF data it raised the Sharpe ratio and cut the drawdown against a discounted baseline under identical code, with two caveats that carry as much weight as the result.


What changes in the Bellman equation

With a discounted sum, the value of a state is a linear expectation of next-period value, and everything downstream, critic targets, advantages, GAE, relies on that linearity. Recursive utility replaces the expectation with a certainty equivalent, a power mean that penalises the bad tail of next-period value, and then aggregates it with current consumption:

CEt=(E[ V(st+1) 1−γ ∣ st])11−γV(st)=max⁡αt, κt[(1−β) Ct ρ+β CEt ρ]1/ρρ=1−1/ψ\begin{aligned} \mathrm{CE}_t &= \Big( \mathbb{E}\big[\, V(s_{t+1})^{\,1-\gamma} \,\big|\, s_t \big] \Big)^{\frac{1}{1-\gamma}} \\[6pt] V(s_t) &= \max_{\alpha_t,\,\kappa_t} \Big[ (1-\beta)\, C_t^{\,\rho} + \beta\, \mathrm{CE}_t^{\,\rho} \Big]^{1/\rho} \\[6pt] \rho &= 1 - 1/\psi \end{aligned}

Here γ\gamma is risk aversion, not the RL discount; β\beta is time preference, ψ\psi is the elasticity of intertemporal substitution, and CtC_t is current consumption. With γ>1\gamma > 1 the certainty equivalent sits below the plain expectation whenever next-period value is dispersed, so the agent is pushed away from allocations with a fat left tail. When γ=1/ψ\gamma = 1/\psi the two parameters collapse into one and you are back to a time-additive objective.

Two practical problems follow. The certainty equivalent has no closed form under real return distributions, and a trading agent does not consume, so CtC_t needs a meaning.


Making it trainable

State and wealth. The state is st=(wt,αt−1)s_t = (w_t, \alpha_{t-1}): log wealth and last period's portfolio weights. Weights live on the simplex, long only and summing to one. The policy outputs increments that are added to the previous weights and projected back, which trains more stably than predicting weights from scratch. Wealth follows the self-financing rule

wt+1=wt+log⁡(1+αt⊤Rt+1)w_{t+1} = w_t + \log\big(1 + \alpha_t^{\top} R_{t+1}\big)

Consumption as an accounting variable. I set Ct=κtWtC_t = \kappa_t W_t with κt\kappa_t a scalar fraction of wealth, learned or fixed. The agent does not spend it. The term exists so that the trade-off between current and future value, which is where the risk-aware behaviour comes from, stays in the objective for a machine that never literally consumes.

Monte Carlo certainty equivalent. Sample KK candidate next states from the return distribution, evaluate the critic on each, and take the empirical power mean:

CE^t=[1K∑k=1KVϕ(st+1(k))1−γ]11−γ\widehat{\mathrm{CE}}_t = \Big[ \tfrac{1}{K} \sum_{k=1}^{K} V_\phi\big(s^{(k)}_{t+1}\big)^{1-\gamma} \Big]^{\frac{1}{1-\gamma}}

It is consistent as KK grows, and it is noisy at any KK you can afford.

Target, critic, advantage. Plug the estimate into the aggregator to get a one-step target, train the critic by squared error to it, and use the Bellman residual as the advantage:

T^t=[(1−β) (κtewt)ρ+β CE^t ρ]1/ρAt=T^t−Vϕ(st)\begin{aligned} \widehat{T}_t &= \Big[ (1-\beta)\, \big(\kappa_t e^{w_t}\big)^{\rho} + \beta\, \widehat{\mathrm{CE}}_t^{\,\rho} \Big]^{1/\rho} \\[6pt] A_t &= \widehat{T}_t - V_\phi(s_t) \end{aligned}

Standard GAE assumes a linear recursion and does not apply here. The paper derives a multi-step version with state-dependent weights, and the runs below use it with λ=0.95\lambda = 0.95. Because VϕV_\phi appears inside the certainty equivalent and in the residual, only critic-based algorithms can use this objective; REINFORCE cannot. The actor is initialised from Campbell–Viceira approximate rules rather than from a random corner of the simplex.


What happened on Korean ETFs

Daily closing prices of 110 Korean ETFs, chosen for listing history and low missingness, split into ten chronological train/test pairs with the training share rising from 50% to 90%. Each split is one trial, and results are the mean and standard deviation over the ten, following the FinRL evaluation protocol. Three objectives ran in the same environment with the same code: a naive discounted return, a per-step Markowitz mean–variance reward, and the recursive target above with γ=5\gamma = 5 and K=10K = 10.

PPO objectiveSharpeMax drawdownCumulative return
Naive1.22 ± 1.0712.3%−6.5%
Markowitz1.43 ± 1.5813.7%−3.7%
Recursive2.07 ± 1.0410.4%8.2%

Volatility was similar across the three, so the Sharpe gain came from the shape of returns and the drawdown, which is what a left-tail-averse objective should buy. Two caveats carry equal weight. The spread across splits is wide and there was one trial per split, so the ranking is suggestive rather than established. And an equal-weight portfolio reached a Sharpe of 2.3 to 2.7 on the same splits, above every PPO run. The claim that survives is recursive versus naive under identical RL code, not RL beating a simple baseline. Ablations showed the results move with KK and with the training window.


What I would keep

Recursive utility is a legitimate way to put finance preferences into an RL objective, and on these splits it behaved the way the theory says it should. It also costs KK extra critic evaluations per step, a custom backup, and a restriction to critic-based methods. If the goal is a portfolio policy, start with a simple reward and a strong simple baseline. If the goal is to separate risk aversion from patience in the objective itself, this is a workable way to do it, and the ETF results are one data point for it rather than a verdict.


References

  1. This paper — Chang, M. (2026). Portfolio Optimization under Recursive Utility via Reinforcement Learning. ICLR 2026 FinAI Workshop, short paper. arXiv:2603.22880.

  2. Recursive utility — Epstein, L. G., & Zin, S. E. (1989). Substitution, Risk Aversion, and the Temporal Behavior of Consumption and Asset Returns: A Theoretical Framework. Econometrica, 57(4), 937–969.

  3. Strategic asset allocation — Campbell, J. Y., & Viceira, L. M. (2002). Strategic Asset Allocation: Portfolio Choice for Long-Term Investors. Oxford University Press.

  4. RL fundamentals — Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.

  5. Actor–critic — Mnih, V., et al. (2016). Asynchronous Methods for Deep Reinforcement Learning. ICML.

  6. PPO — Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. Preprint.

  7. FinRL — Liu, X.-Y., et al. (2020). FinRL: A Deep Reinforcement Learning Library for Automated Stock Trading in Quantitative Finance. ACM International Conference on AI in Finance (ICAIF).

Other posts