Loss Functions

The loss module provides the objective functions minimized during neural network training. Each loss wraps a BellmanPeriod and encodes a different optimality criterion: static reward maximization, estimated discounted lifetime reward, Bellman equation residuals, or Euler equation residuals.

class skagent.loss.BellmanEquationLoss(bellman_period, *, value_function, agent=None, foc_weight=0.0)

Creates a Bellman equation loss function for the Maliar method.

The Bellman equation is: V(s) = max_c { u(s,c,ε) + β E_ε’[V(s’)] } where s’ = f(s,c,ε) is the next state given current state s, control c, and shock ε, and the expectation E_ε’ is taken over future shock realizations ε’.

This function expects the input grid to contain two independent shock realizations: - {shock_sym}_0: shocks for period t (used for immediate reward and transitions) - {shock_sym}_1: shocks for period t+1 (used for continuation value evaluation)

Parameters:
  • bellman_period (BellmanPeriod) – The model block containing dynamics, rewards, and shocks

  • value_function (Union[dict[str, Callable], Callable]) – A value function that takes state variables and returns value estimates

  • parameters (dict, optional) – Model parameters for calibration

  • agent (str | None) – Agent identifier for rewards

  • foc_weight (float)

class skagent.loss.CustomLoss(loss_function, bellman_period, *, agent=None, other_dr=None)

A custom loss function that computes the negative reward for a block, assuming it is executed just once (a non-dynamic model)

Parameters:
  • loss_function (callable) – loss_function(bellman_period, dr, states, shocks=, parameters=, agent=) returning a per-sample reward, whose negative is the loss.

  • bellman_period (BellmanPeriod) – The period being trained. Its calibration is what the loss is evaluated at.

  • agent (str, optional) – Whose payoff to maximize. See StaticRewardLoss.

  • other_dr (dict of callable, optional) – Decision rules for the controls this loss is not training, held fixed.

class skagent.loss.EstimatedDiscountedLifetimeRewardLoss(bellman_period, *, big_t, agent=None)

A loss function for a Block that computes the discounted lifetime reward for T time periods.

Parameters:
  • bellman_period (BellmanPeriod) – The period being trained. Its calibration is what the loss is evaluated at.

  • big_t (int) – The number of time steps to compute reward for.

  • agent (str, optional) – Whose payoff to maximize: the sum of the reward symbols that agent owns, discounted over big_t periods. Required on a block whose utilities have more than one owner, since without it the loss maximizes the sum of every agent’s reward – a planner’s objective, and no player’s.

class skagent.loss.EulerEquationLoss(bellman_period, *, agent=None, weight=1.0, constrained=False)

Creates an Euler equation loss function for the Maliar method.

The Euler equation is the first-order condition from the Bellman equation, relating marginal rewards across periods. For a DSOP with control \(x_t\), arrival states \(s_t\), and pre-decision states \(m_t\), this loss function computes the Euler equation residual:

\[f = u'(x_t) + \beta \cdot u'(x_{t+1}) \cdot \sum_s \left[ \frac{\partial s_{t+1}}{\partial x_t} \cdot \frac{\partial m'}{\partial s_{t+1}} \right]\]

where \(f\) is the residual that equals zero at optimality, \(s_{t+1}\) is the next-period arrival state, and \(m'\) is the pre-decision state.

The discount factor \(\beta\) is obtained from the BellmanPeriod via bellman_period.discount_variable, so it adapts to the model’s calibration.

Multi-control support:

For models with \(J\) control variables, a separate Euler residual is computed per control. The loss sums over all controls: \(L = \sum_j w \cdot f_j^2\).

Handling Inequality Constraints (Fischer-Burmeister):

When constrained=True, a control’s declared bounds are turned into a smooth complementarity residual via the Fischer-Burmeister function (Maliar et al. 2021, equation 25), \(\text{FB}(a, b) = a + b - \sqrt{a^2 + b^2 + \varepsilon}\), which is zero (up to the regularizer \(\sqrt{\varepsilon}\)) exactly where \(a \geq 0\), \(b \geq 0\), \(a \cdot b = 0\).

The sign convention is that the Euler residual \(f\) is \(\geq 0\) when an upper bound binds and \(\leq 0\) when a lower bound binds. Writing \(s_u = \overline{x} - x\) and \(s_l = x - \underline{x}\) for the upper and lower slacks, the per-control residual is

  • upper bound only: \(\text{FB}(f, s_u)\);

  • lower bound only: \(\text{FB}(-f, s_l)\);

  • both bounds: \(\text{FB}(s_u, -\text{FB}(s_l, -f))\), a two-sided form that reduces to either one-sided residual when the opposite bound is slack, so a control with a non-binding floor and a binding ceiling matches the upper-only residual;

  • no bound: the one-sided fallback \(\text{relu}(-f)\), penalizing only violations of \(f \geq 0\).

Parameters:
  • bellman_period (BellmanPeriod) – The model block containing dynamics, rewards, and shocks.

  • parameters (dict, optional) – Model parameters for calibration.

  • agent (str | None) – Agent identifier for rewards.

  • weight (float) – Exogenous weight for combining multiple optimality conditions (default: 1.0). This corresponds to the vector \(v\) in equation (12) of the paper.

  • constrained (bool) – If True, use Fischer-Burmeister or one-sided loss for upper-bound constrained controls (default: False).

Examples

>>> bp = BellmanPeriod(block, "beta", calibration={"R": 1.04, "beta": 0.95})
>>> loss_fn = EulerEquationLoss(bp, parameters={"R": 1.04, "beta": 0.95})
class skagent.loss.StaticRewardLoss(bellman_period, *, agent=None, other_dr=None)

A loss function that computes the negative reward for a block, assuming it is executed just once (a non-dynamic model)

Parameters:
  • bellman_period (BellmanPeriod) – The period whose reward is maximized.

  • other_dr (dict of callable, optional) – Decision rules for the controls this loss is not training, held fixed.

  • agent (str, optional) – Whose payoff to maximize: the sum of the reward symbols that agent owns. Required on a block whose utilities have more than one owner, since without it the loss maximizes the sum of every agent’s reward – a planner’s objective, and no player’s. Use skagent.block.Block.deciding_agent() to read it off the control being trained.

skagent.loss.static_reward(bellman_period, dr, states, shocks=None, parameters=None, agent=None)

Returns the reward for an agent for a block, given a decision rule, states, shocks, and calibration.

The reward is the SUM of the reward symbols in scope: an agent owning several reward variables has a payoff of all of them, not of one.

Parameters:
  • bellman_period (BellmanPeriod) – The Bellman period object containing the model.

  • dr (dict or callable) – Decision rules (dict of functions), or a decision function.

  • states (dict) – Initial states, symbols to values.

  • shocks (dict, optional) – Shock variable values.

  • parameters (dict, optional) – Calibration parameters.

  • agent (str or None, optional) – Name of reference agent for rewards. When omitted, every reward symbol in the block is summed, which is one agent’s payoff only if the block has one agent; on a block whose utilities have several owners it is the sum of their payoffs and no agent’s objective.

Raises:

ValueError – If any reward symbol in scope is NaN.