Models¶
The skagent.models subpackage contains predefined economic models. For a
narrative overview of the benchmark registry and how to call it, see the
Benchmark Models guide; for runnable, plotted walkthroughs
see the Examples.
Consumer Models¶
Benchmark Registry¶
The benchmark registry catalogues discrete-time dynamic programming problems with known closed-form policies (plus a few that are kept for numerical validation). Its public helpers and analytical-policy functions are listed below.
Analytically Solvable Consumption-Savings Models
This module implements a collection of discrete-time consumption-savings dynamic programming problems for which the literature has succeeded in writing down true closed-form policies. These represent well-known benchmark problems from the economic literature with established analytical solutions.
THEORETICAL FOUNDATION¶
An entry qualifies for inclusion ONLY if: (i) The problem is a bona-fide dynamic programming problem (ii) The optimal c_t (and any other control) can be written in closed form with no recursive objects left implicit
Standard Timing Convention (Adopted Throughout)¶
t ∈ {0,1,2,…} : period index A_{t-1} : beginning-of-period assets (arrival state, before interest) y_t : non-capital income (realized in period t) R : gross return on assets (R = 1 + r > 1) m_t = R*A_{t-1} + y_t : cash-on-hand (market resources available for consumption) c_t : consumption (control variable) A_t = m_t - c_t : end-of-period assets (state for next period) H_t = E_t[∑_{s=1}^∞ R^{-s} y_{t+s}] : human wealth (present value of future income) W_t = m_t + H_t : total wealth (cash-on-hand plus human wealth) u(c) : period utility function β : discount factor TVC : lim_{T→∞} E_0[β^T u’(c_T) A_T] = 0 (transversality condition)
Registry access¶
- skagent.models.benchmarks.list_benchmark_models()¶
List all discrete-time benchmark models.
Note: Most models have analytical solutions, but some (e.g., U-3 buffer stock) require numerical solution due to borrowing constraints + income uncertainty.
- skagent.models.benchmarks.get_benchmark_model(model_id)¶
Get benchmark model by ID (D-1 through D-4, U-1 through U-3).
Returns an independent deep copy of the registered block. The registered blocks are module-level singletons whose shock specs are mutated in place by
construct_shocks(the(class, args)tuples are replaced by distribution objects). Returning a copy keeps that mutation local to the caller, so constructing or recalibrating one block never leaks into another caller (or another test) holding “the same” model.
- skagent.models.benchmarks.get_benchmark_calibration(model_id)¶
Get benchmark calibration by model ID
- skagent.models.benchmarks.get_analytical_policy(model_id)¶
Get analytical policy function by model ID.
Raises ValueError if the model does not have an analytical policy (e.g., U-3 buffer stock model requires numerical solution).
- skagent.models.benchmarks.has_analytical_policy(model_id)¶
Return whether the registry exposes a closed-form policy for
model_id.Most benchmark entries pair their
DBlockwith an analytical decision function; a few (for exampleU-3, the buffer-stock model) are registered without one because a borrowing constraint combined with income uncertainty leaves them with no closed form. This predicate lets callers branch on that distinction without catching theValueErrorthatget_analytical_policy()raises.- Parameters:
model_id (
str) – Registry key, e.g."D-1"or"U-3".- Returns:
Trueif a closed-form policy is available,Falseotherwise.- Return type:
- Raises:
ValueError – If
model_idis not a registered benchmark model.
- skagent.models.benchmarks.get_reference_policy(model_id)¶
Get a numerical reference policy for a model lacking a closed form.
Some benchmarks (currently D-4) have no analytical policy because a binding constraint precludes one, but do have an independent numerical oracle (e.g. value-function iteration) suitable for validating trained policies.
Raises ValueError if the model has no reference policy registered.
- skagent.models.benchmarks.get_test_states(model_id, test_points=10)¶
Get test states for model validation by model ID
Validation¶
- skagent.models.benchmarks.validate_analytical_solution(model_id, test_points=10, tolerance=1e-08)¶
Validate analytical solution satisfies optimality conditions and budget constraints
- skagent.models.benchmarks.get_analytical_lifetime_reward(model_id, *args, **kwargs)¶
Get analytical lifetime reward for a benchmark model.
Analytical policies¶
- skagent.models.benchmarks.d1_analytical_policy(states, shocks, parameters)¶
Optimal policy for D-1: finite-horizon log-utility consumption.
With log utility and a deterministic \(T\)-period horizon, the agent solves
\[V_T(W_T) = \log W_T, \qquad V_t(W_t) = \max_{c_t \in (0,\, W_t]} \, \log c_t + \beta \, V_{t+1}\bigl((W_t - c_t)\, R\bigr),\]where \(W_t\) is wealth at the start of period \(t\), \(R\) is the gross return, and \(\beta < 1\) is the discount factor. The value function takes the form \(V_t(W) = \alpha_t + \log W\) for a time-varying additive constant \(\alpha_t\), and the first-order condition gives the remaining-horizon rule
\[c_t \;=\; \frac{1 - \beta}{\,1 - \beta^{T - t}\,} \, W_t.\]In the terminal period (\(T - t = 1\)) the formula simplifies to \(c_t = W_t\) since \((1-\beta)/(1-\beta) = 1\); the implementation handles this case directly to avoid the \(0/0\) form that would arise once \(T - t = 0\). As \(T - t \to \infty\), the rule converges to the infinite-horizon constant-MPC policy \(c_t = (1 - \beta) W_t\), the \(\sigma = 1\) special case of D-2.
- Parameters:
- Returns:
{"c": c_optimal}whose dtype matches the input wealth.- Return type:
- Raises:
ValueError – If \(\beta \geq 1\), since the closed form is undefined.
- skagent.models.benchmarks.d2_analytical_policy(states, shocks, parameters)¶
Optimal policy for D-2: infinite-horizon CRRA perfect foresight.
The canonical perfect-foresight consumption problem solves
\[\max_{\{c_t\}} \, \sum_{t=0}^{\infty} \beta^t \, \frac{c_t^{\,1-\sigma}}{1 - \sigma} \quad \text{s.t.} \quad A_t = R\, A_{t-1} + y - c_t,\]with constant labor income \(y > 0\), gross return \(R > 1\), and the transversality condition \(\lim_{T\to\infty} \beta^T \, u'(c_T)\, A_T = 0\). Equivalently, cash-on-hand \(m_t = R\, A_{t-1} + y\) evolves with end-of-period assets \(A_t = m_t - c_t\). Under return-impatience \((\beta R)^{1/\sigma} < R\), the closed-form policy is linear in total wealth,
\[c_t \;=\; \kappa \, W_t, \qquad \kappa \;=\; \frac{R - (\beta R)^{1/\sigma}}{R}, \qquad W_t \;=\; m_t + H,\]where \(m_t = R\, A_{t-1} + y\) is cash-on-hand at time \(t\), \(H = y / r\) is human wealth (the present value of the constant future income stream), and \(r = R - 1\). The marginal propensity to consume \(\kappa\) depends only on deep parameters, not on the state.
- Parameters:
- Returns:
{"c": c_optimal}.- Return type:
- Raises:
ValueError – If return-impatience is violated, or if \(R \leq 1\).
See also
d3_analytical_policySame model with i.i.d. survival risk.
- skagent.models.benchmarks.d3_analytical_policy(states, shocks, parameters)¶
Optimal policy for D-3: Blanchard (1985) discrete-time mortality.
Extends D-2 by giving the agent i.i.d. survival probability \(s \in (0, 1)\) per period. The objective becomes
\[\max \, \sum_{t=0}^{\infty} (s\beta)^t \, \frac{c_t^{\,1-\sigma}}{1 - \sigma},\]with the same budget constraint as D-2. Mortality is observationally equivalent to scaling the discount factor from \(\beta\) to \(s\beta\), so the linear consumption rule survives:
\[c_t \;=\; \kappa_s \, (m_t + H), \qquad \kappa_s \;=\; \frac{R - (s\beta R)^{1/\sigma}}{R},\]with \(H = y / r\) as before. The MPC \(\kappa_s > \kappa\) because mortality erodes effective patience: the agent consumes a larger share of wealth each period.
- Parameters:
states (dict) – Must contain
"a"(arrival assets). The"liv"alive indicator is part of the DBlock simulator dynamics but is not read by the analytical policy.shocks (dict) – May contain
"live"(Bernoulli survival shock); unused for the analytical policy itself.parameters (dict) – Must contain
"DiscFac","R","CRRA","y", and"SurvivalProb"(\(s\)).
- Returns:
{"c": c_optimal}.- Return type:
- Raises:
ValueError – If mortality-adjusted return-impatience is violated, or if \(R \leq 1\).
References
Blanchard, O.J. (1985). “Debt, deficits, and finite horizons.” Journal of Political Economy, 93(2), 223-247.
See also
d2_analytical_policyUnderlying perfect-foresight model without mortality.
- skagent.models.benchmarks.u1_analytical_policy(states, shocks, parameters)¶
Optimal policy for U-1: Hall (1978) random-walk consumption.
With quadratic utility \(u(c) = ac - bc^2/2\) and the neutral stochastic discount factor \(\beta R = 1\), the Euler equation collapses to the martingale property
\[\mathbb{E}_t[c_{t+1}] = c_t,\]so consumption follows a random walk regardless of the income process. Hall’s contribution was to derive this implication and confront it with consumption data.
The decision rule consistent with this Euler equation, plus transversality, is the Permanent Income Hypothesis: consume the annuity value of total wealth,
\[c_t \;=\; \frac{r}{R} \, (m_t + H),\]where \(m_t = R \, A_{t-1} + y_t\) is cash-on-hand, \(H = \mathbb{E}_t y / r\) is the present value of the expected future income stream, and \(r = R - 1\). The martingale property is a consequence of this PIH policy, not the policy itself.
- Parameters:
- Returns:
{"c": c_optimal}.- Return type:
Notes
Logs a warning via the module logger if \(|\beta R - 1| > 10^{-6}\), since the PIH derivation hinges on \(\beta R = 1\) exactly. With high income variance, transversality may also fail.
References
Hall, R.E. (1978). “Stochastic implications of the life cycle-permanent income hypothesis: Theory and evidence.” Journal of Political Economy, 86(6), 971-987.
- skagent.models.benchmarks.u2_analytical_policy(states, shocks, parameters)¶
Optimal policy for U-2: log utility with permanent income shocks (normalized).
The buffer-stock problem with log utility, geometric random-walk permanent income, and no borrowing constraint admits a closed-form policy in normalized variables. Dividing every level variable by permanent income \(P_t\) yields the lowercase ratios \(m = M/P\), \(c = C/P\), \(a = A/P\), with normalized transition
\[m_{t+1} = \frac{R}{\psi_{t+1}} \, a_t + 1,\]where \(\psi_{t+1}\) is the mean-one permanent income shock. The constant
+1represents normalized transitory income, which is identically one in U-2; the more general two-shock case (with a stochastic transitory component) is U-3. Under log utility the closed form is\[c_t \;=\; (1 - \beta) \, (m_t + h), \qquad h \;=\; 1/r,\]independent of the realized shock path. The MPC \((1 - \beta)\) is the limiting MPC for any unconstrained CRRA agent; log utility (the \(\sigma = 1\) case) makes it exact rather than asymptotic.
- Parameters:
- Returns:
{"c": c_optimal}(normalized consumption).- Return type:
Notes
Setting
"sigma_psi": 0makes \(\psi \equiv 1\), so the PIH analytical solution holds exactly. TheControlupper bound \(0.1\, m + 2\) is loose enough for the analytical policy \(c \approx 0.04\, m + 1.33\) but rules out Ponzi-scheme solutions that satisfy the Euler equation while violating transversality.See also
u3_blockSame problem with a binding borrowing constraint and CRRA utility, which has no closed form.
- skagent.models.benchmarks.crra_utility(c, gamma)¶
CRRA utility: u(c) = c^(1-gamma)/(1-gamma) for gamma != 1, log(c) for gamma == 1
Fisher Two-Period Model¶
Fisher (1930) two-period intertemporal consumption.
The simplest dynamic programming problem with a closed-form solution. An agent receives income \(y\) each period, borrows or saves at gross rate \(R\), and chooses consumption \(c_0, c_1\) to maximize
subject to the lifetime budget constraint \(c_0 + c_1/R = m_0 + y/R\), with \(m_0 = R\, a_{-1} + y\). With CRRA utility \(u(c) = c^{1-\sigma}/(1-\sigma)\), the Euler equation \(u'(c_0) = \beta R \, u'(c_1)\) together with the budget constraint gives the closed form
The two-period horizon makes the model an exact analogue of the
intertemporal-choice diagram in introductory macroeconomics, while the
recursive form is the simplest non-trivial test case for value-function
iteration and Euler-equation solvers in skagent.
The horizon is part of the model, not of the solver call: the block above is one period, and the two-period problem is that period iterated exactly twice against a terminal continuation. A solver that iterates it to a fixed point is answering a different question – the infinite-horizon one – and the closed form above is not its answer.
Notes
The math above uses \(R\) for the gross return; the block parameter
key is Rfree.
References
Fisher, I. (1930). The Theory of Interest. New York: Macmillan.
- skagent.models.fisher.T = 2¶
Number of periods. The closed form below is the period-0 rule under this horizon; the terminal rule is \(c_1 = m_1\).
- skagent.models.fisher.analytical_policy(states, shocks, parameters)¶
Optimal period-0 consumption for the two-period problem.
- Parameters:
- Returns:
{"c": c_0}whose dtype follows the input assets.- Return type:
- Raises:
ValueError – If
CRRA <= 0orRfree <= 0, for which the closed form is undefined.
Perfect Foresight Models¶
Perfect-foresight consumption-savings with stochastic survival and permanent income growth.
Block representation of the canonical perfect-foresight problem with i.i.d. survival probability \(s \in (0, 1)\) and gross permanent income growth \(G\). The agent solves
with permanent income growing as \(P_{t+1} = G\, P_t\). Two
conditions are needed for the closed form: mortality-adjusted
return-impatience \((s\beta R)^{1/\sigma} < R\) (so consumption does
not explode), and \(R > G\) (so human wealth is finite). Under both,
the consumption rule is linear in total wealth \(W_t = m_t + H_t\),
with human wealth \(H_t = G\, P_t / (R - G)\) and MPC
\(\kappa_s = (R - (s\beta R)^{1/\sigma})/R\). The companion module
skagent.models.perfect_foresight_normalized solves the same
problem in variables divided by \(P_t\), which collapses the state
space and is the form used in the buffer-stock literature.
Notes
The math above uses \(R\), \(G\), and \(s\) for the gross
return, permanent income growth, and survival probability; the
corresponding block parameter keys are Rfree, PermGroFac, and
LivPrb.
References
Carroll, C.D. (2024). Solution Methods for Solving Microeconomic Dynamic Stochastic Optimization Problems. https://llorracc.github.io/SolvingMicroDSOPs/
Perfect Foresight Models (Normalized)¶
Perfect-foresight consumption-savings in normalized variables.
The same problem as skagent.models.perfect_foresight, but with
every level variable divided by permanent income \(P_t\). Lowercase
ratios \(m = M/P\), \(c = C/P\), \(a = A/P\) evolve via the
effective return \(R_{\text{eff}} = R / G\),
with normalized income identically equal to one. Normalization reduces the state space from \((M, P)\) to the single ratio \(m\), which matters both for analytical tractability and for neural-network solvers, where the network learns a one-dimensional function \(c(m)\) instead of a two-dimensional \(c(M, P)\).
The closed-form normalized policy is \(c_t = \kappa_s\, (m_t + h)\) with \(\kappa_s = (R - (s\beta R)^{1/\sigma})/R\) and \(h = 1 / (R_{\text{eff}} - 1)\). Although the normalized Bellman discounts future utility by \(s\beta G^{1-\sigma}\) rather than \(s\beta\), the algebra collapses \((s\beta G^{1-\sigma} R_{\text{eff}})^{1/\sigma}\) to \((s\beta R)^{1/\sigma}/G\), so the level and normalized MPCs agree.
Notes
The math above uses \(R\), \(G\), and \(s\) for the gross
return, permanent income growth, and survival probability; the
corresponding block parameter keys are Rfree, PermGroFac, and
LivPrb.
References
Carroll, C.D. (2024). Solution Methods for Solving Microeconomic Dynamic Stochastic Optimization Problems. https://llorracc.github.io/SolvingMicroDSOPs/
Resource Extraction¶
Resource Extraction with Optimal Escapement (Reed 1979)¶
Resource extraction models analyze the optimal management of renewable or depletable resources over time. These models appear in:
Fisheries management: Determining sustainable harvest rates for fish populations Forestry: Optimal timber harvesting schedules Environmental economics: Managing renewable natural resources (water, wildlife) Energy: Optimal depletion of oil fields and mineral deposits Finance: Portfolio liquidation and asset drawdown strategies
The core problem involves balancing immediate extraction (profit now) against preserving the resource stock (profit later), accounting for natural growth dynamics and environmental uncertainty.
This implementation follows the model from:
Reed, W.J. (1979). “Optimal escapement levels in stochastic and deterministic harvesting models.” Journal of Environmental Economics and Management, 6(4), 350-363.
Reed showed that under multiplicative environmental shocks and stock-dependent harvesting costs, the optimal policy has a simple “constant escapement” form: maintain a target stock level S* and harvest any surplus above it. This optimal escapement level S* can be computed analytically without solving the full dynamic programming problem, making it an excellent benchmark for testing reinforcement learning algorithms.
Mathematical Model¶
State: \(x_t\) = resource stock at time \(t\)
Control: \(u_t\) = harvest, constrained by \(0 \leq u_t \leq x_t\)
Dynamics:
where \(r > 1\) is the growth rate, \(\epsilon_t\) is a mean-one log-normal shock, and \((x_t - u_t)\) is the escapement (remaining stock).
Profit:
where \(p\) is price and \(c_0/x_t\) is the stock-dependent unit cost.
Objective: Maximize \(\mathbb{E}\left[\sum_{t=0}^{\infty} \delta^t \pi(u_t, x_t)\right]\)
Optimal Policy¶
Reed (1979) [Reed1979] proved the optimal policy has constant escapement form:
where the optimal escapement level is:
This requires the impatience condition \(\delta r < 1\).
This model provides an excellent benchmark for RL algorithms since \(S^*\) can be computed analytically without dynamic programming.
References
Reed, W.J. (1979). “Optimal escapement levels in stochastic and deterministic harvesting models.” Journal of Environmental Economics and Management, 6(4), 350-363.
- skagent.models.resource_extraction.df_u(states, shocks, parameters)¶
Decision function compatible with DBlock interface.
- skagent.models.resource_extraction.dr_u(x)¶
Optimal harvest under constant escapement policy.
- skagent.models.resource_extraction.make_optimal_decision_rule(parameters)¶
Compute the optimal constant-escapement policy from Reed (1979).
Reed showed that for the model with multiplicative shocks and stock-dependent costs \(c(x) = c_0/x\), the optimal policy has the form:
\[u^*(x) = \max(0, x - S^*)\]where \(S^*\) is the optimal escapement level (target stock to maintain). This can be computed analytically from the first-order condition without solving the full dynamic programming problem.
- Parameters:
parameters (dict) –
Model parameters including:
r: float, growth ratep: float, price per unitc_0: float, cost parameterDiscFac: float, discount factor
- Returns:
decision_rule (callable) – Maps stock x to optimal harvest \(u^*(x) = \max(0, x - S^*)\)
decision_function (callable) – Compatible with DBlock interface:
decision_function(states, shocks, parameters)
Notes
The optimal escapement \(S^*\) satisfies the first-order condition:
\[p - \frac{c_0}{S^*} = \delta r \left(p - \frac{c_0}{r S^*}\right)\]where \(\delta\) is the discount factor. Rearranging gives:
\[S^* = \frac{c_0 (1 - \delta)}{p (1 - \delta r)}\]This requires the “impatience condition” \(\delta r < 1\), which ensures the agent prefers extraction over indefinite accumulation.
Cournot¶
Quantity competition among several firms: the smallest model with an entity class, an agent role and an aggregation. Its decision profiles are supplied rather than solved for, since the model has a crossing.
Cournot (1838) quantity competition among firms.
The smallest model that exercises an entity class, an agent role and an aggregation. Several firms each draw a production cost, each choose a quantity, and the market price falls with the average quantity supplied:
The model is static: there are no arrival states, so simulating it for \(T\) periods is \(T\) independent plays of the same one-shot game.
Because the price is a function of the mean rather than the sum, the slope on total quantity is \(b/N\). That keeps the two readings of \(N\) below in one model without changing an equation.
- skagent.models.cournot.A = 10.0¶
Demand intercept, shared by both calibrations.
- skagent.models.cournot.B = 1.0¶
Demand slope on the average quantity, shared by both calibrations.
- skagent.models.cournot.PROFILES = {'cournot-nash': [4.5, 4.5, 4.5], 'joint-monopoly': [3.0, 3.0, 3.0], 'one-defects': [6.0, 3.0, 3.0]}¶
Three hand-derivable profiles at
collusion_calibration().Supplied, not found. Colluding pays jointly (27.0 against 20.25 in total profit), the cartel is unstable because deviating pays the deviator (12 against 9) at its partners’ expense (6 against 9), and Nash is stable.
- skagent.models.cournot.analytic_moments(low=2.0, high=6.0)¶
The heterogeneous-cost market’s analytic aggregate and expected profit.
Holds under the rule
competitive_rule(), in the limit of many firms.- Returns:
Q,PandE[u].- Return type:
- skagent.models.cournot.collusion_calibration(size=3, cost=4.0)¶
A market of size firms that all have cost cost.
- skagent.models.cournot.competitive_rule(c)¶
The supplied rule the heterogeneous-cost claims are stated under.
- skagent.models.cournot.heterogeneous_calibration(size=200, low=2.0, high=6.0)¶
A market of size firms whose costs are uniform on
[low, high].
- skagent.models.cournot.monopoly_quantity(cost=4.0)¶
The quantity per firm that maximizes the industry’s total profit.
- skagent.models.cournot.nash_quantity(size=3, cost=4.0)¶
The symmetric Cournot-Nash quantity per firm.
From the first-order condition
A - bq - c - qb/N = 0, which givesq = N(A - c) / (b(N + 1)).
- skagent.models.cournot.profile_rules(quantities)¶
One fixed decision rule per firm, so a profile may be asymmetric.
The asymmetry is in the ASSIGNMENT of rules to instances, not in the model: each rule is a constant function, and none reads a firm’s position.
- Parameters:
quantities (sequence of float) – One quantity per firm, in instance order.
- Returns:
Callables, one per firm, suitable as a
Simulatordecision rule.- Return type:
Lemons¶
Akerlof’s market for adverse selection, at the paper’s own numbers: sellers know what they hold and buyers can price only the average of what is offered, so the better cars leave the market first. Eight leaf blocks compose into several markets, which differ in when the price is set – anticipated by the sellers, posted from the round before, or committed to by a buyer – and in whether quality is a spread or Akerlof’s two types. All of them have an acyclic relevance graph, and only some can be solved a decision at a time.
Akerlof (1970) adverse selection: the market for lemons.
Sellers know the quality of what they hold; buyers do not, and can price only the average quality of what is offered. Each seller holds one item of quality \(\theta\), values it at \(\theta\), and parts with it only if the price covers that. A buyer values the same item at a premium \(m\), so the price that clears a competitive market is the buyer valuation of what actually trades:
The defaults here are the paper’s own numbers [Akerlof1970]: quality uniform
on \([0, 2]\), and buyers who value a car at \(3/2\) of what its owner
does. So lemons_calibration() with no arguments is the model of that
paper’s section II, and its answer is the paper’s – no trade at all, although
every car is worth more to a buyer than to the seller holding it.
The module builds its markets out of eight leaf blocks. Each states one thing about the market, and every version below is a different composition of them, so no equation is written twice:
spread_quality_block quality is uniform between the ends of the range two_type_quality_block a car is a peach or a lemon, and nothing between supply_block the seller decides, seeing only its own quality offer_block the seller decides, seeing the price on the table seller_payoff_block the seller’s surplus, at whatever price is current market_block the competitive price, given what was offered bid_block a buyer names the price instead surplus_block that buyer’s surplus, per seller
The first two are alternatives to each other, and so are the next two: a market takes one quality distribution and one seller decision. The last two are the monopsony market’s, and no other version composes them.
Where the payoff block sits determines which price the seller is paid at, and where the market block sits determines whether the price is known when the seller decides. Declaration order is therefore the whole of the timing, and there is no timing annotation anywhere in the model. The three timings differ in nothing else.
In lemons_block and peaches_block the sellers anticipate the
price their own supply induces. They decide, the market clears on what they
offered, and they are paid at that price. Nothing is lagged and no seller reads
a price, so a seller’s rule is a best response to the price that rule induces
once every seller follows it. The equilibrium is a fixed point in rules, and a
solver has to find it.
In naive_lemons_block and naive_peaches_block the sellers
respond to a price already posted. The payoff block comes before the market
block, so sellers are paid at the price they responded to and the market then
clears at what their decisions imply. The price is read before it is written,
which makes it an arrival state: period \(t\)’s sellers face period
\(t-1\)’s clearing price. Simulating \(T\) periods therefore runs
\(T\) rounds of the clearing map, and no solver is needed to watch the same
equilibrium arrive.
In monopsony_block a buyer commits to a price before supply. The price
becomes a decision, owned by a buyer who names it to maximize its own surplus
and cannot see quality:
The buyer’s payoff is written per seller rather than as a total, so the closed forms below do not depend on how many sellers there are.
A buyer that named a price only after the cars had been offered would name the lowest price it was allowed to. Once the pool is fixed, its surplus is \(v(m\bar\theta - p)\) in a volume and a grade it can no longer change, so the surplus falls in the price at every price and the best response is the lower bound. Sellers anticipate that and offer nothing, so the market shuts for a reason that has nothing to do with adverse selection. That is why the buyer here commits first, and why a competitive price is a condition on the market rather than any one participant’s choice: competition among buyers is what holds the price up to the value of what trades.
The three timings need three different treatments, and the relevance graph
tells them apart. Where the price is anticipated, a seller’s payoff runs
through it to every other seller’s decision, so the model is a strategic fixed
point among instances of one class and relies_on("S", "S") is True:
one decision, and it is its own predecessor, so no order solves it. Where the
price is posted, the same call is False and correct, since this period’s
payoff turns on last period’s price and within a period there is no reliance to
find. Where a buyer commits, the graph reports two nodes and one edge, from the
price to the sell decision, and its topological order is backward induction.
The market behaves differently under the two quality distributions. The
uniform of section II makes the clearing map exactly linear, since
\(E[\theta \mid \theta \leq p]\) is \((\ell + p)/2\), so the only prices
it can reproduce are no trade, a corner, or – at \(m = 2\) exactly – every
price at once. The paper’s automobiles are two types rather than a spread, and
that is where a second price becomes possible: only lemons trade at
\(m\ell\), and if peaches are common enough there is a second price,
\(m E[\theta]\), at which everything trades. Which one the market reaches
depends on where it starts. See MARKETS and PEACH_MARKETS for
the named configurations, and clearing_fixed_points() and
peaches_fixed_points() for the general answers.
Within the uniform family the quality floor is what decides whether the unravelling has anywhere to stop. With no floor and a premium below 2 the map is a contraction onto \(p = 0\). With a floor the same map has a second, higher fixed point at \(m\ell/(2 - m)\), and the market unravels down to it rather than away: the top of the market is destroyed and the bottom keeps trading. Above the threshold \(m = 2\) the collapse is unstable instead, and the price runs up to \(m(\ell + h)/2\), where every item trades.
clearing_map() and peaches_clearing_map() are each the same function
under two readings. Iterated, they are the price path a posted-price market
simulates; read once, they give the price an anticipated price induces. An
anticipating market’s equilibria are therefore exactly their fixed points.
Signalling is the paper’s own answer to adverse selection, and this model is
shaped to reach it. Section IV of [Akerlof1970] names guarantees, brand names
and licensing as the institutions that counteract it, and all of them work the
same way: the seller takes a costly action whose cost is lower for higher quality, so
that taking it is credible. Two things have to change here. The first is
already in place, since a separating equilibrium separates types and
two_type_quality_block supplies them. The second is that the price stops
being one number: the buyer commits to a pricing rule, taking the seller’s
signal as its information set, and that rule is applied once per seller. Each
seller is then paid at the price its own signal earns. That adds a second
control to supply_block, which is why the seller’s information set is
its quality alone, so a signalling decision drops in beside the sell decision.
When the price is below what every seller thinks their own car is worth, nothing is offered at all, and the clearing price would be an average over an empty market. This model answers zero. That is a choice rather than a definition, and it is what makes no trade an equilibrium rather than an error: a market at a price of zero has nothing offered to it, so it clears at zero again and stays there. A different market could reasonably answer something else, so the value is written into the clearing equation rather than supplied by the library.
Binary decisions are relaxed to continuous [0, 1] controls, pending
discrete-action support, following the convention of
skagent.models.macid. seller_rule() and supply_rule() return
exact 0 and 1, so the relaxation costs the supplied equilibria nothing.
References
Akerlof, G.A. (1970). “The Market for ‘Lemons’: Quality Uncertainty and the Market Mechanism.” The Quarterly Journal of Economics, 84(3), 488-500. https://doi.org/10.2307/1879431
- skagent.models.lemons.MARKETS = {'akerlof': {'low': 0.0, 'premium': 1.5}, 'knife-edge': {'low': 0.0, 'premium': 2.0}, 'no-collapse': {'low': 0.0, 'premium': 2.5}, 'partial-collapse': {'low': 0.4, 'premium': 1.5}}¶
Four configurations of
lemons_calibration(), on quality up to 2.Each is a pair of arguments rather than a whole calibration, so the number of sellers stays the caller’s:
lemons_calibration(size=1000, **MARKETS["akerlof"]).name competitive prices monopsony price and payoff akerlof 0 0, and 0 partial-collapse 0 and 1.2 0.8, and 0.025 knife-edge every price up to 2 any, and 0 no-collapse 0 (unstable) and 2.5 2, and 0.5
akerlofis the paper’s own section II. Underpartial-collapsethe two competitive prices are both stable and the quality floor divides their basins: a market that starts below the floor has nobody willing to sell and stays collapsed, and one that starts above it unravels down to 1.2 rather than to nothing.
- skagent.models.lemons.PEACH_MARKETS = {'lemons-only': {'share': 0.3}, 'two-prices': {'share': 0.7}}¶
Two configurations of
peaches_calibration(), at its default qualities.name competitive prices lemons-only 0 and 0.6 two-prices 0 and 0.6 and 2.28
A lemon is worth 0.4 and a peach 2.0, so at a price of 0.6 the lemons all sell and no peach does, whatever peaches are worth and however many there are. The second price exists only where peaches are common enough to carry the average, which is
peach_share_for_trade()and is 0.583 here; above it a market that starts high stays high and one that starts low still collapses to the lemons.
- skagent.models.lemons.PREMIUM = 1.5¶
How much more a buyer values an item than the seller holding it.
Akerlof’s own 3/2: his sellers value a car at its quality, his buyers at three halves of it.
- skagent.models.lemons.PREMIUM_THRESHOLD = 2.0¶
The premium at which the uniform market’s clearing map has slope exactly 1.
Below it the map contracts and the market collapses; above it no trade is an unstable fixed point and the price runs up until every item trades. The threshold does not depend on the quality range.
- skagent.models.lemons.QUALITY_HIGH = 2.0¶
uniform on [0, 2].
- Type:
The ends of the quality range, and Akerlof’s own
- skagent.models.lemons.bid_rule(price)¶
A buyer that always names price, as a decision rule for
p.- Parameters:
price (float)
- Returns:
A rule taking no arguments, since the buyer’s information set is empty.
- Return type:
callable
- skagent.models.lemons.buyer_payoff(p, premium=1.5, low=0.0, high=2.0)¶
The monopsonist’s surplus per seller at price p, on uniform quality.
- Parameters:
p (float) – The price the buyer names.
premium (float, optional) – As in
lemons_calibration().low (float, optional) – As in
lemons_calibration().high (float, optional) – As in
lemons_calibration().
- Return type:
- skagent.models.lemons.clearing_fixed_points(premium=1.5, low=0.0, high=2.0)¶
The prices a uniform-quality competitive market reproduces.
These are the equilibria of
lemons_blockand the rest points ofnaive_lemons_block. No trade is always one of them. A second appears either where the unravelling stops at the quality floor, or, abovePREMIUM_THRESHOLD, where every item trades.- Parameters:
premium (float, optional) – As in
lemons_calibration().low (float, optional) – As in
lemons_calibration().high (float, optional) – As in
lemons_calibration().
- Returns:
In increasing order. At the threshold with no quality floor the map is the identity below high and the two values returned are the ends of a continuum of fixed points rather than two isolated ones.
- Return type:
- skagent.models.lemons.clearing_map(p, premium=1.5, low=0.0, high=2.0)¶
The price that sellers facing p bring about, on a uniform quality range.
- Parameters:
p (float) – The price the sellers act on.
premium (float, optional) – As in
lemons_calibration().low (float, optional) – As in
lemons_calibration().high (float, optional) – As in
lemons_calibration().
- Return type:
- skagent.models.lemons.clearing_path(p0, periods, mapping=None, **market)¶
A clearing map iterated from p0, one price per round.
- Parameters:
p0 (float) – The starting price.
periods (int) – How many rounds.
mapping (callable, optional) – The clearing map. Defaults to
clearing_map().**market – Passed to mapping.
- Returns:
Length periods, the price after each round.
- Return type:
- skagent.models.lemons.clearing_price(theta, S, premium)¶
The competitive price: buyer valuation of the average item that sold.
A weighted mean rather than a mean over a selection, which is what keeps it usable on every path the library solves on: there is no boolean index, no branch on the data and no dynamic shape, so it differentiates and batches under
torchas readily as it evaluates undernumpy. Where the sell decision is 0 or 1 the two forms agree exactly; where it is relaxed to a probability of selling, the weighted form is the expected quality of what trades, which is the mixed extension rather than an approximation of it.
- skagent.models.lemons.lemons_calibration(size=10000, low=0.0, high=2.0, premium=1.5)¶
A market of size sellers whose quality is uniform on
[low, high].The defaults are Akerlof’s section II.
- Parameters:
size (int, optional) – How many sellers. The clearing price is an average over those who sold, so this is the accuracy of the fixed point rather than part of the model.
low (float, optional) – The ends of the quality range.
lowis the quality floor.high (float, optional) – The ends of the quality range.
lowis the quality floor.premium (float, optional) – How much more a buyer values an item than the seller holding it.
- Return type:
- skagent.models.lemons.monopsony_price(premium=1.5, low=0.0, high=2.0)¶
The price that maximizes
buyer_payoff().- Parameters:
premium (float, optional) – As in
lemons_calibration().low (float, optional) – As in
lemons_calibration().high (float, optional) – As in
lemons_calibration().
- Returns:
At and above
PREMIUM_THRESHOLDthe buyer takes the whole market at high. Below it the buyer bids the floor up by a factor of1 / (2 - premium), which with no floor is no trade at all.- Return type:
The share of peaches above which a market in peaches exists at all.
Below it, no price a buyer will pay for the average car is enough to bring a peach out, whatever the market does.
- Parameters:
premium (float, optional) – As in
peaches_calibration().low (float, optional) – As in
peaches_calibration().high (float, optional) – As in
peaches_calibration().
- Returns:
A share, which may exceed 1 where no share of peaches is enough.
- Return type:
- skagent.models.lemons.peaches_calibration(size=10000, low=0.4, high=2.0, premium=1.5, share=0.5)¶
A market of size sellers holding a peach or a lemon and nothing between.
- Parameters:
size (int, optional) – How many sellers.
low (float, optional) – The quality of a lemon and of a peach.
high (float, optional) – The quality of a lemon and of a peach.
premium (float, optional) – How much more a buyer values an item than the seller holding it.
share (float, optional) – The fraction of cars that are peaches.
- Return type:
- skagent.models.lemons.peaches_clearing_map(p, premium=1.5, low=0.4, high=2.0, share=0.5)¶
The price that sellers facing p bring about, in a two-type market.
Below a lemon’s worth nothing is offered; between the two qualities only lemons are, and their average quality does not depend on the price at all; at or above a peach’s worth everything is.
- Parameters:
p (float) – The price the sellers act on.
premium (float, optional) – As in
peaches_calibration().low (float, optional) – As in
peaches_calibration().high (float, optional) – As in
peaches_calibration().share (float, optional) – As in
peaches_calibration().
- Return type:
- skagent.models.lemons.peaches_fixed_points(premium=1.5, low=0.4, high=2.0, share=0.5)¶
The prices a two-type competitive market reproduces.
Up to three, and the middle one is what the uniform market cannot have: a price at which trade survives and every peach is withheld.
- Parameters:
premium (float, optional) – As in
peaches_calibration().low (float, optional) – As in
peaches_calibration().high (float, optional) – As in
peaches_calibration().share (float, optional) – As in
peaches_calibration().
- Returns:
In increasing order.
- Return type:
- skagent.models.lemons.seller_rule(theta, p)¶
Sell if and only if the posted price covers the seller’s own valuation.
- Parameters:
theta (array) – Quality, one per seller.
p (float) – The price the sellers can see.
- Returns:
1.0 where the item sells, 0.0 where it does not.
- Return type:
array
- skagent.models.lemons.supply_rule(price)¶
Supply as though the market will clear at price, seeing only quality.
The equilibrium rule of an anticipating market is this one at a fixed point of that market’s clearing map.
- Parameters:
price (float) – The price the sellers anticipate.
- Returns:
A rule of quality alone.
- Return type:
callable
Aiyagari¶
Many households saving out of labour income, whose average assets are the economy’s capital and therefore set the interest rate and wage they all face. Under a fixed savings rate the aggregate has a closed form, so the model checks the arithmetic of a dynamic path through an entity class rather than only its shape.
Aiyagari (1994) [Aiyagari1994]: a cross-section that prices its own capital.
Many households each save out of labour income, and what they save between them is the economy’s capital stock. Capital sets the interest rate and the wage, those set what each household has to spend, and what each household leaves over is next period’s capital. The households never interact directly; they meet only through an average.
Symbols¶
One value per household:
theta |
the labour endowment drawn this period, which has mean one |
a |
assets carried in from last period, the model’s arrival state |
z |
cash on hand: labour income plus assets and the interest on them |
c |
consumption, the household’s one decision |
u |
the household’s payoff from consuming that much |
One value for the whole economy, and capitalised for it:
K |
capital per head, the average of |
R |
the net interest rate, so a household earns |
W |
the wage paid per unit of labour endowment |
Calibration parameters:
alpha |
capital’s share of output |
delta |
the fraction of capital that wears out each period |
CRRA |
relative risk aversion in the household’s utility |
DiscFac |
the household’s discount factor |
sigma_theta |
how dispersed the labour endowment is |
household |
how many households there are |
The model¶
The two prices are the marginal products of Cobb-Douglas production in its intensive form, \(Y = K^{\alpha} L^{1-\alpha}\), with labour normalised to one unit per head: the endowment has mean one and every household supplies it whatever the wage, so \(L\) drops out of both formulas. Factors are paid their marginal products, so output is exactly exhausted before depreciation, \((R + \delta) K + W = K^{\alpha}\), and capital’s share of it is \(\alpha\).
A fraction \(\delta\) of the capital stock wears out each period, so \(R\) is the marginal product of capital net of that, and a household’s assets earn \((1 + R)\). This is the rate [Aiyagari1994] works with. Depreciation is also what keeps the model’s scale sensible: a household’s resources \(W + (1+R)K\) come to output plus the capital that survived, \(K^{\alpha} + (1-\delta)K\), so the stationary capital-output ratio is exactly \(s/(1 - s(1-\delta))\). With no depreciation that is \(s/(1-s)\), which reaches 9 at plausible savings rates, about three times what an economy shows, because nothing ever wears out.
The labour endowment is drawn independently each period. In [Aiyagari1994] it follows a persistent AR(1) process, and that persistence is what drives the paper’s precautionary saving. Under a fixed savings rate persistence leaves the aggregate’s law of motion unchanged, since the class still averages an endowment of mean one, so this model draws it fresh each period.
The savings rate is a fraction of cash on hand rather than of income, so it is
not the textbook savings rate and picking one by eye is misleading.
savings_rate_for() inverts the relationship instead, returning the rate at
which the economy settles at a given interest rate.
The market block is declared before the households, so \(K\) is computed from the assets the households arrived with rather than the ones they are about to choose. That makes \(a\) an arrival state, and simulating \(T\) periods runs \(T\) rounds of the aggregate’s own law of motion.
Two symbols differ from the design note this model was written from. It writes
the labour endowment \(l\), which is theta here, both because a bare
l is not a legal identifier under the project’s linter and because theta
is what the rest of the library calls a mean-one transitory shock. And it writes
capital per head \(k\), which is K here, since the library capitalises a
variable that stands for the whole economy rather than for one member of it.
What the model is for¶
Under a fixed savings rate the aggregate has a closed form, and that is what makes this model a test rather than only a demonstration. If every household consumes \((1-s)z\), then \(a' = s z\), and averaging over the class with \(E[\theta] = 1\):
so the stationary capital solves \(K^{1-\alpha} = s / (1 - s(1-\delta))\), giving \(R^{*} = \alpha (1 - s(1-\delta))/s - \delta\) exactly. The map is increasing and concave, passes through the origin, and leaves it steeper than the 45-degree line, so it crosses that line at exactly one positive capital. From any positive start the path therefore moves monotonically toward that crossing without overshooting, and the convergence needs no damping. The map’s slope at the crossing is \(\alpha + s(1-\delta)(1-\alpha)\), below 1 whenever \(s(1-\delta) < 1\); it sets the speed of the last stretch, each period closing a fraction one minus the slope of the remaining gap. Away from the crossing the slope can exceed 1, so the map is not a contraction on the whole half-line, and the global convergence rests on its shape.
A finite class averages an endowment only close to one, so each simulated
period adds \(s\,W(\bar\theta - 1)\) to the map, where \(\bar\theta\) is
that period’s average draw. capital_map() takes that average as its
endowment argument, and with it the simulated path matches the closed form
to rounding error.
That claim is cross-sectional rather than optimal. The savings rate is a rule of thumb, no household is solving anything, and the aggregate still arrives where the arithmetic says it should – which is what makes it a check on the simulator’s treatment of a class and its average rather than on a solver.
References
Aiyagari, S.R. (1994). “Uninsured Idiosyncratic Risk and Aggregate Saving.” The Quarterly Journal of Economics, 109(3), 659-684. https://doi.org/10.2307/2118417
- skagent.models.aiyagari.CAPITAL_SHARE = 0.36¶
Capital’s share of output, the exponent on capital in the production function.
- skagent.models.aiyagari.DEPRECIATION = 0.08¶
The fraction of the capital stock that wears out each period.
A conventional value for an annual calibration. Setting it to zero gives a model in which capital lasts forever, whose capital-output ratio is about three times an economy’s.
- skagent.models.aiyagari.aiyagari_calibration(size=1000, alpha=0.36, delta=0.08, crra=2.0, sigma=1.0, beta=0.96)¶
An economy of size households.
- Parameters:
size (int, optional) – How many households. The capital stock is an average over them, so this is the accuracy of the aggregate rather than part of the model.
alpha (float, optional) – Capital’s share of output.
delta (float, optional) – The fraction of the capital stock that wears out each period.
crra (float, optional) – Relative risk aversion in the household’s utility, with 1 meaning log utility. It does not enter the aggregate under a fixed savings rate.
sigma (float, optional) – Dispersion of the labour endowment, which has mean one whatever this is.
beta (float, optional) – The household’s discount factor. It does not enter the aggregate under a fixed savings rate.
- Return type:
- skagent.models.aiyagari.capital_map(capital, rate, alpha=0.36, delta=0.08, endowment=1.0)¶
The capital an economy holding capital per head leaves for next period.
Iterated, this is the path
aiyagari_blocksimulates undersavings_rule(); its fixed point isstationary_capital().- Parameters:
capital (float) – Capital per head this period.
rate (float) – The savings rate the households follow.
alpha (float, optional) – Capital’s share of output.
delta (float, optional) – The fraction of the capital stock that wears out each period.
endowment (float, optional) – The households’ average labour endowment this period. It is one in expectation, the default; a finite class draws an average only close to one, and passing the realized average makes the map exact.
- Return type:
- skagent.models.aiyagari.convergence_rate(rate, alpha=0.36, delta=0.08)¶
The slope of
capital_map()at its fixed point.Below 1 whenever
rate * (1 - delta)is. Near the fixed point each period closes a fraction one minus this slope of the remaining gap. Convergence from any starting capital is a separate fact: the map is increasing and concave through the origin, so the path approaches the fixed point monotonically, even from capital low enough that the slope there exceeds 1.
- skagent.models.aiyagari.savings_rate_for(interest, alpha=0.36, delta=0.08)¶
The savings rate at which the economy settles at interest rate interest.
The inverse of
stationary_prices()’ first entry. Cash on hand includes a household’s whole asset position, so a savings rate here is not the textbook fraction of income and is easier to choose by the interest rate it implies than by eye.
- skagent.models.aiyagari.savings_rule(rate)¶
Consume all but rate of what is on hand, as a decision rule for
c.- Parameters:
rate (float) – The fraction saved, in
(0, 1).- Returns:
A rule of cash on hand alone.
- Return type:
callable
- skagent.models.aiyagari.stationary_capital(rate, alpha=0.36, delta=0.08)¶
The capital per head that reproduces itself under rate.
- skagent.models.aiyagari.stationary_prices(rate, alpha=0.36, delta=0.08)¶
The interest rate and wage at
stationary_capital().
Differential Privacy¶
Benthall and Cummings’ two causal games for tuning a differential-privacy parameter: data subjects decide whether to share, an analyst estimates a population mean from the reports that arrive, and a designer above both chooses how much noise the mechanism adds. The two games are the local and central trust models, which differ only in whether each subject privatizes its own report or the analyst privatizes the estimate. Both agents’ equilibrium rules have closed forms, so the model is an oracle for the machinery that finds them: an exact backup returns the subjects’ threshold rule exactly, while the analyst’s estimator, whose information set is the whole class, is supplied.
Tuning a differential-privacy parameter as a causal game.
An analyst wants to estimate a population mean and offers a differential-privacy guarantee to the people it collects data from. Each data subject decides whether to share. Sharing helps the estimate, which the subject values, and costs the subject privacy, which it does not. How much noise the mechanism adds is neither agent’s choice: a designer sets it, knowing that the agents will then play a best response to it. So the noise scale is a design parameter above an equilibrium rather than a parameter of anyone’s problem, and choosing it is the mechanism being designed.
The model is Benthall and Cummings (2026) [BC2026], sections 4 to 6.
Two games, and they differ in one place¶
The local model has each subject add its own noise before sharing, so what reaches the analyst is already private. The central model has subjects share their true data with a trusted analyst, which adds noise once to the estimate it publishes. Every equation below is shared except where the noise enters:
with \(a\) the population mean, \(\zeta_i \sim U(-\Delta/2, \Delta/2)\) individual variation, \(p_i \sim N(0, \sigma_p^2)\) how much subject \(i\) minds sharing, \(q_i\) what it gains by sharing, and \(\sigma\) the noise scale. The report is \(d_i = c_i(b_i + \gamma_i)\) locally, with one \(\gamma_i \sim N(0, \sigma^2)\) per subject, and \(d_i = c_i b_i\) centrally, where a single \(\gamma\) is added to the estimate instead.
What each agent decides, and what is supplied¶
Both decisions have closed forms in the paper, so this module supplies both as rules and the blocks declare them as decisions:
A subject shares iff sharing is worth more than it costs. The utility is linear in \(c_i\), so the optimum is at a vertex whatever the control’s bounds: \(c_i = 1\) iff \(p_i < q_i \sigma\), which is
sharing_rule(). A solver on the relaxed \([0, 1]\) control returns that rule rather than an approximation of it.The analyst’s estimate is the mean of the reports it received, which the paper proves is the minimum-variance unbiased estimator of \(a\), so maximizing its accuracy is computing that mean. It is
analyst_rule().
The analyst’s information set is the whole subject class, which no policy network
in this library is shaped for – a rule over a class is a permutation-invariant
architecture – so this decision is represented and supplied rather than learned.
The subjects’ is not: their information set is per-instance, and
sharing_rule() is what a solver should reproduce.
Two reductions written as weighted means¶
The analyst averages over the subjects who shared, and averages a prior when
nobody did. Both are written as weighted sums over the whole class rather than as
a mean over a selection and a branch on the count, so the estimate carries no
boolean index and no branch on the data: it differentiates and batches under
torch as readily as it evaluates under numpy. A non-sharer contributes
nothing to either the numerator or the denominator, which is what makes a null
report and a zero weight the same thing here, and the prior enters as one
pseudo-observation of negligible weight, which is what makes the empty database
return the prior exactly.
The designer’s problem¶
The designer maximizes an objective of its own over \(\sigma\), given that the agents play their equilibrium at each \(\sigma\):
That is a sweep over equilibria rather than a solution concept or a method, so it
is a loop a caller writes; this module supplies what the loop evaluates.
designer_objective() is the paper’s own example of an \(h\), and it
uses the bounded subject utility of the paper’s Figure 5,
\(u_i = c_i(1 - e^{p_i}/(1 + \sigma))\), under which more noise always helps
a subject. The base utility above is unbounded as \(\sigma \to 0\): a subject
who positively enjoys sharing gains \(-p_i/\sigma\), without limit.
The closed forms here are exact, and the accuracy ones are the paper’s equation (3) and its central analogue. The number of subjects who share is \(n' \sim \mathrm{Binomial}(n, F_p(q\sigma))\), and the error is a variance divided by that count, so the count’s distribution has to be summed over rather than replaced by its mean.
What the estimator rests on¶
The empirical mean is unbiased because who shares is independent of what is
being measured: \(p_i\) and \(b_i\) are drawn independently, so
selecting on \(p_i\) selects a random subset of the data. That is an
assumption about the population rather than a property of the mechanism, and
analyst_rule() is the minimum-variance unbiased estimator only while it
holds.
Benthall, S. and Cummings, R. (2026). “Principled Differential Privacy Parameter Tuning with Causal Games.” FAccT ‘26. https://doi.org/10.1145/3805689.3806530
- skagent.models.privacy.DELTA = 1.0¶
The width of individual variation, so
zeta ~ U(-1/2, 1/2).Bounded rather than normal because it is what bounds the sensitivity of the statistic being estimated, which is what the mechanism’s noise is calibrated to.
- skagent.models.privacy.Q = 1.0¶
What a subject gains by sharing, and the paper’s own value throughout.
- skagent.models.privacy.SIGMA_P = 1.0¶
p ~ N(0, 1).- Type:
The spread of privacy concern, and the paper’s own
- skagent.models.privacy.SUBJECTS = 1000¶
How many data subjects, and the size the paper’s figures are computed at.
- skagent.models.privacy.analyst_rule(a_mean)¶
estimate()against a fixed prior mean, as a decision rule forf.- Parameters:
a_mean (float) – The mean of the analyst’s prior over the population mean.
- Returns:
A rule of
(c, d), the class’s decisions and reports.- Return type:
callable
- skagent.models.privacy.bounded_local_block = RBlock(name='local_dp', description='', blocks=[DBlock(name='population', description='', shocks={'a': (<class 'skagent.distributions.Normal'>, {'mu': 'a_mean', 'sigma': 'a_sd'})}, dynamics={}, reward={}, entity=None), RBlock(name='subjects', description='', blocks=[DBlock(name='data', description='', shocks={'zeta': (<class 'skagent.distributions.Uniform'>, {'low': '-delta/2', 'high': 'delta/2'}), 'p': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma_p'})}, dynamics={'b': <function <lambda>>}, reward={}, entity=None), DBlock(name='sharing', description='', shocks={}, dynamics={'c': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='bounded_privacy_cost', description='', shocks={}, dynamics={'u': <function <lambda>>}, reward={'u': 'subject'}, entity=None), DBlock(name='local_noise', description='', shocks={'gamma': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma'})}, dynamics={}, reward={}, entity=None), DBlock(name='local_report', description='', shocks={}, dynamics={'d': <function <lambda>>}, reward={}, entity=None)], entity=Entity(name='subject')), DBlock(name='estimate', description='', shocks={}, dynamics={'f': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='local_accuracy', description='', shocks={}, dynamics={'g': <function <lambda>>}, reward={'g': 'analyst'}, entity=None)], entity=None)¶
The local model under the bounded utility
designer_objective()uses.
The fraction who share under
bounded_sharing_rule().- Parameters:
sigma (float) – As in
calibration().sigma_p (float) – As in
calibration().
- Return type:
- skagent.models.privacy.bounded_sharing_rule(sigma)¶
Share iff a lognormal privacy cost is below the benefit of sharing.
The rule for the bounded utility
c * (1 - exp(p) / (1 + sigma))thatdesigner_objective()is written against.- Parameters:
sigma (float) – The mechanism’s noise scale.
- Returns:
A rule of
(b, p), returning 1.0 where the subject shares.- Return type:
callable
- skagent.models.privacy.bounded_subject_utility(sigma, sigma_p=1.0)¶
A subject’s expected utility under the bounded, lognormal-cost utility.
E[c (1 - exp(p) / (1 + sigma))]underbounded_sharing_rule(). Rises in sigma and is bounded above by the benefit of sharing, which is what makes the designer’s objective have an interior maximum.- Parameters:
sigma (float) – As in
calibration().sigma_p (float) – As in
calibration().
- Return type:
- skagent.models.privacy.calibration(sigma, size=1000, q=1.0, sigma_p=1.0, delta=1.0, a_mean=0.0, a_sd=1.0)¶
A calibration of either game at noise scale sigma.
- Parameters:
sigma (float) – The mechanism’s noise scale, and the designer’s decision. Must be positive: the base utility divides by it.
size (int, optional) – How many data subjects.
q (float, optional) – What a subject gains by sharing.
sigma_p (float, optional) – The spread of privacy concern.
delta (float, optional) – The width of individual variation, which bounds the sensitivity of the statistic and so what the mechanism’s noise is calibrated to.
a_mean (float, optional) – The analyst’s prior over the population mean.
a_sdis what an estimate falls back on when nobody shares, so it is part of the accuracy the designer weighs and not only a simulation detail.a_sd (float, optional) – The analyst’s prior over the population mean.
a_sdis what an estimate falls back on when nobody shares, so it is part of the accuracy the designer weighs and not only a simulation detail.
- Return type:
- skagent.models.privacy.central_block = RBlock(name='central_dp', description='', blocks=[DBlock(name='population', description='', shocks={'a': (<class 'skagent.distributions.Normal'>, {'mu': 'a_mean', 'sigma': 'a_sd'})}, dynamics={}, reward={}, entity=None), DBlock(name='dp_noise', description='', shocks={'gamma': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma'})}, dynamics={}, reward={}, entity=None), RBlock(name='subjects', description='', blocks=[DBlock(name='data', description='', shocks={'zeta': (<class 'skagent.distributions.Uniform'>, {'low': '-delta/2', 'high': 'delta/2'}), 'p': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma_p'})}, dynamics={'b': <function <lambda>>}, reward={}, entity=None), DBlock(name='sharing', description='', shocks={}, dynamics={'c': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='privacy_cost', description='', shocks={}, dynamics={'u': <function <lambda>>}, reward={'u': 'subject'}, entity=None), DBlock(name='central_report', description='', shocks={}, dynamics={'d': <function <lambda>>}, reward={}, entity=None)], entity=Entity(name='subject')), DBlock(name='estimate', description='', shocks={}, dynamics={'f': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='central_accuracy', description='', shocks={}, dynamics={'f_dp': <function <lambda>>, 'g': <function <lambda>>}, reward={'g': 'analyst'}, entity=None)], entity=None)¶
a trusted analyst adds noise once, to the estimate.
- Type:
The central model
- skagent.models.privacy.central_error(sigma, size=1000, q=1.0, sigma_p=1.0, delta=1.0, a_sd=1.0)¶
The analyst’s expected squared error in the central model.
The noise is added once rather than per report, so it is not averaged down: it enters as
sigma^2whatever the number of subjects, while what is averaged down is individual variation alone. That is why the two models make different designs optimal at the same privacy guarantee.- Parameters:
sigma (float) – As in
calibration().size (float) – As in
calibration().q (float) – As in
calibration().sigma_p (float) – As in
calibration().delta (float) – As in
calibration().a_sd (float) – As in
calibration().
- Returns:
E[-g], so smaller is better.- Return type:
- skagent.models.privacy.designer_objective(sigma, size=1000, sigma_p=1.0, delta=1.0, a_sd=1.0)¶
h = E[g] + E[u_i]: the paper’s example of a designer’s objective.The analyst’s accuracy plus a subject’s welfare, under the bounded utility and the local model. Accuracy falls in sigma and welfare rises in it, so the maximum is interior: a designer weighing both chooses a noise scale that is neither zero nor unbounded, which is the paper’s result.
A designer with a different objective is a different function here, and that is the point of the argument rather than a limitation of it: what is being designed is a trade between the parties, so the trade has to be stated.
- Parameters:
sigma (float) – As in
calibration().size (float) – As in
calibration().sigma_p (float) – As in
calibration().delta (float) – As in
calibration().a_sd (float) – As in
calibration().
- Return type:
See also
bounded_subject_utilitythe welfare term.
local_errorthe accuracy term, negated here.
- skagent.models.privacy.estimate(c, d, a_mean)¶
The mean of the reports received, and the prior mean when there are none.
The minimum-variance unbiased estimator of the population mean, given that who shares is independent of what is being measured.
A weighted sum over the whole class rather than a mean over a selection: a subject that did not share carries weight 0 and contributes to neither the numerator nor the denominator, so a null report and a zero weight are the same thing. The prior enters as one pseudo-observation of weight
_NOBODY, which leaves a non-empty database unmoved and is the whole estimate when the database is empty – the paper’sn' = 0branch, without a branch.
- skagent.models.privacy.expected_subject_utility(sigma, q=1.0, sigma_p=1.0)¶
A subject’s expected utility under the base utility and its rule.
Unbounded as sigma falls: the subjects who share at a small noise scale are those who most dislike privacy loss, and the utility credits them
-p / sigmafor sharing anyway.- Parameters:
sigma (float) – As in
calibration().q (float) – As in
calibration().sigma_p (float) – As in
calibration().
- Return type:
- skagent.models.privacy.local_block = RBlock(name='local_dp', description='', blocks=[DBlock(name='population', description='', shocks={'a': (<class 'skagent.distributions.Normal'>, {'mu': 'a_mean', 'sigma': 'a_sd'})}, dynamics={}, reward={}, entity=None), RBlock(name='subjects', description='', blocks=[DBlock(name='data', description='', shocks={'zeta': (<class 'skagent.distributions.Uniform'>, {'low': '-delta/2', 'high': 'delta/2'}), 'p': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma_p'})}, dynamics={'b': <function <lambda>>}, reward={}, entity=None), DBlock(name='sharing', description='', shocks={}, dynamics={'c': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='privacy_cost', description='', shocks={}, dynamics={'u': <function <lambda>>}, reward={'u': 'subject'}, entity=None), DBlock(name='local_noise', description='', shocks={'gamma': (<class 'skagent.distributions.Normal'>, {'mu': 0.0, 'sigma': 'sigma'})}, dynamics={}, reward={}, entity=None), DBlock(name='local_report', description='', shocks={}, dynamics={'d': <function <lambda>>}, reward={}, entity=None)], entity=Entity(name='subject')), DBlock(name='estimate', description='', shocks={}, dynamics={'f': <skagent.block.Control object>}, reward={}, entity=None), DBlock(name='local_accuracy', description='', shocks={}, dynamics={'g': <function <lambda>>}, reward={'g': 'analyst'}, entity=None)], entity=None)¶
each subject privatizes its own report.
- Type:
The local model
- skagent.models.privacy.local_error(sigma, size=1000, q=1.0, sigma_p=1.0, delta=1.0, a_sd=1.0)¶
The analyst’s expected squared error in the local model.
Every report carries the mechanism’s noise as well as individual variation, so the variance being averaged down is
sigma^2 + delta^2 / 12. More noise therefore costs accuracy directly and buys it back by drawing more subjects in, which is the trade the designer is making.- Parameters:
sigma (float) – As in
calibration().size (float) – As in
calibration().q (float) – As in
calibration().sigma_p (float) – As in
calibration().delta (float) – As in
calibration().a_sd (float) – As in
calibration().
- Returns:
E[-g], so smaller is better.- Return type:
- skagent.models.privacy.privacy_epsilon(sigma, delta=1.0, dp_delta=1e-05)¶
The
epsilonthe Gaussian mechanism gives at noise scale sigma.What the designer’s choice means as a privacy guarantee: the Gaussian mechanism is
(epsilon, dp_delta)-differentially private forepsilon = sensitivity * sqrt(2 log(1.25 / dp_delta)) / sigma, and the sensitivity of a mean whose entries vary by at most delta is delta.
The fraction of subjects who share at noise scale sigma.
F_p(q * sigma), the privacy-concern CDF at the benefit of sharing.- Parameters:
sigma (float) – As in
calibration().q (float) – As in
calibration().sigma_p (float) – As in
calibration().
- Return type:
- skagent.models.privacy.sharing_rule(sigma, q=1.0)¶
Share iff privacy concern is below what sharing is worth.
The subject’s utility
c * (q - p / sigma)is linear in the decision, so the optimum is at a vertex: share when the bracket is positive.- Parameters:
- Returns:
A rule of
(b, p), returning 1.0 where the subject shares. It ignoresb: the subject’s own data does not enter its utility, which is what makes the sharing decision uninformative about the data.- Return type:
callable
Multi-Agent Influence Diagrams (MAIDs)¶
Game-theoretic influence diagrams from the literature, encoded to illustrate
strategic-relevance analysis (Block.relevance_graph / Block.relies_on).
These are not solved for a policy; only their graphical structure is used.
Multi-agent influence diagram (MAID) illustration models.
These are game-theoretic influence diagrams from the literature, encoded as
scikit-agent blocks to illustrate and exercise strategic-relevance analysis
(Block.relevance_graph / Block.relies_on). Unlike the consumption-saving
models in benchmarks.py, these are games rather than dynamic programs: what
they pin down is graphical structure (information sets, agent ownership,
dependencies), and their functional forms and payoff magnitudes are illustrative.
The magnitudes are nonetheless chosen so that each game is strategically
non-degenerate – no decision has an optimum that is independent of the others –
so that a solver exercises the structure rather than sidestepping it.
skagent.solver.solve_in_relevance_order solves acyclic components once and
iterates simultaneous best responses for cyclic components, returning a
pure-strategy fixed point or reporting non-convergence.
Encoding conventions (a deliberate departure from the source presentations):
Chance nodes are given in structural-causal form rather than as conditional probability distributions P(node | parents): each chance node is a deterministic mechanism of its endogenous parents plus an explicit exogenous noise variable (a shock). This is equivalent in distribution (any CPD can be written as a function of its parents plus independent noise) but makes the noise a first-class graph node, matching scikit-agent’s shock/dynamics vocabulary. Because the noise nodes are single-child exogenous roots, they cannot lie on any d-connecting path and so do not change the relevance graph.
The relaxed Prisoner’s Dilemma uses continuous
[0, 1]controls.0means cooperate and1means defect; intermediate values use the multilinear extension of the standard payoff matrix. Its discrete variant limits both controls to{0, 1}. The iterated model’s utilities are per-round payoffs; accumulation or discounting across rounds belongs to the simulator or solver using the block.
AI-safety influence diagrams¶
Influence diagrams from the AI-safety literature, encoded to exercise the
graphical criteria in skagent.relevance. Like the MAIDs above, these are used
for their structure rather than solved for a policy.
Influence-diagram models from the AI-safety literature.
Models drawn from papers that use causal influence diagrams to state safety and
fairness properties, encoded as scikit-agent blocks and exercised through the
graphical criteria in skagent.relevance. Each module cites the paper its
diagrams come from and names the figures.
Kept separate from skagent.models.benchmarks, which collects economic
models with known solutions. Nothing here is a benchmark: what these models pin
down is graphical structure – who observes what, who owns which payoff, what
influences what – and their functional forms are illustrative.
Single-decision incentive diagrams (Everitt, Carey, Langlois, Ortega & Legg).
The four diagrams “Agent Incentives: A Causal Perspective” (AAAI-21,
35(13):11487-11495; arXiv:2102.01685) draws as its running examples, encoded as
scikit-agent blocks: grade prediction (Figs. 1a, 3a, 3b) and content
recommendation (Figs. 1b, 2, 4a, 4b). They exercise the criteria in
skagent.relevance – admits_voi, admits_ri, admits_voc and
admits_ici.
Each example is a PAIR: a diagram that admits an incentive, and a redesign that does not. One edge separates the members of a pair, and that edge flips one criterion while leaving another fixed, which is what makes the pairs a test of the criteria rather than an illustration of them.
Encoding conventions, following skagent.models.macid:
Chance nodes are given in structural-causal form: a deterministic mechanism of the node’s endogenous parents plus an explicit exogenous noise shock. Utility nodes are deterministic functions of their parents, as influence diagrams require.
The paper’s variables are finite-domain; here they are relaxed to continuous
[0, 1]quantities, pending discrete-variable support.Control bounds are declared as constants, so a control’s bounds do not depend on what it observes.
The relaxation is a re-encoding, not an approximation of the result: the criteria these models exercise are purely graphical, reading the diagram’s edges and never the mechanisms, so no assertion about these diagrams depends on the functional forms. The forms are nonetheless chosen so that every declared dependency has some effect, which the graphical criteria do not need but a quantitative reading of the same diagrams would.
Beside the blocks are two helpers for reading them, following the convention of
skagent.models.benchmarks, where a model ships the analysis its own shape
calls for. Both are shared by more than one caller and neither is a permanent
fixture:
print_incentive_tablerenders all four criteria for every node of one of these diagrams. The rendering conventions are local to this module – the"P"decision, theu_-prefixed noise filter, the glyphs – which is why it lives here rather than inskagent.relevance. When that module gains a general table returning{node: {criterion: bool | None}}, this function keeps its signature and its body becomes a loop over that result.draw_shocksdraws each shock of a diagram, seeded, for the numerical checks below. The diagrams are module-level values andBlock.construct_shocksleaves them alone, so nothing here needs a copy.
- skagent.models.safety.incentives.content_recommender_block = DBlock(name='content_recommender', description='', shocks={'O': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_M': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_I': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0})}, dynamics={'M': <function _content_recommender.<locals>.<lambda>>, 'P': <skagent.block.Control object>, 'I': <function _content_recommender.<locals>.<lambda>>, 'C': <function <lambda>>}, reward={'C': 'recommender'}, entity=None)¶
the user clicks on posts that match what they now think.
- Type:
Fig. 4a. Clicks depend on the posts and the user’s influenced opinions
- skagent.models.safety.incentives.content_recommender_redesign_block = DBlock(name='content_recommender_redesign', description='', shocks={'O': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_M': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_I': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0})}, dynamics={'M': <function _content_recommender.<locals>.<lambda>>, 'P': <skagent.block.Control object>, 'I': <function _content_recommender.<locals>.<lambda>>, 'C': <function <lambda>>}, reward={'C': 'recommender'}, entity=None)¶
predicted clicks, computed from the model of the user’s original opinions rather than from the opinions the posts produced.
- Type:
Fig. 4b. The redesign
- skagent.models.safety.incentives.draw_shocks(block, n=200000, seed=0)¶
n draws of each of block’s shocks.
- Returns:
A mapping from shock symbol to its draws.
- Return type:
- skagent.models.safety.incentives.grade_predictor_block = DBlock(name='grade_predictor', description='', shocks={'R': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'Ge': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_HS': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_E': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_Gr': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0})}, dynamics={'HS': <function _grade_predictor.<locals>.<lambda>>, 'E': <function _grade_predictor.<locals>.<lambda>>, 'Gr': <function _grade_predictor.<locals>.<lambda>>, 'P': <skagent.block.Control object>, 'Ac': <function _grade_predictor.<locals>.<lambda>>}, reward={'Ac': 'university'}, entity=None)¶
Fig. 3a. The prediction observes the high school and gender.
- skagent.models.safety.incentives.grade_predictor_redesign_block = DBlock(name='grade_predictor_redesign', description='', shocks={'R': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'Ge': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_HS': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_E': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0}), 'u_Gr': (<class 'skagent.distributions.Uniform'>, {'low': 0.0, 'high': 1.0})}, dynamics={'HS': <function _grade_predictor.<locals>.<lambda>>, 'E': <function _grade_predictor.<locals>.<lambda>>, 'Gr': <function _grade_predictor.<locals>.<lambda>>, 'P': <skagent.block.Control object>, 'Ac': <function _grade_predictor.<locals>.<lambda>>}, reward={'Ac': 'university'}, entity=None)¶
the prediction no longer observes the high school.
- Type:
Fig. 3b. The redesign
- skagent.models.safety.incentives.print_incentive_table(block, decision='P', criteria=None)¶
Print criteria for every node of block’s influence diagram.
Rows are the diagram’s nodes, excluding the exogenous noise shocks this module names with a
u_prefix, which answer the criteria themselves but say nothing about the diagram.n/amarks a query a criterion refuses as outside its domain rather than answers: value of information asks about observing a variable, which is undefined for the decision itself or for anything downstream of it.- Parameters:
block (DBlock) – One of this module’s diagrams.
decision (str) – The control the criteria are read against.
criteria (sequence of str, optional) – Which criteria to print, as labels drawn from
"VoI","RI","VoC"and"ICI". Defaults to all four. Narrowing it keeps a table to the criteria under discussion, since the four are independent and a reader met with all of them has three to ignore.
- Returns:
The influence graph the table was computed from, so that a caller can go on to query it.
- Return type:
- Raises:
KeyError – If criteria names something that is not one of the four.