WYE 公开读物 · 经济常用模型攻略
EN ES FR 下载完整版 · 需注册 WYE 账号
经济学核心 15 模型 · Fifteen Core Models in Economics

经济学核心 15 模型 · 交互精读Fifteen Core Models in Economics — an interactive reader

微观 5 · 宏观 5 · 计量 5。每个模型给出:一句话定位、核心方程的逐步推导(默认折叠)、可拖参数的交互图、带具体数字的教学算例、经典文献出处、以及最容易栽跟头的三处。

Five in micro, five in macro, five in econometrics. Each model comes with a one-line placement, a step-by-step derivation of its core equations (collapsed by default), an interactive chart with draggable parameters, a worked example with real numbers, the original sources, and the three places it most often goes wrong.

算例数字 85 项已用 sympy / numpy 独立验算 文献出处经联网核实(作者·年份·期刊·卷页) 2025 年诺奖(Aghion–Howitt 创造性破坏)与 DID 方法论修正
微观 · 01Microeconomics · 01

消费者最优与对偶三角Consumer optimum and the duality triangle

Utility maximisation, duality, and the Slutsky decomposition

整个需求理论只有一个优化问题,从四个角度看就成了四组结论。

Demand theory contains exactly one optimisation problem. Viewed from four angles it becomes four sets of results.

消费者理论看上去有一堆函数——马歇尔需求、补偿需求、间接效用、支出函数—— 其实它们是同一个优化问题的四个投影。搞清楚谁是谁的什么,后面的比较静态、福利分析、 生产理论全都是同一套数学换层皮。
Consumer theory appears to be a pile of separate functions — Marshallian demand, compensated demand, indirect utility, expenditure — but they are four projections of a single optimisation problem. Once you see which is which, comparative statics, welfare analysis and the whole of production theory turn out to be the same mathematics wearing different clothes.

核心方程Core equations

原问题与对偶问题
The primal and its dual
$$\max_{x,y}\ U(x,y)\ \text{ s.t. }\ p_xx+p_yy=I \qquad\Longleftrightarrow\qquad \min_{x,y}\ p_xx+p_yy\ \text{ s.t. }\ U(x,y)=\bar U$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 写拉格朗日并取一阶条件
    Set up the Lagrangian and take first-order conditions
    $$\mathcal{L}=U(x,y)+\lambda(I-p_xx-p_yy);\quad U_x=\lambda p_x,\ \ U_y=\lambda p_y$$
    两式相除消掉 \(\lambda\),得 \(MRS=\dfrac{U_x}{U_y}=\dfrac{p_x}{p_y}\)——主观替代率等于市场替代率。
    Dividing one condition by the other cancels \(\lambda\) and gives \(MRS=\dfrac{U_x}{U_y}=\dfrac{p_x}{p_y}\): the subjective rate of substitution equals the market rate.
  2. 解出马歇尔需求
    Solve for Marshallian demand
    $$x=x(p_x,p_y,I),\qquad y=y(p_x,p_y,I)$$
    代回效用得间接效用函数 \(V(p_x,p_y,I)\):给定价格与收入能达到的最高效用。
    Substituting back into utility yields the indirect utility function \(V(p_x,p_y,I)\): the highest utility attainable at given prices and income.
  3. 对偶问题给出补偿需求与支出函数
    The dual delivers compensated demand and the expenditure function
    $$x^c(p_x,p_y,\bar U),\qquad E(p_x,p_y,\bar U)$$
    \(V\) 与 \(E\) 互为反函数:\(E(p,V(p,I))=I\)、\(V(p,E(p,\bar U))=\bar U\)。这是对偶的全部内容。
    \(V\) and \(E\) are inverses of one another: \(E(p,V(p,I))=I\) and \(V(p,E(p,\bar U))=\bar U\). That is all duality amounts to.
两条恒等式
Two identities
$$\underbrace{x^c=\frac{\partial E}{\partial p_x}}_{\text{Shephard lemma}} \qquad\qquad \underbrace{x=-\frac{\partial V/\partial p_x}{\partial V/\partial I}}_{\text{Roy identity}}$$
展开逐步推导(2 步)Show the 2-step derivation
  1. Shephard 引理来自包络定理
    Shephard's lemma follows from the envelope theorem
    $$\frac{\partial E}{\partial p_x}=\frac{\partial}{\partial p_x}\bigl[p_xx^c+p_yy^c\bigr]=x^c$$
    只对参数求偏导,内生量的变化不用管——这正是包络定理。多算链式项是这里的头号错误。
    Differentiate with respect to the parameter only; the response of the endogenous variables can be ignored. That is precisely what the envelope theorem buys you, and adding the chain-rule terms anyway is the commonest error here.
  2. Roy 恒等式来自对 V 全微分
    Roy's identity follows from totally differentiating V
    $$\frac{\partial V}{\partial p_x}=-\lambda x,\qquad \frac{\partial V}{\partial I}=\lambda$$
    两式相除,\(\lambda\) 约掉。负号来自价格上升使效用下降,写掉负号会得到负的需求。
    Divide one by the other and \(\lambda\) cancels. The minus sign comes from the fact that a price rise lowers utility; drop it and demand comes out negative.
斯勒茨基方程
The Slutsky equation
$$\frac{\partial x}{\partial p_x} =\underbrace{\left.\frac{\partial x}{\partial p_x}\right|_{U=\text{const}}}_{\text{substitution}\ \le 0} -\underbrace{x\frac{\partial x}{\partial I}}_{\text{income}}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 从恒等式 xᶜ(p, Ū) = x(p, E(p, Ū)) 出发,对 pₓ 求导
    Start from the identity xᶜ(p, Ū) = x(p, E(p, Ū)) and differentiate with respect to pₓ
    $$\frac{\partial x^c}{\partial p_x}=\frac{\partial x}{\partial p_x}+\frac{\partial x}{\partial I}\cdot\frac{\partial E}{\partial p_x}$$
    右边第二项用了链式法则:\(E\) 也依赖 \(p_x\)。
    The second term on the right uses the chain rule, since \(E\) depends on \(p_x\) as well.
  2. 把 Shephard 引理代进去
    Substitute Shephard's lemma
    $$\frac{\partial E}{\partial p_x}=x^c=x \quad(\text{equal at the initial point})$$
    补偿需求与马歇尔需求只在初始点重合,这一步是整个推导的枢纽。
    Compensated and Marshallian demand coincide only at the initial point. This step is the pivot of the whole derivation.
  3. 移项
    Rearrange
    $$\frac{\partial x}{\partial p_x}=\frac{\partial x^c}{\partial p_x}-x\frac{\partial x}{\partial I}$$
    替代效应恒非正(凹性保证);收入效应符号取决于正常品还是低档品。吉芬品要求低档且收入效应压过替代效应。
    The substitution effect is never positive (concavity guarantees it); the sign of the income effect depends on whether the good is normal or inferior. A Giffen good requires an inferior good whose income effect outweighs the substitution effect.

交互图Interactive chart

价格上涨的替代效应与收入效应
Substitution and income effects of a price rise
默认参数正好复现下方算例的三个数。拖动 α 看份额如何决定效应大小。
The default settings reproduce the three numbers in the worked example below. Drag α to see how the budget share governs the size of each effect.

教学算例Worked example

\(u=\alpha\ln q+(1-\alpha)\ln Q,\ \alpha=0.3,\ y=100,\ p_y=5\);咖啡价格 \(p_x\) 由 2 涨到 4。

\(u=\alpha\ln q+(1-\alpha)\ln Q,\ \alpha=0.3,\ y=100,\ p_y=5\); the price of coffee \(p_x\) rises from 2 to 4.

原需求 \(q_0=\alpha y/p_x\)Initial demand \(q_0=\alpha y/p_x\)15.000份额固定:花在咖啡上的钱恒为 30The share is fixed: spending on coffee stays at 30
新需求 \(q_1\)New demand \(q_1\)7.500价格翻倍、支出不变 ⇒ 数量减半Price doubles, spending unchanged, so quantity halves
原效用 \(v_0\)Initial utility \(v_0\)2.659755\(=\ln y-\alpha\ln p_x-(1-\alpha)\ln p_y+\alpha\ln\alpha+(1-\alpha)\ln(1-\alpha)\)\(=\ln y-\alpha\ln p_x-(1-\alpha)\ln p_y+\alpha\ln\alpha+(1-\alpha)\ln(1-\alpha)\)
补偿需求 \(q^c\) @ 新价Compensated demand \(q^c\) at the new price9.233583把收入补到能维持 \(v_0\) 时会买的量What she would buy if income were topped up to hold \(v_0\)
替代效应Substitution effect−5.766沿同一条无差异曲线滑动,恒为负A slide along one indifference curve, always negative
收入效应Income effect−1.734实际购买力下降;咖啡是正常品故同向Real purchasing power falls; coffee is normal, so the effect runs the same way
总效应Total effect−7.500−5.766 + (−1.734),斯勒茨基恒等式成立−5.766 + (−1.734); the Slutsky identity holds
替代效应占了总效应的 77%。份额 α 越小,收入效应越小——这是「份额即权重」的直接含义。
The substitution effect accounts for 77% of the total. The smaller α is, the smaller the income effect — which is what "the share is the weight" means in practice.

经典文献Original sources

  • Slutsky, E. (1915). "Sulla teoria del bilancio del consumatore." Giornale degli Economisti.替代–收入分解的原始文献,被埋没近 20 年才由 Hicks 与 Allen 重新发现。The original source of the substitution–income decomposition. It lay buried for nearly twenty years until Hicks and Allen rediscovered it.
  • Hicks, J. R. & Allen, R. G. D. (1934). "A Reconsideration of the Theory of Value." Economica.把序数效用与替代弹性带进主流,Slutsky 分解由此广为人知。Brought ordinal utility and the elasticity of substitution into the mainstream, and with them the Slutsky decomposition.
  • Roy, R. (1947). "La distribution du revenu entre les divers biens." Econometrica 15, 205–225.Roy 恒等式。Roy's identity.
  • Shephard, R. W. (1953). Cost and Production Functions. Princeton University Press.Shephard 引理,原本是为成本函数写的——见模型 2,两者是同一件事。Shephard's lemma, originally written for cost functions — see Model 2, where it is the same result.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把补偿需求与马歇尔需求全程当成相等。它们只在初始点重合。斯勒茨基推导的枢纽正是「在初始点 \(x^c=x\)」这一步,用早了或用晚了结论都会错。Treating compensated and Marshallian demand as equal throughout. They coincide only at the initial point. The pivot of the Slutsky derivation is exactly the step where \(x^c=x\) holds; invoke it too early or too late and the result is wrong.
  2. Roy 恒等式忘负号。\(\partial V/\partial p_x\lt0\),分母 \(\partial V/\partial I\gt0\),不加负号会得到负需求。Dropping the minus sign in Roy's identity. \(\partial V/\partial p_x\lt0\) while \(\partial V/\partial I\gt0\); without the minus sign demand comes out negative.
  3. 用包络定理时多算链式项。\(\partial E/\partial p_x\) 只对显式出现的 \(p_x\) 求导,\(x^c\) 随 \(p_x\) 的变化不用管——这是包络定理的全部好处,多算就白用了。Adding chain-rule terms when using the envelope theorem. \(\partial E/\partial p_x\) differentiates only the \(p_x\) that appears explicitly; the response of \(x^c\) to \(p_x\) can be ignored. That is the entire benefit of the theorem, and adding the terms throws it away.
知识勾连 Cross-links 本节的 \(MRS\)、\(\sigma\)、无差异曲线,与模型 2 的 \(RTS\)、等产量线是同一套数学换层皮。The \(MRS\), \(\sigma\) and indifference curves of this section are the same mathematics as the \(RTS\) and isoquants of Model 2, wearing different clothes.
微观 · 02Microeconomics · 02

生产、成本与 CES 替代弹性Production, cost duality and the CES family

Production, cost duality, and the CES family

一个参数 σ 就把 Leontief、柯布–道格拉斯、线性生产函数串成一条谱。

One parameter, σ, strings Leontief, Cobb–Douglas and linear technology onto a single spectrum.

生产理论几乎是消费者理论的镜像:等产量线对无差异曲线、RTS 对 MRS、 成本函数对支出函数、Shephard 引理两边通用。真正的新东西只有一个:替代弹性 σ, 它把所有常见生产函数排成一条连续的谱。
Production theory is almost a mirror of consumer theory: isoquants for indifference curves, RTS for MRS, cost function for expenditure function, and Shephard's lemma working on both sides. Only one genuinely new object appears — the elasticity of substitution σ — and it arranges every familiar production function along a continuous spectrum.

核心方程Core equations

CES 生产函数与替代弹性
The CES production function and the elasticity of substitution
$$q=\bigl(aK^{\rho}+bL^{\rho}\bigr)^{1/\rho},\qquad \sigma=\frac{1}{1-\rho}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 成本最小化的一阶条件
    First-order condition for cost minimisation
    $$\frac{q_K}{q_L}=\frac{r}{w}\ \Longrightarrow\ \frac{a}{b}\Bigl(\frac{K}{L}\Bigr)^{\rho-1}=\frac{r}{w}$$
    \(RTS\) 等于要素价格比——与消费者的 \(MRS=p_x/p_y\) 一模一样。
    The \(RTS\) equals the factor price ratio, exactly as \(MRS=p_x/p_y\) does for the consumer.
  2. 两边取对数,把指数拿下来
    Take logs of both sides to bring the exponent down
    $$(\rho-1)\ln\frac{K}{L}=\ln\frac{r}{w}-\ln\frac{a}{b}$$
    用 \(\ln x^n=n\ln x\)。这一步是为了让下一步能直接读出弹性。
    Using \(\ln x^n=n\ln x\). The point of this step is to let the next one read the elasticity straight off.
  3. 解出对数比值,再对 ln(w/r) 求导
    Solve for the log ratio, then differentiate with respect to ln(w/r)
    $$\ln\frac{K}{L}=\frac{1}{\rho-1}\Bigl(\ln\frac{r}{w}-\ln\frac{a}{b}\Bigr) \ \Longrightarrow\ \frac{d\ln(K/L)}{d\ln(w/r)}=\frac{1}{1-\rho}=\sigma$$
    \(\ln(r/w)=-\ln(w/r)\),负号与 \(\rho-1\) 的负号相消。σ 的定义就是这个对数导数,不是别的。
    \(\ln(r/w)=-\ln(w/r)\), and that minus sign cancels the one in \(\rho-1\). σ is defined as this log derivative and as nothing else.
σ 谱上的四个点
Four points on the σ spectrum
$$\sigma=0\ (\text{Leontief})\ \to\ \sigma=1\ (\text{Cobb–Douglas}) \ \to\ \sigma\gt1\ \to\ \sigma\to\infty\ (\text{linear})$$
展开逐步推导(2 步)Show the 2-step derivation
  1. ρ→0 的极限是柯布–道格拉斯
    The ρ→0 limit is Cobb–Douglas
    $$\lim_{\rho\to0}\bigl(aK^{\rho}+bL^{\rho}\bigr)^{1/\rho}=K^{a}L^{b}\quad(a+b=1)$$
    先确认是 \(1^{\infty}\) 型未定式,取对数化成 \(0/0\) 再用洛必达。跳过这一步直接写结论是这节最常见的偷工。
    Confirm first that this is a \(1^{\infty}\) indeterminate form, take logs to turn it into \(0/0\), then apply LHôpital. Skipping straight to the answer is the standard shortcut taken here — and the standard place marks are lost.
  2. ρ→−∞ 给 Leontief,ρ→1 给线性
    ρ→−∞ gives Leontief, ρ→1 gives the linear case
    $$\sigma=\tfrac{1}{1-\rho}:\quad \rho\to-\infty\Rightarrow\sigma\to0;\quad \rho\to1\Rightarrow\sigma\to\infty$$
    σ 的经济含义:要素价格比变动 1%,要素投入比变动 σ%。σ=0 表示完全不可替代。
    What σ means economically: a 1% change in the factor price ratio changes the factor input ratio by σ%. σ=0 means the factors cannot be substituted at all.

交互图Interactive chart

σ 决定等产量线的形状
σ determines the shape of the isoquant
σ→0 折成直角(Leontief),σ=1 是柯布–道格拉斯,σ→∞ 拉成直线。同一条谱上的四个熟面孔。
σ→0 folds into a right angle (Leontief), σ=1 is Cobb–Douglas, σ→∞ straightens into a line. Four familiar faces on one spectrum.

教学算例Worked example

\(a=b=\tfrac12\)。工资租金比 \(w/r\) 翻倍,问资本–劳动比 \(K/L\) 变多少。

\(a=b=\tfrac12\). The wage–rental ratio \(w/r\) doubles: by how much does \(K/L\) change?

ρ = −1 ⇒ σ = 0.5ρ = −1 ⇒ σ = 0.5×1.414\(2^{0.5}\):替代困难,比例只涨 41%\(2^{0.5}\): substitution is hard, so the ratio rises only 41%
ρ = 0 ⇒ σ = 1ρ = 0 ⇒ σ = 1×2.000柯布–道格拉斯:等比例替代Cobb–Douglas: proportional substitution
ρ = 0.5 ⇒ σ = 2ρ = 0.5 ⇒ σ = 2×4.000\(2^{2}\):替代容易,比例翻两番\(2^{2}\): substitution is easy, so the ratio quadruples
ρ → 0 的极限The ρ → 0 limit\(\sqrt{KL}\)sympy 验证:CES 确实退化成柯布–道格拉斯Verified in sympy: CES does collapse to Cobb–Douglas
σ 就是「K/L 对 w/r 的弹性」,不是别的比值。记住这一句,四种函数形式就不用死记。
σ is the elasticity of K/L with respect to w/r and not any other ratio. Hold on to that one sentence and the four functional forms need not be memorised.

经典文献Original sources

  • Cobb, C. W. & Douglas, P. H. (1928). "A Theory of Production." American Economic Review 18(1), 139–165.柯布–道格拉斯函数的原始文献,用 1899–1922 年美国制造业数据拟合。The original Cobb–Douglas paper, fitted to US manufacturing data for 1899–1922.
  • Arrow, K. J., Chenery, H. B., Minhas, B. S. & Solow, R. M. (1961). "Capital-Labor Substitution and Economic Efficiency." Review of Economics and Statistics 43(3), 225–250.CES 生产函数的出处(常简称 ACMS)。四位作者里两位后来得了诺奖。The source of the CES production function, usually abbreviated ACMS. Two of the four authors later won the Nobel Prize.
  • Shephard, R. W. (1953). Cost and Production Functions.\(\partial C/\partial w=L^c\):与消费者那边的 \(\partial E/\partial p=x^c\) 是同一条引理。\(\partial C/\partial w=L^c\) is the same lemma as \(\partial E/\partial p=x^c\) on the consumer side.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. σ 与 ρ 的关系记成 \(1/(1+\rho)\)。正确是 \(\sigma=\dfrac{1}{1-\rho}\);\(1/(1+\rho)\) 是把 CES 写成 \(\rho\) 在分母那个版本时的形式,两种写法混用必错。Remembering the relation as \(\sigma=1/(1+\rho)\). It is \(\sigma=\dfrac{1}{1-\rho}\); \(1/(1+\rho)\) belongs to the version of CES written with \(\rho\) in the denominator, and mixing the two conventions guarantees an error.
  2. ρ→0 直接断言等于柯布–道格拉斯。这是 \(1^{\infty}\) 未定式,必须取对数化成 \(0/0\) 再用洛必达。省掉这一步,遇到「证明」类题就丢分。Asserting the ρ→0 limit without proof. It is a \(1^{\infty}\) indeterminate form and needs logs plus LHôpital. Skip the step and any "show that" question is lost.
  3. 把 σ 理解成「产出对要素的弹性」。σ 是要素比对价格比的弹性,与产出弹性(柯布–道格拉斯里的 \(a,b\))是两回事。Reading σ as an output elasticity. σ is the elasticity of the factor ratio with respect to the price ratio, quite distinct from the output elasticities \(a\) and \(b\) in Cobb–Douglas.
知识勾连 Cross-links 与模型 1 完全同构:等产量线↔无差异曲线、\(RTS\)↔\(MRS\)、成本函数↔支出函数。学过一边,另一边只需换名词。Fully isomorphic to Model 1: isoquant ↔ indifference curve, \(RTS\) ↔ \(MRS\), cost function ↔ expenditure function. Learn one side and the other needs only new nouns.
微观 · 03Microeconomics · 03

一般均衡与两个福利定理General equilibrium and the two welfare theorems

General equilibrium and the two welfare theorems

把所有市场同时出清写成一个不动点问题——这是「看不见的手」的严格版本。

Writing "all markets clear at once" as a fixed-point problem — the rigorous version of the invisible hand.

局部均衡假装一个市场的变化不影响别的市场。一般均衡不作这个假设: 所有价格同时决定,所有市场同时出清。它的两条福利定理是整个经济学最强的规范性结论, 也是所有政策讨论的底线参照。
Partial equilibrium pretends that what happens in one market leaves the others alone. General equilibrium makes no such assumption: all prices are determined together and all markets clear together. Its two welfare theorems are the strongest normative results in economics and the benchmark against which every policy argument is measured.

核心方程Core equations

竞争均衡的定义(必须写全)
Competitive equilibrium (state all three conditions)
$$\bigl(\{x^i\},\{y^j\},p\bigr):\quad \text{(1) consumers optimise}\quad \text{(2) firms maximise profit}\quad \text{(3) markets clear}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 消费者问题
    The consumer's problem
    $$x^i\in\arg\max\ u^i(x)\ \text{ s.t. }\ p\cdot x\le p\cdot\omega^i+\textstyle\sum_j\theta^{ij}\pi^j$$
    注意收入不是外生的:它由禀赋的市场价值加上企业利润的份额内生决定。
    Income is not exogenous: it is the market value of the endowment plus a share of firm profits, both determined within the model.
  2. 企业问题
    The firm's problem
    $$y^j\in\arg\max\ p\cdot y\ \text{ s.t. }\ y\in Y^j$$
    \(Y^j\) 是生产可能集。
    \(Y^j\) is the production possibility set.
  3. 市场出清
    Market clearing
    $$\sum_i x^i=\sum_i\omega^i+\sum_j y^j$$
    三条缺一条就不是均衡,而不是「写得简略」。这是判断题与论述题的固定扣分点。
    Drop any one of the three and it is not an equilibrium — not merely a condensed statement of one. Examiners take marks off for this with great regularity.
瓦尔拉斯定律与两个福利定理
Walras' law and the two welfare theorems
$$p\cdot z(p)\equiv 0\quad\forall p \qquad\Longrightarrow\qquad \text{only }n-1\text{ markets need clear}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 瓦尔拉斯定律
    Walras' law
    $$z(p)=\textstyle\sum_i x^i(p)-\sum_i\omega^i-\sum_j y^j(p),\qquad p\cdot z(p)=0$$
    来自每个人的预算约束取等号。后果:一个市场出清 ⇒ 另一个自动出清(两商品情形),所以只能定相对价格,绝对价格水平不定(需要标准化)。
    It follows from each budget constraint holding with equality. Consequence: if one market clears the other clears automatically (with two goods), so only relative prices are determined and the absolute level needs a normalisation.
  2. 第一福利定理
    First welfare theorem
    $$\text{competitive equilibrium}\ \Longrightarrow\ \text{Pareto optimal}$$
    只需要偏好局部非饱和。不需要凸性、不需要连续。这是「看不见的手」的严格陈述。
    All it requires is local non-satiation of preferences — not convexity, not continuity. This is the invisible hand stated rigorously.
  3. 第二福利定理
    Second welfare theorem
    $$\text{Pareto optimal}\ \Longrightarrow\ \text{prices + lump-sum transfers}$$
    要求凸性(偏好凸、技术凸)。它把「效率」与「公平」分开:先转移禀赋,再让市场干活。
    This one does require convexity of preferences and technology. It separates efficiency from equity: redistribute endowments first, then let the market work.

交互图Interactive chart

埃奇沃斯盒:禀赋、契约曲线与均衡
Edgeworth box: endowment, contract curve and equilibrium
一个点是禀赋,另一个是均衡;一条虚线是预算线,另一条是契约曲线(所有帕累托最优点)。
One marker is the endowment, the other the equilibrium; one dashed line is the budget line, the other the contract curve (the locus of Pareto optima).

教学算例Worked example

两人两货,\(u^A=u^B=\sqrt{xy}\);A 禀赋 \((10,0)\),B 禀赋 \((0,10)\)。令 \(p_y=1\),求 \(p=p_x\)。

Two agents, two goods, \(u^A=u^B=\sqrt{xy}\); A is endowed with \((10,0)\), B with \((0,10)\). Normalise \(p_y=1\) and solve for \(p=p_x\).

A 的收入A's income\(10p\)禀赋的市场价值The market value of the endowment
A 的需求A's demand\(x_A=5,\ y_A=5p\)份额 ½ ⇒ 各花一半With a share of ½, half of income goes to each good
B 的需求B's demand\(x_B=5/p,\ y_B=5\)
x 市场出清Clearing the x market\(5+5/p=10\)解得 p = 1gives p = 1
y 市场The y market\(5p+5=10\ \checkmark\)自动出清 —— 瓦尔拉斯定律clears automatically — Walras' law
均衡配置Equilibrium allocationA=(5,5), B=(5,5)
A 的效用增益A's utility gain0 → 5交易前 \(\sqrt{10\times0}=0\),交易后 5Before trade \(\sqrt{10\times0}=0\); after trade, 5
交易前双方效用都是 0,交易后都是 5。这就是第一福利定理最赤裸的版本:价格把两个人从饿死边缘带到帕累托最优。
Both agents have utility 0 before trade and 5 after. This is the first welfare theorem at its barest: prices carry two people from the edge of starvation to a Pareto optimum.

经典文献Original sources

  • Arrow, K. J. & Debreu, G. (1954). "Existence of an Equilibrium for a Competitive Economy." Econometrica 22(3), 265–290.存在性证明,用角谷不动点定理。两位作者分别在 1972、1983 年获诺奖。The existence proof, via Kakutani's fixed-point theorem. The authors received the Nobel Prize in 1972 and 1983 respectively.
  • Debreu, G. (1959). Theory of Value. Cowles Foundation Monograph 17.公理化的完整体系,全书不到 100 页,至今是范本。The complete axiomatic treatment, under 100 pages, still a model of exposition.
  • Sonnenschein (1972), Mantel (1974), Debreu (1974).SMD 定理:超额需求函数几乎可以是任意形状 ⇒ 一般均衡不保证唯一性与稳定性。这是对该框架最重要的限定。The SMD theorem: excess demand functions can take almost any shape, so general equilibrium guarantees neither uniqueness nor stability. This is the most important qualification the framework carries.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 均衡只写「供给=需求」。竞争均衡是「一组配置 一组价格」,且必须同时满足消费者最优、企业最优、市场出清三条。少写一条就不是均衡的定义。Writing equilibrium as "supply equals demand". A competitive equilibrium is an allocation together with a price vector, satisfying consumer optimisation, firm optimisation and market clearing simultaneously. Omit one and you have not stated the definition.
  2. 忘了价格只能定到相对水平。瓦尔拉斯定律使方程组少一个独立方程,必须标准化(令某个价格为 1),否则会以为「方程比未知数少一个,无解」。Forgetting that prices are determined only up to scale. Walras' law removes one independent equation, so a normalisation is required (set one price to 1). Without it the system looks underdetermined.
  3. 把第二福利定理当成「市场自动实现公平」。它要求一次性转移支付先把禀赋挪好,而现实中的转移几乎总是扭曲性的。这是理论与政策之间最大的裂缝。Reading the second welfare theorem as "markets deliver fairness on their own". It requires lump-sum transfers to move endowments first, and real-world transfers are almost always distortionary. This is the widest gap between the theory and any policy built on it.
知识勾连 Cross-links 模型 4 放弃价格接受者假设 ⇒ 均衡不再帕累托最优;模型 5 放弃信息对称 ⇒ 市场可能直接消失。Model 4 drops price-taking and the equilibrium ceases to be Pareto optimal; Model 5 drops symmetric information and the market may vanish altogether.
微观 · 04Microeconomics · 04

纳什均衡:古诺与伯特兰Nash equilibrium: Cournot and Bertrand

Nash equilibrium, Cournot and Bertrand competition

每个人对别人的策略做最优反应,且所有人的信念自洽——交点就是均衡。

Everyone best-responds to everyone else and all beliefs are mutually consistent — the intersection is the equilibrium.

完全竞争与垄断是两个极端。中间地带(寡头)没法用「价格接受者」处理, 因为每家的决策会改变别家的最优选择。纳什均衡就是让这种相互依赖自洽的解概念, 它是现代产业组织、拍卖、契约、宏观协调失灵模型的公共语言。
Perfect competition and monopoly are the two extremes. The ground between them — oligopoly — cannot be handled with price-taking, because each firm's decision shifts every other firm's best choice. Nash equilibrium is the solution concept that makes this interdependence consistent, and it is the shared language of modern industrial organisation, auction theory, contract theory and macroeconomic coordination-failure models.

核心方程Core equations

纳什均衡与最优反应
Nash equilibrium and best responses
$$s^*\ \text{is Nash}\iff \forall i:\ s_i^*\in\arg\max_{s_i}\ \pi_i(s_i,s_{-i}^*)$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 定义最优反应函数
    Define the best-response function
    $$BR_i(s_{-i})=\arg\max_{s_i}\pi_i(s_i,s_{-i})$$
    均衡就是所有最优反应函数的不动点:\(s^*\in BR(s^*)\)。
    An equilibrium is a fixed point of the best-response correspondence: \(s^*\in BR(s^*)\).
  2. 存在性
    Existence
    $$\text{finite game}\ \Longrightarrow\ \text{mixed-strategy Nash equilibrium exists}$$
    Nash (1950) 用角谷不动点定理证明——与 Arrow–Debreu 用的是同一件工具
    Nash (1950) proved it with Kakutani's fixed-point theorem — the same instrument Arrow and Debreu used.
古诺双寡头的完整推导
The Cournot duopoly, worked through
$$P=a-b(q_1+q_2),\qquad c_1=c_2=c$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 写企业 1 的利润
    Write down firm 1's profit
    $$\pi_1=\bigl[a-b(q_1+q_2)-c\bigr]q_1=(a-c)q_1-bq_1^2-bq_1q_2$$
    展开是为了下一步求导时不出错——括号里含 \(q_1\),直接对乘积求导容易漏项。
    Expanding first avoids slips at the next step: \(q_1\) sits inside the bracket, and differentiating the product directly invites a dropped term.
  2. 对 q₁ 求导得最优反应
    Differentiate with respect to q₁ for the best response
    $$\frac{\partial\pi_1}{\partial q_1}=(a-c)-2bq_1-bq_2=0\ \Longrightarrow\ q_1=\frac{a-c-bq_2}{2b}$$
    \(-2bq_1\) 里的 2 来自 \(q_1^2\) 求导。这个 2 是古诺与垄断差别的全部来源。
    The 2 in \(-2bq_1\) comes from differentiating \(q_1^2\). That single 2 is the entire difference between Cournot and monopoly.
  3. 对称性 ⇒ 令 q₁ = q₂ = q
    Impose symmetry: q₁ = q₂ = q
    $$q=\frac{a-c-bq}{2b}\ \Longrightarrow\ 2bq+bq=a-c\ \Longrightarrow\ q^*=\frac{a-c}{3b}$$
    〔移项:两边同乘 \(2b\),把 \(bq\) 挪到左边合并成 \(3bq\)〕分母上的 3 = n+1,\(n\) 家企业时 \(q^*=\dfrac{a-c}{(n+1)b}\),\(n\to\infty\) 退化到完全竞争。
    〔Rearranging: multiply through by \(2b\), move \(bq\) to the left and collect into \(3bq\)〕The 3 in the denominator is n+1: with \(n\) firms \(q^*=\dfrac{a-c}{(n+1)b}\), which tends to the competitive outcome as \(n\to\infty\).
伯特兰悖论
The Bertrand paradox
$$\text{homogeneous goods, price competition, no capacity limit}\ \Longrightarrow\ p^*=c$$
展开逐步推导(1 步)Show the 1-step derivation
  1. 为什么
    Why
    $$p_i\gt c\ \Rightarrow\ \text{rival cuts by }\varepsilon\text{ and takes the whole market}$$
    古诺(选产量)与伯特兰(选价格)给出完全不同的结论——这说明结论对「策略变量是什么」极度敏感,而不是对「有几家企业」敏感。
    Cournot (quantity setting) and Bertrand (price setting) give completely different answers. The conclusion is acutely sensitive to what the strategic variable is, and hardly at all to how many firms there are.

交互图Interactive chart

最优反应曲线的交点
Where the best-response curves cross
两条向下倾斜的直线是各自的最优反应,交点是纳什均衡。成本不对称时低成本方产量更大。
The two downward-sloping lines are the best responses; their intersection is the Nash equilibrium. With asymmetric costs the low-cost firm produces more.

教学算例Worked example

线性需求,两家对称,边际成本 10。

Linear demand, two symmetric firms, marginal cost 10.

古诺:每家产量Cournot: output per firm30\((a-c)/3b=90/3\)\((a-c)/3b=90/3\)
古诺:总产量 / 价格Cournot: total output / price60 / 40
古诺:每家利润Cournot: profit per firm900\((40-10)\times30\)\((40-10)\times30\)
垄断:产量 / 价格Monopoly: output / price45 / 55\((a-c)/2b\)\((a-c)/2b\)
垄断:总利润Monopoly: total profit2025比两家古诺加总的 1800 更高higher than the 1800 the two Cournot firms make between them
完全竞争:产量 / 价格Perfect competition: output / price90 / 10\(P=c\)\(P=c\)
古诺产量(60)夹在垄断(45)与竞争(90)之间。垄断总利润 2025 > 古诺两家合计 1800,正是卡特尔有动机、又难以维持的根源:合谋有利可图,但每家都想偷偷多产。
Cournot output (60) sits between monopoly (45) and competition (90). Monopoly profit of 2025 exceeds the Cournot pair's 1800, which is exactly why cartels are tempting and unstable at once: collusion pays, and every member wants to cheat on it.

经典文献Original sources

  • Cournot, A. A. (1838). Recherches sur les principes mathématiques de la théorie des richesses.比纳什早 112 年就写出了这个均衡——数量竞争的原型。Written 112 years before Nash, and already the prototype of quantity competition.
  • Bertrand, J. (1883). "Théorie mathématique de la richesse sociale." Journal des Savants.对古诺的书评,指出改选价格结论就翻转。A review of Cournot pointing out that switching to price competition reverses the conclusion.
  • Nash, J. F. (1950). "Equilibrium Points in N-Person Games." PNAS 36(1), 48–49;(1951) Annals of Mathematics 54, 286–295.两页的 PNAS 短文加一篇年鉴长文,奠定了整个非合作博弈论。1994 年诺奖。A two-page note in PNAS and a paper in the Annals founded the whole of non-cooperative game theory. Nobel Prize 1994.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把最优反应的一阶条件当成均衡。\(q_1=BR_1(q_2)\) 只是一条曲线;均衡是两条曲线的交点,必须联立。只解一条就报答案是最常见的失分。Mistaking a first-order condition for the equilibrium. \(q_1=BR_1(q_2)\) is only a curve; the equilibrium is where two curves cross, and the pair must be solved jointly. Reporting one curve as the answer is the commonest way marks are lost.
  2. 混合策略均衡不用「让对手无差异」的条件。求混合策略时,均衡概率由使对手在其纯策略间无差异的条件决定,而不是由自己的收益最大化决定——这个反直觉之处是考试重灾区。Not using the indifference condition for mixed strategies. Equilibrium probabilities are pinned down by making the opponent indifferent across their pure strategies, not by maximising your own payoff. The counter-intuitiveness of this is why it goes wrong so often.
  3. 以为「企业越多越竞争」是普遍规律。伯特兰里两家就够把价格压到边际成本;古诺里要 \(n\to\infty\)。结论取决于策略变量,不取决于家数。Believing that more firms always means more competition. Two suffice under Bertrand to drive price to marginal cost; under Cournot it takes \(n\to\infty\). The answer depends on the strategic variable, not the headcount.
知识勾连 Cross-links 与模型 3 对照:放弃价格接受者假设后,均衡仍存在(Nash),但第一福利定理失效——古诺均衡不是帕累托最优。Set against Model 3: once price-taking is dropped an equilibrium still exists (Nash), but the first welfare theorem fails — the Cournot equilibrium is not Pareto optimal.
微观 · 05Microeconomics · 05

信息不对称:柠檬市场与信号Asymmetric information: lemons and signalling

Adverse selection, signalling, and screening

当一方比另一方更了解商品质量,市场可能不是定价偏了,而是整个消失。

When one side knows more about quality than the other, the market may not merely misprice — it may disappear.

前四个模型都假设信息对称。放松这一条,结论不是「效率损失一点」, 而是市场可能完全崩溃——即使人人理性、人人守法、交易明明有利可图。 这一支撑起了保险、劳动力市场、信贷、二手车、医疗几乎所有现实市场的分析。
The first four models all assume symmetric information. Relax that assumption and the result is not a modest efficiency loss but the possible collapse of the market altogether — with everyone rational, everyone honest, and gains from trade plainly available. This one strand underpins the analysis of insurance, labour markets, credit, used cars and health care.

核心方程Core equations

逆向选择:柠檬市场
Adverse selection: the market for lemons
$$\theta\sim U[0,1];\quad \text{seller value}=\theta,\quad \text{buyer value}=k\theta\ (k\gt1)$$
展开逐步推导(4 步)Show the 4-step derivation
  1. 给定价格 p,谁愿意卖
    At a price p, who is willing to sell
    $$\text{sell}\iff \theta\le p$$
    保留值低于价格才卖。关键:愿意卖的恰恰是质量差的那批——这就是「逆向」二字的由来。
    Only those whose reservation value falls below the price. The sellers who come forward are precisely the ones with the worst goods — hence the word "adverse".
  2. 成交品的平均质量
    Average quality of what actually trades
    $$\mathbb{E}[\theta\mid\theta\le p]=\frac{p}{2}$$
    均匀分布截断后的均值是区间中点。
    Truncating a uniform distribution leaves the midpoint of the remaining interval.
  3. 买方的支付意愿
    The buyer's willingness to pay
    $$\text{WTP}=k\cdot\frac{p}{2}$$
    买方理性预期到自己买到的是次品,据此出价。
    The buyer rationally anticipates receiving a poor unit and bids accordingly.
  4. 市场存活的条件
    The condition for the market to survive
    $$k\cdot\frac{p}{2}\ge p\iff k\ge 2$$
    即使 \(k\gt1\)(每一笔交易都有得益),只要 \(k\lt2\) 市场就全面崩溃。不是萎缩,是唯一均衡为零交易。
    Even with \(k\gt1\), so that every single trade creates surplus, any \(k\lt2\) collapses the market entirely. Not shrinks it — the unique equilibrium is zero trade.
信号与筛选
Signalling and screening
$$\text{signalling: informed moves first}\qquad \text{screening: uninformed moves first}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 单交叉条件是信号有效的前提
    Single crossing is what makes a signal work
    $$\frac{\partial}{\partial\theta}\Bigl(\frac{\partial c(e,\theta)/\partial e}{\partial u/\partial w}\Bigr)\lt0$$
    高能力者获取信号的边际成本更低。没有这一条,任何信号都无法分离类型——教育之所以能当信号,靠的是「学得动」而不是「学到了什么」。
    The high type must find the signal cheaper at the margin. Without that condition no signal can separate types — education works as a signal because the able find it easier, not because of anything they learn.
  2. 分离均衡
    Separating equilibrium
    $$e^*:\ w_H-c(e^*,\theta_L)\le w_L$$
    分离均衡通常有连续多个(不唯一),且社会性浪费:教育本身不提高生产率,纯粹烧钱证明自己。
    Separating equilibria are typically a continuum rather than unique, and they are socially wasteful: education raises no one\'s productivity here and is pure expenditure on proving a point.
  3. Rothschild–Stiglitz 的负面结论
    The negative result of Rothschild–Stiglitz
    $$\text{no pooling equilibrium; separating may fail to exist}$$
    竞争性保险市场可能根本没有均衡——这是对「市场总能出清」最锋利的反例。
    A competitive insurance market may have no equilibrium at all — the sharpest available counterexample to the presumption that markets clear.

交互图Interactive chart

柠檬市场:买方估值倍数 k 与市场存活
Lemons: the buyer's valuation multiple k and market survival
一条是价格 p(卖方要价),另一条是买方在理性预期下的支付意愿 kp/2。后者在前者之上市场才存活。
One line is the price p that sellers ask; the other is the buyer's willingness to pay, kp/2, under rational expectations. The market survives only while the second lies above the first.

教学算例Worked example

质量 \(\theta\sim U[0,1]\),卖方保留值 \(\theta\),买方估值 \(k\theta\)。

Quality \(\theta\sim U[0,1]\), seller reservation value \(\theta\), buyer valuation \(k\theta\).

k = 1.5k = 1.5WTP = 0.75p < p市场崩溃——尽管每笔交易都创造 50% 的剩余Market collapses — even though every trade creates 50% surplus
k = 2.0k = 2.0WTP = 1.00p = p临界:任意 p 都是均衡(退化)Knife-edge: any p is an equilibrium (degenerate)
k = 2.5k = 2.5WTP = 1.25p > p市场存活,全部质量都能成交Market survives; every quality level trades
判据是 k ≥ 2,不是 k > 1。交易有得益(k>1)远不足以让市场存在。这个「1 与 2 之间的鸿沟」就是信息不对称的全部代价。
The criterion is k ≥ 2, not k > 1. Gains from trade are nowhere near sufficient for a market to exist. The gap between 1 and 2 is the whole cost of asymmetric information.

经典文献Original sources

  • Akerlof, G. A. (1970). "The Market for Lemons: Quality Uncertainty and the Market Mechanism." Quarterly Journal of Economics 84(3), 488–500.被数家顶刊拒稿后才发表,理由包括「太琐碎」。2001 年诺奖。Rejected by several leading journals before publication, on grounds that included triviality. Nobel Prize 2001.
  • Spence, M. (1973). "Job Market Signaling." Quarterly Journal of Economics 87(3), 355–374.教育作为信号:即使教育完全不提高生产率,高能力者仍愿意购买它。Education as a signal: the able buy it even when it raises productivity not at all.
  • Rothschild, M. & Stiglitz, J. (1976). "Equilibrium in Competitive Insurance Markets." QJE 90(4), 629–649.筛选(screening):不知情方设计合同菜单让对方自选。给出了均衡可能不存在的著名结果。Screening: the uninformed side designs a menu of contracts and lets the informed side sort itself. Contains the celebrated non-existence result.
  • Holmström, B. (1979). "Moral Hazard and Observability." Bell Journal of Economics 10(1), 74–91.道德风险与充分统计量原理——逆向选择之外的另一半信息经济学。2016 年诺奖。Moral hazard and the sufficient statistic result — the other half of information economics. Nobel Prize 2016.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把逆向选择与道德风险混为一谈。逆向选择是签约前的隐藏信息(不知道对方是什么类型);道德风险是签约后的隐藏行动(看不见对方做了什么)。对策完全不同:前者靠信号/筛选,后者靠激励合同。Conflating adverse selection with moral hazard. Adverse selection is hidden information before contracting (you do not know the type); moral hazard is hidden action after contracting (you cannot see what they do). The remedies differ entirely: signalling and screening for the first, incentive contracts for the second.
  2. 以为「有交易得益市场就会存在」。柠檬模型的核心正是反例:\(k=1.5\) 时每笔交易都有 50% 的剩余,市场却完全消失。存在得益 ≠ 存在均衡。Assuming that gains from trade imply a market. The lemons model is the standing counterexample: at \(k=1.5\) every trade would create 50% surplus and the market vanishes anyway. Gains from trade ≠ existence of equilibrium.
  3. 忘记买方是理性预期的。常见错误是让买方按平均质量 \(\mathbb{E}[\theta]=0.5\) 出价。不对——买方知道只有 \(\theta\le p\) 的会卖,所以按截断后的条件均值 \(p/2\) 出价。这一步是整个模型的引擎。Forgetting that the buyer has rational expectations. The usual slip is to let the buyer bid on unconditional average quality \(\mathbb{E}[\theta]=0.5\). No — the buyer knows only \(\theta\le p\) is offered and bids on the truncated conditional mean \(p/2\). That step is the engine of the whole model.
知识勾连 Cross-links 模型 3 的第一福利定理默认信息对称。这里说明:信息不对称不是给它打个折扣,而是让「均衡存在」本身都成问题。The first welfare theorem of Model 3 quietly assumes symmetric information. What this model shows is not a discount on that result but a threat to the existence of equilibrium itself.
宏观 · 06Macroeconomics · 06

Solow 增长模型The Solow growth model

The Solow–Swan growth model

资本积累有收益递减,所以长期增长必须来自技术——这是增长论的出发点与自我否定。

Capital accumulation runs into diminishing returns, so long-run growth must come from technology — the starting point of growth theory and its own refutation.

Solow 模型的伟大之处在于它证明了自己不够用:在收益递减下, 储蓄率再高也只能提高收入水平、不能提高长期增长率。长期人均增长率完全由外生的技术进步 g 决定。 这个「余值」逼出了后来的内生增长理论(模型 8)。
What makes Solow's model great is that it demonstrates its own insufficiency: under diminishing returns, however high the saving rate, it can raise the level of income but not the long-run growth rate, which is set entirely by exogenous technical progress g. That residual is what forced the endogenous growth literature into existence (Model 8).

核心方程Core equations

资本积累方程
The capital accumulation equation
$$\dot k=sf(k)-(\delta+n+g)k,\qquad f(k)=k^{\alpha}$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 从总量到有效人均
    From aggregates to effective labour units
    $$k\equiv\frac{K}{AL},\qquad \frac{\dot k}{k}=\frac{\dot K}{K}-n-g$$
    口径必须钉死:\(k\) 是「有效劳动人均资本」,不是人均资本,也不是总资本。差一个口径,后面的 \(\delta+n+g\) 就写不对。
    Fix the units before anything else: \(k\) is capital per unit of effective labour, not per worker and not the aggregate. Get the units wrong and \(\delta+n+g\) comes out wrong.
  2. 写出净积累
    Write net accumulation
    $$\dot k=\underbrace{sf(k)}_{\text{investment}}-\underbrace{(\delta+n+g)k}_{\text{break-even}}$$
    \(\delta\) 折旧、\(n\) 人口稀释、\(g\) 技术稀释——三者都是「要维持 k 不变必须补上的量」,所以并列相加。
    Depreciation \(\delta\), population dilution \(n\) and technological dilution \(g\) are all amounts that must be replaced merely to hold k constant, which is why they simply add.
稳态与黄金律
Steady state and the golden rule
$$k^*=\Bigl(\frac{s}{\delta+n+g}\Bigr)^{\frac{1}{1-\alpha}}, \qquad f'(k_{\text{gold}})=\delta+n+g\ \Longrightarrow\ s_{\text{gold}}=\alpha$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 令 k̇ = 0 解稳态
    Set k̇ = 0 and solve
    $$sk^{\alpha}=(\delta+n+g)k\ \Longrightarrow\ k^{1-\alpha}=\frac{s}{\delta+n+g}$$
    〔两边同除 \(k^{\alpha}\),指数相减:\(k^{1}/k^{\alpha}=k^{1-\alpha}\)〕再取 \(\frac{1}{1-\alpha}\) 次幂。
    〔Divide both sides by \(k^{\alpha}\) and subtract exponents: \(k^{1}/k^{\alpha}=k^{1-\alpha}\)〕then raise to the power \(\frac{1}{1-\alpha}\).
  2. 黄金律:最大化稳态消费
    Golden rule: maximise steady-state consumption
    $$c^*=f(k^*)-(\delta+n+g)k^*,\qquad \frac{dc^*}{dk^*}=f'(k^*)-(\delta+n+g)=0$$
    对柯布–道格拉斯,\(f'(k)=\alpha k^{\alpha-1}\),代入得 黄金律储蓄率就等于资本份额 α。这是一个漂亮到不像话的结果。
    For Cobb–Douglas, \(f'(k)=\alpha k^{\alpha-1}\), which gives a golden-rule saving rate equal to the capital share α — a result almost too neat to be true.
  3. 收敛速度
    Speed of convergence
    $$\dot k\approx-(1-\alpha)(\delta+n+g)(k-k^*)$$
    缺口每年缩小 \((1-\alpha)(\delta+n+g)\)。这正是实证「条件收敛速度约 2%/年」的理论对应物。
    The gap closes at \((1-\alpha)(\delta+n+g)\) a year. This is the theoretical counterpart of the empirical finding that conditional convergence runs at roughly 2% a year.

交互图Interactive chart

Solow 图:储蓄曲线与持平投资线
The Solow diagram: saving curve against break-even investment
交点是稳态 k*。把 s 拖到 α(默认 0.33)时 k* 正好落在黄金律那条虚线上。
The intersection is the steady state k*. Drag s up to α (0.33 by default) and k* lands exactly on the golden-rule line.

教学算例Worked example

持平投资率 \(\delta+n+g=0.08\)。

Break-even investment rate \(\delta+n+g=0.08\).

稳态资本 \(k^*\)Steady-state capital \(k^*\)3.9528\((0.2/0.08)^{1.5}\)\((0.2/0.08)^{1.5}\)
稳态产出 \(y^*\)Steady-state output \(y^*\)1.5811\(=k^{*1/3}\)\(=k^{*1/3}\)
稳态消费 \(c^*\)Steady-state consumption \(c^*\)1.2649\(=(1-s)y^*\)\(=(1-s)y^*\)
黄金律资本Golden-rule capital8.5052口算 8.5046 是错的——脚本算出 8.50517The mental arithmetic gave 8.5046 and was wrong; the script returns 8.50517
黄金律储蓄率Golden-rule saving rate1/3 = 0.3333等于资本份额 αEqual to the capital share α
黄金律消费Golden-rule consumption1.3608比当前 c* 高 7.6%7.6% above the current c*
收敛速度Speed of convergence5.33%/年\((1-\alpha)(\delta+n+g)\)\((1-\alpha)(\delta+n+g)\)
缺口半衰期Half-life of the gap13.0 年\(\ln2/0.0533\)\(\ln2/0.0533\)
s = 0.20 < α = 0.33 ⇒ 储蓄不足(动态有效)。多储蓄能提高长期消费,但要先经历一段消费下降的过渡期——这正是「代际公平」争论的技术核心。
s = 0.20 < α = 0.33, so the economy saves too little (it is dynamically efficient). Saving more would raise long-run consumption, but only after a transition during which consumption falls — which is the technical core of every argument about intergenerational fairness.

经典文献Original sources

  • Solow, R. M. (1956). "A Contribution to the Theory of Economic Growth." Quarterly Journal of Economics 70(1), 65–94.1987 年诺奖。Nobel Prize 1987.
  • Swan, T. W. (1956). "Economic Growth and Capital Accumulation." Economic Record 32(2), 334–361.同年独立提出,故正式称 Solow–Swan 模型。Published independently the same year, which is why the model is properly Solow–Swan.
  • Phelps, E. S. (1961). "The Golden Rule of Accumulation: A Fable for Growthmen." American Economic Review 51(4), 638–643.黄金律。2006 年诺奖。The golden rule. Nobel Prize 2006.
  • Mankiw, N. G., Romer, D. & Weil, D. N. (1992). "A Contribution to the Empirics of Economic Growth." QJE 107(2), 407–437.加入人力资本的扩展 Solow,实证上把跨国收入差异解释力提到约 80%。Solow augmented with human capital; empirically it raises the share of cross-country income differences explained to around 80%.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把 k 的口径搞混。有效劳动人均 \(K/(AL)\)、人均 \(K/L\)、总量 \(K\) 三者的积累方程分母不同:依次是 \(\delta+n+g\)、\(\delta+n\)、\(\delta\)。写方程前先声明口径,这是本模型最高频的错。Losing track of the units of k. Per effective worker \(K/(AL)\), per worker \(K/L\) and aggregate \(K\) carry different accumulation terms: \(\delta+n+g\), \(\delta+n\) and \(\delta\) respectively. State the units before writing the equation — this is the single most frequent error in the model.
  2. 以为提高储蓄率能提高长期增长率。提高 \(s\) 只提高水平 \(k^*,y^*\),长期增长率恒等于外生的 \(g\)。这是 Solow 模型最反直觉、也最重要的结论,模型 8 正是为解决这一点而生。Believing a higher saving rate raises the long-run growth rate. Raising \(s\) raises the levels \(k^*\) and \(y^*\); the long-run growth rate is exogenous \(g\) and nothing else. This is the most counter-intuitive and most important result in the model, and Model 8 exists to answer it.
  3. 认为储蓄越多越好。\(s\gt\alpha\) 时经济动态无效:资本过多,减少储蓄可以让每一代人的消费都上升。Assuming more saving is always better. When \(s\gt\alpha\) the economy is dynamically inefficient: there is too much capital, and cutting saving would raise consumption for every generation.
知识勾连 Cross-links 模型 7 把 s 从外生参数变成家庭最优选择;模型 8 把 g 从外生技术变成研发的结果。Model 7 turns s from an exogenous parameter into a household choice; Model 8 turns g from exogenous technology into the output of research.
宏观 · 07Macroeconomics · 07

Ramsey–Cass–Koopmans 最优增长Ramsey–Cass–Koopmans optimal growth

Optimal growth: the Ramsey–Cass–Koopmans model

储蓄率不再是天上掉下来的参数,而是家庭在无限期上权衡出来的。

The saving rate stops being a parameter handed down from above and becomes something households trade off over an infinite horizon.

Solow 的储蓄率是外生的,因此无法回答「储蓄多少才对」。 Ramsey 让家庭最大化无限期的贴现效用,储蓄率成为内生结果。 代价是必须处理动态最优化:欧拉方程、相图、鞍点路径、横截性条件—— 这套工具是现代宏观(RBC、DSGE、新凯恩斯)的通用语言。
Because Solow's saving rate is exogenous, the model cannot answer how much saving is right. Ramsey has households maximise discounted utility over an infinite horizon, which makes the saving rate an outcome. The price is that dynamic optimisation must now be handled: Euler equations, phase diagrams, saddle paths, transversality conditions. That toolkit is the common language of modern macroeconomics — RBC, DSGE and New Keynesian alike.

核心方程Core equations

家庭问题与欧拉方程
The household problem and the Euler equation
$$\max\int_0^{\infty}e^{-\rho t}\frac{c^{1-\theta}-1}{1-\theta}dt \ \text{ s.t. }\ \dot k=f(k)-c-(\delta+n)k$$
展开逐步推导(4 步)Show the 4-step derivation
  1. 构造现值汉密尔顿函数
    Form the present-value Hamiltonian
    $$H=\frac{c^{1-\theta}-1}{1-\theta}+\mu\bigl[f(k)-c-(\delta+n)k\bigr]$$
    \(\mu\) 是资本的影子价格。控制变量是 \(c\),状态变量是 \(k\)。
    \(\mu\) is the shadow price of capital. The control is \(c\), the state is \(k\).
  2. 一阶条件与协态方程
    First-order and costate conditions
    $$\frac{\partial H}{\partial c}=c^{-\theta}-\mu=0;\qquad \dot\mu=\rho\mu-\frac{\partial H}{\partial k}$$
    第一条把 \(\mu\) 与消费的边际效用绑定;第二条是资产定价方程。
    The first ties \(\mu\) to the marginal utility of consumption; the second is an asset pricing equation.
  3. 消掉 μ 得欧拉方程
    Eliminate μ to get the Euler equation
    $$\frac{\dot c}{c}=\frac{1}{\theta}\bigl[f'(k)-\delta-\rho\bigr]$$
    对 \(c^{-\theta}=\mu\) 取对数再求导:\(-\theta\dfrac{\dot c}{c}=\dfrac{\dot\mu}{\mu}\)。1/θ 是跨期替代弹性:资本回报超过贴现率多少,消费就以多快的速度增长。
    Take logs of \(c^{-\theta}=\mu\) and differentiate: \(-\theta\dfrac{\dot c}{c}=\dfrac{\dot\mu}{\mu}\). 1/θ is the elasticity of intertemporal substitution: it converts the excess of the return on capital over the discount rate into a rate of consumption growth.
  4. 横截性条件
    Transversality condition
    $$\lim_{t\to\infty}e^{-\rho t}\mu(t)k(t)=0$$
    这一条最常被遗忘。没有它,欧拉方程有无穷多条解路径(可以永远借下去),鞍点路径的唯一性正是靠它钉住的。
    The most frequently forgotten condition in the model. Without it the Euler equation admits infinitely many paths (one may borrow for ever); it is what pins the solution down to the unique saddle path.
稳态与修正黄金律
Steady state and the modified golden rule
$$f'(k^*)=\rho+\delta\qquad\text{vs.}\qquad f'(k_{\text{gold}})=\delta+n+g$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 令 ċ = 0
    Set ċ = 0
    $$f'(k^*)=\rho+\delta\ \Longrightarrow\ k^*=\Bigl(\frac{\alpha}{\rho+\delta}\Bigr)^{1/(1-\alpha)}$$
    贴现率 \(\rho\) 进了分母,ρ 越大稳态资本越少——不耐心的社会积累得少。
    The discount rate \(\rho\) enters the denominator, so a larger ρ means a smaller steady-state capital stock: an impatient society accumulates less.
  2. 与黄金律比较
    Compare with the golden rule
    $$\rho\gt n+g\ \Longrightarrow\ k^*\lt k_{\text{gold}}$$
    Ramsey 经济永远不会动态无效。这是与 Solow 最重要的差别:Solow 的 \(s\) 可以随便设得过高,Ramsey 里理性家庭绝不会那么做。
    A Ramsey economy is never dynamically inefficient. That is the important contrast with Solow: an exogenous \(s\) can be set too high, whereas optimising households never choose to.

交互图Interactive chart

相图:两条零变化线与鞍点路径
Phase diagram: two loci and the saddle path
竖线是 ċ=0(由 ρ 决定),拱形是 k̇=0。交点为稳态,虚线是唯一收敛的鞍点路径。
The vertical line is ċ=0 (set by ρ); the hump is k̇=0. They cross at the steady state, and the dashed line is the unique convergent saddle path.

教学算例Worked example

与模型 6 用同一个生产函数,便于直接对照。

Same production function as Model 6, so the two are directly comparable.

稳态 \(k^*\)Steady state \(k^*\)7.1278\((\alpha/(\rho+\delta))^{1.5}\)\((\alpha/(\rho+\delta))^{1.5}\)
\(f'(k^*)\)\(f'(k^*)\)0.09\(=\rho+\delta\) ✓\(=\rho+\delta\) ✓
黄金律 \(k_{gold}\)(模型 6)Golden-rule \(k_{gold}\) (Model 6)8.5052
贴现造成的缺口Gap created by discounting1.3774理性家庭主动选择低于黄金律的资本Optimising households deliberately choose less capital than the golden rule
k=5 处的消费增长率Consumption growth at k=51.20%/年\((f'(5)-\delta-\rho)/\theta\),为正 ⇒ 仍在积累\((f'(5)-\delta-\rho)/\theta\), positive, so accumulation continues
k* = 7.13 < k_gold = 8.51。差额不是「失误」而是最优:多积累的那部分资本,其带来的未来消费不足以补偿当下的忍耐。这就是「修正」黄金律里「修正」二字的含义。
k* = 7.13 < k_gold = 8.51. The shortfall is optimal rather than a mistake: the extra capital would deliver future consumption not worth the present abstinence. That is what the word "modified" is doing in "modified golden rule".

经典文献Original sources

  • Ramsey, F. P. (1928). "A Mathematical Theory of Saving." Economic Journal 38(152), 543–559.26 岁写就,两年后去世。Keynes 称之为「数学经济学有史以来最杰出的贡献之一」。Written at 26; the author died two years later. Keynes called it one of the most remarkable contributions ever made to mathematical economics.
  • Cass, D. (1965). "Optimum Growth in an Aggregative Model of Capital Accumulation." Review of Economic Studies 32(3), 233–240.把 Ramsey 问题嵌入新古典增长框架。Embedded Ramsey's problem in the neoclassical growth framework.
  • Koopmans, T. C. (1965). "On the Concept of Optimal Economic Growth."与 Cass 同年独立完成,故三人并称。1975 年诺奖。Completed independently in the same year, hence the three names. Nobel Prize 1975.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 忘掉横截性条件。只有欧拉方程时,任何一条初始消费都能生成一条满足欧拉方程的路径;横截性条件把它们筛到唯一一条鞍点路径上。论述题里漏掉这一条几乎必然扣分。Omitting the transversality condition. With only the Euler equation, any initial consumption generates a path satisfying it; the transversality condition is what selects the unique saddle path. Leaving it out of a written answer costs marks almost every time.
  2. 把 1/θ 说成风险规避系数。\(\theta\) 是相对风险规避系数,\(1/\theta\) 是跨期替代弹性。在 CRRA 效用下二者互为倒数,这是一个受批评的巧合(Epstein–Zin 偏好就是为拆开它们而生)。Calling 1/θ the coefficient of risk aversion. \(\theta\) is relative risk aversion; \(1/\theta\) is the elasticity of intertemporal substitution. Under CRRA they are reciprocals, a coincidence that has drawn steady criticism — Epstein–Zin preferences exist precisely to prise them apart.
  3. 以为稳态一定优于黄金律。\(k^*\lt k_{\text{gold}}\) 意味着稳态消费低于黄金律消费——但这是最优的,因为达到黄金律要求当下牺牲太多。「消费更高」不等于「更好」。Assuming the steady state must beat the golden rule. \(k^*\lt k_{\text{gold}}\) means steady-state consumption is lower than at the golden rule — and that is optimal, because reaching the golden rule would cost too much today. Higher consumption is not the same thing as better.
知识勾连 Cross-links 欧拉方程是模型 9(RBC)与模型 10(新凯恩斯 IS 曲线)的共同祖先——三者的动态核心是同一条式子。The Euler equation is the common ancestor of Model 9 (RBC) and the IS curve of Model 10 — the dynamic core of all three is one equation.
宏观 · 08Macroeconomics · 08

内生增长与创造性破坏Endogenous growth and creative destruction

Endogenous growth: AK, Romer, and Schumpeterian models

增长率不再外生:它由研发的私人回报决定,而回报又被下一次创新所摧毁。

The growth rate is no longer exogenous: it is set by the private return to research, and that return is destroyed by the next innovation.

Solow 把长期增长归给外生的 g,等于承认「增长论解释不了增长」。 内生增长把 g 变成经济决策的结果。其中熊彼特路线(Aghion–Howitt)最有味道: 创新的价值被未来的创新摧毁,这个自我拆台的机制既是增长的引擎,也是它的刹车。 2025 年诺贝尔经济学奖授予 Aghion、Howitt 与 Mokyr,正是表彰这一支。
By attributing long-run growth to an exogenous g, Solow effectively conceded that growth theory could not explain growth. Endogenous growth makes g an outcome of economic decisions. The Schumpeterian branch (Aghion–Howitt) is the most interesting of them: the value of an innovation is destroyed by future innovations, a self-undermining mechanism that is at once the engine of growth and its brake. The 2025 Nobel Prize went to Aghion, Howitt and Mokyr for exactly this line of work.

核心方程Core equations

AK 模型:去掉收益递减
AK: remove diminishing returns
$$Y=AK\ \Longrightarrow\ g=sA-\delta$$
展开逐步推导(1 步)Show the 1-step derivation
  1. 为什么没有收敛
    Why there is no convergence
    $$\frac{\dot K}{K}=s A-\delta\quad(\text{independent of }K)$$
    \(f'(k)=A\) 是常数,收益不再递减,所以储蓄率永久影响增长率。代价是这个结论对「α 恰好等于 1」极度敏感——稍微小一点就退回 Solow。
    \(f'(k)=A\) is constant, so returns no longer diminish and the saving rate permanently affects the growth rate. The cost is that the result is acutely sensitive to α being exactly 1; anything less and the model reverts to Solow.
熊彼特增长:创造性破坏
Schumpeterian growth: creative destruction
$$g=\lambda n\ln\gamma,\qquad V=\frac{\pi}{r+\lambda n}$$
展开逐步推导(4 步)Show the 4-step derivation
  1. 创新的到达与幅度
    Arrival rate and size of innovations
    $$\text{Poisson rate}=\lambda n,\qquad \text{productivity}\times\gamma\ (\gamma\gt1)$$
    \(n\) 是投入研发的人数。单位时间的对数产出增长 = 到达率 × 每次的对数跳幅。
    \(n\) is the number of researchers. Log output growth per unit time is the arrival rate times the log size of each jump.
  2. 增长率
    The growth rate
    $$g=\lambda n\cdot\ln\gamma$$
    取对数是因为增长是乘法的:\(\ln(\gamma^m)=m\ln\gamma\)。
    Logs appear because growth is multiplicative: \(\ln(\gamma^m)=m\ln\gamma\).
  3. 创新的价值——关键一步
    The value of an innovation — the crucial step
    $$V=\frac{\pi}{r+\lambda n}$$
    分母里的 λn 就是创造性破坏:现任垄断者以 \(\lambda n\) 的速率被下一个创新者取代,所以它的租金必须按 \(r+\lambda n\) 而不是 \(r\) 来贴现。研发越热,每项创新越不值钱。
    The λn in the denominator is creative destruction. The incumbent monopolist is displaced at rate \(\lambda n\), so its rents must be discounted at \(r+\lambda n\) rather than \(r\). The hotter the research race, the less any single innovation is worth.
  4. 研究套利条件
    Research arbitrage
    $$w=\lambda V\ \Longrightarrow\ n^*$$
    研发者的工资等于其边际产出(创新到达率 × 创新价值)。这条方程与劳动市场出清一起,把 \(n^*\) 定下来,从而定住增长率。
    A researcher's wage equals their marginal product (arrival rate times the value of an innovation). Together with labour market clearing this pins down \(n^*\), and with it the growth rate.

交互图Interactive chart

研发强度的两难:增长 vs 创新价值
The research dilemma: growth against the value of an innovation
上升的那条是 g=λn·lnγ,下降的那条是 V=π/(r+λn)。研发越热,增长越快,但每项创新越不值钱。
The rising line is g=λn·lnγ, the falling one V=π/(r+λn). More research means faster growth and a smaller reward per innovation.

教学算例Worked example

到达率 \(\lambda n=0.2\)/年。

Arrival rate \(\lambda n=0.2\) per year.

增长率 gGrowth rate g8.11%\(0.1\times2\times\ln1.5\)\(0.1\times2\times\ln1.5\)
创新价值 VValue of an innovation V4.00\(\pi/(r+\lambda n)=1/0.25\)\(\pi/(r+\lambda n)=1/0.25\)
若无创造性破坏Value absent creative destruction20.00\(\pi/r=1/0.05\)\(\pi/r=1/0.05\)
破坏造成的价值折损Value lost to destruction80%4 / 204 / 20
AK 对照:A=0.5, s=0.2, δ=0.05AK for comparison: A=0.5, s=0.2, δ=0.05g = 5%\(sA-\delta\)\(sA-\delta\)
创造性破坏让创新的私人价值缩水 80%。这是内生增长理论里最重要的外部性:私人研发回报低于社会回报(知识溢出,倾向不足),但同时又高估了自己(商业窃取效应,倾向过度)——两个外部性方向相反,政策上并不必然「补贴研发」。
Creative destruction cuts the private value of an innovation by 80%. This is the central externality of the literature: the private return to research falls short of the social return (knowledge spillovers, so too little research) while simultaneously exceeding it (business stealing, so too much) — two externalities pointing opposite ways, which is why "subsidise R&D" does not follow automatically.

经典文献Original sources

  • Romer, P. M. (1990). "Endogenous Technological Change." Journal of Political Economy 98(5), S71–S102.知识的非竞争性 + 垄断竞争,水平创新(品种增加)。2018 年诺奖。Non-rivalry of knowledge plus monopolistic competition; horizontal innovation (more varieties). Nobel Prize 2018.
  • Aghion, P. & Howitt, P. (1992). "A Model of Growth Through Creative Destruction." Econometrica 60(2), 323–351.垂直创新(质量阶梯)。1987 年投稿,历时五年才发表。2025 年诺贝尔经济学奖(与 Joel Mokyr 共享,表彰「解释创新驱动的经济增长」)。Vertical innovation (quality ladders). Submitted in 1987 and five years in press. Nobel Prize in Economic Sciences 2025, shared with Joel Mokyr, for explaining innovation-driven economic growth.
  • Grossman, G. M. & Helpman, E. (1991). Innovation and Growth in the Global Economy.把内生增长嵌入开放经济与贸易。Endogenous growth embedded in open economies and trade.
  • Jones, C. I. (1995). "R&D-Based Models of Economic Growth." JPE 103(4), 759–784.尖锐的实证批评:研发人员数十年间增长数十倍,增长率却没变 ⇒ 一代内生增长模型的「规模效应」被数据否定。A sharp empirical objection: research employment rose many times over across decades while growth rates did not, which refutes the scale effect built into the first generation of models.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把创新价值写成 π/r。分母必须是 \(r+\lambda n\)。漏掉 \(\lambda n\) 等于假设垄断租金永远持续,那样就没有「创造性破坏」了——这是整个模型的名字所在。Writing the value of an innovation as π/r. The denominator must be \(r+\lambda n\). Dropping \(\lambda n\) assumes monopoly rents last for ever, which removes the creative destruction the model is named after.
  2. 认为内生增长理论必然支持补贴研发。模型里有两个方向相反的外部性:知识溢出(研发不足)与商业窃取(研发过度)。净效应取决于参数,不是定论。Assuming the theory implies research subsidies. Two externalities point in opposite directions: knowledge spillovers (too little research) and business stealing (too much). The net effect depends on parameters and is not settled.
  3. 忽视 Jones (1995) 的批评。一代内生增长模型隐含「研发人数翻倍则增长率翻倍」的规模效应,与二战后数据严重冲突。半内生增长模型正是为修补这一点而生。Ignoring Jones (1995). The first generation of models implies that doubling research employment doubles the growth rate, which post-war data contradict outright. Semi-endogenous growth models exist to repair exactly this.
知识勾连 Cross-links 这是对模型 6「g 外生」的直接回应;模型 9 用的技术冲击 z 则把 g 的波动而非水平当成研究对象。A direct answer to the exogenous g of Model 6; Model 9 takes the technology shock z and studies the fluctuations in g rather than its level.
宏观 · 09Macroeconomics · 09

实际经济周期(RBC)Real business cycles

Real business cycles

把 Ramsey 模型加上随机技术冲击,波动就成了最优反应而非市场失灵。

Add stochastic technology shocks to the Ramsey model and fluctuations become an optimal response rather than a market failure.

RBC 的挑衅性在于:它用一个没有任何摩擦、没有货币、没有失业的模型, 复制出了美国战后产出波动的大部分二阶矩。无论同不同意它的结论, 它确立了现代宏观的方法论标准——写下微观基础、校准参数、模拟、与数据的矩对比。 今天所有 DSGE 模型都是它的后代。
What made RBC provocative is that a model with no frictions, no money and no unemployment reproduces most of the second moments of post-war US output. Whether or not one accepts the conclusion, it set the methodological standard of modern macroeconomics: write down microfoundations, calibrate, simulate, compare moments with the data. Every DSGE model in use today descends from it.

核心方程Core equations

闭式可解的特例(Long–Plosser)
The closed-form special case (Long–Plosser)
$$u=\ln c,\ \delta=1,\ y=zk^{\alpha}\ \Longrightarrow\ k_{t+1}=\alpha\beta y_t,\quad c_t=(1-\alpha\beta)y_t$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 写欧拉方程(离散时间)
    Write the Euler equation in discrete time
    $$\frac{1}{c_t}=\beta\,\mathbb{E}_t\Bigl[\frac{1}{c_{t+1}}\cdot\alpha z_{t+1}k_{t+1}^{\alpha-1}\Bigr]$$
    与模型 7 的连续时间欧拉方程是同一条式子的离散版:今天少吃一口的边际损失 = 明天多吃的贴现边际收益
    The same equation as the continuous-time Euler equation of Model 7: the marginal loss from eating one less unit today equals the discounted marginal gain tomorrow.
  2. 猜解:储蓄率为常数
    Guess a constant saving rate
    $$k_{t+1}=\sigma y_t\ \Longrightarrow\ c_t=(1-\sigma)y_t$$
    对数效用 + 全折旧 + 柯布–道格拉斯这三条同时成立时,收入效应与替代效应正好抵消,储蓄率为常数。
    With log utility, full depreciation and Cobb–Douglas holding together, income and substitution effects cancel exactly and the saving rate is constant.
  3. 代回欧拉方程定 σ
    Substitute back to pin down σ
    $$\frac{1}{(1-\sigma)y_t}=\frac{\alpha\beta}{(1-\sigma)\sigma y_t} \ \Longrightarrow\ \sigma=\alpha\beta$$
    〔约分:\(y_{t+1}\) 上下相消,\(k_{t+1}=\sigma y_t\) 代入分母〕储蓄率 αβ 与冲击无关,所以才有闭式解。
    〔Cancelling: \(y_{t+1}\) divides out and \(k_{t+1}=\sigma y_t\) goes into the denominator〕The saving rate αβ is independent of the shock, which is what makes the closed form possible.
技术冲击与传导
Technology shocks and propagation
$$\ln z_t=\rho_z\ln z_{t-1}+\varepsilon_t,\qquad \varepsilon_t\sim N(0,\sigma_{\varepsilon}^2)$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 当期产出弹性
    Impact elasticity of output
    $$\frac{\partial\ln y_t}{\partial\ln z_t}=1$$
    \(k_t\) 是前定变量(上一期决定的),所以 TFP 冲击当期一比一传到产出。
    \(k_t\) is predetermined (it was set last period), so a TFP shock passes one-for-one into output on impact.
  2. 内在传导:资本积累
    Internal propagation through capital
    $$\hat k_{t+1}=\hat y_t=\hat z_t+\alpha\hat k_t$$
    这就是 RBC 的传导机制——冲击本身是 AR(1),但资本积累把它拉长成更持久的产出波动。
    This is the RBC propagation mechanism: the shock itself is an AR(1), and capital accumulation stretches it into a more persistent path for output.

交互图Interactive chart

1% TFP 冲击的脉冲响应
Impulse response to a 1% TFP shock
ρ_z 越接近 1,冲击衰减越慢,产出的持续期越长。这就是「内在传导」的可视化。
The closer ρ_z is to 1, the more slowly the shock decays and the longer output stays away from trend. This is internal propagation made visible.

教学算例Worked example

季度校准,标准 RBC 参数。

Quarterly calibration, standard RBC parameters.

储蓄率 αβSaving rate αβ0.3564常数,不随冲击变化Constant, invariant to the shock
消费份额Consumption share0.6436
TFP 路径 t=0…5TFP path, t=0…51.000, 0.950, 0.903, 0.857, 0.815, 0.774ρ_z 的幂Powers of ρ_z
产出当期弹性Impact elasticity of output1.00k 前定 ⇒ 一比一k predetermined, hence one-for-one
TFP 半衰期TFP half-life13.5 季度\(\ln0.5/\ln0.95\)\(\ln0.5/\ln0.95\)
冲击本身衰减很慢(ρ=0.95),产出的持续性主要来自冲击而非模型内生机制。这正是 RBC 最受诟病之处:要匹配数据,得先假设一个高度持久的外生冲击,等于把要解释的东西塞进了假设里。
The shock itself decays slowly (ρ=0.95), so most of the persistence in output comes from the shock rather than from anything internal to the model. That is the standing objection to RBC: matching the data requires assuming a highly persistent exogenous process, which puts the thing to be explained into the assumptions.

经典文献Original sources

  • Kydland, F. E. & Prescott, E. C. (1982). "Time to Build and Aggregate Fluctuations." Econometrica 50(6), 1345–1370.RBC 的开山之作,「校准」方法论也由此确立。2004 年诺奖。The founding paper of RBC, and the origin of the calibration methodology. Nobel Prize 2004.
  • Long, J. B. & Plosser, C. I. (1983). "Real Business Cycles." Journal of Political Economy 91(1), 39–69.给出了上面那个闭式可解的特例,是理解 RBC 机制最好的入口。Contains the closed-form special case above, which remains the best way into the mechanism.
  • Summers, L. H. (1986). "Some Skeptical Observations on Real Business Cycle Theory." Minneapolis Fed Quarterly Review.最著名的批评:Solow 余值究竟是技术冲击,还是要素利用率的测量误差?The best-known criticism: is the Solow residual a technology shock, or mismeasured factor utilisation?
  • Smets, F. & Wouters, R. (2007). "Shocks and Frictions in US Business Cycles." AER 97(3), 586–606.RBC 的现代后裔:加上价格粘性、工资粘性、习惯形成后贝叶斯估计,成为各国央行的主力模型。The modern descendant: add sticky prices, sticky wages and habit formation, estimate by Bayesian methods, and the result is the workhorse model of central banks.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把「校准」当成「估计」。校准是从微观证据或长期均值挑参数值,不做统计推断;它的辩护理由是「模型必错,估计出的标准误没有意义」。这是方法论立场,不是技术细节。Treating calibration as estimation. Calibration takes parameter values from microeconomic evidence or long-run averages and performs no inference; its defence is that since the model is certainly false, standard errors around its estimates mean little. This is a methodological position, not a technical detail.
  2. 以为 RBC 说「衰退是好事」。模型说的是在给定冲击下,观察到的波动是最优反应,因此稳定政策无益。它不说冲击本身是好事。这个区分在论述题里价值很高。Reading RBC as the claim that recessions are good. The claim is that given the shocks, observed fluctuations are an optimal response, and therefore stabilisation policy does not help. It says nothing about the shocks themselves being desirable. The distinction is worth marks in any written answer.
  3. 把 Solow 余值直接当技术冲击。余值里混着要素利用率、加成率、测量误差。Summers (1986) 与 Basu–Fernald–Kimball (2006) 的修正表明,纠正后的技术冲击效应与原始 RBC 结论方向相反。Taking the Solow residual as a technology shock. The residual mixes in factor utilisation, markups and measurement error. The corrections in Summers (1986) and Basu–Fernald–Kimball (2006) reverse the sign of the estimated response.
知识勾连 Cross-links 把模型 7 的欧拉方程加上随机项就是这里;模型 10 在此基础上加价格粘性,货币政策才重新有效。Add a stochastic term to the Euler equation of Model 7 and you arrive here; Model 10 adds price stickiness on top, at which point monetary policy matters again.
宏观 · 10Macroeconomics · 10

新凯恩斯三方程The three-equation New Keynesian model

The three-equation New Keynesian model

一条动态 IS、一条菲利普斯曲线、一条利率规则——现代央行的思维骨架。

A dynamic IS curve, a Phillips curve and an interest rate rule — the skeleton of how central banks now think.

RBC 里货币中性、政策无效。加一条「价格不能天天调」的摩擦, 整个结论翻转:需求冲击会造成真实波动,货币政策能且应当反应。 三方程模型是今天几乎所有货币政策讨论的最小公共语言。
In RBC money is neutral and policy is powerless. Add the single friction that prices cannot be reset every period and the conclusions reverse: demand shocks generate real fluctuations, and monetary policy both can and should respond. The three-equation model is the minimum shared language of virtually every monetary policy discussion today.

核心方程Core equations

三条方程
The three equations
$$\begin{aligned} x_t&=\mathbb{E}_tx_{t+1}-\tfrac{1}{\sigma}\bigl(i_t-\mathbb{E}_t\pi_{t+1}-r_t^n\bigr)\\ \pi_t&=\beta\mathbb{E}_t\pi_{t+1}+\kappa x_t\\ i_t&=r^n+\phi_{\pi}\pi_t+\phi_xx_t\end{aligned}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. IS 曲线就是欧拉方程
    The IS curve is the Euler equation
    $$\frac{1}{c_t}=\beta\mathbb{E}_t\Bigl[\frac{1+i_t}{1+\pi_{t+1}}\frac{1}{c_{t+1}}\Bigr]$$
    与模型 7、9 是同一条欧拉方程,只是写成了缺口形式。\(1/\sigma\) 又是跨期替代弹性——实际利率越高,越愿意把消费推后。
    The same Euler equation as Models 7 and 9, written in gap form. Once again \(1/\sigma\) is the elasticity of intertemporal substitution: the higher the real rate, the more willing the household is to postpone consumption.
  2. 菲利普斯曲线来自 Calvo 定价
    The Phillips curve comes from Calvo pricing
    $$\kappa=\frac{(1-\theta)(1-\beta\theta)}{\theta}(\sigma+\varphi)$$
    \(\theta\) 是每期不能调价的概率。\(\theta\to0\)(价格完全灵活)时 \(\kappa\to\infty\),曲线变垂直,货币重归中性——新凯恩斯模型退化成 RBC
    \(\theta\) is the probability of not being able to reset the price. As \(\theta\to0\) (fully flexible prices) \(\kappa\to\infty\), the curve becomes vertical and money is neutral again — the New Keynesian model collapses back into RBC.
  3. 泰勒原则
    The Taylor principle
    $$\phi_{\pi}\gt1$$
    名义利率对通胀的反应必须大于一比一,实际利率才会上升、才能压住通胀。\(\phi_{\pi}\lt1\) 时均衡不确定,会出现自我实现的通胀波动——这被广泛用来解释 1970 年代的大通胀。
    The nominal rate must respond more than one-for-one to inflation, otherwise the real rate does not rise and inflation is not restrained. With \(\phi_{\pi}\lt1\) the equilibrium is indeterminate and self-fulfilling inflation becomes possible — the standard account of the Great Inflation of the 1970s.

交互图Interactive chart

泰勒规则与泰勒原则
The Taylor rule and the Taylor principle
一条是名义利率、一条是实际利率对通胀的反应。φ_π>1 时实际利率线才向上倾斜——这就是泰勒原则。
One line is the nominal rate, the other the real rate, both as functions of inflation. Only when φ_π>1 does the real-rate line slope upward — that is the Taylor principle.

教学算例Worked example

\(i=2+\pi+0.5(\pi-2)+0.5\cdot\text{gap}\),即 \(r^*=2\%\)、\(\pi^*=2\%\)、\(\phi_{\pi}=1.5\)、\(\phi_x=0.5\)。

\(i=2+\pi+0.5(\pi-2)+0.5\cdot\text{gap}\), that is \(r^*=2\%\), \(\pi^*=2\%\), \(\phi_{\pi}=1.5\), \(\phi_x=0.5\).

π=2%, gap=0π=2%, gap=0i = 4.0%通胀达标、产出达标 ⇒ 中性名义利率 = r*+π*Inflation and output both on target, so the neutral nominal rate is r*+π*
π=4%, gap=+1%π=4%, gap=+1%i = 7.5%通胀超标 2pp ⇒ 名义利率多加 3ppInflation 2pp above target draws a 3pp rise in the nominal rate
π=1%, gap=−2%π=1%, gap=−2%i = 1.5%通缩 + 衰退 ⇒ 大幅宽松Disinflation plus recession calls for substantial easing
实际利率对通胀的斜率Slope of the real rate in inflation+0.5\(\phi_{\pi}-1=0.5\gt0\) ⇒ 满足泰勒原则\(\phi_{\pi}-1=0.5\gt0\), so the Taylor principle holds
关键不是 φ_π 的绝对大小,而是它是否大于 1。π 上升 1pp 时名义利率上升 1.5pp,实际利率才上升 0.5pp,需求才会被压住。φ_π<1 的央行会让通胀自我强化。
What matters is not how large φ_π is but whether it exceeds 1. A 1pp rise in inflation raises the nominal rate 1.5pp so that the real rate rises 0.5pp, and only then is demand restrained. A central bank with φ_π<1 lets inflation feed on itself.

经典文献Original sources

  • Calvo, G. A. (1983). "Staggered Prices in a Utility-Maximizing Framework." Journal of Monetary Economics 12(3), 383–398.「每期以固定概率能调价」的定价假设,因为可解性极好而成为业界标准。The assumption that a firm may reset its price with fixed probability each period. Its tractability made it the industry standard.
  • Taylor, J. B. (1993). "Discretion versus Policy Rules in Practice." Carnegie-Rochester Conference Series on Public Policy 39, 195–214.泰勒规则。原意是描述 1987–92 年美联储的实际行为,后来变成规范性基准。The Taylor rule. Intended as a description of what the Federal Reserve actually did between 1987 and 1992, later adopted as a normative benchmark.
  • Clarida, R., Galí, J. & Gertler, M. (1999). "The Science of Monetary Policy: A New Keynesian Perspective." Journal of Economic Literature 37(4), 1661–1707.把三方程整理成今天教科书的形式,并给出最优政策分析。Assembled the three equations into their textbook form and worked out optimal policy.
  • Galí, J. (2015). Monetary Policy, Inflation, and the Business Cycle (2nd ed.). Princeton University Press.标准研究生教材,三方程的完整微观基础推导都在这里。The standard graduate text; the full microfoundations of the three equations are here.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把 NKPC 当成传统菲利普斯曲线。传统版本是「通胀 vs 失业」的可利用的权衡;新凯恩斯版本是前瞻的:\(\pi_t\) 取决于预期未来通胀,不存在长期权衡,而且理论上「反通胀可以无成本」(只要央行可信)——这与经验事实的紧张是该模型的老问题。Reading the NKPC as the traditional Phillips curve. The traditional version offers an exploitable trade-off between inflation and unemployment; the New Keynesian version is forward-looking: \(\pi_t\) depends on expected future inflation, there is no long-run trade-off, and in principle disinflation is costless given a credible central bank — a prediction whose tension with the evidence is a long-standing problem for the model.
  2. 忘记泰勒规则里的利率是名义利率。判断政策松紧要看实际利率 \(i-\mathbb{E}\pi\)。名义利率上升但通胀上升更多,实际是在放松。这是新闻评论里最常见的错误。Forgetting that the rate in the Taylor rule is nominal. Whether policy is tight or loose depends on the real rate \(i-\mathbb{E}\pi\). A nominal rate that rises by less than inflation is an easing. This is the commonest error in press commentary.
  3. 以为 φ_π 大于 1 就万事大吉。在零利率下限(ZLB)处泰勒规则无法执行,不确定性重新出现——这正是 2009 年后前瞻指引与量化宽松的理论动机。Assuming φ_π > 1 settles the matter. At the zero lower bound the rule cannot be implemented and indeterminacy returns — which is precisely the motivation for forward guidance and quantitative easing after 2009.
知识勾连 Cross-links θ→0(价格灵活)时本模型退化为模型 9 的 RBC;IS 曲线则直接来自模型 7 的欧拉方程。As θ→0 (flexible prices) this model collapses into the RBC model of Model 9; the IS curve comes straight from the Euler equation of Model 7.
计量 · 11Econometrics · 11

OLS、FWL 定理与遗漏变量偏误OLS, the Frisch–Waugh–Lovell theorem and omitted variable bias

OLS, the Frisch–Waugh–Lovell theorem, and omitted variable bias

「控制变量」到底在做什么?FWL 定理给出了一句话的答案:把它们的影响先剔掉。

What exactly does "controlling for X" do? FWL answers in one sentence: it strips X out first.

几乎所有实证论文都在说「我们控制了 X」。FWL 定理精确说明了这句话的含义: 多元回归里某个系数,等于把其他变量的影响从 y 和该变量中都剔除后,用残差做一元回归。 理解了它,遗漏变量偏误、固定效应、去趋势、部分线性模型全都是同一件事。
Almost every empirical paper says it controls for X. The FWL theorem states precisely what that means: a coefficient in a multiple regression equals the coefficient from a simple regression of one residual on another, once the other regressors have been partialled out of both. Grasp it and omitted variable bias, fixed effects, detrending and partially linear models all turn out to be the same operation.

核心方程Core equations

OLS 与高斯–马尔可夫
OLS and Gauss–Markov
$$\hat\beta=(X'X)^{-1}X'y,\qquad \mathbb{E}[\hat\beta]=\beta\ \text{ if }\ \mathbb{E}[u\mid X]=0$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 一阶条件
    First-order conditions
    $$\min_{\beta}(y-X\beta)'(y-X\beta)\ \Longrightarrow\ X'(y-X\hat\beta)=0$$
    正规方程的含义是「残差与每个回归元正交」,这是几何投影,与统计假设无关。
    The normal equations say that the residual is orthogonal to every regressor. That is geometry — a projection — and holds regardless of any statistical assumption.
  2. 无偏性需要外生性
    Unbiasedness requires exogeneity
    $$\hat\beta=\beta+(X'X)^{-1}X'u$$
    \(\mathbb{E}[u\mid X]=0\) 才能让第二项期望为零。同方差与无自相关只影响有效性与标准误,不影响无偏性——这两件事常被混为一谈。
    \(\mathbb{E}[u\mid X]=0\) is what makes the second term have zero expectation. Homoskedasticity and no autocorrelation bear on efficiency and standard errors, not on unbiasedness — the two are constantly conflated.
FWL 定理
The FWL theorem
$$\hat\beta_1=\frac{\tilde x_1'\tilde y}{\tilde x_1'\tilde x_1},\qquad \tilde\cdot=(I-X_2(X_2'X_2)^{-1}X_2')\,\cdot$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 三步做法
    Three steps
    $$\text{(1) }y\ \text{on}\ X_2\to\tilde y;\quad \text{(2) }x_1\ \text{on}\ X_2\to\tilde x_1;\quad \text{(3) }\tilde y\ \text{on}\ \tilde x_1$$
    得到的系数与三元回归里 \(x_1\) 的系数逐位相同(不是近似)。
    The coefficient obtained is numerically identical to the one from the full regression, not merely close to it.
  2. 这就是「控制」的定义
    This is what "controlling for" means
    $$\text{control }X_2\ \equiv\ \text{use only the part of }x_1\text{ orthogonal to }X_2$$
    固定效应就是控制个体虚拟变量的 FWL:等价于组内去均值。去趋势、季节调整同理,全是同一定理的应用。
    Fixed effects are FWL applied to a set of unit dummies: exactly equivalent to demeaning within groups. Detrending and seasonal adjustment are the same theorem again.
遗漏变量偏误
Omitted variable bias
$$\hat\beta_1^{\text{short}}=\hat\beta_1+\hat\beta_2\hat\delta$$
展开逐步推导(1 步)Show the 1-step derivation
  1. 符号判断
    Signing the bias
    $$\operatorname{sign}(\text{bias})=\operatorname{sign}(\beta_2)\times\operatorname{sign}(\delta)$$
    遗漏变量与被解释变量正相关、且与关键回归元正相关 ⇒ 高估。这是审稿人第一个会问的问题,也是最容易口头推断的诊断。
    If the omitted variable is positively related to the outcome and positively related to the regressor of interest, the estimate is too large. It is the first question a referee asks and the easiest diagnostic to run in your head.

交互图Interactive chart

FWL:控制掉 x₂ 之后的散点
FWL: the scatter once x₂ has been partialled out
左:y 对 x₁ 的原始散点(斜率被污染)。右:两个残差的散点(斜率就是多元回归系数)。
Left: the raw scatter of y on x₁, whose slope is contaminated. Right: residual against residual, whose slope is the multiple regression coefficient.

教学算例Worked example

真值 \(y=1+2x_1-1.5x_2+\varepsilon\),其中 \(x_1=0.7x_2+\nu\)。

True model \(y=1+2x_1-1.5x_2+\varepsilon\), with \(x_1=0.7x_2+\nu\).

三元回归的 \(\hat\beta_1\)\(\hat\beta_1\) from the full regression2.0231
FWL 两步法FWL in two steps2.0231逐位相同(差小于 1e−10)Identical to the digit (difference below 1e−10)
漏掉 \(x_2\) 的估计Estimate omitting \(x_2\)1.2535
\(\hat\beta_1+\hat\beta_2\hat\delta\)\(\hat\beta_1+\hat\beta_2\hat\delta\)1.2535遗漏变量偏误公式,恒等成立The omitted variable bias formula, an exact identity
偏误大小Size of the bias−0.7696\(\hat\beta_2\lt0\) 且 \(\hat\delta\gt0\) ⇒ 低估\(\hat\beta_2\lt0\) and \(\hat\delta\gt0\), so the estimate is too small
偏误公式里必须用样本估计 β̂₂ 而不是真值 −1.5。用真值只在概率极限意义上成立;有限样本里想让等式精确成立,必须用 β̂₂。这一点写验算脚本时才会暴露。
The formula requires the sample estimate β̂₂, not the true value −1.5. Using the true value holds only in probability limit; for the identity to be exact in a finite sample the estimate is required. Writing the verification script is what exposes this.

经典文献Original sources

  • Frisch, R. & Waugh, F. V. (1933). "Partial Time Regressions as Compared with Individual Trends." Econometrica 1(4), 387–401.原本是关于去趋势的争论:先去趋势再回归 vs 把趋势当回归元,二者等价。Originally an argument about detrending: removing a trend first and including it as a regressor turn out to be equivalent.
  • Lovell, M. C. (1963). "Seasonal Adjustment of Economic Time Series." JASA 58(304), 993–1010.推广到季节调整,故称 FWL 定理。Extended to seasonal adjustment, whence the name FWL.
  • Angrist, J. D. & Pischke, J.-S. (2009). Mostly Harmless Econometrics. Princeton University Press.把 FWL、遗漏变量偏误、IV、DID、RD 串成一条「因果识别」主线的现代标准入门。The modern standard introduction, running FWL, omitted variable bias, IV, DID and RD together as one argument about causal identification.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把「无偏」与「有效」搞混。高斯–马尔可夫的无偏性只要外生性;最小方差才需要同方差 + 无自相关。异方差不会让 OLS 有偏,只会让默认标准误错——所以对策是稳健标准误,而不是换估计量。Confusing unbiasedness with efficiency. The unbiasedness half of Gauss–Markov needs only exogeneity; minimum variance needs homoskedasticity and no autocorrelation. Heteroskedasticity does not bias OLS, it only invalidates the default standard errors — so the remedy is robust standard errors, not a different estimator.
  2. 只做一步残差化。FWL 要求 \(y\) 与 \(x_1\) 对 \(X_2\) 残差化。只残差化一边,自由度与标准误都会错。Residualising only one side. FWL requires both \(y\) and \(x_1\) to be residualised against \(X_2\). Doing only one leaves the degrees of freedom and standard errors wrong.
  3. 用真值套遗漏变量偏误公式。有限样本里恒等式要用 \(\hat\beta_2\hat\delta\);\(\beta_2\delta\) 只在 plim 意义下成立。Plugging true values into the omitted variable bias formula. In a finite sample the identity requires \(\hat\beta_2\hat\delta\); \(\beta_2\delta\) holds only in probability limit.
知识勾连 Cross-links 本节的「控制」在观测数据里往往不够——模型 12–14 是三种在无法控制时仍能识别因果的设计。In observational data, controlling is often not enough — Models 12 to 14 are three designs that identify causal effects when it is not.
计量 · 12Econometrics · 12

工具变量与两阶段最小二乘Instrumental variables and two-stage least squares

Instrumental variables and 2SLS

找一个只通过 x 影响 y 的外生变异源,用它那一部分变异来识别因果。

Find a source of exogenous variation that reaches y only through x, and identify the causal effect from that part alone.

当 x 与误差项相关(遗漏变量、测量误差、双向因果),OLS 无救。 IV 换一条路:不用 x 的全部变异,只用其中由外生工具 z 驱动的那一部分。 这是「可信性革命」的核心技术,也是最容易被滥用的技术。
When x is correlated with the error — omitted variables, measurement error, reverse causation — OLS cannot be salvaged. IV takes another route: use not all the variation in x but only the part driven by an exogenous instrument z. This is the central technique of the credibility revolution, and the one most easily abused.

核心方程Core equations

识别条件
Identifying assumptions
$$\text{(1) relevance: }\operatorname{Cov}(z,x)\ne0\qquad \text{(2) exclusion: }\operatorname{Cov}(z,u)=0$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 简单 IV 估计量
    The simple IV estimator
    $$\hat\beta_{IV}=\frac{\operatorname{Cov}(z,y)}{\operatorname{Cov}(z,x)}$$
    把 \(y=\beta x+u\) 代入:\(\dfrac{\beta\operatorname{Cov}(z,x)+\operatorname{Cov}(z,u)}{\operatorname{Cov}(z,x)}=\beta\),第二项为零全靠外生性
    Substituting \(y=\beta x+u\) gives \(\dfrac{\beta\operatorname{Cov}(z,x)+\operatorname{Cov}(z,u)}{\operatorname{Cov}(z,x)}=\beta\), and the second term vanishes only by exogeneity.
  2. 两阶段最小二乘
    Two-stage least squares
    $$\text{(1) }x=\pi z+v\ \to\ \hat x;\qquad \text{(2) }y=\beta\hat x+e$$
    第一阶段把 \(x\) 投影到 \(z\) 的空间;第二阶段只用这块干净的变异。标准误必须用 2SLS 公式,手工分两步跑 OLS 会低估标准误。
    The first stage projects \(x\) onto the space spanned by \(z\); the second uses only that clean variation. The standard errors must come from the 2SLS formula — running two OLS regressions by hand understates them.
  3. 排他性约束不可检验
    The exclusion restriction cannot be tested
    $$\operatorname{Cov}(z,u)=0\ \text{is not testable when just-identified}$$
    这是 IV 最脆弱的地方:它是一个论证,不是一个检验。过度识别时的 Sargan/Hansen J 检验也只检验「若至少一个工具有效则其余有效」,不能自举。
    This is where IV is weakest: it is an argument, not a test. Even the Sargan/Hansen J test under over-identification only asks whether the instruments agree with one another, which cannot lift itself by its own bootstraps.
弱工具与 LATE
Weak instruments and LATE
$$\text{bias}\approx\frac{\operatorname{Cov}(z,u)}{\pi\operatorname{Var}(z)}; \qquad \hat\beta_{IV}\xrightarrow{p}\ \mathbb{E}[Y_1-Y_0\mid\text{complier}]$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 弱工具的后果
    What weak instruments do
    $$\pi\to0\ \Longrightarrow\ \text{bias}\to\infty,\ \text{variance}\to\infty$$
    弱工具下 IV 比 OLS 更糟,且渐近分布不再正态。经验法则:第一阶段 F 统计量大于 10;但近年研究指出这个门槛远远不够。
    A weak instrument makes IV worse than OLS, and the asymptotic distribution ceases to be normal. The rule of thumb is a first-stage F above 10; recent work finds that threshold far too generous.
  2. 异质性效应下识别的是 LATE
    Under heterogeneity, IV identifies the LATE
    $$\text{only for compliers}$$
    不是平均处理效应 ATE,也不是对全体人群的效应。换一个工具就换一批 complier,因此换一个 LATE——这解释了为什么不同 IV 研究的估计值差异很大。
    Not the average treatment effect, and not the effect for the population. A different instrument recruits a different set of compliers and therefore identifies a different LATE — which is why estimates across IV studies diverge so widely.

交互图Interactive chart

工具强度与 IV 的表现
Instrument strength and the behaviour of IV
柱状是 IV 估计的抽样分布;两条竖线分别是真值 1.0 与 OLS 的期望 1.49。把 π 拖到 0.05 看弱工具如何崩坏。
The bars are the sampling distribution of the IV estimate; the two vertical lines mark the true value of 1.0 and the OLS expectation of 1.49. Drag π down to 0.05 to watch a weak instrument fall apart.

教学算例Worked example

\(y=x+u\),\(x=\pi z+0.8u+\text{噪声}\) ⇒ x 内生。

\(y=x+u\) with \(x=\pi z+0.8u+\text{noise}\), so x is endogenous.

OLS(强工具情形)OLS (strong-instrument case)1.4881偏高 49%——内生性偏误49% too high — the endogeneity bias
IV,π=0.8(强)IV, π=0.8 (strong)0.9984接近真值Close to the truth
IV 标准差,π=0.8IV standard deviation, π=0.80.0410
IV,π=0.05(弱)IV, π=0.05 (weak)−0.0888完全崩坏,符号都反了Complete breakdown; even the sign is wrong
IV 标准差,π=0.05IV standard deviation, π=0.0519.61方差放大 478 倍Variance inflated 478-fold
弱工具不是「精度差一点」,是彻底不能用。π 从 0.8 降到 0.05,点估计从 1.00 掉到 −0.09,标准差涨了近 500 倍。先看第一阶段 F 值,再看第二阶段结果。
A weak instrument is not slightly less precise; it is unusable. Taking π from 0.8 to 0.05 moves the point estimate from 1.00 to −0.09 and multiplies the standard deviation by nearly 500. Read the first-stage F before reading the second-stage result.

经典文献Original sources

  • Wright, P. G. (1928). The Tariff on Animal and Vegetable Oils, Appendix B.IV 的最早出处。作者身份(是 Philip 还是其子 Sewall)至今仍有争议。The earliest known use of IV. Whether the author was Philip or his son Sewall is still disputed.
  • Angrist, J. D. & Krueger, A. B. (1991). "Does Compulsory School Attendance Affect Schooling and Earnings?" QJE 106(4), 979–1014.用出生季度作教育年限的工具,1980 年人口普查约 32.9 万名 1930–39 年出生男性。经典中的经典,也是弱工具批评的经典靶子。Quarter of birth as an instrument for years of schooling, on roughly 329,000 men born 1930–39 in the 1980 US census. A classic, and equally a classic target for the weak-instrument critique.
  • Bound, J., Jaeger, D. A. & Baker, R. M. (1995). "Problems with Instrumental Variables Estimation..." JASA 90(430), 443–450.用随机数当工具复现了 Angrist–Krueger 的结果——弱工具问题的著名演示。Reproduced the Angrist–Krueger results using randomly generated instruments — the celebrated demonstration of the weak-instrument problem.
  • Imbens, G. W. & Angrist, J. D. (1994). "Identification and Estimation of Local Average Treatment Effects." Econometrica 62(2), 467–475.LATE 定理。Imbens 与 Angrist 因因果推断方法获 2021 年诺奖(与 Card 共享)。The LATE theorem. Imbens and Angrist shared the 2021 Nobel Prize with Card for work on causal inference.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 手工跑两阶段 OLS 然后报第二阶段的标准误。那个标准误是错的(低估),因为它没考虑第一阶段的估计误差。必须用 2SLS/GMM 的联合公式。Running the two stages by hand and reporting the second-stage standard errors. Those standard errors are wrong (too small) because they ignore first-stage estimation error. Use the joint 2SLS or GMM formula.
  2. 把排他性约束当成可检验的假设。恰好识别时它完全不可检验,只能靠制度知识论证。「过度识别检验通过了」不等于工具有效。Treating the exclusion restriction as testable. Under just-identification it is not testable at all and can only be argued from institutional knowledge. Passing an over-identification test is not evidence that the instruments are valid.
  3. 把 LATE 当成 ATE 汇报。在效应异质时 IV 只识别 complier 的效应。政策外推前必须说清楚 complier 是谁。Reporting a LATE as though it were an ATE. Under heterogeneous effects IV identifies the effect for compliers only. Say who the compliers are before extrapolating to policy.
知识勾连 Cross-links 模型 13、14 是另外两条识别路径:DID 用时间维度的可比对照,RD 用断点附近的准随机分配。Models 13 and 14 are the other two routes: DID uses a comparable control group across time, RD uses quasi-random assignment at a threshold.
计量 · 13Econometrics · 13

双重差分:从 2×2 到交错处理Difference-in-differences: from 2×2 to staggered adoption

Difference-in-differences: from 2×2 to staggered adoption

用「对照组的变化」当作「处理组本来会发生的变化」——以及这个做法近年是怎么被推翻重建的。

Use the change in the control group as the change the treated group would have seen — and see how that practice was overturned and rebuilt in recent years.

DID 是应用微观最常用的设计,Card–Krueger 的最低工资研究把它推向主流。 但 2018 年以后一系列论文发现:当处理时点交错、效应随时间变化时, 标准双向固定效应(TWFE)估计量会给出连符号都可能错的结果。 这是近十年计量经济学最重要的方法论修正,不了解它会直接用错工具。
DID is the most widely used design in applied microeconomics, and Card and Krueger's minimum wage study pushed it into the mainstream. Since 2018, however, a series of papers has shown that when treatment timing is staggered and effects vary over time, the standard two-way fixed effects estimator can return an answer with the wrong sign. This is the most important methodological correction in econometrics of the past decade, and not knowing it means reaching for the wrong tool.

核心方程Core equations

2×2 DID 与平行趋势
2×2 DID and parallel trends
$$\widehat{\text{DID}}=(\bar y_{T,1}-\bar y_{T,0})-(\bar y_{C,1}-\bar y_{C,0})$$
展开逐步推导(2 步)Show the 2-step derivation
  1. 识别假设
    The identifying assumption
    $$\mathbb{E}[Y_{T,1}(0)-Y_{T,0}(0)]=\mathbb{E}[Y_{C,1}(0)-Y_{C,0}(0)]$$
    平行趋势针对的是反事实 Y(0),不是观测到的趋势。「处理前趋势平行」是支持性证据,不是假设本身——这个区分极其重要。
    Parallel trends is an assumption about the counterfactual Y(0), not about observed trends. Parallel pre-trends are supporting evidence, not the assumption itself — a distinction that matters enormously.
  2. 回归形式
    The regression form
    $$y_{it}=\alpha_i+\lambda_t+\tau D_{it}+\varepsilon_{it}$$
    个体固定效应吸收水平差异,时间固定效应吸收共同冲击。这就是模型 11 的 FWL 定理:\(\tau\) 用的是双向去均值后的残差变异。
    Unit fixed effects absorb level differences, time fixed effects absorb common shocks. This is the FWL theorem of Model 11: \(\tau\) is estimated from residual variation after two-way demeaning.
交错处理下 TWFE 的崩坏(Goodman-Bacon 分解)
Where TWFE breaks down under staggered timing (Goodman-Bacon)
$$\hat\tau^{TWFE}=\sum_k w_k\widehat{DID}_k,\qquad \textstyle\sum_k w_k=1,\ \text{some }w_k\lt0$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 分解
    The decomposition
    $$\text{TWFE}=\text{weighted average of all 2}\times\text{2 comparisons}$$
    其中包含一类「后处理组 vs 早处理组」的比较——把已经受处理的单位当成了对照组。
    Among those comparisons is a class that uses later-treated units against earlier-treated ones — that is, units that have already been treated serve as controls.
  2. 为什么会错
    Why this goes wrong
    $$\text{if effects grow over time, early-treated }y\text{ also rises}$$
    早处理组自己的增长被当成「共同趋势」减掉了。结果是系统性低估,权重甚至可能为负,导致符号翻转——即使每个个体的真实效应都为正。
    The growth in the early-treated group is subtracted out as though it were a common trend. The result is systematic understatement, with weights that can even be negative and flip the sign — while every individual treatment effect is positive.
  3. 现代做法
    What to do instead
    $$ATT(g,t):\ \text{group }g\text{ (first treated)}\times\text{time }t$$
    Callaway–Sant'Anna (2021) 只用干净的比较(对照组在两期都未受处理),因此在任意效应异质下仍然一致。事件研究图应当基于它,而不是简单的 TWFE 带 leads/lags。
    Callaway and Sant'Anna (2021) use only clean comparisons, in which the control group is untreated in both periods, and remain consistent under arbitrary heterogeneity. Event study plots should be built on this rather than on TWFE with leads and lags.

交互图Interactive chart

2×2 DID 与平行趋势
2×2 DID and parallel trends
实线是观测数据,虚线是处理组的反事实。把「违背程度」拖离 0,看 DID 估计如何偏离真值。
Solid lines are observed data, the dashed line is the counterfactual for the treated group. Drag the violation away from 0 to see the DID estimate depart from the truth.

教学算例Worked example

1992 年新泽西最低工资由 4.25 美元升至 5.05 美元,宾夕法尼亚不变。快餐店全职等价(FTE)就业。

In 1992 New Jersey raised its minimum wage from $4.25 to $5.05 while Pennsylvania did not. Outcome: full-time-equivalent employment at fast-food restaurants.

新泽西 前 → 后New Jersey, before → after20.44 → 21.03变化 +0.59Change +0.59
宾夕法尼亚 前 → 后Pennsylvania, before → after23.33 → 21.17变化 −2.16Change −2.16
DID 估计DID estimate+2.75就业不降反升,与竞争性劳动市场模型的预测相反Employment rose rather than fell, contrary to the competitive labour market prediction
—— 交错处理模拟 ———— staggered-timing simulation ——
效应恒定时的 TWFETWFE with constant effects1.000与真值一致,无偏Matches the truth; unbiased
效应随时间增长时的 TWFETWFE with effects growing over time1.750
同一情形的真实 ATTTrue ATT in the same setting2.417TWFE 低估 28%TWFE understates by 28%
Card–Krueger 的正号引发了三十年的最低工资论战,也直接催生了「可信性革命」(Card 因此获 2021 年诺奖)。而下半张表说明:同样的 DID 思路,一旦处理时点交错,标准做法就会系统性低估。
The positive Card–Krueger estimate set off thirty years of argument about the minimum wage and launched the credibility revolution outright (Card received the 2021 Nobel Prize for it). The lower half of the table makes the second point: the same DID logic, applied under staggered timing, understates systematically.

经典文献Original sources

  • Card, D. & Krueger, A. B. (1994). "Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania." American Economic Review 84(4), 772–793.Card 因劳动经济学的实证贡献获 2021 年诺奖,此文是引用核心。Card received the 2021 Nobel Prize for empirical work in labour economics, with this paper at the centre of the citation.
  • Goodman-Bacon, A. (2021). "Difference-in-Differences with Variation in Treatment Timing." Journal of Econometrics 225(2), 254–277.把 TWFE 分解成所有 2×2 比较的加权平均,指出「后处理 vs 早处理」这一类比较的危害。Decomposes TWFE into a weighted average of all 2×2 comparisons and identifies the damage done by later-versus-earlier comparisons.
  • Callaway, B. & Sant'Anna, P. H. C. (2021). "Difference-in-Differences with Multiple Time Periods." Journal of Econometrics 225(2), 200–230.group-time ATT(g,t) 估计量,只用干净比较,是目前的默认做法(R 包 did;Stata 命令 csdid)。The group-time ATT(g,t) estimator, built only from clean comparisons; now the default (R: did; Stata: csdid).
  • de Chaisemartin, C. & D'Haultfœuille, X. (2020). "Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects." AER 110(9), 2964–2996.给出负权重的诊断方法,并提出替代估计量。Provides diagnostics for negative weights and an alternative estimator.
  • Sun, L. & Abraham, S. (2021). "Estimating Dynamic Treatment Effects in Event Studies..." Journal of Econometrics 225(2), 175–199.事件研究图里 leads/lags 系数被污染的机制与修正。The mechanism by which leads and lags in event study plots are contaminated, and how to fix it.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 把「处理前趋势平行」当成平行趋势假设成立的证明。假设针对的是处理后的反事实,本质不可检验。预趋势检验的功效往往很低,「看起来平行」的说服力被系统性高估。Treating parallel pre-trends as proof that parallel trends holds. The assumption concerns the post-treatment counterfactual and is fundamentally untestable. Pre-trend tests are often badly underpowered, and "it looks parallel" carries far less weight than it is given.
  2. 在交错处理下直接跑 TWFE 并报事件研究图。这是 2020 年以前的标准做法,现在已知有偏。应改用 Callaway–Sant'Anna 或 Sun–Abraham,并报告 Goodman-Bacon 分解看权重结构。Running TWFE under staggered timing and reporting the event study plot. Standard practice before 2020, now known to be biased. Use Callaway–Sant'Anna or Sun–Abraham instead, and report the Goodman-Bacon decomposition to inspect the weights.
  3. 对处理组数量很少的情形用常规聚类标准误。聚类数少于 40 左右时严重低估,应使用 wild cluster bootstrap 或随机化推断。Using conventional clustered standard errors with few treated clusters. Below roughly 40 clusters they understate badly; use a wild cluster bootstrap or randomisation inference.
知识勾连 Cross-links 与模型 12 相比:IV 靠外生工具,DID 靠时间维度的可比对照,模型 14 靠断点附近的准随机分配。三者是同一目标的三条路。Set against Model 12: IV relies on an exogenous instrument, DID on a comparable control group across time, Model 14 on quasi-random assignment at a threshold. Three routes to one destination.
计量 · 14Econometrics · 14

断点回归设计Regression discontinuity design

Regression discontinuity design

在分数线两侧一分之差的人几乎一样——把这条线当作一次自然的随机分配。

People one mark either side of a cutoff are nearly identical — so treat the cutoff as a natural randomisation.

RD 是所有准实验设计里识别假设最弱、最可信的一种: 只需要「除了处理之外,一切在断点处连续」。它的代价是外部有效性极窄—— 估计的只是断点附近那批人的效应。近年它在教育、医保、选举、信贷监管领域被大量使用。
RD makes the weakest and most credible identifying assumption of any quasi-experimental design: everything other than treatment must be continuous at the cutoff. The price is very narrow external validity — what is estimated is the effect for those near the threshold and no one else. It is now used heavily in education, health insurance, elections and financial regulation.

核心方程Core equations

识别
Identification
$$\tau_{RD}=\lim_{x\downarrow c}\mathbb{E}[Y\mid X=x]-\lim_{x\uparrow c}\mathbb{E}[Y\mid X=x]$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 识别假设
    The identifying assumption
    $$\mathbb{E}[Y(0)\mid X=x],\ \mathbb{E}[Y(1)\mid X=x]\ \text{continuous at }x=c$$
    只要求潜在结果连续,不要求随机化。这是 RD 可信度高的根源:跳跃只能来自处理,因为别的一切都连续。
    Only continuity of potential outcomes is required, not randomisation. That is the source of the credibility: a jump can only come from treatment, because everything else is continuous.
  2. 局部随机化的解释
    The local randomisation reading
    $$\text{no precise manipulation of }X\ \Rightarrow\ \text{as-if random near }c$$
    Lee (2008) 的论证。对应的检验是 McCrary 密度检验:看 \(X\) 的密度在断点处是否有跳跃,有跳跃说明有人在操纵。
    Lee's (2008) argument. The corresponding check is the McCrary density test: a jump in the density of \(X\) at the cutoff indicates that someone is manipulating it.
  3. 模糊 RD
    Fuzzy RD
    $$\tau_{FRD}=\frac{\text{jump in outcome}}{\text{jump in treatment probability}}$$
    这就是以「过线与否」为工具的 IV——分子分母都是跳跃,形式与 Wald 估计量一致。
    This is IV with "crossing the threshold" as the instrument — numerator and denominator are both jumps, in the form of a Wald estimator.
带宽:偏误–方差权衡
Bandwidth: the bias–variance trade-off
$$\text{MSE}(h)=\underbrace{O(h^{4})}_{\text{bias}^2}+\underbrace{O(1/nh)}_{\text{variance}} \ \Longrightarrow\ h^*\propto n^{-1/5}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 窄带宽
    Narrow bandwidth
    $$h\downarrow:\ \text{bias}\downarrow,\ \text{variance}\uparrow$$
    样本少了,噪声大。
    Fewer observations, hence more noise.
  2. 宽带宽
    Wide bandwidth
    $$h\uparrow:\ \text{bias}\uparrow,\ \text{variance}\downarrow$$
    把远处的点也拿来外推,函数弯曲就会被当成跳跃。
    Distant points are brought in and extrapolated, so curvature in the function is read as a jump.
  3. 一个反直觉的细节
    A counter-intuitive detail
    $$\text{equal curvature on both sides}\ \Rightarrow\ \text{extrapolation bias cancels}$$
    所以偏误来自「两侧曲率不同」,不是来自「有曲率」。这一点是写验算脚本时发现的:用对称的 \(0.8x^2\) 做模拟,宽带宽根本看不出偏误。
    So the bias arises from curvature that differs across the two sides, not from curvature as such. This surfaced while writing the verification script: with a symmetric \(0.8x^2\), a wide bandwidth shows no bias at all.

交互图Interactive chart

带宽如何影响断点估计
How bandwidth affects the estimated discontinuity
灰点是带宽外的样本,彩色点是带宽内的,两条直线是各侧的局部线性拟合,中间竖段就是估计的跳跃。
Grey points lie outside the bandwidth, coloured points inside; the two lines are local linear fits on each side, and the vertical segment between them is the estimated jump.

教学算例Worked example

\(x\sim U[-1,1]\),\(n=4000\),噪声 \(\sigma=0.5\);两侧曲率不同(左 0.4,右 2.5)。

\(x\sim U[-1,1]\), \(n=4000\), noise \(\sigma=0.5\); curvature differs across sides (0.4 left, 2.5 right).

h = 0.10h = 0.100.804 (se 0.104, n=381)方差大High variance
h = 0.30h = 0.300.919 (se 0.060, n=1149)偏误与方差的较好折中The better compromise between bias and variance
h = 0.90h = 0.900.689 (se 0.035, n=3587)标准误最小,但偏误最大Smallest standard error, largest bias
若两侧曲率相同With equal curvature on both sides宽带宽也几乎无偏外推偏误相互抵消Even a wide bandwidth is nearly unbiased; the extrapolation biases cancel
最小的标准误对应最差的点估计。这是带宽选择的全部难点,也是为什么必须用 MSE 最优带宽(Imbens–Kalyanaraman 2012、Calonico–Cattaneo–Titiunik 2014)而不是「看着顺眼」。
The smallest standard error goes with the worst point estimate. That is the whole difficulty of bandwidth choice, and the reason an MSE-optimal bandwidth (Imbens–Kalyanaraman 2012, Calonico–Cattaneo–Titiunik 2014) is required rather than one that merely looks reasonable.

经典文献Original sources

  • Thistlethwaite, D. L. & Campbell, D. T. (1960). "Regression-Discontinuity Analysis: An Alternative to the Ex Post Facto Experiment." Journal of Educational Psychology 51(6), 309–317.RD 的原始文献,研究国家优秀学生奖学金对后续学业的影响。此后沉寂了近 40 年。The original RD paper, on the effect of National Merit awards on later academic outcomes. It then lay dormant for close to forty years.
  • Hahn, J., Todd, P. & van der Klaauw, W. (2001). "Identification and Estimation of Treatment Effects with a Regression-Discontinuity Design." Econometrica 69(1), 201–209.现代 RD 的理论基础:给出连续性识别条件与局部线性估计。The theoretical basis of modern RD: continuity-based identification and local linear estimation.
  • Lee, D. S. (2008). "Randomized Experiments from Non-random Selection in U.S. House Elections." Journal of Econometrics 142(2), 675–697.局部随机化解释;美国众议院现任优势的经典 RD 应用。The local randomisation reading, and the classic RD application to incumbency advantage in US House elections.
  • Calonico, S., Cattaneo, M. D. & Titiunik, R. (2014). "Robust Nonparametric Confidence Intervals for Regression-Discontinuity Designs." Econometrica 82(6), 2295–2326.偏误修正的稳健置信区间,现在的标准做法(rdrobust)。Bias-corrected robust confidence intervals, now standard practice (rdrobust).
  • McCrary, J. (2008). "Manipulation of the Running Variable in the Regression Discontinuity Design." Journal of Econometrics 142(2), 698–714.密度检验,判断有没有人在操纵分数。The density test for manipulation of the running variable.

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 用高阶全局多项式拟合。Gelman–Imbens (2019) 证明三次以上的全局多项式会给出荒谬的权重,导致伪跳跃。现在的标准是局部线性 + MSE 最优带宽 + 偏误修正置信区间Fitting high-order global polynomials. Gelman and Imbens (2019) show that global polynomials of order three and above produce absurd implicit weights and spurious jumps. The standard is now local linear estimation with an MSE-optimal bandwidth and bias-corrected confidence intervals.
  2. 不做 McCrary 密度检验与协变量平衡检验。如果个体能精确操纵 running variable,断点两侧就不再可比,RD 的全部可信度立刻消失。Omitting the McCrary density test and covariate balance checks. If units can manipulate the running variable precisely, the two sides of the cutoff are no longer comparable and the credibility of RD evaporates at once.
  3. 把 RD 估计外推到全体人群。它识别的只是断点处的处理效应。分数线附近的学生与远离分数线的学生,效应可能完全不同。Extrapolating the RD estimate to the population. What is identified is the treatment effect at the cutoff. Students near the threshold and students far from it may respond quite differently.
知识勾连 Cross-links 模糊 RD 在数学上就是模型 12 的 IV,工具是「是否过线」;与模型 13 相比,它不需要平行趋势,只需要连续性。Fuzzy RD is mathematically the IV of Model 12, instrumented by crossing the threshold; unlike Model 13 it needs no parallel trends, only continuity.
计量 · 15Econometrics · 15

GMM 与时间序列:平稳性、单位根、协整GMM and time series: stationarity, unit roots, cointegration

GMM, stationarity, unit roots, and cointegration

所有估计量都是矩条件的解;而时间序列的第一个问题永远是:这个序列平稳吗?

Every estimator is the solution to a set of moment conditions; and the first question in time series is always whether the series is stationary.

GMM 把 OLS、IV、2SLS、极大似然全部统一成「让样本矩接近零」的一个框架, 是现代计量的通用语法。而在时间序列里,平稳性是一切渐近理论的前提: 不平稳的序列做回归会产生伪回归,R² 很高、t 值很大、结论全错。
GMM unifies OLS, IV, 2SLS and maximum likelihood into a single framework — drive the sample moments to zero — and is the common grammar of modern econometrics. In time series, stationarity is the premise of all asymptotic theory: regressing non-stationary series on one another produces spurious regressions, with a high R² and large t-statistics and conclusions that are entirely false.

核心方程Core equations

GMM 框架
The GMM framework
$$\mathbb{E}[g(w_i,\theta_0)]=0\ \Longrightarrow\ \hat\theta=\arg\min_{\theta}\ \bar g(\theta)'W\bar g(\theta)$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 常见估计量都是特例
    Familiar estimators as special cases
    $$\text{OLS: }\mathbb{E}[x_iu_i]=0;\quad \text{IV: }\mathbb{E}[z_iu_i]=0;\quad \text{MLE: }\mathbb{E}[\partial\ln f/\partial\theta]=0$$
    换矩条件就换估计量,这是 GMM 的全部威力。
    Change the moment conditions and you change the estimator. That is the whole power of GMM.
  2. 最优权重矩阵
    The optimal weighting matrix
    $$W^*=\bigl(\mathbb{E}[gg']\bigr)^{-1}=\Omega^{-1}$$
    给方差小的矩条件更大权重。两步 GMM:先用 \(W=I\) 得一致估计,再估 \(\Omega\),再重估。
    Moment conditions with smaller variance receive more weight. Two-step GMM: obtain a consistent estimate with \(W=I\), estimate \(\Omega\), re-estimate.
  3. 过度识别检验
    The over-identification test
    $$J=n\,\bar g(\hat\theta)'\hat\Omega^{-1}\bar g(\hat\theta)\ \xrightarrow{d}\ \chi^2_{m-k}$$
    \(m\) 个矩条件、\(k\) 个参数 ⇒ 自由度 \(m-k\)。恰好识别时 J 恒等于 0,检验不存在——这就是为什么排他性约束在恰好识别的 IV 里不可检验。
    \(m\) moment conditions and \(k\) parameters give \(m-k\) degrees of freedom. Under just-identification J is identically zero and the test does not exist — which is exactly why the exclusion restriction is untestable in a just-identified IV.
单位根与协整
Unit roots and cointegration
$$y_t=\rho y_{t-1}+\varepsilon_t:\quad |\rho|\lt1\ \text{stationary};\quad \rho=1\ \text{unit root}$$
展开逐步推导(3 步)Show the 3-step derivation
  1. 平稳时的方差与半衰期
    Variance and half-life under stationarity
    $$\operatorname{Var}(y)=\frac{\sigma^2}{1-\rho^2},\qquad \text{half-life}=\frac{\ln 0.5}{\ln\rho}$$
    \(\rho\to1\) 时方差发散,冲击永不消退。
    As \(\rho\to1\) the variance diverges and shocks never die away.
  2. 为什么不能用常规 t 检验
    Why the usual t test fails
    $$\text{under }\rho=1,\ \hat\rho\ \text{has a non-normal (Wiener) limit}$$
    所以要用 Dickey–Fuller 的专用临界值,比常规 t 临界值更负。用常规临界值会过度拒绝单位根。
    Hence the dedicated Dickey–Fuller critical values, which lie further into the negative tail. Using conventional critical values over-rejects the unit root.
  3. 协整
    Cointegration
    $$y_t,x_t\sim I(1)\ \text{but}\ y_t-\beta x_t\sim I(0)$$
    两个各自游走的序列被一条长期关系拴住。Granger 表示定理:协整 ⟺ 存在误差修正模型 \(\Delta y_t=\alpha(y_{t-1}-\beta x_{t-1})+\dots\)。
    Two series that each wander are tied together by a long-run relation. Granger's representation theorem: cointegration holds if and only if an error correction model exists, \(\Delta y_t=\alpha(y_{t-1}-\beta x_{t-1})+\dots\).

交互图Interactive chart

ρ 决定序列的性格
ρ determines the character of the series
同一串随机扰动,只改 ρ。ρ 小于 1 时序列被拉回均值;ρ=1 时变成随机游走,永不回头。
One and the same sequence of disturbances, with only ρ changed. Below 1 the series is pulled back to its mean; at 1 it becomes a random walk and never returns.

教学算例Worked example

\(y_t=\rho y_{t-1}+\varepsilon_t\),\(\varepsilon\sim N(0,1)\),固定随机种子。

\(y_t=\rho y_{t-1}+\varepsilon_t\), \(\varepsilon\sim N(0,1)\), fixed random seed.

ρ = 0.5ρ = 0.5样本方差 1.12理论值 \(1/(1-0.25)=1.33\)Theoretical value \(1/(1-0.25)=1.33\)
ρ = 0.9ρ = 0.9样本方差 3.47理论值 \(1/(1-0.81)=5.26\)Theoretical value \(1/(1-0.81)=5.26\)
ρ = 1.0ρ = 1.0样本方差 15.01理论方差不存在——随 T 增长而发散No theoretical variance exists — it diverges with T
ρ=0.9 的冲击半衰期Half-life of a shock at ρ=0.96.58 期\(\ln0.5/\ln0.9\)\(\ln0.5/\ln0.9\)
3 个矩条件、1 个参数3 moment conditions, 1 parameterJ 检验自由度 = 2J test with 2 degrees of freedom
注意 ρ=0.9 时样本方差 3.47 明显低于理论值 5.26。T=200 对高持续性序列根本不够——这正是单位根检验功效低下的直接体现:ρ=0.9 与 ρ=1 在有限样本里几乎分不开。
Note that at ρ=0.9 the sample variance of 3.47 falls well short of the theoretical 5.26. T=200 is simply not enough for a highly persistent series — which is the low power of unit root tests made concrete: ρ=0.9 and ρ=1 are barely distinguishable in a finite sample.

经典文献Original sources

  • Hansen, L. P. (1982). "Large Sample Properties of Generalized Method of Moments Estimators." Econometrica 50(4), 1029–1054.GMM 的奠基文献。2013 年诺奖。The founding paper of GMM. Nobel Prize 2013.
  • Dickey, D. A. & Fuller, W. A. (1979). "Distribution of the Estimators for Autoregressive Time Series with a Unit Root." JASA 74(366), 427–431.单位根检验与非标准渐近分布。Unit root tests and their non-standard asymptotic distribution.
  • Granger, C. W. J. & Newbold, P. (1974). "Spurious Regressions in Econometrics." Journal of Econometrics 2(2), 111–120.两个独立随机游走互相回归,R² 却很高——伪回归问题的著名演示。Two independent random walks regressed on each other, with a high R² — the celebrated demonstration of spurious regression.
  • Engle, R. F. & Granger, C. W. J. (1987). "Co-integration and Error Correction." Econometrica 55(2), 251–276.协整与误差修正模型。二人共获 2003 年诺奖(Engle 另因 ARCH)。Cointegration and error correction. The two shared the 2003 Nobel Prize (Engle also for ARCH).

易错点Where it goes wrong

⚠ 最容易栽的三处
⚠ The three commonest errors
  1. 对不平稳序列直接跑回归。两个独立的随机游走互相回归会得到很高的 R² 与很大的 t 值,全是假的。先做单位根检验,再决定是差分还是建协整/误差修正模型。Regressing non-stationary series directly. Two independent random walks regressed on one another yield a high R² and large t-statistics, all of them false. Test for unit roots first, then decide between differencing and a cointegration or error correction model.
  2. 差分掉一切以求平稳。若序列本来协整,差分会丢掉长期关系的信息,得到只有短期动态的模型。协整时正确做法是误差修正模型,不是无脑差分。Differencing everything in pursuit of stationarity. If the series are cointegrated, differencing discards the long-run relation and leaves a model of short-run dynamics alone. Under cointegration the correct response is an error correction model, not differencing on reflex.
  3. 在恰好识别时报告 J 检验。自由度为 0,\(J\) 恒等于 0,这个「检验」没有任何信息量。Reporting a J test under just-identification. With zero degrees of freedom \(J\) is identically zero and the test carries no information whatever.
知识勾连 Cross-links 与模型 12 呼应:IV 是 GMM 的特例,而「恰好识别时排他性不可检验」在这里得到了 J 检验自由度为 0 的精确解释。This closes the loop with Model 12: IV is a special case of GMM, and "the exclusion restriction is untestable under just-identification" now has an exact explanation — the J test has zero degrees of freedom.