The Normal Distribution, Revisited · 再看正态分布
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| normal distribution/ˈnɔːml ˌdɪstrɪˈbjuːʃn/ | 正态分布 | zhèng tài fēn bù |
| z-score/zed skɔː/ | 标准分数 | biāo zhǔn fēn shù |
| standard normal/ˈstændəd ˈnɔːml/ | 标准正态 | biāo zhǔn zhèng tài |
| statistical inference/stəˈtɪstɪkl ˈɪnfərəns/ | 统计推断 | tǒng jì tuī duàn |
The bell curve returns
- Many sampling distributions are well modeled by the normal distribution 正态分布.
- It's the familiar symmetric bell, described by a mean and a standard deviation.
- When a statistic's sampling distribution is normal, we can compute probabilities for it.
- This is why the normal model is the workhorse of inference.
钟形曲线归来
- 许多抽样分布能被正态分布很好地建模。
- 它就是那条熟悉的对称钟形,由一个均值和一个标准差描述。
- 当一个统计量的抽样分布是正态的,我们就能为它计算概率。
- 这就是正态模型成为推断主力的原因。
z-scores for a statistic
- A $z$-score 标准分数 says how many standard deviations a value sits from the mean:
-
$$z = \frac{\text{statistic} - \text{mean}}{\text{standard deviation}}$$
- For a sample statistic, the "standard deviation" is that of its sampling distribution.
- The $z$-score converts any normal value onto one common scale.
统计量的 z 分数
- $z$ 分数说明一个值离均值有多少个标准差:
-
$$z = \frac{\text{statistic} - \text{mean}}{\text{standard deviation}}$$
- 对样本统计量,这里的“标准差”是它的抽样分布的标准差。
- $z$ 分数把任何正态值转换到同一个公共刻度上。
The standard normal
- The standard normal 标准正态 distribution has mean $0$ and standard deviation $1$.
- Every $z$-score lives on it, so one table (or calculator) handles all normal problems.
- Probabilities: area under the curve left/right of a $z$.
- Critical values: the $z$ that cuts off a chosen tail area (e.g. $z^* = 1.96$ for the middle $95\%$).
标准正态
- 标准正态分布均值为 $0$、标准差为 $1$。
- 每个 $z$ 分数都落在它上面,所以一张表(或一个计算器)就能处理所有正态问题。
- **概率:**曲线下、某个 $z$ 左侧/右侧的面积。
- **临界值:**切出某个所选尾部面积的那个 $z$(如中间 $95\%$ 对应 $z^* = 1.96$)。
Why it matters for inference
- Inference asks: "how surprising is my sample, if a claim were true?"
- The normal model turns that into an area (a probability) under the bell.
- A statistic far out in the tail (large $|z|$) is surprising evidence.
- Every confidence interval and test in Units 6–9 runs on this machinery.
它为何对推断重要
- 推断要问:“如果某个主张为真,我的样本有多意外?”
- 正态模型把它变成钟形下的一块面积(一个概率)。
- 一个远在尾部的统计量(大的 $|z|$)是令人意外的证据。
- 第 6–9 单元里每个置信区间和检验都靠这套机器运转。
The $z$-score's denominator is the standard deviation of the sampling distribution, not the spread of one sample's raw data. Mixing these up is a classic slip. And remember the empirical rule benchmarks — about $68\%$ within $1$ SD, $95\%$ within $2$, $99.7\%$ within $3$ — for a quick sense of how extreme a $z$ is.
$z$ 分数的分母是抽样分布的标准差,而不是单个样本原始数据的分散。把这两者弄混是经典失误。并记住经验法则基准——约 $68\%$ 在 $1$ 个标准差内、$95\%$ 在 $2$ 个内、$99.7\%$ 在 $3$ 个内——用来快速判断一个 $z$ 有多极端。
A statistic is normal with mean $50$ and SD $4$. You observe a value of $58$.
- $z = \dfrac{58 - 50}{4} = 2$ — two standard deviations above the mean.
- By the empirical rule, only about $2.5\%$ of values exceed $z = 2$.
- So a value of $58$ is fairly surprising under this model.
一个统计量服从均值 $50$、标准差 $4$ 的正态分布。你观察到一个值 $58$。
- $z = \dfrac{58 - 50}{4} = 2$——比均值高两个标准差。
- 由经验法则,只有约 $2.5\%$ 的值超过 $z = 2$。
- 所以在这个模型下,$58$ 这个值相当意外。
The normal distribution models many sampling distributions. A $z$-score $z = \frac{\text{stat} - \text{mean}}{\text{SD}}$ maps a value onto the standard normal (mean $0$, SD $1$), where probabilities are areas and critical values cut off tail areas. This normal machinery powers all statistical inference.
正态分布为许多抽样分布建模。$z$ 分数 $z = \frac{\text{stat} - \text{mean}}{\text{SD}}$ 把一个值映射到标准正态(均值 $0$、标准差 $1$)上,那里概率是面积、临界值切出尾部面积。这套正态机器驱动所有统计推断。
The standard normal, shaded by z · 按 z 着色的标准正态
Areas under the curve give probabilities; z marks the cutoff. · 曲线下的面积给出概率;z 标出临界位置。
A statistic is normal with mean 50 and SD 4. Find the z-score of the value 58. · 一个统计量服从均值 50、标准差 4 的正态分布。求值 58 的 z 分数。
z = (58 − 50)/4 = 2. · z = (58 − 50)/4 = 2。
The standard normal distribution has... · 标准正态分布有……
By definition, mean 0 and SD 1. · 按定义,均值 0、标准差 1。
The critical value z* for the middle 95% of the standard normal is about... (two decimals) · 标准正态中间 95% 对应的临界值 z* 约为……(保留两位小数)
z* = 1.96 leaves 2.5% in each tail. · z* = 1.96 在每条尾部各留 2.5%。
The z-score for a sample statistic uses the standard deviation of its sampling distribution. · 样本统计量的 z 分数使用的是它的抽样分布的标准差。
Not the spread of one sample's raw data — the sampling distribution's SD. · 不是单个样本原始数据的分散——而是抽样分布的标准差。
Under a normal model, a statistic with a large |z| is... · 在正态模型下,一个 |z| 很大的统计量是……
Large |z| = far from center = surprising evidence. · 大 |z| = 远离中心 = 令人意外的证据。