The Central Limit Theorem · 中心极限定理
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| Central Limit Theorem/ˈsentrəl ˈlɪmɪt ˈθɪərəm/ | 中心极限定理 | zhōng xīn jí xiàn dìng lǐ |
| population distribution/ˌpɒpjʊˈleɪʃn ˌdɪstrɪˈbjuːʃn/ | 总体分布 | zǒng tǐ fēn bù |
The most important theorem
- The Central Limit Theorem (CLT) 中心极限定理 is the engine behind most inference.
- It says: the sampling distribution of the sample mean $\bar{x}$ is approximately normal for a large sample.
- This holds no matter what the population's own shape is.
- A near-magical result — normality emerges from averaging.
最重要的定理
- **中心极限定理(CLT)**是大多数推断背后的引擎。
- 它说:对于大样本,样本均值 $\bar{x}$ 的抽样分布近似正态。
- 这无论总体本身是什么形状都成立。
- 一个近乎神奇的结果——正态性从求平均中涌现出来。
Works even for weird populations
- The population distribution 总体分布 can be skewed, bimodal, or bizarre.
- Yet the distribution of $\bar{x}$ over many samples still turns out roughly normal.
- Averaging "smooths out" the population's odd shape.
- Only a large enough sample is needed — not a normal population.
即使总体很古怪也有效
- 总体分布可以是偏斜的、双峰的,或稀奇古怪的。
- 然而 $\bar{x}$ 在许多样本上的分布仍然大致是正态的。
- 求平均把总体古怪的形状“抹平”了。
- 只需一个足够大的样本——而不需要正态的总体。
The n ≥ 30 guideline
- A common rule of thumb: $n \ge 30$ is usually enough for the CLT to kick in.
- For a roughly symmetric population, smaller $n$ can already be fine.
- For a very skewed population, you may want $n$ even larger.
- The guideline is practical, not an exact law.
n ≥ 30 准则
- 一条常用的经验法则:$n \ge 30$ 通常足以让 CLT 生效。
- 对大致对称的总体,更小的 $n$ 可能就已经够了。
- 对非常偏斜的总体,你可能需要更大的 $n$。
- 这条准则是实用的,而非精确的定律。
Bigger n, better normality
- The larger the sample size, the more nearly normal the sampling distribution of $\bar{x}$.
- Larger $n$ also makes the sampling distribution narrower (less spread).
- So big samples give estimates that are both more normal and more precise.
- This is why sample size matters so much in study design.
n 越大,正态性越好
- 样本量越大,$\bar{x}$ 的抽样分布越接近正态。
- 更大的 $n$ 也让抽样分布更窄(分散更小)。
- 所以大样本给出的估计既更正态、又更精确。
- 这就是样本量在研究设计中如此重要的原因。
The CLT is about the sampling distribution of $\bar{x}$, not the data itself. A large sample does not make the population or the raw data normal — it makes the distribution of the sample mean approximately normal. If the population is already normal, $\bar{x}$ is exactly normal for any $n$; the CLT only matters when it isn't.
CLT 讲的是 $\bar{x}$ 的抽样分布,而非数据本身。大样本不会让总体或原始数据变正态——它让样本均值的分布近似正态。如果总体本就是正态的,那么对任何 $n$,$\bar{x}$ 都恰好正态;CLT 只在总体不正态时才重要。
Incomes are strongly right-skewed (a few very high earners).
- One sample's incomes look skewed — the CLT says nothing about that.
- But average the incomes of $n = 50$ people, and repeat: those $\bar{x}$ values form a roughly normal distribution.
- So we can use a normal model for $\bar{x}$ even though incomes aren't normal.
收入强烈右偏(少数极高收入者)。
- 单个样本的收入看起来偏斜——CLT 对此不置一词。
- 但把 $n = 50$ 个人的收入求平均,并重复:那些 $\bar{x}$ 值会形成一个大致正态的分布。
- 所以即使收入不正态,我们也能对 $\bar{x}$ 使用正态模型。
The Central Limit Theorem says the sampling distribution of $\bar{x}$ is approximately normal for a large sample size, even when the population distribution isn't normal. A common guideline is $n \ge 30$; a larger sample makes the sampling distribution both more normal and narrower.
中心极限定理说,对于大样本,$\bar{x}$ 的抽样分布近似正态,即使总体分布不正态。一条常用准则是 $n \ge 30$;更大的样本让抽样分布既更正态、又更窄。
Normality emerges from averaging · 正态性从求平均中涌现
For large n, x-bar is approximately normal even from a skewed population. · 对大 n,即使来自偏斜总体,x-bar 也近似正态。
The Central Limit Theorem says that for a large sample, the sampling distribution of x-bar is... · 中心极限定理说,对大样本,x-bar 的抽样分布……
The CLT delivers approximate normality of x-bar for large n. · CLT 对大 n 给出 x-bar 的近似正态性。
The CLT requires the population itself to be normal. · CLT 要求总体本身是正态的。
It works even for skewed populations — that's the point. · 它即使对偏斜总体也有效——这正是重点。
What sample size is the common rule-of-thumb minimum for applying the CLT? · 应用 CLT 的常用经验法则最小样本量是多少?
n ≥ 30 is the usual guideline. · n ≥ 30 是常用准则。
As the sample size increases, the sampling distribution of x-bar becomes... · 随着样本量增大,x-bar 的抽样分布变得……
Bigger n → more normal and less spread. · n 越大 → 越正态、分散越小。
The CLT makes the raw data values normal, not just the distribution of x-bar. · CLT 让原始数据值变正态,而不仅仅是 x-bar 的分布。
It's about the sampling distribution of the mean, not the raw data. · 它讲的是均值的抽样分布,而非原始数据。