Summary Statistics for a Quantitative Variable · 定量变量的汇总统计量
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| mean/miːn/ | 均值 | jūn zhí |
| median/ˈmiːdiːən/ | 中位数 | zhōng wèi shù |
| Range/reɪndʒ/ | 极差 | jí chà |
| Standard deviation/ˈstændəd ˌdiːvɪˈeɪʃn/ | 标准差 | biāo zhǔn chà |
| resistant/rɪˈzɪstənt/ | 稳健 | wěn jiàn |
Numbers that summarize the data
- A graph shows the shape; summary statistics boil a distribution down to a few key numbers.
- Two jobs: measure the center (a typical value) and the spread (how variable).
- Which numbers you use depends on the shape — especially whether there are outliers.
- These summaries are the backbone of every later inference.
汇总数据的数字
- 图显示形状;汇总统计量把一个分布浓缩成几个关键数字。
- 两项任务:度量中心(典型值)和分散(多变异)。
- 用哪些数字取决于形状——尤其是否有离群值。
- 这些汇总是之后每次推断的骨干。
Measures of center
- The mean 均值 $\bar{x}$ is the arithmetic average: add all values, divide by how many.
- The median 中位数 is the middle value when the data are sorted (average the two middle if even).
- For a symmetric distribution they're close; for a skewed one they differ.
- The mean gets pulled toward a long tail; the median stays put.
中心的度量
- 平均数 $\bar{x}$ 是算术平均:把所有值相加,除以个数。
- 中位数是数据排序后中间的值(偶数个则取中间两个的平均)。
- 对对称分布它们接近;对偏斜分布它们不同。
- 平均数被长尾拉过去;中位数留在原地。
Center and spread of data · 数据的中心和离散度
The mean and median mark the center; the range, IQR, and standard deviation measure the spread. · 均值和中位数标示中心;极差、四分位距和标准差衡量离散度。
Find the median of $4, 5, 6, 7, 100$. · 求 $4, 5, 6, 7, 100$ 的中位数。
The middle value of the sorted data is $6$. · 排序后数据的中间值是 $6$。
In $4,5,6,7,100$, the mean ($24.4$) is dragged upward by the value $100$. · 在 $4,5,6,7,100$ 中,均值 ($24.4$) 被值 $100$ 向上拉动。
The mean is sensitive to extreme values. · 均值对极端值敏感。
Measures of spread
- Range 极差 = maximum − minimum (simple, but sensitive to extremes).
- IQR (interquartile range) $=Q_3-Q_1$ — the spread of the middle $50\%$.
- Standard deviation 标准差 $s_x$ — the typical distance of values from the mean.
- Bigger spread numbers mean more variability.
分散的度量
- 极差 = 最大值 − 最小值(简单,但对极端值敏感)。
- 四分位距(IQR)$=Q_3-Q_1$——中间 $50\%$ 的分散。
- 标准差 $s_x$——值到平均数的典型距离。
- 分散数字越大意味着变异越多。
For a data set with $Q_1=5$ and $Q_3=7$, find the IQR. · 对于含 $Q_1=5$ 和 $Q_3=7$ 的数据集,求 IQR。
$\text{IQR}=Q_3-Q_1=7-5=2$.
Resistant vs. sensitive
- Median and IQR are resistant 稳健 — outliers barely move them.
- Mean and standard deviation are sensitive — one extreme value can shift them a lot.
- So for skewed data or data with outliers, report the median and IQR.
- Outlier rule: a value is an outlier if it's below $Q_1-1.5\times\text{IQR}$ or above $Q_3+1.5\times\text{IQR}$.
抗离群 vs 敏感
- 中位数和 IQR 是抗离群的——离群值几乎不动它们。
- 平均数和标准差是敏感的——一个极端值能大幅移动它们。
- 所以对偏斜数据或有离群值的数据,报告中位数和 IQR。
- 离群值规则: 若一个值低于 $Q_1-1.5\times\text{IQR}$ 或高于 $Q_3+1.5\times\text{IQR}$,它就是离群值。
Which measures are resistant to outliers? · 哪些度量对离群值是稳健的?
Median and IQR resist outliers; mean and SD do not. · 中位数和 IQR 对离群值稳健;均值和 SD 不稳健。
By the $1.5\times\text{IQR}$ rule, a high outlier is any value above... · 根据 $1.5\times\text{IQR}$ 规则,高离群值是指任何高于...的值
Measure from the quartile: $Q_3+1.5\,\text{IQR}$. · 从四分位数测量:$Q_3+1.5\,\text{IQR}$。
For strongly skewed data, the best center/spread summary is... · 对于强偏态数据,最佳的中心/离散度总结是...
Resistant measures suit skewed data. · 稳健度量适合偏态数据。
Match the summary to the shape: for skewed data or outliers, use the resistant median and IQR — the mean and standard deviation get dragged by extremes. And the $1.5\times\text{IQR}$ rule measures from the quartiles ($Q_1,Q_3$), not the mean: flag values below $Q_1-1.5\,\text{IQR}$ or above $Q_3+1.5\,\text{IQR}$.
让汇总与形状匹配:对偏斜数据或离群值,用抗离群的中位数和 IQR——平均数和标准差会被极端值拖动。而 $1.5\times\text{IQR}$ 规则从四分位数($Q_1,Q_3$)而非平均数量起:标记低于 $Q_1-1.5\,\text{IQR}$ 或高于 $Q_3+1.5\,\text{IQR}$ 的值。
Data: $4, 5, 6, 7, 100$.
- Mean $=\tfrac{4+5+6+7+100}{5}=24.4$ (dragged up by $100$). Median $=6$ (unaffected).
- The median ($6$) better represents the typical value here.
- With $Q_1=5$, $Q_3=7$, $\text{IQR}=2$: $100>7+1.5(2)=10$, so $100$ is an outlier.
数据:$4, 5, 6, 7, 100$。
- 平均数 $=\tfrac{4+5+6+7+100}{5}=24.4$(被 $100$ 拉高)。中位数 $=6$(不受影响)。
- 这里中位数($6$)更好地代表典型值。
- 用 $Q_1=5$、$Q_3=7$、$\text{IQR}=2$:$100>7+1.5(2)=10$,所以 $100$ 是离群值。
Summarize center with the mean $\bar{x}$ or median, and spread with the range, IQR $=Q_3-Q_1$, or standard deviation $s_x$. The median and IQR are resistant to outliers; the mean and SD are not — so prefer them for skewed data. Flag outliers with the $1.5\times\text{IQR}$ rule.
用平均数 $\bar{x}$ 或中位数汇总中心,用极差、IQR $=Q_3-Q_1$ 或标准差 $s_x$ 汇总分散。中位数和 IQR 抗离群值;平均数和标准差则不然——所以对偏斜数据优先用它们。用 $1.5\times\text{IQR}$ 规则标记离群值。