Exploring One-Variable Data
AP Statistics Topic 1 9:03 English narration · English + 中文 subtitles burned in
Chapters
Transcript
Measure the same thing twenty times and you will not get the same answer twenty times.
同一个东西量二十次,你不会得到二十个相同的答案。
Twenty students measure the same table. The values vary.
二十个学生量同一张桌子,数值各不相同。
That is not a mistake, and it is not a failure of the ruler.
这不是出错,也不是尺子的问题。
Variation is the normal state of real data, and statistics is the science of learning from it.
变异是真实数据的常态,而统计学就是从变异中学习的科学。
So statistics does not ask, how long is the table.
所以统计学问的不是"这张桌子有多长"。
It asks a statistical question, one that expects an answer built from data that vary.
它问的是一个统计问题, 一个预期由变动的数据来回答的问题:典型的测量值是多少?
What is a typical measurement, and how spread out are they?
它们分散得有多开?
Unit one is about describing a single variable well.
第一单元讲的是如何把单个变量描述清楚。
Let's begin.
让我们开始吧。
Two distinctions run through the whole course, and the exam is strict about the words.
有两个区分贯穿整门课程,而且考试对用词非常严格。
First. A number that summarises the whole population is a parameter. A number that summarises a sample is a statistic.
第一, 概括整个总体的数叫参数,概括一个样本的数叫统计量。
We can almost never measure the population, so we compute the statistic to estimate the parameter.
我们几乎永远无法测量整个总体,所以我们算出统计量,用它来估计参数。
Notation keeps them apart. Greek letters for parameters, Roman letters for statistics.
记号能帮你分清:参数用希腊字母,统计量用罗马字母。
Second, the two halves of the course.
第二,这就是本课程的两半。
Descriptive statistics summarise the data you have. Inferential statistics use a sample to make claims about the population.
描述统计只概括你手上真正拥有的数据; 推断统计则用一个样本,对更大的总体做出并检验论断。
A variable is a characteristic that can differ between individuals.
变量是个体之间可能不同的某种特征。 它有两类,而你接下来做的一切都取决于你手上是哪一类。
There are two kinds, and everything you do next depends on which one you have.
分类变量取的是标签或组别,比如眼睛颜色、手机品牌。
A categorical variable takes labels or groups. Eye colour. Brand of phone.
你可以数每一组里有多少个,但不能求平均。
You can count how many fall in each group, but you cannot average them.
定量变量取的是可以做算术的数字,比如身高、年龄。
A quantitative variable takes numbers you can do arithmetic on, like height and age. And it splits again. Discrete means countable.
定量变量还能再分:离散是可以数出来的,比如兄弟姐妹的个数; 连续是在一个刻度上量出来的,比如身高。
Continuous means measured on a scale. Here is the test that never fails. If averaging the values would be nonsense, it is categorical.
有一个百试不爽的判据: 如果把这些值平均起来毫无意义,那这个变量就是分类变量。
For a categorical variable, start with a frequency table, which lists each category and its count.
对分类变量,先做一张频数表,把每个类别和它的计数列出来。
Then a relative frequency table divides each count by the total, giving a proportion.
然后相对频数表把每个计数除以总数,得到比例。 为什么要多做第二张?
Why bother? Because it lets you compare groups of different sizes fairly.
因为它能让不同大小的组公平地比较。
Forty out of fifty is a bigger share than sixty out of a hundred, even though sixty is the bigger number.
五十里的四十,比一百里的六十占比更大, 尽管六十这个数更大。
To display it, a bar chart draws each category as a separated bar; a pie chart shows each category's share of the whole.
要展示它,条形图把每个类别画成一根分开的条, 高度让你一眼就能比较各类别。 饼图显示每个类别占整体的份额。
Keep the bars separated.
记住条与条之间要分开。
Touching bars mean something else, as you are about to see.
挨在一起的条表示别的东西,你马上就会看到。
For a quantitative variable you have three displays.
对定量变量,你有三种图。
A dotplot puts one dot above each value, so you keep every individual data point.
点图在每个值的上方点一个点,所以每一个个体数据都保留下来。
A stem and leaf plot splits each number into a leading part and a trailing digit, so the numbers stay readable.
茎叶图把每个数拆成前面的一部分和末位数字,所以数字本身仍然读得出来。
And a histogram groups the values into intervals called bins.
直方图把数值分到叫做"组"的区间里,在每个区间上画一根条。
All three show the distribution, meaning how the values spread out.
这三种图都显示分布,也就是数值是怎么散开的。
But a histogram hides a choice. The bin width changes the picture.
但直方图里藏着一个选择: 组距会改变整张图的样子。
Too wide and the shape flattens. Too narrow and it turns into noise.
太宽,形状就被压平得什么都看不出;太窄,就变成一堆噪声。
Choose the width that reveals the shape.
要选出能把形状显现出来的组距。
Now, how do you describe a distribution?
那么该怎么描述一个分布?
Say four things, and remember them with the word SOCS. Shape. Outliers.
说四件事,用 SOCS 这个词来记:形状、离群值、中心、离散程度。
Centre. Spread.
先说形状。
Shape first. Is it symmetric, or skewed?
它是对称的,还是偏斜的?
Here is the rule students reverse every year. The skew is named for where the tail points, not where the pile of data sits.
这里有一条学生每年都会记反的规则: 偏斜的方向按尾巴指向哪边来命名,不是按数据堆在哪边。
A long tail stretching right is skewed right, even though most of the data are on the left.
一条长尾巴伸向右边就叫右偏,尽管大部分数据在左边。
Also count the peaks. One is unimodal, two is bimodal, roughly flat is uniform.
还要数峰的个数: 一个主峰叫单峰,两个叫双峰,大致平坦叫均匀。
Then Outliers, Centre and Spread.
然后是离群值、中心和离散程度。
And say all of it in context, with units.
而且不管你说什么,都要结合情境、带上单位。 光写数字得不到分。
Now put numbers on centre and spread. Two measures of centre.
现在给中心和离散程度配上数字。
The mean is the total divided by how many values there are.
中心有两个度量。 均值是总和除以数值的个数。
The median is the middle value once sorted.
中位数是排好序之后正中间的那个值。
They behave very differently. Drag one point far to the right and the mean is pulled toward the tail, while the median barely moves.
它们的表现很不一样: 把一个点拖到很远的右边,均值会被拉向尾巴,而中位数几乎不动。
Three measures of spread.
离散程度有三个度量。
The range is the largest minus the smallest.
极差就是最大值减最小值。
The interquartile range is the third quartile minus the first, covering the middle fifty percent.
四分位距是第三四分位数减第一四分位数, 覆盖中间百分之五十。
And the standard deviation is a typical distance from the mean — its square is the variance.
标准差是与均值的典型距离,它的平方就是方差。
Now the pairing rule. Median and interquartile range are resistant, so use them for skewed data. Use the mean and standard deviation when the data are roughly symmetric.
现在讲配对规则:中位数和四分位距是稳健的,所以偏斜数据用它们; 数据大致对称时,用均值和标准差。
The percentile of a value is the percent of the data at or below it.
一个值的百分位数,是小于或等于它的数据所占的百分比。
So the median is the fiftieth percentile, and the first quartile is the twenty fifth.
所以中位数是第五十百分位数,第一四分位数是第二十五百分位数。
There is a graph built to read these off directly.
有一种图就是专门用来直接读出这些的:累积相对频率图。
A cumulative relative frequency graph plots, for each value, the proportion of the data at or below it.
它对每一个值,画出小于或等于它的数据所占的比例。
It starts at zero, rises to one, and never goes down.
它总是从零开始,总是升到一,而且永远不会下降。
Read it in both directions.
要会双向读它。
Go up from a value to the curve, then across to the axis, and you have its percentile.
从某个值往上走到曲线,再横着走到坐标轴,你就得到它的百分位数。
Reverse the steps to get the value at any percentile.
把步骤反过来,你就能得到任意给定百分位数所对应的值。
Five numbers describe a distribution surprisingly well.
有五个数,能把一个分布描述得出奇地好:最小值、第一四分位数、中位数、第三四分位数、最大值。
The minimum, the first quartile, the median, the third quartile, and the maximum.
这就是五数概括,而箱线图就是把它画出来。
That is the five number summary, and a boxplot draws it.
箱子从第一四分位数延伸到第三四分位数,里面在中位数处画一条线, 两条须伸向不是离群值的最极端的数值。
A box runs from the first quartile to the third, with a line inside at the median, and two whiskers reach out to the most extreme values that are not outliers.
这就带出一个问题:什么算离群值?
So what counts as an outlier?
用这条规则: 如果一个点超出某个四分位数达到四分位距的一点五倍以上,它就是离群值。
A point is an outlier if it lies more than one and a half times the interquartile range beyond a quartile.
做一道。
Work one.
假设第一四分位数是 20,第三四分位数是 32。
Suppose the first quartile is twenty and the third is thirty two.
四分位距是 12, 1.5 乘以 12 等于 18。
The interquartile range is twelve, and one and a half times twelve is eighteen.
所以两道界线在 2 和 50。
So the fences sit at two and at fifty.
低于 2 或高于 50 的都会被标记出来。
Boxplots really earn their keep when you compare groups.
箱线图真正发挥价值,是在你比较不同组的时候。
Draw them side by side on the same scale, and differences in centre and spread jump out.
把它们在同一个刻度上并排画出来, 中心和离散程度的差别就一目了然。
But the marks are for the writing, not the drawing.
但分数给的是文字,不是图。
Compare shape, centre and spread, and mention outliers.
要比较形状、比较中心、比较离散程度,还要提到离群值。
And use comparative words. Group A has a higher median than group B. Group A is more spread out than group B.
而且要用比较性的词语: A 组的中位数比 B 组高;A 组比 B 组更分散。
Here is the mistake that costs the mark every year. Describing each group separately and never making the comparison explicit.
这里有一个每年都在丢分的错误:把每一组分别描述一遍,一组接着一组, 却从来没有把比较明确地写出来。
Two descriptions are not a comparison.
两段描述并不等于一次比较。
One shape appears so often that it gets its own model.
有一种形状出现得太频繁,以至于它有了自己的模型。
The normal distribution is a symmetric bell shape, and two numbers describe it completely.
正态分布是一个对称的钟形, 两个数就能把它完全确定:均值决定中心在哪里,标准差决定这口钟有多宽。
The mean fixes where the centre sits, and the standard deviation fixes how wide the bell is.
标准差变大,曲线就变平变开;标准差变小,曲线就收紧变高。
Increase the standard deviation and the curve flattens. Decrease it and the curve pulls in and grows tall.
有一个想法让其余一切都能运作:在这条曲线下,比例就是面积。
And one idea makes everything else work. Under this curve, a proportion is an area, and the total area is exactly one.
落在任意范围内的数值所占的比例,等于曲线在那个范围上方所围出的面积,而总面积恰好是一。
For a normal curve you can do a lot with one memorised rule.
对正态曲线,只要记住一条规则就能做很多事。
About sixty eight percent of the values lie within one standard deviation of the mean.
大约百分之六十八的数值落在离均值一个标准差以内。 大约百分之九十五落在两个标准差以内。
About ninety five percent lie within two.
大约百分之九十九点七落在三个标准差以内。
And about ninety nine point seven percent lie within three. That is the empirical rule.
这就是经验法则,常常直接叫做 68、95、99.7。
The trick that makes it useful is symmetry. Whatever is left over, split it evenly between the two tails.
让它真正好用的窍门是对称性: 剩下的部分要在两条尾巴之间平均分配。
If ninety five percent is inside two standard deviations, then two point five percent sits in each tail.
如果两个标准差以内是百分之九十五, 那么以外就是百分之五,所以每条尾巴里各有百分之二点五。
But the empirical rule only helps at whole standard deviations.
但经验法则只在整数个标准差处有用。
For anything else, standardise.
其他情况就要标准化。
A z-score says how many standard deviations a value sits from the mean.
标准分数表示一个值离均值有多少个标准差:减去均值,再除以标准差。
Subtract the mean, then divide by the standard deviation. Positive is above the mean, negative is below.
正的标准分数在均值上方,负的在下方,标准分数为零就正好在均值处。
Work one.
做一道。
Test scores are normal with mean five hundred and standard deviation one hundred.
考试成绩服从正态分布,均值是 500,标准差是 100。
A score of seven hundred is two hundred above the mean, which is two standard deviations, so its z-score is two.
700 分比均值高 200 分, 也就是两个标准差,所以它的标准分数是 2。
By the empirical rule, two point five percent of scores lie above it, so that score is at about the ninety seven point fifth percentile. And you can run it backwards.
根据经验法则, 有百分之二点五的成绩高于它,所以这个分数大约在第 97.5 百分位。
Start from the percentile, find the z-score it matches, then multiply by the standard deviation and add the mean.
而且你可以反过来做:从百分位出发,找出标准分数,再乘以标准差、加上均值。
Three marks students lose.
学生最常丢分的三个地方。
First, when a question says describe the distribution, give all four letters of SOCS, not just the centre.
第一,题目说"描述这个分布"时, 要把 SOCS 四个字母都写全,不能只说中心。
Second, the tail names the skew, so a long right tail is skewed right.
第二,偏斜按尾巴命名,所以长尾巴在右边就是右偏。
Third, every answer goes in context, with units and with the variable named.
第三,每个答案都要结合情境,带上单位并说出变量名。
A bare number is not an answer.
光写一个数不算答案。