Skip to content

Inference for Categorical Data: Chi-Square

AP Statistics Topic 8 7:49 English narration · English + 中文 subtitles burned in

space play · ←/→ 5s · j/l 10s · f fullscreen · ,/. speed

Chapters

Transcript
Roll a fair die sixty times. 把一枚公平的骰子掷六十次。
Every face should come up ten times. 每一面本该出现十次。
But it never does — eight here, thirteen there. 但它从来不会正好这样——这一面八次, 那一面十三次。
The gaps are always there. 差距总是存在。
So roll another sixty, and again, and again, and each time measure how far the counts sit from ten. 那就再掷六十次,一次又一次,每一次都量一量这些次数 离十有多远。
Those measurements pile up into a shape: a curve that leans hard to the right. 这些量出来的值会堆成一个形状:一条向右倾斜得很厉害的曲线。
Small gaps are ordinary. 小的差距很平常。
Only far out in that long tail are the gaps too big to blame on luck. 只有远在那条长尾里,差距才大到不能用运气来解释。
Unit eight is about counts. 第八单元讲的是计数数据。
One variable against a claimed pattern, or two variables in a table. 可以是一个变量对照某个声称的分布,也可以是表格里的两个变量。
One statistic, three names. 同一个统计量,三种名称。
Let's begin. 让我们开始吧。
Here is the whole idea in one picture. 整个想法就在这一张图里。
A shop counts its customers on five weekdays. 一家商店统计了五个工作日的顾客人数。
The blue bars are what happened — the observed counts. 蓝色柱是实际发生的——观测频数。
The grey bars are what you would expect if every day were the same: twenty eight each. 灰色柱是在每天都一样这个前提下应有的结果: 每天二十八。
Monday and Friday run high, Thursday runs low. 周一和周五偏高,周四偏低。
Chi-square turns all five gaps into a single number. 卡方把这五个差距变成一个数。
Look at each piece. 逐个来看。
First, square the gap. 第一,把差距平方。
Squaring makes it positive, so under and over both count, and a big gap counts far more. 平方让它变成正的,所以偏低和偏高都算数, 而且大差距的分量重得多。
Second, divide by the expected count: a gap of five is serious when you expected ten, but nothing when you expected five hundred. 第二,除以期望频数:当你期望十的时候,差五很严重; 当你期望五百的时候,差五就什么也不算。
Third, add one term for every category — including the ones that fit perfectly and add zero. 第三,每个类别都加一项—— 包括那些完全吻合、加零的类别。
The total is the chi-square statistic. 总和就是卡方统计量。
So how big is big? 那么多大才算大呢?
Compare the statistic with its own distribution. 把统计量和它自己的分布比较。
On the left is the family of chi-square curves. 左边是卡方曲线族。
Every one starts at zero and leans right, and the only thing that changes the shape is the degrees of freedom. 每一条都从零开始,都向右倾斜,唯一改变形状的就是自由度。
On the right, the p-value is the area in the tail beyond your statistic. 右边, P 值就是统计量之外那块尾部面积。
Past the critical value that area falls below five percent, and you reject the claim. 越过临界值,这块面积就小于百分之五, 于是你拒绝该论断。
The first of the three tests is goodness of fit, or GOF. 三种检验中的第一种是拟合优度,也叫 GOF。
It takes one categorical variable and checks it against a claimed distribution. 它拿一个分类变量,去对照某个声称的分布。
The null hypothesis says the distribution is exactly as claimed; the alternative says at least one proportion is different, without saying which. 原假设说这个分布就和声称的一样;备择假设说至少有一个比例不同,但不指明是哪一个。
Each expected count is the sample size times the claimed proportion, every expected count must be at least five, and the degrees of freedom are the number of categories minus one. 每个期望频数等于样本量乘以声称的比例,每个期望频数都必须至少是五, 自由度等于类别数减一。
Let's do one in full. 我们完整地做一道。
A die is rolled sixty times: eight, ten, twelve, nine, eleven, ten. 一枚骰子掷六十次:八、十、十二、九、十一、十。
The hypotheses: the die is fair, every face has probability one sixth, against at least one face being different. 假设是:骰子是公平的,每一面的概率都是六分之一;备择假设是至少有一面不同。
The expected counts are sixty times one sixth — ten for every face, all above five. 期望频数是六十乘以六分之一——每一面都是十,都大于五。
Now one term per face. 现在每一面写一项。
Eight is two away from ten, so four divided by ten, zero point four. 八离十差二,所以是四除以十,零点四。
Twelve gives zero point four as well. 十二同样给出零点四。
Nine and eleven give zero point one each, and the two tens give exactly zero. 九和十一各给出零点一, 两个十正好给出零。
Add them up: the statistic is one point zero, with five degrees of freedom. 加起来:统计量是一点零,自由度是五。
Now finish it. 现在把它做完。
With five degrees of freedom, where does one point zero land? 自由度是五时,一点零落在哪里?
Right under the fat part of the curve. 就在曲线最厚的地方。
Almost the whole area lies to its right, so the p-value is about zero point nine six. 几乎全部面积都在它右边,所以 P 值大约是零点九六。
That is far above zero point zero five, so we fail to reject the null hypothesis. 这远大于零点零五, 所以我们不拒绝原假设。
In context: there is not convincing evidence that the die is unfair. 结合情境说:没有令人信服的证据表明这枚骰子不公平。
Careful — that is not the same as proving it is fair. 要小心——这和证明它公平并不是一回事。
Now the second setting: a two-way table, where every person is counted in one row and one column. 现在是第二种情形:双向表,每个人都被记在某一行和某一列里。
The expected counts no longer come from a claim — they come from the table's own totals. 期望频数不再来自某个论断——它们来自表格自身的合计。
For each cell, multiply its row total by its column total and divide by the grand total. 对每个格子, 用它的行合计乘以列合计,再除以总计。
That is the count you would see if row and column had nothing to do with each other. 这就是当行与列彼此无关时你会看到的计数。
The degrees of freedom are rows minus one, times columns minus one. 自由度等于行数减一乘以列数减一。
Two different tests share that arithmetic, and the design decides which one you are doing. 有两种检验共用这套算法,而是由研究设计决定你在做哪一种。
Several separate samples — men and women surveyed separately — with one variable measured in each: that is homogeneity, and the null hypothesis says the distribution is the same across several populations or groups. 几个彼此独立的样本——比如分别调查男性和女性——每个样本里测量同一个变量: 这是同质性检验,原假设说若干个总体或组的分布相同。
One sample with two variables recorded on every person: that is independence, and the null hypothesis says the two variables are not associated within a single population. 只抽一个样本, 在每个人身上记录两个变量:这是独立性检验,原假设说这两个变量在同一个总体内没有关联。
Here is a two-way example. 来看一个双向表的例子。
Two hundred Grade eleven and two hundred Grade twelve students said how they travel to school: walk, bus, bike or car. 两百名十一年级学生和两百名十二年级学生说出了自己上学的方式: 步行、公交、自行车或小汽车。
Two separate samples, one variable — so this is a test for homogeneity, and the null hypothesis says the two grades travel the same way. 两个彼此独立的样本,一个变量——所以这是同质性检验, 原假设说两个年级的出行方式相同。
Add the margins first: one hundred walk, one hundred and forty take the bus, eighty bike, eighty go by car, four hundred students in all. 先把合计加出来:步行一百人,公交一百四十人, 骑车八十人,小汽车八十人,一共四百人。
Now the expected counts. 现在算期望频数。
For Grade eleven walkers, two hundred times one hundred over four hundred is fifty. 十一年级步行的格子: 两百乘一百再除以四百,等于五十。
Do the same for all eight cells. 八个格子都这样做。
The degrees of freedom are one times three, which is three. 自由度是一乘三,等于三。
Now the same sum as before: one term for every cell. 现在还是同样的求和:每个格子一项。
Grade eleven walkers — sixty observed against fifty expected, a gap of ten, squared is one hundred, divided by fifty gives two point zero. 十一年级步行的格子——观测六十,期望五十, 差十,平方是一百,除以五十得到二点零。
Work through all eight cells the same way. 八个格子都用同样的方法算一遍。
The walking terms are the biggest, so that is where the two grades differ most. 步行这两项最大,说明两个年级在这里差别最大。
Add every term and the statistic is nine point three six, with three degrees of freedom. 把所有项加起来,统计量是九点三六, 自由度是三。
Three degrees of freedom gives us this curve. 自由度为三对应的就是这条曲线。
At the five percent level the critical value is seven point eight one five, and everything past it is the reject region. 在百分之五的水平上,临界值是七点八一五, 越过它的部分就是拒绝域。
Our nine point three six sits inside it, so the p-value is below five percent — about zero point zero two five. 我们的九点三六落在里面,所以 P 值小于百分之五—— 大约是零点零二五。
So we reject the null hypothesis. 所以我们拒绝原假设。
Now write it the way the exam wants, in context: because the p-value is below zero point zero five, there is convincing evidence that Grade eleven and Grade twelve travel to school in different ways. 现在按考试要求,结合情境把它写出来:因为 P 值低于零点零五, 有令人信服的证据表明十一年级和十二年级上学的方式不同。
Watch the wording. 注意措辞。
Homogeneity finds a difference between the groups; independence finds an association between two variables. 同质性找到的是各组之间的差异;独立性找到的是两个变量之间的关联。
And a chi-square test never shows cause — it cannot tell you why. 而且卡方检验从不说明因果——它无法告诉你原因。
Which procedure? 该用哪种方法?
Count the samples and the variables. 数一数样本数和变量数。
One sample, one variable against a claimed distribution — goodness of fit. 一个样本、一个变量、对照某个声称的分布—— 拟合优度。
Several samples, one variable — homogeneity. 若干个样本、一个变量——同质性。
One sample, two variables — independence. 一个样本、两个变量——独立性。
If you are only comparing two proportions, the two-sample test and a chi-square test agree exactly — but only when the claim has no direction, a two-tailed alternative. 如果你只是在比较两个比例,双样本检验和卡方检验的结论完全一致—— 但前提是论断不带方向,也就是双尾备择。
Chi-square looks at the right tail alone, so it can never say which group is bigger. 卡方只看右尾,所以它无法说明哪一组更大。
Three marks students lose every year. 学生每年都会丢的三分。
First, always divide by the expected count, never the observed one — and check every expected count is at least five before you start. 第一,永远除以期望频数,不要除以观测频数—— 而且开始之前先检查每个期望频数至少是五。
Second, get the degrees of freedom right: categories minus one for goodness of fit, rows minus one times columns minus one for a table. 第二,自由度要写对: 拟合优度是类别数减一,表格是行数减一乘以列数减一。
Third, remember chi-square has no direction. 第三,记住卡方没有方向。
Get these right, and this unit is yours. 把这三点做对,这个单元就是你的了。

Log in or create account

IGCSE, A-Level & AP