Skip to content

Inference for Categorical Data: Proportions

AP Statistics Topic 6 7:52 English narration · English + 中文 subtitles burned in

space play · ←/→ 5s · j/l 10s · f fullscreen · ,/. speed

Chapters

Transcript
A city wants to know how many people support a new policy. 一座城市想知道有多少市民支持一项新政策。
We cannot ask everyone, so we ask two hundred at random. 我们不可能问遍所有人,于是随机问二百个人。
A hundred and twenty say yes — sixty percent. 一百二十人说支持——百分之六十。
Ask a different two hundred and the answer moves: fifty seven, then sixty three. 换另外二百个人再问,答案就变了:先是百分之五十七, 再是百分之六十三。
Keep sampling, and the answers pile up into a shape: approximately normally distributed, a normal curve, centred on the true proportion. 不停地抽样,这些答案会堆成一个形状:近似正态分布, 一条正态曲线,以真实比例为中心。
Make each sample bigger, and the pile tightens. 让每个样本更大,这堆答案就会收紧。
Unit six turns one sample into a confident statement about a whole population, and tests a claim. 第六单元教你把一个样本变成关于整个总体的可靠结论,并检验一个论断。
Let's begin. 让我们开始吧。
We never know the true proportion, so we give a range: an estimate, plus or minus a margin of error. 我们永远不知道真实比例,所以给出一个区间:一个估计值,加上或减去误差幅度。
The estimate is the sample proportion. 估计值就是样本比例。
The margin of error is a critical value times the standard error — for ninety five percent confidence, one point nine six. 误差幅度等于临界值乘以标准误——在百分之九十五的置信水平下, 是一点九六。
But first, three conditions. The sample must be random, under ten percent of the population, and the counts must be large: ten successes, ten failures. 但要先看三个条件:样本必须随机,不超过总体的百分之十, 计数还要够大:十个成功,十个失败。
The sample proportion is a hundred and twenty over two hundred, or zero point six. 样本比例是一百二十除以二百,等于零点六。
Check the counts: a hundred and twenty and eighty, both above ten. 检查计数:一百二十和八十,都大于十。
The standard error is zero point zero three four six. 标准误是零点零三四六。
The margin is one point nine six times that: zero point zero six eight. 误差幅度是一点九六乘以它:零点零六八。
The interval runs from zero point five three two to zero point six six eight. 这个区间从零点五三二到零点六六八。
Say it properly: we are ninety five percent confident the true proportion is between fifty three point two and sixty six point eight percent. 要这样表述:我们有百分之九十五的把握,认为真实比例在百分之五十三点二 到百分之六十六点八之间。
So what does ninety five percent confidence mean? 那么百分之九十五的置信水平到底是什么意思?
Here are twenty samples, each with its own interval. 这里有二十个样本,每个都给出自己的区间。
Nineteen capture the true proportion; one misses. 其中十九个抓住了真实比例,有一个落空了。
The ninety five percent describes the method, not this one interval — and you will never know which kind yours is. 这个百分之九十五描述的是方法, 而不是你手上这一个区间——而且你永远不会知道你的属于哪一种。
An interval settles arguments. 区间还能用来判定论断。
Ours runs from zero point five three two to zero point six six eight. 我们的区间从零点五三二到零点六六八。
Someone claims exactly half the city supports it: zero point five is outside, so that is evidence against the claim. 有人声称正好一半市民支持它:零点五落在区间之外,所以这是反对该论断的证据。
Someone else claims sixty five percent: that is inside, so the data agree. 另一个人声称是百分之六十五:这个值在区间之内,所以数据与它一致。
Two rules — a bigger sample narrows the interval, a higher level widens it. 两条规律——样本越大,区间越窄;置信水平越高,区间越宽。
Now the reverse question: how many people must you survey? 现在反过来问:你必须调查多少人?
The margin of error must be at most zero point zero three, at ninety five percent. 误差幅度最多只能是零点零三, 置信水平是百分之九十五。
Rearrange the margin formula for the sample size. 把误差幅度的公式变形,解出样本量。
One problem — you have not sampled yet. 有个问题——你还没有抽样。
Use zero point five: it makes that product biggest, giving the safe largest required sample size. 那就用零点五:它让那个乘积达到最大,给出最保险的最大所需样本量。
Put the numbers in: about one thousand and sixty seven. 代入数字:大约一千零六十七。
Always round up — survey one thousand and sixty eight. 永远向上取整——调查一千零六十八人。
Now the other half of inference. 现在讲推断的另一半:检验一个论断。
A significance test weighs evidence against a claim. 显著性检验衡量的是反对某个论断的证据。
Write two hypotheses: the null says the proportion equals a claimed value; the alternative says it is different. 写出两个假设:原假设说比例等于某个被声称的值;备择假设说它不一样。
Then count how many standard errors your sample sits from that value — the test statistic. 然后数一数,你的样本离那个值有多少个标准误——这就是检验统计量。
The trap: for the conditions and the standard error, use the claimed value, not the sample proportion. 陷阱在于:检查条件和计算标准误时,要用被声称的值,而不是样本比例。
A company claims ninety percent are satisfied; in a hundred customers, eighty four are. 一家公司声称有百分之九十的顾客满意;在一百位顾客里,八十四人满意。
The null hypothesis: the true proportion is zero point nine. 原假设:真实比例是零点九。
The alternative: it is not. 备择假设:不是零点九。
Check the conditions with the claimed value — ninety and ten, both at least ten. 用被声称的值检查条件——九十和十,都至少是十。
Now the statistic. 现在算统计量。
The sample proportion is zero point eight four, so the top is minus zero point zero six, and the standard error is zero point zero three. 样本比例是零点八四,所以分子是负零点零六,标准误是零点零三。
Divide, and you get minus two. 相除,得到负二。
The sample sits two standard errors below the claim. 这个样本比论断低了两个标准误。
Where does minus two land? 负二落在哪里?
At the five percent level the middle ninety five percent is the do-not-reject region, with edges at minus and plus one point nine six. 在百分之五的水平上,中间的百分之九十五是不拒绝的区域, 边界在负一点九六和正一点九六。
Minus two is just past the left edge. 负二刚好越过了左边那条边界。
So we reject the claim. 所以我们拒绝这个论断。
Now the p-value. 现在说 P 值。
If the null hypothesis were true, results would scatter like this curve — ours sits out here. 如果原假设为真,结果就会像这条曲线这样散布—— 而我们的结果落在这里。
The p-value is the tail area beyond it: the chance of a result as extreme or more extreme, assuming H zero is true. P 值就是它之外那块尾部面积:在假定原假设为真的前提下, 出现同样极端或更极端结果的概率。
Each tail here holds two point three percent, and ours is two-sided, so count both: zero point zero four six. 这里每条尾巴占百分之二点三, 而我们是双侧的,所以两边都要算:零点零四六。
A small p-value is evidence against the claim, not the probability that it is true. P 值小,是反对该论断的证据,而不是它为真的概率。
Now finish properly. 现在把结论写完整。
Compare the p-value with the significance level, alpha — usually zero point zero five. 把 P 值和显著性水平阿尔法比较——通常是零点零五。
If the p-value is at most alpha, reject the null hypothesis; if bigger, fail to reject. 如果 P 值不超过阿尔法,就拒绝原假设;如果更大,就不拒绝。
Ours was zero point zero four six, so we reject. 我们的是零点零四六,所以拒绝。
The sentence: there is convincing evidence the satisfaction rate is not ninety percent. 结论句这样写:有令人信服的证据表明, 真实的满意率不是百分之九十。
Never write that you accept the null hypothesis. 永远不要写你接受原假设。
No test is perfect: the data are random. 没有哪个检验是完美的:数据本身是随机的。
Two worlds — the null hypothesis true, and false. 这里有两个世界——原假设为真,和为假。
Wherever you put the cutoff, you choose which mistake to make. 无论你把界线放在哪里,你都在选择犯哪一种错误。
Move it right and false alarms become rare, but you miss real effects. 把它右移,虚惊会变少,但你会漏掉真实的效应。
Move it left and you catch everything — including things that were never there. 把它左移,你什么都能抓到—— 包括那些根本不存在的东西。
A Type one error is rejecting a null hypothesis that was true — a false alarm, with probability alpha. 第一类错误是拒绝了一个本来为真的原假设——一次虚惊,概率是阿尔法。
A Type two error is failing to reject one that was false — a missed detection, with probability beta. 第二类错误是没能拒绝一个本来为假的原假设——一次漏检,概率叫贝塔。
Power is one minus beta, the chance of catching a real effect; it rises with a larger sample, a bigger effect, or a larger alpha. 检验效能等于一减贝塔,也就是抓到真实效应的概率;样本越大、效应越大、 阿尔法越大,效能就越高。
Often a question compares two groups. 很多题目要比较两个组。
The estimate is now the difference between the two sample proportions, and the standard error adds a piece from each sample under one square root. 现在的估计值是两个样本比例之差; 标准误在同一个根号下把来自两个样本的两部分加起来。
The conditions must hold in both samples, and the samples must be independent. 两个样本都必须满足条件,而且两个样本必须相互独立。
A difference interval comes down to one number: zero, which means no difference at all. 读懂一个差的区间,归根结底看一个数:零,它表示完全没有差别。
If the whole interval sits above zero, the first proportion is bigger. 如果整个区间都在零的上方,第一个比例更大。
If it sits entirely below zero, the second is bigger. 如果整个区间都在零的下方, 第二个比例更大。
But if the interval contains zero, no difference is plausible, so you have no evidence. 但如果区间包含零,"没有差别"就是合理的,于是你没有证据。
Whichever it is, state the direction in context, and name the two groups. 不管是哪一种,都要结合情境说明方向,并说出是哪两个组。
The null hypothesis says the two proportions are equal — so if it is true, there is really only one proportion. 原假设说这两个比例相等——所以如果它为真, 其实只有一个比例。
So add all the successes, divide by the total sample size: the combined, or pooled, proportion. 那就把所有成功数加起来,除以总样本量, 这就是合并比例。
Put it inside the standard error. 把它放进标准误里。
Try one: a hundred and twenty of two hundred against a hundred and thirty of two hundred and fifty. 试一道:二百人里有一百二十人, 对二百五十人里有一百三十人。
The pooled proportion is zero point five five six, and the statistic is one point seven. 合并比例是零点五五六,统计量是一点七。
That gives about zero point zero nine — bigger than alpha, so we fail to reject. 算出来大约是零点零九——比阿尔法大,所以我们不拒绝原假设。
There is not enough evidence of a real difference. 没有足够的证据说明存在真实差别。
Three marks students lose every year. 学生每年都会丢的三分。
First, state the conditions before any proportion inference, and in a test use the claimed value. 第一,做任何比例推断之前先写条件, 而且在检验里要用被声称的值。
Second, ninety five percent confidence describes the method, not your one interval. 第二,百分之九十五的置信水平描述的是方法, 而不是你手上那一个区间。
Third, never write that you accept the null hypothesis. 第三,永远不要写你接受原假设。
Get these right, and this unit is yours. 把这三点做对,这个单元就是你的了。

Log in or create account

IGCSE, A-Level & AP