Interval for a Difference of Means · Intervalo para a Diferença de Médias
| English | Português |
|---|---|
| paired/peəd/ | pareados |
| independent samples/ˌɪndɪˈpendənt ˈsæmplz/ | amostras independentes |
Pair before calculating the uncertainty
- The same three runners have before/after times (12,10), (15,12) and (11,10) minutes. Take before minus after.
- The paired differences are 2, 3 and 1 minutes, giving a mean difference of 2. The observational units for this analysis are three pairs, not six independent measurements.
Independent or paired?
- Comparing two means starts with one crucial question about the design. Independent samples 独立样本: two separate · separar groups (e.g. treatment vs. control).
- Paired (matched-pairs) 配对: two measurements on the same · mesmo or · ou matched subjects (before/after). The design decides which procedure you use — get it right first.
The two-sample interval
- For · A favor independent samples, use a two-sample $t$-interval for $\mu_1 - \mu_2$: $$(\bar{x}_1 - \bar{x}_2) \pm t^{*}\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$
- The variances add · adicionar under the root (each group's $\frac{s^2}{n}$). Check random design, independence between groups, the 10% condition when sampling without replacement, and the data shape in each group.
The paired interval
- For · A favor paired data, first compute each pair's difference $d = x_1 - x_2$. Then run a one-sample $t$-interval on the differences:
- $$\bar{d} \pm t^{*}\frac{s_d}{\sqrt{n}}$$$n$ is the number of pairs; $df = n - 1$. Check random design and independence across pairs, and inspect the differences for skew/outliers.
Measuring the same runners before and after training is which design?
Same subjects measured twice → paired.
For paired data, the correct interval is...
Collapse pairs to differences, then one-sample t.
For a two-sample t-interval, the two variances are added under the square root.
SE = √(s1²/n1 + s2²/n2).
Analyzing paired data as two independent samples gives the correct standard error.
It ignores dependence within pairs. The resulting standard error need not always be larger; analyse the paired differences.
The key question that decides the procedure is...
Linked → paired; unlinked → two-sample.
Match the design to the shape check for its t analysis.
Paired analysis works with differences as the observations. Marginal normal-looking groups do not automatically establish a suitable differences distribution.
Which are needed for a paired t analysis?
The within-pair relationship is the reason to use differences. Independence is required across the observational pairs.
Match the method to the design
- Paired data analyzed as two independent samples → wrong standard error, wrong answer. Independent data forced into a paired test → also wrong.
- The tell: is each value in group $1$ naturally linked to one in group $2$? If yes → paired; if no → two-sample.
The design dictates the procedure. If subjects are matched or measured twice (before/after), it's paired — collapse to differences and use a one-sample $t$-interval. If the two groups are separate and unlinked, it's a two-sample $t$-interval. Using the two-sample formula on paired data ignores the within-pair dependence and generally gives the wrong standard error. The direction of the error depends on the relationship within pairs.
Measure each runner's time before · antes and · e after · depois training.
- Same runners, two measurements → paired.
- Compute differences $d = \text{before} - \text{after}$ for each runner.
- Build a one-sample $t$-interval on the $d$'s (not a two-sample interval).
Carry the reasoning to a new case
- The paired standard error uses the variability of the differences.
- Ignoring pairing does not always inflate the standard error; its effect depends on the within-pair relationship.
For the three stated pairs, compute the mean before-minus-after difference in minutes.
The differences 2, 3 and 1 sum to 6, so the mean across three pairs is 2.
Comparing two means: independent samples use a two-sample $t$-interval $(\bar{x}_1 - \bar{x}_2) \pm t^{*}\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$; paired data collapse to differences and use a one-sample $t$-interval $\bar{d} \pm t^{*}\frac{s_d}{\sqrt{n}}$. Match the method to the design, checking conditions for both groups.