Sampling and a large data set · Échantillonnage et grand ensemble de données
| English | Français |
|---|---|
| sampling frame/ˈsæmplɪŋ freɪm/ | cadre d'échantillonnage |
Who is missing from the data?
- A weather database has a missing entry and several stations. Treating each row as identical can distort a comparison.
- This lesson studies sampling frame 抽样框: The list or population definition from which a sample is selected.
Choose the mathematical structure
- Identify the population, sampling unit and frame. Distinguish random, systematic, stratified, quota and opportunity sampling. Missing data is not zero; verify units, dates and variable definitions before comparing samples.
- State the allowed inputs and units before calculating. An equation should express the relationship, not just record a calculator entry.
Which description correctly defines sampling frame? · Quelle description définit correctement la trame d'échantillonnage ?
The list or population definition from which a sample is selected. · La liste ou la définition de population à partir de laquelle un échantillon est sélectionné.
Work through a checked case
- Check the result against the starting quantities. Substitute into the original relation, or compare the graph and numerical answer where appropriate.
A population has 120 students in one group and 80 in another. A proportional stratified sample of 30 needs 30×120/200=18 from the first and 12 from the second. Random selection is then needed within each group.
Sampling and a large data set · Échantillonnage et grand ensemble de données
Identify the population, sampling unit and frame · Identifier la population, l'unité d'échantillonnage et la trame
Allocate proportionally, then explain why random selection within each group still matters. · Allouer proportionnellement, puis expliquer pourquoi la sélection aléatoire au sein de chaque groupe reste importante.
Find the first-group allocation when sample size is 30 and group sizes are 120 and 80. · Trouver l'allocation du premier groupe lorsque la taille de l'échantillon est 30 et que les tailles des groupes sont 120 et 80.
Allocate 30×120/200=18, then select randomly within that group. · Allouer 30×120/200=18, puis sélectionner aléatoirement au sein de ce groupe.
Test a tempting shortcut
- A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values.
- When a shortcut fails, identify the assumption it breaks. Keep an exact value until the requested final rounding.
Doubling a biased sample automatically removes its selection bias. This claim is false. Explain which definition or assumption it violates.
Find the second-group allocation for that sample. · Trouver l'allocation du deuxième groupe pour cet échantillon.
Allocate 30×80/200=12. Check that 18+12=30. · Allouer 30×80/200=12. Vérifier que 18+12=30.
Doubling a biased sample automatically removes its selection bias. · Doubler un échantillon biaisé ne supprime pas automatiquement son biais de sélection.
A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values. · Un grand échantillon biaisé reste biaisé. La stratification n'est pas identique à sélectionner qui que ce soit qui est disponible dans chaque groupe. La familiarité avec les grands ensembles de données AQA nécessite les données réelles fournies et leurs métadonnées, et non des valeurs météorologiques inventées.
Interpret a new situation
- For AQA, use the official large data set in a supervised spreadsheet task: identify a variable, justify a comparison, inspect missing values, create a display and explain a limitation. Save the decisions with the analysis.
- A complete solution gives the mathematical result and explains what it means. Check that it is possible in the stated context.
A sample includes 45 of 300 people. Find its percentage. · Un échantillon comprend 45 parmi 300 personnes. Trouver son pourcentage.
Percentage=45/300×100=15%. · Pourcentage=45/300×100=15%.
Match each part of a complete solution to its purpose. · Associer chaque partie d'une solution complète à son but.
An assumption justifies the model; a check tests the result; interpretation connects it to the question. · Une hypothèse justifie le modèle ; une vérification teste le résultat ; l'interprétation le relie à la question.
Use this in your course
- 7357 · A-level · K. Match the target tier and specification before assigning extensions.
- Give the method before the final answer, and use the paper's calculator and formula rules. Review a wrong answer by locating the first invalid step.
The list or population definition from which a sample is selected. Choose the relationship, show the method, check its assumptions and interpret the result.