Data science
| English | Français |
|---|---|
| data science/ˈdeɪtə ˈsaɪəns/ | science des données |
| descriptive statistics/dɪˈskrɪptɪv stəˈtɪstɪks/ | statistiques descriptives |
| visualisation/ˌvɪʒuːəlaɪˈzeɪʃn/ | visualisation |
| correlation is not causation/ˌkɒrɪˈleɪʃn ɪz nɒt kɔːˈseɪʃn/ | la corrélation n'implique pas la causalité |
Five students answer, but only four give a time
- A fictional commute survey has five rows: 10, 20, 30, 40 and one blank, measured in minutes. Calling the blank zero changes the result.
- Data science 数据科学 combines a question, data preparation, analysis and interpretation. Begin by stating which students and which journey the question concerns.
Check what a suspicious value means
- Inspect units, impossible values, repeated rows and missing fields. A repeated value is not automatically a duplicate person, and an unusual value is not automatically an error.
- Keep a record of each correction or exclusion. If one field is blank, other usable fields in that row may still support a different analysis.
Deleting every row with any blank field is always a neutral cleaning decision.
Useful fields may remain, and missingness can bias which observations are retained.
Two equal commute times prove that one row is a duplicated respondent.
Different respondents can give the same time. Investigate identity and collection records before deleting data.
Use the denominator that matches the question
- Descriptive statistics 描述性统计 summarize the observed data. Here the four recorded times total 100 minutes: mean 25 and median 25, with four valid times from five respondents.
- A visualisation 可视化 should state its units and population. Do not label the mean as the mean for all five students unless the missing value is resolved.
Recorded times are 10, 20, 30 and 40 minutes, with one additional blank. What is the mean of the recorded times?
The total 100 is divided by four valid recorded times, giving 25 minutes.
Choose the chart for the comparison
- Use a bar chart to compare categories, a line chart for an ordered time series, and a scatter plot to examine two numeric variables together.
- Give the chart a useful title, labelled axes, units and the relevant sample size. A viewer must know what each value represents.
Match each comparison to a suitable starting chart.
Choose a chart that expresses the data relationship.
An unlabeled axis still lets a reader know the unit and measured variable reliably.
Provide the variable and units; a reader should not have to guess.
Separate a pattern from its cause
- Correlation is not causation 相关不等于因果: an observed relationship alone does not show that one variable produced the other. Other factors and the study design matter.
- If a survey finds longer commuters sleep less, describe the association within the respondents. Do not conclude that changing commute time alone will change sleep.
Survey respondents with longer commutes report less sleep. Which conclusion is supported?
The observation does not isolate a causal effect or justify a wider population claim.
Report a finding with its limits
- Report: “The four recorded commute times have mean 25 minutes; one of five respondents did not give a time.” Missingness can affect how representative the result is.
- State the sample and collection limits before recommending action. A one-class survey does not establish a result for every student in the city.
Question → checked data → transparent calculation → labelled chart → limited conclusion. Missing values and sample boundaries belong in the explanation.
Write one limit of a survey of 60 students from one class.
One class can differ from other classes, so the result should not automatically be generalized to the city.
Which reports the commute result transparently?
State the valid denominator and missing count.