Using Programs with Data · 用程序处理数据
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| processing/ˈprəʊsesɪŋ/ | 处理 | chǔ lǐ |
| iteration/ˌɪtəˈreɪʃn/ | 迭代 | dié dài |
| element/ˈelɪmənt/ | 元素 | yuán sù |
| filter/ˈfɪltə/ | 过滤器 | guò lǜ qì |
| subset/ˈsʌbset/ | 子集 | zi jí |
| combine/kəmˈbaɪn/ | 合并 | hé bìng |
| transform/trænsˈfɔːm/ | 转换 | zhuǎn huàn |
| Cleaning/ˈkliːnɪŋ/ | 清洗 | qīng xǐ |
Programs that answer questions
- The last step is to use a program to actually work with data.
- A program does this by processing 处理 — taking data in, doing steps, giving back new information.
- The input might be a list of records, like students and their test scores.
- The output might be a new fact the data never directly stated — like the class average.
回答问题的程序
- 最后一步是用一个程序真正地处理数据。
- 程序通过处理(processing)做到这点——接收数据、执行步骤、返回新信息。
- 输入可能是一个记录列表,像学生和他们的考试分数。
- 输出可能是数据从未直接陈述的新事实——像班级平均分。
A program that takes scores in and gives back the average is doing: · 一个接收分数并返回平均的程序在做:
Input → processing → output produces a new fact, the average. · 输入 → 处理 → 输出产生一个新事实,平均。
Iteration visits every element
- To look at a whole data set, a program uses iteration 迭代 — repeating the same steps for each item.
- A loop visits each element 元素 (each item) of the list, one after another.
- So the program examines every record without a separate line for each one.
Average of five scores. A loop runs through $70, 85, 90, 60, 95$, adding each to a running total. After the loop the total is $400$; dividing by 5 gives the average $400 \div 5 = 80$ — information the raw list never stated.
迭代访问每个元素
- 要查看整个数据集,程序使用迭代(iteration)——为每一项重复同样的步骤。
- 一个循环一个接一个地访问列表的每个元素(element,每一项)。
- 于是程序检查每条记录,而不必为每一条写一行。
五个分数的平均。 一个循环遍历 $70, 85, 90, 60, 95$,把每个加到运行总和上。循环后总和是 $400$;除以 5 得平均 $400 \div 5 = 80$——原始列表从未陈述的信息。
Trace a loop that sums scores · 追踪一个求和分数的循环
A loop iterates through the list, adding each score to a running total. After the last pass the total is 400; dividing by 5 gives the average, 80. · 一个循环遍历列表,把每个分数加到运行总和。最后一遍后总和是 400;除以 5 得平均 80。
Repeating the same steps for each item of a list is called . · 为列表的每一项重复同样的步骤叫。
A loop visits each element in turn. · 一个循环依次访问每个元素。
A loop sums 70, 85, 90, 60, 95 to 400. What is the average of the 5 scores? · 一个循环把 70、85、90、60、95 加到 400。这 5 个分数的平均是多少?
400 ÷ 5 = 80. · 400 ÷ 5 = 80。
Filtering to a subset
- Often we want only some records. A filter 过滤器 keeps records that meet a condition and drops the rest.
- A condition is a test that is true or false for each record, like "score is 80 or higher."
- The filter produces a smaller subset 子集 — only the records that pass.
- Worked idea. Filtering $70, 85, 90, 60, 95$ with "score $\geq 80$" keeps $85, 90, 95$ — 3 students.
筛选出一个子集
- 我们常常只想要一些记录。一个过滤器(filter)保留满足条件的记录,丢弃其余的。
- 条件是对每条记录为真或假的测试,像"分数是 80 或更高"。
- 过滤器产生一个更小的子集(subset)——只有通过的记录。
- 例题。 用"分数 $\geq 80$"过滤 $70, 85, 90, 60, 95$,保留 $85, 90, 95$——3 名学生。
From 70, 85, 90, 60, 95, how many scores pass the filter "score ≥ 80"? · 从 70、85、90、60、95 中,有多少个分数通过筛选"分数 ≥ 80"?
85, 90, 95 pass — a subset of 3. · 85、90、95 通过——一个 3 个的子集。
The smaller group of records that pass a filter's condition is a . · 通过过滤器条件的那一小组记录是一个。
A filter produces a subset of the original records. · 过滤器产生原始记录的一个子集。
Why clean the data before calculating? · 为什么在计算前清洗数据?
Dirty data gives a wrong answer, so clean first for reliable results. · 脏数据给出错误答案,所以先清洗以获得可靠结果。
Combine, transform, and clean
- A program can combine 合并 data from several sources — matching names to scores by a shared ID.
- Or transform 转换 it — converting Celsius to Fahrenheit, or raw scores to percentages.
- Cleaning 清洗 inside the program (remove duplicates, skip blanks) keeps results reliable.
- Dirty data misleads: a duplicate skews the average, and a blank could crash the calculation.
合并、转换与清洗
- 程序能合并(combine)多个来源的数据——通过共享的 ID 把名字与分数匹配。
- 或转换(transform)它——把摄氏转华氏,或把原始分数转百分比。
- 程序内的清洗(cleaning,去重复、跳过空白)让结果可靠。
- 脏数据会误导:一个重复项使平均偏斜,一个空白可能让计算崩溃。
A program processes data: input records in, new information out. Iteration visits every element to compute totals and averages; a filter keeps a subset meeting a condition. Programs also combine sources, transform values, and clean the data first — because dirty data gives a wrong answer.
程序处理数据:记录输入,新信息输出。迭代访问每个元素来计算总和与平均;过滤器保留满足条件的子集。程序还合并来源、转换值,并先清洗数据——因为脏数据给出错误答案。