Extracting Information from Data · 从数据中提取信息
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| information/ˌɪnfəˈmeɪʃn/ | 信息 | xìn xī |
| data sets/ˈdeɪtə sets/ | 数据集 | shù jù jí |
| patterns/ˈpætnz/ | 模式 | mó shì |
| trends/trendz/ | 趋势 | qū shì |
| Filtering/ˈfɪltərɪŋ/ | 筛选 | shāi xuǎn |
| Cleaning/ˈkliːnɪŋ/ | 清洗 | qīng xǐ |
| correlation/ˌkɒrɪˈleɪʃn/ | 相关 | xiāng guān |
| Causation/kɔːˈseɪʃn/ | 因果 | yīn guǒ |
| Metadata/ˌmetəˈdeɪtə/ | 元数据 | yuán shù jù |
| visualization/ˌvɪʒuːəlaɪˈzeɪʃn/ | 可视化 | kě shì huà |
| bias/ˈbaɪəs/ | 偏差 | piān chā |
From raw data to information
- Raw data by itself is just numbers and text.
- The goal is to turn it into useful information 信息 — something that answers a real question.
- Programs can scan huge data sets 数据集 far faster than a person.
- By scanning, they find patterns 模式 (things that repeat) and trends 趋势 (changes over time).
从原始数据到信息
- 原始数据本身只是数字和文本。
- 目标是把它变成有用的信息(information)——回答一个真实问题的东西。
- 程序能比人快得多地扫描巨大的数据集(data sets)。
- 通过扫描,它们找到模式(patterns,重复的东西)和趋势(trends,随时间的变化)。
Turning raw data into something that answers a real question produces: · 把原始数据变成回答真实问题的东西,产生:
Information is data made useful — patterns and trends answering a question. · 信息是变得有用的数据——回答问题的模式和趋势。
Match each term to its meaning. · 把每个词与其含义配对。
Programs scan data sets to find both. · 程序扫描数据集来找到两者。
Preparing messy data
- Real data is messy, so it must be prepared first.
- Filtering 筛选 keeps only the rows you care about and hides the rest.
- Cleaning 清洗 fixes problems — removing duplicates, filling missing values, correcting mistakes.
- Preparing data this way makes the later analysis meaningful and trustworthy.
准备杂乱的数据
- 真实数据很杂乱,所以必须先准备。
- 筛选(filtering)只保留你关心的行,隐藏其余的。
- 清洗(cleaning)修复问题——去除重复、填补缺失值、更正错误。
- 这样准备数据让后面的分析有意义、可信。
Filtering or cleaning? · 筛选还是清洗?
Filtering keeps only the rows you care about; cleaning fixes problems in the data (duplicates, blanks, mistakes) before analysis. · 筛选只保留你关心的行;清洗在分析前修复数据中的问题(重复、空白、错误)。
Removing duplicate rows and filling missing values is part of cleaning data. · 去除重复行和填补缺失值是清洗数据的一部分。
Cleaning fixes problems so the analysis can be trusted. · 清洗修复问题,让分析可信。
Correlation is not causation
- A correlation 相关 means two things change together (ice-cream sales and sunburns both rise in summer).
- Causation 因果 means one thing actually causes the other.
- A correlation does not prove causation.
Don't confuse the two. Ice cream does not cause sunburns — hot, sunny weather causes both. Two things rising together can share a hidden cause, so a correlation alone never proves that one causes the other.
相关不是因果
- 相关(correlation)意味着两件事一起变化(冰淇淋销量和晒伤在夏天都上升)。
- 因果(causation)意味着一件事真正导致另一件。
- 相关不证明因果。
别混淆两者。 冰淇淋不导致晒伤——炎热晴朗的天气导致两者。两件事一起上升可能共享一个隐藏的原因,所以仅有相关永远不能证明一件事导致另一件。
Ice-cream sales and sunburns both rise in summer. This shows: · 冰淇淋销量和晒伤在夏天都上升。这说明:
Hot weather causes both — a shared cause, not one causing the other. · 炎热天气导致两者——一个共同的原因,而非一个导致另一个。
"Data about data", like a photo's date and location, is called . · "关于数据的数据",如照片的日期和位置,叫。
Metadata helps organize, search, and locate information. · 元数据帮助组织、搜索和定位信息。
Surveying only people who downloaded an app about whether the app is popular introduces: · 只调查下载了应用的人该应用是否受欢迎,会引入:
The collection method slants the result — a biased sample. · 收集方法使结果倾斜——一个有偏差的样本。
Metadata, visualization, and bias
- Metadata 元数据 is "data about data" — a photo's date, camera, and location. It helps organize and search.
- A visualization 可视化 turns data into a picture (a chart or graph) so a pattern is obvious at a glance.
- Always ask how the data was collected: bias 偏差 is a slant that pushes conclusions one way.
Biased survey. A company asks "Is our app popular?" but surveys only people who already downloaded it. Almost all say yes — but unhappy users never took the survey, so the happy result is misleading.
元数据、可视化与偏差
- 元数据(metadata)是"关于数据的数据"——一张照片的日期、相机和位置。它帮助组织和搜索。
- 一个可视化(visualization)把数据变成一幅图(图表),让模式一眼就明显。
- 总要问数据是怎么收集的:偏差(bias)是把结论推向一个方向的倾斜。
有偏差的调查。 一家公司问"我们的应用受欢迎吗?",但只调查已经下载它的人。几乎所有人都说是——但不满意的用户从不参加调查,所以这个令人高兴的结果具有误导性。
Programs turn raw data sets into information, finding patterns and trends. Prepare data by filtering (keep relevant rows) and cleaning (fix problems). Beware: a correlation never proves causation, and bias in how data was collected can mislead. Metadata and visualization help organize and reveal insight.
程序把原始数据集变成信息,找到模式和趋势。通过筛选(保留相关行)和清洗(修复问题)准备数据。当心:相关永远不证明因果,数据收集方式中的偏差能误导。元数据和可视化帮助组织和揭示洞见。