Skip to content · ⁨דלג לתוכן⁩

Exploring Two-Variable Data · ⁨חקר נתוני שני משתנים⁩

AP Statistics · ⁨סטטיסטיקה - AP⁩ · Topic 2 · ⁨נושא 2⁩

Video lesson for this topic · ⁨שיעור וידאו לנושא זה⁩ Open the video page · ⁨פתח את עמוד הוידאו⁩
7:49

חקר נתוני שני משתנים

שלושים ילדים מאחת מבתי הספר היסודיים. עבור כל ילד, שני מספרים: מידת נעליים, ותוצאת מבחן קריאה. צייר נקודה אחת לכל ילד. הדפוס קשה להבין…

English narration · English + 中文 subtitles burned in · ⁨קריאת קול באנגלית · תרגום אנגלי + סינית שרוף בתוך הסרטון⁩

2.1

Are Two Variables Related? · ⁨האם שני משתנים קשורים?⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (VAR-1): Given that variation may be random or not, conclusions are uncertain.

Learning Objective VAR-1.D: Identify questions to be answered about possible relationships in data. [Skill 1.A]

  • VAR-1.D.1 Apparent patterns and associations in data may be random or not.
עברית

הבנה מתמשכת (VAR-1): בשל כך ששינוי עשוי להיות אקראי או לא, המסקנות הן לא וודאות.

מטרות למידה VAR-1.D: לזהות שאלות יש לענות על קשרים אפשריים בנתונים. [מיומנות 1.A]

  • VAR-1.D.1 דפוסים וקשרים נראים בנתונים עשויים להיות מקריים או לא.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

Two-variable data let us ask whether two characteristics are associated 关联 – whether knowing one tells you something about the other. An explanatory variable 解释变量 (the "input") may help predict a response variable 响应变量 (the "output"). Association is not the same as causation.

עברית

נתונים דו-משתניים מאפשרים לשאול האם שתי מאפיינים קשורים – האם ידע אחד מספק מידע על השני. משתנה הסבר (ה"קלט") עשוי לסייע בחיזוי משתנה תגובה (ה"יציאה"). קשר אינו זהה לחיבור סיבתי.

2.2

Two Categorical Variables · ⁨שני משתנים קטגוריאליים⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (UNC-1): Graphical representations and statistics allow us to identify and represent key features of data.

Learning Objective UNC-1.P: Compare numerical and graphical representations for two categorical variables. [Skill 2.D]

  • UNC-1.P.1 Side-by-side bar graphs, segmented bar graphs, and mosaic plots are examples of bar graphs for one categorical variable, broken down by categories of another categorical variable.
  • UNC-1.P.2 Graphical representations of two categorical variables can be used to compare distributions and/or determine if variables are associated.
  • UNC-1.P.3 A two-way table, also called a contingency table, is used to summarize two categorical variables. The entries in the cells can be frequency counts or relative frequencies.
  • UNC-1.P.4 A joint relative frequency is a cell frequency divided by the total for the entire table.
עברית

הבנה מתמשכת (UNC-1): ייצוגים גרפיים וסטטיסטיקה מאפשרים לזהות ולייצג תכונות מרכזיות של נתונים.

מטרות למידה UNC-1.P: להשוות ייצוגים מספריים וגרפיים לשני משתנים קטגוריאליים. [מיומנות 2.D]

  • UNC-1.P.1 גרפי עמודות בצדדים זה לזה, גרפי עמודות מחולקים (segmented bar graphs) ותמונות מוזאיק הם דוגמאות לגרפי עמודות למשתנה קטגוריאלי אחד, המפורקים לפי קטגוריות של משתנה קטגוריאלי אחר.
  • UNC-1.P.2 ייצוגים גרפיים של שני משתנים קטגוריאליים יכולים לשמש להשוואת התפלגויות ו/או לקבוע האם משתנים קשורים.
  • UNC-1.P.3 טבלה דו-כיוונית, הנקראת גם טבלת תלות, משמשת לסיכום שני משתנים קטגוריאליים. הכניסות בתאים יכולות להיות ספירות תדירות או תדירויות יחסיות.
  • UNC-1.P.4 תדירות יחסית משותפת היא תדירות תא המחולקת בסך הכל של הטבלה כולה.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

A two-way table 双向表 (contingency table) counts individuals by two categorical variables at once. A marginal distribution 边缘分布 is a row or column total written as a fraction of the grand total (the totals themselves are just counts). Comparing the inside cells shows whether the variables are related.

עברית

טבלה דו-כיוונית (טבלת תלות) סופרת פרטים לפי שני משתנים קטגוריאליים בו-זמנית. התפלגות שוליים היא סכום שורה או עמודה המוצג כשבר מהסך הכללי (הסכומים עצמם הם פשוט ספירות). השוואת התאים הפנימיים מראה האם המשתנים קשורים.

2.3

Comparing Groups with Conditional Distributions · ⁨השוואת קבוצות באמצעות התפלגויות תנאי⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (UNC-1): Graphical representations and statistics allow us to identify and represent key features of data.

Learning Objective UNC-1.Q: Calculate statistics for two categorical variables. [Skill 2.C]

  • UNC-1.Q.1 The marginal relative frequencies are the row and column totals in a two-way table divided by the total for the entire table.
  • UNC-1.Q.2 A conditional relative frequency is a relative frequency for a specific part of the contingency table (e.g., cell frequencies in a row divided by the total for that row).

Learning Objective UNC-1.R: Compare statistics for two categorical variables. [Skill 2.D]

  • UNC-1.R.1 Summary statistics for two categorical variables can be used to compare distributions and/or determine if variables are associated.
עברית

הבנה מתמשכת (UNC-1): ייצוגים גרפיים וסטטיסטיקה מאפשרים לזהות ולייצג תכונות מרכזיות של נתונים.

מטרות למידה UNC-1.Q: לחשב סטטיסטיקות לשני משתנים קטגוריאליים. [מיומנות 2.C]

  • UNC-1.Q.1 התדירויות היחסיות השוליות הן סכומי שורות ועמודות בטבלה דו-צירית מחולקים בסך הכל של הטבלה כולה.
  • UNC-1.Q.2 תדירות יחסית מותנית היא תדירות יחסית עבור חלק ספציפי בטבלת העלילה (למשל, תדירויות תאים בשורה מחולפות בסך הכל של אותה שורה).

מטרת למידה UNC-1.R: השוואת סטטיסטיקה לשני משתנים קטגוריאליים. [Skill 2.D]

  • UNC-1.R.1 סטטיסטיקת סיכום לשני משתנים קטגוריאליים יכולה לשמש להשוואת ההתפלגויות ו/או לקבוע האם המשתנים קשורים זה לזה.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

A conditional distribution 条件分布 is the distribution of one variable within a fixed category of the other (found by dividing each cell by its row or column total). If the conditional distributions differ across groups, the two variables are associated; if they are the same, there is no association. Segmented bar charts 分段条形图 or mosaic plots display them.

עברית

התפלגות תנאית היא ההתפלגות של משתנה אחד בתוך קטגוריה קבועה של המשתנה השני (נמצאת על ידי חלוקת כל תא בסכום השורה או העמודה שלו). אם ההתפלגויות התנאיות שונות בין קבוצות, שני המשתנים קשורים; אם הן זהות, אין קשר. גרפי עמודות מקוטעים או גרפי פיסות מציגים אותן.

2.4

Scatterplots for Two Quantitative Variables · ⁨גרפי פיזור לשני משתנים כמותיים⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (UNC-1): Graphical representations and statistics allow us to identify and represent key features of data.

Learning Objective UNC-1.S: Represent bivariate quantitative data using scatterplots. [Skill 2.B]

  • UNC-1.S.1 A bivariate quantitative data set consists of observations of two different quantitative variables made on individuals in a sample or population.
  • UNC-1.S.2 A scatterplot shows two numeric values for each observation, one corresponding to the value on the $x$-axis and one corresponding to the value on the $y$-axis.
  • UNC-1.S.3 An explanatory variable is a variable whose values are used to explain or predict corresponding values for the response variable.

Enduring Understanding (DAT-1): Regression models may allow us to predict responses to changes in an explanatory variable.

Learning Objective DAT-1.A: Describe the characteristics of a scatter plot. [Skill 2.A]

  • DAT-1.A.1 A description of a scatter plot includes form, direction, strength, and unusual features.
  • DAT-1.A.2 The direction of the association shown in a scatterplot, if any, can be described as positive or negative.
  • DAT-1.A.3 A positive association means that as values of one variable increase, the values of the other variable tend to increase. A negative association means that as values of one variable increase, values of the other variable tend to decrease.
  • DAT-1.A.4 The form of the association shown in a scatterplot, if any, can be described as linear or non-linear to varying degrees.
  • DAT-1.A.5 The strength of the association is how closely the individual points follow a specific pattern, e.g., linear, and can be shown in a scatterplot. Strength can be described as strong, moderate, or weak.
  • DAT-1.A.6 Unusual features of a scatter plot include clusters of points or points with relatively large discrepancies between the value of the response variable and a predicted value for the response variable.
עברית

הבנה מתמשכת (UNC-1): ייצוגים גרפיים וסטטיסטיקה מאפשרים לזהות ולייצג תכונות מרכזיות של נתונים.

מטרת למידה UNC-1.S: נייגון נתונים כמותיים בי-משתניים באמצעות גרפי פיזור. [Skill 2.B]

  • UNC-1.S.1 ערכת נתונים כמותיים בי-משתניים מורכבת מצפייה על שני משתנים כותיים שונים שנערכו על פרטים בדגימה או באוכלוסייה.
  • UNC-1.S.2 גרף פיזור מראה שני ערכים מספריים לכל צפייה, אחד המתאים לערך על ציר ה $x$ ואחד המתאים לערך על ציר ה $y$.
  • UNC-1.S.3 משתנה הסבר הוא משתנה whose ערכיו משמשים להסביר או לחזות ערכים תואמים במשתנה התגובה.

הבנה מתמשכת (DAT-1): מודלים רגרסיה עשויים לאפשר לנו לחזות תגובות לשינויים במשתנה הסבר.

מטרת למידה DAT-1.A: לתאר את מאפייני גרף פיזור. [Skill 2.A]

  • DAT-1.A.1 תיאור גרף פיזור כולל צורה, כיוון, חוזק ותכונות בלתי רגילות.
  • DAT-1.A.2 הכיוון של ההקשר המוצג בגרף פיזור, אם קיים, ניתן לתאר כחיובי או שלילי.
  • DAT-1.A.3 הקשר חיובי פירושו שככל שערכי משתנה אחד עולים, ערכי המשתנה השני נוטים לעלות. הקשר שלילי פירושו שככל שערכי משתנה אחד עולים, ערכי המשתנה השני נוטים לרדת.
  • DAT-1.A.4 הצורה של ההקשר המוצג בגרף פיזור, אם קיים, ניתן לתאר כליניארי או לא-ליניארי ברמת התייחסות שונות.
  • DAT-1.A.5 חוזק ההקשר הוא עד כמה הנקודות הבודדות עוקבות אחר דפוס ספציפי, לדוגמה ליניארי, והוא יכול להיות מוצג בגרף פיזור. חוזק יכול להתאפיין בחזק, בינוני או חלש.
  • DAT-1.A.6 תכונות בלתי רגילות בגרף פיזור כוללות אשכולות של נקודות או נקודות עם אי-התאמות גדולות יחסית בין ערך משתנה התגובה לערך החזוי עבור משתנה התגובה.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

A scatterplot 散点图 plots each individual as a point, explanatory variable on the $x$-axis and response on the $y$-axis. Describe it with DUFS: Direction (positive/negative), Unusual features (outliers, clusters), Form (linear or curved), and Strength (how tightly the points follow the pattern) – always in context.

עברית

תרפיס פיזור מצייר כל פרט כנקודה, משתנה הסבר על ציר ה-$x$ ומשתנה התגובה על ציר ה-$y$. מתארים אותו באמצעות DUFS: כיוון (חיובי/שלילי), מאפיינים בלתי רגילים (חריגים, אשכולות), צורה (ליניארית או מעוקלת), וחוזק (כמה הנקודות עוקבות אחרי הדפוס) – תמיד בהקשר.

קו התאמה מיטבית עובר באמצע הנקודות המפוזרות
קו התאמה מיטבית עובר באמצע הנקודות המפוזרות
2.5

Correlation · ⁨קורלציה⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (DAT-1): Regression models may allow us to predict responses to changes in an explanatory variable.

Learning Objective DAT-1.B: Determine the correlation for a linear relationship. [Skill 2.C]

  • DAT-1.B.1 The correlation, $r$, gives the direction and quantifies the strength of the linear association between two quantitative variables.
  • DAT-1.B.2 The correlation coefficient can be calculated by: $r = \dfrac{1}{n-1} \sum \left( \dfrac{x_i - \bar{x}}{s_x} \right) \left( \dfrac{y_i - \bar{y}}{s_y} \right)$. However, the most common way to determine $r$ is by using technology.
  • DAT-1.B.3 A correlation coefficient close to 1 or $-1$ does not necessarily mean that a linear model is appropriate.

Learning Objective DAT-1.C: Interpret the correlation for a linear relationship. [Skill 4.B]

  • DAT-1.C.1 The correlation, $r$, is unit-free, and always between $-1$ and 1, inclusive. A value of $r = 0$ indicates that there is no linear association. A value of $r = 1$ or $r = -1$ indicates that there is a perfect linear association.
  • DAT-1.C.2 A perceived or real relationship between two variables does not mean that changes in one variable cause changes in the other. That is, correlation does not necessarily imply causation.
עברית

הבנה מתמשכת (DAT-1): מודלים רגרסיה עשויים לאפשר לנו לחזות תגובות לשינויים במשתנה הסבר.

מטרת למידה DAT-1.B: לקבוע את הקורלציה לקשר ליניארי. [Skill 2.C]

  • DAT-1.B.1 המקדם, $r$, מספק את הכיוון ומכמת את חוזק ההקשר הליניארי בין שני משתנים כותיים.
  • DAT-1.B.2 מקדם המתאם ניתן לחישוב על ידי: $r = \dfrac{1}{n-1} \sum \left( \dfrac{x_i - \bar{x}}{s_x} \right) \left( \dfrac{y_i - \bar{y}}{s_y} \right)$. עם זאת, הדרך הנפוצה ביותר לקבוע $r$ היא באמצעות שימוש בטכנולוגיה.
  • DAT-1.B.3 מקדם מתאם הקרוב ל-1 או ל$-1$ אינו בהכרח מצביע על כך שדגם ליניארי הוא מתאים.

מטרת למידה DAT-1.C: פרשן את המקדם במתאם ליניארי. [מיומנות 4.B]

  • DAT-1.C.1 המקדם, $r$, חסר יחידות, תמיד נמצא בין $-1$ ל-1, כולל קצוות. ערך של $r = 0$ מעיד על כך שאין מתאם ליניארי. ערך של $r = 1$ או $r = -1$ מעיד על כך שקיים מתאם ליניארי מושלם.
  • DAT-1.C.2 קשר נתפס או אמיתי בין משתנים שניים אינו אומר שהשינויים במשתנה אחד גורמים לשינויים במשתנה השני. כלומר, מתאם אינו implies בהכרח סיבתיות.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English
What r actually measures

The correlation coefficient 相关系数 $r$ measures the strength and direction of a linear relationship. It runs from $-1$ to $1$: near $\pm 1$ is strong linear, near $0$ is weak linear. $r$ has no units and does not change if you swap the variables. Warnings: $r$ only measures linear strength, it is not resistant to outliers, and a strong $r$ does not prove causation.

עברית
מה r מודד בפועל

המקדם הקורלציה $r$ מדד את חוזק וכיוון הקשר הליניארי. הוא נע בין $-1$ ל$1$: קרוב ל$\pm 1$ הוא ליניארי חזק, וקרוב ל$0$ הוא ליניארי חלש. $r$ אין לו יחידות מידה ואינו משתנה אם מחליפים את המשתנים. אזהרות: $r$ מדד רק חוזק ליניארי, הוא אינו עמיד בפני ערכים קיצוניים, וחוזק גבוה ב$r$ לא מוכיח סיבתיות.

קורלציה חיובית עולה יחד; קורלציה שלילית זזה בכיוונים הפוכים
קורלציה חיובית עולה יחד; קורלציה שלילית זזה בכיוונים הפוכים
Explore · ⁨חקור⁩

Strength of a linear relationship · ⁨חוזק הקשר הליניארי⁩

Correlation $r$ runs from $-1$ to $1$: near $\pm1$ the points hug a line, near 0 they scatter. Change it and watch the cloud tighten or spread. · ⁨קורלציה $r$ נעה בין $-1$ ל$1$: קרוב ל$\pm1$ הנקודות צמודות לקו, קרוב ל-0 הן מפוזרות. שנה אותו והרה את הענן מתכווץ או מתפשט.⁩

2.6

Linear Regression Models · ⁨דגמי רגרסיה ליניארית⁩

Syllabus · ⁨סיילבוס⁩
Enduring UnderstandingLearning ObjectiveEssential Knowledge

DAT-1
Regression models may allow us to predict responses to changes in an explanatory variable.

DAT-1.D
Calculate a predicted response value using a linear regression model. [Skill 2.C]

  • DAT-1.D.1 A simple linear regression model is an equation that uses an explanatory variable, $x$, to predict the response variable, $y$.
  • DAT-1.D.2 The predicted response value, denoted by $\hat{y}$, is calculated as $\hat{y} = a + bx$, where $a$ is the $y$-intercept and $b$ is the slope of the regression line, and $x$ is the value of the explanatory variable.
  • DAT-1.D.3 Extrapolation is predicting a response value using a value for the explanatory variable that is beyond the interval of $x$-values used to determine the regression line. The predicted value is less reliable as an estimate the further we extrapolate.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

The least-squares regression line 最小二乘回归线 predicts the response: $\hat{y}=a+bx$, where $\hat{y}$ is the predicted response. The slope 斜率 $b$ is the predicted change in $y$ per one-unit increase in $x$; the $y$-intercept 截距 $a$ is the predicted $y$ when $x=0$. Interpret both in context and with units – a graded skill. Avoid extrapolation 外推 (predicting far outside the data).

Worked example. A study of hours studied ($x$) and test score ($y$) gives $\hat{y}=20+3x$. The slope means each extra hour of study is associated with a predicted $3$-point increase. A student who studies $5$ hours is predicted to score $\hat{y}=20+3(5)=35$.

עברית

קו הרגרסיה לפחות-הריבועים מנבא את התגובה: $\hat{y}=a+bx$, כאשר $\hat{y}$ היא התגובה המוכחת. שיפוע $b$ הוא השינוי המוכחת ב$y$ לכל עלייה של יחידה אחת ב$x$; ה$y$-חתך $a$ הוא ה$y$ המוכחת כשה$x=0$. יש לפרש שניהם בהקשר ובכלליות עם יחידות מידה – מיומנות מוערכת. יש להימנע מאקסטרפולציה (הכרת תחזית הרחוקה מאוד מהנתונים).

דוגמה מופרדת. מחקר על שעות לימוד ($x$) ותוצאת מבחן ($y$) נתנו $\hat{y}=20+3x$. השיפוע אומר שכל שעת לימוד נוספת קשורה לעלייה מוכחת של $3$ נקודות. סטודנט שלימד $5$ שעות מוכתב שתציין $\hat{y}=20+3(5)=35$.

Explore · ⁨חקור⁩

Fit a least-squares line · ⁨התאמת קו פחות-הריבועים⁩

A regression line is the best straight-line fit, minimising the squared vertical distances. Its slope predicts how $y$ changes per unit of $x$. · ⁨קו רגרסיה הוא ההתאמה הישרה הטובה ביותר, המזערת את הריבועים של המרחקים האנכיים. שיפועו מתنبא כיצד $y$ משתנה ליחידה אחת של $x$.⁩

2.7

Residuals · ⁨שאריות⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (DAT-1): Regression models may allow us to predict responses to changes in an explanatory variable.

Learning Objective DAT-1.E: Represent differences between measured and predicted responses using residual plots. [Skill 2.B]

  • DAT-1.E.1 The residual is the difference between the actual value and the predicted value: $\text{residual} = y - \hat{y}$.
  • DAT-1.E.2 A residual plot is a plot of residuals versus explanatory variable values or predicted response values.

Learning Objective DAT-1.F: Describe the form of association of bivariate data using residual plots. [Skill 2.A]

  • DAT-1.F.1 Apparent randomness in a residual plot for a linear model is evidence of a linear form to the association between the variables.
  • DAT-1.F.2 Residual plots can be used to investigate the appropriateness of a selected model.
עברית

הבנה מתמשכת (DAT-1): מודלים רגרסיה עשויים לאפשר לנו לחזות תגובות לשינויים במשתנה הסבר.

מטרת למידה DAT-1.E: נציג הבדלים בין תגובות נמדדות לתגובות צפויות באמצעות גרפי שאריות. [מיומנות 2.B]

  • DAT-1.E.1 השארית היא ההפרש בין הערך בפועל לערך הצפוי: $\text{residual} = y - \hat{y}$.
  • DAT-1.E.2 גרף שאריות הוא גרף של שאריות נגד ערכי משתנה ההסבר או ערכי התגובה הצפויים.

מטרת למידה DAT-1.F: תאר את צורת המתאם של נתונים דו-משתנים באמצעות גרפי שאריות. [מיומנות 2.A]

  • DAT-1.F.1 אקראיות apparent בגרף שאריות לדגם ליניארי היא עדות לצורה ליניארית במתאם בין המשתנים.
  • DAT-1.F.2 גרפי שאריות יכולים לשמש לבדיקת ההתאמה של דגם שנבחר.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English
Least-squares regression

A residual 残差 is actual minus predicted, $y-\hat{y}$: how far a point sits above (+) or below (−) the line. A residual plot 残差图 graphs residuals against $x$. If it shows no pattern (random scatter), a linear model is appropriate; a curved or fanning pattern means the linear model is a poor fit.

Worked example. Continuing the study above, a student who studied $5$ hours actually scored $40$. The residual is $y-\hat{y}=40-35=+5$: the line under-predicted by $5$ points, so this point sits above the line.

עברית
רגרסיה פחותים מרבעים

שארית היא האמת פחות הכרעה, $y-\hat{y}$: כמה נקודה נמצאת מעל (+) או מתחת (−) לקו. גרף שאריות מצייר שאריות מול $x$. אם הוא מראה אין דפוס (פיזור אקראי), דגם ליניארי מתאים; דפוס מעוגל או מתפשט מסמן שהדגם הליניארי הוא התאמה גרועה.

דוגמה מופרדת. בהמשך למחקר לעיל, סטודנט שלימד $5$ שעות ציין בפועל $40$. השארית היא $y-\hat{y}=40-35=+5$: הקו הכרה נמוכה ב$5$ נקודות, ולכן הנקודה הזו נמצאת מעל הקו.

ארבעה מאגרים נתונים עם r זהה וקו רגרסיה זהה אך ארבעה צורות שונות
אזהרה לגבי ⟨$r$⟩ והקו: כל ארבעת מאגרי הנתונים הם אותו ⟨$r=0.82$⟩ ואותו ⟨$\hat{y}=3.0+0.5x$⟩, אך רק הראשון הוא ליניארי אמיתי. הגרפים מפוזרים קשה להבחין ביניהם — גרף השאריות בתחתית כל אחד הוא מה חושף את העיקום, הערך הקיצוני ונקודת ההשפעה הגבוהה.
2.8

Least-Squares Regression and Its Fit · ⁨רגרסיה לפחות-הריבועים והתאמתה⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (DAT-1): Regression models may allow us to predict responses to changes in an explanatory variable.

Learning Objective DAT-1.G: Estimate parameters for the least-squares regression line model. [Skill 2.C]

  • DAT-1.G.1 The least-squares regression model minimizes the sum of the squares of the residuals and contains the point $(\bar{x}, \bar{y})$.
  • DAT-1.G.2 The slope, $b$, of the regression line can be calculated as $b = r \left( \dfrac{s_y}{s_x} \right)$ where $r$ is the correlation between $x$ and $y$, $s_y$ is the sample standard deviation of the response variable, $y$, and $s_x$ is the sample standard deviation of the explanatory variable, $x$.
  • DAT-1.G.3 Sometimes, the $y$-intercept of the line does not have a logical interpretation in context.
  • DAT-1.G.4 In simple linear regression, $r^2$ is the square of the correlation, $r$. It is also called the coefficient of determination. $r^2$ is the proportion of variation in the response variable that is explained by the explanatory variable in the model.

Learning Objective DAT-1.H: Interpret coefficients for the least-squares regression line model. [Skill 4.B]

  • DAT-1.H.1 The coefficients of the least-squares regression model are the estimated slope and $y$-intercept.
  • DAT-1.H.2 The slope is the amount that the predicted $y$-value changes for every unit increase in $x$.
  • DAT-1.H.3 The $y$-intercept value is the predicted value of the response variable when the explanatory variable is equal to $0$. The formula for the $y$-intercept, $a$, is $a = \bar{y} - b\bar{x}$.
עברית

הבנה מתמשכת (DAT-1): מודלים רגרסיה עשויים לאפשר לנו לחזות תגובות לשינויים במשתנה הסבר.

מטרת למידה DAT-1.G: העריך פרמטרים עבור דגם קו הרגרסיה של סכומי רבועים מינימליים. [מיומנות 2.C]

  • DAT-1.G.1 דגם הרגרסיה של סכומי רבועים מינימליים ממזער את סכום רבועות השאריות ומכיל את הנקודה $(\bar{x}, \bar{y})$.
  • DAT-1.G.2 השיפוע, $b$, של קו הרגרסיה ניתן לחישוב כ $b = r \left( \dfrac{s_y}{s_x} \right)$ כאשר $r$ הוא הקורלציה בין $x$ ל $y$, $s_y$ היא סטיית התקן המדגמית של משתנה התגובה, $y$, ו $s_x$ היא סטיית התקן המדגמית של משתנה ההסבר, $x$.
  • DAT-1.G.3 לעיתים, החיתוך עם ציר ה $y$ של הקו אינו בעל פרשנות לוגית בהקשר נתון.
  • DAT-1.G.4 ברגרסיה ליניארית פשוטה, $r^2$ הוא ריבוע הקורלציה, $r$. הוא מכונה גם מקדם הקביעה. $r^2$ הוא היחס שבהתאם אליו השונות במשתנה התגובה מוסברת על ידי משתנה ההסבר במודל.

מטרת למידה DAT-1.H: פרש מקדמים במודל קו הרגרסיה לפחות ריבועים. [מיומנות 4.B]

  • DAT-1.H.1 המקדמים במודל הרגרסיה לפחות ריבועים הם השיפוע המוערך וחיתוך ציר ה $y$.
  • DAT-1.H.2 השיפוע הוא הכמות שבה ישתנה הערך המוצף של $y$ עבור כל עלייה של יחידה אחת ב $x$.
  • DAT-1.H.3 ערך חיתוך ציר ה $y$ הוא הערך המוצף של משתנה התגובה כאשר משתנה ההסבר שווה ל $0$. הנוסחה לחיתוך ציר ה $y$, $a$, היא $a = \bar{y} - b\bar{x}$.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

The line minimizes the sum of squared residuals. Its fit is measured by:

  • $s$, the standard deviation of the residuals – the typical prediction error, in the response's units.
  • $r^2$, the coefficient of determination 决定系数 – the proportion of the variation in $y$ that the linear model explains (a value between $0$ and $1$; multiply by $100$ to state it as a percent). Report it in context: "$r^2 = 0.81$ means 81% of the variation in $y$ is explained by the linear relationship with $x$."
עברית
הקו לפחות-הריבועים ממזער את סכום השאריות המרובעות
הקו לפחות-הריבועים ממזער את סכום השאריות המרובעות

הקו ממזער את סכום השאריות המרובעות. ההתאמה שלו נמדדת על ידי:

  • ⟨$s$⟩, סטיית התקן של השאריות – טעות הכרעה טיפוסית, ביחידות התגובה.
  • $r^2$, מקדם הקביעה – החלק מהשינוי ב$y$ שהמודל הליניארי מסביר (ערך בין $0$ ל$1$; כפול ב-$100$ כדי לבטאו באחוזים). דווח במרחב: "$r^2 = 0.81$ פירושו ש-81% מהשינוי ב$y$ מוסבר על ידי הקשר הליניארי עם $x$."
Vocabulary · ⁨מילון מונחים⁩ Train · ⁨אימון⁩
English עברית
associated/əˈsəʊsɪeɪtɪd/ קשור
explanatory variable/ekˈsplænətəri ˈveərɪəbl/ משתנה הסבר
response variable/rɪˈspɒns ˈveərɪəbl/ משתנה תגובה
two-way table/tuː weɪ ˈteɪbl/ טבלה דו-כיוונית
marginal distributions/ˈmɑːdʒɪnl ˌdɪstrɪˈbjuːʃnz/ התפלגויות שוליים
conditional distribution/kənˈdɪʃənl ˌdɪstrɪˈbjuːʃn/ התפלגות תנאיית
Segmented bar charts/seɡˈmentɪd bɑː tʃɑːts/ גרפי עמודות מחולקים
scatterplot/ˈskætəplɒt/ תרשים פיזור
correlation coefficient/ˌkɒrɪˈleɪʃn ˌkəʊɪˈfɪʃənt/ מקדם קורלציה
least-squares regression line/liːst skweəz rɪˈɡreʃn laɪn/ ישר רגרסיה של פחותים מרובעים
slope/sləʊp/ שיפוע
y-intercept/waɪ ˌɪntəˈsept/ נקודת חיתוך עם ציר ה-y
extrapolation/ekˈstræpəleɪʃn/ אקסטרפולציה
residual/rɪˈsɪdʒuːəl/ שארית
residual plot/rɪˈsɪdʒuːəl plɒt/ גרף שאריות
coefficient of determination/ˌkəʊɪˈfɪʃənt ɒv dɪˌtɜːmɪˈneɪʃn/ מקדם הקביעה
high-leverage/haɪ ˈliːvərɪdʒ/ בעל השפעה גבוהה
influential/ˌɪnfluːˈenʃl/ משפיע
2.9

Departures from Linearity · ⁨סטייה מליניאריות⁩

Syllabus · ⁨סיילבוס⁩
English

Enduring Understanding (DAT-1): Regression models may allow us to predict responses to changes in an explanatory variable.

Learning Objective DAT-1.I: Identify influential points in regression. [Skill 2.A]

  • DAT-1.I.1 An outlier in regression is a point that does not follow the general trend shown in the rest of the data and has a large residual when the Least Squares Regression Line (LSRL) is calculated.
  • DAT-1.I.2 A high-leverage point in regression has a substantially larger or smaller $x$-value than the other observations have.
  • DAT-1.I.3 An influential point in regression is any point that, if removed, changes the relationship substantially. Examples include much different slope, $y$-intercept, and/or correlation. Outliers and high leverage points are often influential.

Learning Objective DAT-1.J: Calculate a predicted response using a least-squares regression line for a transformed data set. [Skill 2.C]

  • DAT-1.J.1 Transformations of variables, such as evaluating the natural logarithm of each value of the response variable or squaring each value of the explanatory variable, can be used to create transformed data sets, which may be more linear in form than the untransformed data.
  • DAT-1.J.2 Increased randomness in residual plots after transformation of data and/or movement of $r^2$ to a value closer to 1 offers evidence that the least-squares regression line for the transformed data is a more appropriate model to use to predict responses to the explanatory variable than the regression line for the untransformed data.
עברית

הבנה מתמשכת (DAT-1): מודלים רגרסיה עשויים לאפשר לנו לחזות תגובות לשינויים במשתנה הסבר.

מטרת למידה DAT-1.I: זיהוי נקודות השפעתיות ברגרסיה. [מיומנות 2.A]

  • DAT-1.I.1 נקודת חריגה ברגרסיה היא נקודה שלא עוקבת אחר מגמה כללית המוצגת בשאר הנתונים ולها שארית גדולה כאשר מחושב קו הרגרסיה לפחות ריבועים (LSRL).
  • DAT-1.I.2 נקודת שילוב גבוהה ברגרסיה היא נקודה שערכה ב $x$ גדול או קטן משמעותית מערכיה של שאר התצפיות.
  • DAT-1.I.3 נקודה השפעתית ברגרסיה היא כל נקודה שהסרתה משנה את הקשר באופן משמעותי. דוגמאות כוללות שיפוע שונה מאוד, חיתוך ציר ה $y$ ו/או קורלציה שונים. נקודות חריגה ונקודות שילוב גבוהות הן לעיתים נקודות השפעתיות.

מטרת למידה DAT-1.J: לחשב תגובה מוצפת באמצעות קו רגרסיה לפחות ריבועים עבור קבץ נתונים מופעל. [מיומנות 2.C]

  • DAT-1.J.1 הפעלות של משתנים, כמו חישוב הלוגריתם הטבעי של כל ערך במשתנה התגובה או הריבוע של כל ערך במשתנה ההסבר, יכולות לשמש ליצירת קבצי נתונים מופעלים, שעשויים להיות יותר ליניאריים בצורתם מאשר הנתונים לא מופעלים.
  • DAT-1.J.2 הגדלת האקראיות בגרפי שאריות לאחר הפעלת נתונים ו/או תנועת $r^2$ לערך הקרוב יותר ל-1 מספקת ראיה לכך שקו הרגרסיה לפחות ריבועים עבור הנתונים המופעלים הוא מודל מתאים יותר לשימוש כדי לחזות תגובות למשתנה ההסבר מאשר קו הרגרסיה עבור הנתונים לא מופעלים.

Source: College Board AP Course and Exam Description · ⁨מקור: תיאור הקורס והמבחן של College Board AP⁩

English

Some points strongly affect the line. A high-leverage 高杠杆 point has an extreme $x$-value; an influential 有影响的 point noticeably changes the slope or $r$ when removed; an outlier here is a point with a large residual. When the pattern is curved, transform a variable (e.g. take a log) to straighten it, then fit a line to the transformed data.

עברית

חלק מהנקודות משפיעות חזק על הקו. נקודה בעלת שקל גבוה היא בעלת ערך קיצוני ב$x$; נקודה משפיעה משנה משמעותית את השיפוע או את ה$r$ כאשר נשלפת; נקודת אנומליה כאן היא נקודה עם שארית גדולה. כאשר הדפוס מעוקם, עבור משתנה (למשל, לקחת לוגaritmo) כדי להחליק אותו, ולאחר מכן התאם ישר לנתונים המעובים.

2.9

Exam tips · ⁨טיפים לבחינות⁩

English
  • On a scatterplot describe direction, form, strength, and outliers; $r$ ranges $-1$ to $1$.
  • Correlation is not causation — a lurking variable can drive both.
  • Interpret the slope of the least-squares line in context ("per one unit of $x$, predicted $y$ changes by $b$").
  • Check a residual plot: no pattern means a line fits; a curve means it does not. Avoid extrapolation.
  • $r^2$ is the fraction of variation in $y$ explained by the model.
עברית
  • בגרף פיזור תאר כיוון, צורה, חוזקה ואנומליות; $r$ נע בין $-1$ ל$1$.
  • קורלציה אינה סיבתיות – משתנה נסתר עשוי להיות הגורם בשניהם.
  • פרש את השיפוע של הישר פחות הריבועים במרחב ("כל שינוי של יחידה אחת ב$x$, השינוי הנבוא ב$y$ הוא $b$")."
  • בדוק גרף שאריות: העדר דפוס מצביע על התאמה טובה לישר; עיקום מצביע שלא. הימנע מהתנבאות מחוץ לטווח.
  • $r^2$ הוא השבר של השינוי ב$y$ שמוסבר על ידי המודל.

Interactive lessons on this topic · ⁨שיעורים אינטראקטיביים בנושא זה⁩

Work through it step by step, with instant-check exercises. · ⁨לעבור על הדברים צעד אחר צעד, עם תרגילים לבדיקה מיידית.⁩

Past Papers · ⁨מבחני עבר⁩

More topics in AP Statistics · ⁨סטטיסטיקה - AP⁩ · ⁨נושאים נוספים בAP Statistics · ⁨סטטיסטיקה - AP⁩⁩

Log in or create account · ⁨היכנס או צור חשבון⁩

IGCSE, A-Level & AP