📖 Crammy · All study guides
AP Statistics · Unit 2

Exploring Bivariate Data: every key term you need

22 flashcard terms for AP Statistics Unit 2, written to match the course framework. Study them here, then drill them as interactive flashcards — free, no account needed.

Study this unit free →

More AP Statistics guides

Exploring Bivariate Data
Examining relationship between two variables; key questions: Is there association? How strong? What's the pattern?
Scatterplot
Plot with one variable on x-axis, other on y-axis; shows association pattern (linear, curved, clustered, none).
Correlation Coefficient (r)
Measures strength/direction of LINEAR relationship: r=-1 (perfect negative), r=0 (no linear), r=+1 (perfect positive).
Correlation Properties
Unitless, ranges -1 to +1. Same r if x/y swapped. r ≠ causation. Sensitive to outliers. Only measures linear relationships.
Causation vs Correlation
Strong correlation ≠ causation. May have: reverse causation, confounding variables, lurking variables, coincidence.
Least Squares Regression Line
Line minimizing squared vertical distances (residuals) from points. Equation: ŷ = a + bx where b=r(sy/sx).
Regression Line Interpretation
Slope b: for each 1-unit increase in x, y increases by b units (on average). Intercept a: predicted y when x=0.
Predictions & Extrapolation
Use regression to predict y for given x (interpolation is OK). Extrapolation (beyond data range) unreliable.
Residuals
Actual y - Predicted ŷ. Residual plot (residuals vs x) should show random scatter (no pattern) if linear model appropriate.
Coefficient of Determination (R²)
R² = r²; proportion of y's variation explained by x. R²=0.85 means 85% of variation in y explained by x.
Outliers & Influential Points
Outlier: unusual y value. Influential: point far from others on x-axis, dramatically changes regression line.
Transformations
When linear model fails, transform data (log, square root) to achieve linearity; analyze transformed data.
Categorical Variables
Use indicator (dummy) variable (0/1) as predictor. Regression slope compares mean outcome between categories.
Simpson's Paradox
Trend reverses when data subdivided by group. Shows importance of examining relationships within subgroups.
Ecological Fallacy
Drawing individual-level conclusions from group-level data (e.g., states' data doesn't apply to individuals in states).
Spurious Correlation
Strong correlation exists due to confounding variable, not direct causation. Classic example: ice cream sales & drownings (temperature).
Two-Way Tables
Frequency table for two categorical variables. Examine marginal (row/column) distributions and conditional distributions.
Conditional Distribution
Distribution of one variable given specific value of another. Shows how relationship differs within subgroups.
Drill these as interactive flashcards →
Independence
Two categorical variables independent if conditional distributions same as marginal distribution (no association).
Chi-Square Association
Measures strength of association between categorical variables. Higher value = stronger association.
Mosaic Plots
Visual display of two-way table; tile size represents cell frequency; useful for categorical relationships.
Unit 2 Key Ideas
Correlation measures linear association; r close to ±1 = strong, r≈0 = weak. Regression predicts y from x but doesn't imply causation. Residual plots reveal if linear model appropriate.
Turn these into flashcards & quizzes →