The key ideas
Lines of best fit
A model like gives predictions, not exact values. Read the slope as "predicted change per unit" and the intercept as "predicted value when x is 0", which only makes sense if 0 is near the data.
A positive residual means the point is above the line. A negative residual means it is below.
Samples and margin of error
If a random sample gives an estimate with a margin of error, the plausible values for the whole population are the estimate minus the margin to the estimate plus the margin. Bigger random samples give smaller margins of error.
| How the study was done | Can it generalize to the population? | Can it show cause and effect? |
|---|---|---|
| Random sample, no random assignment (survey) | Yes, to the population sampled | No, only an association |
| Volunteers, randomly assigned to groups (experiment) | No, only to people like the volunteers | Yes, for people like them |
| Random sample and random assignment | Yes | Yes |
| Neither | No | No |
Worked examples
Example 1: use and explain a line of best fit
Problem For a group of plants, the line of best fit for height h, in centimeters, after w weeks is . What height does the model predict at 10 weeks, and what does 2.4 mean?
- Put into the model.
- 2.4 is the slope: the predicted height goes up 2.4 cm for each additional week.
Answer 55 cm. The model predicts growth of 2.4 cm per week.
Example 2: a residual
Problem One plant in Example 1 was actually 51 cm tall at 10 weeks. What is its residual?
- Actual minus predicted.
- The residual is negative, so this plant is 4 cm below the line of best fit.
Answer cm
Example 3: margin of error
Problem A random sample of 500 voters in a city finds that 46% support a new park, with a margin of error of 3 percentage points. What range of values is plausible for the whole city?
- Lower end.
- Upper end.
- Every plausible value is below 50%, so the data suggests less than half of the city's voters support the park. It does not tell you about voters in other cities.
Answer 43% to 49%
Example 4 (SAT-hard): what can the study prove?
Problem A researcher recruits 300 student volunteers. She randomly assigns half to use a flashcard app for a month and half to study as usual. The app group's average quiz score is clearly higher. Which conclusion is supported?
- Random assignment: yes. So the difference can be credited to the app, for students like these.
- Random sample: no, they are volunteers. So the result cannot be extended to all students.
Answer The app likely caused higher scores for students similar to the volunteers, but the result cannot be generalized to all students.
Common mistakes
- Computing the residual backwards. Predicted minus actual flips the sign. Fix: residual is always actual minus predicted.
- Treating predictions as exact. A model value is an estimate. Fix: say "the model predicts", not "the plant will be".
- Claiming cause from a survey. A survey can show that two things go together, not that one causes the other. Fix: look for random assignment.
- Generalizing past the sample. A random sample of one school says nothing certain about all schools. Fix: the conclusion stops at the population that was sampled.
- Thinking a bigger margin of error is better. A smaller margin means a more precise estimate. Fix: larger random samples shrink the margin.
Quick method
Practice
5 SAT-style questions
A line of best fit for cups of lemonade sold, y, at a price of x dollars is . How many cups does the model predict at a price of $8?
- 26
- 28
- 32
- 52
Show answer
Answer: 28
. 52 adds the 12 instead of subtracting it.
Using the model in question 1, a stand actually sold 31 cups at $8. What is the residual?
Show answer
Answer:
Actual minus predicted: . subtracts in the wrong order. A positive residual means the point is above the line.
A random sample of students at a school sleep an average of 6.2 hours a night, with a margin of error of 0.4 hours. Which value is a plausible average for all students at the school?
- 5.7 hours
- 6.5 hours
- 6.7 hours
- 7.0 hours
Show answer
Answer: 6.5 hours
Plausible values run from to hours. Only 6.5 is in that range.
A survey of 200 randomly selected students at Lincoln High School finds that 62% bike to school at least once a week. To which group can this result be generalized?
- All high school students in the US
- All students at Lincoln High School
- Only the 200 students surveyed
- Students who like biking
Show answer
Answer: All students at Lincoln High School
The sample was random from Lincoln High, so the estimate applies to that school. It says nothing reliable about other schools, and it is more than a fact about the 200 students because the sample was random.
Student-produced response: in a random sample of 250 students from a school of 2,000 students, 60 walk to school. Based on the sample, about how many students at the school walk to school?
Show answer
Answer: 480
The sample rate is . Apply it to the whole school: .
Frequently asked questions
What does the slope of a line of best fit mean?
It is the predicted change in y for each increase of 1 in x. For example, a slope of 2.4 in a height model means the model predicts 2.4 more centimeters each week. It is a prediction, not a promise for any one data point.
How does sample size affect margin of error?
A larger random sample gives a smaller margin of error, because the estimate is more precise. A biased sample, like only volunteers, stays biased no matter how large it is.
When can a study show cause and effect?
Only when subjects are randomly assigned to groups, like a treatment group and a control group. Surveys and observational studies can show that two things are related, but something else might explain the link.