Line of best fit and regression
Fit a line of best fit by eye and identify the least-squares regression line; interpret the gradient and y-intercept in context; use the regression equation for prediction; distinguish between interpolation and extrapolation and understand the risks of extrapolating beyond the data range.
Worked examples
Using a regression equation to make a prediction
Straightforward
Problem
The least-squares regression line for the relationship between daily temperature (, °C) and electricity usage (, kWh) for a household is:
A day is forecast to reach 25°C. Predict the household's electricity usage on that day.
A day is forecast to reach 25°C. Predict the household's electricity usage on that day.
1
Identify the known value and what is being predicted.
Known: (temperature in °C)
Find: (predicted electricity usage in kWh)
Find: (predicted electricity usage in kWh)
2
Substitute into the regression equation.
3
Interpret the result in context.
A predicted usage of kWh makes sense: the negative gradient () shows that hotter days reduce electricity usage, presumably because heating is needed less.
The -intercept kWh would be the predicted usage on a °C day.
The -intercept kWh would be the predicted usage on a °C day.
Answer
The model predicts the household will use kWh of electricity on a 25°C day.
Interpreting gradient and y-intercept; identifying interpolation vs extrapolation
Moderate
Problem
A regression line fitted to data for 30 school canteens (with between 200 and 800 students enrolled) is:
where is the number of students enrolled and is the number of lunch orders per day.
(a) Interpret the gradient and -intercept in context.
(b) Predict the number of lunch orders for a canteen with 500 students.
(c) A large secondary school has 1200 students. Is using this model to predict their lunch orders reliable? Explain.
where is the number of students enrolled and is the number of lunch orders per day.
(a) Interpret the gradient and -intercept in context.
(b) Predict the number of lunch orders for a canteen with 500 students.
(c) A large secondary school has 1200 students. Is using this model to predict their lunch orders reliable? Explain.
1
Interpret the gradient.
Gradient : for each additional student enrolled, the model predicts more lunch orders per day.
2
Interpret the -intercept.
-intercept : the model predicts 120 lunch orders when enrolment is 0.
Note: this is outside the data range (200–800 students), so the intercept does not have a meaningful practical interpretation here.
Note: this is outside the data range (200–800 students), so the intercept does not have a meaningful practical interpretation here.
3
Predict lunch orders for 500 students.
is within the data range (200 to 800), so this is **interpolation**.
Predicted orders: 345 per day.
Predicted orders: 345 per day.
4
Assess the reliability of predicting for 1200 students.
is well outside the data range of 200–800. This is **extrapolation**.
The linear relationship between enrolment and orders may not continue to hold beyond 800 students (e.g. canteen capacity limits, different school structures). The prediction is unreliable.
The linear relationship between enrolment and orders may not continue to hold beyond 800 students (e.g. canteen capacity limits, different school structures). The prediction is unreliable.
Answer
(a) Gradient: for each additional student, there are 0.45 more predicted lunch orders. -intercept: 120 predicted orders with zero enrolment (not practically meaningful).
(b) 345 lunch orders predicted for 500 students.
(c) Not reliable — 1200 students is outside the data range (extrapolation), and the linear model may not hold.
(b) 345 lunch orders predicted for 500 students.
(c) Not reliable — 1200 students is outside the data range (extrapolation), and the linear model may not hold.
Practise
Q1·Straightforward
A regression line for the relationship between hours of study () and exam mark () is . Predict the exam mark for a student who studies for 6 hours.
Explanation
Substitute :
The model predicts an exam mark of for a student who studies for 6 hours.
The model predicts an exam mark of for a student who studies for 6 hours.
Q2·Straightforward
A regression line has equation , where is the number of staff absent and is daily production output (units). Predict the output when 8 staff are absent.
Explanation
Substitute :
The model predicts a daily output of units when 8 staff are absent.
The model predicts a daily output of units when 8 staff are absent.
Q3·Straightforward
A regression line for weekly advertising spend (, in $) and weekly revenue (, in $) is . What does the gradient represent?
Explanation
The gradient of a regression line is the **rate of change**: for every additional $1 increase in advertising spend (), the predicted weekly revenue () increases by $12.
This is an *average* effect — the regression line models the typical relationship across all observations.
This is an *average* effect — the regression line models the typical relationship across all observations.
Q4·Straightforward
In the regression equation , where is weekly advertising spend and is weekly revenue (both in $), what does the -intercept represent?
Explanation
The -intercept is the predicted value of when . Here, it means the model predicts weekly revenue of $800 even when nothing is spent on advertising.
This could represent baseline revenue from repeat customers, but care should be taken: if is outside the data range, this intercept may not be meaningful to interpret literally.
This could represent baseline revenue from repeat customers, but care should be taken: if is outside the data range, this intercept may not be meaningful to interpret literally.
Q5·Moderate
A regression line is , where is temperature (°C) and is daily ice cream sales. A shop owner wants to sell exactly ice creams. What temperature does the model predict this will occur at? Give your answer to the nearest whole number.
Explanation
Set and solve for :
The model predicts ice cream sales of 78 units will occur at a temperature of °C.
The model predicts ice cream sales of 78 units will occur at a temperature of °C.
Q6·Moderate
A regression line for arm span (, cm) versus height (, cm) is . A student is cm tall and has an arm span of cm. What is the residual (actual minus predicted) for this student? Give your answer to 1 decimal place.
Explanation
Find the predicted arm span:
Calculate the residual:
The positive residual means this student's arm span is cm longer than the model predicts.
Calculate the residual:
The positive residual means this student's arm span is cm longer than the model predicts.
Q7·Moderate
A regression line was fitted to data for students aged 12 to 17, relating weekly reading hours () to vocabulary score (). Which prediction is more reliable, and why?
Explanation
**Interpolation** — predicting within the data range — is more reliable than **extrapolation** — predicting outside it.
The regression line was fitted to students aged 12–17 reading up to a certain number of hours. Predicting for a student who reads 6 hours per week (within the range) is a valid use of the model.
Predicting for an adult or for 50 hours per week would be extrapolating: the relationship may not hold beyond the original data range, and the model could give misleading results.
The regression line was fitted to students aged 12–17 reading up to a certain number of hours. Predicting for a student who reads 6 hours per week (within the range) is a valid use of the model.
Predicting for an adult or for 50 hours per week would be extrapolating: the relationship may not hold beyond the original data range, and the model could give misleading results.
Q8·Moderate
Two points that lie on a regression line are and . Find the gradient of the regression line.
Explanation
Apply the gradient formula:
The gradient of the regression line is , meaning for each 1-unit increase in , the predicted increases by .
The gradient of the regression line is , meaning for each 1-unit increase in , the predicted increases by .
Q9·Moderate
The least-squares regression line for a dataset is , where is hours studied per day and is exam score (out of 100). Which statement correctly interprets the gradient ?
Explanation
The gradient means: for each additional hour of study per day, the predicted exam score increases by marks (on average).
The -intercept is the predicted score when study time is zero — though this should only be interpreted if is within the range of the data.
The -intercept is the predicted score when study time is zero — though this should only be interpreted if is within the range of the data.
Q10·Challenging
A regression line for years of experience () and annual salary in $1000s () is . A worker with 12 years of experience earns $72\,000. Find the residual (actual minus predicted salary) in thousands of dollars, to 1 decimal place.
Explanation
The actual salary is $72\,000 = (in thousands).
Find the predicted salary:
Calculate the residual:
This worker earns $5200 more per year than the regression model predicts for someone with 12 years of experience.
Find the predicted salary:
Calculate the residual:
This worker earns $5200 more per year than the regression model predicts for someone with 12 years of experience.
Q11·Challenging
A regression model for house price (, in $1000s) versus floor area (, in m²) was built using data for houses between 80 m² and 250 m². Two predictions are made: Prediction A is for a 160 m² house; Prediction B is for a 400 m² house. Which statement is correct?
Explanation
The regression model was built on data from 80 m² to 250 m².
- **Prediction A** (160 m²): within the data range — this is **interpolation** and is more reliable.
- **Prediction B** (400 m²): well outside the data range — this is **extrapolation**. The linear relationship may not hold at such extreme values, so this prediction is much less reliable.
Using a regression model beyond its data range can produce misleading or meaningless results.
- **Prediction A** (160 m²): within the data range — this is **interpolation** and is more reliable.
- **Prediction B** (400 m²): well outside the data range — this is **extrapolation**. The linear relationship may not hold at such extreme values, so this prediction is much less reliable.
Using a regression model beyond its data range can produce misleading or meaningless results.
Q12·Challenging
A regression line for vehicle mass (, tonnes) and fuel consumption (, L/100 km) is . A car weighs 1.5 tonnes and actually uses 12.6 L/100 km. Find the residual to 1 decimal place.
Explanation
Find the predicted fuel consumption:
Calculate the residual:
The positive residual means this car uses L/100 km more than the model predicts for a vehicle of its mass.
Calculate the residual:
The positive residual means this car uses L/100 km more than the model predicts for a vehicle of its mass.
Open Math
Line of best fit and regression
Statistical Analysis · MS-S4
Name:
Date:
Q1Straightforward
A regression line for the relationship between hours of study () and exam mark () is . Predict the exam mark for a student who studies for 6 hours.
Q2Straightforward
A regression line has equation , where is the number of staff absent and is daily production output (units). Predict the output when 8 staff are absent.
Q3Straightforward
A regression line for weekly advertising spend (, in $) and weekly revenue (, in $) is . What does the gradient represent?
- A.The predicted revenue when advertising spend is zero
- B.For every additional $1 spent on advertising, predicted revenue increases by $12
- C.The maximum possible weekly revenue
- D.The number of advertising campaigns per week
Q4Straightforward
In the regression equation , where is weekly advertising spend and is weekly revenue (both in $), what does the -intercept represent?
- A.The advertising spend that gives the maximum revenue
- B.The predicted weekly revenue when there is no advertising spend ()
- C.The minimum revenue the business can earn
- D.The number of customers per week
Q5Moderate
A regression line is , where is temperature (°C) and is daily ice cream sales. A shop owner wants to sell exactly ice creams. What temperature does the model predict this will occur at? Give your answer to the nearest whole number.
Q6Moderate
A regression line for arm span (, cm) versus height (, cm) is . A student is cm tall and has an arm span of cm. What is the residual (actual minus predicted) for this student? Give your answer to 1 decimal place.
Q7Moderate
A regression line was fitted to data for students aged 12 to 17, relating weekly reading hours () to vocabulary score (). Which prediction is more reliable, and why?
- A.Predicting vocabulary for a student who reads 6 hours per week, because this is within the data range (interpolation)
- B.Predicting vocabulary for an adult who reads 6 hours per week, because adults read more carefully
- C.Predicting vocabulary for a student who reads 50 hours per week, because more data is always better
- D.Both predictions are equally reliable because they use the same equation
Q8Moderate
Two points that lie on a regression line are and . Find the gradient of the regression line.
Q9Moderate
The least-squares regression line for a dataset is , where is hours studied per day and is exam score (out of 100). Which statement correctly interprets the gradient ?
- A.A student with no study time is predicted to score
- B.For each additional hour of study per day, the predicted exam score increases by marks
- C.The maximum possible exam score is more than the minimum
- D. of students studied for more than one hour per day
Q10Challenging
A regression line for years of experience () and annual salary in $1000s () is . A worker with 12 years of experience earns $72\,000. Find the residual (actual minus predicted salary) in thousands of dollars, to 1 decimal place.
Q11Challenging
A regression model for house price (, in $1000s) versus floor area (, in m²) was built using data for houses between 80 m² and 250 m². Two predictions are made: Prediction A is for a 160 m² house; Prediction B is for a 400 m² house. Which statement is correct?
- A.Prediction A is more reliable because 160 m² is within the data range (interpolation)
- B.Prediction B is more reliable because the house is larger and worth more
- C.Both predictions are equally reliable because they use the same regression equation
- D.Prediction A is less reliable because it is closer to the middle of the data range
Q12Challenging
A regression line for vehicle mass (, tonnes) and fuel consumption (, L/100 km) is . A car weighs 1.5 tonnes and actually uses 12.6 L/100 km. Find the residual to 1 decimal place.
Worked solutions and answers at openmath.au/year-12/standard-2/bivariate-data-analysis/line-of-best-fit-and-regression