sqoolpapers.co.za

Study Guide For Grade

11

Mathematics

Term 4
Paper 2

Statistics

Statistics is the collection, organisation, representation, analysis and interpretation of data. Data may be presented as a list, frequency table, graph or diagram. In examinations, always arrange raw data in ascending order before calculating the median, quartiles or five number summary.

Measures of central tendency

Measures of central tendency describe a value around which the data is centred. The three main measures are the mean, median and mode.

The most suitable measure depends on the data:

  • The mean uses every value, but is affected by extreme values.
  • The median is the middle value and is less affected by extreme values.
  • The mode is the value that occurs most often and can be used for numerical or categorical data.

Mean, median and mode

Mean

For ungrouped data, the mean is calculated by dividing the sum of all the values by the number of values.

x¯=xn

Example

Calculate the mean of 2, 4, 4, 5, 7 and 8.

x¯=2+4+4+5+7+86
x¯=306
x¯=5

Median

The median is the middle value when the data is arranged in ascending order.

If there is an odd number of values, the median is the middle value. If there is an even number of values, the median is the mean of the two middle values.

Example

Determine the median of 2, 4, 4, 5, 7 and 8.

There are six values, so use the third and fourth values.

Median=4+52
Median=4,5

Mode

The mode is the value that occurs most frequently. In the data set 2, 4, 4, 5, 7 and 8, the value 4 occurs twice, while every other value occurs once. Therefore, the mode is 4.

A data set may have one mode, more than one mode or no mode.

Grouped and ungrouped data

Ungrouped data consists of individual observations, such as 3, 5, 7 and 9. Grouped data is organised into class intervals.

Consider the following grouped data:

Class interval: 0 to less than 10; frequency: 3

Class interval: 10 to less than 20; frequency: 5

Class interval: 20 to less than 30; frequency: 8

Class interval: 30 to less than 40; frequency: 4

For grouped data, use each interval’s midpoint as an estimate of the values in that interval.

Midpoint=lower boundary+upper boundary2

The midpoints are 5, 15, 25 and 35.

The estimated mean of grouped data is:

x¯=fxf

Example

Calculate the estimated mean of the grouped data above.

f=3+5+8+4
f=20
fx=(3)(5)+(5)(15)+(8)(25)+(4)(35)
fx=15+75+200+140
fx=430
x¯=43020
x¯=21,5

The answer is an estimate because the exact values in each interval are unknown.

Modal interval

The modal interval is the class interval with the highest frequency.

In the grouped data above, the highest frequency is 8. Therefore, the modal interval is:

20x<30

Common Mistake

Do not give the midpoint as the modal interval. The modal interval must be written as the complete class interval.

Measures of dispersion

Measures of dispersion describe how spread out the data is. Two data sets may have the same mean but very different spreads.

Common measures of dispersion include:

  • Range
  • Interquartile range
  • Variance
  • Standard deviation

The range is the difference between the maximum and minimum values.

Range=maximumminimum

For 2, 4, 4, 5, 7 and 8:

Range=82
Range=6

Variance

Variance measures the average squared distance of the data values from the mean.

For a population:

σ2=(xx¯)2n

Example

Calculate the variance of 2, 4, 4, 5, 7 and 8. The mean is 5.

First calculate the squared deviations from the mean.

(xx¯)2=(25)2+(45)2+(45)2+(55)2+(75)2+(85)2
(xx¯)2=9+1+1+0+4+9
(xx¯)2=24

Now calculate the variance.

σ2=246
σ2=4

Variance is expressed in squared units.

Standard deviation

Standard deviation is the positive square root of the variance.

σ=(xx¯)2n

For the previous example:

σ=4
σ=2

A small standard deviation shows that the data values are close to the mean. A large standard deviation shows that the values are widely spread around the mean.

Exam Tip

When using a calculator, first check whether the question refers to a population or a sample. In most school data-handling questions, the population standard deviation is used unless stated otherwise.

Five number summary

The five number summary gives an overview of the position and spread of ordered data. It consists of:

  • Minimum value
  • Lower quartile
  • Median
  • Upper quartile
  • Maximum value

The lower quartile is the median of the lower half of the data. The upper quartile is the median of the upper half.

Example

Determine the five number summary of 2, 4, 4, 5, 7 and 8.

The median is:

Q2=4+52=4,5

The lower half is 2, 4, 4, so:

Q1=4

The upper half is 5, 7, 8, so:

Q3=7

The five number summary is:

2;4;4,5;7;8

The interquartile range measures the spread of the middle half of the data.

IQR=Q3Q1
IQR=74
IQR=3

Box and whisker diagrams

A box and whisker diagram is drawn from the five number summary.

  • The box extends from the lower quartile to the upper quartile.
  • A line inside the box shows the median.
  • The whiskers extend towards the minimum and maximum values, unless outliers are plotted separately.

Box and whisker diagrams are useful for comparing the centres and spreads of different data sets. A longer box indicates a larger interquartile range. An off-centre median may indicate skewness.

The diagram must be drawn on a labelled number line using a consistent scale.

Histograms

A histogram represents continuous grouped data. The horizontal axis shows class intervals, and the vertical axis normally shows frequency.

The bars touch because the intervals are continuous. There are no gaps between consecutive bars.

For equal class widths, the height of each bar represents its frequency. If class widths are unequal, frequency density should be used.

Frequency density=frequencyclass width

Common Mistake

A histogram is not a bar graph. Histogram bars touch, while bar graph bars are separated.

Frequency polygons

A frequency polygon represents grouped data by plotting the midpoint of each class interval against its frequency.

For the grouped data used earlier, the points are:

(5;3),(15;5),(25;8),(35;4)

Join consecutive points with straight-line segments. Extra points with frequency zero may be added before the first interval and after the last interval to close the polygon.

Frequency polygons are useful when comparing two distributions on the same set of axes.

Ogives and cumulative frequency curves

An ogive is a cumulative frequency curve. Cumulative frequency is the running total of the frequencies.

For frequencies 3, 5, 8 and 4, the cumulative frequencies are calculated as follows:

3
3+5=8
8+8=16
16+4=20

Therefore, the cumulative frequencies are 3, 8, 16 and 20.

Plot each cumulative frequency at the upper boundary of its class interval. The curve normally begins at the lower boundary of the first interval with cumulative frequency zero.

An ogive can be used to estimate the median and quartiles. If there are 20 observations, locate the following cumulative frequencies:

Q1:204=5
Q2:202=10
Q3:3(20)4=15

Read the corresponding data values from the horizontal axis.

Symmetric and skewed data

A symmetric distribution has approximately the same shape on both sides of its centre. For a perfectly symmetric unimodal distribution:

mean=median=mode

A positively skewed distribution has a long tail extending to the right. Large values pull the mean towards the right.

mode<median<mean

A negatively skewed distribution has a long tail extending to the left. Small values pull the mean towards the left.

mean<median<mode

The median is usually a better measure of central tendency for skewed data because it is less affected by extreme values.

Identifying outliers

An outlier is a value that lies unusually far from the other observations. Outliers can strongly affect the mean, range and standard deviation.

Use the interquartile range to calculate the lower and upper fences.

Lower fence=Q11,5(IQR)
Upper fence=Q3+1,5(IQR)

Any value below the lower fence or above the upper fence is an outlier.

Example

Identify any outliers in 2, 4, 5, 6, 7, 8, 9 and 25.

The median is:

Q2=6+72=6,5

The lower and upper quartiles are:

Q1=4+52=4,5
Q3=8+92=8,5

Calculate the interquartile range.

IQR=8,54,5
IQR=4

Calculate the fences.

Lower fence=4,51,5(4)
Lower fence=1,5
Upper fence=8,5+1,5(4)
Upper fence=14,5

Since 25 is greater than 14,5, it is an outlier.

Remember

Always show the calculation of the interquartile range and both fences. Do not identify a value as an outlier merely because it appears much larger or smaller than the other values.