(Must be 250 or more words)
The larger the sample, the more reliable the results. Do you agree or disagree with the
statement and why?
Total bare minimum word count for the initial discussion post is 250 words with a minimum of
two peer-reviewed sources from the Bethel University Databases, plus your textbook located in
the read section. Further, when conducting outside research limit the search to the past two
years. In addition, respond to two classmates with a minimum of 75 words each. We will
keep this word count format for the discussion section of all five units.
CHAPTER TWO
PROBABILITY & THE NORMAL DISTRIBUTION
W
I
What comes to mind when you hear the word
L probability? Most people think of gambling
and it is understandable since you can never
L bet on a sure thing. But what is probability
really? It is the chance or likelihood of an
I event occurring. Whether you are deciding to
take an umbrella because you just heard there would be a 50% chance of rain or you are
S
making the decision to take another card in blackjack based on your total and a dealer’s
up card or you are choosing which route, to take home based on likely traffic from past
Basics of Probability
experience, probability is a part of our daily lives. We unknowingly use probability to
make routine decisions, but a better understanding
of the concepts could actually help in
K
your making the best decisions.
A
S (payoffs); Kentucky Derby and Pro Sports
Probability is NOT the same as betting odds
sets odds to insure they make money and S
the odds depend on how the betting goes. Lower
odds does not mean more likely to win, just
A that more people think it is true. Vegas tables
set payoffs lower than true odds so they always
come out ahead.
N
D
Mathematically, probability equals the number of events meeting the specified condition
R
divided by the number of possibilities. With a deck of playing cards, there are 52 possible
A ace from a deck of cards = the number of cards
events. Thus, the probability of drawing an
meeting the condition of being an ace (i.e., 4) divided by the number of possibilities (i.e.,
52) = 4 / 52.
2
1
An event is just an uncertain outcome, the
6 result of an experiment; an experiment is the
process of making an observation. If you flip a coin (experiment), there are two possible
1
events à heads or tails. If you flip a coin several times, you expect 1/2 Heads, 1/2 Tails. If
you flip it 10 times you expect 5 Heads, 5TTails, but what if you get 10 Tails in a row? Has
the probability changed? Not really. TheSprobability is still 50% of getting Heads or Tails
because what happened in the past has nothing to do with the future (when dealing with an
independent event like a coin flip).
Probability is always between 0 and 1 (0% – 100%) inclusive. It is NEVER less than 0%
and NEVER greater than 100%.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-1
Chapter Two: Probability & the Normal Distribution
Solving Probability Problems
Probability problems can easily be solved with tables. Setting up the table is usually the tricky
part, but once it is done, the rest is easy. With a deck of cards, you would make a table with the
values serving as the columns (A, 2, 3, …, J, Q, K) and the suits serving as the rows (diamonds,
hearts, clubs, spades). Basically, we just employ rules of Set Theory in which you circle the objects in question and count what is in the relevant circle.
Suppose 100 people were asked about their political affiliation and the following shows their responses:
W
Total
I
Male
10
25
5
L 40
Female
30
20
10
L 60
Total
40
45
15
I 100
S
Now if you were asked the following, you can use simple rules to solve them.
,
Dem
1.
Rep
Ind
Prob[Female]
K
A
3.
Prob[Female and Democrat]
S
S
4.
Prob[Female or Democrat]
A
5.
Prob[Female / Democrat] (i.e., prob of female given Democrat)
N
6.
Prob[Democrat / Female] (i.e., prob
D of Democrat given female)
R
A
Let’s take them one at a time.
2.
Prob[Republican]
2 = 60/100 = 60%
1. Prob[Female] = total females / total people
Dem
Rep
Male
Female
10
30
25
20
Total
40
45
Ind1
56
101
15T
S
Total
40
60
100
2. Prob[Republican] = total republicans / total people = 45/100 = 45%
Dem
Rep
Ind
Total
Male
Female
10
30
25
20
5
10
40
60
Total
40
45
15
100
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-2
Chapter Two: Probability & the Normal Distribution
3. Prob[Female and Dem] = intersection of females and
democrats / total people = 30/100 = 30%
Dem
Rep
Ind
Total
Male
Female
10
30
25
20
5
10
40
60
Total
40
45
15
100
4. Prob[Female or Dem] = union of females and
W
dem / total people = (10+30+20+10)/100 =70/100 = 70%
Dem
Rep
Male
Female
10
30
25
20
Total
40
45
I
L
5
L
10I
15S
,
Ind
Total
40
60
100
5. Prob[Female/Dem] = percentage of democrats who are female (i.e., prob of
female given Dem) = # of female democrats
K / total democrats = 30/40 = 75%
Dem
A
S
Male
10
S
Female
30
A
Total
40
N
D
6. Prob[Dem/Female] = percentage of females who are democrat (i.e., prob of
R / total females = 30/60 = 50%
Dem given female) = # of female democrats
A
Dem
Rep
Ind
Total
2
Female
60
30
20
101
6
1
T
The screenshot on the next page from Joint_Probability.xls shows the results of entering the disS
cussed political affiliation data.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-3
W
I
L
L
I
S
,
Excel will solve this for you, but it is important to understand how the answers are derived. Is political
K
party independent of gender? To be independent, Prob (A given B) = Prob (A). In other words, the probA
ability of an event occurring is unchanged by the existence
of the other event being known. The Prob
S
[Female] = 60 / 100 = 60%. The Prob [ Female given Democrat] = 30 / 40 = 75%. Knowing the person’s
S so the two variables are dependent.
political party alters the probability in regards to gender,
A
An example of independent variables would be the suit
N and value of a playing card. The probability of
drawing an Ace from a deck of cards is 4 / 52. Given that the card drawn is known to be a diamond, the
D
probability that the card is an Ace is now 1 / 13. Since 4 / 52 = 1 / 13, the two variables must be indepenRvalue is independent of the suit (i.e., knowing the
dent. Thus, the probability of a card being a specific
A being a specific suit is independent of its value (it
suit will not help you), AND the probability of a card
works in both directions).
2
1
The Normal Distribution
6
If you were to look at a 5’8” man, you’d probably say he was short. If you were to look at a 5’7”” woman,
1
you’d probably say she was tall, despite the fact that she’s shorter than the man. Why? It is because we
T
tend to compare each data point to its respective mean. If we wanted to solve problems regarding heights,
S (one for men and one for women), but it is not true.
you would think we would need two separate formulas
Thanks to the Normal Probability Distribution, we can put everything on the same scale. The Normal
curve is a bell-shaped curve which peaks in the middle at the mean. Units on the curve are measured in
terms of the standard deviation; one standard deviation in both directions from the mean captures 68% of
the data, two standard deviations in both directions captures 95% of the data, and three standard deviations
in both directions captures 99.7% of the data. As expected, the bulk of the data is close to the mean.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-4
Chapter Two: Probability & the Normal Distribution
The Standard Normal Distribution has values mainly between -3 and +3 as measured by a z-score, with the
z-score being the number of standard deviations a value is from the mean. The formula is
Thus the male height of 5’8” is about one standard deviation below average (z = -1) and the female height
of 5’7” is about two standard deviations above average (z = +2). The female would show up to the right
of the male on this standard scale.
The Standard Normal Distribution is commonly used. You often hear about standardized testing, but this
is what it means. The SAT scores, for example, are computed by the comparing one’s raw score with
W
the mean raw score and dividing by the standard deviation
(thus giving the z-score); then the z-score is
converted to the SAT scale, which forces the mean toI be 500 and the standard deviation to be 100. If the
mean raw score on the math section were 40 and the L
standard deviation was 10, then a person who scored
60 would have scored 20 points higher than the mean,
L and since the standard deviation is 10, that person
scored 2 standard deviations above average (z = 60-40/10 = +2.00). On the SAT system, this translates to
I
200 points above the average of 500, so the SAT score would be 700. Since it is normally distributed, we
know that about 68% of students score between 400Sand 600, 95% score between 300 and 700, and the
,
remaining 5% score below 300 or above 700.
One of the great features of using the Normal Distribution is that we can compute probabilities very easily.
K
We only need the mean and standard deviation to completely define the curve. If you know that you
A
scored 400, and you know the mean = 500 and the standard
deviation = 100, you can compute that your
S
score is 1 standard deviation below the mean.
S
A
N
D
R
A
2
1
6
1
T
————–|————–|————–|————–|————–|————-S
2.5%
13.5%
34%
34%
13.5%
2.5%
|————68%———–|
|—————————95%————————–|
|—————————————-99.7%————————————–|
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-5
Chapter Two: Probability & the Normal Distribution
If you have ever asked a teacher to grade on a curve, be careful. A true grading curve means
that any score more than 1 standard deviation about the mean gets an A, less than one standard
deviation about the mean gets a B, up to one standard deviation below the mean gets a C, one to
two standard deviations below the mean gets a D, and more than two standard deviations below
the mean gets an F. This results in approximately 16% A’s, 34% B’s, 34% C’s, 13.5% D’s, and
2.5% F’s. Some teachers adjust the curve a bit so that the cutoff for a B is even higher so there
are more C’s and fewer B’s, but the concept is the same. And this means one’s grade depends
on the grades of others. While a 90% is an A or A- in most schools, if the class mean were 95%,
then this score is below average, and depending on the standard deviation, a 90% could be in the
C or D range (although we only seem to recall the situations where our poor scores got raised).
Standardizing grades mostly benefits those who W
performed worst (and rarely those who did very
well), and the irony is that the curve killer in a class
is not the top student but rather the bottom
I
one. A student scoring 100% no doubt raises the L
class average a bit, but a student scoring 0% will
drop the average a bit more while raising the standard
L deviation an astronomical amount. So in a
class of 20 students where the average is 80% and the standard deviation is 10%, the cutoff for an
I
A would be 90%, a B would be 80%, a C would be 70% and a D would be 60%. A grade of 100
S one point to 81% while barely impacting the
added to the mix would likely raise the mean by about
standard deviation; the cutoffs for A, B, C and D ,might be 91%, 81%, 71% and 61% respectively.
But if a grade of 0 were added to the mix instead, the mean would likely drop to around 76% while
the standard deviation might increase to over 20%;
K you might then see a cutoff of 102% for an A,
76% for a B, 50% for a C and 24% for a D. ThusA
an A would be impossible, a B would be lowered
a bit, a C would be lowered a lot, and a student with a solid F can now pass the course, courtesy
S
of a poor student who deprived the A students from getting what they deserve. Rather unfair, but
such is the result of standardized scores for smallSsamples (but SATs and other standardized tests
are given to over a million students, making this A
a non-issue).
N
D
R
A
2
1
6
1
T
S
In the Normal_Probability.xls file, you can do multiple computations at once, given just a mean
and a standard deviation. For example, we know the mean SAT math score is 500 and the standard
deviation is 100. We can then compute the probability of scoring above a number, below a number,
or between two numbers (this can also be read as the percent of people who scored in those
ranges). We can also do the reverse and determine what score corresponds to a certain percentile
(in the case above, to score in the top 5% and be at the 95th percentile, you would need to score at
least 664.49).
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-6
Chapter Two: Probability & the Normal Distribution
W
I
L
L
I
S
The Normal_Probability.xls file also includes tabs, that show each of the computations graphically.
Here we see a computation for the probability of scoring less than 650 on the SAT Math. This corresponds to a z-score of 1.50 (650 – 500 = 150, 150
K divided by 100 = 1.50), as 650 is 1.50 standard
deviations above the mean. The area under the standard
normal curve = 1 or 100%, and the area
A
shaded in red corresponds to the area less than a z of 1.50. The red area accounts for 93.32% of
S
the total area, which means that 93.32% of those taking the SAT score less than 650 on the Math
S
section.
A
N
The Central Limit Theorem
D
When a distribution is not Normal, we cannot compute
probability as is. However, if we take
R
samples from the data, the means of the samples
will
appear
normal if the sample size is large
A
enough. Regardless of the look of the original distribution, larger samples will result in the curve
starting to appear Normal. As a general rule, a sample of at least 30 guarantees that the distribution
2 mean of the new distribution is the same as the
of sample means will be normally distributed. The
1
original mean, but the spread is cut down dramatically
and is known as the standard error (which
is computed from the standard deviation and the6 sample size). So instead of solving problems
involving individual values, we solve problems involving
means, but essentially they are similar.
1
When solving problems regarding the probabilityTof a sample mean being in a specified range, you
only need the mean, the standard deviation ANDSthe sample size. So in the case of the SAT, you
might want to know what the probability is of taking a sample of 4 students and getting a mean
greater than 600.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-7
Chapter Two: Probability & the Normal Distribution
W
I
L
L
I
Here we see in the Normal-CLT tab of the Normal_Probability.xls
file that 2.28% of the time,
S
a sample of 4 students will have a sample mean greater
than
600.
This
doesn’t mean that 2.28%
,
of the students score above 600. If we used a sample greater than 4, the probability would get
smaller because as you take a larger and larger sample, there is a greater likelihood that you will
K 200 to 800 on the exam and only 16%
be closer to the true mean. With scores ranging from
scoring above 600, it is hard to imagine taking a A
large sample and having the mean be greater
than 600.
S
S
The Central Limit Theorem is quite powerful. With a measure of central tendency and a measure
A and solve any probability problems; the
of dispersion, we can completely define a distribution
N
only catch is that the data must be normally distributed,
but if it isn’t, there are ways to overcome
this and make the data work for you. As we will D
see in later units, this theorem will form the
basis for hypothesis testing and we will take advantage
of working with data, whether or not it is
R
normally distributed.
A
2
1
6
1
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-8
CHAPTER TWO KNOWLEDGE ASSESSMENT
Probability & the Normal Distribution
Discussion Questions
DISCUSSION QUESTION 1
SAMPLING VS. RELIABILITY: “The larger the sample, the more reliable the results.” Do you
agree or disagree with this statement? Explain.
DISCUSSION QUESTION 2
W
SEEING THROUGH THE CLAIM: An auto manufacturer
advertises that “90% of the cars we’ve
I
ever made are still on the road.” Assuming this L
is literally true, how can it be explained? What
facts / statistics would you need to know to expose this misleading claim?
Practice Problems: Real
L
I
S
,
Estate
Solutions are provided to practice problems so you
K can check your work.
Use the Real_Estate.xls file which consists of 100A
homes purchased in 2007 and appraised in 2008.
S
It includes variables regarding the number of bedrooms,
number of bathrooms, whether the house
has a pool or garage, the age, size and price of the
Shome, what the house is constructed from, and
the appraisals from two agents.
A
N
PRACTICE PROBLEM 1:
D
Create a pivot table of Pool and Garage. Then complete the Joint Probability table so you can
R
answer the following:
a)
b)
c)
d)
A
What is the probability of randomly choosing a home that has no pool?
What is the probability of randomly choosing a home that has a pool AND no garage?
2
What is the probability of randomly choosing a home that has a pool OR a garage?
1
Given that the home you selected has a pool, what is the probability it also has a garage?
6
1
PRACTICE PROBLEM 2:
T entire population. We know the mean and
Let’s assume that the Real_Estate.xls file was the
standard deviation of home sizes to be 2212 sq. ft.Sand 235 sq. ft., respectively. Using the Normal_
Probability.xls file, compute the percentage of homes that are
a) smaller than 2500 sq. ft.?
b) larger than 2300 sq. ft.?
c) between 2000 and 2100 sq. ft.?
d) The largest 20% of homes are greater than what size?
In each case, compare the computed results to the truth as found in the actual data file.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-9
CHAPTER TWO KNOWLEDGE ASSESSMENT
Probability & the Normal Distribution
PRACTICE PROBLEM 3:
Knowing the mean and standard deviation of home sizes to be 2212 sq. ft. and 235 sq. ft., respectively, if we were to visit 4 homes at random every day and compute the mean size, what is the
probability that the mean would be
a) smaller than 2500 sq. ft.?
b) larger than 2300 sq. ft.?
c) between 2000 and 2100 sq. ft.?
d) 20% of the time what would you expect the mean to be above?
W
I
L
Assigned Problems: Student
L Data
I
Use the Student_Data.xls file which consists of 200
S MBA at Whatsamattu U. It includes variables
regarding their age, gender, major, GPA, Bachelors
, GPA, course load, English speaking status,
family, weekly hours spent studying.
K
ASSIGNED PROBLEM 1:
A
Create a pivot table of Gender and Major. Then complete the Joint Probability table so you can
S
answer the following:
S
A
What is the probability of randomly choosing
a Male AND Finance major?
N
What is the probability of randomly choosing
a Female OR Leadership major?
D
Given that the home you selected is a R
Male, what is the probability he has no major?
Given that the home you selected hasAno major, what is the probability the student is
a) What is the probability of randomly choosing a Female?
b)
c)
d)
e)
male?
2
1
ASSIGNED PROBLEM 2:
Let’s assume that the Student_Data.xls file was 6the entire population. We know the mean and
standard deviation of student ages to be 42.3 and 8.9,
1 respectively. Using the Normal_ Probability.
xls file, compute the percentage of students that are
T older than 50, younger than 40, between 41 and
46, and oldest 10% are at what age? Then compare to the truth as found in the actual file.
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
2-10
APPENDIX
PIVOT TABLE TOOL TUTORIAL
In this tutorial, you will learn to utilize the Pivot Table tool built into Excel. This tool allows you to create
W
crosstabulations and to dig into a data file very deeply, grouping the data as you wish, and even analyzing
I
with statistical options and graphs.
L
We will use the Real Estate data from the course, which
L consists of 100 homes. First, open the file and
select the data with your mouse (to row 101). It is imperative that the variable names be in the first row.
I
S
,
K
A
S
S
A
N
D
R
A
2
1
6
1
T
S
Under the INSERT menu, choose the PIVOT TABLE icon.
You will then see an option for PIVOT TABLE and PIVOT
CHART – choose the PIVOT TABLE.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-1
Appendix: Pivot Table Tutorial
The CREATE PIVOT TABLE box pops up with the range already entered (if you highlighted the data first).
It is also defaulted to create a new worksheet. Click OK.
W
I
L
L
I
S
,
You will now see a grid on the left where you can drop the row, column
K the right, you will see a PIVOT TABLE FIELD
and data items. On
LIST based on theAvariables in your data file.
S
S
A
N
D
R
A
2
1
6
1
T
S
Click on CONSTRUCTION and it will automatically populate the ROW
LABELS box. Now click on SIZE and while holding the mouse button
down, drag the label into the VALUES box.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-2
Appendix: Pivot Table Tutorial
You will now see that it shows SUM OF SIZE. This
means that it adds up the sizes, which we do not want.
This is a default with numerical data and is easily
changed. Click on the arrow next to SUM OF SIZE and
then choose VALUE FIELD SETTINGS. This will open
the box below.
W
I
Change to COUNT
L
and click OK.
L
Truthfully it doesn’t matter what variable you used for
I values when you are just doing a COUNT as long
as it is a variable that doesn’t have missing values. You now see a table showing that of the 100 homes, 40
S
were made of Brick, 35 of Stucco and 25 of Wood.
,
Here’s something interesting that you should see and can use at any time in any pivot table. Double click
on cell B7 which shows a 25 in it. Instantaneously K
you get a new spreadsheet showing all variables for
those 25 Wood homes. It is on its own worksheet, so you can save it separately if you wish. For now let’s
A
go back to the Pivot Table.
S
S
A
N
D
R
A
2
1
6
1
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-3
Appendix: Pivot Table Tutorial
Click on the arrow next to COUNT OF SIZE and
then choose VALUE FIELD SETTINGS. This
time when the box pops up, go to the tab labeled
SHOW VALUES AS.
W
I
L
L
I
S
Now click on the arrow next to NORMAL
,
and change to % OF TOTAL. Then click OK.
K
A
S
S
A
N
Note that the Pivot Table now shows the totals in percents. So
D
we see that 40% of the homes are made of Brick, 35% Stucco
R
and 25% Wood.
A
Let’s go back to the VALUE FIELD SETTINGS but this time show the data as NORMAL again and change
from a COUNT to an AVERAGE. Then click OK. 2
1
6
1
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-4
Appendix: Pivot Table Tutorial
Now you can see the average Size for the different constructions.
A Brick home has an average size of 2239.65 square feet, for
instance.
W
I
L
L
I
Let’s add another
S variable to the table. Drag POOL into the
COLUMN LABELS.
,
Now let’s reset it back to a COUNT using the VALUE FIELD SETTINGS.
K
We now have a two-dimensional crosstabulation. Note that 0
A and 1 means POOL. You can actually type
means NO POOL
right over the S
0 and 1 to change these labels (as shown below).
S
A
N
D
R
A
2
1
6
1
T
S
Here we see that 10 of the Brick homes did not have a pool while 30 did. And the Stucco homes
were almost evenly split on having or not having a pool. At this point you can go back to displaying
percentages if you wish, and even choose to show the percents by row, by column or by the overall
total (e.g., if shown by row, you would see 25% of the Brick homes did not have a pool and 75%
did).
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-5
Appendix: Pivot Table Tutorial
Now click on any of the numbers in the Pivot Table. Then click on the COLUMN option on the
toolbar (under INSERT in the CHARTS section). For now just choose the chart in the top row on
the left side (under 2-D Charts).
W
I
L
L
You can instantly convert your Pivot Table into aI useful graph. If you wish to edit it further, you
need only click on the graph. For now, click on the
S outer frame of the chart and hit the DELETE
key.
,
While we needn’t do it here, click on the arrow next to CONSTRUCTION. Notice the check
boxes. If you unchecked Stucco, you would thenKonly be displaying the Brick and Wood homes.
And if you remove the CONSTRUCTION variable
A from the Pivot Table and put it back later,
Stucco would still be omitted until you restore it.S
S
A
Now for one last but important tool / grouping. Often
N we have a situation where we want to recode
data into groups for various reasons and rather than altering the data file, we can do it so easily in
D
a Pivot Table. A common example is to take survey responses and recode (group) Strongly Agree
R
and Agree responses into a category titled POSITIVE,
and other responses into a category titled
A variables, let’s use the SALES PRICE.
NEGATIVE. Since our data set doesn’t have such
Under COLUMN LABELS and ROW LABELS (leave the COUNT OF SIZE alone), click on the
2
arrows next to each line and choose REMOVE FIELD. This will blank out your Pivot Table. You
1 and dragged it out.
could have also grabbed the variable from the table
6
1
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-6
Appendix: Pivot Table Tutorial
Now click on SALEPRICE and drag it into the ROW LABELS.
There are 74 different prices in the 100 homes. Such a display is not very
helpful since it is mostly 1’s with a few 2’s. But we can turn this scale
data into an ordinal variable by putting the data into ranges.
W
I
L
L
I
S
,
K
A
S
S
First highlight cells
A A5 through A31, capturing the 125000 to the
199000. Then right
N click your mouse and choose the GROUP
option.
D
R
A new variable was created called (SALEPRICE2) with a value
A
of Group1 under it. If you click on the Group1, you can rename
it (change it to 250k.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-7
Appendix: Pivot Table Tutorial
And then drag SALEPRICE out of the table (or click on the arrow next to the variable name and
choose REMOVE FIELD). You now have a table with a new variable, which actually appears in
the variable list on the right of the screen. From the table you can see that 41 of the homes sold for
less than $200,000 and 23 sold for over $250,000.
W
I
We can go on further, as this is not the depth of Pivot Tables, but it covers the basics and even some
L
advanced stuff. By playing with it further, you will likely discover even more, but most of what
L
you would do with this wonderful tool has been addressed
here. With little effort you will see that
I
this tool is fun to use and quite amazing.
S
,
K
A
S
S
A
N
D
R
A
2
1
6
1
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
A-8
CHAPTER THREE
CONFIDENCE INTERVALS & SAMPLE SIZES
W
I
Statistics means never having to say you’re
L certain. Or at least that how it feels since we
discuss everything in terms of probabilityLand error. Remember that statistics are all about
using samples to infer the truth about populations.
When we take a sample, we would love
I
to believe that whatever is true about the sample is true about the population, but that will
S
rarely ever happen, as there is almost always some difference (known as a margin of error
or sampling error). The better the sample,is, the smaller the margin of error should be, and
Sampling
so understanding the nature of sampling is critical to understanding statistics.
K
Let’s start with why we sample at all. Wouldn’t
you rather have all of the data in the
A
population and thereby have no sampling error? Of course, but we choose not to for
S
several reasons.
S
1. Cost – it costs money to purchase A
mailing lists or phone lists or even email lists. It
costs even more to hire people to mail
N / phone / email the surveys, and going doorto-door to survey everyone is obscenely expensive. Add to that the fact that phone
D
calls are not free, and mailing surveys involves postage for the survey itself and the
R
return envelope. Then factor in the low return rate on surveys, meaning that you
A people or mailing out second and third rounds
would have to keep calling the same
of surveys. Clearly not an inexpensive venture, which is part of the reason we only
conduct a U.S. census once every 2ten years.
1
2. Time – To gather 400 responses for a phone survey, a marketing research firm will
6 the nonresponses they will likely get. And
often call over 4,000 people, factoring
1 hopes of reaching them adds more time to the
calling people multiple times in the
process. It could take a week for T
a typical sized firm to conduct such a poll. Now
imagine the time involved if theyShad to survey the population, and then proceed
to analyze all of that data which must be loaded into a computer. Our census
takes most of the year to conduct and they have a lot of people working full time
on nothing but the census. So even if money is no object, time is still a major
consideration when polling a large population.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-1
Chapter Three: Confidence Intervals & Sample Sizes
3. Accuracy – While it may seem to defy logic, samples are often more accurate than
populations. With a sample, we accept that it is not perfect and there is a built-in margin of
error. And if we choose a fairly sized representative sample, the error is minimized. Yet,
if we aim to get the entire population, we are expecting perfection. Do you truly believe
that the U.S. Census has the exact average age in the U.S.A., for example? That not one
person failed to receive a form, that everyone completed their forms, that none of the
forms got lost in the mail en route to the Census Bureau, and that all the data was entered
in a database without error? If even one form is missing or one number is mistyped, the
population parameters are wrong. But in entering a sample of 1,000, for example, we
can easily triple-check our work and even contact the respondent to make sure the data
W to report a statistic with a small margin of
is accurately entered. It is more preferable
error we feel confident in than to report Ia parameter with no alleged error that we have
little confidence in. So unless collecting
L the population data involves little more than
downloading the data from the web (e.g., the population of NYSE stock prices on a given
L
day), a well-chosen sample is often more preferable for the sake of accuracy.
I
4. Availability of population – Whether orSnot time and money are on your side doesn’t
matter if the population is not available. While
, some populations may be a captive audience
(e.g., all students attending a grade school – can be surveyed in school over several days
to capture any absentees), others are not quite as easy to reach. Do you really believe
K
the Census captures everyone, including all of the homeless, all of those living in the
mountains away from civilization, all ofAthose visiting abroad, all of those hospitalized,
S zone? Let’s not forget those who are home
and all of those in the military amidst a war
but refuse to answer their phones or openStheir mail. If you cannot reach everyone in the
population, you cannot possibly computeAthe parameter, and if you mistakenly think you
captured the population, you are misrepresenting
your results.
N
Dthis is not an issue with simple surveys, some
5. Destructiveness of observations – While
data collection involves the destruction orRconsumption of the item being studied. To test
A would mean running them until them and
the life of all batteries produced by Duracell
recording the lifespan, and then having nothing left to sell. To taste test cookies would
mean taking a bite and packaging the partially
2 chewed cookie for sale (not very appetizing
for the customer). For many products, it is not practical or advisable to test an entire
1
population, even at the risk of selling a customer a defective product.
6
Now that it is apparent why we choose to sample,1it is important that the sample be a good one or
else the data is meaningless. Selecting a sample T
involves six stages.
S
1. Define the target population – Before proceeding to get any data, you must define the
group that the data is representing. An election poll would be meaningless if the people
surveyed were nonvoters or those from the wrong voting district or even those who are
eligible but not registered voters. As an example, if a study were being done at a company
to see how satisfied employees were with the performance review process, the target
population in question would be only those employees who participated in the process in
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-2
Chapter Three: Confidence Intervals & Sample Sizes
the past year (excluding newer employees and those who haven’t been involved in a recent
performance review for whatever reason). Once we know the population we are targeting,
we can then figure out how best to represent them without surveying all of them.
2. Select a sampling frame – A sampling frame is merely the source from which we get our
sample. This could include a physical list, such as a phone book or mailing list, or it could
be more of a situation, such as stating that you will select every 10th person who shows
up in line (i.e., the sampling frame is being in line). With every sampling frame, there is
likely to be some problems, such as missing people or people that should not be included.
Judging a sampling frame comes down to asking whether it really misrepresents the truth.
Is a residential phone book a poor choice
Wof a sampling frame… it depends! And it is
not because not everyone is in it. If you are
I using a phone book to ask people about their
professions, then it is poor since certain professions,
such as law enforcement and doctors
L
and judges, are often unlisted and the results will be misleading. Yet if you use the same
L
phone book to survey people about whether they prefer smooth or crunchy peanut butter,
the question becomes whether you think Ithat the results of the survey would be different
if unlisted people were included (like doSunlisted people differ significantly from listed
people in their preference for peanut butter?).
, Since it is hardly likely to bias or results, we
would happily use the phone book and assume the results of the survey reflect the entire
population.
K
3. Choose a probability or non-probabilityAsampling method – Samples come in two major
S known as probability and non-probability).
categories, random and non-random (otherwise
A random sample is essentially one in which
S all members of a population have an equal
chance of selection and is representative of
A the population, while a non-random sample is
selected without regard to their probability of occurrence and is only representative of the
N
sample itself. Non-random samples come in many forms and have the advantage of being
D
cheaper and faster than random ones, but are not generalizeable to the population, lack
R for bias:
objectivity and introduce more opportunities
A
•
Convenience sample – this is essentially a mall survey in which people are chosen
out of convenience because they are there. Images of people carrying clipboards by
2
the mall escalator probably come to mind. They don’t care who they survey or that it
1 want warm bodies to complete surveys. While
is a representative sample; they merely
6 they are often a useful option. If you wanted
not typically the best option for a survey,
to know what customers think of a 1
new line of coffee at your restaurant, the most
practical way to get this data is to offer
T samples to those in line, but you don’t care if
it is representative of the population (you
S merely want to know if it is worth selling in
your store).
•
Judgment sample – this is a sample in which experts choose the members. Consider
the Dow Jones Industrial Average in which 30 stocks are chosen to represent the New
York Stock Exchange (NYSE), or the Consumer Price Index (CPI) in which a few select
products are chosen to determine the rate of inflation. Clearly not an ideal sample, but
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-3
Chapter Three: Confidence Intervals & Sample Sizes
because of the expertise used, the sample size can be kept small and the value of the
sample is regularly monitored (note: a mere 30 stocks does a fair job at gauging the
entire NYSE, and it has been in use for over a century).
•
Snowball sample – this is a sample in which it literally grows by snowballing. Think
of all of the political surveys sent out over the web; if the respondent feels it is important
(like asking if you would favor a tax increase), that person would likely forward the
link to many friends to also participate, and they would do likewise. The result would
be that more people respond to it than were actually solicited to respond. And while
we like to see large sample sizes, since the person conducting the survey has no way
of knowing who participated, it wouldW
be inappropriate to try to draw inferences to the
population, and the statistics are not Ias meaningful as they would have been had the
original sample responded and kept itLto themselves.
•
Voluntary sample – this is a sample L
in which respondents throw themselves into the
I is on a website for anyone to answer, there is
study without an invitation. When a poll
no control in place as to who participates
S or how often, and the results are questionable.
Such is the case with the American Idol
, voting too, as they are unsolicited votes from a
voluntary sample, often with repeated votes (but the folks at AI don’t really care about
the best person winning or the one that most viewers want, but that they get a lot of
K
audience participation since they are in the business of making money and selling ads).
A
Random samples can also take many forms
S but they all have a similar goal in mind of
representing the population. Compared to
S the non-random samples, they are more costly
and more time-consuming, but minimize bias and have more generalizeable results.
A
N
• Simple random sample – this is essentially
Lotto, in which all members of the
population are tossed in the mix and D
a sample is chosen, one at a time. Truly a fair
sample with everyone having an equalRchance at selection, but it can be slow to choose
one at a time (unless it is automated onAa computer). There are also some inherent risks
of leaving out a subgroup by random chance (e.g., randomly choosing 10 people from
a class and somehow choosing all men or all women even though there are 50 of each
2 though it is less likely with larger samples).
in the room – can certainly happen, even
•
1
Stratified random sample – To combat
6 the problem of accidentally not representing
a subgroup, you can stratify the data 1into subgroups and then choose randomly from
each. So if a class consisted of 60 men and 40 women, and you wanted to sample
T
10 people, you would randomly choose 6 of the 60 men, and separately choose 4 of
S have a sample that mirrored the population
the 40 women. As a result you would
proportionally and wouldn’t have problems with a group being omitted. You can even
subgroup further, like making sure to get a representative group of Male freshmen,
Female freshmen, Male sophomores, etc.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-4
Chapter Three: Confidence Intervals & Sample Sizes
•
Systematic random sample – Instead of pulling one item at a time from a group, a far
simpler approach to getting the data randomly is to take every nth item. For example, if
you wanted to sample 50 from a group of 100, you could flip a coin and either take the
1st, 3rd, 5th, etc. or the 2nd, 4th, 6th, etc. If you only wanted 25, you could randomly pick
one of them and then take every 4th thereafter. This can be done on the whole group
or each subgroup (so you could stratify your data and pick the 6 men from the 60 by
choosing every 10th and pick the 4 women from the 40 women by choosing every 10th).
Just be careful not to get carried away with too many subgroups or it gets challenging
(such as gender and class year and hair color and major all at once – many subgroups
may only have 1 or 0 in them and you won’t get proportional results but will be forced
W
to choose an entire subgroup).
•
Cluster random sample – This is truly
L the quick-and-dirty method of getting a random
sample. If the data is in clusters of sorts, you can randomly choose the clusters and then
L
all members of those clusters are included in the sample. So if you wanted to conduct
a survey at a university, and wished Ito survey 1,000 students, instead of mailing out
1,000 surveys and hoping for the best S
(at great expense) or picking a couple of students
per class and disrupting almost every, class, you could randomly choose 50 course
sections (assuming an average of 20 per class) and then survey the entire class in each
case. This is cheaper and faster but there
K is a risk of misrepresentation if you choose
very few clusters (e.g., you might miss out on the freshmen classes or the evening
A
classes easily). You could once again combine methods by choosing clusters within
S courses, Y of the evening courses, etc. and
each strata, and choose X of the freshmen
make sure that you get an even better S
cross-section of the university.
I
A
4. Determine the sample size – While the methods used to compute the sample size will
N
be addressed later in this unit, it would help to deal with any misconceptions first. Ask
D
yourself whether you would trust the results of a sample of 500 chosen from a population
of 10,000 students (5% of the population).RNow ask yourself if you would trust the results
of a sample of 500 chosen from 1,000,000Ain your city (0.05%). And how about a sample
of 500 chosen from across the USA’s 300,000,000 (0.00017%). Would you be surprised
to learn that they are EQUALLY GOOD?2 That’s right…the size of the population means
nothing here. If this doesn’t make sense, then think about this. You are making a two1
gallon pot of five alarm chili for your family and then found out that your grandmother was
6
coming over, so you decided to also make a separate one-quart pot of mild chili. Now you
wish to see if they are ready to serve. You1take a spoon and stir the mild chili and taste it…
T many spoonfuls of the other chili would you
mmm good. Now here’s the question … how
need to taste? There is 8 times as much chili
S in the larger pot, and yet one spoonful will tell
you as much about the larger pot as it will about the smaller pot. And the key here is that
you first stirred the pot. As long as the chili is well-mixed, the sample is representative of
the population and you don’t need a large sample to do the job. Most polls are conducted
with samples of 400 – 1,000 and we put a lot of trust in the results.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-5
Chapter Three: Confidence Intervals & Sample Sizes
5. Choose a data collection technique – Data is collected in so many ways, and the method
varies with the budget, audience and type of data being collected. We now collect most
data by phone, mail, email, in person, on the web, and by text messaging. To help in
making the decision as to the best approach, look at the table below which compares and
contrasts several techniques.
Personal
Phone
Mail
E-Mail
High
Medium
High
No
Yes
Yes
High
Low
Yes
Medium
Low
Medium
Maybe
Maybe
Maybe
Medium
Low
No
Low
High
Low
Yes
No
No
Low
High
Maybe
Very Low
Low
Medium
Yes
Yes
No
Low
High
Yes
Costs
Time required
Data quantity per respondent
Reaches widespread sample
Reaches special locations
Interaction with respondents
Degree of interviewer bias
Severity of non-response bias
Presentation of visual stimuli
W
I
L
L
I
S
,
K
6. Select your sample — The stage is set, so now you can choose your sample to minimize
error and maximize representativeness. A
S
S
Confidence Intervals
A
Sampling is the essence of statistics. We normally
N cannot get access to an entire population and
must make do with a sample, but we know that a sample is not exactly the same as the population.
D
There is going to be some difference between the results of a sample and the truth, and if we took
R Ideally, we would like for the results of the
another sample, the results would be different again.
sample to be relatively close to the truth, and forA
us to have confidence in the results. Remember
that the real objective in statistics can best be summarized in the definition of inferential statistics
— using a sample to make inferences about a population.
2
1
When we make inferences, we are estimating. We rarely know the value of a population mean,
6
but we can compute a sample mean. And no matter
how many times you draw a sample, the
1
numbers will vary and the sample mean will not likely equal the population mean exactly. You
can randomly stop at 10 gas stations, and compute
T the mean gas price, but it is doubtful the mean
will equal the population mean; yet if you visit gasSstations in different locations and with different
brands, you will likely come closer to the true mean than if you just picked 10 nearest your home or
10 that were all of the same brand. Considering the sample mean by itself is not accurate and may
be too large or too small, we use an interval to estimate. This “confidence” interval is basically the
sample mean + / – a margin of error (often called a sampling error since it is the difference between
the results from a sample and the actual truth of the population).
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-6
Chapter Three: Confidence Intervals & Sample Sizes
To better understand how this works, imagine being in a war zone in 1942 and watching an enemy
aircraft fly over and drop a bomb from the cockpit, aiming for some target unknown to you (let’s
call him 5 o’clock Charlie, based on a classic MASH episode). Once he hits his target, he leaves
the area and doesn’t return. The bomb hits a Starbucks, but was that the target? Was Charlie the
bombardier so perfect with his accuracy that on his first attempt he hit the target? Or did he have
poor aim and / or wind conditions which caused him to miss by a lot? Or perhaps he barely missed
the target? Without knowing more about Charlie and the conditions, our sample of one bomb leads
us to infer that the target was on Starbucks give or take a large margin of error (perhaps a mile in
any direction). Suppose he came back at 5 o’clock the next day and dropped a single bomb that
hit a Caribou Coffee Shop a half mile East of the Starbucks. Does Charlie have something against
coffee? Was this his actual target which he nowW
right on the second try? Or did he overcorrect
from the first drop and miss by a bunch in the other
I direction? Who knows for sure, but if I had
to guess, I would estimate the target is mid-way between
the two coffee shops give/take a smaller
L
margin of error (perhaps ¾ of a mile) based on the
fact
that
his first two bombs were a half mile
L
apart. This pattern would continue – after 10 bombs, I would estimate the target to be in the center
I
of the 10 drops give/take a margin of error based on the spread of the drops. Basically you would
S and your margin of error would be based on
estimate the location based on the mean of the sample,
, and how sure you wish to be about capturing
the spread of the sample data, the size of the sample,
the true target. And such is how confidence intervals work.
K
Here’s another illustration that might help further. Imagine if you were asked to throw a dart at the
A
bullseye on a dartboard from 12 feet away. If you were given one throw, you probably wouldn’t
S
have much confidence in your ability to hit the target.
If the target were expanded to include the
S grow. If the target were expanded further to
ring around the bullseye, your confidence would
include the entire dartboard, your confidence would
A grow more. Essentially, the larger the margin
of error, the greater your confidence at hitting the N
target. Now what if you were aiming for the ring
around the bullseye again and you were given several
D practice throws. If you hit the target 8 out
of 10 times and you were asked where you expect the next dart to hit, you would probably draw
R
a small target. The expected target is probably much smaller than it was prior to your first throw.
Essentially, the larger the sample size, the smallerAthe margin of error.
Now for the math. We express a confidence interval
by stating that the population mean = the
2
sample mean +/- the sampling error. The sampling error is computed as the z-score (based on the
1
level of confidence) times the standard deviation divided by the square root of the sample size.
6
S.E.= z *1S.D. / √n
T
We are essentially saying that we are X% confident that the true population mean is within the
S
confidence interval.
With a 95% confidence interval, we use a z of 1.96 since 95% of the data in a normal distribution
is within 1.96 standard deviations of the mean. For a 90% confidence interval we use a z-score of
1.645. For a 99% confidence interval we use 2.576. Thus the higher the level of confidence, the
larger the margin of error, and thus the wider the confidence interval.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-7
Chapter Three: Confidence Intervals & Sample Sizes
The sample size is in the denominator of the sampling error, so the larger the sample, the smaller
the error and thus the smaller the confidence interval. This makes sense since more data would
result in having more information about the population and being closer to the truth.
The z-score corresponds to the standard normal distribution. A 95% confidence interval captures
the middle 95% of the normal curve, so the z-score is 1.96 since the area between -1.96 and +1.96
equals 95%. A 90% confidence interval captures a smaller area and has a z-score of 1.645. A 99%
confidence interval captures a larger area and has a z-score of 2.576. Thus, the larger the level of
confidence, the larger the margin of error.
The remaining element in the formula is the standard
W deviation. While you have the freedom to
choose your level of confidence and your sample size, you have no freedom to choose a standard
I
deviation, as it is computed from the data, and it is what it is. Yet, it has an impact. If the data is
L
naturally spread out, the confidence interval will be spread out more, and vice versa. If I teach a
daytime undergraduate class at a university and anLevening MBA class at the same university, and I
asked you to guess the average age of my classesIwithin one year, you would need some data first.
If I told you the ages of 5 randomly chosen students
S from my daytime class was 19, 20, 19, 21,
19, you would have no problem in guessing the mean
, age within one year since the spread of ages
is apparently very slim (and understandably so). Now if I told you the ages of some of my MBA
students was 28, 43, 55, 32, 39, you wouldn’t know where to begin to guess within one year (you
K
probably wouldn’t even feel comfortable guessing within five years). You clearly would need a lot
A
more data to accomplish this task.
S
Confidence intervals can also be computed for proportions.
They are essentially the same in that
S
there is still a confidence level (translated into a z-score) and sample size, but there are no standard
A
deviations. Whenever you see an election poll that reports a margin of error, you are looking at a
N candidate is shown to have 52% of the vote
confidence interval. For example, if the Republican
D X% confident the candidate truly has between
with a 3-point margin of error, it means that we are
49% and 55% of the vote. The 52% is based on aRsample, but the confidence interval implies that
if the entire population were polled, the results would
A likely be in that range. Our certainty that the
truth is in that range is the level of confidence. Notice how often the newspapers report the margin
of error but not the level of confidence!!! Anyway, a range of 49 – 55% means that the candidate
2
might win or might lose, and the race should be labeled as undecided. You shouldn’t focus on
1 clearly a victory or a loss, it is straddling the
the fact that it leans toward the positive; if it is not
line and a decision should not be made. Making6a conclusion based on being close is like stating
1 about it). If the confidence interval were 51 –
a woman is almost pregnant (there is no “almost”
55%, then you could state that the candidate is winning,
and you are X% confident of it.
T
S
Let’s look at how we can solve Confidence Intervals problems using Excel. Suppose a sample
of 25 students at Whatsamatta U. was found to have a mean GMAT score of 550 with a standard
deviation of 100. What is the 95% confidence interval of the mean GMAT score for all students
at that university?
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-8
Chapter Three: Confidence Intervals & Sample Sizes
Figure 1: Screen display from Confidence_Interval.xls
W
I
L
L
I
S
,
Plugging in the data appropriately in the Confidence_Interval.xls file yields a 95% confidence
interval of 508.722 – 591.278 or 550 +/- 41.278.
K This means we are 95% confident that the
population mean GMAT score at Whatsamatta U. is somewhere between 508.722 and 591.278.
A
Now let’s see what happens when we change some of the factors.
Figure 2: Screen display from Confidence_Interval.xls
S
S
A
N
D
R
A
2
1 size? Here you can see that a sample of 50
What if the mean was based on a larger sample
generated a tighter confidence interval, with a width
6 of +/- 27.718.
1
Figure 3: Screen display from Confidence_Interval.xls
T
S
And what if we wished to create a 99% confidence interval instead of a 95% confidence interval?
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-9
Chapter Three: Confidence Intervals & Sample Sizes
The larger level of confidence results in an increase in the sampling error.
So you can clearly see the effects of changes in the confidence level or sample size on the size
of a confidence interval. If by chance there is a limited population size, such as only 1000 in the
entire MBA program, plug that in to the FINITE POPULATION SIZE box and you will see an
even tighter interval, depending on how small the population. Notice how a population of 1000
didn’t really cut the interval size down much, and populations of 10,000 and up will likely show
no difference, but very small populations will cut the interval significantly.
Figure 4: Screen display from Confidence_Interval.xls
W
I
L
L
I
S
,
In Figure 4 above, a population size of 51 was entered, and the sampling error was cut to +/- 5.152;
note that at this point you have a sample which consists
of 50 people out of the 51, so only one is
K
missing, and cannot impact the mean by much. A
Figure 5: Screen display from Confidence_Interval.xls
S
S
A
N
D
R
A
2
Note what happens when we enter a population size of 50, the same as the sample size. The
1 your sample equals the population, so there
sampling error is now zero, which makes sense since
6 the population.
can be no sampling error, as you know the truth about
1
If you should have raw data, you can use the RAW
T DATA tab in the file which will compute the
sample mean, standard deviation and sample size for you (as seen in Figure 6). You still choose
S
the confidence level, and the rest is the same.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-10
Chapter Three: Confidence Intervals & Sample Sizes
Figure 6: Screen display from Confidence_Interval.xls
W
I
L
L
I
S
,
Now let’s look at a proportion problem. If 200 people are polled and 110 reveal they voted for
candidate X, that is 55%.
Figure 7: Screen display from Confidence_Interval.xls
K
A
S
S
A
N
D
R
A
The 95% confidence interval shows a margin of error of .069, meaning that we are 95% confident
that the true proportion is between 48.1% and 61.9%. We would not be able to conclusively
2
declare the candidate a winner since he could be below or above 50%. The election would be
1
undecided at this point.
Figure 8: Screen display from Confidence_Interval.xls
Copyright 2011, Savant Learning SystemsTM
6
1
T
S
Introduction to Statistics by Jim Mirabella
3-11
Chapter Three: Confidence Intervals & Sample Sizes
Now if our sample were twice as large and we had 220 out of 400 declaring their vote for Candidate
X, he still has 55% of the sample vote, but the margin of error is now 4.9%. Thus we are 95%
confident his actual vote count will be somewhere between 50.1% and 59.9%, which means he will
likely win since the entire interval is above 50%.
Sample Size
Probably the most common question often asked of a statistician is “how large of a sample do I
need to take…” To understand the answer to this, it is important to look back a bit. When data
is normally distributed, you can compute probabilities if only you knew the mean and standard
W know the population mean? In that case, the
deviation, and life is good. But what if you don’t
best estimate of a population mean is a sample mean.
I So take a sample and compute the mean and
you have an estimate. But how good is that estimate?
That’s what margins of error are for, based
L
on your level of confidence. But what if you compute
your confidence interval and it is just too
L
large (after all, what good is it to estimate that the presidential candidate has 52% of the vote with
I
a margin of error of 20%)? If the confidence interval is too large, you can always take a larger
S much is enough? Wouldn’t it be nice to know
sample to tighten the confidence interval, but how
exactly how large a sample you need to get the, confidence interval you desire? Thanks to the
sample size formula, that is possible.
K
To compute the minimum sample size, you need to answer two questions: how much error are
A
you willing to tolerate AND how confident do you want to be about the truth being within that
S the larger the sample you’ll need. The greater
margin of error. The more confident you want to be,
S sample you will need. To illustrate, if I gave
the margin of error you can tolerate, the smaller the
you a dart and asked you to hit the bullseye on theAfirst throw, you probably wouldn’t bet on being
successful. If asked how many darts you would need
N to throw before you were 90% confident you
would hit the bullseye, you would request several.
D If asked how many it would take for you to
be 99% confident of hitting the bullseye, you would request even more. The more confident you
R
need to be of hitting the target, the larger the sample you would need. If asked about how close
A you might ask for a lot of room for error. If
you would come to the bullseye with just one dart,
given two darts you would probably feel better at getting closer to the bullseye. The more darts
you throw, the closer you feel you will get to the bullseye.
And so the closer you need to get to the
2
bullseye (i.e., the less error you can tolerate), the1larger the sample size you need.
6
Normally when asked about the tolerance for error and the confidence level, those not in the
1
know is to say they want 0% error and 100% confidence,
but that would require sampling the
T involves incurring a degree of error, but the
entire population. Anything less than the population
question is “how much is acceptable?” Once youS
decide on a margin of error and confidence level,
the sample size is computed, but what if you need to sample 2,000 people? If you do a phone
survey, you may have to call 20,000 people just to get 2,000 responses, but if you cut the sample
size you also cut the confidence and increase the margin of error. It is truly a balancing game, but
you are in control.
With a knowledge of confidence intervals and sample sizes, you now have an insider’s knowledge
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-12
Chapter Three: Confidence Intervals & Sample Sizes
of interpreting survey results, and what a margin of error truly means, and why we prefer the error
to be small, and how larger sample sizes translate to higher levels of confidence with smaller
margins of error. This is often a misused concept, especially by the media, but a little knowledge
of it gives you an advantage over the majority of Americans.
Figure 9: Screen display from Sample_Sizes.xls
W
I
L
L
I
Now computing a minimum sample size is merelySa shortcut to getting the confidence interval you
desire. If you wished to know the mean GMAT score
, at a university within 25 points and with 95%
confidence, according to this computation you would need 62. That means a sample of 62 GMAT
scores would result in a 95% confidence interval with a 25-point sampling error.
Figure 10: Screen display from Sample_Sizes.xls
K
A
S
S
A
N
D
R
A
2
If you were willing to tolerate more error, you wouldn’t
need as large a sample. Here the tolerable
sampling error was doubled and the minimum sample
size
was reduced 75%.
1
Figure 11: Screen display from Sample_Sizes.xls
Copyright 2011, Savant Learning SystemsTM
6
1
T
S
Introduction to Statistics by Jim Mirabella
3-13
Chapter Three: Confidence Intervals & Sample Sizes
If you needed to be more confident in your findings, notice how the minimum sample size increases.
Figure 12: Screen display from Confidence_Interval.xls
W
And so if you sampled 42 GMAT scores, and then
I created a 99% confidence interval of the mean,
you would find it to have a sampling error of +/- L
40 points, which is demonstrated above.
Figure 13: Screen display from Sample_Sizes.xls
L
I
S
,
K
A
S
S
A
N
While procedurally the same as for means, sample
D size computations for proportions are more
commonly performed, especially since polling is
R so popular. Suppose you wanted to conduct
a poll and be 95% confident of capturing the truth within 5%, accordingly you would need to
A
survey 385 people. And as with the means, an increase in tolerable error will decrease the sample
size needed, and vice versa. And an increase in confidence levels will increase the sample size
needed, and vice versa. Notice the ESTIMATE2OF TRUE PROPORTION value of .50; if you
1
have historical data to support what the true proportion
probably is, enter it here. For example, if
you are sampling for defective widgets, and historically
you expect to find 12% defectives, then
6
use .12 for this value; values here range from 0.00
to
1.00,
and the further the value is from .50,
1
the smaller the sample size is. In the absence of any information, use .50 as the default, as it will
T
result in the largest possible sample.
S
As for the sample sizes for means, the results have the same implications. Here, if you take a
sample of 385 people, your 95% confidence interval of the proportion will have a sampling error
of +/- 5%.
So now you have an inside knowledge into the methods used for election polls and what triggers
the decision to declare a state as red, blue or undecided.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-14
CHAPTER THREE KNOWLEDGE ASSESSMENT
Confidence Intervals & Sample Sizes
Discussion Questions
The results of voter surveys and polls guided the political campaigns of many presidents and their
opponents and advised news groups on which political strategies are working. While polls have
an important role in the political process, they are anything but perfect tools for measuring and
forecasting voter behavior. Public opinion often changes daily or weekly (or so the polls report this
to be the case). And even worse, competing polls report contradictory results over the same period
of time. How are you to know which polls to trust,
W if any?
I
Look at http://www.ncpp.org/?q=node/4 (20 Questions a Journalist Should Ask About Poll Results)
L
and http://www.pollingreport.com (Polling Report). Answer the following two questions.
L
I
DISCUSSION QUESTION 1
SIZE VS. VALIDITY: In evaluating a poll, aSlarger sample size would be considered more
favorably, but what consideration would you consider
even important than the sample size before
,
trusting the validity of a poll?
K
DISCUSSION QUESTION 2
A
POLLING RESULTS:Choose any ONE of the polls from Polling Report. Compare the many polls
S differences in results.
used to address a single question and explain possible
S
Please don’t turn this into a heated debate — just A
have fun with it.
N
D
Practice Problems: Real Estate
R
Solutions are provided to practice problems so you
A can check your work.
Use the Real_Estate.xls file which consists of 100 homes purchased in 2007 and appraised in 2008.
2
It includes variables regarding the number of bedrooms, number of bathrooms, whether the house
has a pool or garage, the age, size and price of the1home, what the house is constructed from, how
far it is to the city center, and the appraisals from6two agents.
1
PRACTICE PROBLEM 1: You wish to know the
T average sale price of a home in 2007. Compute
the 95% confidence interval of the mean using the sample of 100 homes. Describe your findings.
S
PRACTICE PROBLEM 2: You wish to know the proportion of homes with a swimming pool
in 2007. Compute the 95% confidence interval of the proportion using the sample of 100 homes.
Describe your findings.
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-15
CHAPTER THREE KNOWLEDGE ASSESSMENT
Confidence Intervals & Sample Sizes
PRACTICE PROBLEM 3: You wish to learn the average appraised price of a home in 2008
within $5000, and with 95% confidence, assuming a standard deviation of $43,250. How large of
a sample should you get? (note: the standard deviation shown here was computed from the 200
appraisals in the sample; unless historical data is provided, the best option is to take an initial
sample of at least 30 to get an estimated standard deviation).
PRACTICE PROBLEM 4: You wish to learn the proportion of homes with a garage within
3%, and with 90% confidence. How large of a sample should you get? (note: the estimate of the
W prior knowledge or an initial sample. Since
true proportion is defaulted to 0.50 unless you have
there is a sample of 100 in the data file and 66 ofI the 100 have a garage, you can use 0.66 as the
estimate. standard deviation shown here was computed
from the 200 appraisals in the sample;
L
unless historical data is provided, the best optionLis to take an initial sample of at least 30 to get
an estimated standard deviation).
Assigned Problems:
I
S
Student
,
Data
Use the Student_Data.xls file which consists of 200
K MBA at Whatsamattu U. It includes variables
regarding their age, gender, major, GPA, Bachelors GPA, course load, English speaking status,
A
family, weekly hours spent studying.
S
S average GPA of MBA students at Whatsamatta
ASSIGNED PROBLEM 1: You wish to know the
U. Compute the 95% confidence interval of the mean
A using the sample of 200 students. Describe
your findings.
N
D proportion of MBA students that are majoring
ASSIGNED PROBLEM 2: You wish to know the
R of the proportion using the sample of 200
in Finance. Compute the 95% confidence interval
A
students. Describe your findings.
ASSIGNED PROBLEM 3: You wish to learn the average age of an MBA student within 2 years
2
and with 99% confidence. How large of a sample should you get?
1
ASSIGNED PROBLEM 4: You wish to learn the
6 proportion of MBA students that are female
within 3%, and with 98% confidence. How large1of a sample should you get?
T
S
Copyright 2011, Savant Learning SystemsTM
Introduction to Statistics by Jim Mirabella
3-16
Essay Writing Service Features
Our Experience
No matter how complex your assignment is, we can find the right professional for your specific task. Achiever Papers is an essay writing company that hires only the smartest minds to help you with your projects. Our expertise allows us to provide students with high-quality academic writing, editing & proofreading services.Free Features
Free revision policy
$10Free bibliography & reference
$8Free title page
$8Free formatting
$8How Our Dissertation Writing Service Works
First, you will need to complete an order form. It's not difficult but, if anything is unclear, you may always chat with us so that we can guide you through it. On the order form, you will need to include some basic information concerning your order: subject, topic, number of pages, etc. We also encourage our clients to upload any relevant information or sources that will help.
Complete the order form
Once we have all the information and instructions that we need, we select the most suitable writer for your assignment. While everything seems to be clear, the writer, who has complete knowledge of the subject, may need clarification from you. It is at that point that you would receive a call or email from us.
Writer’s assignment
As soon as the writer has finished, it will be delivered both to the website and to your email address so that you will not miss it. If your deadline is close at hand, we will place a call to you to make sure that you receive the paper on time.
Completing the order and download