DEMO Portfolio PSY7065Group Comparisons and Correlations
Important note: this portfolio would not achieve the maximal score. More depth could have
been provided when displaying understanding and graphical data summaries could have
been inserted.
Colour coding:
Basic Study Description
Displaying understanding
Testing assumptions
Reporting
Study 1
This study examined the attitude towards ‘killer clowns‘, a relatively recent phenomenon of
people dressing up as scary clowns and harassing/startling innocent people. An experiment
was conducted to examine whether men and women differ in their attitudes towards these
‘killer clowns’. The experimenter recruited 20 males and 20 females to express their attitude
towards ‘killer clowns’ in two ways: using a 7-point Likert scale and using a fictive interview
scenario (setting up chairs in a room in preparation for an interview with a killer clown; the
closer the chair were placed, the more positive the attitude).
Given that this assessment concerns Gender, this represents a between-groups or
independent-samples design. A single item Likert scale is ordinal measurement scale (i.e.,
equal intervals cannot be assumed) and should not be expected to yield normally
distributed data; indeed, the Shapiro-Wilk test showed that for both men and women the
data distribution differed significantly from a normal distribution (Men: p = 0.003; Women:
p = 0.003). Non-parametric statistics were thus required, in this case the Mann-Whitney U
test. For the Likert scale measurement of attitude towards ‘killer clowns’ there was no
significant effect of Gender, U = 240, p = 0.289. Chair distance measure, however, does have
equal intervals and even a meaningful 0, which means it is ratio data. This, in combination
with the fact that the chair distances were normally distributed (Shapiro-Wilk tests: Men: p
= 0.306; Women: p = 0.228), means an independent-samples t-test is preferred. It also
confirmed that the two groups had equal variances, or Homogeneity (Levene’s test: p =
0.502; Hanna & Dempster, 2012). For chair distance, there was a large, significant effect of
Gender, t(38) = 2.265, p = 0.029, d = 0.72. Males selected a larger distance between chairs
(Mean = 0.93m; SD = 0.40m) than females (Mean = 0.67m; SD = 0.34m), suggesting males
had a more negative attitude towards killer clowns.
To assess the relationship between the Likert-scale and chair distance measures of
attitude, a correlational analysis is appropriate. Pearson’s correlation would be preferred
due to its higher statistical power than its non-parametric equivalents (i.e., it is better at
flagging relationships as significant (the test yielding a p-value below the alpha-level) in
cases where there is indeed a relationship). However, the assumptions underlying this test
are not met (i.e., the Likert scale data is nominal, but should be interval or ratio, and it is not
normally distributed, as mentioned above; it therefore also is not needed anymore to check
for outliers). It is thus appropriate to use a non-parametric, rank-based correlation, such as
Spearman’s correlation or Kendall’s correlation. The latter is preferred if there are tied
ranks, which commonly happens for Likert scales (Field, 2013; Hanna & Dempster, 2012);
this correlation measure is thus used here. There was a large negative significant correlation
between the attitude measures, ρ(38) = -0.66, p < 0.001; higher scores on the anxiety
measure were associated with smaller chair distances. Thus, while both measures seemed
to tap into a similar aspect (likely attitude), only the chair distance measure highlighted a
Gender difference in attitudes towards killer clowns. This is likely at least partly associated
with the higher power of parametric tests. Interval measures are more fine grained and can
thus be used to detect smaller differences than ordinal/rank-based measures (i.e., when
ranking interval data, information about the size of the differences between scores is lost)
(Field, 2013; Gravetter & Walnau, 2013; Hanna & Dempster, 2012; Norman & Streiner,
2008). This is further illustrated for these data by the fact that a Mann-Whitney U test for
the chair distances also failed to reach significance, U = 132, p = 0.07 (while the
independent-samples t-test did reveal significance). Since the t-test has a higher power than
this test, we base our conclusions on the independent-samples t-test.
Study 2
This study examined whether knowledge of physics influences performance in a task that
could depend on physical principles. An experiment was conducted to examine this:
students who were about to start a basic physics module were recruited for a netball aiming
task. They were again asked to complete this task after they had finished their module.
Aiming performance was measured through the distance the ball was from the centre of the
bucket as it passed the top of the bucket, averaged across 100 throws. The effect of release
height on aiming performance was also assessed.
Since participants contributed scores to both conditions (Pre- and Post-module), this
study involves a within-subjects or paired-samples design. The average distance measure
has equal intervals, and the Pre-Post difference is normally distributed (as shown by a
Shapiro-Wilk test, Distance Pre-Post: p = 0.641). This means that a paired-samples t-test is
preferred over a non-parametric variance (e.g., Wilcoxon Signed-Rank test). This test
showed a large, significant difference in aiming performance between the tests before and
after the physics module, t(24) = 3.42, p = 0.002, d = 0.68. Importantly, the distance was
smaller after the physics module (Mean = 0.56m, SD = 0.29m) than before the physics
module (Mean = 0.97, SD = 0.41), suggesting that knowledge of physics indeed improved the
performance of the participants. One caveat to mention is that there could of course be
other factors that result in improved aiming performance during an extended period of
time. Some participants may for instance have joined a sports club.
A second aspect to-be-tested was the relationship between aiming performance and
the height at which the ball was released. This requires a correlational analysis. Since the
distance scores as well as the release height scores are ratio (i.e., have equal intervals and a
meaningful 0), are all normally distributed (Shapiro-Wilk tests, Distance, Pre: p = 0.231,
Distance, Post: Post: p = 0.827, Release Height, Pre: p = 0.195, Release Height, Post: Post: p =
0.744), and contain no outliers (as checked using scatter plots), conducting a Pearson
correlation is most adequate (two correlations were calculated, so a Bonferroni adjusted
alpha-level of 0.025 was employed). Before the physics module, there was a strong,
significant correlation between release height and aiming performance, r(23) = 0.54, p =
0.006. A similar strong, significant correlation between release height and aiming
performance was found after the physics module, r(23) = 0.50, p = 0.011. In both cases,
higher release heights were associated with higher distances, that is, worse aiming
performance. This finding is opposite to the expectations.
References
Field, A. (2013). Discovering statistics using IBM SPSS statistics (and sex and drugs and rock 'n' roll)
(4th Ed). London: Sage.
Gravetter, F. J. & Wallnau, L. B. (2013). Statistics for the behavioural sciences, International Edition
(9th Edition). Hampshire, UK: Cengage.
Hanna, D., & Dempster, M. (2012) Psychology Statistics For Dummies. New York: John Wiley & Sons.
Norman, G. R. & Streiner, D. L. (2008). Biostatistics: The Bare Essentials, (3rd Ed.). Hamilton: BC
Decker.
DEMO Portfolio PSY7065
Group Comparisons and Correlations
Important note: this portfolio would not achieve the maximal score. More depth could have
been provided when displaying understanding and graphical data summaries could have
been inserted.
Colour coding:
Basic Study Description
Displaying understanding
Testing assumptions
Reporting
Study 1
This study examined the attitude towards ‘killer clowns‘, a relatively recent phenomenon of
people dressing up as scary clowns and harassing/startling innocent people. An experiment
was conducted to examine whether men and women differ in their attitudes towards these
‘killer clowns’. The experimenter recruited 20 males and 20 females to express their attitude
towards ‘killer clowns’ in two ways: using a 7-point Likert scale and using a fictive interview
scenario (setting up chairs in a room in preparation for an interview with a killer clown; the
closer the chair were placed, the more positive the attitude).
Given that this assessment concerns Gender, this represents a between-groups or
independent-samples design. A single item Likert scale is ordinal measurement scale (i.e.,
equal intervals cannot be assumed) and should not be expected to yield normally
distributed data; indeed, the Shapiro-Wilk test showed that for both men and women the
data distribution differed significantly from a normal distribution (Men: p = 0.003; Women:
p = 0.003). Non-parametric statistics were thus required, in this case the Mann-Whitney U
test. For the Likert scale measurement of attitude towards ‘killer clowns’ there was no
significant effect of Gender, U = 240, p = 0.289. Chair distance measure, however, does have
equal intervals and even a meaningful 0, which means it is ratio data. This, in combination
with the fact that the chair distances were normally distributed (Shapiro-Wilk tests: Men: p
= 0.306; Women: p = 0.228), means an independent-samples t-test is preferred. It also
confirmed that the two groups had equal variances, or Homogeneity (Levene’s test: p =
0.502; Hanna & Dempster, 2012). For chair distance, there was a large, significant effect of
Gender, t(38) = 2.265, p = 0.029, d = 0.72. Males selected a larger distance between chairs
(Mean = 0.93m; SD = 0.40m) than females (Mean = 0.67m; SD = 0.34m), suggesting males
had a more negative attitude towards killer clowns.
To assess the relationship between the Likert-scale and chair distance measures of
attitude, a correlational analysis is appropriate. Pearson’s correlation would be preferred
due to its higher statistical power than its non-parametric equivalents (i.e., it is better at
flagging relationships as significant (the test yielding a p-value below the alpha-level) in
cases where there is indeed a relationship). However, the assumptions underlying this test
are not met (i.e., the Likert scale data is nominal, but should be interval or ratio, and it is not
normally distributed, as mentioned above; it therefore also is not needed anymore to check
for outliers). It is thus appropriate to use a non-parametric, rank-based correlation, such as
Spearman’s correlation or Kendall’s correlation. The latter is preferred if there are tied
ranks, which commonly happens for Likert scales (Field, 2013; Hanna & Dempster, 2012);
this correlation measure is thus used here. There was a large negative significant correlation
between the attitude measures, ρ(38) = -0.66, p < 0.001; higher scores on the anxiety
measure were associated with smaller chair distances. Thus, while both measures seemed
to tap into a similar aspect (likely attitude), only the chair distance measure highlighted a
Gender difference in attitudes towards killer clowns. This is likely at least partly associated
with the higher power of parametric tests. Interval measures are more fine grained and can
thus be used to detect smaller differences than ordinal/rank-based measures (i.e., when
ranking interval data, information about the size of the differences between scores is lost)
(Field, 2013; Gravetter & Walnau, 2013; Hanna & Dempster, 2012; Norman & Streiner,
2008). This is further illustrated for these data by the fact that a Mann-Whitney U test for
the chair distances also failed to reach significance, U = 132, p = 0.07 (while the
independent-samples t-test did reveal significance). Since the t-test has a higher power than
this test, we base our conclusions on the independent-samples t-test.
Study 2
This study examined whether knowledge of physics influences performance in a task that
could depend on physical principles. An experiment was conducted to examine this:
students who were about to start a basic physics module were recruited for a netball aiming
task. They were again asked to complete this task after they had finished their module.
Aiming performance was measured through the distance the ball was from the centre of the
bucket as it passed the top of the bucket, averaged across 100 throws. The effect of release
height on aiming performance was also assessed.
Since participants contributed scores to both conditions (Pre- and Post-module), this
study involves a within-subjects or paired-samples design. The average distance measure
has equal intervals, and the Pre-Post difference is normally distributed (as shown by a
Shapiro-Wilk test, Distance Pre-Post: p = 0.641). This means that a paired-samples t-test is
preferred over a non-parametric variance (e.g., Wilcoxon Signed-Rank test). This test
showed a large, significant difference in aiming performance between the tests before and
after the physics module, t(24) = 3.42, p = 0.002, d = 0.68. Importantly, the distance was
smaller after the physics module (Mean = 0.56m, SD = 0.29m) than before the physics
module (Mean = 0.97, SD = 0.41), suggesting that knowledge of physics indeed improved the
performance of the participants. One caveat to mention is that there could of course be
other factors that result in improved aiming performance during an extended period of
time. Some participants may for instance have joined a sports club.
A second aspect to-be-tested was the relationship between aiming performance and
the height at which the ball was released. This requires a correlational analysis. Since the
distance scores as well as the release height scores are ratio (i.e., have equal intervals and a
meaningful 0), are all normally distributed (Shapiro-Wilk tests, Distance, Pre: p = 0.231,
Distance, Post: Post: p = 0.827, Release Height, Pre: p = 0.195, Release Height, Post: Post: p =
0.744), and contain no outliers (as checked using scatter plots), conducting a Pearson
correlation is most adequate (two correlations were calculated, so a Bonferroni adjusted
alpha-level of 0.025 was employed). Before the physics module, there was a strong,
significant correlation between release height and aiming performance, r(23) = 0.54, p =
0.006. A similar strong, significant correlation between release height and aiming
performance was found after the physics module, r(23) = 0.50, p = 0.011. In both cases,
higher release heights were associated with higher distances, that is, worse aiming
performance. This finding is opposite to the expectations.
References
Field, A. (2013). Discovering statistics using IBM SPSS statistics (and sex and drugs and rock 'n' roll)
(4th Ed). London: Sage.
Gravetter, F. J. & Wallnau, L. B. (2013). Statistics for the behavioural sciences, International Edition
(9th Edition). Hampshire, UK: Cengage.
Hanna, D., & Dempster, M. (2012) Psychology Statistics For Dummies. New York: John Wiley & Sons.
Norman, G. R. & Streiner, D. L. (2008). Biostatistics: The Bare Essentials, (3rd Ed.). Hamilton: BC
Decker.
DEMO Coursework (NOT TO BE SUBMITTED!!!)
Study 1
A researcher was interested in Gender differences in attitudes towards ‘killer clowns’. She
thus obtained measures of this attitude using a 7-point Likert scale for 20 males and 20
females (higher means more positive attitude). However, because the researcher wondered
whether the single 7-point Likert scale was the most appropriate measure for attitude, she
devised an alternative test. This test required participants to prepare a room for an
interview with a ‘killer clown’, by setting up two chairs. The selected distance between the
chairs was taken as a measure of attitude (lower distance suggesting a more positive
attitude). The researcher asked the same 40 participants to complete this test. The data for
this study is provided in the accompanying SPSS data file:
PSY7065_DEMO_Portfolio_Study1.sav
1. Assess whether, according to both measures, there is a Gender difference in the
attitude towards killer clowns.
2. Assess the relationship between the two measures.
Study 2
A researcher studied whether knowledge of physics influences aiming in netball. He thus
recruited participants among students of a basic physics module and measured their aiming
performance prior to and directly after completing a module; aiming performance was
quantified as the average distance the ball was from the centre of the bucket (as it passed
the top of the bucket) across 100 throws (lower distance means better aiming). The
researcher also hypothesized that aiming performance is better when balls are released
higher. He thus measured the release height as well (in metres, relative to the ground). The
data for this study is provided in the accompanying SPSS data file:
PSY7065_DEMO_Portfolio_Study2.sav
1. Assess whether aiming improves with obtaining physics knowledge.
2. Assess whether release height relates to aiming performance.
PSY7065
Quantitative Data Analysis
Multiple Linear Regression
Gary McKeown
1
Overview
• Multiple linear regression (the basics)
• Multiple linear regression (advanced)
• SPSS practice in the Lab
2
Confidence Intervals
• A single number is a point estimate
– a midpoint (e.g. mean, median, beta values - today)
• We can also have interval estimates
– upper and lower limits
• How often, in the long run, does a given interval from
a sample contain the true (population) value of the
parameter
• Typically 95% confidence intervals
• Like a complement to the standard error
3
Confidence Intervals
• Work out the upper and lower
limits
• It is within limits we are interested
in this time not outside the limit
• As we are interested in sampling
we use SE and the mean to find the
limits.
• There are two equations
Confidence Intervals
• This is a long run probability so for any given confidence
interval the population mean (or other parameter) is
either within (probability of 1) or not within (probability of
0) the limits
• Therefore this does not mean there is a 95% probability that
the population mean (or other parameter) lies within the
limits
• It means if we take 100 samples from the population approx.
95 of those samples will contain the population mean (or
other parameter)
Confidence Intervals
https://shiny.rit.albany.edu/stat/confidence/
• Click the “Confidence Interval Graph Plus Sampling
Distribution of the Mean” radio button
• Change the sample size and see what happens
– remember larger sample sizes lead to smaller
standard errors
– Watch the sampling distribution of the mean
graph beneath too
• Put the sample size to 100 and repeatedly press
the resample button repeatedly — count the
number of red confidence intervals these are
where the mean is not within the limits
6
Quick recap: correlation
• Capturing relationship between two random variables
• Variables typically measured
• Any departure of non-dependence qualifies as
correlation
– Linearity not a necessity for the general concept, but a requirement for
the most frequently used correlation measure (Pearson’s)
• Correlation may be positive or negative
7
Correlation may indicate causation
• Causal relationships: A causally leads to B -> A
and B may be correlated as a result
– Example: practice makes perfect: amount of hours spent
practicing a skill is correlated with skill level
• However, this cannot be turned around:
correlation does not imply causation!
• Experimental design is needed to say
something about causality – not the statistics
http://www.tylervigen.com/spurious-correlations
8
From correlation to regression
• Concepts very strongly related
• Practical difference:
– Correlation: relationship between two random variables (typically
both measured), DV-DV
– Regression: relationship between random and fixed variable, DV-IV
• Example: practice makes perfect
– Measure skill levels at set times: apply regression to skill level as a
function of time
• Linear regression: model relationship between
explanatory variable (IV) and response variable (DV)
by fitting a linear equation/function
9
Some terminology
• Linear regression captures (linear) relationship
between measured variable and controlled variable.
Outcome variable
Dependent variable
– Measured variable = ‘dependent variable’ = ‘outcome
variable’
– Controlled variable = ‘independent variable’ = ‘predictor
variable’
Predictor variable
Independent variable
10
Think about
𝑦 = 𝑚𝑥 + 𝑐
11
Think about – Matrices
12
What is linear about linear regression?
When one variable is linearly related to
another variable, one can describe relationship
with what is called a ‘linear function’:
𝑦 = 𝑚𝑥 + 𝑐
𝑐 = constant/intercept/baseline (i.e., 𝑦 at 𝑥 = 0)
𝑚 = slope/regression coefficient
13
The model is a line
All of these capture the same relationship
c or b0 = intercept/baseline (i.e., DV at IV = 0)
m or b1 = slope/regression coefficient
error = residuals = part of data not captured by linear relation = the stuff left
over after the model has tried its best to explain the relationship
Residuals
– Residuals are the error, the bit left over after the model
15
Example
A participant trains to improve a skill and completes
a skill test multiple times prior to training. A
researcher predicts that skill score should increase
linearly (at least over days considered)
𝑆𝑘𝑖𝑙𝑙𝑠𝑐𝑜𝑟𝑒 = 𝑏0 + 𝑏1 × 𝐷𝑎𝑦 + 𝑒𝑟𝑟𝑜𝑟
b0: baseline Skill score (at Day 0)
b1: measure of learning rate (change in skill score
per day)
error: deviation scores ( = variation in performance
around regression line)
16
Skill score = Intercept + b1 ∙ Day
Different intercepts (b0) = different baseline skill levels
Different
slopes (b1)
=
Different
regression
coefficients
=
different
learning rates
17
Skill score = Intercept + b1 ∙ Day + error
Different intercepts (b0) = different baseline skill levels
Different
slopes (b1)
=
Different
regression
coefficients
=
different
learning rates
18
Skill score = Intercept + b1 ∙ Day + error
9 Datasets
Ideal line (population)
Data
Best fitting line (for data)
How well each
sample/dataset matches
population depends on
amount of error (=random
variation) and number of
data points.
19
Skill score = Intercept + b1 ∙ Day + error
9 Datasets
Ideal line (population)
Data
Best fitting line (for data)
9 sets of randomly selected values
→minor variations in ‘fitted’ line
(the linear equation), that is, in its
associated regression parameters
(intercept/slope)
20
Find a Fit
• https://shinyapps.org/showapp.php?app=https://tellmi.psy.
lmu.de/felix/lmfit&by=Felix%20Schönbrodt&title=Find-afit!&shorttitle=Find-a-fit!
• Be your own Ordinary Least Squares algorithm
•
(c) by Felix Schönbrodt
21
Hypothesis testing with linear regression
• There are a few different hypothesis test that occur
within linear models
– The hypothesis test of the model – the goodness of fit
– The hypothesis tests of the predictors and the
intercept
• The model hypothesis tests against the null that there is
no relationship
• In correlation
+1 = perfect positive relationship
– 1 = perfect negative relationship
0 = no relationship
22
Hypothesis testing with linear regression
Testing the goodness of fit
0 = no relationship is very important
This is the null hypothesis
no relationship is a flat line
This is just the intercept
The null model is then:
𝑅𝑒𝑠𝑝𝑜𝑛𝑠𝑒 = (𝑏0 ) + 𝑒𝑟𝑟𝑜𝑟
𝑦 = (𝑏0 ) + 𝜖
𝑦 = (𝑏0 ) + 0 × 𝑥 + 𝜖
Often called the baseline model
Field (2017)
23
Total Sum of Squares
SST
The sum of the squares
of the baseline model
when there is no
relationship
between x and y
“how good the mean is
as a model of the
observed outcome
scores”
Field (2017)
24
Total Sum of Squares
SST
The sum of the squares
of the baseline model
when there is no
relationship
between x and y
“how good the mean is
as a model of the
observed outcome
scores”
Field (2017)
25
Mean Residuals
Squared
-5.78
33.4084
-4.88
23.8144
7.02
49.2804
-6.98
48.7204
-11.28
127.2384
-5.08
25.8064
-17.88
319.6944
-4.48
20.0704
-9.18
84.2724
3.42
11.6964
-10.18
103.6324
-10.68
114.0624
1.22
1.4884
-2.48
6.1504
-6.98
48.7204
3.12
9.7344
7.02
49.2804
7.62
58.0644
3.02
9.1204
5.02
25.2004
15.02
225.6004
0.92
0.8464
16.32
266.3424
13.62
185.5044
12.52
156.7504
sum =
2004.5
25
𝑆𝑆𝑇𝑜𝑡𝑎𝑙 = ∑ 𝑚𝑒𝑎𝑛𝑟𝑒𝑠𝑖𝑑𝑖2 = 2004.5
𝑖=1
26
Residual Sum of
Squares
SSR
The sum of the squares
of the residuals once
the model has been
fitted using Ordinary
Least Squares
(OLS) – typically
“represents the degree
of inaccuracy when the
best model is fitted to
the data”
Field (2017)
27
Residual Sum of
Squares
SSR
The sum of the squares
of the residuals once
the model has been
fitted using Ordinary
Least Squares
(OLS) – typically
“represents the degree
of inaccuracy when the
best model is fitted to
the data”
Field (2017)
28
Residuals
Squared
4.68
21.86
4.70
22.13
15.73
247.53
0.86
0.74
-4.31
18.57
1.02
1.04
-12.65
160.08
-0.12
0.02
-5.69
32.43
6.03
36.41
-8.44
71.19
-9.81
96.21
1.22
1.49
-3.35
11.23
-8.72
76.08
0.51
0.26
3.53
12.49
3.26
10.65
-2.21
4.87
-1.08
1.16
8.05
64.80
-6.92
47.91
7.61
57.87
4.04
16.29
2.06
4.26
sum =
1017.57
25
𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙 = ∑ 𝑟𝑒𝑠𝑖𝑑𝑖2 = 1017.57
𝑖=1
29
Model Sum of Squares
SSM
The sum of the squares
of the differences
between the mean
value of y and the
model
(Sometimes called
regression or explained
sum of squares)
“This difference shows
us the reduction in the
inaccuracy of the model
resulting from fitting
the regression model to
the data”
Field (2017)
30
Model Sum of Squares
SSM
The sum of the squares
of the differences
between the mean
value of y and the
model
(Sometimes called
regression or explained
sum of squares)
“This difference shows
us the reduction in the
inaccuracy of the model
resulting from fitting
the regression model to
the data”
Field (2017)
31
Improvement
Squared
-10.46
109.32
-9.58
91.86
-8.71
75.92
-7.84
61.49
-6.97
48.59
-6.10
37.20
-5.23
27.33
-4.36
18.98
-3.49
12.15
-2.61
6.83
-1.74
3.04
-0.87
0.76
0.00
0.00
0.87
0.76
1.74
3.04
2.61
6.83
3.49
12.15
4.36
18.98
5.23
27.33
6.10
37.20
6.97
48.59
7.84
61.49
8.71
75.92
9.58
91.86
10.46
109.32
sum =
986.93
25
𝑆𝑆𝑀𝑜𝑑𝑒𝑙 = ∑ 𝑖𝑚𝑝𝑟𝑜𝑣𝑒𝑖2 = 986.93
𝑖=1
32
Goodness of fit – 𝑅
2
• 𝑅2 – the coefficient of determination
𝑅
2
𝑆𝑆𝑀𝑜𝑑𝑒𝑙
𝑆𝑆𝑀
2
=
=𝑅 =
𝑆𝑆𝑇𝑜𝑡𝑎𝑙
𝑆𝑆𝑇
2
𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙
𝑆𝑆𝑅
2
=1−
=𝑅 =1−
𝑆𝑆𝑇𝑜𝑡𝑎𝑙
𝑆𝑆𝑇
or 𝑅
• 𝑅2 tells you the proportion of variance explained by the
model.
• Multiply by 100 to get the percentage
• In a simple linear regression this is the correlation r squared
33
Goodness of fit – 𝑅
2
• 𝑅2 – the coefficient of determination
𝑅
2
𝑆𝑆𝑀𝑜𝑑𝑒𝑙
𝑆𝑆𝑀
986.93
2
=
=𝑅 =
=
= 0.492
𝑆𝑆𝑇𝑜𝑡𝑎𝑙
𝑆𝑆𝑇
2004.5
𝑆𝑆
1017.57
or 𝑅2 = 1 − 𝑆𝑆𝑆𝑆𝑅𝑒𝑠𝑖𝑑𝑢𝑎𝑙
=
𝑅2 = 1 − 𝑅 = 1 −
= 1 − 0.508 = 0.492
𝑆𝑆𝑇
2004.5
𝑇𝑜𝑡𝑎𝑙
• 𝑅2 tells you the proportion of variance explained by the
model.
• Multiply by 100 to get the percentage
• In a simple linear regression this is the correlation r squared
34
Goodness of fit – R2
• How do we know if this is a significant difference
• Do an F-test
• Calculate the Mean Squares to compute
• One for each of SSModel and SSResiduals
• Divide them by the degrees of freedom
• F has a probability distribution we can use to get a p-value
Goodness of fit – R2
• How do we know if this is a significant difference
• Do an F-test
• Calculate the Mean Squares to compute
• One for each of SSModel and SSResiduals
• Divide them by the degrees of freedom
• F has a probability distribution we can use to get a p-value
Improvement
Step
-10.46
-9.58
0.87
-8.71
0.87
-7.84
0.87
-6.97
0.87
-6.10
0.87
-5.23
0.87
-4.36
0.87
-3.49
0.87
-2.61
0.87
-1.74
0.87
-0.87
0.87
0.00
0.87
0.87
0.87
1.74
0.87
2.61
0.87
3.49
0.87
4.36
0.87
5.23
0.87
6.10
0.87
6.97
0.87
7.84
0.87
8.71
Coefficients
𝑦𝑖 = (𝑏0 + 𝑏1 𝑥𝑖 ) + 𝜖𝑖
𝑦𝑖 = (𝑏0 + 𝑏1 𝑥𝑖 ) + 𝜖𝑖
0.87
𝑦5 = (𝑏0 + 𝑏1 𝑥5 ) + 𝜖5
𝑦7 = (𝑏0 + 𝑏1 𝑥7 ) + 𝜖7
9.58
0.87
𝑦7 = 45.053 + (.87 × 7) + −12.65 = 38.493
𝑦5 = 45.053 + (.87 × 5) + −4.31 = 45.013
10.46
0.87
37
Coefficient
• We interpret the coefficient as:
For every increase of 1 unit in the x variable
y increases by the coefficient
• In our model for every 1 increase in Time, Skill increase by 0.87
• If we continue our line to point where it crosses the y axis that
is the intercept, here 45.053 (b0 or c)
• The model part is equal to:
𝑦
𝑆𝑘𝑖𝑙𝑙
= 𝑚𝑥 + 𝑐
= 0.87 × 𝑇𝑖𝑚𝑒 + 45.053
• Add on the error/residuals to get back to our observed data
38
Hypothesis testing with linear regression
Testing individual predictors
0 = no relationship is very important again this is the null hypothesis for
each predictor
is each individual predictor a flat line = 0
or is it significantly different from 0
Tested using the t-statistic – how big b1 is compared to the standard error
𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 − 𝑏𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑
𝑡=
=
𝑆𝐸𝑏
𝑆𝐸𝑏
𝑏𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 is 0 as the expected here is the null hypothesis and it drops out of the equation
Field (2017)
39
Hypothesis testing with linear regression
Testing individual predictors
0 = no relationship is very important again this is the null hypothesis for
each predictor
is each individual predictor a flat line = 0
or is it significantly different from 0
Tested using the t-statistic – how big b1 is compared to the standard error
𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 − 𝑏𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 .871
𝑡=
=
=
= 4.723
𝑆𝐸𝑏
𝑆𝐸𝑏
.184
𝑏𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 is 0 as the expected here is the null hypothesis and it drops out of the equation
Field (2017)
40
Example
• Let’s see what this looks like for the data in SPSS
• PSY7065_LinearRegression_example1.sav
Key basic outputs
• ANOVA table
• F-value/p-value/dfs for null-hypothesis significance testing
• Basically: does the full regression model account for more variance than expected of
a model with regression predictor coefficients of 0 (i.e., DV = intercept + b1*IV versus
DV = intercept + 0*IV)
• Coefficients table
• Regression coefficients (B), their standard error, standardized coefficients, and
associated t-value and p-value
• t-value/p-value? Significance test for model DV = intercept + b1*IV versus H0 (→ DV =
intercept + 0*IV)
41
Example
• PSY7065_LinearRegression_example1.sav
Example
• PSY7065_LinearRegression_example1.sav
Example
• PSY7065_LinearRegression_example1.sav
𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 − 𝑏𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑏𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 .871
𝑡=
=
=
= 4.723
𝑆𝐸𝑏
𝑆𝐸𝑏
.184
There is a rounding error in this one strictly speaking it is 0.871308/0.184479 = 4.723074
44
Example
• PSY7065_LinearRegression_example1.sav
Design Matrix
Design Matrix
It is all Matrix Algebra – properly Linear Algebra
Example
• Try: PSY7065_LinearRegression_example2.sav
48
Multiple linear regression
• Same principles, but now with more than one
IV/predictor variable
Design Matrix Multiple Regression
Overview
• Multiple linear regression (the basics)
• Multiple linear regression (advanced)
• SPSS practice
51
Bias and Assumptions
The linear model has assumptions and we want to avoid
bias
• As all the statistical techniques are based on the linear
model the assumptions are the same across the
techniques
• Benefit of learning the linear model route
• Simple & Multiple linear regression are appropriate for
Outcome/DV: Continuous — interval ratio
Predictor(s)/IV: Continuous or Nominal
• but more than two categories can be dummy coded
as multiple predictors
52
Residuals
Squared
4.68
21.86
4.70
22.13
23.30
542.89
0.86
0.74
-4.31
18.57
1.02
1.04
-12.65
160.08
-0.12
0.02
-5.69
32.43
6.03
36.41
-8.44
71.19
-9.81
96.21
1.22
1.49
-3.35
11.23
-8.72
76.08
0.51
0.26
3.53
12.49
3.26
10.65
-2.21
4.87
-1.08
1.16
8.05
64.80
-6.92
47.91
7.61
57.87
4.04
16.29
2.06
4.26
sum =
1312.93
Outliers
• As with the mean outliers are a problem.
• But much more so as we use the sum of
squares
• Extreme values carry more weight
53
No extreme outliers
Outlier by
distance:
regression line
not greatly
affected
Outlier by
influence:
regression line
affected
54
Outliers
• What to do?
• In general, we can use histograms and box plots to spot
outliers
• In regression
– We can create standardised residuals and remove
those over 3.29 or under -3.29 (Field suggests 3)
– Remove influential cases
• Adjusted predicted value – leave one out
• Cook’s Distance (>1 is a concern)
• Leverage (investigate case with more than twice
the average leverage value)
55
Linear and Additive
• Relationship between DV and IV should be linear (a
straight line), not curvilinear (e.g. quadratic, sinusoid) or
some other wiggly non-linear function.
• There may be no relationship, but if there is one we are
assuming it will be a straight line.
• Linear relationship can be checked through visual
inspection of scatter plots
• We can use transformations to bring data back into shape
• Additive means that different predictor’s effect can be
combined by adding, The + in the linear model. They do
not interact (although we can extend to add interactions).
56
Independent Residuals
• There should be no correlation between residuals
• Assumption means that there is no correlation between any of
residuals – they are random
𝜖 ≈ 𝑖. 𝑖. 𝑑. 𝑁(0, 𝜎 2 )
• i.i.d. means every residual is independent and identically distributed
• identically distributed means there is equal likelihood they will be
drawn in a sample from the population distribution
• 𝑁(0, 𝜎 2 ) means normally distributed
• Typically arises with time (time series) or, with consecutive residuals
being related
• Also important in spatial (geographic data)
– imagine two postcodes or longitude and latitudes positions
• Can arise elsewhere
• Tested with the Durbin–Watson test
57
Normally Distributed Residuals
𝜖 ≈ 𝑖. 𝑖. 𝑑. 𝑁(0, 𝜎 2 )
• 𝑁(0, 𝜎 2 ) means normally distributed
• Often people make the mistake of thinking the data itself has to be
normally distributed — it is the residuals.
• The errors cancel out and are just noise and are mostly close to zero
with extreme instances being less common.
• Can be tested with histogram of the residuals.
• Formal tests of normality exist (Kolmogorov-Smirnov and ShapiroWilk)
58
Homoscedasticity vs
heteroscedasticity
Ideal line (population)
Data
Best fitting line (for data)
With a heteroscedasticity (here: error
increases as Time increases), estimate of
best fitting line becomes more variable.
This affects null-hypothesis significance
testing for both intercept and slope.
Tested by plotting Standardized
Residuals as a function of Standardized
Predicted Values (should not show
fanning in/out)
59
Homoscedasticity
Normal plot with residuals
Residuals plotted against the fitted values
Fanning out of the residuals evident in both plots
but clearer in the second
Solution use weighted least squares regression or robust regression
http://www.sthda.com/ 60
Little or no multicollinearity
Predictors should not be highly correlated
Correlation between predictors is called
multicollinearity
• Your predictors should not be strongly correlated
• Example:
Perfect correlation between b1 and b2
– 𝑦 = 𝑏0 + 𝑏1 × 𝑥1 + 𝑏2 × 𝑥2 = 𝑏0 + 𝑏1 × 𝑥2 + 𝑏2 × 𝑥1
– For any data point one cannot say whether its specific value is
due to x1 or x2: B1 and B2 cannot be uniquely determined
• Check VIF/Tolerance
• Solution: Ridge Regression
61
Hierarchical regression
• Determining whether an predictor should be included in
model
• Build your model from theory using hierarchical regression
– Enter known predictors in order of importance for
predicting the outcome
– After known predictors enter new ones (most important
first)
– Forced entry (Enter) includes them all at the same time
– Use model comparison techniques to let you know the
value of a model and the benefits of predictors.
62
Stepwise methods for models with
multiple predictors
• Determining whether an predictor should be included in model
• Avoid these. Build your model from theory using hierarchical
regression process rather than stepwise regression
– Forward: start with simplest model, without predictors
• add variable yielding the most significant improvement in model
• Rerun extended model
• Repeat
– Backward: start with full model
• exclude variables that least reduce model’s explanatory value
• Rerun reduced model
• Repeat
– Bidirectional
• Combination of above 2 methods
63
Reporting the Results
• APA 7 or APA 6 is fine
• APA 7 abbreviations
b, bi
in regression and multiple regression analyses, estimated values of raw (unstandardized) regression
coefficients
b*, bi*
estimated values of standardized regression coefficients in regression and multiple regression analyses
df
degrees of freedom
ES
effect size
F
F distribution
MS
mean square
r
estimate of the Pearson product–moment correlation coefficient
r2
coefficient of determination; measure of strength of relationship; estimate Pearson correlation squared
SE
standard error
SS
Sum of Squares
t
Student’s t distribution; a statistical test based on the Student t distribution; the sample value of the t-test
statistic
z
a standardized score; the value of a statistic divided by its standard error
64
Reporting the Results
• APA 7
• “The results of multiple regression, including
mediation and moderation analyses, can be presented
in a variety of ways depending on the purpose of the
table and need for detail. Clearly label the regression
type (e.g., hierarchical) and the type of regression
coefficients (raw or standardized) being reported.”
65
Reporting the Results
• 7.16 Confidence Intervals in Tables
• When a table includes point estimates—for example,
means, correlations, or regression slopes—it should also,
when possible, include confidence intervals. Report
confidence intervals in tables either by using square
brackets, as in the text (see Section 6.9) and in Table 7.16
in Section 7.21, or by giving lower and upper limits in
separate columns, as in Table 7.17 in Section 7.21. In every
table that includes confidence intervals, state the
confidence level (e.g., 95% or 99%). It is usually best to use
the same confidence level throughout a paper.
66
Reporting the Results
• Table 7.15 Sample Regression Table, Without Confidence Intervals
67
Reporting the Results
• Table 7.17 Sample Regression Table, With Confidence Intervals in Separate Columns
68
Reporting the Results
• Table 7.18 Sample Hierarchical Multiple Regression Table
69
Reporting the Results
SJPlot
in R
Papaja
in R
70
Reporting the Results
• Table 7.18 Sample Hierarchical Multiple Regression Table
APA 7
F ratios:
For immediate recognition, the omnibus test of the main effect of sentence format
was significant, F(2, 177) = 6.30, p = .002, est ω2 = .07.
t values:
The one-degree-of-freedom contrast of primary interest was significant at the
specified p < .05 level, t(177) = 3.51, p < .001, d = 0.65, 95% CI [0.35, 0.95].
Hierarchical and other sequential regression statistics:
High school GPA predicted college mathematics performance, R2 = .12, F(1, 148) =
20.18, p < .001, 95% CI [.02, .22].
71
PSY7065 Quantitative Data Analysis 1
Logistic Regression Coursework
A researcher studied decision making in childhood using a variant of the marshmallow test.
In the regular marshmallow test (Mischel & Ebbesen, 1970), children are given a
marshmallow on a plate and promised two marshmallows in reward for not eating the
marshmallow while the experimenter is out of the room (For an example, see:
https://www.youtube.com/watch?v=QX_oy9614HQ.)
The variant studied here concerned a shared reward. In this ‘Shared-Reward’ variant
children were told beforehand that one of the two promised marshmallows had to be
shared with a friend. They thus only got to keep one marshmallow.
The researcher registered participants’ decisions (Wait=0 or Eat=1) in this scenario and also
recorded several other factors thought to be relevant:
•
•
•
Gender (Male=0, Female=1)
Self-control (0-100, measured through a validated questionnaire; higher means more
self-control)
Ability to have fun (0-30, measured through a questionnaire, completed by the
child’s carers; higher means more ability to have fun).
The data for 194 participants is presented in the accompanying file
PSY7065_LogisticRegression_Portfolio2022.sav. Conduct a logistic regression analyses to
examine the relationship between decisions as a function Gender, Self-control, and Ability
to have fun. Provide the motivation that you are allowed to conduct the analyses (i.e.,
ensure the assumptions are met) and conduct the analyses. Report information on the
model fit, amount of variance explained, which individual variables act as significant
predictors and the relationship between the individual variables and whether or not they
increase the odds of the child deciding to eat the marshmallow.
•
•
•
•
•
•
•
The data has been created for assessment purposes; it may not reflect reality.
In your submitted document, avoid simply cutting and pasting prolifically from SPSS
output, choose your evidence to submit carefully. The only output that can be
included concerns the data summary and/or the assumption checks (although, when
possible, reporting statistical values within the text is preferred); for any figures,
include appropriate captions.
Submit this coursework together with the other two assignments in a single portfolio
(i.e., a single file). This section should be word processed, written in full paragraphs,
and should not be longer than 3 sides of A4.
Report findings in the appropriate statistical format.
Cite references to support your answers where relevant. This only concerns
references about statistics, not about the fictive topic.
Interpret and explain your results; display understanding of the statistical principles
underlying the test.
Do not copy out lecture notes or material from the internet!
PSY7065 Quantitative Data Analysis 1
•
All coursework will be checked for plagiarism.
References
Mischel, W., & Ebbesen, E. B. (1970). Attention in delay of gratification. Journal of Personality and
Social Psychology, 16(2), 329–337. https://doi.org/10.1037/h0029815
PSY7065 Quantitative Data Analysis 1
ANOVA Coursework
A researcher studied the effectiveness of a new virtual reality-based treatment for
Agoraphobia (fear of open spaces). She asked people who suffered from this phobia to
participate in her study. The researcher was interested in the direct effect of the treatment
as well as its longer-term effects; she thus measured agoraphobia prior to (PreTest), directly
after (PostTest) and 1 week after the virtual reality treatment (OneWeek). The degree of
experienced agoraphobia was quantified using the average heart rate measured in a real-life
test on the university main square (higher means more fear).
Two variants of the virtual reality treatment were considered, and participants were divided
into two groups to explore the effects of these variants: one involving participants standing
in different types of open spaces (level called Standing) and one in which they were moving
around in these spaces (level called Moving). In the data file 1 = Standing and 2 = Moving.
The data for 20 participants is presented in the accompanying file
PSY7065_ANOVA_Portfolio2022.sav
1) Conduct and report the appropriate analysis, including post-hoc analyses, after
confirming that its assumptions have not been violated (i.e., also report evidence of the
associated checks).
•
•
•
•
•
•
•
•
The data has been created for assessment purposes; it may not reflect reality.
In your submitted document, avoid simply cutting and pasting prolifically from SPSS
output, choose your evidence to submit carefully. The only output that can be
included concerns the data summary and/or the assumption checks (although, when
possible, reporting statistical values within the text is preferred); for any figures,
include appropriate captions.
Submit this coursework together with the other two assignments in a single portfolio
(i.e., a single file). This section should be word processed, written in full paragraphs,
and should not be longer than 3 sides of A4.
Report findings in the appropriate statistical format.
Cite references to support your answers where relevant. This only concerns
references about statistics, not about the fictive topic.
Interpret and explain your results; display understanding of the statistical principles
underlying the test.
Do not copy out lecture notes or material from the internet!
All coursework will be checked for plagiarism.
PSY7065 Quantitative Data Analysis 1
Logistic Regression Coursework
A researcher studied decision making in childhood using a variant of the marshmallow test.
In the regular marshmallow test (Mischel & Ebbesen, 1970), children are given a
marshmallow on a plate and promised two marshmallows in reward for not eating the
marshmallow while the experimenter is out of the room (For an example, see:
https://www.youtube.com/watch?v=QX_oy9614HQ.)
The variant studied here concerned a shared reward. In this ‘Shared-Reward’ variant
children were told beforehand that one of the two promised marshmallows had to be
shared with a friend. They thus only got to keep one marshmallow.
The researcher registered participants’ decisions (Wait=0 or Eat=1) in this scenario and also
recorded several other factors thought to be relevant:
•
•
•
Gender (Male=0, Female=1)
Self-control (0-100, measured through a validated questionnaire; higher means more
self-control)
Ability to have fun (0-30, measured through a questionnaire, completed by the
child’s carers; higher means more ability to have fun).
The data for 194 participants is presented in the accompanying file
PSY7065_LogisticRegression_Portfolio2022.sav. Conduct a logistic regression analyses to
examine the relationship between decisions as a function Gender, Self-control, and Ability
to have fun. Provide the motivation that you are allowed to conduct the analyses (i.e.,
ensure the assumptions are met) and conduct the analyses. Report information on the
model fit, amount of variance explained, which individual variables act as significant
predictors and the relationship between the individual variables and whether or not they
increase the odds of the child deciding to eat the marshmallow.
•
•
•
•
•
•
•
The data has been created for assessment purposes; it may not reflect reality.
In your submitted document, avoid simply cutting and pasting prolifically from SPSS
output, choose your evidence to submit carefully. The only output that can be
included concerns the data summary and/or the assumption checks (although, when
possible, reporting statistical values within the text is preferred); for any figures,
include appropriate captions.
Submit this coursework together with the other two assignments in a single portfolio
(i.e., a single file). This section should be word processed, written in full paragraphs,
and should not be longer than 3 sides of A4.
Report findings in the appropriate statistical format.
Cite references to support your answers where relevant. This only concerns
references about statistics, not about the fictive topic.
Interpret and explain your results; display understanding of the statistical principles
underlying the test.
Do not copy out lecture notes or material from the internet!
PSY7065 Quantitative Data Analysis 1
•
All coursework will be checked for plagiarism.
References
Mischel, W., & Ebbesen, E. B. (1970). Attention in delay of gratification. Journal of Personality and
Social Psychology, 16(2), 329–337. https://doi.org/10.1037/h0029815
Essay Writing Service Features
Our Experience
No matter how complex your assignment is, we can find the right professional for your specific task. Achiever Papers is an essay writing company that hires only the smartest minds to help you with your projects. Our expertise allows us to provide students with high-quality academic writing, editing & proofreading services.Free Features
Free revision policy
$10Free bibliography & reference
$8Free title page
$8Free formatting
$8How Our Dissertation Writing Service Works
First, you will need to complete an order form. It's not difficult but, if anything is unclear, you may always chat with us so that we can guide you through it. On the order form, you will need to include some basic information concerning your order: subject, topic, number of pages, etc. We also encourage our clients to upload any relevant information or sources that will help.
Complete the order form
Once we have all the information and instructions that we need, we select the most suitable writer for your assignment. While everything seems to be clear, the writer, who has complete knowledge of the subject, may need clarification from you. It is at that point that you would receive a call or email from us.
Writer’s assignment
As soon as the writer has finished, it will be delivered both to the website and to your email address so that you will not miss it. If your deadline is close at hand, we will place a call to you to make sure that you receive the paper on time.
Completing the order and download