数学|MATH2010 – Statistical Modelling I Coursework 1

联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981

MATH2010 – Statistical Modelling I
Coursework 1
This coursework accounts for 20% of the module assessment. The total number of marks available is 20. Your
solutions should be submitted via Blackboard by 4pm on Friday 17 March.
Your solutions should take the form of an RMarkdown document, that produces all requested analyses and plots.
You must annotate your RMarkdown code using markdown blocks. These annotations should explain what you
are doing, and provide insight into the analysis.
Please name your solution file math2010 cwk student number.Rmd, where student number is replaced by your
student number. Do not put your name anywhere in the file.
Questions about loading data into R will be answered in the computing labs in Week 6 (beginning 6 March).
All other questions should be asked via the Coursework forum on Blackboard.
[2 marks] will be awarded for presentation and reproducibility. The following list describes the things
we will be looking for:
We should be able to run your script on our machines from start to finish without modification. You
should check that you can run your script from a clean R environment (by using the command Run >
Restart R and Run All Chunks as shown in lectures).
(Relatedly) we should not need to remove or change the working directory. Your notebook should assume
that the data is located in the same directory as the notebook.
You must use set.seed(1) at the start of your notebook to make your results reproducible. (i.e. if we
run your notebook twice, it should produce the same results).
It must be possible to identify which question you are answering with each piece of code.
Explain your results clearly with text, inbetween your code blocks.
All graphs should have axis labels.
Question 1 [9 marks]
Load the data q1 data.csv from Blackboard. This consists of a single explanatory variable x1 (with column
name x), but does not contain a response variable.
(a) [1 mark] Consider the simple linear regression model
Yi = β0 + β1xi1 + _x005f_x000f_ i
where i
iid~ N (0, σ2
). Using the R function rnorm, generate a vector of response variables y for the model
above, with β0 = β1 = σ = 1.
(b) [2 marks] Fit a simple linear regression model with the variable generated in Part (a) as the response and
x1 as the explanatory variable. Comment on the significance of x1.
(c) [1 mark] Compute a 95% confidence interval for β1. Does the true value of β1 (i.e. β1 = 1) lie in this
confidence interval
(d) [4 marks] Simulate 1,000 different instances of the response variable y. For each of these, fit a simple
linear regression model. Calculate in what proportion of your simulations does the parameter β1 lie in a
95% confidence interval from the model. (See the hint below for help.)
(e) [1 marks] Explain these results, referring to your understanding of what a confidence interval represents.
1
Question 2 [9 marks]
In the second question we will build a regression model for modelling mosquito populations in Seoul, South
Korea. The dataset mosquito.csv contains measurements taken daily measurements of the following quantities,
taken between 2016 and 2019:
mosquito — an indication of the number of mosquitoes in a particular area on that day. (This is the
response variable).
rain — rainfall (in mm).
min T, mean T, max T — the minimum, mean and maximum temperatures measured, respectively, in
degrees celsius.
(a) [2 marks] Load the data into R. Visualise the data: plot four scatter plots, each having the response
variable on the y-axis and a different explanatory variable on the x-axis. Comment on the potential for
fitting a linear regression model.
(b) [2 marks] Fit a linear regression model for the response against all of the explanatory variables. For each
of the explanatory variables, perform a hypothesis test at the 95% significance level for the null hypothesis
H0 : βi = 0, with alternative hypothesis H1 : βi 6 = 0.
(c) [2 marks] Produce an Anscombe and normal probability plot, and use them to comment on any departures
from the standard assumptions in a linear regression model.
Explanatory variables are said to be collinear if they are highly correlated with one another. In linear regression
collinearity can lead to strange phenomena, for example explanatory variables being reported as insignificant
when they have an evident correlation with the response when plotted on a graph, or having negative coefficients
when there is clearly a positive correlation when plotted.
(d) [2 marks] Compute the sample correlation between each pair of explanatory variables. Which of the
variables in this dataset do you believe to suffer from the highest degree of collinearity
(e) [1 marks] Considering the graphs plotted in (a) and the model fitted in (b), what impact do you think
this has had on your model (if any)
Hints
There are a lot of different ways to do 1(d). My suggestion is to write a for loop in R. The syntax for this looks
like
n = 10
for(i in 1:n) {
print(paste0(“Hello “, i))
}
The code for(i in 1:n) tells R to create a variable called i and run the code between the { and } with i set
to each value in the vector 1:n in turn. (You can probably guess what the vector 1:n contains, but typing 1:n
into an R prompt will make this clear!) Running the code above should produce the output
[1] “Hello 1”
[1] “Hello 2”
2
[1] “Hello 3”
[1] “Hello 4”
[1] “Hello 5”
[1] “Hello 6”
[1] “Hello 7”
[1] “Hello 8”
[1] “Hello 9”
[1] “Hello 10”
Obviously you will need to write some slightly more sophisticated code between { and } to answer the question!
An even nicer way of accomplishing this is to write an R function which performs the task, and use the function
sapply to run the function multiple times (this is how I solved the problem). However, which approach you
use to solve this won’t affect your mark.
3

发表评论

了解 KJESSAY历史案例 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读