联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981
Module Code: LUBS5346M01
Page 1 of 18 Turn the page over
Module Title: Data Analytics for Human
Resources
UNIVERSITY OF LEEDS
Leeds University Business School Semester Two 2022/2023
Exam information:
There are 20 pages to this exam.
There will be 48 hours to complete this exam. We anticipate that this exam
should take students approximately 2 hours to complete.
This exam paper contains 3 sections.
Section A has ten questions and is worth 50 marks. Answer all questions in
section A.
Section B has four questions and is worth 40 marks. Answer all questions in
section B.
Section C has two questions, answer one question only. The word limit for section
C is 1000 words.
You should allow a minimum of 30 minutes to upload/submit your work.
The deadline date for this assessment is 09:00 am on Wednesday 24th May
2023 (UK time).
An electronic copy of your completed answers must be submitted to the exam
submission area within the module resource on the Blackboard MINERVA website no
later than the date and time stated above.
Faxed, emailed or hard copies of your exam answers will not be accepted.
Late submissions will not be accepted.
Failure to meet this deadline will result in you being marked as absent from
this assessment.
Module Code: LUBS5346M01
Page 2 of 20 Turn the page over
Producing your assessment
We would prefer your examination answers to be typed up and submitted using Microsoft
Word.
If in any answer, it is not possible to complete it electronically (eg calculations, graphs,
formulae) these can be handwritten. Please take a photograph of your workings and save
the image. The image should then be inserted into your examination paper at the correct
point within your answer. It is your responsibility to ensure that any image that is inserted is
readable. High quality images can impact the size of the file and causes issues with
uploading.
Further information on how to insert images can be found on this page
Submission
Please ensure that you leave sufficient time to complete the online submission process, as
upload times can vary. Accessing the submission link before the deadline does NOT
constitute completion of submission. You MUST click the ‘CONFIRM’ button before the date
and time stated above for your answers to be classed as submitted on time.
It is important that any file submitted follows the conventions stated below:
File name
The name of the file that you upload must be your student ID only.
Submission title
During the submission process the system will ask you to enter the title of your submission.
This should also be your student ID only.
Front cover
The first page of your submission should always be the Online Exam Coversheet, which is
provided in the same Minerva area as this exam paper.
Student name
You should NOT include your name anywhere on your submission.
Word limit
You are required to adhere to the word limit specified and state an accurate word count on the
cover page of your assignment. Your declared word count must be accurate, and should not
mislead. Making a fraudulent statement concerning the work submitted for assessment could
be considered academic malpractice and investigated as such. If the amount of work
submitted is higher than that specified by the word limit or that declared on your word count,
this may be reflected in the mark awarded and noted through individual feedback given to you.
Questions during your exam
If you have a question about this examination paper during the exam, please direct all
queries to the assessment team on LUBSassessment@leeds.ac.uk between the hours of
9am – 5pm (UK Time), Monday to Friday. All questions should go through this route. Do
not contact your module leader directly.
Module Code: LUBS5346M01
Page 3 of 20 Turn the page over
Section A
There are 50 marks in total for section A. There are 10 questions in total. Answer all
questions in the script book provided.
1. Suppose you wish to study the relationship between weekly wage and education.
For your research, you collect data on random sample of individuals and then run
a regression of wages on education (measured in number of years) and other
variables, namely, age, gender, IQ, and years of experience. Briefly suggest one
reason why you might be interested to transform the continuous variables using zscores in the regression model.
(2 marks)
2. Briefly describe hierarchical clustering and provide one HR application on how it
can be used.
(3 marks)
3. As a consultant for a government sponsored project, you are studying the role of
manager-employee relations and job satisfaction and productivity in the UK public
sector. For this study, you are aiming to collect data on a sample of public sector
employees’ background information, job characteristics, salary, and their
perceptions of their working environment, workplace procedural justice, and the
relationship with their manager. As a follow-up, the project also aims to collect a
repeated data on the same set of employees 2 years after the first baseline sample
data was collected.
i. Briefly outline two types of sampling biases that can occur when collecting the
data.
(2 marks)
ii. For the study, you analyse statistics, such as, mean job satisfaction, median
labour productivity etc. from the sample. Briefly explain what is meant by the
standard error of the mean job satisfaction
(2 marks)
Module Code: LUBS5346M01
Page 4 of 20 Turn the page over
4. Figure 1 shows the distribution in the use of HR analytics for a sample of European
firms from the 2019 European Company Survey data. The adoption of HR analytics
is measured using a dummy variable, where 1 = if a firm has adopted HR analytics
to monitor performance and 0 otherwise. There are 20,000 firms in the data and
the percentage of firms that have adopted HR analytics to monitor employee
performance is 30%.
i. From the information provided, find the number of firms that don’t use HR
analytics to monitor employee performance.
(1 mark)
ii. You are interested to study the impact of company’s use of HR analytics to
monitor employee performance – a binary categorical variable – on employee
productivity. You control for firm size (measured by number of employees), firm
age, managerial experience, and other firm controls. The regression model is
specified below:
= 0 + 1 + 2 + 3
+ +
where i = 1, 2, 3, …, 20000. Assume that all the continuous variables are
unstandardized. Briefly explain why the coefficients of the binary variable on the
use of HR analytics by the firm, 1 are often larger than the coefficients of
numeric variables, such as, company age, in the regression model
(2 marks)
Module Code: LUBS5346M01
Page 5 of 20 Turn the page over
Figure 1: Use of HR Analytics to Monitor Performance
5. Figure 2 below produces the absenteeism rate amongst NHS Staff from the years
2010/11 and 2018/19. Table 2 provides the summary statistics over the same
period. The data are obtained from NHS Digital.
(7 marks)
Figure 2
Module Code: LUBS5346M01
Page 6 of 20 Turn the page over
Table 1: Summary Statistics for NHS Staff Absenteeism Rate
Year Minimum
25th
Percentile Median Mean
75th
Percentile Maximum
Std.
Dev
2010/11 2.63 3.52 3.89 3.90 4.25 5.24 0.55
2011/12 2.42 3.52 3.86 3.89 4.23 5.37 0.56
2012/13 2.58 3.68 4.01 4.03 4.48 5.57 0.59
2013/14 2.52 3.47 3.88 3.86 4.18 5.26 0.56
2014/15 2.42 3.58 4.04 4.04 4.49 5.83 0.63
2015/16 2.43 3.57 3.99 4.00 4.41 5.80 0.60
2016/17 2.45 3.64 4.11 4.05 4.38 5.59 0.59
2017/18 2.79 3.76 4.14 4.07 4.40 5.52 0.57
2018/19 2.74 3.71 4.16 4.10 4.50 5.76 0.61
i. For the year 2018/19 from Table 1, briefly explain the result of the median
absenteeism rate.
(1 mark)
ii. For the year 2018/19 from Table 1, briefly explain the 75th percentile statistic
for NHS staff absenteeism rate and provide a brief comment on this statistic
over the years.
(3 marks)
iii. Briefly explain and comment on the standard deviation of NHS Staff
absenteeism rate.
(2 marks)
iv. Briefly explain why the standard deviation is a more intuitive measure than the
variance of a distribution.
(1 mark)
Module Code: LUBS5346M01
Page 7 of 20 Turn the page over
6. The company that you work for is concerned that there may be a gender-wage gap
between its female and male employees with respect to performance evaluations.
As a senior Data Scientist for the company, you have been asked to investigate if
such evidence for ‘glass-ceiling’ exists overall and by specific performance
evaluation score. Figure 3 below is a boxplot, where the y-axis measures total
salary, sum of base salary and bonus, and the x-axis contains the 5 categories of
performance evaluation scores from 1-to-5, where 1 is the lowest possible
performance score and 5 is the highest possible.
i. Briefly explain what the boxplot tells us about gender pay gaps.
(3 marks)
ii. Suppose you wish to examine whether gender-gap exists across the
performance scores in the company. Briefly explain the analytical approach you
would take.
(4 marks)
Figure 3: Gender Gap
Module Code: LUBS5346M01
Page 8 of 20 Turn the page over
7. Using a data set on a random sample of 209 CEOs from Bloomberg Businessweek,
you regress CEO salary on company’s return on equity (ROE) using an OLS
regression. ROE is measured as return on investment as a percentage and CEO
salary is measured in terms thousand dollars. The estimated relationship is given
below.
(4 marks)
= 963.19 + 18.50
i. The standard error of the estimate for ROE is 11.12. Find the t-test statistic
and test whether the null hypothesis of no effect of ROE on CEO salary can
be rejected at 5% level of significance.
(2 marks)
ii. The
2
from the simple OLS regression is 0.008. Provide the interpretation
for this
2 value.
(1 mark)
iii. Why is the
2 very low
(1 mark)
Module Code: LUBS5346M01
Page 9 of 20 Turn the page over
8. Figure 4 below plots a simple scatter graph between UNDP’s overall gender gap
index against globalization index. The higher the value of gender gap index, the
lower the inequality between women and men. Whereas, the data for globalization
index is from the KOF Swiss Economic Institute and the greater its value, the
greater is the level of trade and economic openness in the society.
(5 marks).
i. The correlation coefficient between these two indices is 0.46. Briefly explain the
plot and interpret the value of the coefficient.
(2 marks)
ii. Briefly explain why the relationship between globalisation and gender-gap index
may be confounded.
(3 marks)
Figure 4: UNDP Gender-Gap Index vs. Globalisation Index
Module Code: LUBS5346M01
Page 10 of 20 Turn the page over
9. Suppose you apply OLS to the following regression model to estimate the
relationship between hot weather and labour productivity from a cross-sectional
sample of employees working for a construction company.
(8 marks)
= 0 + 1 + +
+ ; = 1,2, … , = 1, 2, … ,
In the above regression model, i, indexes individual and c indexes the country.
Temp, here, is a dummy variable that is equal to 1, if the temperate exceeds 35°C
and 0 otherwise. The following is the plot of residuals and fitted values from the
model:
Figure 5: Residuals against Fitted Values
Module Code: LUBS5346M01
Page 11 of 20 Turn the page over
i. Based on Figure 5, discuss whether the regression model satisfies the
assumption of homoscedasticity, that is, constant variance.
(2 marks)
ii. How will the presence of non-constant variance of the error term going to
affect the regression analysis
(2 marks)
iii. Suggest one visual diagnostic tool that can be used to examine if the error
terms are normally distributed and suggest one solution if the normality
assumption is not met.
(2 marks)
iv. Briefly discuss whether autocorrelation is likely to be issue for the above
regression model in Question 9.
(2 marks)
10. As in studies by Card (1995) and Joshua and Krueger (1991), you collect data to
estimate the following structural equation
= 0 + 1 +
The subscript, i, indexes individual and is the error term. Due to concerns of
endogeneity, you perform an instrumental variable regression using the birth order
of the individual as an instrument. You also decide to add other controls, such as,
age, gender, and training.
i. Briefly explain how you would test to examine if the instrument, birth order,
satisfies the instrument relevance assumption.
(3 marks)
ii. Briefly discuss if you expect the ‘birth order of the individual’ variable to be
uncorrelated with the error term and, therefore, satisfy the exogeneity
assumption of the instrument variable.
(4 marks)
Module Code: LUBS5346M01
Page 12 of 20 Turn the page over
Section B
There are 40 marks total for Section B. There are 4 questions. Each question
consists of 10 marks in total. Answer all questions in the script book provided.
1. The following is a fitted regression line using logistic regression obtained by
regressing a dummy variable for Swiss labour force participation (that is,
employment) on age, age squared, income, education, does the individual have
young and old kids, and is a foreigner. The data is from 1981 Swiss Health Survey
Project. Income is the logarithm of non-labour income; age is measured as age in
years; Education is years of formal education; YoungKids is number of young
children (under 7 years old); OldKids is number of old children (over 7 years of
age); and Foreign is a dummy/categorical variable to indicate whether the
individual is a foreigner, that is, not a Swiss national. All the variables are
statistically significant at 5% level of significance except for education.
= 6.19 1.10 + 0.34 0.005
2 + 0.03
1.19 0.24 + 1.17
i. Describe the effects of the variables, income, young kids, and foreign in
terms of log-odds.
(4 marks)
ii. Describe the effects of the variables, income, young kids, and foreign in
terms of odds.
(3 marks)
iii. Given the results for the variables age and its squared term, briefly describe
the relationship between labour force participation and individuals’ age. The
mean for the age variable is 39.96, 1st quartile is 32, and 3rd quartile is 48.
(3 marks)
Hint: This is an example of a second order polynomial regression. To
interpret the coefficient for age, use a special value for age, for example, the
mean or the 1st quartile.
Module Code: LUBS5346M01
Page 13 of 20 Turn the page over
2. Table 2 provides ordinary least squares (OLS) regression results using data on
firm-level data on performance and management practices for six EU countries,
namely, Austria, Belgium, France, Germany, Luxembourg, and Netherlands from
the July 2022 World Bank Enterprise Survey data. The main dependent variable is
HR score, which is a composite measure of performance bonuses for managers
and non-managers, identifying and retaining talent. The greater the HR score, the
greater or more innovative is the quality of HR management at the firm. The score
is converted to a z-score that is, standardised. Four performance measures were
used: sales measured as total annual sales for all products and services in the last
fiscal year measured in euros; labour productivity measured as sales divided by
the number of full-time workers, employment measured as the number of
permanent full-time workers in the last fiscal year; and innovation measured as
whether a new product or service was introduced by the firm in the last three years.
Log sales, log labour productivity and log employment were used in the regression
in models (1) to (3), whereas innovation is a dummy variable (Innovation) if a firm
has introduced a new product or service (Model 4). All the regression models
control for firm characteristics, such as, size, export status, foreign ownership, state
ownership etc. and country and sector dummies.
Table 2: Regression Results for Firm Performance and Human Resource
Management Practices
Log of Sales Log Labour
Productivity
Log
Employment Innovation
(1) (2) (3) (4)
HR Score (Z-scale) 0.222*** 0.151*** 0.077*** 0.068***
(0.022) (0.023) (0.015) (0.011)
Controls:
Firm Yes Yes Yes Yes
Country Yes Yes Yes Yes
Sector Yes Yes Yes Yes
————–
Observations 2,252 2,250 2,328 2,329
Adjusted R2 0.674 0.249 0.821 0.114
F Statistic 126.533*** 21.107*** 289.306*** 9.105***
Note: *p<0.10 **p<0.05 ***p<0.01
Module Code: LUBS5346M01
Page 14 of 20 Turn the page over
i. Assume that all the continuous control variables are measured at their
respective mean values and categorical control variables assume the value
of highest occurring category. Briefly explain the marginal effects of the
standardised HR score variable for each of the performance measures
corresponding to regression models from (1) to (4).
(4 marks)
ii. Explain two key challenges that are likely to result when running a regression
of firm innovation dummy on HR score in model (4) of Table 1 using OLS.
(2 marks)
iii. Based on the regression results, discuss to what extent the regression
results provided in Table 1 are unbiased.
(3 marks)
iv. Suppose models from (1)-(4) suffer from heteroscedasticity. Suggest one
option that can be used to correct for heteroscedasticity.
(1 mark)
Module Code: LUBS5346M01
Page 15 of 20 Turn the page over
3. Suppose in 2021, the UK government raised the maximum weekly employee
compensation benefits on the time until an injured worker returns to work following
an injury at the workplace. The UK government wants to conduct a study to know
if the introduction of this new policy has led to some workers spending more time
unemployed one year after its introduction.
If the raise is not generous enough, then the workers could sue companies for
workplace related injuries. On the contrary, if the new benefits are very generous,
then this may affect workers’ incentives to avoid injuries, and, therefore, may
induce some workers to be more reckless and filing for compensation benefits for
any given job related injury. This is known as the moral hazard problem. Suppose
the policy was designed in such a way that it affected the high income working
population – the treatment group – but not the low income segment of the
population – the control group. The main outcome variable is duration in weeks of
workers’ compensation benefits. To study the effect of this policy, the following
Difference-in-Difference (DID) model is estimated:
= 0 + 0 21 + 1 + 1
( × 21) +
The outcome variable, duration, is measured in logarithms. 21 is a dummy
variable that is equal to 1 if the year is post the year of intervention, that is, 2021.
The dummy variable, , is the treatment dummy that is equal to 1 if
workers belong to high-income group and 0 for low-income group. Figure 6 shows
the effect of the policy change.
Table 3 provides the DID regression results. Model (1) in Table 3 is a simple DID
model, whereas Model (2) includes the list of full set of controls. Model (2) controls
for Male dummy variable (1 = male employee; 0 = female employee), Married (1 =
‘married’; 0 = ‘single’), age at the time of injury (Age), dummy variable if employee
was hospitalised from the injury (1 = ‘Yes’; 0 = ‘No’), and log of employee wage.
Model (2) also controls for 7 different injury types (head, neck, lower back etc.) and
type of industry (manufacturing, construction, and other)
Module Code: LUBS5346M01
Page 16 of 20 Turn the page over
i. Based on the DID results in Table 3, interpret the results of the coefficients of
and 21 variables in models (1) and (2), and comment on their
significance.
(3 marks)
ii. Based on the results in Table 3, explain if the introduction of the policy of
raising the level of benefits had an effect on the duration of employee
compensation benefits.
(3 marks)
iii. Discuss the necessary assumptions required for the DID regression estimate
to yield a causal effect of the policy on the duration of employee compensation
benefits.
(4 marks)
Figure 6: Policy Change
Module Code: LUBS5346M01
Page 17 of 20 Turn the page over
Table 3: DID Estimates
Simple Full
(1) (2)
HighIncome 0.215***
-0.327***
(0.043) (0.067)
post21 0.024 0.058
(0.040) (0.037)
Male -0.140***
(0.039)
Married: Single -0.026
(0.033)
Age 0.008***
(0.001)
Hospital: Yes 1.098***
(0.033)
Log Wage 0.462***
(0.057)
HighIncome X post21 0.188*** 0.162***
(0.063) (0.059)
Controls:
Industry No Yes
Injury Type No Yes
--------------
Observations 7,150 6,822
Adjusted R2 0.015 0.180
F Statistic 38.342*** 89.112***
Note: *p<0.10 **p<0.05 ***p<0.01
Module Code: LUBS5346M01
Page 18 of 20 Turn the page over
4. As a Data Scientist working for an IT start-up, you are interested to use machine
learning models to predict employee attrition. Attrition is a dummy variable and
equal to 1 if an employee has left the company and 0 otherwise. The mean
employee attrition rate is 16%. Table 4 provides names, definitions, and brief
summary statistics for some of the variables used to predict attrition.
(10 marks)
Table 4: Summary Statistics for Employee Attrition Model
Variables Definition N Mean SD
Minimu
m
25th
Quartil
e
75th
Quartile
Maximu
m
Age
Age of the
employee 1,470 36.924 9.135 18 30 43 60
Education
Years of
education 1,470 2.913 1.024 1 2 4 5
Gender
Gender = 1 if Male
and 0 = Female 1,470 0.60
Monthly Income
Monthly Income in
$ 1,470 6,503 4,708 1,009 2,911 8,379 19,999
Performance Rating
Employee
Performance
Ratings 1,470 3.154 0.361 3 3 3 4
Stock Option Level
Number of
company stocks
employee own 1,470 0.794 0.852 0 0 1 3
Years In Current Role
The number of
years employee
has been working
in the current role 1,470 4.229 3.623 0 2 7 18
Environment
Satisfaction
Employee
satisfaction's
score with the
company's
working
environment
(1='lowest
satisfaction score',
4='highest
satisfaction score') 1,470 2.722 1.093 1 2 4 4
Job Satisfaction
Employee job
satisfaction score
(1='lowest
satisfaction score',
4='highest
satisfaction score') 1,470 2.729 1.103 1 2 4 4
To this end, you have decided to use decision trees as your machine learning
algorithm. Figure 7 provides the result of classification tree.
Module Code: LUBS5346M01
Page 19 of 20 Turn the page over
Figure 7: Classification Tree for Employee Attrition
i. Based on information from Table 4, explain the output of the classification
tree in Figure 7.
(4 marks)
ii. Briefly explain how can obtain the result of the classification tree that leads
to best possible predictive accuracy.
(2 marks)
iii. Briefly discuss some of the limitations of using decision trees for predictive
modelling and suggest some alternatives.
(4 marks)
Module Code: LUBS5346M01
Page 20 of 20 End
Section C
This is an essay style question section. There are 10 marks in total for section C.
There are two questions. Answer one question only in the script book provided.
1. Employee turnover or attrition is one of the most significant and expensive
challenges faced by employers across virtually all industries. Suppose your
company wants to build a prediction model for the purposes of attrition analytics.
As a senior HR Data Scientist, you have been provided with employee data on
their work-related characteristics, such as, level of job autonomy, job satisfaction
score etc. and information on their individual background, for example, age, tenure,
education, gender etc. Your role is to identify those features or characteristics that
form some of the most important predictors of employee attrition.
i. Discuss why machine learning algorithms would be better suited to such
prediction tasks then regression using OLS.
ii. Explain how would assess the predictive performance of your model using
machine learning algorithms, such as, random forest.
2. As an HR consultant, you are interested to study the relationship between NHS
managers’ interpersonal skills also known as people management skills and
nursing turnover across hospitals in the UK. You hypothesize that, managers with
better people management skills, such as, the ability to communicate effectively
with co-workers builds trust and, in turn, is expected to lower attrition rate. You are
interested in estimating the following equation:
= 0 + 1 +
In the above equation, h, indexes hospital, , is the error term.
i. Discuss why using OLS to fit the above equation is likely to lead to biased
estimates for the impact of people management quality.
ii. Discuss the extent to which an unbiased estimate on the impact of people
management on nursing turnover can be obtained by using a panel data.


发表评论