BUSANA 7003 – Business Analytics Project

联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981

BUSANA 7003 – Business Analytics Project
Last course additional assessment
(1) Exploratory data analysis, visualisations, descriptive statistics
Examine the dataset Stock_data_part1.xlsx and characterise the performance of US securities during
the period 20/08/2019 – 20/08/2020. Specifically, discuss the following:
– How does the average number of trades differ for stocks vs ETFs over the period 20/08/2019 –
20/08/2020
– Use sprtrn variable to compute monthly returns. Compare daily returns to monthly returns. How
does the calculation frequency affect the mean and standard deviation of S&P 500 returns
How does the time series pattern change
– Run a t-test to compare the average daily returns for stocks during the month of March 2020,
vs the month of July 2020.
In your output, please report the appropriate statistics to back your arguments. For example, you may
consider:
A table with descriptive statistics (mean, median, 25th percentile, 75th percentile, standard
deviation, min, max).
The time series of the key variables of interest.
A table with t-statistics, means, and standard deviations.
(2) OLS regressions
Use the dataset Stock_data_part1.xlsx and compute security-level averages for all variables you’ll need
in this analysis. This analysis will be based on cross-sectional data. Compute the relative bid-ask spread
(RelSpread=(ask-bid) / midpoint) for all securities, and run regression analysis to help explain
RelSpread for stocks vs ETFs. You may use any variables in the existing dataset, or variables outside
the provided datasets (from external sources). Please, start by using volatility (Volatility = (ASKHI –
BIDLO)/ [0.5*(ASKHI + BIDLO)]) as the key explanatory variable. Then, you may add other variables
of interest, which you consider relevant to explaining RelSpread.
Make sure to split the dataset you use for analysis into training and testing samples, and comment on
the model accuracy. In your analysis, you should present the following output:
Regression coefficients from at least 4 different regression models (each model with a
different set of explanatory variables). Comment on whether your models explain
RelSpread better for stocks or for ETFs.
Model evaluation metrics: MSE, RMSE, MAE. Comment on how these metrics differ for
training vs testing sample.
Your assessment of which factors are most important for explaining RelSpread. Comment
on how well your models perform in the testing sample.
BUSANA 7003 – Business Analytics Project
(3) Probit regressions
Use the dataset from the previous task, and evaluate the probability of observing RelSpread above 0.01.
Fit a probit model, and comment on the following:
– What is the role of returns in observing RelSpread above 0.01
– Does the sign on returns matter for observing RelSpread above 0.01
– Does the type of security (stock or ETF) matter for observing RelSpread above 0.01
– Does market capitalisation matter for observing RelSpread above 0.01
– Can you reliably predict the probability of observing RelSpread above 0.01, if you account for
returns, market capitalisation, volatility, and dollar volume traded
Make sure to split the dataset you use for analysis into training and testing samples, and comment on
the model accuracy. In your analysis, you should present the following output:
Regression coefficients from at least 4 different regression models (each model with a
different set of explanatory variables).
Model evaluation metrics: Accuracy, Precision, Recall. Comment on how these metrics
differ for training vs testing sample.
Your assessment of which factors are most important for predicting RelSpread above 0.01.
Comment on how well your models perform in the testing sample.

发表评论

了解 KJESSAY历史案例 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读