联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981
Machine Learning Assignment: Using Supervised Machine Learning to Predict Sepsis Diagnosis
Aim
To apply learned data exploration and machine learning skills to design and implement an end-to-end machine learning pipeline for predicting a binary outcomes from a highly dimensional real-world dataset.
Weight:
This assignment will carry 70% of your module’s assessment grade.
Indicative Timetable
The following table shows a breakdown of the activities for this week to help you in prioritising your time schedule. The actual time needed to complete each task is dependent on your implementation/analysis speed and personal circumstances. The timetable provided below is merely representative of the average time needed to complete each task and the time needed for each relative to other tasks. The timetable is also provided to ensure that you know that the assignment is time-consuming and requires thoughtful planning to complete all required tasks.
Activity Tentative Duration
Data Cleaning
Data Aggregation
Data Exploration & Visualisation
Classifier Implementation & Hyperparameter Tuning
Model Evaluation
Analysis and Reflection
Writing
Assignment Details
Now that you have been equipped with the skills to use different Machine Learning algorithms, you will have the opportunity to practice and apply it to a practical problem using real-world hospital data. The dataset provided in the assignment folder is a .csv version of the PhysioNet/Computing in Cardiology Challenge 2019, which is concerned with predicting sepsis diagnosis from clinical data.
Below is a link to the challenge:
https://archive.physionet.org/challenge/2019/
The dataset comprises:
Vital signs (columns 1-8)
Laboratory tests (columns 9-34)
Demographics (columns 35-38)
Outcome columns (column 39)
The dataset consists of repeated measurements of the vital signs, laboratory test results and outcomes 27186 patients recorded over 24-48 hours. In this assignment, you will complete and submit a Jupyter notebook containing an end-to-end supervised learning pipeline using the XGBoost algorithm to predict sepsis diagnosis from aggregates of patient vital signs and laboratory test results given in the PneumoniaTimeSeries.csv dataset.
Description of the Fields
The Patient_id field contains the patient identifiers. The outcome field (SepsisLabel) contains the outcome diagnosis for each patient. The fields in each record are described in Table 1 of the challenge page: https://physionet.org/content/challenge-2019/1.0.0/
Instructions:
Task 1: Understand & Cleaning the Data (22 marks)
1.Examining the distribution of the features in the dataset can identify potential errors and discrepancies. Use boxplots to visually inspect the data for nonsensical values, replacing any erroneous-looking values found the average value of the feature (12 marks)
2.The next task is to use summary statistics & plots to understand the data properties and describe its structure & distribution with respect to demographical features, outcomes etc…e.g. how are the demographics (e. g. gender) and other variables (e.g. blood pressure) distributed across sepsis categories (i.e. sepsis = 0 and sepsis = 1) (10 marks)
Task 2: Data Processing (20 marks)
The dataset comprises a time-series of patient measurements over 24-48 hours. Prior to implementing a supervised learning model, it is essential to transform the data into an appropriate matrix that can be fed into a supervised learning algorithm. We will be using XGBoost. Like most other supervised ML algorithms we covered in our module, XGBoost does not take into account the temporal nature of the data. Moreover, solving the time-series classification will require more time than allocated to your assignment. Therefore, we will start with aggregating the measurements so that each person will end up with one row in the dataset. Each patient will then have one outcome point in the outcome vector y.
3.For each measurement m (e.g. heart rate), you will extract Δm, which measures the difference between the value of the measurement during hour 1 and the value of the measurement during the last hour recorded for the patient (adjusted to accommodate possible missingness during the two intervals). What is the dimension of the resulting matrix After the aggregation, perform any other preprocessing tasks you deem necessary to yield a dataset ready for the next steps.
Special Note on the outcome variable (SepsisLabel): The dataset documents the diagnosis of a patient at a given hour. For some patients, you will see that SepsisLabel starts with 0 then becomes 1 at a later stage. Because our aim is to distinguish patients who end up with a sepsis diagnosis from those who do not, it is sufficient for you to aggregate SepsisLabel to a value of 1 if the patient has 1 at any given hour, and 0 otherwise.
Task 3: Classifier Implementation (33 marks)
You will implement your classification system using the XGBoost algorithm.. XGBoost has several Python implementations, but is also a part of the sklearn library (sklearn.ensemble.GradientBoostingClassifier).
The following should be adhered to while building your model:
Hyperparameter Tuning (12 marks)
4.Use XGBoost’s documentation to identify the model’s hyperparameters.
5.By understanding the meaning of each of XGBoost’s hyperparameters, identify a tentative space of candidate hyperparameters you’d like to run the model on, as well as their potential values.
6.Hyperparameter tuning should be performed via a grid search over the classifier’s hyperparameter space.
Data Resampling (21 marks)
7.Data resampling techniques used should be a) appropriate and b) justified during the different stages of the pipeline.
Task 4: Performance & Interpretation (25 marks)
Performance Metrics (10 marks)
8.Evaluate the model’s performance using the confusion matrix while reporting the following performance metrics: accuracy, sensitivity (recall), specificity, precision, and f1-score. Plot the ROC figure.
9.Is there any discrepancy between the performance metrics What is the applicability of each metric to the problem What does each performance metric tell us about out model Explain in details.
Understanding the classifier’s output (15 marks)
10.How does the model perform Can you explain the prediction on a global and a local level (using the Shapley library and/or variable importance measures) Comment on the robustness of the interpretation generated.
Note: You will not be penalised if the performance is not great. However, it is essential that you follow a stepped & disciplined approach to developing your pipeline. Unjustified, unexplored, or missed decision points will be penalised.
What to Submit:
You are to submit to files:
1.A well-documented and fully executable Jupyter notebook containing the final Machine Learning pipeline.
2.An HTML version of your Jupyter notebook.
3.A report 1000-1300 word report detailing the following:
a.Aim of your pipeline
b.Description of the dataset with respect to the outcome and features.
c.All processing steps
d.XGBoost implementation, hyperparmetertuning and class imbalance techniques implemented
e.Model evaluation and interpretation.
f.Conclusions
Marking Scheme:
The total marks are 100 and are distributed across the sections as described in each task’s section heading. The marks for each section will be equally distributed among the following requirements (each requirement will be checked against the Jupyter notebook or the final report as underlined here):
1.Correct, complete and well-documented code (Jupyter Notebook). All documentation is to be performed in Jupyter Notebook markdown cells. A few comments in code cells are acceptable where necessity is obvious.
2.Minimalist-style coding (Jupyter Notebook).
a.No redundant lines or ambiguous variable names.
b.Use loops to iterate over features instead of processing each feature individually.
c.Points will be deducted when unnecessary repetition of steps due to the lack of control structures (e.g. repeating the same step for all features instead of using a loop control structure).
3.Detailed description and Justification of the task (Report). Each step must be fully described in detail and justified. E.g. why did you select the set of hyperparameters shown in the notebook How did you what functionality did you use to aggregate the features Etc..
4.Introduction, Reflection and Summary/Conclusion (Report): Your report must be self-contained. It should present a fully-specified document that can be used to understand your project in full. Start your report with an introductory paragraph to put the reader in context. At each step, implications of your design choices and implementation must be examined analytically at each step.
5.Visualisation (Report+Jupyter Notebook). You may duplicate figures in your report for readability purposes. Figures must be fully described via a comprehensive caption.
The table below shows a tentative marking rubric that can guide you through the design, implementation, discussion and reflection. Please note that the report serves to support your submitted ‘pipelines’ and is therefore an integral part of it. Generally:
– You will not get report marks if your code does not work.
– You will get code marks if your report is non-existent. In this scenario, you may also getreport marks if your jupyter notebook contains sufficient information to understand your design choices and interpret your results.
Section
Excellent
5
Very Good
4
Good
3
Room for Improve-ment
2 Significant Room for Improve-ment or Section is Missing 1
Design and implementation -Clear, correct and well-justified design choices
– provision of a full account of the choices made based on material covered or online sources.
-Full acknowledgement of missing steps or unexplored design choices, fully justifying the absence and reflection on the effect on the results.
– Compact, clear and well-documented fully-working code.
-Clear, correct and well-justified design choices
– provision of a full account of the choices made based on material covered or online sources.
-Some acknowledgement of unexplored design choices and their effect on the results.
-Clear and well-documented fully-working code.
-Clearly-stated design choices with some justification of design choices.
– provision of some account of the choices made based on material covered or online sources.
-Some acknowledgement of unexplored design choices and some discussion on their effect on the results.
–Clear and well-documented fully-working code.
-Design choices ambiguous or unjustified.
-No acknowledgement of unexplored design choices and their effect on the results.
–code unclear, not well-documented but fully working. -Design choices are neither stated nor unjustified.
– Code unclear, not well-documented and erroneous.
Results and reflection – Well-documented reflection on the generated output; provision of a full critical assessment of the steps taken and relation to the output produced is discussed.
– Results are fully described and contextualized and possible explanations for unexpected results are given, using appropriate technical and scientific language.
– Results are fully and comprehensively described using accurate technical and scientific language.
– Some critical assessment and reflection on the generated output;
– Results are fully described using at least some appropriate technical and scientific language.
– Unexpected results are acknowledged.
– Minimal critical assessment of the results.
– Results are fully described. May be slightly lacking in completeness, clarity or use of technical language.
– Results are partially described or no critical assessment is provided
– Technical language may be used inappropriately. A description of the results is missing, incomplete or incomprehensible and wrong use of technical language in presenting the results.


发表评论