INF6028 Mining and Evaluating a Structured Dataset

联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981

INF6028 Coursework 2022/23
Mining and Evaluating a Structured Dataset
1. Introduction
The assessment for INF6028 Data Mining consists of a single piece of individual coursework to assess your
ability to understand key data mining, analysis and evaluation concepts. You will be assigned a single
dataset with an associated data mining problem to solve (e.g., a regression problem). You should first use
data exploration techniques to explore the data, conduct appropriate data preparation, and then choose
two supervised data mining techniques available in KNIME to predict certain data values and evaluate and
compare their performance. You will need to select appropriate techniques, justify your choices made at
different stages of your workflow, and demonstrate that you have knowledge of the necessary underlying
data mining techniques.
You should write a 2,500 word structured report (see Section 3) that includes the following headings
(more details on how the report will be assessed are provided below):
Introduction – introduce the prediction problem.
Data mining theory – provide a theoretical description of the two supervised data mining methods
used in the workflow (for example, the classification or regression techniques that have been used),
why they are appropriate to the prediction task, and how their performance can be assessed. This
should include citations to relevant prior literature.
Data exploration and preparation – describe the approaches used in the workflow to explore the
data; and perform feature selection, transformation and normalisation, where appropriate.
Experimental setup – describe the experimental setup and the evaluation measures used in the
workflow and how the data has been handled to ensure that the models were not over-fitted. You
should explain which nodes were used in KNIME and provide a rationale for the various parameter
settings that were used. You should not, however, simply list all the modules in your workflow and
their parameters – be selective and discuss the modules most critical to solving the data mining task.
Results – present the results for each data mining method and compare the performance of the
different methods using graphical and tabular methods. What insights can you gain from the
models For example, which are the most important features, are there any outliers in the
predictions
Conclusion and reflections – summarise the main findings of your report and reflect on the methods
used.
Charts and tables (and their associated captions), references and appendices are not included in the word
count.
Remember: your report should be a critical evaluation of the workflow in the context of the data mining
problem posed, it should not be merely a description of what was done.
This assessment is worth 100% of the overall module mark for INF6028. A pass mark of 50 is required to pass
the module. Submission deadline: 1
st of June at 4 PM, via Turnitin. See Section 4 for more general
information about Coursework Submission Requirements within the Information School.
2. The Datasets
You will choose a single dataset to base your analyses and report on. Please choose one of the two datasets
below and ensure before you start working on the assessment that you are using the correct dataset.
The datasets have been derived from Kaggle competitions and are downloadable from Blackboard in the
Assessment section. A brief description of the attributes in each dataset is given at the end of this document.
Note that in both cases the data are different to the standard Kaggle datasets – they have been
extensively modified for this year’s run of INF6028. Do not attempt to use the datasets from Kaggle or
to use/copy any of the workbooks available there – this would constitute unfair means
Titanic Dataset (Binary Classification)
The data is split across two files, each of which contains 1,204 entries representing 1,204 passengers,
although it should be noted that the passengers are not necessarily the same in the two files. The two files
are titanic_ticket_data.csv and titanic_personal_data.csv
The aim of this challenge is to build a model that is able to predict whether or not a passenger will survive
the sinking of the Titanic.
Song Popularity Dataset (Regression)
The data is split across two files, each of which contains 603 entries representing 603 popular songs from
the Spotify platform. The two files are song_details.csv and song_acoustic_analysis.csv.
The aim of the challenge is to build a model to predict the popularity of each song on Spotify.
3. Report Structure
You are required to produce a structured report that includes all the sections detailed in Table 1. You must
state the word count somewhere in the report. As there is a word count limit you should aim to make your
writing as concise and informative as possible. The emphasis of the report should be on the clarity, accuracy
and quality in communicating your findings. Where helpful, you may wish to state specifically which KNIME
nodes you have used but you should avoid simply listing nodes used and their settings – be selective.
Table 1: Required content of the structured report.
Section Description Maximum allocated marks
Structured abstract This should provide a summary of your report in
a structured manner. This is not included in the
word count.
Required, but 0 marks
Introduction This section should introduce the data mining
task that is addressed in the report. You should
10 marks
indicate the property/data value that is
predicted and give a brief overview of the
dataset and methods used.
Data Mining Theory This section should provide an overview of the
algorithms for predictive data mining used in
the workflow from a theoretical aspect. Explain
why they are relevant to the prediction problem.
Support your rationale by providing references
to the literature where the techniques have been
applied to similar problems.
Include a short discussion of the most
appropriate methods for evaluating the
performance of these data mining methods.
25 marks
Data Exploration
and Preparation
This section should provide a brief description of
the data and of the approaches used to pre_x005f process the data. You should present an
investigation of the attributes (including the
data value to be predicted) and describe any
data cleaning employed, including handling of
missing data, data transformations and data
aggregations.
10 marks
Experimental Setup This section should describe the experimental
design in the workflow.
You should describe the process followed in
order to find the best performing model for each
method and how this was validated.
For example, which KNIME nodes were used
How were they configured Was any cross validation or a separate validation set used and
why
20 marks
Results and
Discussion
Present the results of the data mining process
including the results of experiments to find the
best model for each data mining method.
Compare the best performance of the different
methods and, if appropriate, consider which
attribute contributes most to each model.
Discuss the advantages and disadvantages of
the data mining methods. Which of the chosen
methods produced the best model and why
20 marks
Conclusion and
reflections
Summarise the main findings of the analysis and
reflect on the choice of methods for the
problem, for example, how might the models be
improved with hindsight Use evidence from the
literature to support your arguments.
15 marks
KNIME workflow You should submit your KNIME workflow(s) as a
.knwf or .knar file. Note that this can consist of
separate workflows but they should all be saved
to one file. Include your best setup for each data
mining method.
Required, but 0 marks.
Note that 5 marks will be
deducted if this is not submitted
and it may make it difficult for
your marker to assess your work.
Appendix 1: Information School Coursework Submission Requirements
It is the student’s responsibility to ensure no aspect of their work is plagiarised or the result of other unfair
means. The University’s and Information School’s Advice on unfair means can be found in your Student
Handbook, available via http://www.sheffield.ac.uk/is/current
Your assignment has a word count limit. A deduction of 3 marks will be applied for coursework that is 10% or
more above or below the word count as specified above or that does not state the word count.
It is your responsibility to ensure your coursework is correctly submitted before the deadline. It is highly
recommended that you submit well before the deadline. Coursework submitted after 10am on the stated
submission date will result in a deduction of 5% of the mark awarded for each working day after the submission
date/time up to a maximum of 5 working days, where ‘working day’ includes Monday to Friday (excluding public
holidays) and runs from 10am to 10am. Coursework submitted after the maximum period will receive zero marks.
Work submitted electronically, including through Turnitin, should be reviewed to ensure it appears as you
intended.
Before the submission deadline, you can submit coursework to Turnitin numerous times. Each submission will
overwrite the previous submission. Only your most recent submission will be assessed. However, after the
submission deadline, the coursework can only be submitted once.
Details about the submission of work via Turnitin can be found at http://youtu.be/C_wO9vHHheo
If you encounter any problems during the electronic submission of your coursework, you should immediately
contact the module coordinator and one of the Information School Teaching Support Team is-teaching support@shef.ac.uk (Julie Priestley 0114 2222839). This does not negate your responsibilities to submit your
coursework on time and correctly.
Appendix 2: Titanic Dataset (Binary Classification)
The titanic data consist of two files that need to be merged.
titanic_ticket_data.csv consists of the following variables:
PassengerId: the identifier
Survived: the value to predict
Ticket: the Ticket Number
Fare: the passenger fare
Cabin: Cabin number
Embarked: Port of embarkation. C = Cherbourg, Q = Queenstown, S = Southampton
titanic_personal_data.csv consists of the following variables:
PassengerId – the identifier
Name: the name of the passenger
Sex: male or female
Age: Age is fractional if less than 1. If the age is estimated, is it in the form of xx.5
SibSp: number of siblings/spouses where family relations are defined as follows:
Sibling = brother, sister, stepbrother, stepsister
Spouse = husband, wife
Parch: number of parent/children where family relations are defined as follows:
Parent = mother, father;
Child = daughter, son, stepdaughter, stepson.
Some children travelled only with a nanny, therefore parch=0 for them
Salary: in dollars
Job: job title
Appendix 3: Song Popularity Dataset (Regression)
The song popularity data set consists of two files that need to be merged.
song_details.csv consists of the following variables:
id – the identifier
title – song title
artist – song artist
genre – song genre
year – release year
bpm – tempo in Beats Per Minute
dB – (average) loudness in decibells
popularity – the value to be predicted (the higher the value, the more popular the song is)
song_acoustic_analysis.csv consists of the following variables:
id – the identifier
energy – the energy of a song – the higher the value, the more energetic the song is. Typically, energetic
tracks feel fast, loud, and noisy.
danceability – how suitable a track is for dancing based on a combination of musical elements including
tempo, rhythm stability, beat strength, and overall regularity. A value of 0.0 is least danceable and 1.0 is
most danceable.
lineness – the higher the value, the more likely the song is a live recording.
valence – a measure from 0.0 to 1.0 describing the musical positiveness conveyed by a track. Tracks with
high valence sound more positive (e.g. happy, cheerful, euphoric), while tracks with low valence sound more
negative (e.g. sad, depressed, angry).
duration – the length of the recording in seconds.
acousticness – how acoustic the song is
speechiness – speechiness detects the presence of spoken words in a track. If the speechiness of a song is
above 0.66, it is probably made of spoken words, a score between 0.33 and 0.66 is a song that may contain
both music and words, and a score below 0.33 means the song does not have any speech.

发表评论

了解 KJESSAY历史案例 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读