
Content
1. A description of the data mining problem;
2. The data preprocessing and transformations you did (if any);
3. How you went about solving the problem.
4. Classification techniques used and summary of the results and parameter settings.
5. The best classifier that you selected – the type, its performance, how it solved the problem (if it makes sense for that type of classifier), and reasons for selecting it.
6. Kaggle Submission Score
1. A description of the data mining problem.
Assessment 3 was tasked with building a classifier to classify “AIRLINE-NAME” that could be used in the data exploration process based on Assessment 2. In this task, but Knime is also a good fit for the visual programming software of this report, so I chose Knime for data mining and prediction. Data mining can not only help analysts understand data, but also help analysts make predictions and judgments through visual analysis and data mining results.
This task provides three data sets: FlightDataset. CSV, UnknownDataset. CSV and Kaggle-Submission-Random- sample.csv. The “training data” data set will be used to predict the data contained in the “Airline_Name” column. After classifying the data set, the prediction result will become the unknown data set of the prediction column, and the generated file format is the same as that of the kaggle-submitt-random-sample.csv data set.
The task requires the use of a variety of methods to predict the data and determine the maximum prediction accuracy of the method, and I will use several methods to predict.
2. The data preprocessing and transformations you did (if any);
First of all, use the CSV Reader node to import the CSV format file we want to use. Then use Rule Engine to achieve the conversion we need. The third step is the “Number to String” node to convert the number type to the string type. The picture above is what I set, I converted the data of Airline Name to string type, because the learner only deals with columns of string type.
I set the training set to 70%, the test set to 30%, and choose stratified sampling and random seed. Setting the random seed will make the sequence number the same every time, and will not cause the samples to be different, and the results of each validation will not be Change.
4. Classification techniques used and summary of the results and parameter settings.
Decision tree induction
A decision tree is a tree structure, each internal node represents an attribute test, each branch represents a test output, and each leaf node represents a category, which is a graphical method through probability analysis. In the above diagram, it shows how the decision tree classifier works and outputs the results to the scorer node and the ROC curve node.
The above three pictures are the configuration of the decision tree learner. Because the goal is to predict the Airline Name, I chose the Airline Name of the class column in the pruning method. I chose MDL because trying to use MDL can effectively improve the prediction accuracy. Then I set the number of threads to 2 and changed the minimum records per node to 8 to further improve the accuracy. The configuration of the decision tree predictor, changing the name of the prediction column to something that matches the task conditions, and then rechecking the box with the other columns of the normalized class distribution, it shows the normalized distribution for each prediction. and the ROC Curve node, I chose P (Airline_Name), which represents the predicted value of “Airline_Name”. Shows the results of equally split nodes, the accuracy of the decision tree model, and the accuracy of the error prediction.
Naive Bayes Predictor
This is a classification technique based on Bayes’ theorem, assuming independence among predictors. In short, a naive Bayesian classifier assumes that the presence of a particular feature in a class is independent of the presence of any other feature. Bayesian models are easy to build and are especially useful for very large datasets. It is well known that in addition to simplicity, Bayesian outperforms even very complex classification methods.
Random Forest
The Random Forest classifier used for classification is shown. Random Forest is composed of multiple decision trees, and the establishment of each tree relies on independent trees extracted independently. M samples are randomly selected from the original training sample set, and then new training samples are generated in the training decision tree set to form a random forest, and finally countless decision trees are generated. Random forest does not require pruning to improve accuracy, it can automatically handle unbalanced data samples, and has strong fit and stability. and got results.
Gradient Boosted Trees
The figure above shows the workflow of a gradient boosting tree classifier. This method reduces the predicted value, and by repeatedly changing the sample difference, the trained model has different decision-making capabilities. When training a model, the smaller the loss of prediction samples, the better, and the underlying model must be a decision tree. and gives the result.
5. The best classifier that you selected – the type, its performance, how it solved the problem (if it makes sense for that type of classifier), and reasons for selecting it.
Among these methods, I think the random forest classifier is the best of the six, so I optimized some of them. To get more accurate predictions, I performed round-robin operations in Random Forest Learner and Random Forest Predictor. For the UnknownData.csv filcreateI use a new file read node to import. Next, I use the same Number to String app node and Normalizer app node to make sure the same preprocessing method is used. I connected the normalize application node to the random forest predictor node and in the random forest predictor changed the predicted column name to Prediction (AIRLINE_NAME), used the column filter to select Row ID and Prediction (AIRLINE_NAME), then in the CSV Name the filename. Write and export files.

担心学业?你还有其他选择!
KJEssay 学年守护计划!
我们是全网首家积极根据新政策优化应对方案的论文服务机构!
全面升级给你最好的防护!
1、远程代劳,资料下载,作业提交,有需要全程代劳!
KJEssay已对目前主流的教学系统Blackboard、ReCap,以及各校的ePortfolio,对全体老师做过专项培训,这方面有困难的学生,可直接授意老师代劳,我们将为你全面服务!
2、考核考试,老师提前充分备考,同程协助,助力满分!
KJEssay 保障学业提供全面服务!专业老师团队先学习了解课程内容,做充足应对,设计方案,
考试时,老师,专业应急团队,客服,同时待命!
老师快速反应,迅速做出最佳答案以及思路!
应急团队集思广益可对重难点迅速突破!
客服居中,全面负责协调沟通,提高效率!
给予及时而效率的全面帮助!
3、远程上课,录屏打卡课程讨论一个不落!
针对目前在线网课,KJEssay做出专项研究,对包括Autodesk、Azure、Skype、Zoom等视频教学软件有着充分熟悉。上网打卡一个不落。
4、保障隐私安全,全程一人全面追踪服务!所有人均签有隐私合同!
全面服务将主要安排在一位老师全面负责,做好对信息情况的充足了解掌握,不假他手!更因为全程彻底的参与,对情况以及考试有更彻底的把握!更能依据情况做出应对!也更易获取更高分!
客服以及第三方,时刻追踪,定期反馈情况。
5、一举一动全面反馈!时刻监控,看得到的全过程!24小时客服待命!
我们一直把沟通反馈,放在重中之重!尤其是代理服务,最了解的肯定还是客户,所以KJEssay会反馈所有的情况,没有客户允许下,不擅专!不乱动!
以最安全的形式,保障拿到最好的成绩!
在上半年的全面代理中,现已取得了优异的成绩与效果。

















新学期,我们应对留学网课,更有经验,更加从容!
关于KJEssay
我们是KJEssay,31639人的选择!




现在就可联系我们

微信->添加朋友->添加企业微信联系人:13262280223
官网:https://www.kjessay.com
邮箱:kaijiewrite@163.com service@kjessay.com
WhatsApp:+44 7410496844(推荐添加)
QQ:1483266981
立即联系我们参与活动吧~

