31250 Ass3

Content

1. A description of the data mining problem;

2. The data preprocessing and transformations you did (if any);

3. How you went about solving the problem.

4. Classification techniques used and summary of the results and parameter settings.

5. The best classifier that you selected – the type, its performance, how it solved the problem (if it makes sense for that type of classifier), and reasons for selecting it.

6. Kaggle Submission Score

1. A description of the data mining problem.

Assessment 3 was tasked with building a classifier to classify “AIRLINE-NAME” that could be used in the data exploration process based on Assessment 2. In this task, but Knime is also a good fit for the visual programming software of this report, so I chose Knime for data mining and prediction. Data mining can not only help analysts understand data, but also help analysts make predictions and judgments through visual analysis and data mining results.

This task provides three data sets: FlightDataset. CSV, UnknownDataset. CSV and Kaggle-Submission-Random- sample.csv. The “training data” data set will be used to predict the data contained in the “Airline_Name” column. After classifying the data set, the prediction result will become the unknown data set of the prediction column, and the generated file format is the same as that of the kaggle-submitt-random-sample.csv data set.

The task requires the use of a variety of methods to predict the data and determine the maximum prediction accuracy of the method, and I will use several methods to predict.

2. The data preprocessing and transformations you did (if any);

First of all, use the CSV Reader node to import the CSV format file we want to use. Then use Rule Engine to achieve the conversion we need. The third step is the “Number to String” node to convert the number type to the string type. The picture above is what I set, I converted the data of Airline Name to string type, because the learner only deals with columns of string type.

I set the training set to 70%, the test set to 30%, and choose stratified sampling and random seed. Setting the random seed will make the sequence number the same every time, and will not cause the samples to be different, and the results of each validation will not be Change.

4. Classification techniques used and summary of the results and parameter settings.

Decision tree induction

A decision tree is a tree structure, each internal node represents an attribute test, each branch represents a test output, and each leaf node represents a category, which is a graphical method through probability analysis. In the above diagram, it shows how the decision tree classifier works and outputs the results to the scorer node and the ROC curve node.

The above three pictures are the configuration of the decision tree learner. Because the goal is to predict the Airline Name, I chose the Airline Name of the class column in the pruning method. I chose MDL because trying to use MDL can effectively improve the prediction accuracy. Then I set the number of threads to 2 and changed the minimum records per node to 8 to further improve the accuracy. The configuration of the decision tree predictor, changing the name of the prediction column to something that matches the task conditions, and then rechecking the box with the other columns of the normalized class distribution, it shows the normalized distribution for each prediction. and the ROC Curve node, I chose P (Airline_Name), which represents the predicted value of “Airline_Name”. Shows the results of equally split nodes, the accuracy of the decision tree model, and the accuracy of the error prediction. 

Naive Bayes Predictor

This is a classification technique based on Bayes’ theorem, assuming independence among predictors. In short, a naive Bayesian classifier assumes that the presence of a particular feature in a class is independent of the presence of any other feature. Bayesian models are easy to build and are especially useful for very large datasets. It is well known that in addition to simplicity, Bayesian outperforms even very complex classification methods.

Random Forest

The Random Forest classifier used for classification is shown. Random Forest is composed of multiple decision trees, and the establishment of each tree relies on independent trees extracted independently. M samples are randomly selected from the original training sample set, and then new training samples are generated in the training decision tree set to form a random forest, and finally countless decision trees are generated. Random forest does not require pruning to improve accuracy, it can automatically handle unbalanced data samples, and has strong fit and stability. and got results.

Gradient Boosted Trees

The figure above shows the workflow of a gradient boosting tree classifier. This method reduces the predicted value, and by repeatedly changing the sample difference, the trained model has different decision-making capabilities. When training a model, the smaller the loss of prediction samples, the better, and the underlying model must be a decision tree. and gives the result.

5. The best classifier that you selected – the type, its performance, how it solved the problem (if it makes sense for that type of classifier), and reasons for selecting it.

Among these methods, I think the random forest classifier is the best of the six, so I optimized some of them. To get more accurate predictions, I performed round-robin operations in Random Forest Learner and Random Forest Predictor. For the UnknownData.csv filcreateI use a new file read node to import. Next, I use the same Number to String app node and Normalizer app node to make sure the same preprocessing method is used. I connected the normalize application node to the random forest predictor node and in the random forest predictor changed the predicted column name to Prediction (AIRLINE_NAME), used the column filter to select Row ID and Prediction (AIRLINE_NAME), then in the CSV Name the filename. Write and export files.

此图片的alt属性为空;文件名为428a5d076a5c26765b1d9471292d46cb.jpg

担心学业?你还有其他选择!

KJEssay 学年守护计划!

我们是全网首家积极根据新政策优化应对方案的论文服务机构!

全面升级给你最好的防护!

1、远程代劳,资料下载,作业提交,有需要全程代劳!

KJEssay已对目前主流的教学系统Blackboard、ReCap,以及各校的ePortfolio,对全体老师做过专项培训,这方面有困难的学生,可直接授意老师代劳,我们将为你全面服务!

2、考核考试,老师提前充分备考,同程协助,助力满分!

KJEssay 保障学业提供全面服务!专业老师团队先学习了解课程内容,做充足应对,设计方案,

考试时,老师,专业应急团队,客服,同时待命!

老师快速反应,迅速做出最佳答案以及思路!

应急团队集思广益可对重难点迅速突破!

客服居中,全面负责协调沟通,提高效率!

给予及时而效率的全面帮助!

3、远程上课,录屏打卡课程讨论一个不落!

针对目前在线网课,KJEssay做出专项研究,对包括Autodesk、Azure、Skype、Zoom等视频教学软件有着充分熟悉。上网打卡一个不落。

4、保障隐私安全,全程一人全面追踪服务!所有人均签有隐私合同!

全面服务将主要安排在一位老师全面负责,做好对信息情况的充足了解掌握,不假他手!更因为全程彻底的参与,对情况以及考试有更彻底的把握!更能依据情况做出应对!也更易获取更高分!

客服以及第三方,时刻追踪,定期反馈情况。

5、一举一动全面反馈!时刻监控,看得到的全过程!24小时客服待命!

我们一直把沟通反馈,放在重中之重!尤其是代理服务,最了解的肯定还是客户,所以KJEssay会反馈所有的情况,没有客户允许下,不擅专!不乱动!

以最安全的形式,保障拿到最好的成绩!

在上半年的全面代理中,现已取得了优异的成绩与效果。

此图片的alt属性为空;文件名为35bbe3364a2ce7c2de4d6716ce45b0aa.jpg
此图片的alt属性为空;文件名为d578155fe7f87334ab9d272e80498c2f.jpg
此图片的alt属性为空;文件名为2f1a120444c69fe7a267ab20060348ec.jpg
此图片的alt属性为空;文件名为21d615ec4d8503ac2380cc30732bee21.jpg
此图片的alt属性为空;文件名为200b367966acea40ae507d5cf05e1975.jpg
此图片的alt属性为空;文件名为74dfb41160baebf4cc5f10bdbc4a78bf.jpg
此图片的alt属性为空;文件名为5de508126b22823e6b5b8210a3dca3cf.jpg
此图片的alt属性为空;文件名为6622ac43c8ef56a4c3bdf5cd61fadc69.jpg
此图片的alt属性为空;文件名为6a29689ca0b185e515c419046ec22d0b.jpg
此图片的alt属性为空;文件名为30c701ceae69f917e42864927b0be21a.jpg
此图片的alt属性为空;文件名为12559bac08b9eea12ca431b874504d93.jpg
此图片的alt属性为空;文件名为fcfb0df3fc9bea3e620604f3b628b4a9.jpg
此图片的alt属性为空;文件名为aabedf10bb2cc61671dc64c733ef3b72.jpg
此图片的alt属性为空;文件名为83d41a3f3b3be772e336c842211ed29f.jpg
此图片的alt属性为空;文件名为1b7aff9f394e28fe9016715e48d344eb.jpg
此图片的alt属性为空;文件名为aed4df7939da959d2df596fc6aa534ac.jpg
此图片的alt属性为空;文件名为635115284cf3fb723f9dd2152a8f825f.jpg

新学期,我们应对留学网课,更有经验,更加从容!

关于KJEssay

我们是KJEssay,31639人的选择!

此图片的alt属性为空;文件名为1e348b0dbfb0fe9bc770d4fd66bf5ea4.jpg
此图片的alt属性为空;文件名为7e7391e590eccaf19579cfa2fce375fe.jpg
此图片的alt属性为空;文件名为50c979b700996121b378e00463d9a33c.jpg
此图片的alt属性为空;文件名为15c412569e730f451ed57fe6212af2ec.jpg

现在就可联系我们

此图片的alt属性为空;文件名为c7bcbc86b3a922d4b4fea9d23e58059b.jpg

微信->添加朋友->添加企业微信联系人:13262280223

官网:https://www.kjessay.com

邮箱:kaijiewrite@163.com  service@kjessay.com

WhatsApp:+44 7410496844(推荐添加)

QQ:1483266981

立即联系我们参与活动吧~

了解 KJESSAY历史案例 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读