Data Mining: Practical Machine Learning Tools and ... - LIDeCC

More documents

Recommendations

Info

90% used in 10-fold cross-validation. To compensate for this, we combine thetest-set error rate with the resubstitution error on the instances in the trainingset. The resubstitution figure, as we warned earlier, gives a very optimistic estimateof the true error and should certainly not be used as an error figure on itsown. But the bootstrap procedure combines it with the test error rate to give afinal estimate e as follows:e = 0. 632 ¥ e + 0. 368 ¥ e .test instances5.5 COMPARING DATA MINING METHODS 153training instancesThen, the whole bootstrap procedure is repeated several times, with differentreplacement samples for the training set, and the results averaged.The bootstrap procedure may be the best way of estimating error for verysmall datasets. However, like leave-one-out cross-validation, it has disadvantagesthat can be illustrated by considering a special, artificial situation. In fact, thevery dataset we considered previously will do: a completely random dataset withtwo classes. The true error rate is 50% for any prediction rule. But a scheme thatmemorized the training set would give a perfect resubstitution score of 100%so that e training instances = 0, and the 0.632 bootstrap will mix this in with a weightof 0.368 to give an overall error rate of only 31.6% (0.632 ¥ 50% + 0.368 ¥ 0%),which is misleadingly optimistic.5.5 Comparing data mining methodsWe often need to compare two different learning methods on the same problemto see which is the better one to use. It seems simple: estimate the error usingcross-validation (or any other suitable estimation procedure), perhaps repeatedseveral times, and choose the scheme whose estimate is smaller. This is quitesufficient in many practical applications: if one method has a lower estimatederror than another on a particular dataset, the best we can do is to use the formermethod’s model. However, it may be that the difference is simply caused by estimationerror, and in some circumstances it is important to determine whetherone scheme is really better than another on a particular problem. This is a standardchallenge for machine learning researchers. If a new learning algorithm isproposed, its proponents must show that it improves on the state of the art forthe problem at hand and demonstrate that the observed improvement is notjust a chance effect in the estimation process.This is a job for a statistical test that gives confidence bounds, the kind wemet previously when trying to predict true performance from a given test-seterror rate. If there were unlimited data, we could use a large amount for trainingand evaluate performance on a large independent test set, obtaining confidencebounds just as before. However, if the difference turns out to be significantwe must ensure that this is not just because of the particular dataset we
154 CHAPTER 5 | CREDIBILITY: EVALUATING WHAT’S BEEN LEARNEDhappened to base the experiment on. What we want to determine is whetherone scheme is better or worse than another on average, across all possible trainingand test datasets that can be drawn from the domain. Because the amountof training data naturally affects performance, all datasets should be the samesize: indeed, the experiment might be repeated with different sizes to obtain alearning curve.For the moment, assume that the supply of data is unlimited. For definiteness,suppose that cross-validation is being used to obtain the error estimates(other estimators, such as repeated cross-validation, are equally viable). For eachlearning method we can draw several datasets of the same size, obtain an accuracyestimate for each dataset using cross-validation, and compute the mean ofthe estimates. Each cross-validation experiment yields a different, independenterror estimate. What we are interested in is the mean accuracy across all possibledatasets of the same size, and whether this mean is greater for one schemeor the other.From this point of view, we are trying to determine whether the mean ofa set of samples—cross-validation estimates for the various datasets that wesampled from the domain—is significantly greater than, or significantly lessthan, the mean of another. This is a job for a statistical device known as the t-test, or Student’s t-test. Because the same cross-validation experiment can beused for both learning methods to obtain a matched pair of results for eachdataset, a more sensitive version of the t-test known as a paired t-test can beused.We need some notation. There is a set of samples x 1 , x 2 ,...,x k obtained bysuccessive 10-fold cross-validations using one learning scheme, and a second setof samples y 1 , y 2 ,...,y k obtained by successive 10-fold cross-validations usingthe other. Each cross-validation estimate is generated using a different dataset(but all datasets are of the same size and from the same domain). We will getthe best results if exactly the same cross-validation partitions are used for bothschemes so that x 1 and y 1 are obtained using the same cross-validation split, asare x 2 and y 2 , and so on. Denote the mean of the first set of samples by x – andthe mean of the second set by y – . We are trying to determine whether x – is significantlydifferent from y – .If there are enough samples, the mean (x – ) of a set of independent samples(x 1 , x 2 ,...,x k ) has a normal (i.e., Gaussian) distribution, regardless of the distributionunderlying the samples themselves. We will call the true value of themean m. If we knew the variance of that normal distribution, so that it could bereduced to have zero mean and unit variance, we could obtain confidence limitson m given the mean of the samples (x – ). However, the variance is unknown, andthe only way we can obtain it is to estimate it from the set of samples.That is not hard to do. The variance of x – can be estimated by dividing thevariance calculated from the samples x 1 , x 2 ,...,x k —call it s 2 x—by k. But the
Page 2:
Data MiningPractical Machine Learni
Page 5 and 6:
Publisher:Publishing Services Manag
Page 7 and 8:
viFOREWORDThis book presents this n
Page 10 and 11:
CONTENTSix4 Algorithms: The basic m
Page 12 and 13:
CONTENTSxiGenerating good rules 202
Page 14 and 15:
CONTENTSxiii8 Moving on: Extensions
Page 16:
CONTENTSxv13 The command-line inter
Page 19 and 20:
xviiiLIST OF FIGURESFigure 4.10 The
Page 21 and 22:
xxLIST OF FIGURESFigure 10.13 Worki
Page 23 and 24:
xxiiLIST OF TABLESTable 5.2 Confide
Page 25 and 26:
xxivPREFACEalchemy. Instead, there
Page 27 and 28:
xxviPREFACEwho interprets them, and
Page 29 and 30:
xxviiiPREFACEin Section 6.3. We hav
Page 31 and 32:
xxxPREFACEration. All who have work
Page 34:
partIMachine Learning Toolsand Tech
Page 37 and 38:
4 CHAPTER 1 | WHAT’S IT ALL ABOUT
Page 39 and 40:
Page 41 and 42:
Page 43 and 44:
10 CHAPTER 1 | WHAT’S IT ALL ABOU
Page 45 and 46:
Page 47 and 48:
Page 49 and 50:
Page 51 and 52:
Page 53 and 54:
Page 55 and 56:
Page 57 and 58:
Page 59 and 60:
Page 61 and 62:
Page 63 and 64:
Page 65 and 66:
Page 67 and 68:
Page 69 and 70:
Page 71 and 72:
Page 74 and 75:
chapter 2Input:Concepts, Instances,
Page 76 and 77:
2.1 WHAT’S A CONCEPT? 43increase
Page 78 and 79:
2.2 WHAT’S IN AN EXAMPLE? 45data
Page 80 and 81:
2.2 WHAT’S IN AN EXAMPLE? 47Table
Page 82 and 83:
2.3 WHAT’S IN AN ATTRIBUTE? 49Tab
Page 84 and 85:
2.3 WHAT’S IN AN ATTRIBUTE? 51Not
Page 86 and 87:
2.4 PREPARING THE INPUT 53cleaned u
Page 88 and 89:
missing values in this dataset). Th
Page 90 and 91:
dividing by the range between the m
Page 92 and 93:
2.4 PREPARING THE INPUT 59a record
Page 94 and 95:
chapter 3Output:Knowledge Represent
Page 96 and 97:
3.2 DECISION TREES 63Alternatively,
Page 98 and 99:
3.3 CLASSIFICATION RULES 65nated by
Page 100 and 101:
3.3 CLASSIFICATION RULES 671abx = 1
Page 102 and 103:
3.4 ASSOCIATION RULES 69yes, then i
Page 104 and 105:
3.5 RULES WITH EXCEPTIONS 71If peta
Page 106 and 107:
3.6 RULES INVOLVING RELATIONS 73ica
Page 108 and 109:
Standard relations include equality
Page 110 and 111:
PRP =-56.1+0.049 MYCT+0.015 MMIN+0.
Page 112 and 113:
3.8 INSTANCE-BASED REPRESENTATION 7
Page 114 and 115:
3.9 Clusters3.9 CLUSTERS 81When clu
Page 116 and 117:
chapter 4Algorithms:The Basic Metho
Page 118 and 119:
4.1 INFERRING RUDIMENTARY RULES 85F
Page 120 and 121:
described overfitting-avoidance bia
Page 122 and 123:
4.2 STATISTICAL MODELING 89Table 4.
Page 124 and 125:
just as we calculated previously. A
Page 126 and 127:
4.2 STATISTICAL MODELING 93Table 4.
Page 128 and 129:
of a document. Instead, a document
Page 130 and 131:
4.3 DIVIDE-AND-CONQUER: CONSTRUCTIN
Page 132 and 133:
Page 134 and 135:
Page 136 and 137: 4.3 DIVIDE-AND-CONQUER: CONSTRUCTIN
Page 138 and 139: 4.4 COVERING ALGORITHMS: CONSTRUCTI
Page 146 and 147: 4.5 MINING ASSOCIATION RULES 113acc
Page 148 and 149: 4.5 MINING ASSOCIATION RULES 115Tab
Page 150 and 151: 4.5 MINING ASSOCIATION RULES 117whi
Page 152 and 153: 4.6 LINEAR MODELS 119through the da
Page 154 and 155: 4.6 LINEAR MODELS 121However, linea
Page 156 and 157: 4.6 LINEAR MODELS 123n( i)i iÂ 1-x
Page 158 and 159: 4.6 LINEAR MODELS 125Set all weight
Page 160 and 161: 4.6 LINEAR MODELS 127While some ins
Page 162 and 163: 4.7 INSTANCE-BASED LEARNING 129When
Page 164 and 165: 4.7 INSTANCE-BASED LEARNING 131Figu
Page 170 and 171: 4.8 CLUSTERING 137As we saw in Sect
Page 172 and 173: 4.9 FURTHER READING 139can be updat
Page 174 and 175: 4.9 FURTHER READING 141Bayes was an
Page 176 and 177: chapter 5Credibility:Evaluating Wha
Page 178 and 179: 5.1 TRAINING AND TESTING 145of each
Page 180 and 181: ather than error rate, so this corr
Page 182 and 183: 5.3 CROSS-VALIDATION 149mediate con
Page 184 and 185: 5.4 OTHER ESTIMATES 151A single 10-
Page 188 and 189: 5.5 COMPARING DATA MINING METHODS 1
Page 190 and 191: 5.6 PREDICTING PROBABILITIES 157In
Page 192 and 193: 5.6 PREDICTING PROBABILITIES 159whe
Page 194 and 195: mental job expected of a loss funct
Page 196 and 197: y the total number of positives, wh
Page 198 and 199: 5.7 COUNTING THE COST 165different
Page 200 and 201: 5.7 COUNTING THE COST 167Table 5.6D
Page 202 and 203: 5.7 COUNTING THE COST 169100%80%tru
Page 204 and 205: 5.7 COUNTING THE COST 171should cho
Page 206 and 207: Different terms are used in differe
Page 208 and 209: 5.7 COUNTING THE COST 1750.5Anormal
Page 210 and 211: 5.8 EVALUATING NUMERIC PREDICTION 1
Page 212 and 213: 5.9 THE MINIMUM DESCRIPTION LENGTH
Page 214 and 215: 5.9 THE MINIMUM DESCRIPTION LENGTH
Page 216 and 217: 5.10 APPLYING THE MDL PRINCIPLE TO
Page 218: 5.11 FURTHER READING 185tion theory
Page 221 and 222: 188 CHAPTER 6 | IMPLEMENTATIONS: RE
Page 237 and 238:
204 CHAPTER 6 | IMPLEMENTATIONS: RE
Page 239:
Page 243 and 244:
Page 245 and 246:
Page 247 and 248:
Page 249 and 250:
Page 251 and 252:
Page 253 and 254:
Page 255 and 256:
Page 257 and 258:
Page 259 and 260:
Page 261 and 262:
Page 263 and 264:
Page 265 and 266:
Page 267 and 268:
Page 269 and 270:
Page 271 and 272:
Page 273 and 274:
Page 275 and 276:
Page 277 and 278:
Page 279 and 280:
Page 281 and 282:
Page 283 and 284:
Page 285 and 286:
Page 287 and 288:
Page 289 and 290:
Page 291 and 292:
Page 293 and 294:
Page 295 and 296:
Page 297 and 298:
Page 299 and 300:
Page 301 and 302:
Page 303 and 304:
Page 305 and 306:
Page 307 and 308:
Page 309 and 310:
Page 311 and 312:
Page 313 and 314:
Page 315 and 316:
Page 318 and 319:
chapter 7Transformations:Engineerin
Page 320 and 321:
7.1 ATTRIBUTE SELECTION 287attribut
Page 322 and 323:
7.1 ATTRIBUTE SELECTION 289and less
Page 324 and 325:
tion—and it is much easier to und
Page 326 and 327:
7.1 ATTRIBUTE SELECTION 293outlook
Page 328 and 329:
7.1 ATTRIBUTE SELECTION 295the t-te
Page 330 and 331:
7.2 DISCRETIZING NUMERIC ATTRIBUTES
Page 332 and 333:
Page 334 and 335:
Page 336 and 337:
Page 338 and 339:
7.3 SOME USEFUL TRANSFORMATIONS 305
Page 340 and 341:
Page 342 and 343:
Page 344 and 345:
Page 346 and 347:
7.4 AUTOMATIC DATA CLEANSING 313Int
Page 348 and 349:
7.5 COMBINING MULTIPLE MODELS 315da
Page 350 and 351:
7.5 COMBINING MULTIPLE MODELS 317pa
Page 352 and 353:
7.5 COMBINING MULTIPLE MODELS 319mo
Page 354 and 355:
7.5 COMBINING MULTIPLE MODELS 321Ra
Page 356 and 357:
7.5 COMBINING MULTIPLE MODELS 323Ho
Page 358 and 359:
7.5 COMBINING MULTIPLE MODELS 325Th
Page 360 and 361:
7.5 COMBINING MULTIPLE MODELS 327of
Page 362 and 363:
7.5 COMBINING MULTIPLE MODELS 329ou
Page 364 and 365:
7.5 COMBINING MULTIPLE MODELS 331ev
Page 366 and 367:
7.5 COMBINING MULTIPLE MODELS 333be
Page 368 and 369:
7.5 COMBINING MULTIPLE MODELS 335Ta
Page 370 and 371:
7.6 Using unlabeled data7.6 USING U
Page 372 and 373:
7.6 USING UNLABELED DATA 339automat
Page 374 and 375:
7.7 FURTHER READING 341deal with we
Page 376:
7.7 FURTHER READING 343Domingos (19
Page 379 and 380:
346 CHAPTER 8 | MOVING ON: EXTENSIO
Page 381 and 382:
Page 383 and 384:
Page 385 and 386:
Page 387 and 388:
Page 389 and 390:
Page 391 and 392:
Page 393 and 394:
Page 395 and 396:
Page 398 and 399:
chapter 9Introduction to WekaExperi
Page 400 and 401:
9.2 How do you use it?9.2 HOW DO YO
Page 402 and 403:
chapter 10The ExplorerWeka’s main
Page 404 and 405:
10.1 GETTING STARTED 371(a)(b)(c)Fi
Page 406 and 407:
10.1 GETTING STARTED 373deviation.
Page 408 and 409:
10.1 GETTING STARTED 375=== Run inf
Page 410 and 411:
10.1 GETTING STARTED 377Doing it ag
Page 412 and 413:
10.1 GETTING STARTED 379(a)(b)Figur
Page 414 and 415:
10.2 EXPLORING THE EXPLORER 381(a)(
Page 416 and 417:
10.2 EXPLORING THE EXPLORER 383(b)(
Page 418 and 419:
10.2 EXPLORING THE EXPLORER 385(a)(
Page 420 and 421:
10.2 EXPLORING THE EXPLORER 387+ 0.
Page 422 and 423:
10.2 EXPLORING THE EXPLORER 389posi
Page 424 and 425:
10.2 EXPLORING THE EXPLORER 391Figu
Page 426 and 427:
10.3 FILTERING ALGORITHMS 393method
Page 428 and 429:
10.3 FILTERING ALGORITHMS 395(b)Fig
Page 430 and 431:
10.3 FILTERING ALGORITHMS 397or fir
Page 432 and 433:
10.3 FILTERING ALGORITHMS 399attrib
Page 434 and 435:
10.3 FILTERING ALGORITHMS 401it int
Page 436 and 437:
10.4 LEARNING ALGORITHMS 403There i
Page 438 and 439:
10.4 LEARNING ALGORITHMS 405Table 1
Page 440 and 441:
10.4 LEARNING ALGORITHMS 407Figure
Page 442 and 443:
10.4 LEARNING ALGORITHMS 409value (
Page 444 and 445:
10.4 LEARNING ALGORITHMS 411Neural
Page 446 and 447:
10.4 LEARNING ALGORITHMS 413there a
Page 448 and 449:
10.5 METALEARNING ALGORITHMS 415Tab
Page 450 and 451:
10.5 METALEARNING ALGORITHMS 417Com
Page 452 and 453:
10.7 ASSOCIATION-RULE LEARNERS 419N
Page 454 and 455:
10.8 ATTRIBUTE SELECTION 421Table 1
Page 456 and 457:
10.8 ATTRIBUTE SELECTION 423tribute
Page 458:
10.8 ATTRIBUTE SELECTION 425with on
Page 461 and 462:
428 CHAPTER 11 | THE KNOWLEDGE FLOW
Page 463 and 464:
Page 465 and 466:
Page 467 and 468:
Page 470 and 471:
chapter 12The ExperimenterThe Explo
Page 472 and 473:
12.1 GETTING STARTED 439Dataset,Run
Page 474 and 475:
12.2 SIMPLE SETUP 441cance test of
Page 476 and 477:
12.4 THE ANALYZE PANEL 443advanced
Page 478 and 479:
12.5 DISTRIBUTING PROCESSING OVER S
Page 480:
12.5 DISTRIBUTING PROCESSING OVER S
Page 483 and 484:
450 CHAPTER 13 | THE COMMAND-LINE I
Page 485 and 486:
Page 487 and 488:
Page 489 and 490:
Page 491 and 492:
Page 494 and 495:
chapter 14Embedded Machine Learning
Page 496 and 497:
14.2 GOING THROUGH THE CODE 463/***
Page 498 and 499:
14.2 GOING THROUGH THE CODE 465// I
Page 500 and 501:
14.2 GOING THROUGH THE CODE 467}//
Page 502:
14.2 GOING THROUGH THE CODE 469filt
Page 505 and 506:
472 CHAPTER 15 | WRITING NEW LEARNI
Page 507 and 508:
Page 509 and 510:
Page 511 and 512:
Page 513 and 514:
Page 515 and 516:
Page 518 and 519:
ReferencesAdriaans, P., and D. Zant
Page 520 and 521:
Bouckaert, R. R. 2004. Bayesian net
Page 522 and 523:
REFERENCES 489Cypher, A., editor. 1
Page 524 and 525:
REFERENCES 491Fix, E., and J. L. Ho
Page 526 and 527:
REFERENCES 493Gennari, J. H., P. La
Page 528 and 529:
Conference on Knowledge Discovery a
Page 530 and 531:
REFERENCES 497Kushmerick, N., D. S.
Page 532 and 533:
Moore, A. W., and M. S. Lee. 1994.
Page 534 and 535:
REFERENCES 501editor, Proceedings o
Page 536:
REFERENCES 503Webb, G. I., J. Bough
Page 539 and 540:
506 INDEXanomaly detection systems,
Page 541 and 542:
508 INDEXcausal relations, 350CfsSu
Page 543 and 544:
510 INDEXCSVLoader, 381cumulative m
Page 545 and 546:
512 INDEXExperimenter, 437-447advan
Page 547 and 548:
514 INDEXimplementation—real-worl
Page 549 and 550:
516 INDEXlistOptions(), 482literary
Page 551 and 552:
518 INDEXnumeric prediction (contin
Page 553 and 554:
520 INDEXrelational data, 49relatio
Page 555 and 556:
522 INDEXsupport vector, 216support
Page 557 and 558:
524 INDEXWinnow, 410Winnow algorith
show all

Data Mining: Practical Machine Learning Tools and ... - LIDeCC

You also want an ePaper? Increase the reach of your titles

Delete template?

Save as template?