Get in Touch

Course Outline

Introduction

This module offers an overview of the appropriate scenarios for employing 'machine learning', key considerations, and its implications, including advantages and limitations. It covers data types (structured, unstructured, static, streamed), data quality and volume, the distinction between data-driven and user-driven analytics, comparisons between statistical and machine learning models, the challenges of unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.

CORE TOPICS

1. Grasping Naive Bayes

  • Foundations of Bayesian methods
  • Probability concepts
  • Joint probability
  • Conditional probability and Bayes' theorem
  • The Naive Bayes algorithm
  • Classification using Naive Bayes
  • The Laplace estimator
  • Handling numerical features with Naive Bayes

2. Grasping Decision Trees

  • The divide-and-conquer approach
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Grasping Neural Networks

  • Transition from biological to artificial neurons
  • Activation functions
  • Network structure
  • Determining the number of layers
  • Direction of data flow
  • Determining nodes per layer
  • Training networks via backpropagation
  • Deep Learning

4. Grasping Support Vector Machines

  • Classification via hyperplanes
  • Maximizing the margin
  • Handling linearly separable data
  • Handling non-linearly separable data
  • Applying kernels for non-linear spaces

5. Grasping Clustering

  • Clustering as a machine learning objective
  • The k-means clustering algorithm
  • Utilizing distance for cluster assignment and updates
  • Determining the optimal number of clusters

6. Evaluating Classification Performance

  • Interpreting classification prediction data
  • Deep dive into confusion matrices
  • Assessing performance using confusion matrices
  • Metrics beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Optimizing Standard Models for Enhanced Performance

  • Leveraging caret for automated parameter tuning
  • Developing a basic tuned model
  • Customizing the tuning workflow
  • Enhancing model output via meta-learning
  • Concepts of ensembles
  • Bagging
  • Boosting
  • Random forests
  • Training random forests
  • Assessing random forest performance

ADDITIONAL TOPICS

8. Classification via Nearest Neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting the appropriate k value
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Classification Rule-Based Methods

  • The separate-and-conquer strategy
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Fundamentals of Regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlations
  • Multiple linear regression

11. Regression and Model Trees

  • Incorporating regression into tree structures

12. Association Rule Learning

  • The Apriori algorithm for association rules
  • Evaluating rule significance – support and confidence
  • Generating rule sets using the Apriori principle

Supplementary Materials

  • Spark, PySpark, MLlib, and Multi-armed bandits

Requirements

Proficiency in Python

 21 Hours

Testimonials (7)

Upcoming Courses

Related Categories