- The Design of HDFS
- Command line interface
- Hadoop File System
- Anatomy of a cluster
- Mater Node / Slave node
- Name Node / Data Node
- Map phase
- Reduce phase
1.2.2Analytics with Map Reduce
- Group-By with MapReduce
- Frequency distributions and sorting with MapReduce
- Plotting results (GNU Plot)
- Histograms with MapReduce
- Scatter plots with MapReduce
- Parsing complex datasets
- Counting with MapReduce and Combiners
- Build reports
- Document Cleaning
- Fuzzy string search
- Record linkage / data deduplication
- Transform and sort event dates
- Validate source reliability
- Trim Outliers
1.2.4Extracting and Transforming Data
- Transforming logs
- Using Apache Pig to filter
- Using Apache Pig to sort
- Using Apache Pig to sessionize
- Joining data in the Mapper using MapReduce
- Joining data using Apache Pig replicated join
- Joining sorted data using Apache Pig merge join
- Joining skewed data using Apache Pig skewed join
- Using a map-side join in Apache Hive
- Using optimized full outer joins in Apache Hive
- Joining data using an external key value store
1.3Performance Diagnosis and Optimization Techniques
- Investigating spikes in input data
- Identifying map-side data skew problems
- Map task throughput
- Small files
- Unsplittable files
- Too few or too many reducers
- Reduce-side data skew problems
- Reduce tasks throughput
- Slow shuffle and sort
- Competing jobs and scheduler throttling
- Stack dumps & unoptimized code
- Hardware failures
- CPU contention
- Extracting and visualizing task execution times
- Profiling your map and reduce tasks
- Avoid the reducer
- Filter and project
- Using the combiner
- Fast sorting with comparators
- Collecting skewed data
- Reduce skew mitigation
Attendees are not required to have any specific skill as the training is focused on end users skills for both the administration and the manipulation of data under Apache Hadoop
The fact that all the data and software was ready to use on an already prepared VM, provided by the trainer in external disks.
I mostly liked the trainer giving real live Examples.
I genuinely enjoyed the big competences of Trainer.
I genuinely enjoyed the many hands-on sessions.
It was very hands-on, we spent half the time actually doing things in Clouded/Hardtop, running different commands, checking the system, and so on. The extra materials (books, websites, etc. .) were really appreciated, we will have to continue to learn. The installations were quite fun, and very handy, the cluster setup from scratch was really good.
Lot of hands-on exercises.
Ambari management tool. Ability to discuss practical Hadoop experiences from other business case than telecom.
The VM I liked very much The Teacher was very knowledgeable regarding the topic as well as other topics, he was very nice and friendly I liked the facility in Dubai.
Safar Alqahtani - Elm Information Security
Training topics and engagement of the trainer
- Izba Administracji Skarbowej w Lublinie
Communication with people attending training.
Andrzej Szewczuk - Izba Administracji Skarbowej w Lublinie
practical things of doing, also theory was served good by Ajay
Dominik Mazur - Capgemini Polska Sp. z o.o.
- Capgemini Polska Sp. z o.o.
usefulness of exercises
- Algomine sp.z.o.o sp.k.
I found the training good, very informative....but could have been spread over 4 or 5 days, allowing us to go into more details on different aspects.
- Veterans Affairs Canada
I really enjoyed the training. Anton has a lot of knowledge and laid out the necessary theory in a very accessible way. It is great that the training was a lot of interesting exercises, so we have been in contact with the technology we know from the very beginning.
Szymon Dybczak - Algomine sp.z.o.o sp.k.
I found this course gave a great overview and quickly touched some areas I wasn't even considering.
- Veterans Affairs Canada
I genuinely liked work exercises with cluster to see performance of nodes across cluster and extended functionality.
The trainers in depth knowledge of the subject
Ajay was a very experienced consultant and was able to answer all our questions and even made suggestions on best practices for the project we are currently engaged on.
That I had it in the first place.
Peter Scales - CACI Ltd
The NIFI workflow excercises
answers to our specific questions
Apache Ambari: Efficiently Manage Hadoop Clusters21 hours
Apache Ambari is an open-source management platform for provisioning, managing, monitoring and securing Apache Hadoop clusters. In this instructor-led live training participants will learn the management tools and practices provided by Ambari to
Administrator Training for Apache Hadoop35 hours
Audience: The course is intended for IT specialists looking for a solution to store and process large data sets in a distributed system environment Goal: Deep knowledge on Hadoop cluster
Hadoop Administration21 hours
The course is dedicated to IT specialists that are looking for a solution to store and process large data sets in distributed system environment Course goal: Getting knowledge regarding Hadoop cluster
Hadoop For Administrators21 hours
Apache Hadoop is the most popular framework for processing Big Data on clusters of servers. In this three (optionally, four) days course, attendees will learn about the business benefits and use cases for Hadoop and its ecosystem, how to plan
Hadoop for Business Analysts21 hours
Apache Hadoop is the most popular framework for processing Big Data. Hadoop provides rich and deep analytics capability, and it is making in-roads in to tradional BI analytics world. This course will introduce an analyst to the core components of
Hadoop for Developers (4 days)28 hours
Apache Hadoop is the most popular framework for processing Big Data on clusters of servers. This course will introduce a developer to various components (HDFS, MapReduce, Pig, Hive and HBase) Hadoop
Advanced Hadoop for Developers21 hours
Apache Hadoop is one of the most popular frameworks for processing Big Data on clusters of servers. This course delves into data management in HDFS, advanced Pig, Hive, and HBase. These advanced programming techniques will be beneficial to
Hadoop for Developers and Administrators21 hours
Hadoop is the most popular Big Data processing framework.
Hadoop for Project Managers14 hours
As more and more software and IT projects migrate from local processing and data management to distributed processing and big data storage, Project Managers are finding the need to upgrade their knowledge and skills to grasp the concepts and
Hadoop Administration on MapR28 hours
Audience: This course is intended to demystify big data/hadoop technology and to show it is not difficult to understand.
HBase for Developers21 hours
This course introduces HBase – a NoSQL store on top of Hadoop. The course is intended for developers who will be using HBase to develop applications, and administrators who will manage HBase clusters. We will walk a developer
Hortonworks Data Platform (HDP) for Administrators21 hours
Hortonworks Data Platform (HDP) is an open-source Apache Hadoop support platform that provides a stable foundation for developing big data solutions on the Apache Hadoop ecosystem. This instructor-led, live training (online or onsite) introduces
Data Analysis with Hive/HiveQL7 hours
This course covers how to use Hive SQL language (AKA: Hive HQL, SQL on Hive, HiveQL) for people who extract data from Hive
Impala for Business Intelligence21 hours
Cloudera Impala is an open source massively parallel processing (MPP) SQL query engine for Apache Hadoop clusters. Impala enables users to issue low-latency SQL queries to data stored in Hadoop Distributed File System and Apache
Apache Avro: Data Serialization for Distributed Applications14 hours
Audience Developers Format of the Course Lectures, hands-on practice, small tests along the way to gauge understanding