Unit 4 of 4 · M.Sc IT Sem 4

Unit 4: Hadoop implementation

Big Data Analytics notes · PTU syllabus (PGCA1947)

3 min read4 topics9 exam questions
On this page
  1. Unit summary
  2. Hadoop cluster components and architecture
  3. The Hadoop ecosystem
  4. Evaluating distributed MapReduce runtimes
  5. Enterprise-grade Hadoop deployment
  6. Key terms
  7. Quick revision
  8. Important questions

Unit summary

Hadoop is the foundation of many big data platforms. This unit covers Hadoop cluster components and architecture, the Hadoop ecosystem, evaluation criteria for distributed MapReduce runtimes and enterprise-grade deployment.

After this unit you can

  • Describe Hadoop cluster architecture
  • Explain components of the Hadoop ecosystem
  • Evaluate distributed MapReduce runtimes
  • Plan an enterprise Hadoop deployment

PTU syllabus topics

  • Hadoop cluster components and architecture
  • Hadoop ecosystem
  • evaluation criteria for distributed MapReduce runtimes
  • enterprise-grade Hadoop deployment and implementation
ProcessHow MapReduce works
  1. 1Input split

    HDFS blocks

  2. 2Map

    Emit key-value pairs

  3. 3Shuffle and sort

    Group by key

  4. 4Reduce

    Aggregate each key

  5. 5Output

    Written to HDFS

1

Topic 1

Hadoop cluster components and architecture

ClassificationHadoop cluster
Master nodes|NameNode, ResourceManager, standby NameNode with JournalNodes and ZooKeeper for high availability
  • Worker nodes

    DataNode and NodeManager on each machine

  • Edge (gateway) nodes

    Client tools, job submission

  • Network

    Rack-aware topology; top-of-rack switches

2

Topic 2

The Hadoop ecosystem

ComparisonEcosystem
Purpose
Notes

Hive

SQL (HiveQL) on Hadoop

Data warehousing

Pig

Data-flow scripting (Pig Latin)

ETL

HBase

Column-family NoSQL on HDFS

Random reads and writes

Sqoop

Bulk transfer RDBMS ↔ HDFS

Now retired; replaced by Spark JDBC and others

Flume and Kafka

Ingest logs and events

Streaming data

Oozie

Workflow scheduler

Airflow is common today

ZooKeeper

Coordination

Leader election, configuration

Spark

In-memory processing

Much faster than MapReduce for iterative jobs

Mahout and Spark MLlib

Machine learning

Scalable algorithms

sql-- HiveQL
CREATE TABLE sales (item STRING, amount DOUBLE) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',';
LOAD DATA INPATH '/data/sales.csv' INTO TABLE sales;
SELECT item, SUM(amount) FROM sales GROUP BY item;
3

Topic 3

Evaluating distributed MapReduce runtimes

Key termsEvaluation criteria
Scalability
Performance as nodes and data grow
Fault tolerance
Recovery from node failure
Performance
Throughput, latency, in-memory support
Programming model
APIs, languages, SQL support
Resource management
Multi-tenancy, scheduling
Ecosystem and support
Tools, community, vendor
Cost
Hardware, licences, cloud spend
4

Topic 4

Enterprise-grade Hadoop deployment

ProcessDeployment
  1. 1

    Define use cases and data volumes

  2. 2

    Size the cluster — storage = data × replication × growth

  3. 3

    Choose on-premises, cloud-managed (EMR, Dataproc, HDInsight) or hybrid

  4. 4

    Secure — Kerberos, Ranger, encryption

  5. 5

    Govern — Atlas metadata and lineage

  6. 6

    Monitor — Ambari or Cloudera Manager

  7. 7

    Back up and plan disaster recovery

Example

100 TB data, replication 3, 25% headroom → about 375 TB raw storage needed.

Key terms

High availability
Standby NameNode avoiding a single point of failure
Hive
SQL engine on Hadoop
HBase
NoSQL database on HDFS
Spark
In-memory distributed processing engine
Kerberos
Authentication protocol used to secure Hadoop

Quick revision

  • Master, worker and edge nodes; HA.
  • Hive, Pig, HBase, Sqoop, Flume, Kafka, Oozie, ZooKeeper, Spark, MLlib.
  • Evaluation criteria.
  • Sizing, cloud options, security, governance, monitoring.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.What is an edge node?
  2. Q2.What is Hive used for?
  3. Q3.Distinguish HBase and HDFS.
  4. Q4.Why is Spark faster than MapReduce?
  5. Q5.Name two criteria for evaluating MapReduce runtimes.
  6. Q6.How is Hadoop secured?

Long-answer questions

  1. Q1.Explain Hadoop cluster architecture.
  2. Q2.Explain the components of the Hadoop ecosystem.
  3. Q3.Explain enterprise-grade Hadoop deployment.

Stuck on this unit?

Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.

WhatsApp us