Unit 4: Hadoop implementation
Big Data Analytics notes · PTU syllabus (PGCA1947)
On this page
Unit summary
Hadoop is the foundation of many big data platforms. This unit covers Hadoop cluster components and architecture, the Hadoop ecosystem, evaluation criteria for distributed MapReduce runtimes and enterprise-grade deployment.
After this unit you can
- Describe Hadoop cluster architecture
- Explain components of the Hadoop ecosystem
- Evaluate distributed MapReduce runtimes
- Plan an enterprise Hadoop deployment
PTU syllabus topics
- Hadoop cluster components and architecture
- Hadoop ecosystem
- evaluation criteria for distributed MapReduce runtimes
- enterprise-grade Hadoop deployment and implementation
- 1Input split
HDFS blocks
- 2Map
Emit key-value pairs
- 3Shuffle and sort
Group by key
- 4Reduce
Aggregate each key
- 5Output
Written to HDFS
Topic 1
Hadoop cluster components and architecture
Worker nodes
DataNode and NodeManager on each machine
Edge (gateway) nodes
Client tools, job submission
Network
Rack-aware topology; top-of-rack switches
Topic 2
The Hadoop ecosystem
Hive
SQL (HiveQL) on Hadoop
Data warehousing
Pig
Data-flow scripting (Pig Latin)
ETL
HBase
Column-family NoSQL on HDFS
Random reads and writes
Sqoop
Bulk transfer RDBMS ↔ HDFS
Now retired; replaced by Spark JDBC and others
Flume and Kafka
Ingest logs and events
Streaming data
Oozie
Workflow scheduler
Airflow is common today
ZooKeeper
Coordination
Leader election, configuration
Spark
In-memory processing
Much faster than MapReduce for iterative jobs
Mahout and Spark MLlib
Machine learning
Scalable algorithms
sql-- HiveQL
CREATE TABLE sales (item STRING, amount DOUBLE) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',';
LOAD DATA INPATH '/data/sales.csv' INTO TABLE sales;
SELECT item, SUM(amount) FROM sales GROUP BY item;Topic 3
Evaluating distributed MapReduce runtimes
- Scalability
- Performance as nodes and data grow
- Fault tolerance
- Recovery from node failure
- Performance
- Throughput, latency, in-memory support
- Programming model
- APIs, languages, SQL support
- Resource management
- Multi-tenancy, scheduling
- Ecosystem and support
- Tools, community, vendor
- Cost
- Hardware, licences, cloud spend
Topic 4
Enterprise-grade Hadoop deployment
- 1
Define use cases and data volumes
- 2
Size the cluster — storage = data × replication × growth
- 3
Choose on-premises, cloud-managed (EMR, Dataproc, HDInsight) or hybrid
- 4
Secure — Kerberos, Ranger, encryption
- 5
Govern — Atlas metadata and lineage
- 6
Monitor — Ambari or Cloudera Manager
- 7
Back up and plan disaster recovery
Example
100 TB data, replication 3, 25% headroom → about 375 TB raw storage needed.
Key terms
- High availability
- Standby NameNode avoiding a single point of failure
- Hive
- SQL engine on Hadoop
- HBase
- NoSQL database on HDFS
- Spark
- In-memory distributed processing engine
- Kerberos
- Authentication protocol used to secure Hadoop
Quick revision
- Master, worker and edge nodes; HA.
- Hive, Pig, HBase, Sqoop, Flume, Kafka, Oozie, ZooKeeper, Spark, MLlib.
- Evaluation criteria.
- Sizing, cloud options, security, governance, monitoring.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is an edge node?
- Q2.What is Hive used for?
- Q3.Distinguish HBase and HDFS.
- Q4.Why is Spark faster than MapReduce?
- Q5.Name two criteria for evaluating MapReduce runtimes.
- Q6.How is Hadoop secured?
Long-answer questions
- Q1.Explain Hadoop cluster architecture.
- Q2.Explain the components of the Hadoop ecosystem.
- Q3.Explain enterprise-grade Hadoop deployment.
Stuck on this unit?
Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.
