Unit 1: Big data fundamentals
Big Data Analytics notes · PTU syllabus (PGCA1947)
On this page
Unit summary
Big data is data too large, fast or varied for traditional systems. This unit covers the evolution and characteristics of big data, best practices and use cases, big data storage, and the high-performance architecture of HDFS, MapReduce and YARN with the MapReduce programming model.
After this unit you can
- Trace the evolution and characteristics of big data
- Describe use cases and best practices
- Explain HDFS, MapReduce and YARN
- Write a MapReduce word count
PTU syllabus topics
- Evolution and characteristics of big data
- best practices
- big data use cases
- understanding big data storage
- high-performance architecture overview — HDFS
- MapReduce and YARN
- MapReduce programming model
Volume
Huge amounts of data
Velocity
Arrives fast
Variety
Structured to unstructured
Veracity
Uncertain quality
Value
Useful insight
Topic 1
Evolution and characteristics of big data
- 1Files and early databases
- 2Relational databases and data warehouses
- 3Web scale (Google File System 2003, MapReduce 2004)
- 4Hadoop (2006) and NoSQL
- 5Spark, cloud data lakes and real-time streaming
- Volume
- Terabytes to petabytes
- Velocity
- Speed of arrival — streams, sensors
- Variety
- Structured, semi-structured (JSON), unstructured (text, video)
- Veracity
- Uncertain quality
- Value
- Business benefit extracted
Topic 2
Best practices
- Start with business questions
- Not technology
- Data governance
- Ownership, quality, privacy, lineage
- Scalable architecture
- Commodity clusters or cloud
- Right tool for the job
- Batch, stream, NoSQL, warehouse
- Skills
- Data engineers and scientists
- Iterate
- Pilot, measure, scale
Topic 3
Big data use cases
Retail
Purchases, clicks
Recommendations, demand forecasting
Banking
Transactions
Real-time fraud detection
Healthcare
Records, devices
Disease prediction
Telecom
Call records
Churn prediction
Government
Aadhaar, GST
Tax evasion analytics
Manufacturing
Sensors
Predictive maintenance
Topic 4
Big data storage
Distributed file system
Large files split into blocks across nodes
HDFS
Object storage
Objects in buckets; data lakes
Amazon S3, Azure Blob
NoSQL key-value and wide-column
Scalable, flexible schema
Cassandra, HBase
Document
JSON documents
MongoDB
Graph
Nodes and edges
Neo4j
Topic 5
HDFS
- NameNode
- Master storing metadata (file → blocks → DataNodes)
- DataNodes
- Store blocks; send heartbeats
- Block
- 128 MB default
- Replication
- 3 copies across racks for fault tolerance
- Secondary NameNode
- Checkpoints metadata (not a hot standby)
- Design: write once, read many; move computation to data.
Topic 6
MapReduce and YARN
- 1
Input split
- 2
Map
Emit (key, value) pairs
- 3
Combine (optional local reduce)
- 4
Shuffle and sort
Group by key
- 5
Reduce
Aggregate values per key
- 6
Output to HDFS
Example
Word count: "big data big" → map emits (big,1), (data,1), (big,1) → shuffle (big,[1,1]), (data,[1]) → reduce (big,2), (data,1).
- ResourceManager
- Allocates cluster resources
- NodeManager
- Manages containers on each node
- ApplicationMaster
- Negotiates resources for one job
- Container
- Bundle of CPU and memory
python# mapper.py (Hadoop Streaming)
import sys
for line in sys.stdin:
for w in line.split(): print(f"{w}\t1")
# reducer.py
import sys
cur, n = None, 0
for line in sys.stdin:
w, c = line.split("\t")
if w != cur:
if cur: print(f"{cur}\t{n}")
cur, n = w, 0
n += int(c)
if cur: print(f"{cur}\t{n}")Key terms
- Big data
- Data with high volume, velocity and variety
- HDFS
- Hadoop Distributed File System
- NameNode
- HDFS master holding metadata
- Shuffle
- Grouping map output by key for reducers
- YARN
- Hadoop's resource manager
Quick revision
- Evolution; five Vs; best practices; use cases.
- Storage: HDFS, object, NoSQL.
- NameNode, DataNodes, 128 MB blocks, replication 3.
- Map, combine, shuffle, reduce; YARN components.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.List the Vs of big data.
- Q2.Give two big data use cases.
- Q3.What does the NameNode do?
- Q4.What is the default HDFS replication factor?
- Q5.What happens in the shuffle phase?
- Q6.Name the YARN components.
Long-answer questions
- Q1.Explain the characteristics and use cases of big data.
- Q2.Explain HDFS architecture.
- Q3.Explain MapReduce with the word count example, and YARN.
Stuck on this unit?
Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.
