Unit 1 of 4 · M.Sc IT Sem 4

Unit 1: Big data fundamentals

Big Data Analytics notes · PTU syllabus (PGCA1947)

3 min read6 topics9 exam questions
On this page
  1. Unit summary
  2. Evolution and characteristics of big data
  3. Best practices
  4. Big data use cases
  5. Big data storage
  6. HDFS
  7. MapReduce and YARN
  8. Key terms
  9. Quick revision
  10. Important questions

Unit summary

Big data is data too large, fast or varied for traditional systems. This unit covers the evolution and characteristics of big data, best practices and use cases, big data storage, and the high-performance architecture of HDFS, MapReduce and YARN with the MapReduce programming model.

After this unit you can

  • Trace the evolution and characteristics of big data
  • Describe use cases and best practices
  • Explain HDFS, MapReduce and YARN
  • Write a MapReduce word count

PTU syllabus topics

  • Evolution and characteristics of big data
  • best practices
  • big data use cases
  • understanding big data storage
  • high-performance architecture overview — HDFS
  • MapReduce and YARN
  • MapReduce programming model
ClassificationThe Vs of big data
Big data
  • Volume

    Huge amounts of data

  • Velocity

    Arrives fast

  • Variety

    Structured to unstructured

  • Veracity

    Uncertain quality

  • Value

    Useful insight

1

Topic 1

Evolution and characteristics of big data

ProcessEvolution
  1. 1Files and early databases
  2. 2Relational databases and data warehouses
  3. 3Web scale (Google File System 2003, MapReduce 2004)
  4. 4Hadoop (2006) and NoSQL
  5. 5Spark, cloud data lakes and real-time streaming
Key termsThe Vs of big data
Volume
Terabytes to petabytes
Velocity
Speed of arrival — streams, sensors
Variety
Structured, semi-structured (JSON), unstructured (text, video)
Veracity
Uncertain quality
Value
Business benefit extracted
2

Topic 2

Best practices

Key termsBest practices
Start with business questions
Not technology
Data governance
Ownership, quality, privacy, lineage
Scalable architecture
Commodity clusters or cloud
Right tool for the job
Batch, stream, NoSQL, warehouse
Skills
Data engineers and scientists
Iterate
Pilot, measure, scale
3

Topic 3

Big data use cases

ComparisonUse cases
Data
Outcome

Retail

Purchases, clicks

Recommendations, demand forecasting

Banking

Transactions

Real-time fraud detection

Healthcare

Records, devices

Disease prediction

Telecom

Call records

Churn prediction

Government

Aadhaar, GST

Tax evasion analytics

Manufacturing

Sensors

Predictive maintenance

4

Topic 4

Big data storage

ComparisonStorage options
Model
Examples

Distributed file system

Large files split into blocks across nodes

HDFS

Object storage

Objects in buckets; data lakes

Amazon S3, Azure Blob

NoSQL key-value and wide-column

Scalable, flexible schema

Cassandra, HBase

Document

JSON documents

MongoDB

Graph

Nodes and edges

Neo4j

5

Topic 5

HDFS

Key termsHDFS architecture
NameNode
Master storing metadata (file → blocks → DataNodes)
DataNodes
Store blocks; send heartbeats
Block
128 MB default
Replication
3 copies across racks for fault tolerance
Secondary NameNode
Checkpoints metadata (not a hot standby)
  • Design: write once, read many; move computation to data.
6

Topic 6

MapReduce and YARN

ProcessMapReduce
  1. 1

    Input split

  2. 2

    Map

    Emit (key, value) pairs

  3. 3

    Combine (optional local reduce)

  4. 4

    Shuffle and sort

    Group by key

  5. 5

    Reduce

    Aggregate values per key

  6. 6

    Output to HDFS

Example

Word count: "big data big" → map emits (big,1), (data,1), (big,1) → shuffle (big,[1,1]), (data,[1]) → reduce (big,2), (data,1).

Key termsYARN
ResourceManager
Allocates cluster resources
NodeManager
Manages containers on each node
ApplicationMaster
Negotiates resources for one job
Container
Bundle of CPU and memory
python# mapper.py (Hadoop Streaming)
import sys
for line in sys.stdin:
    for w in line.split(): print(f"{w}\t1")
# reducer.py
import sys
cur, n = None, 0
for line in sys.stdin:
    w, c = line.split("\t")
    if w != cur:
        if cur: print(f"{cur}\t{n}")
        cur, n = w, 0
    n += int(c)
if cur: print(f"{cur}\t{n}")

Key terms

Big data
Data with high volume, velocity and variety
HDFS
Hadoop Distributed File System
NameNode
HDFS master holding metadata
Shuffle
Grouping map output by key for reducers
YARN
Hadoop's resource manager

Quick revision

  • Evolution; five Vs; best practices; use cases.
  • Storage: HDFS, object, NoSQL.
  • NameNode, DataNodes, 128 MB blocks, replication 3.
  • Map, combine, shuffle, reduce; YARN components.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.List the Vs of big data.
  2. Q2.Give two big data use cases.
  3. Q3.What does the NameNode do?
  4. Q4.What is the default HDFS replication factor?
  5. Q5.What happens in the shuffle phase?
  6. Q6.Name the YARN components.

Long-answer questions

  1. Q1.Explain the characteristics and use cases of big data.
  2. Q2.Explain HDFS architecture.
  3. Q3.Explain MapReduce with the word count example, and YARN.

Stuck on this unit?

Message SBS on WhatsApp for help with Big Data Analytics, or to ask about studying M.Sc IT at Synetic.

WhatsApp us