Unit 1 of 1 · BCA Sem 3

Unit 1: Feature engineering in Python

Feature Engineering Laboratory notes · PTU syllabus (UGDSE202)

3 min read5 topics8 exam questions
On this page
  1. Unit summary
  2. Cleaning and normalisation
  3. Exploratory data analysis
  4. Creating new features
  5. Text and image preprocessing
  6. Time series and PCA
  7. Key terms
  8. Quick revision
  9. Important questions

Unit summary

This lab performs feature engineering in Python with pandas, NumPy and scikit-learn: cleaning, scaling, exploratory analysis, binning, polynomial features, transformations, encoding, text preprocessing, image augmentation, time-series decomposition and PCA.

After this unit you can

  • Clean data and normalise it with pandas and scikit-learn
  • Explore data with histograms, boxplots and correlation matrices
  • Create binned, polynomial, log-transformed and one-hot encoded features
  • Preprocess text and images, decompose a time series and apply PCA

PTU syllabus topics

  • Handling missing values and invalid entries
  • Min-Max normalization
  • exploratory data analysis (histograms, boxplots, correlation matrix)
  • binning numerical data
  • polynomial/interaction features
  • logarithmic transformation
  • one-hot encoding
  • text preprocessing (tokenization, stemming, lemmatization, Bag-of-Words, TF-IDF)
  • image augmentation (resizing, normalization, rotation, translation)
  • time-series decomposition
  • Principal Component Analysis and visualization
ProcessA feature engineering workflow
  1. 1Explore

    Histograms, boxplots, correlation

  2. 2Clean

    Missing and invalid values

  3. 3Transform

    Scale, log-transform, bin

  4. 4Encode

    One-hot or label

  5. 5Reduce

    Select features or apply PCA

1

Topic 1

Cleaning and normalisation

pythonimport pandas as pd
from sklearn.preprocessing import MinMaxScaler
df = pd.read_csv("data.csv")
df = df.drop_duplicates()
df["age"] = df["age"].fillna(df["age"].median())
df["city"] = df["city"].fillna("Unknown")
df = df[df["age"].between(0, 100)]            # remove invalid entries
df[["age", "income"]] = MinMaxScaler().fit_transform(df[["age", "income"]])
2

Topic 2

Exploratory data analysis

pythonimport matplotlib.pyplot as plt, seaborn as sns
df["income"].hist(bins=20)
sns.boxplot(x=df["income"])                  # shows outliers
sns.heatmap(df.corr(numeric_only=True), annot=True)
plt.show()
3

Topic 3

Creating new features

pythonimport numpy as np
from sklearn.preprocessing import PolynomialFeatures
df["age_group"] = pd.cut(df["age"], bins=[0, 18, 35, 60, 100],
                         labels=["teen", "young", "middle", "senior"])
df["log_income"] = np.log1p(df["income"])
df = pd.get_dummies(df, columns=["city"])    # one-hot encoding
poly = PolynomialFeatures(degree=2, include_bias=False)
X_poly = poly.fit_transform(df[["x1", "x2"]])
4

Topic 4

Text and image preprocessing

  • Text: tokenise (split into words), remove stop words, stem (running → run) or lemmatise (better → good), then represent with Bag-of-Words (CountVectorizer) or TF-IDF (TfidfVectorizer).
  • Images: resize to a fixed size, normalise pixel values to 0–1, and augment with rotation, flips and translation to create more training examples.
ComparisonBag-of-Words vs TF-IDF
Bag-of-Words
TF-IDF

Value

Count of each word

Count weighted by rarity across documents

Common words

Get high values

Get down-weighted

Use

Simple text models

Better for search and classification

5

Topic 5

Time series and PCA

pythonfrom statsmodels.tsa.seasonal import seasonal_decompose
result = seasonal_decompose(sales, model="additive", period=12)
result.plot()                                # trend, seasonal, residual

from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
X = StandardScaler().fit_transform(features)
pca = PCA(n_components=2)
X2 = pca.fit_transform(X)
print(pca.explained_variance_ratio_)

Key terms

pd.cut
pandas function for binning values
get_dummies
pandas function for one-hot encoding
Lemmatisation
Reducing words to their dictionary form
TF-IDF
Term frequency–inverse document frequency weighting
Explained variance ratio
Share of variance captured by each principal component

Quick revision

  • fillna for missing values; MinMaxScaler for 0–1 scaling.
  • Boxplots show outliers; heatmaps show correlations.
  • pd.cut bins, np.log1p transforms, get_dummies encodes.
  • Standardise before PCA.

Important exam questions

Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).

Short-answer questions

  1. Q1.Which pandas function fills missing values?
  2. Q2.What does np.log1p do?
  3. Q3.Differentiate between stemming and lemmatisation.
  4. Q4.Why is image augmentation used?
  5. Q5.Why must data be standardised before PCA?

Long-answer questions

  1. Q1.Write a Python program to clean a data set, handle missing values and normalise numerical columns.
  2. Q2.Write a program to create binned, polynomial and one-hot encoded features.
  3. Q3.Apply PCA to a data set and visualise the first two principal components.

Stuck on this unit?

Message SBS on WhatsApp for help with Feature Engineering Laboratory, or to ask about studying BCA at Synetic.

WhatsApp us