Unit 1: Feature engineering in Python
Feature Engineering Laboratory notes · PTU syllabus (UGDSE202)
On this page
Unit summary
This lab performs feature engineering in Python with pandas, NumPy and scikit-learn: cleaning, scaling, exploratory analysis, binning, polynomial features, transformations, encoding, text preprocessing, image augmentation, time-series decomposition and PCA.
After this unit you can
- Clean data and normalise it with pandas and scikit-learn
- Explore data with histograms, boxplots and correlation matrices
- Create binned, polynomial, log-transformed and one-hot encoded features
- Preprocess text and images, decompose a time series and apply PCA
PTU syllabus topics
- Handling missing values and invalid entries
- Min-Max normalization
- exploratory data analysis (histograms, boxplots, correlation matrix)
- binning numerical data
- polynomial/interaction features
- logarithmic transformation
- one-hot encoding
- text preprocessing (tokenization, stemming, lemmatization, Bag-of-Words, TF-IDF)
- image augmentation (resizing, normalization, rotation, translation)
- time-series decomposition
- Principal Component Analysis and visualization
- 1Explore
Histograms, boxplots, correlation
- 2Clean
Missing and invalid values
- 3Transform
Scale, log-transform, bin
- 4Encode
One-hot or label
- 5Reduce
Select features or apply PCA
Topic 1
Cleaning and normalisation
pythonimport pandas as pd
from sklearn.preprocessing import MinMaxScaler
df = pd.read_csv("data.csv")
df = df.drop_duplicates()
df["age"] = df["age"].fillna(df["age"].median())
df["city"] = df["city"].fillna("Unknown")
df = df[df["age"].between(0, 100)] # remove invalid entries
df[["age", "income"]] = MinMaxScaler().fit_transform(df[["age", "income"]])Topic 2
Exploratory data analysis
pythonimport matplotlib.pyplot as plt, seaborn as sns
df["income"].hist(bins=20)
sns.boxplot(x=df["income"]) # shows outliers
sns.heatmap(df.corr(numeric_only=True), annot=True)
plt.show()Topic 3
Creating new features
pythonimport numpy as np
from sklearn.preprocessing import PolynomialFeatures
df["age_group"] = pd.cut(df["age"], bins=[0, 18, 35, 60, 100],
labels=["teen", "young", "middle", "senior"])
df["log_income"] = np.log1p(df["income"])
df = pd.get_dummies(df, columns=["city"]) # one-hot encoding
poly = PolynomialFeatures(degree=2, include_bias=False)
X_poly = poly.fit_transform(df[["x1", "x2"]])Topic 4
Text and image preprocessing
- Text: tokenise (split into words), remove stop words, stem (running → run) or lemmatise (better → good), then represent with Bag-of-Words (
CountVectorizer) or TF-IDF (TfidfVectorizer). - Images: resize to a fixed size, normalise pixel values to 0–1, and augment with rotation, flips and translation to create more training examples.
Value
Count of each word
Count weighted by rarity across documents
Common words
Get high values
Get down-weighted
Use
Simple text models
Better for search and classification
Topic 5
Time series and PCA
pythonfrom statsmodels.tsa.seasonal import seasonal_decompose
result = seasonal_decompose(sales, model="additive", period=12)
result.plot() # trend, seasonal, residual
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
X = StandardScaler().fit_transform(features)
pca = PCA(n_components=2)
X2 = pca.fit_transform(X)
print(pca.explained_variance_ratio_)Key terms
- pd.cut
- pandas function for binning values
- get_dummies
- pandas function for one-hot encoding
- Lemmatisation
- Reducing words to their dictionary form
- TF-IDF
- Term frequency–inverse document frequency weighting
- Explained variance ratio
- Share of variance captured by each principal component
Quick revision
- fillna for missing values; MinMaxScaler for 0–1 scaling.
- Boxplots show outliers; heatmaps show correlations.
- pd.cut bins, np.log1p transforms, get_dummies encodes.
- Standardise before PCA.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.Which pandas function fills missing values?
- Q2.What does np.log1p do?
- Q3.Differentiate between stemming and lemmatisation.
- Q4.Why is image augmentation used?
- Q5.Why must data be standardised before PCA?
Long-answer questions
- Q1.Write a Python program to clean a data set, handle missing values and normalise numerical columns.
- Q2.Write a program to create binned, polynomial and one-hot encoded features.
- Q3.Apply PCA to a data set and visualise the first two principal components.
Stuck on this unit?
Message SBS on WhatsApp for help with Feature Engineering Laboratory, or to ask about studying BCA at Synetic.
