Unit 1: Introduction to feature engineering
Feature Engineering notes · PTU syllabus (UGDSE201)
On this page
Unit summary
A machine learning model learns only from the features (input columns) it is given. Good features often matter more than the choice of algorithm. This unit explains why features are important, the different data and feature types, and the basic preprocessing every data set needs: handling missing data, cleaning, scaling, normalisation and transformation.
After this unit you can
- Explain the role of features in machine learning
- Classify features as numerical, categorical, ordinal, discrete, continuous, interval or ratio
- Handle missing data and clean a data set
- Apply feature scaling, normalisation and transformations
PTU syllabus topics
- Importance of features in machine learning
- data and feature types (numerical, categorical, ordinal, discrete, continuous, interval, ratio)
- basic preprocessing — handling missing data
- data cleaning
- feature scaling
- normalization and transformation
Numerical: discrete
Counts, e.g. number of orders
Numerical: continuous
Measurements, e.g. height
Categorical: nominal
No order, e.g. city
Categorical: ordinal
Ordered, e.g. low/medium/high
Topic 1
Why features matter
A feature is an individual measurable property used as input to a model (for example a house's area, location and number of rooms). Feature engineering is the process of creating, transforming and selecting features to improve model performance.
- Better features → simpler models, faster training and higher accuracy.
- "Garbage in, garbage out": poor features produce poor predictions however advanced the algorithm.
- 1Collect raw data
- 2Clean and preprocess
- 3Engineer features
Create, transform, encode, select
- 4Train the model
- 5Evaluate and iterate
Topic 2
Data and feature types
| Type | Meaning | Example |
|---|---|---|
| Numerical — discrete | Countable whole numbers | Number of children |
| Numerical — continuous | Any value in a range | Height, salary |
| Categorical — nominal | Categories with no order | City, colour |
| Categorical — ordinal | Categories with an order | Low / medium / high |
| Interval | Equal gaps, no true zero | Temperature in °C |
| Ratio | Equal gaps and a true zero | Weight, income |
Exam tip
The type decides the technique: nominal features need one-hot encoding, ordinal features need label (ordered) encoding, and numerical features may need scaling.
Topic 3
Handling missing data and cleaning
- Deletion: drop rows (or columns) with many missing values — simple but loses data.
- Imputation: fill with the mean or median (numerical), the mode (categorical), a constant ("Unknown"), or a value predicted from other features.
- Indicator feature: add a column flagging that the value was missing.
Cleaning also removes duplicates, fixes inconsistent labels ("Delhi", "delhi", "New Delhi"), corrects wrong types and treats outliers.
Topic 4
Feature scaling and normalisation
Many algorithms (k-NN, SVM, gradient descent, PCA) are sensitive to the scale of features — a salary in lakhs would swamp an age in years.
Min-max normalisation
x' = (x − min) / (max − min)
Range 0 to 1
Standardisation (Z-score)
z = (x − μ) / σ
Mean 0, SD 1
Robust scaling
(x − median) / IQR
Resistant to outliers
Example
Ages 20, 30, 40 with min 20 and max 40 become 0, 0.5 and 1 after min-max normalisation.
Topic 5
Feature transformation
- Log transformation reduces right skew (incomes, prices): x' = log(x + 1).
- Square-root and Box-Cox transformations make data more normal.
- Date features: extract day, month, weekday or "is weekend" from a date.
- Tree-based models usually do not need scaling; distance-based models do.
Key terms
- Feature
- An input variable used by a model
- Feature engineering
- Creating and transforming features to improve models
- Imputation
- Filling missing values with estimates
- Normalisation
- Rescaling values into a fixed range such as 0–1
- Standardisation
- Rescaling to mean 0 and standard deviation 1
Quick revision
- Nominal → one-hot; ordinal → ordered labels; numerical → scale if needed.
- Missing data: delete, impute (mean/median/mode) or flag.
- Min-max: (x − min)/(max − min); Z-score: (x − μ)/σ.
- Log transform fixes right skew.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is a feature in machine learning?
- Q2.Differentiate between nominal and ordinal features.
- Q3.List three ways to handle missing values.
- Q4.Differentiate between normalisation and standardisation.
- Q5.Why is feature scaling needed?
Long-answer questions
- Q1.Explain the types of data and features with examples.
- Q2.Explain techniques for handling missing data and cleaning a data set.
- Q3.Explain feature scaling and transformation methods with formulas and examples.
Stuck on this unit?
Message SBS on WhatsApp for help with Feature Engineering, or to ask about studying BCA at Synetic.
