Unit 1: MongoDB and data mining tools
Database Management Systems-II Laboratory notes · PTU syllabus (UGCC2524)
On this page
Unit summary
This lab has two parts: working with MongoDB documents (insert, query, update, delete, projection and sorting), and using data mining tools such as Weka or R for preprocessing, Apriori association rules and decision-tree classification.
After this unit you can
- Insert, query, update and delete MongoDB documents
- Use projection, sorting and operators in MongoDB
- Preprocess data in Weka or R
- Run Apriori and decision-tree classification and interpret results
PTU syllabus topics
- MongoDB document insertion (insertOne/insertMany)
- selection and filtering queries
- updates and deletes
- projection and sorting
- Weka/R tool installation and components
- fundamental programming
- data preprocessing
- Apriori algorithm implementation
- decision-tree classification
Insert
INSERT INTO students ...
db.students.insertOne({...})
Read
SELECT * FROM students
db.students.find()
Update
UPDATE students SET ...
db.students.updateOne(...)
Delete
DELETE FROM students ...
db.students.deleteOne(...)
Topic 1
MongoDB operations
javascriptuse college
db.students.insertMany([
{ name: "Ana", course: "BCA", marks: 82, city: "Ludhiana" },
{ name: "Ravi", course: "BBA", marks: 68, city: "Jalandhar" }
])
db.students.find({ course: "BCA", marks: { $gte: 75 } })
db.students.find({}, { name: 1, marks: 1, _id: 0 }).sort({ marks: -1 }).limit(5)
db.students.updateMany({ course: "BBA" }, { $inc: { marks: 2 } })
db.students.deleteMany({ marks: { $lt: 40 } })
db.students.countDocuments({ city: "Ludhiana" })Topic 2
Weka and R basics
- Weka: open the Explorer, load an ARFF or CSV file in the Preprocess tab, apply filters (remove attributes, replace missing values, discretise), then use the Associate and Classify tabs.
- R:
data <- read.csv("data.csv"),summary(data),na.omit(data); packagesarules(Apriori) andrpart(decision trees).
Topic 3
Apriori and decision tree
rlibrary(arules)
tx <- read.transactions("baskets.csv", format = "basket", sep = ",")
rules <- apriori(tx, parameter = list(supp = 0.3, conf = 0.6))
inspect(sort(rules, by = "lift"))
library(rpart)
model <- rpart(Result ~ ., data = students, method = "class")
plot(model); text(model)In Weka, choose Apriori under Associate (set minimum support and confidence) and J48 (C4.5 decision tree) under Classify with 10-fold cross-validation; read accuracy and the confusion matrix.
Interface
GUI, no coding
Scripting language
Learning curve
Easy
Steeper
Flexibility
Limited to built-in tools
Very flexible with packages
Key terms
- Projection
- Choosing which fields a query returns
- ARFF
- Attribute-Relation File Format used by Weka
- J48
- Weka's implementation of the C4.5 decision tree
- Cross-validation
- Testing a model on rotating subsets of data
Quick revision
- find(filter, projection).sort().limit() in MongoDB.
- $gte, $lt, $inc, $set are common operators.
- Weka: Preprocess → Associate (Apriori) → Classify (J48).
- R: arules for Apriori, rpart for trees.
Important exam questions
Practice questions written to the PTU exam pattern for this unit's syllabus: short answers (Section A style) and long answers (Sections B and C style).
Short-answer questions
- Q1.What is projection in MongoDB?
- Q2.Write a query to sort documents by marks in descending order.
- Q3.What is an ARFF file?
- Q4.What is J48 in Weka?
- Q5.What does lift indicate in association rules?
Long-answer questions
- Q1.Perform CRUD operations on a students collection in MongoDB with filters, projection and sorting.
- Q2.Run the Apriori algorithm on a transaction data set in Weka or R and interpret the rules.
- Q3.Build a decision-tree classifier in Weka and explain the confusion matrix.
Stuck on this unit?
Message SBS on WhatsApp for help with Database Management Systems-II Laboratory, or to ask about studying BCA at Synetic.
