


3 screenshots, click to enlarge
Overview
An NLP machine learning project that classifies emails as spam or not spam using text preprocessing and classification algorithms. Trained on the Enron email dataset with multiple model comparisons.
Machine Learning
6 technologies · 5 key features · 3 outcomes
Challenge & Approach
Email spam is a persistent problem that wastes time and reduces productivity. Manual filtering is inefficient. Organizations need automated, accurate classification systems.
Built a text classification pipeline with tokenization, stopword removal, TF-IDF vectorization, and multiple classifier comparison. Achieved high accuracy with Naive Bayes and SVM models.
Key Features
Technologies Used
Results & Impact
Achieved 97.8% accuracy with SVM classifier
Low false positive rate (legitimate emails kept safe)
Reusable text classification pipeline
Lessons Learned
NLP preprocessing quality directly impacts classification accuracy
SVM outperforms Naive Bayes on larger text datasets
Hire Me
Want to build something like this?
Open to freelance projects, SaaS builds, data products, and business websites.