Video summary

Project 36 : Resume Categorization Using Machine Learning

Main summary

Key takeaways

Technology

Resume Categorization Using Machine Learning (Python)

This video builds an end-to-end resume categorization application using Python + NLP + machine learning, along with a Streamlit web UI. The workflow lets you upload resumes (PDF), classify them into job categories, save results, and download predictions as CSV.


1) Application (Streamlit) demo: how categorization works

  • Users upload one or multiple resumes via the UI (expects PDF resumes).
  • The app:
    • Reads each PDF
    • Extracts text
    • Cleans the text (removes URL/email/special characters + stopwords)
  • It then predicts the resume’s category (examples shown: Data Science, Java Developer, Business Analyst/Analyst, DevOps engineer, etc.).

Outputs and workflow

  • Creates output folders per category inside a chosen output directory (e.g., categorized resume/Java developer, .../data science).
  • Optionally downloads predictions as a CSV containing:
    • file name
    • category
  • Supports updating automatically for newly uploaded files and allows deleting/resubmitting resumes within the app workflow.
  • Allows bulk selection with a size limit noted in the UI (up to 200MB).

2) Dataset and model training approach

  • Training data comes from Kaggle resume datasets with a CSV structure:
    • Columns: category and resume
  • Datasets referenced:
    • resume_dataset.csv
    • updated_resume_dataset.csv (used in the tutorial)

Framing the task

  • The problem is set up as a supervised multi-class classification task.
  • There are many categories (examples include Data Science, HR Advocate, Arts, Web Designing, Mechanical Engineer, Sales, Health and Fitness, Civil Engineer, Java Developer, Business Analyst, etc.).
  • Dataset size mentioned: ~962 rows, 2 columns (category + resume).

3) Core NLP preprocessing (text cleaning)

Using regex plus NLP tooling (e.g., NLTK stopwords):

  • Libraries include:
    • re
    • nltk stopwords, etc.
  • The cleaning process removes:
    • URLs
    • emails
    • special characters (via regex)
    • stopwords (NLTK English stopwords)

A clean() function is defined and applied across the dataset, for example:

  • df['resume'].apply(lambda x: clean(x))

The cleaned text is then used for vectorization and classification.


4) Feature extraction and label encoding

  • Label encoding converts categorical labels (e.g., “data science”, “Java developer”) into numeric IDs using LabelEncoder.
  • TF-IDF vectorization converts the cleaned resume text into numeric feature vectors using TfidfVectorizer (TF-IDF).

5) Train/test split

  • Data is split using train_test_split:
    • 80% training
    • 20% testing
    • random_state=42

6) Model training + comparison (multiple classifiers)

The tutorial trains several models and compares test accuracy:

  • K-Nearest Neighbors (KNN)
  • Logistic Regression
  • Random Forest
  • Support Vector Classifier (SVC)
  • Multinomial Naive Bayes
  • A One-vs-Rest approach wrapping an estimator for multi-class handling (one-vs-rest logic is evaluated and debugged after an “estimator” usage issue)

Approximate observed accuracies

  • KNN: around 0.98
  • Logistic Regression: around 0.99
  • Random Forest: around 0.98–0.99
  • SVC: around 0.99
  • Multinomial NB: around 0.97
  • One-vs-rest: also evaluated and compared

7) Evaluation and prediction mapping back to category names

  • The selected model predicts a sample resume (based on extracted skills/text).
  • Predictions initially return numeric IDs.
  • A category_map (ID → label) maps numeric outputs back to category names (otherwise defaults to “unknown”).

8) Persisting artifacts (model + vectorizer)

Both the trained model and the TF-IDF vectorizer are saved using pickle:

  • model.pkl
  • tf-idf.pkl

The Streamlit app later loads these artifacts to run predictions on uploaded resumes.


9) Streamlit app implementation details

  • Run command:
    • streamlit run app.py

Key imports mentioned:

  • pickle (to load model + TF-IDF)
  • PyPDF2 (to extract text from uploaded PDFs)
  • re (regex cleaning patterns)
  • streamlit as st

UI elements

  • File upload:
    • st.file_uploader(..., type="pdf", accept_multiple_files=True)
  • Output directory:
    • selection via st.text_input
  • Button to trigger categorization

Per-file processing (for each uploaded PDF)

  • Extract text (first page noted as preferred/simplified)
  • Clean resume text
  • TF-IDF transform
  • Predict category
  • Write outputs:
    • files into corresponding category folders
    • results into a DataFrame for CSV download

10) Handling non-PDF sources (docs → PDF conversion)

The tutorial briefly addresses Kaggle dataset format issues:

  • If resumes are in DOCS format and need conversion to PDF, it provides a utility that:
    • iterates a directory
    • checks file extensions
    • converts using a “convert” method from a docs-to-PDF module

Main speakers / sources

  • Speaker: The video narrator/author (unnamed in subtitles), presenting a “multiverse of 100+ data science project series”.
  • Sources referenced:
    • Kaggle resume datasets (resume_dataset.csv, updated_resume_dataset.csv)
    • Libraries such as NLTK, scikit-learn, PyPDF2, and Streamlit.

Original video