JobAnalyze_6k / model /MODEL.md
akshaybabu06's picture
Files
ab331fa verified
|
Raw
History Blame Contribute Delete
6.43 kB

JobAnalyze 6k v1.0 Skill Classifier

A lightweight multi-label PyTorch model that predicts a fixed set of skills/keywords from a job description plus a provided role and job type.

This README documents the exact artifacts and behavior implemented in:

  • model/model.py (training definition)
  • model/pred.py (inference wrapper)
  • model/prep/data_prep.py (feature creation)

Sample runs and evaluation numbers referenced from:

  • data/sample_data/test.txt
  • data/sample_data/eval.txt
  • data/sample_data/cli.txt

Model summary

Task type

  • Multi-label classification (each skill is predicted independently)

Inputs

  • job_desc (job description text)
  • role (free text, appended)
  • job_type / type (free text, appended)

These are concatenated during prediction as:

{job_desc} {role} {job_type}

Features

  • TF-IDF features created by model/prep/data_prep.py using:
    • TfidfVectorizer(max_features=150, stop_words='english', ngram_range=(1, 2), min_df=2)
  • Artifacts:
    • model/prep/vectorizer.pkl
    • model/prep/label_vocab.json
    • model/prep/prepared_data.npz

Labels

  • A fixed vocabulary of skills/keywords stored in model/prep/label_vocab.json
  • In training artifacts:
    • NUM_LABELS = len(VOCAB)
    • In sample data: 48 Keywords/labels

Network architecture (the “6k / 6000 parameter” model)

model/model.py defines a small feed-forward network:

  • Linear(input_dim → hidden_dim=32)
  • ReLU
  • Dropout(p=0.3)
  • Linear(hidden_dim=32 → num_labels)

The file name and training printout refer to the total parameter count computed at runtime.


Training (model/model.py)

Do not run model/model.py directly for day-to-day use. It is designed to be executed via pipeline.py (and/or the notebooks).

Training uses:

  • Loss: torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight)
    • pos_weight is computed per-label from the training set (class imbalance handling)
    • Clamped with max=10.0
  • Optimizer: Adam(lr=1e-3, weight_decay=1e-4)
  • Epochs: 300

Outputs (saved artifacts)

At the end of training, the following are written to model_out/:

  • model_out/skill_classifier.pt
  • model_out/training_history.json

The test suite asserts these exist (see test/test_model.py).


Inference (model/pred.py)

model/pred.py exposes a prediction helper:

JobAnalyze_6k(job_desc, role="", job_type="", top_k=50) -> List[(str, float)]

Key behavior:

  • Loads artifacts from repo-relative paths:
    • model/prep/label_vocab.json
    • model/prep/vectorizer.pkl
    • model_out/skill_classifier.pt
  • Vectorizes the concatenated text with TF-IDF and produces logits through the trained network.
  • Converts logits to probabilities with sigmoid.
  • Ranks all labels by probability descending and returns the top results.

Note: top_k is capped by the number of labels effectively returned from the ranked list (in the code, slicing is applied directly).


Evaluation snapshot (from data/sample_data)

Micro/Macro F1

From data/sample_data/eval.txt:

  • Micro-F1: 0.624
  • Macro-F1: 0.420

Baseline comparison

Also from data/sample_data/eval.txt, baseline always predicts a fixed set of frequent labels:

  • Baseline Micro-F1: 0.538
  • Baseline Macro-F1: 0.152

The evaluation script prints:

“Model meaningfully beats the naive baseline.”

Per-label and “trap” analysis

The evaluation output includes per-label precision/recall/F1 and an additional heuristic:

  • “trap?” flags labels where the model does not perform better than a trivial always-zero expectation (within a small tolerance).

From data/sample_data/eval.txt:

  • Right: 15
  • Wrong: 33
  • Total labels: 48
  • Keyword Accuracy: 31.25%

Example: single inference test (data/sample_data/test.txt)

data/sample_data/test.txt contains a job description plus:

  • Role: AI Engineer
  • Type: Junior

It lists 15 keys to be predicted, including:

  • Python, LLMs, LangGraph, MCP, GenAI, VectorDB, SQL, APIs, Docker, Agents, Github, CI/CD, Git, AWS/Azure, Prompt Engineering

Reported performance:

  • Accuracy (recall): 93.34% (14/15)

Example: CLI output format (data/sample_data/cli.txt)

cli/jobauto.py uses JobAnalyze_6k from model.pred and prints:

  • The job description provided
  • Role and Type
  • A ranked “TOP Skills” list as {label} {probability} plus a text bar

From data/sample_data/cli.txt, the top-ranked skills include (with example probabilities):

  • apis ~0.78
  • langgraph ~0.78
  • vectordb ~0.76
  • mcp ~0.75
  • langchain ~0.74
  • rag ~0.72

Data prep artifacts (model/prep/data_prep.py)

model/prep/data_prep.py creates all required inputs for training and inference.

It:

  1. Loads cleaned job description dataset
  2. Normalizes and fixes known skill spelling issues (SKILLS_FIX)
  3. Applies synonym replacement using model/prep/sym_map.py
  4. Builds a multi-hot label vector for the vocabulary
  5. Vectorizes job text with TF-IDF
  6. Splits into train/test and saves:
    • prepared_data.npz
    • label_vocab.json
    • vectorizer.pkl

Reproducibility / how to use

Requirements (high level)

See requirements.txt and pyproject.toml.

Minimum artifacts required for prediction

For model/pred.py to work, these must exist:

  • model/prep/label_vocab.json
  • model/prep/vectorizer.pkl
  • model_out/skill_classifier.pt

If any are missing, model/pred.py raises a FileNotFoundError with guidance.

Recommended workflow

  • Run the full pipeline via pipeline.py (which coordinates data prep + training).
  • Use cli/jobauto.py for interactive predictions.

Notes on “JobAnalyze 6k”

Despite the “6k / 6000 parameter” naming, the true parameter count is computed dynamically in model/model.py.

The implementation is intentionally small:

  • TF-IDF input features
  • 1 hidden layer with 32 units
  • Multi-label BCE loss with class imbalance reweighting

This design keeps inference fast and model size small while still providing meaningful gains over the naive baseline (see evaluation snapshot above).