JobAnalyze 6k v1.0 Skill Classifier
A lightweight multi-label PyTorch model that predicts a fixed set of skills/keywords from a job description plus a provided role and job type.
This README documents the exact artifacts and behavior implemented in:
model/model.py(training definition)model/pred.py(inference wrapper)model/prep/data_prep.py(feature creation)
Sample runs and evaluation numbers referenced from:
data/sample_data/test.txtdata/sample_data/eval.txtdata/sample_data/cli.txt
Model summary
Task type
- Multi-label classification (each skill is predicted independently)
Inputs
job_desc(job description text)role(free text, appended)job_type/type(free text, appended)
These are concatenated during prediction as:
{job_desc} {role} {job_type}
Features
- TF-IDF features created by
model/prep/data_prep.pyusing:TfidfVectorizer(max_features=150, stop_words='english', ngram_range=(1, 2), min_df=2)
- Artifacts:
model/prep/vectorizer.pklmodel/prep/label_vocab.jsonmodel/prep/prepared_data.npz
Labels
- A fixed vocabulary of skills/keywords stored in
model/prep/label_vocab.json - In training artifacts:
NUM_LABELS = len(VOCAB)- In sample data: 48 Keywords/labels
Network architecture (the “6k / 6000 parameter” model)
model/model.py defines a small feed-forward network:
- Linear(input_dim → hidden_dim=32)
- ReLU
- Dropout(p=0.3)
- Linear(hidden_dim=32 → num_labels)
The file name and training printout refer to the total parameter count computed at runtime.
Training (model/model.py)
Do not run model/model.py directly for day-to-day use. It is designed to be executed via pipeline.py (and/or the notebooks).
Training uses:
- Loss:
torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight)pos_weightis computed per-label from the training set (class imbalance handling)- Clamped with
max=10.0
- Optimizer:
Adam(lr=1e-3, weight_decay=1e-4) - Epochs:
300
Outputs (saved artifacts)
At the end of training, the following are written to model_out/:
model_out/skill_classifier.ptmodel_out/training_history.json
The test suite asserts these exist (see test/test_model.py).
Inference (model/pred.py)
model/pred.py exposes a prediction helper:
JobAnalyze_6k(job_desc, role="", job_type="", top_k=50) -> List[(str, float)]
Key behavior:
- Loads artifacts from repo-relative paths:
model/prep/label_vocab.jsonmodel/prep/vectorizer.pklmodel_out/skill_classifier.pt
- Vectorizes the concatenated text with TF-IDF and produces logits through the trained network.
- Converts logits to probabilities with
sigmoid. - Ranks all labels by probability descending and returns the top results.
Note:
top_kis capped by the number of labels effectively returned from the ranked list (in the code, slicing is applied directly).
Evaluation snapshot (from data/sample_data)
Micro/Macro F1
From data/sample_data/eval.txt:
- Micro-F1: 0.624
- Macro-F1: 0.420
Baseline comparison
Also from data/sample_data/eval.txt, baseline always predicts a fixed set of frequent labels:
- Baseline Micro-F1: 0.538
- Baseline Macro-F1: 0.152
The evaluation script prints:
“Model meaningfully beats the naive baseline.”
Per-label and “trap” analysis
The evaluation output includes per-label precision/recall/F1 and an additional heuristic:
- “trap?” flags labels where the model does not perform better than a trivial always-zero expectation (within a small tolerance).
From data/sample_data/eval.txt:
- Right: 15
- Wrong: 33
- Total labels: 48
- Keyword Accuracy: 31.25%
Example: single inference test (data/sample_data/test.txt)
data/sample_data/test.txt contains a job description plus:
- Role: AI Engineer
- Type: Junior
It lists 15 keys to be predicted, including:
- Python, LLMs, LangGraph, MCP, GenAI, VectorDB, SQL, APIs, Docker, Agents, Github, CI/CD, Git, AWS/Azure, Prompt Engineering
Reported performance:
- Accuracy (recall): 93.34% (14/15)
Example: CLI output format (data/sample_data/cli.txt)
cli/jobauto.py uses JobAnalyze_6k from model.pred and prints:
- The job description provided
- Role and Type
- A ranked “TOP Skills” list as
{label} {probability}plus a text bar
From data/sample_data/cli.txt, the top-ranked skills include (with example probabilities):
- apis ~0.78
- langgraph ~0.78
- vectordb ~0.76
- mcp ~0.75
- langchain ~0.74
- rag ~0.72
Data prep artifacts (model/prep/data_prep.py)
model/prep/data_prep.py creates all required inputs for training and inference.
It:
- Loads cleaned job description dataset
- Normalizes and fixes known skill spelling issues (
SKILLS_FIX) - Applies synonym replacement using
model/prep/sym_map.py - Builds a multi-hot label vector for the vocabulary
- Vectorizes job text with TF-IDF
- Splits into train/test and saves:
prepared_data.npzlabel_vocab.jsonvectorizer.pkl
Reproducibility / how to use
Requirements (high level)
See requirements.txt and pyproject.toml.
Minimum artifacts required for prediction
For model/pred.py to work, these must exist:
model/prep/label_vocab.jsonmodel/prep/vectorizer.pklmodel_out/skill_classifier.pt
If any are missing, model/pred.py raises a FileNotFoundError with guidance.
Recommended workflow
- Run the full pipeline via
pipeline.py(which coordinates data prep + training). - Use
cli/jobauto.pyfor interactive predictions.
Notes on “JobAnalyze 6k”
Despite the “6k / 6000 parameter” naming, the true parameter count is computed dynamically in model/model.py.
The implementation is intentionally small:
- TF-IDF input features
- 1 hidden layer with 32 units
- Multi-label BCE loss with class imbalance reweighting
This design keeps inference fast and model size small while still providing meaningful gains over the naive baseline (see evaluation snapshot above).