JobAnalyze_6k / model /MODEL.md
akshaybabu06's picture
Files
ab331fa verified
|
Raw
History Blame Contribute Delete
6.43 kB
# JobAnalyze 6k v1.0 Skill Classifier
A lightweight **multi-label** PyTorch model that predicts a fixed set of **skills/keywords** from a job description plus a provided **role** and **job type**.
This README documents the exact artifacts and behavior implemented in:
- `model/model.py` (training definition)
- `model/pred.py` (inference wrapper)
- `model/prep/data_prep.py` (feature creation)
Sample runs and evaluation numbers referenced from:
- `data/sample_data/test.txt`
- `data/sample_data/eval.txt`
- `data/sample_data/cli.txt`
---
## Model summary
### Task type
- **Multi-label classification** (each skill is predicted independently)
### Inputs
- `job_desc` (job description text)
- `role` (free text, appended)
- `job_type` / `type` (free text, appended)
These are concatenated during prediction as:
```text
{job_desc} {role} {job_type}
```
### Features
- TF-IDF features created by `model/prep/data_prep.py` using:
- `TfidfVectorizer(max_features=150, stop_words='english', ngram_range=(1, 2), min_df=2)`
- Artifacts:
- `model/prep/vectorizer.pkl`
- `model/prep/label_vocab.json`
- `model/prep/prepared_data.npz`
### Labels
- A fixed vocabulary of skills/keywords stored in `model/prep/label_vocab.json`
- In training artifacts:
- `NUM_LABELS = len(VOCAB)`
- In sample data: **48 Keywords/labels**
### Network architecture (the “6k / 6000 parameter” model)
`model/model.py` defines a small feed-forward network:
- Linear(input_dim → hidden_dim=32)
- ReLU
- Dropout(p=0.3)
- Linear(hidden_dim=32 → num_labels)
The file name and training printout refer to the total parameter count computed at runtime.
---
## Training (model/model.py)
**Do not run `model/model.py` directly for day-to-day use.** It is designed to be executed via `pipeline.py` (and/or the notebooks).
Training uses:
- Loss: `torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight)`
- `pos_weight` is computed per-label from the training set (class imbalance handling)
- Clamped with `max=10.0`
- Optimizer: `Adam(lr=1e-3, weight_decay=1e-4)`
- Epochs: `300`
### Outputs (saved artifacts)
At the end of training, the following are written to `model_out/`:
- `model_out/skill_classifier.pt`
- `model_out/training_history.json`
The test suite asserts these exist (see `test/test_model.py`).
---
## Inference (model/pred.py)
`model/pred.py` exposes a prediction helper:
```python
JobAnalyze_6k(job_desc, role="", job_type="", top_k=50) -> List[(str, float)]
```
Key behavior:
- Loads artifacts from repo-relative paths:
- `model/prep/label_vocab.json`
- `model/prep/vectorizer.pkl`
- `model_out/skill_classifier.pt`
- Vectorizes the concatenated text with TF-IDF and produces logits through the trained network.
- Converts logits to probabilities with `sigmoid`.
- Ranks all labels by probability descending and returns the top results.
> Note: `top_k` is capped by the number of labels effectively returned from the ranked list (in the code, slicing is applied directly).
---
## Evaluation snapshot (from data/sample_data)
### Micro/Macro F1
From `data/sample_data/eval.txt`:
- **Micro-F1: 0.624**
- **Macro-F1: 0.420**
### Baseline comparison
Also from `data/sample_data/eval.txt`, baseline always predicts a fixed set of frequent labels:
- Baseline Micro-F1: **0.538**
- Baseline Macro-F1: **0.152**
The evaluation script prints:
> “Model meaningfully beats the naive baseline.”
### Per-label and “trap” analysis
The evaluation output includes per-label precision/recall/F1 and an additional heuristic:
- “trap?” flags labels where the model does not perform better than a trivial always-zero expectation (within a small tolerance).
From `data/sample_data/eval.txt`:
- Right: **15**
- Wrong: **33**
- Total labels: **48**
- Keyword Accuracy: **31.25%**
---
## Example: single inference test (data/sample_data/test.txt)
`data/sample_data/test.txt` contains a job description plus:
- Role: **AI Engineer**
- Type: **Junior**
It lists **15 keys to be predicted**, including:
- Python, LLMs, LangGraph, MCP, GenAI, VectorDB, SQL, APIs, Docker, Agents, Github, CI/CD, Git, AWS/Azure, Prompt Engineering
Reported performance:
- Accuracy (recall): **93.34%** (14/15)
---
## Example: CLI output format (data/sample_data/cli.txt)
`cli/jobauto.py` uses `JobAnalyze_6k` from `model.pred` and prints:
- The job description provided
- Role and Type
- A ranked “TOP Skills” list as `{label} {probability}` plus a text bar
From `data/sample_data/cli.txt`, the top-ranked skills include (with example probabilities):
- apis ~0.78
- langgraph ~0.78
- vectordb ~0.76
- mcp ~0.75
- langchain ~0.74
- rag ~0.72
---
## Data prep artifacts (model/prep/data_prep.py)
`model/prep/data_prep.py` creates all required inputs for training and inference.
It:
1. Loads cleaned job description dataset
2. Normalizes and fixes known skill spelling issues (`SKILLS_FIX`)
3. Applies synonym replacement using `model/prep/sym_map.py`
4. Builds a multi-hot label vector for the vocabulary
5. Vectorizes job text with TF-IDF
6. Splits into train/test and saves:
- `prepared_data.npz`
- `label_vocab.json`
- `vectorizer.pkl`
---
## Reproducibility / how to use
### Requirements (high level)
See `requirements.txt` and `pyproject.toml`.
### Minimum artifacts required for prediction
For `model/pred.py` to work, these must exist:
- `model/prep/label_vocab.json`
- `model/prep/vectorizer.pkl`
- `model_out/skill_classifier.pt`
If any are missing, `model/pred.py` raises a `FileNotFoundError` with guidance.
### Recommended workflow
- Run the full pipeline via `pipeline.py` (which coordinates data prep + training).
- Use `cli/jobauto.py` for interactive predictions.
---
## Notes on “JobAnalyze 6k”
Despite the “6k / 6000 parameter” naming, the true parameter count is computed dynamically in `model/model.py`.
The implementation is intentionally small:
- TF-IDF input features
- 1 hidden layer with 32 units
- Multi-label BCE loss with class imbalance reweighting
This design keeps inference fast and model size small while still providing meaningful gains over the naive baseline (see evaluation snapshot above).