Arabic dialect classifier (18 countries)
Fine-tuned AraBERTv0.2-Twitter that guesses which of 18 Arab countries a tweet's author is from. Labels: OM, SD, SA, KW, QA, LB, JO, SY, IQ, MA, EG, PL, YE, BH, DZ, AE, TN, LY.
How to use
Tweets must go through the same cleaning as in training, then the model.
# pip install transformers torch arabert farasapy
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from arabert.preprocess import ArabertPreprocessor
repo = "Majellan/arabic-dialect-classifier"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
prep = ArabertPreprocessor(model_name="aubmindlab/bert-base-arabertv02-twitter")
text = "..." # an Arabic tweet
enc = tok(prep.preprocess(text), return_tensors="pt", truncation=True, max_length=64)
with torch.no_grad():
probs = torch.softmax(model(**enc).logits, dim=-1)[0]
print(model.config.id2label[probs.argmax().item()], float(probs.max()))
Training
- Data: a Hugging Face copy of QADI (
Abdelrahman-Rezk/Arabic_Dialect_Identification): 440,021 train / 9,164 validation / 8,981 test tweets (31 empty train tweets dropped). - Preprocessing:
ArabertPreprocessorfor the Twitter model, max length 64 tokens. - 4 epochs, learning rate 2e-5, batch size 32, fp16, one T4 GPU.
- Best validation macro F1 was at epoch 3 (0.645). This is the epoch-4 model (validation 0.643).
Results (test set, evaluated once)
| Model | Macro F1 |
|---|---|
| TF-IDF + LinearSVC (untuned baseline) | 0.524 |
| This model | 0.631 (accuracy 0.655) |
With my own grouping of the 18 countries into 5 regions, region accuracy is 0.846. Strongest: EG 0.856, LY 0.795, LB 0.756. Weakest: YE 0.454, BH 0.498, OM 0.515, JO 0.530. The QADI paper reports 0.606 on its own official test set (about 3,300 tweets). That is a different test set from this one, so the numbers are not directly comparable.
Limitations
- Where it makes mistakes. It picks the right region (my own grouping of the 18 countries into 5) for 84.6% of test tweets, but the right country for only 65.5%. The biggest mix-up is Jordan and Palestine (104 and 99 tweets), and six of the eight lowest-scoring countries are Gulf countries (counting Yemen as Gulf).
- Iraq and the Gulf. Iraq is the weakest of the five regions (F1 0.643; 181 of 304 tweets correct), but 10th of 18 as a country. 59% of its 123 mistakes go to Gulf countries (the Gulf is 41% of the other tweets), and 64% of the tweets wrongly predicted as Iraqi are Gulf tweets. I read 20 Iraqi tweets that the model sent to the Gulf: 15 sounded Gulf to me (I couldn't tell which country) and 5 had a clear Iraqi clue, so the model's guess was often reasonable. Two explanations fit, and I can't tell them apart: real overlap between Iraqi and Gulf speech (as a native Arabic speaker from eastern Syria, I know Iraqi dialects vary and parts are close to Kuwaiti and Bahraini), or label noise (see the next point). This is a small, subjective sample of mistakes only.
- QADI labels come from users' account descriptions, not from the tweets. The QADI authors estimate 91.5% label accuracy, so some labels are wrong.
- Scores are model outputs, not calibrated probabilities.
- Trained on tweets only. Short texts, formal Arabic and tweets without dialect clues are guesses.
- Do not use it to make decisions about people.
License and data
No license set. Intended for research and education only. I could not find a clear license for the training data (QADI), so please check it yourself before any other use. Sources: QADI, "Arabic Dialect Identification in the Wild" (arXiv 2005.06557); AraBERT, https://huggingface.co/aubmindlab/bert-base-arabertv02-twitter
- Downloads last month
- 35
Model tree for Majellan/arabic-dialect-classifier
Base model
aubmindlab/bert-base-arabertv02-twitter