Arabic dialect classifier (18 countries)

Fine-tuned AraBERTv0.2-Twitter that guesses which of 18 Arab countries a tweet's author is from. Labels: OM, SD, SA, KW, QA, LB, JO, SY, IQ, MA, EG, PL, YE, BH, DZ, AE, TN, LY.

How to use

Tweets must go through the same cleaning as in training, then the model.

# pip install transformers torch arabert farasapy
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from arabert.preprocess import ArabertPreprocessor

repo = "Majellan/arabic-dialect-classifier"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
prep = ArabertPreprocessor(model_name="aubmindlab/bert-base-arabertv02-twitter")

text = "..."  # an Arabic tweet
enc = tok(prep.preprocess(text), return_tensors="pt", truncation=True, max_length=64)
with torch.no_grad():
    probs = torch.softmax(model(**enc).logits, dim=-1)[0]
print(model.config.id2label[probs.argmax().item()], float(probs.max()))

Training

  • Data: a Hugging Face copy of QADI (Abdelrahman-Rezk/Arabic_Dialect_Identification): 440,021 train / 9,164 validation / 8,981 test tweets (31 empty train tweets dropped).
  • Preprocessing: ArabertPreprocessor for the Twitter model, max length 64 tokens.
  • 4 epochs, learning rate 2e-5, batch size 32, fp16, one T4 GPU.
  • Best validation macro F1 was at epoch 3 (0.645). This is the epoch-4 model (validation 0.643).

Results (test set, evaluated once)

Model Macro F1
TF-IDF + LinearSVC (untuned baseline) 0.524
This model 0.631 (accuracy 0.655)

With my own grouping of the 18 countries into 5 regions, region accuracy is 0.846. Strongest: EG 0.856, LY 0.795, LB 0.756. Weakest: YE 0.454, BH 0.498, OM 0.515, JO 0.530. The QADI paper reports 0.606 on its own official test set (about 3,300 tweets). That is a different test set from this one, so the numbers are not directly comparable.

Limitations

    • Where it makes mistakes. It picks the right region (my own grouping of the 18 countries into 5) for 84.6% of test tweets, but the right country for only 65.5%. The biggest mix-up is Jordan and Palestine (104 and 99 tweets), and six of the eight lowest-scoring countries are Gulf countries (counting Yemen as Gulf).
  • Iraq and the Gulf. Iraq is the weakest of the five regions (F1 0.643; 181 of 304 tweets correct), but 10th of 18 as a country. 59% of its 123 mistakes go to Gulf countries (the Gulf is 41% of the other tweets), and 64% of the tweets wrongly predicted as Iraqi are Gulf tweets. I read 20 Iraqi tweets that the model sent to the Gulf: 15 sounded Gulf to me (I couldn't tell which country) and 5 had a clear Iraqi clue, so the model's guess was often reasonable. Two explanations fit, and I can't tell them apart: real overlap between Iraqi and Gulf speech (as a native Arabic speaker from eastern Syria, I know Iraqi dialects vary and parts are close to Kuwaiti and Bahraini), or label noise (see the next point). This is a small, subjective sample of mistakes only.
  • QADI labels come from users' account descriptions, not from the tweets. The QADI authors estimate 91.5% label accuracy, so some labels are wrong.
  • Scores are model outputs, not calibrated probabilities.
  • Trained on tweets only. Short texts, formal Arabic and tweets without dialect clues are guesses.
  • Do not use it to make decisions about people.

License and data

No license set. Intended for research and education only. I could not find a clear license for the training data (QADI), so please check it yourself before any other use. Sources: QADI, "Arabic Dialect Identification in the Wild" (arXiv 2005.06557); AraBERT, https://huggingface.co/aubmindlab/bert-base-arabertv02-twitter

Downloads last month
35
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Majellan/arabic-dialect-classifier

Finetuned
(46)
this model