ProseLens-ModernBERT-L

ProseLens-ModernBERT-L is a prose-level detector: it reads the document itself and judges who wrote the words. It is the ModernBERT counterpart of ProseLens and the comparison point for the idea-level detectors.

Model ModernBERT-large with a sequence-classification head
Reads the document text itself
Training data WildOutlines, train split
Output P(human); a document is flagged as AI when P(human) is below a cut
Default cut 0.18790 (global, 1% false-positive rate)
Hardware any GPU; a CPU works for small jobs

Usage

Score documents directly; no outline is needed. The idealens package (PyPI) applies this model and the thresholds in this repo:

pip install "idealens[hf]"
idealens score docs.jsonl -o scores.jsonl --model ProseLens-ModernBERT-L

Input is JSONL with a text field per document.

In Python:

import idealens as il

with il.Detector("ProseLens-ModernBERT-L") as det:
    records = det.score_documents([open("document.txt").read()])
print(records[0]["p_human"], records[0]["verdict"]["ai"])

Thresholds

thresholds.json holds this model's cuts at 0.1%, 0.5%, 1%, 2% and 5% false-positive rates, fitted on the 80,000 human documents of WildOutlines' calibration split: one global cut per rate, plus per-format cuts. The package applies them. A cut fitted for one model does not transfer to another model's scores. For documents unlike English web text, fit cuts on human documents from your own domain with idealens calibrate.

Related

License

CC BY-NC-SA 4.0: free to share and adapt for non-commercial purposes, with attribution and under the same license. Built on answerdotai/ModernBERT-large (Apache 2.0).

Citation

@article{idealens2026,
  title         = {IdeaLens: Detecting AI Ideas in Long-form Writing},
  author        = {Rajendhran, Rishanth and Choi, Minjoon and Russell, Jenna and Namuduri, Ramya and B{\"o}l{\"o}ni-Turgut, Deniz and Karpinska, Marzena and Wieting, John and Iyyer, Mohit},
  journal       = {arXiv preprint arXiv:2610.06778},
  year          = {2026},
  eprint        = {2610.06778},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2610.06778}
}
Downloads last month
31
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rishanthrajendhran/ProseLens-ModernBERT-L

Finetuned
(395)
this model

Dataset used to train rishanthrajendhran/ProseLens-ModernBERT-L

Collection including rishanthrajendhran/ProseLens-ModernBERT-L

Paper for rishanthrajendhran/ProseLens-ModernBERT-L