Instructions to use rishanthrajendhran/ProseLens-ModernBERT-L with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rishanthrajendhran/ProseLens-ModernBERT-L with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="rishanthrajendhran/ProseLens-ModernBERT-L")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("rishanthrajendhran/ProseLens-ModernBERT-L") model = AutoModelForSequenceClassification.from_pretrained("rishanthrajendhran/ProseLens-ModernBERT-L", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ProseLens-ModernBERT-L
ProseLens-ModernBERT-L is a prose-level detector: it reads the document itself and judges who wrote the words. It is the ModernBERT counterpart of ProseLens and the comparison point for the idea-level detectors.
| Model | ModernBERT-large with a sequence-classification head |
| Reads | the document text itself |
| Training data | WildOutlines, train split |
| Output | P(human); a document is flagged as AI when P(human) is below a cut |
| Default cut | 0.18790 (global, 1% false-positive rate) |
| Hardware | any GPU; a CPU works for small jobs |
Usage
Score documents directly; no outline is needed. The idealens package (PyPI) applies this model and the thresholds in this repo:
pip install "idealens[hf]"
idealens score docs.jsonl -o scores.jsonl --model ProseLens-ModernBERT-L
Input is JSONL with a text field per document.
In Python:
import idealens as il
with il.Detector("ProseLens-ModernBERT-L") as det:
records = det.score_documents([open("document.txt").read()])
print(records[0]["p_human"], records[0]["verdict"]["ai"])
Thresholds
thresholds.json holds this model's cuts at 0.1%, 0.5%, 1%, 2% and 5% false-positive rates, fitted on the 80,000 human
documents of WildOutlines' calibration split: one global cut per rate, plus per-format cuts. The package applies them. A cut
fitted for one model does not transfer to another model's scores. For documents unlike English web text, fit cuts on
human documents from your own domain with idealens calibrate.
Related
- IdeaLens: the main idea-level detector, with full documentation
- ProseLens: its prose-level counterpart
- idealens: the Python package that runs these models
- WildOutlines: the training corpus
- IdeaLens collection: every model and dataset in one place
License
CC BY-NC-SA 4.0: free to share and adapt for non-commercial purposes, with attribution and under the same license. Built on answerdotai/ModernBERT-large (Apache 2.0).
Citation
@article{idealens2026,
title = {IdeaLens: Detecting AI Ideas in Long-form Writing},
author = {Rajendhran, Rishanth and Choi, Minjoon and Russell, Jenna and Namuduri, Ramya and B{\"o}l{\"o}ni-Turgut, Deniz and Karpinska, Marzena and Wieting, John and Iyyer, Mohit},
journal = {arXiv preprint arXiv:2610.06778},
year = {2026},
eprint = {2610.06778},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.06778}
}
- Downloads last month
- 31
Model tree for rishanthrajendhran/ProseLens-ModernBERT-L
Base model
answerdotai/ModernBERT-large