--- library_name: peft pipeline_tag: text-generation base_model: Qwen/Qwen2.5-1.5B-Instruct license: mit tags: - co-lmlm - annotation - lora --- # CoLMLM-Question-Generator The **question generator** used to build the training corpora for [**Co-LMLM: Continuous-Query Limited Memory Language Models**](https://arxiv.org/abs/2607.07707). Co-LMLM is trained on text in which each factual span carries the question it answers. Producing those questions with a frontier LLM is far too expensive to run over a pretraining-scale corpus, so this model distills that step: given a document whose fact spans are already marked and numbered, and the id of one of them, it emits the question that span answers plus a paraphrased answer. It is the second stage of a two-stage annotation pipeline. The first stage, [CoLMLM-Fact-Span-Annotator](https://huggingface.co/lil-lab/CoLMLM-Fact-Span-Annotator), marks the spans this model is asked about. This repository contains a **LoRA adapter**, not a standalone model — the base weights are loaded from `Qwen/Qwen2.5-1.5B-Instruct` at inference time. ## Model details | | | | ------------------- | ------------------------------------------------------------------------------------------------------- | | Base model | [`Qwen/Qwen2.5-1.5B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) | | Adaptation | LoRA, r=64, α=64, dropout 0.05, on all linear projections (`q,k,v,o,gate,up,down`) | | Also trained | embeddings for the 8 added annotation tokens (``, ``, ``, ``, ``, ``, ``, ``) | | Precision | bfloat16 | | Sequence length | 8192 tokens | ## Prompt format The user message is the numbered document, then ``, then the question for one fact id. The chat template is Qwen's default (no system prompt is supplied, so Qwen's default system block is used — matching training). ``` Nspan tags> What are the question and paraphrased answer for N? ``` The model responds with `......`. ## Usage For more details and the full annotation pipeline, see the code repository: 👉 **[github.com/lil-lab/Co-LMLM](https://github.com/lil-lab/Co-LMLM)** Standalone: ```python import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer adapter_id = "lil-lab/CoLMLM-Question-Generator" base_id = "Qwen/Qwen2.5-1.5B-Instruct" tokenizer = AutoTokenizer.from_pretrained(adapter_id) model = AutoModelForCausalLM.from_pretrained(base_id, dtype=torch.bfloat16) model = PeftModel.from_pretrained(model, adapter_id).eval() context = ("Marie Curie was born in 1Warsaw in " "21867 and won 3two Nobel Prizes.") fact_id = 1 user = (f"{context}\n\n" f"What are the question and paraphrased answer for {fact_id}?\n") prompt = tokenizer.apply_chat_template([{"role": "user", "content": user}], add_generation_prompt=True, tokenize=False) inputs = tokenizer(prompt, return_tensors="pt").to(model.device) with torch.no_grad(): output = model.generate(**inputs, max_new_tokens=64, do_sample=False) print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=False)) # Where was Marie Curie born?Warsaw<|im_end|> ``` This model is part of the [**Co-LMLM** collection](https://huggingface.co/collections/lil-lab/co-lmlm-6a4e8216d55eae83af348f57). ## Citation ```bibtex @misc{feldman2026colmlmcontinuousquerylimitedmemory, title={Co-LMLM: Continuous-Query Limited Memory Language Models}, author={Yair Feldman and Linxi Zhao and Nathan Godey and Dongyoung Go and Yilun Hua and Kilian Q. Weinberger and Jennifer J. Sun and Yoav Artzi}, year={2026}, eprint={2607.07707}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2607.07707}, } ```