Text Generation
Transformers
Safetensors
PyTorch
English
French
Spanish
lfm2
classification
inference-only
structured-generation
constrained-decoding
apple-silicon
conversational
Instructions to use notnotsamuel/LFM2.5-350M-RLCD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use notnotsamuel/LFM2.5-350M-RLCD with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="notnotsamuel/LFM2.5-350M-RLCD") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("notnotsamuel/LFM2.5-350M-RLCD") model = AutoModelForCausalLM.from_pretrained("notnotsamuel/LFM2.5-350M-RLCD", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use notnotsamuel/LFM2.5-350M-RLCD with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "notnotsamuel/LFM2.5-350M-RLCD" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "notnotsamuel/LFM2.5-350M-RLCD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/notnotsamuel/LFM2.5-350M-RLCD
- SGLang
How to use notnotsamuel/LFM2.5-350M-RLCD with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "notnotsamuel/LFM2.5-350M-RLCD" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "notnotsamuel/LFM2.5-350M-RLCD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "notnotsamuel/LFM2.5-350M-RLCD" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "notnotsamuel/LFM2.5-350M-RLCD", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use notnotsamuel/LFM2.5-350M-RLCD with Docker Model Runner:
docker model run hf.co/notnotsamuel/LFM2.5-350M-RLCD
| # Execution notes | |
| The final measurements were made on 2026-09-16. No model training or weight saving was performed. GPU measurements used transient L40S and H100 benchmark jobs. | |
| Initial diagnostic measurements were retained under `results/exploratory/`. The final engine adds JSON Schema meta-validation at request time so malformed enum definitions are rejected before inference. Final benchmark runs include that overhead equally in both methods. The final harness additionally records full dependency inventories and includes the separate scaling suite. The initial and final task labels, model revision, precision, prompts and candidate-scoring method are unchanged. | |
| During validation, a test expecting ValueError for an unsupported `allOf: []` schema instead received jsonschema.SchemaError because the schema itself was invalid. This was a test-fixture issue, not a model/cache failure. The corrected test uses a valid but unsupported `allOf: [{}]` schema and separately tests rejection of a malformed enum schema. The failed GPU jobs aborted before benchmarking. All final device checks pass. | |
| The larger Mac probes show sizeable per-repeat variation. We retain every measurement and report means without discarding outliers or selecting faster runs. The final tables use exactly the final run per device/suite, not the faster of the exploratory and final measurements. Inspect raw latency rows before interpreting the aggregated ratios. | |
| Final GPU source hashes match the packaged inference and benchmark code. Numerical comparison tolerances account for FP16 differences from batch and kernel ordering. No optimized causal-conv1d kernel was installed; both paths use the reference implementation. This can materially affect absolute latency and relative speedups versus a fully optimized serving stack. | |
| The 28-field constrained decision differs by one field on L40S compared with MPS/H100 (64.3% versus 60.7% field accuracy). FP16 decision margins can be sensitive to backend numerical differences. Each device was internally consistent across its three repeats; bitwise cross-device equivalence is not claimed. | |
| After the benchmarks, the original model weights, configuration and tokenizer were bundled byte-for-byte from the same pinned revision. This packaging update does not change the measured inference code, model parameters, or results. SHA-256 checksums are recorded in `BASE_MODEL_MANIFEST.json`. | |