Text Generation
GGUF
Japanese
japanese
instruction-tuning
little-language-model
tiny-language-model
edge-ai
embedded-ai
ex-word
llama-cpp
lm-studio
custom-code
conversational
Instructions to use ToTo-40417/EXLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- Ollama
How to use ToTo-40417/EXLLM with Ollama:
ollama run hf.co/ToTo-40417/EXLLM:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM:F16
- Lemonade
How to use ToTo-40417/EXLLM with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM:F16
Run and chat with the model
lemonade run user.EXLLM-F16
List all available models
lemonade list
- Atomic Chat
Download tools/export_exq12.py from ToTo-40417/EXLLM: direct link, hf CLI and curl.
- Browser
- Download file 3.52 kB
-
https://huggingface.co/ToTo-40417/EXLLM/resolve/main/tools/export_exq12.py
- Command line
-
hf download hf://ToTo-40417/EXLLM/tools/export_exq12.py
-
curl -L -o export_exq12.py https://huggingface.co/ToTo-40417/EXLLM/resolve/main/tools/export_exq12.py
3.52 kB
| #!/usr/bin/env python3 | |
| """Convert the published EXLLM8 artifact to the EX-word EXQ12 format.""" | |
| import argparse | |
| import hashlib | |
| import json | |
| import struct | |
| from pathlib import Path | |
| MAGIC_IN = b"EXLLM8\0\0" | |
| MAGIC_OUT = b"EXQ12\0\0\0" | |
| def read_exllm8(path): | |
| raw = path.read_bytes() | |
| if raw[:8] != MAGIC_IN: | |
| raise ValueError("input is not EXLLM8") | |
| offset = 8 | |
| version, count = struct.unpack_from("<II", raw, offset) | |
| offset += 8 | |
| if version != 1: | |
| raise ValueError(f"unsupported EXLLM8 version: {version}") | |
| records = [] | |
| for _ in range(count): | |
| name_len, ndim, qtype = struct.unpack_from("<HBB", raw, offset) | |
| offset += 4 | |
| name = raw[offset : offset + name_len] | |
| offset += name_len | |
| dims = struct.unpack_from("<" + "I" * ndim, raw, offset) | |
| offset += 4 * ndim | |
| scale_count, payload_bytes = struct.unpack_from("<II", raw, offset) | |
| offset += 8 | |
| scales = struct.unpack_from("<" + "f" * scale_count, raw, offset) if scale_count else () | |
| offset += 4 * scale_count | |
| payload = raw[offset : offset + payload_bytes] | |
| offset += payload_bytes | |
| records.append((name, qtype, dims, scales, payload)) | |
| if offset != len(raw): | |
| raise ValueError("trailing bytes in EXLLM8 input") | |
| return records | |
| def convert(source, output): | |
| records = read_exllm8(source) | |
| with output.open("wb") as stream: | |
| stream.write(MAGIC_OUT) | |
| stream.write(struct.pack("<II", 1, len(records))) | |
| for name, qtype, dims, scales, payload in records: | |
| if qtype == 1: | |
| fixed_scales = [max(1, min(0x7FFFFFFF, round(scale * (1 << 20)))) for scale in scales] | |
| output_qtype = 1 | |
| elif qtype == 2: | |
| values = struct.unpack("<" + "e" * (len(payload) // 2), payload) | |
| payload = struct.pack( | |
| "<" + "h" * len(values), | |
| *(max(-32768, min(32767, round(value * 4096))) for value in values), | |
| ) | |
| fixed_scales = [] | |
| output_qtype = 3 | |
| else: | |
| raise ValueError(f"unsupported EXLLM8 qtype: {qtype}") | |
| stream.write(struct.pack("<HBB", len(name), len(dims), output_qtype)) | |
| stream.write(name) | |
| stream.write(struct.pack("<" + "I" * len(dims), *dims)) | |
| stream.write(struct.pack("<II", len(fixed_scales), len(payload))) | |
| if fixed_scales: | |
| stream.write(struct.pack("<" + "i" * len(fixed_scales), *fixed_scales)) | |
| stream.write(payload) | |
| return len(records) | |
| def main(): | |
| parser = argparse.ArgumentParser() | |
| parser.add_argument("--input", type=Path, default=Path("weights/EXLLM-v1.1-5m-int8.bin")) | |
| parser.add_argument("--output", type=Path, default=Path("weights/model.q12")) | |
| parser.add_argument( | |
| "--expect-sha256", | |
| default="d64037fde791e5c0e48101bc1a8ab366a3287f1f36b4464879e36495ea7e5a53", | |
| ) | |
| args = parser.parse_args() | |
| args.output.parent.mkdir(parents=True, exist_ok=True) | |
| tensor_count = convert(args.input, args.output) | |
| digest = hashlib.sha256(args.output.read_bytes()).hexdigest() | |
| result = {"output": str(args.output), "bytes": args.output.stat().st_size, "sha256": digest, "tensors": tensor_count} | |
| print(json.dumps(result, indent=2)) | |
| if args.expect_sha256 and digest != args.expect_sha256: | |
| raise SystemExit("EXQ12 SHA-256 mismatch") | |
| if __name__ == "__main__": | |
| main() | |