Text Generation
GGUF
Japanese
japanese
instruction-tuning
little-language-model
tiny-language-model
edge-ai
embedded-ai
ex-word
llama-cpp
lm-studio
custom-code
conversational
Instructions to use ToTo-40417/EXLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM:F16
- Ollama
How to use ToTo-40417/EXLLM with Ollama:
ollama run hf.co/ToTo-40417/EXLLM:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM:F16
- Lemonade
How to use ToTo-40417/EXLLM with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM:F16
Run and chat with the model
lemonade run user.EXLLM-F16
List all available models
lemonade list
- Atomic Chat
ToTo-40417 commited on
Raise LM Studio runtime context to 512
Browse files- EXLLM-0.005B-LMStudio-F16.gguf +1 -1
- README.md +6 -6
- lmstudio/training_manifest.json +2 -1
EXLLM-0.005B-LMStudio-F16.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 10915168
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0e3c519d4a044fac45be84dc607eb8a43de4a7c60083983bf246c5b9c8a43472
|
| 3 |
size 10915168
|
README.md
CHANGED
|
@@ -120,16 +120,16 @@ These are internal regression tests for the intended application scope. They are
|
|
| 120 |
|
| 121 |
### LM Studio / llama.cpp
|
| 122 |
|
| 123 |
-
In LM Studio, search for `ToTo-40417/EXLLM`, download `EXLLM-0.005B-LMStudio-F16.gguf`, and start a chat. The required single-turn chat template is embedded in the GGUF.
|
| 124 |
|
| 125 |
With llama.cpp:
|
| 126 |
|
| 127 |
```bash
|
| 128 |
llama-cli -m EXLLM-0.005B-LMStudio-F16.gguf \
|
| 129 |
-
-p "RAMとは何ですか?" -n 48 --temp 0 --ctx-size
|
| 130 |
```
|
| 131 |
|
| 132 |
-
The GGUF companion uses a Llama-compatible decoder, a 1,024-piece SentencePiece tokenizer with byte fallback, 6 layers, hidden size 288, intermediate size 608, 9 attention/KV heads, and a
|
| 133 |
|
| 134 |
### Original EXLLM reference runtime
|
| 135 |
|
|
@@ -166,7 +166,7 @@ print(answer(model, tokenizer, "RAMとは何ですか?", temperature=0.0))
|
|
| 166 |
| `weights/EXLLM-v1.1-5m-release3.pt` | 21,526,737 bytes (20.53 MiB; 0.021526737 GB) | `48ad00c6640689595a253fc949cad83f3000e14cccdf69275fcf245e85be3ceb` | fp32 PyTorch checkpoint |
|
| 167 |
| `weights/EXLLM-v1.1-5m-int8.bin` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `3b2bf2d9f103cba714bc98976709c5ad34860c1ff55fed2fce91fc95afb5b34a` | EXLLM8 interchange artifact |
|
| 168 |
| `weights/model.q12` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `d64037fde791e5c0e48101bc1a8ab366a3287f1f36b4464879e36495ea7e5a53` | EX-word EXQ12 artifact |
|
| 169 |
-
| `EXLLM-0.005B-LMStudio-F16.gguf` | 10,915,168 bytes (10.41 MiB; 0.010915168 GB) | `
|
| 170 |
|
| 171 |
EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
|
| 172 |
|
|
@@ -181,7 +181,7 @@ EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
|
|
| 181 |
## Limitations
|
| 182 |
|
| 183 |
- The model is not a general-purpose assistant.
|
| 184 |
-
- The
|
| 185 |
- Knowledge coverage is restricted to the project-generated training scope.
|
| 186 |
- Unseen concepts, long instructions, translation, and multi-step reasoning are unreliable.
|
| 187 |
- No RLHF, DPO, tool use, retrieval, web access, or current-information source is included.
|
|
@@ -199,7 +199,7 @@ EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
|
|
| 199 |
|
| 200 |
EXLLM-0.005B-Instructは、低メモリ組込み機器でのローカル推論を目的とした、**0.005377824Bパラメータ**の日本語instruction modelです。主な実装対象はCASIO EX-word XD-B4800です。同じプロジェクトデータから別途学習した、LM Studio / llama.cpp向けの**0.005441184Bパラメータ**GGUF companionも収録しています。
|
| 201 |
|
| 202 |
-
電子辞書用checkpointは独自のdecoder-only Transformerとtokenizerを使用するため、GGUFに直接変換したもの
|
| 203 |
|
| 204 |
RTX 3060でのfp32参照実行は、40回のwarm-up後に120回測定し、TTFT中央値0.003990秒、生成速度239.082 token/sでした。EX-word整数runtimeでは、0.024184 GHzのSH-4A-class CPU上でTTFT中央値23.711秒、生成速度0.52–0.55 token/sでした。
|
| 205 |
|
|
|
|
| 120 |
|
| 121 |
### LM Studio / llama.cpp
|
| 122 |
|
| 123 |
+
In LM Studio, search for `ToTo-40417/EXLLM`, download `EXLLM-0.005B-LMStudio-F16.gguf`, and start a chat. The required single-turn chat template is embedded in the GGUF. Set the loaded context length to **512 tokens** and use greedy decoding or a low temperature; the model is intentionally tiny and is intended for short Japanese prompts. Training sequences were limited to 128 tokens; the larger runtime window reserves space for LM Studio's chat wrapper and short history.
|
| 124 |
|
| 125 |
With llama.cpp:
|
| 126 |
|
| 127 |
```bash
|
| 128 |
llama-cli -m EXLLM-0.005B-LMStudio-F16.gguf \
|
| 129 |
+
-p "RAMとは何ですか?" -n 48 --temp 0 --ctx-size 512 --single-turn
|
| 130 |
```
|
| 131 |
|
| 132 |
+
The GGUF companion uses a Llama-compatible decoder, a 1,024-piece SentencePiece tokenizer with byte fallback, 6 layers, hidden size 288, intermediate size 608, 9 attention/KV heads, 128-token training sequences, and a 512-token runtime context. It was trained separately for five epochs (3,325 optimizer steps) on 85,005 project-data records; best validation loss was 0.0390303. Exact metadata is in [`lmstudio/training_manifest.json`](lmstudio/training_manifest.json).
|
| 133 |
|
| 134 |
### Original EXLLM reference runtime
|
| 135 |
|
|
|
|
| 166 |
| `weights/EXLLM-v1.1-5m-release3.pt` | 21,526,737 bytes (20.53 MiB; 0.021526737 GB) | `48ad00c6640689595a253fc949cad83f3000e14cccdf69275fcf245e85be3ceb` | fp32 PyTorch checkpoint |
|
| 167 |
| `weights/EXLLM-v1.1-5m-int8.bin` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `3b2bf2d9f103cba714bc98976709c5ad34860c1ff55fed2fce91fc95afb5b34a` | EXLLM8 interchange artifact |
|
| 168 |
| `weights/model.q12` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `d64037fde791e5c0e48101bc1a8ab366a3287f1f36b4464879e36495ea7e5a53` | EX-word EXQ12 artifact |
|
| 169 |
+
| `EXLLM-0.005B-LMStudio-F16.gguf` | 10,915,168 bytes (10.41 MiB; 0.010915168 GB) | `0e3c519d4a044fac45be84dc607eb8a43de4a7c60083983bf246c5b9c8a43472` | Separately trained Llama-compatible F16 companion for LM Studio / llama.cpp |
|
| 170 |
|
| 171 |
EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
|
| 172 |
|
|
|
|
| 181 |
## Limitations
|
| 182 |
|
| 183 |
- The model is not a general-purpose assistant.
|
| 184 |
+
- The original embedded checkpoint and the GGUF training sequences use 0.000128M tokens. The GGUF advertises a 0.000512M-token runtime window for LM Studio framing and short history; quality beyond the trained 128-token range is not guaranteed.
|
| 185 |
- Knowledge coverage is restricted to the project-generated training scope.
|
| 186 |
- Unseen concepts, long instructions, translation, and multi-step reasoning are unreliable.
|
| 187 |
- No RLHF, DPO, tool use, retrieval, web access, or current-information source is included.
|
|
|
|
| 199 |
|
| 200 |
EXLLM-0.005B-Instructは、低メモリ組込み機器でのローカル推論を目的とした、**0.005377824Bパラメータ**の日本語instruction modelです。主な実装対象はCASIO EX-word XD-B4800です。同じプロジェクトデータから別途学習した、LM Studio / llama.cpp向けの**0.005441184Bパラメータ**GGUF companionも収録しています。
|
| 201 |
|
| 202 |
+
電子辞書用checkpointは独自のdecoder-only Transformerとtokenizerを使用するため、GGUFに直接変換したもの���はありません。LM Studioでは`ToTo-40417/EXLLM`を検索し、`EXLLM-0.005B-LMStudio-F16.gguf`を選択してください。チャット用templateはGGUFに内蔵済みで、LM Studioのロード時contextは512 tokensに設定します。学習時の系列長は128 tokensであり、追加領域はチャット制御情報と短い履歴のための余白です。このGGUFはプロジェクトデータで別途学習したPC用companionであり、電子辞書版と同一重みではありません。
|
| 203 |
|
| 204 |
RTX 3060でのfp32参照実行は、40回のwarm-up後に120回測定し、TTFT中央値0.003990秒、生成速度239.082 token/sでした。EX-word整数runtimeでは、0.024184 GHzのSH-4A-class CPU上でTTFT中央値23.711秒、生成速度0.52–0.55 token/sでした。
|
| 205 |
|
lmstudio/training_manifest.json
CHANGED
|
@@ -9,7 +9,8 @@
|
|
| 9 |
"intermediate_size": 608,
|
| 10 |
"attention_heads": 9,
|
| 11 |
"kv_heads": 9,
|
| 12 |
-
"
|
|
|
|
| 13 |
"vocabulary_size": 1024
|
| 14 |
},
|
| 15 |
"training_records_including_overlaps": 85005,
|
|
|
|
| 9 |
"intermediate_size": 608,
|
| 10 |
"attention_heads": 9,
|
| 11 |
"kv_heads": 9,
|
| 12 |
+
"training_sequence_length": 128,
|
| 13 |
+
"runtime_context_length": 512,
|
| 14 |
"vocabulary_size": 1024
|
| 15 |
},
|
| 16 |
"training_records_including_overlaps": 85005,
|