ToTo-40417 commited on
Commit
160045e
·
unverified ·
1 Parent(s): 8ac1edb

Raise LM Studio runtime context to 512

Browse files
EXLLM-0.005B-LMStudio-F16.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6f649e74fd006c4c487a141307a7598531a1ff1ce8ba5e45d45f9c590f399759
3
  size 10915168
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e3c519d4a044fac45be84dc607eb8a43de4a7c60083983bf246c5b9c8a43472
3
  size 10915168
README.md CHANGED
@@ -120,16 +120,16 @@ These are internal regression tests for the intended application scope. They are
120
 
121
  ### LM Studio / llama.cpp
122
 
123
- In LM Studio, search for `ToTo-40417/EXLLM`, download `EXLLM-0.005B-LMStudio-F16.gguf`, and start a chat. The required single-turn chat template is embedded in the GGUF. Keep the context length at **128 tokens** and use greedy decoding or a low temperature; the model is intentionally tiny and is intended for short Japanese prompts.
124
 
125
  With llama.cpp:
126
 
127
  ```bash
128
  llama-cli -m EXLLM-0.005B-LMStudio-F16.gguf \
129
- -p "RAMとは何ですか?" -n 48 --temp 0 --ctx-size 128 --single-turn
130
  ```
131
 
132
- The GGUF companion uses a Llama-compatible decoder, a 1,024-piece SentencePiece tokenizer with byte fallback, 6 layers, hidden size 288, intermediate size 608, 9 attention/KV heads, and a 128-token context. It was trained separately for five epochs (3,325 optimizer steps) on 85,005 project-data records; best validation loss was 0.0390303. Exact metadata is in [`lmstudio/training_manifest.json`](lmstudio/training_manifest.json).
133
 
134
  ### Original EXLLM reference runtime
135
 
@@ -166,7 +166,7 @@ print(answer(model, tokenizer, "RAMとは何ですか?", temperature=0.0))
166
  | `weights/EXLLM-v1.1-5m-release3.pt` | 21,526,737 bytes (20.53 MiB; 0.021526737 GB) | `48ad00c6640689595a253fc949cad83f3000e14cccdf69275fcf245e85be3ceb` | fp32 PyTorch checkpoint |
167
  | `weights/EXLLM-v1.1-5m-int8.bin` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `3b2bf2d9f103cba714bc98976709c5ad34860c1ff55fed2fce91fc95afb5b34a` | EXLLM8 interchange artifact |
168
  | `weights/model.q12` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `d64037fde791e5c0e48101bc1a8ab366a3287f1f36b4464879e36495ea7e5a53` | EX-word EXQ12 artifact |
169
- | `EXLLM-0.005B-LMStudio-F16.gguf` | 10,915,168 bytes (10.41 MiB; 0.010915168 GB) | `6f649e74fd006c4c487a141307a7598531a1ff1ce8ba5e45d45f9c590f399759` | Separately trained Llama-compatible F16 companion for LM Studio / llama.cpp |
170
 
171
  EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
172
 
@@ -181,7 +181,7 @@ EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
181
  ## Limitations
182
 
183
  - The model is not a general-purpose assistant.
184
- - The context limit is 0.000128M tokens.
185
  - Knowledge coverage is restricted to the project-generated training scope.
186
  - Unseen concepts, long instructions, translation, and multi-step reasoning are unreliable.
187
  - No RLHF, DPO, tool use, retrieval, web access, or current-information source is included.
@@ -199,7 +199,7 @@ EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
199
 
200
  EXLLM-0.005B-Instructは、低メモリ組込み機器でのローカル推論を目的とした、**0.005377824Bパラメータ**の日本語instruction modelです。主な実装対象はCASIO EX-word XD-B4800です。同じプロジェクトデータから別途学習した、LM Studio / llama.cpp向けの**0.005441184Bパラメータ**GGUF companionも収録しています。
201
 
202
- 電子辞書用checkpointは独自のdecoder-only Transformerとtokenizerを使用するため、GGUFに直接変換したものではありません。LM Studioでは`ToTo-40417/EXLLM`を検索し、`EXLLM-0.005B-LMStudio-F16.gguf`を選択してください。チャット用templateはGGUFに内蔵済みで、contextは128 tokensです。このGGUFはプロジェクトデータで別途学習したPC用companionであり、電子辞書版と同一重みではありません。
203
 
204
  RTX 3060でのfp32参照実行は、40回のwarm-up後に120回測定し、TTFT中央値0.003990秒、生成速度239.082 token/sでした。EX-word整数runtimeでは、0.024184 GHzのSH-4A-class CPU上でTTFT中央値23.711秒、生成速度0.52–0.55 token/sでした。
205
 
 
120
 
121
  ### LM Studio / llama.cpp
122
 
123
+ In LM Studio, search for `ToTo-40417/EXLLM`, download `EXLLM-0.005B-LMStudio-F16.gguf`, and start a chat. The required single-turn chat template is embedded in the GGUF. Set the loaded context length to **512 tokens** and use greedy decoding or a low temperature; the model is intentionally tiny and is intended for short Japanese prompts. Training sequences were limited to 128 tokens; the larger runtime window reserves space for LM Studio's chat wrapper and short history.
124
 
125
  With llama.cpp:
126
 
127
  ```bash
128
  llama-cli -m EXLLM-0.005B-LMStudio-F16.gguf \
129
+ -p "RAMとは何ですか?" -n 48 --temp 0 --ctx-size 512 --single-turn
130
  ```
131
 
132
+ The GGUF companion uses a Llama-compatible decoder, a 1,024-piece SentencePiece tokenizer with byte fallback, 6 layers, hidden size 288, intermediate size 608, 9 attention/KV heads, 128-token training sequences, and a 512-token runtime context. It was trained separately for five epochs (3,325 optimizer steps) on 85,005 project-data records; best validation loss was 0.0390303. Exact metadata is in [`lmstudio/training_manifest.json`](lmstudio/training_manifest.json).
133
 
134
  ### Original EXLLM reference runtime
135
 
 
166
  | `weights/EXLLM-v1.1-5m-release3.pt` | 21,526,737 bytes (20.53 MiB; 0.021526737 GB) | `48ad00c6640689595a253fc949cad83f3000e14cccdf69275fcf245e85be3ceb` | fp32 PyTorch checkpoint |
167
  | `weights/EXLLM-v1.1-5m-int8.bin` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `3b2bf2d9f103cba714bc98976709c5ad34860c1ff55fed2fce91fc95afb5b34a` | EXLLM8 interchange artifact |
168
  | `weights/model.q12` | 5,443,105 bytes (5.19 MiB; 0.005443105 GB) | `d64037fde791e5c0e48101bc1a8ab366a3287f1f36b4464879e36495ea7e5a53` | EX-word EXQ12 artifact |
169
+ | `EXLLM-0.005B-LMStudio-F16.gguf` | 10,915,168 bytes (10.41 MiB; 0.010915168 GB) | `0e3c519d4a044fac45be84dc607eb8a43de4a7c60083983bf246c5b9c8a43472` | Separately trained Llama-compatible F16 companion for LM Studio / llama.cpp |
170
 
171
  EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
172
 
 
181
  ## Limitations
182
 
183
  - The model is not a general-purpose assistant.
184
+ - The original embedded checkpoint and the GGUF training sequences use 0.000128M tokens. The GGUF advertises a 0.000512M-token runtime window for LM Studio framing and short history; quality beyond the trained 128-token range is not guaranteed.
185
  - Knowledge coverage is restricted to the project-generated training scope.
186
  - Unseen concepts, long instructions, translation, and multi-step reasoning are unreliable.
187
  - No RLHF, DPO, tool use, retrieval, web access, or current-information source is included.
 
199
 
200
  EXLLM-0.005B-Instructは、低メモリ組込み機器でのローカル推論を目的とした、**0.005377824Bパラメータ**の日本語instruction modelです。主な実装対象はCASIO EX-word XD-B4800です。同じプロジェクトデータから別途学習した、LM Studio / llama.cpp向けの**0.005441184Bパラメータ**GGUF companionも収録しています。
201
 
202
+ 電子辞書用checkpointは独自のdecoder-only Transformerとtokenizerを使用するため、GGUFに直接変換したもの���はありません。LM Studioでは`ToTo-40417/EXLLM`を検索し、`EXLLM-0.005B-LMStudio-F16.gguf`を選択してください。チャット用templateはGGUFに内蔵済みで、LM Studioのロード時contextは512 tokensに設定します。学習時の系列長は128 tokensであり、追加領域はチャット制御情報と短い履歴のための余白です。このGGUFはプロジェクトデータで別途学習したPC用companionであり、電子辞書版と同一重みではありません。
203
 
204
  RTX 3060でのfp32参照実行は、40回のwarm-up後に120回測定し、TTFT中央値0.003990秒、生成速度239.082 token/sでした。EX-word整数runtimeでは、0.024184 GHzのSH-4A-class CPU上でTTFT中央値23.711秒、生成速度0.52–0.55 token/sでした。
205
 
lmstudio/training_manifest.json CHANGED
@@ -9,7 +9,8 @@
9
  "intermediate_size": 608,
10
  "attention_heads": 9,
11
  "kv_heads": 9,
12
- "context_length": 128,
 
13
  "vocabulary_size": 1024
14
  },
15
  "training_records_including_overlaps": 85005,
 
9
  "intermediate_size": 608,
10
  "attention_heads": 9,
11
  "kv_heads": 9,
12
+ "training_sequence_length": 128,
13
+ "runtime_context_length": 512,
14
  "vocabulary_size": 1024
15
  },
16
  "training_records_including_overlaps": 85005,