Instructions to use logic65/whittle-dev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use logic65/whittle-dev with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: llama cli -hf logic65/whittle-dev:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: llama cli -hf logic65/whittle-dev:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf logic65/whittle-dev:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf logic65/whittle-dev:Q8_0
Use Docker
docker model run hf.co/logic65/whittle-dev:Q8_0
- LM Studio
- Jan
- Ollama
How to use logic65/whittle-dev with Ollama:
ollama run hf.co/logic65/whittle-dev:Q8_0
- Unsloth Desktop
- Pi
How to use logic65/whittle-dev with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "logic65/whittle-dev:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use logic65/whittle-dev with Docker Model Runner:
docker model run hf.co/logic65/whittle-dev:Q8_0
- Lemonade
How to use logic65/whittle-dev with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull logic65/whittle-dev:Q8_0
Run and chat with the model
lemonade run user.whittle-dev-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use logic65/whittle-dev with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default logic65/whittle-dev:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use logic65/whittle-dev with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "logic65/whittle-dev:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
lw3 s3711: MATH-60, parity, CE curve, logs
Browse files- lw3/eval/convert_lw3.log +0 -0
- lw3/eval/export_lw3s3711.log +6 -0
- lw3/eval/lw3_ce_curve.txt +99 -0
- lw3/eval/lw_eval_lw3s3711_held.json +38 -0
- lw3/eval/lw_eval_lw3s3711_train.json +38 -0
- lw3/eval/lw_lw3s3711_held.log +13 -0
- lw3/eval/lw_lw3s3711_train.log +13 -0
- lw3/eval/math60_replies.jsonl +0 -0
- lw3/eval/math60_score.txt +1 -0
lw3/eval/convert_lw3.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
lw3/eval/export_lw3s3711.log
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
merged LoRA/shared into 330 tensors
|
| 2 |
+
carried 612 tensors from /content/work/train/base-sigmoid, folded 81 norms into HC
|
| 3 |
+
/content/work/kit/code/wt/export_next36.py:93: UserWarning: The given NumPy array is not writable, and PyTorch does not support non-writable tensors. This means writing to this tensor will result in undefined behavior. You may want to copy the array to protect its data or make it writable before converting it to a tensor. This type of warning will be suppressed for the rest of this program. (Triggered internally at /pytorch/torch/csrc/utils/tensor_numpy.cpp:213.)
|
| 4 |
+
t = torch.from_numpy(np.ascontiguousarray(a))
|
| 5 |
+
PLE: 8 heads x ~4,880,000 rows x 256 = 9.99B params in 5 shards, width 2048 into hidden 2048, injected at layer 2
|
| 6 |
+
EXPORT_DONE /dev/shm/export_lw3 tensors=979 shards=14 gate=sigmoid ple_layer=2
|
lw3/eval/lw3_ce_curve.txt
ADDED
|
@@ -0,0 +1,99 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[18:37:29 + 4.5m] PLE gain @ 0: SFT +0.0047 | OOD +0.3397 (CE with memory OFF minus ON on 8+8 held rows; positive = the table helps)
|
| 2 |
+
[18:37:35 + 4.6m] layer-wise row 164: 1022 positions, mean KL 0.3693 = S35<-T59 0.595 | S39<-T63 0.144
|
| 3 |
+
[18:37:41 + 4.7m] layer-wise row 1852: 1022 positions, mean KL 0.2489 = S35<-T59 0.385 | S39<-T63 0.113
|
| 4 |
+
[18:48:16 + 15.3m] layer-wise row 902: 1022 positions, mean KL 0.2366 = S35<-T59 0.383 | S39<-T63 0.090
|
| 5 |
+
[18:48:36 + 15.7m] held-out CE @ 200: SFT 1.2133 (start 1.2138) | OOD 2.1311 (start 2.1317)
|
| 6 |
+
[18:48:36 + 15.7m] PoSE @ 200: 238 rows offset | highest position index 130,411 of 131,072
|
| 7 |
+
[18:49:21 + 16.4m] PLE gain @ 200: SFT +0.7674 | OOD +1.0803 | IN-corpus +4.0056 | extra +3.4148 (CE 1.2524) | extra2 +0.5580 (CE 0.7409) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 8 |
+
[18:49:21 + 16.4m] injection @ 200: 132 rows | PLE/residual ratio mean 0.2991 min 0.1784 max 0.4051 | mean use term 0.0000
|
| 9 |
+
[18:49:21 + 16.4m] dependence @ 200: 38 rows since last eval | mean gap CE_off-CE_on +1.9240 | frac meeting margin 0.97 | mean term 0.0020
|
| 10 |
+
[19:00:40 + 27.7m] held-out CE @ 400: SFT 1.2153 (start 1.2138) | OOD 2.1341 (start 2.1317)
|
| 11 |
+
[19:00:40 + 27.7m] PoSE @ 400: 478 rows offset | highest position index 130,674 of 131,072
|
| 12 |
+
[19:01:26 + 28.5m] PLE gain @ 400: SFT +0.9952 | OOD +1.6398 | IN-corpus +5.6298 | extra +4.9067 (CE 1.2504) | extra2 +0.9471 (CE 0.7660) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 13 |
+
[19:01:26 + 28.5m] injection @ 400: 158 rows | PLE/residual ratio mean 0.2984 min 0.2238 max 0.4108 | mean use term 0.0000
|
| 14 |
+
[19:01:26 + 28.5m] dependence @ 400: 40 rows since last eval | mean gap CE_off-CE_on +2.0561 | frac meeting margin 0.97 | mean term 0.0018
|
| 15 |
+
[19:12:32 + 39.6m] held-out CE @ 600: SFT 1.2158 (start 1.2138) | OOD 2.1342 (start 2.1317)
|
| 16 |
+
[19:12:32 + 39.6m] PoSE @ 600: 711 rows offset | highest position index 130,674 of 131,072
|
| 17 |
+
[19:13:22 + 40.4m] PLE gain @ 600: SFT +0.6825 | OOD +1.3738 | IN-corpus +5.0787 | extra +4.2743 (CE 1.2534) | extra2 +0.8201 (CE 0.7361) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 18 |
+
[19:13:22 + 40.4m] injection @ 600: 145 rows | PLE/residual ratio mean 0.2931 min 0.2047 max 0.4669 | mean use term 0.0000
|
| 19 |
+
[19:13:22 + 40.4m] dependence @ 600: 33 rows since last eval | mean gap CE_off-CE_on +2.3760 | frac meeting margin 1.00 | mean term 0.0000
|
| 20 |
+
[19:25:11 + 52.2m] held-out CE @ 800: SFT 1.2125 (start 1.2138) | OOD 2.1334 (start 2.1317)
|
| 21 |
+
[19:25:11 + 52.2m] PoSE @ 800: 952 rows offset | highest position index 130,955 of 131,072
|
| 22 |
+
[19:25:56 + 53.0m] PLE gain @ 800: SFT +1.1591 | OOD +1.8637 | IN-corpus +6.4735 | extra +5.8430 (CE 1.2522) | extra2 +1.4713 (CE 0.7465) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 23 |
+
[19:25:56 + 53.0m] injection @ 800: 144 rows | PLE/residual ratio mean 0.2944 min 0.1995 max 0.4084 | mean use term 0.0000
|
| 24 |
+
[19:25:56 + 53.0m] dependence @ 800: 41 rows since last eval | mean gap CE_off-CE_on +3.4760 | frac meeting margin 0.98 | mean term 0.0053
|
| 25 |
+
[19:37:31 + 64.6m] held-out CE @ 1000: SFT 1.2126 (start 1.2138) | OOD 2.1302 (start 2.1317)
|
| 26 |
+
[19:37:31 + 64.6m] PoSE @ 1000: 1,188 rows offset | highest position index 130,955 of 131,072
|
| 27 |
+
[19:38:16 + 65.3m] PLE gain @ 1000: SFT +0.5153 | OOD +1.2069 | IN-corpus +5.1250 | extra +4.0858 (CE 1.2594) | extra2 +0.8563 (CE 0.7565) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 28 |
+
[19:38:16 + 65.3m] injection @ 1000: 142 rows | PLE/residual ratio mean 0.2908 min 0.2249 max 0.3984 | mean use term 0.0000
|
| 29 |
+
[19:38:16 + 65.3m] dependence @ 1000: 36 rows since last eval | mean gap CE_off-CE_on +2.6581 | frac meeting margin 1.00 | mean term 0.0000
|
| 30 |
+
[19:48:28 + 75.5m] layer-wise row 1486: 1022 positions, mean KL 0.2221 = S35<-T59 0.397 | S39<-T63 0.047
|
| 31 |
+
[19:48:48 + 75.9m] held-out CE @ 1200: SFT 1.2168 (start 1.2138) | OOD 2.1322 (start 2.1317)
|
| 32 |
+
[19:48:48 + 75.9m] PoSE @ 1200: 1,411 rows offset | highest position index 130,955 of 131,072
|
| 33 |
+
[19:49:33 + 76.6m] PLE gain @ 1200: SFT +1.0990 | OOD +1.4824 | IN-corpus +5.1794 | extra +4.1689 (CE 1.2579) | extra2 +0.8264 (CE 0.7565) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 34 |
+
[19:49:33 + 76.6m] injection @ 1200: 152 rows | PLE/residual ratio mean 0.2972 min 0.1990 max 0.5407 | mean use term 0.0000
|
| 35 |
+
[19:49:33 + 76.6m] dependence @ 1200: 23 rows since last eval | mean gap CE_off-CE_on +2.9236 | frac meeting margin 0.96 | mean term 0.0085
|
| 36 |
+
[20:01:27 + 88.5m] held-out CE @ 1400: SFT 1.2199 (start 1.2138) | OOD 2.1365 (start 2.1317)
|
| 37 |
+
[20:01:27 + 88.5m] PoSE @ 1400: 1,655 rows offset | highest position index 130,955 of 131,072
|
| 38 |
+
[20:02:12 + 89.3m] PLE gain @ 1400: SFT +0.3702 | OOD +0.7709 | IN-corpus +4.0843 | extra +2.9723 (CE 1.2621) | extra2 +0.5127 (CE 0.7171) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 39 |
+
[20:02:12 + 89.3m] injection @ 1400: 154 rows | PLE/residual ratio mean 0.2993 min 0.1827 max 0.4169 | mean use term 0.0000
|
| 40 |
+
[20:02:12 + 89.3m] dependence @ 1400: 44 rows since last eval | mean gap CE_off-CE_on +2.0227 | frac meeting margin 1.00 | mean term 0.0000
|
| 41 |
+
[20:13:43 +100.8m] held-out CE @ 1600: SFT 1.2153 (start 1.2138) | OOD 2.1311 (start 2.1317)
|
| 42 |
+
[20:13:43 +100.8m] PoSE @ 1600: 1,884 rows offset | highest position index 130,955 of 131,072
|
| 43 |
+
[20:14:28 +101.5m] PLE gain @ 1600: SFT +1.6227 | OOD +1.8616 | IN-corpus +5.7615 | extra +5.1352 (CE 1.2594) | extra2 +1.3110 (CE 0.7310) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 44 |
+
[20:14:28 +101.5m] injection @ 1600: 146 rows | PLE/residual ratio mean 0.2954 min 0.1902 max 0.5098 | mean use term 0.0000
|
| 45 |
+
[20:14:28 +101.5m] dependence @ 1600: 29 rows since last eval | mean gap CE_off-CE_on +1.7480 | frac meeting margin 0.97 | mean term 0.0006
|
| 46 |
+
[20:25:36 +112.7m] held-out CE @ 1800: SFT 1.2146 (start 1.2138) | OOD 2.1343 (start 2.1317)
|
| 47 |
+
[20:25:36 +112.7m] PoSE @ 1800: 2,128 rows offset | highest position index 130,955 of 131,072
|
| 48 |
+
[20:26:21 +113.4m] PLE gain @ 1800: SFT +1.8236 | OOD +2.1390 | IN-corpus +7.0003 | extra +6.5872 (CE 1.2633) | extra2 +1.4485 (CE 0.7521) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 49 |
+
[20:26:21 +113.4m] injection @ 1800: 153 rows | PLE/residual ratio mean 0.3015 min 0.2205 max 0.5204 | mean use term 0.0000
|
| 50 |
+
[20:26:21 +113.4m] dependence @ 1800: 44 rows since last eval | mean gap CE_off-CE_on +2.3682 | frac meeting margin 0.98 | mean term 0.0001
|
| 51 |
+
[20:36:50 +123.9m] layer-wise row 366: 1022 positions, mean KL 0.4187 = S35<-T59 0.636 | S39<-T63 0.202
|
| 52 |
+
[20:37:10 +124.2m] held-out CE @ 2000: SFT 1.2206 (start 1.2138) | OOD 2.1392 (start 2.1317)
|
| 53 |
+
[20:37:10 +124.2m] PoSE @ 2000: 2,357 rows offset | highest position index 130,955 of 131,072
|
| 54 |
+
[20:37:56 +125.0m] PLE gain @ 2000: SFT +1.5347 | OOD +1.9197 | IN-corpus +7.2523 | extra +6.4835 (CE 1.2677) | extra2 +1.4159 (CE 0.7508) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 55 |
+
[20:37:56 +125.0m] injection @ 2000: 154 rows | PLE/residual ratio mean 0.3044 min 0.2177 max 0.4653 | mean use term 0.0000
|
| 56 |
+
[20:37:56 +125.0m] dependence @ 2000: 29 rows since last eval | mean gap CE_off-CE_on +2.2464 | frac meeting margin 0.97 | mean term 0.0012
|
| 57 |
+
[20:49:52 +136.9m] held-out CE @ 2200: SFT 1.2193 (start 1.2138) | OOD 2.1401 (start 2.1317)
|
| 58 |
+
[20:49:52 +136.9m] PoSE @ 2200: 2,600 rows offset | highest position index 130,955 of 131,072
|
| 59 |
+
[20:50:38 +137.7m] PLE gain @ 2200: SFT +0.6062 | OOD +1.0096 | IN-corpus +4.6089 | extra +3.2036 (CE 1.2613) | extra2 +0.5614 (CE 0.7840) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 60 |
+
[20:50:38 +137.7m] injection @ 2200: 154 rows | PLE/residual ratio mean 0.3112 min 0.1746 max 0.4870 | mean use term 0.0000
|
| 61 |
+
[20:50:38 +137.7m] dependence @ 2200: 43 rows since last eval | mean gap CE_off-CE_on +3.5534 | frac meeting margin 1.00 | mean term 0.0000
|
| 62 |
+
[21:01:51 +148.9m] held-out CE @ 2400: SFT 1.2258 (start 1.2138) | OOD 2.1420 (start 2.1317)
|
| 63 |
+
[21:01:51 +148.9m] PoSE @ 2400: 2,834 rows offset | highest position index 130,955 of 131,072
|
| 64 |
+
[21:02:36 +149.7m] PLE gain @ 2400: SFT +0.2468 | OOD +0.6322 | IN-corpus +3.7653 | extra +2.2639 (CE 1.2596) | extra2 +0.1475 (CE 0.7674) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 65 |
+
[21:02:36 +149.7m] injection @ 2400: 148 rows | PLE/residual ratio mean 0.3068 min 0.2304 max 0.4569 | mean use term 0.0000
|
| 66 |
+
[21:02:36 +149.7m] dependence @ 2400: 34 rows since last eval | mean gap CE_off-CE_on +1.8114 | frac meeting margin 1.00 | mean term 0.0000
|
| 67 |
+
[21:14:00 +161.1m] layer-wise row 725: 1022 positions, mean KL 0.3130 = S35<-T59 0.459 | S39<-T63 0.167
|
| 68 |
+
[21:14:20 +161.4m] held-out CE @ 2600: SFT 1.2227 (start 1.2138) | OOD 2.1379 (start 2.1317)
|
| 69 |
+
[21:14:20 +161.4m] PoSE @ 2600: 3,071 rows offset | highest position index 130,955 of 131,072
|
| 70 |
+
[21:15:06 +162.2m] PLE gain @ 2600: SFT +0.5261 | OOD +0.9650 | IN-corpus +4.1110 | extra +3.7646 (CE 1.2618) | extra2 +0.3936 (CE 0.7544) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 71 |
+
[21:15:06 +162.2m] injection @ 2600: 145 rows | PLE/residual ratio mean 0.3054 min 0.2282 max 0.4501 | mean use term 0.0000
|
| 72 |
+
[21:15:06 +162.2m] dependence @ 2600: 37 rows since last eval | mean gap CE_off-CE_on +2.3101 | frac meeting margin 0.97 | mean term 0.0007
|
| 73 |
+
[21:27:12 +174.2m] held-out CE @ 2800: SFT 1.2231 (start 1.2138) | OOD 2.1442 (start 2.1317)
|
| 74 |
+
[21:27:12 +174.2m] PoSE @ 2800: 3,313 rows offset | highest position index 130,955 of 131,072
|
| 75 |
+
[21:27:57 +175.0m] PLE gain @ 2800: SFT +0.7078 | OOD +1.2372 | IN-corpus +4.5593 | extra +4.3177 (CE 1.2623) | extra2 +0.7899 (CE 0.7768) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 76 |
+
[21:27:57 +175.0m] injection @ 2800: 164 rows | PLE/residual ratio mean 0.3096 min 0.2239 max 0.4460 | mean use term 0.0000
|
| 77 |
+
[21:27:57 +175.0m] dependence @ 2800: 42 rows since last eval | mean gap CE_off-CE_on +2.7516 | frac meeting margin 0.98 | mean term 0.0008
|
| 78 |
+
[21:39:36 +186.7m] held-out CE @ 3000: SFT 1.2243 (start 1.2138) | OOD 2.1445 (start 2.1317)
|
| 79 |
+
[21:39:36 +186.7m] PoSE @ 3000: 3,546 rows offset | highest position index 130,955 of 131,072
|
| 80 |
+
[21:40:21 +187.4m] PLE gain @ 3000: SFT +0.2893 | OOD +0.7889 | IN-corpus +3.5836 | extra +3.1686 (CE 1.2559) | extra2 +0.2485 (CE 0.7768) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 81 |
+
[21:40:21 +187.4m] injection @ 3000: 154 rows | PLE/residual ratio mean 0.3072 min 0.1704 max 0.4286 | mean use term 0.0000
|
| 82 |
+
[21:40:21 +187.4m] dependence @ 3000: 33 rows since last eval | mean gap CE_off-CE_on +1.9253 | frac meeting margin 1.00 | mean term 0.0000
|
| 83 |
+
[21:51:58 +199.0m] held-out CE @ 3200: SFT 1.2242 (start 1.2138) | OOD 2.1462 (start 2.1317)
|
| 84 |
+
[21:51:58 +199.0m] PoSE @ 3200: 3,784 rows offset | highest position index 130,955 of 131,072
|
| 85 |
+
[21:52:44 +199.8m] PLE gain @ 3200: SFT +0.5958 | OOD +1.0680 | IN-corpus +3.7487 | extra +3.6534 (CE 1.2651) | extra2 +0.5994 (CE 0.7865) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 86 |
+
[21:52:44 +199.8m] injection @ 3200: 143 rows | PLE/residual ratio mean 0.3111 min 0.2277 max 0.4397 | mean use term 0.0000
|
| 87 |
+
[21:52:44 +199.8m] dependence @ 3200: 38 rows since last eval | mean gap CE_off-CE_on +1.9921 | frac meeting margin 0.97 | mean term 0.0056
|
| 88 |
+
[22:03:56 +211.0m] held-out CE @ 3400: SFT 1.2246 (start 1.2138) | OOD 2.1431 (start 2.1317)
|
| 89 |
+
[22:03:56 +211.0m] PoSE @ 3400: 4,036 rows offset | highest position index 130,955 of 131,072
|
| 90 |
+
[22:04:41 +211.7m] PLE gain @ 3400: SFT +1.1368 | OOD +1.5550 | IN-corpus +5.2589 | extra +5.2941 (CE 1.2673) | extra2 +1.1339 (CE 0.7528) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 91 |
+
[22:04:41 +211.7m] injection @ 3400: 155 rows | PLE/residual ratio mean 0.3127 min 0.2434 max 0.5199 | mean use term 0.0000
|
| 92 |
+
[22:04:41 +211.7m] dependence @ 3400: 52 rows since last eval | mean gap CE_off-CE_on +1.3096 | frac meeting margin 0.98 | mean term 0.0054
|
| 93 |
+
[22:16:29 +223.5m] held-out CE @ 3600: SFT 1.2286 (start 1.2138) | OOD 2.1500 (start 2.1317)
|
| 94 |
+
[22:16:29 +223.5m] PoSE @ 3600: 4,278 rows offset | highest position index 130,955 of 131,072
|
| 95 |
+
[22:17:14 +224.3m] PLE gain @ 3600: SFT +0.6549 | OOD +1.2428 | IN-corpus +4.2679 | extra +4.1953 (CE 1.2595) | extra2 +0.5401 (CE 0.7549) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 96 |
+
[22:17:14 +224.3m] injection @ 3600: 149 rows | PLE/residual ratio mean 0.3157 min 0.2173 max 0.4824 | mean use term 0.0000
|
| 97 |
+
[22:17:14 +224.3m] dependence @ 3600: 42 rows since last eval | mean gap CE_off-CE_on +1.9727 | frac meeting margin 1.00 | mean term 0.0000
|
| 98 |
+
[22:22:57 +230.0m] TIME BUDGET: stopping at step 3712
|
| 99 |
+
[22:24:01 +231.1m] DONE best step 1000 OOD 2.1302 (start 2.1317) | checkpoints in /content/work/lw3/ckpts | table /content/work/lw3/ngram_table.npy
|
lw3/eval/lw_eval_lw3s3711_held.json
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"tag": "lw3s3711_held",
|
| 3 |
+
"pairs": [
|
| 4 |
+
[
|
| 5 |
+
35,
|
| 6 |
+
59
|
| 7 |
+
],
|
| 8 |
+
[
|
| 9 |
+
39,
|
| 10 |
+
63
|
| 11 |
+
]
|
| 12 |
+
],
|
| 13 |
+
"cache": "/content/work/probe/teacher_L59_L63_n400.npz",
|
| 14 |
+
"row0": 0,
|
| 15 |
+
"nseq": 8,
|
| 16 |
+
"seqlen": 512,
|
| 17 |
+
"topk": 32,
|
| 18 |
+
"positions": 480,
|
| 19 |
+
"model_dir": "/dev/shm/export_lw3",
|
| 20 |
+
"results": [
|
| 21 |
+
{
|
| 22 |
+
"student": 35,
|
| 23 |
+
"teacher": 59,
|
| 24 |
+
"agree": 0.4583333333333333,
|
| 25 |
+
"dprob": 0.21632331609725952,
|
| 26 |
+
"teacher_mass": 0.7221164703369141,
|
| 27 |
+
"student_mass": 0.7757255434989929
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"student": 39,
|
| 31 |
+
"teacher": 63,
|
| 32 |
+
"agree": 0.8796875,
|
| 33 |
+
"dprob": 0.07888820767402649,
|
| 34 |
+
"teacher_mass": 0.9730892777442932,
|
| 35 |
+
"student_mass": 0.9773320555686951
|
| 36 |
+
}
|
| 37 |
+
]
|
| 38 |
+
}
|
lw3/eval/lw_eval_lw3s3711_train.json
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"tag": "lw3s3711_train",
|
| 3 |
+
"pairs": [
|
| 4 |
+
[
|
| 5 |
+
35,
|
| 6 |
+
59
|
| 7 |
+
],
|
| 8 |
+
[
|
| 9 |
+
39,
|
| 10 |
+
63
|
| 11 |
+
]
|
| 12 |
+
],
|
| 13 |
+
"cache": "/content/work/probe/teacher_L59_L63.npz",
|
| 14 |
+
"row0": 0,
|
| 15 |
+
"nseq": 8,
|
| 16 |
+
"seqlen": 512,
|
| 17 |
+
"topk": 32,
|
| 18 |
+
"positions": 480,
|
| 19 |
+
"model_dir": "/dev/shm/export_lw3",
|
| 20 |
+
"results": [
|
| 21 |
+
{
|
| 22 |
+
"student": 35,
|
| 23 |
+
"teacher": 59,
|
| 24 |
+
"agree": 0.47421875,
|
| 25 |
+
"dprob": 0.22691109776496887,
|
| 26 |
+
"teacher_mass": 0.7626522183418274,
|
| 27 |
+
"student_mass": 0.7950366735458374
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"student": 39,
|
| 31 |
+
"teacher": 63,
|
| 32 |
+
"agree": 0.8401041666666667,
|
| 33 |
+
"dprob": 0.09364508837461472,
|
| 34 |
+
"teacher_mass": 0.9724294543266296,
|
| 35 |
+
"student_mass": 0.976559042930603
|
| 36 |
+
}
|
| 37 |
+
]
|
| 38 |
+
}
|
lw3/eval/lw_lw3s3711_held.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[ 0s] batch (8, 512) from teacher_L59_L63_n400.npz rows 0..8 | pairs [(35, 59), (39, 63)]
|
| 2 |
+
[ 0s] indexer budget 262,144 -> 512 for this 512-token probe (same visible tokens)
|
| 3 |
+
[ 0s] lw3s3711_held: 40 layers, hidden 2048, hc_count 4, attention at [3, 7, 11, 15, 19, 23, 27, 31, 35, 39]
|
| 4 |
+
[ 1s] streams (8, 512, 8192) | masks + rope built
|
| 5 |
+
[ 2s] .. layer 0 done
|
| 6 |
+
[ 4s] layer 1: n-gram memory streamed 39,040,000 rows into (39040640, 256) on cuda
|
| 7 |
+
[ 8s] .. layer 8 done
|
| 8 |
+
[ 11s] .. layer 16 done
|
| 9 |
+
[ 14s] .. layer 24 done
|
| 10 |
+
[ 17s] .. layer 32 done
|
| 11 |
+
[ 21s] lw3s3711_held S35 <- T59: top-1 agree 45.8% | dprob 0.216 | teacher-mass 72.2% | student-mass 77.6% (3,840 positions)
|
| 12 |
+
[ 25s] lw3s3711_held S39 <- T63: top-1 agree 88.0% | dprob 0.079 | teacher-mass 97.3% | student-mass 97.7% (3,840 positions)
|
| 13 |
+
[ 25s] LW_EVAL DONE lw3s3711_held: S35<-T59 agree 45.8% mass 72.2% | S39<-T63 agree 88.0% mass 97.3%
|
lw3/eval/lw_lw3s3711_train.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[ 0s] batch (8, 512) from teacher_L59_L63.npz rows 0..8 | pairs [(35, 59), (39, 63)]
|
| 2 |
+
[ 0s] indexer budget 262,144 -> 512 for this 512-token probe (same visible tokens)
|
| 3 |
+
[ 0s] lw3s3711_train: 40 layers, hidden 2048, hc_count 4, attention at [3, 7, 11, 15, 19, 23, 27, 31, 35, 39]
|
| 4 |
+
[ 1s] streams (8, 512, 8192) | masks + rope built
|
| 5 |
+
[ 1s] .. layer 0 done
|
| 6 |
+
[ 3s] layer 1: n-gram memory streamed 39,040,000 rows into (39040640, 256) on cuda
|
| 7 |
+
[ 6s] .. layer 8 done
|
| 8 |
+
[ 9s] .. layer 16 done
|
| 9 |
+
[ 12s] .. layer 24 done
|
| 10 |
+
[ 15s] .. layer 32 done
|
| 11 |
+
[ 19s] lw3s3711_train S35 <- T59: top-1 agree 47.4% | dprob 0.227 | teacher-mass 76.3% | student-mass 79.5% (3,840 positions)
|
| 12 |
+
[ 23s] lw3s3711_train S39 <- T63: top-1 agree 84.0% | dprob 0.094 | teacher-mass 97.2% | student-mass 97.7% (3,840 positions)
|
| 13 |
+
[ 23s] LW_EVAL DONE lw3s3711_train: S35<-T59 agree 47.4% mass 76.3% | S39<-T63 agree 84.0% mass 97.2%
|
lw3/eval/math60_replies.jsonl
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
lw3/eval/math60_score.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
/content/work/lw3_gen/gen.jsonl: 39/60 correct (65%) | by level L2:16/20 L3:13/20 L4:10/20 | no_boxed 6 | cap_hit 6 | format {'ok': 54, 'cap_hit': 6} | mean 1483 tok, think ~376 tok (54 closed)
|