Instructions to use logic65/whittle-dev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use logic65/whittle-dev with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: llama cli -hf logic65/whittle-dev:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: llama cli -hf logic65/whittle-dev:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf logic65/whittle-dev:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf logic65/whittle-dev:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf logic65/whittle-dev:Q8_0
Use Docker
docker model run hf.co/logic65/whittle-dev:Q8_0
- LM Studio
- Jan
- Ollama
How to use logic65/whittle-dev with Ollama:
ollama run hf.co/logic65/whittle-dev:Q8_0
- Unsloth Desktop
- Pi
How to use logic65/whittle-dev with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "logic65/whittle-dev:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use logic65/whittle-dev with Docker Model Runner:
docker model run hf.co/logic65/whittle-dev:Q8_0
- Lemonade
How to use logic65/whittle-dev with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull logic65/whittle-dev:Q8_0
Run and chat with the model
lemonade run user.whittle-dev-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use logic65/whittle-dev with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default logic65/whittle-dev:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use logic65/whittle-dev with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf logic65/whittle-dev:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "logic65/whittle-dev:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
lw5 s1642: MATH-60, parity, CE curve, logs
Browse files- lw5/eval/convert_lw5.log +0 -0
- lw5/eval/export_lw5s1642.log +6 -0
- lw5/eval/lw5_ce_curve.txt +46 -0
- lw5/eval/lw_eval_lw5s1642_held.json +38 -0
- lw5/eval/lw_eval_lw5s1642_train.json +38 -0
- lw5/eval/lw_lw5s1642_held.log +13 -0
- lw5/eval/lw_lw5s1642_train.log +13 -0
- lw5/eval/math60_replies.jsonl +0 -0
- lw5/eval/math60_score.txt +1 -0
lw5/eval/convert_lw5.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
lw5/eval/export_lw5s1642.log
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
merged LoRA/shared into 330 tensors
|
| 2 |
+
carried 612 tensors from /content/work/train/base-sigmoid, folded 81 norms into HC
|
| 3 |
+
/content/work/kit/code/wt/export_next36.py:93: UserWarning: The given NumPy array is not writable, and PyTorch does not support non-writable tensors. This means writing to this tensor will result in undefined behavior. You may want to copy the array to protect its data or make it writable before converting it to a tensor. This type of warning will be suppressed for the rest of this program. (Triggered internally at /pytorch/torch/csrc/utils/tensor_numpy.cpp:213.)
|
| 4 |
+
t = torch.from_numpy(np.ascontiguousarray(a))
|
| 5 |
+
PLE: 8 heads x ~4,880,000 rows x 256 = 9.99B params in 5 shards, width 2048 into hidden 2048, injected at layer 2
|
| 6 |
+
EXPORT_DONE /dev/shm/export_lw5 tensors=979 shards=14 gate=sigmoid ple_layer=2
|
lw5/eval/lw5_ce_curve.txt
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[10:19:27 + 2.9m] PLE gain @ 0: SFT +0.0066 | OOD +0.3424 (CE with memory OFF minus ON on 8+8 held rows; positive = the table helps)
|
| 2 |
+
[10:19:41 + 3.1m] layer-wise row 800: 1022 positions, mean KL 0.1506 = S35<-T59 0.241 | S39<-T63 0.061
|
| 3 |
+
[10:30:53 + 14.3m] held-out CE @ 200: SFT 1.2139 (start 1.2144) | OOD 2.1312 (start 2.1320)
|
| 4 |
+
[10:30:53 + 14.3m] PoSE @ 200: 240 rows offset | highest position index 130,633 of 131,072
|
| 5 |
+
[10:31:38 + 15.1m] PLE gain @ 200: SFT +3.3540 | OOD +3.9652 | IN-corpus +11.2230 | extra +12.4708 (CE 1.2529) | extra2 +3.0735 (CE 0.7602) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 6 |
+
[10:31:38 + 15.1m] injection @ 200: 80 rows | PLE/residual ratio mean 0.2986 min 0.2198 max 0.3998 | mean use term 0.0000
|
| 7 |
+
[10:31:38 + 15.1m] dependence @ 200: 40 rows since last eval | mean gap CE_off-CE_on +4.2412 | frac meeting margin 0.90 | mean term 0.0361
|
| 8 |
+
[10:44:57 + 28.4m] held-out CE @ 400: SFT 1.2116 (start 1.2144) | OOD 2.1284 (start 2.1320)
|
| 9 |
+
[10:44:57 + 28.4m] PoSE @ 400: 493 rows offset | highest position index 131,025 of 131,072
|
| 10 |
+
[10:45:42 + 29.1m] PLE gain @ 400: SFT +3.0849 | OOD +4.7276 | IN-corpus +12.7540 | extra +14.2183 (CE 1.2551) | extra2 +2.2868 (CE 0.7321) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 11 |
+
[10:45:42 + 29.1m] injection @ 400: 90 rows | PLE/residual ratio mean 0.2875 min 0.2076 max 0.4420 | mean use term 0.0000
|
| 12 |
+
[10:45:42 + 29.1m] dependence @ 400: 53 rows since last eval | mean gap CE_off-CE_on +8.5142 | frac meeting margin 0.98 | mean term 0.0005
|
| 13 |
+
[10:57:41 + 41.1m] layer-wise row 1749: 1022 positions, mean KL 0.2133 = S35<-T59 0.333 | S39<-T63 0.093
|
| 14 |
+
[10:58:02 + 41.5m] held-out CE @ 600: SFT 1.2065 (start 1.2144) | OOD 2.1251 (start 2.1320)
|
| 15 |
+
[10:58:02 + 41.5m] PoSE @ 600: 738 rows offset | highest position index 131,025 of 131,072
|
| 16 |
+
[10:58:47 + 42.2m] PLE gain @ 600: SFT +2.3054 | OOD +3.5673 | IN-corpus +10.3617 | extra +11.9346 (CE 1.2545) | extra2 +1.6659 (CE 0.7508) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 17 |
+
[10:58:47 + 42.2m] injection @ 600: 76 rows | PLE/residual ratio mean 0.2984 min 0.2054 max 0.5249 | mean use term 0.0000
|
| 18 |
+
[10:58:47 + 42.2m] dependence @ 600: 45 rows since last eval | mean gap CE_off-CE_on +5.5111 | frac meeting margin 1.00 | mean term 0.0000
|
| 19 |
+
[11:10:27 + 53.9m] layer-wise row 1974: 1022 positions, mean KL 0.1070 = S35<-T59 0.190 | S39<-T63 0.024
|
| 20 |
+
[11:10:49 + 54.2m] held-out CE @ 800: SFT 1.2052 (start 1.2144) | OOD 2.1297 (start 2.1320)
|
| 21 |
+
[11:10:49 + 54.2m] PoSE @ 800: 984 rows offset | highest position index 131,025 of 131,072
|
| 22 |
+
[11:11:34 + 55.0m] PLE gain @ 800: SFT +2.2802 | OOD +3.2870 | IN-corpus +8.6690 | extra +10.4940 (CE 1.2582) | extra2 +1.2679 (CE 0.7618) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 23 |
+
[11:11:34 + 55.0m] injection @ 800: 84 rows | PLE/residual ratio mean 0.2870 min 0.1626 max 0.4513 | mean use term 0.0000
|
| 24 |
+
[11:11:34 + 55.0m] dependence @ 800: 46 rows since last eval | mean gap CE_off-CE_on +3.8568 | frac meeting margin 0.98 | mean term 0.0073
|
| 25 |
+
[11:23:05 + 66.5m] held-out CE @ 1000: SFT 1.2022 (start 1.2144) | OOD 2.1339 (start 2.1320)
|
| 26 |
+
[11:23:05 + 66.5m] PoSE @ 1000: 1,216 rows offset | highest position index 131,025 of 131,072
|
| 27 |
+
[11:23:51 + 67.3m] PLE gain @ 1000: SFT +1.8558 | OOD +3.0573 | IN-corpus +8.1644 | extra +9.9091 (CE 1.2570) | extra2 +0.9752 (CE 0.7807) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 28 |
+
[11:23:51 + 67.3m] injection @ 1000: 75 rows | PLE/residual ratio mean 0.2839 min 0.2247 max 0.4424 | mean use term 0.0000
|
| 29 |
+
[11:23:51 + 67.3m] dependence @ 1000: 32 rows since last eval | mean gap CE_off-CE_on +4.0904 | frac meeting margin 1.00 | mean term 0.0000
|
| 30 |
+
[11:36:36 + 80.0m] held-out CE @ 1200: SFT 1.2059 (start 1.2144) | OOD 2.1289 (start 2.1320)
|
| 31 |
+
[11:36:36 + 80.0m] PoSE @ 1200: 1,467 rows offset | highest position index 131,025 of 131,072
|
| 32 |
+
[11:37:21 + 80.8m] PLE gain @ 1200: SFT +1.6350 | OOD +2.5600 | IN-corpus +5.5254 | extra +7.0162 (CE 1.2495) | extra2 +0.2386 (CE 0.8021) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 33 |
+
[11:37:21 + 80.8m] injection @ 1200: 81 rows | PLE/residual ratio mean 0.2733 min 0.2059 max 0.3497 | mean use term 0.0000
|
| 34 |
+
[11:37:21 + 80.8m] dependence @ 1200: 51 rows since last eval | mean gap CE_off-CE_on +4.4319 | frac meeting margin 0.98 | mean term 0.0024
|
| 35 |
+
[11:49:42 + 93.1m] held-out CE @ 1400: SFT 1.2039 (start 1.2144) | OOD 2.1309 (start 2.1320)
|
| 36 |
+
[11:49:42 + 93.1m] PoSE @ 1400: 1,713 rows offset | highest position index 131,025 of 131,072
|
| 37 |
+
[11:50:27 + 93.9m] PLE gain @ 1400: SFT +2.9396 | OOD +3.8990 | IN-corpus +8.7007 | extra +11.1144 (CE 1.2478) | extra2 +0.7373 (CE 0.8054) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 38 |
+
[11:50:27 + 93.9m] injection @ 1400: 86 rows | PLE/residual ratio mean 0.2718 min 0.2241 max 0.4252 | mean use term 0.0000
|
| 39 |
+
[11:50:27 + 93.9m] dependence @ 1400: 46 rows since last eval | mean gap CE_off-CE_on +3.7728 | frac meeting margin 0.98 | mean term 0.0076
|
| 40 |
+
[12:02:52 +106.3m] held-out CE @ 1600: SFT 1.2055 (start 1.2144) | OOD 2.1299 (start 2.1320)
|
| 41 |
+
[12:02:52 +106.3m] PoSE @ 1600: 1,949 rows offset | highest position index 131,025 of 131,072
|
| 42 |
+
[12:03:37 +107.1m] PLE gain @ 1600: SFT +2.9867 | OOD +3.9038 | IN-corpus +8.2404 | extra +10.7882 (CE 1.2511) | extra2 +1.4362 (CE 0.7955) (CE off minus on; IN = training text the table may have seen) | live rows -1
|
| 43 |
+
[12:03:37 +107.1m] injection @ 1600: 80 rows | PLE/residual ratio mean 0.2673 min 0.1988 max 0.3578 | mean use term 0.0000
|
| 44 |
+
[12:03:37 +107.1m] dependence @ 1600: 36 rows since last eval | mean gap CE_off-CE_on +5.3818 | frac meeting margin 0.94 | mean term 0.0393
|
| 45 |
+
[12:06:35 +110.0m] TIME BUDGET: stopping at step 1643
|
| 46 |
+
[12:07:36 +111.0m] DONE best step 600 OOD 2.1251 (start 2.1320) | checkpoints in /content/work/lw5/ckpts | table /content/work/lw5/ngram_table.npy
|
lw5/eval/lw_eval_lw5s1642_held.json
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"tag": "lw5s1642_held",
|
| 3 |
+
"pairs": [
|
| 4 |
+
[
|
| 5 |
+
35,
|
| 6 |
+
59
|
| 7 |
+
],
|
| 8 |
+
[
|
| 9 |
+
39,
|
| 10 |
+
63
|
| 11 |
+
]
|
| 12 |
+
],
|
| 13 |
+
"cache": "/content/work/probe/teacher_L59_L63_n400.npz",
|
| 14 |
+
"row0": 0,
|
| 15 |
+
"nseq": 8,
|
| 16 |
+
"seqlen": 512,
|
| 17 |
+
"topk": 32,
|
| 18 |
+
"positions": 480,
|
| 19 |
+
"model_dir": "/dev/shm/export_lw5",
|
| 20 |
+
"results": [
|
| 21 |
+
{
|
| 22 |
+
"student": 35,
|
| 23 |
+
"teacher": 59,
|
| 24 |
+
"agree": 0.42994791666666665,
|
| 25 |
+
"dprob": 0.2268601357936859,
|
| 26 |
+
"teacher_mass": 0.710132360458374,
|
| 27 |
+
"student_mass": 0.7634059190750122
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"student": 39,
|
| 31 |
+
"teacher": 63,
|
| 32 |
+
"agree": 0.8786458333333333,
|
| 33 |
+
"dprob": 0.08035383373498917,
|
| 34 |
+
"teacher_mass": 0.9730727076530457,
|
| 35 |
+
"student_mass": 0.9764617085456848
|
| 36 |
+
}
|
| 37 |
+
]
|
| 38 |
+
}
|
lw5/eval/lw_eval_lw5s1642_train.json
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"tag": "lw5s1642_train",
|
| 3 |
+
"pairs": [
|
| 4 |
+
[
|
| 5 |
+
35,
|
| 6 |
+
59
|
| 7 |
+
],
|
| 8 |
+
[
|
| 9 |
+
39,
|
| 10 |
+
63
|
| 11 |
+
]
|
| 12 |
+
],
|
| 13 |
+
"cache": "/content/work/probe/teacher_L59_L63.npz",
|
| 14 |
+
"row0": 0,
|
| 15 |
+
"nseq": 8,
|
| 16 |
+
"seqlen": 512,
|
| 17 |
+
"topk": 32,
|
| 18 |
+
"positions": 480,
|
| 19 |
+
"model_dir": "/dev/shm/export_lw5",
|
| 20 |
+
"results": [
|
| 21 |
+
{
|
| 22 |
+
"student": 35,
|
| 23 |
+
"teacher": 59,
|
| 24 |
+
"agree": 0.45286458333333335,
|
| 25 |
+
"dprob": 0.24202226102352142,
|
| 26 |
+
"teacher_mass": 0.7444586157798767,
|
| 27 |
+
"student_mass": 0.7800483107566833
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"student": 39,
|
| 31 |
+
"teacher": 63,
|
| 32 |
+
"agree": 0.8354166666666667,
|
| 33 |
+
"dprob": 0.09558665007352829,
|
| 34 |
+
"teacher_mass": 0.971971333026886,
|
| 35 |
+
"student_mass": 0.9753647446632385
|
| 36 |
+
}
|
| 37 |
+
]
|
| 38 |
+
}
|
lw5/eval/lw_lw5s1642_held.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[ 0s] batch (8, 512) from teacher_L59_L63_n400.npz rows 0..8 | pairs [(35, 59), (39, 63)]
|
| 2 |
+
[ 0s] indexer budget 262,144 -> 512 for this 512-token probe (same visible tokens)
|
| 3 |
+
[ 0s] lw5s1642_held: 40 layers, hidden 2048, hc_count 4, attention at [3, 7, 11, 15, 19, 23, 27, 31, 35, 39]
|
| 4 |
+
[ 1s] streams (8, 512, 8192) | masks + rope built
|
| 5 |
+
[ 2s] .. layer 0 done
|
| 6 |
+
[ 4s] layer 1: n-gram memory streamed 39,040,000 rows into (39040640, 256) on cuda
|
| 7 |
+
[ 7s] .. layer 8 done
|
| 8 |
+
[ 11s] .. layer 16 done
|
| 9 |
+
[ 14s] .. layer 24 done
|
| 10 |
+
[ 17s] .. layer 32 done
|
| 11 |
+
[ 21s] lw5s1642_held S35 <- T59: top-1 agree 43.0% | dprob 0.227 | teacher-mass 71.0% | student-mass 76.3% (3,840 positions)
|
| 12 |
+
[ 25s] lw5s1642_held S39 <- T63: top-1 agree 87.9% | dprob 0.080 | teacher-mass 97.3% | student-mass 97.6% (3,840 positions)
|
| 13 |
+
[ 25s] LW_EVAL DONE lw5s1642_held: S35<-T59 agree 43.0% mass 71.0% | S39<-T63 agree 87.9% mass 97.3%
|
lw5/eval/lw_lw5s1642_train.log
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[ 0s] batch (8, 512) from teacher_L59_L63.npz rows 0..8 | pairs [(35, 59), (39, 63)]
|
| 2 |
+
[ 0s] indexer budget 262,144 -> 512 for this 512-token probe (same visible tokens)
|
| 3 |
+
[ 0s] lw5s1642_train: 40 layers, hidden 2048, hc_count 4, attention at [3, 7, 11, 15, 19, 23, 27, 31, 35, 39]
|
| 4 |
+
[ 1s] streams (8, 512, 8192) | masks + rope built
|
| 5 |
+
[ 1s] .. layer 0 done
|
| 6 |
+
[ 3s] layer 1: n-gram memory streamed 39,040,000 rows into (39040640, 256) on cuda
|
| 7 |
+
[ 6s] .. layer 8 done
|
| 8 |
+
[ 9s] .. layer 16 done
|
| 9 |
+
[ 12s] .. layer 24 done
|
| 10 |
+
[ 14s] .. layer 32 done
|
| 11 |
+
[ 19s] lw5s1642_train S35 <- T59: top-1 agree 45.3% | dprob 0.242 | teacher-mass 74.4% | student-mass 78.0% (3,840 positions)
|
| 12 |
+
[ 23s] lw5s1642_train S39 <- T63: top-1 agree 83.5% | dprob 0.096 | teacher-mass 97.2% | student-mass 97.5% (3,840 positions)
|
| 13 |
+
[ 23s] LW_EVAL DONE lw5s1642_train: S35<-T59 agree 45.3% mass 74.4% | S39<-T63 agree 83.5% mass 97.2%
|
lw5/eval/math60_replies.jsonl
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
lw5/eval/math60_score.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
/content/work/lw5_gen/gen.jsonl: 43/60 correct (72%) | by level L2:19/20 L3:14/20 L4:10/20 | no_boxed 6 | cap_hit 6 | format {'ok': 54, 'cap_hit': 6} | mean 1485 tok, think ~373 tok (55 closed)
|