Zaras-M commited on
Commit
a7e0283
·
verified ·
1 Parent(s): 4b5abeb

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +88 -16
README.md CHANGED
@@ -28,11 +28,11 @@ tags:
28
  - prediction-markets
29
  ---
30
 
31
- # Olas prediction-market forecaster (14B LoRA)
32
 
33
- A LoRA adapter for DeepSeek-R1-Distill-Qwen-14B, fine-tuned to forecast the outcome of binary prediction-market questions. Given the full question prompt with background and retrieved evidence, the model reasons in a `<think>` block and then returns a probability that the market resolves YES.
34
 
35
- On a held-out test set of 14,344 markets it matches the balanced accuracy of the GPT-4.1-based tool it was built to replace, with lower calibration error. See [Evaluation](#evaluation).
36
 
37
  ## Model details
38
 
@@ -41,25 +41,31 @@ On a held-out test set of 14,344 markets it matches the balanced accuracy of the
41
  | Base model | [unsloth/DeepSeek-R1-Distill-Qwen-14B](https://huggingface.co/unsloth/DeepSeek-R1-Distill-Qwen-14B) (mirror of [deepseek-ai/DeepSeek-R1-Distill-Qwen-14B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B)) |
42
  | Adapter type | LoRA, rank 16, alpha 16, dropout 0 |
43
  | Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
44
- | Trainable parameters | ~69M (adapter file is 275 MB in bf16) |
45
- | Training precision | Base loaded in 4-bit (bitsandbytes) via Unsloth, adapter in bf16 |
 
 
46
  | Context | 8192 tokens total, prompts capped at 2048 tokens |
47
  | Language | English |
48
 
49
- The adapter file was trained against the 4-bit base, so `adapter_config.json` lists `unsloth/deepseek-r1-distill-qwen-14b-unsloth-bnb-4bit` as `base_model_name_or_path`. It applies cleanly to the bf16 base as well, and that is how it was evaluated. Load the base explicitly rather than relying on the config value.
50
 
51
  ## How it was trained
52
 
53
  Training data: historical prediction-market questions served by the Olas mech, each stored with the exact prompt the production tool sent to its LLM and the market's final resolution. Two stages:
54
 
55
- 1. **SFT warm start.** Step-by-step reasoning traces were generated for past markets and filtered for leakage. The base model was fine-tuned on 5,000 of them for one epoch (lr 2e-4) so it learns the answer format and reasoning style before reinforcement learning.
56
- 2. **GRPO.** Starting from the warm-started adapter, one epoch of GRPO over 20% of the training split (lr 2e-5, 4 sampled completions per prompt). The reward is the Brier score of the model's `p_yes` against the resolved outcome, relative to the market base rate, with a penalty for malformed answers. The checkpoint with the best balanced accuracy on the eval slice (step 100) was kept.
57
 
58
  Frameworks: Unsloth, TRL, PEFT, on a single A100 80GB.
59
 
60
  ## Evaluation
61
 
62
- Held-out test split of 14,344 markets the model never saw, prompts under 1,900 tokens, greedy decoding (temperature 0). The accuracy threshold is the base rate (0.245).
 
 
 
 
63
 
64
  | System | Brier (lower is better) | ECE (lower is better) | Balanced accuracy (higher is better) |
65
  |---|---|---|---|
@@ -67,13 +73,61 @@ Held-out test split of 14,344 markets the model never saw, prompts under 1,900 t
67
  | Production tool (GPT-4.1-based) | 0.308 | 0.284 | 0.612 |
68
  | **This model** | 0.252 | **0.204** | **0.614** |
69
 
70
- - Balanced accuracy is the discrimination metric: 0.5 is a coin flip.
71
- - Brier is the squared error of the probability; ECE is the calibration error.
72
- - Malformed outputs: 6 out of 8,615 in an earlier run of the same recipe, versus 6,700 for the base model, which often ignores the answer format.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
 
74
  ## Input and output format
75
 
76
- Send one user message containing the full forecaster prompt, with no system message. The prompt is the production one: the question, a `<background>` block, an `<additional_information>` block with retrieved evidence, and the answer instructions. Anything else changes the input distribution the model was trained on.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
77
 
78
  The model replies with a `<think>...</think>` block followed by a JSON answer:
79
 
@@ -81,7 +135,11 @@ The model replies with a `<think>...</think>` block followed by a JSON answer:
81
  {"p_yes": 0.27, "p_no": 0.73, "confidence": 0.8, "info_utility": 0.6}
82
  ```
83
 
84
- Take `p_yes` from the text after the last `</think>`. Use `max_tokens` of at least 1024 so the reasoning is not cut off.
 
 
 
 
85
 
86
  ## How to run
87
 
@@ -111,7 +169,7 @@ resp = client.chat.completions.create(
111
  print(resp.choices[0].message.content)
112
  ```
113
 
114
- Use `bfloat16`, not `float16`. Half precision degrades quality on this model.
115
 
116
  ### Transformers + PEFT
117
 
@@ -142,7 +200,21 @@ To ship a single checkpoint, call `model.merge_and_unload()` and save, then serv
142
  | bf16 base + adapter | ~28 GB weights, 40-50 GB with KV cache at 8k context | A100 40/80GB, H100, L40S 48GB |
143
  | 4-bit base (`--quantization bitsandbytes --load-format bitsandbytes`) | ~14 GB | RTX 4090, A10, L4 |
144
 
145
- 4-bit serving costs roughly 10-20% latency and slightly lower quality. On an A100 80GB with vLLM continuous batching, one instance handles dozens of concurrent requests.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
146
 
147
  ## License
148
 
 
28
  - prediction-markets
29
  ---
30
 
31
+ # Olas Predict R1 14B — forecasting LoRA adapter
32
 
33
+ Developed by [Valory](https://valory.xyz), a core contributor to [Olas](https://olas.network), this LoRA adapter specialises DeepSeek-R1-Distill-Qwen-14B in forecasting binary prediction-market outcomes. It requires the base model and a prompt containing the market question, background information and retrieved evidence. It generates reasoning followed by a JSON answer containing the estimated probability of a YES outcome. The model does not retrieve evidence itself.
34
 
35
+ In a separate evaluation of 2,628 markets that closed after the training period, this model reduced Brier score by 20.3% relative to the base model, from 0.2249 to 0.1792, and increased accuracy from 71.4% to 75.8%. GPT-4.1, given the same evidence, achieved a Brier score of 0.1914 and accuracy of 75.4%. These results support comparable forecasting performance under the evaluation conditions. See [Evaluation on later markets](#evaluation-on-later-markets).
36
 
37
  ## Model details
38
 
 
41
  | Base model | [unsloth/DeepSeek-R1-Distill-Qwen-14B](https://huggingface.co/unsloth/DeepSeek-R1-Distill-Qwen-14B) (mirror of [deepseek-ai/DeepSeek-R1-Distill-Qwen-14B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B)) |
42
  | Adapter type | LoRA, rank 16, alpha 16, dropout 0 |
43
  | Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
44
+ | Trainable parameters | 68,812,800 (approximately 69 million) |
45
+ | Adapter download size | 275 MB (`adapter_model.safetensors`) |
46
+ | Stored adapter precision | float32 (all 672 tensors) |
47
+ | Training precision | Base loaded in 4-bit (bitsandbytes) via Unsloth, adapter trained in bf16 with float32 master weights |
48
  | Context | 8192 tokens total, prompts capped at 2048 tokens |
49
  | Language | English |
50
 
51
+ The adapter was trained against the 4-bit base, so `adapter_config.json` lists `unsloth/deepseek-r1-distill-qwen-14b-unsloth-bnb-4bit` as `base_model_name_or_path`. It applies cleanly to the bf16 base as well, and that is how it was evaluated. Load the base explicitly rather than relying on the config value.
52
 
53
  ## How it was trained
54
 
55
  Training data: historical prediction-market questions served by the Olas mech, each stored with the exact prompt the production tool sent to its LLM and the market's final resolution. Two stages:
56
 
57
+ 1. **Supervised fine-tuning (SFT).** An initial reinforcement-learning checkpoint (GRPO from the base model, same recipe as stage 2) generated reasoning traces for a separate set of markets carved out of the training data by market id, sampling four candidates per prompt at temperature 0.7. We retained traces with a readable probability that outperformed a constant base-rate prediction (Brier skill above zero against a base rate of 0.25). Outcomes were used to select traces, but were not supplied to the model generating them. All passing traces were kept rather than only the most confident one per market, to avoid teaching over-confidence. We then trained an adapter on 5,000 retained examples for one epoch, using a learning rate of 2e-4, before the final reinforcement-learning stage.
58
+ 2. **Reinforcement learning (GRPO).** Starting from the SFT adapter, we trained for one epoch on 20% of the training split, using a learning rate of 2e-5 and four sampled completions per prompt. The reward is the Brier skill score against a constant base-rate prediction: `(base_rate − y)² − (p_yes − y)²`, where `y` is the resolved outcome and `base_rate` is the YES rate of the training split. Answers without a parseable probability receive a fixed penalty of −2.0, below the worst achievable valid reward. A secondary format reward (weight 0.1) rewards a closed `<think>` block with a parseable answer.
59
 
60
  Frameworks: Unsloth, TRL, PEFT, on a single A100 80GB.
61
 
62
  ## Evaluation
63
 
64
+ ### Internal held-out evaluation
65
+
66
+ The development dataset contains 214,529 prediction examples across 5,116 resolved markets, split by market to prevent the same market appearing in training and testing. The test split contains 21,583 examples across 513 markets. Excluding 7,239 examples that exceed the 1,900-token prompt limit leaves 14,344 evaluated examples across 467 markets. Every example carries equal weight.
67
+
68
+ Decoding is greedy (temperature 0) with a 1,024-token output budget. The classification threshold for accuracy is the YES rate of the evaluated examples (0.244). Its balanced accuracy and calibration results are separate from the later-market evaluation reported below.
69
 
70
  | System | Brier (lower is better) | ECE (lower is better) | Balanced accuracy (higher is better) |
71
  |---|---|---|---|
 
73
  | Production tool (GPT-4.1-based) | 0.308 | 0.284 | 0.612 |
74
  | **This model** | 0.252 | **0.204** | **0.614** |
75
 
76
+ - Balanced accuracy averages the fraction of YES outcomes and the fraction of NO outcomes correctly classified at the stated threshold. A score of 0.5 represents chance-level balanced accuracy; higher is better.
77
+ - Brier score measures the mean squared difference between predicted YES probabilities and actual outcomes. Lower is better. In this internal test, the adapter improves Brier score relative to the production tool, but not relative to the base model. A constant prediction at the training-split YES rate (0.216) would score 0.185 on this test set, below all three systems, because the base model's malformed answers and the models' over-confidence both cost more than a flat prior on this metric.
78
+ - Expected calibration error (ECE) measures discrepancies between predicted probabilities and observed outcome frequencies within probability bins. We calculated ECE using 10 equal-width bins. Lower is better.
79
+
80
+ **Formatting reliability.** On the internal test, the released adapter produced 11 malformed answers out of 14,344 examples, compared with 11,138 for the base model, which mostly ignores the answer format. An answer was considered invalid if no probability between 0 and 1 could be parsed from the text after the last `</think>` tag, in either the JSON or the XML-tag answer format. Invalid answers were scored as a prediction of 0.5 and included in all reported metrics.
81
+
82
+ ### Evaluation on later markets
83
+
84
+ A separate evaluation covered 1,247 Omen markets and 1,381 Polymarket markets that closed between July 16 and August 27, 2026, after the training data ended.
85
+
86
+ All models received the same frozen evidence collected before market resolution, with no web access or later news. The evaluation contained 4,347 evidence contexts. Scores were averaged within each market so that every market had equal weight. Answers without a usable probability received a Brier score of 1. Open models had a 1,024-token output allowance; GPT-4.1 had 4,096 tokens.
87
+
88
+ | System | Brier score (lower is better) | Accuracy (higher is better) |
89
+ |---|---|---|
90
+ | Base model | 0.2249 | 71.4% |
91
+ | GRPO without SFT warm start | 0.1947 | 72.9% |
92
+ | **SFT warm start followed by GRPO (this model)** | **0.1792** | **75.8%** |
93
+ | GPT-4.1, same evidence | 0.1914 | 75.4% |
94
+ | Market price at evidence capture | 0.1264 | 83.0% |
95
+
96
+ This model's Brier-score difference relative to the base model was −0.0457, with a 95% confidence interval of [−0.0539, −0.0375]. Relative to GPT-4.1, the difference was −0.0122 [−0.0241, −0.0005]. Confidence intervals used 10,000 bootstrap resamples, grouping markets linked to the same real-world event.
97
+
98
+ The improvement over the base model was clear on both platforms. The advantage over GPT-4.1 was smaller and varied by platform and metric. Market prices achieved better aggregate forecasting scores than every model.
99
+
100
+ These results use uncalibrated model predictions. Platt scaling improved the internal test results but did not carry over to the separate evaluation, so it was not retained.
101
 
102
  ## Input and output format
103
 
104
+ Send one user message with no system message, using the complete prompt template at https://github.com/valory-xyz/mech-predict/blob/v0.21.32/packages/valory/customs/finetuned_prediction/finetuned_prediction.py. Replace the market question, background and retrieved-evidence fields while preserving the answer instructions. The `PROMPT` variable in the examples below refers to this completed template.
105
+
106
+ Two production prompt variants appear in the training data. One wraps the evidence in a `<background>` block. The other wraps it in an `<additional_information>` block, followed by a first-stage reasoning section, and ends with the answer instructions. The skeleton of the second variant:
107
+
108
+ ```text
109
+ Here is the user's question: <question>
110
+ Here is some additional information that may be relevant to answering the question: <additional_information> ARTICLE 0, URL: ..., CONTENT: ...
111
+
112
+ ARTICLE 1, URL: ..., CONTENT: ... </additional_information>
113
+
114
+ Please carefully read the user's question and the additional information provided. Think through the problem step-by-step ...
115
+ <reasoning></reasoning>
116
+ ////
117
+ You will be evaluating the likelihood of an event based on a user's question and reasoning provided by another AI.
118
+ The user's question is: <user_input> <question> </user_input>
119
+
120
+ The reasoning from the other AI is: <first-stage reasoning>
121
+
122
+ Carefully consider the user's question and the provided reasoning. Then, think through the following:
123
+ - The probability that the event specified in the user's question will happen (p_yes)
124
+ - The probability that the event will not happen (p_no)
125
+ - Your confidence level in your prediction
126
+ - How useful the reasoning was in helping you make your prediction (info_utility)
127
+ ...
128
+ ```
129
+
130
+ The internal evaluation used this prompt format. Performance with other formats should be assessed separately.
131
 
132
  The model replies with a `<think>...</think>` block followed by a JSON answer:
133
 
 
135
  {"p_yes": 0.27, "p_no": 0.73, "confidence": 0.8, "info_utility": 0.6}
136
  ```
137
 
138
+ Parse the JSON answer after the closing `</think>` tag and use `p_yes` as the YES probability. Validate that `p_yes` and `p_no` fall between 0 and 1 and sum to 1 within numerical tolerance. Treat missing, incomplete or invalid JSON as a failed forecast.
139
+
140
+ The additional fields follow the prompt's own definitions: `confidence` is the model's self-reported confidence in its prediction, and `info_utility` is its self-reported rating of how useful the provided evidence and reasoning were. Neither field is used by the training reward or the evaluation, and their quality has not been assessed.
141
+
142
+ The examples allow up to 1,024 generated tokens. This limit does not guarantee a complete answer; detect truncation (no `</think>` tag or no JSON after it) and handle it explicitly.
143
 
144
  ## How to run
145
 
 
169
  print(resp.choices[0].message.content)
170
  ```
171
 
172
+ Use bfloat16 to reproduce the reported evaluation. Results for float16 serving were not evaluated.
173
 
174
  ### Transformers + PEFT
175
 
 
200
  | bf16 base + adapter | ~28 GB weights, 40-50 GB with KV cache at 8k context | A100 40/80GB, H100, L40S 48GB |
201
  | 4-bit base (`--quantization bitsandbytes --load-format bitsandbytes`) | ~14 GB | RTX 4090, A10, L4 |
202
 
203
+ Four-bit serving reduces weight memory requirements. Its effects on forecast quality, latency and throughput depend on hardware and serving configuration and were not measured. The model-card evaluation used the bf16 base with the adapter.
204
+
205
+ Memory and concurrency requirements depend on input length, generated tokens and batch size. The hardware figures above are approximate planning estimates, not measured benchmarks.
206
+
207
+ ## Limitations
208
+
209
+ This adapter is intended for English-language binary forecasting questions supplied with background information and retrieved evidence. Performance depends on the relevance, quality and timing of that evidence.
210
+
211
+ The results do not establish general equivalence to GPT-4.1 or performance across all question types and domains. The later-market evaluation selected researchable binary questions and excluded several categories.
212
+
213
+ Trading results in the accompanying study are simulations, not live performance. Forecasting scores alone do not establish profitability, which also depends on market prices, fees, execution and strategy. Market prices achieved better aggregate forecasting scores than every evaluated model.
214
+
215
+ The model remains over-confident on the internal test (ECE 0.20). No calibration layer is shipped.
216
+
217
+ Lower serving costs relative to the production tool have not yet been demonstrated in the supplied evidence.
218
 
219
  ## License
220