Claude commited on
Commit
3339913
·
unverified ·
1 Parent(s): a84fd7d

Add algotrader 2.0: a backtester that tries to prove itself wrong

Browse files

Most backtesting tools answer "how much would this have made?". This adds a
second, harder question: how much of that was luck?

New `algotrader` package (additive — v1's agentic_ai_system, its tests, CI and
Docker setup are untouched):

- engine.py: vectorised backtester with an explicit no-look-ahead contract.
A strategy emits a target exposure at bar t's close; the engine holds
position[t] = target[t - lag] with lag >= 1, so a signal computed on a bar
cannot earn that bar's move. Costs are charged on exposure changes, plus a
borrow fee on shorts.
- validation/: the point of the release.
- permutation.py — Monte-Carlo permutation test. Shuffles bar order while
preserving each bar's anatomy, then re-runs the same rule on hundreds of
structure-free markets. Block-bootstrap mode preserves vol clustering.
- deflated_sharpe.py — Bailey & Lopez de Prado's PSR/DSR, charging for every
parameter variant tried and for skew and fat tails.
- pbo.py — probability of backtest overfitting via CSCV.
- walkforward.py — re-tune, trade forward blind, measure the decay.
- verdict.py: combines the panel into a 0-100 Reality Score and plain-English
warnings. Losing money, drawdowns past 50%, and fewer than 20 trades cap the
score regardless of how good the statistics look.
- data.py: yfinance -> cache -> deterministic simulator, so the app never shows
a stack trace on a cold click. The simulator uses regime switching, Student-t
innovations and persistent volatility; naive GBM flatters strategies.
- strategies.py: 11 strategies with parameter grids. Buy & hold and a coin flip
are permanent controls — a leaderboard without a control group is marketing.
- lab.py, charts.py, cli.py: pipeline shared by the app and the command line.

app.py is a Gradio Space with a Lab tab, an Arena leaderboard ranked by evidence
rather than return, and a How-it-works tab. Ships with SPACE_README.md (HF
frontmatter), a minimal requirements-space.txt so the Space builds in under a
minute, a deploy script and a sync workflow.

Two bugs found while testing: a constant return stream reported a Sharpe of
7.3e16 because the zero-variance guard tested `sd == 0` against a float that was
actually 1.4e-19, and the same pattern sat in PBO's column ranking. Both now use
an absolute floor.

111 tests, no network required. The validation tests check both directions —
the statistics reject noise and detect a genuine edge — since a suite that
always says "no edge" is as useless as one that always says "great edge".

.github/workflows/sync-hf-space.yml ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ name: Sync Hugging Face Space
2
+
3
+ # Publishes app.py + the algotrader package to a Hugging Face Space.
4
+ #
5
+ # Setup (one time):
6
+ # 1. Create the Space at https://huggingface.co/new-space with SDK "Gradio".
7
+ # 2. Add repository secret HF_TOKEN — a write token from
8
+ # https://huggingface.co/settings/tokens
9
+ # 3. Add repository variable HF_SPACE_ID, e.g. "your-username/backtest-reality-check".
10
+ #
11
+ # Without those set the job is skipped, so forks and PRs are unaffected.
12
+
13
+ on:
14
+ push:
15
+ branches: [main]
16
+ paths:
17
+ - 'app.py'
18
+ - 'algotrader/**'
19
+ - 'SPACE_README.md'
20
+ - 'requirements-space.txt'
21
+ - 'tests/test_v2_*.py'
22
+ - '.github/workflows/sync-hf-space.yml'
23
+ workflow_dispatch:
24
+
25
+ jobs:
26
+ sync:
27
+ runs-on: ubuntu-latest
28
+ if: vars.HF_SPACE_ID != ''
29
+ steps:
30
+ - uses: actions/checkout@v4
31
+
32
+ - uses: actions/setup-python@v5
33
+ with:
34
+ python-version: '3.11'
35
+
36
+ - name: Verify the Space actually runs before publishing it
37
+ run: |
38
+ pip install --quiet -r requirements-space.txt pytest
39
+ python -m pytest tests/test_v2_engine.py tests/test_v2_validation.py \
40
+ tests/test_v2_strategies.py -q
41
+ ALGOTRADER_OFFLINE=1 python -c "import app; app.build_app(); print('Space builds OK')"
42
+
43
+ - name: Push to the Space
44
+ env:
45
+ HF_TOKEN: ${{ secrets.HF_TOKEN }}
46
+ run: |
47
+ if [ -z "$HF_TOKEN" ]; then
48
+ echo "HF_TOKEN secret is not set; skipping publish." >&2
49
+ exit 0
50
+ fi
51
+ chmod +x scripts/deploy_hf_space.sh
52
+ ./scripts/deploy_hf_space.sh "${{ vars.HF_SPACE_ID }}"
README.md CHANGED
@@ -1,130 +1,179 @@
1
- # Algorithmic Trading
2
 
3
- FinRL reinforcement-learning trading with Alpaca execution, plus optional Yahoo Finance OHLCV for unlabeled real-price research. Parallel LLC.
4
 
5
- This is **research and paper-trading infrastructure**. Live capital requires a separate evaluation contract, feature-parity tests, and a rewritten execution path. Do not treat `paper_trading: false` as a promotion gate.
 
 
6
 
7
- ---
8
-
9
- ## 1. Title and Summary
10
-
11
- **Algorithmic Trading**
12
- Northwestern-trained data-engineering practice applied to a trading loop: ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
13
-
14
- GitHub `main` is the FinRL / Docker / Streamlit tree. `dev` is the integration branch. Yahoo is an additive `data_source.type`, not a replacement for Alpaca or FinRL.
15
-
16
- **Design themes**
17
 
18
- * Four ingest paths: CSV replay, synthetic GBM, Alpaca REST, Yahoo (`yfinance>=1.0`)
19
- * FinRL policies (PPO, A2C, DDPG, TD3) on a Gymnasium-style environment
20
- * Alpaca for authenticated market data and order routing (paper by default)
21
- * Yahoo for delayed public bars when no broker key is available
22
- * Secrets from environment (`ALPACA_API_KEY`, `ALPACA_SECRET_KEY`), never committed
23
- * Tests and Docker/CI as already present on this tree
24
 
25
  ---
26
 
27
- ## 2. Concepts and Methods
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28
 
29
- ### Market data
30
 
31
- | Source | When to use | Failure modes |
32
- | ------ | ----------- | ------------- |
33
- | **CSV** | Offline replay; default in `config.yaml` | Missing path or OHLCV columns → `None` |
34
- | **Synthetic** | Unit tests and demos | GBM is not tradable edge |
35
- | **Alpaca** | Authenticated bars and live/paper orders | Auth, feed, and rate-limit failures |
36
- | **Yahoo** | Real Close without a broker account | Unofficial API, ~15 min delay, interval lookback caps (1m ≈ 7 days). Pin `yfinance>=1.0`; 0.2.x fails against the current chart API |
37
 
38
- `load_data` dispatches on `data_source.type`. Existing `alpaca` / `csv` / `synthetic` branches are unchanged.
 
 
 
 
 
 
 
 
39
 
40
- ### Strategy and FinRL
41
 
42
- * `StrategyAgent`: SMA, RSI, Bollinger, MACD on Close; teaching rule, not an alpha claim
43
- * `FinRLAgent`: PPO / A2C / DDPG / TD3 via Stable-Baselines3; persist under `models/`
44
- * `ExecutionAgent` / `AlpacaBroker`: paper simulation or Alpaca market/limit orders
 
 
 
45
 
46
- Backtests in this repo are in-sample passes unless you add a purged walk-forward yourself. Leakage is the null hypothesis.
 
47
 
48
- ---
49
 
50
- ## 3. Stack
 
 
51
 
52
- | Layer | Tools |
53
- | ----- | ----- |
54
- | Language | Python 3.11 (CI); 3.8+ stated for local |
55
- | RL | FinRL / Stable-Baselines3, Gym/Gymnasium, PyTorch |
56
- | Broker | alpaca-py |
57
- | Market data | Alpaca REST; yfinance ≥ 1.0 (Yahoo) |
58
- | Tabular | pandas, NumPy, scikit-learn |
59
- | UI | Streamlit, Dash, Jupyter widgets |
60
- | Deploy | Docker Compose, GitHub Actions |
61
- | Tests | pytest, pytest-cov |
62
 
63
- ---
64
 
65
- ## 4. Structure
66
 
67
- ```
68
- algorithmic_trading/
69
- ├── agentic_ai_system/ # ingest, strategy, FinRL, Alpaca, Yahoo
70
- ├── ui/ # Streamlit, Dash, Jupyter, WebSocket
71
- ├── tests/
72
- ├── models/ # trained artifacts (gitignored bodies)
73
- ├── data/ # generated CSV (gitignored)
74
- ├── scripts/ # Docker / deploy helpers
75
- ├── .github/workflows/ # CI/CD, release, backtesting
76
- ├── config.yaml
77
- ├── requirements.txt
78
- ├── Dockerfile
79
- └── docker-compose*.yml
80
- ```
81
 
82
- Branch policy: **`main`** (protected) and **`dev`** only. Do not re-enable Dependabot or the Monday `dependency-updates` workflow; those created extra branches.
83
 
84
- ---
85
 
86
- ## 5. Quick start
 
87
 
88
  ```bash
89
- git clone https://github.com/ParallelLLC/algorithmic_trading.git
90
- cd algorithmic_trading
91
- python -m venv .venv && source .venv/bin/activate
92
- pip install -r requirements.txt
93
- cp .env.example .env # Alpaca keys if using alpaca ingest or orders
94
  ```
95
 
96
- Default ingest is CSV. For Yahoo daily bars without a broker:
 
 
97
 
98
- ```yaml
99
- data_source:
100
- type: 'yahoo'
101
- trading:
102
- symbol: 'AAPL'
103
- timeframe: '1d'
104
- ```
105
 
106
  ```bash
107
- python demo.py
108
- python -m agentic_ai_system.main --mode backtest --start-date 2024-01-01 --end-date 2024-12-31
109
- pytest tests/ -q
110
  ```
111
 
112
- UI launchers and Docker are documented in `UI_SETUP.md` and `DOCKER_HUB_SETUP.md`. Paper-trade before live. Yahoo is not a SIP tape.
 
 
 
 
113
 
114
- ---
115
 
116
- ## 6. Configuration (additive Yahoo keys)
 
 
 
117
 
118
- | Key | Meaning |
119
- | --- | ------- |
120
- | `data_source.type` | `csv` \| `synthetic` \| `alpaca` \| `yahoo` |
121
- | `yahoo.start_date` / `end_date` | Historical window; clamped per Yahoo interval limits |
122
- | `yahoo.auto_adjust` | Passed to `yfinance` |
123
- | `execution.broker_api` | `paper` \| `alpaca_paper` \| `alpaca_live` |
124
- | `finrl.algorithm` | PPO, A2C, DDPG, TD3 |
125
-
126
- ---
127
 
128
- **License:** Apache License 2.0
129
- **Organization:** [Parallel LLC](https://github.com/ParallelLLC)
130
- **Repository:** <https://github.com/ParallelLLC/algorithmic_trading>
 
1
+ # Backtest Reality Check
2
 
3
+ **algotrader 2.0 — a backtester that tries to prove itself wrong.**
4
 
5
+ Most backtesting tools answer *"how much would this have made?"*. That is the easy
6
+ question, and the answer is almost always flattering. This one answers the question
7
+ you actually need before risking money: **how much of that was luck?**
8
 
9
+ ```bash
10
+ pip install -r requirements-space.txt
11
+ python app.py # the Gradio app on localhost:7860
12
+ python -m algotrader.cli lab --symbol SPY --strategy sma_cross
13
+ ```
 
 
 
 
 
14
 
15
+ <sub>The v1 agentic trading system (FinRL, Alpaca, Yahoo ingest, Streamlit/Dash UIs) is
16
+ unchanged and still lives here — see [docs/AGENTIC_SYSTEM_V1.md](docs/AGENTIC_SYSTEM_V1.md).</sub>
 
 
 
 
17
 
18
  ---
19
 
20
+ ## The four ways a backtest lies
21
+
22
+ | The lie | The test | Where |
23
+ |---|---|---|
24
+ | The market had no structure to find | Monte-Carlo **permutation test** — re-run your rule on hundreds of shuffled markets | `algotrader/validation/permutation.py` |
25
+ | You tried 200 things and reported the best | **Deflated Sharpe Ratio** — charge for every variant you tried | `algotrader/validation/deflated_sharpe.py` |
26
+ | The parameters were fitted to the past | **PBO** (CSCV) and **walk-forward** | `algotrader/validation/pbo.py`, `walkforward.py` |
27
+ | The edge is smaller than the costs | **Cost stress test** at 3× friction | `algotrader/lab.py` |
28
+
29
+ Each feeds a single **Reality Score** out of 100 with a grade from A to F:
30
+
31
+ | Weight | Component | What it measures |
32
+ |---:|---|---|
33
+ | 30% | Significance | How far outside the shuffled-market null the result sits |
34
+ | 25% | Selection | Deflated Sharpe — does it clear the best-of-N bar |
35
+ | 20% | Walk-forward | How much of the tuned Sharpe survived trading forward |
36
+ | 15% | Overfitting | 1 − PBO |
37
+ | 10% | Robustness | Sharpe retained when costs triple |
38
+
39
+ The scale is deliberately harsh. On most markets, plain buy & hold beats every
40
+ strategy in the arena on evidence, and the built-in coin-flip control out-ranks
41
+ several respectable-looking rules. That is the finding, not a bug.
42
+
43
+ ## The permutation test, concretely
44
+
45
+ We take the real price series and shuffle it. Each bar's gap, high, low, body and
46
+ volume are kept intact, but their **order** is destroyed. The result is a market
47
+ with the same volatility and the same fat tails, and no exploitable structure at
48
+ all. Then we re-run *your exact rule* on hundreds of these shuffled markets.
49
+
50
+ If your Sharpe sits inside that cloud, your rule found nothing that a coin-flip
51
+ market would not also have handed it. The p-value is the share of shuffled markets
52
+ that did as well or better.
53
+
54
+ Block mode resamples contiguous chunks instead of single bars, preserving
55
+ short-horizon momentum and volatility clustering — a harder null that trend
56
+ strategies deserve to be held to.
57
+
58
+ ## No look-ahead, by construction
59
+
60
+ A strategy emits a target exposure at each bar's close using only data up to that
61
+ bar. The engine holds `position[t] = target[t - lag]` with `lag >= 1`, so a signal
62
+ computed on Tuesday's close cannot earn Tuesday's move.
63
+
64
+ That is the single line where look-ahead could enter, and the test suite asserts it
65
+ from four directions — including that truncating the data never changes the equity
66
+ curve before the cut, and that a `lag=0` request is refused outright.
67
+
68
+ ## Python API
69
+
70
+ ```python
71
+ from algotrader import LabConfig, run_lab
72
+
73
+ report = run_lab(LabConfig(
74
+ symbol="SPY",
75
+ start="2015-01-01",
76
+ strategy="sma_cross",
77
+ params={"fast": 20, "slow": 100},
78
+ commission_bps=1.0,
79
+ slippage_bps=2.0,
80
+ n_permutations=500,
81
+ ))
82
+
83
+ print(report.verdict["grade"], report.verdict["score"])
84
+ print("p-value ", report.permutation.p_value)
85
+ print("deflated Sharpe ", report.dsr["dsr"])
86
+ print("overfit prob. ", report.pbo["pbo"])
87
+ print("walk-forward eff.", report.walkforward["efficiency"])
88
+ for flag in report.verdict["flags"]:
89
+ print(" !", flag)
90
+ ```
91
 
92
+ Lower-level pieces compose on their own:
93
 
94
+ ```python
95
+ from algotrader import load_ohlcv, run_backtest, get_strategy
96
+ from algotrader.types import CostModel
 
 
 
97
 
98
+ market = load_ohlcv("BTC-USD", "2018-01-01")
99
+ strategy = get_strategy("donchian_breakout")
100
+ result = run_backtest(
101
+ market.df,
102
+ strategy.generate(market.df, {"window": 55}),
103
+ costs=CostModel(commission_bps=1, slippage_bps=5, short_borrow_bps=50),
104
+ )
105
+ print(result.metrics["sharpe"], result.metrics["max_drawdown"])
106
+ ```
107
 
108
+ ## CLI
109
 
110
+ ```bash
111
+ python -m algotrader.cli strategies # list the zoo
112
+ python -m algotrader.cli lab --symbol NVDA --strategy rsi_reversion --permutations 500
113
+ python -m algotrader.cli lab --symbol SPY --param fast=10 --param slow=50 --json
114
+ python -m algotrader.cli arena --symbol BTC-USD --start 2018-01-01
115
+ ```
116
 
117
+ `--source synthetic` forces the offline simulator, which makes runs fully
118
+ deterministic and network-free.
119
 
120
+ ## The strategy zoo
121
 
122
+ `buy_and_hold` · `sma_cross` · `ema_cross` · `macd_trend` · `rsi_reversion` ·
123
+ `bollinger_reversion` · `donchian_breakout` · `momentum` · `vol_target_momentum` ·
124
+ `channel_trend` · `coin_flip`
125
 
126
+ Buy & hold and the coin flip are controls, and they stay in the arena on purpose: a
127
+ leaderboard without a control group is marketing, not measurement.
 
 
 
 
 
 
 
 
128
 
129
+ Adding one is a function and a registry entry — see `algotrader/strategies.py`.
130
 
131
+ ## Data
132
 
133
+ Live prices come from Yahoo Finance. When the network is unavailable or rate-limited,
134
+ the app falls back to a deterministic market simulator — regime switching, Student-t
135
+ innovations, persistent volatility — and says so on every result. Naive geometric
136
+ Brownian motion flatters strategies; this simulator does not.
 
 
 
 
 
 
 
 
 
 
137
 
138
+ Set `ALGOTRADER_OFFLINE=1` to skip network access entirely.
139
 
140
+ ## Deploying the Hugging Face Space
141
 
142
+ The Space ships `app.py` plus the `algotrader` package and nothing else, so it builds
143
+ in well under a minute:
144
 
145
  ```bash
146
+ HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <your-username>/backtest-reality-check
 
 
 
 
147
  ```
148
 
149
+ Or set the `HF_TOKEN` secret and `HF_SPACE_ID` variable on the repository and let
150
+ `.github/workflows/sync-hf-space.yml` publish on every push to `main`. The workflow
151
+ runs the test suite and builds the app before it publishes anything.
152
 
153
+ `SPACE_README.md` is the Space card (with the Hugging Face YAML frontmatter);
154
+ `requirements-space.txt` is its dependency set. The root `requirements.txt` still
155
+ carries the full v1 stack for CI, Docker and the FinRL agents.
156
+
157
+ ## Tests
 
 
158
 
159
  ```bash
160
+ python -m pytest tests/test_v2_*.py -q # 111 tests, ~15s, no network
 
 
161
  ```
162
 
163
+ The validation tests check both directions, which is the part that matters: the
164
+ statistics must reject noise **and** detect a real edge. They build a market with
165
+ genuine serial correlation and assert that the permutation test finds it, that PBO
166
+ stays near 0.5 on pure noise and drops below 0.15 when one variant is genuinely
167
+ better, and that walk-forward efficiency survives.
168
 
169
+ ## References
170
 
171
+ - Bailey & López de Prado (2014), *The Deflated Sharpe Ratio: Correcting for Selection
172
+ Bias, Backtest Overfitting and Non-Normality*
173
+ - Bailey, Borwein, López de Prado & Zhu (2016), *The Probability of Backtest Overfitting*
174
+ - Masters (2018), *Permutation and Randomization Tests for Trading System Development*
175
 
176
+ ## License
 
 
 
 
 
 
 
 
177
 
178
+ Apache-2.0. Research tooling, not investment advice. Nothing here is a
179
+ recommendation to trade.
 
SPACE_README.md ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: Backtest Reality Check
3
+ emoji: 🎲
4
+ colorFrom: blue
5
+ colorTo: gray
6
+ sdk: gradio
7
+ sdk_version: 5.49.1
8
+ app_file: app.py
9
+ pinned: true
10
+ license: apache-2.0
11
+ short_description: Your backtest is probably lying to you. This proves it.
12
+ tags:
13
+ - finance
14
+ - quantitative-finance
15
+ - algorithmic-trading
16
+ - backtesting
17
+ - statistics
18
+ - time-series
19
+ ---
20
+
21
+ # Backtest Reality Check
22
+
23
+ **Your backtest is probably lying to you.**
24
+
25
+ Pick a market and a trading rule. This Space runs the backtest — and then spends
26
+ the rest of its effort trying to prove the result was luck.
27
+
28
+ Most backtesting tools answer *"how much would this have made?"*. That is the easy
29
+ question, and the answer is almost always flattering. This one answers the question
30
+ you need before risking money: **how much of that was luck?**
31
+
32
+ ## The four ways a backtest lies, and the test for each
33
+
34
+ | The lie | The test |
35
+ |---|---|
36
+ | The market had no structure to find | **Permutation test** — re-run your rule on hundreds of shuffled markets |
37
+ | You tried 200 things and reported the best | **Deflated Sharpe Ratio** — charge for every variant you tried |
38
+ | The parameters were fitted to the past | **PBO + walk-forward** — does the in-sample winner keep winning? |
39
+ | The edge is smaller than the costs | **Cost stress test** — triple the friction and see what survives |
40
+
41
+ Each contributes to a single **Reality Score** out of 100, with a grade from A to F.
42
+ The scale is deliberately harsh. Most strategies people post online score below 40.
43
+
44
+ ## Try this first
45
+
46
+ Run the **Arena** tab on `SPY`. On most markets and most date ranges, plain
47
+ **buy & hold** tops the leaderboard, and the **coin flip** control out-ranks
48
+ several respectable-looking strategies. That is not a bug in the app — it is the
49
+ finding.
50
+
51
+ ## How the permutation test works
52
+
53
+ We take the real price series and shuffle it. Each bar's gap, high, low, body and
54
+ volume are kept intact, but their **order** is destroyed. The result is a market
55
+ with the same volatility and the same fat tails, and no exploitable structure at
56
+ all. Then we re-run *your exact rule* on hundreds of these shuffled markets.
57
+
58
+ If your Sharpe ratio sits comfortably inside that cloud, your rule found nothing
59
+ a coin-flip market would not also have handed it.
60
+
61
+ ## No look-ahead, by construction
62
+
63
+ A strategy emits a target exposure at each bar's close using only data up to that
64
+ bar. The engine holds `position[t] = target[t - lag]` with `lag >= 1`, so a signal
65
+ computed on Tuesday's close cannot earn Tuesday's move. That is the single line
66
+ where look-ahead could enter, and the test suite asserts it directly.
67
+
68
+ ## Use it from Python
69
+
70
+ ```python
71
+ from algotrader import LabConfig, run_lab
72
+
73
+ report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross"))
74
+ print(report.verdict["grade"], report.verdict["score"])
75
+ print(report.permutation.p_value, report.dsr["dsr"], report.pbo["pbo"])
76
+ ```
77
+
78
+ Or from the command line:
79
+
80
+ ```bash
81
+ python -m algotrader.cli lab --symbol SPY --strategy donchian_breakout --permutations 500
82
+ python -m algotrader.cli arena --symbol BTC-USD
83
+ ```
84
+
85
+ ## Data
86
+
87
+ Live prices come from Yahoo Finance. When the network is unavailable or rate-limited,
88
+ the app falls back to a deterministic market simulator with regime switching, fat
89
+ tails and volatility clustering — and says so, clearly, on every result. The
90
+ statistics remain valid; they are just measured on a simulated market.
91
+
92
+ ## References
93
+
94
+ - Bailey & López de Prado (2014), *The Deflated Sharpe Ratio*
95
+ - Bailey, Borwein, López de Prado & Zhu (2016), *The Probability of Backtest Overfitting*
96
+ - Masters (2018), *Permutation and Randomization Tests for Trading System Development*
97
+
98
+ ---
99
+
100
+ Apache-2.0. Research tooling, not investment advice. Nothing here is a
101
+ recommendation to trade.
algotrader/__init__.py ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """algotrader 2.0 — a backtester that tries to prove itself wrong.
2
+
3
+ Most backtesting libraries answer "how much would this have made?". This one
4
+ answers the question that actually matters before you risk money: "how much of
5
+ that was luck?"
6
+
7
+ Quick start::
8
+
9
+ from algotrader import LabConfig, run_lab
10
+
11
+ report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross"))
12
+ print(report.verdict["verdict"])
13
+ """
14
+
15
+ from .data import load_ohlcv, simulate_ohlcv
16
+ from .engine import run_backtest
17
+ from .lab import LabConfig, LabReport, run_arena, run_lab
18
+ from .metrics import compute_metrics
19
+ from .strategies import REGISTRY, get_strategy, list_strategies
20
+ from .types import BacktestResult, CostModel, MarketData
21
+ from .verdict import reality_score
22
+
23
+ __version__ = "2.0.0"
24
+
25
+ __all__ = [
26
+ "__version__",
27
+ "LabConfig",
28
+ "LabReport",
29
+ "run_lab",
30
+ "run_arena",
31
+ "run_backtest",
32
+ "compute_metrics",
33
+ "load_ohlcv",
34
+ "simulate_ohlcv",
35
+ "get_strategy",
36
+ "list_strategies",
37
+ "REGISTRY",
38
+ "BacktestResult",
39
+ "CostModel",
40
+ "MarketData",
41
+ "reality_score",
42
+ ]
algotrader/charts.py ADDED
@@ -0,0 +1,286 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Plotly figures for the Lab.
2
+
3
+ Colour system (dark surface, validated for CVD separation):
4
+
5
+ * **blue** is always *your strategy's honest result* — the realised equity
6
+ curve, the out-of-sample fold, the observed Sharpe.
7
+ * **orange** is always *the thing it is measured against* — buy & hold, the
8
+ in-sample fold, the null distribution.
9
+
10
+ Holding that mapping across every figure means a reader learns it once.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ from typing import Dict, Optional
16
+
17
+ import numpy as np
18
+ import pandas as pd
19
+ import plotly.graph_objects as go
20
+
21
+ SURFACE = "#1a1a19"
22
+ PAGE = "#0d0d0d"
23
+ INK = "#ffffff"
24
+ INK_SECONDARY = "#c3c2b7"
25
+ INK_MUTED = "#898781"
26
+ GRID = "#2c2c2a"
27
+ AXIS = "#383835"
28
+
29
+ SUBJECT = "#3987e5" # categorical slot 1
30
+ REFERENCE = "#d95926" # categorical slot 2
31
+ NEGATIVE = "#e66767" # negative arm of the diverging pair (drawdowns)
32
+
33
+ FONT = 'system-ui, -apple-system, "Segoe UI", sans-serif'
34
+
35
+ _EMPTY_NOTE = "Run an analysis to populate this chart."
36
+
37
+
38
+ def _base_layout(title: str, height: int = 340, **kwargs) -> dict:
39
+ return dict(
40
+ title=dict(text=title, font=dict(size=15, color=INK), x=0, xanchor="left", pad=dict(b=8)),
41
+ paper_bgcolor=PAGE,
42
+ plot_bgcolor=SURFACE,
43
+ font=dict(family=FONT, size=12, color=INK_SECONDARY),
44
+ height=height,
45
+ margin=dict(l=56, r=24, t=48, b=40),
46
+ hovermode="x unified",
47
+ hoverlabel=dict(bgcolor=SURFACE, bordercolor=AXIS, font=dict(color=INK, family=FONT)),
48
+ xaxis=dict(gridcolor=GRID, linecolor=AXIS, zeroline=False, tickfont=dict(color=INK_MUTED)),
49
+ yaxis=dict(gridcolor=GRID, linecolor=AXIS, zeroline=False, tickfont=dict(color=INK_MUTED)),
50
+ legend=dict(
51
+ orientation="h", yanchor="bottom", y=1.02, xanchor="left", x=0,
52
+ font=dict(color=INK_SECONDARY, size=11), bgcolor="rgba(0,0,0,0)",
53
+ ),
54
+ **kwargs,
55
+ )
56
+
57
+
58
+ def empty_figure(message: str = _EMPTY_NOTE, height: int = 340) -> go.Figure:
59
+ fig = go.Figure()
60
+ fig.update_layout(**_base_layout("", height=height))
61
+ fig.update_xaxes(visible=False)
62
+ fig.update_yaxes(visible=False)
63
+ fig.add_annotation(
64
+ text=message, showarrow=False, xref="paper", yref="paper", x=0.5, y=0.5,
65
+ font=dict(color=INK_MUTED, size=13),
66
+ )
67
+ return fig
68
+
69
+
70
+ def equity_chart(report) -> go.Figure:
71
+ """Strategy equity against buy & hold, both indexed to the same start."""
72
+ bt = report.backtest
73
+ strat = bt.equity / bt.equity.iloc[0] * 100.0
74
+ bench = bt.benchmark_equity / bt.benchmark_equity.iloc[0] * 100.0
75
+
76
+ fig = go.Figure()
77
+ fig.add_trace(
78
+ go.Scatter(
79
+ x=bench.index, y=bench.to_numpy(), name="Buy & hold", mode="lines",
80
+ line=dict(color=REFERENCE, width=2, dash="dash"),
81
+ hovertemplate="Buy & hold %{y:.1f}<extra></extra>",
82
+ )
83
+ )
84
+ fig.add_trace(
85
+ go.Scatter(
86
+ x=strat.index, y=strat.to_numpy(), name=report.strategy.name, mode="lines",
87
+ line=dict(color=SUBJECT, width=2),
88
+ hovertemplate=report.strategy.name + " %{y:.1f}<extra></extra>",
89
+ )
90
+ )
91
+
92
+ # Direct-label the two endpoints; the axis and tooltip carry everything else.
93
+ for series, color, label in ((strat, SUBJECT, report.strategy.name), (bench, REFERENCE, "Buy & hold")):
94
+ fig.add_annotation(
95
+ x=series.index[-1], y=float(series.iloc[-1]),
96
+ text=f" {label}: {series.iloc[-1]:.0f}", showarrow=False,
97
+ xanchor="left", font=dict(color=color, size=11),
98
+ )
99
+
100
+ fig.update_layout(**_base_layout("Growth of 100 (net of costs)", height=360))
101
+ fig.update_layout(margin=dict(l=56, r=140, t=48, b=40))
102
+ return fig
103
+
104
+
105
+ def drawdown_chart(report) -> go.Figure:
106
+ """Underwater plot — how deep, and for how long."""
107
+ from .metrics import drawdown_series
108
+
109
+ dd = drawdown_series(report.backtest.equity) * 100.0
110
+ fig = go.Figure(
111
+ go.Scatter(
112
+ x=dd.index, y=dd.to_numpy(), mode="lines", name="Drawdown",
113
+ line=dict(color=NEGATIVE, width=2), fill="tozeroy",
114
+ fillcolor="rgba(230,103,103,0.18)",
115
+ hovertemplate="Drawdown %{y:.1f}%<extra></extra>",
116
+ )
117
+ )
118
+ trough = float(dd.min())
119
+ fig.add_annotation(
120
+ x=dd.idxmin(), y=trough, text=f"worst {trough:.1f}%", showarrow=True,
121
+ arrowhead=0, arrowcolor=AXIS, ay=24, font=dict(color=INK_SECONDARY, size=11),
122
+ )
123
+ fig.update_layout(**_base_layout("Drawdown", height=240, showlegend=False))
124
+ fig.update_yaxes(ticksuffix="%")
125
+ return fig
126
+
127
+
128
+ def permutation_chart(report) -> go.Figure:
129
+ """The headline chart: your Sharpe against Sharpes from shuffled markets."""
130
+ perm = report.permutation
131
+ if perm is None or perm.null.size == 0:
132
+ return empty_figure("Permutation test was skipped.", height=320)
133
+
134
+ null = perm.null
135
+ fig = go.Figure()
136
+ fig.add_trace(
137
+ go.Histogram(
138
+ x=null, name="Shuffled markets (no real edge)", nbinsx=44,
139
+ marker=dict(color="rgba(217,89,38,0.55)", line=dict(color=REFERENCE, width=1)),
140
+ hovertemplate="Sharpe %{x:.2f}<br>%{y} shuffles<extra></extra>",
141
+ )
142
+ )
143
+
144
+ top = np.histogram(null, bins=44)[0].max() if null.size else 1
145
+ fig.add_trace(
146
+ go.Scatter(
147
+ x=[perm.observed, perm.observed], y=[0, top * 1.08], mode="lines",
148
+ name="Your strategy", line=dict(color=SUBJECT, width=2),
149
+ hovertemplate="Your Sharpe %{x:.2f}<extra></extra>",
150
+ )
151
+ )
152
+ fig.add_annotation(
153
+ x=perm.observed, y=top * 1.08, text=f" your Sharpe {perm.observed:.2f}",
154
+ showarrow=False, xanchor="left", font=dict(color=SUBJECT, size=11),
155
+ )
156
+
157
+ beats = (null >= perm.observed).mean() * 100.0
158
+ fig.update_layout(
159
+ **_base_layout(
160
+ f"Permutation test — {beats:.0f}% of structure-free markets did this well or better "
161
+ f"(p = {perm.p_value:.3f})",
162
+ height=320,
163
+ )
164
+ )
165
+ fig.update_layout(hovermode="closest", bargap=0.02)
166
+ fig.update_xaxes(title=dict(text="Annualised Sharpe ratio", font=dict(color=INK_MUTED, size=11)))
167
+ # Headroom so the "your Sharpe" label never collides with the plot edge.
168
+ fig.update_yaxes(
169
+ title=dict(text="Shuffled markets", font=dict(color=INK_MUTED, size=11)),
170
+ range=[0, top * 1.28],
171
+ )
172
+ return fig
173
+
174
+
175
+ def walkforward_chart(report) -> go.Figure:
176
+ """In-sample vs out-of-sample Sharpe, fold by fold."""
177
+ folds = report.walkforward.get("folds") or []
178
+ if not folds:
179
+ return empty_figure(report.walkforward.get("note") or _EMPTY_NOTE, height=300)
180
+
181
+ labels = [f"Fold {f['fold']}<br><span style='font-size:10px'>{f['test_start'][:7]}</span>" for f in folds]
182
+ fig = go.Figure()
183
+ fig.add_trace(
184
+ go.Bar(
185
+ x=labels, y=[f["is_sharpe"] for f in folds], name="In-sample (tuned)",
186
+ marker=dict(color=REFERENCE, line=dict(color=SURFACE, width=2)),
187
+ hovertemplate="In-sample Sharpe %{y:.2f}<extra></extra>",
188
+ )
189
+ )
190
+ fig.add_trace(
191
+ go.Bar(
192
+ x=labels, y=[f["oos_sharpe"] for f in folds], name="Out-of-sample (blind)",
193
+ marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
194
+ hovertemplate="Out-of-sample Sharpe %{y:.2f}<extra></extra>",
195
+ )
196
+ )
197
+ eff = report.walkforward.get("efficiency", 0.0)
198
+ fig.update_layout(
199
+ **_base_layout(f"Walk-forward — {eff:.0%} of the tuned Sharpe survived out of sample", height=300)
200
+ )
201
+ fig.update_layout(barmode="group", bargap=0.35, bargroupgap=0.08, hovermode="x unified")
202
+ fig.add_hline(y=0, line=dict(color=AXIS, width=1))
203
+ return fig
204
+
205
+
206
+ def score_chart(verdict: Dict[str, object]) -> go.Figure:
207
+ """The five components behind the Reality Score."""
208
+ components = verdict.get("components") or {}
209
+ if not components:
210
+ return empty_figure(height=260)
211
+
212
+ pretty = {
213
+ "significance": "Beats shuffled markets",
214
+ "selection": "Survives selection bias",
215
+ "walk_forward": "Holds up walking forward",
216
+ "overfitting": "Not overfit (PBO)",
217
+ "robustness": "Survives 3x costs",
218
+ }
219
+ keys = list(pretty)
220
+ values = [float(components.get(k, 0.0)) for k in keys]
221
+
222
+ fig = go.Figure(
223
+ go.Bar(
224
+ x=values, y=[pretty[k] for k in keys], orientation="h",
225
+ marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
226
+ text=[f"{v:.0f}" for v in values], textposition="outside",
227
+ textfont=dict(color=INK_SECONDARY, size=11),
228
+ hovertemplate="%{y}: %{x:.0f}/100<extra></extra>",
229
+ )
230
+ )
231
+ fig.update_layout(**_base_layout("Where the score comes from", height=260, showlegend=False))
232
+ fig.update_layout(margin=dict(l=190, r=48, t=48, b=32), hovermode="closest")
233
+ fig.update_xaxes(range=[0, 108], tickvals=[0, 25, 50, 75, 100])
234
+ fig.update_yaxes(autorange="reversed")
235
+ return fig
236
+
237
+
238
+ def arena_chart(table: pd.DataFrame) -> go.Figure:
239
+ """Leaderboard bars. One measure, one colour — the table carries the rest."""
240
+ if table is None or table.empty:
241
+ return empty_figure(height=380)
242
+
243
+ ordered = table.iloc[::-1]
244
+ fig = go.Figure(
245
+ go.Bar(
246
+ x=ordered["Sharpe"].to_numpy(), y=ordered["Strategy"].tolist(), orientation="h",
247
+ marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
248
+ customdata=np.column_stack([ordered["p-value"].to_numpy(), ordered["DSR"].to_numpy()]),
249
+ hovertemplate="%{y}<br>Sharpe %{x:.2f}<br>p = %{customdata[0]:.3f}"
250
+ "<br>Deflated Sharpe %{customdata[1]:.2f}<extra></extra>",
251
+ )
252
+ )
253
+ # Direct-label only what matters: the ones that actually cleared significance.
254
+ for _, row in ordered.iterrows():
255
+ if np.isfinite(row["p-value"]) and row["p-value"] < 0.05:
256
+ fig.add_annotation(
257
+ x=row["Sharpe"], y=row["Strategy"], text=" p &lt; 0.05", showarrow=False,
258
+ xanchor="left" if row["Sharpe"] >= 0 else "right",
259
+ font=dict(color=INK_SECONDARY, size=10),
260
+ )
261
+ fig.update_layout(
262
+ **_base_layout(
263
+ "Strategy arena — Sharpe ratio, ordered by strength of evidence",
264
+ height=max(300, 42 * len(table)),
265
+ showlegend=False,
266
+ )
267
+ )
268
+ fig.update_layout(margin=dict(l=180, r=96, t=48, b=32), hovermode="closest")
269
+ fig.add_vline(x=0, line=dict(color=AXIS, width=1))
270
+ return fig
271
+
272
+
273
+ def exposure_chart(report) -> go.Figure:
274
+ """What the strategy was actually holding, over time."""
275
+ pos = report.backtest.position
276
+ fig = go.Figure(
277
+ go.Scatter(
278
+ x=pos.index, y=pos.to_numpy(), mode="lines", name="Exposure",
279
+ line=dict(color=SUBJECT, width=2, shape="hv"), fill="tozeroy",
280
+ fillcolor="rgba(57,135,229,0.16)",
281
+ hovertemplate="Exposure %{y:.2f}x<extra></extra>",
282
+ )
283
+ )
284
+ fig.update_layout(**_base_layout("Position held", height=200, showlegend=False))
285
+ fig.add_hline(y=0, line=dict(color=AXIS, width=1))
286
+ return fig
algotrader/cli.py ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Command-line interface.
2
+
3
+ python -m algotrader.cli lab --symbol SPY --strategy sma_cross
4
+ python -m algotrader.cli arena --symbol BTC-USD --start 2018-01-01
5
+ python -m algotrader.cli strategies
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import argparse
11
+ import json
12
+ import sys
13
+ from typing import Dict, List
14
+
15
+ from . import __version__
16
+ from .lab import LabConfig, run_arena, run_lab
17
+ from .strategies import REGISTRY, get_strategy
18
+
19
+
20
+ def _parse_params(pairs: List[str] | None) -> Dict[str, float]:
21
+ params: Dict[str, float] = {}
22
+ for pair in pairs or []:
23
+ if "=" not in pair:
24
+ raise SystemExit(f"--param expects name=value, got '{pair}'")
25
+ name, _, value = pair.partition("=")
26
+ params[name.strip()] = float(value)
27
+ return params
28
+
29
+
30
+ def _add_common(parser: argparse.ArgumentParser) -> None:
31
+ parser.add_argument("--symbol", default="SPY")
32
+ parser.add_argument("--start", default="2015-01-01")
33
+ parser.add_argument("--end", default=None)
34
+ parser.add_argument("--interval", default="1d")
35
+ parser.add_argument(
36
+ "--source", default="auto", choices=["auto", "cache", "synthetic"],
37
+ help="'auto' downloads and falls back offline; 'synthetic' forces the simulator.",
38
+ )
39
+ parser.add_argument("--commission-bps", type=float, default=1.0)
40
+ parser.add_argument("--slippage-bps", type=float, default=2.0)
41
+ parser.add_argument("--no-short", action="store_true", help="Long/flat only.")
42
+
43
+
44
+ def _config_from(args: argparse.Namespace, **overrides) -> LabConfig:
45
+ return LabConfig(
46
+ symbol=args.symbol,
47
+ start=args.start,
48
+ end=args.end,
49
+ interval=args.interval,
50
+ source=args.source,
51
+ commission_bps=args.commission_bps,
52
+ slippage_bps=args.slippage_bps,
53
+ allow_short=not args.no_short,
54
+ **overrides,
55
+ )
56
+
57
+
58
+ def _cmd_lab(args: argparse.Namespace) -> int:
59
+ cfg = _config_from(
60
+ args,
61
+ strategy=args.strategy,
62
+ params=_parse_params(args.param),
63
+ n_permutations=args.permutations,
64
+ permutation_method=args.null,
65
+ wf_folds=args.folds,
66
+ )
67
+ progress = None if args.quiet else (lambda f, m: print(f" [{f:5.0%}] {m}", file=sys.stderr))
68
+ report = run_lab(cfg, progress=progress)
69
+
70
+ if args.json:
71
+ payload = {
72
+ "symbol": report.market.symbol,
73
+ "source": report.market.source,
74
+ "strategy": report.strategy.key,
75
+ "params": report.params,
76
+ "metrics": report.backtest.metrics,
77
+ "benchmark_metrics": report.backtest.benchmark_metrics,
78
+ "p_value": report.permutation.p_value if report.permutation else None,
79
+ "deflated_sharpe": report.dsr.get("dsr"),
80
+ "pbo": report.pbo.get("pbo"),
81
+ "walkforward_efficiency": report.walkforward.get("efficiency"),
82
+ "cost_stress": report.cost_stress,
83
+ "verdict": {k: v for k, v in report.verdict.items()},
84
+ }
85
+ print(json.dumps(payload, indent=2, default=str))
86
+ return 0
87
+
88
+ v, m, b = report.verdict, report.backtest.metrics, report.backtest.benchmark_metrics
89
+ bar = "=" * 66
90
+ print(f"\n{bar}")
91
+ print(f" {report.strategy.name} on {report.market.symbol} [{report.market.source} data]")
92
+ print(f" {report.market.start.date()} to {report.market.end.date()} · {len(report.market.df):,} bars")
93
+ print(bar)
94
+ print(f" REALITY SCORE {v['score']:.1f} / 100 GRADE {v['grade']}")
95
+ print(f" {v['headline']}")
96
+ print(bar)
97
+ print(f" Total return {m['total_return']:>9.1%} buy & hold {b['total_return']:>8.1%}")
98
+ print(f" CAGR {m['cagr']:>9.1%} buy & hold {b['cagr']:>8.1%}")
99
+ print(f" Sharpe {m['sharpe']:>9.2f} buy & hold {b['sharpe']:>8.2f}")
100
+ print(f" Max drawdown {m['max_drawdown']:>9.1%}")
101
+ print(f" Trades {int(m.get('n_trades', 0)):>9,}")
102
+ print(bar)
103
+ if report.permutation:
104
+ print(f" Permutation p {report.permutation.p_value:>9.3f} ({report.permutation.n_permutations} shuffled markets)")
105
+ print(f" Deflated Sharpe {report.dsr.get('dsr', 0):>9.2f} (after {report.trials.get('n', 1)} variants)")
106
+ pbo = report.pbo.get("pbo")
107
+ print(f" Overfit prob. {pbo:>9.2f}" if pbo == pbo else " Overfit prob. n/a")
108
+ print(f" Walk-forward eff. {report.walkforward.get('efficiency', 0):>9.2f}")
109
+ print(f" Sharpe at 3x cost {report.cost_stress.get('sharpe_3x', 0):>9.2f}")
110
+ print(bar)
111
+ for flag in v["flags"]:
112
+ print(f" ! {flag}")
113
+ if v["flags"]:
114
+ print(bar)
115
+ print(f" {v['verdict']}\n")
116
+ return 0
117
+
118
+
119
+ def _cmd_arena(args: argparse.Namespace) -> int:
120
+ cfg = _config_from(args)
121
+ progress = None if args.quiet else (lambda f, m: print(f" [{f:5.0%}] {m}", file=sys.stderr))
122
+ table, market, _ = run_arena(cfg, n_permutations=args.permutations, progress=progress)
123
+
124
+ if args.json:
125
+ print(table.to_json(orient="records", indent=2))
126
+ return 0
127
+
128
+ print(f"\n {market.symbol} [{market.source} data] "
129
+ f"{market.start.date()} to {market.end.date()}\n")
130
+ display = table.drop(columns=["key"]).copy()
131
+ for col in ("Return", "CAGR", "MaxDD"):
132
+ display[col] = display[col].map("{:.1%}".format)
133
+ for col in ("Sharpe", "DSR", "Evidence"):
134
+ display[col] = display[col].map("{:.2f}".format)
135
+ display["p-value"] = display["p-value"].map(lambda v: "—" if v != v else f"{v:.3f}")
136
+ print(display.to_string(index=False))
137
+ print("\n Ranked by evidence = (1 - p) x deflated Sharpe, not by return.\n")
138
+ return 0
139
+
140
+
141
+ def _cmd_strategies(args: argparse.Namespace) -> int:
142
+ for key, strategy in REGISTRY.items():
143
+ params = ", ".join(f"{p.name}={p.default:g}" for p in strategy.params) or "no parameters"
144
+ print(f" {key:<22} {strategy.name:<26} [{strategy.family}]")
145
+ print(f" {'':<22} {strategy.description}")
146
+ print(f" {'':<22} defaults: {params}\n")
147
+ return 0
148
+
149
+
150
+ def main(argv: List[str] | None = None) -> int:
151
+ parser = argparse.ArgumentParser(
152
+ prog="algotrader",
153
+ description="Backtest a trading rule, then try to prove the result was luck.",
154
+ )
155
+ parser.add_argument("--version", action="version", version=f"algotrader {__version__}")
156
+ sub = parser.add_subparsers(dest="command", required=True)
157
+
158
+ lab = sub.add_parser("lab", help="Full reality check for one strategy.")
159
+ _add_common(lab)
160
+ lab.add_argument("--strategy", default="sma_cross", choices=sorted(REGISTRY))
161
+ lab.add_argument("--param", action="append", metavar="NAME=VALUE",
162
+ help="Override a strategy parameter. Repeatable.")
163
+ lab.add_argument("--permutations", type=int, default=250)
164
+ lab.add_argument("--null", default="permute", choices=["permute", "block"])
165
+ lab.add_argument("--folds", type=int, default=5)
166
+ lab.add_argument("--json", action="store_true")
167
+ lab.add_argument("--quiet", "-q", action="store_true")
168
+ lab.set_defaults(func=_cmd_lab)
169
+
170
+ arena = sub.add_parser("arena", help="Race every strategy on one market.")
171
+ _add_common(arena)
172
+ arena.add_argument("--permutations", type=int, default=120)
173
+ arena.add_argument("--json", action="store_true")
174
+ arena.add_argument("--quiet", "-q", action="store_true")
175
+ arena.set_defaults(func=_cmd_arena)
176
+
177
+ listing = sub.add_parser("strategies", help="List the strategy zoo.")
178
+ listing.set_defaults(func=_cmd_strategies)
179
+
180
+ args = parser.parse_args(argv)
181
+ try:
182
+ return args.func(args)
183
+ except KeyboardInterrupt:
184
+ return 130
185
+ except Exception as exc: # noqa: BLE001
186
+ print(f"error: {exc}", file=sys.stderr)
187
+ return 1
188
+
189
+
190
+ if __name__ == "__main__":
191
+ raise SystemExit(main())
algotrader/data.py ADDED
@@ -0,0 +1,246 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Market data loading with a three-tier fallback.
2
+
3
+ Order of preference: live Yahoo download -> on-disk cache -> a deterministic
4
+ simulator. The fallback exists because a Hugging Face Space that shows a
5
+ stack trace on the first click is a Space nobody shares. When the simulator is
6
+ used, :class:`~algotrader.types.MarketData` says so and the UI shows it.
7
+ """
8
+
9
+ from __future__ import annotations
10
+
11
+ import hashlib
12
+ import logging
13
+ import os
14
+ from dataclasses import dataclass
15
+ from pathlib import Path
16
+ from typing import Optional
17
+
18
+ import numpy as np
19
+ import pandas as pd
20
+
21
+ from .types import OHLCV_COLUMNS, MarketData
22
+
23
+ logger = logging.getLogger(__name__)
24
+
25
+ CACHE_DIR = Path(os.environ.get("ALGOTRADER_CACHE", Path.home() / ".cache" / "algotrader"))
26
+ NETWORK_ENABLED = os.environ.get("ALGOTRADER_OFFLINE", "").lower() not in ("1", "true", "yes")
27
+
28
+ # Popular tickers get hand-set simulation parameters so the offline demo is at
29
+ # least in the right postcode: annual drift, annual vol, and a starting price.
30
+ @dataclass(frozen=True)
31
+ class SimProfile:
32
+ drift: float
33
+ vol: float
34
+ price: float
35
+
36
+
37
+ SIM_PROFILES: dict[str, SimProfile] = {
38
+ "AAPL": SimProfile(0.24, 0.29, 190.0),
39
+ "MSFT": SimProfile(0.25, 0.27, 410.0),
40
+ "NVDA": SimProfile(0.55, 0.52, 120.0),
41
+ "TSLA": SimProfile(0.30, 0.58, 250.0),
42
+ "AMZN": SimProfile(0.22, 0.33, 180.0),
43
+ "GOOGL": SimProfile(0.20, 0.31, 170.0),
44
+ "META": SimProfile(0.26, 0.40, 500.0),
45
+ "SPY": SimProfile(0.10, 0.16, 550.0),
46
+ "QQQ": SimProfile(0.14, 0.21, 480.0),
47
+ "BTC-USD": SimProfile(0.45, 0.65, 65000.0),
48
+ "ETH-USD": SimProfile(0.35, 0.75, 3000.0),
49
+ "GLD": SimProfile(0.07, 0.14, 200.0),
50
+ "TLT": SimProfile(0.01, 0.15, 95.0),
51
+ }
52
+
53
+ DEFAULT_UNIVERSE = ["SPY", "AAPL", "NVDA", "MSFT", "TSLA", "QQQ", "BTC-USD", "GLD"]
54
+
55
+
56
+ def _seed_for(symbol: str) -> int:
57
+ """Stable per-symbol seed so a given ticker always simulates identically."""
58
+ digest = hashlib.sha256(symbol.upper().encode()).digest()
59
+ return int.from_bytes(digest[:4], "big")
60
+
61
+
62
+ def _normalise(df: pd.DataFrame) -> pd.DataFrame:
63
+ """Coerce any loader's output into a clean lowercase OHLCV frame."""
64
+ if isinstance(df.columns, pd.MultiIndex):
65
+ df = df.copy()
66
+ df.columns = [str(c[0]) for c in df.columns]
67
+ df = df.rename(columns={c: str(c).strip().lower().replace(" ", "_") for c in df.columns})
68
+ if "adj_close" in df.columns and "close" not in df.columns:
69
+ df = df.rename(columns={"adj_close": "close"})
70
+ missing = [c for c in OHLCV_COLUMNS if c not in df.columns]
71
+ for col in missing:
72
+ if col == "volume":
73
+ df["volume"] = 0.0
74
+ elif "close" in df.columns:
75
+ df[col] = df["close"]
76
+ else:
77
+ raise ValueError(f"Price data is missing required column: {col}")
78
+ df = df.loc[:, list(OHLCV_COLUMNS)].astype(float)
79
+ if not isinstance(df.index, pd.DatetimeIndex):
80
+ df.index = pd.to_datetime(df.index)
81
+ df.index = df.index.tz_localize(None) if df.index.tz is not None else df.index
82
+ df = df[~df.index.duplicated(keep="last")].sort_index()
83
+ df = df[df["close"] > 0].dropna(subset=["close"])
84
+ return df
85
+
86
+
87
+ def _cache_path(symbol: str, interval: str) -> Path:
88
+ safe = symbol.upper().replace("/", "_")
89
+ return CACHE_DIR / f"{safe}_{interval}.csv"
90
+
91
+
92
+ def _read_cache(symbol: str, interval: str) -> Optional[pd.DataFrame]:
93
+ path = _cache_path(symbol, interval)
94
+ if not path.exists():
95
+ return None
96
+ try:
97
+ return _normalise(pd.read_csv(path, index_col=0, parse_dates=True))
98
+ except Exception as exc: # pragma: no cover - corrupted cache is not worth failing over
99
+ logger.warning("Ignoring unreadable cache %s: %s", path, exc)
100
+ return None
101
+
102
+
103
+ def _write_cache(symbol: str, interval: str, df: pd.DataFrame) -> None:
104
+ try:
105
+ CACHE_DIR.mkdir(parents=True, exist_ok=True)
106
+ df.to_csv(_cache_path(symbol, interval))
107
+ except Exception as exc: # pragma: no cover - a read-only FS must not break the app
108
+ logger.warning("Could not write cache for %s: %s", symbol, exc)
109
+
110
+
111
+ def _download(symbol: str, start: str, end: str | None, interval: str) -> Optional[pd.DataFrame]:
112
+ if not NETWORK_ENABLED:
113
+ return None
114
+ try:
115
+ import yfinance as yf
116
+ except ImportError:
117
+ logger.info("yfinance not installed; using offline data")
118
+ return None
119
+ try:
120
+ raw = yf.download(
121
+ symbol,
122
+ start=start,
123
+ end=end,
124
+ interval=interval,
125
+ progress=False,
126
+ auto_adjust=True,
127
+ threads=False,
128
+ )
129
+ except Exception as exc:
130
+ logger.warning("Download failed for %s: %s", symbol, exc)
131
+ return None
132
+ if raw is None or len(raw) == 0:
133
+ logger.warning("Download for %s returned no rows", symbol)
134
+ return None
135
+ try:
136
+ return _normalise(raw)
137
+ except Exception as exc:
138
+ logger.warning("Could not normalise download for %s: %s", symbol, exc)
139
+ return None
140
+
141
+
142
+ def simulate_ohlcv(
143
+ symbol: str = "SIM",
144
+ start: str = "2015-01-01",
145
+ end: str | None = None,
146
+ interval: str = "1d",
147
+ seed: Optional[int] = None,
148
+ ) -> pd.DataFrame:
149
+ """Generate a deterministic but realistic-looking OHLCV series.
150
+
151
+ This is not geometric Brownian motion with a straight face: it uses a
152
+ two-state (calm / stressed) regime switch, Student-t innovations and
153
+ GARCH-ish vol persistence, so the resulting series has fat tails and
154
+ volatility clustering. That matters, because a strategy tested against
155
+ naive GBM looks far better than it deserves to.
156
+ """
157
+ profile = SIM_PROFILES.get(symbol.upper(), SimProfile(0.08, 0.25, 100.0))
158
+ rng = np.random.default_rng(_seed_for(symbol) if seed is None else seed)
159
+
160
+ freq = {"1d": "B", "1wk": "W-FRI", "1h": "h"}.get(interval, "B")
161
+ index = pd.date_range(start=start, end=end or pd.Timestamp.today().normalize(), freq=freq)
162
+ n = len(index)
163
+ if n < 50:
164
+ raise ValueError("Simulated range is too short to backtest")
165
+
166
+ ppy = 252 if freq in ("B", "h") else 52
167
+ mu = profile.drift / ppy
168
+ base_vol = profile.vol / np.sqrt(ppy)
169
+
170
+ # Regime chain: calm state is sticky, stressed state is short and violent.
171
+ p_calm_to_stress, p_stress_to_calm = 0.01, 0.06
172
+ regime = np.zeros(n, dtype=int)
173
+ for i in range(1, n):
174
+ flip = rng.random()
175
+ if regime[i - 1] == 0:
176
+ regime[i] = 1 if flip < p_calm_to_stress else 0
177
+ else:
178
+ regime[i] = 0 if flip < p_stress_to_calm else 1
179
+
180
+ # Persistent vol around a regime-dependent level.
181
+ vol = np.empty(n)
182
+ level = np.where(regime == 1, base_vol * 2.4, base_vol * 0.9)
183
+ vol[0] = level[0]
184
+ for i in range(1, n):
185
+ vol[i] = 0.92 * vol[i - 1] + 0.08 * level[i]
186
+
187
+ shocks = rng.standard_t(df=4, size=n) / np.sqrt(2.0) # unit-ish variance, fat tails
188
+ drift = np.where(regime == 1, mu - 3.0 * base_vol**2, mu)
189
+ log_ret = drift + vol * shocks
190
+ close = profile.price * np.exp(np.cumsum(log_ret))
191
+ close = close * (profile.price / close[-1]) # end near the quoted level
192
+
193
+ intrabar = vol * rng.uniform(0.3, 1.1, size=n)
194
+ open_ = close * np.exp(-log_ret * rng.uniform(0.2, 0.8, size=n))
195
+ high = np.maximum(open_, close) * np.exp(np.abs(intrabar))
196
+ low = np.minimum(open_, close) * np.exp(-np.abs(intrabar))
197
+ volume = rng.lognormal(mean=15.5, sigma=0.45, size=n) * (1.0 + 3.0 * regime)
198
+
199
+ return _normalise(
200
+ pd.DataFrame(
201
+ {"open": open_, "high": high, "low": low, "close": close, "volume": volume},
202
+ index=index,
203
+ )
204
+ )
205
+
206
+
207
+ def load_ohlcv(
208
+ symbol: str = "SPY",
209
+ start: str = "2015-01-01",
210
+ end: str | None = None,
211
+ interval: str = "1d",
212
+ source: str = "auto",
213
+ ) -> MarketData:
214
+ """Load OHLCV for ``symbol``, never raising for a merely-unreachable network.
215
+
216
+ ``source`` is one of ``auto`` (download, then cache, then simulate),
217
+ ``cache``, or ``synthetic``.
218
+ """
219
+ symbol = (symbol or "SPY").strip().upper()
220
+
221
+ if source == "synthetic":
222
+ df = simulate_ohlcv(symbol, start, end, interval)
223
+ return MarketData(symbol, df, "synthetic", interval, "Simulated prices (requested).")
224
+
225
+ if source in ("auto", "live"):
226
+ df = _download(symbol, start, end, interval)
227
+ if df is not None and len(df) > 50:
228
+ _write_cache(symbol, interval, df)
229
+ return MarketData(symbol, df, "yfinance", interval, "Live data from Yahoo Finance.")
230
+
231
+ cached = _read_cache(symbol, interval)
232
+ if cached is not None and len(cached) > 50:
233
+ window = cached.loc[str(start) : str(end)] if end else cached.loc[str(start) :]
234
+ if len(window) > 50:
235
+ return MarketData(symbol, window, "bundled", interval, "Cached data (network unavailable).")
236
+
237
+ df = simulate_ohlcv(symbol, start, end, interval)
238
+ return MarketData(
239
+ symbol,
240
+ df,
241
+ "synthetic",
242
+ interval,
243
+ f"Live data for {symbol} was unavailable, so this run uses a deterministic "
244
+ "market simulator with fat tails and volatility clustering. The statistics "
245
+ "below are still valid — they are just measured on a simulated market.",
246
+ )
algotrader/engine.py ADDED
@@ -0,0 +1,133 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Vectorised, look-ahead-free backtest engine.
2
+
3
+ Contract
4
+ --------
5
+ A strategy emits ``target[t]``: the exposure it wants, decided using only
6
+ information available at the close of bar ``t``. The engine holds
7
+ ``position[t] = target[t - lag]`` during bar ``t`` and credits it with that
8
+ bar's close-to-close return. With the default ``lag=1`` this means "decide on
9
+ today's close, hold the position through tomorrow" -- the single place where
10
+ look-ahead could sneak in, and it is one line.
11
+
12
+ Costs are charged on exposure *changes*, so a strategy that flips daily pays
13
+ for it. Short exposure additionally accrues a borrow fee.
14
+ """
15
+
16
+ from __future__ import annotations
17
+
18
+ from typing import Optional
19
+
20
+ import numpy as np
21
+ import pandas as pd
22
+
23
+ from .metrics import compute_metrics, infer_periods_per_year
24
+ from .types import BacktestResult, CostModel
25
+
26
+ __all__ = ["run_backtest", "bars_to_returns"]
27
+
28
+
29
+ def bars_to_returns(df: pd.DataFrame) -> pd.Series:
30
+ """Close-to-close simple returns."""
31
+ return df["close"].astype(float).pct_change().fillna(0.0)
32
+
33
+
34
+ def run_backtest(
35
+ df: pd.DataFrame,
36
+ target: pd.Series,
37
+ costs: CostModel | None = None,
38
+ lag: int = 1,
39
+ max_leverage: float = 1.0,
40
+ allow_short: bool = True,
41
+ initial_capital: float = 100_000.0,
42
+ periods_per_year: Optional[int] = None,
43
+ rf: float = 0.0,
44
+ meta: Optional[dict] = None,
45
+ ) -> BacktestResult:
46
+ """Run one backtest and return equity, returns and the full metric bundle."""
47
+ if df.empty:
48
+ raise ValueError("Cannot backtest an empty price frame")
49
+ if lag < 1:
50
+ raise ValueError("lag must be >= 1; lag=0 would trade on unavailable information")
51
+
52
+ costs = costs or CostModel()
53
+ ppy = periods_per_year or infer_periods_per_year(df.index)
54
+
55
+ asset_ret = bars_to_returns(df)
56
+
57
+ target = target.reindex(df.index).astype(float).fillna(0.0)
58
+ lower = -max_leverage if allow_short else 0.0
59
+ target = target.clip(lower, max_leverage)
60
+
61
+ position = target.shift(lag).fillna(0.0)
62
+
63
+ gross = position * asset_ret
64
+
65
+ traded = position.diff()
66
+ traded.iloc[0] = position.iloc[0]
67
+ trade_cost = traded.abs() * (costs.one_way_bps / 1e4)
68
+
69
+ borrow_cost = position.clip(upper=0.0).abs() * (costs.short_borrow_bps / 1e4) / ppy
70
+ total_cost = trade_cost + borrow_cost
71
+
72
+ net = gross - total_cost
73
+ equity = initial_capital * (1.0 + net).cumprod()
74
+ benchmark_equity = initial_capital * (1.0 + asset_ret).cumprod()
75
+
76
+ result = BacktestResult(
77
+ equity=equity,
78
+ returns=net,
79
+ gross_returns=gross,
80
+ position=position,
81
+ target=target,
82
+ costs=total_cost,
83
+ benchmark_equity=benchmark_equity,
84
+ metrics=compute_metrics(net, equity, position, ppy, rf),
85
+ benchmark_metrics=compute_metrics(asset_ret, benchmark_equity, None, ppy, rf),
86
+ meta={
87
+ "lag": lag,
88
+ "commission_bps": costs.commission_bps,
89
+ "slippage_bps": costs.slippage_bps,
90
+ "short_borrow_bps": costs.short_borrow_bps,
91
+ "max_leverage": max_leverage,
92
+ "allow_short": allow_short,
93
+ "initial_capital": initial_capital,
94
+ "periods_per_year": ppy,
95
+ **(meta or {}),
96
+ },
97
+ )
98
+ result.metrics["cost_drag_ann"] = float(total_cost.sum() / max(result.metrics.get("years", 1e-9), 1e-9))
99
+ result.metrics["gross_sharpe"] = float(
100
+ compute_metrics(gross, initial_capital * (1.0 + gross).cumprod(), None, ppy, rf).get("sharpe", 0.0)
101
+ )
102
+ return result
103
+
104
+
105
+ def fast_sharpe(
106
+ asset_ret: np.ndarray,
107
+ target: np.ndarray,
108
+ one_way_bps: float,
109
+ lag: int,
110
+ periods_per_year: int,
111
+ ) -> float:
112
+ """Numpy-only Sharpe for hot loops (permutation tests, PBO grids).
113
+
114
+ Mirrors :func:`run_backtest` exactly for the no-borrow case; it exists only
115
+ because building a DataFrame 1000 times is the difference between a Space
116
+ that answers in 4 seconds and one nobody waits for.
117
+ """
118
+ n = asset_ret.size
119
+ position = np.empty(n, dtype=float)
120
+ position[:lag] = 0.0
121
+ position[lag:] = target[:-lag] if lag else target
122
+ gross = position * asset_ret
123
+ traded = np.empty(n, dtype=float)
124
+ traded[0] = position[0]
125
+ traded[1:] = np.diff(position)
126
+ net = gross - np.abs(traded) * (one_way_bps / 1e4)
127
+ net = net[np.isfinite(net)]
128
+ if net.size < 2:
129
+ return 0.0
130
+ sd = net.std(ddof=1)
131
+ if not np.isfinite(sd) or sd < 1e-12:
132
+ return 0.0
133
+ return float(net.mean() / sd * np.sqrt(periods_per_year))
algotrader/indicators.py ADDED
@@ -0,0 +1,102 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Vectorised technical indicators.
2
+
3
+ Every function takes and returns pandas objects aligned to the input index, and
4
+ every one of them is causal: the value at bar ``t`` uses only data up to and
5
+ including ``t``. That property is what makes the backtest engine's single
6
+ ``shift`` enough to guarantee no look-ahead.
7
+ """
8
+
9
+ from __future__ import annotations
10
+
11
+ import numpy as np
12
+ import pandas as pd
13
+
14
+ __all__ = [
15
+ "sma",
16
+ "ema",
17
+ "rsi",
18
+ "macd",
19
+ "bollinger",
20
+ "atr",
21
+ "donchian",
22
+ "zscore",
23
+ "roc",
24
+ "realised_vol",
25
+ ]
26
+
27
+
28
+ def sma(series: pd.Series, window: int) -> pd.Series:
29
+ return series.rolling(window, min_periods=window).mean()
30
+
31
+
32
+ def ema(series: pd.Series, window: int) -> pd.Series:
33
+ return series.ewm(span=window, adjust=False, min_periods=window).mean()
34
+
35
+
36
+ def rsi(series: pd.Series, window: int = 14) -> pd.Series:
37
+ """Wilder's RSI."""
38
+ delta = series.diff()
39
+ gain = delta.clip(lower=0.0)
40
+ loss = -delta.clip(upper=0.0)
41
+ avg_gain = gain.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
42
+ avg_loss = loss.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
43
+ rs = avg_gain / avg_loss.replace(0.0, np.nan)
44
+ out = 100.0 - (100.0 / (1.0 + rs))
45
+ # avg_loss == 0 leaves rs undefined: an all-gain window is RSI 100, and a
46
+ # perfectly flat window (no gains either) is RSI 50.
47
+ flat = (avg_gain == 0.0) & (avg_loss == 0.0)
48
+ out = out.mask((avg_loss == 0.0) & (avg_gain > 0.0), 100.0)
49
+ out = out.mask(flat, 50.0)
50
+ return out.where(avg_gain.notna())
51
+
52
+
53
+ def macd(
54
+ series: pd.Series, fast: int = 12, slow: int = 26, signal: int = 9
55
+ ) -> tuple[pd.Series, pd.Series, pd.Series]:
56
+ """Returns ``(macd_line, signal_line, histogram)``."""
57
+ macd_line = ema(series, fast) - ema(series, slow)
58
+ signal_line = macd_line.ewm(span=signal, adjust=False, min_periods=signal).mean()
59
+ return macd_line, signal_line, macd_line - signal_line
60
+
61
+
62
+ def bollinger(
63
+ series: pd.Series, window: int = 20, k: float = 2.0
64
+ ) -> tuple[pd.Series, pd.Series, pd.Series]:
65
+ """Returns ``(lower, middle, upper)``."""
66
+ mid = sma(series, window)
67
+ sd = series.rolling(window, min_periods=window).std(ddof=0)
68
+ return mid - k * sd, mid, mid + k * sd
69
+
70
+
71
+ def atr(df: pd.DataFrame, window: int = 14) -> pd.Series:
72
+ prev_close = df["close"].shift(1)
73
+ tr = pd.concat(
74
+ [
75
+ df["high"] - df["low"],
76
+ (df["high"] - prev_close).abs(),
77
+ (df["low"] - prev_close).abs(),
78
+ ],
79
+ axis=1,
80
+ ).max(axis=1)
81
+ return tr.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
82
+
83
+
84
+ def donchian(df: pd.DataFrame, window: int = 20) -> tuple[pd.Series, pd.Series]:
85
+ """Rolling channel excluding the current bar, so a breakout test is causal."""
86
+ upper = df["high"].rolling(window, min_periods=window).max().shift(1)
87
+ lower = df["low"].rolling(window, min_periods=window).min().shift(1)
88
+ return lower, upper
89
+
90
+
91
+ def zscore(series: pd.Series, window: int = 20) -> pd.Series:
92
+ mean = series.rolling(window, min_periods=window).mean()
93
+ sd = series.rolling(window, min_periods=window).std(ddof=0)
94
+ return (series - mean) / sd.replace(0.0, np.nan)
95
+
96
+
97
+ def roc(series: pd.Series, window: int = 20) -> pd.Series:
98
+ return series.pct_change(window)
99
+
100
+
101
+ def realised_vol(returns: pd.Series, window: int = 20, periods_per_year: int = 252) -> pd.Series:
102
+ return returns.rolling(window, min_periods=window).std(ddof=0) * np.sqrt(periods_per_year)
algotrader/lab.py ADDED
@@ -0,0 +1,328 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The Lab: one call that runs a backtest and then tries to disprove it.
2
+
3
+ This is the module both the Gradio Space and the CLI drive. Keeping the whole
4
+ pipeline here means the app and the command line can never disagree about what
5
+ a Reality Score means.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import logging
11
+ from dataclasses import dataclass, field
12
+ from typing import Callable, Dict, List, Optional
13
+
14
+ import numpy as np
15
+ import pandas as pd
16
+
17
+ from .data import load_ohlcv
18
+ from .engine import run_backtest
19
+ from .metrics import infer_periods_per_year
20
+ from .strategies import Strategy, get_strategy, list_strategies
21
+ from .types import BacktestResult, CostModel, MarketData
22
+ from .validation.deflated_sharpe import deflated_sharpe_ratio, min_track_record_length
23
+ from .validation.pbo import probability_of_backtest_overfitting
24
+ from .validation.permutation import PermutationResult, permutation_test
25
+ from .validation.walkforward import walk_forward
26
+ from .verdict import reality_score
27
+
28
+ logger = logging.getLogger(__name__)
29
+
30
+ __all__ = ["LabConfig", "LabReport", "run_lab", "run_arena"]
31
+
32
+ ProgressFn = Optional[Callable[[float, str], None]]
33
+
34
+
35
+ @dataclass
36
+ class LabConfig:
37
+ symbol: str = "SPY"
38
+ start: str = "2015-01-01"
39
+ end: Optional[str] = None
40
+ interval: str = "1d"
41
+ source: str = "auto"
42
+
43
+ strategy: str = "sma_cross"
44
+ params: Dict[str, float] = field(default_factory=dict)
45
+
46
+ commission_bps: float = 1.0
47
+ slippage_bps: float = 2.0
48
+ short_borrow_bps: float = 50.0
49
+ lag: int = 1
50
+ allow_short: bool = True
51
+ max_leverage: float = 1.0
52
+ capital: float = 100_000.0
53
+
54
+ n_permutations: int = 250
55
+ permutation_method: str = "permute"
56
+ block_size: int = 20
57
+ wf_folds: int = 5
58
+ pbo_splits: int = 8
59
+ grid_limit: int = 40
60
+ seed: int = 0
61
+
62
+ def costs(self, multiplier: float = 1.0) -> CostModel:
63
+ return CostModel(
64
+ commission_bps=self.commission_bps * multiplier,
65
+ slippage_bps=self.slippage_bps * multiplier,
66
+ short_borrow_bps=self.short_borrow_bps * multiplier,
67
+ )
68
+
69
+
70
+ @dataclass
71
+ class LabReport:
72
+ config: LabConfig
73
+ market: MarketData
74
+ strategy: Strategy
75
+ params: Dict[str, float]
76
+ backtest: BacktestResult
77
+ permutation: Optional[PermutationResult] = None
78
+ dsr: Dict[str, float] = field(default_factory=dict)
79
+ pbo: Dict[str, object] = field(default_factory=dict)
80
+ walkforward: Dict[str, object] = field(default_factory=dict)
81
+ trials: Dict[str, object] = field(default_factory=dict)
82
+ verdict: Dict[str, object] = field(default_factory=dict)
83
+ cost_stress: Dict[str, float] = field(default_factory=dict)
84
+ benchmark_correlation: float = float("nan")
85
+
86
+
87
+ def _trial_matrix(
88
+ df: pd.DataFrame,
89
+ strategy: Strategy,
90
+ cfg: LabConfig,
91
+ progress: ProgressFn = None,
92
+ ) -> tuple[np.ndarray, List[float], List[str]]:
93
+ """Backtest every parameter combination a researcher would plausibly try.
94
+
95
+ The resulting ``T x N`` return matrix feeds both the Deflated Sharpe (how
96
+ many variants were tried, and how spread out were they) and PBO.
97
+ """
98
+ grid = strategy.grid(limit=cfg.grid_limit)
99
+ costs = cfg.costs()
100
+ columns, sharpes, labels = [], [], []
101
+
102
+ for i, params in enumerate(grid):
103
+ target = strategy.generate(df, params)
104
+ result = run_backtest(
105
+ df, target, costs=costs, lag=cfg.lag,
106
+ max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
107
+ )
108
+ columns.append(result.returns.to_numpy(dtype=float))
109
+ sharpes.append(result.sharpe)
110
+ labels.append(", ".join(f"{k}={v}" for k, v in params.items()) or "default")
111
+ if progress is not None and i % 5 == 0:
112
+ progress((i + 1) / max(len(grid), 1), f"Variant {i + 1}/{len(grid)}")
113
+
114
+ matrix = np.column_stack(columns) if columns else np.zeros((len(df), 0))
115
+ return matrix, sharpes, labels
116
+
117
+
118
+ def run_lab(cfg: LabConfig, progress: ProgressFn = None) -> LabReport:
119
+ """Run the full honesty pipeline for one strategy on one symbol."""
120
+
121
+ def step(fraction: float, message: str) -> None:
122
+ if progress is not None:
123
+ progress(min(max(fraction, 0.0), 1.0), message)
124
+
125
+ step(0.02, "Loading market data")
126
+ market = load_ohlcv(cfg.symbol, cfg.start, cfg.end, cfg.interval, cfg.source)
127
+ df = market.df
128
+ if len(df) < 120:
129
+ raise ValueError(
130
+ f"Only {len(df)} bars available for {cfg.symbol}. "
131
+ "Widen the date range — anything shorter cannot be validated."
132
+ )
133
+
134
+ strategy = get_strategy(cfg.strategy)
135
+ params = strategy.clean(cfg.params)
136
+ ppy = infer_periods_per_year(df.index)
137
+
138
+ step(0.10, "Running the backtest")
139
+ target = strategy.generate(df, params)
140
+ backtest = run_backtest(
141
+ df,
142
+ target,
143
+ costs=cfg.costs(),
144
+ lag=cfg.lag,
145
+ max_leverage=cfg.max_leverage,
146
+ allow_short=cfg.allow_short,
147
+ initial_capital=cfg.capital,
148
+ periods_per_year=ppy,
149
+ meta={"symbol": market.symbol, "strategy": strategy.key, "params": params},
150
+ )
151
+
152
+ step(0.16, "Stress-testing costs")
153
+ stressed = run_backtest(
154
+ df, target, costs=cfg.costs(3.0), lag=cfg.lag,
155
+ max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
156
+ periods_per_year=ppy,
157
+ )
158
+ base_sharpe = backtest.sharpe
159
+ cost_stress_ratio = float(stressed.sharpe / base_sharpe) if base_sharpe > 1e-9 else 0.0
160
+ cost_stress = {
161
+ "sharpe_1x": base_sharpe,
162
+ "sharpe_3x": stressed.sharpe,
163
+ "ratio": cost_stress_ratio,
164
+ "return_3x": float(stressed.metrics.get("total_return", 0.0)),
165
+ }
166
+
167
+ step(0.22, "Backtesting every parameter variant")
168
+ matrix, trial_sharpes, labels = _trial_matrix(
169
+ df, strategy, cfg, lambda f, m: step(0.22 + 0.18 * f, m)
170
+ )
171
+ n_trials = max(len(trial_sharpes), 1)
172
+
173
+ step(0.42, "Deflating the Sharpe ratio for selection bias")
174
+ dsr = deflated_sharpe_ratio(
175
+ backtest.returns.to_numpy(dtype=float),
176
+ sharpe_annual=base_sharpe,
177
+ periods_per_year=ppy,
178
+ n_trials=n_trials,
179
+ trial_sharpes=trial_sharpes if n_trials > 1 else None,
180
+ )
181
+ mtrl = min_track_record_length(
182
+ dsr["sr_per_period"], dsr["n_obs"], dsr["skew"], dsr["kurtosis"],
183
+ benchmark=dsr["threshold_sr_per_period"],
184
+ )
185
+ dsr["min_track_record_bars"] = mtrl
186
+ dsr["min_track_record_years"] = float(mtrl / ppy) if np.isfinite(mtrl) else float("inf")
187
+
188
+ step(0.46, "Measuring backtest overfitting")
189
+ pbo = probability_of_backtest_overfitting(matrix, n_splits=cfg.pbo_splits, labels=labels)
190
+
191
+ step(0.50, "Shuffling the market")
192
+ permutation = None
193
+ if cfg.n_permutations > 0:
194
+ permutation = permutation_test(
195
+ df,
196
+ lambda frame: strategy.generate(frame, params),
197
+ n_permutations=cfg.n_permutations,
198
+ method=cfg.permutation_method,
199
+ block=cfg.block_size,
200
+ costs=cfg.costs(),
201
+ lag=cfg.lag,
202
+ max_leverage=cfg.max_leverage,
203
+ allow_short=cfg.allow_short,
204
+ seed=cfg.seed,
205
+ observed=base_sharpe,
206
+ progress=lambda f, m: step(0.50 + 0.32 * f, m),
207
+ )
208
+
209
+ step(0.84, "Walking the strategy forward")
210
+ wf = walk_forward(
211
+ df, strategy, n_folds=cfg.wf_folds, costs=cfg.costs(), lag=cfg.lag,
212
+ max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
213
+ grid_limit=min(cfg.grid_limit, 24),
214
+ progress=lambda f, m: step(0.84 + 0.12 * f, m),
215
+ )
216
+
217
+ bench_corr = float(
218
+ pd.Series(backtest.returns).corr(backtest.benchmark_equity.pct_change().fillna(0.0))
219
+ )
220
+
221
+ step(0.98, "Grading")
222
+ verdict = reality_score(
223
+ metrics=backtest.metrics,
224
+ benchmark_metrics=backtest.benchmark_metrics,
225
+ p_value=permutation.p_value if permutation else None,
226
+ dsr=dsr.get("dsr"),
227
+ pbo=pbo.get("pbo"),
228
+ wf_efficiency=wf.get("efficiency"),
229
+ wf_win_rate=wf.get("oos_win_rate"),
230
+ cost_stress_ratio=cost_stress_ratio,
231
+ benchmark_correlation=bench_corr,
232
+ )
233
+
234
+ step(1.0, "Done")
235
+ return LabReport(
236
+ config=cfg,
237
+ market=market,
238
+ strategy=strategy,
239
+ params=params,
240
+ backtest=backtest,
241
+ permutation=permutation,
242
+ dsr=dsr,
243
+ pbo=pbo,
244
+ walkforward=wf,
245
+ trials={"n": n_trials, "sharpes": trial_sharpes, "labels": labels, "matrix_shape": matrix.shape},
246
+ verdict=verdict,
247
+ cost_stress=cost_stress,
248
+ benchmark_correlation=bench_corr,
249
+ )
250
+
251
+
252
+ def run_arena(
253
+ cfg: LabConfig,
254
+ strategy_keys: Optional[List[str]] = None,
255
+ n_permutations: int = 120,
256
+ progress: ProgressFn = None,
257
+ ) -> tuple[pd.DataFrame, MarketData, Dict[str, BacktestResult]]:
258
+ """Race every strategy on the same market, ranked by evidence not returns.
259
+
260
+ Buy & hold and the coin flip stay in the field on purpose: a leaderboard
261
+ without a control group is marketing, not measurement.
262
+ """
263
+ market = load_ohlcv(cfg.symbol, cfg.start, cfg.end, cfg.interval, cfg.source)
264
+ df = market.df
265
+ ppy = infer_periods_per_year(df.index)
266
+ costs = cfg.costs()
267
+
268
+ keys = strategy_keys or [s.key for s in list_strategies()]
269
+ rows, curves = [], {}
270
+
271
+ for i, key in enumerate(keys):
272
+ strategy = get_strategy(key)
273
+ params = strategy.defaults()
274
+ target = strategy.generate(df, params)
275
+ result = run_backtest(
276
+ df, target, costs=costs, lag=cfg.lag, max_leverage=cfg.max_leverage,
277
+ allow_short=cfg.allow_short, initial_capital=cfg.capital, periods_per_year=ppy,
278
+ )
279
+ curves[key] = result
280
+
281
+ p_value = None
282
+ if n_permutations > 0:
283
+ p_value = permutation_test(
284
+ df,
285
+ lambda frame, s=strategy, p=params: s.generate(frame, p),
286
+ n_permutations=n_permutations,
287
+ method=cfg.permutation_method,
288
+ block=cfg.block_size,
289
+ costs=costs,
290
+ lag=cfg.lag,
291
+ max_leverage=cfg.max_leverage,
292
+ allow_short=cfg.allow_short,
293
+ seed=cfg.seed,
294
+ observed=result.sharpe,
295
+ ).p_value
296
+
297
+ grid_size = len(strategy.grid(limit=cfg.grid_limit))
298
+ dsr = deflated_sharpe_ratio(
299
+ result.returns.to_numpy(dtype=float),
300
+ sharpe_annual=result.sharpe,
301
+ periods_per_year=ppy,
302
+ n_trials=grid_size,
303
+ )
304
+
305
+ rows.append(
306
+ {
307
+ "Strategy": strategy.name,
308
+ "key": key,
309
+ "Family": strategy.family,
310
+ "Return": result.metrics.get("total_return", 0.0),
311
+ "CAGR": result.metrics.get("cagr", 0.0),
312
+ "Sharpe": result.sharpe,
313
+ "MaxDD": result.metrics.get("max_drawdown", 0.0),
314
+ "Trades": int(result.metrics.get("n_trades", 0)),
315
+ "p-value": p_value if p_value is not None else float("nan"),
316
+ "DSR": dsr["dsr"],
317
+ }
318
+ )
319
+ if progress is not None:
320
+ progress((i + 1) / len(keys), f"{strategy.name} ({i + 1}/{len(keys)})")
321
+
322
+ table = pd.DataFrame(rows)
323
+ if not table.empty:
324
+ # Rank by evidence: a high Sharpe with a p-value of 0.4 is not a win.
325
+ table["Evidence"] = (1.0 - table["p-value"].fillna(0.5)) * table["DSR"]
326
+ table = table.sort_values("Evidence", ascending=False).reset_index(drop=True)
327
+ table.insert(0, "#", table.index + 1)
328
+ return table, market, curves
algotrader/metrics.py ADDED
@@ -0,0 +1,157 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Performance and risk metrics.
2
+
3
+ All ratios are computed from *net* per-bar returns and annualised with the
4
+ periodicity inferred from the index, so daily / hourly / minute series all get
5
+ comparable numbers.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ from typing import Dict
11
+
12
+ import numpy as np
13
+ import pandas as pd
14
+
15
+ __all__ = [
16
+ "infer_periods_per_year",
17
+ "sharpe_ratio",
18
+ "sortino_ratio",
19
+ "max_drawdown",
20
+ "drawdown_series",
21
+ "compute_metrics",
22
+ ]
23
+
24
+ _SECONDS_PER_YEAR = 365.25 * 24 * 3600
25
+ _TRADING_DAYS = 252
26
+
27
+ # A return stream with dispersion below this is constant to floating-point
28
+ # noise. Without an absolute floor, a flat series divides by ~1e-19 and reports
29
+ # a Sharpe of 1e16 -- the exact kind of nonsense number this project exists to
30
+ # catch, so it must not originate here.
31
+ _DEGENERATE_SD = 1e-12
32
+
33
+
34
+ def infer_periods_per_year(index: pd.Index) -> int:
35
+ """Guess bars-per-year from an index, defaulting to daily trading bars."""
36
+ if not isinstance(index, pd.DatetimeIndex) or len(index) < 3:
37
+ return _TRADING_DAYS
38
+ nanos = index.to_numpy(dtype="datetime64[ns]").astype("int64")
39
+ deltas = np.diff(nanos) / 1e9 # seconds
40
+ deltas = deltas[deltas > 0]
41
+ if deltas.size == 0:
42
+ return _TRADING_DAYS
43
+ step = float(np.median(deltas))
44
+ if step >= 20 * 3600: # daily or slower -> use trading-day convention
45
+ days = step / 86400.0
46
+ return max(1, int(round(_TRADING_DAYS / max(days / 1.4, 1.0))))
47
+ # Intraday: assume a 6.5h session, 252 days a year.
48
+ bars_per_session = (6.5 * 3600) / step
49
+ return max(1, int(round(bars_per_session * _TRADING_DAYS)))
50
+
51
+
52
+ def _clean(returns: pd.Series) -> np.ndarray:
53
+ arr = np.asarray(returns, dtype=float)
54
+ return arr[np.isfinite(arr)]
55
+
56
+
57
+ def sharpe_ratio(returns: pd.Series, periods_per_year: int, rf: float = 0.0) -> float:
58
+ """Annualised Sharpe. ``rf`` is an annual risk-free rate."""
59
+ arr = _clean(returns)
60
+ if arr.size < 2:
61
+ return 0.0
62
+ excess = arr - rf / periods_per_year
63
+ sd = excess.std(ddof=1)
64
+ if not np.isfinite(sd) or sd < _DEGENERATE_SD:
65
+ return 0.0
66
+ return float(excess.mean() / sd * np.sqrt(periods_per_year))
67
+
68
+
69
+ def sortino_ratio(returns: pd.Series, periods_per_year: int, rf: float = 0.0) -> float:
70
+ arr = _clean(returns)
71
+ if arr.size < 2:
72
+ return 0.0
73
+ excess = arr - rf / periods_per_year
74
+ downside = excess[excess < 0]
75
+ if downside.size == 0:
76
+ return float("inf") if excess.mean() > 0 else 0.0
77
+ dd = np.sqrt(np.mean(downside**2))
78
+ if not np.isfinite(dd) or dd < _DEGENERATE_SD:
79
+ return 0.0
80
+ return float(excess.mean() / dd * np.sqrt(periods_per_year))
81
+
82
+
83
+ def drawdown_series(equity: pd.Series) -> pd.Series:
84
+ peak = equity.cummax()
85
+ return equity / peak - 1.0
86
+
87
+
88
+ def max_drawdown(equity: pd.Series) -> float:
89
+ if equity.empty:
90
+ return 0.0
91
+ return float(drawdown_series(equity).min())
92
+
93
+
94
+ def _time_under_water(equity: pd.Series, periods_per_year: int) -> float:
95
+ """Longest stretch below a prior peak, in years."""
96
+ if equity.empty:
97
+ return 0.0
98
+ dd = drawdown_series(equity).to_numpy()
99
+ longest = current = 0
100
+ for value in dd:
101
+ current = current + 1 if value < 0 else 0
102
+ longest = max(longest, current)
103
+ return longest / periods_per_year
104
+
105
+
106
+ def compute_metrics(
107
+ returns: pd.Series,
108
+ equity: pd.Series,
109
+ position: pd.Series | None = None,
110
+ periods_per_year: int | None = None,
111
+ rf: float = 0.0,
112
+ ) -> Dict[str, float]:
113
+ """Full metric bundle for one equity curve."""
114
+ ppy = periods_per_year or infer_periods_per_year(returns.index)
115
+ arr = _clean(returns)
116
+ n = arr.size
117
+ if n == 0 or equity.empty:
118
+ return {"periods_per_year": float(ppy)}
119
+
120
+ years = n / ppy
121
+ total_return = float(equity.iloc[-1] / equity.iloc[0] - 1.0)
122
+ cagr = float((equity.iloc[-1] / equity.iloc[0]) ** (1.0 / years) - 1.0) if years > 0 else 0.0
123
+ vol = float(arr.std(ddof=1) * np.sqrt(ppy))
124
+ mdd = max_drawdown(equity)
125
+ sr = sharpe_ratio(returns, ppy, rf)
126
+
127
+ out: Dict[str, float] = {
128
+ "total_return": total_return,
129
+ "cagr": cagr,
130
+ "ann_vol": vol,
131
+ "sharpe": sr,
132
+ "sortino": sortino_ratio(returns, ppy, rf),
133
+ "calmar": float(cagr / abs(mdd)) if mdd < 0 else 0.0,
134
+ "max_drawdown": mdd,
135
+ "time_under_water_yrs": _time_under_water(equity, ppy),
136
+ "hit_rate": float((arr > 0).mean()),
137
+ "skew": float(pd.Series(arr).skew()) if n > 2 else 0.0,
138
+ "kurtosis": float(pd.Series(arr).kurtosis()) if n > 3 else 0.0,
139
+ "var_95": float(np.percentile(arr, 5)),
140
+ "cvar_95": float(arr[arr <= np.percentile(arr, 5)].mean()) if n > 20 else 0.0,
141
+ "best_bar": float(arr.max()),
142
+ "worst_bar": float(arr.min()),
143
+ "n_bars": float(n),
144
+ "years": float(years),
145
+ "periods_per_year": float(ppy),
146
+ }
147
+
148
+ if position is not None and not position.empty:
149
+ pos = position.fillna(0.0)
150
+ turnover = pos.diff().abs().fillna(pos.abs().iloc[0] if len(pos) else 0.0)
151
+ out["exposure"] = float(pos.abs().mean())
152
+ out["long_share"] = float((pos > 0).mean())
153
+ out["short_share"] = float((pos < 0).mean())
154
+ out["turnover_ann"] = float(turnover.sum() / years) if years > 0 else 0.0
155
+ # A "trade" is any change in sign or size of exposure.
156
+ out["n_trades"] = float((turnover > 1e-9).sum())
157
+ return out
algotrader/strategies.py ADDED
@@ -0,0 +1,345 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The strategy zoo.
2
+
3
+ Every strategy is a pure function ``(df, **params) -> target exposure series``
4
+ in ``[-1, 1]``, causal by construction. They never see costs, capital or
5
+ execution — that is the engine's job — which is what lets the same function be
6
+ re-run thousands of times inside the permutation and PBO machinery.
7
+ """
8
+
9
+ from __future__ import annotations
10
+
11
+ import itertools
12
+ from dataclasses import dataclass
13
+ from typing import Callable, Dict, Iterable, List, Sequence
14
+
15
+ import numpy as np
16
+ import pandas as pd
17
+
18
+ from . import indicators as ind
19
+
20
+ __all__ = ["Strategy", "ParamSpec", "REGISTRY", "get_strategy", "list_strategies"]
21
+
22
+
23
+ @dataclass(frozen=True)
24
+ class ParamSpec:
25
+ name: str
26
+ label: str
27
+ default: float
28
+ grid: Sequence[float]
29
+ kind: str = "int"
30
+ minimum: float | None = None
31
+ maximum: float | None = None
32
+ step: float | None = None
33
+
34
+ def cast(self, value):
35
+ return int(value) if self.kind == "int" else float(value)
36
+
37
+
38
+ @dataclass(frozen=True)
39
+ class Strategy:
40
+ key: str
41
+ name: str
42
+ family: str
43
+ description: str
44
+ fn: Callable[..., pd.Series]
45
+ params: tuple[ParamSpec, ...] = ()
46
+
47
+ def defaults(self) -> Dict[str, float]:
48
+ return {p.name: p.cast(p.default) for p in self.params}
49
+
50
+ def clean(self, params: Dict[str, float] | None) -> Dict[str, float]:
51
+ """Fill in missing params and coerce types, ignoring unknown keys."""
52
+ merged = self.defaults()
53
+ for spec in self.params:
54
+ if params and spec.name in params and params[spec.name] is not None:
55
+ merged[spec.name] = spec.cast(params[spec.name])
56
+ return merged
57
+
58
+ def generate(self, df: pd.DataFrame, params: Dict[str, float] | None = None) -> pd.Series:
59
+ target = self.fn(df, **self.clean(params))
60
+ return target.reindex(df.index).astype(float).fillna(0.0).clip(-1.0, 1.0)
61
+
62
+ def grid(self, limit: int | None = None) -> List[Dict[str, float]]:
63
+ """Cartesian product of the per-parameter grids (the 'trials' a
64
+ researcher would realistically run before picking a winner)."""
65
+ if not self.params:
66
+ return [{}]
67
+ names = [p.name for p in self.params]
68
+ combos = [
69
+ dict(zip(names, values))
70
+ for values in itertools.product(*[p.grid for p in self.params])
71
+ ]
72
+ combos = [c for c in combos if self._valid(c)]
73
+ if limit is not None and len(combos) > limit:
74
+ step = len(combos) / limit
75
+ combos = [combos[int(i * step)] for i in range(limit)]
76
+ return combos
77
+
78
+ def _valid(self, combo: Dict[str, float]) -> bool:
79
+ """Reject nonsensical combinations (a fast MA slower than the slow one)."""
80
+ if "fast" in combo and "slow" in combo and combo["fast"] >= combo["slow"]:
81
+ return False
82
+ if "lower" in combo and "upper" in combo and combo["lower"] >= combo["upper"]:
83
+ return False
84
+ return True
85
+
86
+
87
+ def _hold_until_flip(raw: pd.Series) -> pd.Series:
88
+ """Turn sparse entry/exit signals into a continuously held position."""
89
+ return raw.ffill().fillna(0.0)
90
+
91
+
92
+ # --------------------------------------------------------------------------
93
+ # Strategy implementations
94
+ # --------------------------------------------------------------------------
95
+
96
+ def _buy_and_hold(df: pd.DataFrame) -> pd.Series:
97
+ return pd.Series(1.0, index=df.index)
98
+
99
+
100
+ def _sma_cross(df: pd.DataFrame, fast: int = 20, slow: int = 100) -> pd.Series:
101
+ f, s = ind.sma(df["close"], fast), ind.sma(df["close"], slow)
102
+ return pd.Series(np.where(f > s, 1.0, -1.0), index=df.index).where(s.notna())
103
+
104
+
105
+ def _ema_cross(df: pd.DataFrame, fast: int = 12, slow: int = 50) -> pd.Series:
106
+ f, s = ind.ema(df["close"], fast), ind.ema(df["close"], slow)
107
+ return pd.Series(np.where(f > s, 1.0, -1.0), index=df.index).where(s.notna())
108
+
109
+
110
+ def _macd_trend(df: pd.DataFrame, fast: int = 12, slow: int = 26, signal: int = 9) -> pd.Series:
111
+ _, _, hist = ind.macd(df["close"], fast, slow, signal)
112
+ return pd.Series(np.sign(hist), index=df.index).where(hist.notna())
113
+
114
+
115
+ def _rsi_reversion(df: pd.DataFrame, window: int = 14, lower: int = 30, upper: int = 70) -> pd.Series:
116
+ r = ind.rsi(df["close"], window)
117
+ raw = pd.Series(np.nan, index=df.index)
118
+ raw[r < lower] = 1.0
119
+ raw[r > upper] = -1.0
120
+ raw[(r > 45) & (r < 55)] = 0.0 # flatten in the middle of the range
121
+ return _hold_until_flip(raw).where(r.notna())
122
+
123
+
124
+ def _bollinger_reversion(df: pd.DataFrame, window: int = 20, k: float = 2.0) -> pd.Series:
125
+ low, mid, high = ind.bollinger(df["close"], window, k)
126
+ close = df["close"]
127
+ raw = pd.Series(np.nan, index=df.index)
128
+ raw[close < low] = 1.0
129
+ raw[close > high] = -1.0
130
+ raw[(close - mid).abs() < 0.1 * (high - mid)] = 0.0
131
+ return _hold_until_flip(raw).where(mid.notna())
132
+
133
+
134
+ def _donchian_breakout(df: pd.DataFrame, window: int = 20) -> pd.Series:
135
+ low, high = ind.donchian(df, window)
136
+ raw = pd.Series(np.nan, index=df.index)
137
+ raw[df["close"] > high] = 1.0
138
+ raw[df["close"] < low] = -1.0
139
+ return _hold_until_flip(raw).where(high.notna())
140
+
141
+
142
+ def _momentum(df: pd.DataFrame, lookback: int = 60) -> pd.Series:
143
+ return np.sign(ind.roc(df["close"], lookback))
144
+
145
+
146
+ def _vol_target_momentum(
147
+ df: pd.DataFrame, lookback: int = 60, vol_window: int = 20, target_vol: float = 15
148
+ ) -> pd.Series:
149
+ """Momentum sized inversely to recent volatility (targets ``target_vol`` %)."""
150
+ signal = np.sign(ind.roc(df["close"], lookback))
151
+ rv = ind.realised_vol(df["close"].pct_change(), vol_window)
152
+ scale = (target_vol / 100.0) / rv.replace(0.0, np.nan)
153
+ return (signal * scale.clip(upper=1.0)).where(rv.notna())
154
+
155
+
156
+ def _channel_trend(df: pd.DataFrame, window: int = 50, atr_window: int = 14, mult: float = 1.0) -> pd.Series:
157
+ """Long above an ATR band around the mean, short below it, flat inside."""
158
+ mid = ind.sma(df["close"], window)
159
+ band = ind.atr(df, atr_window) * mult
160
+ raw = pd.Series(np.nan, index=df.index)
161
+ raw[df["close"] > mid + band] = 1.0
162
+ raw[df["close"] < mid - band] = -1.0
163
+ raw[(df["close"] - mid).abs() < 0.25 * band] = 0.0
164
+ return _hold_until_flip(raw).where(mid.notna() & band.notna())
165
+
166
+
167
+ def _coin_flip(df: pd.DataFrame, hold: int = 5, seed: int = 7) -> pd.Series:
168
+ """A deliberately worthless strategy: the control group.
169
+
170
+ If your clever rule cannot beat this on the validation panel, that is the
171
+ single most useful thing this app can tell you.
172
+ """
173
+ rng = np.random.default_rng(int(seed))
174
+ n = len(df)
175
+ draws = rng.choice([-1.0, 1.0], size=int(np.ceil(n / max(hold, 1))))
176
+ return pd.Series(np.repeat(draws, max(hold, 1))[:n], index=df.index)
177
+
178
+
179
+ REGISTRY: Dict[str, Strategy] = {}
180
+
181
+
182
+ def _register(strategy: Strategy) -> Strategy:
183
+ REGISTRY[strategy.key] = strategy
184
+ return strategy
185
+
186
+
187
+ _register(
188
+ Strategy(
189
+ key="buy_and_hold",
190
+ name="Buy & Hold",
191
+ family="benchmark",
192
+ description="Own the asset, do nothing. The bar every other strategy has to clear.",
193
+ fn=_buy_and_hold,
194
+ )
195
+ )
196
+
197
+ _register(
198
+ Strategy(
199
+ key="sma_cross",
200
+ name="SMA Crossover",
201
+ family="trend",
202
+ description="Long when the fast simple moving average is above the slow one, short when below.",
203
+ fn=_sma_cross,
204
+ params=(
205
+ ParamSpec("fast", "Fast MA", 20, (5, 10, 20, 30, 50), "int", 2, 100, 1),
206
+ ParamSpec("slow", "Slow MA", 100, (50, 100, 150, 200), "int", 10, 300, 5),
207
+ ),
208
+ )
209
+ )
210
+
211
+ _register(
212
+ Strategy(
213
+ key="ema_cross",
214
+ name="EMA Crossover",
215
+ family="trend",
216
+ description="Same idea as the SMA cross but with exponential averages, so it turns faster.",
217
+ fn=_ema_cross,
218
+ params=(
219
+ ParamSpec("fast", "Fast EMA", 12, (5, 8, 12, 21, 34), "int", 2, 100, 1),
220
+ ParamSpec("slow", "Slow EMA", 50, (34, 50, 89, 144, 200), "int", 10, 300, 1),
221
+ ),
222
+ )
223
+ )
224
+
225
+ _register(
226
+ Strategy(
227
+ key="macd_trend",
228
+ name="MACD Trend",
229
+ family="trend",
230
+ description="Follow the sign of the MACD histogram.",
231
+ fn=_macd_trend,
232
+ params=(
233
+ ParamSpec("fast", "Fast", 12, (8, 12, 16), "int", 2, 60, 1),
234
+ ParamSpec("slow", "Slow", 26, (21, 26, 34, 50), "int", 10, 200, 1),
235
+ ParamSpec("signal", "Signal", 9, (5, 9, 13), "int", 2, 50, 1),
236
+ ),
237
+ )
238
+ )
239
+
240
+ _register(
241
+ Strategy(
242
+ key="rsi_reversion",
243
+ name="RSI Mean Reversion",
244
+ family="mean-reversion",
245
+ description="Buy oversold, sell overbought, flatten in the middle of the range.",
246
+ fn=_rsi_reversion,
247
+ params=(
248
+ ParamSpec("window", "RSI window", 14, (7, 14, 21), "int", 2, 60, 1),
249
+ ParamSpec("lower", "Oversold", 30, (20, 25, 30, 35), "int", 5, 49, 1),
250
+ ParamSpec("upper", "Overbought", 70, (65, 70, 75, 80), "int", 51, 95, 1),
251
+ ),
252
+ )
253
+ )
254
+
255
+ _register(
256
+ Strategy(
257
+ key="bollinger_reversion",
258
+ name="Bollinger Reversion",
259
+ family="mean-reversion",
260
+ description="Fade moves outside the Bollinger bands and exit back at the middle band.",
261
+ fn=_bollinger_reversion,
262
+ params=(
263
+ ParamSpec("window", "Window", 20, (10, 20, 30, 50), "int", 5, 120, 1),
264
+ ParamSpec("k", "Band width (σ)", 2.0, (1.5, 2.0, 2.5, 3.0), "float", 0.5, 4.0, 0.1),
265
+ ),
266
+ )
267
+ )
268
+
269
+ _register(
270
+ Strategy(
271
+ key="donchian_breakout",
272
+ name="Donchian Breakout",
273
+ family="breakout",
274
+ description="The classic turtle rule: buy new highs, sell new lows.",
275
+ fn=_donchian_breakout,
276
+ params=(ParamSpec("window", "Channel", 20, (10, 20, 40, 55, 100), "int", 5, 250, 1),),
277
+ )
278
+ )
279
+
280
+ _register(
281
+ Strategy(
282
+ key="momentum",
283
+ name="Time-Series Momentum",
284
+ family="momentum",
285
+ description="Hold long if the asset is up over the lookback, short if it is down.",
286
+ fn=_momentum,
287
+ params=(ParamSpec("lookback", "Lookback", 60, (5, 10, 20, 60, 120, 250), "int", 2, 500, 1),),
288
+ )
289
+ )
290
+
291
+ _register(
292
+ Strategy(
293
+ key="vol_target_momentum",
294
+ name="Vol-Targeted Momentum",
295
+ family="momentum",
296
+ description="Momentum sized down when markets get volatile, so risk stays roughly constant.",
297
+ fn=_vol_target_momentum,
298
+ params=(
299
+ ParamSpec("lookback", "Lookback", 60, (20, 60, 120, 250), "int", 5, 500, 1),
300
+ ParamSpec("vol_window", "Vol window", 20, (10, 20, 60), "int", 5, 120, 1),
301
+ ParamSpec("target_vol", "Target vol %", 15, (10, 15, 20), "float", 2, 60, 1),
302
+ ),
303
+ )
304
+ )
305
+
306
+ _register(
307
+ Strategy(
308
+ key="channel_trend",
309
+ name="ATR Channel Trend",
310
+ family="trend",
311
+ description="Trade with the trend only once price clears an ATR band around its mean.",
312
+ fn=_channel_trend,
313
+ params=(
314
+ ParamSpec("window", "Mean window", 50, (20, 50, 100, 200), "int", 5, 300, 1),
315
+ ParamSpec("atr_window", "ATR window", 14, (7, 14, 28), "int", 2, 60, 1),
316
+ ParamSpec("mult", "ATR multiple", 1.0, (0.5, 1.0, 1.5, 2.0), "float", 0.1, 5.0, 0.1),
317
+ ),
318
+ )
319
+ )
320
+
321
+ _register(
322
+ Strategy(
323
+ key="coin_flip",
324
+ name="Coin Flip (control)",
325
+ family="control",
326
+ description="Random positions. The control group — anything that cannot beat this is noise.",
327
+ fn=_coin_flip,
328
+ params=(
329
+ ParamSpec("hold", "Bars per flip", 5, (1, 5, 10, 20), "int", 1, 60, 1),
330
+ ParamSpec("seed", "Seed", 7, (1, 7, 42, 123), "int", 0, 9999, 1),
331
+ ),
332
+ )
333
+ )
334
+
335
+
336
+ def get_strategy(key: str) -> Strategy:
337
+ try:
338
+ return REGISTRY[key]
339
+ except KeyError:
340
+ raise KeyError(f"Unknown strategy '{key}'. Available: {', '.join(sorted(REGISTRY))}") from None
341
+
342
+
343
+ def list_strategies(exclude: Iterable[str] = ()) -> List[Strategy]:
344
+ skip = set(exclude)
345
+ return [s for k, s in REGISTRY.items() if k not in skip]
algotrader/types.py ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Core data types shared across the algotrader 2.0 stack.
2
+
3
+ Everything downstream (engine, validation, UI) speaks these types, so they are
4
+ deliberately small, immutable-ish and free of framework dependencies.
5
+ """
6
+
7
+ from __future__ import annotations
8
+
9
+ from dataclasses import dataclass, field
10
+ from typing import Any, Dict, Optional
11
+
12
+ import pandas as pd
13
+
14
+ OHLCV_COLUMNS = ("open", "high", "low", "close", "volume")
15
+
16
+
17
+ @dataclass(frozen=True)
18
+ class MarketData:
19
+ """A validated OHLCV series plus provenance.
20
+
21
+ Provenance matters here: the app is about honesty, so the UI always tells
22
+ the user whether they are looking at real prices or a simulation.
23
+ """
24
+
25
+ symbol: str
26
+ df: pd.DataFrame
27
+ source: str # "yfinance" | "bundled" | "synthetic"
28
+ interval: str = "1d"
29
+ note: str = ""
30
+
31
+ @property
32
+ def is_real(self) -> bool:
33
+ return self.source in ("yfinance", "bundled")
34
+
35
+ @property
36
+ def start(self) -> pd.Timestamp:
37
+ return self.df.index[0]
38
+
39
+ @property
40
+ def end(self) -> pd.Timestamp:
41
+ return self.df.index[-1]
42
+
43
+ def __len__(self) -> int: # pragma: no cover - trivial
44
+ return len(self.df)
45
+
46
+
47
+ @dataclass(frozen=True)
48
+ class CostModel:
49
+ """Round-trip friction. All values are one-way, in basis points."""
50
+
51
+ commission_bps: float = 1.0
52
+ slippage_bps: float = 2.0
53
+ short_borrow_bps: float = 50.0 # annualised, charged on short exposure
54
+
55
+ @property
56
+ def one_way_bps(self) -> float:
57
+ return self.commission_bps + self.slippage_bps
58
+
59
+
60
+ @dataclass
61
+ class BacktestResult:
62
+ """Output of a single backtest run."""
63
+
64
+ equity: pd.Series
65
+ returns: pd.Series # net of costs
66
+ gross_returns: pd.Series
67
+ position: pd.Series # exposure actually held during each bar
68
+ target: pd.Series # exposure requested by the strategy
69
+ costs: pd.Series
70
+ benchmark_equity: pd.Series
71
+ metrics: Dict[str, float] = field(default_factory=dict)
72
+ benchmark_metrics: Dict[str, float] = field(default_factory=dict)
73
+ meta: Dict[str, Any] = field(default_factory=dict)
74
+
75
+ @property
76
+ def sharpe(self) -> float:
77
+ return float(self.metrics.get("sharpe", 0.0))
78
+
79
+ @property
80
+ def n_trades(self) -> int:
81
+ return int(self.metrics.get("n_trades", 0))
82
+
83
+
84
+ @dataclass
85
+ class ValidationReport:
86
+ """Everything we know about how much of a backtest is luck."""
87
+
88
+ permutation_p_value: Optional[float] = None
89
+ permutation_null: Optional[Any] = None # np.ndarray of null Sharpes
90
+ deflated_sharpe: Optional[float] = None
91
+ probabilistic_sharpe: Optional[float] = None
92
+ min_track_record_years: Optional[float] = None
93
+ n_trials: int = 1
94
+ pbo: Optional[float] = None
95
+ pbo_detail: Dict[str, Any] = field(default_factory=dict)
96
+ walkforward: Dict[str, Any] = field(default_factory=dict)
97
+ reality_score: float = 0.0
98
+ grade: str = "?"
99
+ verdict: str = ""
100
+ flags: list = field(default_factory=list)
algotrader/validation/__init__.py ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Statistical tests that ask "is this edge real, or did we just get lucky?"."""
2
+
3
+ from .deflated_sharpe import deflated_sharpe_ratio, min_track_record_length, probabilistic_sharpe_ratio
4
+ from .pbo import probability_of_backtest_overfitting
5
+ from .permutation import permutation_test
6
+ from .walkforward import walk_forward
7
+
8
+ __all__ = [
9
+ "deflated_sharpe_ratio",
10
+ "probabilistic_sharpe_ratio",
11
+ "min_track_record_length",
12
+ "probability_of_backtest_overfitting",
13
+ "permutation_test",
14
+ "walk_forward",
15
+ ]
algotrader/validation/deflated_sharpe.py ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Probabilistic and Deflated Sharpe Ratios.
2
+
3
+ Bailey & López de Prado (2014), "The Deflated Sharpe Ratio: Correcting for
4
+ Selection Bias, Backtest Overfitting and Non-Normality".
5
+
6
+ The intuition: if you try 200 strategy variants, the best one will show a
7
+ handsome Sharpe *even when none of them has any edge*. The Deflated Sharpe
8
+ Ratio asks whether the winner beats what the luckiest of 200 coin-flippers
9
+ would have produced, and it charges extra for fat tails and negative skew --
10
+ exactly the return shapes that make naive Sharpe ratios flatter.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import numpy as np
16
+ from scipy import stats
17
+
18
+ __all__ = [
19
+ "probabilistic_sharpe_ratio",
20
+ "expected_max_sharpe",
21
+ "deflated_sharpe_ratio",
22
+ "min_track_record_length",
23
+ ]
24
+
25
+ _EULER = 0.5772156649015329
26
+
27
+
28
+ def _moments(returns: np.ndarray) -> tuple[float, float]:
29
+ """Sample skew and *non-excess* kurtosis, as the PSR formula expects."""
30
+ arr = np.asarray(returns, dtype=float)
31
+ arr = arr[np.isfinite(arr)]
32
+ if arr.size < 4:
33
+ return 0.0, 3.0
34
+ return float(stats.skew(arr, bias=False)), float(stats.kurtosis(arr, bias=False) + 3.0)
35
+
36
+
37
+ def probabilistic_sharpe_ratio(
38
+ sharpe: float,
39
+ n_obs: int,
40
+ skew: float = 0.0,
41
+ kurtosis: float = 3.0,
42
+ benchmark: float = 0.0,
43
+ ) -> float:
44
+ """P(true Sharpe > ``benchmark``) given the observed Sharpe and its shape.
45
+
46
+ ``sharpe`` and ``benchmark`` are per-observation (i.e. *not* annualised).
47
+ """
48
+ if n_obs < 3:
49
+ return 0.5
50
+ denom = 1.0 - skew * sharpe + ((kurtosis - 1.0) / 4.0) * sharpe**2
51
+ if denom <= 0:
52
+ return 0.5
53
+ z = (sharpe - benchmark) * np.sqrt(n_obs - 1) / np.sqrt(denom)
54
+ return float(stats.norm.cdf(z))
55
+
56
+
57
+ def expected_max_sharpe(n_trials: int, variance_of_trials: float) -> float:
58
+ """Expected maximum Sharpe across ``n_trials`` *skill-free* strategies.
59
+
60
+ This is the bar the winner has to clear to be interesting. It grows with
61
+ the number of things you tried — which is why "I found a strategy with
62
+ Sharpe 2" means nothing until you say how many you looked at.
63
+ """
64
+ n = max(int(n_trials), 1)
65
+ if n == 1 or variance_of_trials <= 0:
66
+ return 0.0
67
+ sd = np.sqrt(variance_of_trials)
68
+ # Bailey & López de Prado's Gumbel-based approximation.
69
+ q1 = stats.norm.ppf(1.0 - 1.0 / n)
70
+ q2 = stats.norm.ppf(1.0 - 1.0 / (n * np.e))
71
+ return float(sd * ((1.0 - _EULER) * q1 + _EULER * q2))
72
+
73
+
74
+ def deflated_sharpe_ratio(
75
+ returns,
76
+ sharpe_annual: float,
77
+ periods_per_year: int,
78
+ n_trials: int,
79
+ trial_sharpes=None,
80
+ variance_of_trials: float | None = None,
81
+ ) -> dict:
82
+ """Deflate an annualised Sharpe for selection bias and non-normality.
83
+
84
+ Returns a dict with the PSR against a zero benchmark, the selection-bias
85
+ threshold, the deflated probability, and the inputs used, so the UI can
86
+ show its working rather than just a number.
87
+ """
88
+ arr = np.asarray(returns, dtype=float)
89
+ arr = arr[np.isfinite(arr)]
90
+ n_obs = arr.size
91
+ sr_per_period = sharpe_annual / np.sqrt(periods_per_year)
92
+ skew, kurt = _moments(arr)
93
+
94
+ if variance_of_trials is None:
95
+ if trial_sharpes is not None and len(trial_sharpes) > 1:
96
+ trials = np.asarray(trial_sharpes, dtype=float) / np.sqrt(periods_per_year)
97
+ trials = trials[np.isfinite(trials)]
98
+ variance_of_trials = float(np.var(trials, ddof=1)) if trials.size > 1 else 0.0
99
+ else:
100
+ # With no trial cloud to measure, fall back to the asymptotic
101
+ # variance of a skill-free Sharpe estimate.
102
+ variance_of_trials = 1.0 / max(n_obs - 1, 1)
103
+
104
+ threshold = expected_max_sharpe(n_trials, variance_of_trials)
105
+
106
+ psr = probabilistic_sharpe_ratio(sr_per_period, n_obs, skew, kurt, 0.0)
107
+ dsr = probabilistic_sharpe_ratio(sr_per_period, n_obs, skew, kurt, threshold)
108
+
109
+ return {
110
+ "psr": float(psr),
111
+ "dsr": float(dsr),
112
+ "sr_per_period": float(sr_per_period),
113
+ "threshold_sr_per_period": float(threshold),
114
+ "threshold_sr_annual": float(threshold * np.sqrt(periods_per_year)),
115
+ "n_obs": int(n_obs),
116
+ "n_trials": int(n_trials),
117
+ "skew": float(skew),
118
+ "kurtosis": float(kurt),
119
+ "variance_of_trials": float(variance_of_trials),
120
+ }
121
+
122
+
123
+ def min_track_record_length(
124
+ sharpe: float,
125
+ n_obs: int,
126
+ skew: float = 0.0,
127
+ kurtosis: float = 3.0,
128
+ benchmark: float = 0.0,
129
+ confidence: float = 0.95,
130
+ ) -> float:
131
+ """Observations needed before the Sharpe is significant at ``confidence``.
132
+
133
+ Inputs are per-observation. Returns ``inf`` when the edge is too small to
134
+ ever clear the bar.
135
+ """
136
+ if sharpe <= benchmark:
137
+ return float("inf")
138
+ z = stats.norm.ppf(confidence)
139
+ denom = (sharpe - benchmark) ** 2
140
+ if denom <= 0:
141
+ return float("inf")
142
+ numer = 1.0 - skew * sharpe + ((kurtosis - 1.0) / 4.0) * sharpe**2
143
+ if numer <= 0:
144
+ return float("inf")
145
+ return float(1.0 + numer * (z / (sharpe - benchmark)) ** 2)
algotrader/validation/pbo.py ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Probability of Backtest Overfitting via Combinatorially Symmetric CV.
2
+
3
+ Bailey, Borwein, López de Prado & Zhu (2016), "The Probability of Backtest
4
+ Overfitting".
5
+
6
+ Take the N parameter variants you tried, cut the timeline into S chunks, and
7
+ for every way of splitting those chunks half-and-half: pick the variant that
8
+ won in-sample, then look up where it ranked out-of-sample. If your selection
9
+ process has skill, the in-sample winner should keep winning. If it is fitting
10
+ noise, the winner lands in the bottom half about as often as not — and PBO
11
+ approaches 0.5.
12
+ """
13
+
14
+ from __future__ import annotations
15
+
16
+ import itertools
17
+ from typing import Dict, Sequence
18
+
19
+ import numpy as np
20
+
21
+ __all__ = ["probability_of_backtest_overfitting"]
22
+
23
+
24
+ def _sharpe_columns(matrix: np.ndarray) -> np.ndarray:
25
+ """Per-column Sharpe (per-period, unannualised — ranks are all we need)."""
26
+ if matrix.shape[0] < 2:
27
+ return np.zeros(matrix.shape[1])
28
+ mean = matrix.mean(axis=0)
29
+ sd = matrix.std(axis=0, ddof=1)
30
+ # Absolute floor, not `> 0`: a flat column's std is float noise, and
31
+ # dividing by it would hand a do-nothing variant an enormous rank.
32
+ with np.errstate(divide="ignore", invalid="ignore"):
33
+ out = np.where(sd > 1e-12, mean / sd, 0.0)
34
+ return np.nan_to_num(out, nan=0.0, posinf=0.0, neginf=0.0)
35
+
36
+
37
+ def probability_of_backtest_overfitting(
38
+ returns_matrix: np.ndarray,
39
+ n_splits: int = 8,
40
+ labels: Sequence[str] | None = None,
41
+ max_combinations: int = 200,
42
+ ) -> Dict[str, object]:
43
+ """Compute PBO for a ``T x N`` matrix of per-period strategy returns.
44
+
45
+ ``n_splits`` must be even. Returns PBO, the logit distribution, the
46
+ in-sample/out-of-sample Sharpe pairs for the selected variants, and the
47
+ rate at which the selected variant actually loses money out of sample.
48
+ """
49
+ matrix = np.asarray(returns_matrix, dtype=float)
50
+ if matrix.ndim != 2:
51
+ raise ValueError("returns_matrix must be 2-D (time x strategy)")
52
+ matrix = np.nan_to_num(matrix, nan=0.0, posinf=0.0, neginf=0.0)
53
+ t_obs, n_strats = matrix.shape
54
+
55
+ if n_strats < 2 or t_obs < 2 * n_splits:
56
+ return {
57
+ "pbo": float("nan"),
58
+ "n_strategies": int(n_strats),
59
+ "n_combinations": 0,
60
+ "logits": np.array([]),
61
+ "is_sharpes": np.array([]),
62
+ "oos_sharpes": np.array([]),
63
+ "prob_oos_loss": float("nan"),
64
+ "performance_degradation": float("nan"),
65
+ "note": "Not enough variants or observations to estimate PBO.",
66
+ }
67
+
68
+ if n_splits % 2:
69
+ n_splits += 1
70
+ chunks = np.array_split(np.arange(t_obs), n_splits)
71
+
72
+ combos = list(itertools.combinations(range(n_splits), n_splits // 2))
73
+ if len(combos) > max_combinations:
74
+ step = len(combos) / max_combinations
75
+ combos = [combos[int(i * step)] for i in range(max_combinations)]
76
+
77
+ logits, is_sr, oos_sr, chosen = [], [], [], []
78
+ for combo in combos:
79
+ is_idx = np.concatenate([chunks[c] for c in combo])
80
+ oos_idx = np.concatenate([chunks[c] for c in range(n_splits) if c not in combo])
81
+
82
+ is_perf = _sharpe_columns(matrix[is_idx])
83
+ oos_perf = _sharpe_columns(matrix[oos_idx])
84
+
85
+ best = int(np.argmax(is_perf))
86
+ chosen.append(best)
87
+ is_sr.append(float(is_perf[best]))
88
+ oos_sr.append(float(oos_perf[best]))
89
+
90
+ # Relative rank of the chosen variant in the OOS ranking, in (0, 1).
91
+ rank = float(np.sum(oos_perf <= oos_perf[best]))
92
+ omega = rank / (n_strats + 1.0)
93
+ omega = min(max(omega, 1e-6), 1.0 - 1e-6)
94
+ logits.append(float(np.log(omega / (1.0 - omega))))
95
+
96
+ logits_arr = np.asarray(logits)
97
+ is_arr, oos_arr = np.asarray(is_sr), np.asarray(oos_sr)
98
+
99
+ # Slope of OOS on IS: negative means better in-sample fits do *worse* live.
100
+ degradation = float("nan")
101
+ if is_arr.size > 2 and np.std(is_arr) > 1e-12:
102
+ degradation = float(np.polyfit(is_arr, oos_arr, 1)[0])
103
+
104
+ counts = np.bincount(chosen, minlength=n_strats)
105
+ most_selected = int(np.argmax(counts))
106
+
107
+ return {
108
+ "pbo": float(np.mean(logits_arr <= 0.0)),
109
+ "n_strategies": int(n_strats),
110
+ "n_combinations": int(len(combos)),
111
+ "logits": logits_arr,
112
+ "is_sharpes": is_arr,
113
+ "oos_sharpes": oos_arr,
114
+ "prob_oos_loss": float(np.mean(oos_arr <= 0.0)),
115
+ "performance_degradation": degradation,
116
+ "most_selected_index": most_selected,
117
+ "most_selected_label": (labels[most_selected] if labels is not None else str(most_selected)),
118
+ "selection_stability": float(counts[most_selected] / max(len(combos), 1)),
119
+ "note": "",
120
+ }
algotrader/validation/permutation.py ADDED
@@ -0,0 +1,187 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Monte-Carlo permutation test for trading rules.
2
+
3
+ The question this answers is not "did the strategy make money" but "would a
4
+ rule of this shape have made this much money on a market with no exploitable
5
+ structure?". We destroy the serial dependence in the price path while keeping
6
+ its distribution of moves intact, re-run the *same* strategy on each shuffled
7
+ market, and see where the real result lands in that null distribution.
8
+
9
+ A strategy whose Sharpe sits comfortably inside the null is not a strategy —
10
+ it is a lottery ticket that happened to win.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ from dataclasses import dataclass
16
+ from typing import Callable, Optional
17
+
18
+ import numpy as np
19
+ import pandas as pd
20
+
21
+ from ..engine import bars_to_returns, run_backtest
22
+ from ..types import CostModel
23
+
24
+ __all__ = ["permutation_test", "PermutationResult", "permute_bars"]
25
+
26
+
27
+ @dataclass
28
+ class PermutationResult:
29
+ observed: float
30
+ null: np.ndarray
31
+ p_value: float
32
+ method: str
33
+ n_permutations: int
34
+
35
+ @property
36
+ def null_mean(self) -> float:
37
+ return float(np.mean(self.null)) if self.null.size else 0.0
38
+
39
+ @property
40
+ def percentile(self) -> float:
41
+ """Where the observed Sharpe sits in the null distribution, 0-100."""
42
+ if not self.null.size:
43
+ return 50.0
44
+ return float((self.null < self.observed).mean() * 100.0)
45
+
46
+
47
+ def _decompose(df: pd.DataFrame) -> tuple[np.ndarray, float]:
48
+ """Split bars into scale-free log moves that can be reshuffled safely."""
49
+ open_ = df["open"].to_numpy(dtype=float)
50
+ high = df["high"].to_numpy(dtype=float)
51
+ low = df["low"].to_numpy(dtype=float)
52
+ close = df["close"].to_numpy(dtype=float)
53
+ volume = df["volume"].to_numpy(dtype=float)
54
+
55
+ gap = np.log(open_[1:] / close[:-1])
56
+ hi = np.log(np.maximum(high[1:], open_[1:]) / open_[1:])
57
+ lo = np.log(np.minimum(low[1:], open_[1:]) / open_[1:])
58
+ body = np.log(close[1:] / open_[1:])
59
+ return np.column_stack([gap, hi, lo, body, volume[1:]]), float(close[0])
60
+
61
+
62
+ def _rebuild(parts: np.ndarray, anchor: float, index: pd.Index, first_row: pd.Series) -> pd.DataFrame:
63
+ gap, hi, lo, body, volume = (parts[:, i] for i in range(5))
64
+ n = parts.shape[0] + 1
65
+
66
+ close = np.empty(n)
67
+ open_ = np.empty(n)
68
+ high = np.empty(n)
69
+ low = np.empty(n)
70
+ vol = np.empty(n)
71
+
72
+ close[0] = anchor
73
+ open_[0] = float(first_row["open"])
74
+ high[0] = float(first_row["high"])
75
+ low[0] = float(first_row["low"])
76
+ vol[0] = float(first_row["volume"])
77
+
78
+ # Cumulative product form: close[i] = close[0] * exp(cumsum(gap + body)).
79
+ close[1:] = anchor * np.exp(np.cumsum(gap + body))
80
+ open_[1:] = close[:-1] * np.exp(gap)
81
+ high[1:] = open_[1:] * np.exp(hi)
82
+ low[1:] = open_[1:] * np.exp(lo)
83
+ vol[1:] = volume
84
+
85
+ return pd.DataFrame(
86
+ {"open": open_, "high": high, "low": low, "close": close, "volume": vol}, index=index
87
+ )
88
+
89
+
90
+ def permute_bars(
91
+ df: pd.DataFrame,
92
+ rng: np.random.Generator,
93
+ method: str = "permute",
94
+ block: int = 20,
95
+ ) -> pd.DataFrame:
96
+ """Return a shuffled market with the same index and bar anatomy.
97
+
98
+ ``permute`` reshuffles individual bars, destroying all serial structure.
99
+ ``block`` resamples contiguous blocks with replacement, which preserves
100
+ short-horizon autocorrelation and volatility clustering — a harder null
101
+ that trend strategies deserve to be tested against.
102
+ """
103
+ parts, anchor = _decompose(df)
104
+ m = parts.shape[0]
105
+ if m < 2:
106
+ return df.copy()
107
+
108
+ if method == "block":
109
+ size = max(2, min(int(block), m))
110
+ starts = rng.integers(0, m, size=int(np.ceil(m / size)))
111
+ order = np.concatenate([(np.arange(s, s + size) % m) for s in starts])[:m]
112
+ else:
113
+ order = rng.permutation(m)
114
+
115
+ return _rebuild(parts[order], anchor, df.index, df.iloc[0])
116
+
117
+
118
+ def permutation_test(
119
+ df: pd.DataFrame,
120
+ signal_fn: Callable[[pd.DataFrame], pd.Series],
121
+ n_permutations: int = 300,
122
+ method: str = "permute",
123
+ block: int = 20,
124
+ costs: Optional[CostModel] = None,
125
+ lag: int = 1,
126
+ max_leverage: float = 1.0,
127
+ allow_short: bool = True,
128
+ seed: int = 0,
129
+ observed: Optional[float] = None,
130
+ progress: Optional[Callable[[float, str], None]] = None,
131
+ ) -> PermutationResult:
132
+ """Run ``signal_fn`` against ``n_permutations`` shuffled markets.
133
+
134
+ ``signal_fn`` must be the strategy's target-exposure generator; it is
135
+ re-evaluated on every synthetic market, which is the whole point — a rule
136
+ that only works because of the specific path it was tuned on will fall
137
+ apart here.
138
+ """
139
+ costs = costs or CostModel()
140
+ rng = np.random.default_rng(seed)
141
+
142
+ def sharpe_on(frame: pd.DataFrame) -> float:
143
+ target = signal_fn(frame)
144
+ result = run_backtest(
145
+ frame,
146
+ target,
147
+ costs=costs,
148
+ lag=lag,
149
+ max_leverage=max_leverage,
150
+ allow_short=allow_short,
151
+ )
152
+ return result.sharpe
153
+
154
+ if observed is None:
155
+ observed = sharpe_on(df)
156
+
157
+ null = np.empty(n_permutations, dtype=float)
158
+ for i in range(n_permutations):
159
+ null[i] = sharpe_on(permute_bars(df, rng, method, block))
160
+ if progress is not None and (i % 25 == 0 or i == n_permutations - 1):
161
+ progress((i + 1) / n_permutations, f"Permutation {i + 1}/{n_permutations}")
162
+
163
+ # +1 in both places: the observed result is itself one draw from the null
164
+ # under H0, which keeps the test from ever reporting an impossible p = 0.
165
+ p_value = float((1 + np.sum(null >= observed)) / (n_permutations + 1))
166
+
167
+ return PermutationResult(
168
+ observed=float(observed),
169
+ null=null,
170
+ p_value=p_value,
171
+ method=method,
172
+ n_permutations=n_permutations,
173
+ )
174
+
175
+
176
+ def bootstrap_return_paths(returns: pd.Series, n: int = 500, seed: int = 0) -> np.ndarray:
177
+ """Bootstrap terminal-wealth outcomes from a realised return stream.
178
+
179
+ Useful for the "how wide is the cone of outcomes?" chart — the same edge
180
+ can produce wildly different equity curves.
181
+ """
182
+ arr = np.asarray(returns.dropna(), dtype=float)
183
+ if arr.size == 0:
184
+ return np.zeros((n, 1))
185
+ rng = np.random.default_rng(seed)
186
+ draws = rng.choice(arr, size=(n, arr.size), replace=True)
187
+ return np.cumprod(1.0 + draws, axis=1)
algotrader/validation/walkforward.py ADDED
@@ -0,0 +1,129 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Walk-forward analysis.
2
+
3
+ Re-tune on a training window, trade the next window blind, roll forward. The
4
+ gap between in-sample and out-of-sample Sharpe is the honest estimate of how
5
+ much of the backtest was curve-fitting.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ from typing import Callable, Dict, List, Optional
11
+
12
+ import numpy as np
13
+ import pandas as pd
14
+
15
+ from ..engine import run_backtest
16
+ from ..strategies import Strategy
17
+ from ..types import CostModel
18
+
19
+ __all__ = ["walk_forward"]
20
+
21
+
22
+ def walk_forward(
23
+ df: pd.DataFrame,
24
+ strategy: Strategy,
25
+ n_folds: int = 5,
26
+ train_ratio: float = 0.7,
27
+ costs: Optional[CostModel] = None,
28
+ lag: int = 1,
29
+ max_leverage: float = 1.0,
30
+ allow_short: bool = True,
31
+ grid_limit: int = 40,
32
+ progress: Optional[Callable[[float, str], None]] = None,
33
+ ) -> Dict[str, object]:
34
+ """Roll a train/test split forward ``n_folds`` times.
35
+
36
+ Each fold picks the best parameters by in-sample Sharpe and reports what
37
+ those parameters then did out of sample. The stitched OOS returns are the
38
+ closest thing to a paper-trading record this repo can produce offline.
39
+ """
40
+ costs = costs or CostModel()
41
+ grid = strategy.grid(limit=grid_limit)
42
+ n = len(df)
43
+
44
+ if n < 250 or n_folds < 2:
45
+ return {"folds": [], "note": "Not enough history for walk-forward analysis."}
46
+
47
+ # Each fold is a contiguous train+test block; blocks advance by test length.
48
+ block = int(n / (1 + (n_folds - 1) * (1 - train_ratio)))
49
+ block = min(block, n)
50
+ train_len = int(block * train_ratio)
51
+ test_len = block - train_len
52
+ if train_len < 100 or test_len < 20:
53
+ return {"folds": [], "note": "Not enough history for walk-forward analysis."}
54
+
55
+ folds: List[Dict[str, object]] = []
56
+ oos_returns: List[pd.Series] = []
57
+
58
+ for k in range(n_folds):
59
+ start = k * test_len
60
+ train = df.iloc[start : start + train_len]
61
+ test = df.iloc[start + train_len : start + train_len + test_len]
62
+ if len(test) < 20:
63
+ break
64
+
65
+ best_params, best_sharpe = None, -np.inf
66
+ for params in grid:
67
+ target = strategy.generate(train, params)
68
+ sr = run_backtest(
69
+ train, target, costs=costs, lag=lag,
70
+ max_leverage=max_leverage, allow_short=allow_short,
71
+ ).sharpe
72
+ if sr > best_sharpe:
73
+ best_params, best_sharpe = params, sr
74
+
75
+ # Generate signals on train+test so indicators are warm at the fold
76
+ # boundary, then evaluate only the test slice.
77
+ combined = df.iloc[start : start + train_len + len(test)]
78
+ target = strategy.generate(combined, best_params).loc[test.index]
79
+ oos = run_backtest(
80
+ test, target, costs=costs, lag=lag,
81
+ max_leverage=max_leverage, allow_short=allow_short,
82
+ )
83
+
84
+ folds.append(
85
+ {
86
+ "fold": k + 1,
87
+ "train_start": str(train.index[0].date()),
88
+ "train_end": str(train.index[-1].date()),
89
+ "test_start": str(test.index[0].date()),
90
+ "test_end": str(test.index[-1].date()),
91
+ "params": best_params,
92
+ "is_sharpe": float(best_sharpe),
93
+ "oos_sharpe": float(oos.sharpe),
94
+ "oos_return": float(oos.metrics.get("total_return", 0.0)),
95
+ "oos_max_dd": float(oos.metrics.get("max_drawdown", 0.0)),
96
+ }
97
+ )
98
+ oos_returns.append(oos.returns)
99
+
100
+ if progress is not None:
101
+ progress((k + 1) / n_folds, f"Walk-forward fold {k + 1}/{n_folds}")
102
+
103
+ if not folds:
104
+ return {"folds": [], "note": "Not enough history for walk-forward analysis."}
105
+
106
+ is_sharpes = np.array([f["is_sharpe"] for f in folds], dtype=float)
107
+ oos_sharpes = np.array([f["oos_sharpe"] for f in folds], dtype=float)
108
+ stitched = pd.concat(oos_returns) if oos_returns else pd.Series(dtype=float)
109
+ stitched = stitched[~stitched.index.duplicated(keep="first")].sort_index()
110
+
111
+ mean_is = float(np.mean(is_sharpes))
112
+ mean_oos = float(np.mean(oos_sharpes))
113
+
114
+ return {
115
+ "folds": folds,
116
+ "mean_is_sharpe": mean_is,
117
+ "mean_oos_sharpe": mean_oos,
118
+ # 1.0 = the edge fully survived; 0.0 = it evaporated out of sample.
119
+ "efficiency": float(mean_oos / mean_is) if mean_is > 1e-9 else 0.0,
120
+ "oos_win_rate": float(np.mean(oos_sharpes > 0)),
121
+ # 1.0 means every fold chose different parameters -- a tuning process
122
+ # that cannot make up its mind is fitting noise.
123
+ "param_instability": float(
124
+ len({str(f["params"]) for f in folds}) / max(len(folds), 1)
125
+ ),
126
+ "oos_returns": stitched,
127
+ "oos_equity": (1.0 + stitched).cumprod() if len(stitched) else stitched,
128
+ "note": "",
129
+ }
algotrader/verdict.py ADDED
@@ -0,0 +1,177 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Turn a pile of statistics into one number and one sentence.
2
+
3
+ The Reality Score is deliberately harsh. Most published backtests would score
4
+ below 40, and that is the point: the score exists to be screenshotted.
5
+ """
6
+
7
+ from __future__ import annotations
8
+
9
+ from typing import Dict, List, Optional
10
+
11
+ import numpy as np
12
+
13
+ __all__ = ["reality_score", "GRADES"]
14
+
15
+ GRADES = [
16
+ (85, "A", "Survives everything we threw at it"),
17
+ (70, "B", "Probably a real edge, with caveats"),
18
+ (55, "C", "Ambiguous — could go either way"),
19
+ (40, "D", "Mostly luck"),
20
+ (0, "F", "Indistinguishable from randomness"),
21
+ ]
22
+
23
+ WEIGHTS = {
24
+ "significance": 0.30,
25
+ "selection": 0.25,
26
+ "walk_forward": 0.20,
27
+ "overfitting": 0.15,
28
+ "robustness": 0.10,
29
+ }
30
+
31
+
32
+ def _ramp(value: float, good: float, bad: float) -> float:
33
+ """Linear 0-100 score where ``good`` maps to 100 and ``bad`` maps to 0."""
34
+ if not np.isfinite(value):
35
+ return 50.0
36
+ if good == bad:
37
+ return 50.0
38
+ scaled = (value - bad) / (good - bad)
39
+ return float(np.clip(scaled, 0.0, 1.0) * 100.0)
40
+
41
+
42
+ def reality_score(
43
+ metrics: Dict[str, float],
44
+ benchmark_metrics: Dict[str, float],
45
+ p_value: Optional[float] = None,
46
+ dsr: Optional[float] = None,
47
+ pbo: Optional[float] = None,
48
+ wf_efficiency: Optional[float] = None,
49
+ wf_win_rate: Optional[float] = None,
50
+ cost_stress_ratio: Optional[float] = None,
51
+ benchmark_correlation: Optional[float] = None,
52
+ ) -> Dict[str, object]:
53
+ """Combine the validation panel into a 0-100 score, a grade and warnings."""
54
+ components: Dict[str, float] = {}
55
+
56
+ components["significance"] = _ramp(p_value, good=0.01, bad=0.50) if p_value is not None else 50.0
57
+ components["selection"] = float(np.clip(dsr, 0.0, 1.0) * 100.0) if dsr is not None else 50.0
58
+
59
+ if wf_efficiency is not None:
60
+ wf = _ramp(wf_efficiency, good=0.8, bad=-0.2)
61
+ if wf_win_rate is not None:
62
+ wf = 0.7 * wf + 0.3 * float(np.clip(wf_win_rate, 0.0, 1.0) * 100.0)
63
+ components["walk_forward"] = wf
64
+ else:
65
+ components["walk_forward"] = 50.0
66
+
67
+ components["overfitting"] = _ramp(pbo, good=0.05, bad=0.50) if pbo is not None and np.isfinite(pbo) else 50.0
68
+ components["robustness"] = (
69
+ _ramp(cost_stress_ratio, good=0.8, bad=0.0) if cost_stress_ratio is not None else 50.0
70
+ )
71
+
72
+ score = float(sum(components[k] * w for k, w in WEIGHTS.items()))
73
+
74
+ flags: List[str] = []
75
+ n_trades = int(metrics.get("n_trades", 0))
76
+ sharpe = float(metrics.get("sharpe", 0.0))
77
+ total_return = float(metrics.get("total_return", 0.0))
78
+ max_dd = float(metrics.get("max_drawdown", 0.0))
79
+ bench_sharpe = float(benchmark_metrics.get("sharpe", 0.0))
80
+
81
+ # The score asks "is the measured edge real?" — if the strategy lost money,
82
+ # there is no edge to validate, whatever the statistics say about it.
83
+ if total_return <= 0:
84
+ flags.append(
85
+ f"The strategy lost money over the test period ({total_return:.1%}). "
86
+ "There is no edge here to validate."
87
+ )
88
+ score = min(score, 50.0)
89
+ if total_return <= 0 < sharpe:
90
+ flags.append(
91
+ "Positive Sharpe with a negative total return: the average bar was profitable but "
92
+ "the compounding was not. Volatility drag ate the arithmetic edge."
93
+ )
94
+ if max_dd < -0.5:
95
+ flags.append(
96
+ f"Peak-to-trough drawdown of {max_dd:.0%} — an account running this would have been "
97
+ "closed long before the recovery arrived."
98
+ )
99
+ score = min(score, 60.0)
100
+
101
+ if n_trades < 20:
102
+ flags.append(
103
+ f"Only {n_trades} position changes — with this few decisions, "
104
+ "the result is a handful of coin flips, not a track record."
105
+ )
106
+ score = min(score, 55.0)
107
+ if p_value is not None and p_value > 0.10:
108
+ flags.append(
109
+ f"Permutation p-value is {p_value:.2f}: roughly {p_value * 100:.0f}% of shuffled, "
110
+ "structure-free markets did this well or better."
111
+ )
112
+ if dsr is not None and dsr < 0.5:
113
+ flags.append(
114
+ f"Deflated Sharpe is {dsr:.2f} — once you account for how many variants were tried, "
115
+ "the edge does not clear the selection-bias bar."
116
+ )
117
+ if pbo is not None and np.isfinite(pbo) and pbo > 0.3:
118
+ flags.append(
119
+ f"Probability of backtest overfitting is {pbo:.0%}: the in-sample winner "
120
+ "usually lands in the bottom half out of sample."
121
+ )
122
+ if wf_efficiency is not None and wf_efficiency < 0.3:
123
+ flags.append(
124
+ f"Walk-forward efficiency is {wf_efficiency:.0%} — most of the in-sample Sharpe "
125
+ "does not survive re-tuning and trading forward."
126
+ )
127
+ if cost_stress_ratio is not None and cost_stress_ratio < 0.5:
128
+ flags.append(
129
+ "Tripling trading costs removes more than half the Sharpe. The edge is "
130
+ "smaller than the friction it has to pay."
131
+ )
132
+ if benchmark_correlation is not None and benchmark_correlation > 0.95:
133
+ flags.append(
134
+ f"Returns are {benchmark_correlation:.0%} correlated with buy & hold — "
135
+ "this is mostly a repackaged long position."
136
+ )
137
+ if sharpe < bench_sharpe:
138
+ flags.append(
139
+ f"Buy & hold beat it on risk-adjusted return ({bench_sharpe:.2f} vs {sharpe:.2f} Sharpe)."
140
+ )
141
+ if float(metrics.get("turnover_ann", 0.0)) > 100:
142
+ flags.append(
143
+ f"Annual turnover of {metrics.get('turnover_ann', 0):.0f}x is far beyond what "
144
+ "retail execution can absorb without moving the modelled fills."
145
+ )
146
+
147
+ grade, headline = next((g, h) for threshold, g, h in GRADES if score >= threshold)
148
+
149
+ if score >= 70:
150
+ verdict = (
151
+ f"Grade {grade}. {headline}. The edge is still there after shuffling the market, "
152
+ "after charging for every variant tried, and after walking it forward."
153
+ )
154
+ elif score >= 55:
155
+ verdict = (
156
+ f"Grade {grade}. {headline}. Parts of the panel hold up and parts do not — "
157
+ "this is the zone where more data, not more tuning, is what settles it."
158
+ )
159
+ elif score >= 40:
160
+ verdict = (
161
+ f"Grade {grade}. {headline}. The backtest looks better than the evidence supports; "
162
+ "the gap between the two is selection bias."
163
+ )
164
+ else:
165
+ verdict = (
166
+ f"Grade {grade}. {headline}. A randomly shuffled market produces results like this "
167
+ "often enough that there is nothing here to trade."
168
+ )
169
+
170
+ return {
171
+ "score": round(score, 1),
172
+ "grade": grade,
173
+ "headline": headline,
174
+ "verdict": verdict,
175
+ "components": {k: round(v, 1) for k, v in components.items()},
176
+ "flags": flags,
177
+ }
app.py ADDED
@@ -0,0 +1,556 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Backtest Reality Check — the Hugging Face Space entrypoint.
2
+
3
+ Point it at a ticker and a trading rule. It runs the backtest, then spends the
4
+ rest of its time trying to prove the result was luck: shuffled markets,
5
+ selection-bias deflation, walk-forward, and a cost stress test.
6
+
7
+ Run locally with ``python app.py``.
8
+ """
9
+
10
+ from __future__ import annotations
11
+
12
+ import logging
13
+ import os
14
+ from typing import Dict, List
15
+
16
+ import gradio as gr
17
+ import pandas as pd
18
+
19
+ from algotrader import __version__
20
+ from algotrader.charts import (
21
+ arena_chart,
22
+ drawdown_chart,
23
+ empty_figure,
24
+ equity_chart,
25
+ exposure_chart,
26
+ permutation_chart,
27
+ score_chart,
28
+ walkforward_chart,
29
+ )
30
+ from algotrader.data import DEFAULT_UNIVERSE
31
+ from algotrader.lab import LabConfig, run_arena, run_lab
32
+ from algotrader.strategies import REGISTRY, get_strategy
33
+
34
+ logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s: %(message)s")
35
+ logger = logging.getLogger("app")
36
+
37
+ MAX_PARAMS = 3
38
+ STRATEGY_CHOICES = [(s.name, key) for key, s in REGISTRY.items()]
39
+
40
+ GRADE_COLORS = {
41
+ "A": "#0ca30c",
42
+ "B": "#3987e5",
43
+ "C": "#fab219",
44
+ "D": "#ec835a",
45
+ "F": "#d03b3b",
46
+ }
47
+
48
+ CSS = """
49
+ .gradio-container { max-width: 1280px !important; }
50
+ #hero h1 { font-size: 2.4rem; line-height: 1.1; margin: 0 0 .4rem 0; letter-spacing: -.02em; }
51
+ #hero p { color: #c3c2b7; margin: 0; font-size: 1.05rem; max-width: 60ch; }
52
+ .score-card {
53
+ display: flex; gap: 24px; align-items: center; padding: 22px 24px;
54
+ border: 1px solid rgba(255,255,255,.10); border-radius: 14px; background: #1a1a19;
55
+ }
56
+ .score-badge {
57
+ min-width: 132px; text-align: center; padding: 14px 10px; border-radius: 12px;
58
+ background: #0d0d0d; border: 1px solid rgba(255,255,255,.10);
59
+ }
60
+ .score-badge .grade { font-size: 3.2rem; font-weight: 700; line-height: 1; }
61
+ .score-badge .num { font-size: .95rem; color: #898781; margin-top: 6px; }
62
+ .score-body h3 { margin: 0 0 6px 0; font-size: 1.15rem; color: #fff; }
63
+ .score-body p { margin: 0; color: #c3c2b7; line-height: 1.55; }
64
+ .tiles { display: grid; grid-template-columns: repeat(auto-fit, minmax(132px,1fr)); gap: 10px; margin-top: 14px; }
65
+ .tile { padding: 12px 14px; border: 1px solid rgba(255,255,255,.10); border-radius: 10px; background: #1a1a19; }
66
+ .tile .label { font-size: .72rem; text-transform: uppercase; letter-spacing: .06em; color: #898781; }
67
+ .tile .value { font-size: 1.5rem; font-weight: 600; color: #fff; margin-top: 4px; }
68
+ .tile .sub { font-size: .75rem; color: #898781; margin-top: 2px; }
69
+ .flags { margin-top: 14px; padding: 0; list-style: none; }
70
+ .flags li {
71
+ padding: 9px 12px; margin-bottom: 7px; border-radius: 8px; background: #1a1a19;
72
+ border-left: 3px solid #ec835a; color: #c3c2b7; font-size: .9rem; line-height: 1.5;
73
+ }
74
+ .provenance { font-size: .82rem; color: #898781; margin-top: 10px; }
75
+ .provenance.sim { color: #fab219; }
76
+ """
77
+
78
+
79
+ def _fmt_pct(value: float) -> str:
80
+ return f"{value * 100:,.1f}%"
81
+
82
+
83
+ def param_controls(strategy_key: str):
84
+ """Re-label the shared sliders to match the selected strategy."""
85
+ strategy = get_strategy(strategy_key)
86
+ updates = []
87
+ for i in range(MAX_PARAMS):
88
+ if i < len(strategy.params):
89
+ spec = strategy.params[i]
90
+ lo = spec.minimum if spec.minimum is not None else min(spec.grid)
91
+ hi = spec.maximum if spec.maximum is not None else max(spec.grid)
92
+ updates.append(
93
+ gr.update(
94
+ visible=True,
95
+ label=spec.label,
96
+ value=spec.cast(spec.default),
97
+ minimum=lo,
98
+ maximum=hi,
99
+ step=spec.step or (1 if spec.kind == "int" else 0.1),
100
+ )
101
+ )
102
+ else:
103
+ updates.append(gr.update(visible=False))
104
+ return tuple(updates)
105
+
106
+
107
+ def collect_params(strategy_key: str, *values) -> Dict[str, float]:
108
+ strategy = get_strategy(strategy_key)
109
+ return {spec.name: spec.cast(values[i]) for i, spec in enumerate(strategy.params[:MAX_PARAMS])}
110
+
111
+
112
+ def _score_card(report) -> str:
113
+ v = report.verdict
114
+ color = GRADE_COLORS.get(v["grade"], "#898781")
115
+ market = report.market
116
+ provenance = (
117
+ f'<div class="provenance">Data: {market.source} · {market.symbol} · '
118
+ f"{market.start.date()} to {market.end.date()} · {len(market.df):,} bars</div>"
119
+ if market.is_real
120
+ else f'<div class="provenance sim">⚠ {market.note}</div>'
121
+ )
122
+ flags = "".join(f"<li>{f}</li>" for f in v["flags"])
123
+ flags_html = f'<ul class="flags">{flags}</ul>' if flags else ""
124
+
125
+ return f"""
126
+ <div class="score-card">
127
+ <div class="score-badge">
128
+ <div class="grade" style="color:{color}">{v['grade']}</div>
129
+ <div class="num">{v['score']} / 100</div>
130
+ </div>
131
+ <div class="score-body">
132
+ <h3>{report.strategy.name} on {market.symbol}</h3>
133
+ <p>{v['verdict']}</p>
134
+ </div>
135
+ </div>
136
+ {flags_html}
137
+ {provenance}
138
+ """
139
+
140
+
141
+ def _tiles(report) -> str:
142
+ m = report.backtest.metrics
143
+ b = report.backtest.benchmark_metrics
144
+ perm = report.permutation
145
+ p_text = f"{perm.p_value:.3f}" if perm else "—"
146
+ p_sub = "vs shuffled markets" if perm else "test skipped"
147
+ pbo = report.pbo.get("pbo")
148
+ pbo_text = f"{pbo:.0%}" if pbo is not None and pbo == pbo else "n/a"
149
+
150
+ cells = [
151
+ ("Total return", _fmt_pct(m.get("total_return", 0)), f"buy & hold {_fmt_pct(b.get('total_return', 0))}"),
152
+ ("CAGR", _fmt_pct(m.get("cagr", 0)), f"over {m.get('years', 0):.1f} years"),
153
+ ("Sharpe", f"{m.get('sharpe', 0):.2f}", f"buy & hold {b.get('sharpe', 0):.2f}"),
154
+ ("Max drawdown", _fmt_pct(m.get("max_drawdown", 0)), f"{m.get('time_under_water_yrs', 0):.1f}y under water"),
155
+ ("p-value", p_text, p_sub),
156
+ ("Deflated Sharpe", f"{report.dsr.get('dsr', 0):.2f}", f"after {report.trials.get('n', 1)} variants"),
157
+ ("Overfit prob.", pbo_text, "in-sample winner fails OOS"),
158
+ ("Trades", f"{int(m.get('n_trades', 0)):,}", f"{m.get('turnover_ann', 0):.0f}x turnover/yr"),
159
+ ]
160
+ tiles = "".join(
161
+ f'<div class="tile"><div class="label">{label}</div>'
162
+ f'<div class="value">{value}</div><div class="sub">{sub}</div></div>'
163
+ for label, value, sub in cells
164
+ )
165
+ return f'<div class="tiles">{tiles}</div>'
166
+
167
+
168
+ def _detail_markdown(report) -> str:
169
+ dsr, wf, pbo, stress = report.dsr, report.walkforward, report.pbo, report.cost_stress
170
+ mtr = dsr.get("min_track_record_years", float("inf"))
171
+ mtr_text = f"{mtr:.1f} years" if mtr == mtr and mtr != float("inf") else "never, at this effect size"
172
+
173
+ lines = [
174
+ "### Reading the evidence",
175
+ "",
176
+ f"**Selection bias.** {report.trials.get('n', 1)} parameter variants of "
177
+ f"*{report.strategy.name}* were backtested. The luckiest skill-free variant of that many "
178
+ f"would be expected to show an annualised Sharpe of about "
179
+ f"**{dsr.get('threshold_sr_annual', 0):.2f}** on its own. Yours was "
180
+ f"**{report.backtest.sharpe:.2f}**, which puts the deflated probability of a real edge at "
181
+ f"**{dsr.get('dsr', 0):.0%}**.",
182
+ "",
183
+ f"**Track record needed.** To call this Sharpe significant at 95% confidence given its "
184
+ f"skew ({dsr.get('skew', 0):.2f}) and kurtosis ({dsr.get('kurtosis', 0):.1f}), you would need "
185
+ f"about **{mtr_text}** of live returns.",
186
+ "",
187
+ f"**Costs.** At the modelled friction the Sharpe is {stress.get('sharpe_1x', 0):.2f}. "
188
+ f"Triple the costs and it becomes {stress.get('sharpe_3x', 0):.2f} "
189
+ f"({stress.get('ratio', 0):.0%} retained).",
190
+ "",
191
+ ]
192
+
193
+ if wf.get("folds"):
194
+ lines += [
195
+ f"**Walk-forward.** Across {len(wf['folds'])} folds the tuned in-sample Sharpe averaged "
196
+ f"{wf.get('mean_is_sharpe', 0):.2f} and the blind out-of-sample Sharpe averaged "
197
+ f"{wf.get('mean_oos_sharpe', 0):.2f} — an efficiency of {wf.get('efficiency', 0):.0%}. "
198
+ f"{wf.get('oos_win_rate', 0):.0%} of folds were profitable out of sample, and the tuner "
199
+ f"picked a different parameter set in {wf.get('param_instability', 0):.0%} of them.",
200
+ "",
201
+ ]
202
+ if pbo.get("note"):
203
+ lines += [f"**Overfitting.** {pbo['note']}", ""]
204
+ elif pbo.get("pbo") == pbo.get("pbo"):
205
+ lines += [
206
+ f"**Overfitting (CSCV).** Over {pbo.get('n_combinations', 0)} in/out splits of "
207
+ f"{pbo.get('n_strategies', 0)} variants, the in-sample winner landed in the bottom half "
208
+ f"out of sample **{pbo.get('pbo', 0):.0%}** of the time, and lost money outright "
209
+ f"{pbo.get('prob_oos_loss', 0):.0%} of the time. The most frequently selected variant was "
210
+ f"`{pbo.get('most_selected_label', 'n/a')}` "
211
+ f"({pbo.get('selection_stability', 0):.0%} of splits).",
212
+ "",
213
+ ]
214
+
215
+ lines += [
216
+ "> Past performance, simulated or otherwise, does not predict future returns. "
217
+ "This is research tooling, not investment advice.",
218
+ ]
219
+ return "\n".join(lines)
220
+
221
+
222
+ def analyse(
223
+ symbol: str,
224
+ start: str,
225
+ end: str,
226
+ strategy_key: str,
227
+ p1: float,
228
+ p2: float,
229
+ p3: float,
230
+ commission: float,
231
+ slippage: float,
232
+ allow_short: bool,
233
+ n_permutations: int,
234
+ perm_method: str,
235
+ wf_folds: int,
236
+ progress=gr.Progress(),
237
+ ):
238
+ """Main Lab handler. Never raises into the UI — it returns a readable message."""
239
+ try:
240
+ cfg = LabConfig(
241
+ symbol=symbol or "SPY",
242
+ start=start or "2015-01-01",
243
+ end=end or None,
244
+ strategy=strategy_key,
245
+ params=collect_params(strategy_key, p1, p2, p3),
246
+ commission_bps=float(commission),
247
+ slippage_bps=float(slippage),
248
+ allow_short=bool(allow_short),
249
+ n_permutations=int(n_permutations),
250
+ permutation_method="block" if perm_method.startswith("Block") else "permute",
251
+ wf_folds=int(wf_folds),
252
+ )
253
+ report = run_lab(cfg, progress=lambda f, m: progress(f, desc=m))
254
+ except Exception as exc: # noqa: BLE001 - the UI must always say something useful
255
+ logger.exception("Lab run failed")
256
+ message = f'<div class="score-card"><div class="score-body"><h3>Could not run that</h3><p>{exc}</p></div></div>'
257
+ blank = empty_figure("No results.")
258
+ return message, "", blank, blank, blank, blank, blank, blank, ""
259
+
260
+ return (
261
+ _score_card(report),
262
+ _tiles(report),
263
+ permutation_chart(report),
264
+ equity_chart(report),
265
+ drawdown_chart(report),
266
+ exposure_chart(report),
267
+ walkforward_chart(report),
268
+ score_chart(report.verdict),
269
+ _detail_markdown(report),
270
+ )
271
+
272
+
273
+ def race(symbol: str, start: str, allow_short: bool, n_permutations: int, progress=gr.Progress()):
274
+ try:
275
+ cfg = LabConfig(
276
+ symbol=symbol or "SPY",
277
+ start=start or "2015-01-01",
278
+ allow_short=bool(allow_short),
279
+ )
280
+ table, market, _ = run_arena(
281
+ cfg,
282
+ n_permutations=int(n_permutations),
283
+ progress=lambda f, m: progress(f, desc=m),
284
+ )
285
+ except Exception as exc: # noqa: BLE001
286
+ logger.exception("Arena run failed")
287
+ return pd.DataFrame({"Error": [str(exc)]}), empty_figure("No results."), ""
288
+
289
+ display = table.copy()
290
+ for col in ("Return", "CAGR", "MaxDD"):
291
+ display[col] = display[col].map(lambda v: f"{v * 100:,.1f}%")
292
+ for col in ("Sharpe", "DSR", "Evidence"):
293
+ display[col] = display[col].map(lambda v: f"{v:.2f}")
294
+ display["p-value"] = display["p-value"].map(lambda v: "—" if v != v else f"{v:.3f}")
295
+ display = display.drop(columns=["key"])
296
+
297
+ provenance = (
298
+ f"Data: {market.source} · {market.symbol} · {market.start.date()} to {market.end.date()}"
299
+ if market.is_real
300
+ else f"⚠ {market.note}"
301
+ )
302
+ note = (
303
+ f"{provenance}\n\nRanked by **evidence** — `(1 − p) × deflated Sharpe` — not by return. "
304
+ "*Buy & Hold* and *Coin Flip* are in the field on purpose: a leaderboard without a "
305
+ "control group is marketing, not measurement."
306
+ )
307
+ return display, arena_chart(table), note
308
+
309
+
310
+ HOW_IT_WORKS = """
311
+ ## Why most backtests are wrong
312
+
313
+ A backtest is a measurement taken with a ruler you built after seeing the thing you
314
+ are measuring. Four failure modes do almost all the damage, and this Space tests for
315
+ each one.
316
+
317
+ ### 1. The market had no structure to find — permutation test
318
+
319
+ We take the real price series and shuffle it: each bar's gap, high, low, body and
320
+ volume are kept intact, but their **order** is destroyed. The result is a market with
321
+ the same volatility and the same fat tails, and no exploitable structure whatsoever.
322
+ Then we re-run *your exact rule* on hundreds of these shuffled markets.
323
+
324
+ If your Sharpe sits inside that cloud of results, your rule found nothing that a
325
+ coin-flip market would not also have handed it. The **p-value** is the share of
326
+ shuffled markets that did as well or better.
327
+
328
+ *Block mode* resamples contiguous chunks instead of single bars, preserving
329
+ short-horizon momentum and volatility clustering. It is a harder null, and trend
330
+ strategies should be held to it.
331
+
332
+ ### 2. You tried 200 things and reported the best — Deflated Sharpe Ratio
333
+
334
+ If you test 200 worthless strategies, the best of them will show a Sharpe near 1.0
335
+ purely by chance. The **Deflated Sharpe Ratio** (Bailey & López de Prado, 2014) works
336
+ out what the luckiest of *N* skill-free variants would have scored, and asks whether
337
+ yours beats that bar — with an extra penalty for negative skew and fat tails, the
338
+ return shapes that flatter naive Sharpe ratios.
339
+
340
+ This Space counts the whole parameter grid as trials, because that is what a
341
+ researcher would really have run.
342
+
343
+ ### 3. The parameters were fitted to the past — PBO and walk-forward
344
+
345
+ **Probability of Backtest Overfitting** (CSCV) cuts the timeline into chunks, and for
346
+ every way of splitting them half in-sample and half out-of-sample, checks whether the
347
+ in-sample winner stayed a winner. If the winner lands in the bottom half about half
348
+ the time, PBO ≈ 50% and your selection process has no skill at all.
349
+
350
+ **Walk-forward** re-tunes on a training window and trades the next window blind,
351
+ rolling forward. Efficiency is out-of-sample Sharpe over in-sample Sharpe: 100% means
352
+ the edge survived intact, 0% means it was entirely curve-fit.
353
+
354
+ ### 4. The edge is smaller than the costs — stress test
355
+
356
+ Every result here is net of commission and slippage charged on exposure changes, plus
357
+ a borrow fee on short positions. We then re-run at **triple** the friction. A real edge
358
+ degrades; a fake one disappears.
359
+
360
+ ---
361
+
362
+ ## The Reality Score
363
+
364
+ | Weight | Component | What it measures |
365
+ |---:|---|---|
366
+ | 30% | Significance | How far outside the shuffled-market null the result sits |
367
+ | 25% | Selection | Deflated Sharpe — does it clear the best-of-N bar |
368
+ | 20% | Walk-forward | How much of the tuned Sharpe survived trading forward |
369
+ | 15% | Overfitting | 1 − PBO, from combinatorially symmetric cross-validation |
370
+ | 10% | Robustness | Sharpe retained when costs triple |
371
+
372
+ Grades: **A** ≥ 85 · **B** ≥ 70 · **C** ≥ 55 · **D** ≥ 40 · **F** below 40.
373
+
374
+ The scale is deliberately harsh. Most strategies people post online score below 40,
375
+ and the honest response to that is not to soften the scale.
376
+
377
+ ## No look-ahead, by construction
378
+
379
+ A strategy emits a target exposure at each bar's close using only data up to that
380
+ bar. The engine holds `position[t] = target[t - lag]` with `lag ≥ 1`, so a signal
381
+ computed on Tuesday's close cannot earn Tuesday's move. That is the single line where
382
+ look-ahead could enter, and the test suite asserts it directly.
383
+
384
+ ## Use it from Python
385
+
386
+ ```python
387
+ from algotrader import LabConfig, run_lab
388
+
389
+ report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross", params={"fast": 20, "slow": 100}))
390
+ print(report.verdict["grade"], report.verdict["score"])
391
+ print(report.permutation.p_value, report.dsr["dsr"], report.pbo["pbo"])
392
+ ```
393
+
394
+ Or from the command line:
395
+
396
+ ```bash
397
+ python -m algotrader.cli lab --symbol SPY --strategy donchian_breakout --permutations 500
398
+ python -m algotrader.cli arena --symbol BTC-USD
399
+ ```
400
+
401
+ ---
402
+
403
+ *Research tooling, not investment advice. Nothing here is a recommendation to trade.*
404
+ """
405
+
406
+
407
+ # Gradio 6 moved `css` and `theme` from the Blocks constructor to launch().
408
+ # Spaces pin their own version, so pass them wherever the installed one wants.
409
+ _GRADIO_MAJOR = int(gr.__version__.split(".")[0])
410
+ _STYLE_KWARGS = {"css": CSS, "theme": gr.themes.Base()}
411
+ _BLOCKS_KWARGS = {} if _GRADIO_MAJOR >= 6 else _STYLE_KWARGS
412
+ # Gradio 6 also dropped launch(show_api=...).
413
+ _LAUNCH_KWARGS = dict(_STYLE_KWARGS) if _GRADIO_MAJOR >= 6 else {"show_api": False}
414
+
415
+
416
+ def build_app() -> gr.Blocks:
417
+ with gr.Blocks(title="Backtest Reality Check", **_BLOCKS_KWARGS) as demo:
418
+ with gr.Column(elem_id="hero"):
419
+ gr.HTML(
420
+ "<h1>Backtest Reality Check</h1>"
421
+ "<p>Your backtest is probably lying to you. Pick a market and a trading rule — "
422
+ "this runs it, then spends the rest of its effort trying to prove the result "
423
+ "was luck.</p>"
424
+ )
425
+
426
+ with gr.Tabs():
427
+ with gr.Tab("The Lab"):
428
+ with gr.Row():
429
+ with gr.Column(scale=1):
430
+ symbol = gr.Dropdown(
431
+ choices=DEFAULT_UNIVERSE, value="SPY", label="Ticker",
432
+ allow_custom_value=True,
433
+ info="Any Yahoo Finance symbol. Falls back to a simulated market if offline.",
434
+ )
435
+ with gr.Row():
436
+ start = gr.Textbox(value="2015-01-01", label="Start", scale=1)
437
+ end = gr.Textbox(value="", label="End (blank = today)", scale=1)
438
+
439
+ strategy = gr.Dropdown(
440
+ choices=STRATEGY_CHOICES, value="sma_cross", label="Strategy"
441
+ )
442
+ strategy_note = gr.Markdown(get_strategy("sma_cross").description)
443
+
444
+ param_sliders = [
445
+ gr.Slider(label=f"Parameter {i + 1}", visible=False, minimum=0, maximum=100)
446
+ for i in range(MAX_PARAMS)
447
+ ]
448
+
449
+ with gr.Accordion("Costs and testing", open=False):
450
+ commission = gr.Slider(0, 20, value=1, step=0.5, label="Commission (bps per trade)")
451
+ slippage = gr.Slider(0, 50, value=2, step=0.5, label="Slippage (bps per trade)")
452
+ allow_short = gr.Checkbox(value=True, label="Allow short positions")
453
+ n_perms = gr.Slider(
454
+ 0, 1000, value=250, step=50, label="Shuffled markets to test against",
455
+ info="More is stricter and slower. 250 is plenty for a first look.",
456
+ )
457
+ perm_method = gr.Radio(
458
+ ["Shuffle bars (standard)", "Block bootstrap (harder)"],
459
+ value="Shuffle bars (standard)", label="Null market",
460
+ )
461
+ wf_folds = gr.Slider(2, 8, value=5, step=1, label="Walk-forward folds")
462
+
463
+ run_button = gr.Button("Run reality check", variant="primary", size="lg")
464
+
465
+ gr.Examples(
466
+ label="Or try one of these",
467
+ examples=[
468
+ ["SPY", "sma_cross"],
469
+ ["BTC-USD", "donchian_breakout"],
470
+ ["NVDA", "rsi_reversion"],
471
+ ["QQQ", "momentum"],
472
+ ["SPY", "coin_flip"],
473
+ ],
474
+ inputs=[symbol, strategy],
475
+ )
476
+
477
+ with gr.Column(scale=2):
478
+ verdict_html = gr.HTML(
479
+ '<div class="score-card"><div class="score-body">'
480
+ "<h3>Nothing tested yet</h3><p>Pick a market and a rule, then hit "
481
+ "<b>Run reality check</b>. A full run is a few seconds.</p>"
482
+ "</div></div>"
483
+ )
484
+ tiles_html = gr.HTML("")
485
+
486
+ # The headline test gets the full width — it is the whole point.
487
+ perm_plot = gr.Plot(value=empty_figure("The headline test appears here.", height=320))
488
+ equity_plot = gr.Plot(value=empty_figure())
489
+ with gr.Row():
490
+ dd_plot = gr.Plot(value=empty_figure(height=240))
491
+ exposure_plot = gr.Plot(value=empty_figure(height=200))
492
+ with gr.Row():
493
+ wf_plot = gr.Plot(value=empty_figure(height=300))
494
+ components_plot = gr.Plot(value=empty_figure(height=260))
495
+ detail_md = gr.Markdown("")
496
+
497
+ strategy.change(
498
+ fn=param_controls, inputs=strategy, outputs=param_sliders
499
+ ).then(
500
+ fn=lambda k: get_strategy(k).description, inputs=strategy, outputs=strategy_note
501
+ )
502
+
503
+ run_button.click(
504
+ fn=analyse,
505
+ inputs=[
506
+ symbol, start, end, strategy, *param_sliders,
507
+ commission, slippage, allow_short, n_perms, perm_method, wf_folds,
508
+ ],
509
+ outputs=[
510
+ verdict_html, tiles_html, perm_plot, equity_plot,
511
+ dd_plot, exposure_plot, wf_plot, components_plot, detail_md,
512
+ ],
513
+ )
514
+
515
+ with gr.Tab("Arena"):
516
+ gr.Markdown(
517
+ "Race every strategy on the same market, ranked by **evidence** rather than "
518
+ "return. Buy & hold and a coin flip stay in the field as controls."
519
+ )
520
+ with gr.Row():
521
+ arena_symbol = gr.Dropdown(
522
+ choices=DEFAULT_UNIVERSE, value="SPY", label="Ticker", allow_custom_value=True
523
+ )
524
+ arena_start = gr.Textbox(value="2015-01-01", label="Start")
525
+ arena_short = gr.Checkbox(value=True, label="Allow shorts")
526
+ arena_perms = gr.Slider(0, 400, value=120, step=20, label="Shuffled markets per strategy")
527
+ arena_button = gr.Button("Run the arena", variant="primary")
528
+ arena_note = gr.Markdown("")
529
+ arena_table = gr.Dataframe(interactive=False, wrap=True)
530
+ arena_plot = gr.Plot(value=empty_figure(height=380))
531
+
532
+ arena_button.click(
533
+ fn=race,
534
+ inputs=[arena_symbol, arena_start, arena_short, arena_perms],
535
+ outputs=[arena_table, arena_plot, arena_note],
536
+ )
537
+
538
+ with gr.Tab("How it works"):
539
+ gr.Markdown(HOW_IT_WORKS)
540
+
541
+ gr.Markdown(
542
+ f"<sub>algotrader {__version__} · Apache-2.0 · "
543
+ "Research tooling, not investment advice.</sub>"
544
+ )
545
+
546
+ demo.load(fn=param_controls, inputs=strategy, outputs=param_sliders)
547
+
548
+ return demo
549
+
550
+
551
+ if __name__ == "__main__":
552
+ build_app().queue(max_size=24).launch(
553
+ server_name="0.0.0.0",
554
+ server_port=int(os.environ.get("PORT", 7860)),
555
+ **_LAUNCH_KWARGS,
556
+ )
docs/AGENTIC_SYSTEM_V1.md ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Algorithmic Trading
2
+
3
+ FinRL reinforcement-learning trading with Alpaca execution, plus optional Yahoo Finance OHLCV for unlabeled real-price research. Parallel LLC.
4
+
5
+ This is **research and paper-trading infrastructure**. Live capital requires a separate evaluation contract, feature-parity tests, and a rewritten execution path. Do not treat `paper_trading: false` as a promotion gate.
6
+
7
+ ---
8
+
9
+ ## 1. Title and Summary
10
+
11
+ **Algorithmic Trading**
12
+ Northwestern-trained data-engineering practice applied to a trading loop: ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
13
+
14
+ GitHub `main` is the FinRL / Docker / Streamlit tree. `dev` is the integration branch. Yahoo is an additive `data_source.type`, not a replacement for Alpaca or FinRL.
15
+
16
+ **Design themes**
17
+
18
+ * Four ingest paths: CSV replay, synthetic GBM, Alpaca REST, Yahoo (`yfinance>=1.0`)
19
+ * FinRL policies (PPO, A2C, DDPG, TD3) on a Gymnasium-style environment
20
+ * Alpaca for authenticated market data and order routing (paper by default)
21
+ * Yahoo for delayed public bars when no broker key is available
22
+ * Secrets from environment (`ALPACA_API_KEY`, `ALPACA_SECRET_KEY`), never committed
23
+ * Tests and Docker/CI as already present on this tree
24
+
25
+ ---
26
+
27
+ ## 2. Concepts and Methods
28
+
29
+ ### Market data
30
+
31
+ | Source | When to use | Failure modes |
32
+ | ------ | ----------- | ------------- |
33
+ | **CSV** | Offline replay; default in `config.yaml` | Missing path or OHLCV columns → `None` |
34
+ | **Synthetic** | Unit tests and demos | GBM is not tradable edge |
35
+ | **Alpaca** | Authenticated bars and live/paper orders | Auth, feed, and rate-limit failures |
36
+ | **Yahoo** | Real Close without a broker account | Unofficial API, ~15 min delay, interval lookback caps (1m ≈ 7 days). Pin `yfinance>=1.0`; 0.2.x fails against the current chart API |
37
+
38
+ `load_data` dispatches on `data_source.type`. Existing `alpaca` / `csv` / `synthetic` branches are unchanged.
39
+
40
+ ### Strategy and FinRL
41
+
42
+ * `StrategyAgent`: SMA, RSI, Bollinger, MACD on Close; teaching rule, not an alpha claim
43
+ * `FinRLAgent`: PPO / A2C / DDPG / TD3 via Stable-Baselines3; persist under `models/`
44
+ * `ExecutionAgent` / `AlpacaBroker`: paper simulation or Alpaca market/limit orders
45
+
46
+ Backtests in this repo are in-sample passes unless you add a purged walk-forward yourself. Leakage is the null hypothesis.
47
+
48
+ ---
49
+
50
+ ## 3. Stack
51
+
52
+ | Layer | Tools |
53
+ | ----- | ----- |
54
+ | Language | Python 3.11 (CI); 3.8+ stated for local |
55
+ | RL | FinRL / Stable-Baselines3, Gym/Gymnasium, PyTorch |
56
+ | Broker | alpaca-py |
57
+ | Market data | Alpaca REST; yfinance ≥ 1.0 (Yahoo) |
58
+ | Tabular | pandas, NumPy, scikit-learn |
59
+ | UI | Streamlit, Dash, Jupyter widgets |
60
+ | Deploy | Docker Compose, GitHub Actions |
61
+ | Tests | pytest, pytest-cov |
62
+
63
+ ---
64
+
65
+ ## 4. Structure
66
+
67
+ ```
68
+ algorithmic_trading/
69
+ ├── agentic_ai_system/ # ingest, strategy, FinRL, Alpaca, Yahoo
70
+ ├── ui/ # Streamlit, Dash, Jupyter, WebSocket
71
+ ├── tests/
72
+ ├── models/ # trained artifacts (gitignored bodies)
73
+ ├── data/ # generated CSV (gitignored)
74
+ ├── scripts/ # Docker / deploy helpers
75
+ ├── .github/workflows/ # CI/CD, release, backtesting
76
+ ├── config.yaml
77
+ ├── requirements.txt
78
+ ├── Dockerfile
79
+ └── docker-compose*.yml
80
+ ```
81
+
82
+ Branch policy: **`main`** (protected) and **`dev`** only. Do not re-enable Dependabot or the Monday `dependency-updates` workflow; those created extra branches.
83
+
84
+ ---
85
+
86
+ ## 5. Quick start
87
+
88
+ ```bash
89
+ git clone https://github.com/ParallelLLC/algorithmic_trading.git
90
+ cd algorithmic_trading
91
+ python -m venv .venv && source .venv/bin/activate
92
+ pip install -r requirements.txt
93
+ cp .env.example .env # Alpaca keys if using alpaca ingest or orders
94
+ ```
95
+
96
+ Default ingest is CSV. For Yahoo daily bars without a broker:
97
+
98
+ ```yaml
99
+ data_source:
100
+ type: 'yahoo'
101
+ trading:
102
+ symbol: 'AAPL'
103
+ timeframe: '1d'
104
+ ```
105
+
106
+ ```bash
107
+ python demo.py
108
+ python -m agentic_ai_system.main --mode backtest --start-date 2024-01-01 --end-date 2024-12-31
109
+ pytest tests/ -q
110
+ ```
111
+
112
+ UI launchers and Docker are documented in `UI_SETUP.md` and `DOCKER_HUB_SETUP.md`. Paper-trade before live. Yahoo is not a SIP tape.
113
+
114
+ ---
115
+
116
+ ## 6. Configuration (additive Yahoo keys)
117
+
118
+ | Key | Meaning |
119
+ | --- | ------- |
120
+ | `data_source.type` | `csv` \| `synthetic` \| `alpaca` \| `yahoo` |
121
+ | `yahoo.start_date` / `end_date` | Historical window; clamped per Yahoo interval limits |
122
+ | `yahoo.auto_adjust` | Passed to `yfinance` |
123
+ | `execution.broker_api` | `paper` \| `alpaca_paper` \| `alpaca_live` |
124
+ | `finrl.algorithm` | PPO, A2C, DDPG, TD3 |
125
+
126
+ ---
127
+
128
+ **License:** Apache License 2.0
129
+ **Organization:** [Parallel LLC](https://github.com/ParallelLLC)
130
+ **Repository:** <https://github.com/ParallelLLC/algorithmic_trading>
requirements-space.txt ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Dependencies for the Hugging Face Space (app.py + the algotrader package).
2
+ #
3
+ # Deliberately minimal: a Space that takes ten minutes to build is a Space
4
+ # nobody waits for. The v1 agentic system's heavier stack (torch,
5
+ # stable-baselines3, dash, alpaca-py) lives in requirements.txt and is not
6
+ # needed to run the Lab.
7
+ gradio>=4.44,<7
8
+ numpy>=1.24
9
+ pandas>=2.0
10
+ scipy>=1.10
11
+ plotly>=5.18
12
+ yfinance>=0.2.40
scripts/deploy_hf_space.sh ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ #
3
+ # Assemble and push the Hugging Face Space.
4
+ #
5
+ # The Space gets only what it needs to run: app.py, the algotrader package, the
6
+ # Space card, and the minimal requirements file. The v1 agentic system and its
7
+ # heavy ML stack stay in this repo, so the Space builds in under a minute.
8
+ #
9
+ # Usage:
10
+ # HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <hf-username>/<space-name>
11
+ #
12
+ # Create the Space first at https://huggingface.co/new-space (SDK: Gradio).
13
+
14
+ set -euo pipefail
15
+
16
+ SPACE_ID="${1:-}"
17
+ if [[ -z "$SPACE_ID" ]]; then
18
+ echo "usage: HF_TOKEN=hf_xxx $0 <hf-username>/<space-name>" >&2
19
+ exit 2
20
+ fi
21
+ if [[ -z "${HF_TOKEN:-}" ]]; then
22
+ echo "error: HF_TOKEN is not set. Create a write token at https://huggingface.co/settings/tokens" >&2
23
+ exit 2
24
+ fi
25
+
26
+ REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
27
+ STAGING="$(mktemp -d)"
28
+ trap 'rm -rf "$STAGING"' EXIT
29
+
30
+ echo "==> Staging Space contents in $STAGING"
31
+ git clone --quiet "https://user:${HF_TOKEN}@huggingface.co/spaces/${SPACE_ID}" "$STAGING/space"
32
+ cd "$STAGING/space"
33
+
34
+ # Replace tracked content wholesale so deletions propagate, but keep .git.
35
+ find . -mindepth 1 -maxdepth 1 ! -name .git -exec rm -rf {} +
36
+
37
+ cp -r "$REPO_ROOT/algotrader" ./algotrader
38
+ cp "$REPO_ROOT/app.py" ./app.py
39
+ cp "$REPO_ROOT/LICENSE" ./LICENSE
40
+ cp "$REPO_ROOT/SPACE_README.md" ./README.md
41
+ cp "$REPO_ROOT/requirements-space.txt" ./requirements.txt
42
+ find ./algotrader -name '__pycache__' -type d -exec rm -rf {} + 2>/dev/null || true
43
+
44
+ # A small test suite ships too: reviewers who check whether the statistics are
45
+ # real are exactly the audience worth convincing.
46
+ mkdir -p tests
47
+ cp "$REPO_ROOT/tests/test_v2_engine.py" \
48
+ "$REPO_ROOT/tests/test_v2_validation.py" \
49
+ "$REPO_ROOT/tests/test_v2_strategies.py" tests/
50
+
51
+ echo "==> Files staged:"
52
+ find . -path ./.git -prune -o -type f -print | sed 's|^\./| |'
53
+
54
+ git add -A
55
+ if git diff --cached --quiet; then
56
+ echo "==> Space is already up to date; nothing to push."
57
+ exit 0
58
+ fi
59
+
60
+ git -c user.email="deploy@localhost" -c user.name="space-deploy" \
61
+ commit --quiet -m "Deploy algotrader $(python3 -c 'import sys; sys.path.insert(0, "'"$REPO_ROOT"'"); import algotrader; print(algotrader.__version__)')"
62
+ git push --quiet origin HEAD:main
63
+
64
+ echo "==> Pushed. Live at https://huggingface.co/spaces/${SPACE_ID}"
tests/test_v2_cli.py ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """CLI surface — argument handling and that each command actually completes."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import json
6
+
7
+ import pytest
8
+
9
+ from algotrader.cli import _parse_params, main
10
+
11
+
12
+ class TestArgumentParsing:
13
+ def test_params_are_parsed_into_floats(self):
14
+ assert _parse_params(["fast=10", "slow=50.5"]) == {"fast": 10.0, "slow": 50.5}
15
+
16
+ def test_params_without_an_equals_sign_are_rejected(self):
17
+ with pytest.raises(SystemExit, match="name=value"):
18
+ _parse_params(["fast10"])
19
+
20
+ def test_no_params_is_an_empty_dict(self):
21
+ assert _parse_params(None) == {}
22
+
23
+ def test_an_unknown_strategy_is_rejected_at_parse_time(self):
24
+ with pytest.raises(SystemExit):
25
+ main(["lab", "--strategy", "nope"])
26
+
27
+ def test_a_missing_subcommand_is_rejected(self):
28
+ with pytest.raises(SystemExit):
29
+ main([])
30
+
31
+
32
+ class TestCommands:
33
+ """All runs force the simulator so the suite never touches the network."""
34
+
35
+ def test_lab_prints_a_report(self, capsys):
36
+ code = main([
37
+ "lab", "--symbol", "SPY", "--source", "synthetic", "--strategy", "sma_cross",
38
+ "--start", "2018-01-01", "--end", "2023-01-01",
39
+ "--permutations", "10", "--folds", "2", "--quiet",
40
+ ])
41
+ out = capsys.readouterr().out
42
+ assert code == 0
43
+ assert "REALITY SCORE" in out
44
+ assert "Deflated Sharpe" in out
45
+
46
+ def test_lab_json_output_is_machine_readable(self, capsys):
47
+ code = main([
48
+ "lab", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
49
+ "--permutations", "10", "--folds", "2", "--quiet", "--json",
50
+ ])
51
+ payload = json.loads(capsys.readouterr().out)
52
+ assert code == 0
53
+ assert 0.0 <= payload["verdict"]["score"] <= 100.0
54
+ assert payload["verdict"]["grade"] in {"A", "B", "C", "D", "F"}
55
+ assert payload["source"] == "synthetic"
56
+
57
+ def test_lab_honours_param_overrides(self, capsys):
58
+ main([
59
+ "lab", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
60
+ "--strategy", "sma_cross", "--param", "fast=5", "--param", "slow=60",
61
+ "--permutations", "0", "--folds", "2", "--quiet", "--json",
62
+ ])
63
+ payload = json.loads(capsys.readouterr().out)
64
+ assert payload["params"] == {"fast": 5, "slow": 60}
65
+ assert payload["p_value"] is None # permutations disabled
66
+
67
+ def test_arena_ranks_the_zoo(self, capsys):
68
+ code = main([
69
+ "arena", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
70
+ "--permutations", "0", "--quiet",
71
+ ])
72
+ out = capsys.readouterr().out
73
+ assert code == 0
74
+ assert "Buy & Hold" in out and "Coin Flip" in out
75
+ assert "Ranked by evidence" in out
76
+
77
+ def test_strategies_lists_the_registry(self, capsys):
78
+ assert main(["strategies"]) == 0
79
+ out = capsys.readouterr().out
80
+ assert "sma_cross" in out and "donchian_breakout" in out
81
+
82
+ def test_errors_are_reported_not_raised(self, capsys):
83
+ code = main([
84
+ "lab", "--source", "synthetic",
85
+ "--start", "2022-01-01", "--end", "2022-02-01", # far too short
86
+ "--permutations", "0", "--quiet",
87
+ ])
88
+ assert code == 1
89
+ assert "error:" in capsys.readouterr().err
tests/test_v2_engine.py ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Engine correctness: alignment, costs, and the no-look-ahead guarantee."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+ import pytest
8
+
9
+ from algotrader.engine import bars_to_returns, run_backtest
10
+ from algotrader.metrics import compute_metrics, infer_periods_per_year, max_drawdown, sharpe_ratio
11
+ from algotrader.types import CostModel
12
+
13
+
14
+ def make_bars(n: int = 400, seed: int = 0) -> pd.DataFrame:
15
+ rng = np.random.default_rng(seed)
16
+ close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.01, n)))
17
+ index = pd.date_range("2020-01-01", periods=n, freq="B")
18
+ return pd.DataFrame(
19
+ {
20
+ "open": close,
21
+ "high": close * 1.005,
22
+ "low": close * 0.995,
23
+ "close": close,
24
+ "volume": 1e6,
25
+ },
26
+ index=index,
27
+ )
28
+
29
+
30
+ class TestNoLookAhead:
31
+ """The one property the whole project rests on."""
32
+
33
+ def test_signal_earns_the_following_bar_not_its_own(self):
34
+ df = make_bars()
35
+ asset_ret = bars_to_returns(df)
36
+
37
+ # A target that knows the *next* bar's direction must be perfect...
38
+ clairvoyant = np.sign(asset_ret.shift(-1)).fillna(0.0)
39
+ good = run_backtest(df, clairvoyant, costs=CostModel(0, 0, 0))
40
+ assert (good.returns.iloc[1:-1] >= -1e-12).all(), "perfect foresight should never lose"
41
+
42
+ # ...and a target that only knows the *current* bar must not be.
43
+ hindsight = np.sign(asset_ret).fillna(0.0)
44
+ meh = run_backtest(df, hindsight, costs=CostModel(0, 0, 0))
45
+ assert (meh.returns < 0).any(), "same-bar signal must not be risk-free"
46
+ assert good.sharpe > meh.sharpe
47
+
48
+ def test_position_is_target_shifted_by_lag(self):
49
+ df = make_bars(200)
50
+ target = pd.Series(np.linspace(-1, 1, len(df)), index=df.index)
51
+ for lag in (1, 2, 5):
52
+ result = run_backtest(df, target, lag=lag)
53
+ expected = target.shift(lag).fillna(0.0)
54
+ pd.testing.assert_series_equal(result.position, expected, check_names=False)
55
+
56
+ def test_lag_zero_is_rejected(self):
57
+ df = make_bars(120)
58
+ with pytest.raises(ValueError, match="lag"):
59
+ run_backtest(df, pd.Series(1.0, index=df.index), lag=0)
60
+
61
+ def test_future_bars_cannot_change_past_equity(self):
62
+ """Truncating the data must not alter the equity curve before the cut."""
63
+ df = make_bars(400)
64
+ target = pd.Series(np.tile([1.0, -1.0], len(df) // 2), index=df.index)
65
+
66
+ full = run_backtest(df, target)
67
+ cut = run_backtest(df.iloc[:250], target.iloc[:250])
68
+ np.testing.assert_allclose(
69
+ full.equity.iloc[:250].to_numpy(), cut.equity.to_numpy(), rtol=1e-12
70
+ )
71
+
72
+
73
+ class TestCosts:
74
+ def test_buy_and_hold_pays_once(self):
75
+ df = make_bars(300)
76
+ target = pd.Series(1.0, index=df.index)
77
+ result = run_backtest(df, target, costs=CostModel(commission_bps=5, slippage_bps=5, short_borrow_bps=0))
78
+ # One entry at 10bps one-way, and no further turnover.
79
+ assert result.costs.sum() == pytest.approx(10 / 1e4, rel=1e-9)
80
+ assert int(result.metrics["n_trades"]) == 1
81
+
82
+ def test_flipping_every_bar_costs_more_than_holding(self):
83
+ df = make_bars(300)
84
+ costs = CostModel(commission_bps=5, slippage_bps=5, short_borrow_bps=0)
85
+ hold = run_backtest(df, pd.Series(1.0, index=df.index), costs=costs)
86
+ flip = run_backtest(df, pd.Series(np.tile([1.0, -1.0], 150), index=df.index), costs=costs)
87
+ assert flip.costs.sum() > 100 * hold.costs.sum()
88
+
89
+ def test_zero_costs_means_gross_equals_net(self):
90
+ df = make_bars(200)
91
+ target = pd.Series(np.tile([1.0, 0.0], 100), index=df.index)
92
+ result = run_backtest(df, target, costs=CostModel(0, 0, 0))
93
+ pd.testing.assert_series_equal(result.returns, result.gross_returns, check_names=False)
94
+
95
+ def test_short_borrow_is_charged_only_on_shorts(self):
96
+ df = make_bars(260)
97
+ costs = CostModel(commission_bps=0, slippage_bps=0, short_borrow_bps=365)
98
+ long_only = run_backtest(df, pd.Series(1.0, index=df.index), costs=costs)
99
+ short_only = run_backtest(df, pd.Series(-1.0, index=df.index), costs=costs)
100
+ assert long_only.costs.sum() == pytest.approx(0.0, abs=1e-12)
101
+ assert short_only.costs.sum() > 0
102
+
103
+ def test_higher_costs_never_improve_returns(self):
104
+ df = make_bars(300)
105
+ target = pd.Series(np.tile([1.0, -1.0], 150), index=df.index)
106
+ cheap = run_backtest(df, target, costs=CostModel(1, 1, 0))
107
+ dear = run_backtest(df, target, costs=CostModel(20, 20, 0))
108
+ assert dear.equity.iloc[-1] < cheap.equity.iloc[-1]
109
+
110
+
111
+ class TestConstraints:
112
+ def test_shorts_are_clipped_when_disallowed(self):
113
+ df = make_bars(150)
114
+ target = pd.Series(-1.0, index=df.index)
115
+ result = run_backtest(df, target, allow_short=False)
116
+ assert (result.target >= 0).all()
117
+ assert (result.position >= 0).all()
118
+
119
+ def test_leverage_is_clipped(self):
120
+ df = make_bars(150)
121
+ result = run_backtest(df, pd.Series(5.0, index=df.index), max_leverage=1.5)
122
+ assert result.target.max() == pytest.approx(1.5)
123
+
124
+ def test_empty_frame_is_rejected(self):
125
+ with pytest.raises(ValueError):
126
+ run_backtest(pd.DataFrame(columns=["open", "high", "low", "close", "volume"]), pd.Series(dtype=float))
127
+
128
+
129
+ class TestMetrics:
130
+ def test_sharpe_of_constant_returns_is_zero_not_infinite(self):
131
+ flat = pd.Series([0.001] * 100)
132
+ assert sharpe_ratio(flat, 252) == 0.0
133
+
134
+ def test_sharpe_scales_with_annualisation(self):
135
+ rng = np.random.default_rng(1)
136
+ returns = pd.Series(rng.normal(0.001, 0.01, 5000))
137
+ assert sharpe_ratio(returns, 252) == pytest.approx(sharpe_ratio(returns, 1) * np.sqrt(252))
138
+
139
+ def test_max_drawdown_matches_a_hand_worked_example(self):
140
+ equity = pd.Series([100.0, 120.0, 60.0, 90.0])
141
+ assert max_drawdown(equity) == pytest.approx(-0.5)
142
+
143
+ def test_buy_and_hold_metrics_match_the_price_series(self):
144
+ df = make_bars(500)
145
+ result = run_backtest(df, pd.Series(1.0, index=df.index), costs=CostModel(0, 0, 0))
146
+ expected = df["close"].iloc[-1] / df["close"].iloc[0] - 1.0
147
+ assert result.metrics["total_return"] == pytest.approx(expected, rel=1e-9)
148
+
149
+ def test_periodicity_inference(self):
150
+ daily = pd.date_range("2020-01-01", periods=300, freq="B")
151
+ assert 200 <= infer_periods_per_year(daily) <= 300
152
+ hourly = pd.date_range("2020-01-01", periods=300, freq="h")
153
+ assert infer_periods_per_year(hourly) > 1000
154
+
155
+ def test_metrics_survive_a_degenerate_series(self):
156
+ empty = pd.Series(dtype=float)
157
+ out = compute_metrics(empty, pd.Series(dtype=float))
158
+ assert "periods_per_year" in out
tests/test_v2_strategies.py ADDED
@@ -0,0 +1,229 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Strategy zoo, data loading, and the end-to-end Lab pipeline."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+ import pytest
8
+
9
+ from algotrader import LabConfig, run_lab
10
+ from algotrader.data import _normalise, load_ohlcv, simulate_ohlcv
11
+ from algotrader.indicators import atr, bollinger, donchian, ema, macd, rsi, sma
12
+ from algotrader.lab import run_arena
13
+ from algotrader.strategies import REGISTRY, get_strategy, list_strategies
14
+ from algotrader.types import MarketData
15
+ from algotrader.verdict import reality_score
16
+
17
+ ALL_KEYS = sorted(REGISTRY)
18
+
19
+
20
+ @pytest.fixture(scope="module")
21
+ def bars() -> pd.DataFrame:
22
+ return simulate_ohlcv("AAPL", "2016-01-01", "2023-01-01")
23
+
24
+
25
+ class TestIndicatorsAreCausal:
26
+ """An indicator at bar t must not move when bars after t arrive."""
27
+
28
+ @pytest.mark.parametrize(
29
+ "fn",
30
+ [
31
+ lambda d: sma(d["close"], 20),
32
+ lambda d: ema(d["close"], 12),
33
+ lambda d: rsi(d["close"], 14),
34
+ lambda d: macd(d["close"])[0],
35
+ lambda d: bollinger(d["close"])[2],
36
+ lambda d: atr(d, 14),
37
+ lambda d: donchian(d, 20)[1],
38
+ ],
39
+ )
40
+ def test_prefix_is_stable(self, bars, fn):
41
+ cut = 800
42
+ full = fn(bars).iloc[:cut]
43
+ partial = fn(bars.iloc[:cut])
44
+ pd.testing.assert_series_equal(full, partial, check_names=False, rtol=1e-9)
45
+
46
+ def test_rsi_stays_in_range(self, bars):
47
+ values = rsi(bars["close"], 14).dropna()
48
+ assert values.between(0, 100).all()
49
+
50
+ def test_donchian_excludes_the_current_bar(self, bars):
51
+ _, upper = donchian(bars, 20)
52
+ # A breakout must be possible: the channel cannot already contain today.
53
+ assert (bars["high"] > upper).any()
54
+
55
+
56
+ class TestStrategies:
57
+ @pytest.mark.parametrize("key", ALL_KEYS)
58
+ def test_output_is_well_formed(self, bars, key):
59
+ target = get_strategy(key).generate(bars)
60
+ assert target.index.equals(bars.index)
61
+ assert target.notna().all()
62
+ assert target.between(-1.0, 1.0).all()
63
+
64
+ @pytest.mark.parametrize("key", ALL_KEYS)
65
+ def test_signals_are_causal(self, bars, key):
66
+ cut = 900
67
+ full = get_strategy(key).generate(bars).iloc[:cut]
68
+ partial = get_strategy(key).generate(bars.iloc[:cut])
69
+ pd.testing.assert_series_equal(full, partial, check_names=False, rtol=1e-9)
70
+
71
+ @pytest.mark.parametrize("key", ALL_KEYS)
72
+ def test_grid_is_non_empty_and_valid(self, key):
73
+ strategy = get_strategy(key)
74
+ grid = strategy.grid(limit=40)
75
+ assert 1 <= len(grid) <= 40
76
+ for combo in grid:
77
+ assert set(combo) <= {p.name for p in strategy.params}
78
+ if "fast" in combo and "slow" in combo:
79
+ assert combo["fast"] < combo["slow"]
80
+
81
+ def test_unknown_strategy_names_the_alternatives(self):
82
+ with pytest.raises(KeyError, match="Available"):
83
+ get_strategy("does_not_exist")
84
+
85
+ def test_clean_ignores_unknown_params_and_casts_types(self):
86
+ strategy = get_strategy("sma_cross")
87
+ cleaned = strategy.clean({"fast": 15.7, "nonsense": 1})
88
+ assert cleaned == {"fast": 15, "slow": 100}
89
+
90
+ def test_buy_and_hold_is_always_fully_invested(self, bars):
91
+ assert (get_strategy("buy_and_hold").generate(bars) == 1.0).all()
92
+
93
+ def test_list_strategies_honours_exclusions(self):
94
+ keys = {s.key for s in list_strategies(exclude=["coin_flip"])}
95
+ assert "coin_flip" not in keys and "sma_cross" in keys
96
+
97
+
98
+ class TestData:
99
+ def test_simulation_is_deterministic_per_symbol(self):
100
+ a = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
101
+ b = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
102
+ pd.testing.assert_frame_equal(a, b)
103
+
104
+ def test_different_symbols_simulate_differently(self):
105
+ a = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
106
+ b = simulate_ohlcv("TSLA", "2018-01-01", "2022-01-01")
107
+ assert not np.allclose(a["close"].to_numpy(), b["close"].to_numpy())
108
+
109
+ def test_simulated_bars_are_internally_consistent(self):
110
+ df = simulate_ohlcv("SPY", "2015-01-01", "2023-01-01")
111
+ assert (df["high"] >= df[["open", "close"]].max(axis=1) - 1e-9).all()
112
+ assert (df["low"] <= df[["open", "close"]].min(axis=1) + 1e-9).all()
113
+ assert (df["close"] > 0).all()
114
+ assert df.index.is_monotonic_increasing
115
+
116
+ def test_simulation_has_fat_tails_and_vol_clustering(self):
117
+ """Naive GBM flatters strategies; the simulator must be harder than that."""
118
+ returns = simulate_ohlcv("SPY", "2005-01-01", "2023-01-01")["close"].pct_change().dropna()
119
+ assert returns.kurtosis() > 1.0
120
+ assert returns.abs().autocorr(1) > 0.05
121
+
122
+ def test_offline_load_falls_back_and_says_so(self, monkeypatch):
123
+ monkeypatch.setattr("algotrader.data._download", lambda *a, **k: None)
124
+ monkeypatch.setattr("algotrader.data._read_cache", lambda *a, **k: None)
125
+ market = load_ohlcv("SPY", "2018-01-01", "2022-01-01")
126
+ assert market.source == "synthetic"
127
+ assert not market.is_real
128
+ assert "unavailable" in market.note
129
+
130
+ def test_normalise_handles_yahoo_style_frames(self):
131
+ index = pd.date_range("2020-01-01", periods=5, tz="UTC")
132
+ raw = pd.DataFrame(
133
+ {"Open": 1.0, "High": 2.0, "Low": 0.5, "Adj Close": 1.5, "Volume": 10},
134
+ index=index,
135
+ )
136
+ out = _normalise(raw)
137
+ assert list(out.columns) == ["open", "high", "low", "close", "volume"]
138
+ assert out.index.tz is None
139
+
140
+ def test_normalise_rejects_frames_with_no_price(self):
141
+ with pytest.raises(ValueError, match="missing required column"):
142
+ _normalise(pd.DataFrame({"volume": [1, 2, 3]}))
143
+
144
+ def test_too_short_a_range_is_rejected(self):
145
+ with pytest.raises(ValueError, match="too short"):
146
+ simulate_ohlcv("SPY", "2020-01-01", "2020-01-10")
147
+
148
+
149
+ class TestVerdict:
150
+ def test_strong_evidence_outranks_weak_evidence(self):
151
+ metrics = {"n_trades": 300, "sharpe": 1.2, "total_return": 0.8, "max_drawdown": -0.2}
152
+ benchmark = {"sharpe": 0.4}
153
+ strong = reality_score(metrics, benchmark, p_value=0.001, dsr=0.99, pbo=0.02,
154
+ wf_efficiency=0.9, wf_win_rate=1.0, cost_stress_ratio=0.9)
155
+ weak = reality_score(metrics, benchmark, p_value=0.45, dsr=0.10, pbo=0.55,
156
+ wf_efficiency=-0.2, wf_win_rate=0.2, cost_stress_ratio=0.1)
157
+ assert strong["score"] > 85 > weak["score"]
158
+ assert strong["grade"] == "A" and weak["grade"] == "F"
159
+
160
+ def test_score_is_always_inside_the_scale(self):
161
+ for p in (0.0, 0.5, 1.0):
162
+ for dsr in (0.0, 1.0):
163
+ out = reality_score(
164
+ {"n_trades": 100, "sharpe": 0.5, "total_return": 0.2, "max_drawdown": -0.1},
165
+ {"sharpe": 0.1}, p_value=p, dsr=dsr, pbo=0.2,
166
+ wf_efficiency=0.5, cost_stress_ratio=0.5,
167
+ )
168
+ assert 0.0 <= out["score"] <= 100.0
169
+
170
+ def test_losing_money_caps_the_score(self):
171
+ out = reality_score(
172
+ {"n_trades": 200, "sharpe": 0.3, "total_return": -0.4, "max_drawdown": -0.6},
173
+ {"sharpe": 0.5}, p_value=0.001, dsr=0.99, pbo=0.01,
174
+ wf_efficiency=1.0, cost_stress_ratio=1.0,
175
+ )
176
+ assert out["score"] <= 50
177
+ assert any("lost money" in f for f in out["flags"])
178
+
179
+ def test_too_few_trades_is_flagged_and_capped(self):
180
+ out = reality_score(
181
+ {"n_trades": 3, "sharpe": 2.5, "total_return": 1.0, "max_drawdown": -0.1},
182
+ {"sharpe": 0.3}, p_value=0.001, dsr=0.99, pbo=0.01,
183
+ wf_efficiency=1.0, cost_stress_ratio=1.0,
184
+ )
185
+ assert out["score"] <= 55
186
+ assert any("coin flips" in f for f in out["flags"])
187
+
188
+ def test_closet_indexing_is_called_out(self):
189
+ out = reality_score(
190
+ {"n_trades": 50, "sharpe": 0.6, "total_return": 0.5, "max_drawdown": -0.2},
191
+ {"sharpe": 0.6}, benchmark_correlation=0.99,
192
+ )
193
+ assert any("repackaged long position" in f for f in out["flags"])
194
+
195
+
196
+ class TestLabEndToEnd:
197
+ def test_full_pipeline_produces_a_complete_report(self, monkeypatch):
198
+ monkeypatch.setattr("algotrader.lab.load_ohlcv", lambda *a, **k: MarketData(
199
+ "SIM", simulate_ohlcv("SPY", "2016-01-01", "2023-01-01"), "synthetic", "1d", "test"
200
+ ))
201
+ report = run_lab(LabConfig(strategy="sma_cross", n_permutations=25, wf_folds=3, grid_limit=12))
202
+
203
+ assert 0.0 <= report.verdict["score"] <= 100.0
204
+ assert report.verdict["grade"] in {"A", "B", "C", "D", "F"}
205
+ assert 0 < report.permutation.p_value <= 1
206
+ assert 0.0 <= report.dsr["dsr"] <= 1.0
207
+ assert report.trials["n"] > 1
208
+ assert len(report.backtest.equity) == len(report.market.df)
209
+ assert report.cost_stress["sharpe_3x"] <= report.cost_stress["sharpe_1x"] + 1e-9
210
+
211
+ def test_too_little_history_gives_a_readable_error(self, monkeypatch):
212
+ short = simulate_ohlcv("SPY", "2020-01-01", "2020-06-01")
213
+ monkeypatch.setattr(
214
+ "algotrader.lab.load_ohlcv",
215
+ lambda *a, **k: MarketData("SIM", short, "synthetic", "1d", "test"),
216
+ )
217
+ with pytest.raises(ValueError, match="Widen the date range"):
218
+ run_lab(LabConfig(n_permutations=0, wf_folds=2))
219
+
220
+ def test_arena_ranks_every_strategy_and_keeps_the_controls(self, monkeypatch):
221
+ monkeypatch.setattr("algotrader.lab.load_ohlcv", lambda *a, **k: MarketData(
222
+ "SIM", simulate_ohlcv("SPY", "2017-01-01", "2022-01-01"), "synthetic", "1d", "test"
223
+ ))
224
+ table, market, curves = run_arena(LabConfig(), n_permutations=0)
225
+
226
+ assert len(table) == len(REGISTRY)
227
+ assert {"buy_and_hold", "coin_flip"} <= set(table["key"])
228
+ assert table["Evidence"].is_monotonic_decreasing
229
+ assert set(curves) == set(REGISTRY)
tests/test_v2_validation.py ADDED
@@ -0,0 +1,179 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Statistical machinery.
2
+
3
+ These tests matter more than the engine's, because a validation suite that
4
+ always says "no edge" is as useless as one that always says "great edge". Each
5
+ class below checks both directions: it must reject noise *and* detect signal.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import numpy as np
11
+ import pandas as pd
12
+ import pytest
13
+
14
+ from algotrader.data import simulate_ohlcv
15
+ from algotrader.strategies import get_strategy
16
+ from algotrader.validation.deflated_sharpe import (
17
+ deflated_sharpe_ratio,
18
+ expected_max_sharpe,
19
+ min_track_record_length,
20
+ probabilistic_sharpe_ratio,
21
+ )
22
+ from algotrader.validation.pbo import probability_of_backtest_overfitting
23
+ from algotrader.validation.permutation import permutation_test, permute_bars
24
+ from algotrader.validation.walkforward import walk_forward
25
+
26
+
27
+ def trending_market(n: int = 2200, phi: float = 0.35, seed: int = 5) -> pd.DataFrame:
28
+ """A market with genuine, exploitable serial correlation."""
29
+ rng = np.random.default_rng(seed)
30
+ returns = np.zeros(n)
31
+ noise = rng.normal(0, 0.01, n)
32
+ for i in range(1, n):
33
+ returns[i] = phi * returns[i - 1] + noise[i]
34
+ close = 100 * np.exp(np.cumsum(returns))
35
+ index = pd.date_range("2012-01-01", periods=n, freq="B")
36
+ return pd.DataFrame(
37
+ {"open": close, "high": close * 1.004, "low": close * 0.996, "close": close, "volume": 1e6},
38
+ index=index,
39
+ )
40
+
41
+
42
+ class TestPermutationMechanics:
43
+ def test_shuffling_preserves_the_distribution_of_moves(self):
44
+ df = simulate_ohlcv("SPY", "2018-01-01", "2023-01-01")
45
+ shuffled = permute_bars(df, np.random.default_rng(0))
46
+
47
+ assert len(shuffled) == len(df)
48
+ assert shuffled.index.equals(df.index)
49
+ original = np.sort(np.log(df["close"] / df["open"]).to_numpy()[1:])
50
+ permuted = np.sort(np.log(shuffled["close"] / shuffled["open"]).to_numpy()[1:])
51
+ np.testing.assert_allclose(original, permuted, rtol=1e-9)
52
+
53
+ def test_shuffling_keeps_bars_internally_valid(self):
54
+ df = simulate_ohlcv("AAPL", "2019-01-01", "2023-01-01")
55
+ shuffled = permute_bars(df, np.random.default_rng(3))
56
+ assert (shuffled["high"] >= shuffled["low"]).all()
57
+ assert (shuffled["high"] >= shuffled["close"]).all()
58
+ assert (shuffled["low"] <= shuffled["close"]).all()
59
+ assert (shuffled["close"] > 0).all()
60
+
61
+ def test_shuffling_destroys_serial_correlation(self):
62
+ df = trending_market()
63
+ real = df["close"].pct_change().dropna().autocorr(1)
64
+ shuffled = permute_bars(df, np.random.default_rng(1))["close"].pct_change().dropna().autocorr(1)
65
+ assert real > 0.2
66
+ assert abs(shuffled) < 0.1
67
+
68
+ def test_block_mode_retains_some_structure(self):
69
+ df = trending_market()
70
+ blocked = permute_bars(df, np.random.default_rng(2), method="block", block=40)
71
+ assert blocked["close"].pct_change().dropna().autocorr(1) > 0.1
72
+
73
+ def test_p_value_can_never_be_zero(self):
74
+ """+1 correction: the observed run is itself a draw from the null."""
75
+ df = trending_market()
76
+ strategy = get_strategy("momentum")
77
+ result = permutation_test(
78
+ df, lambda f: strategy.generate(f, {"lookback": 5}), n_permutations=30, seed=0
79
+ )
80
+ assert result.p_value >= 1 / 31
81
+ assert 0 < result.p_value <= 1
82
+
83
+
84
+ class TestPermutationPower:
85
+ def test_real_edge_is_detected(self):
86
+ df = trending_market()
87
+ strategy = get_strategy("momentum")
88
+ result = permutation_test(
89
+ df, lambda f: strategy.generate(f, {"lookback": 5}), n_permutations=200, seed=1
90
+ )
91
+ assert result.observed > result.null.mean()
92
+ assert result.p_value < 0.05
93
+
94
+ def test_random_strategy_on_a_structureless_market_is_not_significant(self):
95
+ df = simulate_ohlcv("SIM", "2010-01-01", "2023-01-01")
96
+ strategy = get_strategy("coin_flip")
97
+ result = permutation_test(
98
+ df, lambda f: strategy.generate(f, {"hold": 5, "seed": 7}), n_permutations=200, seed=2
99
+ )
100
+ assert result.p_value > 0.05
101
+
102
+
103
+ class TestDeflatedSharpe:
104
+ def test_selection_bar_rises_with_the_number_of_trials(self):
105
+ low = expected_max_sharpe(10, 0.01)
106
+ high = expected_max_sharpe(1000, 0.01)
107
+ assert 0 < low < high
108
+
109
+ def test_a_single_trial_has_no_selection_bar(self):
110
+ assert expected_max_sharpe(1, 0.01) == 0.0
111
+
112
+ def test_more_trials_lowers_the_deflated_sharpe(self):
113
+ rng = np.random.default_rng(7)
114
+ returns = rng.normal(0.0006, 0.01, 2000)
115
+ few = deflated_sharpe_ratio(returns, 1.0, 252, n_trials=2, variance_of_trials=0.01)
116
+ many = deflated_sharpe_ratio(returns, 1.0, 252, n_trials=500, variance_of_trials=0.01)
117
+ assert many["dsr"] < few["dsr"]
118
+ assert many["psr"] == pytest.approx(few["psr"]) # PSR ignores selection
119
+
120
+ def test_psr_rises_with_track_record_length(self):
121
+ short = probabilistic_sharpe_ratio(0.05, 100)
122
+ long = probabilistic_sharpe_ratio(0.05, 5000)
123
+ assert 0.5 < short < long < 1.0
124
+
125
+ def test_negative_skew_and_fat_tails_are_penalised(self):
126
+ clean = probabilistic_sharpe_ratio(0.06, 1000, skew=0.0, kurtosis=3.0)
127
+ nasty = probabilistic_sharpe_ratio(0.06, 1000, skew=-1.5, kurtosis=12.0)
128
+ assert nasty < clean
129
+
130
+ def test_track_record_requirement_is_infinite_below_the_bar(self):
131
+ assert min_track_record_length(0.01, 500, benchmark=0.05) == float("inf")
132
+ assert np.isfinite(min_track_record_length(0.10, 500, benchmark=0.02))
133
+
134
+
135
+ class TestPBO:
136
+ def test_pure_noise_scores_near_one_half(self):
137
+ rng = np.random.default_rng(11)
138
+ matrix = rng.normal(0, 0.01, size=(1200, 30)) # 30 skill-free variants
139
+ result = probability_of_backtest_overfitting(matrix, n_splits=8)
140
+ assert 0.3 < result["pbo"] < 0.7
141
+
142
+ def test_a_genuinely_better_variant_is_not_flagged(self):
143
+ rng = np.random.default_rng(12)
144
+ matrix = rng.normal(0, 0.01, size=(1200, 20))
145
+ matrix[:, 3] += 0.004 # column 3 has a persistent, real edge
146
+ result = probability_of_backtest_overfitting(matrix, n_splits=8)
147
+ assert result["pbo"] < 0.15
148
+ assert result["most_selected_index"] == 3
149
+ assert result["selection_stability"] > 0.9
150
+
151
+ def test_too_few_variants_returns_nan_not_a_crash(self):
152
+ rng = np.random.default_rng(13)
153
+ result = probability_of_backtest_overfitting(rng.normal(0, 0.01, size=(500, 1)))
154
+ assert np.isnan(result["pbo"])
155
+ assert result["note"]
156
+
157
+ def test_odd_split_counts_are_made_even(self):
158
+ rng = np.random.default_rng(14)
159
+ result = probability_of_backtest_overfitting(rng.normal(0, 0.01, (800, 10)), n_splits=7)
160
+ assert result["n_combinations"] > 0
161
+
162
+
163
+ class TestWalkForward:
164
+ def test_a_real_edge_survives_out_of_sample(self):
165
+ result = walk_forward(trending_market(), get_strategy("momentum"), n_folds=4)
166
+ assert result["folds"]
167
+ assert result["mean_oos_sharpe"] > 0
168
+ assert result["efficiency"] > 0.3
169
+
170
+ def test_folds_do_not_overlap_train_and_test(self):
171
+ result = walk_forward(trending_market(), get_strategy("sma_cross"), n_folds=4)
172
+ for fold in result["folds"]:
173
+ assert fold["train_end"] <= fold["test_start"]
174
+
175
+ def test_short_history_degrades_gracefully(self):
176
+ df = trending_market(n=150)
177
+ result = walk_forward(df, get_strategy("momentum"), n_folds=5)
178
+ assert result["folds"] == []
179
+ assert result["note"]