Add algotrader 2.0: a backtester that tries to prove itself wrong
Browse filesMost backtesting tools answer "how much would this have made?". This adds a
second, harder question: how much of that was luck?
New `algotrader` package (additive — v1's agentic_ai_system, its tests, CI and
Docker setup are untouched):
- engine.py: vectorised backtester with an explicit no-look-ahead contract.
A strategy emits a target exposure at bar t's close; the engine holds
position[t] = target[t - lag] with lag >= 1, so a signal computed on a bar
cannot earn that bar's move. Costs are charged on exposure changes, plus a
borrow fee on shorts.
- validation/: the point of the release.
- permutation.py — Monte-Carlo permutation test. Shuffles bar order while
preserving each bar's anatomy, then re-runs the same rule on hundreds of
structure-free markets. Block-bootstrap mode preserves vol clustering.
- deflated_sharpe.py — Bailey & Lopez de Prado's PSR/DSR, charging for every
parameter variant tried and for skew and fat tails.
- pbo.py — probability of backtest overfitting via CSCV.
- walkforward.py — re-tune, trade forward blind, measure the decay.
- verdict.py: combines the panel into a 0-100 Reality Score and plain-English
warnings. Losing money, drawdowns past 50%, and fewer than 20 trades cap the
score regardless of how good the statistics look.
- data.py: yfinance -> cache -> deterministic simulator, so the app never shows
a stack trace on a cold click. The simulator uses regime switching, Student-t
innovations and persistent volatility; naive GBM flatters strategies.
- strategies.py: 11 strategies with parameter grids. Buy & hold and a coin flip
are permanent controls — a leaderboard without a control group is marketing.
- lab.py, charts.py, cli.py: pipeline shared by the app and the command line.
app.py is a Gradio Space with a Lab tab, an Arena leaderboard ranked by evidence
rather than return, and a How-it-works tab. Ships with SPACE_README.md (HF
frontmatter), a minimal requirements-space.txt so the Space builds in under a
minute, a deploy script and a sync workflow.
Two bugs found while testing: a constant return stream reported a Sharpe of
7.3e16 because the zero-variance guard tested `sd == 0` against a float that was
actually 1.4e-19, and the same pattern sat in PBO's column ranking. Both now use
an absolute floor.
111 tests, no network required. The validation tests check both directions —
the statistics reject noise and detect a genuine edge — since a suite that
always says "no edge" is as useless as one that always says "great edge".
- .github/workflows/sync-hf-space.yml +52 -0
- README.md +144 -95
- SPACE_README.md +101 -0
- algotrader/__init__.py +42 -0
- algotrader/charts.py +286 -0
- algotrader/cli.py +191 -0
- algotrader/data.py +246 -0
- algotrader/engine.py +133 -0
- algotrader/indicators.py +102 -0
- algotrader/lab.py +328 -0
- algotrader/metrics.py +157 -0
- algotrader/strategies.py +345 -0
- algotrader/types.py +100 -0
- algotrader/validation/__init__.py +15 -0
- algotrader/validation/deflated_sharpe.py +145 -0
- algotrader/validation/pbo.py +120 -0
- algotrader/validation/permutation.py +187 -0
- algotrader/validation/walkforward.py +129 -0
- algotrader/verdict.py +177 -0
- app.py +556 -0
- docs/AGENTIC_SYSTEM_V1.md +130 -0
- requirements-space.txt +12 -0
- scripts/deploy_hf_space.sh +64 -0
- tests/test_v2_cli.py +89 -0
- tests/test_v2_engine.py +158 -0
- tests/test_v2_strategies.py +229 -0
- tests/test_v2_validation.py +179 -0
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
name: Sync Hugging Face Space
|
| 2 |
+
|
| 3 |
+
# Publishes app.py + the algotrader package to a Hugging Face Space.
|
| 4 |
+
#
|
| 5 |
+
# Setup (one time):
|
| 6 |
+
# 1. Create the Space at https://huggingface.co/new-space with SDK "Gradio".
|
| 7 |
+
# 2. Add repository secret HF_TOKEN — a write token from
|
| 8 |
+
# https://huggingface.co/settings/tokens
|
| 9 |
+
# 3. Add repository variable HF_SPACE_ID, e.g. "your-username/backtest-reality-check".
|
| 10 |
+
#
|
| 11 |
+
# Without those set the job is skipped, so forks and PRs are unaffected.
|
| 12 |
+
|
| 13 |
+
on:
|
| 14 |
+
push:
|
| 15 |
+
branches: [main]
|
| 16 |
+
paths:
|
| 17 |
+
- 'app.py'
|
| 18 |
+
- 'algotrader/**'
|
| 19 |
+
- 'SPACE_README.md'
|
| 20 |
+
- 'requirements-space.txt'
|
| 21 |
+
- 'tests/test_v2_*.py'
|
| 22 |
+
- '.github/workflows/sync-hf-space.yml'
|
| 23 |
+
workflow_dispatch:
|
| 24 |
+
|
| 25 |
+
jobs:
|
| 26 |
+
sync:
|
| 27 |
+
runs-on: ubuntu-latest
|
| 28 |
+
if: vars.HF_SPACE_ID != ''
|
| 29 |
+
steps:
|
| 30 |
+
- uses: actions/checkout@v4
|
| 31 |
+
|
| 32 |
+
- uses: actions/setup-python@v5
|
| 33 |
+
with:
|
| 34 |
+
python-version: '3.11'
|
| 35 |
+
|
| 36 |
+
- name: Verify the Space actually runs before publishing it
|
| 37 |
+
run: |
|
| 38 |
+
pip install --quiet -r requirements-space.txt pytest
|
| 39 |
+
python -m pytest tests/test_v2_engine.py tests/test_v2_validation.py \
|
| 40 |
+
tests/test_v2_strategies.py -q
|
| 41 |
+
ALGOTRADER_OFFLINE=1 python -c "import app; app.build_app(); print('Space builds OK')"
|
| 42 |
+
|
| 43 |
+
- name: Push to the Space
|
| 44 |
+
env:
|
| 45 |
+
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
| 46 |
+
run: |
|
| 47 |
+
if [ -z "$HF_TOKEN" ]; then
|
| 48 |
+
echo "HF_TOKEN secret is not set; skipping publish." >&2
|
| 49 |
+
exit 0
|
| 50 |
+
fi
|
| 51 |
+
chmod +x scripts/deploy_hf_space.sh
|
| 52 |
+
./scripts/deploy_hf_space.sh "${{ vars.HF_SPACE_ID }}"
|
|
@@ -1,130 +1,179 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
Northwestern-trained data-engineering practice applied to a trading loop: ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
|
| 13 |
-
|
| 14 |
-
GitHub `main` is the FinRL / Docker / Streamlit tree. `dev` is the integration branch. Yahoo is an additive `data_source.type`, not a replacement for Alpaca or FinRL.
|
| 15 |
-
|
| 16 |
-
**Design themes**
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
* Alpaca for authenticated market data and order routing (paper by default)
|
| 21 |
-
* Yahoo for delayed public bars when no broker key is available
|
| 22 |
-
* Secrets from environment (`ALPACA_API_KEY`, `ALPACA_SECRET_KEY`), never committed
|
| 23 |
-
* Tests and Docker/CI as already present on this tree
|
| 24 |
|
| 25 |
---
|
| 26 |
|
| 27 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
| **Synthetic** | Unit tests and demos | GBM is not tradable edge |
|
| 35 |
-
| **Alpaca** | Authenticated bars and live/paper orders | Auth, feed, and rate-limit failures |
|
| 36 |
-
| **Yahoo** | Real Close without a broker account | Unofficial API, ~15 min delay, interval lookback caps (1m ≈ 7 days). Pin `yfinance>=1.0`; 0.2.x fails against the current chart API |
|
| 37 |
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
-
##
|
| 41 |
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
-
|
|
|
|
| 47 |
|
| 48 |
-
|
| 49 |
|
| 50 |
-
|
|
|
|
|
|
|
| 51 |
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
| Language | Python 3.11 (CI); 3.8+ stated for local |
|
| 55 |
-
| RL | FinRL / Stable-Baselines3, Gym/Gymnasium, PyTorch |
|
| 56 |
-
| Broker | alpaca-py |
|
| 57 |
-
| Market data | Alpaca REST; yfinance ≥ 1.0 (Yahoo) |
|
| 58 |
-
| Tabular | pandas, NumPy, scikit-learn |
|
| 59 |
-
| UI | Streamlit, Dash, Jupyter widgets |
|
| 60 |
-
| Deploy | Docker Compose, GitHub Actions |
|
| 61 |
-
| Tests | pytest, pytest-cov |
|
| 62 |
|
| 63 |
-
|
| 64 |
|
| 65 |
-
##
|
| 66 |
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
├── tests/
|
| 72 |
-
├── models/ # trained artifacts (gitignored bodies)
|
| 73 |
-
├── data/ # generated CSV (gitignored)
|
| 74 |
-
├── scripts/ # Docker / deploy helpers
|
| 75 |
-
├── .github/workflows/ # CI/CD, release, backtesting
|
| 76 |
-
├── config.yaml
|
| 77 |
-
├── requirements.txt
|
| 78 |
-
├── Dockerfile
|
| 79 |
-
└── docker-compose*.yml
|
| 80 |
-
```
|
| 81 |
|
| 82 |
-
|
| 83 |
|
| 84 |
-
|
| 85 |
|
| 86 |
-
|
|
|
|
| 87 |
|
| 88 |
```bash
|
| 89 |
-
|
| 90 |
-
cd algorithmic_trading
|
| 91 |
-
python -m venv .venv && source .venv/bin/activate
|
| 92 |
-
pip install -r requirements.txt
|
| 93 |
-
cp .env.example .env # Alpaca keys if using alpaca ingest or orders
|
| 94 |
```
|
| 95 |
|
| 96 |
-
|
|
|
|
|
|
|
| 97 |
|
| 98 |
-
``
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
timeframe: '1d'
|
| 104 |
-
```
|
| 105 |
|
| 106 |
```bash
|
| 107 |
-
python
|
| 108 |
-
python -m agentic_ai_system.main --mode backtest --start-date 2024-01-01 --end-date 2024-12-31
|
| 109 |
-
pytest tests/ -q
|
| 110 |
```
|
| 111 |
|
| 112 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 113 |
|
| 114 |
-
|
| 115 |
|
| 116 |
-
|
|
|
|
|
|
|
|
|
|
| 117 |
|
| 118 |
-
|
| 119 |
-
| --- | ------- |
|
| 120 |
-
| `data_source.type` | `csv` \| `synthetic` \| `alpaca` \| `yahoo` |
|
| 121 |
-
| `yahoo.start_date` / `end_date` | Historical window; clamped per Yahoo interval limits |
|
| 122 |
-
| `yahoo.auto_adjust` | Passed to `yfinance` |
|
| 123 |
-
| `execution.broker_api` | `paper` \| `alpaca_paper` \| `alpaca_live` |
|
| 124 |
-
| `finrl.algorithm` | PPO, A2C, DDPG, TD3 |
|
| 125 |
-
|
| 126 |
-
---
|
| 127 |
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
**Repository:** <https://github.com/ParallelLLC/algorithmic_trading>
|
|
|
|
| 1 |
+
# Backtest Reality Check
|
| 2 |
|
| 3 |
+
**algotrader 2.0 — a backtester that tries to prove itself wrong.**
|
| 4 |
|
| 5 |
+
Most backtesting tools answer *"how much would this have made?"*. That is the easy
|
| 6 |
+
question, and the answer is almost always flattering. This one answers the question
|
| 7 |
+
you actually need before risking money: **how much of that was luck?**
|
| 8 |
|
| 9 |
+
```bash
|
| 10 |
+
pip install -r requirements-space.txt
|
| 11 |
+
python app.py # the Gradio app on localhost:7860
|
| 12 |
+
python -m algotrader.cli lab --symbol SPY --strategy sma_cross
|
| 13 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
+
<sub>The v1 agentic trading system (FinRL, Alpaca, Yahoo ingest, Streamlit/Dash UIs) is
|
| 16 |
+
unchanged and still lives here — see [docs/AGENTIC_SYSTEM_V1.md](docs/AGENTIC_SYSTEM_V1.md).</sub>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
---
|
| 19 |
|
| 20 |
+
## The four ways a backtest lies
|
| 21 |
+
|
| 22 |
+
| The lie | The test | Where |
|
| 23 |
+
|---|---|---|
|
| 24 |
+
| The market had no structure to find | Monte-Carlo **permutation test** — re-run your rule on hundreds of shuffled markets | `algotrader/validation/permutation.py` |
|
| 25 |
+
| You tried 200 things and reported the best | **Deflated Sharpe Ratio** — charge for every variant you tried | `algotrader/validation/deflated_sharpe.py` |
|
| 26 |
+
| The parameters were fitted to the past | **PBO** (CSCV) and **walk-forward** | `algotrader/validation/pbo.py`, `walkforward.py` |
|
| 27 |
+
| The edge is smaller than the costs | **Cost stress test** at 3× friction | `algotrader/lab.py` |
|
| 28 |
+
|
| 29 |
+
Each feeds a single **Reality Score** out of 100 with a grade from A to F:
|
| 30 |
+
|
| 31 |
+
| Weight | Component | What it measures |
|
| 32 |
+
|---:|---|---|
|
| 33 |
+
| 30% | Significance | How far outside the shuffled-market null the result sits |
|
| 34 |
+
| 25% | Selection | Deflated Sharpe — does it clear the best-of-N bar |
|
| 35 |
+
| 20% | Walk-forward | How much of the tuned Sharpe survived trading forward |
|
| 36 |
+
| 15% | Overfitting | 1 − PBO |
|
| 37 |
+
| 10% | Robustness | Sharpe retained when costs triple |
|
| 38 |
+
|
| 39 |
+
The scale is deliberately harsh. On most markets, plain buy & hold beats every
|
| 40 |
+
strategy in the arena on evidence, and the built-in coin-flip control out-ranks
|
| 41 |
+
several respectable-looking rules. That is the finding, not a bug.
|
| 42 |
+
|
| 43 |
+
## The permutation test, concretely
|
| 44 |
+
|
| 45 |
+
We take the real price series and shuffle it. Each bar's gap, high, low, body and
|
| 46 |
+
volume are kept intact, but their **order** is destroyed. The result is a market
|
| 47 |
+
with the same volatility and the same fat tails, and no exploitable structure at
|
| 48 |
+
all. Then we re-run *your exact rule* on hundreds of these shuffled markets.
|
| 49 |
+
|
| 50 |
+
If your Sharpe sits inside that cloud, your rule found nothing that a coin-flip
|
| 51 |
+
market would not also have handed it. The p-value is the share of shuffled markets
|
| 52 |
+
that did as well or better.
|
| 53 |
+
|
| 54 |
+
Block mode resamples contiguous chunks instead of single bars, preserving
|
| 55 |
+
short-horizon momentum and volatility clustering — a harder null that trend
|
| 56 |
+
strategies deserve to be held to.
|
| 57 |
+
|
| 58 |
+
## No look-ahead, by construction
|
| 59 |
+
|
| 60 |
+
A strategy emits a target exposure at each bar's close using only data up to that
|
| 61 |
+
bar. The engine holds `position[t] = target[t - lag]` with `lag >= 1`, so a signal
|
| 62 |
+
computed on Tuesday's close cannot earn Tuesday's move.
|
| 63 |
+
|
| 64 |
+
That is the single line where look-ahead could enter, and the test suite asserts it
|
| 65 |
+
from four directions — including that truncating the data never changes the equity
|
| 66 |
+
curve before the cut, and that a `lag=0` request is refused outright.
|
| 67 |
+
|
| 68 |
+
## Python API
|
| 69 |
+
|
| 70 |
+
```python
|
| 71 |
+
from algotrader import LabConfig, run_lab
|
| 72 |
+
|
| 73 |
+
report = run_lab(LabConfig(
|
| 74 |
+
symbol="SPY",
|
| 75 |
+
start="2015-01-01",
|
| 76 |
+
strategy="sma_cross",
|
| 77 |
+
params={"fast": 20, "slow": 100},
|
| 78 |
+
commission_bps=1.0,
|
| 79 |
+
slippage_bps=2.0,
|
| 80 |
+
n_permutations=500,
|
| 81 |
+
))
|
| 82 |
+
|
| 83 |
+
print(report.verdict["grade"], report.verdict["score"])
|
| 84 |
+
print("p-value ", report.permutation.p_value)
|
| 85 |
+
print("deflated Sharpe ", report.dsr["dsr"])
|
| 86 |
+
print("overfit prob. ", report.pbo["pbo"])
|
| 87 |
+
print("walk-forward eff.", report.walkforward["efficiency"])
|
| 88 |
+
for flag in report.verdict["flags"]:
|
| 89 |
+
print(" !", flag)
|
| 90 |
+
```
|
| 91 |
|
| 92 |
+
Lower-level pieces compose on their own:
|
| 93 |
|
| 94 |
+
```python
|
| 95 |
+
from algotrader import load_ohlcv, run_backtest, get_strategy
|
| 96 |
+
from algotrader.types import CostModel
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
+
market = load_ohlcv("BTC-USD", "2018-01-01")
|
| 99 |
+
strategy = get_strategy("donchian_breakout")
|
| 100 |
+
result = run_backtest(
|
| 101 |
+
market.df,
|
| 102 |
+
strategy.generate(market.df, {"window": 55}),
|
| 103 |
+
costs=CostModel(commission_bps=1, slippage_bps=5, short_borrow_bps=50),
|
| 104 |
+
)
|
| 105 |
+
print(result.metrics["sharpe"], result.metrics["max_drawdown"])
|
| 106 |
+
```
|
| 107 |
|
| 108 |
+
## CLI
|
| 109 |
|
| 110 |
+
```bash
|
| 111 |
+
python -m algotrader.cli strategies # list the zoo
|
| 112 |
+
python -m algotrader.cli lab --symbol NVDA --strategy rsi_reversion --permutations 500
|
| 113 |
+
python -m algotrader.cli lab --symbol SPY --param fast=10 --param slow=50 --json
|
| 114 |
+
python -m algotrader.cli arena --symbol BTC-USD --start 2018-01-01
|
| 115 |
+
```
|
| 116 |
|
| 117 |
+
`--source synthetic` forces the offline simulator, which makes runs fully
|
| 118 |
+
deterministic and network-free.
|
| 119 |
|
| 120 |
+
## The strategy zoo
|
| 121 |
|
| 122 |
+
`buy_and_hold` · `sma_cross` · `ema_cross` · `macd_trend` · `rsi_reversion` ·
|
| 123 |
+
`bollinger_reversion` · `donchian_breakout` · `momentum` · `vol_target_momentum` ·
|
| 124 |
+
`channel_trend` · `coin_flip`
|
| 125 |
|
| 126 |
+
Buy & hold and the coin flip are controls, and they stay in the arena on purpose: a
|
| 127 |
+
leaderboard without a control group is marketing, not measurement.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
|
| 129 |
+
Adding one is a function and a registry entry — see `algotrader/strategies.py`.
|
| 130 |
|
| 131 |
+
## Data
|
| 132 |
|
| 133 |
+
Live prices come from Yahoo Finance. When the network is unavailable or rate-limited,
|
| 134 |
+
the app falls back to a deterministic market simulator — regime switching, Student-t
|
| 135 |
+
innovations, persistent volatility — and says so on every result. Naive geometric
|
| 136 |
+
Brownian motion flatters strategies; this simulator does not.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
+
Set `ALGOTRADER_OFFLINE=1` to skip network access entirely.
|
| 139 |
|
| 140 |
+
## Deploying the Hugging Face Space
|
| 141 |
|
| 142 |
+
The Space ships `app.py` plus the `algotrader` package and nothing else, so it builds
|
| 143 |
+
in well under a minute:
|
| 144 |
|
| 145 |
```bash
|
| 146 |
+
HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <your-username>/backtest-reality-check
|
|
|
|
|
|
|
|
|
|
|
|
|
| 147 |
```
|
| 148 |
|
| 149 |
+
Or set the `HF_TOKEN` secret and `HF_SPACE_ID` variable on the repository and let
|
| 150 |
+
`.github/workflows/sync-hf-space.yml` publish on every push to `main`. The workflow
|
| 151 |
+
runs the test suite and builds the app before it publishes anything.
|
| 152 |
|
| 153 |
+
`SPACE_README.md` is the Space card (with the Hugging Face YAML frontmatter);
|
| 154 |
+
`requirements-space.txt` is its dependency set. The root `requirements.txt` still
|
| 155 |
+
carries the full v1 stack for CI, Docker and the FinRL agents.
|
| 156 |
+
|
| 157 |
+
## Tests
|
|
|
|
|
|
|
| 158 |
|
| 159 |
```bash
|
| 160 |
+
python -m pytest tests/test_v2_*.py -q # 111 tests, ~15s, no network
|
|
|
|
|
|
|
| 161 |
```
|
| 162 |
|
| 163 |
+
The validation tests check both directions, which is the part that matters: the
|
| 164 |
+
statistics must reject noise **and** detect a real edge. They build a market with
|
| 165 |
+
genuine serial correlation and assert that the permutation test finds it, that PBO
|
| 166 |
+
stays near 0.5 on pure noise and drops below 0.15 when one variant is genuinely
|
| 167 |
+
better, and that walk-forward efficiency survives.
|
| 168 |
|
| 169 |
+
## References
|
| 170 |
|
| 171 |
+
- Bailey & López de Prado (2014), *The Deflated Sharpe Ratio: Correcting for Selection
|
| 172 |
+
Bias, Backtest Overfitting and Non-Normality*
|
| 173 |
+
- Bailey, Borwein, López de Prado & Zhu (2016), *The Probability of Backtest Overfitting*
|
| 174 |
+
- Masters (2018), *Permutation and Randomization Tests for Trading System Development*
|
| 175 |
|
| 176 |
+
## License
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 177 |
|
| 178 |
+
Apache-2.0. Research tooling, not investment advice. Nothing here is a
|
| 179 |
+
recommendation to trade.
|
|
|
|
@@ -0,0 +1,101 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: Backtest Reality Check
|
| 3 |
+
emoji: 🎲
|
| 4 |
+
colorFrom: blue
|
| 5 |
+
colorTo: gray
|
| 6 |
+
sdk: gradio
|
| 7 |
+
sdk_version: 5.49.1
|
| 8 |
+
app_file: app.py
|
| 9 |
+
pinned: true
|
| 10 |
+
license: apache-2.0
|
| 11 |
+
short_description: Your backtest is probably lying to you. This proves it.
|
| 12 |
+
tags:
|
| 13 |
+
- finance
|
| 14 |
+
- quantitative-finance
|
| 15 |
+
- algorithmic-trading
|
| 16 |
+
- backtesting
|
| 17 |
+
- statistics
|
| 18 |
+
- time-series
|
| 19 |
+
---
|
| 20 |
+
|
| 21 |
+
# Backtest Reality Check
|
| 22 |
+
|
| 23 |
+
**Your backtest is probably lying to you.**
|
| 24 |
+
|
| 25 |
+
Pick a market and a trading rule. This Space runs the backtest — and then spends
|
| 26 |
+
the rest of its effort trying to prove the result was luck.
|
| 27 |
+
|
| 28 |
+
Most backtesting tools answer *"how much would this have made?"*. That is the easy
|
| 29 |
+
question, and the answer is almost always flattering. This one answers the question
|
| 30 |
+
you need before risking money: **how much of that was luck?**
|
| 31 |
+
|
| 32 |
+
## The four ways a backtest lies, and the test for each
|
| 33 |
+
|
| 34 |
+
| The lie | The test |
|
| 35 |
+
|---|---|
|
| 36 |
+
| The market had no structure to find | **Permutation test** — re-run your rule on hundreds of shuffled markets |
|
| 37 |
+
| You tried 200 things and reported the best | **Deflated Sharpe Ratio** — charge for every variant you tried |
|
| 38 |
+
| The parameters were fitted to the past | **PBO + walk-forward** — does the in-sample winner keep winning? |
|
| 39 |
+
| The edge is smaller than the costs | **Cost stress test** — triple the friction and see what survives |
|
| 40 |
+
|
| 41 |
+
Each contributes to a single **Reality Score** out of 100, with a grade from A to F.
|
| 42 |
+
The scale is deliberately harsh. Most strategies people post online score below 40.
|
| 43 |
+
|
| 44 |
+
## Try this first
|
| 45 |
+
|
| 46 |
+
Run the **Arena** tab on `SPY`. On most markets and most date ranges, plain
|
| 47 |
+
**buy & hold** tops the leaderboard, and the **coin flip** control out-ranks
|
| 48 |
+
several respectable-looking strategies. That is not a bug in the app — it is the
|
| 49 |
+
finding.
|
| 50 |
+
|
| 51 |
+
## How the permutation test works
|
| 52 |
+
|
| 53 |
+
We take the real price series and shuffle it. Each bar's gap, high, low, body and
|
| 54 |
+
volume are kept intact, but their **order** is destroyed. The result is a market
|
| 55 |
+
with the same volatility and the same fat tails, and no exploitable structure at
|
| 56 |
+
all. Then we re-run *your exact rule* on hundreds of these shuffled markets.
|
| 57 |
+
|
| 58 |
+
If your Sharpe ratio sits comfortably inside that cloud, your rule found nothing
|
| 59 |
+
a coin-flip market would not also have handed it.
|
| 60 |
+
|
| 61 |
+
## No look-ahead, by construction
|
| 62 |
+
|
| 63 |
+
A strategy emits a target exposure at each bar's close using only data up to that
|
| 64 |
+
bar. The engine holds `position[t] = target[t - lag]` with `lag >= 1`, so a signal
|
| 65 |
+
computed on Tuesday's close cannot earn Tuesday's move. That is the single line
|
| 66 |
+
where look-ahead could enter, and the test suite asserts it directly.
|
| 67 |
+
|
| 68 |
+
## Use it from Python
|
| 69 |
+
|
| 70 |
+
```python
|
| 71 |
+
from algotrader import LabConfig, run_lab
|
| 72 |
+
|
| 73 |
+
report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross"))
|
| 74 |
+
print(report.verdict["grade"], report.verdict["score"])
|
| 75 |
+
print(report.permutation.p_value, report.dsr["dsr"], report.pbo["pbo"])
|
| 76 |
+
```
|
| 77 |
+
|
| 78 |
+
Or from the command line:
|
| 79 |
+
|
| 80 |
+
```bash
|
| 81 |
+
python -m algotrader.cli lab --symbol SPY --strategy donchian_breakout --permutations 500
|
| 82 |
+
python -m algotrader.cli arena --symbol BTC-USD
|
| 83 |
+
```
|
| 84 |
+
|
| 85 |
+
## Data
|
| 86 |
+
|
| 87 |
+
Live prices come from Yahoo Finance. When the network is unavailable or rate-limited,
|
| 88 |
+
the app falls back to a deterministic market simulator with regime switching, fat
|
| 89 |
+
tails and volatility clustering — and says so, clearly, on every result. The
|
| 90 |
+
statistics remain valid; they are just measured on a simulated market.
|
| 91 |
+
|
| 92 |
+
## References
|
| 93 |
+
|
| 94 |
+
- Bailey & López de Prado (2014), *The Deflated Sharpe Ratio*
|
| 95 |
+
- Bailey, Borwein, López de Prado & Zhu (2016), *The Probability of Backtest Overfitting*
|
| 96 |
+
- Masters (2018), *Permutation and Randomization Tests for Trading System Development*
|
| 97 |
+
|
| 98 |
+
---
|
| 99 |
+
|
| 100 |
+
Apache-2.0. Research tooling, not investment advice. Nothing here is a
|
| 101 |
+
recommendation to trade.
|
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""algotrader 2.0 — a backtester that tries to prove itself wrong.
|
| 2 |
+
|
| 3 |
+
Most backtesting libraries answer "how much would this have made?". This one
|
| 4 |
+
answers the question that actually matters before you risk money: "how much of
|
| 5 |
+
that was luck?"
|
| 6 |
+
|
| 7 |
+
Quick start::
|
| 8 |
+
|
| 9 |
+
from algotrader import LabConfig, run_lab
|
| 10 |
+
|
| 11 |
+
report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross"))
|
| 12 |
+
print(report.verdict["verdict"])
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
from .data import load_ohlcv, simulate_ohlcv
|
| 16 |
+
from .engine import run_backtest
|
| 17 |
+
from .lab import LabConfig, LabReport, run_arena, run_lab
|
| 18 |
+
from .metrics import compute_metrics
|
| 19 |
+
from .strategies import REGISTRY, get_strategy, list_strategies
|
| 20 |
+
from .types import BacktestResult, CostModel, MarketData
|
| 21 |
+
from .verdict import reality_score
|
| 22 |
+
|
| 23 |
+
__version__ = "2.0.0"
|
| 24 |
+
|
| 25 |
+
__all__ = [
|
| 26 |
+
"__version__",
|
| 27 |
+
"LabConfig",
|
| 28 |
+
"LabReport",
|
| 29 |
+
"run_lab",
|
| 30 |
+
"run_arena",
|
| 31 |
+
"run_backtest",
|
| 32 |
+
"compute_metrics",
|
| 33 |
+
"load_ohlcv",
|
| 34 |
+
"simulate_ohlcv",
|
| 35 |
+
"get_strategy",
|
| 36 |
+
"list_strategies",
|
| 37 |
+
"REGISTRY",
|
| 38 |
+
"BacktestResult",
|
| 39 |
+
"CostModel",
|
| 40 |
+
"MarketData",
|
| 41 |
+
"reality_score",
|
| 42 |
+
]
|
|
@@ -0,0 +1,286 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Plotly figures for the Lab.
|
| 2 |
+
|
| 3 |
+
Colour system (dark surface, validated for CVD separation):
|
| 4 |
+
|
| 5 |
+
* **blue** is always *your strategy's honest result* — the realised equity
|
| 6 |
+
curve, the out-of-sample fold, the observed Sharpe.
|
| 7 |
+
* **orange** is always *the thing it is measured against* — buy & hold, the
|
| 8 |
+
in-sample fold, the null distribution.
|
| 9 |
+
|
| 10 |
+
Holding that mapping across every figure means a reader learns it once.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
from typing import Dict, Optional
|
| 16 |
+
|
| 17 |
+
import numpy as np
|
| 18 |
+
import pandas as pd
|
| 19 |
+
import plotly.graph_objects as go
|
| 20 |
+
|
| 21 |
+
SURFACE = "#1a1a19"
|
| 22 |
+
PAGE = "#0d0d0d"
|
| 23 |
+
INK = "#ffffff"
|
| 24 |
+
INK_SECONDARY = "#c3c2b7"
|
| 25 |
+
INK_MUTED = "#898781"
|
| 26 |
+
GRID = "#2c2c2a"
|
| 27 |
+
AXIS = "#383835"
|
| 28 |
+
|
| 29 |
+
SUBJECT = "#3987e5" # categorical slot 1
|
| 30 |
+
REFERENCE = "#d95926" # categorical slot 2
|
| 31 |
+
NEGATIVE = "#e66767" # negative arm of the diverging pair (drawdowns)
|
| 32 |
+
|
| 33 |
+
FONT = 'system-ui, -apple-system, "Segoe UI", sans-serif'
|
| 34 |
+
|
| 35 |
+
_EMPTY_NOTE = "Run an analysis to populate this chart."
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def _base_layout(title: str, height: int = 340, **kwargs) -> dict:
|
| 39 |
+
return dict(
|
| 40 |
+
title=dict(text=title, font=dict(size=15, color=INK), x=0, xanchor="left", pad=dict(b=8)),
|
| 41 |
+
paper_bgcolor=PAGE,
|
| 42 |
+
plot_bgcolor=SURFACE,
|
| 43 |
+
font=dict(family=FONT, size=12, color=INK_SECONDARY),
|
| 44 |
+
height=height,
|
| 45 |
+
margin=dict(l=56, r=24, t=48, b=40),
|
| 46 |
+
hovermode="x unified",
|
| 47 |
+
hoverlabel=dict(bgcolor=SURFACE, bordercolor=AXIS, font=dict(color=INK, family=FONT)),
|
| 48 |
+
xaxis=dict(gridcolor=GRID, linecolor=AXIS, zeroline=False, tickfont=dict(color=INK_MUTED)),
|
| 49 |
+
yaxis=dict(gridcolor=GRID, linecolor=AXIS, zeroline=False, tickfont=dict(color=INK_MUTED)),
|
| 50 |
+
legend=dict(
|
| 51 |
+
orientation="h", yanchor="bottom", y=1.02, xanchor="left", x=0,
|
| 52 |
+
font=dict(color=INK_SECONDARY, size=11), bgcolor="rgba(0,0,0,0)",
|
| 53 |
+
),
|
| 54 |
+
**kwargs,
|
| 55 |
+
)
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
def empty_figure(message: str = _EMPTY_NOTE, height: int = 340) -> go.Figure:
|
| 59 |
+
fig = go.Figure()
|
| 60 |
+
fig.update_layout(**_base_layout("", height=height))
|
| 61 |
+
fig.update_xaxes(visible=False)
|
| 62 |
+
fig.update_yaxes(visible=False)
|
| 63 |
+
fig.add_annotation(
|
| 64 |
+
text=message, showarrow=False, xref="paper", yref="paper", x=0.5, y=0.5,
|
| 65 |
+
font=dict(color=INK_MUTED, size=13),
|
| 66 |
+
)
|
| 67 |
+
return fig
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
def equity_chart(report) -> go.Figure:
|
| 71 |
+
"""Strategy equity against buy & hold, both indexed to the same start."""
|
| 72 |
+
bt = report.backtest
|
| 73 |
+
strat = bt.equity / bt.equity.iloc[0] * 100.0
|
| 74 |
+
bench = bt.benchmark_equity / bt.benchmark_equity.iloc[0] * 100.0
|
| 75 |
+
|
| 76 |
+
fig = go.Figure()
|
| 77 |
+
fig.add_trace(
|
| 78 |
+
go.Scatter(
|
| 79 |
+
x=bench.index, y=bench.to_numpy(), name="Buy & hold", mode="lines",
|
| 80 |
+
line=dict(color=REFERENCE, width=2, dash="dash"),
|
| 81 |
+
hovertemplate="Buy & hold %{y:.1f}<extra></extra>",
|
| 82 |
+
)
|
| 83 |
+
)
|
| 84 |
+
fig.add_trace(
|
| 85 |
+
go.Scatter(
|
| 86 |
+
x=strat.index, y=strat.to_numpy(), name=report.strategy.name, mode="lines",
|
| 87 |
+
line=dict(color=SUBJECT, width=2),
|
| 88 |
+
hovertemplate=report.strategy.name + " %{y:.1f}<extra></extra>",
|
| 89 |
+
)
|
| 90 |
+
)
|
| 91 |
+
|
| 92 |
+
# Direct-label the two endpoints; the axis and tooltip carry everything else.
|
| 93 |
+
for series, color, label in ((strat, SUBJECT, report.strategy.name), (bench, REFERENCE, "Buy & hold")):
|
| 94 |
+
fig.add_annotation(
|
| 95 |
+
x=series.index[-1], y=float(series.iloc[-1]),
|
| 96 |
+
text=f" {label}: {series.iloc[-1]:.0f}", showarrow=False,
|
| 97 |
+
xanchor="left", font=dict(color=color, size=11),
|
| 98 |
+
)
|
| 99 |
+
|
| 100 |
+
fig.update_layout(**_base_layout("Growth of 100 (net of costs)", height=360))
|
| 101 |
+
fig.update_layout(margin=dict(l=56, r=140, t=48, b=40))
|
| 102 |
+
return fig
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def drawdown_chart(report) -> go.Figure:
|
| 106 |
+
"""Underwater plot — how deep, and for how long."""
|
| 107 |
+
from .metrics import drawdown_series
|
| 108 |
+
|
| 109 |
+
dd = drawdown_series(report.backtest.equity) * 100.0
|
| 110 |
+
fig = go.Figure(
|
| 111 |
+
go.Scatter(
|
| 112 |
+
x=dd.index, y=dd.to_numpy(), mode="lines", name="Drawdown",
|
| 113 |
+
line=dict(color=NEGATIVE, width=2), fill="tozeroy",
|
| 114 |
+
fillcolor="rgba(230,103,103,0.18)",
|
| 115 |
+
hovertemplate="Drawdown %{y:.1f}%<extra></extra>",
|
| 116 |
+
)
|
| 117 |
+
)
|
| 118 |
+
trough = float(dd.min())
|
| 119 |
+
fig.add_annotation(
|
| 120 |
+
x=dd.idxmin(), y=trough, text=f"worst {trough:.1f}%", showarrow=True,
|
| 121 |
+
arrowhead=0, arrowcolor=AXIS, ay=24, font=dict(color=INK_SECONDARY, size=11),
|
| 122 |
+
)
|
| 123 |
+
fig.update_layout(**_base_layout("Drawdown", height=240, showlegend=False))
|
| 124 |
+
fig.update_yaxes(ticksuffix="%")
|
| 125 |
+
return fig
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
def permutation_chart(report) -> go.Figure:
|
| 129 |
+
"""The headline chart: your Sharpe against Sharpes from shuffled markets."""
|
| 130 |
+
perm = report.permutation
|
| 131 |
+
if perm is None or perm.null.size == 0:
|
| 132 |
+
return empty_figure("Permutation test was skipped.", height=320)
|
| 133 |
+
|
| 134 |
+
null = perm.null
|
| 135 |
+
fig = go.Figure()
|
| 136 |
+
fig.add_trace(
|
| 137 |
+
go.Histogram(
|
| 138 |
+
x=null, name="Shuffled markets (no real edge)", nbinsx=44,
|
| 139 |
+
marker=dict(color="rgba(217,89,38,0.55)", line=dict(color=REFERENCE, width=1)),
|
| 140 |
+
hovertemplate="Sharpe %{x:.2f}<br>%{y} shuffles<extra></extra>",
|
| 141 |
+
)
|
| 142 |
+
)
|
| 143 |
+
|
| 144 |
+
top = np.histogram(null, bins=44)[0].max() if null.size else 1
|
| 145 |
+
fig.add_trace(
|
| 146 |
+
go.Scatter(
|
| 147 |
+
x=[perm.observed, perm.observed], y=[0, top * 1.08], mode="lines",
|
| 148 |
+
name="Your strategy", line=dict(color=SUBJECT, width=2),
|
| 149 |
+
hovertemplate="Your Sharpe %{x:.2f}<extra></extra>",
|
| 150 |
+
)
|
| 151 |
+
)
|
| 152 |
+
fig.add_annotation(
|
| 153 |
+
x=perm.observed, y=top * 1.08, text=f" your Sharpe {perm.observed:.2f}",
|
| 154 |
+
showarrow=False, xanchor="left", font=dict(color=SUBJECT, size=11),
|
| 155 |
+
)
|
| 156 |
+
|
| 157 |
+
beats = (null >= perm.observed).mean() * 100.0
|
| 158 |
+
fig.update_layout(
|
| 159 |
+
**_base_layout(
|
| 160 |
+
f"Permutation test — {beats:.0f}% of structure-free markets did this well or better "
|
| 161 |
+
f"(p = {perm.p_value:.3f})",
|
| 162 |
+
height=320,
|
| 163 |
+
)
|
| 164 |
+
)
|
| 165 |
+
fig.update_layout(hovermode="closest", bargap=0.02)
|
| 166 |
+
fig.update_xaxes(title=dict(text="Annualised Sharpe ratio", font=dict(color=INK_MUTED, size=11)))
|
| 167 |
+
# Headroom so the "your Sharpe" label never collides with the plot edge.
|
| 168 |
+
fig.update_yaxes(
|
| 169 |
+
title=dict(text="Shuffled markets", font=dict(color=INK_MUTED, size=11)),
|
| 170 |
+
range=[0, top * 1.28],
|
| 171 |
+
)
|
| 172 |
+
return fig
|
| 173 |
+
|
| 174 |
+
|
| 175 |
+
def walkforward_chart(report) -> go.Figure:
|
| 176 |
+
"""In-sample vs out-of-sample Sharpe, fold by fold."""
|
| 177 |
+
folds = report.walkforward.get("folds") or []
|
| 178 |
+
if not folds:
|
| 179 |
+
return empty_figure(report.walkforward.get("note") or _EMPTY_NOTE, height=300)
|
| 180 |
+
|
| 181 |
+
labels = [f"Fold {f['fold']}<br><span style='font-size:10px'>{f['test_start'][:7]}</span>" for f in folds]
|
| 182 |
+
fig = go.Figure()
|
| 183 |
+
fig.add_trace(
|
| 184 |
+
go.Bar(
|
| 185 |
+
x=labels, y=[f["is_sharpe"] for f in folds], name="In-sample (tuned)",
|
| 186 |
+
marker=dict(color=REFERENCE, line=dict(color=SURFACE, width=2)),
|
| 187 |
+
hovertemplate="In-sample Sharpe %{y:.2f}<extra></extra>",
|
| 188 |
+
)
|
| 189 |
+
)
|
| 190 |
+
fig.add_trace(
|
| 191 |
+
go.Bar(
|
| 192 |
+
x=labels, y=[f["oos_sharpe"] for f in folds], name="Out-of-sample (blind)",
|
| 193 |
+
marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
|
| 194 |
+
hovertemplate="Out-of-sample Sharpe %{y:.2f}<extra></extra>",
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
eff = report.walkforward.get("efficiency", 0.0)
|
| 198 |
+
fig.update_layout(
|
| 199 |
+
**_base_layout(f"Walk-forward — {eff:.0%} of the tuned Sharpe survived out of sample", height=300)
|
| 200 |
+
)
|
| 201 |
+
fig.update_layout(barmode="group", bargap=0.35, bargroupgap=0.08, hovermode="x unified")
|
| 202 |
+
fig.add_hline(y=0, line=dict(color=AXIS, width=1))
|
| 203 |
+
return fig
|
| 204 |
+
|
| 205 |
+
|
| 206 |
+
def score_chart(verdict: Dict[str, object]) -> go.Figure:
|
| 207 |
+
"""The five components behind the Reality Score."""
|
| 208 |
+
components = verdict.get("components") or {}
|
| 209 |
+
if not components:
|
| 210 |
+
return empty_figure(height=260)
|
| 211 |
+
|
| 212 |
+
pretty = {
|
| 213 |
+
"significance": "Beats shuffled markets",
|
| 214 |
+
"selection": "Survives selection bias",
|
| 215 |
+
"walk_forward": "Holds up walking forward",
|
| 216 |
+
"overfitting": "Not overfit (PBO)",
|
| 217 |
+
"robustness": "Survives 3x costs",
|
| 218 |
+
}
|
| 219 |
+
keys = list(pretty)
|
| 220 |
+
values = [float(components.get(k, 0.0)) for k in keys]
|
| 221 |
+
|
| 222 |
+
fig = go.Figure(
|
| 223 |
+
go.Bar(
|
| 224 |
+
x=values, y=[pretty[k] for k in keys], orientation="h",
|
| 225 |
+
marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
|
| 226 |
+
text=[f"{v:.0f}" for v in values], textposition="outside",
|
| 227 |
+
textfont=dict(color=INK_SECONDARY, size=11),
|
| 228 |
+
hovertemplate="%{y}: %{x:.0f}/100<extra></extra>",
|
| 229 |
+
)
|
| 230 |
+
)
|
| 231 |
+
fig.update_layout(**_base_layout("Where the score comes from", height=260, showlegend=False))
|
| 232 |
+
fig.update_layout(margin=dict(l=190, r=48, t=48, b=32), hovermode="closest")
|
| 233 |
+
fig.update_xaxes(range=[0, 108], tickvals=[0, 25, 50, 75, 100])
|
| 234 |
+
fig.update_yaxes(autorange="reversed")
|
| 235 |
+
return fig
|
| 236 |
+
|
| 237 |
+
|
| 238 |
+
def arena_chart(table: pd.DataFrame) -> go.Figure:
|
| 239 |
+
"""Leaderboard bars. One measure, one colour — the table carries the rest."""
|
| 240 |
+
if table is None or table.empty:
|
| 241 |
+
return empty_figure(height=380)
|
| 242 |
+
|
| 243 |
+
ordered = table.iloc[::-1]
|
| 244 |
+
fig = go.Figure(
|
| 245 |
+
go.Bar(
|
| 246 |
+
x=ordered["Sharpe"].to_numpy(), y=ordered["Strategy"].tolist(), orientation="h",
|
| 247 |
+
marker=dict(color=SUBJECT, line=dict(color=SURFACE, width=2)),
|
| 248 |
+
customdata=np.column_stack([ordered["p-value"].to_numpy(), ordered["DSR"].to_numpy()]),
|
| 249 |
+
hovertemplate="%{y}<br>Sharpe %{x:.2f}<br>p = %{customdata[0]:.3f}"
|
| 250 |
+
"<br>Deflated Sharpe %{customdata[1]:.2f}<extra></extra>",
|
| 251 |
+
)
|
| 252 |
+
)
|
| 253 |
+
# Direct-label only what matters: the ones that actually cleared significance.
|
| 254 |
+
for _, row in ordered.iterrows():
|
| 255 |
+
if np.isfinite(row["p-value"]) and row["p-value"] < 0.05:
|
| 256 |
+
fig.add_annotation(
|
| 257 |
+
x=row["Sharpe"], y=row["Strategy"], text=" p < 0.05", showarrow=False,
|
| 258 |
+
xanchor="left" if row["Sharpe"] >= 0 else "right",
|
| 259 |
+
font=dict(color=INK_SECONDARY, size=10),
|
| 260 |
+
)
|
| 261 |
+
fig.update_layout(
|
| 262 |
+
**_base_layout(
|
| 263 |
+
"Strategy arena — Sharpe ratio, ordered by strength of evidence",
|
| 264 |
+
height=max(300, 42 * len(table)),
|
| 265 |
+
showlegend=False,
|
| 266 |
+
)
|
| 267 |
+
)
|
| 268 |
+
fig.update_layout(margin=dict(l=180, r=96, t=48, b=32), hovermode="closest")
|
| 269 |
+
fig.add_vline(x=0, line=dict(color=AXIS, width=1))
|
| 270 |
+
return fig
|
| 271 |
+
|
| 272 |
+
|
| 273 |
+
def exposure_chart(report) -> go.Figure:
|
| 274 |
+
"""What the strategy was actually holding, over time."""
|
| 275 |
+
pos = report.backtest.position
|
| 276 |
+
fig = go.Figure(
|
| 277 |
+
go.Scatter(
|
| 278 |
+
x=pos.index, y=pos.to_numpy(), mode="lines", name="Exposure",
|
| 279 |
+
line=dict(color=SUBJECT, width=2, shape="hv"), fill="tozeroy",
|
| 280 |
+
fillcolor="rgba(57,135,229,0.16)",
|
| 281 |
+
hovertemplate="Exposure %{y:.2f}x<extra></extra>",
|
| 282 |
+
)
|
| 283 |
+
)
|
| 284 |
+
fig.update_layout(**_base_layout("Position held", height=200, showlegend=False))
|
| 285 |
+
fig.add_hline(y=0, line=dict(color=AXIS, width=1))
|
| 286 |
+
return fig
|
|
@@ -0,0 +1,191 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Command-line interface.
|
| 2 |
+
|
| 3 |
+
python -m algotrader.cli lab --symbol SPY --strategy sma_cross
|
| 4 |
+
python -m algotrader.cli arena --symbol BTC-USD --start 2018-01-01
|
| 5 |
+
python -m algotrader.cli strategies
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
import argparse
|
| 11 |
+
import json
|
| 12 |
+
import sys
|
| 13 |
+
from typing import Dict, List
|
| 14 |
+
|
| 15 |
+
from . import __version__
|
| 16 |
+
from .lab import LabConfig, run_arena, run_lab
|
| 17 |
+
from .strategies import REGISTRY, get_strategy
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
def _parse_params(pairs: List[str] | None) -> Dict[str, float]:
|
| 21 |
+
params: Dict[str, float] = {}
|
| 22 |
+
for pair in pairs or []:
|
| 23 |
+
if "=" not in pair:
|
| 24 |
+
raise SystemExit(f"--param expects name=value, got '{pair}'")
|
| 25 |
+
name, _, value = pair.partition("=")
|
| 26 |
+
params[name.strip()] = float(value)
|
| 27 |
+
return params
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def _add_common(parser: argparse.ArgumentParser) -> None:
|
| 31 |
+
parser.add_argument("--symbol", default="SPY")
|
| 32 |
+
parser.add_argument("--start", default="2015-01-01")
|
| 33 |
+
parser.add_argument("--end", default=None)
|
| 34 |
+
parser.add_argument("--interval", default="1d")
|
| 35 |
+
parser.add_argument(
|
| 36 |
+
"--source", default="auto", choices=["auto", "cache", "synthetic"],
|
| 37 |
+
help="'auto' downloads and falls back offline; 'synthetic' forces the simulator.",
|
| 38 |
+
)
|
| 39 |
+
parser.add_argument("--commission-bps", type=float, default=1.0)
|
| 40 |
+
parser.add_argument("--slippage-bps", type=float, default=2.0)
|
| 41 |
+
parser.add_argument("--no-short", action="store_true", help="Long/flat only.")
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def _config_from(args: argparse.Namespace, **overrides) -> LabConfig:
|
| 45 |
+
return LabConfig(
|
| 46 |
+
symbol=args.symbol,
|
| 47 |
+
start=args.start,
|
| 48 |
+
end=args.end,
|
| 49 |
+
interval=args.interval,
|
| 50 |
+
source=args.source,
|
| 51 |
+
commission_bps=args.commission_bps,
|
| 52 |
+
slippage_bps=args.slippage_bps,
|
| 53 |
+
allow_short=not args.no_short,
|
| 54 |
+
**overrides,
|
| 55 |
+
)
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
def _cmd_lab(args: argparse.Namespace) -> int:
|
| 59 |
+
cfg = _config_from(
|
| 60 |
+
args,
|
| 61 |
+
strategy=args.strategy,
|
| 62 |
+
params=_parse_params(args.param),
|
| 63 |
+
n_permutations=args.permutations,
|
| 64 |
+
permutation_method=args.null,
|
| 65 |
+
wf_folds=args.folds,
|
| 66 |
+
)
|
| 67 |
+
progress = None if args.quiet else (lambda f, m: print(f" [{f:5.0%}] {m}", file=sys.stderr))
|
| 68 |
+
report = run_lab(cfg, progress=progress)
|
| 69 |
+
|
| 70 |
+
if args.json:
|
| 71 |
+
payload = {
|
| 72 |
+
"symbol": report.market.symbol,
|
| 73 |
+
"source": report.market.source,
|
| 74 |
+
"strategy": report.strategy.key,
|
| 75 |
+
"params": report.params,
|
| 76 |
+
"metrics": report.backtest.metrics,
|
| 77 |
+
"benchmark_metrics": report.backtest.benchmark_metrics,
|
| 78 |
+
"p_value": report.permutation.p_value if report.permutation else None,
|
| 79 |
+
"deflated_sharpe": report.dsr.get("dsr"),
|
| 80 |
+
"pbo": report.pbo.get("pbo"),
|
| 81 |
+
"walkforward_efficiency": report.walkforward.get("efficiency"),
|
| 82 |
+
"cost_stress": report.cost_stress,
|
| 83 |
+
"verdict": {k: v for k, v in report.verdict.items()},
|
| 84 |
+
}
|
| 85 |
+
print(json.dumps(payload, indent=2, default=str))
|
| 86 |
+
return 0
|
| 87 |
+
|
| 88 |
+
v, m, b = report.verdict, report.backtest.metrics, report.backtest.benchmark_metrics
|
| 89 |
+
bar = "=" * 66
|
| 90 |
+
print(f"\n{bar}")
|
| 91 |
+
print(f" {report.strategy.name} on {report.market.symbol} [{report.market.source} data]")
|
| 92 |
+
print(f" {report.market.start.date()} to {report.market.end.date()} · {len(report.market.df):,} bars")
|
| 93 |
+
print(bar)
|
| 94 |
+
print(f" REALITY SCORE {v['score']:.1f} / 100 GRADE {v['grade']}")
|
| 95 |
+
print(f" {v['headline']}")
|
| 96 |
+
print(bar)
|
| 97 |
+
print(f" Total return {m['total_return']:>9.1%} buy & hold {b['total_return']:>8.1%}")
|
| 98 |
+
print(f" CAGR {m['cagr']:>9.1%} buy & hold {b['cagr']:>8.1%}")
|
| 99 |
+
print(f" Sharpe {m['sharpe']:>9.2f} buy & hold {b['sharpe']:>8.2f}")
|
| 100 |
+
print(f" Max drawdown {m['max_drawdown']:>9.1%}")
|
| 101 |
+
print(f" Trades {int(m.get('n_trades', 0)):>9,}")
|
| 102 |
+
print(bar)
|
| 103 |
+
if report.permutation:
|
| 104 |
+
print(f" Permutation p {report.permutation.p_value:>9.3f} ({report.permutation.n_permutations} shuffled markets)")
|
| 105 |
+
print(f" Deflated Sharpe {report.dsr.get('dsr', 0):>9.2f} (after {report.trials.get('n', 1)} variants)")
|
| 106 |
+
pbo = report.pbo.get("pbo")
|
| 107 |
+
print(f" Overfit prob. {pbo:>9.2f}" if pbo == pbo else " Overfit prob. n/a")
|
| 108 |
+
print(f" Walk-forward eff. {report.walkforward.get('efficiency', 0):>9.2f}")
|
| 109 |
+
print(f" Sharpe at 3x cost {report.cost_stress.get('sharpe_3x', 0):>9.2f}")
|
| 110 |
+
print(bar)
|
| 111 |
+
for flag in v["flags"]:
|
| 112 |
+
print(f" ! {flag}")
|
| 113 |
+
if v["flags"]:
|
| 114 |
+
print(bar)
|
| 115 |
+
print(f" {v['verdict']}\n")
|
| 116 |
+
return 0
|
| 117 |
+
|
| 118 |
+
|
| 119 |
+
def _cmd_arena(args: argparse.Namespace) -> int:
|
| 120 |
+
cfg = _config_from(args)
|
| 121 |
+
progress = None if args.quiet else (lambda f, m: print(f" [{f:5.0%}] {m}", file=sys.stderr))
|
| 122 |
+
table, market, _ = run_arena(cfg, n_permutations=args.permutations, progress=progress)
|
| 123 |
+
|
| 124 |
+
if args.json:
|
| 125 |
+
print(table.to_json(orient="records", indent=2))
|
| 126 |
+
return 0
|
| 127 |
+
|
| 128 |
+
print(f"\n {market.symbol} [{market.source} data] "
|
| 129 |
+
f"{market.start.date()} to {market.end.date()}\n")
|
| 130 |
+
display = table.drop(columns=["key"]).copy()
|
| 131 |
+
for col in ("Return", "CAGR", "MaxDD"):
|
| 132 |
+
display[col] = display[col].map("{:.1%}".format)
|
| 133 |
+
for col in ("Sharpe", "DSR", "Evidence"):
|
| 134 |
+
display[col] = display[col].map("{:.2f}".format)
|
| 135 |
+
display["p-value"] = display["p-value"].map(lambda v: "—" if v != v else f"{v:.3f}")
|
| 136 |
+
print(display.to_string(index=False))
|
| 137 |
+
print("\n Ranked by evidence = (1 - p) x deflated Sharpe, not by return.\n")
|
| 138 |
+
return 0
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
def _cmd_strategies(args: argparse.Namespace) -> int:
|
| 142 |
+
for key, strategy in REGISTRY.items():
|
| 143 |
+
params = ", ".join(f"{p.name}={p.default:g}" for p in strategy.params) or "no parameters"
|
| 144 |
+
print(f" {key:<22} {strategy.name:<26} [{strategy.family}]")
|
| 145 |
+
print(f" {'':<22} {strategy.description}")
|
| 146 |
+
print(f" {'':<22} defaults: {params}\n")
|
| 147 |
+
return 0
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
def main(argv: List[str] | None = None) -> int:
|
| 151 |
+
parser = argparse.ArgumentParser(
|
| 152 |
+
prog="algotrader",
|
| 153 |
+
description="Backtest a trading rule, then try to prove the result was luck.",
|
| 154 |
+
)
|
| 155 |
+
parser.add_argument("--version", action="version", version=f"algotrader {__version__}")
|
| 156 |
+
sub = parser.add_subparsers(dest="command", required=True)
|
| 157 |
+
|
| 158 |
+
lab = sub.add_parser("lab", help="Full reality check for one strategy.")
|
| 159 |
+
_add_common(lab)
|
| 160 |
+
lab.add_argument("--strategy", default="sma_cross", choices=sorted(REGISTRY))
|
| 161 |
+
lab.add_argument("--param", action="append", metavar="NAME=VALUE",
|
| 162 |
+
help="Override a strategy parameter. Repeatable.")
|
| 163 |
+
lab.add_argument("--permutations", type=int, default=250)
|
| 164 |
+
lab.add_argument("--null", default="permute", choices=["permute", "block"])
|
| 165 |
+
lab.add_argument("--folds", type=int, default=5)
|
| 166 |
+
lab.add_argument("--json", action="store_true")
|
| 167 |
+
lab.add_argument("--quiet", "-q", action="store_true")
|
| 168 |
+
lab.set_defaults(func=_cmd_lab)
|
| 169 |
+
|
| 170 |
+
arena = sub.add_parser("arena", help="Race every strategy on one market.")
|
| 171 |
+
_add_common(arena)
|
| 172 |
+
arena.add_argument("--permutations", type=int, default=120)
|
| 173 |
+
arena.add_argument("--json", action="store_true")
|
| 174 |
+
arena.add_argument("--quiet", "-q", action="store_true")
|
| 175 |
+
arena.set_defaults(func=_cmd_arena)
|
| 176 |
+
|
| 177 |
+
listing = sub.add_parser("strategies", help="List the strategy zoo.")
|
| 178 |
+
listing.set_defaults(func=_cmd_strategies)
|
| 179 |
+
|
| 180 |
+
args = parser.parse_args(argv)
|
| 181 |
+
try:
|
| 182 |
+
return args.func(args)
|
| 183 |
+
except KeyboardInterrupt:
|
| 184 |
+
return 130
|
| 185 |
+
except Exception as exc: # noqa: BLE001
|
| 186 |
+
print(f"error: {exc}", file=sys.stderr)
|
| 187 |
+
return 1
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
if __name__ == "__main__":
|
| 191 |
+
raise SystemExit(main())
|
|
@@ -0,0 +1,246 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Market data loading with a three-tier fallback.
|
| 2 |
+
|
| 3 |
+
Order of preference: live Yahoo download -> on-disk cache -> a deterministic
|
| 4 |
+
simulator. The fallback exists because a Hugging Face Space that shows a
|
| 5 |
+
stack trace on the first click is a Space nobody shares. When the simulator is
|
| 6 |
+
used, :class:`~algotrader.types.MarketData` says so and the UI shows it.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from __future__ import annotations
|
| 10 |
+
|
| 11 |
+
import hashlib
|
| 12 |
+
import logging
|
| 13 |
+
import os
|
| 14 |
+
from dataclasses import dataclass
|
| 15 |
+
from pathlib import Path
|
| 16 |
+
from typing import Optional
|
| 17 |
+
|
| 18 |
+
import numpy as np
|
| 19 |
+
import pandas as pd
|
| 20 |
+
|
| 21 |
+
from .types import OHLCV_COLUMNS, MarketData
|
| 22 |
+
|
| 23 |
+
logger = logging.getLogger(__name__)
|
| 24 |
+
|
| 25 |
+
CACHE_DIR = Path(os.environ.get("ALGOTRADER_CACHE", Path.home() / ".cache" / "algotrader"))
|
| 26 |
+
NETWORK_ENABLED = os.environ.get("ALGOTRADER_OFFLINE", "").lower() not in ("1", "true", "yes")
|
| 27 |
+
|
| 28 |
+
# Popular tickers get hand-set simulation parameters so the offline demo is at
|
| 29 |
+
# least in the right postcode: annual drift, annual vol, and a starting price.
|
| 30 |
+
@dataclass(frozen=True)
|
| 31 |
+
class SimProfile:
|
| 32 |
+
drift: float
|
| 33 |
+
vol: float
|
| 34 |
+
price: float
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
SIM_PROFILES: dict[str, SimProfile] = {
|
| 38 |
+
"AAPL": SimProfile(0.24, 0.29, 190.0),
|
| 39 |
+
"MSFT": SimProfile(0.25, 0.27, 410.0),
|
| 40 |
+
"NVDA": SimProfile(0.55, 0.52, 120.0),
|
| 41 |
+
"TSLA": SimProfile(0.30, 0.58, 250.0),
|
| 42 |
+
"AMZN": SimProfile(0.22, 0.33, 180.0),
|
| 43 |
+
"GOOGL": SimProfile(0.20, 0.31, 170.0),
|
| 44 |
+
"META": SimProfile(0.26, 0.40, 500.0),
|
| 45 |
+
"SPY": SimProfile(0.10, 0.16, 550.0),
|
| 46 |
+
"QQQ": SimProfile(0.14, 0.21, 480.0),
|
| 47 |
+
"BTC-USD": SimProfile(0.45, 0.65, 65000.0),
|
| 48 |
+
"ETH-USD": SimProfile(0.35, 0.75, 3000.0),
|
| 49 |
+
"GLD": SimProfile(0.07, 0.14, 200.0),
|
| 50 |
+
"TLT": SimProfile(0.01, 0.15, 95.0),
|
| 51 |
+
}
|
| 52 |
+
|
| 53 |
+
DEFAULT_UNIVERSE = ["SPY", "AAPL", "NVDA", "MSFT", "TSLA", "QQQ", "BTC-USD", "GLD"]
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
def _seed_for(symbol: str) -> int:
|
| 57 |
+
"""Stable per-symbol seed so a given ticker always simulates identically."""
|
| 58 |
+
digest = hashlib.sha256(symbol.upper().encode()).digest()
|
| 59 |
+
return int.from_bytes(digest[:4], "big")
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def _normalise(df: pd.DataFrame) -> pd.DataFrame:
|
| 63 |
+
"""Coerce any loader's output into a clean lowercase OHLCV frame."""
|
| 64 |
+
if isinstance(df.columns, pd.MultiIndex):
|
| 65 |
+
df = df.copy()
|
| 66 |
+
df.columns = [str(c[0]) for c in df.columns]
|
| 67 |
+
df = df.rename(columns={c: str(c).strip().lower().replace(" ", "_") for c in df.columns})
|
| 68 |
+
if "adj_close" in df.columns and "close" not in df.columns:
|
| 69 |
+
df = df.rename(columns={"adj_close": "close"})
|
| 70 |
+
missing = [c for c in OHLCV_COLUMNS if c not in df.columns]
|
| 71 |
+
for col in missing:
|
| 72 |
+
if col == "volume":
|
| 73 |
+
df["volume"] = 0.0
|
| 74 |
+
elif "close" in df.columns:
|
| 75 |
+
df[col] = df["close"]
|
| 76 |
+
else:
|
| 77 |
+
raise ValueError(f"Price data is missing required column: {col}")
|
| 78 |
+
df = df.loc[:, list(OHLCV_COLUMNS)].astype(float)
|
| 79 |
+
if not isinstance(df.index, pd.DatetimeIndex):
|
| 80 |
+
df.index = pd.to_datetime(df.index)
|
| 81 |
+
df.index = df.index.tz_localize(None) if df.index.tz is not None else df.index
|
| 82 |
+
df = df[~df.index.duplicated(keep="last")].sort_index()
|
| 83 |
+
df = df[df["close"] > 0].dropna(subset=["close"])
|
| 84 |
+
return df
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
def _cache_path(symbol: str, interval: str) -> Path:
|
| 88 |
+
safe = symbol.upper().replace("/", "_")
|
| 89 |
+
return CACHE_DIR / f"{safe}_{interval}.csv"
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
def _read_cache(symbol: str, interval: str) -> Optional[pd.DataFrame]:
|
| 93 |
+
path = _cache_path(symbol, interval)
|
| 94 |
+
if not path.exists():
|
| 95 |
+
return None
|
| 96 |
+
try:
|
| 97 |
+
return _normalise(pd.read_csv(path, index_col=0, parse_dates=True))
|
| 98 |
+
except Exception as exc: # pragma: no cover - corrupted cache is not worth failing over
|
| 99 |
+
logger.warning("Ignoring unreadable cache %s: %s", path, exc)
|
| 100 |
+
return None
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
def _write_cache(symbol: str, interval: str, df: pd.DataFrame) -> None:
|
| 104 |
+
try:
|
| 105 |
+
CACHE_DIR.mkdir(parents=True, exist_ok=True)
|
| 106 |
+
df.to_csv(_cache_path(symbol, interval))
|
| 107 |
+
except Exception as exc: # pragma: no cover - a read-only FS must not break the app
|
| 108 |
+
logger.warning("Could not write cache for %s: %s", symbol, exc)
|
| 109 |
+
|
| 110 |
+
|
| 111 |
+
def _download(symbol: str, start: str, end: str | None, interval: str) -> Optional[pd.DataFrame]:
|
| 112 |
+
if not NETWORK_ENABLED:
|
| 113 |
+
return None
|
| 114 |
+
try:
|
| 115 |
+
import yfinance as yf
|
| 116 |
+
except ImportError:
|
| 117 |
+
logger.info("yfinance not installed; using offline data")
|
| 118 |
+
return None
|
| 119 |
+
try:
|
| 120 |
+
raw = yf.download(
|
| 121 |
+
symbol,
|
| 122 |
+
start=start,
|
| 123 |
+
end=end,
|
| 124 |
+
interval=interval,
|
| 125 |
+
progress=False,
|
| 126 |
+
auto_adjust=True,
|
| 127 |
+
threads=False,
|
| 128 |
+
)
|
| 129 |
+
except Exception as exc:
|
| 130 |
+
logger.warning("Download failed for %s: %s", symbol, exc)
|
| 131 |
+
return None
|
| 132 |
+
if raw is None or len(raw) == 0:
|
| 133 |
+
logger.warning("Download for %s returned no rows", symbol)
|
| 134 |
+
return None
|
| 135 |
+
try:
|
| 136 |
+
return _normalise(raw)
|
| 137 |
+
except Exception as exc:
|
| 138 |
+
logger.warning("Could not normalise download for %s: %s", symbol, exc)
|
| 139 |
+
return None
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def simulate_ohlcv(
|
| 143 |
+
symbol: str = "SIM",
|
| 144 |
+
start: str = "2015-01-01",
|
| 145 |
+
end: str | None = None,
|
| 146 |
+
interval: str = "1d",
|
| 147 |
+
seed: Optional[int] = None,
|
| 148 |
+
) -> pd.DataFrame:
|
| 149 |
+
"""Generate a deterministic but realistic-looking OHLCV series.
|
| 150 |
+
|
| 151 |
+
This is not geometric Brownian motion with a straight face: it uses a
|
| 152 |
+
two-state (calm / stressed) regime switch, Student-t innovations and
|
| 153 |
+
GARCH-ish vol persistence, so the resulting series has fat tails and
|
| 154 |
+
volatility clustering. That matters, because a strategy tested against
|
| 155 |
+
naive GBM looks far better than it deserves to.
|
| 156 |
+
"""
|
| 157 |
+
profile = SIM_PROFILES.get(symbol.upper(), SimProfile(0.08, 0.25, 100.0))
|
| 158 |
+
rng = np.random.default_rng(_seed_for(symbol) if seed is None else seed)
|
| 159 |
+
|
| 160 |
+
freq = {"1d": "B", "1wk": "W-FRI", "1h": "h"}.get(interval, "B")
|
| 161 |
+
index = pd.date_range(start=start, end=end or pd.Timestamp.today().normalize(), freq=freq)
|
| 162 |
+
n = len(index)
|
| 163 |
+
if n < 50:
|
| 164 |
+
raise ValueError("Simulated range is too short to backtest")
|
| 165 |
+
|
| 166 |
+
ppy = 252 if freq in ("B", "h") else 52
|
| 167 |
+
mu = profile.drift / ppy
|
| 168 |
+
base_vol = profile.vol / np.sqrt(ppy)
|
| 169 |
+
|
| 170 |
+
# Regime chain: calm state is sticky, stressed state is short and violent.
|
| 171 |
+
p_calm_to_stress, p_stress_to_calm = 0.01, 0.06
|
| 172 |
+
regime = np.zeros(n, dtype=int)
|
| 173 |
+
for i in range(1, n):
|
| 174 |
+
flip = rng.random()
|
| 175 |
+
if regime[i - 1] == 0:
|
| 176 |
+
regime[i] = 1 if flip < p_calm_to_stress else 0
|
| 177 |
+
else:
|
| 178 |
+
regime[i] = 0 if flip < p_stress_to_calm else 1
|
| 179 |
+
|
| 180 |
+
# Persistent vol around a regime-dependent level.
|
| 181 |
+
vol = np.empty(n)
|
| 182 |
+
level = np.where(regime == 1, base_vol * 2.4, base_vol * 0.9)
|
| 183 |
+
vol[0] = level[0]
|
| 184 |
+
for i in range(1, n):
|
| 185 |
+
vol[i] = 0.92 * vol[i - 1] + 0.08 * level[i]
|
| 186 |
+
|
| 187 |
+
shocks = rng.standard_t(df=4, size=n) / np.sqrt(2.0) # unit-ish variance, fat tails
|
| 188 |
+
drift = np.where(regime == 1, mu - 3.0 * base_vol**2, mu)
|
| 189 |
+
log_ret = drift + vol * shocks
|
| 190 |
+
close = profile.price * np.exp(np.cumsum(log_ret))
|
| 191 |
+
close = close * (profile.price / close[-1]) # end near the quoted level
|
| 192 |
+
|
| 193 |
+
intrabar = vol * rng.uniform(0.3, 1.1, size=n)
|
| 194 |
+
open_ = close * np.exp(-log_ret * rng.uniform(0.2, 0.8, size=n))
|
| 195 |
+
high = np.maximum(open_, close) * np.exp(np.abs(intrabar))
|
| 196 |
+
low = np.minimum(open_, close) * np.exp(-np.abs(intrabar))
|
| 197 |
+
volume = rng.lognormal(mean=15.5, sigma=0.45, size=n) * (1.0 + 3.0 * regime)
|
| 198 |
+
|
| 199 |
+
return _normalise(
|
| 200 |
+
pd.DataFrame(
|
| 201 |
+
{"open": open_, "high": high, "low": low, "close": close, "volume": volume},
|
| 202 |
+
index=index,
|
| 203 |
+
)
|
| 204 |
+
)
|
| 205 |
+
|
| 206 |
+
|
| 207 |
+
def load_ohlcv(
|
| 208 |
+
symbol: str = "SPY",
|
| 209 |
+
start: str = "2015-01-01",
|
| 210 |
+
end: str | None = None,
|
| 211 |
+
interval: str = "1d",
|
| 212 |
+
source: str = "auto",
|
| 213 |
+
) -> MarketData:
|
| 214 |
+
"""Load OHLCV for ``symbol``, never raising for a merely-unreachable network.
|
| 215 |
+
|
| 216 |
+
``source`` is one of ``auto`` (download, then cache, then simulate),
|
| 217 |
+
``cache``, or ``synthetic``.
|
| 218 |
+
"""
|
| 219 |
+
symbol = (symbol or "SPY").strip().upper()
|
| 220 |
+
|
| 221 |
+
if source == "synthetic":
|
| 222 |
+
df = simulate_ohlcv(symbol, start, end, interval)
|
| 223 |
+
return MarketData(symbol, df, "synthetic", interval, "Simulated prices (requested).")
|
| 224 |
+
|
| 225 |
+
if source in ("auto", "live"):
|
| 226 |
+
df = _download(symbol, start, end, interval)
|
| 227 |
+
if df is not None and len(df) > 50:
|
| 228 |
+
_write_cache(symbol, interval, df)
|
| 229 |
+
return MarketData(symbol, df, "yfinance", interval, "Live data from Yahoo Finance.")
|
| 230 |
+
|
| 231 |
+
cached = _read_cache(symbol, interval)
|
| 232 |
+
if cached is not None and len(cached) > 50:
|
| 233 |
+
window = cached.loc[str(start) : str(end)] if end else cached.loc[str(start) :]
|
| 234 |
+
if len(window) > 50:
|
| 235 |
+
return MarketData(symbol, window, "bundled", interval, "Cached data (network unavailable).")
|
| 236 |
+
|
| 237 |
+
df = simulate_ohlcv(symbol, start, end, interval)
|
| 238 |
+
return MarketData(
|
| 239 |
+
symbol,
|
| 240 |
+
df,
|
| 241 |
+
"synthetic",
|
| 242 |
+
interval,
|
| 243 |
+
f"Live data for {symbol} was unavailable, so this run uses a deterministic "
|
| 244 |
+
"market simulator with fat tails and volatility clustering. The statistics "
|
| 245 |
+
"below are still valid — they are just measured on a simulated market.",
|
| 246 |
+
)
|
|
@@ -0,0 +1,133 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Vectorised, look-ahead-free backtest engine.
|
| 2 |
+
|
| 3 |
+
Contract
|
| 4 |
+
--------
|
| 5 |
+
A strategy emits ``target[t]``: the exposure it wants, decided using only
|
| 6 |
+
information available at the close of bar ``t``. The engine holds
|
| 7 |
+
``position[t] = target[t - lag]`` during bar ``t`` and credits it with that
|
| 8 |
+
bar's close-to-close return. With the default ``lag=1`` this means "decide on
|
| 9 |
+
today's close, hold the position through tomorrow" -- the single place where
|
| 10 |
+
look-ahead could sneak in, and it is one line.
|
| 11 |
+
|
| 12 |
+
Costs are charged on exposure *changes*, so a strategy that flips daily pays
|
| 13 |
+
for it. Short exposure additionally accrues a borrow fee.
|
| 14 |
+
"""
|
| 15 |
+
|
| 16 |
+
from __future__ import annotations
|
| 17 |
+
|
| 18 |
+
from typing import Optional
|
| 19 |
+
|
| 20 |
+
import numpy as np
|
| 21 |
+
import pandas as pd
|
| 22 |
+
|
| 23 |
+
from .metrics import compute_metrics, infer_periods_per_year
|
| 24 |
+
from .types import BacktestResult, CostModel
|
| 25 |
+
|
| 26 |
+
__all__ = ["run_backtest", "bars_to_returns"]
|
| 27 |
+
|
| 28 |
+
|
| 29 |
+
def bars_to_returns(df: pd.DataFrame) -> pd.Series:
|
| 30 |
+
"""Close-to-close simple returns."""
|
| 31 |
+
return df["close"].astype(float).pct_change().fillna(0.0)
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def run_backtest(
|
| 35 |
+
df: pd.DataFrame,
|
| 36 |
+
target: pd.Series,
|
| 37 |
+
costs: CostModel | None = None,
|
| 38 |
+
lag: int = 1,
|
| 39 |
+
max_leverage: float = 1.0,
|
| 40 |
+
allow_short: bool = True,
|
| 41 |
+
initial_capital: float = 100_000.0,
|
| 42 |
+
periods_per_year: Optional[int] = None,
|
| 43 |
+
rf: float = 0.0,
|
| 44 |
+
meta: Optional[dict] = None,
|
| 45 |
+
) -> BacktestResult:
|
| 46 |
+
"""Run one backtest and return equity, returns and the full metric bundle."""
|
| 47 |
+
if df.empty:
|
| 48 |
+
raise ValueError("Cannot backtest an empty price frame")
|
| 49 |
+
if lag < 1:
|
| 50 |
+
raise ValueError("lag must be >= 1; lag=0 would trade on unavailable information")
|
| 51 |
+
|
| 52 |
+
costs = costs or CostModel()
|
| 53 |
+
ppy = periods_per_year or infer_periods_per_year(df.index)
|
| 54 |
+
|
| 55 |
+
asset_ret = bars_to_returns(df)
|
| 56 |
+
|
| 57 |
+
target = target.reindex(df.index).astype(float).fillna(0.0)
|
| 58 |
+
lower = -max_leverage if allow_short else 0.0
|
| 59 |
+
target = target.clip(lower, max_leverage)
|
| 60 |
+
|
| 61 |
+
position = target.shift(lag).fillna(0.0)
|
| 62 |
+
|
| 63 |
+
gross = position * asset_ret
|
| 64 |
+
|
| 65 |
+
traded = position.diff()
|
| 66 |
+
traded.iloc[0] = position.iloc[0]
|
| 67 |
+
trade_cost = traded.abs() * (costs.one_way_bps / 1e4)
|
| 68 |
+
|
| 69 |
+
borrow_cost = position.clip(upper=0.0).abs() * (costs.short_borrow_bps / 1e4) / ppy
|
| 70 |
+
total_cost = trade_cost + borrow_cost
|
| 71 |
+
|
| 72 |
+
net = gross - total_cost
|
| 73 |
+
equity = initial_capital * (1.0 + net).cumprod()
|
| 74 |
+
benchmark_equity = initial_capital * (1.0 + asset_ret).cumprod()
|
| 75 |
+
|
| 76 |
+
result = BacktestResult(
|
| 77 |
+
equity=equity,
|
| 78 |
+
returns=net,
|
| 79 |
+
gross_returns=gross,
|
| 80 |
+
position=position,
|
| 81 |
+
target=target,
|
| 82 |
+
costs=total_cost,
|
| 83 |
+
benchmark_equity=benchmark_equity,
|
| 84 |
+
metrics=compute_metrics(net, equity, position, ppy, rf),
|
| 85 |
+
benchmark_metrics=compute_metrics(asset_ret, benchmark_equity, None, ppy, rf),
|
| 86 |
+
meta={
|
| 87 |
+
"lag": lag,
|
| 88 |
+
"commission_bps": costs.commission_bps,
|
| 89 |
+
"slippage_bps": costs.slippage_bps,
|
| 90 |
+
"short_borrow_bps": costs.short_borrow_bps,
|
| 91 |
+
"max_leverage": max_leverage,
|
| 92 |
+
"allow_short": allow_short,
|
| 93 |
+
"initial_capital": initial_capital,
|
| 94 |
+
"periods_per_year": ppy,
|
| 95 |
+
**(meta or {}),
|
| 96 |
+
},
|
| 97 |
+
)
|
| 98 |
+
result.metrics["cost_drag_ann"] = float(total_cost.sum() / max(result.metrics.get("years", 1e-9), 1e-9))
|
| 99 |
+
result.metrics["gross_sharpe"] = float(
|
| 100 |
+
compute_metrics(gross, initial_capital * (1.0 + gross).cumprod(), None, ppy, rf).get("sharpe", 0.0)
|
| 101 |
+
)
|
| 102 |
+
return result
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def fast_sharpe(
|
| 106 |
+
asset_ret: np.ndarray,
|
| 107 |
+
target: np.ndarray,
|
| 108 |
+
one_way_bps: float,
|
| 109 |
+
lag: int,
|
| 110 |
+
periods_per_year: int,
|
| 111 |
+
) -> float:
|
| 112 |
+
"""Numpy-only Sharpe for hot loops (permutation tests, PBO grids).
|
| 113 |
+
|
| 114 |
+
Mirrors :func:`run_backtest` exactly for the no-borrow case; it exists only
|
| 115 |
+
because building a DataFrame 1000 times is the difference between a Space
|
| 116 |
+
that answers in 4 seconds and one nobody waits for.
|
| 117 |
+
"""
|
| 118 |
+
n = asset_ret.size
|
| 119 |
+
position = np.empty(n, dtype=float)
|
| 120 |
+
position[:lag] = 0.0
|
| 121 |
+
position[lag:] = target[:-lag] if lag else target
|
| 122 |
+
gross = position * asset_ret
|
| 123 |
+
traded = np.empty(n, dtype=float)
|
| 124 |
+
traded[0] = position[0]
|
| 125 |
+
traded[1:] = np.diff(position)
|
| 126 |
+
net = gross - np.abs(traded) * (one_way_bps / 1e4)
|
| 127 |
+
net = net[np.isfinite(net)]
|
| 128 |
+
if net.size < 2:
|
| 129 |
+
return 0.0
|
| 130 |
+
sd = net.std(ddof=1)
|
| 131 |
+
if not np.isfinite(sd) or sd < 1e-12:
|
| 132 |
+
return 0.0
|
| 133 |
+
return float(net.mean() / sd * np.sqrt(periods_per_year))
|
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Vectorised technical indicators.
|
| 2 |
+
|
| 3 |
+
Every function takes and returns pandas objects aligned to the input index, and
|
| 4 |
+
every one of them is causal: the value at bar ``t`` uses only data up to and
|
| 5 |
+
including ``t``. That property is what makes the backtest engine's single
|
| 6 |
+
``shift`` enough to guarantee no look-ahead.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from __future__ import annotations
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
import pandas as pd
|
| 13 |
+
|
| 14 |
+
__all__ = [
|
| 15 |
+
"sma",
|
| 16 |
+
"ema",
|
| 17 |
+
"rsi",
|
| 18 |
+
"macd",
|
| 19 |
+
"bollinger",
|
| 20 |
+
"atr",
|
| 21 |
+
"donchian",
|
| 22 |
+
"zscore",
|
| 23 |
+
"roc",
|
| 24 |
+
"realised_vol",
|
| 25 |
+
]
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
def sma(series: pd.Series, window: int) -> pd.Series:
|
| 29 |
+
return series.rolling(window, min_periods=window).mean()
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
def ema(series: pd.Series, window: int) -> pd.Series:
|
| 33 |
+
return series.ewm(span=window, adjust=False, min_periods=window).mean()
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def rsi(series: pd.Series, window: int = 14) -> pd.Series:
|
| 37 |
+
"""Wilder's RSI."""
|
| 38 |
+
delta = series.diff()
|
| 39 |
+
gain = delta.clip(lower=0.0)
|
| 40 |
+
loss = -delta.clip(upper=0.0)
|
| 41 |
+
avg_gain = gain.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
|
| 42 |
+
avg_loss = loss.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
|
| 43 |
+
rs = avg_gain / avg_loss.replace(0.0, np.nan)
|
| 44 |
+
out = 100.0 - (100.0 / (1.0 + rs))
|
| 45 |
+
# avg_loss == 0 leaves rs undefined: an all-gain window is RSI 100, and a
|
| 46 |
+
# perfectly flat window (no gains either) is RSI 50.
|
| 47 |
+
flat = (avg_gain == 0.0) & (avg_loss == 0.0)
|
| 48 |
+
out = out.mask((avg_loss == 0.0) & (avg_gain > 0.0), 100.0)
|
| 49 |
+
out = out.mask(flat, 50.0)
|
| 50 |
+
return out.where(avg_gain.notna())
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def macd(
|
| 54 |
+
series: pd.Series, fast: int = 12, slow: int = 26, signal: int = 9
|
| 55 |
+
) -> tuple[pd.Series, pd.Series, pd.Series]:
|
| 56 |
+
"""Returns ``(macd_line, signal_line, histogram)``."""
|
| 57 |
+
macd_line = ema(series, fast) - ema(series, slow)
|
| 58 |
+
signal_line = macd_line.ewm(span=signal, adjust=False, min_periods=signal).mean()
|
| 59 |
+
return macd_line, signal_line, macd_line - signal_line
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def bollinger(
|
| 63 |
+
series: pd.Series, window: int = 20, k: float = 2.0
|
| 64 |
+
) -> tuple[pd.Series, pd.Series, pd.Series]:
|
| 65 |
+
"""Returns ``(lower, middle, upper)``."""
|
| 66 |
+
mid = sma(series, window)
|
| 67 |
+
sd = series.rolling(window, min_periods=window).std(ddof=0)
|
| 68 |
+
return mid - k * sd, mid, mid + k * sd
|
| 69 |
+
|
| 70 |
+
|
| 71 |
+
def atr(df: pd.DataFrame, window: int = 14) -> pd.Series:
|
| 72 |
+
prev_close = df["close"].shift(1)
|
| 73 |
+
tr = pd.concat(
|
| 74 |
+
[
|
| 75 |
+
df["high"] - df["low"],
|
| 76 |
+
(df["high"] - prev_close).abs(),
|
| 77 |
+
(df["low"] - prev_close).abs(),
|
| 78 |
+
],
|
| 79 |
+
axis=1,
|
| 80 |
+
).max(axis=1)
|
| 81 |
+
return tr.ewm(alpha=1.0 / window, adjust=False, min_periods=window).mean()
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
def donchian(df: pd.DataFrame, window: int = 20) -> tuple[pd.Series, pd.Series]:
|
| 85 |
+
"""Rolling channel excluding the current bar, so a breakout test is causal."""
|
| 86 |
+
upper = df["high"].rolling(window, min_periods=window).max().shift(1)
|
| 87 |
+
lower = df["low"].rolling(window, min_periods=window).min().shift(1)
|
| 88 |
+
return lower, upper
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
def zscore(series: pd.Series, window: int = 20) -> pd.Series:
|
| 92 |
+
mean = series.rolling(window, min_periods=window).mean()
|
| 93 |
+
sd = series.rolling(window, min_periods=window).std(ddof=0)
|
| 94 |
+
return (series - mean) / sd.replace(0.0, np.nan)
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def roc(series: pd.Series, window: int = 20) -> pd.Series:
|
| 98 |
+
return series.pct_change(window)
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
def realised_vol(returns: pd.Series, window: int = 20, periods_per_year: int = 252) -> pd.Series:
|
| 102 |
+
return returns.rolling(window, min_periods=window).std(ddof=0) * np.sqrt(periods_per_year)
|
|
@@ -0,0 +1,328 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The Lab: one call that runs a backtest and then tries to disprove it.
|
| 2 |
+
|
| 3 |
+
This is the module both the Gradio Space and the CLI drive. Keeping the whole
|
| 4 |
+
pipeline here means the app and the command line can never disagree about what
|
| 5 |
+
a Reality Score means.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
import logging
|
| 11 |
+
from dataclasses import dataclass, field
|
| 12 |
+
from typing import Callable, Dict, List, Optional
|
| 13 |
+
|
| 14 |
+
import numpy as np
|
| 15 |
+
import pandas as pd
|
| 16 |
+
|
| 17 |
+
from .data import load_ohlcv
|
| 18 |
+
from .engine import run_backtest
|
| 19 |
+
from .metrics import infer_periods_per_year
|
| 20 |
+
from .strategies import Strategy, get_strategy, list_strategies
|
| 21 |
+
from .types import BacktestResult, CostModel, MarketData
|
| 22 |
+
from .validation.deflated_sharpe import deflated_sharpe_ratio, min_track_record_length
|
| 23 |
+
from .validation.pbo import probability_of_backtest_overfitting
|
| 24 |
+
from .validation.permutation import PermutationResult, permutation_test
|
| 25 |
+
from .validation.walkforward import walk_forward
|
| 26 |
+
from .verdict import reality_score
|
| 27 |
+
|
| 28 |
+
logger = logging.getLogger(__name__)
|
| 29 |
+
|
| 30 |
+
__all__ = ["LabConfig", "LabReport", "run_lab", "run_arena"]
|
| 31 |
+
|
| 32 |
+
ProgressFn = Optional[Callable[[float, str], None]]
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
@dataclass
|
| 36 |
+
class LabConfig:
|
| 37 |
+
symbol: str = "SPY"
|
| 38 |
+
start: str = "2015-01-01"
|
| 39 |
+
end: Optional[str] = None
|
| 40 |
+
interval: str = "1d"
|
| 41 |
+
source: str = "auto"
|
| 42 |
+
|
| 43 |
+
strategy: str = "sma_cross"
|
| 44 |
+
params: Dict[str, float] = field(default_factory=dict)
|
| 45 |
+
|
| 46 |
+
commission_bps: float = 1.0
|
| 47 |
+
slippage_bps: float = 2.0
|
| 48 |
+
short_borrow_bps: float = 50.0
|
| 49 |
+
lag: int = 1
|
| 50 |
+
allow_short: bool = True
|
| 51 |
+
max_leverage: float = 1.0
|
| 52 |
+
capital: float = 100_000.0
|
| 53 |
+
|
| 54 |
+
n_permutations: int = 250
|
| 55 |
+
permutation_method: str = "permute"
|
| 56 |
+
block_size: int = 20
|
| 57 |
+
wf_folds: int = 5
|
| 58 |
+
pbo_splits: int = 8
|
| 59 |
+
grid_limit: int = 40
|
| 60 |
+
seed: int = 0
|
| 61 |
+
|
| 62 |
+
def costs(self, multiplier: float = 1.0) -> CostModel:
|
| 63 |
+
return CostModel(
|
| 64 |
+
commission_bps=self.commission_bps * multiplier,
|
| 65 |
+
slippage_bps=self.slippage_bps * multiplier,
|
| 66 |
+
short_borrow_bps=self.short_borrow_bps * multiplier,
|
| 67 |
+
)
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
@dataclass
|
| 71 |
+
class LabReport:
|
| 72 |
+
config: LabConfig
|
| 73 |
+
market: MarketData
|
| 74 |
+
strategy: Strategy
|
| 75 |
+
params: Dict[str, float]
|
| 76 |
+
backtest: BacktestResult
|
| 77 |
+
permutation: Optional[PermutationResult] = None
|
| 78 |
+
dsr: Dict[str, float] = field(default_factory=dict)
|
| 79 |
+
pbo: Dict[str, object] = field(default_factory=dict)
|
| 80 |
+
walkforward: Dict[str, object] = field(default_factory=dict)
|
| 81 |
+
trials: Dict[str, object] = field(default_factory=dict)
|
| 82 |
+
verdict: Dict[str, object] = field(default_factory=dict)
|
| 83 |
+
cost_stress: Dict[str, float] = field(default_factory=dict)
|
| 84 |
+
benchmark_correlation: float = float("nan")
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
def _trial_matrix(
|
| 88 |
+
df: pd.DataFrame,
|
| 89 |
+
strategy: Strategy,
|
| 90 |
+
cfg: LabConfig,
|
| 91 |
+
progress: ProgressFn = None,
|
| 92 |
+
) -> tuple[np.ndarray, List[float], List[str]]:
|
| 93 |
+
"""Backtest every parameter combination a researcher would plausibly try.
|
| 94 |
+
|
| 95 |
+
The resulting ``T x N`` return matrix feeds both the Deflated Sharpe (how
|
| 96 |
+
many variants were tried, and how spread out were they) and PBO.
|
| 97 |
+
"""
|
| 98 |
+
grid = strategy.grid(limit=cfg.grid_limit)
|
| 99 |
+
costs = cfg.costs()
|
| 100 |
+
columns, sharpes, labels = [], [], []
|
| 101 |
+
|
| 102 |
+
for i, params in enumerate(grid):
|
| 103 |
+
target = strategy.generate(df, params)
|
| 104 |
+
result = run_backtest(
|
| 105 |
+
df, target, costs=costs, lag=cfg.lag,
|
| 106 |
+
max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
|
| 107 |
+
)
|
| 108 |
+
columns.append(result.returns.to_numpy(dtype=float))
|
| 109 |
+
sharpes.append(result.sharpe)
|
| 110 |
+
labels.append(", ".join(f"{k}={v}" for k, v in params.items()) or "default")
|
| 111 |
+
if progress is not None and i % 5 == 0:
|
| 112 |
+
progress((i + 1) / max(len(grid), 1), f"Variant {i + 1}/{len(grid)}")
|
| 113 |
+
|
| 114 |
+
matrix = np.column_stack(columns) if columns else np.zeros((len(df), 0))
|
| 115 |
+
return matrix, sharpes, labels
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def run_lab(cfg: LabConfig, progress: ProgressFn = None) -> LabReport:
|
| 119 |
+
"""Run the full honesty pipeline for one strategy on one symbol."""
|
| 120 |
+
|
| 121 |
+
def step(fraction: float, message: str) -> None:
|
| 122 |
+
if progress is not None:
|
| 123 |
+
progress(min(max(fraction, 0.0), 1.0), message)
|
| 124 |
+
|
| 125 |
+
step(0.02, "Loading market data")
|
| 126 |
+
market = load_ohlcv(cfg.symbol, cfg.start, cfg.end, cfg.interval, cfg.source)
|
| 127 |
+
df = market.df
|
| 128 |
+
if len(df) < 120:
|
| 129 |
+
raise ValueError(
|
| 130 |
+
f"Only {len(df)} bars available for {cfg.symbol}. "
|
| 131 |
+
"Widen the date range — anything shorter cannot be validated."
|
| 132 |
+
)
|
| 133 |
+
|
| 134 |
+
strategy = get_strategy(cfg.strategy)
|
| 135 |
+
params = strategy.clean(cfg.params)
|
| 136 |
+
ppy = infer_periods_per_year(df.index)
|
| 137 |
+
|
| 138 |
+
step(0.10, "Running the backtest")
|
| 139 |
+
target = strategy.generate(df, params)
|
| 140 |
+
backtest = run_backtest(
|
| 141 |
+
df,
|
| 142 |
+
target,
|
| 143 |
+
costs=cfg.costs(),
|
| 144 |
+
lag=cfg.lag,
|
| 145 |
+
max_leverage=cfg.max_leverage,
|
| 146 |
+
allow_short=cfg.allow_short,
|
| 147 |
+
initial_capital=cfg.capital,
|
| 148 |
+
periods_per_year=ppy,
|
| 149 |
+
meta={"symbol": market.symbol, "strategy": strategy.key, "params": params},
|
| 150 |
+
)
|
| 151 |
+
|
| 152 |
+
step(0.16, "Stress-testing costs")
|
| 153 |
+
stressed = run_backtest(
|
| 154 |
+
df, target, costs=cfg.costs(3.0), lag=cfg.lag,
|
| 155 |
+
max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
|
| 156 |
+
periods_per_year=ppy,
|
| 157 |
+
)
|
| 158 |
+
base_sharpe = backtest.sharpe
|
| 159 |
+
cost_stress_ratio = float(stressed.sharpe / base_sharpe) if base_sharpe > 1e-9 else 0.0
|
| 160 |
+
cost_stress = {
|
| 161 |
+
"sharpe_1x": base_sharpe,
|
| 162 |
+
"sharpe_3x": stressed.sharpe,
|
| 163 |
+
"ratio": cost_stress_ratio,
|
| 164 |
+
"return_3x": float(stressed.metrics.get("total_return", 0.0)),
|
| 165 |
+
}
|
| 166 |
+
|
| 167 |
+
step(0.22, "Backtesting every parameter variant")
|
| 168 |
+
matrix, trial_sharpes, labels = _trial_matrix(
|
| 169 |
+
df, strategy, cfg, lambda f, m: step(0.22 + 0.18 * f, m)
|
| 170 |
+
)
|
| 171 |
+
n_trials = max(len(trial_sharpes), 1)
|
| 172 |
+
|
| 173 |
+
step(0.42, "Deflating the Sharpe ratio for selection bias")
|
| 174 |
+
dsr = deflated_sharpe_ratio(
|
| 175 |
+
backtest.returns.to_numpy(dtype=float),
|
| 176 |
+
sharpe_annual=base_sharpe,
|
| 177 |
+
periods_per_year=ppy,
|
| 178 |
+
n_trials=n_trials,
|
| 179 |
+
trial_sharpes=trial_sharpes if n_trials > 1 else None,
|
| 180 |
+
)
|
| 181 |
+
mtrl = min_track_record_length(
|
| 182 |
+
dsr["sr_per_period"], dsr["n_obs"], dsr["skew"], dsr["kurtosis"],
|
| 183 |
+
benchmark=dsr["threshold_sr_per_period"],
|
| 184 |
+
)
|
| 185 |
+
dsr["min_track_record_bars"] = mtrl
|
| 186 |
+
dsr["min_track_record_years"] = float(mtrl / ppy) if np.isfinite(mtrl) else float("inf")
|
| 187 |
+
|
| 188 |
+
step(0.46, "Measuring backtest overfitting")
|
| 189 |
+
pbo = probability_of_backtest_overfitting(matrix, n_splits=cfg.pbo_splits, labels=labels)
|
| 190 |
+
|
| 191 |
+
step(0.50, "Shuffling the market")
|
| 192 |
+
permutation = None
|
| 193 |
+
if cfg.n_permutations > 0:
|
| 194 |
+
permutation = permutation_test(
|
| 195 |
+
df,
|
| 196 |
+
lambda frame: strategy.generate(frame, params),
|
| 197 |
+
n_permutations=cfg.n_permutations,
|
| 198 |
+
method=cfg.permutation_method,
|
| 199 |
+
block=cfg.block_size,
|
| 200 |
+
costs=cfg.costs(),
|
| 201 |
+
lag=cfg.lag,
|
| 202 |
+
max_leverage=cfg.max_leverage,
|
| 203 |
+
allow_short=cfg.allow_short,
|
| 204 |
+
seed=cfg.seed,
|
| 205 |
+
observed=base_sharpe,
|
| 206 |
+
progress=lambda f, m: step(0.50 + 0.32 * f, m),
|
| 207 |
+
)
|
| 208 |
+
|
| 209 |
+
step(0.84, "Walking the strategy forward")
|
| 210 |
+
wf = walk_forward(
|
| 211 |
+
df, strategy, n_folds=cfg.wf_folds, costs=cfg.costs(), lag=cfg.lag,
|
| 212 |
+
max_leverage=cfg.max_leverage, allow_short=cfg.allow_short,
|
| 213 |
+
grid_limit=min(cfg.grid_limit, 24),
|
| 214 |
+
progress=lambda f, m: step(0.84 + 0.12 * f, m),
|
| 215 |
+
)
|
| 216 |
+
|
| 217 |
+
bench_corr = float(
|
| 218 |
+
pd.Series(backtest.returns).corr(backtest.benchmark_equity.pct_change().fillna(0.0))
|
| 219 |
+
)
|
| 220 |
+
|
| 221 |
+
step(0.98, "Grading")
|
| 222 |
+
verdict = reality_score(
|
| 223 |
+
metrics=backtest.metrics,
|
| 224 |
+
benchmark_metrics=backtest.benchmark_metrics,
|
| 225 |
+
p_value=permutation.p_value if permutation else None,
|
| 226 |
+
dsr=dsr.get("dsr"),
|
| 227 |
+
pbo=pbo.get("pbo"),
|
| 228 |
+
wf_efficiency=wf.get("efficiency"),
|
| 229 |
+
wf_win_rate=wf.get("oos_win_rate"),
|
| 230 |
+
cost_stress_ratio=cost_stress_ratio,
|
| 231 |
+
benchmark_correlation=bench_corr,
|
| 232 |
+
)
|
| 233 |
+
|
| 234 |
+
step(1.0, "Done")
|
| 235 |
+
return LabReport(
|
| 236 |
+
config=cfg,
|
| 237 |
+
market=market,
|
| 238 |
+
strategy=strategy,
|
| 239 |
+
params=params,
|
| 240 |
+
backtest=backtest,
|
| 241 |
+
permutation=permutation,
|
| 242 |
+
dsr=dsr,
|
| 243 |
+
pbo=pbo,
|
| 244 |
+
walkforward=wf,
|
| 245 |
+
trials={"n": n_trials, "sharpes": trial_sharpes, "labels": labels, "matrix_shape": matrix.shape},
|
| 246 |
+
verdict=verdict,
|
| 247 |
+
cost_stress=cost_stress,
|
| 248 |
+
benchmark_correlation=bench_corr,
|
| 249 |
+
)
|
| 250 |
+
|
| 251 |
+
|
| 252 |
+
def run_arena(
|
| 253 |
+
cfg: LabConfig,
|
| 254 |
+
strategy_keys: Optional[List[str]] = None,
|
| 255 |
+
n_permutations: int = 120,
|
| 256 |
+
progress: ProgressFn = None,
|
| 257 |
+
) -> tuple[pd.DataFrame, MarketData, Dict[str, BacktestResult]]:
|
| 258 |
+
"""Race every strategy on the same market, ranked by evidence not returns.
|
| 259 |
+
|
| 260 |
+
Buy & hold and the coin flip stay in the field on purpose: a leaderboard
|
| 261 |
+
without a control group is marketing, not measurement.
|
| 262 |
+
"""
|
| 263 |
+
market = load_ohlcv(cfg.symbol, cfg.start, cfg.end, cfg.interval, cfg.source)
|
| 264 |
+
df = market.df
|
| 265 |
+
ppy = infer_periods_per_year(df.index)
|
| 266 |
+
costs = cfg.costs()
|
| 267 |
+
|
| 268 |
+
keys = strategy_keys or [s.key for s in list_strategies()]
|
| 269 |
+
rows, curves = [], {}
|
| 270 |
+
|
| 271 |
+
for i, key in enumerate(keys):
|
| 272 |
+
strategy = get_strategy(key)
|
| 273 |
+
params = strategy.defaults()
|
| 274 |
+
target = strategy.generate(df, params)
|
| 275 |
+
result = run_backtest(
|
| 276 |
+
df, target, costs=costs, lag=cfg.lag, max_leverage=cfg.max_leverage,
|
| 277 |
+
allow_short=cfg.allow_short, initial_capital=cfg.capital, periods_per_year=ppy,
|
| 278 |
+
)
|
| 279 |
+
curves[key] = result
|
| 280 |
+
|
| 281 |
+
p_value = None
|
| 282 |
+
if n_permutations > 0:
|
| 283 |
+
p_value = permutation_test(
|
| 284 |
+
df,
|
| 285 |
+
lambda frame, s=strategy, p=params: s.generate(frame, p),
|
| 286 |
+
n_permutations=n_permutations,
|
| 287 |
+
method=cfg.permutation_method,
|
| 288 |
+
block=cfg.block_size,
|
| 289 |
+
costs=costs,
|
| 290 |
+
lag=cfg.lag,
|
| 291 |
+
max_leverage=cfg.max_leverage,
|
| 292 |
+
allow_short=cfg.allow_short,
|
| 293 |
+
seed=cfg.seed,
|
| 294 |
+
observed=result.sharpe,
|
| 295 |
+
).p_value
|
| 296 |
+
|
| 297 |
+
grid_size = len(strategy.grid(limit=cfg.grid_limit))
|
| 298 |
+
dsr = deflated_sharpe_ratio(
|
| 299 |
+
result.returns.to_numpy(dtype=float),
|
| 300 |
+
sharpe_annual=result.sharpe,
|
| 301 |
+
periods_per_year=ppy,
|
| 302 |
+
n_trials=grid_size,
|
| 303 |
+
)
|
| 304 |
+
|
| 305 |
+
rows.append(
|
| 306 |
+
{
|
| 307 |
+
"Strategy": strategy.name,
|
| 308 |
+
"key": key,
|
| 309 |
+
"Family": strategy.family,
|
| 310 |
+
"Return": result.metrics.get("total_return", 0.0),
|
| 311 |
+
"CAGR": result.metrics.get("cagr", 0.0),
|
| 312 |
+
"Sharpe": result.sharpe,
|
| 313 |
+
"MaxDD": result.metrics.get("max_drawdown", 0.0),
|
| 314 |
+
"Trades": int(result.metrics.get("n_trades", 0)),
|
| 315 |
+
"p-value": p_value if p_value is not None else float("nan"),
|
| 316 |
+
"DSR": dsr["dsr"],
|
| 317 |
+
}
|
| 318 |
+
)
|
| 319 |
+
if progress is not None:
|
| 320 |
+
progress((i + 1) / len(keys), f"{strategy.name} ({i + 1}/{len(keys)})")
|
| 321 |
+
|
| 322 |
+
table = pd.DataFrame(rows)
|
| 323 |
+
if not table.empty:
|
| 324 |
+
# Rank by evidence: a high Sharpe with a p-value of 0.4 is not a win.
|
| 325 |
+
table["Evidence"] = (1.0 - table["p-value"].fillna(0.5)) * table["DSR"]
|
| 326 |
+
table = table.sort_values("Evidence", ascending=False).reset_index(drop=True)
|
| 327 |
+
table.insert(0, "#", table.index + 1)
|
| 328 |
+
return table, market, curves
|
|
@@ -0,0 +1,157 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Performance and risk metrics.
|
| 2 |
+
|
| 3 |
+
All ratios are computed from *net* per-bar returns and annualised with the
|
| 4 |
+
periodicity inferred from the index, so daily / hourly / minute series all get
|
| 5 |
+
comparable numbers.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
from typing import Dict
|
| 11 |
+
|
| 12 |
+
import numpy as np
|
| 13 |
+
import pandas as pd
|
| 14 |
+
|
| 15 |
+
__all__ = [
|
| 16 |
+
"infer_periods_per_year",
|
| 17 |
+
"sharpe_ratio",
|
| 18 |
+
"sortino_ratio",
|
| 19 |
+
"max_drawdown",
|
| 20 |
+
"drawdown_series",
|
| 21 |
+
"compute_metrics",
|
| 22 |
+
]
|
| 23 |
+
|
| 24 |
+
_SECONDS_PER_YEAR = 365.25 * 24 * 3600
|
| 25 |
+
_TRADING_DAYS = 252
|
| 26 |
+
|
| 27 |
+
# A return stream with dispersion below this is constant to floating-point
|
| 28 |
+
# noise. Without an absolute floor, a flat series divides by ~1e-19 and reports
|
| 29 |
+
# a Sharpe of 1e16 -- the exact kind of nonsense number this project exists to
|
| 30 |
+
# catch, so it must not originate here.
|
| 31 |
+
_DEGENERATE_SD = 1e-12
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def infer_periods_per_year(index: pd.Index) -> int:
|
| 35 |
+
"""Guess bars-per-year from an index, defaulting to daily trading bars."""
|
| 36 |
+
if not isinstance(index, pd.DatetimeIndex) or len(index) < 3:
|
| 37 |
+
return _TRADING_DAYS
|
| 38 |
+
nanos = index.to_numpy(dtype="datetime64[ns]").astype("int64")
|
| 39 |
+
deltas = np.diff(nanos) / 1e9 # seconds
|
| 40 |
+
deltas = deltas[deltas > 0]
|
| 41 |
+
if deltas.size == 0:
|
| 42 |
+
return _TRADING_DAYS
|
| 43 |
+
step = float(np.median(deltas))
|
| 44 |
+
if step >= 20 * 3600: # daily or slower -> use trading-day convention
|
| 45 |
+
days = step / 86400.0
|
| 46 |
+
return max(1, int(round(_TRADING_DAYS / max(days / 1.4, 1.0))))
|
| 47 |
+
# Intraday: assume a 6.5h session, 252 days a year.
|
| 48 |
+
bars_per_session = (6.5 * 3600) / step
|
| 49 |
+
return max(1, int(round(bars_per_session * _TRADING_DAYS)))
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
def _clean(returns: pd.Series) -> np.ndarray:
|
| 53 |
+
arr = np.asarray(returns, dtype=float)
|
| 54 |
+
return arr[np.isfinite(arr)]
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def sharpe_ratio(returns: pd.Series, periods_per_year: int, rf: float = 0.0) -> float:
|
| 58 |
+
"""Annualised Sharpe. ``rf`` is an annual risk-free rate."""
|
| 59 |
+
arr = _clean(returns)
|
| 60 |
+
if arr.size < 2:
|
| 61 |
+
return 0.0
|
| 62 |
+
excess = arr - rf / periods_per_year
|
| 63 |
+
sd = excess.std(ddof=1)
|
| 64 |
+
if not np.isfinite(sd) or sd < _DEGENERATE_SD:
|
| 65 |
+
return 0.0
|
| 66 |
+
return float(excess.mean() / sd * np.sqrt(periods_per_year))
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
def sortino_ratio(returns: pd.Series, periods_per_year: int, rf: float = 0.0) -> float:
|
| 70 |
+
arr = _clean(returns)
|
| 71 |
+
if arr.size < 2:
|
| 72 |
+
return 0.0
|
| 73 |
+
excess = arr - rf / periods_per_year
|
| 74 |
+
downside = excess[excess < 0]
|
| 75 |
+
if downside.size == 0:
|
| 76 |
+
return float("inf") if excess.mean() > 0 else 0.0
|
| 77 |
+
dd = np.sqrt(np.mean(downside**2))
|
| 78 |
+
if not np.isfinite(dd) or dd < _DEGENERATE_SD:
|
| 79 |
+
return 0.0
|
| 80 |
+
return float(excess.mean() / dd * np.sqrt(periods_per_year))
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
def drawdown_series(equity: pd.Series) -> pd.Series:
|
| 84 |
+
peak = equity.cummax()
|
| 85 |
+
return equity / peak - 1.0
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
def max_drawdown(equity: pd.Series) -> float:
|
| 89 |
+
if equity.empty:
|
| 90 |
+
return 0.0
|
| 91 |
+
return float(drawdown_series(equity).min())
|
| 92 |
+
|
| 93 |
+
|
| 94 |
+
def _time_under_water(equity: pd.Series, periods_per_year: int) -> float:
|
| 95 |
+
"""Longest stretch below a prior peak, in years."""
|
| 96 |
+
if equity.empty:
|
| 97 |
+
return 0.0
|
| 98 |
+
dd = drawdown_series(equity).to_numpy()
|
| 99 |
+
longest = current = 0
|
| 100 |
+
for value in dd:
|
| 101 |
+
current = current + 1 if value < 0 else 0
|
| 102 |
+
longest = max(longest, current)
|
| 103 |
+
return longest / periods_per_year
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
def compute_metrics(
|
| 107 |
+
returns: pd.Series,
|
| 108 |
+
equity: pd.Series,
|
| 109 |
+
position: pd.Series | None = None,
|
| 110 |
+
periods_per_year: int | None = None,
|
| 111 |
+
rf: float = 0.0,
|
| 112 |
+
) -> Dict[str, float]:
|
| 113 |
+
"""Full metric bundle for one equity curve."""
|
| 114 |
+
ppy = periods_per_year or infer_periods_per_year(returns.index)
|
| 115 |
+
arr = _clean(returns)
|
| 116 |
+
n = arr.size
|
| 117 |
+
if n == 0 or equity.empty:
|
| 118 |
+
return {"periods_per_year": float(ppy)}
|
| 119 |
+
|
| 120 |
+
years = n / ppy
|
| 121 |
+
total_return = float(equity.iloc[-1] / equity.iloc[0] - 1.0)
|
| 122 |
+
cagr = float((equity.iloc[-1] / equity.iloc[0]) ** (1.0 / years) - 1.0) if years > 0 else 0.0
|
| 123 |
+
vol = float(arr.std(ddof=1) * np.sqrt(ppy))
|
| 124 |
+
mdd = max_drawdown(equity)
|
| 125 |
+
sr = sharpe_ratio(returns, ppy, rf)
|
| 126 |
+
|
| 127 |
+
out: Dict[str, float] = {
|
| 128 |
+
"total_return": total_return,
|
| 129 |
+
"cagr": cagr,
|
| 130 |
+
"ann_vol": vol,
|
| 131 |
+
"sharpe": sr,
|
| 132 |
+
"sortino": sortino_ratio(returns, ppy, rf),
|
| 133 |
+
"calmar": float(cagr / abs(mdd)) if mdd < 0 else 0.0,
|
| 134 |
+
"max_drawdown": mdd,
|
| 135 |
+
"time_under_water_yrs": _time_under_water(equity, ppy),
|
| 136 |
+
"hit_rate": float((arr > 0).mean()),
|
| 137 |
+
"skew": float(pd.Series(arr).skew()) if n > 2 else 0.0,
|
| 138 |
+
"kurtosis": float(pd.Series(arr).kurtosis()) if n > 3 else 0.0,
|
| 139 |
+
"var_95": float(np.percentile(arr, 5)),
|
| 140 |
+
"cvar_95": float(arr[arr <= np.percentile(arr, 5)].mean()) if n > 20 else 0.0,
|
| 141 |
+
"best_bar": float(arr.max()),
|
| 142 |
+
"worst_bar": float(arr.min()),
|
| 143 |
+
"n_bars": float(n),
|
| 144 |
+
"years": float(years),
|
| 145 |
+
"periods_per_year": float(ppy),
|
| 146 |
+
}
|
| 147 |
+
|
| 148 |
+
if position is not None and not position.empty:
|
| 149 |
+
pos = position.fillna(0.0)
|
| 150 |
+
turnover = pos.diff().abs().fillna(pos.abs().iloc[0] if len(pos) else 0.0)
|
| 151 |
+
out["exposure"] = float(pos.abs().mean())
|
| 152 |
+
out["long_share"] = float((pos > 0).mean())
|
| 153 |
+
out["short_share"] = float((pos < 0).mean())
|
| 154 |
+
out["turnover_ann"] = float(turnover.sum() / years) if years > 0 else 0.0
|
| 155 |
+
# A "trade" is any change in sign or size of exposure.
|
| 156 |
+
out["n_trades"] = float((turnover > 1e-9).sum())
|
| 157 |
+
return out
|
|
@@ -0,0 +1,345 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The strategy zoo.
|
| 2 |
+
|
| 3 |
+
Every strategy is a pure function ``(df, **params) -> target exposure series``
|
| 4 |
+
in ``[-1, 1]``, causal by construction. They never see costs, capital or
|
| 5 |
+
execution — that is the engine's job — which is what lets the same function be
|
| 6 |
+
re-run thousands of times inside the permutation and PBO machinery.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from __future__ import annotations
|
| 10 |
+
|
| 11 |
+
import itertools
|
| 12 |
+
from dataclasses import dataclass
|
| 13 |
+
from typing import Callable, Dict, Iterable, List, Sequence
|
| 14 |
+
|
| 15 |
+
import numpy as np
|
| 16 |
+
import pandas as pd
|
| 17 |
+
|
| 18 |
+
from . import indicators as ind
|
| 19 |
+
|
| 20 |
+
__all__ = ["Strategy", "ParamSpec", "REGISTRY", "get_strategy", "list_strategies"]
|
| 21 |
+
|
| 22 |
+
|
| 23 |
+
@dataclass(frozen=True)
|
| 24 |
+
class ParamSpec:
|
| 25 |
+
name: str
|
| 26 |
+
label: str
|
| 27 |
+
default: float
|
| 28 |
+
grid: Sequence[float]
|
| 29 |
+
kind: str = "int"
|
| 30 |
+
minimum: float | None = None
|
| 31 |
+
maximum: float | None = None
|
| 32 |
+
step: float | None = None
|
| 33 |
+
|
| 34 |
+
def cast(self, value):
|
| 35 |
+
return int(value) if self.kind == "int" else float(value)
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
@dataclass(frozen=True)
|
| 39 |
+
class Strategy:
|
| 40 |
+
key: str
|
| 41 |
+
name: str
|
| 42 |
+
family: str
|
| 43 |
+
description: str
|
| 44 |
+
fn: Callable[..., pd.Series]
|
| 45 |
+
params: tuple[ParamSpec, ...] = ()
|
| 46 |
+
|
| 47 |
+
def defaults(self) -> Dict[str, float]:
|
| 48 |
+
return {p.name: p.cast(p.default) for p in self.params}
|
| 49 |
+
|
| 50 |
+
def clean(self, params: Dict[str, float] | None) -> Dict[str, float]:
|
| 51 |
+
"""Fill in missing params and coerce types, ignoring unknown keys."""
|
| 52 |
+
merged = self.defaults()
|
| 53 |
+
for spec in self.params:
|
| 54 |
+
if params and spec.name in params and params[spec.name] is not None:
|
| 55 |
+
merged[spec.name] = spec.cast(params[spec.name])
|
| 56 |
+
return merged
|
| 57 |
+
|
| 58 |
+
def generate(self, df: pd.DataFrame, params: Dict[str, float] | None = None) -> pd.Series:
|
| 59 |
+
target = self.fn(df, **self.clean(params))
|
| 60 |
+
return target.reindex(df.index).astype(float).fillna(0.0).clip(-1.0, 1.0)
|
| 61 |
+
|
| 62 |
+
def grid(self, limit: int | None = None) -> List[Dict[str, float]]:
|
| 63 |
+
"""Cartesian product of the per-parameter grids (the 'trials' a
|
| 64 |
+
researcher would realistically run before picking a winner)."""
|
| 65 |
+
if not self.params:
|
| 66 |
+
return [{}]
|
| 67 |
+
names = [p.name for p in self.params]
|
| 68 |
+
combos = [
|
| 69 |
+
dict(zip(names, values))
|
| 70 |
+
for values in itertools.product(*[p.grid for p in self.params])
|
| 71 |
+
]
|
| 72 |
+
combos = [c for c in combos if self._valid(c)]
|
| 73 |
+
if limit is not None and len(combos) > limit:
|
| 74 |
+
step = len(combos) / limit
|
| 75 |
+
combos = [combos[int(i * step)] for i in range(limit)]
|
| 76 |
+
return combos
|
| 77 |
+
|
| 78 |
+
def _valid(self, combo: Dict[str, float]) -> bool:
|
| 79 |
+
"""Reject nonsensical combinations (a fast MA slower than the slow one)."""
|
| 80 |
+
if "fast" in combo and "slow" in combo and combo["fast"] >= combo["slow"]:
|
| 81 |
+
return False
|
| 82 |
+
if "lower" in combo and "upper" in combo and combo["lower"] >= combo["upper"]:
|
| 83 |
+
return False
|
| 84 |
+
return True
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
def _hold_until_flip(raw: pd.Series) -> pd.Series:
|
| 88 |
+
"""Turn sparse entry/exit signals into a continuously held position."""
|
| 89 |
+
return raw.ffill().fillna(0.0)
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
# --------------------------------------------------------------------------
|
| 93 |
+
# Strategy implementations
|
| 94 |
+
# --------------------------------------------------------------------------
|
| 95 |
+
|
| 96 |
+
def _buy_and_hold(df: pd.DataFrame) -> pd.Series:
|
| 97 |
+
return pd.Series(1.0, index=df.index)
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
def _sma_cross(df: pd.DataFrame, fast: int = 20, slow: int = 100) -> pd.Series:
|
| 101 |
+
f, s = ind.sma(df["close"], fast), ind.sma(df["close"], slow)
|
| 102 |
+
return pd.Series(np.where(f > s, 1.0, -1.0), index=df.index).where(s.notna())
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def _ema_cross(df: pd.DataFrame, fast: int = 12, slow: int = 50) -> pd.Series:
|
| 106 |
+
f, s = ind.ema(df["close"], fast), ind.ema(df["close"], slow)
|
| 107 |
+
return pd.Series(np.where(f > s, 1.0, -1.0), index=df.index).where(s.notna())
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
def _macd_trend(df: pd.DataFrame, fast: int = 12, slow: int = 26, signal: int = 9) -> pd.Series:
|
| 111 |
+
_, _, hist = ind.macd(df["close"], fast, slow, signal)
|
| 112 |
+
return pd.Series(np.sign(hist), index=df.index).where(hist.notna())
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def _rsi_reversion(df: pd.DataFrame, window: int = 14, lower: int = 30, upper: int = 70) -> pd.Series:
|
| 116 |
+
r = ind.rsi(df["close"], window)
|
| 117 |
+
raw = pd.Series(np.nan, index=df.index)
|
| 118 |
+
raw[r < lower] = 1.0
|
| 119 |
+
raw[r > upper] = -1.0
|
| 120 |
+
raw[(r > 45) & (r < 55)] = 0.0 # flatten in the middle of the range
|
| 121 |
+
return _hold_until_flip(raw).where(r.notna())
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
def _bollinger_reversion(df: pd.DataFrame, window: int = 20, k: float = 2.0) -> pd.Series:
|
| 125 |
+
low, mid, high = ind.bollinger(df["close"], window, k)
|
| 126 |
+
close = df["close"]
|
| 127 |
+
raw = pd.Series(np.nan, index=df.index)
|
| 128 |
+
raw[close < low] = 1.0
|
| 129 |
+
raw[close > high] = -1.0
|
| 130 |
+
raw[(close - mid).abs() < 0.1 * (high - mid)] = 0.0
|
| 131 |
+
return _hold_until_flip(raw).where(mid.notna())
|
| 132 |
+
|
| 133 |
+
|
| 134 |
+
def _donchian_breakout(df: pd.DataFrame, window: int = 20) -> pd.Series:
|
| 135 |
+
low, high = ind.donchian(df, window)
|
| 136 |
+
raw = pd.Series(np.nan, index=df.index)
|
| 137 |
+
raw[df["close"] > high] = 1.0
|
| 138 |
+
raw[df["close"] < low] = -1.0
|
| 139 |
+
return _hold_until_flip(raw).where(high.notna())
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def _momentum(df: pd.DataFrame, lookback: int = 60) -> pd.Series:
|
| 143 |
+
return np.sign(ind.roc(df["close"], lookback))
|
| 144 |
+
|
| 145 |
+
|
| 146 |
+
def _vol_target_momentum(
|
| 147 |
+
df: pd.DataFrame, lookback: int = 60, vol_window: int = 20, target_vol: float = 15
|
| 148 |
+
) -> pd.Series:
|
| 149 |
+
"""Momentum sized inversely to recent volatility (targets ``target_vol`` %)."""
|
| 150 |
+
signal = np.sign(ind.roc(df["close"], lookback))
|
| 151 |
+
rv = ind.realised_vol(df["close"].pct_change(), vol_window)
|
| 152 |
+
scale = (target_vol / 100.0) / rv.replace(0.0, np.nan)
|
| 153 |
+
return (signal * scale.clip(upper=1.0)).where(rv.notna())
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
def _channel_trend(df: pd.DataFrame, window: int = 50, atr_window: int = 14, mult: float = 1.0) -> pd.Series:
|
| 157 |
+
"""Long above an ATR band around the mean, short below it, flat inside."""
|
| 158 |
+
mid = ind.sma(df["close"], window)
|
| 159 |
+
band = ind.atr(df, atr_window) * mult
|
| 160 |
+
raw = pd.Series(np.nan, index=df.index)
|
| 161 |
+
raw[df["close"] > mid + band] = 1.0
|
| 162 |
+
raw[df["close"] < mid - band] = -1.0
|
| 163 |
+
raw[(df["close"] - mid).abs() < 0.25 * band] = 0.0
|
| 164 |
+
return _hold_until_flip(raw).where(mid.notna() & band.notna())
|
| 165 |
+
|
| 166 |
+
|
| 167 |
+
def _coin_flip(df: pd.DataFrame, hold: int = 5, seed: int = 7) -> pd.Series:
|
| 168 |
+
"""A deliberately worthless strategy: the control group.
|
| 169 |
+
|
| 170 |
+
If your clever rule cannot beat this on the validation panel, that is the
|
| 171 |
+
single most useful thing this app can tell you.
|
| 172 |
+
"""
|
| 173 |
+
rng = np.random.default_rng(int(seed))
|
| 174 |
+
n = len(df)
|
| 175 |
+
draws = rng.choice([-1.0, 1.0], size=int(np.ceil(n / max(hold, 1))))
|
| 176 |
+
return pd.Series(np.repeat(draws, max(hold, 1))[:n], index=df.index)
|
| 177 |
+
|
| 178 |
+
|
| 179 |
+
REGISTRY: Dict[str, Strategy] = {}
|
| 180 |
+
|
| 181 |
+
|
| 182 |
+
def _register(strategy: Strategy) -> Strategy:
|
| 183 |
+
REGISTRY[strategy.key] = strategy
|
| 184 |
+
return strategy
|
| 185 |
+
|
| 186 |
+
|
| 187 |
+
_register(
|
| 188 |
+
Strategy(
|
| 189 |
+
key="buy_and_hold",
|
| 190 |
+
name="Buy & Hold",
|
| 191 |
+
family="benchmark",
|
| 192 |
+
description="Own the asset, do nothing. The bar every other strategy has to clear.",
|
| 193 |
+
fn=_buy_and_hold,
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
|
| 197 |
+
_register(
|
| 198 |
+
Strategy(
|
| 199 |
+
key="sma_cross",
|
| 200 |
+
name="SMA Crossover",
|
| 201 |
+
family="trend",
|
| 202 |
+
description="Long when the fast simple moving average is above the slow one, short when below.",
|
| 203 |
+
fn=_sma_cross,
|
| 204 |
+
params=(
|
| 205 |
+
ParamSpec("fast", "Fast MA", 20, (5, 10, 20, 30, 50), "int", 2, 100, 1),
|
| 206 |
+
ParamSpec("slow", "Slow MA", 100, (50, 100, 150, 200), "int", 10, 300, 5),
|
| 207 |
+
),
|
| 208 |
+
)
|
| 209 |
+
)
|
| 210 |
+
|
| 211 |
+
_register(
|
| 212 |
+
Strategy(
|
| 213 |
+
key="ema_cross",
|
| 214 |
+
name="EMA Crossover",
|
| 215 |
+
family="trend",
|
| 216 |
+
description="Same idea as the SMA cross but with exponential averages, so it turns faster.",
|
| 217 |
+
fn=_ema_cross,
|
| 218 |
+
params=(
|
| 219 |
+
ParamSpec("fast", "Fast EMA", 12, (5, 8, 12, 21, 34), "int", 2, 100, 1),
|
| 220 |
+
ParamSpec("slow", "Slow EMA", 50, (34, 50, 89, 144, 200), "int", 10, 300, 1),
|
| 221 |
+
),
|
| 222 |
+
)
|
| 223 |
+
)
|
| 224 |
+
|
| 225 |
+
_register(
|
| 226 |
+
Strategy(
|
| 227 |
+
key="macd_trend",
|
| 228 |
+
name="MACD Trend",
|
| 229 |
+
family="trend",
|
| 230 |
+
description="Follow the sign of the MACD histogram.",
|
| 231 |
+
fn=_macd_trend,
|
| 232 |
+
params=(
|
| 233 |
+
ParamSpec("fast", "Fast", 12, (8, 12, 16), "int", 2, 60, 1),
|
| 234 |
+
ParamSpec("slow", "Slow", 26, (21, 26, 34, 50), "int", 10, 200, 1),
|
| 235 |
+
ParamSpec("signal", "Signal", 9, (5, 9, 13), "int", 2, 50, 1),
|
| 236 |
+
),
|
| 237 |
+
)
|
| 238 |
+
)
|
| 239 |
+
|
| 240 |
+
_register(
|
| 241 |
+
Strategy(
|
| 242 |
+
key="rsi_reversion",
|
| 243 |
+
name="RSI Mean Reversion",
|
| 244 |
+
family="mean-reversion",
|
| 245 |
+
description="Buy oversold, sell overbought, flatten in the middle of the range.",
|
| 246 |
+
fn=_rsi_reversion,
|
| 247 |
+
params=(
|
| 248 |
+
ParamSpec("window", "RSI window", 14, (7, 14, 21), "int", 2, 60, 1),
|
| 249 |
+
ParamSpec("lower", "Oversold", 30, (20, 25, 30, 35), "int", 5, 49, 1),
|
| 250 |
+
ParamSpec("upper", "Overbought", 70, (65, 70, 75, 80), "int", 51, 95, 1),
|
| 251 |
+
),
|
| 252 |
+
)
|
| 253 |
+
)
|
| 254 |
+
|
| 255 |
+
_register(
|
| 256 |
+
Strategy(
|
| 257 |
+
key="bollinger_reversion",
|
| 258 |
+
name="Bollinger Reversion",
|
| 259 |
+
family="mean-reversion",
|
| 260 |
+
description="Fade moves outside the Bollinger bands and exit back at the middle band.",
|
| 261 |
+
fn=_bollinger_reversion,
|
| 262 |
+
params=(
|
| 263 |
+
ParamSpec("window", "Window", 20, (10, 20, 30, 50), "int", 5, 120, 1),
|
| 264 |
+
ParamSpec("k", "Band width (σ)", 2.0, (1.5, 2.0, 2.5, 3.0), "float", 0.5, 4.0, 0.1),
|
| 265 |
+
),
|
| 266 |
+
)
|
| 267 |
+
)
|
| 268 |
+
|
| 269 |
+
_register(
|
| 270 |
+
Strategy(
|
| 271 |
+
key="donchian_breakout",
|
| 272 |
+
name="Donchian Breakout",
|
| 273 |
+
family="breakout",
|
| 274 |
+
description="The classic turtle rule: buy new highs, sell new lows.",
|
| 275 |
+
fn=_donchian_breakout,
|
| 276 |
+
params=(ParamSpec("window", "Channel", 20, (10, 20, 40, 55, 100), "int", 5, 250, 1),),
|
| 277 |
+
)
|
| 278 |
+
)
|
| 279 |
+
|
| 280 |
+
_register(
|
| 281 |
+
Strategy(
|
| 282 |
+
key="momentum",
|
| 283 |
+
name="Time-Series Momentum",
|
| 284 |
+
family="momentum",
|
| 285 |
+
description="Hold long if the asset is up over the lookback, short if it is down.",
|
| 286 |
+
fn=_momentum,
|
| 287 |
+
params=(ParamSpec("lookback", "Lookback", 60, (5, 10, 20, 60, 120, 250), "int", 2, 500, 1),),
|
| 288 |
+
)
|
| 289 |
+
)
|
| 290 |
+
|
| 291 |
+
_register(
|
| 292 |
+
Strategy(
|
| 293 |
+
key="vol_target_momentum",
|
| 294 |
+
name="Vol-Targeted Momentum",
|
| 295 |
+
family="momentum",
|
| 296 |
+
description="Momentum sized down when markets get volatile, so risk stays roughly constant.",
|
| 297 |
+
fn=_vol_target_momentum,
|
| 298 |
+
params=(
|
| 299 |
+
ParamSpec("lookback", "Lookback", 60, (20, 60, 120, 250), "int", 5, 500, 1),
|
| 300 |
+
ParamSpec("vol_window", "Vol window", 20, (10, 20, 60), "int", 5, 120, 1),
|
| 301 |
+
ParamSpec("target_vol", "Target vol %", 15, (10, 15, 20), "float", 2, 60, 1),
|
| 302 |
+
),
|
| 303 |
+
)
|
| 304 |
+
)
|
| 305 |
+
|
| 306 |
+
_register(
|
| 307 |
+
Strategy(
|
| 308 |
+
key="channel_trend",
|
| 309 |
+
name="ATR Channel Trend",
|
| 310 |
+
family="trend",
|
| 311 |
+
description="Trade with the trend only once price clears an ATR band around its mean.",
|
| 312 |
+
fn=_channel_trend,
|
| 313 |
+
params=(
|
| 314 |
+
ParamSpec("window", "Mean window", 50, (20, 50, 100, 200), "int", 5, 300, 1),
|
| 315 |
+
ParamSpec("atr_window", "ATR window", 14, (7, 14, 28), "int", 2, 60, 1),
|
| 316 |
+
ParamSpec("mult", "ATR multiple", 1.0, (0.5, 1.0, 1.5, 2.0), "float", 0.1, 5.0, 0.1),
|
| 317 |
+
),
|
| 318 |
+
)
|
| 319 |
+
)
|
| 320 |
+
|
| 321 |
+
_register(
|
| 322 |
+
Strategy(
|
| 323 |
+
key="coin_flip",
|
| 324 |
+
name="Coin Flip (control)",
|
| 325 |
+
family="control",
|
| 326 |
+
description="Random positions. The control group — anything that cannot beat this is noise.",
|
| 327 |
+
fn=_coin_flip,
|
| 328 |
+
params=(
|
| 329 |
+
ParamSpec("hold", "Bars per flip", 5, (1, 5, 10, 20), "int", 1, 60, 1),
|
| 330 |
+
ParamSpec("seed", "Seed", 7, (1, 7, 42, 123), "int", 0, 9999, 1),
|
| 331 |
+
),
|
| 332 |
+
)
|
| 333 |
+
)
|
| 334 |
+
|
| 335 |
+
|
| 336 |
+
def get_strategy(key: str) -> Strategy:
|
| 337 |
+
try:
|
| 338 |
+
return REGISTRY[key]
|
| 339 |
+
except KeyError:
|
| 340 |
+
raise KeyError(f"Unknown strategy '{key}'. Available: {', '.join(sorted(REGISTRY))}") from None
|
| 341 |
+
|
| 342 |
+
|
| 343 |
+
def list_strategies(exclude: Iterable[str] = ()) -> List[Strategy]:
|
| 344 |
+
skip = set(exclude)
|
| 345 |
+
return [s for k, s in REGISTRY.items() if k not in skip]
|
|
@@ -0,0 +1,100 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Core data types shared across the algotrader 2.0 stack.
|
| 2 |
+
|
| 3 |
+
Everything downstream (engine, validation, UI) speaks these types, so they are
|
| 4 |
+
deliberately small, immutable-ish and free of framework dependencies.
|
| 5 |
+
"""
|
| 6 |
+
|
| 7 |
+
from __future__ import annotations
|
| 8 |
+
|
| 9 |
+
from dataclasses import dataclass, field
|
| 10 |
+
from typing import Any, Dict, Optional
|
| 11 |
+
|
| 12 |
+
import pandas as pd
|
| 13 |
+
|
| 14 |
+
OHLCV_COLUMNS = ("open", "high", "low", "close", "volume")
|
| 15 |
+
|
| 16 |
+
|
| 17 |
+
@dataclass(frozen=True)
|
| 18 |
+
class MarketData:
|
| 19 |
+
"""A validated OHLCV series plus provenance.
|
| 20 |
+
|
| 21 |
+
Provenance matters here: the app is about honesty, so the UI always tells
|
| 22 |
+
the user whether they are looking at real prices or a simulation.
|
| 23 |
+
"""
|
| 24 |
+
|
| 25 |
+
symbol: str
|
| 26 |
+
df: pd.DataFrame
|
| 27 |
+
source: str # "yfinance" | "bundled" | "synthetic"
|
| 28 |
+
interval: str = "1d"
|
| 29 |
+
note: str = ""
|
| 30 |
+
|
| 31 |
+
@property
|
| 32 |
+
def is_real(self) -> bool:
|
| 33 |
+
return self.source in ("yfinance", "bundled")
|
| 34 |
+
|
| 35 |
+
@property
|
| 36 |
+
def start(self) -> pd.Timestamp:
|
| 37 |
+
return self.df.index[0]
|
| 38 |
+
|
| 39 |
+
@property
|
| 40 |
+
def end(self) -> pd.Timestamp:
|
| 41 |
+
return self.df.index[-1]
|
| 42 |
+
|
| 43 |
+
def __len__(self) -> int: # pragma: no cover - trivial
|
| 44 |
+
return len(self.df)
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
@dataclass(frozen=True)
|
| 48 |
+
class CostModel:
|
| 49 |
+
"""Round-trip friction. All values are one-way, in basis points."""
|
| 50 |
+
|
| 51 |
+
commission_bps: float = 1.0
|
| 52 |
+
slippage_bps: float = 2.0
|
| 53 |
+
short_borrow_bps: float = 50.0 # annualised, charged on short exposure
|
| 54 |
+
|
| 55 |
+
@property
|
| 56 |
+
def one_way_bps(self) -> float:
|
| 57 |
+
return self.commission_bps + self.slippage_bps
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
@dataclass
|
| 61 |
+
class BacktestResult:
|
| 62 |
+
"""Output of a single backtest run."""
|
| 63 |
+
|
| 64 |
+
equity: pd.Series
|
| 65 |
+
returns: pd.Series # net of costs
|
| 66 |
+
gross_returns: pd.Series
|
| 67 |
+
position: pd.Series # exposure actually held during each bar
|
| 68 |
+
target: pd.Series # exposure requested by the strategy
|
| 69 |
+
costs: pd.Series
|
| 70 |
+
benchmark_equity: pd.Series
|
| 71 |
+
metrics: Dict[str, float] = field(default_factory=dict)
|
| 72 |
+
benchmark_metrics: Dict[str, float] = field(default_factory=dict)
|
| 73 |
+
meta: Dict[str, Any] = field(default_factory=dict)
|
| 74 |
+
|
| 75 |
+
@property
|
| 76 |
+
def sharpe(self) -> float:
|
| 77 |
+
return float(self.metrics.get("sharpe", 0.0))
|
| 78 |
+
|
| 79 |
+
@property
|
| 80 |
+
def n_trades(self) -> int:
|
| 81 |
+
return int(self.metrics.get("n_trades", 0))
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
@dataclass
|
| 85 |
+
class ValidationReport:
|
| 86 |
+
"""Everything we know about how much of a backtest is luck."""
|
| 87 |
+
|
| 88 |
+
permutation_p_value: Optional[float] = None
|
| 89 |
+
permutation_null: Optional[Any] = None # np.ndarray of null Sharpes
|
| 90 |
+
deflated_sharpe: Optional[float] = None
|
| 91 |
+
probabilistic_sharpe: Optional[float] = None
|
| 92 |
+
min_track_record_years: Optional[float] = None
|
| 93 |
+
n_trials: int = 1
|
| 94 |
+
pbo: Optional[float] = None
|
| 95 |
+
pbo_detail: Dict[str, Any] = field(default_factory=dict)
|
| 96 |
+
walkforward: Dict[str, Any] = field(default_factory=dict)
|
| 97 |
+
reality_score: float = 0.0
|
| 98 |
+
grade: str = "?"
|
| 99 |
+
verdict: str = ""
|
| 100 |
+
flags: list = field(default_factory=list)
|
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Statistical tests that ask "is this edge real, or did we just get lucky?"."""
|
| 2 |
+
|
| 3 |
+
from .deflated_sharpe import deflated_sharpe_ratio, min_track_record_length, probabilistic_sharpe_ratio
|
| 4 |
+
from .pbo import probability_of_backtest_overfitting
|
| 5 |
+
from .permutation import permutation_test
|
| 6 |
+
from .walkforward import walk_forward
|
| 7 |
+
|
| 8 |
+
__all__ = [
|
| 9 |
+
"deflated_sharpe_ratio",
|
| 10 |
+
"probabilistic_sharpe_ratio",
|
| 11 |
+
"min_track_record_length",
|
| 12 |
+
"probability_of_backtest_overfitting",
|
| 13 |
+
"permutation_test",
|
| 14 |
+
"walk_forward",
|
| 15 |
+
]
|
|
@@ -0,0 +1,145 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Probabilistic and Deflated Sharpe Ratios.
|
| 2 |
+
|
| 3 |
+
Bailey & López de Prado (2014), "The Deflated Sharpe Ratio: Correcting for
|
| 4 |
+
Selection Bias, Backtest Overfitting and Non-Normality".
|
| 5 |
+
|
| 6 |
+
The intuition: if you try 200 strategy variants, the best one will show a
|
| 7 |
+
handsome Sharpe *even when none of them has any edge*. The Deflated Sharpe
|
| 8 |
+
Ratio asks whether the winner beats what the luckiest of 200 coin-flippers
|
| 9 |
+
would have produced, and it charges extra for fat tails and negative skew --
|
| 10 |
+
exactly the return shapes that make naive Sharpe ratios flatter.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
import numpy as np
|
| 16 |
+
from scipy import stats
|
| 17 |
+
|
| 18 |
+
__all__ = [
|
| 19 |
+
"probabilistic_sharpe_ratio",
|
| 20 |
+
"expected_max_sharpe",
|
| 21 |
+
"deflated_sharpe_ratio",
|
| 22 |
+
"min_track_record_length",
|
| 23 |
+
]
|
| 24 |
+
|
| 25 |
+
_EULER = 0.5772156649015329
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
def _moments(returns: np.ndarray) -> tuple[float, float]:
|
| 29 |
+
"""Sample skew and *non-excess* kurtosis, as the PSR formula expects."""
|
| 30 |
+
arr = np.asarray(returns, dtype=float)
|
| 31 |
+
arr = arr[np.isfinite(arr)]
|
| 32 |
+
if arr.size < 4:
|
| 33 |
+
return 0.0, 3.0
|
| 34 |
+
return float(stats.skew(arr, bias=False)), float(stats.kurtosis(arr, bias=False) + 3.0)
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def probabilistic_sharpe_ratio(
|
| 38 |
+
sharpe: float,
|
| 39 |
+
n_obs: int,
|
| 40 |
+
skew: float = 0.0,
|
| 41 |
+
kurtosis: float = 3.0,
|
| 42 |
+
benchmark: float = 0.0,
|
| 43 |
+
) -> float:
|
| 44 |
+
"""P(true Sharpe > ``benchmark``) given the observed Sharpe and its shape.
|
| 45 |
+
|
| 46 |
+
``sharpe`` and ``benchmark`` are per-observation (i.e. *not* annualised).
|
| 47 |
+
"""
|
| 48 |
+
if n_obs < 3:
|
| 49 |
+
return 0.5
|
| 50 |
+
denom = 1.0 - skew * sharpe + ((kurtosis - 1.0) / 4.0) * sharpe**2
|
| 51 |
+
if denom <= 0:
|
| 52 |
+
return 0.5
|
| 53 |
+
z = (sharpe - benchmark) * np.sqrt(n_obs - 1) / np.sqrt(denom)
|
| 54 |
+
return float(stats.norm.cdf(z))
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def expected_max_sharpe(n_trials: int, variance_of_trials: float) -> float:
|
| 58 |
+
"""Expected maximum Sharpe across ``n_trials`` *skill-free* strategies.
|
| 59 |
+
|
| 60 |
+
This is the bar the winner has to clear to be interesting. It grows with
|
| 61 |
+
the number of things you tried — which is why "I found a strategy with
|
| 62 |
+
Sharpe 2" means nothing until you say how many you looked at.
|
| 63 |
+
"""
|
| 64 |
+
n = max(int(n_trials), 1)
|
| 65 |
+
if n == 1 or variance_of_trials <= 0:
|
| 66 |
+
return 0.0
|
| 67 |
+
sd = np.sqrt(variance_of_trials)
|
| 68 |
+
# Bailey & López de Prado's Gumbel-based approximation.
|
| 69 |
+
q1 = stats.norm.ppf(1.0 - 1.0 / n)
|
| 70 |
+
q2 = stats.norm.ppf(1.0 - 1.0 / (n * np.e))
|
| 71 |
+
return float(sd * ((1.0 - _EULER) * q1 + _EULER * q2))
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def deflated_sharpe_ratio(
|
| 75 |
+
returns,
|
| 76 |
+
sharpe_annual: float,
|
| 77 |
+
periods_per_year: int,
|
| 78 |
+
n_trials: int,
|
| 79 |
+
trial_sharpes=None,
|
| 80 |
+
variance_of_trials: float | None = None,
|
| 81 |
+
) -> dict:
|
| 82 |
+
"""Deflate an annualised Sharpe for selection bias and non-normality.
|
| 83 |
+
|
| 84 |
+
Returns a dict with the PSR against a zero benchmark, the selection-bias
|
| 85 |
+
threshold, the deflated probability, and the inputs used, so the UI can
|
| 86 |
+
show its working rather than just a number.
|
| 87 |
+
"""
|
| 88 |
+
arr = np.asarray(returns, dtype=float)
|
| 89 |
+
arr = arr[np.isfinite(arr)]
|
| 90 |
+
n_obs = arr.size
|
| 91 |
+
sr_per_period = sharpe_annual / np.sqrt(periods_per_year)
|
| 92 |
+
skew, kurt = _moments(arr)
|
| 93 |
+
|
| 94 |
+
if variance_of_trials is None:
|
| 95 |
+
if trial_sharpes is not None and len(trial_sharpes) > 1:
|
| 96 |
+
trials = np.asarray(trial_sharpes, dtype=float) / np.sqrt(periods_per_year)
|
| 97 |
+
trials = trials[np.isfinite(trials)]
|
| 98 |
+
variance_of_trials = float(np.var(trials, ddof=1)) if trials.size > 1 else 0.0
|
| 99 |
+
else:
|
| 100 |
+
# With no trial cloud to measure, fall back to the asymptotic
|
| 101 |
+
# variance of a skill-free Sharpe estimate.
|
| 102 |
+
variance_of_trials = 1.0 / max(n_obs - 1, 1)
|
| 103 |
+
|
| 104 |
+
threshold = expected_max_sharpe(n_trials, variance_of_trials)
|
| 105 |
+
|
| 106 |
+
psr = probabilistic_sharpe_ratio(sr_per_period, n_obs, skew, kurt, 0.0)
|
| 107 |
+
dsr = probabilistic_sharpe_ratio(sr_per_period, n_obs, skew, kurt, threshold)
|
| 108 |
+
|
| 109 |
+
return {
|
| 110 |
+
"psr": float(psr),
|
| 111 |
+
"dsr": float(dsr),
|
| 112 |
+
"sr_per_period": float(sr_per_period),
|
| 113 |
+
"threshold_sr_per_period": float(threshold),
|
| 114 |
+
"threshold_sr_annual": float(threshold * np.sqrt(periods_per_year)),
|
| 115 |
+
"n_obs": int(n_obs),
|
| 116 |
+
"n_trials": int(n_trials),
|
| 117 |
+
"skew": float(skew),
|
| 118 |
+
"kurtosis": float(kurt),
|
| 119 |
+
"variance_of_trials": float(variance_of_trials),
|
| 120 |
+
}
|
| 121 |
+
|
| 122 |
+
|
| 123 |
+
def min_track_record_length(
|
| 124 |
+
sharpe: float,
|
| 125 |
+
n_obs: int,
|
| 126 |
+
skew: float = 0.0,
|
| 127 |
+
kurtosis: float = 3.0,
|
| 128 |
+
benchmark: float = 0.0,
|
| 129 |
+
confidence: float = 0.95,
|
| 130 |
+
) -> float:
|
| 131 |
+
"""Observations needed before the Sharpe is significant at ``confidence``.
|
| 132 |
+
|
| 133 |
+
Inputs are per-observation. Returns ``inf`` when the edge is too small to
|
| 134 |
+
ever clear the bar.
|
| 135 |
+
"""
|
| 136 |
+
if sharpe <= benchmark:
|
| 137 |
+
return float("inf")
|
| 138 |
+
z = stats.norm.ppf(confidence)
|
| 139 |
+
denom = (sharpe - benchmark) ** 2
|
| 140 |
+
if denom <= 0:
|
| 141 |
+
return float("inf")
|
| 142 |
+
numer = 1.0 - skew * sharpe + ((kurtosis - 1.0) / 4.0) * sharpe**2
|
| 143 |
+
if numer <= 0:
|
| 144 |
+
return float("inf")
|
| 145 |
+
return float(1.0 + numer * (z / (sharpe - benchmark)) ** 2)
|
|
@@ -0,0 +1,120 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Probability of Backtest Overfitting via Combinatorially Symmetric CV.
|
| 2 |
+
|
| 3 |
+
Bailey, Borwein, López de Prado & Zhu (2016), "The Probability of Backtest
|
| 4 |
+
Overfitting".
|
| 5 |
+
|
| 6 |
+
Take the N parameter variants you tried, cut the timeline into S chunks, and
|
| 7 |
+
for every way of splitting those chunks half-and-half: pick the variant that
|
| 8 |
+
won in-sample, then look up where it ranked out-of-sample. If your selection
|
| 9 |
+
process has skill, the in-sample winner should keep winning. If it is fitting
|
| 10 |
+
noise, the winner lands in the bottom half about as often as not — and PBO
|
| 11 |
+
approaches 0.5.
|
| 12 |
+
"""
|
| 13 |
+
|
| 14 |
+
from __future__ import annotations
|
| 15 |
+
|
| 16 |
+
import itertools
|
| 17 |
+
from typing import Dict, Sequence
|
| 18 |
+
|
| 19 |
+
import numpy as np
|
| 20 |
+
|
| 21 |
+
__all__ = ["probability_of_backtest_overfitting"]
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
def _sharpe_columns(matrix: np.ndarray) -> np.ndarray:
|
| 25 |
+
"""Per-column Sharpe (per-period, unannualised — ranks are all we need)."""
|
| 26 |
+
if matrix.shape[0] < 2:
|
| 27 |
+
return np.zeros(matrix.shape[1])
|
| 28 |
+
mean = matrix.mean(axis=0)
|
| 29 |
+
sd = matrix.std(axis=0, ddof=1)
|
| 30 |
+
# Absolute floor, not `> 0`: a flat column's std is float noise, and
|
| 31 |
+
# dividing by it would hand a do-nothing variant an enormous rank.
|
| 32 |
+
with np.errstate(divide="ignore", invalid="ignore"):
|
| 33 |
+
out = np.where(sd > 1e-12, mean / sd, 0.0)
|
| 34 |
+
return np.nan_to_num(out, nan=0.0, posinf=0.0, neginf=0.0)
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def probability_of_backtest_overfitting(
|
| 38 |
+
returns_matrix: np.ndarray,
|
| 39 |
+
n_splits: int = 8,
|
| 40 |
+
labels: Sequence[str] | None = None,
|
| 41 |
+
max_combinations: int = 200,
|
| 42 |
+
) -> Dict[str, object]:
|
| 43 |
+
"""Compute PBO for a ``T x N`` matrix of per-period strategy returns.
|
| 44 |
+
|
| 45 |
+
``n_splits`` must be even. Returns PBO, the logit distribution, the
|
| 46 |
+
in-sample/out-of-sample Sharpe pairs for the selected variants, and the
|
| 47 |
+
rate at which the selected variant actually loses money out of sample.
|
| 48 |
+
"""
|
| 49 |
+
matrix = np.asarray(returns_matrix, dtype=float)
|
| 50 |
+
if matrix.ndim != 2:
|
| 51 |
+
raise ValueError("returns_matrix must be 2-D (time x strategy)")
|
| 52 |
+
matrix = np.nan_to_num(matrix, nan=0.0, posinf=0.0, neginf=0.0)
|
| 53 |
+
t_obs, n_strats = matrix.shape
|
| 54 |
+
|
| 55 |
+
if n_strats < 2 or t_obs < 2 * n_splits:
|
| 56 |
+
return {
|
| 57 |
+
"pbo": float("nan"),
|
| 58 |
+
"n_strategies": int(n_strats),
|
| 59 |
+
"n_combinations": 0,
|
| 60 |
+
"logits": np.array([]),
|
| 61 |
+
"is_sharpes": np.array([]),
|
| 62 |
+
"oos_sharpes": np.array([]),
|
| 63 |
+
"prob_oos_loss": float("nan"),
|
| 64 |
+
"performance_degradation": float("nan"),
|
| 65 |
+
"note": "Not enough variants or observations to estimate PBO.",
|
| 66 |
+
}
|
| 67 |
+
|
| 68 |
+
if n_splits % 2:
|
| 69 |
+
n_splits += 1
|
| 70 |
+
chunks = np.array_split(np.arange(t_obs), n_splits)
|
| 71 |
+
|
| 72 |
+
combos = list(itertools.combinations(range(n_splits), n_splits // 2))
|
| 73 |
+
if len(combos) > max_combinations:
|
| 74 |
+
step = len(combos) / max_combinations
|
| 75 |
+
combos = [combos[int(i * step)] for i in range(max_combinations)]
|
| 76 |
+
|
| 77 |
+
logits, is_sr, oos_sr, chosen = [], [], [], []
|
| 78 |
+
for combo in combos:
|
| 79 |
+
is_idx = np.concatenate([chunks[c] for c in combo])
|
| 80 |
+
oos_idx = np.concatenate([chunks[c] for c in range(n_splits) if c not in combo])
|
| 81 |
+
|
| 82 |
+
is_perf = _sharpe_columns(matrix[is_idx])
|
| 83 |
+
oos_perf = _sharpe_columns(matrix[oos_idx])
|
| 84 |
+
|
| 85 |
+
best = int(np.argmax(is_perf))
|
| 86 |
+
chosen.append(best)
|
| 87 |
+
is_sr.append(float(is_perf[best]))
|
| 88 |
+
oos_sr.append(float(oos_perf[best]))
|
| 89 |
+
|
| 90 |
+
# Relative rank of the chosen variant in the OOS ranking, in (0, 1).
|
| 91 |
+
rank = float(np.sum(oos_perf <= oos_perf[best]))
|
| 92 |
+
omega = rank / (n_strats + 1.0)
|
| 93 |
+
omega = min(max(omega, 1e-6), 1.0 - 1e-6)
|
| 94 |
+
logits.append(float(np.log(omega / (1.0 - omega))))
|
| 95 |
+
|
| 96 |
+
logits_arr = np.asarray(logits)
|
| 97 |
+
is_arr, oos_arr = np.asarray(is_sr), np.asarray(oos_sr)
|
| 98 |
+
|
| 99 |
+
# Slope of OOS on IS: negative means better in-sample fits do *worse* live.
|
| 100 |
+
degradation = float("nan")
|
| 101 |
+
if is_arr.size > 2 and np.std(is_arr) > 1e-12:
|
| 102 |
+
degradation = float(np.polyfit(is_arr, oos_arr, 1)[0])
|
| 103 |
+
|
| 104 |
+
counts = np.bincount(chosen, minlength=n_strats)
|
| 105 |
+
most_selected = int(np.argmax(counts))
|
| 106 |
+
|
| 107 |
+
return {
|
| 108 |
+
"pbo": float(np.mean(logits_arr <= 0.0)),
|
| 109 |
+
"n_strategies": int(n_strats),
|
| 110 |
+
"n_combinations": int(len(combos)),
|
| 111 |
+
"logits": logits_arr,
|
| 112 |
+
"is_sharpes": is_arr,
|
| 113 |
+
"oos_sharpes": oos_arr,
|
| 114 |
+
"prob_oos_loss": float(np.mean(oos_arr <= 0.0)),
|
| 115 |
+
"performance_degradation": degradation,
|
| 116 |
+
"most_selected_index": most_selected,
|
| 117 |
+
"most_selected_label": (labels[most_selected] if labels is not None else str(most_selected)),
|
| 118 |
+
"selection_stability": float(counts[most_selected] / max(len(combos), 1)),
|
| 119 |
+
"note": "",
|
| 120 |
+
}
|
|
@@ -0,0 +1,187 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Monte-Carlo permutation test for trading rules.
|
| 2 |
+
|
| 3 |
+
The question this answers is not "did the strategy make money" but "would a
|
| 4 |
+
rule of this shape have made this much money on a market with no exploitable
|
| 5 |
+
structure?". We destroy the serial dependence in the price path while keeping
|
| 6 |
+
its distribution of moves intact, re-run the *same* strategy on each shuffled
|
| 7 |
+
market, and see where the real result lands in that null distribution.
|
| 8 |
+
|
| 9 |
+
A strategy whose Sharpe sits comfortably inside the null is not a strategy —
|
| 10 |
+
it is a lottery ticket that happened to win.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
from __future__ import annotations
|
| 14 |
+
|
| 15 |
+
from dataclasses import dataclass
|
| 16 |
+
from typing import Callable, Optional
|
| 17 |
+
|
| 18 |
+
import numpy as np
|
| 19 |
+
import pandas as pd
|
| 20 |
+
|
| 21 |
+
from ..engine import bars_to_returns, run_backtest
|
| 22 |
+
from ..types import CostModel
|
| 23 |
+
|
| 24 |
+
__all__ = ["permutation_test", "PermutationResult", "permute_bars"]
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
@dataclass
|
| 28 |
+
class PermutationResult:
|
| 29 |
+
observed: float
|
| 30 |
+
null: np.ndarray
|
| 31 |
+
p_value: float
|
| 32 |
+
method: str
|
| 33 |
+
n_permutations: int
|
| 34 |
+
|
| 35 |
+
@property
|
| 36 |
+
def null_mean(self) -> float:
|
| 37 |
+
return float(np.mean(self.null)) if self.null.size else 0.0
|
| 38 |
+
|
| 39 |
+
@property
|
| 40 |
+
def percentile(self) -> float:
|
| 41 |
+
"""Where the observed Sharpe sits in the null distribution, 0-100."""
|
| 42 |
+
if not self.null.size:
|
| 43 |
+
return 50.0
|
| 44 |
+
return float((self.null < self.observed).mean() * 100.0)
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
def _decompose(df: pd.DataFrame) -> tuple[np.ndarray, float]:
|
| 48 |
+
"""Split bars into scale-free log moves that can be reshuffled safely."""
|
| 49 |
+
open_ = df["open"].to_numpy(dtype=float)
|
| 50 |
+
high = df["high"].to_numpy(dtype=float)
|
| 51 |
+
low = df["low"].to_numpy(dtype=float)
|
| 52 |
+
close = df["close"].to_numpy(dtype=float)
|
| 53 |
+
volume = df["volume"].to_numpy(dtype=float)
|
| 54 |
+
|
| 55 |
+
gap = np.log(open_[1:] / close[:-1])
|
| 56 |
+
hi = np.log(np.maximum(high[1:], open_[1:]) / open_[1:])
|
| 57 |
+
lo = np.log(np.minimum(low[1:], open_[1:]) / open_[1:])
|
| 58 |
+
body = np.log(close[1:] / open_[1:])
|
| 59 |
+
return np.column_stack([gap, hi, lo, body, volume[1:]]), float(close[0])
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def _rebuild(parts: np.ndarray, anchor: float, index: pd.Index, first_row: pd.Series) -> pd.DataFrame:
|
| 63 |
+
gap, hi, lo, body, volume = (parts[:, i] for i in range(5))
|
| 64 |
+
n = parts.shape[0] + 1
|
| 65 |
+
|
| 66 |
+
close = np.empty(n)
|
| 67 |
+
open_ = np.empty(n)
|
| 68 |
+
high = np.empty(n)
|
| 69 |
+
low = np.empty(n)
|
| 70 |
+
vol = np.empty(n)
|
| 71 |
+
|
| 72 |
+
close[0] = anchor
|
| 73 |
+
open_[0] = float(first_row["open"])
|
| 74 |
+
high[0] = float(first_row["high"])
|
| 75 |
+
low[0] = float(first_row["low"])
|
| 76 |
+
vol[0] = float(first_row["volume"])
|
| 77 |
+
|
| 78 |
+
# Cumulative product form: close[i] = close[0] * exp(cumsum(gap + body)).
|
| 79 |
+
close[1:] = anchor * np.exp(np.cumsum(gap + body))
|
| 80 |
+
open_[1:] = close[:-1] * np.exp(gap)
|
| 81 |
+
high[1:] = open_[1:] * np.exp(hi)
|
| 82 |
+
low[1:] = open_[1:] * np.exp(lo)
|
| 83 |
+
vol[1:] = volume
|
| 84 |
+
|
| 85 |
+
return pd.DataFrame(
|
| 86 |
+
{"open": open_, "high": high, "low": low, "close": close, "volume": vol}, index=index
|
| 87 |
+
)
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def permute_bars(
|
| 91 |
+
df: pd.DataFrame,
|
| 92 |
+
rng: np.random.Generator,
|
| 93 |
+
method: str = "permute",
|
| 94 |
+
block: int = 20,
|
| 95 |
+
) -> pd.DataFrame:
|
| 96 |
+
"""Return a shuffled market with the same index and bar anatomy.
|
| 97 |
+
|
| 98 |
+
``permute`` reshuffles individual bars, destroying all serial structure.
|
| 99 |
+
``block`` resamples contiguous blocks with replacement, which preserves
|
| 100 |
+
short-horizon autocorrelation and volatility clustering — a harder null
|
| 101 |
+
that trend strategies deserve to be tested against.
|
| 102 |
+
"""
|
| 103 |
+
parts, anchor = _decompose(df)
|
| 104 |
+
m = parts.shape[0]
|
| 105 |
+
if m < 2:
|
| 106 |
+
return df.copy()
|
| 107 |
+
|
| 108 |
+
if method == "block":
|
| 109 |
+
size = max(2, min(int(block), m))
|
| 110 |
+
starts = rng.integers(0, m, size=int(np.ceil(m / size)))
|
| 111 |
+
order = np.concatenate([(np.arange(s, s + size) % m) for s in starts])[:m]
|
| 112 |
+
else:
|
| 113 |
+
order = rng.permutation(m)
|
| 114 |
+
|
| 115 |
+
return _rebuild(parts[order], anchor, df.index, df.iloc[0])
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def permutation_test(
|
| 119 |
+
df: pd.DataFrame,
|
| 120 |
+
signal_fn: Callable[[pd.DataFrame], pd.Series],
|
| 121 |
+
n_permutations: int = 300,
|
| 122 |
+
method: str = "permute",
|
| 123 |
+
block: int = 20,
|
| 124 |
+
costs: Optional[CostModel] = None,
|
| 125 |
+
lag: int = 1,
|
| 126 |
+
max_leverage: float = 1.0,
|
| 127 |
+
allow_short: bool = True,
|
| 128 |
+
seed: int = 0,
|
| 129 |
+
observed: Optional[float] = None,
|
| 130 |
+
progress: Optional[Callable[[float, str], None]] = None,
|
| 131 |
+
) -> PermutationResult:
|
| 132 |
+
"""Run ``signal_fn`` against ``n_permutations`` shuffled markets.
|
| 133 |
+
|
| 134 |
+
``signal_fn`` must be the strategy's target-exposure generator; it is
|
| 135 |
+
re-evaluated on every synthetic market, which is the whole point — a rule
|
| 136 |
+
that only works because of the specific path it was tuned on will fall
|
| 137 |
+
apart here.
|
| 138 |
+
"""
|
| 139 |
+
costs = costs or CostModel()
|
| 140 |
+
rng = np.random.default_rng(seed)
|
| 141 |
+
|
| 142 |
+
def sharpe_on(frame: pd.DataFrame) -> float:
|
| 143 |
+
target = signal_fn(frame)
|
| 144 |
+
result = run_backtest(
|
| 145 |
+
frame,
|
| 146 |
+
target,
|
| 147 |
+
costs=costs,
|
| 148 |
+
lag=lag,
|
| 149 |
+
max_leverage=max_leverage,
|
| 150 |
+
allow_short=allow_short,
|
| 151 |
+
)
|
| 152 |
+
return result.sharpe
|
| 153 |
+
|
| 154 |
+
if observed is None:
|
| 155 |
+
observed = sharpe_on(df)
|
| 156 |
+
|
| 157 |
+
null = np.empty(n_permutations, dtype=float)
|
| 158 |
+
for i in range(n_permutations):
|
| 159 |
+
null[i] = sharpe_on(permute_bars(df, rng, method, block))
|
| 160 |
+
if progress is not None and (i % 25 == 0 or i == n_permutations - 1):
|
| 161 |
+
progress((i + 1) / n_permutations, f"Permutation {i + 1}/{n_permutations}")
|
| 162 |
+
|
| 163 |
+
# +1 in both places: the observed result is itself one draw from the null
|
| 164 |
+
# under H0, which keeps the test from ever reporting an impossible p = 0.
|
| 165 |
+
p_value = float((1 + np.sum(null >= observed)) / (n_permutations + 1))
|
| 166 |
+
|
| 167 |
+
return PermutationResult(
|
| 168 |
+
observed=float(observed),
|
| 169 |
+
null=null,
|
| 170 |
+
p_value=p_value,
|
| 171 |
+
method=method,
|
| 172 |
+
n_permutations=n_permutations,
|
| 173 |
+
)
|
| 174 |
+
|
| 175 |
+
|
| 176 |
+
def bootstrap_return_paths(returns: pd.Series, n: int = 500, seed: int = 0) -> np.ndarray:
|
| 177 |
+
"""Bootstrap terminal-wealth outcomes from a realised return stream.
|
| 178 |
+
|
| 179 |
+
Useful for the "how wide is the cone of outcomes?" chart — the same edge
|
| 180 |
+
can produce wildly different equity curves.
|
| 181 |
+
"""
|
| 182 |
+
arr = np.asarray(returns.dropna(), dtype=float)
|
| 183 |
+
if arr.size == 0:
|
| 184 |
+
return np.zeros((n, 1))
|
| 185 |
+
rng = np.random.default_rng(seed)
|
| 186 |
+
draws = rng.choice(arr, size=(n, arr.size), replace=True)
|
| 187 |
+
return np.cumprod(1.0 + draws, axis=1)
|
|
@@ -0,0 +1,129 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Walk-forward analysis.
|
| 2 |
+
|
| 3 |
+
Re-tune on a training window, trade the next window blind, roll forward. The
|
| 4 |
+
gap between in-sample and out-of-sample Sharpe is the honest estimate of how
|
| 5 |
+
much of the backtest was curve-fitting.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
from typing import Callable, Dict, List, Optional
|
| 11 |
+
|
| 12 |
+
import numpy as np
|
| 13 |
+
import pandas as pd
|
| 14 |
+
|
| 15 |
+
from ..engine import run_backtest
|
| 16 |
+
from ..strategies import Strategy
|
| 17 |
+
from ..types import CostModel
|
| 18 |
+
|
| 19 |
+
__all__ = ["walk_forward"]
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
def walk_forward(
|
| 23 |
+
df: pd.DataFrame,
|
| 24 |
+
strategy: Strategy,
|
| 25 |
+
n_folds: int = 5,
|
| 26 |
+
train_ratio: float = 0.7,
|
| 27 |
+
costs: Optional[CostModel] = None,
|
| 28 |
+
lag: int = 1,
|
| 29 |
+
max_leverage: float = 1.0,
|
| 30 |
+
allow_short: bool = True,
|
| 31 |
+
grid_limit: int = 40,
|
| 32 |
+
progress: Optional[Callable[[float, str], None]] = None,
|
| 33 |
+
) -> Dict[str, object]:
|
| 34 |
+
"""Roll a train/test split forward ``n_folds`` times.
|
| 35 |
+
|
| 36 |
+
Each fold picks the best parameters by in-sample Sharpe and reports what
|
| 37 |
+
those parameters then did out of sample. The stitched OOS returns are the
|
| 38 |
+
closest thing to a paper-trading record this repo can produce offline.
|
| 39 |
+
"""
|
| 40 |
+
costs = costs or CostModel()
|
| 41 |
+
grid = strategy.grid(limit=grid_limit)
|
| 42 |
+
n = len(df)
|
| 43 |
+
|
| 44 |
+
if n < 250 or n_folds < 2:
|
| 45 |
+
return {"folds": [], "note": "Not enough history for walk-forward analysis."}
|
| 46 |
+
|
| 47 |
+
# Each fold is a contiguous train+test block; blocks advance by test length.
|
| 48 |
+
block = int(n / (1 + (n_folds - 1) * (1 - train_ratio)))
|
| 49 |
+
block = min(block, n)
|
| 50 |
+
train_len = int(block * train_ratio)
|
| 51 |
+
test_len = block - train_len
|
| 52 |
+
if train_len < 100 or test_len < 20:
|
| 53 |
+
return {"folds": [], "note": "Not enough history for walk-forward analysis."}
|
| 54 |
+
|
| 55 |
+
folds: List[Dict[str, object]] = []
|
| 56 |
+
oos_returns: List[pd.Series] = []
|
| 57 |
+
|
| 58 |
+
for k in range(n_folds):
|
| 59 |
+
start = k * test_len
|
| 60 |
+
train = df.iloc[start : start + train_len]
|
| 61 |
+
test = df.iloc[start + train_len : start + train_len + test_len]
|
| 62 |
+
if len(test) < 20:
|
| 63 |
+
break
|
| 64 |
+
|
| 65 |
+
best_params, best_sharpe = None, -np.inf
|
| 66 |
+
for params in grid:
|
| 67 |
+
target = strategy.generate(train, params)
|
| 68 |
+
sr = run_backtest(
|
| 69 |
+
train, target, costs=costs, lag=lag,
|
| 70 |
+
max_leverage=max_leverage, allow_short=allow_short,
|
| 71 |
+
).sharpe
|
| 72 |
+
if sr > best_sharpe:
|
| 73 |
+
best_params, best_sharpe = params, sr
|
| 74 |
+
|
| 75 |
+
# Generate signals on train+test so indicators are warm at the fold
|
| 76 |
+
# boundary, then evaluate only the test slice.
|
| 77 |
+
combined = df.iloc[start : start + train_len + len(test)]
|
| 78 |
+
target = strategy.generate(combined, best_params).loc[test.index]
|
| 79 |
+
oos = run_backtest(
|
| 80 |
+
test, target, costs=costs, lag=lag,
|
| 81 |
+
max_leverage=max_leverage, allow_short=allow_short,
|
| 82 |
+
)
|
| 83 |
+
|
| 84 |
+
folds.append(
|
| 85 |
+
{
|
| 86 |
+
"fold": k + 1,
|
| 87 |
+
"train_start": str(train.index[0].date()),
|
| 88 |
+
"train_end": str(train.index[-1].date()),
|
| 89 |
+
"test_start": str(test.index[0].date()),
|
| 90 |
+
"test_end": str(test.index[-1].date()),
|
| 91 |
+
"params": best_params,
|
| 92 |
+
"is_sharpe": float(best_sharpe),
|
| 93 |
+
"oos_sharpe": float(oos.sharpe),
|
| 94 |
+
"oos_return": float(oos.metrics.get("total_return", 0.0)),
|
| 95 |
+
"oos_max_dd": float(oos.metrics.get("max_drawdown", 0.0)),
|
| 96 |
+
}
|
| 97 |
+
)
|
| 98 |
+
oos_returns.append(oos.returns)
|
| 99 |
+
|
| 100 |
+
if progress is not None:
|
| 101 |
+
progress((k + 1) / n_folds, f"Walk-forward fold {k + 1}/{n_folds}")
|
| 102 |
+
|
| 103 |
+
if not folds:
|
| 104 |
+
return {"folds": [], "note": "Not enough history for walk-forward analysis."}
|
| 105 |
+
|
| 106 |
+
is_sharpes = np.array([f["is_sharpe"] for f in folds], dtype=float)
|
| 107 |
+
oos_sharpes = np.array([f["oos_sharpe"] for f in folds], dtype=float)
|
| 108 |
+
stitched = pd.concat(oos_returns) if oos_returns else pd.Series(dtype=float)
|
| 109 |
+
stitched = stitched[~stitched.index.duplicated(keep="first")].sort_index()
|
| 110 |
+
|
| 111 |
+
mean_is = float(np.mean(is_sharpes))
|
| 112 |
+
mean_oos = float(np.mean(oos_sharpes))
|
| 113 |
+
|
| 114 |
+
return {
|
| 115 |
+
"folds": folds,
|
| 116 |
+
"mean_is_sharpe": mean_is,
|
| 117 |
+
"mean_oos_sharpe": mean_oos,
|
| 118 |
+
# 1.0 = the edge fully survived; 0.0 = it evaporated out of sample.
|
| 119 |
+
"efficiency": float(mean_oos / mean_is) if mean_is > 1e-9 else 0.0,
|
| 120 |
+
"oos_win_rate": float(np.mean(oos_sharpes > 0)),
|
| 121 |
+
# 1.0 means every fold chose different parameters -- a tuning process
|
| 122 |
+
# that cannot make up its mind is fitting noise.
|
| 123 |
+
"param_instability": float(
|
| 124 |
+
len({str(f["params"]) for f in folds}) / max(len(folds), 1)
|
| 125 |
+
),
|
| 126 |
+
"oos_returns": stitched,
|
| 127 |
+
"oos_equity": (1.0 + stitched).cumprod() if len(stitched) else stitched,
|
| 128 |
+
"note": "",
|
| 129 |
+
}
|
|
@@ -0,0 +1,177 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Turn a pile of statistics into one number and one sentence.
|
| 2 |
+
|
| 3 |
+
The Reality Score is deliberately harsh. Most published backtests would score
|
| 4 |
+
below 40, and that is the point: the score exists to be screenshotted.
|
| 5 |
+
"""
|
| 6 |
+
|
| 7 |
+
from __future__ import annotations
|
| 8 |
+
|
| 9 |
+
from typing import Dict, List, Optional
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
|
| 13 |
+
__all__ = ["reality_score", "GRADES"]
|
| 14 |
+
|
| 15 |
+
GRADES = [
|
| 16 |
+
(85, "A", "Survives everything we threw at it"),
|
| 17 |
+
(70, "B", "Probably a real edge, with caveats"),
|
| 18 |
+
(55, "C", "Ambiguous — could go either way"),
|
| 19 |
+
(40, "D", "Mostly luck"),
|
| 20 |
+
(0, "F", "Indistinguishable from randomness"),
|
| 21 |
+
]
|
| 22 |
+
|
| 23 |
+
WEIGHTS = {
|
| 24 |
+
"significance": 0.30,
|
| 25 |
+
"selection": 0.25,
|
| 26 |
+
"walk_forward": 0.20,
|
| 27 |
+
"overfitting": 0.15,
|
| 28 |
+
"robustness": 0.10,
|
| 29 |
+
}
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
def _ramp(value: float, good: float, bad: float) -> float:
|
| 33 |
+
"""Linear 0-100 score where ``good`` maps to 100 and ``bad`` maps to 0."""
|
| 34 |
+
if not np.isfinite(value):
|
| 35 |
+
return 50.0
|
| 36 |
+
if good == bad:
|
| 37 |
+
return 50.0
|
| 38 |
+
scaled = (value - bad) / (good - bad)
|
| 39 |
+
return float(np.clip(scaled, 0.0, 1.0) * 100.0)
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def reality_score(
|
| 43 |
+
metrics: Dict[str, float],
|
| 44 |
+
benchmark_metrics: Dict[str, float],
|
| 45 |
+
p_value: Optional[float] = None,
|
| 46 |
+
dsr: Optional[float] = None,
|
| 47 |
+
pbo: Optional[float] = None,
|
| 48 |
+
wf_efficiency: Optional[float] = None,
|
| 49 |
+
wf_win_rate: Optional[float] = None,
|
| 50 |
+
cost_stress_ratio: Optional[float] = None,
|
| 51 |
+
benchmark_correlation: Optional[float] = None,
|
| 52 |
+
) -> Dict[str, object]:
|
| 53 |
+
"""Combine the validation panel into a 0-100 score, a grade and warnings."""
|
| 54 |
+
components: Dict[str, float] = {}
|
| 55 |
+
|
| 56 |
+
components["significance"] = _ramp(p_value, good=0.01, bad=0.50) if p_value is not None else 50.0
|
| 57 |
+
components["selection"] = float(np.clip(dsr, 0.0, 1.0) * 100.0) if dsr is not None else 50.0
|
| 58 |
+
|
| 59 |
+
if wf_efficiency is not None:
|
| 60 |
+
wf = _ramp(wf_efficiency, good=0.8, bad=-0.2)
|
| 61 |
+
if wf_win_rate is not None:
|
| 62 |
+
wf = 0.7 * wf + 0.3 * float(np.clip(wf_win_rate, 0.0, 1.0) * 100.0)
|
| 63 |
+
components["walk_forward"] = wf
|
| 64 |
+
else:
|
| 65 |
+
components["walk_forward"] = 50.0
|
| 66 |
+
|
| 67 |
+
components["overfitting"] = _ramp(pbo, good=0.05, bad=0.50) if pbo is not None and np.isfinite(pbo) else 50.0
|
| 68 |
+
components["robustness"] = (
|
| 69 |
+
_ramp(cost_stress_ratio, good=0.8, bad=0.0) if cost_stress_ratio is not None else 50.0
|
| 70 |
+
)
|
| 71 |
+
|
| 72 |
+
score = float(sum(components[k] * w for k, w in WEIGHTS.items()))
|
| 73 |
+
|
| 74 |
+
flags: List[str] = []
|
| 75 |
+
n_trades = int(metrics.get("n_trades", 0))
|
| 76 |
+
sharpe = float(metrics.get("sharpe", 0.0))
|
| 77 |
+
total_return = float(metrics.get("total_return", 0.0))
|
| 78 |
+
max_dd = float(metrics.get("max_drawdown", 0.0))
|
| 79 |
+
bench_sharpe = float(benchmark_metrics.get("sharpe", 0.0))
|
| 80 |
+
|
| 81 |
+
# The score asks "is the measured edge real?" — if the strategy lost money,
|
| 82 |
+
# there is no edge to validate, whatever the statistics say about it.
|
| 83 |
+
if total_return <= 0:
|
| 84 |
+
flags.append(
|
| 85 |
+
f"The strategy lost money over the test period ({total_return:.1%}). "
|
| 86 |
+
"There is no edge here to validate."
|
| 87 |
+
)
|
| 88 |
+
score = min(score, 50.0)
|
| 89 |
+
if total_return <= 0 < sharpe:
|
| 90 |
+
flags.append(
|
| 91 |
+
"Positive Sharpe with a negative total return: the average bar was profitable but "
|
| 92 |
+
"the compounding was not. Volatility drag ate the arithmetic edge."
|
| 93 |
+
)
|
| 94 |
+
if max_dd < -0.5:
|
| 95 |
+
flags.append(
|
| 96 |
+
f"Peak-to-trough drawdown of {max_dd:.0%} — an account running this would have been "
|
| 97 |
+
"closed long before the recovery arrived."
|
| 98 |
+
)
|
| 99 |
+
score = min(score, 60.0)
|
| 100 |
+
|
| 101 |
+
if n_trades < 20:
|
| 102 |
+
flags.append(
|
| 103 |
+
f"Only {n_trades} position changes — with this few decisions, "
|
| 104 |
+
"the result is a handful of coin flips, not a track record."
|
| 105 |
+
)
|
| 106 |
+
score = min(score, 55.0)
|
| 107 |
+
if p_value is not None and p_value > 0.10:
|
| 108 |
+
flags.append(
|
| 109 |
+
f"Permutation p-value is {p_value:.2f}: roughly {p_value * 100:.0f}% of shuffled, "
|
| 110 |
+
"structure-free markets did this well or better."
|
| 111 |
+
)
|
| 112 |
+
if dsr is not None and dsr < 0.5:
|
| 113 |
+
flags.append(
|
| 114 |
+
f"Deflated Sharpe is {dsr:.2f} — once you account for how many variants were tried, "
|
| 115 |
+
"the edge does not clear the selection-bias bar."
|
| 116 |
+
)
|
| 117 |
+
if pbo is not None and np.isfinite(pbo) and pbo > 0.3:
|
| 118 |
+
flags.append(
|
| 119 |
+
f"Probability of backtest overfitting is {pbo:.0%}: the in-sample winner "
|
| 120 |
+
"usually lands in the bottom half out of sample."
|
| 121 |
+
)
|
| 122 |
+
if wf_efficiency is not None and wf_efficiency < 0.3:
|
| 123 |
+
flags.append(
|
| 124 |
+
f"Walk-forward efficiency is {wf_efficiency:.0%} — most of the in-sample Sharpe "
|
| 125 |
+
"does not survive re-tuning and trading forward."
|
| 126 |
+
)
|
| 127 |
+
if cost_stress_ratio is not None and cost_stress_ratio < 0.5:
|
| 128 |
+
flags.append(
|
| 129 |
+
"Tripling trading costs removes more than half the Sharpe. The edge is "
|
| 130 |
+
"smaller than the friction it has to pay."
|
| 131 |
+
)
|
| 132 |
+
if benchmark_correlation is not None and benchmark_correlation > 0.95:
|
| 133 |
+
flags.append(
|
| 134 |
+
f"Returns are {benchmark_correlation:.0%} correlated with buy & hold — "
|
| 135 |
+
"this is mostly a repackaged long position."
|
| 136 |
+
)
|
| 137 |
+
if sharpe < bench_sharpe:
|
| 138 |
+
flags.append(
|
| 139 |
+
f"Buy & hold beat it on risk-adjusted return ({bench_sharpe:.2f} vs {sharpe:.2f} Sharpe)."
|
| 140 |
+
)
|
| 141 |
+
if float(metrics.get("turnover_ann", 0.0)) > 100:
|
| 142 |
+
flags.append(
|
| 143 |
+
f"Annual turnover of {metrics.get('turnover_ann', 0):.0f}x is far beyond what "
|
| 144 |
+
"retail execution can absorb without moving the modelled fills."
|
| 145 |
+
)
|
| 146 |
+
|
| 147 |
+
grade, headline = next((g, h) for threshold, g, h in GRADES if score >= threshold)
|
| 148 |
+
|
| 149 |
+
if score >= 70:
|
| 150 |
+
verdict = (
|
| 151 |
+
f"Grade {grade}. {headline}. The edge is still there after shuffling the market, "
|
| 152 |
+
"after charging for every variant tried, and after walking it forward."
|
| 153 |
+
)
|
| 154 |
+
elif score >= 55:
|
| 155 |
+
verdict = (
|
| 156 |
+
f"Grade {grade}. {headline}. Parts of the panel hold up and parts do not — "
|
| 157 |
+
"this is the zone where more data, not more tuning, is what settles it."
|
| 158 |
+
)
|
| 159 |
+
elif score >= 40:
|
| 160 |
+
verdict = (
|
| 161 |
+
f"Grade {grade}. {headline}. The backtest looks better than the evidence supports; "
|
| 162 |
+
"the gap between the two is selection bias."
|
| 163 |
+
)
|
| 164 |
+
else:
|
| 165 |
+
verdict = (
|
| 166 |
+
f"Grade {grade}. {headline}. A randomly shuffled market produces results like this "
|
| 167 |
+
"often enough that there is nothing here to trade."
|
| 168 |
+
)
|
| 169 |
+
|
| 170 |
+
return {
|
| 171 |
+
"score": round(score, 1),
|
| 172 |
+
"grade": grade,
|
| 173 |
+
"headline": headline,
|
| 174 |
+
"verdict": verdict,
|
| 175 |
+
"components": {k: round(v, 1) for k, v in components.items()},
|
| 176 |
+
"flags": flags,
|
| 177 |
+
}
|
|
@@ -0,0 +1,556 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Backtest Reality Check — the Hugging Face Space entrypoint.
|
| 2 |
+
|
| 3 |
+
Point it at a ticker and a trading rule. It runs the backtest, then spends the
|
| 4 |
+
rest of its time trying to prove the result was luck: shuffled markets,
|
| 5 |
+
selection-bias deflation, walk-forward, and a cost stress test.
|
| 6 |
+
|
| 7 |
+
Run locally with ``python app.py``.
|
| 8 |
+
"""
|
| 9 |
+
|
| 10 |
+
from __future__ import annotations
|
| 11 |
+
|
| 12 |
+
import logging
|
| 13 |
+
import os
|
| 14 |
+
from typing import Dict, List
|
| 15 |
+
|
| 16 |
+
import gradio as gr
|
| 17 |
+
import pandas as pd
|
| 18 |
+
|
| 19 |
+
from algotrader import __version__
|
| 20 |
+
from algotrader.charts import (
|
| 21 |
+
arena_chart,
|
| 22 |
+
drawdown_chart,
|
| 23 |
+
empty_figure,
|
| 24 |
+
equity_chart,
|
| 25 |
+
exposure_chart,
|
| 26 |
+
permutation_chart,
|
| 27 |
+
score_chart,
|
| 28 |
+
walkforward_chart,
|
| 29 |
+
)
|
| 30 |
+
from algotrader.data import DEFAULT_UNIVERSE
|
| 31 |
+
from algotrader.lab import LabConfig, run_arena, run_lab
|
| 32 |
+
from algotrader.strategies import REGISTRY, get_strategy
|
| 33 |
+
|
| 34 |
+
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s: %(message)s")
|
| 35 |
+
logger = logging.getLogger("app")
|
| 36 |
+
|
| 37 |
+
MAX_PARAMS = 3
|
| 38 |
+
STRATEGY_CHOICES = [(s.name, key) for key, s in REGISTRY.items()]
|
| 39 |
+
|
| 40 |
+
GRADE_COLORS = {
|
| 41 |
+
"A": "#0ca30c",
|
| 42 |
+
"B": "#3987e5",
|
| 43 |
+
"C": "#fab219",
|
| 44 |
+
"D": "#ec835a",
|
| 45 |
+
"F": "#d03b3b",
|
| 46 |
+
}
|
| 47 |
+
|
| 48 |
+
CSS = """
|
| 49 |
+
.gradio-container { max-width: 1280px !important; }
|
| 50 |
+
#hero h1 { font-size: 2.4rem; line-height: 1.1; margin: 0 0 .4rem 0; letter-spacing: -.02em; }
|
| 51 |
+
#hero p { color: #c3c2b7; margin: 0; font-size: 1.05rem; max-width: 60ch; }
|
| 52 |
+
.score-card {
|
| 53 |
+
display: flex; gap: 24px; align-items: center; padding: 22px 24px;
|
| 54 |
+
border: 1px solid rgba(255,255,255,.10); border-radius: 14px; background: #1a1a19;
|
| 55 |
+
}
|
| 56 |
+
.score-badge {
|
| 57 |
+
min-width: 132px; text-align: center; padding: 14px 10px; border-radius: 12px;
|
| 58 |
+
background: #0d0d0d; border: 1px solid rgba(255,255,255,.10);
|
| 59 |
+
}
|
| 60 |
+
.score-badge .grade { font-size: 3.2rem; font-weight: 700; line-height: 1; }
|
| 61 |
+
.score-badge .num { font-size: .95rem; color: #898781; margin-top: 6px; }
|
| 62 |
+
.score-body h3 { margin: 0 0 6px 0; font-size: 1.15rem; color: #fff; }
|
| 63 |
+
.score-body p { margin: 0; color: #c3c2b7; line-height: 1.55; }
|
| 64 |
+
.tiles { display: grid; grid-template-columns: repeat(auto-fit, minmax(132px,1fr)); gap: 10px; margin-top: 14px; }
|
| 65 |
+
.tile { padding: 12px 14px; border: 1px solid rgba(255,255,255,.10); border-radius: 10px; background: #1a1a19; }
|
| 66 |
+
.tile .label { font-size: .72rem; text-transform: uppercase; letter-spacing: .06em; color: #898781; }
|
| 67 |
+
.tile .value { font-size: 1.5rem; font-weight: 600; color: #fff; margin-top: 4px; }
|
| 68 |
+
.tile .sub { font-size: .75rem; color: #898781; margin-top: 2px; }
|
| 69 |
+
.flags { margin-top: 14px; padding: 0; list-style: none; }
|
| 70 |
+
.flags li {
|
| 71 |
+
padding: 9px 12px; margin-bottom: 7px; border-radius: 8px; background: #1a1a19;
|
| 72 |
+
border-left: 3px solid #ec835a; color: #c3c2b7; font-size: .9rem; line-height: 1.5;
|
| 73 |
+
}
|
| 74 |
+
.provenance { font-size: .82rem; color: #898781; margin-top: 10px; }
|
| 75 |
+
.provenance.sim { color: #fab219; }
|
| 76 |
+
"""
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
def _fmt_pct(value: float) -> str:
|
| 80 |
+
return f"{value * 100:,.1f}%"
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
def param_controls(strategy_key: str):
|
| 84 |
+
"""Re-label the shared sliders to match the selected strategy."""
|
| 85 |
+
strategy = get_strategy(strategy_key)
|
| 86 |
+
updates = []
|
| 87 |
+
for i in range(MAX_PARAMS):
|
| 88 |
+
if i < len(strategy.params):
|
| 89 |
+
spec = strategy.params[i]
|
| 90 |
+
lo = spec.minimum if spec.minimum is not None else min(spec.grid)
|
| 91 |
+
hi = spec.maximum if spec.maximum is not None else max(spec.grid)
|
| 92 |
+
updates.append(
|
| 93 |
+
gr.update(
|
| 94 |
+
visible=True,
|
| 95 |
+
label=spec.label,
|
| 96 |
+
value=spec.cast(spec.default),
|
| 97 |
+
minimum=lo,
|
| 98 |
+
maximum=hi,
|
| 99 |
+
step=spec.step or (1 if spec.kind == "int" else 0.1),
|
| 100 |
+
)
|
| 101 |
+
)
|
| 102 |
+
else:
|
| 103 |
+
updates.append(gr.update(visible=False))
|
| 104 |
+
return tuple(updates)
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def collect_params(strategy_key: str, *values) -> Dict[str, float]:
|
| 108 |
+
strategy = get_strategy(strategy_key)
|
| 109 |
+
return {spec.name: spec.cast(values[i]) for i, spec in enumerate(strategy.params[:MAX_PARAMS])}
|
| 110 |
+
|
| 111 |
+
|
| 112 |
+
def _score_card(report) -> str:
|
| 113 |
+
v = report.verdict
|
| 114 |
+
color = GRADE_COLORS.get(v["grade"], "#898781")
|
| 115 |
+
market = report.market
|
| 116 |
+
provenance = (
|
| 117 |
+
f'<div class="provenance">Data: {market.source} · {market.symbol} · '
|
| 118 |
+
f"{market.start.date()} to {market.end.date()} · {len(market.df):,} bars</div>"
|
| 119 |
+
if market.is_real
|
| 120 |
+
else f'<div class="provenance sim">⚠ {market.note}</div>'
|
| 121 |
+
)
|
| 122 |
+
flags = "".join(f"<li>{f}</li>" for f in v["flags"])
|
| 123 |
+
flags_html = f'<ul class="flags">{flags}</ul>' if flags else ""
|
| 124 |
+
|
| 125 |
+
return f"""
|
| 126 |
+
<div class="score-card">
|
| 127 |
+
<div class="score-badge">
|
| 128 |
+
<div class="grade" style="color:{color}">{v['grade']}</div>
|
| 129 |
+
<div class="num">{v['score']} / 100</div>
|
| 130 |
+
</div>
|
| 131 |
+
<div class="score-body">
|
| 132 |
+
<h3>{report.strategy.name} on {market.symbol}</h3>
|
| 133 |
+
<p>{v['verdict']}</p>
|
| 134 |
+
</div>
|
| 135 |
+
</div>
|
| 136 |
+
{flags_html}
|
| 137 |
+
{provenance}
|
| 138 |
+
"""
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
def _tiles(report) -> str:
|
| 142 |
+
m = report.backtest.metrics
|
| 143 |
+
b = report.backtest.benchmark_metrics
|
| 144 |
+
perm = report.permutation
|
| 145 |
+
p_text = f"{perm.p_value:.3f}" if perm else "—"
|
| 146 |
+
p_sub = "vs shuffled markets" if perm else "test skipped"
|
| 147 |
+
pbo = report.pbo.get("pbo")
|
| 148 |
+
pbo_text = f"{pbo:.0%}" if pbo is not None and pbo == pbo else "n/a"
|
| 149 |
+
|
| 150 |
+
cells = [
|
| 151 |
+
("Total return", _fmt_pct(m.get("total_return", 0)), f"buy & hold {_fmt_pct(b.get('total_return', 0))}"),
|
| 152 |
+
("CAGR", _fmt_pct(m.get("cagr", 0)), f"over {m.get('years', 0):.1f} years"),
|
| 153 |
+
("Sharpe", f"{m.get('sharpe', 0):.2f}", f"buy & hold {b.get('sharpe', 0):.2f}"),
|
| 154 |
+
("Max drawdown", _fmt_pct(m.get("max_drawdown", 0)), f"{m.get('time_under_water_yrs', 0):.1f}y under water"),
|
| 155 |
+
("p-value", p_text, p_sub),
|
| 156 |
+
("Deflated Sharpe", f"{report.dsr.get('dsr', 0):.2f}", f"after {report.trials.get('n', 1)} variants"),
|
| 157 |
+
("Overfit prob.", pbo_text, "in-sample winner fails OOS"),
|
| 158 |
+
("Trades", f"{int(m.get('n_trades', 0)):,}", f"{m.get('turnover_ann', 0):.0f}x turnover/yr"),
|
| 159 |
+
]
|
| 160 |
+
tiles = "".join(
|
| 161 |
+
f'<div class="tile"><div class="label">{label}</div>'
|
| 162 |
+
f'<div class="value">{value}</div><div class="sub">{sub}</div></div>'
|
| 163 |
+
for label, value, sub in cells
|
| 164 |
+
)
|
| 165 |
+
return f'<div class="tiles">{tiles}</div>'
|
| 166 |
+
|
| 167 |
+
|
| 168 |
+
def _detail_markdown(report) -> str:
|
| 169 |
+
dsr, wf, pbo, stress = report.dsr, report.walkforward, report.pbo, report.cost_stress
|
| 170 |
+
mtr = dsr.get("min_track_record_years", float("inf"))
|
| 171 |
+
mtr_text = f"{mtr:.1f} years" if mtr == mtr and mtr != float("inf") else "never, at this effect size"
|
| 172 |
+
|
| 173 |
+
lines = [
|
| 174 |
+
"### Reading the evidence",
|
| 175 |
+
"",
|
| 176 |
+
f"**Selection bias.** {report.trials.get('n', 1)} parameter variants of "
|
| 177 |
+
f"*{report.strategy.name}* were backtested. The luckiest skill-free variant of that many "
|
| 178 |
+
f"would be expected to show an annualised Sharpe of about "
|
| 179 |
+
f"**{dsr.get('threshold_sr_annual', 0):.2f}** on its own. Yours was "
|
| 180 |
+
f"**{report.backtest.sharpe:.2f}**, which puts the deflated probability of a real edge at "
|
| 181 |
+
f"**{dsr.get('dsr', 0):.0%}**.",
|
| 182 |
+
"",
|
| 183 |
+
f"**Track record needed.** To call this Sharpe significant at 95% confidence given its "
|
| 184 |
+
f"skew ({dsr.get('skew', 0):.2f}) and kurtosis ({dsr.get('kurtosis', 0):.1f}), you would need "
|
| 185 |
+
f"about **{mtr_text}** of live returns.",
|
| 186 |
+
"",
|
| 187 |
+
f"**Costs.** At the modelled friction the Sharpe is {stress.get('sharpe_1x', 0):.2f}. "
|
| 188 |
+
f"Triple the costs and it becomes {stress.get('sharpe_3x', 0):.2f} "
|
| 189 |
+
f"({stress.get('ratio', 0):.0%} retained).",
|
| 190 |
+
"",
|
| 191 |
+
]
|
| 192 |
+
|
| 193 |
+
if wf.get("folds"):
|
| 194 |
+
lines += [
|
| 195 |
+
f"**Walk-forward.** Across {len(wf['folds'])} folds the tuned in-sample Sharpe averaged "
|
| 196 |
+
f"{wf.get('mean_is_sharpe', 0):.2f} and the blind out-of-sample Sharpe averaged "
|
| 197 |
+
f"{wf.get('mean_oos_sharpe', 0):.2f} — an efficiency of {wf.get('efficiency', 0):.0%}. "
|
| 198 |
+
f"{wf.get('oos_win_rate', 0):.0%} of folds were profitable out of sample, and the tuner "
|
| 199 |
+
f"picked a different parameter set in {wf.get('param_instability', 0):.0%} of them.",
|
| 200 |
+
"",
|
| 201 |
+
]
|
| 202 |
+
if pbo.get("note"):
|
| 203 |
+
lines += [f"**Overfitting.** {pbo['note']}", ""]
|
| 204 |
+
elif pbo.get("pbo") == pbo.get("pbo"):
|
| 205 |
+
lines += [
|
| 206 |
+
f"**Overfitting (CSCV).** Over {pbo.get('n_combinations', 0)} in/out splits of "
|
| 207 |
+
f"{pbo.get('n_strategies', 0)} variants, the in-sample winner landed in the bottom half "
|
| 208 |
+
f"out of sample **{pbo.get('pbo', 0):.0%}** of the time, and lost money outright "
|
| 209 |
+
f"{pbo.get('prob_oos_loss', 0):.0%} of the time. The most frequently selected variant was "
|
| 210 |
+
f"`{pbo.get('most_selected_label', 'n/a')}` "
|
| 211 |
+
f"({pbo.get('selection_stability', 0):.0%} of splits).",
|
| 212 |
+
"",
|
| 213 |
+
]
|
| 214 |
+
|
| 215 |
+
lines += [
|
| 216 |
+
"> Past performance, simulated or otherwise, does not predict future returns. "
|
| 217 |
+
"This is research tooling, not investment advice.",
|
| 218 |
+
]
|
| 219 |
+
return "\n".join(lines)
|
| 220 |
+
|
| 221 |
+
|
| 222 |
+
def analyse(
|
| 223 |
+
symbol: str,
|
| 224 |
+
start: str,
|
| 225 |
+
end: str,
|
| 226 |
+
strategy_key: str,
|
| 227 |
+
p1: float,
|
| 228 |
+
p2: float,
|
| 229 |
+
p3: float,
|
| 230 |
+
commission: float,
|
| 231 |
+
slippage: float,
|
| 232 |
+
allow_short: bool,
|
| 233 |
+
n_permutations: int,
|
| 234 |
+
perm_method: str,
|
| 235 |
+
wf_folds: int,
|
| 236 |
+
progress=gr.Progress(),
|
| 237 |
+
):
|
| 238 |
+
"""Main Lab handler. Never raises into the UI — it returns a readable message."""
|
| 239 |
+
try:
|
| 240 |
+
cfg = LabConfig(
|
| 241 |
+
symbol=symbol or "SPY",
|
| 242 |
+
start=start or "2015-01-01",
|
| 243 |
+
end=end or None,
|
| 244 |
+
strategy=strategy_key,
|
| 245 |
+
params=collect_params(strategy_key, p1, p2, p3),
|
| 246 |
+
commission_bps=float(commission),
|
| 247 |
+
slippage_bps=float(slippage),
|
| 248 |
+
allow_short=bool(allow_short),
|
| 249 |
+
n_permutations=int(n_permutations),
|
| 250 |
+
permutation_method="block" if perm_method.startswith("Block") else "permute",
|
| 251 |
+
wf_folds=int(wf_folds),
|
| 252 |
+
)
|
| 253 |
+
report = run_lab(cfg, progress=lambda f, m: progress(f, desc=m))
|
| 254 |
+
except Exception as exc: # noqa: BLE001 - the UI must always say something useful
|
| 255 |
+
logger.exception("Lab run failed")
|
| 256 |
+
message = f'<div class="score-card"><div class="score-body"><h3>Could not run that</h3><p>{exc}</p></div></div>'
|
| 257 |
+
blank = empty_figure("No results.")
|
| 258 |
+
return message, "", blank, blank, blank, blank, blank, blank, ""
|
| 259 |
+
|
| 260 |
+
return (
|
| 261 |
+
_score_card(report),
|
| 262 |
+
_tiles(report),
|
| 263 |
+
permutation_chart(report),
|
| 264 |
+
equity_chart(report),
|
| 265 |
+
drawdown_chart(report),
|
| 266 |
+
exposure_chart(report),
|
| 267 |
+
walkforward_chart(report),
|
| 268 |
+
score_chart(report.verdict),
|
| 269 |
+
_detail_markdown(report),
|
| 270 |
+
)
|
| 271 |
+
|
| 272 |
+
|
| 273 |
+
def race(symbol: str, start: str, allow_short: bool, n_permutations: int, progress=gr.Progress()):
|
| 274 |
+
try:
|
| 275 |
+
cfg = LabConfig(
|
| 276 |
+
symbol=symbol or "SPY",
|
| 277 |
+
start=start or "2015-01-01",
|
| 278 |
+
allow_short=bool(allow_short),
|
| 279 |
+
)
|
| 280 |
+
table, market, _ = run_arena(
|
| 281 |
+
cfg,
|
| 282 |
+
n_permutations=int(n_permutations),
|
| 283 |
+
progress=lambda f, m: progress(f, desc=m),
|
| 284 |
+
)
|
| 285 |
+
except Exception as exc: # noqa: BLE001
|
| 286 |
+
logger.exception("Arena run failed")
|
| 287 |
+
return pd.DataFrame({"Error": [str(exc)]}), empty_figure("No results."), ""
|
| 288 |
+
|
| 289 |
+
display = table.copy()
|
| 290 |
+
for col in ("Return", "CAGR", "MaxDD"):
|
| 291 |
+
display[col] = display[col].map(lambda v: f"{v * 100:,.1f}%")
|
| 292 |
+
for col in ("Sharpe", "DSR", "Evidence"):
|
| 293 |
+
display[col] = display[col].map(lambda v: f"{v:.2f}")
|
| 294 |
+
display["p-value"] = display["p-value"].map(lambda v: "—" if v != v else f"{v:.3f}")
|
| 295 |
+
display = display.drop(columns=["key"])
|
| 296 |
+
|
| 297 |
+
provenance = (
|
| 298 |
+
f"Data: {market.source} · {market.symbol} · {market.start.date()} to {market.end.date()}"
|
| 299 |
+
if market.is_real
|
| 300 |
+
else f"⚠ {market.note}"
|
| 301 |
+
)
|
| 302 |
+
note = (
|
| 303 |
+
f"{provenance}\n\nRanked by **evidence** — `(1 − p) × deflated Sharpe` — not by return. "
|
| 304 |
+
"*Buy & Hold* and *Coin Flip* are in the field on purpose: a leaderboard without a "
|
| 305 |
+
"control group is marketing, not measurement."
|
| 306 |
+
)
|
| 307 |
+
return display, arena_chart(table), note
|
| 308 |
+
|
| 309 |
+
|
| 310 |
+
HOW_IT_WORKS = """
|
| 311 |
+
## Why most backtests are wrong
|
| 312 |
+
|
| 313 |
+
A backtest is a measurement taken with a ruler you built after seeing the thing you
|
| 314 |
+
are measuring. Four failure modes do almost all the damage, and this Space tests for
|
| 315 |
+
each one.
|
| 316 |
+
|
| 317 |
+
### 1. The market had no structure to find — permutation test
|
| 318 |
+
|
| 319 |
+
We take the real price series and shuffle it: each bar's gap, high, low, body and
|
| 320 |
+
volume are kept intact, but their **order** is destroyed. The result is a market with
|
| 321 |
+
the same volatility and the same fat tails, and no exploitable structure whatsoever.
|
| 322 |
+
Then we re-run *your exact rule* on hundreds of these shuffled markets.
|
| 323 |
+
|
| 324 |
+
If your Sharpe sits inside that cloud of results, your rule found nothing that a
|
| 325 |
+
coin-flip market would not also have handed it. The **p-value** is the share of
|
| 326 |
+
shuffled markets that did as well or better.
|
| 327 |
+
|
| 328 |
+
*Block mode* resamples contiguous chunks instead of single bars, preserving
|
| 329 |
+
short-horizon momentum and volatility clustering. It is a harder null, and trend
|
| 330 |
+
strategies should be held to it.
|
| 331 |
+
|
| 332 |
+
### 2. You tried 200 things and reported the best — Deflated Sharpe Ratio
|
| 333 |
+
|
| 334 |
+
If you test 200 worthless strategies, the best of them will show a Sharpe near 1.0
|
| 335 |
+
purely by chance. The **Deflated Sharpe Ratio** (Bailey & López de Prado, 2014) works
|
| 336 |
+
out what the luckiest of *N* skill-free variants would have scored, and asks whether
|
| 337 |
+
yours beats that bar — with an extra penalty for negative skew and fat tails, the
|
| 338 |
+
return shapes that flatter naive Sharpe ratios.
|
| 339 |
+
|
| 340 |
+
This Space counts the whole parameter grid as trials, because that is what a
|
| 341 |
+
researcher would really have run.
|
| 342 |
+
|
| 343 |
+
### 3. The parameters were fitted to the past — PBO and walk-forward
|
| 344 |
+
|
| 345 |
+
**Probability of Backtest Overfitting** (CSCV) cuts the timeline into chunks, and for
|
| 346 |
+
every way of splitting them half in-sample and half out-of-sample, checks whether the
|
| 347 |
+
in-sample winner stayed a winner. If the winner lands in the bottom half about half
|
| 348 |
+
the time, PBO ≈ 50% and your selection process has no skill at all.
|
| 349 |
+
|
| 350 |
+
**Walk-forward** re-tunes on a training window and trades the next window blind,
|
| 351 |
+
rolling forward. Efficiency is out-of-sample Sharpe over in-sample Sharpe: 100% means
|
| 352 |
+
the edge survived intact, 0% means it was entirely curve-fit.
|
| 353 |
+
|
| 354 |
+
### 4. The edge is smaller than the costs — stress test
|
| 355 |
+
|
| 356 |
+
Every result here is net of commission and slippage charged on exposure changes, plus
|
| 357 |
+
a borrow fee on short positions. We then re-run at **triple** the friction. A real edge
|
| 358 |
+
degrades; a fake one disappears.
|
| 359 |
+
|
| 360 |
+
---
|
| 361 |
+
|
| 362 |
+
## The Reality Score
|
| 363 |
+
|
| 364 |
+
| Weight | Component | What it measures |
|
| 365 |
+
|---:|---|---|
|
| 366 |
+
| 30% | Significance | How far outside the shuffled-market null the result sits |
|
| 367 |
+
| 25% | Selection | Deflated Sharpe — does it clear the best-of-N bar |
|
| 368 |
+
| 20% | Walk-forward | How much of the tuned Sharpe survived trading forward |
|
| 369 |
+
| 15% | Overfitting | 1 − PBO, from combinatorially symmetric cross-validation |
|
| 370 |
+
| 10% | Robustness | Sharpe retained when costs triple |
|
| 371 |
+
|
| 372 |
+
Grades: **A** ≥ 85 · **B** ≥ 70 · **C** ≥ 55 · **D** ≥ 40 · **F** below 40.
|
| 373 |
+
|
| 374 |
+
The scale is deliberately harsh. Most strategies people post online score below 40,
|
| 375 |
+
and the honest response to that is not to soften the scale.
|
| 376 |
+
|
| 377 |
+
## No look-ahead, by construction
|
| 378 |
+
|
| 379 |
+
A strategy emits a target exposure at each bar's close using only data up to that
|
| 380 |
+
bar. The engine holds `position[t] = target[t - lag]` with `lag ≥ 1`, so a signal
|
| 381 |
+
computed on Tuesday's close cannot earn Tuesday's move. That is the single line where
|
| 382 |
+
look-ahead could enter, and the test suite asserts it directly.
|
| 383 |
+
|
| 384 |
+
## Use it from Python
|
| 385 |
+
|
| 386 |
+
```python
|
| 387 |
+
from algotrader import LabConfig, run_lab
|
| 388 |
+
|
| 389 |
+
report = run_lab(LabConfig(symbol="SPY", strategy="sma_cross", params={"fast": 20, "slow": 100}))
|
| 390 |
+
print(report.verdict["grade"], report.verdict["score"])
|
| 391 |
+
print(report.permutation.p_value, report.dsr["dsr"], report.pbo["pbo"])
|
| 392 |
+
```
|
| 393 |
+
|
| 394 |
+
Or from the command line:
|
| 395 |
+
|
| 396 |
+
```bash
|
| 397 |
+
python -m algotrader.cli lab --symbol SPY --strategy donchian_breakout --permutations 500
|
| 398 |
+
python -m algotrader.cli arena --symbol BTC-USD
|
| 399 |
+
```
|
| 400 |
+
|
| 401 |
+
---
|
| 402 |
+
|
| 403 |
+
*Research tooling, not investment advice. Nothing here is a recommendation to trade.*
|
| 404 |
+
"""
|
| 405 |
+
|
| 406 |
+
|
| 407 |
+
# Gradio 6 moved `css` and `theme` from the Blocks constructor to launch().
|
| 408 |
+
# Spaces pin their own version, so pass them wherever the installed one wants.
|
| 409 |
+
_GRADIO_MAJOR = int(gr.__version__.split(".")[0])
|
| 410 |
+
_STYLE_KWARGS = {"css": CSS, "theme": gr.themes.Base()}
|
| 411 |
+
_BLOCKS_KWARGS = {} if _GRADIO_MAJOR >= 6 else _STYLE_KWARGS
|
| 412 |
+
# Gradio 6 also dropped launch(show_api=...).
|
| 413 |
+
_LAUNCH_KWARGS = dict(_STYLE_KWARGS) if _GRADIO_MAJOR >= 6 else {"show_api": False}
|
| 414 |
+
|
| 415 |
+
|
| 416 |
+
def build_app() -> gr.Blocks:
|
| 417 |
+
with gr.Blocks(title="Backtest Reality Check", **_BLOCKS_KWARGS) as demo:
|
| 418 |
+
with gr.Column(elem_id="hero"):
|
| 419 |
+
gr.HTML(
|
| 420 |
+
"<h1>Backtest Reality Check</h1>"
|
| 421 |
+
"<p>Your backtest is probably lying to you. Pick a market and a trading rule — "
|
| 422 |
+
"this runs it, then spends the rest of its effort trying to prove the result "
|
| 423 |
+
"was luck.</p>"
|
| 424 |
+
)
|
| 425 |
+
|
| 426 |
+
with gr.Tabs():
|
| 427 |
+
with gr.Tab("The Lab"):
|
| 428 |
+
with gr.Row():
|
| 429 |
+
with gr.Column(scale=1):
|
| 430 |
+
symbol = gr.Dropdown(
|
| 431 |
+
choices=DEFAULT_UNIVERSE, value="SPY", label="Ticker",
|
| 432 |
+
allow_custom_value=True,
|
| 433 |
+
info="Any Yahoo Finance symbol. Falls back to a simulated market if offline.",
|
| 434 |
+
)
|
| 435 |
+
with gr.Row():
|
| 436 |
+
start = gr.Textbox(value="2015-01-01", label="Start", scale=1)
|
| 437 |
+
end = gr.Textbox(value="", label="End (blank = today)", scale=1)
|
| 438 |
+
|
| 439 |
+
strategy = gr.Dropdown(
|
| 440 |
+
choices=STRATEGY_CHOICES, value="sma_cross", label="Strategy"
|
| 441 |
+
)
|
| 442 |
+
strategy_note = gr.Markdown(get_strategy("sma_cross").description)
|
| 443 |
+
|
| 444 |
+
param_sliders = [
|
| 445 |
+
gr.Slider(label=f"Parameter {i + 1}", visible=False, minimum=0, maximum=100)
|
| 446 |
+
for i in range(MAX_PARAMS)
|
| 447 |
+
]
|
| 448 |
+
|
| 449 |
+
with gr.Accordion("Costs and testing", open=False):
|
| 450 |
+
commission = gr.Slider(0, 20, value=1, step=0.5, label="Commission (bps per trade)")
|
| 451 |
+
slippage = gr.Slider(0, 50, value=2, step=0.5, label="Slippage (bps per trade)")
|
| 452 |
+
allow_short = gr.Checkbox(value=True, label="Allow short positions")
|
| 453 |
+
n_perms = gr.Slider(
|
| 454 |
+
0, 1000, value=250, step=50, label="Shuffled markets to test against",
|
| 455 |
+
info="More is stricter and slower. 250 is plenty for a first look.",
|
| 456 |
+
)
|
| 457 |
+
perm_method = gr.Radio(
|
| 458 |
+
["Shuffle bars (standard)", "Block bootstrap (harder)"],
|
| 459 |
+
value="Shuffle bars (standard)", label="Null market",
|
| 460 |
+
)
|
| 461 |
+
wf_folds = gr.Slider(2, 8, value=5, step=1, label="Walk-forward folds")
|
| 462 |
+
|
| 463 |
+
run_button = gr.Button("Run reality check", variant="primary", size="lg")
|
| 464 |
+
|
| 465 |
+
gr.Examples(
|
| 466 |
+
label="Or try one of these",
|
| 467 |
+
examples=[
|
| 468 |
+
["SPY", "sma_cross"],
|
| 469 |
+
["BTC-USD", "donchian_breakout"],
|
| 470 |
+
["NVDA", "rsi_reversion"],
|
| 471 |
+
["QQQ", "momentum"],
|
| 472 |
+
["SPY", "coin_flip"],
|
| 473 |
+
],
|
| 474 |
+
inputs=[symbol, strategy],
|
| 475 |
+
)
|
| 476 |
+
|
| 477 |
+
with gr.Column(scale=2):
|
| 478 |
+
verdict_html = gr.HTML(
|
| 479 |
+
'<div class="score-card"><div class="score-body">'
|
| 480 |
+
"<h3>Nothing tested yet</h3><p>Pick a market and a rule, then hit "
|
| 481 |
+
"<b>Run reality check</b>. A full run is a few seconds.</p>"
|
| 482 |
+
"</div></div>"
|
| 483 |
+
)
|
| 484 |
+
tiles_html = gr.HTML("")
|
| 485 |
+
|
| 486 |
+
# The headline test gets the full width — it is the whole point.
|
| 487 |
+
perm_plot = gr.Plot(value=empty_figure("The headline test appears here.", height=320))
|
| 488 |
+
equity_plot = gr.Plot(value=empty_figure())
|
| 489 |
+
with gr.Row():
|
| 490 |
+
dd_plot = gr.Plot(value=empty_figure(height=240))
|
| 491 |
+
exposure_plot = gr.Plot(value=empty_figure(height=200))
|
| 492 |
+
with gr.Row():
|
| 493 |
+
wf_plot = gr.Plot(value=empty_figure(height=300))
|
| 494 |
+
components_plot = gr.Plot(value=empty_figure(height=260))
|
| 495 |
+
detail_md = gr.Markdown("")
|
| 496 |
+
|
| 497 |
+
strategy.change(
|
| 498 |
+
fn=param_controls, inputs=strategy, outputs=param_sliders
|
| 499 |
+
).then(
|
| 500 |
+
fn=lambda k: get_strategy(k).description, inputs=strategy, outputs=strategy_note
|
| 501 |
+
)
|
| 502 |
+
|
| 503 |
+
run_button.click(
|
| 504 |
+
fn=analyse,
|
| 505 |
+
inputs=[
|
| 506 |
+
symbol, start, end, strategy, *param_sliders,
|
| 507 |
+
commission, slippage, allow_short, n_perms, perm_method, wf_folds,
|
| 508 |
+
],
|
| 509 |
+
outputs=[
|
| 510 |
+
verdict_html, tiles_html, perm_plot, equity_plot,
|
| 511 |
+
dd_plot, exposure_plot, wf_plot, components_plot, detail_md,
|
| 512 |
+
],
|
| 513 |
+
)
|
| 514 |
+
|
| 515 |
+
with gr.Tab("Arena"):
|
| 516 |
+
gr.Markdown(
|
| 517 |
+
"Race every strategy on the same market, ranked by **evidence** rather than "
|
| 518 |
+
"return. Buy & hold and a coin flip stay in the field as controls."
|
| 519 |
+
)
|
| 520 |
+
with gr.Row():
|
| 521 |
+
arena_symbol = gr.Dropdown(
|
| 522 |
+
choices=DEFAULT_UNIVERSE, value="SPY", label="Ticker", allow_custom_value=True
|
| 523 |
+
)
|
| 524 |
+
arena_start = gr.Textbox(value="2015-01-01", label="Start")
|
| 525 |
+
arena_short = gr.Checkbox(value=True, label="Allow shorts")
|
| 526 |
+
arena_perms = gr.Slider(0, 400, value=120, step=20, label="Shuffled markets per strategy")
|
| 527 |
+
arena_button = gr.Button("Run the arena", variant="primary")
|
| 528 |
+
arena_note = gr.Markdown("")
|
| 529 |
+
arena_table = gr.Dataframe(interactive=False, wrap=True)
|
| 530 |
+
arena_plot = gr.Plot(value=empty_figure(height=380))
|
| 531 |
+
|
| 532 |
+
arena_button.click(
|
| 533 |
+
fn=race,
|
| 534 |
+
inputs=[arena_symbol, arena_start, arena_short, arena_perms],
|
| 535 |
+
outputs=[arena_table, arena_plot, arena_note],
|
| 536 |
+
)
|
| 537 |
+
|
| 538 |
+
with gr.Tab("How it works"):
|
| 539 |
+
gr.Markdown(HOW_IT_WORKS)
|
| 540 |
+
|
| 541 |
+
gr.Markdown(
|
| 542 |
+
f"<sub>algotrader {__version__} · Apache-2.0 · "
|
| 543 |
+
"Research tooling, not investment advice.</sub>"
|
| 544 |
+
)
|
| 545 |
+
|
| 546 |
+
demo.load(fn=param_controls, inputs=strategy, outputs=param_sliders)
|
| 547 |
+
|
| 548 |
+
return demo
|
| 549 |
+
|
| 550 |
+
|
| 551 |
+
if __name__ == "__main__":
|
| 552 |
+
build_app().queue(max_size=24).launch(
|
| 553 |
+
server_name="0.0.0.0",
|
| 554 |
+
server_port=int(os.environ.get("PORT", 7860)),
|
| 555 |
+
**_LAUNCH_KWARGS,
|
| 556 |
+
)
|
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Algorithmic Trading
|
| 2 |
+
|
| 3 |
+
FinRL reinforcement-learning trading with Alpaca execution, plus optional Yahoo Finance OHLCV for unlabeled real-price research. Parallel LLC.
|
| 4 |
+
|
| 5 |
+
This is **research and paper-trading infrastructure**. Live capital requires a separate evaluation contract, feature-parity tests, and a rewritten execution path. Do not treat `paper_trading: false` as a promotion gate.
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## 1. Title and Summary
|
| 10 |
+
|
| 11 |
+
**Algorithmic Trading**
|
| 12 |
+
Northwestern-trained data-engineering practice applied to a trading loop: ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
|
| 13 |
+
|
| 14 |
+
GitHub `main` is the FinRL / Docker / Streamlit tree. `dev` is the integration branch. Yahoo is an additive `data_source.type`, not a replacement for Alpaca or FinRL.
|
| 15 |
+
|
| 16 |
+
**Design themes**
|
| 17 |
+
|
| 18 |
+
* Four ingest paths: CSV replay, synthetic GBM, Alpaca REST, Yahoo (`yfinance>=1.0`)
|
| 19 |
+
* FinRL policies (PPO, A2C, DDPG, TD3) on a Gymnasium-style environment
|
| 20 |
+
* Alpaca for authenticated market data and order routing (paper by default)
|
| 21 |
+
* Yahoo for delayed public bars when no broker key is available
|
| 22 |
+
* Secrets from environment (`ALPACA_API_KEY`, `ALPACA_SECRET_KEY`), never committed
|
| 23 |
+
* Tests and Docker/CI as already present on this tree
|
| 24 |
+
|
| 25 |
+
---
|
| 26 |
+
|
| 27 |
+
## 2. Concepts and Methods
|
| 28 |
+
|
| 29 |
+
### Market data
|
| 30 |
+
|
| 31 |
+
| Source | When to use | Failure modes |
|
| 32 |
+
| ------ | ----------- | ------------- |
|
| 33 |
+
| **CSV** | Offline replay; default in `config.yaml` | Missing path or OHLCV columns → `None` |
|
| 34 |
+
| **Synthetic** | Unit tests and demos | GBM is not tradable edge |
|
| 35 |
+
| **Alpaca** | Authenticated bars and live/paper orders | Auth, feed, and rate-limit failures |
|
| 36 |
+
| **Yahoo** | Real Close without a broker account | Unofficial API, ~15 min delay, interval lookback caps (1m ≈ 7 days). Pin `yfinance>=1.0`; 0.2.x fails against the current chart API |
|
| 37 |
+
|
| 38 |
+
`load_data` dispatches on `data_source.type`. Existing `alpaca` / `csv` / `synthetic` branches are unchanged.
|
| 39 |
+
|
| 40 |
+
### Strategy and FinRL
|
| 41 |
+
|
| 42 |
+
* `StrategyAgent`: SMA, RSI, Bollinger, MACD on Close; teaching rule, not an alpha claim
|
| 43 |
+
* `FinRLAgent`: PPO / A2C / DDPG / TD3 via Stable-Baselines3; persist under `models/`
|
| 44 |
+
* `ExecutionAgent` / `AlpacaBroker`: paper simulation or Alpaca market/limit orders
|
| 45 |
+
|
| 46 |
+
Backtests in this repo are in-sample passes unless you add a purged walk-forward yourself. Leakage is the null hypothesis.
|
| 47 |
+
|
| 48 |
+
---
|
| 49 |
+
|
| 50 |
+
## 3. Stack
|
| 51 |
+
|
| 52 |
+
| Layer | Tools |
|
| 53 |
+
| ----- | ----- |
|
| 54 |
+
| Language | Python 3.11 (CI); 3.8+ stated for local |
|
| 55 |
+
| RL | FinRL / Stable-Baselines3, Gym/Gymnasium, PyTorch |
|
| 56 |
+
| Broker | alpaca-py |
|
| 57 |
+
| Market data | Alpaca REST; yfinance ≥ 1.0 (Yahoo) |
|
| 58 |
+
| Tabular | pandas, NumPy, scikit-learn |
|
| 59 |
+
| UI | Streamlit, Dash, Jupyter widgets |
|
| 60 |
+
| Deploy | Docker Compose, GitHub Actions |
|
| 61 |
+
| Tests | pytest, pytest-cov |
|
| 62 |
+
|
| 63 |
+
---
|
| 64 |
+
|
| 65 |
+
## 4. Structure
|
| 66 |
+
|
| 67 |
+
```
|
| 68 |
+
algorithmic_trading/
|
| 69 |
+
├── agentic_ai_system/ # ingest, strategy, FinRL, Alpaca, Yahoo
|
| 70 |
+
├── ui/ # Streamlit, Dash, Jupyter, WebSocket
|
| 71 |
+
├── tests/
|
| 72 |
+
├── models/ # trained artifacts (gitignored bodies)
|
| 73 |
+
├── data/ # generated CSV (gitignored)
|
| 74 |
+
├── scripts/ # Docker / deploy helpers
|
| 75 |
+
├── .github/workflows/ # CI/CD, release, backtesting
|
| 76 |
+
├── config.yaml
|
| 77 |
+
├── requirements.txt
|
| 78 |
+
├── Dockerfile
|
| 79 |
+
└── docker-compose*.yml
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
Branch policy: **`main`** (protected) and **`dev`** only. Do not re-enable Dependabot or the Monday `dependency-updates` workflow; those created extra branches.
|
| 83 |
+
|
| 84 |
+
---
|
| 85 |
+
|
| 86 |
+
## 5. Quick start
|
| 87 |
+
|
| 88 |
+
```bash
|
| 89 |
+
git clone https://github.com/ParallelLLC/algorithmic_trading.git
|
| 90 |
+
cd algorithmic_trading
|
| 91 |
+
python -m venv .venv && source .venv/bin/activate
|
| 92 |
+
pip install -r requirements.txt
|
| 93 |
+
cp .env.example .env # Alpaca keys if using alpaca ingest or orders
|
| 94 |
+
```
|
| 95 |
+
|
| 96 |
+
Default ingest is CSV. For Yahoo daily bars without a broker:
|
| 97 |
+
|
| 98 |
+
```yaml
|
| 99 |
+
data_source:
|
| 100 |
+
type: 'yahoo'
|
| 101 |
+
trading:
|
| 102 |
+
symbol: 'AAPL'
|
| 103 |
+
timeframe: '1d'
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
+
```bash
|
| 107 |
+
python demo.py
|
| 108 |
+
python -m agentic_ai_system.main --mode backtest --start-date 2024-01-01 --end-date 2024-12-31
|
| 109 |
+
pytest tests/ -q
|
| 110 |
+
```
|
| 111 |
+
|
| 112 |
+
UI launchers and Docker are documented in `UI_SETUP.md` and `DOCKER_HUB_SETUP.md`. Paper-trade before live. Yahoo is not a SIP tape.
|
| 113 |
+
|
| 114 |
+
---
|
| 115 |
+
|
| 116 |
+
## 6. Configuration (additive Yahoo keys)
|
| 117 |
+
|
| 118 |
+
| Key | Meaning |
|
| 119 |
+
| --- | ------- |
|
| 120 |
+
| `data_source.type` | `csv` \| `synthetic` \| `alpaca` \| `yahoo` |
|
| 121 |
+
| `yahoo.start_date` / `end_date` | Historical window; clamped per Yahoo interval limits |
|
| 122 |
+
| `yahoo.auto_adjust` | Passed to `yfinance` |
|
| 123 |
+
| `execution.broker_api` | `paper` \| `alpaca_paper` \| `alpaca_live` |
|
| 124 |
+
| `finrl.algorithm` | PPO, A2C, DDPG, TD3 |
|
| 125 |
+
|
| 126 |
+
---
|
| 127 |
+
|
| 128 |
+
**License:** Apache License 2.0
|
| 129 |
+
**Organization:** [Parallel LLC](https://github.com/ParallelLLC)
|
| 130 |
+
**Repository:** <https://github.com/ParallelLLC/algorithmic_trading>
|
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Dependencies for the Hugging Face Space (app.py + the algotrader package).
|
| 2 |
+
#
|
| 3 |
+
# Deliberately minimal: a Space that takes ten minutes to build is a Space
|
| 4 |
+
# nobody waits for. The v1 agentic system's heavier stack (torch,
|
| 5 |
+
# stable-baselines3, dash, alpaca-py) lives in requirements.txt and is not
|
| 6 |
+
# needed to run the Lab.
|
| 7 |
+
gradio>=4.44,<7
|
| 8 |
+
numpy>=1.24
|
| 9 |
+
pandas>=2.0
|
| 10 |
+
scipy>=1.10
|
| 11 |
+
plotly>=5.18
|
| 12 |
+
yfinance>=0.2.40
|
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
#
|
| 3 |
+
# Assemble and push the Hugging Face Space.
|
| 4 |
+
#
|
| 5 |
+
# The Space gets only what it needs to run: app.py, the algotrader package, the
|
| 6 |
+
# Space card, and the minimal requirements file. The v1 agentic system and its
|
| 7 |
+
# heavy ML stack stay in this repo, so the Space builds in under a minute.
|
| 8 |
+
#
|
| 9 |
+
# Usage:
|
| 10 |
+
# HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <hf-username>/<space-name>
|
| 11 |
+
#
|
| 12 |
+
# Create the Space first at https://huggingface.co/new-space (SDK: Gradio).
|
| 13 |
+
|
| 14 |
+
set -euo pipefail
|
| 15 |
+
|
| 16 |
+
SPACE_ID="${1:-}"
|
| 17 |
+
if [[ -z "$SPACE_ID" ]]; then
|
| 18 |
+
echo "usage: HF_TOKEN=hf_xxx $0 <hf-username>/<space-name>" >&2
|
| 19 |
+
exit 2
|
| 20 |
+
fi
|
| 21 |
+
if [[ -z "${HF_TOKEN:-}" ]]; then
|
| 22 |
+
echo "error: HF_TOKEN is not set. Create a write token at https://huggingface.co/settings/tokens" >&2
|
| 23 |
+
exit 2
|
| 24 |
+
fi
|
| 25 |
+
|
| 26 |
+
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
| 27 |
+
STAGING="$(mktemp -d)"
|
| 28 |
+
trap 'rm -rf "$STAGING"' EXIT
|
| 29 |
+
|
| 30 |
+
echo "==> Staging Space contents in $STAGING"
|
| 31 |
+
git clone --quiet "https://user:${HF_TOKEN}@huggingface.co/spaces/${SPACE_ID}" "$STAGING/space"
|
| 32 |
+
cd "$STAGING/space"
|
| 33 |
+
|
| 34 |
+
# Replace tracked content wholesale so deletions propagate, but keep .git.
|
| 35 |
+
find . -mindepth 1 -maxdepth 1 ! -name .git -exec rm -rf {} +
|
| 36 |
+
|
| 37 |
+
cp -r "$REPO_ROOT/algotrader" ./algotrader
|
| 38 |
+
cp "$REPO_ROOT/app.py" ./app.py
|
| 39 |
+
cp "$REPO_ROOT/LICENSE" ./LICENSE
|
| 40 |
+
cp "$REPO_ROOT/SPACE_README.md" ./README.md
|
| 41 |
+
cp "$REPO_ROOT/requirements-space.txt" ./requirements.txt
|
| 42 |
+
find ./algotrader -name '__pycache__' -type d -exec rm -rf {} + 2>/dev/null || true
|
| 43 |
+
|
| 44 |
+
# A small test suite ships too: reviewers who check whether the statistics are
|
| 45 |
+
# real are exactly the audience worth convincing.
|
| 46 |
+
mkdir -p tests
|
| 47 |
+
cp "$REPO_ROOT/tests/test_v2_engine.py" \
|
| 48 |
+
"$REPO_ROOT/tests/test_v2_validation.py" \
|
| 49 |
+
"$REPO_ROOT/tests/test_v2_strategies.py" tests/
|
| 50 |
+
|
| 51 |
+
echo "==> Files staged:"
|
| 52 |
+
find . -path ./.git -prune -o -type f -print | sed 's|^\./| |'
|
| 53 |
+
|
| 54 |
+
git add -A
|
| 55 |
+
if git diff --cached --quiet; then
|
| 56 |
+
echo "==> Space is already up to date; nothing to push."
|
| 57 |
+
exit 0
|
| 58 |
+
fi
|
| 59 |
+
|
| 60 |
+
git -c user.email="deploy@localhost" -c user.name="space-deploy" \
|
| 61 |
+
commit --quiet -m "Deploy algotrader $(python3 -c 'import sys; sys.path.insert(0, "'"$REPO_ROOT"'"); import algotrader; print(algotrader.__version__)')"
|
| 62 |
+
git push --quiet origin HEAD:main
|
| 63 |
+
|
| 64 |
+
echo "==> Pushed. Live at https://huggingface.co/spaces/${SPACE_ID}"
|
|
@@ -0,0 +1,89 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""CLI surface — argument handling and that each command actually completes."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
import json
|
| 6 |
+
|
| 7 |
+
import pytest
|
| 8 |
+
|
| 9 |
+
from algotrader.cli import _parse_params, main
|
| 10 |
+
|
| 11 |
+
|
| 12 |
+
class TestArgumentParsing:
|
| 13 |
+
def test_params_are_parsed_into_floats(self):
|
| 14 |
+
assert _parse_params(["fast=10", "slow=50.5"]) == {"fast": 10.0, "slow": 50.5}
|
| 15 |
+
|
| 16 |
+
def test_params_without_an_equals_sign_are_rejected(self):
|
| 17 |
+
with pytest.raises(SystemExit, match="name=value"):
|
| 18 |
+
_parse_params(["fast10"])
|
| 19 |
+
|
| 20 |
+
def test_no_params_is_an_empty_dict(self):
|
| 21 |
+
assert _parse_params(None) == {}
|
| 22 |
+
|
| 23 |
+
def test_an_unknown_strategy_is_rejected_at_parse_time(self):
|
| 24 |
+
with pytest.raises(SystemExit):
|
| 25 |
+
main(["lab", "--strategy", "nope"])
|
| 26 |
+
|
| 27 |
+
def test_a_missing_subcommand_is_rejected(self):
|
| 28 |
+
with pytest.raises(SystemExit):
|
| 29 |
+
main([])
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
class TestCommands:
|
| 33 |
+
"""All runs force the simulator so the suite never touches the network."""
|
| 34 |
+
|
| 35 |
+
def test_lab_prints_a_report(self, capsys):
|
| 36 |
+
code = main([
|
| 37 |
+
"lab", "--symbol", "SPY", "--source", "synthetic", "--strategy", "sma_cross",
|
| 38 |
+
"--start", "2018-01-01", "--end", "2023-01-01",
|
| 39 |
+
"--permutations", "10", "--folds", "2", "--quiet",
|
| 40 |
+
])
|
| 41 |
+
out = capsys.readouterr().out
|
| 42 |
+
assert code == 0
|
| 43 |
+
assert "REALITY SCORE" in out
|
| 44 |
+
assert "Deflated Sharpe" in out
|
| 45 |
+
|
| 46 |
+
def test_lab_json_output_is_machine_readable(self, capsys):
|
| 47 |
+
code = main([
|
| 48 |
+
"lab", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
|
| 49 |
+
"--permutations", "10", "--folds", "2", "--quiet", "--json",
|
| 50 |
+
])
|
| 51 |
+
payload = json.loads(capsys.readouterr().out)
|
| 52 |
+
assert code == 0
|
| 53 |
+
assert 0.0 <= payload["verdict"]["score"] <= 100.0
|
| 54 |
+
assert payload["verdict"]["grade"] in {"A", "B", "C", "D", "F"}
|
| 55 |
+
assert payload["source"] == "synthetic"
|
| 56 |
+
|
| 57 |
+
def test_lab_honours_param_overrides(self, capsys):
|
| 58 |
+
main([
|
| 59 |
+
"lab", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
|
| 60 |
+
"--strategy", "sma_cross", "--param", "fast=5", "--param", "slow=60",
|
| 61 |
+
"--permutations", "0", "--folds", "2", "--quiet", "--json",
|
| 62 |
+
])
|
| 63 |
+
payload = json.loads(capsys.readouterr().out)
|
| 64 |
+
assert payload["params"] == {"fast": 5, "slow": 60}
|
| 65 |
+
assert payload["p_value"] is None # permutations disabled
|
| 66 |
+
|
| 67 |
+
def test_arena_ranks_the_zoo(self, capsys):
|
| 68 |
+
code = main([
|
| 69 |
+
"arena", "--source", "synthetic", "--start", "2018-01-01", "--end", "2023-01-01",
|
| 70 |
+
"--permutations", "0", "--quiet",
|
| 71 |
+
])
|
| 72 |
+
out = capsys.readouterr().out
|
| 73 |
+
assert code == 0
|
| 74 |
+
assert "Buy & Hold" in out and "Coin Flip" in out
|
| 75 |
+
assert "Ranked by evidence" in out
|
| 76 |
+
|
| 77 |
+
def test_strategies_lists_the_registry(self, capsys):
|
| 78 |
+
assert main(["strategies"]) == 0
|
| 79 |
+
out = capsys.readouterr().out
|
| 80 |
+
assert "sma_cross" in out and "donchian_breakout" in out
|
| 81 |
+
|
| 82 |
+
def test_errors_are_reported_not_raised(self, capsys):
|
| 83 |
+
code = main([
|
| 84 |
+
"lab", "--source", "synthetic",
|
| 85 |
+
"--start", "2022-01-01", "--end", "2022-02-01", # far too short
|
| 86 |
+
"--permutations", "0", "--quiet",
|
| 87 |
+
])
|
| 88 |
+
assert code == 1
|
| 89 |
+
assert "error:" in capsys.readouterr().err
|
|
@@ -0,0 +1,158 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Engine correctness: alignment, costs, and the no-look-ahead guarantee."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
import numpy as np
|
| 6 |
+
import pandas as pd
|
| 7 |
+
import pytest
|
| 8 |
+
|
| 9 |
+
from algotrader.engine import bars_to_returns, run_backtest
|
| 10 |
+
from algotrader.metrics import compute_metrics, infer_periods_per_year, max_drawdown, sharpe_ratio
|
| 11 |
+
from algotrader.types import CostModel
|
| 12 |
+
|
| 13 |
+
|
| 14 |
+
def make_bars(n: int = 400, seed: int = 0) -> pd.DataFrame:
|
| 15 |
+
rng = np.random.default_rng(seed)
|
| 16 |
+
close = 100 * np.exp(np.cumsum(rng.normal(0.0003, 0.01, n)))
|
| 17 |
+
index = pd.date_range("2020-01-01", periods=n, freq="B")
|
| 18 |
+
return pd.DataFrame(
|
| 19 |
+
{
|
| 20 |
+
"open": close,
|
| 21 |
+
"high": close * 1.005,
|
| 22 |
+
"low": close * 0.995,
|
| 23 |
+
"close": close,
|
| 24 |
+
"volume": 1e6,
|
| 25 |
+
},
|
| 26 |
+
index=index,
|
| 27 |
+
)
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
class TestNoLookAhead:
|
| 31 |
+
"""The one property the whole project rests on."""
|
| 32 |
+
|
| 33 |
+
def test_signal_earns_the_following_bar_not_its_own(self):
|
| 34 |
+
df = make_bars()
|
| 35 |
+
asset_ret = bars_to_returns(df)
|
| 36 |
+
|
| 37 |
+
# A target that knows the *next* bar's direction must be perfect...
|
| 38 |
+
clairvoyant = np.sign(asset_ret.shift(-1)).fillna(0.0)
|
| 39 |
+
good = run_backtest(df, clairvoyant, costs=CostModel(0, 0, 0))
|
| 40 |
+
assert (good.returns.iloc[1:-1] >= -1e-12).all(), "perfect foresight should never lose"
|
| 41 |
+
|
| 42 |
+
# ...and a target that only knows the *current* bar must not be.
|
| 43 |
+
hindsight = np.sign(asset_ret).fillna(0.0)
|
| 44 |
+
meh = run_backtest(df, hindsight, costs=CostModel(0, 0, 0))
|
| 45 |
+
assert (meh.returns < 0).any(), "same-bar signal must not be risk-free"
|
| 46 |
+
assert good.sharpe > meh.sharpe
|
| 47 |
+
|
| 48 |
+
def test_position_is_target_shifted_by_lag(self):
|
| 49 |
+
df = make_bars(200)
|
| 50 |
+
target = pd.Series(np.linspace(-1, 1, len(df)), index=df.index)
|
| 51 |
+
for lag in (1, 2, 5):
|
| 52 |
+
result = run_backtest(df, target, lag=lag)
|
| 53 |
+
expected = target.shift(lag).fillna(0.0)
|
| 54 |
+
pd.testing.assert_series_equal(result.position, expected, check_names=False)
|
| 55 |
+
|
| 56 |
+
def test_lag_zero_is_rejected(self):
|
| 57 |
+
df = make_bars(120)
|
| 58 |
+
with pytest.raises(ValueError, match="lag"):
|
| 59 |
+
run_backtest(df, pd.Series(1.0, index=df.index), lag=0)
|
| 60 |
+
|
| 61 |
+
def test_future_bars_cannot_change_past_equity(self):
|
| 62 |
+
"""Truncating the data must not alter the equity curve before the cut."""
|
| 63 |
+
df = make_bars(400)
|
| 64 |
+
target = pd.Series(np.tile([1.0, -1.0], len(df) // 2), index=df.index)
|
| 65 |
+
|
| 66 |
+
full = run_backtest(df, target)
|
| 67 |
+
cut = run_backtest(df.iloc[:250], target.iloc[:250])
|
| 68 |
+
np.testing.assert_allclose(
|
| 69 |
+
full.equity.iloc[:250].to_numpy(), cut.equity.to_numpy(), rtol=1e-12
|
| 70 |
+
)
|
| 71 |
+
|
| 72 |
+
|
| 73 |
+
class TestCosts:
|
| 74 |
+
def test_buy_and_hold_pays_once(self):
|
| 75 |
+
df = make_bars(300)
|
| 76 |
+
target = pd.Series(1.0, index=df.index)
|
| 77 |
+
result = run_backtest(df, target, costs=CostModel(commission_bps=5, slippage_bps=5, short_borrow_bps=0))
|
| 78 |
+
# One entry at 10bps one-way, and no further turnover.
|
| 79 |
+
assert result.costs.sum() == pytest.approx(10 / 1e4, rel=1e-9)
|
| 80 |
+
assert int(result.metrics["n_trades"]) == 1
|
| 81 |
+
|
| 82 |
+
def test_flipping_every_bar_costs_more_than_holding(self):
|
| 83 |
+
df = make_bars(300)
|
| 84 |
+
costs = CostModel(commission_bps=5, slippage_bps=5, short_borrow_bps=0)
|
| 85 |
+
hold = run_backtest(df, pd.Series(1.0, index=df.index), costs=costs)
|
| 86 |
+
flip = run_backtest(df, pd.Series(np.tile([1.0, -1.0], 150), index=df.index), costs=costs)
|
| 87 |
+
assert flip.costs.sum() > 100 * hold.costs.sum()
|
| 88 |
+
|
| 89 |
+
def test_zero_costs_means_gross_equals_net(self):
|
| 90 |
+
df = make_bars(200)
|
| 91 |
+
target = pd.Series(np.tile([1.0, 0.0], 100), index=df.index)
|
| 92 |
+
result = run_backtest(df, target, costs=CostModel(0, 0, 0))
|
| 93 |
+
pd.testing.assert_series_equal(result.returns, result.gross_returns, check_names=False)
|
| 94 |
+
|
| 95 |
+
def test_short_borrow_is_charged_only_on_shorts(self):
|
| 96 |
+
df = make_bars(260)
|
| 97 |
+
costs = CostModel(commission_bps=0, slippage_bps=0, short_borrow_bps=365)
|
| 98 |
+
long_only = run_backtest(df, pd.Series(1.0, index=df.index), costs=costs)
|
| 99 |
+
short_only = run_backtest(df, pd.Series(-1.0, index=df.index), costs=costs)
|
| 100 |
+
assert long_only.costs.sum() == pytest.approx(0.0, abs=1e-12)
|
| 101 |
+
assert short_only.costs.sum() > 0
|
| 102 |
+
|
| 103 |
+
def test_higher_costs_never_improve_returns(self):
|
| 104 |
+
df = make_bars(300)
|
| 105 |
+
target = pd.Series(np.tile([1.0, -1.0], 150), index=df.index)
|
| 106 |
+
cheap = run_backtest(df, target, costs=CostModel(1, 1, 0))
|
| 107 |
+
dear = run_backtest(df, target, costs=CostModel(20, 20, 0))
|
| 108 |
+
assert dear.equity.iloc[-1] < cheap.equity.iloc[-1]
|
| 109 |
+
|
| 110 |
+
|
| 111 |
+
class TestConstraints:
|
| 112 |
+
def test_shorts_are_clipped_when_disallowed(self):
|
| 113 |
+
df = make_bars(150)
|
| 114 |
+
target = pd.Series(-1.0, index=df.index)
|
| 115 |
+
result = run_backtest(df, target, allow_short=False)
|
| 116 |
+
assert (result.target >= 0).all()
|
| 117 |
+
assert (result.position >= 0).all()
|
| 118 |
+
|
| 119 |
+
def test_leverage_is_clipped(self):
|
| 120 |
+
df = make_bars(150)
|
| 121 |
+
result = run_backtest(df, pd.Series(5.0, index=df.index), max_leverage=1.5)
|
| 122 |
+
assert result.target.max() == pytest.approx(1.5)
|
| 123 |
+
|
| 124 |
+
def test_empty_frame_is_rejected(self):
|
| 125 |
+
with pytest.raises(ValueError):
|
| 126 |
+
run_backtest(pd.DataFrame(columns=["open", "high", "low", "close", "volume"]), pd.Series(dtype=float))
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
class TestMetrics:
|
| 130 |
+
def test_sharpe_of_constant_returns_is_zero_not_infinite(self):
|
| 131 |
+
flat = pd.Series([0.001] * 100)
|
| 132 |
+
assert sharpe_ratio(flat, 252) == 0.0
|
| 133 |
+
|
| 134 |
+
def test_sharpe_scales_with_annualisation(self):
|
| 135 |
+
rng = np.random.default_rng(1)
|
| 136 |
+
returns = pd.Series(rng.normal(0.001, 0.01, 5000))
|
| 137 |
+
assert sharpe_ratio(returns, 252) == pytest.approx(sharpe_ratio(returns, 1) * np.sqrt(252))
|
| 138 |
+
|
| 139 |
+
def test_max_drawdown_matches_a_hand_worked_example(self):
|
| 140 |
+
equity = pd.Series([100.0, 120.0, 60.0, 90.0])
|
| 141 |
+
assert max_drawdown(equity) == pytest.approx(-0.5)
|
| 142 |
+
|
| 143 |
+
def test_buy_and_hold_metrics_match_the_price_series(self):
|
| 144 |
+
df = make_bars(500)
|
| 145 |
+
result = run_backtest(df, pd.Series(1.0, index=df.index), costs=CostModel(0, 0, 0))
|
| 146 |
+
expected = df["close"].iloc[-1] / df["close"].iloc[0] - 1.0
|
| 147 |
+
assert result.metrics["total_return"] == pytest.approx(expected, rel=1e-9)
|
| 148 |
+
|
| 149 |
+
def test_periodicity_inference(self):
|
| 150 |
+
daily = pd.date_range("2020-01-01", periods=300, freq="B")
|
| 151 |
+
assert 200 <= infer_periods_per_year(daily) <= 300
|
| 152 |
+
hourly = pd.date_range("2020-01-01", periods=300, freq="h")
|
| 153 |
+
assert infer_periods_per_year(hourly) > 1000
|
| 154 |
+
|
| 155 |
+
def test_metrics_survive_a_degenerate_series(self):
|
| 156 |
+
empty = pd.Series(dtype=float)
|
| 157 |
+
out = compute_metrics(empty, pd.Series(dtype=float))
|
| 158 |
+
assert "periods_per_year" in out
|
|
@@ -0,0 +1,229 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Strategy zoo, data loading, and the end-to-end Lab pipeline."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
import numpy as np
|
| 6 |
+
import pandas as pd
|
| 7 |
+
import pytest
|
| 8 |
+
|
| 9 |
+
from algotrader import LabConfig, run_lab
|
| 10 |
+
from algotrader.data import _normalise, load_ohlcv, simulate_ohlcv
|
| 11 |
+
from algotrader.indicators import atr, bollinger, donchian, ema, macd, rsi, sma
|
| 12 |
+
from algotrader.lab import run_arena
|
| 13 |
+
from algotrader.strategies import REGISTRY, get_strategy, list_strategies
|
| 14 |
+
from algotrader.types import MarketData
|
| 15 |
+
from algotrader.verdict import reality_score
|
| 16 |
+
|
| 17 |
+
ALL_KEYS = sorted(REGISTRY)
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
@pytest.fixture(scope="module")
|
| 21 |
+
def bars() -> pd.DataFrame:
|
| 22 |
+
return simulate_ohlcv("AAPL", "2016-01-01", "2023-01-01")
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
class TestIndicatorsAreCausal:
|
| 26 |
+
"""An indicator at bar t must not move when bars after t arrive."""
|
| 27 |
+
|
| 28 |
+
@pytest.mark.parametrize(
|
| 29 |
+
"fn",
|
| 30 |
+
[
|
| 31 |
+
lambda d: sma(d["close"], 20),
|
| 32 |
+
lambda d: ema(d["close"], 12),
|
| 33 |
+
lambda d: rsi(d["close"], 14),
|
| 34 |
+
lambda d: macd(d["close"])[0],
|
| 35 |
+
lambda d: bollinger(d["close"])[2],
|
| 36 |
+
lambda d: atr(d, 14),
|
| 37 |
+
lambda d: donchian(d, 20)[1],
|
| 38 |
+
],
|
| 39 |
+
)
|
| 40 |
+
def test_prefix_is_stable(self, bars, fn):
|
| 41 |
+
cut = 800
|
| 42 |
+
full = fn(bars).iloc[:cut]
|
| 43 |
+
partial = fn(bars.iloc[:cut])
|
| 44 |
+
pd.testing.assert_series_equal(full, partial, check_names=False, rtol=1e-9)
|
| 45 |
+
|
| 46 |
+
def test_rsi_stays_in_range(self, bars):
|
| 47 |
+
values = rsi(bars["close"], 14).dropna()
|
| 48 |
+
assert values.between(0, 100).all()
|
| 49 |
+
|
| 50 |
+
def test_donchian_excludes_the_current_bar(self, bars):
|
| 51 |
+
_, upper = donchian(bars, 20)
|
| 52 |
+
# A breakout must be possible: the channel cannot already contain today.
|
| 53 |
+
assert (bars["high"] > upper).any()
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
class TestStrategies:
|
| 57 |
+
@pytest.mark.parametrize("key", ALL_KEYS)
|
| 58 |
+
def test_output_is_well_formed(self, bars, key):
|
| 59 |
+
target = get_strategy(key).generate(bars)
|
| 60 |
+
assert target.index.equals(bars.index)
|
| 61 |
+
assert target.notna().all()
|
| 62 |
+
assert target.between(-1.0, 1.0).all()
|
| 63 |
+
|
| 64 |
+
@pytest.mark.parametrize("key", ALL_KEYS)
|
| 65 |
+
def test_signals_are_causal(self, bars, key):
|
| 66 |
+
cut = 900
|
| 67 |
+
full = get_strategy(key).generate(bars).iloc[:cut]
|
| 68 |
+
partial = get_strategy(key).generate(bars.iloc[:cut])
|
| 69 |
+
pd.testing.assert_series_equal(full, partial, check_names=False, rtol=1e-9)
|
| 70 |
+
|
| 71 |
+
@pytest.mark.parametrize("key", ALL_KEYS)
|
| 72 |
+
def test_grid_is_non_empty_and_valid(self, key):
|
| 73 |
+
strategy = get_strategy(key)
|
| 74 |
+
grid = strategy.grid(limit=40)
|
| 75 |
+
assert 1 <= len(grid) <= 40
|
| 76 |
+
for combo in grid:
|
| 77 |
+
assert set(combo) <= {p.name for p in strategy.params}
|
| 78 |
+
if "fast" in combo and "slow" in combo:
|
| 79 |
+
assert combo["fast"] < combo["slow"]
|
| 80 |
+
|
| 81 |
+
def test_unknown_strategy_names_the_alternatives(self):
|
| 82 |
+
with pytest.raises(KeyError, match="Available"):
|
| 83 |
+
get_strategy("does_not_exist")
|
| 84 |
+
|
| 85 |
+
def test_clean_ignores_unknown_params_and_casts_types(self):
|
| 86 |
+
strategy = get_strategy("sma_cross")
|
| 87 |
+
cleaned = strategy.clean({"fast": 15.7, "nonsense": 1})
|
| 88 |
+
assert cleaned == {"fast": 15, "slow": 100}
|
| 89 |
+
|
| 90 |
+
def test_buy_and_hold_is_always_fully_invested(self, bars):
|
| 91 |
+
assert (get_strategy("buy_and_hold").generate(bars) == 1.0).all()
|
| 92 |
+
|
| 93 |
+
def test_list_strategies_honours_exclusions(self):
|
| 94 |
+
keys = {s.key for s in list_strategies(exclude=["coin_flip"])}
|
| 95 |
+
assert "coin_flip" not in keys and "sma_cross" in keys
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
class TestData:
|
| 99 |
+
def test_simulation_is_deterministic_per_symbol(self):
|
| 100 |
+
a = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
|
| 101 |
+
b = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
|
| 102 |
+
pd.testing.assert_frame_equal(a, b)
|
| 103 |
+
|
| 104 |
+
def test_different_symbols_simulate_differently(self):
|
| 105 |
+
a = simulate_ohlcv("NVDA", "2018-01-01", "2022-01-01")
|
| 106 |
+
b = simulate_ohlcv("TSLA", "2018-01-01", "2022-01-01")
|
| 107 |
+
assert not np.allclose(a["close"].to_numpy(), b["close"].to_numpy())
|
| 108 |
+
|
| 109 |
+
def test_simulated_bars_are_internally_consistent(self):
|
| 110 |
+
df = simulate_ohlcv("SPY", "2015-01-01", "2023-01-01")
|
| 111 |
+
assert (df["high"] >= df[["open", "close"]].max(axis=1) - 1e-9).all()
|
| 112 |
+
assert (df["low"] <= df[["open", "close"]].min(axis=1) + 1e-9).all()
|
| 113 |
+
assert (df["close"] > 0).all()
|
| 114 |
+
assert df.index.is_monotonic_increasing
|
| 115 |
+
|
| 116 |
+
def test_simulation_has_fat_tails_and_vol_clustering(self):
|
| 117 |
+
"""Naive GBM flatters strategies; the simulator must be harder than that."""
|
| 118 |
+
returns = simulate_ohlcv("SPY", "2005-01-01", "2023-01-01")["close"].pct_change().dropna()
|
| 119 |
+
assert returns.kurtosis() > 1.0
|
| 120 |
+
assert returns.abs().autocorr(1) > 0.05
|
| 121 |
+
|
| 122 |
+
def test_offline_load_falls_back_and_says_so(self, monkeypatch):
|
| 123 |
+
monkeypatch.setattr("algotrader.data._download", lambda *a, **k: None)
|
| 124 |
+
monkeypatch.setattr("algotrader.data._read_cache", lambda *a, **k: None)
|
| 125 |
+
market = load_ohlcv("SPY", "2018-01-01", "2022-01-01")
|
| 126 |
+
assert market.source == "synthetic"
|
| 127 |
+
assert not market.is_real
|
| 128 |
+
assert "unavailable" in market.note
|
| 129 |
+
|
| 130 |
+
def test_normalise_handles_yahoo_style_frames(self):
|
| 131 |
+
index = pd.date_range("2020-01-01", periods=5, tz="UTC")
|
| 132 |
+
raw = pd.DataFrame(
|
| 133 |
+
{"Open": 1.0, "High": 2.0, "Low": 0.5, "Adj Close": 1.5, "Volume": 10},
|
| 134 |
+
index=index,
|
| 135 |
+
)
|
| 136 |
+
out = _normalise(raw)
|
| 137 |
+
assert list(out.columns) == ["open", "high", "low", "close", "volume"]
|
| 138 |
+
assert out.index.tz is None
|
| 139 |
+
|
| 140 |
+
def test_normalise_rejects_frames_with_no_price(self):
|
| 141 |
+
with pytest.raises(ValueError, match="missing required column"):
|
| 142 |
+
_normalise(pd.DataFrame({"volume": [1, 2, 3]}))
|
| 143 |
+
|
| 144 |
+
def test_too_short_a_range_is_rejected(self):
|
| 145 |
+
with pytest.raises(ValueError, match="too short"):
|
| 146 |
+
simulate_ohlcv("SPY", "2020-01-01", "2020-01-10")
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
class TestVerdict:
|
| 150 |
+
def test_strong_evidence_outranks_weak_evidence(self):
|
| 151 |
+
metrics = {"n_trades": 300, "sharpe": 1.2, "total_return": 0.8, "max_drawdown": -0.2}
|
| 152 |
+
benchmark = {"sharpe": 0.4}
|
| 153 |
+
strong = reality_score(metrics, benchmark, p_value=0.001, dsr=0.99, pbo=0.02,
|
| 154 |
+
wf_efficiency=0.9, wf_win_rate=1.0, cost_stress_ratio=0.9)
|
| 155 |
+
weak = reality_score(metrics, benchmark, p_value=0.45, dsr=0.10, pbo=0.55,
|
| 156 |
+
wf_efficiency=-0.2, wf_win_rate=0.2, cost_stress_ratio=0.1)
|
| 157 |
+
assert strong["score"] > 85 > weak["score"]
|
| 158 |
+
assert strong["grade"] == "A" and weak["grade"] == "F"
|
| 159 |
+
|
| 160 |
+
def test_score_is_always_inside_the_scale(self):
|
| 161 |
+
for p in (0.0, 0.5, 1.0):
|
| 162 |
+
for dsr in (0.0, 1.0):
|
| 163 |
+
out = reality_score(
|
| 164 |
+
{"n_trades": 100, "sharpe": 0.5, "total_return": 0.2, "max_drawdown": -0.1},
|
| 165 |
+
{"sharpe": 0.1}, p_value=p, dsr=dsr, pbo=0.2,
|
| 166 |
+
wf_efficiency=0.5, cost_stress_ratio=0.5,
|
| 167 |
+
)
|
| 168 |
+
assert 0.0 <= out["score"] <= 100.0
|
| 169 |
+
|
| 170 |
+
def test_losing_money_caps_the_score(self):
|
| 171 |
+
out = reality_score(
|
| 172 |
+
{"n_trades": 200, "sharpe": 0.3, "total_return": -0.4, "max_drawdown": -0.6},
|
| 173 |
+
{"sharpe": 0.5}, p_value=0.001, dsr=0.99, pbo=0.01,
|
| 174 |
+
wf_efficiency=1.0, cost_stress_ratio=1.0,
|
| 175 |
+
)
|
| 176 |
+
assert out["score"] <= 50
|
| 177 |
+
assert any("lost money" in f for f in out["flags"])
|
| 178 |
+
|
| 179 |
+
def test_too_few_trades_is_flagged_and_capped(self):
|
| 180 |
+
out = reality_score(
|
| 181 |
+
{"n_trades": 3, "sharpe": 2.5, "total_return": 1.0, "max_drawdown": -0.1},
|
| 182 |
+
{"sharpe": 0.3}, p_value=0.001, dsr=0.99, pbo=0.01,
|
| 183 |
+
wf_efficiency=1.0, cost_stress_ratio=1.0,
|
| 184 |
+
)
|
| 185 |
+
assert out["score"] <= 55
|
| 186 |
+
assert any("coin flips" in f for f in out["flags"])
|
| 187 |
+
|
| 188 |
+
def test_closet_indexing_is_called_out(self):
|
| 189 |
+
out = reality_score(
|
| 190 |
+
{"n_trades": 50, "sharpe": 0.6, "total_return": 0.5, "max_drawdown": -0.2},
|
| 191 |
+
{"sharpe": 0.6}, benchmark_correlation=0.99,
|
| 192 |
+
)
|
| 193 |
+
assert any("repackaged long position" in f for f in out["flags"])
|
| 194 |
+
|
| 195 |
+
|
| 196 |
+
class TestLabEndToEnd:
|
| 197 |
+
def test_full_pipeline_produces_a_complete_report(self, monkeypatch):
|
| 198 |
+
monkeypatch.setattr("algotrader.lab.load_ohlcv", lambda *a, **k: MarketData(
|
| 199 |
+
"SIM", simulate_ohlcv("SPY", "2016-01-01", "2023-01-01"), "synthetic", "1d", "test"
|
| 200 |
+
))
|
| 201 |
+
report = run_lab(LabConfig(strategy="sma_cross", n_permutations=25, wf_folds=3, grid_limit=12))
|
| 202 |
+
|
| 203 |
+
assert 0.0 <= report.verdict["score"] <= 100.0
|
| 204 |
+
assert report.verdict["grade"] in {"A", "B", "C", "D", "F"}
|
| 205 |
+
assert 0 < report.permutation.p_value <= 1
|
| 206 |
+
assert 0.0 <= report.dsr["dsr"] <= 1.0
|
| 207 |
+
assert report.trials["n"] > 1
|
| 208 |
+
assert len(report.backtest.equity) == len(report.market.df)
|
| 209 |
+
assert report.cost_stress["sharpe_3x"] <= report.cost_stress["sharpe_1x"] + 1e-9
|
| 210 |
+
|
| 211 |
+
def test_too_little_history_gives_a_readable_error(self, monkeypatch):
|
| 212 |
+
short = simulate_ohlcv("SPY", "2020-01-01", "2020-06-01")
|
| 213 |
+
monkeypatch.setattr(
|
| 214 |
+
"algotrader.lab.load_ohlcv",
|
| 215 |
+
lambda *a, **k: MarketData("SIM", short, "synthetic", "1d", "test"),
|
| 216 |
+
)
|
| 217 |
+
with pytest.raises(ValueError, match="Widen the date range"):
|
| 218 |
+
run_lab(LabConfig(n_permutations=0, wf_folds=2))
|
| 219 |
+
|
| 220 |
+
def test_arena_ranks_every_strategy_and_keeps_the_controls(self, monkeypatch):
|
| 221 |
+
monkeypatch.setattr("algotrader.lab.load_ohlcv", lambda *a, **k: MarketData(
|
| 222 |
+
"SIM", simulate_ohlcv("SPY", "2017-01-01", "2022-01-01"), "synthetic", "1d", "test"
|
| 223 |
+
))
|
| 224 |
+
table, market, curves = run_arena(LabConfig(), n_permutations=0)
|
| 225 |
+
|
| 226 |
+
assert len(table) == len(REGISTRY)
|
| 227 |
+
assert {"buy_and_hold", "coin_flip"} <= set(table["key"])
|
| 228 |
+
assert table["Evidence"].is_monotonic_decreasing
|
| 229 |
+
assert set(curves) == set(REGISTRY)
|
|
@@ -0,0 +1,179 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Statistical machinery.
|
| 2 |
+
|
| 3 |
+
These tests matter more than the engine's, because a validation suite that
|
| 4 |
+
always says "no edge" is as useless as one that always says "great edge". Each
|
| 5 |
+
class below checks both directions: it must reject noise *and* detect signal.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
import numpy as np
|
| 11 |
+
import pandas as pd
|
| 12 |
+
import pytest
|
| 13 |
+
|
| 14 |
+
from algotrader.data import simulate_ohlcv
|
| 15 |
+
from algotrader.strategies import get_strategy
|
| 16 |
+
from algotrader.validation.deflated_sharpe import (
|
| 17 |
+
deflated_sharpe_ratio,
|
| 18 |
+
expected_max_sharpe,
|
| 19 |
+
min_track_record_length,
|
| 20 |
+
probabilistic_sharpe_ratio,
|
| 21 |
+
)
|
| 22 |
+
from algotrader.validation.pbo import probability_of_backtest_overfitting
|
| 23 |
+
from algotrader.validation.permutation import permutation_test, permute_bars
|
| 24 |
+
from algotrader.validation.walkforward import walk_forward
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
def trending_market(n: int = 2200, phi: float = 0.35, seed: int = 5) -> pd.DataFrame:
|
| 28 |
+
"""A market with genuine, exploitable serial correlation."""
|
| 29 |
+
rng = np.random.default_rng(seed)
|
| 30 |
+
returns = np.zeros(n)
|
| 31 |
+
noise = rng.normal(0, 0.01, n)
|
| 32 |
+
for i in range(1, n):
|
| 33 |
+
returns[i] = phi * returns[i - 1] + noise[i]
|
| 34 |
+
close = 100 * np.exp(np.cumsum(returns))
|
| 35 |
+
index = pd.date_range("2012-01-01", periods=n, freq="B")
|
| 36 |
+
return pd.DataFrame(
|
| 37 |
+
{"open": close, "high": close * 1.004, "low": close * 0.996, "close": close, "volume": 1e6},
|
| 38 |
+
index=index,
|
| 39 |
+
)
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
class TestPermutationMechanics:
|
| 43 |
+
def test_shuffling_preserves_the_distribution_of_moves(self):
|
| 44 |
+
df = simulate_ohlcv("SPY", "2018-01-01", "2023-01-01")
|
| 45 |
+
shuffled = permute_bars(df, np.random.default_rng(0))
|
| 46 |
+
|
| 47 |
+
assert len(shuffled) == len(df)
|
| 48 |
+
assert shuffled.index.equals(df.index)
|
| 49 |
+
original = np.sort(np.log(df["close"] / df["open"]).to_numpy()[1:])
|
| 50 |
+
permuted = np.sort(np.log(shuffled["close"] / shuffled["open"]).to_numpy()[1:])
|
| 51 |
+
np.testing.assert_allclose(original, permuted, rtol=1e-9)
|
| 52 |
+
|
| 53 |
+
def test_shuffling_keeps_bars_internally_valid(self):
|
| 54 |
+
df = simulate_ohlcv("AAPL", "2019-01-01", "2023-01-01")
|
| 55 |
+
shuffled = permute_bars(df, np.random.default_rng(3))
|
| 56 |
+
assert (shuffled["high"] >= shuffled["low"]).all()
|
| 57 |
+
assert (shuffled["high"] >= shuffled["close"]).all()
|
| 58 |
+
assert (shuffled["low"] <= shuffled["close"]).all()
|
| 59 |
+
assert (shuffled["close"] > 0).all()
|
| 60 |
+
|
| 61 |
+
def test_shuffling_destroys_serial_correlation(self):
|
| 62 |
+
df = trending_market()
|
| 63 |
+
real = df["close"].pct_change().dropna().autocorr(1)
|
| 64 |
+
shuffled = permute_bars(df, np.random.default_rng(1))["close"].pct_change().dropna().autocorr(1)
|
| 65 |
+
assert real > 0.2
|
| 66 |
+
assert abs(shuffled) < 0.1
|
| 67 |
+
|
| 68 |
+
def test_block_mode_retains_some_structure(self):
|
| 69 |
+
df = trending_market()
|
| 70 |
+
blocked = permute_bars(df, np.random.default_rng(2), method="block", block=40)
|
| 71 |
+
assert blocked["close"].pct_change().dropna().autocorr(1) > 0.1
|
| 72 |
+
|
| 73 |
+
def test_p_value_can_never_be_zero(self):
|
| 74 |
+
"""+1 correction: the observed run is itself a draw from the null."""
|
| 75 |
+
df = trending_market()
|
| 76 |
+
strategy = get_strategy("momentum")
|
| 77 |
+
result = permutation_test(
|
| 78 |
+
df, lambda f: strategy.generate(f, {"lookback": 5}), n_permutations=30, seed=0
|
| 79 |
+
)
|
| 80 |
+
assert result.p_value >= 1 / 31
|
| 81 |
+
assert 0 < result.p_value <= 1
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
class TestPermutationPower:
|
| 85 |
+
def test_real_edge_is_detected(self):
|
| 86 |
+
df = trending_market()
|
| 87 |
+
strategy = get_strategy("momentum")
|
| 88 |
+
result = permutation_test(
|
| 89 |
+
df, lambda f: strategy.generate(f, {"lookback": 5}), n_permutations=200, seed=1
|
| 90 |
+
)
|
| 91 |
+
assert result.observed > result.null.mean()
|
| 92 |
+
assert result.p_value < 0.05
|
| 93 |
+
|
| 94 |
+
def test_random_strategy_on_a_structureless_market_is_not_significant(self):
|
| 95 |
+
df = simulate_ohlcv("SIM", "2010-01-01", "2023-01-01")
|
| 96 |
+
strategy = get_strategy("coin_flip")
|
| 97 |
+
result = permutation_test(
|
| 98 |
+
df, lambda f: strategy.generate(f, {"hold": 5, "seed": 7}), n_permutations=200, seed=2
|
| 99 |
+
)
|
| 100 |
+
assert result.p_value > 0.05
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
class TestDeflatedSharpe:
|
| 104 |
+
def test_selection_bar_rises_with_the_number_of_trials(self):
|
| 105 |
+
low = expected_max_sharpe(10, 0.01)
|
| 106 |
+
high = expected_max_sharpe(1000, 0.01)
|
| 107 |
+
assert 0 < low < high
|
| 108 |
+
|
| 109 |
+
def test_a_single_trial_has_no_selection_bar(self):
|
| 110 |
+
assert expected_max_sharpe(1, 0.01) == 0.0
|
| 111 |
+
|
| 112 |
+
def test_more_trials_lowers_the_deflated_sharpe(self):
|
| 113 |
+
rng = np.random.default_rng(7)
|
| 114 |
+
returns = rng.normal(0.0006, 0.01, 2000)
|
| 115 |
+
few = deflated_sharpe_ratio(returns, 1.0, 252, n_trials=2, variance_of_trials=0.01)
|
| 116 |
+
many = deflated_sharpe_ratio(returns, 1.0, 252, n_trials=500, variance_of_trials=0.01)
|
| 117 |
+
assert many["dsr"] < few["dsr"]
|
| 118 |
+
assert many["psr"] == pytest.approx(few["psr"]) # PSR ignores selection
|
| 119 |
+
|
| 120 |
+
def test_psr_rises_with_track_record_length(self):
|
| 121 |
+
short = probabilistic_sharpe_ratio(0.05, 100)
|
| 122 |
+
long = probabilistic_sharpe_ratio(0.05, 5000)
|
| 123 |
+
assert 0.5 < short < long < 1.0
|
| 124 |
+
|
| 125 |
+
def test_negative_skew_and_fat_tails_are_penalised(self):
|
| 126 |
+
clean = probabilistic_sharpe_ratio(0.06, 1000, skew=0.0, kurtosis=3.0)
|
| 127 |
+
nasty = probabilistic_sharpe_ratio(0.06, 1000, skew=-1.5, kurtosis=12.0)
|
| 128 |
+
assert nasty < clean
|
| 129 |
+
|
| 130 |
+
def test_track_record_requirement_is_infinite_below_the_bar(self):
|
| 131 |
+
assert min_track_record_length(0.01, 500, benchmark=0.05) == float("inf")
|
| 132 |
+
assert np.isfinite(min_track_record_length(0.10, 500, benchmark=0.02))
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
class TestPBO:
|
| 136 |
+
def test_pure_noise_scores_near_one_half(self):
|
| 137 |
+
rng = np.random.default_rng(11)
|
| 138 |
+
matrix = rng.normal(0, 0.01, size=(1200, 30)) # 30 skill-free variants
|
| 139 |
+
result = probability_of_backtest_overfitting(matrix, n_splits=8)
|
| 140 |
+
assert 0.3 < result["pbo"] < 0.7
|
| 141 |
+
|
| 142 |
+
def test_a_genuinely_better_variant_is_not_flagged(self):
|
| 143 |
+
rng = np.random.default_rng(12)
|
| 144 |
+
matrix = rng.normal(0, 0.01, size=(1200, 20))
|
| 145 |
+
matrix[:, 3] += 0.004 # column 3 has a persistent, real edge
|
| 146 |
+
result = probability_of_backtest_overfitting(matrix, n_splits=8)
|
| 147 |
+
assert result["pbo"] < 0.15
|
| 148 |
+
assert result["most_selected_index"] == 3
|
| 149 |
+
assert result["selection_stability"] > 0.9
|
| 150 |
+
|
| 151 |
+
def test_too_few_variants_returns_nan_not_a_crash(self):
|
| 152 |
+
rng = np.random.default_rng(13)
|
| 153 |
+
result = probability_of_backtest_overfitting(rng.normal(0, 0.01, size=(500, 1)))
|
| 154 |
+
assert np.isnan(result["pbo"])
|
| 155 |
+
assert result["note"]
|
| 156 |
+
|
| 157 |
+
def test_odd_split_counts_are_made_even(self):
|
| 158 |
+
rng = np.random.default_rng(14)
|
| 159 |
+
result = probability_of_backtest_overfitting(rng.normal(0, 0.01, (800, 10)), n_splits=7)
|
| 160 |
+
assert result["n_combinations"] > 0
|
| 161 |
+
|
| 162 |
+
|
| 163 |
+
class TestWalkForward:
|
| 164 |
+
def test_a_real_edge_survives_out_of_sample(self):
|
| 165 |
+
result = walk_forward(trending_market(), get_strategy("momentum"), n_folds=4)
|
| 166 |
+
assert result["folds"]
|
| 167 |
+
assert result["mean_oos_sharpe"] > 0
|
| 168 |
+
assert result["efficiency"] > 0.3
|
| 169 |
+
|
| 170 |
+
def test_folds_do_not_overlap_train_and_test(self):
|
| 171 |
+
result = walk_forward(trending_market(), get_strategy("sma_cross"), n_folds=4)
|
| 172 |
+
for fold in result["folds"]:
|
| 173 |
+
assert fold["train_end"] <= fold["test_start"]
|
| 174 |
+
|
| 175 |
+
def test_short_history_degrades_gracefully(self):
|
| 176 |
+
df = trending_market(n=150)
|
| 177 |
+
result = walk_forward(df, get_strategy("momentum"), n_folds=5)
|
| 178 |
+
assert result["folds"] == []
|
| 179 |
+
assert result["note"]
|