Edwin Salguero Cursor commited on
Commit
3d257f3
·
1 Parent(s): 2d2e42a

feat: default ingest to Yahoo and restore a full README

Browse files

Use Yahoo Finance as the default tape for config, CLI, Gradio, and UIs; fail closed instead of silently simulating. Merge the algotrader 2.0 lab docs with the v1 FinRL/agentic README and drop the Northwestern line.

Co-authored-by: Cursor <cursoragent@cursor.com>

README.md CHANGED
@@ -1,127 +1,87 @@
1
- # Backtest Reality Check
2
 
3
- **algotrader 2.0 — a backtester that tries to prove itself wrong.**
4
 
5
- Most backtesting tools answer *"how much would this have made?"*. That is the easy
6
- question, and the answer is almost always flattering. This one answers the question
7
- you actually need before risking money: **how much of that was luck?**
8
 
9
- ```bash
10
- pip install -r requirements-space.txt
11
- python app.py # the Gradio app on localhost:7860
12
- python -m algotrader.cli lab --symbol SPY --strategy sma_cross
13
- ```
14
-
15
- <sub>The v1 agentic trading system (FinRL, Alpaca, Yahoo ingest, Streamlit/Dash UIs) is
16
- unchanged and still lives here — see [docs/AGENTIC_SYSTEM_V1.md](docs/AGENTIC_SYSTEM_V1.md).</sub>
17
 
18
  ---
19
 
20
- ## Two labs
21
 
22
- **The Lab** validates a timing rule on one asset. **The Portfolio Lab** validates a
23
- cross-sectional book that ranks many names — and it asks three harder questions,
24
- because a long-short book fails in ways a timing rule cannot.
25
 
26
- ## The four ways a backtest lies
27
 
28
- | The lie | The test | Where |
29
- |---|---|---|
30
- | The market had no structure to find | Monte-Carlo **permutation test** — re-run your rule on hundreds of shuffled markets | `algotrader/validation/permutation.py` |
31
- | You tried 200 things and reported the best | **Deflated Sharpe Ratio** — charge for every variant you tried | `algotrader/validation/deflated_sharpe.py` |
32
- | The parameters were fitted to the past | **PBO** (CSCV) and **walk-forward** | `algotrader/validation/pbo.py`, `walkforward.py` |
33
- | The edge is smaller than the costs | **Cost stress test** at 3× friction | `algotrader/lab.py` |
34
-
35
- Each feeds a single **Reality Score** out of 100 with a grade from A to F:
36
-
37
- | Weight | Component | What it measures |
38
- |---:|---|---|
39
- | 30% | Significance | How far outside the shuffled-market null the result sits |
40
- | 25% | Selection | Deflated Sharpe — does it clear the best-of-N bar |
41
- | 20% | Walk-forward | How much of the tuned Sharpe survived trading forward |
42
- | 15% | Overfitting | 1 − PBO |
43
- | 10% | Robustness | Sharpe retained when costs triple |
44
-
45
- The scale is deliberately harsh. On most markets, plain buy & hold beats every
46
- strategy in the arena on evidence, and the built-in coin-flip control out-ranks
47
- several respectable-looking rules. That is the finding, not a bug.
48
-
49
- ## The permutation test, concretely
50
-
51
- We take the real price series and shuffle it. Each bar's gap, high, low, body and
52
- volume are kept intact, but their **order** is destroyed. The result is a market
53
- with the same volatility and the same fat tails, and no exploitable structure at
54
- all. Then we re-run *your exact rule* on hundreds of these shuffled markets.
55
-
56
- If your Sharpe sits inside that cloud, your rule found nothing that a coin-flip
57
- market would not also have handed it. The p-value is the share of shuffled markets
58
- that did as well or better.
59
-
60
- Block mode resamples contiguous chunks instead of single bars, preserving
61
- short-horizon momentum and volatility clustering — a harder null that trend
62
- strategies deserve to be held to.
63
-
64
- ## Cross-sectional books get a harder null
65
-
66
- Shuffling the price path is the right null for a timing rule and the *wrong* one for
67
- a book that ranks names: it destroys the market's whole correlation structure, and
68
- almost any long-short book clears a null that weak.
69
-
70
- So the Portfolio Lab permutes the **weights across assets within each date**. Every
71
- calendar effect survives. Every correlation between names survives. Each date's gross
72
- exposure, net exposure and position count survive *exactly*. The only thing destroyed
73
- is the link between the strategy's choice and the asset it chose.
74
-
75
- A book that beats that null is picking names. One that doesn't was being paid for
76
- market exposure or a style tilt — which the factor regression measures directly:
77
-
78
- | Question | Test |
79
- |---|---|
80
- | Did it pick the right names? | Within-date weight permutation |
81
- | Is it alpha, or beta you can buy for 3bps? | Style regression (market, momentum, low-vol, reversal, liquidity) with White standard errors |
82
- | Does the universe contain the losers? | Survivorship measured, not assumed |
83
-
84
- That last one is not optional. A universe where every name is still trading after ten
85
- years was chosen after the fact, and every result computed on it is an upper bound.
86
- The Panel measures survival directly and the Reality Score caps at 60 when it finds
87
- none.
88
 
89
- ```python
90
- from algotrader import PortfolioLabConfig, run_portfolio_lab
 
 
 
 
91
 
92
- report = run_portfolio_lab(PortfolioLabConfig(
93
- symbols=["SPY", "QQQ", "AAPL", "MSFT", "NVDA", "GLD", "TLT"],
94
- strategy="xs_momentum",
95
- rebalance="M",
96
- ))
97
- print(report.verdict["grade"], report.permutation.p_value)
98
- print(report.attribution["note"])
99
- print(report.survivorship.note)
 
 
100
  ```
101
 
102
  ```bash
103
- python -m algotrader.cli portfolio --symbols SPY,QQQ,AAPL,MSFT,NVDA --strategy xs_momentum
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
104
  ```
105
 
106
- ## Turnover is measured against drift, not against the last target
 
 
 
 
 
 
 
 
 
 
107
 
108
- Holding 50% of a book in a name that doubles leaves you at 67% without trading. A
109
- backtest that charges turnover as `|target[t] - target[t-1]|` understates the cost of
110
- doing nothing and overstates the cost of rebalancing. Both engines measure turnover
111
- against the *drifted* weight instead, and a rebalance schedule (`D`/`W`/`M`/`Q`) lets a
112
- monthly book drift between dates rather than paying daily to stand still.
113
 
114
- ## No look-ahead, by construction
 
 
 
 
 
115
 
116
- A strategy emits a target exposure at each bar's close using only data up to that
117
- bar. The engine holds `position[t] = target[t - lag]` with `lag >= 1`, so a signal
118
- computed on Tuesday's close cannot earn Tuesday's move.
119
 
120
- That is the single line where look-ahead could enter, and the test suite asserts it
121
- from four directions — including that truncating the data never changes the equity
122
- curve before the cut, and that a `lag=0` request is refused outright.
123
 
124
- ## Python API
125
 
126
  ```python
127
  from algotrader import LabConfig, run_lab
@@ -131,114 +91,108 @@ report = run_lab(LabConfig(
131
  start="2015-01-01",
132
  strategy="sma_cross",
133
  params={"fast": 20, "slow": 100},
134
- commission_bps=1.0,
135
- slippage_bps=2.0,
136
  n_permutations=500,
137
  ))
 
 
138
 
139
- print(report.verdict["grade"], report.verdict["score"])
140
- print("p-value ", report.permutation.p_value)
141
- print("deflated Sharpe ", report.dsr["dsr"])
142
- print("overfit prob. ", report.pbo["pbo"])
143
- print("walk-forward eff.", report.walkforward["efficiency"])
144
- for flag in report.verdict["flags"]:
145
- print(" !", flag)
146
  ```
147
 
148
- Lower-level pieces compose on their own:
149
 
150
- ```python
151
- from algotrader import load_ohlcv, run_backtest, get_strategy
152
- from algotrader.types import CostModel
153
-
154
- market = load_ohlcv("BTC-USD", "2018-01-01")
155
- strategy = get_strategy("donchian_breakout")
156
- result = run_backtest(
157
- market.df,
158
- strategy.generate(market.df, {"window": 55}),
159
- costs=CostModel(commission_bps=1, slippage_bps=5, short_borrow_bps=50),
160
- )
161
- print(result.metrics["sharpe"], result.metrics["max_drawdown"])
162
- ```
163
 
164
- ## CLI
165
 
166
- ```bash
167
- python -m algotrader.cli strategies # list the zoo
168
- python -m algotrader.cli lab --symbol NVDA --strategy rsi_reversion --permutations 500
169
- python -m algotrader.cli lab --symbol SPY --param fast=10 --param slow=50 --json
170
- python -m algotrader.cli arena --symbol BTC-USD --start 2018-01-01
171
- ```
172
 
173
- `--source synthetic` forces the offline simulator, which makes runs fully
174
- deterministic and network-free.
175
 
176
- ## The strategy zoo
177
 
178
- Single asset: `buy_and_hold` · `sma_cross` · `ema_cross` · `macd_trend` ·
179
- `rsi_reversion` · `bollinger_reversion` · `donchian_breakout` · `momentum` ·
180
- `vol_target_momentum` · `channel_trend` · `coin_flip`
181
 
182
- Cross-sectional: `equal_weight` · `xs_momentum` · `xs_reversal` · `low_volatility` ·
183
- `xs_value_proxy` · `xs_random`
 
 
 
 
184
 
185
- Buy & hold and the coin flip are controls, and they stay in the arena on purpose: a
186
- leaderboard without a control group is marketing, not measurement.
187
 
188
- Adding one is a function and a registry entry — see `algotrader/strategies.py`.
 
 
189
 
190
- ## Data
191
 
192
- Live prices come from Yahoo Finance. When the network is unavailable or rate-limited,
193
- the app falls back to a deterministic market simulator — regime switching, Student-t
194
- innovations, persistent volatility — and says so on every result. Naive geometric
195
- Brownian motion flatters strategies; this simulator does not.
196
 
197
- Set `ALGOTRADER_OFFLINE=1` to skip network access entirely.
198
 
199
- ## Deploying the Hugging Face Space
 
 
 
 
 
 
 
 
 
200
 
201
- The Space ships `app.py` plus the `algotrader` package and nothing else, so it builds
202
- in well under a minute:
203
 
204
- ```bash
205
- HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <your-username>/backtest-reality-check
 
 
 
 
 
 
 
 
 
 
 
 
206
  ```
207
 
208
- Or set the `HF_TOKEN` secret and `HF_SPACE_ID` variable on the repository and let
209
- `.github/workflows/sync-hf-space.yml` publish on every push to `main`. The workflow
210
- runs the test suite and builds the app before it publishes anything.
211
 
212
- `SPACE_README.md` is the Space card (with the Hugging Face YAML frontmatter);
213
- `requirements-space.txt` is its dependency set. The root `requirements.txt` still
214
- carries the full v1 stack for CI, Docker and the FinRL agents.
215
 
216
- ## Tests
 
 
 
 
 
 
 
 
217
 
218
- ```bash
219
- python -m pytest tests/test_v2_*.py -q # 162 tests, ~25s, no network
220
- ```
221
-
222
- The validation tests check both directions, which is the part that matters: the
223
- statistics must reject noise **and** detect a real edge. They build a market with
224
- genuine serial correlation and assert that the permutation test finds it, that PBO
225
- stays near 0.5 on pure noise and drops below 0.15 when one variant is genuinely
226
- better, and that walk-forward efficiency survives.
227
 
228
- The cross-sectional null is held to the same standard, and it is calibrated: on a
229
- universe with no cross-sectional structure it returns p ≈ 0.5, and its power rises
230
- monotonically with the size of the injected effect. The tests also assert the
231
- permutation preserves each date's gross exposure, net exposure and position count
232
- exactly — if it did not, the null would be testing something else.
233
 
234
- ## References
 
 
 
235
 
236
- - Bailey & López de Prado (2014), *The Deflated Sharpe Ratio: Correcting for Selection
237
- Bias, Backtest Overfitting and Non-Normality*
238
- - Bailey, Borwein, López de Prado & Zhu (2016), *The Probability of Backtest Overfitting*
239
- - Masters (2018), *Permutation and Randomization Tests for Trading System Development*
240
 
241
- ## License
242
 
243
- Apache-2.0. Research tooling, not investment advice. Nothing here is a
244
- recommendation to trade.
 
 
1
+ # Algorithmic Trading
2
 
3
+ Parallel LLC. Two layers in one repository:
4
 
5
+ 1. **algotrader 2.0** (`algotrader/`, `app.py`): a backtester that tries to prove a rule was luck (permutation, deflated Sharpe, PBO, walk-forward, cost stress).
6
+ 2. **Agentic v1** (`agentic_ai_system/`): FinRL policies, Yahoo or Alpaca ingest, paper/live execution, Streamlit/Dash/Jupyter UIs, Docker.
 
7
 
8
+ Default market data is **Yahoo Finance** (`yfinance>=1.0`), not simulated prices. The simulator exists for offline tests (`--source synthetic` or `ALGOTRADER_OFFLINE=1` with `source=auto`). Live capital still needs a separate evaluation contract. This is research tooling, not investment advice.
 
 
 
 
 
 
 
9
 
10
  ---
11
 
12
+ ## 1. Title and Summary
13
 
14
+ **Algorithmic Trading**
15
+ Ingest real OHLCV, test whether a timing or cross-sectional rule survives a hostile null, optionally train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
 
16
 
17
+ GitHub keeps two branches: `main` (protected) and `dev` (integration).
18
 
19
+ **Design themes**
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
 
21
+ * Yahoo as the default public tape (delayed, unofficial, lookback-limited)
22
+ * Validation before belief: permutation, DSR, PBO/CSCV, walk-forward, 3× cost stress
23
+ * FinRL (PPO, A2C, DDPG, TD3) unchanged on the v1 path
24
+ * Alpaca optional for authenticated bars and orders; keys from the environment
25
+ * Synthetic GBM / regime simulator only when requested
26
+ * Secrets never in git
27
 
28
+ ---
29
+
30
+ ## 2. Quick start
31
+
32
+ ```bash
33
+ git clone https://github.com/ParallelLLC/algorithmic_trading.git
34
+ cd algorithmic_trading
35
+ python -m venv .venv && source .venv/bin/activate
36
+ pip install -r requirements-space.txt # algotrader + Gradio
37
+ # or: pip install -r requirements.txt # full v1 stack (FinRL, Dash, Docker CI)
38
  ```
39
 
40
  ```bash
41
+ python app.py # Gradio, localhost:7860, Yahoo by default
42
+ python -m algotrader.cli lab --symbol SPY --strategy sma_cross
43
+ python -m algotrader.cli lab --symbol NVDA --strategy rsi_reversion --permutations 500
44
+ python -m agentic_ai_system.main --mode backtest --start-date 2024-01-01 --end-date 2024-12-31
45
+ ```
46
+
47
+ `config.yaml` defaults:
48
+
49
+ ```yaml
50
+ data_source:
51
+ type: 'yahoo'
52
+ trading:
53
+ symbol: 'AAPL'
54
+ timeframe: '1d' # Yahoo 1m history is ~7 days; use 1d for multi-year windows
55
+ yahoo:
56
+ auto_adjust: true # raw Close turns splits into fake crashes
57
  ```
58
 
59
+ Alpaca is opt-in: `ALPACA_API_KEY` / `ALPACA_SECRET_KEY` and `data_source.type: alpaca` or `execution.broker_api: alpaca_paper`.
60
+
61
+ ---
62
+
63
+ ## 3. algotrader 2.0 (validation lab)
64
+
65
+ Most backtests answer "how much would this have made?" This one asks **how much of that was luck?**
66
+
67
+ ### Two labs
68
+
69
+ **The Lab** validates a timing rule on one asset. **The Portfolio Lab** validates a cross-sectional book that ranks many names.
70
 
71
+ ### The four ways a backtest lies
 
 
 
 
72
 
73
+ | The lie | The test | Where |
74
+ | --- | --- | --- |
75
+ | The market had no structure to find | Monte-Carlo permutation (shuffle bar order, keep gap/high/low/body/volume) | `algotrader/validation/permutation.py` |
76
+ | You tried 200 things and reported the best | Deflated Sharpe Ratio | `algotrader/validation/deflated_sharpe.py` |
77
+ | Parameters were fitted to the past | PBO (CSCV) and walk-forward | `algotrader/validation/pbo.py`, `walkforward.py` |
78
+ | The edge is smaller than the costs | Cost stress at 3× friction | `algotrader/lab.py` |
79
 
80
+ Reality Score (0–100, grades A–F): significance 30%, selection 25%, walk-forward 20%, overfitting 15%, robustness 10%. The scale is harsh on purpose. Buy-and-hold and a coin-flip stay in the arena as controls.
 
 
81
 
82
+ Cross-sectional books use a **within-date weight permutation** so market correlation survives; path-shuffle is the wrong null for a long-short ranker. Survivorship is measured. Style regression (market, momentum, low-vol, reversal, liquidity) with White standard errors.
 
 
83
 
84
+ Look-ahead: `position[t] = target[t - lag]` with `lag >= 1`. Turnover is measured against drifted weights, not `|target[t]-target[t-1]|`.
85
 
86
  ```python
87
  from algotrader import LabConfig, run_lab
 
91
  start="2015-01-01",
92
  strategy="sma_cross",
93
  params={"fast": 20, "slow": 100},
94
+ source="yahoo",
 
95
  n_permutations=500,
96
  ))
97
+ print(report.verdict["grade"], report.permutation.p_value, report.dsr["dsr"])
98
+ ```
99
 
100
+ ```bash
101
+ python -m algotrader.cli strategies
102
+ python -m algotrader.cli lab --symbol SPY --source yahoo
103
+ python -m algotrader.cli portfolio --symbols SPY,QQQ,AAPL,MSFT,NVDA --strategy xs_momentum
104
+ python -m algotrader.cli lab --source synthetic # offline tests only
 
 
105
  ```
106
 
107
+ Single-asset zoo: `buy_and_hold`, `sma_cross`, `ema_cross`, `macd_trend`, `rsi_reversion`, `bollinger_reversion`, `donchian_breakout`, `momentum`, `vol_target_momentum`, `channel_trend`, `coin_flip`.
108
 
109
+ Cross-sectional: `equal_weight`, `xs_momentum`, `xs_reversal`, `low_volatility`, `xs_value_proxy`, `xs_random`.
 
 
 
 
 
 
 
 
 
 
 
 
110
 
111
+ **Data:** `load_ohlcv(..., source="yahoo")` downloads from Yahoo and **raises** if the download is empty. `source="auto"` is the Space fallback (cache, then simulator). `ALGOTRADER_OFFLINE=1` disables the network.
112
 
113
+ **HF Space:** `HF_TOKEN=hf_xxx ./scripts/deploy_hf_space.sh <user>/backtest-reality-check`. Card is `SPACE_README.md`. Tests: `python -m pytest tests/test_v2_*.py -q`.
 
 
 
 
 
114
 
115
+ References: Bailey & López de Prado (2014) DSR; Bailey et al. (2016) PBO; Masters (2018) permutation tests for trading systems.
 
116
 
117
+ ---
118
 
119
+ ## 4. Concepts and methods (v1 ingest and execution)
 
 
120
 
121
+ | Source | Default? | Failure modes |
122
+ | ------ | -------- | ------------- |
123
+ | **Yahoo** | Yes (`config.yaml`, algotrader CLI, Gradio) | Unofficial API, ~15 min delay, 1m ≈ 7 days, split-adjustment required (`auto_adjust: true`) |
124
+ | **Alpaca** | Optional | Auth, feed, rate limits |
125
+ | **CSV** | Replay | Missing path or OHLCV columns |
126
+ | **Synthetic** | Tests / `--source synthetic` | Not tradable edge |
127
 
128
+ `agentic_ai_system.data_ingestion.load_data` dispatches on `data_source.type`. Yahoo stream: `yahoo_data_stream.py` (clamped lookback, no incomplete bars by default).
 
129
 
130
+ * `StrategyAgent`: SMA, RSI, Bollinger, MACD on Close (teaching rule, not an alpha claim)
131
+ * `FinRLAgent`: PPO / A2C / DDPG / TD3 via Stable-Baselines3
132
+ * `ExecutionAgent` / `AlpacaBroker`: paper simulation or Alpaca orders
133
 
134
+ v1 `run_backtest` is a single in-sample pass unless you use algotrader walk-forward. Leakage is the null hypothesis.
135
 
136
+ ---
 
 
 
137
 
138
+ ## 5. Stack
139
 
140
+ | Layer | Tools |
141
+ | ----- | ----- |
142
+ | Language | Python 3.11 (CI) |
143
+ | Validation | algotrader (permutation, DSR, PBO, walk-forward) |
144
+ | RL | FinRL / Stable-Baselines3, Gym/Gymnasium, PyTorch |
145
+ | Market data | yfinance ≥ 1.0 (default); alpaca-py optional |
146
+ | Tabular | pandas, NumPy, scikit-learn |
147
+ | UI | Gradio (`app.py`); Streamlit, Dash, Jupyter (v1) |
148
+ | Deploy | Docker Compose, GitHub Actions, Hugging Face Space |
149
+ | Tests | pytest |
150
 
151
+ ---
 
152
 
153
+ ## 6. Structure
154
+
155
+ ```
156
+ algorithmic_trading/
157
+ ├── algotrader/ # 2.0 lab, engine, validation, strategies
158
+ ├── app.py # Gradio Reality Check
159
+ ├── agentic_ai_system/ # v1 FinRL, Yahoo/Alpaca ingest, execution
160
+ ├── ui/ # Streamlit, Dash, Jupyter, WebSocket
161
+ ├── tests/
162
+ ├── docs/AGENTIC_SYSTEM_V1.md # v1 notes
163
+ ├── config.yaml # default data_source.type: yahoo
164
+ ├── requirements-space.txt # Space / algotrader
165
+ ├── requirements.txt # full v1 + CI
166
+ └── scripts/deploy_hf_space.sh
167
  ```
168
 
169
+ ---
 
 
170
 
171
+ ## 7. Configuration
 
 
172
 
173
+ | Key | Meaning |
174
+ | --- | ------- |
175
+ | `data_source.type` | `yahoo` (default) \| `csv` \| `synthetic` \| `alpaca` |
176
+ | `trading.timeframe` | Mapped to Yahoo intervals; use `1d` for multi-year history |
177
+ | `yahoo.auto_adjust` | Split/dividend adjust (keep true) |
178
+ | `yahoo.emit_incomplete_bars` | Default false; forming bars are not closes |
179
+ | `execution.broker_api` | `paper` \| `alpaca_paper` \| `alpaca_live` |
180
+ | `finrl.algorithm` | PPO, A2C, DDPG, TD3 |
181
+ | algotrader `--source` | `yahoo` (default) \| `auto` \| `cache` \| `synthetic` |
182
 
183
+ ---
 
 
 
 
 
 
 
 
184
 
185
+ ## 8. Tests and ops
 
 
 
 
186
 
187
+ ```bash
188
+ python -m pytest tests/test_v2_*.py -q
189
+ python -m pytest tests/test_yahoo_data_stream.py tests/test_data_ingestion.py -q
190
+ ```
191
 
192
+ UI launchers and Docker: `UI_SETUP.md`, `DOCKER_HUB_SETUP.md`. Branch policy: `main` and `dev` only. Do not re-enable Dependabot.
 
 
 
193
 
194
+ ---
195
 
196
+ **License:** Apache License 2.0
197
+ **Organization:** [Parallel LLC](https://github.com/ParallelLLC)
198
+ **Repository:** <https://github.com/ParallelLLC/algorithmic_trading>
algotrader/cli.py CHANGED
@@ -33,8 +33,8 @@ def _add_common(parser: argparse.ArgumentParser) -> None:
33
  parser.add_argument("--end", default=None)
34
  parser.add_argument("--interval", default="1d")
35
  parser.add_argument(
36
- "--source", default="auto", choices=["auto", "cache", "synthetic"],
37
- help="'auto' downloads and falls back offline; 'synthetic' forces the simulator.",
38
  )
39
  parser.add_argument("--commission-bps", type=float, default=1.0)
40
  parser.add_argument("--slippage-bps", type=float, default=2.0)
 
33
  parser.add_argument("--end", default=None)
34
  parser.add_argument("--interval", default="1d")
35
  parser.add_argument(
36
+ "--source", default="yahoo", choices=["yahoo", "live", "auto", "cache", "synthetic"],
37
+ help="'yahoo' requires a Yahoo download. 'synthetic' is offline tests only. 'auto' falls back to the simulator.",
38
  )
39
  parser.add_argument("--commission-bps", type=float, default=1.0)
40
  parser.add_argument("--slippage-bps", type=float, default=2.0)
algotrader/data.py CHANGED
@@ -209,12 +209,12 @@ def load_ohlcv(
209
  start: str = "2015-01-01",
210
  end: str | None = None,
211
  interval: str = "1d",
212
- source: str = "auto",
213
  ) -> MarketData:
214
- """Load OHLCV for ``symbol``, never raising for a merely-unreachable network.
215
 
216
- ``source`` is one of ``auto`` (download, then cache, then simulate),
217
- ``cache``, or ``synthetic``.
218
  """
219
  symbol = (symbol or "SPY").strip().upper()
220
 
@@ -222,11 +222,17 @@ def load_ohlcv(
222
  df = simulate_ohlcv(symbol, start, end, interval)
223
  return MarketData(symbol, df, "synthetic", interval, "Simulated prices (requested).")
224
 
225
- if source in ("auto", "live"):
226
  df = _download(symbol, start, end, interval)
227
  if df is not None and len(df) > 50:
228
  _write_cache(symbol, interval, df)
229
  return MarketData(symbol, df, "yfinance", interval, "Live data from Yahoo Finance.")
 
 
 
 
 
 
230
 
231
  cached = _read_cache(symbol, interval)
232
  if cached is not None and len(cached) > 50:
 
209
  start: str = "2015-01-01",
210
  end: str | None = None,
211
  interval: str = "1d",
212
+ source: str = "yahoo",
213
  ) -> MarketData:
214
+ """Load OHLCV for ``symbol``.
215
 
216
+ Default ``source='yahoo'`` requires a Yahoo download. ``auto`` still falls
217
+ back to cache then the simulator (Hugging Face Space). ``synthetic`` is tests only.
218
  """
219
  symbol = (symbol or "SPY").strip().upper()
220
 
 
222
  df = simulate_ohlcv(symbol, start, end, interval)
223
  return MarketData(symbol, df, "synthetic", interval, "Simulated prices (requested).")
224
 
225
+ if source in ("yahoo", "live", "auto"):
226
  df = _download(symbol, start, end, interval)
227
  if df is not None and len(df) > 50:
228
  _write_cache(symbol, interval, df)
229
  return MarketData(symbol, df, "yfinance", interval, "Live data from Yahoo Finance.")
230
+ if source in ("yahoo", "live"):
231
+ raise RuntimeError(
232
+ f"Yahoo returned no usable bars for {symbol}. "
233
+ "Check the ticker, date range, and network. "
234
+ "Pass source='synthetic' only for offline tests."
235
+ )
236
 
237
  cached = _read_cache(symbol, interval)
238
  if cached is not None and len(cached) > 50:
algotrader/lab.py CHANGED
@@ -38,7 +38,7 @@ class LabConfig:
38
  start: str = "2015-01-01"
39
  end: Optional[str] = None
40
  interval: str = "1d"
41
- source: str = "auto"
42
 
43
  strategy: str = "sma_cross"
44
  params: Dict[str, float] = field(default_factory=dict)
 
38
  start: str = "2015-01-01"
39
  end: Optional[str] = None
40
  interval: str = "1d"
41
+ source: str = "yahoo"
42
 
43
  strategy: str = "sma_cross"
44
  params: Dict[str, float] = field(default_factory=dict)
algotrader/portfolio_lab.py CHANGED
@@ -52,7 +52,7 @@ class PortfolioLabConfig:
52
  start: str = "2015-01-01"
53
  end: Optional[str] = None
54
  interval: str = "1d"
55
- source: str = "auto"
56
 
57
  strategy: str = "xs_momentum"
58
  params: Dict[str, float] = field(default_factory=dict)
 
52
  start: str = "2015-01-01"
53
  end: Optional[str] = None
54
  interval: str = "1d"
55
+ source: str = "yahoo"
56
 
57
  strategy: str = "xs_momentum"
58
  params: Dict[str, float] = field(default_factory=dict)
app.py CHANGED
@@ -342,6 +342,7 @@ def analyse_portfolio(
342
  symbols=universe,
343
  start=start or "2015-01-01",
344
  end=end or None,
 
345
  strategy=strategy_key,
346
  params=params,
347
  commission_bps=float(commission),
@@ -450,6 +451,7 @@ def analyse(
450
  symbol=symbol or "SPY",
451
  start=start or "2015-01-01",
452
  end=end or None,
 
453
  strategy=strategy_key,
454
  params=collect_params(strategy_key, p1, p2, p3),
455
  commission_bps=float(commission),
@@ -484,6 +486,7 @@ def race(symbol: str, start: str, allow_short: bool, n_permutations: int, progre
484
  cfg = LabConfig(
485
  symbol=symbol or "SPY",
486
  start=start or "2015-01-01",
 
487
  allow_short=bool(allow_short),
488
  )
489
  table, market, _ = run_arena(
 
342
  symbols=universe,
343
  start=start or "2015-01-01",
344
  end=end or None,
345
+ source="yahoo",
346
  strategy=strategy_key,
347
  params=params,
348
  commission_bps=float(commission),
 
451
  symbol=symbol or "SPY",
452
  start=start or "2015-01-01",
453
  end=end or None,
454
+ source="yahoo",
455
  strategy=strategy_key,
456
  params=collect_params(strategy_key, p1, p2, p3),
457
  commission_bps=float(commission),
 
486
  cfg = LabConfig(
487
  symbol=symbol or "SPY",
488
  start=start or "2015-01-01",
489
+ source="yahoo",
490
  allow_short=bool(allow_short),
491
  )
492
  table, market, _ = run_arena(
config.yaml CHANGED
@@ -1,11 +1,11 @@
1
  # Configuration file for the agentic AI trading system
2
  data_source:
3
- type: 'csv'
4
  path: 'data/market_data.csv'
5
 
6
  trading:
7
  symbol: 'AAPL'
8
- timeframe: '1m'
9
  capital: 100000
10
 
11
  risk:
@@ -29,7 +29,7 @@ alpaca:
29
  websocket_url: 'wss://stream.data.alpaca.markets/v2/iex' # WebSocket URL
30
  account_type: 'paper' # 'paper' or 'live'
31
 
32
- # Yahoo Finance (optional). Set data_source.type: 'yahoo' to use.
33
  # Unofficial API, typically delayed; 1m lookback is ~7 days.
34
  yahoo:
35
  poll_interval_seconds: 60
 
1
  # Configuration file for the agentic AI trading system
2
  data_source:
3
+ type: 'yahoo'
4
  path: 'data/market_data.csv'
5
 
6
  trading:
7
  symbol: 'AAPL'
8
+ timeframe: '1d'
9
  capital: 100000
10
 
11
  risk:
 
29
  websocket_url: 'wss://stream.data.alpaca.markets/v2/iex' # WebSocket URL
30
  account_type: 'paper' # 'paper' or 'live'
31
 
32
+ # Yahoo Finance (default ingest). Unofficial API, typically delayed; 1m lookback is ~7 days.
33
  # Unofficial API, typically delayed; 1m lookback is ~7 days.
34
  yahoo:
35
  poll_interval_seconds: 60
docs/AGENTIC_SYSTEM_V1.md CHANGED
@@ -1,6 +1,6 @@
1
  # Algorithmic Trading
2
 
3
- FinRL reinforcement-learning trading with Alpaca execution, plus optional Yahoo Finance OHLCV for unlabeled real-price research. Parallel LLC.
4
 
5
  This is **research and paper-trading infrastructure**. Live capital requires a separate evaluation contract, feature-parity tests, and a rewritten execution path. Do not treat `paper_trading: false` as a promotion gate.
6
 
@@ -9,9 +9,9 @@ This is **research and paper-trading infrastructure**. Live capital requires a s
9
  ## 1. Title and Summary
10
 
11
  **Algorithmic Trading**
12
- Northwestern-trained data-engineering practice applied to a trading loop: ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
13
 
14
- GitHub `main` is the FinRL / Docker / Streamlit tree. `dev` is the integration branch. Yahoo is an additive `data_source.type`, not a replacement for Alpaca or FinRL.
15
 
16
  **Design themes**
17
 
 
1
  # Algorithmic Trading
2
 
3
+ FinRL reinforcement-learning trading with Alpaca execution, plus Yahoo Finance OHLCV as the default public tape. Parallel LLC.
4
 
5
  This is **research and paper-trading infrastructure**. Live capital requires a separate evaluation contract, feature-parity tests, and a rewritten execution path. Do not treat `paper_trading: false` as a promotion gate.
6
 
 
9
  ## 1. Title and Summary
10
 
11
  **Algorithmic Trading**
12
+ Ingest OHLCV, compute indicators or train a FinRL policy, size orders under position and drawdown caps, route to paper or live Alpaca.
13
 
14
+ GitHub `main` is the FinRL / Docker / Streamlit tree plus algotrader 2.0. `dev` is the integration branch. Yahoo is the default `data_source.type`.
15
 
16
  **Design themes**
17
 
tests/test_v2_strategies.py CHANGED
@@ -119,10 +119,15 @@ class TestData:
119
  assert returns.kurtosis() > 1.0
120
  assert returns.abs().autocorr(1) > 0.05
121
 
 
 
 
 
 
122
  def test_offline_load_falls_back_and_says_so(self, monkeypatch):
123
  monkeypatch.setattr("algotrader.data._download", lambda *a, **k: None)
124
  monkeypatch.setattr("algotrader.data._read_cache", lambda *a, **k: None)
125
- market = load_ohlcv("SPY", "2018-01-01", "2022-01-01")
126
  assert market.source == "synthetic"
127
  assert not market.is_real
128
  assert "unavailable" in market.note
 
119
  assert returns.kurtosis() > 1.0
120
  assert returns.abs().autocorr(1) > 0.05
121
 
122
+ def test_yahoo_source_does_not_silently_simulate(self, monkeypatch):
123
+ monkeypatch.setattr("algotrader.data._download", lambda *a, **k: None)
124
+ with pytest.raises(RuntimeError, match="Yahoo returned no usable bars"):
125
+ load_ohlcv("SPY", "2018-01-01", "2022-01-01", source="yahoo")
126
+
127
  def test_offline_load_falls_back_and_says_so(self, monkeypatch):
128
  monkeypatch.setattr("algotrader.data._download", lambda *a, **k: None)
129
  monkeypatch.setattr("algotrader.data._read_cache", lambda *a, **k: None)
130
+ market = load_ohlcv("SPY", "2018-01-01", "2022-01-01", source="auto")
131
  assert market.source == "synthetic"
132
  assert not market.is_real
133
  assert "unavailable" in market.note
ui/dash_app.py CHANGED
@@ -158,11 +158,12 @@ class TradingDashApp:
158
  dbc.Select(
159
  id="data-source-select",
160
  options=[
 
161
  {"label": "CSV File", "value": "csv"},
162
  {"label": "Alpaca API", "value": "alpaca"},
163
  {"label": "Synthetic Data", "value": "synthetic"}
164
  ],
165
- value="csv"
166
  )
167
  ], width=4),
168
  dbc.Col([
 
158
  dbc.Select(
159
  id="data-source-select",
160
  options=[
161
+ {"label": "Yahoo Finance", "value": "yahoo"},
162
  {"label": "CSV File", "value": "csv"},
163
  {"label": "Alpaca API", "value": "alpaca"},
164
  {"label": "Synthetic Data", "value": "synthetic"}
165
  ],
166
+ value="yahoo"
167
  )
168
  ], width=4),
169
  dbc.Col([
ui/jupyter_widgets.py CHANGED
@@ -62,8 +62,8 @@ class TradingJupyterUI:
62
 
63
  # Data widgets
64
  self.data_source = widgets.Dropdown(
65
- options=['csv', 'alpaca', 'synthetic'],
66
- value='csv',
67
  description='Data Source:',
68
  style={'description_width': '120px'}
69
  )
@@ -76,7 +76,7 @@ class TradingJupyterUI:
76
 
77
  self.timeframe_input = widgets.Dropdown(
78
  options=['1m', '5m', '15m', '1h', '1d'],
79
- value='1m',
80
  description='Timeframe:',
81
  style={'description_width': '120px'}
82
  )
 
62
 
63
  # Data widgets
64
  self.data_source = widgets.Dropdown(
65
+ options=['yahoo', 'csv', 'alpaca', 'synthetic'],
66
+ value='yahoo',
67
  description='Data Source:',
68
  style={'description_width': '120px'}
69
  )
 
76
 
77
  self.timeframe_input = widgets.Dropdown(
78
  options=['1m', '5m', '15m', '1h', '1d'],
79
+ value='1d',
80
  description='Timeframe:',
81
  style={'description_width': '120px'}
82
  )