CodeDevX commited on
Commit
4ab75b3
·
verified ·
1 Parent(s): 99a121f

Initial release: 7-domain LSTM prediction system with chat interface

Browse files
README.md ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: pytorch
4
+ tags:
5
+ - time-series
6
+ - forecasting
7
+ - lstm
8
+ - finance
9
+ - weather
10
+ - multi-domain
11
+ language:
12
+ - en
13
+ metrics:
14
+ - mae
15
+ - mape
16
+ pipeline_tag: time-series-forecasting
17
+ ---
18
+
19
+ # Future Prediction Models (Multi-Domain LSTM)
20
+
21
+ An end-to-end multi-domain forecasting system. For each of **7 topics** it automatically fetches a real dataset, trains a 2-layer LSTM, evaluates on a held-out test set (no leakage), and generates ChatGPT-style future predictions.
22
+
23
+ ## Topics & validated accuracy
24
+
25
+ | Topic | Asset | Source | H1 MAPE | vs naive |
26
+ |---|---|---|---|---|
27
+ | AI | NVIDIA daily close | Yahoo Finance | 1.86% | ~naive |
28
+ | Economy | S&P 500 daily close | Yahoo Finance | 0.64% | beats naive 1.8% |
29
+ | Energy | WTI crude oil close | Yahoo Finance | 2.83% | ~naive |
30
+ | Finance | Bitcoin BTC-USD | Yahoo Finance | 1.52% | ~naive |
31
+ | Programming | `react` npm downloads | npm registry | 7.12% | **beats naive 77.5%** |
32
+ | Sports | ATP world #1 Elo | Hugging Face tennis | 0.06% | ~naive |
33
+ | Weather | Daily mean temperature | Open-Meteo | 0.18% | beats naive 1.9% |
34
+
35
+ > Markets behave near a random walk, so 100% accuracy is impossible — these are honest, validated numbers.
36
+
37
+ ## Quick start
38
+
39
+ ```bash
40
+ pip install -r requirements.txt
41
+ python data.py --topic finance --refresh # fetch real data
42
+ python train.py --topic finance # train the LSTM
43
+ python ask.py "what will bitcoin do next week?"
44
+ python ask.py "compare all topics"
45
+ ```
46
+
47
+ ## Architecture
48
+
49
+ ```
50
+ user question -> intent detection (response.py)
51
+ -> prediction_engine.py (structured metrics, prints nothing)
52
+ -> response_formatter.py (natural summary + Markdown)
53
+ -> ask.py (chat CLI, stdout)
54
+ ```
55
+
56
+ The engine returns clearly separated metrics — never conflating them:
57
+ - `net_change_pct` (first → last forecast along the path)
58
+ - `forecast_range_pct` (max − min of the forecast path)
59
+ - `current_to_forecast_pct` (final forecast vs latest observed value)
60
+ - `direction` (rising / falling / sideways, configurable threshold)
61
+
62
+ Only the final assistant response reaches the user; internal logs go to stderr and only when `FORECAST_DEBUG=1`.
63
+
64
+ ## Files
65
+
66
+ - `model_<topic>.pt` — trained PyTorch LSTM checkpoint per topic
67
+ - `config_<topic>.json` — hyperparameters + validation metrics per topic
68
+ - `data.py`, `train.py`, `predict.py`, `model.py` — data + training + inference core
69
+ - `ask.py`, `response.py`, `response_formatter.py`, `prediction_engine.py`, `debug_logger.py` — the chat/response layer
70
+ - `weather_cities.py` — city registry for location-specific weather models
71
+
72
+ ## Weather for any city
73
+
74
+ ```bash
75
+ export WEATHER_NAME='Chennai, India' WEATHER_LAT=13.0827 WEATHER_LON=80.2707
76
+ python data.py --topic weather --refresh
77
+ python train.py --topic weather
78
+ python ask.py "predict weather in chennai tomorrow"
79
+ ```
80
+
81
+ ## Add a new domain
82
+
83
+ 1. Add an entry to `TOPICS` in `data.py` (label, unit, asset)
84
+ 2. Add a fetcher returning `{source, series}` to `FETCHERS`
85
+ 3. Run `python train.py --topic <name>` — everything else is automatic
86
+
87
+ ## License
88
+
89
+ MIT
ask.py ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Chat interface. Only the final assistant response reaches stdout.
2
+
3
+ Usage:
4
+ python ask.py # interactive chat
5
+ python ask.py "what will bitcoin do next week?"
6
+ python ask.py --debug "compare all topics" # internal logs on stderr
7
+ """
8
+
9
+ import argparse
10
+ import sys
11
+ import time
12
+
13
+ import debug_logger
14
+ import response
15
+
16
+
17
+ def main():
18
+ parser = argparse.ArgumentParser()
19
+ parser.add_argument("question", nargs="*", help="one-shot question (omit for interactive chat)")
20
+ parser.add_argument("--stream", action="store_true", help="stream output line by line")
21
+ parser.add_argument("--debug", action="store_true", help="show internal logs on stderr")
22
+ parser.add_argument("--detailed", action="store_true", help="include day-by-day table")
23
+ parser.add_argument("--device", default="cpu")
24
+ args = parser.parse_args()
25
+
26
+ if args.debug:
27
+ import os
28
+ os.environ["FORECAST_DEBUG"] = "1"
29
+
30
+ config = response.GenerationConfig(debug=args.debug, detailed=args.detailed)
31
+ ctx = response.ConversationContext(max_turns=8)
32
+
33
+ def ask(question: str) -> str:
34
+ answer = debug_logger.strip_internal(
35
+ response.generate_response(question, device=args.device, config=config))
36
+ if args.stream:
37
+ for chunk in response.stream_response(question, pregenerated=answer):
38
+ sys.stdout.write(chunk)
39
+ sys.stdout.flush()
40
+ time.sleep(config.chunk_delay)
41
+ if answer and not answer.endswith("\n"):
42
+ print()
43
+ else:
44
+ print(answer)
45
+ ctx.add("user", question)
46
+ ctx.add("assistant", answer)
47
+ return answer
48
+
49
+ if args.question:
50
+ ask(" ".join(args.question))
51
+ return
52
+
53
+ print("Forecast assistant - ask about AI, Programming, Finance, Sports,")
54
+ print("Weather, Economy, or Energy. Example: 'what will bitcoin do next week?'")
55
+ while True:
56
+ try:
57
+ q = input("you> ").strip()
58
+ except (EOFError, KeyboardInterrupt):
59
+ print()
60
+ break
61
+ if not q:
62
+ continue
63
+ if q.lower() in ("quit", "exit", "q"):
64
+ break
65
+ ask(q)
66
+
67
+
68
+ if __name__ == "__main__":
69
+ main()
config_ai.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "ai",
3
+ "topic_label": "AI",
4
+ "unit": "USD",
5
+ "source": "yfinance:NVDA",
6
+ "lookback": 90,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 128,
10
+ "num_layers": 3,
11
+ "dropout": 0.1,
12
+ "val_mae": 6.727893872130523,
13
+ "val_rel_mae": 0.03455442935228348,
14
+ "val_h1_mape": 0.018591665790424962,
15
+ "naive_baseline_mae": 6.687884876155594,
16
+ "last_date": "2026-08-14",
17
+ "last_value": 225.16000366210935
18
+ }
config_economy.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "economy",
3
+ "topic_label": "Economy",
4
+ "unit": "USD",
5
+ "source": "yfinance:^GSPC",
6
+ "lookback": 60,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 64,
10
+ "num_layers": 2,
11
+ "dropout": 0.1,
12
+ "val_mae": 85.50677981465654,
13
+ "val_rel_mae": 0.012141019105911255,
14
+ "val_h1_mape": 0.00639636122128309,
15
+ "naive_baseline_mae": 87.09085990816496,
16
+ "last_date": "2026-08-14",
17
+ "last_value": 7785.759765625
18
+ }
config_energy.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "energy",
3
+ "topic_label": "Energy",
4
+ "unit": "USD",
5
+ "source": "yfinance:CL=F(WTI)",
6
+ "lookback": 60,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 64,
10
+ "num_layers": 2,
11
+ "dropout": 0.1,
12
+ "val_mae": 4.7798134313437055,
13
+ "val_rel_mae": 0.05749673768877983,
14
+ "val_h1_mape": 0.02831358343785865,
15
+ "naive_baseline_mae": 4.782200717277816,
16
+ "last_date": "2026-08-18",
17
+ "last_value": 84.04000091552734
18
+ }
config_finance.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "finance",
3
+ "topic_label": "Finance",
4
+ "unit": "USD",
5
+ "source": "yfinance:BTC-USD",
6
+ "lookback": 90,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 128,
10
+ "num_layers": 3,
11
+ "dropout": 0.1,
12
+ "val_mae": 1982.9512579096927,
13
+ "val_rel_mae": 0.028562922030687332,
14
+ "val_h1_mape": 0.0152207905176555,
15
+ "naive_baseline_mae": 1969.4407217279584,
16
+ "last_date": "2026-08-17",
17
+ "last_value": 63229.1796875
18
+ }
config_programming.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "programming",
3
+ "topic_label": "Programming",
4
+ "unit": "downloads",
5
+ "source": "npm-api:react-daily-downloads",
6
+ "lookback": 90,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 128,
10
+ "num_layers": 3,
11
+ "dropout": 0.1,
12
+ "val_mae": 1173766.7827189027,
13
+ "val_rel_mae": 0.059618908911943436,
14
+ "val_h1_mape": 0.07118344794329942,
15
+ "naive_baseline_mae": 5225264.470864349,
16
+ "last_date": "2026-08-17",
17
+ "last_value": 15849863.0
18
+ }
config_sports.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "sports",
3
+ "topic_label": "Sports",
4
+ "unit": "Elo",
5
+ "source": "huggingface:davidtadediji/tennis-atp->world#1-elo",
6
+ "lookback": 90,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 128,
10
+ "num_layers": 3,
11
+ "dropout": 0.1,
12
+ "val_mae": 3.3986784698540626,
13
+ "val_rel_mae": 0.0015747122233733535,
14
+ "val_h1_mape": 0.0005584839231664478,
15
+ "naive_baseline_mae": 3.1888383318009375,
16
+ "last_date": "2024-12-18",
17
+ "last_value": 2217.7739373895524
18
+ }
config_weather_chennai_india.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "weather",
3
+ "topic_label": "Weather",
4
+ "unit": "K",
5
+ "source": "open-meteo:Chennai, India",
6
+ "lookback": 60,
7
+ "horizon_steps": 7,
8
+ "horizon_days": 7,
9
+ "hidden_size": 64,
10
+ "num_layers": 2,
11
+ "dropout": 0.1,
12
+ "val_mae": 0.9925612108358967,
13
+ "val_rel_mae": 0.0032603859435766935,
14
+ "val_h1_mape": 0.0018425700744697428,
15
+ "naive_baseline_mae": 1.0116436048493131,
16
+ "last_date": "2026-08-18",
17
+ "last_value": 303.25
18
+ }
data.py ADDED
@@ -0,0 +1,475 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import argparse
2
+ import json
3
+ import os
4
+ import re
5
+ from pathlib import Path
6
+
7
+ import numpy as np
8
+ import pandas as pd
9
+
10
+ os.environ.setdefault("HF_HUB_DOWNLOAD_TIMEOUT", "30")
11
+
12
+ BASE_DIR = Path(__file__).resolve().parent
13
+ DATA_DIR = BASE_DIR / "data"
14
+
15
+ WEATHER_NAME = os.environ.get("WEATHER_NAME", "New York")
16
+ WEATHER_LAT = os.environ.get("WEATHER_LAT", "40.7128")
17
+ WEATHER_LON = os.environ.get("WEATHER_LON", "-74.0060")
18
+ _WEATHER_NAME = WEATHER_NAME
19
+ _WEATHER_LAT = WEATHER_LAT
20
+ _WEATHER_LON = WEATHER_LON
21
+
22
+ TOPICS = {
23
+ "ai": {"label": "AI", "unit": "USD", "asset": "NVIDIA stock",
24
+ "desc": "NVIDIA stock close price (AI market leader)"},
25
+ "programming": {"label": "Programming", "unit": "downloads", "asset": "react downloads",
26
+ "desc": "npm downloads for 'react' (JS ecosystem demand)"},
27
+ "finance": {"label": "Finance", "unit": "USD", "asset": "Bitcoin",
28
+ "desc": "Bitcoin BTC-USD daily close"},
29
+ "sports": {"label": "Sports", "unit": "Elo", "asset": "ATP world #1 Elo rating",
30
+ "desc": "ATP world #1 Elo rating, computed from full tennis match history"},
31
+ "weather": {"label": "Weather", "unit": "K", "asset": f"{WEATHER_NAME} temperature",
32
+ "desc": f"Daily mean temperature for {WEATHER_NAME} (Open-Meteo archive, stored in Kelvin)"},
33
+ "economy": {"label": "Economy", "unit": "USD", "asset": "S&P 500 index",
34
+ "desc": "S&P 500 daily close (US market benchmark)"},
35
+ "energy": {"label": "Energy", "unit": "USD", "asset": "crude oil price (WTI)",
36
+ "desc": "WTI crude oil futures daily close"},
37
+ }
38
+
39
+ DATE_COL_RE = re.compile(r"(date|time|stamp|day|period|^dt$)", re.IGNORECASE)
40
+ TARGET_CANDIDATES = ["close", "price", "value", "usd", "amount", "open", "target", "y"]
41
+
42
+
43
+ def _strip_tz(parsed):
44
+ if getattr(parsed.dt, "tz", None) is not None:
45
+ parsed = parsed.dt.tz_localize(None)
46
+ return parsed
47
+
48
+
49
+ def _find_date_col(df):
50
+ for col in df.columns:
51
+ if DATE_COL_RE.search(str(col)):
52
+ parsed = pd.to_datetime(df[col], errors="coerce")
53
+ if parsed.notna().mean() > 0.8:
54
+ return col, _strip_tz(parsed)
55
+ for col in df.columns:
56
+ if str(df[col].dtype).startswith("datetime"):
57
+ return col, _strip_tz(pd.to_datetime(df[col]))
58
+ for col in df.columns:
59
+ parsed = pd.to_datetime(df[col], errors="coerce")
60
+ if parsed.notna().mean() > 0.8:
61
+ return col, _strip_tz(parsed)
62
+ return None, None
63
+
64
+
65
+ def _find_target_col(df, date_col):
66
+ numeric_cols = [c for c in df.select_dtypes(include=[np.number]).columns if c != date_col]
67
+ series_map = {c: df[c] for c in numeric_cols}
68
+ for col in df.columns:
69
+ if col == date_col or col in series_map:
70
+ continue
71
+ coerced = pd.to_numeric(df[col], errors="coerce")
72
+ if coerced.notna().mean() > 0.6:
73
+ series_map[col] = coerced
74
+ if not series_map:
75
+ return None, None
76
+ for name in TARGET_CANDIDATES:
77
+ for col in series_map:
78
+ if name in str(col).lower():
79
+ return col, series_map[col]
80
+ best = max(series_map, key=lambda c: series_map[c].notna().mean())
81
+ return best, series_map[best]
82
+
83
+
84
+ def extract_series(df, min_points=300):
85
+ date_col, dates = _find_date_col(df)
86
+ if date_col is None:
87
+ return None
88
+ target_col, values = _find_target_col(df, date_col)
89
+ if target_col is None:
90
+ return None
91
+ out = pd.DataFrame({"date": dates, "value": pd.to_numeric(values, errors="coerce")})
92
+ out = out[np.isfinite(out["value"])].dropna().sort_values("date")
93
+ if len(out) < min_points:
94
+ return None
95
+ out = out.set_index("date")
96
+ gaps = out.index.to_series().diff().dropna().dt.total_seconds()
97
+ if len(gaps) and gaps.median() < 6 * 3600:
98
+ out = out.resample("1D").mean().dropna()
99
+ out = out[~out.index.duplicated(keep="last")].sort_index()
100
+ out = out.reset_index()
101
+ out["date"] = out["date"].dt.strftime("%Y-%m-%d")
102
+ if len(out) < min_points or out["value"].std() <= 0:
103
+ return None
104
+ return out
105
+
106
+
107
+ def _hf_search_best(queries):
108
+ try:
109
+ from huggingface_hub import HfApi, hf_hub_download
110
+ except Exception as exc:
111
+ print(f"[data] huggingface_hub unavailable: {exc}")
112
+ return None
113
+ api = HfApi()
114
+ repo_ids = []
115
+ for query in queries:
116
+ try:
117
+ results = list(api.list_datasets(search=query, limit=8))
118
+ results.sort(key=lambda i: -(getattr(i, "downloads", 0) or 0))
119
+ for info in results:
120
+ if info.id not in repo_ids:
121
+ repo_ids.append(info.id)
122
+ except Exception as exc:
123
+ print(f"[data] search failed for {query!r}: {exc}")
124
+ candidates = []
125
+ for repo in repo_ids:
126
+ try:
127
+ files = api.list_repo_files(repo, repo_type="dataset")
128
+ except Exception:
129
+ continue
130
+ data_files = [f for f in files if f.endswith((".csv", ".parquet")) and "readme" not in f.lower()]
131
+ data_files.sort(key=lambda f: 0 if f.endswith(".csv") else 1)
132
+ for f in data_files[:3]:
133
+ try:
134
+ local = hf_hub_download(repo, f, repo_type="dataset")
135
+ df = pd.read_csv(local) if f.endswith(".csv") else pd.read_parquet(local)
136
+ series = extract_series(df)
137
+ if series is not None:
138
+ candidates.append((pd.Timestamp(series["date"].iloc[-1]), len(series),
139
+ f"huggingface:{repo}::{f}", series))
140
+ except Exception:
141
+ continue
142
+ if not candidates:
143
+ return None
144
+ candidates.sort(key=lambda c: (c[0], c[1]), reverse=True)
145
+ end_date, n_rows, source, series = candidates[0]
146
+ print(f"[data] best HF dataset: {source} (ends {end_date.date()}, {n_rows} pts; {len(candidates)} usable)")
147
+ return {"source": source, "series": series}
148
+
149
+
150
+ def _yf_fetch(symbol, start):
151
+ import yfinance as yf
152
+ df = yf.download(symbol, start=start, auto_adjust=True, progress=False)
153
+ if df is None or df.empty:
154
+ raise RuntimeError(f"no data for {symbol}")
155
+ if isinstance(df.columns, pd.MultiIndex):
156
+ df.columns = df.columns.get_level_values(0)
157
+ df = df.reset_index().rename(columns={"Date": "date", "Close": "close"})
158
+ return extract_series(df)
159
+
160
+
161
+ def fetch_finance():
162
+ candidates = []
163
+ hf = _hf_search_best(("bitcoin price", "cryptocurrency price", "stock price daily"))
164
+ if hf:
165
+ candidates.append(hf)
166
+ try:
167
+ series = _yf_fetch("BTC-USD", "2015-01-01")
168
+ if series is not None:
169
+ candidates.append({"source": "yfinance:BTC-USD", "series": series})
170
+ except Exception as exc:
171
+ print(f"[data] yfinance BTC-USD failed: {exc}")
172
+ return candidates
173
+
174
+
175
+ def fetch_ai():
176
+ candidates = []
177
+ try:
178
+ series = _yf_fetch("NVDA", "1999-01-01")
179
+ if series is not None:
180
+ candidates.append({"source": "yfinance:NVDA", "series": series})
181
+ except Exception as exc:
182
+ print(f"[data] yfinance NVDA failed: {exc}")
183
+ hf = _hf_search_best(("nvidia stock price", "ai stock price", "tech stock daily"))
184
+ if hf:
185
+ candidates.append(hf)
186
+ return candidates
187
+
188
+
189
+ def fetch_programming():
190
+ import urllib.request
191
+ today = pd.Timestamp.today().strftime("%Y-%m-%d")
192
+ url = f"https://api.npmjs.org/downloads/range/2015-01-01:{today}/react"
193
+ with urllib.request.urlopen(url, timeout=30) as resp:
194
+ payload = json.loads(resp.read().decode())
195
+ rows = payload.get("downloads", [])
196
+ series = pd.DataFrame(rows)
197
+ if series.empty:
198
+ return []
199
+ series["date"] = pd.to_datetime(series["day"], errors="coerce")
200
+ series["value"] = pd.to_numeric(series.get("downloads"), errors="coerce")
201
+ series = series[np.isfinite(series["value"])]
202
+ series = series.dropna(subset=["date"]).set_index("date")
203
+ series = series[~series.index.duplicated(keep="last")].sort_index()
204
+ series = series.resample("1D").asfreq().ffill()
205
+ series["value"] = series["value"].replace(0, np.nan).ffill().bfill()
206
+ out = series.reset_index()
207
+ out["date"] = out["date"].dt.strftime("%Y-%m-%d")
208
+ out = out[["date", "value"]]
209
+ if len(out) < 300 or out["value"].std() <= 0:
210
+ return []
211
+ return [{"source": "npm-api:react-daily-downloads", "series": out}]
212
+
213
+
214
+ def _clubelo_series(club, tries=3):
215
+ import io
216
+ import time
217
+ import urllib.request
218
+ url = f"http://api.clubelo.com/{club}"
219
+ req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0 (forecast-bot)"})
220
+ for attempt in range(tries):
221
+ try:
222
+ with urllib.request.urlopen(req, timeout=25) as resp:
223
+ text = resp.read().decode()
224
+ df = pd.read_csv(io.StringIO(text))
225
+ if "From" not in df.columns or "Elo" not in df.columns:
226
+ return None
227
+ s = pd.DataFrame({"date": pd.to_datetime(df["From"], errors="coerce"),
228
+ "value": pd.to_numeric(df["Elo"], errors="coerce")})
229
+ s = s.dropna().sort_values("date")
230
+ s = s[s["date"] <= pd.Timestamp.today().normalize()]
231
+ if len(s) < 300:
232
+ return None
233
+ s = s.set_index("date")
234
+ s = s[~s.index.duplicated(keep="last")]
235
+ s = s.resample("1D").ffill().ffill().dropna()
236
+ out = s.reset_index()
237
+ out["date"] = out["date"].dt.strftime("%Y-%m-%d")
238
+ return out[["date", "value"]] if len(out) >= 300 else None
239
+ except Exception as exc:
240
+ print(f"[data] ClubElo {club} attempt {attempt + 1} failed: {exc}")
241
+ time.sleep(1.5)
242
+ return None
243
+
244
+
245
+ def _atp_world1_elo():
246
+ try:
247
+ from huggingface_hub import HfApi, hf_hub_download
248
+ except Exception:
249
+ return None
250
+ try:
251
+ api = HfApi()
252
+ files = api.list_repo_files("davidtadediji/tennis-atp", repo_type="dataset")
253
+ except Exception as exc:
254
+ print(f"[data] ATP repo listing failed: {exc}")
255
+ return None
256
+ import re as _re
257
+ year_files = {}
258
+ for f in files:
259
+ m = _re.match(r"atp_matches_(\d{4})\.csv$", f)
260
+ if m:
261
+ year_files[int(m.group(1))] = f
262
+ if not year_files:
263
+ return None
264
+ years = sorted(y for y in year_files if y >= 1995)
265
+ records = []
266
+ for y in years:
267
+ try:
268
+ local = hf_hub_download("davidtadediji/tennis-atp", year_files[y], repo_type="dataset")
269
+ df = pd.read_csv(local, usecols=["tourney_date", "winner_id", "loser_id"], low_memory=False)
270
+ records.append(df)
271
+ print(f"[data] ATP {y}: {len(df)} matches")
272
+ except Exception:
273
+ print(f"[data] ATP {y}: download/parse failed")
274
+ continue
275
+ if not records:
276
+ return None
277
+ m = pd.concat(records, ignore_index=True)
278
+ m["date"] = pd.to_datetime(m["tourney_date"].astype(str), format="%Y%m%d", errors="coerce")
279
+ m["winner_id"] = pd.to_numeric(m["winner_id"], errors="coerce")
280
+ m["loser_id"] = pd.to_numeric(m["loser_id"], errors="coerce")
281
+ m = m.dropna(subset=["date", "winner_id", "loser_id"]).astype({"winner_id": "int64", "loser_id": "int64"})
282
+ m = m.sort_values("date")
283
+ print(f"[data] ATP: {len(m)} valid matches across years {years[0]}..{years[-1]}")
284
+
285
+ K, ratings = 32.0, {}
286
+ top_daily = []
287
+ for d, grp in m.groupby("date"):
288
+ for _, row in grp.iterrows():
289
+ w, l = row["winner_id"], row["loser_id"]
290
+ if pd.isna(w) or pd.isna(l) or w == l:
291
+ continue
292
+ rw, rl = ratings.get(w, 1500.0), ratings.get(l, 1500.0)
293
+ ew = 1.0 / (1.0 + 10 ** ((rl - rw) / 400.0))
294
+ ratings[w] = rw + K * (1 - ew)
295
+ ratings[l] = rl + K * (0 - (1 - ew))
296
+ top_daily.append((d, max(ratings.values())))
297
+ if len(top_daily) < 300:
298
+ return None
299
+ s = pd.DataFrame(top_daily, columns=["date", "value"]).set_index("date")
300
+ s = s.resample("1D").ffill().ffill().dropna()
301
+ s = s[s.index <= pd.Timestamp.today().normalize()]
302
+ out = s.reset_index()
303
+ out["date"] = out["date"].dt.strftime("%Y-%m-%d")
304
+ if len(out) < 300 or out["value"].std() <= 0:
305
+ return None
306
+ print(f"[data] sports source: ATP world #1 Elo ({len(out)} daily points, "
307
+ f"{years[0]}-{years[-1]})")
308
+ return [{"source": "huggingface:davidtadediji/tennis-atp->world#1-elo", "series": out[["date", "value"]]}]
309
+
310
+
311
+ def fetch_sports():
312
+ candidates = _atp_world1_elo()
313
+ if candidates:
314
+ return candidates
315
+ club = None
316
+ for c in ["Liverpool", "RealMadrid", "BayernMunich", "Chelsea"]:
317
+ club = _clubelo_series(c, tries=1)
318
+ if club is not None:
319
+ return [{"source": f"clubelo:{c}", "series": club}]
320
+ hf = _hf_search_best(("elo rating", "chess rating history"))
321
+ if hf:
322
+ return [hf]
323
+ return []
324
+
325
+
326
+ def weather_slug() -> str:
327
+ return re.sub(r"[^a-z0-9]+", "_", _WEATHER_NAME.lower()).strip("_")
328
+
329
+
330
+ def set_weather_location(name: str, lat, lon):
331
+ global _WEATHER_NAME, _WEATHER_LAT, _WEATHER_LON, WEATHER_NAME, WEATHER_LAT, WEATHER_LON
332
+ _WEATHER_NAME = WEATHER_NAME = name
333
+ _WEATHER_LAT = WEATHER_LAT = str(lat)
334
+ _WEATHER_LON = WEATHER_LON = str(lon)
335
+ TOPICS["weather"]["asset"] = f"{name} temperature"
336
+ TOPICS["weather"]["desc"] = f"Daily mean temperature for {name} (Open-Meteo archive, stored in Kelvin)"
337
+
338
+
339
+ def artifact_suffix(topic: str) -> str:
340
+ """Per-domain variant suffix so different instances (e.g. cities) of the
341
+ same topic keep separate artifacts."""
342
+ if topic == "weather":
343
+ return f"_{weather_slug()}"
344
+ return ""
345
+
346
+
347
+ def fetch_weather():
348
+ import json as _json
349
+ import urllib.request
350
+ end = pd.Timestamp.today().strftime("%Y-%m-%d")
351
+ url = ("https://archive-api.open-meteo.com/v1/archive"
352
+ f"?latitude={_WEATHER_LAT}&longitude={_WEATHER_LON}"
353
+ f"&start_date=2015-01-01&end_date={end}&daily=temperature_2m_mean&timezone=auto")
354
+ req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0 (forecast-bot)"})
355
+ try:
356
+ with urllib.request.urlopen(req, timeout=45) as resp:
357
+ payload = _json.loads(resp.read().decode())
358
+ daily = payload.get("daily", {})
359
+ dates, temps = daily.get("time", []), daily.get("temperature_2m_mean", [])
360
+ s = pd.DataFrame({"date": pd.to_datetime(dates, errors="coerce"),
361
+ "value": pd.to_numeric(temps, errors="coerce")})
362
+ s = s.dropna().sort_values("date")
363
+ s = s[s["date"] <= pd.Timestamp.today().normalize()]
364
+ if len(s) < 300:
365
+ return []
366
+ s = s.set_index("date")
367
+ s = s[~s.index.duplicated(keep="last")]
368
+ s = s.resample("1D").mean().ffill().dropna()
369
+ s["value"] = s["value"] + 273.15
370
+ out = s.reset_index()
371
+ out["date"] = out["date"].dt.strftime("%Y-%m-%d")
372
+ out = out[["date", "value"]]
373
+ if len(out) >= 300 and out["value"].std() > 0:
374
+ print(f"[data] weather source: Open-Meteo {_WEATHER_NAME} ({len(out)} daily points, Kelvin)")
375
+ return [{"source": f"open-meteo:{_WEATHER_NAME}", "series": out}]
376
+ except Exception as exc:
377
+ print(f"[data] Open-Meteo failed: {exc}")
378
+ return []
379
+
380
+
381
+ def fetch_economy():
382
+ try:
383
+ series = _yf_fetch("^GSPC", "2000-01-01")
384
+ if series is not None:
385
+ return [{"source": "yfinance:^GSPC", "series": series}]
386
+ except Exception as exc:
387
+ print(f"[data] yfinance S&P500 failed: {exc}")
388
+ return []
389
+
390
+
391
+ def fetch_energy():
392
+ try:
393
+ series = _yf_fetch("CL=F", "2000-01-01")
394
+ if series is not None:
395
+ return [{"source": "yfinance:CL=F(WTI)", "series": series}]
396
+ except Exception as exc:
397
+ print(f"[data] yfinance WTI failed: {exc}")
398
+ return []
399
+
400
+
401
+ FETCHERS = {
402
+ "ai": fetch_ai,
403
+ "programming": fetch_programming,
404
+ "finance": fetch_finance,
405
+ "sports": fetch_sports,
406
+ "weather": fetch_weather,
407
+ "economy": fetch_economy,
408
+ "energy": fetch_energy,
409
+ }
410
+
411
+
412
+ def generate_synthetic(topic="finance", n=1500):
413
+ import zlib
414
+ seed = zlib.crc32(topic.encode()) % (2**32)
415
+ rng = np.random.default_rng(seed)
416
+ t = np.arange(n)
417
+ trend = 100 * np.exp(0.0005 * t)
418
+ season = 8 * np.sin(2 * np.pi * t / 365.25) + 3 * np.sin(2 * np.pi * t / 30.4)
419
+ ar = np.zeros(n)
420
+ for i in range(1, n):
421
+ ar[i] = 0.94 * ar[i - 1] + rng.normal(0, 1.5)
422
+ value = trend + season + ar + rng.normal(0, 0.5, n)
423
+ dates = pd.date_range("2020-01-01", periods=n, freq="D")
424
+ return {
425
+ "source": f"synthetic:{topic}",
426
+ "series": pd.DataFrame({"date": dates.strftime("%Y-%m-%d"), "value": value}),
427
+ }
428
+
429
+
430
+ def load(topic="finance", refresh=True):
431
+ if topic not in TOPICS:
432
+ raise SystemExit(f"Unknown topic {topic!r}. Choose from: {', '.join(TOPICS)}")
433
+ DATA_DIR.mkdir(exist_ok=True)
434
+ cache_csv = DATA_DIR / f"{topic}.csv"
435
+ cache_meta = DATA_DIR / f"{topic}.meta.json"
436
+ if not refresh and cache_csv.exists():
437
+ series = pd.read_csv(cache_csv)
438
+ meta = json.loads(cache_meta.read_text()) if cache_meta.exists() else {}
439
+ return series, meta.get("source", "cache")
440
+ candidates = FETCHERS[topic]()
441
+ if not candidates:
442
+ print(f"[data] online sources failed for '{topic}', generating synthetic dataset...")
443
+ result = generate_synthetic(topic)
444
+ else:
445
+ result = max(candidates, key=lambda r: (pd.Timestamp(r["series"]["date"].iloc[-1]), len(r["series"])))
446
+ print(f"[data] selected '{result['source']}' (ends {result['series']['date'].iloc[-1]}, "
447
+ f"{len(result['series'])} pts)")
448
+ series = result["series"].reset_index(drop=True)
449
+ series.to_csv(cache_csv, index=False)
450
+ gaps = pd.to_datetime(series["date"]).diff().dropna().dt.days
451
+ meta = {
452
+ "topic": topic,
453
+ "source": result["source"],
454
+ "rows": len(series),
455
+ "freq_days": float(gaps.median()) if len(gaps) else 1.0,
456
+ "start": str(series["date"].iloc[0]),
457
+ "end": str(series["date"].iloc[-1]),
458
+ }
459
+ cache_meta.write_text(json.dumps(meta, indent=2))
460
+ return series, result["source"]
461
+
462
+
463
+ def main():
464
+ parser = argparse.ArgumentParser()
465
+ parser.add_argument("--topic", default="finance", choices=list(TOPICS))
466
+ parser.add_argument("--refresh", action="store_true")
467
+ args = parser.parse_args()
468
+ series, source = load(topic=args.topic, refresh=args.refresh)
469
+ print(f"[data] topic: {TOPICS[args.topic]['label']} ({TOPICS[args.topic]['desc']})")
470
+ print(f"[data] source: {source}")
471
+ print(f"[data] points: {len(series)} range: {series['date'].iloc[0]} .. {series['date'].iloc[-1]}")
472
+
473
+
474
+ if __name__ == "__main__":
475
+ main()
data/ai.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/ai.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "ai",
3
+ "source": "yfinance:NVDA",
4
+ "rows": 6933,
5
+ "freq_days": 1.0,
6
+ "start": "1999-01-22",
7
+ "end": "2026-08-14"
8
+ }
data/economy.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/economy.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "economy",
3
+ "source": "yfinance:^GSPC",
4
+ "rows": 6694,
5
+ "freq_days": 1.0,
6
+ "start": "2000-01-03",
7
+ "end": "2026-08-14"
8
+ }
data/energy.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/energy.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "energy",
3
+ "source": "yfinance:CL=F(WTI)",
4
+ "rows": 6524,
5
+ "freq_days": 1.0,
6
+ "start": "2000-08-23",
7
+ "end": "2026-08-18"
8
+ }
data/finance.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/finance.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "finance",
3
+ "source": "yfinance:BTC-USD",
4
+ "rows": 4247,
5
+ "freq_days": 1.0,
6
+ "start": "2015-01-01",
7
+ "end": "2026-08-17"
8
+ }
data/programming.csv ADDED
@@ -0,0 +1,548 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ date,value
2
+ 2025-02-17,5122863.0
3
+ 2025-02-18,5723233.0
4
+ 2025-02-19,5827613.0
5
+ 2025-02-20,5787132.0
6
+ 2025-02-21,5036029.0
7
+ 2025-02-22,1639322.0
8
+ 2025-02-23,1575462.0
9
+ 2025-02-24,5563188.0
10
+ 2025-02-25,6040485.0
11
+ 2025-02-26,5951804.0
12
+ 2025-02-27,5968319.0
13
+ 2025-02-28,5133960.0
14
+ 2025-03-01,1796871.0
15
+ 2025-03-02,1760465.0
16
+ 2025-03-03,5672966.0
17
+ 2025-03-04,5785770.0
18
+ 2025-03-05,5844434.0
19
+ 2025-03-06,5851743.0
20
+ 2025-03-07,5227560.0
21
+ 2025-03-08,1921002.0
22
+ 2025-03-09,1725033.0
23
+ 2025-03-10,5917383.0
24
+ 2025-03-11,6269431.0
25
+ 2025-03-12,6340771.0
26
+ 2025-03-13,6111425.0
27
+ 2025-03-14,5144759.0
28
+ 2025-03-15,1815727.0
29
+ 2025-03-16,1731580.0
30
+ 2025-03-17,6015919.0
31
+ 2025-03-18,6279906.0
32
+ 2025-03-19,6228094.0
33
+ 2025-03-20,6055479.0
34
+ 2025-03-21,5526374.0
35
+ 2025-03-22,1909331.0
36
+ 2025-03-23,1788948.0
37
+ 2025-03-24,6130551.0
38
+ 2025-03-25,6395177.0
39
+ 2025-03-26,6203063.0
40
+ 2025-03-27,6210284.0
41
+ 2025-03-28,5459772.0
42
+ 2025-03-29,1900124.0
43
+ 2025-03-30,1803578.0
44
+ 2025-03-31,5954186.0
45
+ 2025-04-01,6661203.0
46
+ 2025-04-02,6576843.0
47
+ 2025-04-03,6888041.0
48
+ 2025-04-04,6670857.0
49
+ 2025-04-05,3071692.0
50
+ 2025-04-06,2919997.0
51
+ 2025-04-07,6712733.0
52
+ 2025-04-08,7212617.0
53
+ 2025-04-09,6750589.0
54
+ 2025-04-10,6579603.0
55
+ 2025-04-11,6203719.0
56
+ 2025-04-12,2453142.0
57
+ 2025-04-13,2098192.0
58
+ 2025-04-14,6567023.0
59
+ 2025-04-15,6749240.0
60
+ 2025-04-16,6874334.0
61
+ 2025-04-17,6521064.0
62
+ 2025-04-18,4355495.0
63
+ 2025-04-19,2219287.0
64
+ 2025-04-20,2190840.0
65
+ 2025-04-21,4966891.0
66
+ 2025-04-22,6906119.0
67
+ 2025-04-23,7058101.0
68
+ 2025-04-24,6836432.0
69
+ 2025-04-25,6015675.0
70
+ 2025-04-26,2157259.0
71
+ 2025-04-27,2275703.0
72
+ 2025-04-28,6425006.0
73
+ 2025-04-29,6733835.0
74
+ 2025-04-30,6651039.0
75
+ 2025-05-01,4753065.0
76
+ 2025-05-02,5271981.0
77
+ 2025-05-03,2291973.0
78
+ 2025-05-04,2402833.0
79
+ 2025-05-05,6291277.0
80
+ 2025-05-06,6992360.0
81
+ 2025-05-07,6928903.0
82
+ 2025-05-08,6566142.0
83
+ 2025-05-09,5914457.0
84
+ 2025-05-10,2190796.0
85
+ 2025-05-11,2277870.0
86
+ 2025-05-12,6741178.0
87
+ 2025-05-13,7192041.0
88
+ 2025-05-14,6994702.0
89
+ 2025-05-15,6705175.0
90
+ 2025-05-16,5998028.0
91
+ 2025-05-17,2140875.0
92
+ 2025-05-18,2410617.0
93
+ 2025-05-19,6975059.0
94
+ 2025-05-20,7315953.0
95
+ 2025-05-21,7457177.0
96
+ 2025-05-22,7130478.0
97
+ 2025-05-23,6282131.0
98
+ 2025-05-24,2543909.0
99
+ 2025-05-25,2575867.0
100
+ 2025-05-26,5777006.0
101
+ 2025-05-27,7328871.0
102
+ 2025-05-28,6910371.0
103
+ 2025-05-29,6200856.0
104
+ 2025-05-30,5962090.0
105
+ 2025-05-31,2234653.0
106
+ 2025-06-01,2493303.0
107
+ 2025-06-02,7213707.0
108
+ 2025-06-03,7394293.0
109
+ 2025-06-04,7406737.0
110
+ 2025-06-05,7473882.0
111
+ 2025-06-06,6272731.0
112
+ 2025-06-07,2274985.0
113
+ 2025-06-08,2233446.0
114
+ 2025-06-09,6181351.0
115
+ 2025-06-10,7123647.0
116
+ 2025-06-11,7654167.0
117
+ 2025-06-12,7098716.0
118
+ 2025-06-13,7076444.0
119
+ 2025-06-14,3350286.0
120
+ 2025-06-15,3354579.0
121
+ 2025-06-16,7812998.0
122
+ 2025-06-17,8188061.0
123
+ 2025-06-18,8892726.0
124
+ 2025-06-19,6626144.0
125
+ 2025-06-20,7157040.0
126
+ 2025-06-21,4059454.0
127
+ 2025-06-22,3474768.0
128
+ 2025-06-23,8076868.0
129
+ 2025-06-24,8858642.0
130
+ 2025-06-25,8441191.0
131
+ 2025-06-26,7979028.0
132
+ 2025-06-27,7203347.0
133
+ 2025-06-28,3099992.0
134
+ 2025-06-29,3355923.0
135
+ 2025-06-30,7900831.0
136
+ 2025-07-01,7699048.0
137
+ 2025-07-02,7331695.0
138
+ 2025-07-03,6882827.0
139
+ 2025-07-04,5412939.0
140
+ 2025-07-05,2536907.0
141
+ 2025-07-06,2754888.0
142
+ 2025-07-07,7472686.0
143
+ 2025-07-08,7659562.0
144
+ 2025-07-09,7060450.0
145
+ 2025-07-10,6846160.0
146
+ 2025-07-11,6487668.0
147
+ 2025-07-12,2703129.0
148
+ 2025-07-13,3412034.0
149
+ 2025-07-14,7251509.0
150
+ 2025-07-15,6966115.0
151
+ 2025-07-16,7150744.0
152
+ 2025-07-17,7552013.0
153
+ 2025-07-18,6213895.0
154
+ 2025-07-19,2370471.0
155
+ 2025-07-20,2352601.0
156
+ 2025-07-21,6716667.0
157
+ 2025-07-22,7224230.0
158
+ 2025-07-23,7381941.0
159
+ 2025-07-24,7056637.0
160
+ 2025-07-25,6554474.0
161
+ 2025-07-26,2729232.0
162
+ 2025-07-27,2496018.0
163
+ 2025-07-28,8060663.0
164
+ 2025-07-29,8300974.0
165
+ 2025-07-30,8507849.0
166
+ 2025-07-31,8767706.0
167
+ 2025-08-01,7324617.0
168
+ 2025-08-02,2655386.0
169
+ 2025-08-03,2599528.0
170
+ 2025-08-04,7102206.0
171
+ 2025-08-05,7810323.0
172
+ 2025-08-06,7633348.0
173
+ 2025-08-07,7541397.0
174
+ 2025-08-08,6699454.0
175
+ 2025-08-09,2430634.0
176
+ 2025-08-10,2613698.0
177
+ 2025-08-11,7076504.0
178
+ 2025-08-12,7453458.0
179
+ 2025-08-13,7511373.0
180
+ 2025-08-14,7542787.0
181
+ 2025-08-15,6026612.0
182
+ 2025-08-16,2542689.0
183
+ 2025-08-17,2637274.0
184
+ 2025-08-18,7329811.0
185
+ 2025-08-19,7780274.0
186
+ 2025-08-20,7827592.0
187
+ 2025-08-21,7819748.0
188
+ 2025-08-22,7202314.0
189
+ 2025-08-23,2949721.0
190
+ 2025-08-24,2773012.0
191
+ 2025-08-25,7429505.0
192
+ 2025-08-26,8137538.0
193
+ 2025-08-27,7907820.0
194
+ 2025-08-28,8355298.0
195
+ 2025-08-29,9086768.0
196
+ 2025-08-30,5783505.0
197
+ 2025-08-31,5817521.0
198
+ 2025-09-01,6216186.0
199
+ 2025-09-02,7599536.0
200
+ 2025-09-03,7851907.0
201
+ 2025-09-04,8232845.0
202
+ 2025-09-05,7037271.0
203
+ 2025-09-06,2745045.0
204
+ 2025-09-07,2873510.0
205
+ 2025-09-08,7949346.0
206
+ 2025-09-09,8227013.0
207
+ 2025-09-10,8034811.0
208
+ 2025-09-11,7985133.0
209
+ 2025-09-12,7031246.0
210
+ 2025-09-13,2720266.0
211
+ 2025-09-14,2681556.0
212
+ 2025-09-15,7808362.0
213
+ 2025-09-16,8050414.0
214
+ 2025-09-17,8242375.0
215
+ 2025-09-18,7925749.0
216
+ 2025-09-19,7139170.0
217
+ 2025-09-20,2624398.0
218
+ 2025-09-21,2601704.0
219
+ 2025-09-22,7570796.0
220
+ 2025-09-23,7753950.0
221
+ 2025-09-24,7995481.0
222
+ 2025-09-25,8117734.0
223
+ 2025-09-26,7273004.0
224
+ 2025-09-27,2785336.0
225
+ 2025-09-28,2840117.0
226
+ 2025-09-29,8031686.0
227
+ 2025-09-30,8380861.0
228
+ 2025-10-01,8319264.0
229
+ 2025-10-02,7860024.0
230
+ 2025-10-03,7071082.0
231
+ 2025-10-04,2859566.0
232
+ 2025-10-05,2747277.0
233
+ 2025-10-06,7959809.0
234
+ 2025-10-07,8451066.0
235
+ 2025-10-08,8279717.0
236
+ 2025-10-09,8255370.0
237
+ 2025-10-10,7601043.0
238
+ 2025-10-11,2901823.0
239
+ 2025-10-12,2815917.0
240
+ 2025-10-13,7407202.0
241
+ 2025-10-14,8297850.0
242
+ 2025-10-15,8455422.0
243
+ 2025-10-16,8376709.0
244
+ 2025-10-17,7905867.0
245
+ 2025-10-18,7905867.0
246
+ 2025-10-19,3381955.0
247
+ 2025-10-20,7753268.0
248
+ 2025-10-21,7753268.0
249
+ 2025-10-22,8727972.0
250
+ 2025-10-23,8727406.0
251
+ 2025-10-24,9152760.0
252
+ 2025-10-25,4017426.0
253
+ 2025-10-26,3310690.0
254
+ 2025-10-27,8725497.0
255
+ 2025-10-28,8813607.0
256
+ 2025-10-29,8652906.0
257
+ 2025-10-30,9958714.0
258
+ 2025-10-31,7583571.0
259
+ 2025-11-01,3156264.0
260
+ 2025-11-02,3109120.0
261
+ 2025-11-03,8234310.0
262
+ 2025-11-04,8872492.0
263
+ 2025-11-05,8933715.0
264
+ 2025-11-06,9048161.0
265
+ 2025-11-07,8004608.0
266
+ 2025-11-08,3439566.0
267
+ 2025-11-09,3541805.0
268
+ 2025-11-10,8635729.0
269
+ 2025-11-11,8324250.0
270
+ 2025-11-12,9096063.0
271
+ 2025-11-13,9010510.0
272
+ 2025-11-14,8286903.0
273
+ 2025-11-15,3342871.0
274
+ 2025-11-16,3501054.0
275
+ 2025-11-17,8776298.0
276
+ 2025-11-18,9445039.0
277
+ 2025-11-19,9881023.0
278
+ 2025-11-20,9408539.0
279
+ 2025-11-21,8396336.0
280
+ 2025-11-22,3733751.0
281
+ 2025-11-23,3855431.0
282
+ 2025-11-24,9128701.0
283
+ 2025-11-25,9999736.0
284
+ 2025-11-26,9956991.0
285
+ 2025-11-27,8944128.0
286
+ 2025-11-28,7606329.0
287
+ 2025-11-29,4400308.0
288
+ 2025-11-30,3950705.0
289
+ 2025-12-01,9737074.0
290
+ 2025-12-02,10790084.0
291
+ 2025-12-03,10383416.0
292
+ 2025-12-04,10826479.0
293
+ 2025-12-05,9233325.0
294
+ 2025-12-06,4617855.0
295
+ 2025-12-07,5631112.0
296
+ 2025-12-08,10796822.0
297
+ 2025-12-09,11294552.0
298
+ 2025-12-10,10553249.0
299
+ 2025-12-11,10099264.0
300
+ 2025-12-12,9461298.0
301
+ 2025-12-13,4476344.0
302
+ 2025-12-14,4484467.0
303
+ 2025-12-15,9954265.0
304
+ 2025-12-16,10678981.0
305
+ 2025-12-17,10291324.0
306
+ 2025-12-18,9869981.0
307
+ 2025-12-19,8484631.0
308
+ 2025-12-20,4027950.0
309
+ 2025-12-21,3863518.0
310
+ 2025-12-22,7832259.0
311
+ 2025-12-23,7405313.0
312
+ 2025-12-24,5627595.0
313
+ 2025-12-25,3835367.0
314
+ 2025-12-26,4336405.0
315
+ 2025-12-27,4102233.0
316
+ 2025-12-28,6002253.0
317
+ 2025-12-29,6183123.0
318
+ 2025-12-30,6100306.0
319
+ 2025-12-31,5148497.0
320
+ 2026-01-01,3745740.0
321
+ 2026-01-02,5323273.0
322
+ 2026-01-03,3533028.0
323
+ 2026-01-04,3917559.0
324
+ 2026-01-05,8666257.0
325
+ 2026-01-06,8882667.0
326
+ 2026-01-07,9417727.0
327
+ 2026-01-08,9686985.0
328
+ 2026-01-09,9123144.0
329
+ 2026-01-10,4118742.0
330
+ 2026-01-11,4107130.0
331
+ 2026-01-12,9646307.0
332
+ 2026-01-13,10391770.0
333
+ 2026-01-14,12308588.0
334
+ 2026-01-15,13372696.0
335
+ 2026-01-16,11619102.0
336
+ 2026-01-17,6615301.0
337
+ 2026-01-18,6420272.0
338
+ 2026-01-19,10950431.0
339
+ 2026-01-20,10385662.0
340
+ 2026-01-21,10953716.0
341
+ 2026-01-22,13677331.0
342
+ 2026-01-23,12123483.0
343
+ 2026-01-24,6279781.0
344
+ 2026-01-25,4425172.0
345
+ 2026-01-26,10593385.0
346
+ 2026-01-27,12293999.0
347
+ 2026-01-28,12581993.0
348
+ 2026-01-29,12297553.0
349
+ 2026-01-30,11281647.0
350
+ 2026-01-31,6705711.0
351
+ 2026-02-01,6791370.0
352
+ 2026-02-02,12651309.0
353
+ 2026-02-03,15027482.0
354
+ 2026-02-04,15290699.0
355
+ 2026-02-05,15471145.0
356
+ 2026-02-06,13766106.0
357
+ 2026-02-07,7479027.0
358
+ 2026-02-08,5910028.0
359
+ 2026-02-09,11795177.0
360
+ 2026-02-10,12671902.0
361
+ 2026-02-11,12214739.0
362
+ 2026-02-12,12027144.0
363
+ 2026-02-13,10942091.0
364
+ 2026-02-14,6703068.0
365
+ 2026-02-15,8035910.0
366
+ 2026-02-16,10343044.0
367
+ 2026-02-17,13219840.0
368
+ 2026-02-18,14257485.0
369
+ 2026-02-19,15881302.0
370
+ 2026-02-20,16453405.0
371
+ 2026-02-21,9517248.0
372
+ 2026-02-22,9818293.0
373
+ 2026-02-23,15903144.0
374
+ 2026-02-24,14584573.0
375
+ 2026-02-25,14099271.0
376
+ 2026-02-26,13491902.0
377
+ 2026-02-27,12164291.0
378
+ 2026-02-28,6829622.0
379
+ 2026-03-01,9014571.0
380
+ 2026-03-02,15154755.0
381
+ 2026-03-03,13812079.0
382
+ 2026-03-04,15558446.0
383
+ 2026-03-05,14240355.0
384
+ 2026-03-06,13554436.0
385
+ 2026-03-07,9026047.0
386
+ 2026-03-08,8398710.0
387
+ 2026-03-09,15164077.0
388
+ 2026-03-10,16252647.0
389
+ 2026-03-11,15374518.0
390
+ 2026-03-12,16057587.0
391
+ 2026-03-13,15243171.0
392
+ 2026-03-14,9237662.0
393
+ 2026-03-15,7813148.0
394
+ 2026-03-16,17281411.0
395
+ 2026-03-17,18185846.0
396
+ 2026-03-18,18326323.0
397
+ 2026-03-19,16674681.0
398
+ 2026-03-20,14885848.0
399
+ 2026-03-21,10066731.0
400
+ 2026-03-22,9621109.0
401
+ 2026-03-23,17035535.0
402
+ 2026-03-24,19126493.0
403
+ 2026-03-25,21034761.0
404
+ 2026-03-26,20983183.0
405
+ 2026-03-27,19271654.0
406
+ 2026-03-28,12477272.0
407
+ 2026-03-29,11936061.0
408
+ 2026-03-30,21773918.0
409
+ 2026-03-31,21548205.0
410
+ 2026-04-01,21248268.0
411
+ 2026-04-02,18379222.0
412
+ 2026-04-03,14276457.0
413
+ 2026-04-04,9591739.0
414
+ 2026-04-05,9081539.0
415
+ 2026-04-06,14788196.0
416
+ 2026-04-07,18790739.0
417
+ 2026-04-08,18819219.0
418
+ 2026-04-09,19529827.0
419
+ 2026-04-10,18476344.0
420
+ 2026-04-11,11249697.0
421
+ 2026-04-12,10819924.0
422
+ 2026-04-13,19230270.0
423
+ 2026-04-14,20922821.0
424
+ 2026-04-15,21066037.0
425
+ 2026-04-16,20601999.0
426
+ 2026-04-17,19411762.0
427
+ 2026-04-18,11613733.0
428
+ 2026-04-19,12341280.0
429
+ 2026-04-20,20855871.0
430
+ 2026-04-21,22346566.0
431
+ 2026-04-22,22694683.0
432
+ 2026-04-23,21791368.0
433
+ 2026-04-24,19408047.0
434
+ 2026-04-25,11369236.0
435
+ 2026-04-26,11743350.0
436
+ 2026-04-27,22110773.0
437
+ 2026-04-28,21906188.0
438
+ 2026-04-29,21688858.0
439
+ 2026-04-30,20123015.0
440
+ 2026-05-01,16230866.0
441
+ 2026-05-02,10806131.0
442
+ 2026-05-03,10867887.0
443
+ 2026-05-04,19462852.0
444
+ 2026-05-05,20987259.0
445
+ 2026-05-06,21251702.0
446
+ 2026-05-07,22011461.0
447
+ 2026-05-08,19818763.0
448
+ 2026-05-09,11751379.0
449
+ 2026-05-10,11210577.0
450
+ 2026-05-11,21870894.0
451
+ 2026-05-12,23284098.0
452
+ 2026-05-13,23002779.0
453
+ 2026-05-14,21557144.0
454
+ 2026-05-15,20356013.0
455
+ 2026-05-16,11975869.0
456
+ 2026-05-17,12071258.0
457
+ 2026-05-18,22061163.0
458
+ 2026-05-19,23111057.0
459
+ 2026-05-20,23035583.0
460
+ 2026-05-21,22527385.0
461
+ 2026-05-22,20513357.0
462
+ 2026-05-23,11751361.0
463
+ 2026-05-24,12354157.0
464
+ 2026-05-25,18363758.0
465
+ 2026-05-26,21548621.0
466
+ 2026-05-27,22256826.0
467
+ 2026-05-28,22557007.0
468
+ 2026-05-29,20251890.0
469
+ 2026-05-30,11664939.0
470
+ 2026-05-31,11524819.0
471
+ 2026-06-01,22288052.0
472
+ 2026-06-02,23866895.0
473
+ 2026-06-03,23866895.0
474
+ 2026-06-04,22888438.0
475
+ 2026-06-05,20895821.0
476
+ 2026-06-06,11719888.0
477
+ 2026-06-07,11742947.0
478
+ 2026-06-08,22158642.0
479
+ 2026-06-09,23410635.0
480
+ 2026-06-10,23937751.0
481
+ 2026-06-11,24528694.0
482
+ 2026-06-12,23419933.0
483
+ 2026-06-13,12928350.0
484
+ 2026-06-14,12156421.0
485
+ 2026-06-15,23213490.0
486
+ 2026-06-16,25978689.0
487
+ 2026-06-17,25461408.0
488
+ 2026-06-18,24329678.0
489
+ 2026-06-19,21580543.0
490
+ 2026-06-20,14499232.0
491
+ 2026-06-21,14869321.0
492
+ 2026-06-22,25186034.0
493
+ 2026-06-23,24376636.0
494
+ 2026-06-24,23812876.0
495
+ 2026-06-25,23542400.0
496
+ 2026-06-26,21890754.0
497
+ 2026-06-27,13852478.0
498
+ 2026-06-28,13586133.0
499
+ 2026-06-29,23709054.0
500
+ 2026-06-30,24084253.0
501
+ 2026-07-01,24059662.0
502
+ 2026-07-02,24255468.0
503
+ 2026-07-03,20313903.0
504
+ 2026-07-04,12649383.0
505
+ 2026-07-05,12488789.0
506
+ 2026-07-06,23666271.0
507
+ 2026-07-07,26440979.0
508
+ 2026-07-08,26299185.0
509
+ 2026-07-09,25414879.0
510
+ 2026-07-10,23799330.0
511
+ 2026-07-11,15207476.0
512
+ 2026-07-12,15207476.0
513
+ 2026-07-13,26705305.0
514
+ 2026-07-14,27460609.0
515
+ 2026-07-15,25957585.0
516
+ 2026-07-16,26232443.0
517
+ 2026-07-17,24016920.0
518
+ 2026-07-18,15711694.0
519
+ 2026-07-19,15139841.0
520
+ 2026-07-20,25608881.0
521
+ 2026-07-21,27007164.0
522
+ 2026-07-22,27571725.0
523
+ 2026-07-23,26888915.0
524
+ 2026-07-24,24781398.0
525
+ 2026-07-25,14627382.0
526
+ 2026-07-26,14866416.0
527
+ 2026-07-27,26770758.0
528
+ 2026-07-28,27755816.0
529
+ 2026-07-29,26997003.0
530
+ 2026-07-30,26772295.0
531
+ 2026-07-31,23864777.0
532
+ 2026-08-01,15321268.0
533
+ 2026-08-02,15105783.0
534
+ 2026-08-03,26850156.0
535
+ 2026-08-04,28227505.0
536
+ 2026-08-05,27778101.0
537
+ 2026-08-06,23966816.0
538
+ 2026-08-07,25765917.0
539
+ 2026-08-08,15082055.0
540
+ 2026-08-09,15412640.0
541
+ 2026-08-10,27438463.0
542
+ 2026-08-11,27438463.0
543
+ 2026-08-12,28221368.0
544
+ 2026-08-13,28442032.0
545
+ 2026-08-14,28442032.0
546
+ 2026-08-15,16059353.0
547
+ 2026-08-16,15849863.0
548
+ 2026-08-17,15849863.0
data/programming.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "programming",
3
+ "source": "npm-api:react-daily-downloads",
4
+ "rows": 547,
5
+ "freq_days": 1.0,
6
+ "start": "2025-02-17",
7
+ "end": "2026-08-17"
8
+ }
data/sports.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/sports.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "sports",
3
+ "source": "huggingface:davidtadediji/tennis-atp->world#1-elo",
4
+ "rows": 10944,
5
+ "freq_days": 1.0,
6
+ "start": "1995-01-02",
7
+ "end": "2024-12-18"
8
+ }
data/weather.csv ADDED
The diff for this file is too large to render. See raw diff
 
data/weather.meta.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": "weather",
3
+ "source": "open-meteo:Chennai, India",
4
+ "rows": 4248,
5
+ "freq_days": 1.0,
6
+ "start": "2015-01-01",
7
+ "end": "2026-08-18"
8
+ }
debug_logger.py ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Debug logging. Internal logs NEVER reach the user-facing output.
2
+
3
+ Enabled only when FORECAST_DEBUG=1 (or config.debug=True). All debug output
4
+ goes to stderr, so stdout stays clean for the final assistant response.
5
+ """
6
+
7
+ import os
8
+ import sys
9
+
10
+ _INTERNAL_PREFIXES = (
11
+ "thought:", "user wants", "$ python", "executing", "running tool",
12
+ "model inference", "the prediction has been computed", "debug:",
13
+ )
14
+
15
+
16
+ def is_debug() -> bool:
17
+ return os.environ.get("FORECAST_DEBUG", "").lower() in ("1", "true", "yes")
18
+
19
+
20
+ def debug(msg: str, enabled: bool = None):
21
+ if enabled is None:
22
+ enabled = is_debug()
23
+ if enabled:
24
+ print(f"[debug] {msg}", file=sys.stderr, flush=True)
25
+
26
+
27
+ def strip_internal(text: str) -> str:
28
+ """Drop any line that looks like an internal execution log."""
29
+ out = []
30
+ for line in text.splitlines():
31
+ if any(line.strip().lower().startswith(p) for p in _INTERNAL_PREFIXES):
32
+ continue
33
+ out.append(line)
34
+ return "\n".join(out).strip()
model.py ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ import torch.nn as nn
3
+
4
+
5
+ class LSTMForecaster(nn.Module):
6
+ def __init__(self, input_size=1, hidden_size=64, num_layers=2, dropout=0.1, horizon=7):
7
+ super().__init__()
8
+ self.horizon = horizon
9
+ self.lstm = nn.LSTM(
10
+ input_size=input_size,
11
+ hidden_size=hidden_size,
12
+ num_layers=num_layers,
13
+ batch_first=True,
14
+ dropout=dropout if num_layers > 1 else 0.0,
15
+ )
16
+ self.head = nn.Sequential(
17
+ nn.Linear(hidden_size, hidden_size),
18
+ nn.ReLU(),
19
+ nn.Dropout(dropout),
20
+ nn.Linear(hidden_size, horizon),
21
+ )
22
+
23
+ def forward(self, x):
24
+ out, _ = self.lstm(x)
25
+ return self.head(out[:, -1, :])
model_ai.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dee7474a03ee54ad7f3520b93cb39fe6a3839202d8e909dfc2d2908f3c8af6b9
3
+ size 1400457
model_economy.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aa6c16a49b9fe6fc324b21daca34830a8b5597227c471c8126716511b5babadc
3
+ size 224939
model_energy.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bdfc2cff1e92cda27a016f48d4ecb47077b050ba405731856b3eebc4d89dde8d
3
+ size 224921
model_finance.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:24f90b7fde4d0914d31a990195a4e99cc05cd2f6be2b3ffc7fae3c4285347d5c
3
+ size 1400567
model_programming.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4d7940f92a8cd48092c50935410229cf8ed918b5892a96a29c7095e5c17aec7a
3
+ size 1400655
model_sports.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e521beb33a8e8e46cfe166b58dcceee967a483967e6e7042cb6375b2581fb37a
3
+ size 1400545
model_weather_chennai_india.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1c99536594bd56b91868633503b129d30d94eec7ac6eeab0e9c50569663e1b5b
3
+ size 224939
multi.py ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import contextlib
2
+ import io
3
+ import json
4
+ import sys
5
+ from pathlib import Path
6
+
7
+ import torch
8
+
9
+ import data
10
+ import debug_logger
11
+ import prediction_engine
12
+ import train
13
+
14
+ BASE_DIR = Path(__file__).resolve().parent
15
+
16
+
17
+ def _quiet():
18
+ """Silence internal stdout unless FORECAST_DEBUG=1."""
19
+ if debug_logger.is_debug():
20
+ return contextlib.nullcontext()
21
+ return contextlib.redirect_stdout(io.StringIO())
22
+
23
+
24
+ def run_topic(topic, epochs=60, horizon_days=7, refresh_data=False, device="cpu"):
25
+ print(f"\nTOPIC: {data.TOPICS[topic]['label']} - {data.TOPICS[topic]['desc']}")
26
+ with _quiet():
27
+ series, source = data.load(topic=topic, refresh=refresh_data)
28
+ train.train_topic(topic, series, source, epochs=epochs,
29
+ horizon_days=horizon_days, device=device)
30
+ pred = prediction_engine.run_prediction(topic, horizon_days=horizon_days,
31
+ device=device, debug=debug_logger.is_debug())
32
+ return pred
33
+
34
+
35
+ def main():
36
+ topics = list(data.TOPICS)
37
+ if len(sys.argv) > 1 and sys.argv[1] not in ("--help",):
38
+ topics = [t for t in sys.argv[1:] if t in data.TOPICS] or topics
39
+ results = []
40
+ for topic in topics:
41
+ try:
42
+ results.append(run_topic(topic))
43
+ except Exception as exc:
44
+ print(f"[multi] topic '{topic}' failed: {exc}")
45
+
46
+ print(f"\n{'=' * 60}\nSUMMARY\n{'=' * 60}")
47
+ for r in results:
48
+ unit = r["unit"]
49
+ beat = r["naive_baseline_mae"] - r["val_mae"]
50
+ verdict = f"beats naive by {beat:,.1f} {unit}" if beat > 0 else f"~naive baseline ({unit})"
51
+ print(f"\n{r['label']} [{r['source']}]")
52
+ print(f" accuracy: MAPE={r['val_h1_mape'] * 100:.2f}% MAE={r['val_mae']:,.2f} {unit} ({verdict})")
53
+ print(f" {r['date_range_start']} -> {r['date_range_end']} center ~{r['center_estimate']:,.2f} "
54
+ f"{unit} outlook={r['direction']}")
55
+ print(f" charts/report: {r['chart']} / {r['report']}")
56
+
57
+ (BASE_DIR / "results.json").write_text(json.dumps(results, indent=2))
58
+ print(f"\n[done] full results written to results.json; charts at prediction_<topic>.png")
59
+
60
+
61
+ if __name__ == "__main__":
62
+ main()
predict.py ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import argparse
2
+ import json
3
+ from pathlib import Path
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+ import torch
8
+ import matplotlib
9
+ matplotlib.use("Agg")
10
+ import matplotlib.pyplot as plt
11
+
12
+ import data
13
+ from model import LSTMForecaster
14
+
15
+ BASE_DIR = Path(__file__).resolve().parent
16
+
17
+
18
+ def write_text_report(rows, cfg, topic):
19
+ unit = cfg.get("unit", "")
20
+ label = cfg.get("topic_label", topic)
21
+ beat = cfg["naive_baseline_mae"] - cfg["val_mae"]
22
+ lines = [
23
+ "=" * 62,
24
+ f"PREDICTION REPORT: {label}",
25
+ "=" * 62,
26
+ "",
27
+ f"Topic : {topic}",
28
+ f"Dataset : {cfg['source']}",
29
+ f"Model : LSTM ({cfg['num_layers']} layers, {cfg['hidden_size']} hidden units)",
30
+ f"Last known value : {cfg['last_date']} = {cfg['last_value']:,.2f} {unit}",
31
+ "",
32
+ "VALIDATION (held-out test set)",
33
+ f" Test MAE : {cfg['val_mae']:,.2f} {unit}",
34
+ f" Relative MAE : {cfg['val_rel_mae'] * 100:.2f}%",
35
+ f" H1 MAPE : {cfg['val_h1_mape'] * 100:.2f}%",
36
+ f" Naive baseline : {cfg['naive_baseline_mae']:,.2f} {unit}"
37
+ + (f" (model beats it by {beat:,.2f} {unit})" if beat > 0 else ""),
38
+ "",
39
+ "FUTURE PREDICTIONS",
40
+ ]
41
+ for r in rows:
42
+ lines.append(f" {r['date']} -> {r['predicted_value']:,.2f} {unit}")
43
+ lines += [
44
+ "",
45
+ "NOTE: forecasts are statistical estimates on a hold-out",
46
+ "validated model; no model can predict the future with 100%",
47
+ "accuracy. Treat these as central-point predictions.",
48
+ f"Generated: {pd.Timestamp.now().strftime('%Y-%m-%d %H:%M:%S')}",
49
+ ]
50
+ txt_path = BASE_DIR / f"prediction_{topic}{data.artifact_suffix(topic)}.txt"
51
+ txt_path.write_text("\n".join(lines), encoding="utf-8")
52
+ return txt_path
53
+
54
+
55
+ def load_artifacts(topic, device="cpu"):
56
+ suffix = data.artifact_suffix(topic)
57
+ model_path = BASE_DIR / f"model_{topic}{suffix}.pt"
58
+ config_path = BASE_DIR / f"config_{topic}{suffix}.json"
59
+ if not model_path.exists() or not config_path.exists():
60
+ raise FileNotFoundError(f"Model for topic '{topic}' not trained yet. Run: python train.py --topic {topic}")
61
+ cfg = json.loads(config_path.read_text())
62
+ model = LSTMForecaster(
63
+ hidden_size=cfg["hidden_size"],
64
+ num_layers=cfg["num_layers"],
65
+ dropout=cfg["dropout"],
66
+ horizon=cfg["horizon_steps"],
67
+ ).to(device)
68
+ model.load_state_dict(torch.load(model_path, map_location=device))
69
+ model.eval()
70
+ return model, cfg
71
+
72
+
73
+ def predict(topic, days=None, device="cpu"):
74
+ model, cfg = load_artifacts(topic, device)
75
+ series = pd.read_csv(BASE_DIR / "data" / f"{topic}.csv")
76
+ vals = series["value"].to_numpy(dtype=np.float64)
77
+ lb = cfg["lookback"]
78
+ freq_days = cfg.get("horizon_days", 7) / cfg["horizon_steps"]
79
+
80
+ want_days = float(days) if days else cfg["horizon_days"]
81
+ n_steps = int(round(want_days / freq_days)) if freq_days else cfg["horizon_steps"]
82
+
83
+ last_val = vals[-1]
84
+ rel = vals[-lb:] / last_val - 1.0
85
+ last_date = pd.Timestamp(cfg["last_date"])
86
+ rows = []
87
+ for step in range(n_steps):
88
+ x = torch.from_numpy(rel[-lb:].astype(np.float32)).unsqueeze(0).unsqueeze(-1).to(device)
89
+ with torch.no_grad():
90
+ pred_rel = model(x).cpu().numpy().flatten()
91
+ cur_val = last_val * (1.0 + pred_rel[0])
92
+ rel = np.concatenate([rel, pred_rel])
93
+ rows.append({
94
+ "date": (last_date + pd.Timedelta(days=freq_days * (step + 1))).strftime("%Y-%m-%d"),
95
+ "predicted_value": round(float(cur_val), 2),
96
+ })
97
+
98
+ recent = series.tail(lb)
99
+ history = pd.DataFrame({"date": recent["date"], "value": recent["value"], "predicted": np.nan})
100
+ future = pd.DataFrame({"date": [r["date"] for r in rows],
101
+ "value": np.nan,
102
+ "predicted": [r["predicted_value"] for r in rows]})
103
+ plot_df = pd.concat([history, future], ignore_index=True)
104
+ plot_df["d"] = pd.to_datetime(plot_df["date"])
105
+
106
+ chart_path = BASE_DIR / f"prediction_{topic}{data.artifact_suffix(topic)}.png"
107
+ fig, ax = plt.subplots(figsize=(11, 5))
108
+ ax.plot(plot_df["d"], plot_df["value"], label="actual", color="#2563eb")
109
+ ax.plot(plot_df["d"], plot_df["predicted"], label="model prediction", color="#dc2626", marker="o", markersize=4)
110
+ ax.axvline(last_date, color="gray", ls="--", lw=1)
111
+ unit = cfg.get("unit", "")
112
+ ax.set_title(f"{cfg.get('topic_label', topic)}: {cfg['source']} "
113
+ f"(val MAE={cfg['val_mae']:,.1f} {unit}, H1 MAPE={cfg['val_h1_mape'] * 100:.2f}%)")
114
+ ax.legend()
115
+ fig.autofmt_xdate()
116
+ fig.tight_layout()
117
+ fig.savefig(chart_path, dpi=130)
118
+ plt.close(fig)
119
+ txt_path = write_text_report(rows, cfg, topic)
120
+ return rows, cfg, chart_path, txt_path
121
+
122
+
123
+ def main():
124
+ parser = argparse.ArgumentParser()
125
+ parser.add_argument("--topic", default="finance", choices=list(data.TOPICS))
126
+ parser.add_argument("--days", type=float, default=None, help="how many days ahead to predict")
127
+ parser.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
128
+ args = parser.parse_args()
129
+ rows, cfg, chart_path, txt_path = predict(args.topic, days=args.days, device=args.device)
130
+ unit = cfg.get("unit", "")
131
+ print(f"Topic : {cfg.get('topic_label', args.topic)} - {cfg['source']}")
132
+ print(f"Last known : {cfg['last_date']} = {cfg['last_value']:,.2f} {unit}")
133
+ print(f"Validated on : MAE={cfg['val_mae']:,.2f}, rel MAE={cfg['val_rel_mae'] * 100:.2f}%, "
134
+ f"H1 MAPE={cfg['val_h1_mape'] * 100:.2f}% (naive baseline MAE={cfg['naive_baseline_mae']:,.2f})")
135
+ print(f"\nFuture predictions:")
136
+ for r in rows:
137
+ print(f" {r['date']} -> {r['predicted_value']:,.2f} {unit}")
138
+ print(f"\nChart saved to {chart_path}")
139
+ print(f"Text report saved to {txt_path}")
140
+
141
+
142
+ if __name__ == "__main__":
143
+ main()
prediction_engine.py ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Prediction engine: runs model inference only. No formatting, no printing.
2
+
3
+ Calculates scientifically meaningful, clearly separated metrics:
4
+ - net_change_pct : change along the forecast path (first -> last forecast)
5
+ - forecast_range_pct : (max - min) of the forecast path / first forecast
6
+ - current_to_forecast_pct : change from the latest OBSERVED value to the final forecast
7
+ - direction : rising/falling/sideways via configurable threshold
8
+ - primary_forecast : the forecast value at the END of the requested horizon
9
+ - center_estimate : mean of the forecast path (a secondary statistic)
10
+
11
+ The response_formatter presents these; this module never formats output.
12
+ Internal logging goes through debug_logger (stderr).
13
+ """
14
+
15
+ import contextlib
16
+ import io
17
+
18
+ import numpy as np
19
+
20
+ import data
21
+ import debug_logger
22
+ import predict
23
+
24
+ DEFAULT_DIRECTION_THRESHOLD_PCT = 0.2 # configurable; can become domain-specific
25
+
26
+
27
+ def _pct(change: float, base: float) -> float:
28
+ return float(change / abs(base) * 100.0) if base != 0 else 0.0
29
+
30
+
31
+ def run_prediction(topic: str, horizon_days: float = 7.0,
32
+ device: str = "cpu", debug: bool = False,
33
+ direction_threshold_pct: float = DEFAULT_DIRECTION_THRESHOLD_PCT) -> dict:
34
+ """Structured prediction result from the trained model. The inference
35
+ call's own stdout (status lines) is captured so nothing internal leaks."""
36
+ debug_logger.debug(f"prediction_engine: topic={topic} horizon={horizon_days}d "
37
+ f"thr={direction_threshold_pct}%", debug)
38
+ buf = io.StringIO()
39
+ with contextlib.redirect_stdout(buf):
40
+ rows, cfg, chart_path, txt_path = predict.predict(topic, days=horizon_days, device=device)
41
+ debug_logger.debug(f"prediction_engine internal: {buf.getvalue()[:300]!r}", debug)
42
+
43
+ vals = np.array([r["predicted_value"] for r in rows], dtype=np.float64)
44
+ if len(vals) == 0 or not np.all(np.isfinite(vals)):
45
+ raise ValueError("model produced an empty or invalid forecast")
46
+
47
+ first = float(vals[0])
48
+ last = float(vals[-1])
49
+ last_known = cfg.get("last_value")
50
+
51
+ net_change_pct = _pct(last - first, first)
52
+ forecast_range_pct = _pct(vals.max() - vals.min(), first)
53
+ current_to_forecast_pct = (
54
+ _pct(last - float(last_known), float(last_known))
55
+ if last_known is not None and float(last_known) != 0 else None
56
+ )
57
+
58
+ if net_change_pct > direction_threshold_pct:
59
+ direction = "rising"
60
+ elif net_change_pct < -direction_threshold_pct:
61
+ direction = "falling"
62
+ else:
63
+ direction = "sideways"
64
+
65
+ result = {
66
+ "topic": topic,
67
+ "label": cfg.get("topic_label", topic),
68
+ "unit": cfg.get("unit", ""),
69
+ "asset": data.TOPICS.get(topic, {}).get("asset", cfg.get("topic_label", topic)),
70
+ "source": cfg["source"],
71
+ "date_range_start": rows[0]["date"],
72
+ "date_range_end": rows[-1]["date"],
73
+ "rows": rows,
74
+ "values": vals.tolist(),
75
+ "primary_forecast": last,
76
+ "center_estimate": float(np.mean(vals)),
77
+ "first_value": first,
78
+ "last_value": last,
79
+ "last_known_value": float(last_known) if last_known is not None else None,
80
+ "last_known_date": cfg.get("last_date"),
81
+ "net_change_pct": net_change_pct,
82
+ "forecast_range_pct": forecast_range_pct,
83
+ "current_to_forecast_pct": current_to_forecast_pct,
84
+ "direction": direction,
85
+ "direction_threshold_pct": direction_threshold_pct,
86
+ "val_mae": cfg["val_mae"],
87
+ "val_h1_mape": cfg["val_h1_mape"],
88
+ "naive_baseline_mae": cfg["naive_baseline_mae"],
89
+ "chart": str(chart_path),
90
+ "report": str(txt_path),
91
+ }
92
+ debug_logger.debug(
93
+ f"prediction_engine: {topic} end-of-horizon={last} direction={direction} "
94
+ f"net={net_change_pct:+.3f}% range={forecast_range_pct:.3f}% "
95
+ f"cur->end={current_to_forecast_pct if current_to_forecast_pct is None else format(current_to_forecast_pct, '+.3f')}%",
96
+ debug)
97
+ return result
requirements.txt ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ torch>=2.0
2
+ numpy
3
+ pandas
4
+ matplotlib
5
+ scikit-learn
6
+ datasets
7
+ huggingface_hub
8
+ yfinance
response.py ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Orchestrator: intent detection -> prediction engine -> response formatter.
2
+
3
+ Only the final formatted answer is returned; internal execution details stay
4
+ in debug_logger (stderr, off by default). Raw tool output never reaches the
5
+ user, and each prediction is presented exactly once.
6
+ """
7
+
8
+ import re
9
+ import sys
10
+ from dataclasses import dataclass, field
11
+ from typing import Iterator, Optional
12
+
13
+ import debug_logger
14
+ import prediction_engine
15
+ import response_formatter
16
+ from data import TOPICS
17
+
18
+ SYSTEM_PROMPT = (
19
+ "You are a forecasting assistant. Present the model's prediction directly "
20
+ "and naturally, exactly once. Never echo the user's prompt, never expose "
21
+ "internal execution details or reasoning, and never present a forecast "
22
+ "as guaranteed."
23
+ )
24
+
25
+ INTENT_PATTERNS = {
26
+ "finance": {"bitcoin", "btc", "crypto", "cryptocurrency"},
27
+ "ai": {"nvidia", "nvda", "gpu"},
28
+ "programming": {"npm", "react", "downloads", "javascript"},
29
+ "sports": {"tennis", "atp", "elo", "sport", "rating"},
30
+ "weather": {"weather", "rain", "raining", "temperature", "snow", "sunny",
31
+ "storm", "humidity", "forecast"},
32
+ "economy": {"economy", "gspc", "sp500", "market", "stocks", "inflation"},
33
+ "energy": {"oil", "wti", "crude", "energy", "gas", "petroleum"},
34
+ }
35
+
36
+
37
+ @dataclass
38
+ class GenerationConfig:
39
+ debug: bool = False
40
+ streaming: bool = False
41
+ chunk_delay: float = 0.02
42
+ max_history_turns: int = 8
43
+ max_input_chars: int = 4000
44
+ default_horizon_days: float = 7.0
45
+ detailed: bool = False
46
+
47
+
48
+ @dataclass
49
+ class ConversationContext:
50
+ turns: list = field(default_factory=list)
51
+ max_turns: int = 8
52
+
53
+ def add(self, role: str, content: str):
54
+ self.turns.append({"role": role, "content": content})
55
+ if len(self.turns) > self.max_turns * 2:
56
+ self.turns = self.turns[-self.max_turns * 2:]
57
+
58
+
59
+ def detect_intent(query: str, config: GenerationConfig) -> dict:
60
+ q = query.lower()
61
+ tokens = set(re.findall(r"[a-z0-9$%-]+", q))
62
+
63
+ unsupported = tokens & {"earthquake", "volcano", "traffic", "election", "lottery",
64
+ "population", "pandemic"}
65
+ topics = [t for t in TOPICS if t in tokens or tokens & INTENT_PATTERNS.get(t, set())]
66
+ explicit_all = any(w in q for w in ("compare", "all topics", "everything", "each topic"))
67
+ greeting = bool(tokens & {"hi", "hello", "hey", "thanks"}) and len(tokens) <= 3
68
+
69
+ m = re.search(r"(\d+)\s*(?:-|\s)?\s*(?:days?|d)\b", q)
70
+ if m:
71
+ horizon = float(m.group(1))
72
+ elif any(w in q for w in ("month", "30 days", "30d")):
73
+ horizon = 30.0
74
+ elif "tomorrow" in tokens:
75
+ horizon = 1.0
76
+ elif "week" in tokens:
77
+ horizon = 7.0
78
+ else:
79
+ horizon = config.default_horizon_days
80
+
81
+ detailed = config.detailed or bool(tokens & {"detail", "detailed", "breakdown", "table", "chart"})
82
+ return {"greeting": greeting, "topics": topics, "compare": explicit_all,
83
+ "horizon": horizon, "detailed": detailed,
84
+ "unsupported": bool(unsupported) and not topics}
85
+
86
+
87
+ def _run_topic(topic: str, horizon: float, device: str, config: GenerationConfig):
88
+ try:
89
+ return prediction_engine.run_prediction(topic, horizon_days=horizon,
90
+ device=device, debug=config.debug)
91
+ except Exception as exc:
92
+ debug_logger.debug(f"engine error for {topic}: {exc!r}", config.debug)
93
+ return None
94
+
95
+
96
+ def generate_response(query: str, device: str = "cpu",
97
+ config: Optional[GenerationConfig] = None) -> str:
98
+ """USER INPUT -> intent -> prediction/engine -> formatted answer."""
99
+ config = config or GenerationConfig()
100
+ debug_logger.debug(f"user: {query[:200]!r}", config.debug)
101
+
102
+ query = (query or "").strip()
103
+ if not query:
104
+ return "Please ask something - for example: *what will bitcoin do next week?*"
105
+ query = query[: config.max_input_chars]
106
+
107
+ try:
108
+ intent = detect_intent(query, config)
109
+ except Exception as exc:
110
+ debug_logger.debug(f"intent error: {exc!r}", config.debug)
111
+ return response_formatter.format_error()
112
+
113
+ if intent["greeting"]:
114
+ return response_formatter.format_greeting()
115
+
116
+ if intent.get("unsupported"):
117
+ return response_formatter.format_unknown()
118
+
119
+ if intent["compare"] or len(intent["topics"]) > 1 or not intent["topics"]:
120
+ topics = intent["topics"] if intent["topics"] and intent["compare"] else list(TOPICS)
121
+ preds = [_run_topic(t, intent["horizon"], device, config) for t in topics]
122
+ if all(p is None for p in preds):
123
+ return response_formatter.format_error()
124
+ return response_formatter.format_comparison([p for p in preds if p is not None],
125
+ debug=config.debug)
126
+
127
+ pred = _run_topic(intent["topics"][0], intent["horizon"], device, config)
128
+ if pred is None:
129
+ return response_formatter.format_error()
130
+ return response_formatter.format_single(pred, detailed=intent["detailed"],
131
+ debug=config.debug)
132
+
133
+
134
+ def stream_response(query: str, device: str = "cpu",
135
+ config: Optional[GenerationConfig] = None,
136
+ pregenerated: Optional[str] = None) -> Iterator[str]:
137
+ """Streaming: the full answer is generated first (or passed in), then
138
+ delivered line by line, so no partial or internal fragments reach the UI."""
139
+ config = config or GenerationConfig()
140
+ full = pregenerated if pregenerated is not None else generate_response(
141
+ query, device=device, config=config)
142
+ for chunk in full.splitlines(True):
143
+ yield chunk
response_formatter.py ADDED
@@ -0,0 +1,182 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Response formatter: presents structured prediction metrics naturally.
2
+
3
+ Never alters predictions, never invents confidence/probability/factors.
4
+ Only displays metrics the prediction engine actually computed. Adapts
5
+ wording to direction and prediction type.
6
+ """
7
+
8
+ import pandas as pd
9
+
10
+ import debug_logger
11
+
12
+ _RISK_NOTES = {
13
+ "USD": "This is a model-generated forecast, and actual market behavior may differ.",
14
+ "Elo": "This is a model-generated forecast; actual results depend on future matches.",
15
+ "downloads": "This is a model-generated estimate; actual download counts may differ.",
16
+ "K": "This is a model-generated forecast; actual weather may differ.",
17
+ }
18
+ _DEFAULT_NOTE = "This is a model-generated estimate, and actual values may differ."
19
+
20
+
21
+ def fmt_range_dates(start: str, end: str) -> str:
22
+ a, b = pd.Timestamp(start), pd.Timestamp(end)
23
+ if a.date() == b.date():
24
+ return f"{a.strftime('%B')} {a.day}, {a.year}"
25
+ if (a.year, a.month) == (b.year, b.month):
26
+ return f"{a.strftime('%B')} {a.day}-{b.day}, {a.year}"
27
+ if a.year == b.year:
28
+ return f"{a.strftime('%B')} {a.day} - {b.strftime('%B')} {b.day}, {b.year}"
29
+ return f"{a.strftime('%B %d, %Y')} - {b.strftime('%B %d, %Y')}"
30
+
31
+
32
+ def fmt_value(v: float, unit: str) -> str:
33
+ if v is None:
34
+ return "n/a"
35
+ if unit == "downloads":
36
+ return f"{v / 1e6:.2f}M" if abs(v) >= 1e6 else f"{v:,.0f}"
37
+ if unit == "USD":
38
+ return f"${v:,.0f}" if abs(v) >= 1000 else f"${v:,.2f}"
39
+ if unit == "K":
40
+ return f"{v - 273.15:.1f} °C"
41
+ if unit:
42
+ return f"{v:,.2f} {unit}"
43
+ return f"{v:,.2f}"
44
+
45
+
46
+ def fmt_pct(pct: float, signed: bool = False) -> str:
47
+ """Human-readable percent; very small values shown as <0.01%."""
48
+ if pct is None:
49
+ return "n/a"
50
+ if abs(pct) < 0.005:
51
+ return "<0.01%"
52
+ return f"{pct:+.2f}%" if signed else f"{pct:.2f}%"
53
+
54
+
55
+ def _horizon_word(start: str, end: str) -> str:
56
+ days = (pd.Timestamp(end) - pd.Timestamp(start)).days + 1
57
+ if days <= 1:
58
+ return "tomorrow"
59
+ if days == 7:
60
+ return "next week"
61
+ if days == 30:
62
+ return "next month"
63
+ return f"over the next {days} days"
64
+
65
+
66
+ _DIRECTION_VERB = {"sideways": "relatively sideways", "rising": "upward", "falling": "downward"}
67
+ _DIRECTION_LABEL = {"sideways": "Sideways", "rising": "Rising", "falling": "Falling"}
68
+
69
+
70
+ def _interpretation(pred: dict, unit: str, horizon_word: str) -> str:
71
+ asset = pred.get("asset", pred["label"])
72
+ direction = pred["direction"]
73
+ end_val = fmt_value(pred["primary_forecast"], unit)
74
+ start_val = fmt_value(pred["first_value"], unit)
75
+
76
+ if direction == "rising":
77
+ return (f"The model forecasts {_DIRECTION_VERB['rising']} movement for {asset} {horizon_word}, "
78
+ f"with the predicted level increasing from {start_val} to approximately **{end_val}** "
79
+ f"({fmt_pct(pred['net_change_pct'], signed=True)} net change).")
80
+ if direction == "falling":
81
+ return (f"The model forecasts {_DIRECTION_VERB['falling']} movement for {asset} {horizon_word}, "
82
+ f"with the predicted level decreasing from {start_val} to approximately **{end_val}** "
83
+ f"({fmt_pct(pred['net_change_pct'], signed=True)} net change).")
84
+ return (f"The model forecasts {_DIRECTION_VERB['sideways']} movement for {asset} {horizon_word}, "
85
+ f"with an estimated level around **{end_val}**.")
86
+
87
+
88
+ def _summary_sentence(pred: dict, unit: str, horizon_word: str) -> str:
89
+ """A single conversational sentence restating the forecast in plain words."""
90
+ asset = pred.get("asset", pred["label"])
91
+ direction = pred["direction"]
92
+ end_val = fmt_value(pred["primary_forecast"], unit)
93
+ change = fmt_pct(pred["net_change_pct"], signed=True)
94
+ if direction == "sideways":
95
+ return (f"In short, the model expects {asset} to stay roughly stable {horizon_word}, "
96
+ f"hovering near {end_val}.")
97
+ if direction == "rising":
98
+ return (f"In short, the model expects {asset} to climb {change} {horizon_word}, "
99
+ f"reaching about {end_val}.")
100
+ return (f"In short, the model expects {asset} to slip {change} {horizon_word}, "
101
+ f"settling around {end_val}.")
102
+
103
+
104
+ def format_single(pred: dict, detailed: bool = False, debug: bool = False) -> str:
105
+ """Summary sentence + title, date range, interpretation, key values, note.
106
+ Only shows metrics that exist and are meaningful. Single coherent answer."""
107
+ unit = pred["unit"]
108
+ asset = pred.get("asset", pred["label"])
109
+ direction = pred["direction"]
110
+ horizon_word = _horizon_word(pred["date_range_start"], pred["date_range_end"])
111
+ end_val = fmt_value(pred["primary_forecast"], unit)
112
+
113
+ lines = [
114
+ _summary_sentence(pred, unit, horizon_word),
115
+ "",
116
+ f"{asset} Forecast",
117
+ "",
118
+ fmt_range_dates(pred["date_range_start"], pred["date_range_end"]),
119
+ "",
120
+ _interpretation(pred, unit, horizon_word),
121
+ "",
122
+ f"Outlook: {_DIRECTION_LABEL[direction]}",
123
+ f"Predicted level ({pred['date_range_end']}): ~{end_val}",
124
+ f"Expected change: {fmt_pct(pred['net_change_pct'], signed=True)}",
125
+ ]
126
+
127
+ if pred.get("forecast_range_pct") is not None:
128
+ lines.append(f"Forecast range: {fmt_pct(abs(pred['forecast_range_pct']))}")
129
+ if pred.get("current_to_forecast_pct") is not None and pred.get("last_known_value") is not None:
130
+ lines.append(
131
+ f"From latest observed value ({fmt_value(pred['last_known_value'], unit)}, "
132
+ f"{pred.get('last_known_date', 'latest')}): {fmt_pct(pred['current_to_forecast_pct'], signed=True)}")
133
+
134
+ if detailed:
135
+ lines += ["", "| Date | Prediction |", "|---|---|"]
136
+ for r in pred["rows"]:
137
+ lines.append(f"| {r['date']} | {fmt_value(r['predicted_value'], unit)} |")
138
+ lines += [
139
+ "",
140
+ f"Path average: {fmt_value(pred['center_estimate'], unit)}. "
141
+ f"Validated accuracy on unseen data: MAPE {pred['val_h1_mape'] * 100:.2f}% "
142
+ f"(MAE {pred['val_mae']:,.2f} {unit}).",
143
+ ]
144
+
145
+ lines += ["", _RISK_NOTES.get(unit, _DEFAULT_NOTE)]
146
+ return debug_logger.strip_internal("\n".join(lines))
147
+
148
+
149
+ def format_comparison(preds: list, debug: bool = False) -> str:
150
+ """One row per topic; shows end-of-horizon forecast + net change."""
151
+ lines = ["Forecast Summary", "",
152
+ "| Topic | Outlook | Predicted (end) | Expected change | Accuracy |",
153
+ "|---|---|---|---|---|"]
154
+ for p in preds:
155
+ if p is None:
156
+ continue
157
+ unit = p["unit"]
158
+ lines.append(
159
+ f"| {p['label']} | {_DIRECTION_LABEL[p['direction']]} | "
160
+ f"~{fmt_value(p['primary_forecast'], unit)} "
161
+ f"({fmt_range_dates(p['date_range_start'], p['date_range_end'])}) | "
162
+ f"{fmt_pct(p['net_change_pct'], signed=True)} | "
163
+ f"MAPE {p['val_h1_mape'] * 100:.2f}% |")
164
+ lines += ["", _DEFAULT_NOTE]
165
+ return debug_logger.strip_internal("\n".join(lines))
166
+
167
+
168
+ def format_error(message: str = "") -> str:
169
+ return "I hit a problem generating that forecast. Please try again in a moment."
170
+
171
+
172
+ def format_greeting() -> str:
173
+ return ("Hello! I generate forecasts from trained models. Ask about **AI**, "
174
+ "**Programming**, **Finance**, **Sports**, **Weather**, **Economy**, "
175
+ "or **Energy** - for example, *what will bitcoin do next week?*")
176
+
177
+
178
+ def format_unknown() -> str:
179
+ import data
180
+ parts = ", ".join(f"**{info['label']}**" for info in data.TOPICS.values())
181
+ return (f"I can forecast these topics: {parts}. "
182
+ f"Try: *predict finance for 14 days* or *what will the weather do tomorrow?*")
run_all.ps1 ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ## Full automation: fetch data -> train -> evaluate -> predict (all topics)
2
+ ## Usage: .\run_all.ps1 [-Days 7] [-Epochs 60] [-Topics ai,programming,finance,sports] [-RefreshData]
3
+
4
+ param(
5
+ [double]$Days = 7,
6
+ [int]$Epochs = 60,
7
+ [string[]]$Topics = @("ai", "programming", "finance", "sports"),
8
+ [switch]$RefreshData
9
+ )
10
+
11
+ $ErrorActionPreference = "Stop"
12
+ $py = Join-Path $PSScriptRoot ".venv\Scripts\python.exe"
13
+ if (-not (Test-Path $py)) { $py = "python" }
14
+
15
+ if ($RefreshData) {
16
+ foreach ($t in $Topics) {
17
+ & $py "$PSScriptRoot\data.py" --topic $t --refresh
18
+ }
19
+ }
20
+
21
+ & $py "$PSScriptRoot\multi.py" @Topics
train.py ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import argparse
2
+ import json
3
+ from pathlib import Path
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+ import torch
8
+ import torch.nn as nn
9
+ from torch.utils.data import DataLoader, TensorDataset
10
+
11
+ import data
12
+ from model import LSTMForecaster
13
+
14
+
15
+ def build_windows(values, lookback, horizon):
16
+ xs, ys, bases = [], [], []
17
+ for i in range(lookback, len(values) - horizon):
18
+ base = values[i - 1]
19
+ if not np.isfinite(base) or base <= 0:
20
+ continue
21
+ xs.append(values[i - lookback:i] / base - 1.0)
22
+ ys.append(values[i:i + horizon] / base - 1.0)
23
+ bases.append(base)
24
+ return (np.array(xs, dtype=np.float32), np.array(ys, dtype=np.float32), np.array(bases, dtype=np.float64))
25
+
26
+
27
+ def evaluate(model, loader, device):
28
+ model.eval()
29
+ preds, actuals, bases = [], [], []
30
+ with torch.no_grad():
31
+ for xb, yb, bb in loader:
32
+ preds.append(model(xb.to(device)).cpu().numpy())
33
+ actuals.append(yb.numpy())
34
+ bases.append(bb.numpy())
35
+ preds = np.concatenate(preds)
36
+ actuals = np.concatenate(actuals)
37
+ bases = np.concatenate(bases)[:, None]
38
+ pred_usd = bases * (1.0 + preds)
39
+ act_usd = bases * (1.0 + actuals)
40
+ mae = float(np.mean(np.abs(pred_usd - act_usd)))
41
+ rel_mae = float(np.mean(np.abs(preds - actuals)))
42
+ h1_mape = float(np.mean(np.abs((act_usd[:, 0] - pred_usd[:, 0]) / act_usd[:, 0])))
43
+ naive_mae = float(np.mean(np.abs(act_usd - bases)))
44
+ return mae, rel_mae, h1_mape, naive_mae, pred_usd, act_usd
45
+
46
+
47
+ def train_model(series, lookback=60, horizon=7, epochs=60, batch_size=32, lr=1e-3, device="cpu", seed=0,
48
+ hidden_size=64, num_layers=2, dropout=0.1, patience=15):
49
+ torch.manual_seed(seed)
50
+ np.random.seed(seed)
51
+ vals = series["value"].to_numpy(dtype=np.float64)
52
+ n = len(vals)
53
+ n_test = min(180, max(horizon * 5, n // 5))
54
+
55
+ x_all, y_all, b_all = build_windows(vals, lookback, horizon)
56
+ train_count = max(lookback, (n - n_test - horizon) - lookback)
57
+ test_start = (n - n_test) - lookback
58
+ x_train, y_train = x_all[:train_count], y_all[:train_count]
59
+ x_test, y_test, b_test = x_all[test_start:], y_all[test_start:], b_all[test_start:]
60
+
61
+ train_ds = TensorDataset(torch.from_numpy(x_train).unsqueeze(-1), torch.from_numpy(y_train))
62
+ train_loader = DataLoader(train_ds, batch_size=batch_size, shuffle=True)
63
+ test_loader = DataLoader(
64
+ TensorDataset(torch.from_numpy(x_test).unsqueeze(-1), torch.from_numpy(y_test), torch.from_numpy(b_test)),
65
+ batch_size=512,
66
+ )
67
+
68
+ model = LSTMForecaster(hidden_size=hidden_size, num_layers=num_layers, dropout=dropout, horizon=horizon).to(device)
69
+ opt = torch.optim.AdamW(model.parameters(), lr=lr, weight_decay=1e-5)
70
+ sched = torch.optim.lr_scheduler.ReduceLROnPlateau(opt, patience=6, factor=0.5)
71
+ crit = nn.HuberLoss()
72
+
73
+ best_mae, best_state, patience_seen = float("inf"), None, 0
74
+ for epoch in range(1, epochs + 1):
75
+ model.train()
76
+ total = 0.0
77
+ for xb, yb in train_loader:
78
+ opt.zero_grad()
79
+ loss = crit(model(xb.to(device)), yb.to(device))
80
+ loss.backward()
81
+ torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
82
+ opt.step()
83
+ total += float(loss.item()) * len(xb)
84
+ mae, rel_mae, mape, naive, _, _ = evaluate(model, test_loader, device)
85
+ sched.step(mae)
86
+ print(f"epoch {epoch:3d}/{epochs} loss={total / len(train_ds):.5f} "
87
+ f"val_MAE=${mae:,.0f} rel_MAE={rel_mae * 100:.2f}% H1_MAPE={mape * 100:.2f}%")
88
+ if mae < best_mae - 1e-6:
89
+ best_mae, patience_seen = mae, 0
90
+ best_state = {k: v.detach().cpu().clone() for k, v in model.state_dict().items()}
91
+ else:
92
+ patience_seen += 1
93
+ if patience_seen >= patience:
94
+ print("[train] early stopping")
95
+ break
96
+
97
+ if best_state is not None:
98
+ model.load_state_dict(best_state)
99
+ mae, rel_mae, h1_mape, naive_mae, _, _ = evaluate(model, test_loader, device)
100
+ return model, {"mae": mae, "rel_mae": rel_mae, "h1_mape": h1_mape, "naive_mae": naive_mae,
101
+ "lookback": lookback, "horizon": horizon}
102
+
103
+
104
+ def train_topic(topic, series, source, epochs=60, horizon_days=7, device="cpu", hidden_size=64, num_layers=2,
105
+ dropout=0.1, lookback=60, patience=15):
106
+ gaps_days = float(pd.to_datetime(series["date"]).diff().dropna().dt.days.median()) if len(series) > 1 else 1.0
107
+ horizon_steps = max(1, int(round(horizon_days / gaps_days))) if gaps_days > 0 else horizon_days
108
+
109
+ print(f"[train] topic={topic} dataset={source} points={len(series)} "
110
+ f"horizon={horizon_steps} steps ({horizon_days} days)")
111
+ model, stats = train_model(series, lookback=lookback, horizon=horizon_steps, epochs=epochs, device=device,
112
+ hidden_size=hidden_size, num_layers=num_layers, dropout=dropout, patience=patience)
113
+
114
+ suffix = data.artifact_suffix(topic)
115
+ model_path = Path(__file__).resolve().parent / f"model_{topic}{suffix}.pt"
116
+ config_path = Path(__file__).resolve().parent / f"config_{topic}{suffix}.json"
117
+ torch.save(model.state_dict(), model_path)
118
+ cfg = {
119
+ "topic": topic,
120
+ "topic_label": data.TOPICS[topic]["label"],
121
+ "unit": data.TOPICS[topic]["unit"],
122
+ "source": source,
123
+ "lookback": stats["lookback"],
124
+ "horizon_steps": horizon_steps,
125
+ "horizon_days": horizon_days,
126
+ "hidden_size": hidden_size,
127
+ "num_layers": num_layers,
128
+ "dropout": dropout,
129
+ "val_mae": stats["mae"],
130
+ "val_rel_mae": stats["rel_mae"],
131
+ "val_h1_mape": stats["h1_mape"],
132
+ "naive_baseline_mae": stats["naive_mae"],
133
+ "last_date": str(series["date"].iloc[-1]),
134
+ "last_value": float(series["value"].iloc[-1]),
135
+ }
136
+ config_path.write_text(json.dumps(cfg, indent=2))
137
+ improvement = (stats["naive_mae"] - stats["mae"]) / stats["naive_mae"] * 100 if stats["naive_mae"] else 0.0
138
+ print(f"\n[result] {topic}: test MAE={stats['mae']:,.2f} rel_MAE={stats['rel_mae'] * 100:.2f}% "
139
+ f"H1 MAPE={stats['h1_mape'] * 100:.2f}%")
140
+ print(f"[result] naive(persistence) MAE={stats['naive_mae']:,.2f} model improvement={improvement:+.1f}%")
141
+ print(f"[result] model -> {model_path.name}, config -> {config_path.name}")
142
+ return model, cfg
143
+
144
+
145
+ def main():
146
+ parser = argparse.ArgumentParser()
147
+ parser.add_argument("--topic", default="finance", choices=list(data.TOPICS))
148
+ parser.add_argument("--refresh-data", action="store_true")
149
+ parser.add_argument("--epochs", type=int, default=60)
150
+ parser.add_argument("--horizon", type=int, default=7)
151
+ parser.add_argument("--hidden-size", type=int, default=64)
152
+ parser.add_argument("--num-layers", type=int, default=2)
153
+ parser.add_argument("--lookback", type=int, default=60)
154
+ parser.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
155
+ args = parser.parse_args()
156
+
157
+ series, source = data.load(topic=args.topic, refresh=args.refresh_data)
158
+ train_topic(args.topic, series, source, epochs=args.epochs, horizon_days=args.horizon, device=args.device,
159
+ hidden_size=args.hidden_size, num_layers=args.num_layers, lookback=args.lookback)
160
+
161
+
162
+ if __name__ == "__main__":
163
+ main()
weather_cities.py ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Weather city registry (name, latitude, longitude). Cities are keyed by
2
+ lowercase token; add entries to expand weather coverage without touching
3
+ any other module."""
4
+
5
+ CITIES = {
6
+ "chennai": {"name": "Chennai, India", "lat": 13.0827, "lon": 80.2707},
7
+ "mumbai": {"name": "Mumbai, India", "lat": 19.076, "lon": 72.8777},
8
+ "delhi": {"name": "Delhi, India", "lat": 28.6139, "lon": 77.209},
9
+ "bengaluru": {"name": "Bengaluru, India", "lat": 12.9716, "lon": 77.5946},
10
+ "bangalore": {"name": "Bengaluru, India", "lat": 12.9716, "lon": 77.5946},
11
+ "hyderabad": {"name": "Hyderabad, India", "lat": 17.385, "lon": 78.4867},
12
+ "kolkata": {"name": "Kolkata, India", "lat": 22.5726, "lon": 88.3639},
13
+ "pune": {"name": "Pune, India", "lat": 18.5204, "lon": 73.8567},
14
+ "new york": {"name": "New York, USA", "lat": 40.7128, "lon": -74.006},
15
+ "nyc": {"name": "New York, USA", "lat": 40.7128, "lon": -74.006},
16
+ "london": {"name": "London, UK", "lat": 51.5072, "lon": -0.1276},
17
+ "tokyo": {"name": "Tokyo, Japan", "lat": 35.6762, "lon": 139.6503},
18
+ "singapore": {"name": "Singapore", "lat": 1.3521, "lon": 103.8198},
19
+ "dubai": {"name": "Dubai, UAE", "lat": 25.2048, "lon": 55.2708},
20
+ "berlin": {"name": "Berlin, Germany", "lat": 52.52, "lon": 13.405},
21
+ "paris": {"name": "Paris, France", "lat": 48.8566, "lon": 2.3522},
22
+ "sydney": {"name": "Sydney, Australia", "lat": -33.8688, "lon": 151.2093},
23
+ }
24
+
25
+
26
+ def match_location(tokens) -> str | None:
27
+ """Return CITIES key present in the token set; None if no known city."""
28
+ for key in CITIES:
29
+ if key in tokens:
30
+ return key
31
+ return None
32
+
33
+
34
+ def display_names() -> list:
35
+ seen, out = set(), []
36
+ for info in CITIES.values():
37
+ if info["name"] not in seen:
38
+ seen.add(info["name"])
39
+ out.append(info["name"].split(",")[0])
40
+ return out