Configurable deep-learning experiments for time-series forecasting and classification. predictor
trains, evaluates and optimizes Keras/TensorFlow forecasting and
classification models — ANN, CNN, LSTM, Transformer, TCN, TFT, N-BEATS, MIMO
and binary/direction classifier variants — through a plugin architecture in
which predictor, optimizer, pipeline, preprocessor and target-calculation
plugins are selected by name from JSON configs. Experiments are organized as
numbered phases under examples/config/, each phase a
reproducible sweep over architectures, dataset sizes and horizons.
This project supports a data-centric research program: characterize the inputs, test temporal preprocessing, and compare learned representations before scaling model search. The doctoral proposal investigates modular temporal representations for forecasting and reinforcement learning. It is a proposal, not a claim that its hypotheses have already been confirmed.
- Current doctoral proposal (PDF) and editable LaTeX.
- Research repository map: data, preprocessing, feature engineering, representation learning, optimization and evaluation.
- Related proposals and earlier formulations.
- Current research plan, latest ML review and active execution orders.
- Data governance: dataset receipts, experiment provenance and results in an analytical warehouse.
Branch scope (2026-09-14): master contains the standalone trainer and
the documents linked above. The larger data-characterization and governed-run
implementation remains on the
published research snapshot.
Its governed-run guide
applies to that revision, not automatically to this branch. Research snapshots
are linked by commit so that development cannot silently change a cited implementation.
Read this README and the selected revision's configuration and entry points. Use an isolated environment. Run CLI help and the bounded CPU example below in a temporary checkout, preserving committed example outputs. Report the effective configuration, exact input/output paths and any failed tests. Do not start a sweep or write to the production cube. For governed runs, follow the linked integration version and keep delivery receipts and every outcome, including failures; do not claim that a profile test is training.
The repository map identifies component ownership. The warehouse service installs separately from training.
Research software under active development. predictor is the offline model-training component. Model serving belongs to prediction_provider. Example results are not evidence of out-of-sample trading profitability.
Disclaimer: all training and evaluation happens offline on historical data (simulation/backtest). Model outputs are research artifacts, not trading signals; nothing in this repository is financial advice, and no real-capital execution happens here.
Role: own model definition, training, evaluation and hyperparameter optimization for time-series prediction, plus the phased experiment configs and their results.
Not responsible for:
- Serving predictions — prediction_provider hosts trained models behind a FastAPI service.
- Feature/label engineering — feature-eng generates technical-indicator features and direction/oracle labels; feature-extractor trains the autoencoder encoders referenced by some configs.
- Generic CSV preprocessing — the standalone preprocessor application.
- Trading environments or RL agents — gym-fx and agent-multi.
- Decentralized optimization infrastructure — doin-node (see below).
JSON config (examples/config/phase_*/...)
│
▼
app/main.py ── app/cli.py / app/config.py / app/config_handler.py
│
▼
pipeline plugin (pipeline_plugins/)
├─ preprocessor plugin (preprocessor_plugins/) sliding windows, STL, features
├─ target plugin (target_plugins/) regression / binary / direction targets
├─ predictor plugin (predictor_plugins/) Keras model build/train/predict
└─ optimizer plugin (optimizer_plugins/) DEAP GA / NEAT hyperparameter search
examples/config/phase_1/— ANN/CNN/LSTM/ Transformer sweeps at 1h over dataset sizes from 1 575 to 50 400 bars, with anoptimization/subdirectory;phase_1_daily/repeats the sweep at 1d and adds TCN+NEAT optimization configs.examples/config/phase_1b_binary/— binary entry/exit classifiers (buy/sell × entry/exit) per architecture, plus champion inference configs.examples/config/phase_1c_direction/— direction classifiers.phase_2…phase_4_3(with_dailyvariants) — progressively deeper experiments culminating in Transformer configs at multiple horizons.
Trained champions are kept as .keras models with JSON metadata (e.g. a
generated, gitignored predictor_model_metadata.json at the repository
root; committed examples live under
examples/results/) so prediction_provider can load
them.
Runtime dependencies are listed in requirements.txt
(TensorFlow, tf-keras, tensorflow-probability, numpy, pandas, scipy, DEAP,
pmdarima, PyWavelets, matplotlib, psycopg2-binary, ...).
setup.py intentionally declares only a minimal
install_requires; treat requirements.txt as authoritative. No
python_requires is declared; the platform is exercised in practice on
Python 3.12 (verified below with Python 3.12.13, TensorFlow 2.21.0). A CUDA
GPU is optional but strongly recommended for training.
git clone https://github.com/harveybc/predictor.git
cd predictor
pip install -r requirements.txt
pip install -e . # installs the `predictor` console scriptUnverified in a clean environment — the commands above are the standard
install; they were not re-executed from scratch for this README. The imports
and CLI below were verified in an existing Python 3.12.13 environment.
The 2026-09-14 publication check used an existing environment and generated
local entry-point metadata with python setup.py egg_info; it did not install
or upgrade packages in a shared environment.
Verified (cheap) — the CLI parses and prints its full usage:
PYTHONPATH=. python app/main.py --help
# observed: "usage: main.py [-h] [--x_train_file X_TRAIN_FILE] ..." with the
# full flag list (plugin, epochs, iterations, load/save config, horizons, ...)Bounded CPU example using the bundled daily dataset (completed with exit 0 in the 2026-09-14 publication check):
CUDA_VISIBLE_DEVICES="" PYTHONPATH=. python app/main.py \
--load_config examples/config/phase_1_daily/phase_1_ann_1575_1d_config.json \
--epochs 2 --max_steps_train 300 --max_steps_test 300 --mc_samples 2This is a mechanics demonstration, not a validated trading experiment. It writes
the configured sample outputs; use a separate clone and fresh output paths for
your own work. Use long-form --flags; short forms do not consistently override
config values. Several 1h configs reference datasets absent from a fresh clone.
predictor.sh simply prepends the checkout to PYTHONPATH
and runs python app/main.py. Training data ship under
examples/data/ (organized by phase) and results are
written under examples/results/; batch drivers live in
examples/scripts/.
predictor is a DOIN domain: the external
doin-plugins package registers
predictor and binary_predictor optimization/inference entry points that
wrap this repository, and doin-node
— the unified participant runtime — runs them collaboratively (candidate
leasing, deduplication, champion migration and blockchain persistence are
doin-node's responsibility). predictor always works locally first; DOIN
extends its optimizers, it does not absorb them. The retired
doin-optimizer/doin-evaluator services are not required. OLAP/ETL helpers
for analyzing experiment databases live under olap/.
Configuration is a flat JSON merged over defaults in
app/config.py; every key can also be passed as a CLI flag
(app/cli.py). Plugins resolve via
app/plugin_loader.py from entry points declared in
setup.py:
| Entry-point group | Plugins (this package) |
|---|---|
predictor.plugins |
ann, cnn, lstm, transformer, tcn, tft, n_beats, mimo, plus binary_* and direction_* variants of each and binary_logistic/direction_logistic (predictor_plugins/) |
optimizer.plugins |
default_optimizer (DEAP GA), neat_optimizer (optimizer_plugins/) |
pipeline.plugins |
default_pipeline, stl_pipeline, binary_pipeline, direction_pipeline (pipeline_plugins/) |
preprocessor.plugins |
default_preprocessor, stl_preprocessor (preprocessor_plugins/) |
target.plugins |
default_target, stl_target, binary_target, direction_target (target_plugins/) |
Note on preprocessor.plugins: the preprocessors under
preprocessor_plugins/ are local to this repository
but registered into the shared preprocessor.plugins entry-point group
also used by gym-fx (which registers
its own default_preprocessor) and the standalone
preprocessor app. Co-installing
those packages mixes the group's contents, so prefer one environment per
application.
python -m pytest tests --collect-only -q
# observed: "3 tests collected, 8 errors in 3.21s"Known limitation of this standalone branch: much of the committed suite predates the
current plugin architecture and fails at import (e.g.
app.autoencoder_manager, load_encoder_decoder_plugins, merge_config no
longer exist). The count above is a recorded baseline observation, not a test
count for the newer research branch. Verified sanity check:
PYTHONPATH=. python -c "from app.plugin_loader import load_plugin; print('plugin_loader OK')"
# observed: "plugin_loader OK"A run writes: predictions CSV (output_file), aggregated metrics
(results_file), training/validation loss plot, optional model plot, the
trained model (save_model, .keras) with metadata JSON, and the fully
merged effective config (save_config) from which the run can be reproduced.
Optimization runs additionally write per-generation statistics, best
hyperparameters and a resumable population state under
examples/results/.
Training operates on local CSV files and requires no credentials. The CLI
retains optional remote config/logging flags (--username, --password,
--remote_log); never embed real credentials in configs or commit them.
Database helpers under olap/ connect to locally provisioned
databases only. Predictions are historical-data research output — not
financial advice and not a live trading system.
- The legacy pytest suite is stale (see Tests) — collection errors are expected until it is rewritten.
setup.pyinstall_requiresis minimal; installing withoutrequirements.txtyields a non-functional environment.- No
python_requiresor dependency version pins are declared. - Top-level package names (
app,*_plugins) are shared conventions across sibling repositories; use a dedicated environment or run from the checkout root (aspredictor.shdoes) so local packages win. - Some root-level artifacts (champion models, sweep scripts, logs) are
working files of ongoing campaigns; treat directories under
examples/as the stable interface.
- prediction_provider — FastAPI service serving models trained here
- feature-eng — feature/label generation upstream of training
- feature-extractor — autoencoder encoder/decoder training
- preprocessor — standalone CSV preprocessing app
- doin-node / doin-plugins — distributed collaborative optimization around this domain
- agent-multi / gym-fx — RL trading side of the stack
MIT — see LICENSE.txt.