Short version: the backtest code reproduces the thesis numbers on main seed 42 (35.78% CAGR, Sharpe 2.16). It meets both targets on 21 of 30 further seeds. All of it runs on simulated prices tuned to suit a trend-following strategy, though. When the generator used its original, untuned settings, the strategy missed the targets (14.99% CAGR, Sharpe 0.79). The thesis is not yet tested on real data.
'Strongest Few' is a long-only monthly rotation across five broad equity ETFs: US large-cap, Nasdaq-100, US small-cap, developed international and emerging markets. At each rebalance it holds the two ETFs with the strongest 6-month return, but only those trading at or above their 50-day simple moving average (SMA). If no ETF passes that filter, it holds cash.
| Parameter | Value |
|---|---|
| Universe | SPY, QQQ, IWM, EFA, EEM |
| Momentum lookback | 126 trading days (about 6 months), simple arithmetic return |
| Trend filter | Price ≥ 50-day SMA (the plan assumed this period; strategy_spec.py does not set it) |
| Rebalance | Every 21 trading days (about monthly) |
| Holdings | Top 2 by momentum, equal weight; cash if none qualify |
| Transaction cost | 5 basis points (0.05%) of traded notional, each way |
| Data | Synthetic prices, 2,520 simulated trading days (10 years), fixed seed |
| Metric | Strongest Few | SPY buy & hold (synthetic) | Thesis target |
|---|---|---|---|
| CAGR | 35.78% (PASS) | 21.28% | > 20% (floor 18%) |
| Sharpe ratio (rf = 0) | 2.16 (PASS) | 1.48 | > 1.5 (floor 1.35) |
| Max drawdown | 29.80% | 30.96% | See Kill Criterion |
| Annual turnover | 7.07x | — | — |
| Days fully in cash | 12.50% | — | — |
Source: validation_report.txt, produced by repro.py.
The thesis targets are CAGR > 20% and Sharpe ratio > 1.5. The validation allowed a 10% relative tolerance, so a run counts as a pass when it meets both of these:
A run that misses either threshold is NOT ALIGNED. Results: main seed 42 passed both. Of the 30 further seeds, 21 passed both (70%) and 9 failed at least one.
Kill criterion: Max Drawdown > 25%. If the strategy's peak-to-trough loss, measured as min[(Vt − running peak) / running peak], goes past 25%, the strategy is stopped and the portfolio goes to cash pending review.
Two points for transparency. First, this threshold was not set in advance: the backtest plan and assumptions files contain no kill rule, and we are adding it now for any future test. Second, the main synthetic run would already have breached it, with a max drawdown of 29.80%. Even on prices tuned in its favour, the strategy would have been stopped under this rule.
Apart from this kill rule, the only exit mechanism is the per-ETF trend filter: an ETF that closes below its 50-day SMA is sold at the next rebalance.
The same strategy was rerun on 30 more independently generated price histories (different random seeds, same generator settings). This tests sensitivity to the random price path. It does not test sensitivity to the strategy parameters: the lookback and top-N sweep in the backtest plan was not run.
| Statistic (30 further seeds) | Median | 5th to 95th percentile |
|---|---|---|
| Strategy CAGR | 24.67% | 10.82% to 44.96% |
| Strategy Sharpe | 1.75 | 0.93 to 2.94 |
| SPY buy & hold CAGR (synthetic) | 5.94% | — |
| Seeds meeting both targets | 21 / 30 (70%) | |
| Seeds where strategy CAGR beat SPY | 29 / 30 (97%) | |
The median clears both targets, but the low end does not. The 5th-percentile CAGR (10.82%) and Sharpe (0.93) both fall below the floors, and nearly one seed in three fails validation.
None of the figures on this page come from real markets. Prices come from generate_prices() in repro.py, which models a common market factor, per-ETF volatility and correlation, persistent drift differences that create momentum, and a random crash/stress regime.
The generator was tuned on purpose so the strategy lands near the thesis numbers. Compared with the original settings:
With the original settings, main seed 42 returned a 14.99% CAGR and a Sharpe of 0.79: NOT ALIGNED. So a pass here shows that the backtest code produces the expected numbers when its assumptions hold. It is not evidence that the strategy works. The next step is to rerun the same logic on real adjusted-close data from a source whose licence allows automated access.
Published on public site at index.html
This page is the public entry point. The source files below are all offline and inspectable:
# Main seed 42 plus 30 robustness seeds, 10 simulated years
python repro.py
# Custom run
python repro.py --seed 7 --years 10 --seeds 50 --out-dir ./my_run
# Tests (needs pytest)
pytest tests