No. of Recommendations: 11
Disclosure up front: I ran this thread through Claude and asked it to check the results rather than extend them. It uses only numbers already posted in this thread so no independent backtesting was done.
To be clear about intent- this isn't meant as an attack on anyone's work. There's been a lot of genuine effort in this thread and I've learned from all of it. I found Claude's critique interesting and thought others might too. Please take what's useful, ignore the rest (or all of it) — no issue either way.
Worth noting how I prompted it, since that shapes the output. I gave it a persona as a quantitative researcher instructed to be strictly objective, to flag look-ahead bias, curve-fitting and unrealistic execution assumptions, and to state flaws directly. A neutral "what do you think of this thread?" would almost certainly have produced something more balanced and much gentler. So treat the tone as a product of the instructions.
What the prompt didn't change is the math. The outlier years, the error bars, the numbers that don't match each other post to post. All of that comes from figures already posted here, and anyone can check it regardless of how I asked the question.
If someone wants to run the opposite prompt and post what comes back, that'd be worth seeing. Here is Claude's take:
1. Four years are the strategy.
Take lizgdal's annual table (post 5611). 1991 (+156.6%), 1998 (+104.5%), 1999 (+185.9%) and 2020 (+118.9%) compound to 32.8x by themselves. Back those out of a 41-year record at ~23% CAGR and the remaining 37 years run at roughly 13-15%, against ~11-12% for the S&P over the same span.
So the honest description is: a 2-3 point gross annual edge, plus four Nasdaq melt-ups. Every one of those four years is a tech inflation episode, and two of them are immediately followed by -40% and -25%. The tail that produces the headline CAGR is the same tail that produces the -66% drawdown. lizgdal called the distribution correctly — log-normal with fat tails — and the fat tail is the result.
2. The error bars swallow every parameter choice made here.
At ~30-35% annualized vol over 40 years, the standard error on CAGR is about 4.7 percentage points. The entire spread between the tested timing variants (22.9% to 27.4%) fits inside one SE. On Sharpe, SE ≈ 0.20 for ~32 years; after a few hundred configurations, the expected best-of-noise sits roughly 0.65 above the true value. An observed 1.06 is fully consistent with a true Sharpe near 0.4 — below SPY's reported 0.49.
Count the knobs turned in this thread: lookback (6/8/9/12/13/15mo), N (1-10), HTD (6/7/10/12/15), timing index (IXIC/NDX/GSPC/SPY/QQQ/ONEQ), SMA length (200/252/325/43wk/52wk/58wk/64wk/65wk), sell band (0/-1/-2/-4%), signal frequency, trade-day offset (-6 to +1), three ADV filters at two thresholds each, trailing stops, a second ranking metric, four ways to combine screens. That's a search space in the hundreds of thousands with, conservatively, several hundred configurations actually evaluated. Nobody recorded how many were rejected, so the number needed to deflate these statistics is unrecoverable.
The timing overlay deserves its own note: it's being fit to 3-5 independent bear markets in 41 years. rayvt says exactly this — "hard to make predictions on something that only happened 3 times in 41 years" — and then reports 64wk/-4%/GSPC as "best in Sortino, STDEV and MaxDD." At n≈4, nothing in that grid is distinguishable from anything else in it.
3. Three results that can't be right.
Day-of-month. -2 vs +1 at 27.61% vs 21.24% post-1999 implies ~50bp/month of free return from shifting execution one trading day. Documented turn-of-month effects are single-digit basis points. An effect this size would be gone in weeks. Either a bug, or a handful of extreme observations landing on one side of a boundary. rayvt's 3-in-21 point stands unanswered.
Known look-ahead, still cited. rayvt found that gprc(1) → gprc(2) moved CAGR 250bp while the next day of lag moved ~10bp. That asymmetry is the textbook signature of signal leakage. He posted in 5698 that the overlaps backtest is wrong. The 32.2% figure from those same links was quoted again in 5799.
Irreproducibility. Same strategy, same period (Oct 93–Apr 26), same timing rule appears as 19.8% (5596), 21.75% (5608), and 27.94% (5627). Post-1999 appears as 18.3%, 19%, and 21.24%. If a pipeline can't reproduce itself within 800bp, no output from it is usable — independent of whether the underlying Norgate data is clean.
Related: the reported figures climb monotonically — 25.6 → 27.6 → 29.1 → 31.8 → 32.2 → 36 → 38 → 41.8 → 43.2 → 45.4 → 48% — over about five weeks of iteration on one dataset. Genuine discovery doesn't do that. Search over noise does.
4. Component-by-component.
Component Verdict
9-12mo cross-sectional momentum Likely real, modest — supported out-of-sample elsewhere, and insensitive to lookback within 8-12mo in their own data
Long-MA trend filter Likely real for drawdown, not return; -66% → -33% is consistent with the literature
Specific SMA parameters Noise — n≈4 events; participants using different vendors get opposite answers
Concentration to 5 names Not alpha — raises vol from ~18% to ~35% while Sharpe barely moves. This is ~2x leverage by another name
PHL vs RS Redundant — rayvt notes ~80% overlap. Two proxies for one quantity
Overlap / sum-of-ranks Noise-mined — four combination rules tried, best reported; 20% of months produce zero picks
ADV volume filters Rejected by their own test — 50.7% hit rate, 5 names = 62.6% of benefit
35% trailing stop Noise — +0.12 CAGR, one of many thresholds tried
Buying more of existing winners Untested and drifts toward a single-name book; rayvt's objection is correct and unanswered
5. Two things being modeled wrong.
Friction. 0.1-0.25% doesn't survive contact with pre-decimalization Nasdaq dealer spreads — and that's precisely where the outperformance lives. On 5 names with near-total monthly turnover, realistic pre-2001 round-trip costs imply 10-20 points a year, not 1-3. The "decay since 1999" everyone has noticed is substantially this. rayvt supplied the live counterexample himself: he couldn't get filled on LITE across multiple limit bumps while it ran +10.9% on the day.
Tax. A -4.8 point drag on 25.6% implies an effective rate near 19% on near-total short-term realization. Realistic drag at top marginal rates is 8-10 points, which consumes the entire remaining edge in taxable accounts. rayvt's flat 25% haircut is the more defensible assumption.
6. The benchmark, and the right way to frame the whole thing.
A 35%-vol strategy shouldn't be compared to unlevered SPY. The honest hurdle is vol-matched: roughly 2x QQQ with the same trend filter, financed. Against that, most of the Sharpe advantage disappears. "3x SPY with lower beta" is a category error — low beta at double the standard deviation means the risk is idiosyncratic and undiversified, which is worse. And the current top-12 is memory/storage/semicap top to bottom: that's one industry cycle, n=1 in factor terms.
rayvt asked the right question in 5503 — why hasn't this been arbitraged away? — and moved past it too fast. The defensible answer is limits to arbitrage: no institutional allocator can hold 5 names at 35% vol through 30-month cash periods and -60% drawdowns, so a premium can persist. But that reframes the finding correctly. It's compensation for accepting risk nobody else will accept, not an inefficiency. And it caps the honest expected excess return at about the 2-3 points that survive the decomposition in section 1.
7. Two research processes are running in this thread, and they aren't equivalent.
The thread reads as one collective effort. It isn't, and the difference matters more than any individual result.
One process leaves artifacts. Spreadsheets with stated formulas, GTR1 links others can run and break, explicit friction assumptions, explicit test periods. Crucially, it produces self-corrections: the gprc lookahead was found and reported, the daily-BCC signal was flagged as unrealistic, and favourable results were repeatedly talked down rather than up — "just a pattern in randomness," "the more rules there are, the more likely they're tuned to that one history." Those disconfirmations are the evidence that the process has a functioning error-detection step.
The other process leaves no artifact. Good raw data, paid for, and two genuinely worthwhile structural questions raised — whether combining two rankings beats either, and whether letting winners run beats rebalancing. But there's no code, no reproducible specification, errors corrected conversationally mid-stream, the same configuration reported at three materially different CAGRs in three posts, and a figure that had already been demonstrated contaminated re-quoted two weeks later. The self-descriptions are accurate and are the problem: "hundreds of hours of conversations with it," "I am constantly trying to improve the strategy." That describes a very long search reporting its best draw at each stage, with the rejected variants unrecorded. Without the rejection count, no reported statistic can be deflated, and every number from it is uninterpretable.
This isn't about effort or good faith; the effort is obvious. It's about what output can support. A result nobody else can reproduce — including the person who generated it — can't be relied on regardless of whether it happens to be true. The correct use of that work is as a list of hypotheses for someone to test properly against a pre-registered split, not as a set of findings to allocate against.
The practical test for anyone reading: for each number quoted in this thread, ask whether you could reproduce it from what's posted. Where the answer is yes, weigh it. Where it's no, treat it as a suggestion of something to check, not as a result.
What would settle it: fit every parameter on 1985-2005 only, apply mechanically to 2006-2026, report whatever comes out. rayvt proposed exactly this in 5796 and then set it aside. It's the one test that hasn't been run.
What survives: momentum on a large-cap universe with a long-MA filter is probably worth a few points a year gross over the index, at roughly double index vol, with -45% to -65% drawdowns and multi-year stretches in T-bills. Defensible at rayvt's suggested 10-25% allocation, in a Roth. Everything above that baseline in this thread is the researcher's fingerprint, not a property of the market.
Worth noting that the most trustworthy output in 140 posts is rayvt's rolling-window tables (5519, 5602) — that's the correct way to present a strategy like this, and it describes something considerably less attractive than the headline numbers. As does his own line: "I am highly doubtful that this 22%+ CAGR is realistic going forward."