Skip to content
HelmFi

Strategy graveyard

The failures
stay on record.

What each strategy claimed, what testing showed, and why it was rejected. Every entry comes from the published research ledger.

18 / 18 records
RejectedKalshiEVTakernegative EV

The claim

EV-gated taker on crypto binaries: cross the spread only when modeled probability edge clears fees, sized by EV. Modeled edge suggested selective positive expectancy on funding/momentum-informed probabilities.

What testing showed

Live paper forward test on the research box: 42 closed positions, 1 win, -588 bps realized. After the entry gate was tightened it stopped trading entirely - every scanned market shows negative EV at the top of book.

Why it was rejected

The taker edge does not exist. Crossing the spread on Kalshi binaries pays the market maker: fees plus adverse selection consumed the modeled edge, and the model's probabilities were not better than the book's. The EV gate did its one job - it stopped the bleeding and proved the strategy empty. Maker-side variants remain in paper with near-zero fills.

Kalshi crypto binary markets (BTC/ETH/SOL/LTC + alts) · 1h
RejectedPolymarketMakerQuoteV1adverse selection

The claim

Execution-first maker strategy: post passive YES/NO bids inside or at the best bid, avoid taker fees, and profit from pair-spread capture when both sides fill.

What testing showed

Pre-registered Stage 1 paper test hit 102 resolved fills across 76 markets: 53W/49L, but -$26.62 PnL and -9.1% ROI before maker rewards.

Why it was rejected

Adverse selection. Both-side fills proved the pair-spread mechanism exists, but one-sided fills were toxic enough to overpower spread capture. The strategy won slightly more often than it lost but still lost dollars.

Polymarket BTC Up/Down · 5m
RejectedPolymarketWalletFollowcopy degradation

The claim

Screened 6,349 public wallets across recent BTC 5m markets and found 25 BUY-heavy candidates with positive historical after-fee ROI.

What testing showed

Pre-registered paper-follow test: 100 resolved copy trades after lock, 59 markets, 10 leaders, 60.0% WR, but -$339.62 PnL and -13.2% ROI after fees and copy degradation.

Why it was rejected

Copy degradation. The public-data follower copied the right side often enough, but leader fills could not be replicated: +$0.02 copy padding, taker fees, and 5-minute market latency turned a 60% WR into negative ROI.

Polymarket BTC Up/Down · 5m
RejectedMomentumProstatistical noise

The claim

Sharpe 0.89, p=0.013 (alpha vs gated null) on 2024–2026 OOS window

What testing showed

IS Sharpe 0.36, p=0.999 (literally every random sign-flip beats the observed alpha)

Why it was rejected

Recency bias. Published number was the 2024–2026 window only. When the 2022 bear is in the test set, the strategy's aggressive-momentum tilt hurts far more than it helps.

13 Kraken USD majors · 5m–1h
RejectedMomentumCorestatistical noise

The claim

Sharpe 0.67, +118% total return over 2021-2026 on the live universe.

What testing showed

Bootstrap 95% Sharpe CI [-0.19, +1.56] CROSSES ZERO — the edge is not statistically distinguishable from zero. MaxDD -45.4%, permutation alpha p=0.066. All 4 walk-forward windows non-negative, but it fails the first gate.

Why it was rejected

Statistical noise. A positive headline Sharpe with a bootstrap CI spanning zero is exactly what the validator's first gate exists to catch. Live-wired but never catalogued; retired 2026-05-29 after its first formal validation run.

11 USD majors (cross-sectional momentum + vol targeting) · 1h, 72h rebalance
RejectedE0V1E_DCAhyperopt overfit

The claim

Aggregate OOS Sharpe 2.70 across 4 walk-forward windows

What testing showed

Decayed to Sharpe -0.09 in the most recent (2025-04 → 2026-02) window

Why it was rejected

Hyperopt overfit. Aggregate-Sharpe number hid a single catastrophic window — the validator's window-by-window check caught it. The base E0V1E without DCA passed at the time and was later superseded by the Helm Index.

27 majors · 5m
RejectedSMAOffsetTunedhyperopt overfit

The claim

Aggregate Sharpe 2.67

What testing showed

Window 2 (Jun24–Apr25) Sharpe -2.09

Why it was rejected

Same pattern as E0V1E_DCA. Hyperopt found parameters that looked great on aggregate but lost money in one specific regime. The un-tuned SMAOffsetProtectOptV1 remains a research candidate and has not cleared the live gate.

14 Kraken USD majors · 5m
RejectedAggressiveScalpV2fee drag

The claim

Gross-of-fees roughly break-even, looked workable

What testing showed

30-day live test: 348 trades, 18.1% win rate, -73.86% net, $15,677 in fees on $10K capital. 22 max consecutive losses.

Why it was rejected

Fee drag, not edge. Avg trade -0.77%, round-trip fee 0.80%. Kraken's 0.4% taker × 2 ate the entire signal. We then retested on OKX with a regime gate; bull window still produced 1,833 trades at 33% WR / -2.06% — the entries themselves don't predict direction.

14 Kraken USD pairs · 5m
RejectedAggressiveTrendV2no edge

The claim

1h trend-pullback with BTC 4h regime gate, ATR/RSI/EMA stack — looked like the right answer for the 'aggressive tier'

What testing showed

12 months / 14 pairs: 44 trades, 25% WR, -24.13% net, profit factor 0.40. All 4 walk-forward windows negative including Q1 (-5.10% in a +7.20% market).

Why it was rejected

Moving timeframe from 5m → 1h didn't fix the entry-quality problem. The gates produce mostly-loser signals. Structural finding extended from 5m to 1h on Kraken USD majors.

14 Kraken USD majors · 1h
RejectedAggressiveDealBotV1architecture flaw

The claim

Bounded DCA / safety-orders / 66.2% win rate / beats hold by 30 pp in a -46% market

What testing showed

12 months: 77 trades, 66.2% WR, -15.71% net, PF 0.56. Critical: ZERO DCA fired across 4 gate configurations.

Why it was rejected

Architecturally a single-entry deal bot with risk-off kill switches in current regime — not a DCA strategy. The safety gates that prevent catastrophic losses also prevent the DCA mechanic from operating. Edge isn't where the design claimed it was.

14 Kraken USD majors · 1h
RejectedAggressiveV2Engineregime shift

The claim

+7.58% Sharpe 2.10 — a bull window that made the strategy look ready to ship

What testing showed

11-window walk-forward: only 4–6 windows pass (36–55%). The original number was 2024 H1 only. Most recent 3 windows (2025 H1, 2025 H2, 2026 Q1) all lose money even with regime gate.

Why it was rejected

Cherry-picked window. Mean-reversion + small TP only works when prices bounce. The 2025+ regime has been grindingly directional in ways that fill all safety orders without hitting the +1.8% TP. Engine mechanics are fine; the bouncy-uptrend assumption isn't.

BTC, ETH, SOL on Kraken · 5m
RejectedichiV1lookahead bias

The claim

Sharpe 16, near-perfect equity curve

What testing showed

Rejected before walk-forward could run

Why it was rejected

Lookahead bias. Ichimoku's senkou span B uses 26-period future data — the indicator reads tomorrow's price to confirm today's signal. Famous failure mode in technical-analysis backtests; the equity curve is fake by construction.

Majors · 1h
RejectedApollo11parameter broken

The claim

Looked like a trend-following candidate

What testing showed

Sharpe -5.35. All trades exit at trailing_stop_loss, avg -$3.50.

Why it was rejected

Every entry trail-stops out. The trailing-stop parameters are mis-tuned for 5m volatility — the noise floor is wider than the trail distance. Strategy literally cannot profit by construction at these settings.

Majors · 5m
RejectedGumbo1code bug

The claim

Community-published strategy with reportedly-good numbers

What testing showed

Sharpe -53

Why it was rejected

Code bug in the published source — buy condition fires every bar (sell condition uses OR instead of AND). Attempted patch didn't fix it. Pattern observed across multiple 2020-2022 community strategies: old API signatures, broken signal logic that never matched the published equity curve.

Majors · 5m
RejectedReinforcedSmoothScalpfee drag

The claim

Smooth equity curve in the published 1m backtest

What testing showed

Sharpe -13

Why it was rejected

1m timeframe + 0.4% taker fee = guaranteed loss. At 1m, signal edge per trade is typically <0.2%; the fee alone is 4× larger. Same lesson as AggressiveScalpV2: high-frequency taker strategies on Kraken are structurally infeasible.

Majors · 1m
RejectedCryptoFrogdrawdown

The claim

Sharpe 0.42 across full backtest

What testing showed

2 negative walk-forward windows, 51% max drawdown

Why it was rejected

Sharpe was marginally positive on aggregate but the validator's per-window check found two losing windows including a 51% drawdown. A live strategy with that drawdown profile is unshippable regardless of aggregate Sharpe.

Majors · 5m
RejectedFreqAI defaultml template trap

The claim

ML feature template marketed as 'just turn it on' for FreqAI

What testing showed

Lost 98.59% in backtest

Why it was rejected

The default FreqAI strategy template is not a working bot — it's a starting skeleton intended to be heavily customized. Running it as published with default hyperparameters produces a near-total wipeout. Lesson: ML strategies don't autosolve; the template-on-default failure mode is on the user, not on FreqAI.

Majors · 5m
RejectedNostalgiaForInfinityX7bag holding

The claim

Catalog listed the conservative tier at Sharpe 2.12; widely treated as proven NFI alpha

What testing showed

Sharpe 0.14, permutation p=0.42 (no better than random), one negative walk-forward window, worst single-trade floating drawdown -89.3% (22% of trades >10% underwater)

Why it was rejected

Honest spot-only validation (BTC/ETH/SOL 2021-2026, run in yearly chunks on a dedicated box because the monolithic 5-year backtest was intractable) failed all three gates. The apparent stability is grinding/bag-holding: NFI averages down and sits up to 89% underwater on open positions while booking small closed gains (+3.1% total over ~5 years). A new floating-drawdown check exposed what closed-trade Sharpe hid; permutation p=0.42 means the edge is indistinguishable from random.

Majors (BTC/ETH/SOL) · 5m

Read the test alongside its limits.

The appropriate benchmark and decision rule depend on the strategy. A historical result is evidence to inspect, not a promise.