Open research · La Forja de Oro

We tested 406,155 strategies. Almost none survived.

We read 398 trading books cover to cover, coded the twenty most-cited strategies in the craft and put them through the hardest test we know. We publish what we found — above all, what did not work.

406,155tests executed
398books read in full
4universes tested
23years of data

What did not survive

Each technique was coded exactly as its original source describes it, tested across decades of data, and then penalised statistically for how many variants were tried. Not one of these was left standing.

That list is, almost exactly, the catalogue every signal group sells. We are not saying it is impossible to trade them. We are saying that when we tested them with statistical honesty, the outcome was indistinguishable from chance.
Assay tray: dozens of blackened metal test samples and only two pieces of polished gold
The most honest way to summarise the work: you put many samples through the fire and almost all of them come out black. The two that shine are the only ones that withstood the full test.

The result that suits us least

On gold, forex and crypto: none survived

This is the market we trade, so it is the finding that binds us most. We tested the twenty textbook strategies on gold, the major currency pairs and the ten most liquid cryptocurrencies — on the daily chart and intraday too, down to fifteen-minute candles.

Not one survived the statistical correction. The only two that did withstand the full test work on US equities on the daily chart, and their edge is one of controlled risk — not of beating the market.

We are not publishing this because it suits us. We are publishing it because it is what we found, and because a group that only publishes what went well is not showing you its research: it is showing you its marketing.

How we tested it

Almost any strategy looks profitable if you test it badly. These are the defences we used, and they are the most important part of the whole exercise.

Rolling window

Every parameter is chosen using only data from before the period being measured. Nothing is ever optimised on the same stretch it is later judged on.

Rules frozen in advance

The selection criteria were written down and sealed with a cryptographic signature before a single result was seen.

Correction for attempts

Given enough tries, something always looks brilliant by pure luck. Every result is penalised according to how many variants were tested.

Sealed data

Two years of market history that no program and no person touched during the entire research, opened exactly once at the end.

Adversarial review

Independent reviewers whose job was to refute every finding. Several of our own mistakes were caught this way, and they are documented.

Reproducibility

Every test re-runs and produces exactly the same result, verified across a sample of more than thirteen thousand tests.

What you can do with this

The use of this research is not that we sell you something. It is that you now have concrete questions to put to anyone offering you a strategy — ourselves included.

If you found this useful, the guide to spotting fake signal groups develops the same ideas from the consumer's side.