The COF Methodology: Research Center, Statistical Validation, and Order Flow
At cofiatrading, a strategy only exists if it survives honest statistical validation on unseen data. FR guide to our method: research center, holdout, DSR, and validated order flow — no promises, just process.
Erwin
Founder cofiatrading
Most of the « trading methods » sold to you rely on one thing: a beautiful backtest curve. The problem is that a beautiful backtest curve is easy to manufacture and easy to fake, often without even intending to. Optimize enough parameters over enough past data, and any noise will eventually look like an edge.
At cofiatrading, we start from the opposite premise. For us, a strategy only exists if it survives honest statistical validation on data it has never seen. No magic curve, no promise of returns. A research process applied with rigor, delivering its result: setups that have passed the filter.
In this article, I explain how we work: the research center, order flow as raw material, holdout separation, and the Deflated Sharpe Ratio to prevent self-deception.
Get an overview of validated setups →
The starting point: order flow, not classic indicators
We do not build our strategies on moving averages or recycled oscillators. We work with the raw material closest to market truth: order flow on centralized volume futures (NQ, ES).
Concretely, we decompose the flow into measurable primitives: delta, CVD, absorption, footprint imbalances, volume profile levels (POC, VAH, VAL), and behavior at the DOM. These are objective facts about who was aggressive and where volume traded — not subjective interpretations.
Why order flow? Because that is where you read the real action of major players before it appears in price. And because, on centralized CME data, it is measurable and backtestable properly — unlike Forex volume which does not exist reliably.
The research center: combine and mine, not guess
An isolated intuition (« absorption at POC works ») counts for nothing until tested. Our research center takes these order flow primitives and combines them systematically, then measures each combination on history.
The principle is simple: instead of starting with an idea we try to confirm (guaranteed confirmation bias), we let the data speak. We test a large number of condition combinations, keeping only those showing statistically credible edge — sufficient winrate, solid profit factor, and most importantly enough occurrences so it isn't luck.
A non-negotiable rule: a setup with 12 historical trades proves nothing. We require a minimum number of occurrences before looking at performance metrics. An edge must repeat itself; otherwise, it is an anecdote.
The first trap: overfitting
Here is the enemy this process fights against: overfitting. When you test hundreds of combinations on the same data, some will appear great by pure chance. Statistically, this is inevitable: search enough, and you will find noise that looks like a signal.
An overfitted strategy shines in the past and collapses in the future, because it merely memorized accidents in the test data rather than capturing real market behavior. This is reason number one why so many magnificent backtests do not survive live trading.
Two safeguards protect against this: holdout separation and the Deflated Sharpe Ratio.
The Holdout: testing on unseen data
The principle is simple and uncompromising. We split history into two parts:
- An training (in-sample) part, where we search and optimize.
- A holdout (out-of-sample) part, set aside and never touched during the entire research phase.
A strategy is only validated if it holds up on holdout — data it has never seen. If a setup shines in training but crumbles in holdout, it is overfitting: it is rejected, no matter how beautiful its training curve was.
It is frustrating: many « good ideas » die here. But that is exactly the goal. Better to kill a false strategy on holdout than on your real capital.
The Deflated Sharpe Ratio: correcting for number of trials
The classic Sharpe Ratio measures risk-adjusted return. Its problem: it does not account for the number of strategies tested. If you try 500 combinations, the best one will have a nice Sharpe even if all are noise — simple effect of large numbers.
The Deflated Sharpe Ratio (DSR) corrects this. It « deflates » the Sharpe based on number of tests performed to answer the real question: is this result credible given everything we tried, or is it just the best draw from a lottery? A setup only passes if its edge remains significant after this correction.
It is our last honest filter. It prevents us from convincing ourselves that we found an edge when we simply searched enough.
What we deliver: the result, not the factory
You do not need to code a mining engine, manage holdouts, or calculate DSRs to benefit. Our work is this process; your benefit is the result: order flow setups that have passed all these steps — sufficient occurrences, held on holdout, significant DSR.
What we do not promise, and never will: guaranteed returns. No method, however rigorous, eliminates loss risk. What rigor brings is to play setups with honestly measured probability rather than faked curves. It is a probabilistic edge applied with discipline — not a money-making machine.
The difference between our approach and an average « method » lies here: we show you how we eliminated our own illusions, not just the curve that survived.
Key takeaways
The COF methodology starts with order flow on futures (reliable CME data), decomposes market into measurable primitives, and combines them in a research center letting data speak. Three safeguards against overfitting: minimum occurrences, validation on holdout (unseen data), and the Deflated Sharpe Ratio (correcting for trials). We deliver the result — validated setups — not the factory. No promise of gain: an honestly measured probabilistic edge applied with discipline.
Trading involves risk of capital loss. Educational content, not investment advice.