Finding Alpha in Random Noise
Ausija Bhattacharjee · 29 September 2026
If none of your signals contain any information, why can one of them still produce a good backtest? A simple experiment in multiple testing, overfitting and out-of-sample validation.
Introduction
Suppose we run a backtest and find a trading signal with a Sharpe ratio of 1.60.
That would normally be enough to make it worth looking at more closely. It has produced fairly consistent returns over several years of data and, at first glance, seems to be picking up something useful.
In this case though, the signal is completely random.
The market it trades is random too. There is no relationship between the signal and future returns because the data has been generated so that none exists.
The question is then fairly simple. If there is no real signal in the data, how did we end up with a backtest that looks quite good?
This is a basic version of a problem that comes up constantly in quantitative research. A backtest tells you what happened in a historical sample. On its own, it cannot tell you whether you found a genuine relationship or whether you simply searched enough variations to find one which happened to work.
The market
We can start with a deliberately simple market.
Daily returns are independent draws from a normal distribution:
$$
r_t \sim N(0,0.01^2)
$$
There is no momentum or mean reversion built into the process. Yesterday's return gives us no information about tomorrow's.
We can generate 3,000 daily returns in Python:
import numpy as np
rng = np.random.default_rng(1)
T = 3000
returns = rng.normal(0, 0.01, T)
If we want something easier to look at, we can turn those returns into a price series:
price = 100 * np.cumprod(1 + returns)
The result will normally look like a perfectly believable financial time series. There will be stretches that look like trends, sell-offs and periods of relative calm.
None of those patterns were deliberately put into the data. They are just the result of random draws.
That is already part of the problem. A pattern being visible does not necessarily mean there is anything behind it.
The signals
Now we can generate 1,000 possible trading signals.
Each signal is another random draw:
$$
x_t^{(j)} \sim N(0,1)
$$
where $j$ identifies each individual signal.
For this example, treat $x_t$ as a signal observed at the start of day $t$, before that day's return is realised. It is generated independently of the return, so by construction it contains no information about what happens next.
We will use a very simple trading rule. If the signal is positive, go long. If it is negative, go short:
$$
s_t^{(j)} =
\begin{cases}
+1 & x_t^{(j)} > 0 \
-1 & x_t^{(j)} < 0
\end{cases}
$$
The strategy return is then:
$$
r_{s,t}^{(j)} = s_t^{(j)}r_t
$$
In Python:
N_SIGNALS = 1000
signals = rng.normal(size=(T, N_SIGNALS))
positions = np.sign(signals)
strategy_returns = positions * returns[:, None]
There is nothing hidden in these signals. Each one is independent of the market return.
So before doing any testing, we already know the answer. None of these strategies has any genuine predictive power and their expected return is zero.
Measuring them
We can compare the strategies using their annualised Sharpe ratios:
$$
SR = \sqrt{252}\frac{\bar r_s}{\sigma_s}
$$
where $\bar r_s$ is the average daily strategy return and $\sigma_s$ is its standard deviation.
For simplicity, I am ignoring the risk-free rate here.
def sharpe(x):
return (
np.sqrt(252)
* x.mean(axis=0)
/ x.std(axis=0, ddof=1)
)
We will use the first 1,500 days as our research sample.
split = 1500
in_sample = strategy_returns[:split]
in_sample_sharpe = sharpe(in_sample)
Now we just take whichever signal performed best:
best = np.argmax(in_sample_sharpe)
print(best + 1)
print(in_sample_sharpe[best])
With the random seed above, signal 894 comes out on top.
Its in-sample Sharpe ratio is approximately:
$$
SR = 1.60
$$
That is a fairly respectable looking result for something we know contains no information at all.
What happens when we search more?
We can make the effect clearer by changing the number of signals available to us.
Using the same generated data:
| Signals searched |
Best in-sample Sharpe |
| 1 |
0.210 |
| 10 |
0.915 |
| 100 |
0.992 |
| 1,000 |
1.600 |
The market has not changed anywhere in this table.
All we have done is increase the number of strategies we are allowed to search through.
With one strategy, the result is not particularly interesting. With 1,000 of them, we have managed to find one which looks much more convincing.
This is the basic issue with data mining in a backtest. Every additional variation gives us another chance to find something which happened to work historically.
That does not necessarily mean testing lots of ideas is wrong. Quant research obviously involves trying different specifications, parameters and models.
The problem is that the final result can look much stronger than the underlying evidence really is if we forget how much searching went into finding it.
Testing it on new data
We still have another 1,500 days which have not been used to select the signal.
Signal 894 was chosen entirely using the first half of the sample. We can now freeze that choice and see how it performs on the second half.
out_sample = strategy_returns[split:]
out_sample_sharpe = sharpe(out_sample)
print(in_sample_sharpe[best])
print(out_sample_sharpe[best])
The result is:
| Sample |
Sharpe |
| In-sample |
1.600 |
| Out-of-sample |
-0.244 |
The performance disappears.
In this example that is not particularly surprising. We already know signal 894 has no relationship with future returns.
It only looked useful because, out of 1,000 random strategies, it happened to produce the strongest result over the first half of the sample.
Once we stop searching and apply that same rule somewhere else, its performance moves back towards what we would expect from a strategy with no edge.
What actually went wrong?
There is nothing wrong with the mechanics of the backtest itself.
The simulated signal is defined before the return it trades, so the result is not coming from look-ahead. The returns and Sharpe ratios are also being calculated as intended.
The problem is how the result is being interpreted.
A Sharpe ratio of 1.60 would be unusual enough to attract attention if we had tested one fixed strategy.
But we did not test one strategy. We tested 1,000 and then deliberately selected the largest result.
Those are different situations.
We are not really asking how likely one useless strategy is to produce a strong Sharpe ratio. We are asking how likely it is that at least one strategy from a large group produces a strong result.
The more things we test, the easier it becomes for random variation to generate something that looks meaningful.
This is a simple example of multiple testing and data snooping, and it is one of the ways a backtest can become overfit.
Parameters count as searching too
We do not need to generate 1,000 completely different strategies for this to matter.
Suppose we start with a simple momentum strategy based on the previous 20 days of returns.
Then we try 10 days. Then 30. Then 50.
After that we change the entry threshold, add volatility scaling, adjust the holding period and perhaps change the group of assets being traded.
Eventually one version is likely to look better than the others.
Again, there is nothing inherently wrong with checking different specifications. You normally have to.
The issue is treating the final version as though it was the only model ever tested.
A parameter can be overfit in much the same way as a whole strategy.
One useful check is to look at the area around the chosen parameter. If a 20-day signal works, it would be reassuring if 18, 22 and 25 days produced broadly similar results.
If exactly 20 days works extremely well and almost everything around it performs badly, that would make me less confident in the result.
Out-of-sample testing
Keeping part of the data away from the research process is one of the simplest checks we can use, although a single train-test split does not solve the problem on its own.
The basic idea is:
$$
\text{research} \rightarrow \text{select} \rightarrow \text{test on unseen data}
$$
The important part is that the final data really is unseen.
If we check the out-of-sample performance, dislike the result, change the strategy and then keep going back to the same test period, we gradually start fitting to that period as well.
At that point it is no longer doing the job we originally set it aside for.
Depending on the problem, researchers can deal with this using simple train-test splits, walk-forward testing, different assets, different market periods or eventually live paper trading.
None of these methods prove that a strategy will work in future.
They just give us another piece of evidence which was not directly involved in building the original model.
Why the idea behind the signal matters
There is another fairly obvious problem with signal 894.
We have no reason to expect it to work.
That matters because statistical results are usually easier to take seriously when there is at least some plausible mechanism behind them.
If we were researching momentum, for example, we could ask whether underreaction or gradual information diffusion might produce some persistence in returns.
That does not prove momentum works. It gives us a hypothesis which can then be tested.
Our random signal has no equivalent explanation.
If we tried to invent one after looking at the backtest, we would simply be building a story around something that happened by chance.
That is another reason it helps to start with a hypothesis rather than only searching for whichever relationship happens to produce the best historical result.
Where this example breaks
Real quant research obviously does not consist of generating 1,000 random columns and choosing the best one.
Real strategies are also often closely related. A 20-day momentum strategy and a 21-day momentum strategy are almost the same model, whereas the random signals used here are much less connected.
The example is deliberately artificial because we know the true process behind the data.
There is no alpha to find.
That makes it easier to see something which is much harder to identify in real market data. A backtest can be completely genuine while the apparent relationship behind it is not.
Out-of-sample testing is not perfect either. A useless strategy can perform well in a second sample by chance, just as it did in the first.
The point is not that one train-test split solves the problem. It is that evidence becomes more convincing when a result survives tests which were not involved in creating it.
I have also ignored transaction costs, slippage and market impact here. Those would make most of these strategies worse, but they are separate issues. The selection problem appears before any of them are added.
What would make the result more convincing?
A strong backtest is useful, but there are a few things I would want to know before putting much weight on it.
Does the strategy still work on data which was not used to build it? Do nearby parameter values produce similar results? Does it work across more than one period or asset? Does it survive reasonable transaction costs? Is there some sensible reason why the relationship might exist?
None of those questions gives us certainty.
They do make it harder for a purely accidental result to survive the research process.
In that sense, finding a good backtest is only part of the job. A lot of quant research is then trying to work out why you should not trust it.
What to try next
-
Change the random seed and repeat the experiment. The winning signal and its Sharpe ratio will change.
-
Increase N_SIGNALS from 10 to 100, 1,000 and 10,000. Compare the best in-sample Sharpe each time.
-
Give one signal a small genuine relationship with future returns and mix it in with the random ones. See how strong the relationship needs to become before the research process reliably identifies it.
-
Replace the random signals with simple momentum or mean-reversion rules and vary their parameters. If the underlying market remains random, any apparent edge is still coming from sampling noise.
-
Try several train-test splits rather than relying on one. If a result only survives one convenient split, that is useful information in itself.
Part of Quant Foundations · Previous: Intro to Market Making
NEFS Quant · About · Blog · News & Events · People · Privacy policy