The value of testing each criterion in isolation?

I have, over some years, collected over 6500 criteria from various posts on the forum and in the free ranking systems.

I also have a ranking system for the USA and Canada that I am satisfied with and which I stress- tested a bit. Then I thought I would test each of the criteria in my XML by gathering as much information as possible for each of the 6500 criteria and then comparing it with my XML, using, infromation from among other methods, rolling tests, screen backtests, and bucket tests for each criterion.

Among other things, I wanted to test if I:

* Have many "misbehaving factors" (Yuval's article).

* If there is a high correlation between some of the factors.

* If it is significant for the overall system that I have a negative CAGR return when I run the single criterion in isolation.

* If the slope of the bucket and its stability have any bearing on finding new good factor combinations.

* If isolated ranking criteria with high CAGR result in exceptionally good systems.

Out of the buckets of 97 factors, only 52% of them are highest at the end or beginning of the 40 buckets; the rest are between 15 and 30. So, some are clearly "misbehaving", even though they constitute part of my ranking system. I will test the possibilities for adjusting them as Yuval addresses in his article.

Some of the criteria clearly have negative returns when tested in the backtest on the same universe with 50 stocks for the period 2006-2026, but are evidently a crucial part of the composition of my ranking systems, because the returns fall significantly if I remove them (and it wasn't just "Size" that had negative returns):

I have attempted to add several different criteria to see if the bucket slope and potential stability in the buckets have any impact on adding extra criteria that would lead to better systems, but for me, it hasn't helped.

Regarding the correlation, I have found 10 out of around 100 criteria that are highly correlated.

I also did some more testing, but it didn't improve my system:

#1 Most Important
L/S Spread + Monotonicity
Strongest return difference between top and bottom.
Spread >15% | Mon >70%

#2 Stability
Low IC-Instability
Ensures the factor worked just as well in more recent times.
|IC1 - IC2| < 0.04

#3 Transformation
Peak Bucket Placement
Top in upper part = Higher. Middle buckets = Method 2/3 or Buy Rules.
Exact Peak Target

#4 Risk-Free Compounding
Sharpe Ratio + Max DD
Avoids bad luck / ruin that destroys compounding.
Sharpe >1.0 | DD >-25%

#5 Rank Strength
IC Spearman + Slope
High and steep predictive power across the entire universe.
Spearman >0.05 | Slope >0.15

#6 Linearity
Bucket R² (Explanatory Power)
Low noise; returns follow a predictable line.
R² > 75%

#7 Net Efficiency
CAGR / Turnover
High turnover eats up theoretical CAGR via friction.
Turnover < 150%

#8 Asymmetry
Up vs Down Excess
Captures the upside and protects on the downside.
Up Excess > Down Excess

#9 Consistency
Skewness & Win Rate
Even monthly gains provide a smoother equity curve.
Skew > 0 | WinRate > 55%

#10 Uniqueness
Alpha vs Benchmark Correlation
True excess return independent of the general market.
Alpha >8% | Correl < 0.6

After all these tests, I believe that extracting a lot of information on how each individual criterion behaves doesn't necessarily say much about the significance it will have in a combined ranking system. Therefore, searching for good individual criteria is just the beginning of what is truly valuable here – the factor combination.

7 Likes