Does Past Performance Predict Future Returns: A look at designer models

To a large extent this is what was easy with the designer model downloads and not necessarily the best method for predicting returns. Still…

There is an interesting discussion about Backtests here: Why the poor live results versus backtesting?

This is related but a slight different topic. So I started a new thread. Summarizing the previous thread, and maybe over-simplifying when doing so, it may be fair to say there is a complex and not fully understood relationship between in-sample and out-of-sample results. Some have suggested that backtested results are not important for predicting future returns. I tend to agree with that.

In any case, that leaves open the question: How do we pick a model? Maybe we want to look at the relative performance of backtests, use validation sample results,, test sample results or even just look at the out-of-sample results of a live port for picking a model?

So lets try the last. I had Claude Fable 5 help me with this. But with download from the designer models of the 1-year and 5-year returns we already have the 1-year returns and the previous 4-years can be calculated from that data. And a clean graph and R^2 or coefficient of determination can be obtained.

Conclusion: Somewhat surprising to me. Enough to be without words at this time. Maybe someone else has comments. Maybe the possibility of survivorship bias is worth mentioning, but otherwise without comment:

My uneducated / unvalidated guess is that the comparison shows that, taken as a group on average, the designer models tended to strongly match the overall market in the latest year’s return. During today’s market, SPY is up 20.26%, RSP is up 16.24% over one year.

It might be helpful to break the models into groups defined by the index they compared themselves to. if there are enough models in any group to make a valid comparison to their index, that analysis might show how difficult or easy it is to improve on stock returns in that segment of the market. Just my thoughts…

Great point!!! I agree that excess returns are better. That is hard to get from the Designer model downloads, however. For 5 years (or longer) that is not possible. But the models are being compared to themselves here, making correlations of one model to itself over different periods valid, I think.

I find essentially zero correlation and essentially zero predictive power from one period to the next in this study interesting.

I would really like to look at the correlation of backtest results to out-of-sample results too, but the designer model dowload does not have enough data for that.

You left the most important thing out. You pick a model by reading the designer's description of it, combining a healthy dose of skepticism with your own opinions about what's likely to work.

I spend a lot of time perusing the Designer Models market.

I'd argue that even more important than reading the description is getting to know the designer — through forum posts, direct messages, email, maybe a call. What I'm gauging is how well they grasp key concepts and how open-minded they are to competing, conflicting, or new ideas and factors. I spend an inordinate amount of time noodling on portfolio construction, hedging, and related concepts — probably to an unhealthy extent — and I always seek out designers who are similarly borderline obsessive about factors and model construction.

Two semi-related thoughts:

1. Can we do away with the 3-month incubation period?

I get the intent, but three months doesn't vet or prove anything — and the live record accrues from launch date, whether or not anyone can subscribe, so the gate adds no informational value. It just delays willing subscribers. There are at least two incubating models right now, from designers I know and trust, that fit a need I have; I'd subscribe to both today. I often subscribe to a model just to get a feel for its holdings, turnover, etc., and see how livable it would be for me — sometimes I trade it, sometimes I watch for a while. Put a warning on incubating models if you must, but let us subscribe. If the real purpose is to keep designers from launching many models and promoting only the winners, there are better ways to get there (e.g., the previously discussed idea of perpetually tracking all of a designer's models, including deleted ones).

2. Why not let us license ranking systems the way we license Designer Models?

Opt-in for designers, priced however they like: locked, nodes hidden, usable in a subscriber's own screens, sims, and books — with whatever safeguards designers' IP requires. It seems like a natural extension of the DM trust model.

I may be misunderstanding how some of this works behind the scenes. If so, apologies in advance — I'd welcome the correction.

“three months doesn't vet or prove anything”

Before the three month rule was instituted, we had designers who overfit systems to maximize backtests. The 3 month rule eliminated it so well that you don’t even see a problem anymore.

I would be very keen to join such a marketplace as an author. Crucially, this approach could circumvent the potential requirement for an investment adviser licence, which can sometimes be an issue when offering a Designer Model.

Right. Backtests don’t really have much, if anything. to do with out-of-sample results.

But that is not a trivial amount of data in the first post in this thread (n = 105). I would definitely like to see more data, and I do not claim this is the final word. I would be nice if we had more data on excess returns for Designer Models for example. Or if could use different periods. Or if out-of-sample results were separated from sim results.

The 4-years to predict the next year is just one study. Affected by not having excess returns or a weird market perhaps the last year perhaps. But even with factor inversion one would expect a negative correlation. These results are pretty remarkable no matter how you try to explain them away.

With the data we have, it is not clear that out of sample results can tell you how a model is going to do going forward, either. There is zero evidence for that.

Let me say ahead of time that there is probably some data, somewhere to show that out-of-sample results will help us pick the better models..I expected to be able to present that study and quantitate how much correlation there actually is (expecting a non-zero number). But the data did not cooperate.

Not seeing the problem IS the problem.

The incentive logic I'll concede — no payday until a live record exists, so overfit-and-market stops paying. That kills my idea of letting us subscribe anyway; the gate should stay.

But the noise was the signal.

A designer can still launch several models, quietly deleting whichever incubate badly while promoting the one that got lucky. The rule arguably raises the payoff, because the survivor exits wearing the credential we hold in highest regard — a live record — even though we all know a 3-month record is statistical noise.

Before the rule, failures at least had witnesses, perhaps victims. Now, a model deleted in week eleven of incubation "never existed."

Which circles back to where I started: picking models by vetting designers. I have less information about how a designer launches, promotes, and abandons models than I otherwise would have — the exact signal I most want to read.

I know there's been talk of a permanent, per-designer record of every launched model — deleted ones included — which may help close the gap.

Without it, we're trading a visible problem for an invisible one, the worst kind to have. I want as much information as possible.

One can vet designers. But there is no longer any story. Marc Gerstein’s models had a story.

But now the story is: “I started with 300 to 600 features of all different types including growth, value, analyst data and even R&D, etc…” The weights of these features come from an algorithm. We can decide if dividing universes using mod() or or bootstrapping is better. Whether randomly adding 2.5% to features to optimize, or using a machine learning algorithm is better. Maybe you like genetic algorithms. That is fine.

But there is seldom a story to the models any more.

One thing that seems pretty certain is that we can see some pretty impressive results in backtests and in periods of one year. Numbers we keep chasing. But the best performer for 5 years has less than 30% CAGR with 105 models having 5 or more years of out-of-sample data. An n of 105 is not small. 30% CARG is probably a realistic ceiling for long-term results. Or at least this should be the base-case (or prior for those who like Bayesian statistics).

I have thought 40% CAGR may be a possible ceiling, long-term, based on some of my data, but I just became much less confident in that. Zero correlation makes luck a much stronger factor in our models, (including mine), no matter how you look at it. I have a wider range on what the true ceiling might be now.

How much luck is a factor will be determined by real data, I hope.

Perhaps place all deleted designer models in a separate category or list by themselves? Or include them in the combined overall list but flag them and allow them to be filtered and sorted as a group? Knowledge could be gained by allowing an analysis of market periods of factor inversion or black swan events common to many models that were originally expected to outperform. Especially if we could compare failed models to those which survived.

This should be a must for transparency @marco

Agree with @regallow and @scifospace , with a possible extension: models deleted during incubation should also be recorded.

I understand if this is complex or too difficult or other reasons we may not do it–I don’t personally care too much whether we do or don’t.

But I think we need to be honest that the 3-month rule isn’t “solving” the problem–it’s hiding the problem and quite possibly exacerbating it.

A designer can choose which of potentially several variants to move forward as vetted with a nice looking live record, quietly eliminating the others. For a savvy buyer, this doesn’t make navigating the marketplace any easier. For a less sophisticated buyer, it gives a false sense of security.

Without the incubation period a model is forced to stand (or fall) on its own, in public view. If we record the outcome for posterity, even better.

And since 3 months isn’t a sufficient sample size to determine pass or fail…..

Just brings me back to my original preference of getting my hands on models I understand, from designers I trust, sooner than later.

“out of sample results can tell you how a model is going to do going forward, either. There is zero evidence for that.”

There is evidence, in a limited way. If I remember correctly, a designer showcased a model with backtests of 75%+ a year and every three month period had huge excess returns. The first three months post launch the model lost money. That’s a form of evidence for a specific type of model. That type of model destroys credibility and is what this holdout period was designed to implement.

And it has done a splendid job. Designers don’t launch those types of models anymore.

So I am with you I think on at least some of the backtests being misleading. I think the data is noisy at best. My study was slightly different though. It actually did not include backtest data for example. We happen to be discussing two different topics.

My study looked at 4 years of out-of-sample data (not just 3 months) to see if it could predict the next year’s return. It was done in a rigorous way with the data we have. The only problem being survivorship bias.

I think many are just asking for the removal of survivorship bias, BTW.

I wonder if 3 months of out-of-sample data would show anything different. But please, redo the study as you wish. I initially made this a Claude Cowork Artifact to share (and fork or customize), but for some strange reason Chat Artifacts can be published but not Cowork Artifacts. I tried to do that so that people could look at several periods if they wished to. If you want me to I will produce an Artifact for a 3 month period. But I think you can get Claude to do that without me independently which may be more convincing whatever you find in the data.

I agree with you on several important points for sure:

  1. You have asked for excess returns (everywhere) before. I agree that would help including with assessing Designer Models as well as machine learning models.

  2. We should have access to more raw data. Maybe let Claude, ChatGPT or Gemini sort it out. The Designer model Excel Spreadsheet could have more data in my opinion.

  3. I am not against not allowing people to invest for 3 months on a model. I am fine with that or actually I do not care. I do not Design models or sub to them so you guys decide that. But the idea of including of all models submitted before the blackout and removed as part of a designers record is important in my opinionas many have said above:

  4. I hope I have given a heart to everyone who wants to asses all of the Designer models without survivorship bias. P123 and FactSet do a lot to avoid survivorship bias in their data. I don’t understand why designer models should be any different.

  5. Above are just a few examples of data we could use. But more raw data of any kind would be beneficial. Here is another request:

TL;DR: I like your ideas and cannot see a single thing I disagree with. I don’t really care what members decide to do with regard to any embargo or blackout. I would prefer that were transparent however.. As for my study, the data did not behave the way I wanted it to. I am truly surprised by it. Make of it what you wish. Or do your own study if you want. I would love to see the results.

That’s possibly my fault, sorry. I have a habit of taking threads where they weren’t going : )

Your comments are very pertinent and I agree with them. And I think they are the same topic really. There is a wealth of information in the Designer model data that we would both like to have access to (including the number of models submitted by a designer).. Maybe different purposes or different hypotheses to test at times.. But I would like to mine that data and/or use it to learn how to select models whether they are Designer models or my private models.

I have done what I can with the data available but it is limited.

Thank you for your comments!!!

I agree with a lot of what is said in the thread. Excess returns being shown would be a nice feature. However, model designers choose the wrong benchmarks sometimes (eg., Sp500 on a small cap model. )

this is a fascinating study, Jrinne

So for those who want to look at the aggregate results of designers, I took those models with 5-year returns and averaged the returns of those models for each designer (45 Designers).

Comments: Survivorship bias seems to have played a huge roll. Those who had never, or rarely, removed any of their models suffered a huge penalty in their average results. At least that is what it looked like to me. I hope I have not commented on any particular designer. Depending on how you interpret a box and whisker’s plot this should give some idea of what to expect with your models and the designer models long-term.

Thanks, this is good. For context SP500 did about 13.3% ann. Congrats to the designer who beat it by 10%ann