Two Pre-Registered Designer Models on OSF

Quantitative model marketplaces often suffer from publication bias, survivorship bias, and p-hacking—assuming any statistics with p-values are presented in the first place. Much of this can be mitigated by pre-registering studies.

I recently launched two Designer Models and have never opened any others. They should become publicly visible in about 3 months, ensuring no survivorship or publication bias.

My hope is that these models will perform well and that the pre-registration will validate the statistical results. But this should add to the discussion of backtesting decay and the effects of survivorship bias in any case.

Both models are pre-registered here: OSF

Edit: I was thinking about why pre-registration helps. The literature on this is pretty dry, but really, it is just like calling your shot in billiards.

I can be aiming at the 7 ball trying to get it in the side pocket, miss entirely, bank the cue ball off the cushion, hit the 9 ball, and send it into the corner pocket. Then I post that shot online claiming I'm great at billiards. I post if did not call my shot ahead of time, I should say.

The analogy is looking at 156 models, finding the single best performer, and claiming it was pure skill. If you had called that specific model ahead of time, it actually would be impressive—or at least the odds of true skill shoot up dramatically.

Looking at Designer Models should be exploratory at first. You may like the designer, the backtest, or the out-of-sample results—all legitimate. But then you need to call your shot going forward. Write it down, specify some statistics, or set a clean benchmark—something like "I think this will beat 95% of Designer Models 3 years from now."

That is dramatically different than picking whichever ball drops after a break and calling it skill. For most players, it wasn't.

If you want the dry statistical version: unless you set a strict end date for the study in advance, over half of all random studies will show statistical significance at some point. And actually, given enough time, the probability approaches 100%

Looking forward to seeing the results, always been interested in seeing a DM from you. And advancing the discussion on DM backtests vs. reality is also a noble goal.

I'm always grateful to see new DM additions — thank you, Jrinne. There are so many of you I would love to see post DMs.

The OSF registration is commendable — more on that below.

Having navigated the DM market for a while, I've found that getting to know the designer and understanding his logic are essential. Most designers rely on the backtest to sell the model, when the things that actually matter — the designer's cred, the model description, basic attributes — are typically all but absent.

For instance: does a model have buy and sell rules beyond rank? What, roughly, does the ranking emphasize? I know designers want to keep their models and ranking systems secret, but even a basic sketch of the approach would help.

As discussed before, the 3-month rule should go away — in favor of full transparency and a launch log. Imagine someone challenged you to show subscribers the best-looking launch possible. What truer gift could you get than 3 months to spawn as many candidates as you like, unbeknownst to potential subscribers, and reveal only the ones that worked? And the window buys nothing in return — 3 months of live returns can't statistically separate skill from luck anyway.

I appreciate Jrinne's approach — pre-registration is the antidote to exactly that scenario, adopted voluntarily. The simplest version, though, wouldn't need OSF at all: remove the 3-month veil and have P123 log every model created, including dead ones–same antidote, built into the platform by default.

The one rationale mentioned for the 3-month rule is reducing "clutter," but clutter is a display problem, with display solutions — filters or default views. The veil solves it by destroying information instead: the launch record that tells you what a surviving model is actually worth. We are removing the denominator–it’s like knowing how many hits, but not how many at-bats. I hope this gets consideration — or that someone can explain what I'm missing about why the current setup serves subscribers better.

Thank you again for the models and for raising the bar, Jrinne.