Hi, Shrewd!        Login  
Shrewd'm.com 
A merry & shrewd investing community
Best Of MIBest OfAll BoardsThe Shrewd’m WeeklyLearn to InvestHow to Become Shrewd
Search
Shrewd'm.com Merry shrewd investors
Search
Best Of MIBest OfAll BoardsThe Shrewd’m WeeklyLearn to InvestHow to Become Shrewd


The week's question
In December 2024, in the thread "Re: BRK: Why Not XOM?", BreckHutHigh asked the members: "What about the long road trips with kids?" This week it is put to everyone again. The button below opens the small thread re-asking it - read what others have said so far, then give your own answer as an ordinary reply.
Answer this questionContinue to Shrewd'mThis note won't appear again
Investment Strategies / Mechanical Investing
Unthreaded | Threaded | Whole Thread (43) |
Author: mungofitch 🐝🐝🐝 GOLD
SHREWD
  😊 😞

Number: of 6132 
Subject: Re: ML for MI
Date: 06/25/24 12:25 PM
Post New | Post Reply | Report Post | Recommend It!
No. of Recommendations: 13
ML evaluates a model differently. They first define an algorithm to take any data and predict the future
(like a neural network or many tree search). Then the break the historical data into shorter subsets. For
each subset you train the predictor model on typically 60 to 70% of the subset and test the results.
The model is judged on the results of all the train-test sequences where no train-test sequence has any
future data. A model that does well did well looking only at the data it had at that point in time.


This is how all data mining is done, when it's done correctly. Same with "classic" MI screens.
(terminology clarification: data mining is a good thing, overmining is a problem)

The problem is this: the process you have described, including seeing which models worked on the withheld out-of-sample validation subset and killing those that didn't, is itself another layer of data mining, another step in a single larger and more complex model building process. That "greater" process has no out-of-sample validation. Once you have culled your set of models by looking at the effectiveness within the validation data set you held back, that validation data set is contaminated and is now in sample, not out of sample.

This might be seen to be "OK" if you only did it once (depending on your strictness), but nobody does this just once. There just isn't enough financial history. In effect all history gets used, and it's all in sample. The only out of sample is the stuff that actually happened after you stopped modelling, and (to be strict) after you stopped culling your set of models.

Combined with machine learning with a lot of parameters, the ability of your final model to have memorized the data set AND the validation set is nigh unbounded. That's not to say it's impossible to find a new and useful insight this way, but it's a very thorny patch in which to be hunting.

Jim
Post New | Post Reply | Report Post | Recommend It!
Print the post
Members reply directly to mungofitch here — and replies get answered. Reading is free; so is joining the conversation. Join Shrewd'm »
This community has written 6,115 posts about Mechanical Investing. The article-length ones it recommended most:
Dividend investing · 52 recs · 2025
Non-Mag7 screen · 34 recs · 2025
OT - Div yields and returns · 32 recs · 2024
Using AI to generate backtesting programs · 30 recs · 2025
Rankings for 19Dec2022 · 29 recs · 2022
Unthreaded | Threaded | Whole Thread (43) |


Announcements
Mechanical Investing FAQ
Contact Shrewd'm
Contact the developer of these message boards.

Best Of MI | Best Of | Favourites & Replies | All Boards | Followed Shrewds | Open Questions | Moving a community