FBRef announced this week they removed all the Advanced Stats provided by Stats Perform by Opta.
Something that took the football analytics community by surprise.
FBRef was maybe one of the first (if not the first) source you could think of to grab data and build high-level analysis. Whether you were a student, analyst, scout, fan, whoever, their web was a must. They these great level of detail in their data that actually made you go down rabbit holes just by looking at the granularity of their metrics.
I used to recommend it as one of the go-to places to get in contact with advanced football data stats → How to Use FBRef for Football Scouting with Python?
This new situation far from being discouraging should make us all football lovers develop our creativity to make the most out of the available sources on the internet and explore other angles. For example: places like SkillCorner Open Tracking & Physical Data or Hudl StatsBomb Event & Physical Open Data are well known data providers where you can obtain top valuable data to run different analyses but with a “detective mind”.
What I mean by that is trying to look at the data as a vehicle to derive close-to-real-world analyses of how technical staff at clubs or pro teams use this data, and also how other fellow colleagues use it.
You can see it on social media. The majority of professionals and data practitioners are active on X or LinkedIn. Perfect places to take a look at how this data is being used and inspire yourself of those signals to craft your own dashboards or reports how they consult other data sources, how they collect, clean, model and even how they present that data.
Having different POVs around how others manage sports data helps you upskill in the data field way faster, which is ultimately what we actually want: be more competent data professionals independently where the data comes from.
I also wouldn’t rule out the use of synthetic data for practicing purposes. With the rise of Generative AI and LLMs, this has become a popular alternative within the industry due to lack of access to closed source data.
Spending some time on crafting comprehensive prompts, using few shot examples, with metrics and well-defined ranges, can help you mimic real-world data for your own practice.
The main thing to remember from this FBRef situation is that it is essential that we adapt by focusing on sharpening skills with what’s available, or by creating our own way of accessing to data sources.
This is a great signal to start creating now.
Till next time,
Ricardo.





The synthetic data angle is more practical than most people realize. I built a few scouting dashboards last season using synthetic event data for proof-of-concept stuff and it worked suprisingly well for practice. StatsBomb open data is solid but limited to specific competitions, synthetic fills gaps when you need broader coverage to test methodolgy before spending on real feeds.