Clustering Playing Styles of Teams in the Top 5 Leagues
Letting the Data Speak
Clustering Playing Styles of Teams in the Top 5 Leagues
Letting the Data Speak

After diving deep in my last piece about defining team style and performance metrics, a more fundamental, almost philosophical question kept nagging at me: How many types of teams actually exist in the top five leagues?
We’ve all got our go-to labels: “possession teams,” “high-pressing teams,” or “direct teams.” But I wanted to move beyond the usual football jargon and see what happens when we let the raw data draw the style boundaries itself.
So for this follow-up, I decided to put on my machine learning hat and explore the tactical universe through clustering.
For this analysis, I focused exclusively on the 12 style metrics from the previous post (bypassing the performance metrics for now). These metrics are the DNA of a team’s tactical identity — they capture how teams build up, progress the ball, defend, press, and move possession around.
The Statistical Processes
Normalizing with Yeo-Johnson
If you’ve spent any time with football data, you know it’s messy. Several of these style metrics aren’t neatly “normal” — some are skewed, some have funky long tails, and others are constrained by how they’re calculated.
Why does this matter? Well, clustering algorithms like K-Means rely on calculating distances. If a metric has extreme, non-normal values, it can throw off the whole process and make that metric unfairly dominant. It’s like measuring distance with a rubber ruler.

To fix this, I applied the Yeo-Johnson transformation. It’s a trusty go-to in the machine-learning world for a few simple reasons:
- It’s versatile! It works for both positive and negative values (unlike its cousin, Box-Cox).
- It calms down the skew and stabilizes the variance, making the features behave nicely for distance calculations.
- Crucially, it makes all the metrics more comparable, which is key for a fair clustering fight.
After this step, the metrics become much more suitable for the next stages, ensuring no single feature dominates the analysis just because it started with bigger, crazier numbers.
Scaling and Dimensionality Reduction
Once the distributions were playing nicely, the next step was to get all 12 metrics on the same scale. Since a ratio, a count, and an inverted value all live in different numerical universes, I used the StandardScaler. This is standard practice: every metric ends up with a mean of 0 and a standard deviation of 1. Again, we’re making sure a metric can’t bully the others simply because its numbers are larger.
With everything normalized and scaled, I performed Principal Component Analysis (PCA). Think of PCA as a brilliant editor. It simplifies the dataset by combining the original 12 metrics into fewer, new components (PC1, PC2, etc.) while keeping the essential tactical story intact.

I included a chart showing how much variance each principal component accounts for, from PC1 through PC12. The analysis clearly showed that the first 8 principal components capture about 90% of the total variance. Why cluster on 12 potentially redundant, noisy features when 8 components can tell you almost the whole story? Using these 8 components helps stabilize the clustering and cuts down on noise.
Choosing the Number of Clusters
How many distinct “flavors” of football are there? To decide on the optimal number of groups (k), I consulted two common methods: the elbow method and the silhouette score.

The silhouette score initially loved k=2. However, while mathematically tidy, two groups felt way too simplistic for the beautiful chaos of the top five leagues — it would basically just split teams into “possession-heavy” vs. “everyone else.” Not exactly a mind-blowing insight!
The next strongest result was k=7, which offered a great balance. But seven clusters felt a touch too granular for a high-level taxonomy, leaving some groups a little too small to confidently interpret.
In the end, I landed on k=6. It felt like the Goldilocks number:
- It had a great silhouette score, very close to the one for k=7.
- It successfully avoided the oversimplification of k=2.
- The resulting clusters were large and distinct enough to interpret without having to squint.
The elbow method wasn’t much help, as the curve didn’t have a definitive “bend,” so the silhouette score and the ultimate practical interpretability were the heroes of the decision.
Cluster Distribution Across Situations

One quick check before diving into the insights: I looked at how teams were distributed across clusters for All, Home, and Away situations. The counts were surprisingly stable! This is great news, as it means the style structure is robust and not heavily skewed by whether teams are playing at Anfield or the Allianz Arena. The cluster interpretations are reliable.
The Insights
To really understand these six groups, I first visualized their profiles. I visualized each cluster using radar charts (both with z-scores and percentile ranks) to see their overall patterns.

Radar profile per cluster using z-scores

Radar profile per cluster using percentile rankings
I asked an AI to help me quickly draft some summaries! 😂
Here’s the breakdown of the six distinct playing styles:
Cluster 0: Shoot-on-Sight Controllers
These teams combine strong territorial dominance with a distinct quick-trigger mentality. They build from the back well and keep opponents pinned (high field tilt), but once they reach dangerous areas, the patience disappears — it’s shoot first, ask questions later. Their shot quality is slightly below average, a reflection of a volume-over-quality approach. They control the flow but lack the refined, methodical chance creation of the elite sides. They often win by overwhelming opponents through constant pressure and repeated attempts rather than surgical precision.
Cluster 1: Disciplined High-Pressers
This group’s identity is forged in defensive organization, not possession. They are masters of the high press, constantly harrying defenders and looking to win the ball high up the pitch. Crucially, they are also the cleanest defending group, boasting the best discipline scores. They don’t seek to dominate the ball; their value comes from disrupting the opponent, forcing turnovers, and then striking quickly in transition. Highly drilled, compact, and dangerous when the ball is recovered.
Cluster 2: Patient Ball Circulators
If patience were a virtue on the pitch, Cluster 2 would be canonized. These teams are extremely deliberate in their ball movement, circulating possession for long periods before even considering a shot. They prefer central progression and maintain control through methodical, low-risk passing. Their pressing and defensive lines are moderate — a true risk-averse philosophy. They wait for the perfect moment, probing and recycling until an opening appears rather than forcing one. Controlled, slow-tempo, and focused on stability.
Cluster 3: Deep Defensive Block (The Bus Parkers)
Cluster 3 is the purest expression of the low-block in the dataset. They happily concede possession and territory, spending most of the match defending deep in their own half. Their low high-line metric confirms they are committed to dropping off. When they finally attack, it’s typically via wide areas, relying on crosses or wing-based transitions. Dribbling is often one of their few outlets for forward relief. These are the teams committed to compactness, survival, and defensive solidity.
Cluster 4: Aggressive Long-Ballers (The Grinders)
Say hello to the most direct teams. They almost never build short; instead, they go long immediately, bypassing the midfield to compete for second balls. This physical approach is reflected in their poor discipline scores — they commit the most fouls. Their attacks often break down before generating high-quality chances, giving them one of the worst shot-quality values. This is the gritty, pragmatic style that relies on chaos, duels, and compensation for a technical disadvantage.
Cluster 5: The Elite Dominators (The Super Teams)
These are the peak performers. Cluster 5 teams excel across virtually every phase. They dominate possession, control territory, and progress through the middle with ease. Their shots are not just frequent — they are also high in quality, reflecting perfectly crafted chances. Defensively, they push the highest line, squeezing the pitch and quickly winning the ball back. They flawlessly combine the patience of Cluster 2 with technical superiority and tactical aggression. These are the title contenders whose metrics reflect total match control.
The Map: Teams in Clusters
Now for the fun part: seeing which teams fell where! I visualized the scatter plots to see which teams fall into each cluster.

Scatter plot for overall situation

Scatter plot for home situation

Scatter plot for away situation
A few fascinating patterns jumped out immediately:
- The Elite Dominators: European giants like PSG (who won the treble last season), Bayern, and Barcelona are exactly where you’d expect them: Cluster 5. Liverpool also sits in Cluster 5 for the overall situation, but so does Tottenham. Spurs are positioned close to Arsenal, but their season outcomes couldn’t have been more different (Tottenham finished 17th while Arsenal finished 2nd). This is a perfect, tangible example that having a similar playing style doesn’t guarantee similar results.
- Serie A’s Tactical Identity: Notably, no Italian teams appear in Cluster 5, highlighting the generally more conservative tactical approach across Serie A. Napoli (the Serie A champions) falls under Cluster 0 (Shoot-on-Sight Controllers) along with Inter, Milan, and Juventus. These teams circulate possession but are more direct and quick to finish their attacks.
- Getafe’s “Haram Ball” is Official: José Bordalás’ Getafe appears in Cluster 4 (Aggressive Long-Ballers). This cluster commits the most fouls and plays the most disruptive, chaotic style. It’s a perfect match for Getafe’s well-earned reputation!
- The Anfield Difference: Some teams shift their identity based on location. Liverpool, for example, is a Cluster 5 Elite Dominator at home but drops to a Cluster 0 Shoot-on-Sight Controller away. Their dominance is clearly context-dependent.
Summary: What I Learned
This has been a genuinely fun project! It was satisfying to see some of my intuitive beliefs about team similarities confirmed by the hard data — and a few results turned out more surprising than expected. Visualizing the clusters really helped illuminate how clubs with nearly identical styles can still have wildly different season outcomes.
Here are the key takeaways from this journey:
- Playstyle ≠ Performance: Tottenham and Arsenal looking similar in style but finishing worlds apart is the most striking example. Style is only the starting point.
- League Identities Matter: The complete absence of Italian teams in the “Elite Dominators” cluster highlights how strong national tactical traditions still shape modern football.
- Home vs. Away Splits Tell a Deeper Story: Liverpool’s shift between clusters depending on the venue shows how team behavior adapts (or is forced to adapt) to context.
- Clustering Simplifies the Chaos: Football data is notoriously noisy, but grouping styles makes it so much easier to understand the broader tactical currents flowing across the top leagues.
Thanks for sticking with me on this deep dive!
메타데이터
- post_id
- d9efd31a5791
- slug
- clustering-playing-styles-of-teams-in-the-top-5-leagues-d9efd31a5791
- url
- https://medium.com/@alf.19x/clustering-playing-styles-of-teams-in-the-top-5-leagues-d9efd31a5791
- canonical_url
- https://medium.com/@alf.19x/clustering-playing-styles-of-teams-in-the-top-5-leagues-d9efd31a5791
- author_url
- https://medium.com/@alf.19x
- status
- ok
- fetched_at
- 2026-07-14 10:19:40