Recruitment case study · Premier League 2015/16
Benteke: A Retrospective
Crystal Palace didn't need more chances. They needed someone to finish them. A statistically-driven case for the centre-forward Palace should sign, built from 9,817 shots and 380 matches of open event data, with an honest account of what the numbers could not see.
This is a recruitment case study written deliberately blind. It takes the position of an analyst sitting at Crystal Palace in June 2016, holding one season of event data and no knowledge of what came next, and asks a single question: which centre-forward should the club sign?
Everything up to the closing section uses only what was observable at the time. No transfer fee, no knowledge of who Palace actually bought, no hindsight about whether it worked. That constraint is the point. A recruitment model is only worth anything if it can be judged on the information it had rather than on the answer it was shown afterwards.
The analysis arrives at Christian Benteke, whom Palace signed weeks later for a club record fee, and whose spell is now generally remembered as poor business. So the closing section removes the blindfold and asks the harder question: if the reasoning held up on the evidence available in 2016, why did the signing fail? The answer has less to do with the player than with everything the data could not encode.
In May 2016 Crystal Palace finished 15th, having scored 38 goals. Only four clubs scored fewer. The obvious reading is that Palace could not create chances. The data says something more specific, and more useful: they created a mid-table volume of chances and converted them worse than almost anyone.
That distinction matters, because it changes what you buy. A team that cannot create needs a creator. A team that creates and cannot convert needs to look at who is taking the chances, and whether anyone is taking enough of them. This report works through that diagnosis, builds a shortlist from it, and then stress-tests the leading name against the one question most public scouting work skips: whether the buying club can actually supply the player.
Act one: the diagnosis
Palace generated 42.4 expected goals and scored 38. That shortfall of 4.4 goals was the 4th worst in the division. They were not starved of the ball in dangerous areas: Palace ranked 10th of 20 for both carries into the penalty area and touches inside it. The chances arrived. The goals did not.
The sharper problem shows up when you ask who Palace's main goal threat actually was. It was Bolasie, a winger, with 4.33 non-penalty expected goals across the season. No other club in the Premier League leaned on its leading chance-taker so lightly. Even Aston Villa, relegated with 25 goals, had a forward generating more.
It is tempting to read that as Palace's forwards being poor, and the totals do mislead. Ranked by expected goals per 90 rather than season totals, Palace's centre-forwards come out as the most productive players in the squad. The two centre-backs sit high on the totals only because they played almost every minute of the season, and on a per-90 basis they fall to the bottom of the same group.
That rescues them against their own team-mates, which is a low bar. Set against the league, once minutes are levelled, they were mediocre. Among 49 forwards with 450 minutes or more, Palace's best rate belonged to Gayle at 0.344 npxG per 90, only the 61st percentile. Their most-used forward, Wickham, managed 0.261, in the bottom quartile at the 24th percentile. Palace's forwards averaged 0.301 against a league median of 0.319.
So the per-90 view corrects one error without excusing the squad. Palace did not have good forwards being wasted. They had forwards who were slightly above average at best, and the one who played most was among the least productive in the division.
Alongside that quality problem sat a starker one about minutes. Palace's most-used centre-forward, Wickham, played 1,366 minutes, fewer than any other club's leading striker in the division. The next lowest was 1,567; Harry Kane played 3,473. Palace shared the role between 7 different centre-forwards, none of whom held it.
Why so few minutes? Not purely by choice, and this is where the event data needs help from outside itself. Wickham injured his ribs in the second match of the season and had made only three appearances by late September; a calf injury in December cost him roughly another month. Palace's striker problem was therefore an availability problem as much as a selection one.
That distinction sharpens the brief rather than softening it. A club whose first-choice forward cannot stay on the pitch does not simply need a better striker. It needs one who is durable enough to hold the position for a season.
The brief that falls out of the data: not simply a better finisher, but a first-choice centre-forward who generates a high volume of good chances and can absorb the minutes Palace spread across seven players.
Act two: the shortlist
Ranking forwards on goals alone rewards whoever had the best season, not whoever is most likely to have another one. The scoring here weights chance volume (npxG per 90) and chance quality (npxG per shot) most heavily, includes box presence, and deliberately caps the influence of finishing over- or under-performance, which is noisy across a single season and regresses hard.
The quality axis is borrowed from basketball. The “Moreyball” insight was that shot selection is a skill distinct from shooting: a player who consistently takes better shots is doing something repeatable. In football that is npxG per shot. A striker averaging 0.16 per attempt is arriving in materially better positions than one averaging 0.08, whatever their conversion rate that year.
One name sits alone in the top right. Christian Benteke produced 10.22 non-penalty expected goals in only 1,634 minutes at Liverpool, the highest npxG per 90 of any forward in the league (0.563), at 0.16 xG per shot. He ranked 3rd overall on the weighted score, behind only Sergio Agüero and Diego Costa, neither of whom Palace could realistically sign.
He was also gettable. Benteke had started fewer than half of Liverpool's league matches under a new manager who plainly did not build around him.
Act three: interrogating the recommendation
A ranking is a hypothesis, not a decision. Before recommending a club record fee, two things need testing: what kind of striker this actually is, and whether the club he would be joining can get the best out of that kind of striker.
What kind of striker is he?
Benteke's profile is unusually narrow. 31.7% of his shots were headers, against a league average of 16.4%. His average shot came from 14.0m, well inside the league's 19.1m. Only 15 of his 63 shots came from regular open play.
Positionally he was where a penalty-box striker should be. Grouping his 2,046 on-ball touches into pitch zones, 20.0% came inside the penalty area, and the single busiest zone in his whole map was the central area of the box at 13.3%.
Against other forwards he was also unusually advanced. His average touch came 83.0 metres up the pitch, against 81.0 for the 42 forwards with 450 minutes or more, making him the 9th highest-operating forward in the division. His average touch also sat left of centre, at 37.6 across a pitch whose centre is 40, with 55.6% of his touches in the left half. Worth noting for later: that left-sided, high starting point is the profile he was bought on.
This is a striker who occupies the box and attacks deliveries. His expected-goals figure is not evidence that he manufactures chances alone. It is evidence that he is exceptional at attacking chances other people manufacture. Which raises the question that decides the transfer.
Can Palace supply him?
A cross-fed striker is only as good as the service he joins. This is where a player-in-isolation shortlist can go badly wrong, so the target's chance profile has to be checked against the buying club's delivery profile.
Palace passed decisively. They attempted 14.7 crosses per match, 2nd highest in the division, ahead of Manchester City and Liverpool. With Zaha and Bolasie wide, Palace were already among the league's most committed crossing teams. They had been building the supply line for a striker they did not have.
On the evidence so far this is a coherent recommendation: the right kind of striker, joining a team that already plays the way he needs it to. Which is the point at which it is worth turning the questions on the analysis itself.
Modelling considerations
Two modelling questions sit underneath the recommendation, and they have opposite answers. Expected goals is a measurement problem: the value is already in the data, its inputs are documented, and it can be rebuilt from scratch and checked. Availability is a prediction problem, and a far harder one, where the literature is openly sceptical it can be done at all.
That asymmetry matters more than it first appears. The part of this analysis that is easiest to verify is not the part the signing actually turned on.
Can the number be trusted?
The entire shortlist rests on expected goals, and StatsBomb's xG model is not documented. The specification states only that the value is “calculated from StatsBomb's own xG model”. If a recruitment decision leans on a number, someone should be able to ask how it is produced.
So I rebuilt it. Using the shot freeze-frames, which are documented, a model was trained on 7,362 shots and tested on 2,455, from features a coach would recognise: distance, the angle of goal available, body part, whether the shot was struck first time, and how many defenders stood in the shooting lane.
| Model | ROC AUC | Log loss | Brier |
|---|---|---|---|
| StatsBomb (baseline) | 0.7894 | 0.2561 | 0.0713 |
| Logistic regression | 0.7815 | 0.2608 | 0.0737 |
| Gradient boosting | 0.7715 | 0.2624 | 0.0734 |
A plain logistic regression on eleven documented features lands within 0.008 AUC of the proprietary model. The gradient booster, tuned across several configurations, did worse than the linear model every time. Chance quality is largely geometric and monotonic, so added flexibility bought variance rather than signal. Reported as found; the simpler model won.
This does not beat a professional model, and is not meant to. It establishes that the metric driving the shortlist is reproducible and inspectable.
The durability it cannot see
The metric holds up. The bigger problem is what the scoring model never looks at, because the brief demanded durability and the model is blind to it. Benteke's 1,634 minutes were not purely a selection decision. He lost time to a thigh injury early in the season, a hamstring problem, and knee ligament trouble that ruled him out of a League Cup tie, then damaged the same knee again on international duty with Belgium in March.
A shortlist built on per-90 rates rewards a player for the minutes he plays and says nothing about the ones he misses. Worse, a low minutes total makes a strong rate look more impressive, so the model reads injury absence as evidence of untapped upside rather than as a risk.
That is not an abstract worry, and the size of it is measurable. Among the 36 forwards on the shortlist, Benteke ranks 1st by npxG per 90 and 10th by total npxG across the season. Nothing about the player changes between those two numbers. Only the assumption about how much he will be on the pitch does, and that single choice moves him nine places.
Could injury risk have been modelled instead?
It is tempting to answer yes and reach for a classifier. The evidence says be careful. Roald Bahr's 2016 review in the British Journal of Sports Medicine argued that screening tests to predict injury “do not work, and probably never will”, and subsequent work developing and validating multivariable models from pre-season screening has broadly agreed that it carries no useful predictive value on its own. A recent scoping review of machine learning for injury risk found football the most-studied sport, with tree-based models performing best, but flagged that studies reporting headline discrimination above 0.9 often rely on wide prediction windows and loose injury definitions that make the clinical usefulness questionable.
So predicting that this striker will injure his knee in March is not a credible ambition. But that was never the question recruitment needed answering. Recruitment does not need the date of the next injury. It needs an expected availability, and a recurrence risk on the injuries a player already has, which is a time-to-event problem rather than a classification one. That reframing is where the more recent literature is more encouraging: work combining survival analysis with decision theory reports better discrimination than standard classifiers, and commercial providers claim strong back-tested accuracy on recurrence and time-out estimates when screening transfer targets, though those figures are vendor-reported rather than independently validated.
The practical fix is smaller than a new model. Rank targets by expected contribution across a season, meaning rate multiplied by the minutes you actually expect to get, rather than by rate alone. Benteke's injury record was public in 2016 and would have pulled his expected availability down, which is the discount the shortlist never applied. Availability is not in the event feed, so a medical file belongs alongside this analysis rather than after it.
What happened next
This is where the blindfold comes off.
On 19 August 2016, Crystal Palace signed Christian Benteke from Liverpool for £27m plus add-ons, a club record. The analysis above reaches the same conclusion the club's recruitment did, from open data and without hindsight of the fee or the outcome.
The first season vindicated it. Benteke scored 15 league goals in 2016/17 and finished as Palace's top scorer, outscoring Jamie Vardy and Roberto Firmino; only two players at clubs outside the top six scored more. The supply-line argument worked exactly as the data suggested it would.
Then it collapsed. Three league goals in 2017/18, followed by a drought lasting almost a year. Any honest case study has to sit with that, so it is worth being precise about what actually changed, because almost none of it was about the player's ability.
The system was dismantled underneath him
Benteke's good season was played under Alan Pardew and then Sam Allardyce, both direct, crossing-oriented managers whose approach matched the profile that made him a fit. Pardew was sacked on 22 December 2016 with Palace 17th, having won four of seventeen; Allardyce replaced him the next day, kept them up in 14th, and left in May 2017.
Palace then appointed Frank de Boer, who set about installing a possession game. He lasted four league matches, all lost, none scored in, the worst start by a top-flight English club in 93 years. Roy Hodgson took over on 12 September 2017 and moved Palace towards a more controlled style again. Inside roughly nine months the club had cycled through three managers and two competing philosophies, and the crossing volume that the recommendation explicitly depended on was no longer the plan.
He stopped standing where he was bought to stand
Sky Sports' analysis of the 2017/18 collapse, built on tracking data this report does not have, found him operating deeper and more centrally than in his good season, working in crowded left-of-centre areas rather than attacking from the high, left-sided starting point above. His physical output had not dropped, with distance covered marginally up and one of his highest sprint counts recorded that season, so this was not a player who had stopped running. He was running in different places.
Set against the 2015/16 baseline, that is the profile the recommendation was built on being inverted. The case for him rested on a striker taking a fifth of his touches inside the penalty area, higher up the pitch than nearly every forward in the league. A deeper, more central role does not make a player worse at that job; it stops him doing it.
Sky's analysis also found no clean relationship between Palace's crossing volume and his scoring, with two of his best games coming on 11 and 13 crosses. That complicates the supply-line argument in Act three, and it is worth being precise about why. The evidence here was about chance type, his 31.7% header share against a league average of 16.4%, not a claim that more crosses mechanically produce more goals. Delivery quality and where he started from mattered more than raw volume, which is a refinement of the argument rather than a refutation, but the original framing was too comfortable.
The injuries then compounded it. Benteke suffered a knee ligament injury in September 2017 that cost him six to eight weeks, arriving precisely as he was trying to adapt to a fourth manager in a year. Palace had also lost Wickham to a ruptured anterior cruciate ligament in November 2016, which removed the competition and the rotation, leaving an ageing, injury-hit striker to carry the line alone.
Then there is the part no dataset holds. Benteke scored four times in his last 58 club matches, a collapse too total to be explained by role or fitness alone, and contemporary accounts describe a striker whose confidence in front of goal had gone, epitomised by overruling the designated penalty taker at Bournemouth and missing. Expected goals measures the quality of a chance. It cannot measure whether the man arriving on it still believes he will score.
So the recommendation was sound conditional on Palace continuing to play the way they played in 2015/16, and that condition expired without anyone revisiting the decision that rested on it. The failure was not that the analysis picked the wrong player for the team as it existed. It was that the team stopped existing, and nobody re-ran the question.
The transferable lesson: a recruitment model ranks players, but signings succeed or fail inside a system, and under a medical file. Fit is a claim about the club as much as the player, with an expiry date attached, and it needs re-testing every time the manager or the style changes. The number that justified a signing in June is not still true in September.
Data and limitations
- Source. StatsBomb Open Data, 2015/16 Premier League: 380 matches, 549 players, 9,817 non-penalty shots. Free for public research; StatsBomb are credited as the data source.
- Why this season. It is the only complete league season in the open data with event-level detail for every club. FBref lost its Opta licence in January 2026 and never carried coordinates; StatsBomb's paid API is not available to individuals. Event-level analysis has to use what is open.
- The season is old. Ten years old, and it constrains the conclusions rather than the method. The pipeline runs unchanged against any StatsBomb competition.
- Candidates are Premier League only. Real recruitment scans dozens of leagues. Every player here is one the open data happens to cover, which is a genuine limitation of the shortlist, not of the approach.
- Derived metrics are ours, not StatsBomb's. Progressive carries and passes are defined as closing at least 25% of the distance to goal with a minimum five-metre gain, excluding set pieces. StatsBomb supply carries and pressures as events; “progressive” is an analyst definition and different sources define it differently.
- Minutes are reconstructed. The open data has no minutes-played field, so minutes come from Starting XI, substitution and red-card events.
- No availability or medical data. The event feed records what a player did on the pitch, never why he was absent from it. Because per-90 rates divide by minutes played, injury absence can make a player look more efficient rather than less reliable. Injury context here was researched separately and sits outside the dataset.
- No transfer-market data. Fees, wages, contracts and age curves are absent, and they decide real transfers as much as performance does.
Built in Python (pandas, scikit-learn, matplotlib). Data source: StatsBomb. Code and the interactive Tableau dashboard are linked from adomasg.com.
Claude Code was used to assist in the production of this report, across the data pipeline, the analysis and the writing. The dataset, the metric definitions and the conclusions were directed and reviewed by the author.
An independent analysis written by a supporter, not affiliated with or endorsed by Crystal Palace Football Club. The club crest is reproduced for identification only and remains the property of the club.