Scout
A ranking, not a rating. It blends 60% the forecast α y 40% the z quality, each
part measured in standard deviations within his position — a +2.0 is two
deviations above the average player in his position.
Only scores players who perform above the average of their league and position (z ≥ 0.5).
Without that floor the ranking fills up with cheap kids who do not play well yet.
Before it was half quality and half bargain, and the result was 28-year-old players
in fourth divisions: genuinely cheap, genuinely useless.
α 12 meses
How much we expect his market value to change in a year, as a percentage. It comes from a model
entrenado sobre 89.983 player-seasons with his performance, the difficulty of his league,
age, starting price, the strength of their club and their international cup matches.
Validated out of sample the right way: trained up to 2023 and tested on 2024 onwards,
with players the model never saw. Rank correlation 0.72 sobre 43.800 casos.
It is an estimate, and in new leagues it is less accurate (0.66 versus 0.77 in the usual ones).
And he is calibrated by data coverage. Measured out of sample, the model
raw promised too much precisely where data is scarcest: in leagues with coverage <60%
forecast +29% where the outcome was +16%, and in the top decile — where the ranking lives
scout-style — promised +33pp too much, while shortchanging the well-covered leagues.
That is why the scout top filled up with weak leagues. Each coverage band is
recalibrates now against what that band actually delivered: the promise of a league
a medio cubrir vale ×0.74, that of a fully covered one
×1.08. After the adjustment, the excess sits at ~0 across the four
bandas.
★ Man of the match
How many times he was the best rated of the 22 on the pitch this season.
SofaScore does not publish that award as data: it goes to the highest rating of the match,
and we already had the match-by-match ratings, so this is arithmetic on top of what
ours, not a new source.
Measures something different from the z: they correlate 0.57. The z is consistency,
these are peaks. Over 89.983 player-seasons it is worth t=+22.6 above
for quality, age, price and year — from 0 awards to 9 or more, the value went from +14% to +45%.
It does not improve the forecast: adding it to the model moved the CI from 0.7232 to
0.7234. The model already extracted that signal from other variables. It is here because it is easy to understand
at a glance, not because it adds precision.
Profile by areas
Each axis is the average of the percentiles we have for that area, measured
within his position family and against the 147 leagues. The problem is that
not every league publishes the same metrics: the Primera Federacion gives one of
all three in Finishing, the Premier all three.
And averaging one metric is not the same as averaging three. With one
one good number is enough to reach 99; with three it takes three. Without correcting it,
el 9,7% of players reached 90 or more on a single metric
in Finishing, against the 0,2% of those who had three — not
because they were better, but because their average was not diluted. Complete profiles came out
castigados.
Now each player is ranked against everyone else measured with the same
metrics, within his position family. That way a 90 means the same in
everywhere: the top 10% among those measured the same way.
We did not pick it by eye. In leagues that publish everything we know the answer, so
that we cover two out of every three metrics precisely in the weakest leagues —which is like the
coverage truly fails— and we look at which rule recovers the number we would have
given with complete information. Across five categories and 106,000 players, ranking
within the group was the only thing that left the tails even across all five
(10,0% contra 10,0%), y en Defensa fue
plus what best reproduced the real ordering between leagues (error 0.022 versus 0.294
of correcting nothing).
What it costs: within a tier, the level differences between
coverage groups flatten out — in Defence the blinded group ends up 5,2 points per
below their truth. Calibrating the shift instead of ranking recovers that
(bias +0,5) and is more accurate player by player, but leaves the one from a poor league in a
0,7% of reaching 90 against the 17,2% del
rest: it gives him back the ceiling, which is the bug we came from. For a number that
compares across profiles we prefer it to mean the same in all of them.
✚ Days out
Days lost to injury in the last two years, of the history of
Transfermarkt, with the type of the longest injury. Covers the 15.167 players
of the platform.
We measure the cost of an injury against 45.538 player-season, y la
honest comparison is lesionado contra lesionado: who has zero days
out can be healthy or can simply not be playing, and mixing them makes a knock
short it may look like suma. Among those who did get injured, doubling the days
fuera cuesta −3,6% of next year value
(t=−12,4). Versus an absence of one month or less: 1-3 months
−6,3%, 3-6 meses −12,7%, more than half a year
−15,9%.
It does not enter the model either: the IC goes from 0.7581 to 0.7590. Minutes already
are in and not playing is the injury channel — with minutes in the regression,
the effect of the +6 months collapses from −7,9% a −0,1%. Y no explica nuestras peleas
with the market: the median is 0 days in all gap bands, and among the
300 the biggest gaps there is a 13,0% with four months lost against 10,5% for the
rest. It is a data point to read, not a discount on the price.
z · calidad
His average SofaScore rating, converted to standard deviations within his own league,
season and position. A 7.2 does not mean the same in the Premier League as in the Peruvian second tier;
this puts it on the same scale.
Before standardizing we discount the quality of his club, because otherwise a good
a goalkeeper on a relegation team came out punished for the goals conceded and a mediocre one at the
champion came out rewarded.
Fair value and gap
El fair value is what their profile justifies charging today: quality, league difficulty,
age (smooth curve, not brackets), minutes, goals and assists, club strength, years of contract remaining,
being a foreigner in his league, a passport from an exporting country, and height where the market pays for it (goalkeepers and centre-backs).
R² — sobre — players of today.
La brecha is the difference with the Transfermarkt price. +50% means that the
market pays him 50% below what the profile justifies. It is re-centered against the position
best paid who can play there: a winger who also plays as a playmaker is compared with playmakers.
Validated where it matters: on the historical panel, players in the most
undervalued went up the following year and those of the most overvalued went down, across all ten deciles without jumps.
But it does not predict that a club will overpay above the list price — we tested that against
traspasos reales y no aguanta.
One single model, and why. Each new data point that arrived —height, versatility, position
fine, man of the match, injuries, and the advanced FotMob metrics (recoveries,
pressing, blocks)— went through the same test: does the gap get better at
anticipate what the market does after, out of sample? They all improve the fit
at the price of today; none improved the prediction, FotMob included (IC 0,359 without, 0,352 with,
over 7.783 historical player-seasons). So they are shown as a data point on the profile and not
enter the model. Only the contract and the passport made it in, and because the market pays for them.
The €60M cap — whom we do NOT value. No valuation for players from
more than €60M, and they are not in the Explorer: above that price, market value is
marketing, scarcity and clauses — things this model does not measure. And history is not enough
to calibrate: in 98.000 player-seasons there is not a single ≥€30M case that the model
called someone 40% overvalued and it can be checked what happened, and the buy signals in
€60–100M got 13% right — worse than a coin flip. Rather than invent a number for
Bellingham, we stay quiet. Our territory is the market where a real club buys:
de €0 a €60M. Between €15M and €60M the player is shown, but a negative gap
is explained rather than asserted (the star premium is not in the regression).
A gap of +200% does not mean the same at 19 as at 31, and the probability
that goes with it knows it: the certainty model receives the gap along with age and price,
and returns what a gap like that, in that profile, did afterwards. In a cheap under-22 P(rises)
usually tops 80%; past 30, with the same gap, it drops below 20%. No caps
no tables either: the same probability for everyone, calibrated.