Alan published his analysis using limited public data available, so please read/revisit here, as my objective is to supplement what he presented.
There has not been a lot of opportunity to benchmark keepers since I started this exercise in summer 2021, but one of the ways I found some value was looking at keepers’ max data sample analyzing the very simple metric of goals saved above average. Basically, look at the distribution of all games comparing goals conceded to the post-shot xG they have faced.
The exercise is very limited, as the post-shot xG models offered by vendors like Opta (FotMob) and Wyscout are limited. This is why I look for max sample size first, with the hope of extracting something reliable signal-wise.
Apologies for the dated/white images, but these were the early-exercise examples for Barkas, Hart, and Frazier Forster:

