pwsqc¶
- class pwsqc.Config(d: float = 10000, n_stat: int = 5, n_int: int = 6, phi_a: float = 0.4, phi_b: float = 10, m_int: int = 4032, m_rain: int = 100, m_match: int = 200, gamma: float = 0.15, beta: float = 0.2, dbc: float = 1.24)[source]
- Parameters:
d – All stations within a range (d) around a given station are selected to compute the median rainfall over the surrounding area.
n_stat – If fewer than
n_statneighboring stations with rainfall measurements are available, the median cannot be calculated and the FZ flag is set to -1n_int – The FZ flag is set to 1 if this median rainfall is larger than zero for at least
n_inttime intervals while the station itself reports zero rainfall. The FZ flag remains 1 until the station reports nonzero rainfall.phi_a – If the median does not exceed a threshold value (phi_a), the HI flag is set to 1 for any rainfall value from the station itself above threshold
phi_b. When the surrounding stations report moderate to heavy rainfall, the threshold becomes variable: for a median ofphi_aor higher, the stations’ HI flag is set to 1 when its measurements exceed median timesphi_b/phi_a. HI flag is set to -1 if fewer thann_statneighboring stations report observations.phi_b – see
phi_a.m_int – To determine whether a station yields nonsensical measurements for that location, it is compared with time series of neighboring stations within a range (d). A previous period of mint intervals, or any longer interval where the station has at least
m_rainintervals of nonzero rainfall measurements, is evaluated. There needs to be at leastn_statstations with at leastm_matchintervals overlapping with the evaluated station to compute the SO flag.m_rain – see
m_int.m_match – see
m_int.gamma – The r (equation (1)) and bias (equation (2)) with all neighboring stations are calculated. If the median of the r values falls short of threshold
gammma, the SO flag is set to 1.beta – If this threshold is exceeded,
BCFnewis computed from the median of the bias values with the neighboring stations. If|log(BCFnew/BCFprev)| > log(1+β), this is deemed a systematic change for that station and BCFprev is replaced with the new value. This is hence a way to dynamically update BCF for individual stations.dbc – The default bias correction to address the fact that the Netatmo rain gauges have a general tendency to underestimate rainfall. DBC is a single-value one-off proxy of the correction needed for the overall PWS network bias and can be determined a priori by comparing network measurements over a period with typical rainfall for the local climate.
-
beta:
float Alias for field number 9
-
d:
float Alias for field number 0
-
dbc:
float Alias for field number 10
-
gamma:
float Alias for field number 8
-
m_int:
int Alias for field number 5
-
m_match:
int Alias for field number 7
-
m_rain:
int Alias for field number 6
-
n_int:
int Alias for field number 2
-
n_stat:
int Alias for field number 1
-
phi_a:
float Alias for field number 3
-
phi_b:
float Alias for field number 4
- pwsqc.apply_flags(data, strict=False, precip_col='precip', flag_cols=('FZflag', 'HIflag', 'SOflag'), bcf_col='BCF', out_col='precip_qc') FrameT[source]
Apply the flags and the bias correction to the rainfall observations.
- Parameters:
data (
TypeVar(FrameT)) – Long format DataFrame with the flag columns of the filters. Either apandas.DataFrameor apolars.DataFrame.strict (
bool) – IfFalseonly the intervals with a flag of 1 are discarded (“filtered flex”), ifTruethe intervals without enough information to determine the flag are discarded as well (“filtered strict”). (default:False)precip_col (
str) – Column name holding the rainfall of the interval in mm. (default:'precip')flag_cols (
tuple[str,...]) – Column names of the flags to apply. (default:('FZflag', 'HIflag', 'SOflag'))bcf_col (
str|None) – Column name of the bias correction factor,Noneto not apply a bias correction. (default:'BCF')out_col (
str) – Name of the column the result is written to. (default:'precip_qc')
- Return type:
TypeVar(FrameT)- Returns:
DataFrame with an additional
precip_qccolumn where the flagged intervals are missing.
- pwsqc.bias_correction(data, dbc=1.24, beta=0.2, id_col='intern_id', date_col='date', so_col='SOflag', bias_col='bias', bcf_col='BCF', n_jobs=None) FrameT[source]
Compute the bias correction factor of every station and interval.
Every station starts out with the default bias correction factor
dbc. Whenever the station is known not to be an outlier, a new factor is derived from the median relative bias with the neighboring stations. If the new factor deviates from the current one by more than a factor of1 + beta, the change is deemed systematic and the new factor is used from the next interval on.The bias corrected rainfall of an interval is
precip * BCF.- Parameters:
data (
TypeVar(FrameT)) – Long format DataFrame with theSOflagandbiascolumns ofpwsqc.station_outlier_filter(). Either apandas.DataFrameor apolars.DataFrame.dbc (
float) – The default bias correction factor. (default:1.24)beta (
float) – Relative threshold a change of the factor has to exceed. (default:0.2)id_col (
str) – Column name for station identifier. (default:'intern_id')date_col (
str) – Column name holding the (interval end) timestamps. (default:'date')so_col (
str) – Column name holding the flags of the SO filter. (default:'SOflag')bias_col (
str) – Column name holding the median relative bias. (default:'bias')bcf_col (
str) – Name of the column the correction factors are written to. (default:'BCF')n_jobs (
int|None) – Number of threads to spread the stations over. Defaults to os.cpu_count(). (default:None)
- Return type:
TypeVar(FrameT)- Returns:
DataFrame with an additional
BCFcolumn.
- pwsqc.faulty_zero_filter(data, neighbors, n_stat=5, n_int=6, id_col='intern_id', date_col='date', precip_col='precip', flag_col='FZflag', n_jobs=None) FrameT[source]
Apply the Faulty Zero (FZ) filter to the precipitation data.
A station is compared with the median of its neighboring stations. The flag is set to 1 once the station reported zero rainfall for more than
n_intconsecutive intervals while the median of the neighbors was larger than zero during more thann_intconsecutive intervals of that period. The flagging continues until the station reports nonzero rainfall again. The flag is -1 whenever fewer thann_statneighboring stations report an observation.- Parameters:
data (
TypeVar(FrameT)) – Long format DataFrame with a regular time series per station, as returned bypwsqc.prepare_timeseries(). Either apandas.DataFrameor apolars.DataFrame.neighbors (
Mapping[int,Sequence[int]]) – Mapping of a station id to the ids of its neighbors, as returned bypwsqc.find_station_neighbors().n_stat (
int) – Minimum number of neighboring stations with an observation. (default:5)n_int (
int) – Number of intervals a station has to report zero rainfall while the surrounding area is wet. (default:6)id_col (
str) – Column name for station identifier. (default:'intern_id')date_col (
str) – Column name holding the (interval end) timestamps. (default:'date')precip_col (
str) – Column name holding the rainfall of the interval in mm. (default:'precip')flag_col (
str) – Name of the column the flags are written to. (default:'FZflag')n_jobs (
int|None) – Number of threads to spread the stations over. Defaults to os.cpu_count(). (default:None)
- Return type:
TypeVar(FrameT)- Returns:
DataFrame with an additional
FZflagcolumn indicating flagged data points.
- pwsqc.find_station_neighbors(station_metadata, d=10000, max_neighbors=None, id_col='intern_id', lat_col='lat', lon_col='lon', n_jobs=None) dict[int, tuple[int, ...]][source]
For each station, find neighboring stations within distance d, computed in parallel.
Stations at a distance of exactly zero are not considered neighbors. This excludes the station itself, but also – as in the reference implementation – any other station that reports the exact same coordinates.
- Parameters:
station_metadata (
Any) – DataFrame containing station metadata with latitude and longitude, either apandas.DataFrameor apolars.DataFrame.d (
float) – Distance threshold in meters to consider stations as neighbors. (default:10000)max_neighbors (
int|None) – If given, only themax_neighborsnearest stations withindare kept. This bounds the runtime of the filters in very dense networks. (default:None)id_col (
str) – Column name for station identifier. (default:'intern_id')lat_col (
str) – Column name for latitude. (default:'lat')lon_col (
str) – Column name for longitude. (default:'lon')n_jobs (
int|None) – Number of worker processes to use. Defaults to os.cpu_count(). (default:None)
- Return type:
- Returns:
Dict mapping each station id to a tuple of neighbor ids, sorted ascending by distance.
- pwsqc.high_influx_filter(data, neighbors, n_stat=5, phi_a=0.4, phi_b=10, id_col='intern_id', date_col='date', precip_col='precip', flag_col='HIflag', n_jobs=None) FrameT[source]
Apply the High Influx (HI) filter to the precipitation data.
A rainfall measurement that is significantly larger than the median of the surrounding stations is flagged. If that median is lower than
phi_a, the observation is flagged when it is larger thanphi_b. If the median is equal to or larger thanphi_a, the observation is flagged when it is larger thanmedian * phi_b / phi_a. The flag is -1 whenever fewer thann_statneighboring stations report an observation.- Parameters:
data (
TypeVar(FrameT)) – Long format DataFrame with a regular time series per station, as returned bypwsqc.prepare_timeseries(). Either apandas.DataFrameor apolars.DataFrame.neighbors (
Mapping[int,Sequence[int]]) – Mapping of a station id to the ids of its neighbors, as returned bypwsqc.find_station_neighbors().n_stat (
int) – Minimum number of neighboring stations with an observation. (default:5)phi_a (
float) – Rainfall threshold of the median of the neighbors in mm. (default:0.4)phi_b (
float) – Rainfall threshold of the station itself in mm. (default:10)id_col (
str) – Column name for station identifier. (default:'intern_id')date_col (
str) – Column name holding the (interval end) timestamps. (default:'date')precip_col (
str) – Column name holding the rainfall of the interval in mm. (default:'precip')flag_col (
str) – Name of the column the flags are written to. (default:'HIflag')n_jobs (
int|None) – Number of threads to spread the stations over. Defaults to os.cpu_count(). (default:None)
- Return type:
TypeVar(FrameT)- Returns:
DataFrame with an additional
HIflagcolumn indicating flagged data points.
- pwsqc.prepare_timeseries(df_raw, freq='5min', id_col='intern_id', date_col='date', precip_col='precip', keep_cols=None) FrameT[source]
Prepare a time series DataFrame to be ready for PWSQC processing.
This includes: - Ensuring the timestamp column is timezone aware. - Resampling the data to a uniform frequency (default: 5 minutes). - sorting the DataFrame by time and station ID.
The rows of the result are not the rows of the input: observations within the same interval are summed up and the intervals a station did not report are added as missing values. An additional column can hence only be carried through unchanged if it is constant within a station, which is the case for station metadata such as a location or a city id.
- Parameters:
df_raw (
TypeVar(FrameT)) – Input DataFrame with a timestamp column and precipitation data, either apandas.DataFrameor apolars.DataFrame.freq (
str) – Frequency string for resampling (default is ‘5min’). Both the pandas aliases and the polars duration strings are understood. (default:'5min')id_col (
str) – Column name for station identifier. (default:'intern_id')date_col (
str) – Column name holding the (interval end) timestamps. (default:'date')precip_col (
str) – Column name holding the rainfall of the interval in mm. (default:'precip')keep_cols (
Sequence[str] |None) – Additional columns to carry through unchanged, they have to be constant within a station. Defaults to every additional column ofdf_raw, pass an empty sequence to drop them all. (default:None)
- Return type:
TypeVar(FrameT)- Returns:
Resampled DataFrame with a uniform time index, of the same type as
df_raw.
- pwsqc.station_outlier_filter(data, neighbors, n_stat=5, m_int=4032, m_rain=100, m_match=200, gamma=0.15, dbc=1.24, id_col='intern_id', date_col='date', precip_col='precip', fz_col='FZflag', hi_col='HIflag', flag_col='SOflag', bias_col='bias', n_jobs=1) FrameT[source]
Apply the Station Outlier (SO) filter to the precipitation data.
A station is an outlier when it shows very different rainfall dynamics than its neighbors. Every interval is compared with the neighboring stations over a previous period of
m_intintervals, or any longer period containing at leastm_rainintervals with nonzero rainfall. Only neighbors with more thanm_matchoverlapping intervals are considered. The flag is set to 1 if the median of the Pearson correlations falls short ofgammaand to -1 if the correlation is known of fewer thann_statneighboring stations.The intervals flagged by the FZ and the HI filter are excluded from the comparison, so both filters have to be applied first.
The median relative bias of the same comparison is written to the
biascolumn, it is the input ofpwsqc.bias_correction().- Parameters:
data (
TypeVar(FrameT)) – Long format DataFrame with a regular time series per station and the flags of the FZ and the HI filter. Either apandas.DataFrameor apolars.DataFrame.neighbors (
Mapping[int,Sequence[int]]) – Mapping of a station id to the ids of its neighbors, as returned bypwsqc.find_station_neighbors().n_stat (
int) – Minimum number of neighboring stations to compare with. (default:5)m_int (
int) – Number of intervals of the comparison window. (default:4032)m_rain (
int) – Number of nonzero rainfall intervals the window must contain. (default:100)m_match (
int) – Minimum number of overlapping intervals of a neighbor. (default:200)gamma (
float) – Threshold of the median correlation with the neighbors. (default:0.15)dbc (
float) – The default bias correction factor. It does not affect the flags, but the relative bias is computed against the corrected neighbors. (default:1.24)id_col (
str) – Column name for station identifier. (default:'intern_id')date_col (
str) – Column name holding the (interval end) timestamps. (default:'date')precip_col (
str) – Column name holding the rainfall of the interval in mm. (default:'precip')fz_col (
str) – Column name holding the flags of the FZ filter. (default:'FZflag')hi_col (
str) – Column name holding the flags of the HI filter. (default:'HIflag')flag_col (
str) – Name of the column the flags are written to. (default:'SOflag')bias_col (
str) – Name of the column the median relative bias is written to. (default:'bias')n_jobs (
int|None) – Number of threads to spread the stations over. Defaults to one: the comparison window sums of a station are far larger than the caches, so this loop is limited by the memory bandwidth rather than by the cores and more threads buy little while every one of them holds another set of those sums. (default:1)
- Return type:
TypeVar(FrameT)- Returns:
DataFrame with an additional
SOflagcolumn indicating flagged data points and abiascolumn with the median relative bias.