Aller au contenu principal
Accès ouvert déclaré 2023 editorial

Discussion of ‘Statistical inference for streamed longitudinal data’

1Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : us, cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

We congratulate the authors for their contribution (Luo et al., 2023) to the development of a timely streaming data analytic. Our discussion focuses on the robustness of their proposed inference against potential outliers or contaminated data cases that are pervasive in practice. Arguably, the issue of data quality is a legitimate concern, which is particularly relevant to streaming data that may arrive voluminously and perpetually. Assuming the absence of abnormal data cases, as considered in Luo et al. (2023), is unrealistic and impractical. Thus, in this commentary we attempt a discussion of the robustness through simulation experiments for the proposed statistical inference, an important methodology aspect that has been ignored by the authors in their paper. We think that statistical robustness is relevant in the proposed methodological framework for the analysis of longitudinal streaming data. This is rooted in the fact that the quadratic inference function technique of Qu et al. (2000) adopted in this paper actually falls within the seminal theory of the generalized method of moments of Hansen (1982), which has been reported as being robust against data contamination (Ronchetti & Trojani, 2001). Specifically, Qu & Song (2004) investigated the robustness property of the quadratic inference function for longitudinal data analysis and found that the quadratic inference function is organically bonded with a certain automatic downweighting scheme that is responsible for robust estimation and inference in the presence of abnormal data cases. Thus, it is worth investigating and confirming such an important property in the context of streaming longitudinal data analytics as studied by Luo et al. (2023). We design a simulation experiment to numerically evaluate the performance of the quadratic inference function’s built-in downweighting scheme in terms of statistical robustness. Adopting the same simulation setting as in § 4.2 of Luo et al. (2023), we generate streaming longitudinal data of 50 batches, each batch consisting of 100 subjects and each subject having a fixed number of 20 repeated measurements. We specify two levels of data contamination severity with respective proportions of outliers: 0.1% (low) and 0.5% (high) of the total sample size, i.e., 100 000. Outliers are distributed by the means of either random or fixed allocation. Random allocation refers to, in the scenario of low contamination, a set-up in which 100 outliers, i.e., the low contamination case, are randomly allocated to 10 randomly selected batches evenly, such that on average each contaminated batch contains 10 outliers that are expected to pollute 10% of subjects within the batch. In the case of high contamination, the random allocation scheme allocates 500 outliers, i.e., the high contamination case, to 10 randomly selected batches evenly, resulting in each contaminated batch containing 50 outliers such that nearly 40%, around 39.7%, of subjects within such a batch are expected to receive at least one outlier. Likewise, using the fixed allocation scheme, we allocate the same amounts of outliers, low or high, to 10 successive target batches, starting at batch 21, a random starting position, in the data streams. As a result, batches 21–30 are polluted by abnormal data cases. Finally, each abnormal outcome is created by a multiplicative type of contamination, 20 × original-y, where such y is randomly selected at the probability of 0.1 (low) or 0.4 (high) within a contaminated batch. We also consider an alternative form of contamination as of a systematic shift, 20 + original-y, and find similar results; thus, we choose to report the results obtained from the multiplicative form of contamination in our discussion below. Figure 1 reports average point estimates of (a) batch-varying parameter β1(b) and (b) interbatch linkage parameter q over 100 rounds of simulations, where the solid curve in (a) is the true parameter curve. We choose to report β^1(b) because this parameter contributes batch heterogeneity as opposed to the other two batch-constant parameters β0 and β2 in the simulation model. Below are a few remarks drawn from this simulation study. Average estimates for (a) batch-varying parameter β1(b) and (b) interbatch linkage parameter q(b) at zero contamination and low and high data contamination levels under fixed allocation (FA) and random allocation (RA). Light contamination: no matter how outliers occur in the data batches, either via fixed or random allocation, they do not exert noticeable estimation bias. This is not surprising as the quadratic inference function is known to tolerate mild data contamination. This confirms the estimation robustness published in the literature for the case of a single longitudinal data batch (Qu & Song, 2004). When the contamination severity increases, estimation biases rise and become apparent. Under random allocation, estimation bias appears to be more pronounced around both peak and valley where the covariate-outcome association exhibits strong changes. Under fixed allocation, estimation bias appears almost exclusively over the range of 21–30 batches, over which overestimation and underestimation biases are somewhat evenly split. We hypothesize that such a structured positive or negative bias sequence may be related to this fact: the proposed method is to achieve global optimization that implicitly enforces certain cancellations in estimation errors over the batches. The proposed method demonstrates strong resilience to local data contamination, judged by the evidence that the estimation bias vanishes quickly after batch 30 under the fixed allocation scheme. In other words, the proposed method seems to have a short memory of data contamination and quickly shrugs off the influence of abnormal data soon after moving out from the data contamination zone. Such empirical evidence of resilience for data contamination suggests a kind of self-correction, which is worth a theoretical investigation. We speculate that such resilience behaviour may have something to do with the interbatch linkage parameter q that characterizes the amount of historical information being used in the current update of model parameters. Figure 1(b) shows that in general the proposed method yields lower q-values for batches falling in the contamination zone, i.e., batches 21–30, regardless of contamination severity. This is somewhat surprising and thus we wonder whether or not the weighting function, Wbj=qtb−tjInj,0<q<1⁠, may play any role in alleviating the influence of contaminated batches in the subsequent updates of model parameters. It would be useful to design a different weighting function that could achieve a robustness guarantee in terms of short memory with contaminated batches. We also study the distribution of the proposed estimator. Table 1 lists the quartiles of 100 estimates for parameter β1(b) at zero contamination and low and high contamination under fixed and random allocation schemes. These statistics of variability are useful to understand the influence of data contamination on statistical inference. Our key findings are as follows. (i) In the case of low contamination, regardless of fixed or random allocation, the estimated quartiles are close to those obtained in the case of no contamination. Together with the previous finding of ignorable estimation bias in these two cases, the proposed method may provide a robust valid inference. (ii) In the case of high contamination, comparing the quartiles from these scenarios suggests notable discrepancies, especially among the Q1 and Q3 estimates. With nonignorable estimation bias in these cases, valid inferences are not given by the proposed method without certain further handling of data contamination. Quartiles of estimation variability over 100 rounds of simulations for model parameter β1(b) at zero contamination and low and high data contamination levels under fixed allocation and random allocation

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Discussion of ‘Statistical inference for streamed longitudinal data’
Date Crossref
15/11/2023
Éditeur
Oxford University Press (OUP)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Statistical Methods and InferenceStatistical Methods and Bayesian InferenceAdvanced Statistical Methods and Models

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.