BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Institute of Applied Statistics and Data Science - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Institute of Applied Statistics and Data Science
X-ORIGINAL-URL:https://isrt.ac.bd
X-WR-CALDESC:Events for Institute of Applied Statistics and Data Science
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:UTC
BEGIN:STANDARD
TZOFFSETFROM:+0000
TZOFFSETTO:+0000
TZNAME:UTC
DTSTART:20250101T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=UTC:20260105T140000
DTEND;TZID=UTC:20260105T150000
DTSTAMP:20260101T042836Z
CREATED:20260101T042836Z
LAST-MODIFIED:20260101T042836Z
UID:8479-1767621600-1767625200@isrt.ac.bd
SUMMARY:Applied Statistics and Data Science Seminar on Monday 5 January 2026
DESCRIPTION:Title: A Moment-Based Generalization To Post-Prediction Inference \nVenue\, date and time: ISRT\, 5 January 2026\, 2 pm \nSpeaker: Awan Afiaz\, PhD candidate at the Department of Biostatistics\, University of Washington Seattle\, WA\, USA and ISRT alumnus \nAbstract: \nAs artificial intelligence (AI) and machine learning (ML) become increasingly integrated into scientific research\, investigators frequently substitute predicted outcomes for expensive or difficult-to-measure data. However\, treating these AI/ML-generated predictions as true observations can lead to biased estimates and anti-conservative inference. While high predictive accuracy is often assumed to ensure valid downstream inference\, statistical challenges in inference with predicted data (IPD) fundamentally reduce to two sources of error: bias\, when predictions systematically distort relationships among variables\, and variance\, when uncertainty from prediction models is inadequately propagated. Wang et al. (2020) introduced post-prediction inference (PostPI)\, a pioneering method that addresses this challenge by modeling the relationship between predicted and observed outcomes in a small gold-standard dataset to calibrate inference in larger unlabeled samples. PostPI has been influential in formalizing the IPD problem and demonstrating how naive approaches fail to appropriately reflect uncertainty. However\, PostPI relies on a critical assumption: that prediction errors are uncorrelated with covariates of interest. In realistic settings where prediction algorithms exhibit systematic errors related to input features\, this assumption is often violated\, leading to biased parameter estimates and inadequate error control. We revisit PostPI in light of recent methodological advances and propose a moment-based generalization that relaxes this restrictive assumption. Our extension explicitly accounts for the covariance between prediction errors and covariates by incorporating an additional correction term estimated from the labeled dataset. This approach yields unbiased point estimates under standard conditions while incorporating a simple scaling factor that appropriately reflects the contribution of relationship model uncertainty regardless of sample size allocation. Through extensive simulations across three data-generating scenarios\, we demonstrate that our method maintains nominal Type-I error rates and achieves proper coverage probability\, even when the labeled sample is substantially smaller than the unlabeled sample settings where both naive approaches and original PostPI fail. Our work illustrates the classic bias-variance trade-off inherent to IPD’s challenges and confirms that there is no free lunch when substituting predicted outcomes for true measurements.
URL:https://isrt.ac.bd/event/applied-statistics-and-data-science-seminar-on-monday-5-january-2026/
CATEGORIES:seminar
END:VEVENT
END:VCALENDAR