Reasoning
Data contamination from AI-generated content is already being documented in academic literature (e.g., 2023 papers on "model collapse" showing quality degradation when models train on synthetic data), and major AI labs have publicly acknowledged this as a concern. By 2028 (4 years away), the volume of AI-generated content will grow exponentially while detection methods remain imperfect, making contamination in training datasets nearly inevitable at scale. Historical precedent suggests that once a technical problem becomes theoretically possible and economically incentivized (cheaper synthetic training data), it becomes documented within 3-5 years; we're already 2+ years into this cycle with preliminary evidence emerging.Key uncertainty
Whether industry-wide adoption of synthetic data filtering and watermarking technologies can scale fast enough to prevent widespread documented quality degradation, or whether competitive pressure to train on all available data will override contamination concerns and accelerate the problem's visibility.