Example: Audit the dataset, identify the highest-risk quality issues, then generate an idempotent SQL cleanup pipeline and a PySpark equivalent.