Air archive maturity
Pilot dataset
3 represented calendar days
Methods register · analysis methods v1.2.0
Delhi Air Health separates measurement, normalization, comparison and interpretation. These methods sit alongside a separately frozen prospective study protocol that defines the primary 2026–27 analysis window, estimand and eligibility rules.
Air archive maturity
Pilot dataset
3 represented calendar days
Meteorology nodes
5/5
Latest hourly capture complete
Meteorology hours
15
Collector 1.1.0 · meteo-1.1.0
Weather-adjusted model
Locked
Requires ≥30 archive days + sufficient meteorology coverage
Primary research question
The primary pollutant outcome is PM2.5. PM10 is secondary. The fundamental observational unit is a station-feed × hour; campus labels are links to nearby ambient reference monitors, not direct campus-exposure measurements.
Open the prospective study protocol →Observation model
Air pollution
One CPCB-linked station feed in one hourly bucket. When multiple source snapshots land in the same feed-hour, the latest source timestamp is selected deterministically.
Missingness
Missing hours remain missing. No temporal interpolation and no substitution from a neighbouring monitor.
Physical sites
Multiple feeds at identical coordinates are preserved as separate source feeds but flagged as the same physical site for independence-sensitive analyses.
Campus interpretation
A campus inherits a nearby monitor reference with explicit distance. It is never treated as an on-campus dosimeter or a personal exposure measurement.
Comparison rules
Paired campus differences
Compare only hours observed at both linked monitors. For hour t: Dₜ = PM2.5(comparator,t) − PM2.5(reference,t).
Diurnal summaries
First calculate the median across reporting Delhi feeds in each Delhi-local hour, then group those network-hour medians by local clock hour across days.
Station correlation
Pearson r uses shared hourly PM2.5 observations. Same-physical-site pairs are excluded from the independent-pair distribution.
No causal language
Correlation and raw paired differences are descriptive. They do not establish a local pollution source or causal campus effect.
Prospective meteorology archive
Beginning with v2.10.0, five fixed Delhi meteorology reference nodes are captured hourly through Open-Meteo using a locked ECMWF IFS HRES 9 km model selection. They are model-derived gridded values, not weather-station measurements. The full provider response and SHA-256 digest are preserved for every collector run; pre-lock Best Match rows are retained but excluded from the protocol-eligible weather series.
Latest model-derived hour
8 Sept 2026, 11:30 am
0 review flags · 85 stored node-hours
| Covariate | Unit | Research role |
|---|---|---|
| Temperature at 2 m | °C | Thermal state and mixing context |
| Relative humidity at 2 m | % | Particle hygroscopic growth and weather context |
| Precipitation | mm/h | Wet scavenging context |
| Surface pressure | hPa | Synoptic context |
| Wind speed at 10 m | km/h | Dispersion / stagnation context |
| Wind direction at 10 m | degrees | Directional transport context |
| Boundary-layer height | m | Vertical dilution / mixing depth |
Spatial design
The five nodes were frozen from a deterministic five-cluster partition of the 43 unique active Delhi monitor coordinates present at v2.10.0. Future station-level models will use the nearest fixed node rather than moving weather reference points after seeing results.
Provider boundary
Open-Meteo's “best match” model selection may change as upstream models evolve. Delhi Air Health therefore stores provider coordinates, collection time, exact requested URL, raw response and digest. Weather values are never described as instrument observations.
Evidence gates
<7 days
Pilot
Pipeline validation and descriptive inspection only. No typical-cycle or persistent spatial claim.
7–29 days
Preliminary
Short-term patterns may be described, but remain highly sensitive to the represented days.
30–89 days
Early longitudinal
Weather-adjusted exploratory models may begin if completeness criteria are also met.
90–364 days
Stronger longitudinal
More stable temporal and spatial summaries become possible; seasonal coverage is still incomplete.
≥365 days
Annual archive
A complete annual cycle can be analysed, with season-specific sensitivity analyses.
Pre-specified weather-adjusted analysis
The first adjusted model is intentionally defined before it is permitted to run. It will begin only after at least 30 represented archive days and adequate meteorology completeness.
Exact encoding, diagnostics, missing-covariate rule and sensitivity analyses must be frozen in a new analysis-spec version before the first model result is published. This equation is a protocol commitment, not a current result.
Current gate
Current archive: 3 air-data days and 15 protocol-eligible meteorology hours. Until the gate opens, weather data are collected prospectively but are not used to manufacture an early adjusted result.
No single imputation. Complete-case/shared-hour denominators are stated with every analysis. A future model's missing-covariate rule must be frozen before use.
QC flags do not automatically delete observations. Any exclusion rule requires a versioned protocol amendment plus sensitivity analysis with flagged data retained.
Early analyses are descriptive. If formal hypothesis testing is introduced, primary contrasts and multiplicity handling must be specified before results are inspected.
Reproducibility boundary
A formal result must identify its software version, analysis-spec version, exact time window, station roster, completeness denominator and preferably a frozen data release. Corrections create a new release identifier rather than mutating an old one.