Methods register · analysis methods v1.2.0

Prospective methods and analysis protocol.

Delhi Air Health separates measurement, normalization, comparison and interpretation. These methods sit alongside a separately frozen prospective study protocol that defines the primary 2026–27 analysis window, estimand and eligibility rules.

Air archive maturity

Pilot dataset

3 represented calendar days

Meteorology nodes

5/5

Latest hourly capture complete

Meteorology hours

15

Collector 1.1.0 · meteo-1.1.0

Weather-adjusted model

Locked

Requires ≥30 archive days + sufficient meteorology coverage

Primary research question

How does ambient particulate pollution vary across Delhi in time and space?

The primary pollutant outcome is PM2.5. PM10 is secondary. The fundamental observational unit is a station-feed × hour; campus labels are links to nearby ambient reference monitors, not direct campus-exposure measurements.

Open the prospective study protocol →
Internal protocol frozenPrepared · not submittedMeteorology collectingPilot gate active

Observation model

What counts as one datum?

Air pollution

One CPCB-linked station feed in one hourly bucket. When multiple source snapshots land in the same feed-hour, the latest source timestamp is selected deterministically.

Missingness

Missing hours remain missing. No temporal interpolation and no substitution from a neighbouring monitor.

Physical sites

Multiple feeds at identical coordinates are preserved as separate source feeds but flagged as the same physical site for independence-sensitive analyses.

Campus interpretation

A campus inherits a nearby monitor reference with explicit distance. It is never treated as an on-campus dosimeter or a personal exposure measurement.

Comparison rules

Avoid comparing different clocks.

Paired campus differences

Compare only hours observed at both linked monitors. For hour t: Dₜ = PM2.5(comparator,t) − PM2.5(reference,t).

Diurnal summaries

First calculate the median across reporting Delhi feeds in each Delhi-local hour, then group those network-hour medians by local clock hour across days.

Station correlation

Pearson r uses shared hourly PM2.5 observations. Same-physical-site pairs are excluded from the independent-pair distribution.

No causal language

Correlation and raw paired differences are descriptive. They do not establish a local pollution source or causal campus effect.

Prospective meteorology archive

Weather is a covariate, not an afterthought.

Beginning with v2.10.0, five fixed Delhi meteorology reference nodes are captured hourly through Open-Meteo using a locked ECMWF IFS HRES 9 km model selection. They are model-derived gridded values, not weather-station measurements. The full provider response and SHA-256 digest are preserved for every collector run; pre-lock Best Match rows are retained but excluded from the protocol-eligible weather series.

Latest model-derived hour

8 Sept 2026, 11:30 am

0 review flags · 85 stored node-hours

CovariateUnitResearch role
Temperature at 2 m°CThermal state and mixing context
Relative humidity at 2 m%Particle hygroscopic growth and weather context
Precipitationmm/hWet scavenging context
Surface pressurehPaSynoptic context
Wind speed at 10 mkm/hDispersion / stagnation context
Wind direction at 10 mdegreesDirectional transport context
Boundary-layer heightmVertical dilution / mixing depth

Spatial design

The five nodes were frozen from a deterministic five-cluster partition of the 43 unique active Delhi monitor coordinates present at v2.10.0. Future station-level models will use the nearest fixed node rather than moving weather reference points after seeing results.

Provider boundary

Open-Meteo's “best match” model selection may change as upstream models evolve. Delhi Air Health therefore stores provider coordinates, collection time, exact requested URL, raw response and digest. Weather values are never described as instrument observations.

Evidence gates

The language should strengthen only when the dataset does.

<7 days

Pilot

Pipeline validation and descriptive inspection only. No typical-cycle or persistent spatial claim.

7–29 days

Preliminary

Short-term patterns may be described, but remain highly sensitive to the represented days.

30–89 days

Early longitudinal

Weather-adjusted exploratory models may begin if completeness criteria are also met.

90–364 days

Stronger longitudinal

More stable temporal and spatial summaries become possible; seasonal coverage is still incomplete.

≥365 days

Annual archive

A complete annual cycle can be analysed, with season-specific sensitivity analyses.

Pre-specified weather-adjusted analysis

Do not fit it yet.

The first adjusted model is intentionally defined before it is permitted to run. It will begin only after at least 30 represented archive days and adequate meteorology completeness.

PM2.5ₜ = β₀ + hour-of-day + day-of-week
+ temperature + humidity + precipitation
+ wind speed + wind direction terms
+ boundary-layer height + station effect + εₜ

Exact encoding, diagnostics, missing-covariate rule and sensitivity analyses must be frozen in a new analysis-spec version before the first model result is published. This equation is a protocol commitment, not a current result.

Current gate

Weather-adjusted modelling remains locked

Current archive: 3 air-data days and 15 protocol-eligible meteorology hours. Until the gate opens, weather data are collected prospectively but are not used to manufacture an early adjusted result.

Missing data

No single imputation. Complete-case/shared-hour denominators are stated with every analysis. A future model's missing-covariate rule must be frozen before use.

Exclusions

QC flags do not automatically delete observations. Any exclusion rule requires a versioned protocol amendment plus sensitivity analysis with flagged data retained.

Multiple testing

Early analyses are descriptive. If formal hypothesis testing is introduced, primary contrasts and multiplicity handling must be specified before results are inspected.

Reproducibility boundary

Live pages are exploratory; frozen releases are citable research objects.

A formal result must identify its software version, analysis-spec version, exact time window, station roster, completeness denominator and preferably a frozen data release. Corrections create a new release identifier rather than mutating an old one.