Google's AI Contrail Avoidance Moves to an Airspace Trial
Operation Blue Skies will test whether AI-guided altitude changes reduce warming contrails across busy North Atlantic airspace, not just selected flights.
Hundreds of commercial flights crossing the North Atlantic are expected to receive small altitude-change instructions during the next two winters. The instructions will come through ordinary air-traffic-control channels, but the regions to avoid will be identified with contrail forecasts that combine machine learning and atmospheric physics.
The 30-month programme, called Operation Blue Skies, is the first attempt to coordinate contrail avoidance across an entire oceanic airspace rather than within one airline or a small set of flights. Google says roughly 10,000 flights will pass through the trial area during test hours each year and only a small percentage will be diverted. The result cannot be inferred from Google’s earlier headline that compliant flights formed 62% fewer observable contrails.
The new experiment has a harder denominator: every relevant flight in the region, including aircraft that cannot change altitude, forecast targets that disappear before arrival, and manoeuvres that move rather than eliminate a warming trail.
The trial moves the decision into air-traffic control
Operation Blue Skies covers Shanwick Oceanic airspace, the eastern half of the North Atlantic corridor managed by NATS. The University of Cambridge describes a £5 million programme with two winter operational trials during 2026–2027 and 2027–2028, about 20 to 40 test days in each winter, and independent assessment of the climate result.
Associated Press reporting adds the expected operational scale: hundreds of flights may deviate by up to 2,000 feet, representing about 1% to 5% of traffic passing through the airspace during the trial. Those are planned parameters, not outcomes. The first operational period has not yet produced a measured reduction.
Moving the intervention from airline dispatch into air-traffic control changes the system being tested. A flight planner can propose a lower-impact route for one carrier. A controller has to preserve safe separation among aircraft from many airlines, account for weather and capacity, communicate a workable instruction, and handle what happens when the preferred altitude is unavailable.
The consortium divides those responsibilities. Google’s announcement assigns NATS the airspace operation, safety assessment, controller training, and coordination; Contrails.org the forecast assessment and trial design; Imperial College London and Cambridge the outcome evaluation; and the Met Office the development and checking of a UK forecast capability. Google is contributing forecast models, satellite analysis, engineering, and computing.
That structure matters more than the label “AI trial.” The UK Jet Zero Taskforce report framed contrail mitigation as an operational problem spanning the aviation community. Operation Blue Skies tests whether those actors can turn a forecast into a safe regional intervention with evidence that survives outside the organization producing the forecast.
From forecast to regional evidence
Operation Blue Skies adds air-traffic control and an airspace-wide outcome to a workflow previously tested inside one airline's flight planning.
- Weather and observations
Forecast weather describes the upper atmosphere; satellite detections provide observed contrail evidence.
- Two-part forecast
A neural network estimates formation probability while a physics model estimates the warming effect.
- Controller decision
NATS must combine forecast regions with traffic, separation, weather, and safety constraints before requesting an altitude change.
- Flight outcome
Selected aircraft may move by up to 2,000 feet; many flights remain unchanged and form part of the regional result.
- Independent assessment
Researchers must connect forecasts, instructions, flown paths, observations, fuel, and operations to net airspace-scale impact.
Did the whole Shanwick trial region produce less warming than it would have without intervention—not merely fewer contrails among flights that completed a requested manoeuvre?
Two models answer different questions
Contrails form when hot, moist aircraft exhaust mixes with sufficiently cold and humid air. Some vanish quickly. Persistent contrails can develop into cirrus clouds, and their net effect is often warming because they trap outgoing infrared radiation, although sunlight reflected by daytime trails complicates the result.
Google’s public forecasting documentation separates the prediction into two parts:
- A deep neural network estimates the probability that a flight waypoint lies in a contrail-likely zone. It consumes numerical weather features such as humidity, temperature, wind, cloud ice, time, latitude, and altitude, and was trained against satellite contrail detections.
- The physics-based Contrail Cirrus Prediction model, or CoCiP, simulates formation, persistence, movement, and radiative forcing—the change in Earth’s energy balance associated with the trail. It uses weather ensembles, aircraft and flight-path information, and cloud microphysics.
The published system multiplies the formation probability by an estimated effective energy forcing, then maps that value to a severity index from zero to four. In plain language, one model asks, “How likely is a persistent trail here?” and the other asks, “If it forms, how much warming could it cause?”
The distinction prevents a common evaluation mistake. Avoiding every likely contrail is not necessarily the best climate strategy because contrails differ in duration and radiative effect. But an impressive forecast score is not the final answer either. The trial still has to show that the ranked regions were accurate at operational lead times and that controllers could act on the highest-value opportunities.
The public documents do not establish that every model version, input, or severity calculation in Google’s Contrails API will be used unchanged during Operation Blue Skies. The Met Office says it is developing a global contrail-impact forecast and will help assess performance and uncertainty. Results should therefore identify the exact forecast version and the contribution of each provider rather than treating “AI-powered” as one stable instrument.
Why 62% is the wrong expectation for a whole airspace
The strongest prior evidence comes from an airline-led randomized trial conducted on scheduled American Airlines flights. Its March 2026 preprint, which has not been peer reviewed, reports 2,400 transatlantic flights in the study workflow. Of 1,232 treatment flights marked eligible for avoidance, observed contrail formation fell 11.6% relative to the control group. Among the 112 flights that actually flew the optimized plan as intended, the reduction was 62%.
Both numbers answer legitimate questions. The 62% per-protocol estimate asks how the manoeuvre performed among the small subset that completed it. The 11.6% intent-to-treat estimate is closer to the operational result because it retains eligible treatment flights whether or not dispatchers released the alternate plan and crews flew it as intended. Our AI evaluation explainer develops the same rule: the denominator has to match the real decision the result is being used to support.
Only 112 of 1,232 intent-to-treat flights entered the study’s strictest per-protocol result. The authors say that take rate was shaped by the particular trial and may not transfer to other operations, but they also identify execution as the main bottleneck to network-scale benefit. No statistically significant fuel-use difference was observed after accounting for aircraft type. That is reassuring evidence for the tested workflow, not a guarantee that different altitude changes across a congested airspace will have no fuel cost.
The study also measured contrail formation primarily with satellite imagery and an automated flight-attribution method. The authors warn that satellites do not observe every contrail and that reductions in total contrail distance or warming could be smaller if the unobserved trails respond differently. Most authors were affiliated with Google, Contrails.org, American Airlines, or Flightkeys, so independent airspace-scale assessment is especially valuable.
| Question | Earlier airline trial | Operation Blue Skies must establish |
|---|---|---|
| Unit of intervention | An alternate plan inside one airline workflow | Controller-coordinated changes among many airlines |
| Main denominator | Eligible treatment flights and compliant subsets | All relevant traffic and conditions in the trial region |
| Observed outcome | Satellite-attributed contrail formation | Net regional climate effect with an explicit counterfactual |
| Operational constraint | Dispatcher release and flight-plan adherence | Separation, capacity, weather, workload, and instruction compliance |
| Climate boundary | Forecast-derived climatological warming estimate | Independent assessment of avoided, shifted, and newly formed trails |
| Cost boundary | No significant fuel difference in the tested groups | Fuel, delay, capacity, training, and repeatable operating cost |
The table is not a claim that the new trial has already solved these design questions. It is the evidence needed to distinguish a regional result from a successful subset.
A regional result needs more than before-and-after pictures
The cleanest airspace-scale evaluation would specify trial days and decision rules before outcomes are known, preserve a credible comparison for what traffic would have done without the intervention, and report both assignment and compliance. Otherwise, favorable weather or unusually light traffic could be mistaken for a forecast-and-control effect.
At minimum, the published result should make six connections auditable:
- forecast regions to the weather data, model version, issue time, and calibration observed on that day;
- eligible flights to the reason each was selected or excluded;
- controller instructions to the path and altitude the aircraft actually flew;
- flown paths to observed and model-estimated contrails, including detection coverage;
- local avoidance to any contrail that appears elsewhere along the route; and
- the regional climate estimate to fuel, delay, capacity, and safety outcomes.
No single metric closes the case. A lower observed contrail rate with worse fuel burn could still be beneficial, but the trade needs a stated climate accounting method. A high instruction-compliance rate could coexist with poor forecasts. A model could accurately identify cold, humid regions yet rank their warming impact badly. Controller workload could make the intervention impractical outside quiet winter test days even if the atmospheric mechanism works.
This is why the move into Shanwick airspace is consequential. The earlier trial provided evidence that forecast-guided manoeuvres can reduce observable contrails when selected flights complete them. Operation Blue Skies can test the missing system: whether forecasting, air-traffic control, airline participation, and independent measurement combine to reduce warming across a region. Its first useful headline will be a protocol and a denominator. The percentage should come later.
Sources
- Google announcement of Operation Blue Skies
- University of Cambridge Operation Blue Skies trial overview
- Met Office Operation Blue Skies forecast and assessment role
- Google Contrails API forecasting documentation
- Efficacy of Scalable Airline-led Contrail Avoidance preprint
- Associated Press report on the North Atlantic trial
- UK Jet Zero Taskforce contrail mitigation report