A test pilot lands an experimental aircraft after a carefully planned series of maneuvers. During the post-flight review, one sensor trace stands out: the aircraft responded differently than predicted during a turn.
What should the team test next? Repeat the turn, try a nearby flight condition, or check the instrumentation? Each choice costs time and resources. The measurements reveal a discrepancy, but they don’t automatically identify the most useful next step.
Across a development campaign, those choices compound. In its 2026 CyPhER Forge solicitation, DARPA describes flight testing as consuming up to one-fourth of an aircraft’s total development cost and more than half of its development time, citing a 2004 RAND study.
DARPA’s CyPhER Forge program targets a tenfold reduction in the test points needed to characterize key quantities of interest to a specified global confidence interval, compared with a conventional campaign flown on the same experimental aircraft.
A Large Physics Model, an AI model trained on physics simulation data, predicts behavior across operating conditions, helping engineers assess new questions before deciding whether another simulation or physical test is needed. Grounded in flight measurements and paired with validated uncertainty estimates, it lets engineers identify where another test would be most valuable. The opportunity is to make each flight improve the plan for the next, reducing the testing burden without replacing mandated tests or other requirements for safety and certification.
The region cleared with model prediction extends well beyond the region cleared by flight alone. Next test cards are chosen from the frontier, where each point removes the most remaining risk per flight hour.
From Measurements to a Better Test Plan
Return to the unexpected turn. Before planning another maneuver, the team needs to understand what that result implies about nearby flight conditions. A Large Physics Model could help engineers assess how the aircraft will respond at nearby flight conditions, provided it incorporates the new measurements and supplies trustworthy uncertainty estimates alongside its revised predictions. Together, those capabilities support a learning digital twin: an updatable model of the aircraft that guides the next test.
For that guidance to matter during a flight, it must arrive before the next maneuver is chosen. CyPhER Forge’s Phase 2 objective targets call for predictions in less than a tenth of a second and model updates in less than a minute.
Separating Model Uncertainty from Sensor Noise
Once the team has an updated twin, the question is what the discrepancy actually signals. Did the model learn too little about the turn’s flight conditions, making another observation at that condition the right next step? Or does sensor noise or changed flight conditions explain the gap, making an instrumentation check or repeated measurements more useful?
The twin’s uncertainty estimates need to distinguish those possibilities, and be validated against observations. An overconfident model leads the team to dismiss the unexpected turn too early. An excessively cautious one recommends repeats after they’ve stopped adding information. The real question is what additional evidence would change the engineering decision.
Recommending the Next Maneuver Without Outrunning the Evidence
Suppose the updated twin points to a nearby flight condition as the highest-value next test. Between sorties, engineers compare that recommendation against a repeat of the original turn, using existing review and approval processes.
An informative maneuver is not necessarily safe to fly. Engineers assess it against approved constraints and the model’s validated operating limits. Out-of-distribution detection and uncertainty quantification flag predictions that fall outside reliable territory; established safety processes take precedence when they do. The model supplies predictions, not permission to act.
After an approved follow-up maneuver, the team reassesses: does the revised model explain the original response, and do the new measurements support its predictions at nearby conditions?
Using Test Data to Make Better Predictions
At Luminary, our SHIFT models use existing simulation data to deliver rapid predictions across operating conditions. Grounded in measurements and validated for the question at hand, those predictions can help engineers answer follow-up questions without repeating the full verification workflow, whether that would involve another simulation or a physical test. For the unexpected turn, the model must account for the measured response and provide trustworthy uncertainty estimates before engineers can decide which additional checks are needed. The goal is to avoid unnecessary repeats, not bypass required verification.
Our work on anchoring Physics AI models to reality explores using wind tunnel pressure measurements to ground predictions in physical evidence. Applying that approach to flight telemetry requires accounting for its specific errors and variability. Faster inference is the starting point. Learning from the aircraft is the essential next step.
At the next post-flight review, success looks like a clearer account of the unexpected turn, a defensible choice about what to test next, and, when the evidence supports it, a reason to close the characterization task rather than schedule another repeat.
For more on how Physics AI is changing the economics of flight test, subscribe to the Physics AI newsletter.