Hailing Ocean · Research

Pilot-grounded failure discovery for autonomous aviation.

A research framework and prospective evaluation protocol connecting pilot judgment, simulation, and reproducible evidence.

Open ResearchHailing Ocean15 references

This paper proposes a method and evaluation plan. No Hailing Ocean experimental results are reported.

Abstract

Autonomous aviation must be evaluated under interacting disturbances, ambiguous observations, and decisions whose consequences unfold over time. Simulation can expose such conditions, but a simulator’s scenario distribution and failure definitions determine what its tests can discover. We propose Hailing Ocean, a framework that connects synchronized pilot telemetry, reviewed decision annotations, and optional voice recordings to constrained simulation search and reproducible evaluation. Its central hypothesis is that pilot evidence can improve the discovery of distinct, operationally plausible failure mechanisms under a fixed testing budget. The framework separates scene generation from flight-dynamics validation, human disagreement from objective failure, and adversarial discovery from operational risk estimation. We specify a data representation, a pilot-informed acquisition heuristic, and a prospective study involving a bounded mountainous-flight scenario with propulsion degradation. Evaluation compares random, space-filling, optimization-based, and sampling-based search, with ablations isolating the value of telemetry, decision annotations, and language. Primary outcomes are adjudicated failure diversity and discovery efficiency; secondary outcomes include held-out policy performance, annotation reliability, and simulator sensitivity. The intended contribution is a falsifiable protocol for converting pilot experience into aviation evaluation evidence. Effectiveness, transfer to real flight, and generalization across aircraft remain unestablished.

Keywords: autonomous aviation; simulation; failure discovery; human judgment; multimodal data; safety evaluation.

Introduction

An aircraft can encounter several individually familiar conditions whose combination is poorly represented in a development dataset. Terrain constraints, changing weather, imperfect instrumentation, and degraded propulsion can interact with delayed decisions. Testing isolated faults or measuring mean prediction error may therefore leave important weaknesses unresolved. The practical research problem is how to choose informative simulated situations and preserve evidence explaining why a system failed.

Hailing Ocean’s proposed data layer centers on the relationship between observations, decisions, and outcomes. Flight telemetry records what happened; pilot annotations can identify which cues were noticed, what alternatives were considered, and where uncertainty affected action. These explanations are fallible observations. They may be incomplete, retrospective, or influenced by the interface. Their value should be demonstrated through controlled comparisons rather than assumed from professional expertise.

Research on black-box safety validation distinguishes finding a counterexample from estimating its likelihood [1]. Recent robotics work develops increasingly effective failure-search and adaptation methods [2], [3], [4]. These methods motivate a narrower question for Hailing Ocean: does adding pilot evidence improve the discovery of distinct aviation failure mechanisms beyond what the same simulator and automated search already provide?

This paper proposes three contributions: (1) a provenance-preserving representation linking telemetry, pilot observations, and simulated counterfactuals; (2) an implementable acquisition strategy for directing simulation toward pilot-relevant disagreements while maintaining physical constraints; and (3) a preregistrable evaluation protocol separating discovery, policy improvement, and risk estimation. These are proposed contributions, not claims of demonstrated algorithmic superiority. A broad commercial ambition to serve flying systems is reduced here to one aircraft configuration and one operational domain.

Failure discovery and simulator mismatch

Dawson and Fan formulate failure discovery and repair through sampling, using a prior over external conditions and a cost that favors adverse outcomes [2]. Their evaluations include an aircraft ground-collision-avoidance example alongside other systems. RADIUM extends this direction to end-to-end robotic systems, including perception [3]. These studies establish relevant search machinery; they do not establish that pilot language improves aviation evaluation. Hailing Ocean should reproduce an appropriate existing baseline before attributing gains to its own data layer.

Parashar et al. combine simulation-derived failure information with a limited budget of hardware demonstrations [4]. This is particularly relevant to the gap between a simulated failure and a failure of the corresponding physical system. Their demonstrations concern manipulation and a small racing vehicle. The present proposal uses this work as motivation for explicitly measuring model discrepancy, without transferring its empirical conclusions to aircraft.

RialTo constructs task environments through a real-to-sim-to-real workflow for robust manipulation [5]. Palatial describes commercial tools for producing simulation assets and environments, including physical properties and integration with robotics simulators [6]. Palatial is a product reference, not peer-reviewed evidence of aviation validity. Its possible role here is scene and asset preparation. Neither visual fidelity nor asset metadata alone establishes correct aerodynamic, propulsion, weather, or sensor behavior.

Recent failure monitoring and aviation evaluation

ARMADA couples online failure detection with human shared control and adaptation in robotic tasks [7]. FIPER studies runtime failure prediction for generative robot policies [8]. These 2025 works motivate recording intervention timing and detector behavior, but intervention itself is not proof of an impending objective failure. An intervention can also reflect preference, discomfort, or excessive conservatism.

Khatiri et al. examine relationships between uncertainty and safety in simulated UAV flights [9]. Their findings motivate treating uncertainty as a candidate search signal rather than a sufficient failure label. Repeatable disturbance injection and explicit constraints, as supported by safe-control-gym, provide useful methodological precedents for fair controller comparisons [10].

Two 2026 studies sharpen the evaluation design. AeroCopilotBench evaluates emergency-procedure execution in an executable cockpit with aircraft-specific rules [11]. This addresses procedural performance rather than independently validating a continuous flight-dynamics model. FLY-EVAL++ separates structured validity, physical consistency, safety constraints, and predictive quality in flight prediction [12]. It motivates auditable checks and disaggregated reporting; its thresholds should not be transplanted into another aircraft or mission. Both references inform the protocol, not a claim that a language model is ready to control an aircraft.

The proposed gap

The distinctive hypothesis is that a pilot can identify decision-relevant ambiguities that are poorly captured by an existing test distribution. This hypothesis may be false: annotations may simply duplicate information already present in telemetry, or bias search toward memorable but unrepresentative events. A meaningful contribution requires demonstrating incremental value after accounting for simulation calls, annotation effort, and access to expert knowledge.

Relevance to Aviation Foundation Models: Archer’s ZEE

What Archer has published

Archer’s July 2026 announcement describes ZEE as an aviation-specific foundation model combining ADS-B, air traffic communications, maps and charts, aircraft state, terrain, and weather. It identifies intended uses in airline operations, airspace management, and copilot assistance, with both on-device and server-hosted execution [13]. These are company-reported capabilities and objectives, rather than independently established results in this paper.

Archer’s August publication reports airport-surface trajectory prediction using conditional flow matching and satellite-image features from a vision transformer, alongside testing at Hawthorne Airport and plans for further validation [14]. Its accompanying technical blog, Zee: Predicting the Airspace, frames the model as decision support for pilots, controllers, and airline operators [15]. The sources reviewed here are corporate publications, not peer-reviewed evidence that ZEE autonomously resolves airborne emergencies.

Hailing Ocean’s proposed complementary role

Our interpretation is that aviation foundation models create a potential use case for Hailing Ocean’s pilot-grounded data and evaluation protocol. Operational traces may reveal which trajectory occurred without fully explaining the alternatives a pilot considered. Reviewed decision evidence could help researchers construct ambiguous situations, label decision checkpoints, and investigate whether a model’s forecasts or recommendations remain useful when familiar signals conflict. Whether this information adds value for ZEE specifically remains an empirical question.

The proposed relationship is an evaluation and data interface. It does not assume access to ZEE’s weights, training corpus, internal architecture, or a public API. Hailing Ocean has not established a partnership, integration, benchmark result, or data-supply arrangement with Archer in this manuscript. The same protocol could be applied to another aviation model when an authorized, documented interface is available.

Proposed application to aviation-model evaluation. These are Hailing Ocean study designs, not reported ZEE deficiencies or integrations.
Evaluation question Proposed Hailing Ocean evidence and measurement
Prediction under ambiguous intent Align observed motion with reviewed pilot-intent annotations; measure forecast error and calibrated coverage across future branches.
Sensitivity to degraded inputs Vary timestamps, missing observations, and conflicting cues within validated bounds; report degradation by condition and modality.
Decision-support usefulness Replay matched scenarios with and without model assistance; measure completion, critical violations, workload, and appropriate rejection of poor advice.
Robustness across environments Hold out airports, source events, and operating conditions; record where performance deteriorates and whether uncertainty responds.

A separate surface-prediction evaluation track

The most direct initial connection to the published ZEE work is an airport-surface study. It should remain separate from the mountainous propulsion-degradation experiment: performance on one cannot establish capability on the other. A prospective surface dataset would synchronize motion histories, map context, available communications, and consented pilot annotations. Only information available at the forecast time enters model inputs; retrospective intent labels and future trajectories are reserved for analysis.

With authorized access, freeze the model version, observation window, prediction horizons, sampling budget, and compute conditions. Compare a motion-extrapolation baseline, a map-constrained predictor, and the accessible aviation model on the same held-out episodes. Report displacement errors by horizon, physical/map constraint violations, and uncertainty calibration. Sample-based models can be assessed with proper scoring rules, such as the energy score, where their output contract permits it. Best-of-many trajectory error is supplementary and uses a fixed sample count; it does not establish a well-calibrated predictive distribution.

For conflict warnings derived from forecasts, fix the downstream warning rule on development data and report false alerts, missed events, and warning lead time separately. No forecast score alone demonstrates improved human decisions. That requires the distinct, counterbalanced assistance study described in the table, including the possibility of automation bias. Pilot explanation quality should be evaluated through reviewed evidence, not inferred from fluent model-generated text.

A subsequent data-value experiment would compare telemetry-only development with the addition of structured pilot annotations and, separately, voice-derived annotations, while holding model and training budgets fixed. Fine-tuning is conditional on permission and technical access; otherwise the study is limited to black-box evaluation. Without access, Hailing Ocean can publish a model-agnostic scenario specification, but must not describe it as an evaluation of ZEE.

Problem Formulation

Let z∈𝒵z\in\mathcal Z encode initial conditions, environmental disturbances, fault timing, and observation imperfections. Let ϕ\phi denote simulator parameters, and let π\pi be a frozen candidate policy. A rollout is τ=S(π,z,ϕ)={xt,ot,at}t=0T,\begin{equation} \tau=S(\pi,z,\phi)=\{x_t,o_t,a_t\}_{t=0}^{T}, \end{equation} where xtx_t is latent simulated state, oto_t is the information exposed to the policy, and ata_t is its action. A policy must not receive latent fault states or future information unavailable to the pilot in the corresponding condition. The first study fixes one policy interface; direct control, high-level decision support, and procedural agents are separate experimental tracks.

Aircraft-specific constraints gk(τ)≤0g_k(\tau)\leq0 define admissibility. Their units, tolerances, temporal windows, and severity classes must be established by qualified reviewers before evaluation. Define F(τ)=𝟏{∃k∈𝒦critical:gk(τ)>0}.\begin{equation} F(\tau)=\mathbf 1\{\exists k\in\mathcal K_{\mathrm{critical}}:g_k(\tau)>0\}. \end{equation} Task completion, procedural deviations, and critical violations are reported separately. A convenient composite reward must not allow task progress to cancel a critical violation.

Pilot evidence is DH={(τi,ci,ri,mi)}D_H=\{(\tau_i,c_i,r_i,m_i)\}, with reviewed cues cic_i, optional rationale rir_i, and provenance mim_i. This evidence informs scenario selection and possible learning targets; it does not redefine FF to make agreement with a pilot equivalent to safety.

Two estimands are deliberately separated. Discovery asks how many distinct reproducible failure mechanisms are found within a budget. Risk estimation asks for Rsim(π;ϕ,pref)=𝔼z∼pref[F(S(π,z,ϕ))],\begin{equation} R_{\mathrm{sim}}(\pi;\phi,p_{\mathrm{ref}}) =\mathbb E_{z\sim p_{\mathrm{ref}}}[F(S(\pi,z,\phi))], \end{equation} under a declared reference distribution. This is conditional simulator risk. If prefp_{\mathrm{ref}} is expert-specified because operational data are inadequate, it must be described as such; it is not an empirically established real-world exposure model.

Pilot-Grounded Data and Simulation Pipeline

Synchronized evidence capture

The proposed interface comprises a pilot simulation client and an administrative replay and annotation view. A desktop presentation is the initial study condition. A standalone headset implementation or paired VR client is optional and requires measured timing and interaction characteristics; a working headset integration is not claimed here.

Every stream retains its native timestamps and a mapping to a common monotonic session clock. Resampling must preserve original records, missingness, and synchronization error estimates. Voice capture is opt-in and separable from telemetry collection. Transcripts retain confidence, speaker role, and correction history. A model-generated summary must remain distinguishable from a human statement.

Proposed episode schema. Fields are requirements, not an existing released dataset.
Layer Required evidence
Provenance Session and source-event IDs; aircraft and simulator versions; consent; participant qualification stratum; real-flight or simulated origin.
State and observations Aircraft state, environmental inputs, instruments actually displayed, sensor quality, units, calibration and missing-data masks.
Actions Control commands, discrete selections, intervention timing, input device, interface latency and sampling rates.
Decision evidence Time-bounded cue annotations, alternatives, uncertainty, voice/transcript links, contemporaneous versus retrospective status.
Outcomes Constraint traces, completion status, adjudicated mechanism, reviewer disagreement and replay seed.
Lineage Parent scenario, generation method, parameter changes, split assignment, review status and content hashes.

Pilot qualification, aircraft familiarity, and recent relevant experience are recorded as study covariates. Novice gameplay is retained as a separate population, not silently pooled with professional-pilot evidence. A comparison of spontaneous speech with retrospective debriefing assesses whether eliciting explanations changes performance. Absence of speech cannot be interpreted as absence of reasoning.

Calibration before expansion

Before automated search, the selected simulator is compared with available reference trajectories, aircraft documentation, and controlled subsystem evidence. Calibration uses designated development data, with separate validation cases. Relevant checks include time response, energy behavior, sensor delays, and disturbance response. Acceptance tolerances are aircraft- and task-specific and remain an expert-review requirement for this draft.

Scenario variables are bounded by a documented validity envelope. Generated assets may change appearance without changing validated physics. A proposed variant is rejected or quarantined if it introduces an unsupported physical regime, an inconsistent fault combination, or inaccessible terrain geometry. Uncertain model parameters are evaluated through declared sensitivity ranges. A failure that exists only for an unsupported parameter value is reported as a modeling concern, not automatically as a controller defect.

Pilot-informed acquisition

The initial implementation uses a gradient-free acquisition heuristic so that usefulness does not depend on a differentiable flight simulator. For candidate scenario zz, define A(z)=wss̃(z)+wuũ(z)+wnñ(z)+wdd̃(z).\begin{equation} A(z)=w_s\widetilde s(z)+w_u\widetilde u(z) +w_n\widetilde n(z)+w_d\widetilde d(z). \end{equation} Here ss estimates severity, uu is surrogate epistemic uncertainty, nn measures distance from previously tested scenario descriptors, and dd predicts disagreement with reviewed pilot decision evidence. Tildes denote normalization fixed using development data. Nonnegative weights sum to one and are chosen before held-out testing. This is a proposed heuristic, not a derived posterior or a convergence guarantee.

For the first implementation, dd concerns discrete decision categories at matched observation checkpoints, avoiding an unsupported assumption that human and automated continuous controls should coincide. A lightweight predictor is trained on observed pilot–policy disagreements. Disagreement is only a prioritization signal: its correlation with actual failures is an outcome to measure. If there are insufficient labels to estimate it reliably, the pilot-informed arm is deferred rather than populated with invented labels.

At each round, candidates are generated within the validity envelope. A fixed fraction of the budget is reserved for samples from prefp_{\mathrm{ref}}, while remaining candidates are ranked by AA. The reserved fraction and batch sizes are preregistered. After execution, traces are checked, candidate mechanisms are clustered, and blinded reviewers adjudicate distinctness. Search updates use development episodes only. This workflow can later incorporate a Bayesian failure sampler, but reproducing established sampling methods remains separate from assessing the pilot signal.

Proposed loopCapture pilot evidence →\rightarrow review and align →\rightarrow calibrate bounded scenarios →\rightarrow search and replay →\rightarrow adjudicate failures →\rightarrow freeze an evaluation set.
Policy updates may use a development archive. The final evaluation archive is isolated from search tuning and policy training.

Prospective Study: Mountainous Diversion under Propulsion Degradation

The first scenario family uses one specified twin-engine aircraft configuration in a mountainous environment. The presentation name may be Honeymoon; the scientific task is defined by aircraft state, environmental constraints, observations, and outcomes. Fictional television events are not a source of flight dynamics or operational procedure.

Variants manipulate bounded propulsion degradation, turbulence, visibility, sensor ambiguity, and event timing. A simulated fire indication is included only if the selected subsystem model supports the intended behavior. Aircraft-specific procedures and experimental limits require qualified review. The study does not prescribe real-flight emergency actions and does not deliberately induce failures in operational aircraft.

The principal comparison holds the candidate controller fixed while varying the search method. Human sessions produce development evidence; automated rollouts then measure discovery. A separate policy-improvement experiment may train on each method’s discoveries using equal training budgets and evaluate on an untouched test set. Improvement in this second experiment cannot be inferred merely from finding more failures in the first.

Gamification is restricted to a separately documented interface layer. Points and progression must not reward speed or risk taking in a way that silently changes the scientific task. Research outcomes come from independent trace checks, not the game score. Desktop and VR sessions are analyzed separately until a randomized, counterbalanced comparison supports pooling. Simulator sickness, display latency, and learning effects are recorded.

Evaluation Protocol

Hypotheses and baselines

H1: Pilot-informed acquisition increases adjudicated failure-mechanism diversity at a fixed simulation budget relative to the strongest non-pilot baseline. H2: Reviewed decision annotations add value beyond telemetry alone. H3: Language adds value beyond structured annotations after controlling for annotation cost. H4: Any discovered advantage persists on held-out scenario families and plausible model perturbations. Each hypothesis admits a null or negative result.

Planned search comparisons. All methods receive the same simulator and validity envelope.
Method Purpose
Reference sampling Independent draws from the declared reference distribution.
Space-filling search Coverage-oriented baseline over the same bounded variables.
Black-box optimization Severity-focused search without pilot annotations.
Failure sampling Reproduction or documented adaptation of an established Bayesian sampling baseline [2].
Pilot-informed search Acquisition above, with telemetry-only, structured-decision, and decision-plus-language ablations.

Report simulation calls, elapsed compute, hardware, and pilot/reviewer minutes. Search tuning consumes a separate, equally reported development budget. Comparable methods share initialization sets and evaluation seeds where appropriate. Giving a pilot-informed method extra expert-authored constraints would confound the comparison; all arms receive the same common constraints.

Splits and statistical design

All derivatives of an originating flight or scenario belong to the same data split. Pilot identities are held out for a dedicated generalization analysis. Search development, policy training, and final evaluation are separately versioned. A final evaluation set is frozen before choosing acquisition weights or selecting a policy checkpoint.

An initial feasibility study estimates measurement variance, annotation agreement, session duration, and simulator throughput. The confirmatory sample size is then set through a documented power analysis for H1, using a prespecified smallest effect of interest and participant attrition allowance. This draft does not claim a powered sample size before those quantities are known. Repeated timesteps are not independent participants. Statistical analyses use episode-level outcomes and account for clustering by pilot and originating scenario, for example through a hierarchical model or cluster bootstrap. Primary comparisons and multiplicity handling must be registered before unblinding.

Outcomes and adjudication

The primary endpoint is the number of distinct, reproducible, expert-adjudicated failure mechanisms discovered by the fixed budget. A mechanism taxonomy is established on development data; raw cluster counts alone are insufficient. Reviewers see de-identified traces without the search-method label, record causal hypotheses and uncertainty, and resolve disagreements through a documented procedure. Novel test mechanisms may extend the taxonomy only through the same blinded rule for every method.

Secondary endpoints include time to first validated failure, discovery curves, severity distribution, replay reproducibility, search sensitivity to simulator parameters, false-alarm burden, and annotation agreement. For a separate failure detector, report precision–recall, detection delay, and calibration on held-out episodes. For a separately trained policy, report task completion and each critical constraint violation independently. Mean trajectory error and preference agreement are supplementary metrics.

Counterfactual replay changes one declared factor at a time where feasible. It can support a causal explanation within the simulator but does not establish causality in actual flight. Sensitivity to multiple coupled factors is retained rather than collapsed into a single convenient label.

Independent risk estimation

Stress-search episodes overrepresent difficult conditions. Their failure fraction cannot be reported as an operational failure rate. For a frozen simulator and policy, an independent reference-sampling experiment estimates R̂sim=1N∑i=1NF(S(π,zi,ϕ)),zi∼iidpref.\begin{equation} \widehat R_{\mathrm{sim}}=\frac{1}{N}\sum_{i=1}^{N}F(S(\pi,z_i,\phi)), \qquad z_i\overset{\mathrm{iid}}{\sim}p_{\mathrm{ref}}. \end{equation} Binomial uncertainty intervals are appropriate only under the stated independence and fixed-model assumptions. For zero observed failures, the one-sided 95% upper bound is 1−0.051/N1-0.05^{1/N}, not zero. Model discrepancy and uncertainty in the reference distribution require separate sensitivity analyses; increasing NN does not eliminate them. Adaptive search samples are excluded from this simple estimator.

Results Status and Reporting Commitments

No participant study, benchmark experiment, dataset release, or transfer evaluation is established by this manuscript. Consequently, there are no Hailing Ocean performance numbers to report. The planned results package contains: a cohort and data-quality table; calibration residuals; discovery curves with uncertainty; adjudicated mechanism counts; ablation comparisons including annotation cost; held-out constraint outcomes; and an independent simulator-risk analysis. Negative results, exclusions, crashes of the test harness, and failed reproductions will be reported alongside successful episodes.

A successful first study would show incremental pilot-data value against a competitive baseline under controlled resources. A failure to outperform telemetry-only search would narrow the product thesis: the useful contribution might be structured evaluation and replay rather than language-informed search. Both outcomes are scientifically informative.

Limitations, Governance, and Reproducibility

Simulator search cannot discover failure mechanisms absent from its models. Pilot populations may be small or unrepresentative, expert opinions may conflict, and retrospective narratives may be rationalizations. A single aircraft and mission family cannot substantiate claims about all autonomous aviation, let alone every flying object. Visually convincing VR can obscure deficiencies in physical fidelity.

Human research requires an appropriate ethics determination before recruitment, clear consent, withdrawal procedures, compensation rules, and a retention plan. Voice and flight traces can identify participants and reveal sensitive information. Access controls, redaction, and purpose-specific consent should be established before collection. Public release may require derived or de-identified artifacts rather than raw recordings. Research data should not be repurposed for employment or licensing judgments without a separately justified process.

Reproducibility artifacts should include simulator and policy versions, parameter envelopes, seeds, units, synchronization diagnostics, split manifests, acquisition settings, constraint code, and adjudication guidance. Released scenarios need license and lineage records. Where proprietary aircraft models prevent full release, the paper should identify exactly which claims an external researcher cannot reproduce. No cited organization is represented as a collaborator or endorser.

Conclusion

Hailing Ocean can formulate a focused scientific contribution around converting pilot evidence into reproducible autonomous-aviation tests. The decisive empirical question is whether that evidence improves failure discovery or held-out performance beyond strong automated baselines. This proposal provides a bounded method and evaluation plan for answering that question while preserving distinctions between human judgment, simulated failure, and real-world safety. Establishing those distinctions makes a subsequent experimental paper more credible and the resulting data more useful.

Declarations

Data and code availability. This draft introduces a proposed protocol; it does not announce a released dataset or implementation.
Authorship, funding, and interests. Individual contributions, funding sources, institutional affiliations, and commercial interests must be completed and reviewed before submission.
Drafting assistance. This manuscript was prepared with AI-assisted literature synthesis and drafting. Named authors must verify technical claims, citations, and the final experimental protocol.
Literature scope. Targeted review through October 6, 2026; not a systematic review. Preprints and product documentation are identified below.

References

  1. A. Corso, R. J. Moss, M. Koren, R. Lee, and M. J. Kochenderfer. A Survey of Algorithms for Black-Box Safety Validation of Cyber-Physical Systems. arXiv:2005.02979, 2020. https://arxiv.org/abs/2005.02979.

  2. C. Dawson and C. Fan. A Bayesian approach to breaking things: efficiently predicting and repairing failure modes via sampling. Conference on Robot Learning, PMLR 229:1706–1722, 2023. https://proceedings.mlr.press/v229/dawson23a.html.

  3. C. Dawson, A. Parashar, and C. Fan. RADIUM: Predicting and Repairing End-to-End Robot Failures using Gradient-Accelerated Sampling. arXiv preprint 2404.03412, 2024. https://arxiv.org/abs/2404.03412.

  4. A. Parashar, K. Garg, J. Zhang, and C. Fan. Failure Prediction from Limited Hardware Demonstrations. arXiv preprint 2410.09249, 2024. https://arxiv.org/abs/2410.09249.

  5. M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal. Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation. Robotics: Science and Systems, 2024. https://arxiv.org/abs/2403.03949.

  6. Palatial. Bring reality into every stage of robot development. Product description, accessed October 6, 2026; not a peer-reviewed publication. https://palatial.cloud/.

  7. W. Yu, J. Lv, Z. Ying, Y. Jin, C. Wen, and C. Lu. ARMADA: Autonomous Online Failure Detection and Human Shared Control Empower Scalable Real-world Deployment and Adaptation. arXiv preprint 2510.02298, 2025. https://arxiv.org/abs/2510.02298.

  8. R. Römer, A. Kobras, L. Worbis, and A. P. Schoellig. Failure Prediction at Runtime for Generative Robot Policies. arXiv preprint 2510.09459, 2025. https://arxiv.org/abs/2510.09459.

  9. S. Khatiri, F. Mohammadi Amin, S. Panichella, and P. Tonella. When Uncertainty Leads to Unsafety: Empirical Insights into the Role of Uncertainty in Unmanned Aerial Vehicle Safety. arXiv:2501.08908, version 2, June 2025. https://arxiv.org/abs/2501.08908v2.

  10. Z. Yuan, A. W. Hall, S. Zhou, L. Brunke, M. Greeff, J. Panerati, and A. P. Schoellig. safe-control-gym: a Unified Benchmark Suite for Safe Learning-based Control and Reinforcement Learning in Robotics. arXiv:2109.06325, version 4, 2022. https://arxiv.org/abs/2109.06325v4.

  11. Y. Yuan, Z. Wu, Y. Li, L. Ma, and K. Li. AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit. arXiv preprint 2608.16349, version 2, September 27, 2026. https://arxiv.org/abs/2608.16349v2.

  12. Y. Wu, J. Fang, J. Wang, H. Liu, Q. Yang, M. Yang, H. Guo, Z. Li, and B. Wang. FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models. arXiv:2609.04021, September 2026. https://arxiv.org/abs/2609.04021.

  13. Archer Aviation. Archer Announces ZEE, AI Foundation Model Purpose-Built for Aviation, a Key Pillar of Its Physical AI Strategy. Corporate announcement, July 15, 2026. https://investors.archer.com/news/news-details/2026/Archer-Announces-Zee-AI-Foundation-Model-Purpose-Built-for-Aviation-a-Key-Pillar-of-Its-Physical-AI-Strategy/default.aspx.

  14. Archer Aviation. Archer’s ZEE AI Foundation Model Achieves Frontier Breakthrough for Aviation Safety: Demonstrating Real-Time Prediction of Airport Surface Trajectories. Corporate announcement, August 5, 2026. https://www.investors.archer.com/news/news-details/2026/Archers-ZEE-AI-Foundation-Model-Achieves-Frontier-Breakthrough-for-Aviation-Safety-Demonstrating-Real-Time-Prediction-of-Airport-Surface-Trajectories/default.aspx.

  15. Archer AI. Zee: Predicting the Airspace. Company technical blog, August 5, 2026. https://archer.com/tech-blog/zee-predicting-the-airspace.