TEVV and the Responsible and Lawful Use of Military AI Systems
Author: Manoj Harjani
Introduction
Reportage and analysis of ongoing conflicts in Ukraine, Gaza and, more recently, Iran are highlighting the growing use of artificial intelligence (AI) in the military domain.[1] In Ukraine, the spotlight has been on how AI has enabled targeting for drones, which have dominated the battlefield there. A case in point is the TFL-1, a module developed by the Ukrainian company The Fourth Law, which leverages machine vision to enable drones to autonomously strike a target that has been identified by an operator.[2] Moving to Gaza, attention has focused on Israel’s use of AI-enabled decision-support systems to identify targets. The Lavender system developed by the Israel Defense Forces (IDF) reportedly generated a list of 37,000 human targets suspected to be members of Hamas and Palestinian Islamic Jihad, using a wide variety of sources.[3] The IDF allegedly utilised this AI-generated list of targets during the early stages of its military response to the October 2023 attack by Hamas, conducting air strikes supported by other AI-enabled systems.[4] Turning to Iran, where a missile strike by the United States in February 2026 on a girls’ school is estimated to have killed up to 168 people,[5] there has been considerable debate around the role of AI in the target selection process.[6] This debate is complicated by confusion over the facts—discussion zeroed in on the use of Claude, a generative AI-enabled chatbot developed by the American company Anthropic, when the targeting for the US campaign against Iran was primarily supported by the Maven system, which leverages machine learning to analyse satellite imagery and sensor data to identify targets.[7]
Meanwhile, multilateral dialogue aimed at creating regulatory guardrails for the use of AI in the military domain has achieved relatively little so far and is moving at a much slower pace in comparison. For instance, at the third Responsible AI in the Military Domain (REAIM) Summit, held in A Coruña, Spain in February 2026, only 44 out of 85 states that attended endorsed the outcome document, ‘Pathways to Action’.[8] Compared to the outcome documents from the first and second REAIM summits, held in 2023 and 2024, which garnered the endorsement of around 50 and 60 states respectively, the drop in support for the latest outcome document also sent a worrying signal regarding a potential loss of interest among states in working to develop regulatory guardrails for the use of AI in the military domain.[9] Moreover, this is the first REAIM outcome document to not have the endorsement of at least one of the two superpowers.[10] China and the United States both endorsed the ‘Call to Action’[11] outcome document from the inaugural REAIM summit, held in 2023, while only the United States supported the ‘Blueprint for Action’[12] outcome document from the second REAIM summit, in 2024. In a similar vein, the US-led Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy (US Political Declaration),[13] launched in 2023—which garnered the support of nearly 60 states[14]—has been quietly set aside by the Trump administration since it took office in 2025.
Taken together, it is clear that facts on the ground are evolving much more quickly and at a scale that multilateral dialogue cannot keep up with. In such circumstances, an important question that arises is around what can be done to ensure the responsible and lawful use of military AI systems. A cynical answer to this question would point to the fact that states will do what they feel is necessary since they are effectively unconstrained in leveraging AI to apply or guide the use of force in the pursuit of political and strategic objectives. Examples from Ukraine, Gaza and Iran appear to validate this point of view, suggesting that states are not particularly interested in the responsible and lawful use of military AI systems. However, this cynical reading also obscures important facts. First, military AI systems do not operate in a vacuum—they are ‘complex sociotechnical systems … [consisting] of material and social components which, by being put into particular kinds of relations, work together in specific ways’.[15] Second, militaries have to reckon with the limitations of AI systems in terms of their predictability and understandability—what is colloquially referred to as the ‘black box dilemma’.[16] Finally, there is a tension between trust and control when it comes to military AI systems, and states face a number of challenges in aiming to reconcile these desired outcomes.[17]
At the same time, even if multilateral dialogue on regulatory guardrails for military AI systems can reach a consensus among states, the way in which its output is framed leaves considerable room for interpretation and significant potential for gaps in accountability. For example, the REAIM Blueprint for Action document makes reference to military AI systems being ‘developed, deployed and used in accordance with international law’[18] but does not delve into how this might be achieved. Additionally, the Blueprint for Action document points to the necessity of military AI systems being ‘ethical and human-centric’[19] but similarly does not outline how this can be concretely implemented in practice. While the REAIM Pathways to Action document sought to provide more detail—for example, by highlighting the importance of a ‘responsible by design’ approach across the lifecycle of military AI systems[20]—it is also framed in largely abstract and aspirational terms. Nevertheless, the Pathways to Action document outlines several recommendations to operationalise the various principles identified across REAIM outcome documents at both the national and international levels. Among these recommendations is the promotion of ‘robust testing, evaluation, validation and verification (TEVV) and [the integration of] TEVV requirements in relevant policies, doctrines, and procurement processes’.[21] TEVV was also highlighted in the US Political Declaration in the context of ensuring ‘that the safety, security, and effectiveness of military AI capabilities are subject to appropriate and rigorous testing and assurance within their well-defined uses and across their entire life-cycles’.[22]
This article seeks to explore the role of TEVV in supporting emerging regulatory guardrails that are aiming to ensure the responsible and lawful use of military AI systems. The first and arguably most critical contention behind this is that TEVV is an essential enabler for governance and regulation of military AI systems, as it could potentially provide a way to translate abstract and aspirational concepts into tangible practices that states can undertake. A second and equally important contention is that TEVV is not a silver bullet when it comes to regulatory guardrails, and comes with its own set of challenges that states must find ways to navigate. Three sets of challenges associated with implementing TEVV to support regulatory guardrails are identified and discussed: (1) the lack of standardised TEVV practices for military AI systems, (2) variations in implementing TEVV, and (3) the private sector’s involvement. A mixture of qualitative methods—documentary analysis and close reading of primary and secondary sources—has been employed alongside a small-N case study design for comparing approaches to TEVV for military AI systems in Australia, the United Kingdom and the United States. The structure of the article is as follows: an overview of TEVV, a discussion of how TEVV can support regulatory guardrails for the responsible and lawful use of military AI systems, a comparison of state-level approaches across Australia, the United Kingdom and the United States, and a discussion of key challenges associated with leveraging TEVV to support regulatory guardrails for military AI systems.
What is TEVV?
‘Testing, evaluation, verification and validation’ (TEVV) refers to a set of practices without a universal definition or set of standards in the military domain.[23] The concept itself and its four elements are not used consistently—testing and evaluation is often treated separately from verification and validation, and the order of the latter two elements is sometimes swapped. Additionally, a fifth element is sometimes associated with verification and validation—namely, accreditation. The various elements within TEVV—regardless of whether they are treated separately as T&E and V&V (or VV&A)—share a common purpose, which is risk reduction.[24] They are also all grounded in the broader set of practices associated with assurance, which for AI could cover a wide range of possible qualities associated with a system’s outcomes, including that they are ‘valid, verified, data-driven, trustworthy and explainable to a layman, ethical in the context of its deployment, unbiased in its learning, and fair to its users’.[25] The ways in which desired outcomes for assurance are identified can vary considerably—this is both essential and problematic when it comes to how TEVV can support regulatory guardrails for the responsible and lawful use of military AI systems. For the purpose of this article, the table below outlines one possible way to define the four elements within TEVV (and assurance), which was chosen for its simplicity and succinctness (see Table 1).
| Testing | Using pass/fail criteria to check a specific, single function. Testing is usually associated with verification |
| Evaluation | Combining multiple tests and experimental use to provide an overall judgement. Evaluation typically relates to validation |
| Verification | Checking against a specification. It asks, ‘Did we build the thing right?’ or ‘Are we doing the thing right?’ |
| Validation | Checking against an end user’s need. It asks, ‘Did we build the right thing?’ or ‘Are we doing the right thing?’ |
| Assurance | Providing a compelling case for the trustworthiness of a system |
Note: Adapted from Defence Science and Technology Laboratory, Assurance of Artificial Intelligence and Autonomous Systems: A Dstl Biscuit Book (London: Ministry of Defence, 2021), pp. 4–5.
Nevertheless, there are other ways of contextualising TEVV. For example, in the civilian domain, TEVV can be framed as one of several paradigms within the practice of AI evaluation, which can be described as a ‘process of measuring and anticipating the behavioural properties of AI systems and their societal impact to inform decisions about their use’.[26] Other paradigms within AI evaluation include benchmarking[27] and measuring real-world impact.[28] Within AI evaluation, each paradigm has a distinct focus, and for TEVV this is ‘ensuring that AI systems behave in a well-defined and predictable way’.[29] AI evaluation is closely linked to the field of AI safety, which has largely been concerned with the existential risks associated with advanced AI systems, although this has created its own challenges in terms of directing attention away from existing work—including through TEVV—on ensuring the robustness and reliability of AI systems.[30] Multilateral dialogue on AI safety has also evolved considerably since the inaugural AI Safety Summit organised by the United Kingdom in 2023. Twenty-nine states—including China and the United States—signed the Bletchley Declaration that was produced through that meeting.[31] In 2024 at the AI Seoul Summit, the framing of the outcome document shifted to encompass not just safety but also innovativeness and inclusion.[32] Similarly, the 2025 AI Action Summit held in Paris changed the focus of the outcome document again to include mention of inclusivity and sustainability.[33] At the 2026 AI Impact Summit held in New Delhi, the scope of the outcome document broadened even further, covering themes ranging from democratising AI to resilience, innovation and efficiency.[34] Little to no attention has been paid to TEVV in the outcome documents produced since the first AI Safety Summit in 2023.
TEVV for Military AI Systems
TEVV for military AI systems[35] does not take place in a vacuum and is shaped by existing practices for other types of systems.[36] This is likely to include the general set of TEVV practices associated with all systems developed or purchased by a military, as well as specific practices depending on the nature of the system.[37] For example, if the system is a helicopter, there will be TEVV practices from the army aviation sector to consider. Similarly, each service branch of a military may have TEVV practices tailored to its needs and influenced by its organisational, engineering and safety cultures. For militaries in states that may not have the capacity or resources to maintain an in-house capability, their TEVV practices will depend heavily on the supplier they are purchasing a system from. Moreover, when it comes to AI systems, there is a need to consider the TEVV practices of the private sector companies developing them, since they are at the forefront of innovation in this field. For the most advanced AI systems, militaries will be dealing with the handful of companies that have the necessary capabilities. All of these considerations point to TEVV for military AI systems being a complex and costly endeavour. This in turn raises the question of whether this ‘premium’ associated with fielding military AI systems will create a clear gap in terms of their impact on the balance of power, which challenges the conventional wisdom suggesting that access to AI is relatively democratised and that more states will have the ability to leverage AI in the military domain.[38]
From a technical perspective, TEVV for military AI systems must also account for the distinctive features of AI-enabled systems more generally. This entails paying attention to (1) the need for continuous testing and monitoring due to the dynamic nature of AI-enabled systems, particularly in terms of their ability to learn and adapt, (2) the need to evaluate unexpected and unanticipated system behaviours arising from the ‘black box dilemma’, and (3) the need to assess the security and resilience of AI-enabled systems given their vulnerability to adversarial attacks.[39] Regarding continuous testing and monitoring, this necessitates a different approach to TEVV that moves away from treating it as a discrete stage prior to fielding a system. It has also prompted exploration of potential ways to automate TEVV.[40] To address the ‘black box dilemma’, one possible way forward is through implementing safety middleware that can ‘supersede the true system’s decisions under certain explicit, well-defined conditions’,[41] thereby making it easier to verify specific conditions under which certain behaviours of the system are occurring. Finally, when it comes to adversarial attacks, these introduce a significant degree of complexity, as there are a wide range of possible attacks, from ‘poisoning’ of training data to the ‘extraction’ of a model through reverse engineering of its output.[42]
TEVV and regulatory guardrails
Turning to the role that TEVV plays in supporting regulatory guardrails for the responsible and lawful use of AI in the military domain, there are three aspects to unpack further: (1) providing the means to tangibly measure desired outcomes of an AI system, (2) enhancing human agency across the lifecycle of AI systems, and (3) enabling the conduct of legal reviews. These three aspects of TEVV’s role were selected to highlight that it is more than a set of practices to narrowly certify a system as ready for use or fit for purpose.
Measuring Desired Outcomes
At the heart of responsible and lawful use of AI in the military domain are the desired outcomes of trust and control.[43] These are typically framed as expectations of military AI systems and of the relationship between these systems and their operators, which can occur in the context of human–machine teams.[44] Put simply, for a military to use an AI system responsibly and lawfully, there must be trust in that system and its operator, but also between the operator and the system. Trust in a military AI system also depends in part on an operator’s ability to control that system to function in specific ways. The ability of a military to meet legal obligations arising from using an AI system also depends on control over that system, particularly through rules that circumscribe the behaviour of the system and its operator. Trust therefore implies ‘a “contract” between humans and AI’,[45] whereas control focuses on ‘the capacity of humans to ensure, in a rather deterministic way, specific (or, at least, sufficiently tightly constrained) behaviours of AI-driven systems’.[46] Reconciling trust and control becomes challenging depending on the circumstances related to a specific application of a military AI system. This is where TEVV can play an important role, particularly in measuring the desired outcomes of trust and control. For example, TEVV can be used to measure trust based on reliability or predictability[47] and to measure control by providing evidence of an AI system’s patterns of behaviour through testing. It can also measure trust based on values, preferences and beliefs[48] where there has been an appropriate translation into expected behaviours and output of an AI system that can be evaluated.
However, TEVV is not a silver bullet. Measuring trust and control in military AI systems faces a number of interrelated challenges. Six of these are identified here, with the caveat that this is not an exhaustive list.[49] First, the appropriate unit of analysis for TEVV in military AI systems can be difficult to define, particularly when the same AI model is used for multiple applications, when it is part of a larger system, and when it is part of a human–machine team.[50] Second, TEVV depends on:
clearly stating [a] purpose and being able to assess whether the tool serves it as well as potential alternatives, and whether ‘ground truth’ benchmarks for evaluating it accurately reflect contexts in which [it] will be used.[51]
But for military AI systems, ‘ground truth’ benchmarks are difficult to define given their unpredictable and inscrutable nature. Moreover, identifying the different contexts in which military AI systems will be used is not a straightforward task—modelling and simulation methods and techniques that are traditionally employed in TEVV may not be sufficient. Third, and related to the previous challenge, is the fact that TEVV will not be able to anticipate all the conditions and environments under which military AI systems will operate, creating a potentially impractical number of scenarios to test and evaluate against.[52] Fourth, military AI systems share the same challenges associated with all AI systems in terms of difficulty in predicting and understanding how they might fail, which complicates the scope of TEVV.[53] Fifth, although military AI systems are likely to comprise software components that were developed by the private sector, the applicability of commercial TEVV tools is not straightforward for a military context.[54] Finally, military AI systems which have models that are updated through data gathered during their use necessitate some parts of the TEVV process to be carried out continuously.[55]
Overcoming these challenges may require militaries to rethink the practice of TEVV from a socio-technical perspective in order to:
enable a broader understanding of AI impacts and the key decisions that happen throughout, and beyond, the AI lifecycle—such as whether technology is even a solution to a given task or problem.[56]
TEVV that leverages a socio-technical perspective could potentially incorporate practices from a wider range of fields in order to provide novel or more effective ways to measure trust and control in military AI systems. However, several of the challenges identified also underscore the point made earlier that TEVV for military AI systems is a complex and costly endeavour. Along with potentially allocating dedicated resources catering to the unique demands of TEVV for military AI systems, militaries are likely to have to invest in developing new TEVV practices and finding ways to adapt existing practices. However, this poses an additional challenge of its own, given the pressure created by militaries’ increasing pace of AI adoption. It will not be straightforward to update existing TEVV practices or introduce new ones under these circumstances where there is an imperative to deploy systems more quickly.
Enhancing Human Agency
The importance of human agency for the responsible and lawful use of military AI systems across all service branches of a military cannot be overstated. In an army context, for example, human agency takes on particular significance given close contact with civilian populations. The REAIM Call to Action document refers to how AI can ‘shape and impact decision making, and [the importance of ensuring] that humans remain responsible and accountable for decisions when using AI in the military domain’.[57] It also specifically mentions the ‘importance of ensuring appropriate safeguards and human oversight of the use of AI systems, bearing in mind human limitations due to constraints in time and capacities’.[58] The REAIM Blueprint for Action document similarly refers to how:
Appropriate human involvement needs to be maintained in the development, deployment and use of AI in the military domain, including appropriate measures that relate to human judgement and control over the use of force.[59]
Meanwhile, the REAIM Pathways to Action document makes specific reference to how the:
nature and degree of human involvement should be appropriate considering, among other factors, the operational context, the function performed, the technical characteristics and capabilities, as well as human factors such as training and fatigue, and the risks and benefits involved.[60]
In the US Political Declaration, there is mention of ‘senior officials effectively and appropriately oversee[ing] the development and deployment of military AI capabilities with high-consequence applications, including, but not limited to, such weapon systems’[61] as well as ensuring that ‘relevant personnel exercise appropriate care in the development, deployment, and use of military AI capabilities, including weapon systems incorporating such capabilities’.[62] What is missing in all of these references to human agency is how it can be implemented in practice.
A gradual evolution in how human agency is being understood—particularly in terms of going beyond a ‘narrow focus on human operators executing military missions at the final tactical level’[63]—has paved the way for its consideration across the entire lifecycle of a military AI system. In the nine-stage lifecycle framework developed by the IEEE Standards Association, TEVV is both a distinct stage and an ongoing activity where human decision-makers play a key role (see Figure 1). As a stage of the lifecycle, TEVV is situated between ‘procurement and acquisition’, and ‘considering the human: education, training, and human-system integration.’ This position reflects TEVV’s role as a bridge between dimensions of an AI system that are technical in nature and those that are dependent on a human user’s needs and interaction with the system. As an ongoing activity taking place throughout each stage of the lifecycle, TEVV is considered together with monitoring, hardware or software updates, interoperability and maintenance, which leans more towards the technical and operational dimensions of AI systems. Specific best practices related to TEVV that can enhance human agency have been identified in a study conducted by the University of Southern Denmark. They include testing the interface of military AI systems not only from a technical and functional perspective but also in terms of human users’ understanding and interaction, verifying that intended users of military AI systems are empowered to use them in ways that are legally and ethically compliant, and engaging in TEVV to independently verify commercial vendors’ claims about their military AI systems.[64]
Note: Adapted from IEEE Standards Association Industry Connections Research Group on Issues of Autonomy and AI for Defence Systems, A Framework for Human Decision-Making Through the Lifecycle of Autonomous and Intelligent Systems in Defence Applications (New York NY: IEEE Standards Association, 2024), p. 21, at: https://ieeexplore.ieee.org/document/10707139.
Enabling Legal Reviews
Both the REAIM Blueprint for Action and Pathways to Action documents make specific references to legal reviews as a regulatory guardrail for military AI systems. In the Blueprint for Action, legal reviews are highlighted alongside trust and confidence-building measures as well as risk-reduction measures,[65] whereas the Pathways to Action document identifies the conduct of legal reviews as a measure to operationalise REAIM principles at the national level.[66] In the US Political Declaration, legal reviews are specifically mentioned as a step that states can take to ‘ensure that their military AI capabilities will be used consistent with their respective obligations under international law, in particular international humanitarian law’.[67] While the exact way in which legal reviews are carried out varies considerably, they are generally understood as an obligation under article 36 of Additional Protocol I to the Geneva Conventions.[68] The starting point for legal reviews for military AI systems is the consensus that has emerged around the applicability of international humanitarian law (IHL), but the challenge has been in determining exactly how it applies. This challenge is exacerbated by a few factors, such as (1) the wide range of technologies embedded within military AI systems, which may be part of other, larger systems involving consideration of other technologies, (2) whether AI systems are fundamentally incompatible with IHL principles due to their unpredictable and inscrutable nature, and (3) the dynamics created through human–machine teams and human–machine interaction.[69]
TEVV has the potential to address some of these factors when integrated within the conduct of legal reviews. This would bring technical and legal assessments of military AI systems together, with significant implications for current expertise, methods and processes related to both TEVV and legal reviews. For instance, IHL principles could be translated into technical criteria that are then ‘embedded into the system during the designing phase. Once the legal standards are learned by the system, compliance with those standards is a technical task’.[70] This approach is aligned with what has been proposed by the Global Commission on Responsible AI in the Military Domain (GC REAIM), which has been encapsulated in the concept of ‘responsible by design’.[71] Although one of the GC REAIM’s core recommendations is that TEVV should inform national policies guaranteeing human responsibility across a military AI system’s lifecycle,[72] the link between TEVV and legal reviews is not made explicit, and they are treated as separate activities and sets of practices. However, the fact that experts and practitioners in both TEVV and legal reviews are increasingly pointing to the need for the two activities and sets of practices to evolve to better address the challenges posed by military AI systems is a good sign. For example, recognising the limitations of conventional legal reviews—particularly when they are conducted as a ‘discrete, one-time requirement performed during the weapon acquisition phase, typically just prior to adoption’[73]—will pave the way for closer alignment between TEVV and legal review processes, which need to be conducted throughout the lifecycle of military AI systems. At the same time, there are some complications that can arise from integrating TEVV and legal reviews. The view of the International Committee of the Red Cross, which is entrusted by the international community of states as the guardian of IHL,[74] is that:
it is essential to preserve human control over tasks and human judgement in decisions that may have serious consequences for people’s lives in armed conflict, especially where they pose risks to life, and where the tasks or decisions are governed by specific rules of international humanitarian law.[75]
Integrating TEVV and legal reviews could therefore run the risk of transferring responsibility for legal determinations away from humans.
Comparing State-Level Approaches
Australia, the United Kingdom and the United States represent a spectrum of different types of actors in the ongoing multilateral dialogue aimed at creating regulatory guardrails for the use of AI in the military domain. The United States sits at one end of the spectrum in terms of the size and capabilities of its military. As treaty allies of the United States, Australia and the United Kingdom not only draw on the capabilities of the United States but also contribute to joint capabilities through arrangements such as AUKUS. Through AUKUS, the three states are also pursuing extensive integration of their defence industrial and technological bases.[76] The United Kingdom and United States are also both permanent members of the United Nations Security Council and founding members of the North Atlantic Treaty Organisation. Australia and the United States are part of the Quad, and all three states are also part of an intelligence-sharing arrangement known as the Five Eyes. The relative postures and positioning of the three states when it comes to AI in the military domain is not a straightforward comparison, but the table below outlines some indicators with respect to regulatory guardrails (see Table 2). These indicators should be viewed with a measured eye—publicly articulated statements in a defence AI policy or strategy and the endorsement of non-binding statements such as the REAIM outcome documents and the US Political Declaration cannot be construed as evidence of a state’s actual position in practice.
| Defence AI Policy / Strategy |
REAIM Call to Action (2023) |
REAIM Blueprint for Action (2024) | REAIM Pathways to Action (2026) | US Political Declaration (2023) | |
|---|---|---|---|---|---|
| Australia | 2026 | ✓ | ✓ | ✓ | ✓ |
| United Kingdom | 2022 | ✓ | ✓ | ✓ | ✓ |
| United States | 2026 | ✓ | ✓ | x | ✓ |
Compiled by the author.
Australia
In March 2026, Australia’s Department of Defence (Defence) released a document titled Policy Settings for Responsible Use of Artificial Intelligence in Defence (Policy Settings),[77] which outlined its approach to responsible AI in the military domain across the lifecycle of military AI systems. The Policy Settings document is explicitly framed as a follow-up to implement Australia’s commitments made through its endorsement of the REAIM Call to Action and Blueprint for Action documents, as well as the US Political Declaration.[78] It anchors Australia’s obligations for responsible use of AI in the military domain on three policy requirements: lawfulness, adherence to values-based principles, and proportionate controls.[79] Lawfulness is based on compliance with domestic law and international legal obligations, clear accountability through identifiable officers, and the conduct of legal reviews.[80] Under adherence to values-based principles, five principles have been identified: accountable, bias and harm mitigation, explainable, human-centric, and reliable and secure.[81] These principles have been developed in line with Australia’s 2024 Policy for the Responsible Use of AI in Government[82] and 2019 AI Ethics Principles.[83] Finally, proportionate controls comprise risk-based ‘layers of policies, processes, training, and procedures, including ongoing assurance and after-action evaluation’.[84] TEVV is included as a control measure, alongside post-incident/impact reviews, and relevant reporting and oversight mechanisms.[85] However, the role of TEVV is otherwise not elaborated in the Policy Settings document.
Defence’s 2021 Test and Evaluation Strategy[86] only mentions AI once in the context of emerging technologies necessitating new approaches to TEVV.[87] This is due to the fact that it is meant to cover all systems regardless of the technologies they comprise. In its 2024 Defence Industry Development Strategy, Defence identified TEVV as a ‘sovereign defence industrial priority’, with specific mention of building skill level and capacity in TEVV across warfighting domains with respect to AI.[88] This was an important signal, both in terms of Defence’s overall assessment of current TEVV capabilities and in terms of its expectations of how Australia’s defence TEVV workforce, processes and practices should evolve. It also built on the 2020 Concept for Robotic and Autonomous Systems,[89] where TEVV is explicitly linked to determining trust in military AI systems.[90] In this document, TEVV also supports the traceability of decisions made by military AI systems, which in turn contributes to trust.[91] Additionally, reform of existing TEVV processes and practices ‘to understand how to verify systems that are the subject of continual improvement while balancing this need against the necessity of providing capability certainty’[92] is highlighted. Moreover, with the Defence AI Centre that was set up in 2024, there is now dedicated bureaucratic capacity for coordination that will also shape the evolution of TEVV, including through initiatives such as developing common standards for military AI systems.[93]
Turning to service-level approaches, in the second version of its Robotic and Autonomous Systems Strategy,[94] the Australian Army mentions TEVV in two lines of effort: first, as part of developing, experimenting and prototyping military AI systems;[95] and second, as part of exploring and integrating military AI systems within force design, specifically in terms of providing assurance for ‘fit-for-purpose capabilities into service at the right time and quantities’.[96] For the first line of effort, TEVV is intended to ‘test and demonstrate key capabilities in representative environments and is critical to ensuring system reliability, robustness and usability in contested environments and to inform Army’s user needs’,[97] whereas the second line of effort sees TEVV supporting future force design.[98] In the Royal Australian Navy’s 2020 RAS-AI Strategy 2040,[99] the need for qualified TEVV practitioners is highlighted,[100] as well as the role that TEVV plays in providing ‘proof that risk is contained within acceptable boundaries’.[101] The Navy’s strategy also outlines how AI can be an opportunity for TEVV in terms of automating testing, documentation, data processing, decision-making and system control, among other areas.[102] While the Royal Australian Air Force does not have a comparable strategy document related to military AI systems, its efforts under the 2015 Plan Jericho[103] led to a dedicated Test and Evaluation Directorate being set up within the Air Warfare Centre inaugurated in 2016.[104]
Australia’s approach to TEVV for military AI systems is generally headed in a positive direction. TEVV has been clearly prioritised at the strategic level centrally in the Australian Defence Force (ADF) through documents such as the 2024 Defence Industry Development Strategy and 2020 Concept for Robotic and Autonomous Systems. Service-level strategies and initiatives for military AI systems also point to TEVV being recognised as a priority, including in the Army’s 2022 Robotic and Autonomous Systems Strategy, the Navy’s RAS-AI Strategy 2040, and the Air Force’s 2015 Plan Jericho. However, the actual impact of this strategic prioritisation on the ADF’s TEVV capabilities remains to be seen. Each service branch has high-profile initiatives involving military AI systems—the Army’s M113 AS4 optionally crewed combat vehicle,[105] the Navy’s Bluebottle uncrewed surface vessel,[106] and the Air Force’s MQ-28A Ghost Bat uncrewed combat aerial vehicle supporting crewed aircraft[107]—which means that their TEVV capabilities, processes and practices will be evolving through experience on the fly. Moreover, collaboration through AUKUS is also likely to shape Australia’s defence TEVV enterprise. An example is the trialling of AI-enabled decision-support systems onboard the Air Force’s P-8A Poseidon aircraft during Exercise Talisman Sabre held from July to August 2025.[108] Although AUKUS has faced some uncertainties,[109] it remains an important platform for TEVV-related activities for military AI given its emphasis on experimentation and technology development, particularly under Pillar II.
United Kingdom
In 2022, the United Kingdom’s Ministry of Defence (MoD) outlined its approach for the delivery of AI-enabled defence capabilities based on three guiding principles: (1) ambition in terms of tools and operational effects, (2) safety, and (3) responsibility.[110] In the same document it also identified five ethical principles for military AI systems: (1) human-centricity, (2) responsibility, (3) understanding, (4) bias and harm mitigation, and (5) reliability.[111] Building on this approach, the 2022 Defence Artificial Intelligence Strategy[112] articulated four objectives: (1) transforming into an ‘AI ready’ organisation, (2) adopting and exploiting AI at pace and scale, (3) strengthening the domestic defence and security AI ecosystem, and (4) shaping global developments to promote security, stability and democratic values.[113] Furthermore, in 2024, the MoD issued the first part of a joint service publication, JSP 936 V1.1: Dependable Artificial Intelligence (AI) in Defence,[114] which contained a directive describing the mandated direction related to the responsible and lawful use of military AI systems. The second part of JSP 936, which contains a defence AI ethical risk assessment toolkit, has not been made public in its entirety. JSP 936 builds on a family of related MoD policies, including on the management of health and safety (JSP 375[115]), acquisition safety (JSP 376[116]), information, knowledge, digital and data (JSP 441[117]), human factors integration (JSP 912[118]), and modelling and simulation (JSP 939[119]).
In January 2026, the MoD issued updated guidance on how its approach to TEVV is changing in a document titled Test and Evaluation (T&E): Future Advantage Through Evaluation (FATE).[120] The challenge posed by AI for TEVV is highlighted in this document, along with AI’s potential to transform TEVV. Crucially, the document also highlights the need for a different approach to TEVV and a slew of initiatives to support the United Kingdom’s TEVV enterprise, including a dedicated innovation fund, increased support for training, and better coordination with industry. The 2022 Defence Artificial Intelligence Strategy outlines specific measures related to TEVV for military AI systems, including new live and virtual test capabilities, establishing a framework for TEVV of military AI systems across their lifecycle that includes both the technical and human dimensions, and driving new technical standards and regulations.[121] JSP 936 also provides detailed guidance on TEVV in the context of AI assurance,[122] which has been complemented by an AI Assurance Framework developed by the MoD’s Defence AI Centre that covers who is responsible, the relevant steps involved and the required documentation for evidence, and identifies enabling technologies.[123] The Defence AI Centre has also published an AI Practitioner’s Handbook[124] with a corresponding resource library containing templates and tools to support TEVV.[125] At the service level, the British Army and Royal Navy have so far published strategies for AI, but only the Army’s 2023 document is available publicly[126] and there is limited mention of TEVV’s role within it.
The United Kingdom’s approach to TEVV for military AI systems appears well organised, with clear prioritisation by the MoD. A review by the House of Commons Defence Committee published in January 2025 did not make any recommendations directly mentioning TEVV, but devoted a section to addressing challenges posed by AI for defence procurement, such as shorter development cycles, evolving contractor–supplier relationships, and the role of primes versus smaller companies.[127] Any changes to address these challenges will also affect the TEVV enterprise. Nevertheless, with JSP 936 and the Defence AI Centre, the MoD is well positioned to drive the evolution of the United Kingdom’s defence TEVV to meet the challenges and needs of military AI systems and to ensure responsible and lawful use. Moreover, since its establishment in 2021, the Defence AI Centre has played a key role in driving TEVV for military AI systems. For example, in November 2025, it launched the AI Model Arena tool to help the MoD redefine its current approach to evaluating and procuring AI.[128] Finally, the Strategic Defence Review published in July 2025 has provided additional momentum through prioritising greater use of autonomy and AI.[129] While there are no recommendations within the document that are directly focused on TEVV for military AI systems, the proposed creation of a new defence research and evaluation organisation will also entail sustaining relevant TEVV capabilities within that organisation.[130]
United States
In January 2026, the United States Department of War (DoW) published a memorandum outlining its new military AI strategy.[131] The first heading in the strategy, ‘Accelerating America’s Military AI Dominance’, is aligned with the intent of Executive Order 14179 issued by US President Donald Trump in January 2025 on ‘Removing Barriers to American Leadership in Artificial Intelligence’[132] and with the AI Action Plan issued by the White House in July 2025.[133] The DoW’s military AI strategy outlines four lines of effort: (1) experimenting with frontier AI models, (2) eliminating bureaucratic barriers to AI integration, (3) focusing investment in AI to leverage existing asymmetric advantages, and (4) carrying out ‘pace-setting projects’ (PSPs) to establish a new execution standard and build out the necessary foundation of enablers to accelerate AI integration.[134] Seven initial PSPs have been identified across warfighting, intelligence and enterprise mission areas: (1) ‘Swarm Forge’, focusing on discovering, testing and scaling with and against AI-enabled capabilities, (2) ‘Agent Network’, to support development and experimentation with AI agents, (3) ‘Ender’s Foundry’, focusing on AI-enabled simulation capabilities, (4) ‘Open Arsenal’, to accelerate the translation of technical intelligence into capability development, (5) ‘Project Grant’, to transform deterrence, (6) ‘GenAI.mil’, to broaden AI experimentation with the latest models, and (7) ‘Enterprise Agents’, focusing on AI agent development and deployment for enterprise workflows.[135] In addition to the memorandum on the DoW’s military AI strategy, two other memoranda were issued: one on restructuring the department’s Advana data platform[136] and the other on transforming the defence innovation ecosystem.[137] The latter memorandum brings all of the various defence innovation organisations under the leadership and responsibility of the Chief Technology Officer, namely the Under Secretary of War for Research and Engineering. Previously these organisations—the Defence Innovation Unit, Strategic Capabilities Office, Defence Advanced Research Projects Agency, Chief Digital and Artificial Intelligence Office (CDAO), Test Resource Management Centre (TRMC), and Office of Strategic Capital—were positioned at different levels of leadership and were part of various steering and working groups.
The DoW’s Acquisition Transformation Strategy, published in November 2025, made three important changes to the department’s approach to TEVV. First, it calls for modernising test infrastructure under the TRMC and creating a minimum set of department-wide test data and quality standards.[138] Second, it aims to reduce central TEVV oversight conducted by the Director, Operational Test & Evaluation (DOT&E), returning responsibility to the service branches.[139] Third, it calls for modernising systems engineering processes and tools informing TEVV ‘to enable agile development, technology insertion, improved technology and manufacturing risk management, and reduced need for testing, rework, and re-testing to certify a system’.[140] In a similar vein, the DoW’s military AI strategy specifically mentions adopting a ‘[w]artime approach to blockers,[141] including in relation to TEVV. It also clarifies the DoW’s approach to responsible military AI, which no longer emphasises ethical principles but rather focuses on what is lawful.[142] As for the set of TEVV frameworks developed by the CDAO for military AI systems in 2024 under the previous administration, their status is unclear. These frameworks cover TEVV for AI models,[143] human–systems integration,[144] systems integration[145] and operational TEVV.[146] They are focused on providing guidance and best practices for formulating a TEVV strategy, rather than binding policy requirements or a step-by-step guide.
The United States’ defence TEVV enterprise will continue to evolve as the DoW’s new military AI and acquisition transformation strategies are implemented, but several aspects remain uncertain. Streamlining reporting lines for all the defence innovation organisations under the DoW’s Chief Technology Officer and minimising central oversight by the DOT&E could produce a more dynamic and synergised TEVV enterprise that draws on a larger pool of capabilities and resources. However, there is also a risk of confusion over roles and responsibilities, which would increase rather than reduce friction. Furthermore, while plans to modernise TEVV infrastructure, processes and tools are uncontroversial, it is unclear whether this will leverage existing work such as the Joint AI Test Infrastructure Capability (JATIC) program initiated by the CDAO in 2023 with funding up to 2029[147] and the CDAO’s TEVV frameworks published in 2024. Moreover, the existing DoW manual on TEVV for military AI systems (5000.101[148]), which was last updated in December 2024, has yet to incorporate the policy changes that have taken place under the Trump administration since 2025. Of greater concern, however, is the sidelining of ethical principles in TEVV supporting responsible and lawful use of military AI systems. The United States did not endorse the recent REAIM Pathways to Action document, which adds to this concern. There is also the concern that the overall desire to ensure speed in the DoW’s military AI strategy could mean that TEVV is marginalised.
Challenges
Six interrelated challenges associated with leveraging TEVV to measure trust and control in military AI systems were identified earlier: (1) determining the appropriate unit of analysis for TEVV, (2) the need for an AI system to have a clearly defined purpose for TEVV to evaluate against, (3) the impracticality of identifying all possible scenarios for TEVV, (4) the unpredictability and inscrutability of AI systems, (5) the fact that private sector TEVV tools and benchmarks may not neatly apply in the military context, and (6) the need for some parts of the TEVV process to be carried out continuously for AI systems that have models which are updated through data gathered during their use.[149] In addition to these challenges, the military domain imposes additional considerations that also affect TEVV. First, the scarcity of high-quality and representative data in operational environments will affect training datasets, the creation of representative test environments, and the ability to minimise bias.[150] Second, the deployment of military systems across different operational environments may necessitate retraining and re-evaluating when it comes to AI.[151] Factoring in consideration of the complexities associated with human–machine teams further muddles what is already a very challenging set of circumstances for TEVV. Nevertheless, all of these challenges and considerations are known and therefore subject to ongoing work in terms of adapting TEVV processes and practices. This section builds on this by discussing challenges that have received comparatively less attention in the literature but are still important in terms of their impact on TEVV’s role in supporting regulatory guardrails.
Lack of Standardised TEVV Practices
There are currently no standardised TEVV practices for AI in the military domain. While the nine-stage lifecycle framework developed by the IEEE Standards Association (see Figure 1) provides a foundation to build up a set of standardised TEVV practices, the fact that TEVV is both a stage within the lifecycle framework and an ongoing activity taking place throughout all stages already adds a layer of complexity. Moreover, the lack of standardised TEVV practices partly stems from the nature of AI systems and how AI is used in the military domain—particularly in terms of the ‘black box dilemma’, as well as in determining the unit of analysis and the system’s purpose for TEVV, which are affected by how military systems can be deployed across different operational environments. Furthermore, where military AI systems are operating as part of human–machine teams, the ability to develop TEVV to account for real-world performance is limited not just to the technical or technological dimension but also to the human dimension.[152] All of this points to a need for flexibility and adaptation in TEVV for military AI systems, which is precisely the dilemma in terms of the lack of standardised practices. Resolving this tension will be important to advance TEVV for military AI systems in such a way that it can then support responsible and lawful use of these systems while accounting for the practicalities of dealing with AI.
Across the case studies, it is clear that this is an ongoing challenge. In Australia’s case, for example, if the service branches appear to be driving the overall approach to TEVV for military AI systems rather than Defence from a central standpoint, then we can already expect variations in practices based on the missions of each service branch. Similarly, for the United Kingdom, even though its MoD can leverage a common standard through JSP 936 and the Defence AI Centre’s AI Assurance Framework, implementation at the service branch level is likely to necessitate some variations in TEVV practices. Turning to the United States, the minimisation of central oversight for TEVV by the DOT&E will mean that there will be a similar challenge in terms of managing the variation across service branches. The next section unpacks these points further in terms of factors shaping the implementation of TEVV, such as a military’s technical capacity and resources, as well as its organisational, engineering and safety cultures. Additionally, the section on the private sector’s involvement highlights another factor which further complicates the ability to develop standardised TEVV practices, namely the extent to which TEVV capabilities exist in-house within a military or depend on outsourcing to the private sector.
Variations in Implementing TEVV
Even if standardised TEVV practices can be identified, their implementation depends on a military’s technical capacity and resources. These will vary considerably—among the case studies in this article, for example, there is a wide gulf between the United States and both Australia and the United Kingdom in terms of technical capacity and resources. The size of the US military’s TEVV workforce was estimated at 10,000 in 2025, with 80 per cent being civilian personnel.[153] For Australia and the United Kingdom, a comparable estimate is much more challenging to make, as TEVV capabilities are distributed across civil servants, military personnel, and private sector contractors. In Australia’s case, however, there was an estimated shortfall of 400 TEVV practitioners identified in 2024, which is expected to grow to 1,000 by 2030.[154] Estimating a military’s investments in TEVV is also not a straightforward exercise, and a distinction must be made between what is allocated for TEVV within a specific capability project versus what is allocated at an organisational level. For example, in May 2025, the United Kingdom’s MoD announced a £1.54 billion extension for its long-term partnering agreement with the British company QinetiQ to continue operating TEVV infrastructure and provide TEVV services.[155] In addition, the Defence and Security Accelerator announced a competition in November 2025 offering up to £1 million in funding for innovative TEVV capabilities as part of the MoD’s Test and Evaluation Transformation Programme.[156] While significant, these developments do not represent the full extent of the MoD’s investment in TEVV.
Moreover, how a military leverages its available technical capacity and resources will be shaped by its organisational, engineering, and safety cultures. For example, since the Trump administration took office in 2025, there has been a concerted attempt to change the culture of the US military and defence bureaucracy, beginning with Executive Order 14347, issued in September 2025, which rebranded the Department of Defense as the Department of War.[157] In a speech to general and flag officers in the same month, Secretary of War Pete Hegseth spoke about restoring the ‘warrior ethos’ of the US military.[158] These changes have filtered down to shape the approach to TEVV for military AI systems. The DoW’s military AI strategy specifically states:
Diversity, Equity, and Inclusion and social ideology have no place in the DoW, so [it] must not employ AI models which incorporate ideological ‘tuning’ that interferes with their ability to provide objectively truthful responses to user prompts.[159]
Moreover, the concern highlighted earlier regarding a desire to ensure speed in the DoW’s military AI strategy potentially leading to TEVV being marginalised is prompted by developments such as the rollout of GenAI.mil in December 2025, which appeared to be rushed, raising the question of whether there was an adequate level of assessment conducted.[160]
Involvement of the Private Sector
Innovation in AI in the military domain is primarily being driven by the private sector within a handful of states. On the surface, this:
suggests fostering unrestrained private competence. On the other hand, there are strong incentives for state control, since military AI is a vital sector for national security, and many applications involve high security risks and significant political costs if left uncontrolled.[161]
There is a trade-off that states and militaries have to manage when it comes to ‘unconstrained expertise’[162] that can support innovation but runs the risk of undesired outcomes, including the diffusion of innovative technologies to rival states. In the context of TEVV, because it is an important ongoing activity throughout the lifecycle of military AI systems (see Figure 1), navigating the trade-off between competence and control poses several challenges. First, it can significantly affect the timeline for capability development and acquisition. In a context where militaries are seeking to reform acquisition processes to keep up with the demands of AI, TEVV could be seen as a stumbling block. Second, the ability to develop in-house or sovereign TEVV capabilities depends on existing technical capacity and available resources. Even if states have the means to build sovereign TEVV capabilities, this is not something that can be achieved easily or within a limited amount of time. Finally, the lack of standardised TEVV practices limits the prospect of conducting TEVV independently, which reduces accountability.
The challenges arising from the competence–control trade-off are particularly evident in the case of the United States. For example, in March 2026, the DoW designated Anthropic as a ‘supply chain risk’, effectively restricting the use of its products and services by the US military and defence bureaucracy.[163] This unprecedented designation—typically reserved to guard against foreign companies from states deemed as adversaries—occurred due to a disagreement over Anthropic’s request to maintain safeguards over the use of its AI models for mass domestic surveillance and fully autonomous weapons.[164] Secretary of War Pete Hegseth attacked the company in a social media post, accusing it of using the ‘sanctimonious rhetoric of “effective altruism” … [to] strong-arm the United States military into submission’.[165] Nonetheless, the DoW’s resort to such a measure is consistent with its military AI strategy, in which it states that the department must ‘utilize models free from usage policy constraints that may limit lawful military applications’.[166] The link to TEVV here is that the strategy directed the CDAO to ‘establish benchmarks for model objectivity as a primary procurement criterion within 90 days’. As of 10 April 2026, Anthropic is in the midst of legally challenging the ‘supply chain risk’ designation on the basis that it was denied due process, but US courts typically do not challenge executive branch decisions related to national security.[167]
Conclusion
This article has aimed to explore the role of TEVV in supporting emerging regulatory guardrails for the responsible and lawful use of military AI systems. It began with an examination of the broader context for military AI, where facts on the ground are evolving much more quickly and at a scale that multilateral dialogue cannot keep up with. This was followed by an overview of TEVV across both the civilian and military domains, which provided a starting point to understand the complexities involved for the latter. These include consideration of practices for all military systems, as well as those for specific types of systems, and tailored practices at the service level, and for AI systems more generally, from a technical standpoint. Building on this, three specific areas were identified where TEVV supports regulatory guardrails: (1) providing the means to tangibly measure desired outcomes of an AI system, (2) enhancing human agency across the lifecycle of AI systems, and (3) enabling the conduct of legal reviews. However, this is not exhaustive—TEVV can also play a role in other ways, such as integrating expertise and assessments that were previously siloed, which can lead to more comprehensive and effective consideration of responsible and lawful use of military AI systems. Another example is the potential for TEVV to encourage cooperation and technical exchanges where such activities may have been constrained.
The three case studies examined—Australia, the United Kingdom, and the United States—demonstrate the range of possible ways to approach TEVV. They also highlight the tension between commitments made at platforms such as the REAIM summits and a state’s pursuit of its interests. The case studies clearly demonstrate that, contrary to what its technical nature might suggest, TEVV is neither objective nor neutral—militaries leverage it in service of larger strategic objectives, and it is as much as a political endeavour as it is a technical one. The case studies also explore differences at the service branch level, and underscore the need for more detailed examination in future research of how differing operational, organisational and strategic contexts can have an impact on TEVV. While this article did not examine how TEVV shapes regulatory guardrails when it is conducted in the context of developing capabilities jointly as in a partnership like AUKUS, this is an additional area worth considering for future research. A further area worth exploring relates to developing strategies and approaches to address the three sets of challenges identified for TEVV’s role in supporting regulatory guardrails: (1) the lack of standardised TEVV practices for military AI systems, (2) variations in implementing TEVV, and (3) the private sector’s involvement. While it may be that some of the challenges cannot be addressed entirely, further discussion regarding what can be achieved will strengthen the ability of TEVV to facilitate responsible and lawful use of military AI systems.
Endnotes
[1] Applications of AI in the military domain range across both operational activities (e.g. command, control and communications) and non-operational activities (e.g. acquisition and procurement). See Global Commission on Responsible AI in the Military Domain, Responsible by Design: Strategic Guidance Report on the Risks, Opportunities, and Governance of Artificial Intelligence in the Military Domain (The Hague: Hague Centre for Strategic Studies, 2025), pp. 11–19.
[2] See ‘TFL-1 Autonomy Modules: Terminal Guidance and Cruise’, The Fourth Law (website), at: https://thefourthlaw.ai/#section4.
[3] Yuval Abraham, ‘“Lavender”: The AI Machine Directing Israel’s Bombing Spree in Gaza’, +972 Magazine, 3 April 2024.
[4] Emelie Andersin, ‘The Use of the “Lavender” in Gaza and the Law of Targeting: AI-Decision Support Systems and Facial Recognition Technology’, Journal of International Humanitarian Legal Studies 16, no. 2 (2025): 336–370.
[5] Tess McClure and Deepa Parent, ‘Minab School Bombing: How the Worst Mass Casualty Event of the Iran War Unfolded—A Visual Guide’, The Guardian, 3 March 2026.
[6] Alex Woodward, ‘Old Intelligence and AI? Behind the Deadly Attack on an Iranian Girls’ School That Left 175 Dead’, The Independent, 12 March 2026.
[7] Kevin T Baker, ‘AI Got the Blame for the Iran School Bombing. The Truth is Far More Worrying’, The Guardian, 26 March 2026.
[8] For the list of endorsing states, see ‘REAIM: Responsible AI in the Military Domain Summit’, REAIM (website), at: www.exteriores.gob.es/en/REAIM2026/Paginas/default.aspx. The number of states endorsing the Pathways to Action document grew from 35 at the summit itself to 44 as of 31 March 2026. The text of the outcome document is REAIM Pathways to Action (A Coruña: Ministry of Foreign Affairs, the European Union and Cooperation, 2026), at: www.exteriores.gob.es/en/REAIM2026/Documents/REAIM%202026%20Pathways%20to%20Action.pdf.
[9] Manoj Harjani and Mei Ching Liu, ‘The Uncertain Future of the REAIM Summit’, IDSS Papers, no. 029–26 (2026).
[10] Ibid.
[11] See REAIM Call to Action (The Hague: Government of the Netherlands, 2023).
[12] See REAIM Blueprint for Action (Seoul: Ministry of Foreign Affairs, 2024).
[13] See Bureau of Arms Control and Nonproliferation, Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy (US Department of State, 2023).
[14] Ibid.
[15] Valerie Hafez, Rania Wazir and Fariba Karimi, AI as Complex Sociotechnical Systems: Problems, Approaches and Reflections (Vienna: Women in AI Austria, 2023), p. 2.
[16] Arthur Holland Michel, The Black Box, Unlocked: Predictability and Understandability in Military AI (Geneva: United Nations Institute for Disarmament Research, 2020), p. 1.
[17] Tim McFarland, ‘Reconciling Trust and Control in the Military Use of Artificial Intelligence’, International Journal of Law and Information Technology 30, no. 4 (2022): 472–483.
[18] REAIM Blueprint for Action, p. 2.
[19] Ibid., p. 2.
[20] REAIM Pathways to Action, p. 3.
[21] Ibid, pp. 4–5.
[22] Bureau of Arms Control and Nonproliferation, Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy.
[23] In the civilian domain, there is a rich body of standards associated with TEVV—for example, ISO/IEC/IEEE 24765:2017, ISO/IEC/IEEE 12207:2026, IEEE 829-2008, and IEEE 1012-2024.
[24] Michael Borowski and Priscilla Glasgow, When Systems Are Simulations—T&E, VV&A, or Both? (McLean VA: MITRE Corporation, 1998), p. 1.
[25] Feras A Batarseh , Laura Freeman and Chih‑Hao Huang, ‘A Survey on Artificial Intelligence Assurance’, Journal of Big Data 8, no. 60 (2021): 2.
[26] John Burden, Marko Tešić, Lorenzo Pacchiardi and José Hernández-Orallo, ‘Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture’, Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, Montreal, Canada, August 2025 (International Joint Conferences on Artificial Intelligence, 2025), p. 10382.
[27] Ibid., p. 10384.
[28] Ibid., p. 10386.
[29] Ibid., p. 10386.
[30] Bálint Gyevnár and Atoosa Kasirzadeh, ‘AI Safety for Everyone’, Nature Machine Intelligence 7 (2025): 531–542.
[31] See ‘The Bletchley Declaration by Countries Attending the AI Safety Summit, 1–2 November 2023’, GOV.UK, at: www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023.
[32] See ‘Seoul Declaration for Safe, Innovative and Inclusive AI by Participants Attending the Leaders’ Session of the AI Seoul Summit, 21st May 2024’, Ministry of Foreign Affairs, Republic of Korea (website), at: www.mofa.go.kr/eng/brd/m_5674/view.do?page=1&seq=321007.
[33] See ‘Statement on Inclusive and Sustainable Artificial Intelligence for People and the Planet’, Permanent Mission of France to the United Nations in New York (website), at: https://onu.delegfrance.org/statement-on-inclusive-and-sustainable-artificial-intelligence-for-people-and.
[34] See ‘AI Impact Summit Declaration, New Delhi’, media release, 21 February 2026, Ministry of External Affairs (website), at: www.mea.gov.in/bilateral-documents?dtl/40809.
[35] There is no universal definition for a military AI system. The term is used here in its broadest sense, encompassing standalone AI applications in the military domain as well as AI-enabled functionality within other types of military systems.
[36] This includes practices from AI in other safety-critical areas, such as air traffic control, outer space, nuclear systems and medicine.
[37] The relevant set of practices depends on the way in which a system is classified and defined. Given the challenges associated with defining military AI systems, this is also acknowledged as a challenge affecting TEVV.
[38] Michael C Horowitz, ‘AI and the Diffusion of Global Power’, in Allison Leonard, Jennifer Goyder and Lynn Schellenberg (eds), Modern Conflict and Artificial Intelligence (Waterloo: Centre for International Governance Innovation, 2020), pp. 36–37.
[39] RS Panwar, Li Qiang and John NT Shanahan (eds), Military Artificial Intelligence Test and Evaluation Model Practices (Geneva: INHR, 2024), pp. 4–5.
[40] Heather M Wojton, Daniel J Porter and John W Dennis, Test & Evaluation of AI-Enabled and Autonomous Systems: A Literature Review (Alexandria VA: Institute for Defence Analyses, 2020), pp. 27–28.
[41] Ibid., p. 19.
[42] Manoj Harjani and Shantanu Sharma, Adversarial Attacks: An Existential Threat to AI, IDSS Paper No. 078 (RSIS, 2023).
[43] McFarland, ‘Reconciling Trust’.
[44] Jean-Marc Rickli, Federico Mantellassi and Quentin Ladetto, What, Why and When? A Review of the Key Issues in the Development of Military Human-Machine Teams (Geneva: Geneva Centre for Security Policy, 2024), pp. 10–16.
[45] McFarland, ‘Reconciling Trust’, p. 474.
[46] Ibid., p. 477.
[47] Heather M Roff and David Danks, ‘“Trust but Verify”: The Difficulty of Trusting Autonomous Weapons Systems’, Journal of Military Ethics 17, no. 1 (2018): 6.
[48] Ibid., pp. 6–7.
[49] Jovana Davidovic and Mitt Regan, ‘Jus Ante Bellum and AI-Enabled Weapons’, in Maria Power and Maggi Savin-Baden (eds), Just War Theory and Artificial Intelligence: Challenges and Consequences (Boca Raton FL: CRC Press, 2026), pp. 4–6.
[50] Ibid., p. 4.
[51] Ibid., p. 4.
[52] Ibid., p. 5.
[53] Ibid., p. 5.
[54] Ibid., p. 5.
[55] Ibid., pp. 5–6.
[56] Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt and Patrick Hall, Towards a Standard for Identifying and Managing Bias in Artificial Intelligence (Gaithersburg, MD: National Institute of Standards and Technology, 2022), p. 10.
[57] REAIM Call to Action, p. 2.
[58] Ibid., p. 2.
[59] REAIM Blueprint for Action, p. 2.
[60] Ibid., p. 4.
[61] Bureau of Arms Control and Nonproliferation, Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy.
[62] Ibid.
[63] Ingvild Bode, Anna Nadibaidze, Tom Watts and Qiaochi Zhang, ‘Ensuring the Exercise of Human Agency in AI-Based Military Systems: Concerns Across the Lifecycle’, Ethics and Information Technology 27, no. 50 (2025): 2.
[64] The AutoPractices Project, Strengthening Human Agency in the Military Domain: Best Practices Toolkit for Policymakers, Developers, and Users of AI Systems (Odense: Centre for War Studies, University of Southern Denmark, 2026), pp. 17–18.
[65] REAIM Blueprint for Action, p. 2.
[66] Ibid., p. 5.
[67] Bureau of Arms Control and Nonproliferation, Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy.
[68] Klaudia Klonowska, Article 36: Review of AI Decision-Support Systems and Other Emerging Technologies of Warfare, Asser Research Paper 2021-02 (Asser Institute, 2021), p. 2.
[69] Klaudia Klonowska and Jonathan Kwik, ‘Artificial Intelligence in Contemporary Conflicts and the Future of Military Law’, in Paul AL Ducheine, Terry D Gill, Peter BMJ Pijpers and Marten C Zwanenburg (eds), A Research Agenda for Military Law (Cheltenham: Edward Elgar Publishing, 2026), p. 105.
[70] Tobias Vestner and Altea Rossi, ‘Legal Reviews of War Algorithms’, International Law Studies 97 (2021): 547.
[71] Global Commission on Responsible AI in the Military Domain, Responsible by Design.
[72] Ibid., p. 42.
[73] Jonathan Kwik, Iterative Assessment for Military Artificial Intelligence (AI) Systems, ASSER Research Paper 2025-03 (Asser Institute, 2025), p. 14.
[74] See ‘The International Committee of the Red Cross as Guardian of International Humanitarian Law’, International Committee of the Red Cross (website), at: www.icrc.org/en/article/guardian-international-humanitarian-law.
[75] Artificial Intelligence and Machine Learning in Armed Conflict: A Human-Centred Approach (Geneva: International Committee of the Red Cross, 2019), p. 7.
[76] Tom Corben and Sophie Mayo, Federation Is Deterrence: The US Defence Industrial and Technology Integration Agenda in the Indo-Pacific (Sydney: The United States Studies Centre, 2025).
[77] See Department of Defence, Policy Settings for Responsible Use of Artificial Intelligence in Defence (Canberra: Commonwealth of Australia, 2025).
[78] Ibid., p. 5.
[79] Ibid., p. 6.
[80] Ibid., p. 7.
[81] Ibid., p. 8.
[82] See ‘Policy for the Responsible Use of AI in Government: Version 2.0’, digital.gov.au, at: www.digital.gov.au/ai/ai-in-government-policy. The policy is, as of April 2026, in its second version since December 2025.
[83] See ‘Australia’s AI Ethics Principles’, Department of Industry, Science and Resources (website), at: www.industry.gov.au/publications/australias-ai-ethics-principles. Eight ethics principles were identified in this framework: (1) human, societal, and environmental wellbeing, (2) human-centred values, (3) fairness, (4) privacy protection and security, (5) reliability and safety, (6) transparency and explainability, (7) contestability and (8) accountability.
[84] Department of Defence, Policy Settings for Responsible Use of Artificial Intelligence, p. 9.
[85] Ibid., p. 9.
[86] See Department of Defence, Defence Test & Evaluation (T&E) Strategy (Canberra: Commonwealth of Australia, 2021).
[87] Ibid., p. 3.
[88] Department of Defence, SDIP 7—Test and Evaluation, Certification and Systems Assurance (Canberra: Commonwealth of Australia, 2024), p. 4.
[89] See Australian Defence Force, Concept for Robotic and Autonomous Systems (Canberra: Commonwealth of Australia, 2020).
[90] Ibid., p. 31.
[91] Ibid., p. 35.
[92] Ibid., p. 59.
[93] Australian Public Service Commission, State of the Service Report 2024–25 (Canberra: Commonwealth of Australia, 2025), p. 112.
[94] See Army, Robotic & Autonomous Systems Strategy 2.0 (Canberra: Commonwealth of Australia, 2022).
[95] Ibid., p. 32.
[96] Ibid., p. 33.
[97] Ibid., p. 34.
[98] Ibid., p. 36.
[99] See Navy, RAS-AI Strategy 2040 (Canberra: Commonwealth of Australia, 2020).
[100] Ibid., p. 19.
[101] Ibid., p. 24.
[102] Ibid., p. 24.
[103] See Royal Australian Air Force, Plan Jericho (2021), at: https://airpower.airforce.gov.au/sites/default/files/2021-03/AF14-Plan-Jericho.pdf.
[104] See ‘Air Warfare Centre’, Royal Australian Air Force (website), at: www.airforce.gov.au/about-us/hq-air-command/air-warfare-centre.
[105] See ‘M113 AS4 Optionally Crewed Combat Vehicle’, BAE Systems (website), at: www.baesystems.com/en/product/m113-occv.
[106] See ‘Bluebottle: Satellites of the Sea’, OCIUS (website), at: https://ocius.com.au/usv/#overview.
[107] See ‘MQ-28 Ghost Bat’, Boeing (website), at: www.boeing.com/defense/autonomous-and-unmanned-systems/mq-28-ghost-bat.
[108] ‘AI Takes Flight on Talisman Sabre’, Department of Defence (website), 11 August 2025, at: www.defence.gov.au/news-events/news/2025-08-11/ai-takes-flight-talisman-sabre.
[109] ‘AUKUS Pillar II Under Pressure’, Strategic Comments 31, no. 35 (2025): 1–4.
[110] Ministry of Defence, Ambitious, Safe, Responsible: Our Approach to the Delivery of AI-Enabled Capability in Defence (London: Ministry of Defence, 2022), p. 1.
[111] Ibid., pp. 10–11.
[112] See Ministry of Defence, Defence Artificial Intelligence Strategy (London: Ministry of Defence, 2022).
[113] Ibid., p. 13.
[114] See Ministry of Defence, JSP 936 v1.1 Dependable Artificial Intelligence (AI) in Defence Part 1: Directive (London: Ministry of Defence, 2024).
[115] See Ministry of Defence, JSP 375 Management of Health and Safety in Defence (London: Ministry of Defence, 2013).
[116] See Ministry of Defence, JSP 376 Defence Acquisition Safety Policy (London: Ministry of Defence, 2023).
[117] See Ministry of Defence, Defence Information, Knowledge, Digital and Data Policy Commitments (London: Ministry of Defence, 2024).
[118] See Ministry of Defence, JSP 912 Human Factors Integration in Defence Systems (London: Ministry of Defence, 2014).
[119] See Ministry of Defence, JSP 939 Defence Policy for Modelling and Simulation (London: Ministry of Defence, 2018).
[120] See Ministry of Defence, Test and Evaluation (T&E): Future Advantage Through Evaluation (FATE) (London: Ministry of Defence, 2026).
[121] Ministry of Defence, Defence Artificial Intelligence Strategy, p. 27.
[122] JSP 936 V1.1 Dependable Artificial Intelligence (AI) in Defence Part 1, p. 30.
[123] See ‘The Defence AI Centre (DAIC) AI Assurance Framework’, Ministry of Defence (website), at: www.digital.mod.uk/policy-rules-standards-and-guidance/aiph.
[124] See ‘AI Practitioner’s Handbook (AIPH)’, Ministry of Defence (website), at: www.digital.mod.uk/policy-rules-standards-and-guidance/aiph.
[125] See ‘Resource Library’, Ministry of Defence (website), at: www.digital.mod.uk/policy-rules-standards-and-guidance/aiph/resource-library.
[126] See Ministry of Defence, British Army’s Approach to Artificial Intelligence (London: Ministry of Defence, 2023).
[127] House of Commons Defence Committee, Developing AI Capacity and Expertise in UK Defence (London: House of Commons, 2025), pp. 26–29.
[128] Ministry of Defence, ‘Launching the AI Model Arena’, GOV.UK, at: www.gov.uk/government/news/launching-the-ai-model-arena.
[129] Ministry of Defence, Strategic Defence Review: Making Britain Safer: Secure at Home, Strong Abroad (London: Ministry of Defence, 2025), p. 96.
[130] Ibid., 56.
[131] See Secretary of War, Artificial Intelligence Strategy for the Department of War (Washington DC: Department of War, 2026).
[132] See President, Removing Barriers to American Leadership in Artificial Intelligence, Executive Order 14179 (Executive Office of the President, 2025).
[133] See Winning the Race: America’s AI Action Plan (Executive Office of the President of the United States, 2025).
[134] Secretary of War, Artificial Intelligence Strategy for the Department of War.
[135] Ibid., pp. 2–3.
[136] See Deputy Secretary of War, Transforming Advana to Accelerate Artificial Intelligence and Enhance Auditability (Washington DC: Department of War, 2026).
[137] See Secretary of War, Transforming the Defense Innovation Ecosystem to Accelerate Warfighting Advantage (Washington DC: Department of War, 2026).
[138] U.S. Department of War, Acquisition Transformation Strategy (Washington DC: Department of War, 2025), pp. 21–22.
[139] Ibid., p. 22.
[140] Ibid., p. 29.
[141] Secretary of War, Artificial Intelligence Strategy for the Department of War, p. 4.
[142] Ibid., p. 5.
[143] See Chief Digital & Artificial Intelligence Office (CDAO), Test and Evaluation of Artificial Intelligence Models (CDAO, 2024).
[144] See Chief Digital & Artificial Intelligence Office (CDAO), Human Systems Integration Test and Evaluation of Artificial Intelligence-Enabled Capabilities (CDAO, 2024).
[145] See Chief Digital & Artificial Intelligence Office (CDAO), Systems Integration Test and Evaluation of Artificial Intelligence-Enabled Capabilities (CDAO, 2024).
[146] See Chief Digital & Artificial Intelligence Office (CDAO), Operational Test and Evaluation of Artificial Intelligence-Enabled Capabilities (CDAO, 2024).
[147] See ‘CDAO JATIC Documentation’, CDAO (website), at: https://cdao.pages.jatic.net/public.
[148] See Department of Defense, DOD Manual 5000.101 Operational Test and Evaluation and Live Fire Test and Evaluation of Artificial Intelligence-Enabled and Autonomous Systems (Department of Defense, 2024).
[149] Davidovic and Regan, ‘Jus Ante Bellum and AI-Enabled Weapons’, pp. 4–6.
[150] Panwar, Qiang and Shanahan (eds), Military Artificial Intelligence Test and Evaluation Model Practices, p. 6.
[151] Ibid., p. 6.
[152] David Helmer, Michael Boardman, S Katie Conroy, Adam Hepworth and Manoj Harjani, ‘Human-Centred Test and Evaluation of Military AI’, arXiv (2024), p. 6, at: https://arxiv.org/abs/2412.01978.
[153] Office of Systems Engineering and Architecture, Test and Evaluation Workforce Report (Washington DC: Office of the Under Secretary of Defence for Research and Engineering, 2025), p. 9.
[154] Department of Defence, SDIP 7—Test and Evaluation, Certification and Systems Assurance, p. 3.
[155] See T3E (website), at: www.t3e.uk/en.
[156] Defence and Security Accelerator, ‘Delivering Future Advantage Through Testing and Evaluation: New £1 Million Themed Competition Launched’, GOV.UK, at: www.gov.uk/government/news/delivering-future-advantage-through-testing-and-evaluation-new-1-million-themed-competition-launched.
[157] See President, Restoring the United States Department of War, Executive Order 14347 (Executive Office of the President, 2025).
[158] ‘Secretary of War Pete Hegseth Addresses General and Flag Officers at Quantico, Virginia’, speech, Quantico VA, 30 September 2025, transcript at: www.war.gov/News/Transcripts/Transcript/article/4318689/secretary-of-war-pete-hegseth-addresses-general-and-flag-officers-at-quantico-v.
[159] Secretary of War, Artificial Intelligence Strategy for the Department of War, p. 5.
[160] Brandi Vincent and Drew F Lawrence, ‘The Era of GenAI.mil Is Here. Users Have Mixed Reactions and Many Questions’, Defense Scoop, at: https://defensescoop.com/2025/12/18/genai-mil-users-have-mixed-reactions-and-many-questions.
[161] Andrea Johansen and Andreas Kruck, ‘The Competence–Control Trade-Off in Military AI Innovation: Autonomous Weapons Systems and Shifting Modes of State Control over Private Experts’, European Journal of International Security (2025): 3.
[162] Ibid., p. 6.
[163] Karen Freifeld and Deepa Seetharaman, ‘Pentagon Designates Anthropic a Supply Chain Risk’, Reuters, 6 March 2026.
[164] ‘Statement from Dario Amodei on Our Discussions with the Department of War’, Anthropic (website), 26 February 2026, at: www.anthropic.com/news/statement-department-of-war.
[165] Pete Hegseth (@SecWar), ‘This week, Anthropic delivered a master class in arrogance and betrayal as well as a textbook case of how not to do business with the United States Government or the Pentagon’, X, 28 February 2026, at: https://x.com/SecWar/status/2027507717469049070.
[166] Secretary of War, Artificial Intelligence Strategy for the Department of War, p. 5.
[167] Jack Queen, ‘Anthropic Has Strong Case Against Pentagon Blacklisting, Legal Experts Say’, Reuters, 11 March 2026.