AI-DSS and the Erosion of Technical Integrity
Rethinking Human Decision-Making in Modern Warfare
Authors: Associate Professor Zena Assaad and Lieutenant Colonel Dr Adam Hepworth
‘Someone decided to compress the kill chain… Calling it an “AI problem” gives those decisions, and those people, a place to hide.’ Kevin T Baker, The Guardian, 26 March 2026
Introduction
The adage of military artificial intelligence (AI) no longer being an imagined reality has become a standard opening remark for much of the discourse on the topic of AI in military operations. Current conflicts in the Ukraine, Palestine and Iran serve as a backdrop for this, demonstrating multiple accounts of AI use in active conflicts. This has shifted the discourse away from hypothetical instantiations to real-world case studies, with aspects of these case studies involving the use of AI decision-support systems (AI-DSS) for targeting operations.
AI-DSS are currently being used in the Israel–Gaza war[1] and the Russia–Ukraine war[2] and most recently in support of Operation Epic Fury between the United States, Israel and Iran.[3] [4] AI-DSS are systems designed to support decision-making by collating and analysing large datasets, shortening decision loops[5] and generating decision-support outputs.[6] These tools have a wide range of uses, which can be categorised into three broad functions:
- description, synthesis and analysis, which involves collecting, organising and presenting data
- prediction and extrapolation, which entails identifying patterns and trends in data, including forecasting possible outcomes and their probabilities
- recommendations, which focus on recommending the most effective course(s) of action.[7]
Specific operational examples of AI-DSS used for targeting include Lavender (target prioritisation), The Gospel (intelligence fusion) and Where’s Daddy (location tracking), which have been used by the Israel Defence Forces (IDF) in Gaza.[8] Lavender uses machine learning to assign a numerical score indicating the likelihood of a person being a member of an armed group. The Gospel processes surveillance data to generate and categorise targets. Where’s Daddy utilises location data from connected devices to detect and locate people identified as military targets.
More recently, there have been reports indicating the use of Palantir’s Maven Smart System, a machine-learning and sensor-fusion command and control platform, as well as Anthropic’s Claude, a frontier large language model, in a decision-support capacity by the US military for targeting operations in Iran as part of Operation Epic Fury.[9] Public reports suggest Claude was being used for ‘intelligence assessments, including for target identification and simulating battle scenarios’ in the lead-up to strikes on Iran.
The IDF claims that the use of AI-DSS has increased operational precision and accuracy,[10] [11] although reports of increased civilian death tolls in Gaza provide contradictory claims.[12] This outcome is, in part, due to the limited accuracy, validation and reliability of these systems in real-world scenarios, as well as the level of procedural integration and oversight of targeting decisions.[13] This highlights the impact of upstream decision-making in capability development intersecting with operational practices. Technology capabilities are not deployed in a vacuum; rather they are implemented into existing military systems and operational procedures, some of which they complement and others with which they conflict. These points of conflict are often identified and mitigated in both acquisition and operational processes; however, as the speed of fielding is prioritised, established processes can be narrowed or even overlooked.
In this paper we argue that the narrowing and omission of critical military processes results in an accumulative degradation of human decision-making over the lifecycle of an AI-enabled system, leading to eroded technical integrity and compromised operational effectiveness. We frame operational effectiveness in this paper through the lens of verification leading to protection, including the safety of non-combatants and combatants, consistent with the international humanitarian law (IHL) principles of distinction (distinguishing between combatants and non-combatants) and proportionality (ensuring disproportionate impacts to civilians during military operations are minimised). Through this lens, an AI-DSS may be functioning ‘as designed’ from an operational verification perspective, while still producing outcomes incompatible with objective IHL distinction and proportionality assessments.
We posit that human decision-making upstream of the point of use—in capability development through requirements setting, acquisition, design, test and evaluation, regulatory and standards compliance including Article 36 legal reviews, and the assignment of accountable officers—is critical to technical integrity and operational effectiveness for AI-DSS when used in operational targeting processes.
This position aligns with Australian responsible AI frameworks as presented in Australia’s Policy Settings for Responsible Use of Artificial Intelligence in Defence (2025), which establishes a values-based approach through three obligations for AI use: lawfulness, adherence to values-based principles, and proportionate controls.[14] The policy requires Defence personnel to consider legal obligations at all stages of decision-making to acquire and use AI. While this is specific to legal obligations, it should be noted that the requirements stage establishes where principles of responsible behaviour must first be embedded, including technical requirements of which some fall under legally mandated regulations. Goussac and Boulanin[15] articulate a similar approach, emphasising a need to consider principles of responsible behaviour throughout the acquisition lifecycle, rather than attempting to ‘retrofit’ them at a later stage.
While this paper focuses on human decision-making, comparative literature on human agency should also be noted. Conn & Bode suggest that upstream agency could be eroded for AI-enabled systems, noting:
[M]ost of these ethical and legal discussions have revolved primarily around the point of use of a hypothetical AIS,[16] and in doing so, one critical component still remains under-appreciated: human decision-making across the full timeline of the AIS lifecycle.
Their proposed corrective action is a ‘policy knot’—as a recurring loop between policy, design and operational practice that must hold across the full lifecycle of an AI-DSS.[17] As the legal-analytical companion to this, Dorsey and Bo state:
[I]ncreased speed and scale in conflict made possible by the integration of AI-DSS, especially within the JTC [joint targeting cycle], can be a tradeoff with other considerations, such as the erosion or loss of human agency.[18]
While these studies focus on human agency, they also highlight the significance of human decision-making across the system lifecycle, motivating the need for a comprehensive discussion of how to preserve human decision-making in AI-DSS beyond the point of procedural use and by what means.[19]
The Institute of Electrical and Electronics Engineers Standards Association (IEEE SA) Research Group on Issues of Autonomy and AI in Defence Systems discusses the significance of technical testing, evaluation, verification and validation alongside human decision-making through the lens of human operators. While this could be characterised as a technical or policy challenge, we frame it also as a procedural challenge focused around the idea of human skills or human readiness level. Conn and Bode outline a need for human readiness levels, modelled on the technical readiness level framework widely in use today. Australian defence policy articulates this need, noting that the policy takes a human-centric approach to consider impacts across all functions. The policy goes on to establish the ‘accountable officer’ framework within the capability lifecycle, uplifting accountability from an operational-only lens to include acquisition. Recent literature[20] highlights that acquisition can be deliberately structured as the implementation lever, with emerging defence acquisition reforms demonstrating progress systemically in Australia.
In section 2 of this paper we introduce two critical military processes, the system lifecycle and the joint targeting process (JTP), highlighting how they intersect. In section 3 we discuss the complications of compromised technical integrity, focusing on three aspects central to AI-DSS in targeting operations: the speed and scale of target generation and nomination; the contention between accuracy and precision; and the explainability of the system outputs. In section 4 we explore the relationship between decision-making and time, and the impacts of AI-DSS on this relationship. In the final section we present recommendations for decision-makers across both the system lifecycle and the targeting operations.
The Cycles of Military Systems and Operations
Military operations are enabled and conducted via iterative and repetitive processes which have been established and refined over years, decades and centuries. These cycles facilitate procedural mechanisms for acquiring military systems and deploying them in operational contexts. There are two cycles we will focus on here, one at the acquisition level and the other at the operational level: the system lifecycle and the JTP respectively.
System Lifecycle
At the acquisition level, systems engineering is a standard framework employed by many states to procure and acquire military systems.[21] [22] [23] Systems engineering is an interdisciplinary approach that enables the successful realisation, use and retirement of engineered systems.[24] This framework is built around the system lifecycle, a sum of phases and activities which bring a system into being, support and enable its use and then dispose of it once it can no longer serve its intended purpose.[25] The system lifecycle is a linear succession of phases that involves cyclical processes within each phase.
Human-made engineered systems come into being through ‘purpose-driven human action’[26] across a system lifecycle. The significance of the human role in the system lifecycle process, particularly the weapon system lifecycle,[27] is critical to understanding the true impacts of AI-DSS used in targeting operations.
Various representations of system lifecycles exist. Within the context of this paper we will refer to the system lifecycle illustrated in Figure 1, adapted from our previous work.[28] We deliberately use an idealised end-to-end representation of a systems lifecycle, which has been adapted from the specific lifecycle in use in Defence and work from the IEEE SA Research Group on Issues of Autonomy and AI in Defense Systems;[29] importantly, the lifecycle has been amended to capture necessary considerations for AI-enabled systems.
This lifecycle involves four stages: three within the acquisition phase and the fourth within the utilisation phase of an engineered system. The four high-level lifecycles stages are:
- Plan. The planning phase includes defining the scope and determining the objectives of a system.
- Design. The design phase involves outlining specific requirements, architectures, functions, interfaces etc. of the system.
- Develop. The development stage involves the actual building of the system, which includes testing, evaluation, verification, and validation (TEVV).
- Deploy. The deployment stage involves integrating the system into operations, which includes maintenance throughout the life of the system.
The iterative and cyclical nature of systems engineering practices is seen in the repetitive application of analysis, synthesis and evaluation across all four lifecycle stages, which progress linearly over time.[30] Analysis resolves what is required and why, synthesis determines how it can be achieved, and evaluation assesses the trade-offs between the what and the how. A short note on terminology: acquisition and procurement are commonly used as synonyms; however, there is a distinction between the two terms. Procurement refers only to purchase, while acquisition encompasses the broader lifecycle process, including the act of procurement.[31]
One of the core phases of a system lifecycle is TEVV, which predominantly, but not exclusively, occurs during the design and development phases. Testing and evaluation involves assessing the performance, reliability and safety of a system. It involves systematically examining and validating these factors to ensure they meet specified requirements and perform as intended. Testing and evaluation has four major categories:
- Preview test and evaluation (T&E) to assist analysis and refinement of major acquisition options for approval
- Development T&E to support design and development efforts at the beginning acquisition stages
- Acceptance T&E, which is formal acceptance testing conducted when the system is more mature
- Operational T&E, which involves testing what the system would be like in operation, using operation-like conditions or scenarios.
Verification refers to the evaluation of whether a product, service or system complies with a regulation, requirement, specification or imposed condition. It is often an internal process. Validation refers to the assurance that a product, service or system meets the needs of the customer and other identified stakeholders. It often involves acceptance and suitability testing with external stakeholders. In essence, verification is the process of ensuring you have built the system right and validation is ensuring you have built the right system.
Targeting Cycle
The JTP is a cyclical and continuous activity consisting of six phases designed to support a commander in achieving an end state (detailed in Table 1).[32] The JTP is the generic targeting methodology that has been used by the Australian military and developed against the US joint targeting cycle, which similarly consists of a six-phase process.[33]
| Phase | Description | |
|---|---|---|
| 1 | Commander’s guidance | Mission, objectives, intent and desired effects |
| 2 | Target development | Intelligence direction, analysis, validation and target list management |
| 3 | Capabilities analysis | Best available means to affect targets or target sets |
| 4 | Force application | Assigning forces, weapons and/or other capabilities to targets |
| 5 | Execution | Applying the force to realise objectives and desired effects |
| 6 | Assessment | Effects assessment, battle damage assessment, weapons effectiveness assessment, collateral assessment and re-attack recommendations |
Targeting refers to the process of ‘selecting and prioritizing targets and matching the appropriate response to them, considering operational requirements and capabilities’.[35] A target is defined as an object, person, organisation or geographic area which can be engaged or influenced in military operations to achieve a political end state.[36]
It should be noted that the JTP represents one specific model for targeting; however, it is not a universally adopted model. What this model represents, more broadly, is the institutional approach to military targeting operations. While not universal, it is emblematic of military operational processes and how weapon systems are deployed in such systematic and cyclical processes.
Within the JTP, there are two categories of targeting: deliberate targeting, which involves planned targets; and dynamic targeting which involves ‘targets of opportunity’—targets not identified or selected in time to be included in deliberate targeting processes.[37] Both of these categories broadly follow the same six-phase targeting process, with the critical differentiator being time. Dynamic targeting deliberately compresses the target development phase with the intention of enabling engagement of ‘targets of opportunity’.
Similar to the system lifecycle, the JTP is an iterative and cyclical[38] process intended to manage decision-making, across multiple human actors and systems, to achieve a desired outcome. These two cycles intersect when an engineered military system, which has been acquired through a systems engineering process, is deployed and utilised in targeting operations which employ the JTP. The cycles of decision-making within the system lifecycle ultimately have flow-on effects on the targeting process. For AI-DSS used in targeting operations, those flow-on effects emerge through an erosion of technical integrity which compromises operational effectiveness and is compounded by temporally compressed decision-making during operations, as is discussed in the next section.
Through this work, we position the systems lifecycle (acquisition) as the upstream determinant of the JTP (operational)—two intrinsically connected processes. While human decision-making must be distributed across the total system, we argue that the most consequential system-level decisions are taken before the system is fielded.[39]
Compromised Technical Integrity
Technical integrity refers to a system’s fitness for service, safety and compliance with regulations.[40] Fit for service differs from fit for purpose: the former refers to meeting technical standards and operational requirements and the latter relates to accomplishing intended functions or outcomes. Fit for purpose is outcome focused, while fit for service is system focused and a pillar of technical integrity.
Technical integrity is a product of human decision-making and a reflection of human intervention in engineered systems. One of the predominant methods for enabling technical integrity is through TEVV, which largely, but not exclusively, occurs during the design and development stages of the system lifecycle. Accelerated development of AI technologies is achieved through accelerated, or in some cases overlooked,[41] systems engineering practices, including TEVV, which underpin technical integrity. Technical regulations are a means of enforcing and assuring TEVV practices, and consequently are an additional pillar of technical integrity;[42] however, the absence of such regulations specific to AI has created an exploitable gap which further compromises technical integrity.
The narrowing and potential omission of these critical processes can lead to an accumulative breakdown of decision-making over the course of the systems lifecycle of AI-DSS, culminating in the erosion of TEVV: the foundation of technical integrity. Technical integrity enables consistency and reliability of systems,[43] with the absence of these attributes resulting in inconsistent and unreliable systems. For example, the IDF’s use of AI-DSS for targeting operations in Gaza has been the subject of substantial reporting on the performance limitations of these systems.[44] This point, more broadly, is structural and well established in the literature.[45] It also alludes to a broader provocation around acceptable risk thresholds for AI-enabled systems in military operations, an open question which is beyond the scope of this paper.
Offering a related critique of the argument in the absence of agreed AI risk thresholds, Khlaaf and Myers West state:
AI technologists have instead advocated for risk tolerances skewed by a purported AI arms race and speculative ‘existential’ risks, taking over the arbitration of risk determinations with life-or-death consequences, subverting democratic processes.[46]
This perspective suggests that ongoing rhetoric about fabricated urgencies can mask irresponsible practices. We offer that this further motivates the need for deliberate treatment of the entire system lifecycle of AI-DSS, expanding beyond the weight of focus on operational-use phases.
The erosion of technical integrity has been identified as complicating three aspects central to AI-DSS use in targeting operations. These are the speed and scale of target generation and nomination; the contention between accuracy and precision; and the explainability of the system outputs. We suggest that all of these three points are manifestations of upstream decision-making which ultimately impact technical integrity and operational effectiveness, as we describe here.
Importantly, the Israel–Gaza war, the Russia–Ukraine war and Operation Epic Fury demonstrate, at present, the most in-depth examples of the use of AI-DSS in targeting operations. For these conflicts—in contrast to previous ones—much of the information on the specifics of use and associated limitations is presented and available in public reporting, as opposed to academic literature; however, many aspects of reported claims are substantiated by technical publications in academic discourse, which we present alongside our argument.[47]
Speed and Scale of Target Generation and Nomination
Multiple investigative reports into the IDF’s use of AI-DSS in targeting operations document the increased speed and scale of target generation and nomination afforded by the use of these systems. Retired Lt Gen Aviv Kohavi, IDF Chief of Staff until 2023, described The Gospel as capable of producing 100 targets per day (over 36,500 annually), against approximately 50 per year for human analyst—a rate increase exceeding two orders of magnitude.[48] [49] [50] [51] It has been reported, although also contested, that human analysts were making decisions regarding a single target within 20 seconds, prior to authorisation. This is consistent with the technical capacity of AI-DSS, which are capable of analysing very large datasets within compressed time frames.[52] Target development, phase two of the JTP, involves intelligence direction, analysis, validation and target list management—tasks traditionally distributed across multiple human analysts. The deployment of AI-DSS in the targeting process compresses these tasks temporally and procedurally, as well as informationally.
While speed is a factor in targeting operations, particularly for dynamic targeting, speed at the expense of operational effectiveness presents a paradoxical challenge. A system is deemed fit for purpose when it accomplishes or supports the accomplishment of intended functions or outcomes. In the case of targeting operations, a system deployed in the JTP will be deemed fit for purpose if targets are correctly identified and engaged in a manner deemed operationally effective; here we take that to mean compliant with IHL principles of distinction and proportionality, as we outlined previously in this paper.
For AI-DSS, the scaled and accelerated identification of targets may allude to a fit-for-purpose system; however, the accuracy of those identifications and the precision of their engagement lends a different perspective. A system can be fit for purpose but not fit for service, an underpinning principle of technical integrity. AI-DSS used in targeting operations are an example of a type of system which could be argued to be fit for purpose but not fit for service.
The erosion of technical integrity which would deem a system fit for service can be the result of the narrowing and omission of critical military processes, leading to an accumulative degradation of decision-making upstream of the point of operational use. The acquisition phase of the system lifecycle—plan, design, develop—traditionally consists of decision-making windows which allow for considerations around objectives of a system, identifying requirements, architectures and functions, legal obligations, and TEVV, among others. The shifting military acquisition landscape has temporally and procedurally compressed these opportunities for intervention. This may be why we are seeing the deployment of AI-DSS in targeting operations which are fit for purpose but not fit for service.
The erosion of technical integrity of AI-DSS used in targeting operations paradoxically complicates the allure of speed and scale of target generation and nomination because the claims of efficiency and effectiveness are contradicted by the reality of operational outcomes. If these systems were as efficient and effective as claimed, it would be logical to assume civilian casualties from targeting operations involving AI-DSS would be minimal, in line with the definition of operational effectiveness presented herein. However, reporting on the civilian death toll in Gaza points to a disproportionately high number of non-combatant casualties.[53] [54] [55]
The use of AI-DSS in targeting operations is not the sole factor contributing to limited operational effectiveness. This outcome is also a product of deploying systems which lack technical integrity into targeting operations, conducted without deliberate, considered and timely decision-making. We note that, specifically regarding the IDF, speed and scale have historically been prioritised in their intelligence gathering and targeting activities.[56] [57]
The degradation of upstream decision-making begins to bleed into operational decision-making through the encouragement of temporally and procedurally compressed decision-making practices during targeting operations. The pursuit of speed and scale begins upstream of operations, during the acquisition phase of the system lifecycle, and is carried forward into operational practices in the JTP.
On the first morning of Operation Epic Fury, US forces struck the Shajareh Tayyebeh primary school in Minab, in southern Iran, killing approximately 156 people, including 120 children, most of them girls aged between seven and twelve.[58] [59] A preliminary US military investigation determined US responsibility for the strikes. As widely reported, two AI systems were used: Palantir’s Maven Smart System for targeting and querying underlying databases, and Anthropic’s Claude large language model in a decision-support capacity.[60] The Commander of US Central Command, US Navy Admiral Brad Cooper, has stated on public record:
these systems help us sift through vast amounts of data in seconds so our leaders can cut through the noise and make smarter decisions faster than the enemy can react. Humans will always make final decisions on what to shoot and what not to shoot and when to shoot.[61]
In 2016, the building had been separated from an adjacent Islamic Revolutionary Guard Corps compound and converted into a school. The change was noted to be visible on both satellite imagery and local business directories.[62] An investigation into this incident found that an analyst had recorded their observation of the change in 2019; however, the analyst’s record was not entered into a digital tool connected to a US military official targeting database.[63] This critical piece of information was therefore missing from the targeting process. A subsequent investigation found repeated decision-making failures over several years, pointing to broader systemic issues. Our focus here is on the disconnect between upstream decision-making and operational outcomes.
The design and development phases of the system lifecycle include activities around implementation compatibility: determining how a system can be robustly and effectively integrated alongside other systems.[64] [65] The investigation into the Minab school bombing revealed a disconnect and a level of incompatibility among the many systems, both technical and operational, utilised in military targeting operations, highlighting the true complexity of system implementation—a complexity which requires its own inquiry that is beyond the scope of this paper. As previously noted, AI systems are not deployed in a vacuum; rather they are implemented into existing military systems and procedures, some of which they complement and others which they conflict with.
The preliminary findings indicate an accumulation of failures across the system lifecycle, of which TEVV practices and consideration of implementation compatibility in the acquisition phase could possibly have minimised the likelihood. A mitigation mechanism such as mandated data quality assurance practices may have emerged during the TEVV process, when limitations of a system are identified, through testing and evaluation, and mitigations are developed to compensate for them. This is how consistency and reliability—that is, technical integrity—can be assured. As processes are being narrowed, or in some cases omitted, the opportunities for human intervention, in the form of decision-making, are also being constrained. Errors can propagate into the targeting process, in which a verification of the target’s status would have been likely to result in a revised assessment.
Table 2 summarises open-source reporting of AI-DSS functions deployed in Operation Epic Fury, and the role of these technologies in decision support. As with the use of AI across a wide range of applications, both civilian and military, we observe that a single, monolithic AI-DSS was not used. Rather, a series of discrete technologies were used to support specific decision outcomes. From available information, it is unclear how decision inputs were handled and the ways in which errors manifested through decision-making processes. However, the accumulation of errors across successive decisions does highlight the implications of accumulative decision failure, and how this can rapidly propagate to lethal outcomes.
| Decision input | AI-DSS function | Technology |
|---|---|---|
| Sensing | Integration of satellite imagery, signals intelligence and uncrewed system sensor data | Maven Smart System, machine learning vision models and machine learning sensor models |
| Target detection and classification | Object recognition and detection confidence scoring | Maven Smart System |
| Intelligence contextualisation | Template matching (algorithmic detections correlated to military facility records) | Human-curated Defense Intelligence Agency database with legacy data (not updated post 2016 changes) |
| Course-of-action development and recommendation | Targeting prioritisation | Maven Smart System |
| Intelligence curation | Search, summarisation and synthesis of intelligence reporting | Large language model, reportedly Anthropic’s Claude |
The emergence of narrowing and omission of critical processes through temporally and procedurally compressed decision-making for AI-DSS in the acquisition phase is likely to be a key contributing factor in technical integrity erosion, which manifests during the deployment of AI systems into targeting operations, leading to compromised operational effectiveness. It is this accumulation of degraded decision-making across the entire system lifecycle—from acquisition through to utilisation—which leads to outcomes such as the example provided.
Three months after Russia’s invasion of Ukraine, Palantir Technologies chief executive Alex Karp arrived in Kyiv to meet President Zelensky, offering to deploy the company’s data and artificial intelligence software in support of Ukraine’s defence.[69] Palantir subsequently entered agreements with at least six Ukrainian agencies, including the Ministry of Defence, the Ministry of Digital Transformation, the Ministry of Economy, and the Ministry of Education.[70] The MetaConstellation platform, now operational, tasks commercial satellite imagery from Maxar Technologies, Planet Labs and Capella Space; integrates signals intelligence; and produces target packages which, in Karp’s own assessment, render Palantir software ‘responsible for most of the targeting in Ukraine’.[71] The specifics of the agreements remain confidential.[72]
This case highlights the emerging commercialisation of military decision-making processes through acquisition, with essential operational and intelligence decisions potentially outsourced to commercial entities. While the need for procurement as implementation here is evident, greater risks may be yet to manifest downstream. The Center for Security and Emerging Technology’s Margarita Konaev has framed the risk as follows:
Most companies operating in Ukraine right now say they align with U.S. national-security goals—but what happens when they don’t? What happens the day after?[73]
We have described technical integrity as a means of enabling consistency and reliability in systems. In the case of AI-DSS, these systems may be capable of scaling and accelerating target identification; however, issues with consistency and reliability emerge in the accuracy and precision of those targets and the method of their engagement.
Contention Between Accuracy and Precision
Accuracy and precision are two terms which mean different things at a technical level. In computer science, precision refers to consistency evident through a lack of, or minimal, variance; whereas accuracy refers to the correctness of outputs.[74] For example, if an AI-DSS correctly identified or classified targets, it would be deemed an accurate system, and if it consistently produced the same identification and classification results for individual targets when prompted multiple times, it would be deemed a precise system. An AI-DSS which produces incorrect target identifications and classifications would be deemed inaccurate; however, if it produced the same output multiple times, despite the output being incorrect, then it would be considered precise. It is for this reason that the distinction between accuracy and precision is important to delineate. The characteristic of precision for software systems is not an indication of accuracy or correctness, and while it may make a system fit for purpose, it does not necessarily make it fit for service.
Precision and accuracy are technical factors which are commonly assessed during TEVV and contribute to the consistency and reliability of a system. What is challenging about AI-DSS used in targeting operations is the dependency on non-technical factors for ensuring technical accuracy. Some of these factors can be captured during the acquisition phase, while others require tailored decision-making at the operational level.
Complexity is an emergent property of AI-DSS,[75] and AI-enabled systems more broadly, wherein certain behaviours or outputs only emerge at the system of systems level, in which there are interconnections and interactions with other elements.[76] One of those elements is data. All AI-enabled systems are data-dependent systems which produce outputs based on training datasets. It is almost impossible to curate a dataset, for any industry or context, which captures every aspect of an operating environment.[77] This is especially true for military operating environments, which are inherently dynamic and often unpredictable. An AI-DSS may be considered precise and accurate prior to deployment, based on its performance against a bounded use case represented in its training dataset, but inaccurate once deployed in a context which exceeds the bounds of its training dataset.
For AI-DSS, or AI systems more broadly, accuracy is neither static nor definitive. This is why a system may be deemed fit for service at the point of deployment but its technical integrity may erode once it is deployed in targeting operations in which accuracy, an operational requirement of the system, becomes compromised. Here, the erosion of technical integrity did not begin at the point of deployment; rather, it began upstream of the point of use as a result of narrowed or omitted TEVV processes, the implication of which was amplified when the system was deployed in the targeting process.
One of the major categories of testing and evaluation is operational testing, which involves testing what the system would be like in operation using operation-like conditions or scenarios. While it is true that the dynamic and unpredictable nature of a military operating environment is almost impossible to completely capture in a dataset, the limitations that come with this reality must be mitigated through both technical interventions (technical controls) and operational practices (procedural controls). One example of an operational practice to mitigate accuracy limitations is verification requirements of identified targets during targeting operations that extend beyond a performative ‘rubber stamp’. This may include integrity checks of information—integrity in this context refers to how much the information can be relied upon (trusted). Mitigation strategies and redundancy measures, at both the system and operational levels, are standard outputs of the TEVV process. Because the TEVV process is being procedurally and temporally compressed with AI-enabled systems, the foresight that traditionally comes with these practices is now evidently missing. This raises questions around the extent to which these systems are truly fit for purpose.
Verification is the process of evaluating whether a system is fit for service against regulations and technical requirements and specifications, and validation determines whether it is fit for purpose against stakeholder needs. The former underpins technical integrity, while the latter supports customer needs and expectations. In this paper we have argued that upstream decision-making erodes technical integrity across the system lifecycle, and this erosion of technical integrity is evidently complicating fitness of purpose during system operation. Different decision-making upstream of the point of use can minimise, but not eliminate, the potential for system inaccuracies through technical interventions where feasible, and can mitigate the impacts of inevitable inaccuracies during operation through operational mitigation and redundancy measures. These opportunities for human intervention, through decision-making during the acquisition phase, can be narrowed and/or omitted, with consequences that manifest during the operational phase when these systems are deployed in targeting operations. The result is systems which are neither fit for service nor fit for purpose.
When implemented in targeting operations, such as the JTP, the accumulative degradation of decision-making is compounded by the absence of operational directives specific to the limitations of the system, which traditionally help to compensate for these limitations. The systemic and procedural nature of the JTP, be it in deliberate or dynamic targeting, may give the semblance of structure to targeting operations which employ AI-DSS; however, the reality of eroded technical integrity is the absence of consistency and reliability, which is evident in the contention between precision and accuracy.
This scenario highlights the intrinsic connection between the acquisition phase and the targeting process; how the contention between precision and accuracy at a technical level complicates the classification of systems as fit for purpose, fit for service, or neither; and the cascading impacts of that on operational effectiveness.
Explainability of System Outputs
The intended purpose of AI-DSS is to support, rather than replace, human decision-making.[78] [79] Ultimately, the decision to engage an identified target sits with a human, regardless of the system or systems deployed within the targeting process to support that decision. A critical part of decision-making is analysing and interpreting information, a task which has been absorbed by AI-DSS because of their capacity to analyse large amounts of information in rapid timeframes. While the analysis and interpretation of information may be delegated to AI-DSS, the requirement for interpreting system outputs to inform decision-making remains with humans.
It is difficult to interpret system outputs when the logic for how those outputs were produced is not clear or transparent.[80] The outputs of AI-DSS are not readily interpretable because the computational process for achieving those outputs is too complex, which is why these systems are labelled ‘black box’ systems.[81] This is where the contention between precision and accuracy further complicates operational effectiveness. Misinterpretation of information can result in inaccurate perceptions or understandings of a military operation, and imprecise or inaccurate AI-DSS present greater opportunities for misinterpretation, particularly when the opacity of these systems makes interrogation of their outputs difficult. If an AI-DSS is precise but inaccurate, it may produce the same incorrect output multiple times, leading to the assumption that the output is correct because the system consistently produces it. In the absence of explainability of the system output, detecting that inaccuracy may not be straightforward.
In addition to consistency and reliability, what technical integrity also enables is trust. In the context of this paper, trust is defined as confidence in the reliability of a system when used in the intended operation of use, a definition introduced in previous work by the authors.[82] While the topic of trust in the context of AI-enabled systems warrants its own paper, here we highlight one of the important elements of trust: trust calibration. Over-trust or under-trust in AI-enabled systems leads to ineffective outcomes, with a calibrated level of trust being the ideal middle ground.[83] When the contention between precision and accuracy is presented against an opaque AI-DSS in which the system outputs are not readily explainable, the result is miscalibrated trust. Over-trust leads to overconfidence in system outputs under the assumption the system is always correct,[84] and under-trust leads to underutilisation of system capabilities through misuse or dismissal of system outputs.[85]
Explainability of system outputs is a means of supporting calibrated levels of trust in AI-DSS, a necessary requirement when these systems are deployed in targeting operations. Over-trust in system outputs can encourage a de-emphasis on the need for human verification of targets, leading to humans acting as a superficial ‘rubber stamp’ during targeting operations.[86] Under-trust in system outputs may lead to the dismissal of identified targets or to missed ‘targets of opportunity’ in dynamic targeting operations.
Explainability of system outputs can be identified as a requirement of AI-DSS during the acquisition phase—noting that expectations around what that explainability looks like must align with technical feasibility. Additionally, robust TEVV practices upstream of the point of use would soften the contention between precision and accuracy and would enforce consistency and reliability in the system in support of technical integrity, resulting in systems which have a greater likelihood of promoting calibrated levels of trust during operation.
The Relationship Between Decision-Making and Time
There is a strong relationship between decision-making and time. The use of AI in the military domain has impacted this relationship by compressing both decision-making timeframes and the opportunities to make decisions. For the system lifecycle, decision-making is static in nature, while the targeting process involves dynamic decision-making. Static decision-making treats the decision task and process as fixed in time, while dynamic decision-making takes into account either the duration of decision-making, the optimal time to make a decision, or changes in the decision structure as a function of time.[87] The system lifecycle involves static decision-making because the sequence of decisions across the linearly progressing lifecycle phases are not subject to constant change, whereas targeting operations involve dynamic decisions resulting from the intrinsically dynamic nature of military operating environments.
The ongoing AI arms race has incited a ‘fabricated fear of falling behind’,[88] which has consequently accelerated the acquisition of AI-enabled military systems. The pace of planning, designing and developing engineered systems in the private technology sector has been condensed to promote technology adoption at unprecedented speed and scale. This compressed process narrows the opportunities for human intervention and critical decision-making windows. AI systems further exploit the factors of speed and scale through their capacity to analyse large datasets in rapid timeframes. When implemented into the targeting processes, this capacity for narrow decision-making timeframes shifts the relationship between decisions and time within targeting operations.
Compromised technical integrity is compounded by temporally compressed, high-stakes decision-making during targeting operations. Although contested, reporting on the IDF’s use of AI-DSS indicates that approximately 20 seconds is devoted to reviewing AI-DSS generated targets before engagement authorisations.[89] This limited timeframe calls into question the purpose of humans in this evolved targeting process. While dynamic targeting is designed to be at a higher tempo in comparison to deliberate targeting, the targeting process, as outlined in Table 1, still requires a minimum viable time to be completed. Accelerated operational decision-making not only has the potential to result in ineffective operational outcomes; it also disrupts broader geopolitical tensions as ‘speed optimised past the pace of human judgment is itself an escalation mechanism’[90]—an escalation mechanism hidden under the veneer of a data-driven approach.[91]
Military operations are often framed as time sensitive or ‘high tempo’;[92] hence the speed characterisation of AI-DSS fits within this framing as a positive intervention. Military operations involve cycles of decision-making, as is outlined in Section 2, and decision-making requires time, a requirement which challenges the purported benefit of AI-DSS. When these two time-dependent constructs are combined, balancing between pursuing speed and preserving time for decision-making becomes a difficult and precarious task. In the case of current deployments of AI-DSS in targeting operations, the scales are tipped in favour of accelerated decision-making, leading to compromised operational effectiveness, as was evident in the Minab school bombing.
Balancing decision-making and time is ultimately a trade-off between the benefits and shortcomings on either side. The inherent safety-critical nature of military operations adds an additional layer of complexity in the form of catastrophic outcomes when things go wrong. The nature of safety-critical operations, such as targeting operations, demand time and consideration in decision-making. However, for targeting operations, there may be circumstances in which that time and consideration cannot be afforded. The issue we are seeing with AI-DSS used in targeting operations in current conflicts is the exploitation of operational time sensitivity which encourages accelerated decision-making as a standard practice. Accelerated decision-making aided by AI-DSS that lack appropriate human oversight and thorough technical integrity may be exposing military forces to inconsistent and unreliable systems informing decisions of life and death.
Recommendations
Recommendation 1: Technical integrity should be a mandatory requirement of AI-DSS intended for use in targeting operations, with evidence of ‘fitness for service’ demonstrated upstream of the point of use.
In this work we have argued that AI-DSS employed for targeting operations are at risk of being fit for purpose but not fit for service. Goussac and Boulanin frame procurement as the mechanism that serves to implement political commitments and legal obligations. Contracting within the acquisition stage is highlighted as:
a critical opportunity for implementing principles of responsible behaviour because it determines what the supplier must deliver and prove (including claims about the results of testing and assurance processes), who is accountable for certain risks and failures, and what obligations and requirements are borne by the parties.
While the Policy Settings for Responsible Use of AI in Defence establish a values-based architecture, we advocate for further technical integrity depth to meaningfully operationalise the policy. This may include the development of fit-for-service evidence, such as contextual TEVV under realistic conditions to draw out system friction and fracture points, draw out failure modes and establish robust mitigation strategies. The caution of Conn and Bode that ‘if the use of the system seems to turn human users into button pushers and therefore risks only securing nominal human control … this needs to be flagged in earlier stages during testing’ applies directly: fitness-for-purpose performance is not a substitute for evidence of fitness for service.
Recommendation 2: Speed of operational decision-making should be assessed against operational effectiveness, which should include IHL obligations.
High-tempo operations do not always lead to operationally effective outcomes. Accordingly, the potential advantages of accelerated operational decision-making should be assessed against operational effectiveness. In this paper, we frame operational effectiveness as verification leading to protection, including the safety of non-combatants and combatants, consistent with the IHL principles of distinction (distinguishing between combatants and non-combatants) and proportionality (ensuring that disproportionate impacts on civilians during military operations are minimised).
Recommendation 3: AI-DSS operational effectiveness for targeting operations should be quantitatively verified against established benchmarks and validated against the commander’s intent prior to deployment.
A complete operational effectiveness assessment for AI-DSS in targeting operations should evaluate performance against four dimensions: mission achievement, force preservation, operational tempo and IHL compliance. Validation against a commander’s intent provides the standard against which trade-offs between these dimensions are resolved, and retains human command and operator engagement throughout critical phases of the targeting cycle. Consideration should be given to the design and interface of AI-DSS to support human engagement, negating potentially irreversible ‘automatic acceptance’ of outputs.[93]
Conclusion
This paper has examined the lifecycle of AI-DSS deployed in military targeting, exploring failure modes evident at the point of use resulting from upstream decisions in acquisition and capability development. The omission of critical military processes results in an accumulative degradation of decision-making across the lifecycle, eroding technical integrity of the system and compromising operational effectiveness in the targeting operations for which such systems are deployed. We have explored the intersection of acquisition and operational processes through a use-case based approach to highlight potential pitfalls and challenges that arise from the integration of AI-DSS. The forward-looking question is not whether AI-DSS will be deployed in targeting; these systems are already deployed. This paper highlights the compounding impacts of narrowed and omitted critical military processes across the system lifecycle and how these manifest when systems are utilised in military operations, specifically targeting operations. The question is whether the institutional governance architecture of acquisition, command accountability and operational TEVV will adapt to evolving military procurement processes.
Endnotes
[1] Michael Biesecker, Sam Mednick and Garance Burke, ‘As Israel Uses US-Made AI Models in War, Concerns Arise About Tech’s Role in Who Lives and Who Dies’, AP News, 18 February 2025.
[2] Alexander Blanchard and Laura Brunn, Autonomous Weapon Systems and AI-Enabled Decision Support Systems in Military Targeting: A Comparison and Recommended Policy Responses (Stockholm International Peace Research Institute, 2025).
[3] Katrina Manson, ‘US Military Relying on AI as Tool to Speed Iran Operations’, Bloomberg, 5 March 2026.
[4] ‘Anthropic’s AI Tool Claude Central to US Campaign in Iran, Amid a Bitter Feud’, The Washington Post, 4 March 2026.
[5] AD Larson and CC Hayes, ‘An Assessment of Weasel: A Decision Support System to Assist in Military Planning’, Proceedings of the Human Factors and Ergonomics Society 49th Annual Meeting 49, no. 3 (2005): 287–291.
[6] ‘Accelerating Decision Making: National Geospatial-Intelligence Agency at AIPCon 5’ (video), Palantir YouTube channel, 19 September 2024, at: www.youtube.com/watch?v=XzKnUt6NAbw.
[7] Anna Nadibaidze, Ingvild Bode and Qiaochu Zhang, AI in Military Decision Support Systems: A Review of Developments and Debates (Odense: Center for War Studies, 2024).
[8] Human Rights Watch, ‘Questions and Answers: Israeli Military’s Use of Digital Tools in Gaza’, Human Rights Watch (website), at: www.hrw.org/news/2024/09/10/questions-and-answers-israeli-militarys-use-digital-tools-gaza.
[9] Gaby Tejeda, ‘AI Integration in Operation Epic Fury and Cascading Effects’, The Soufan Center (website), 3 March 2026, at: https://thesoufancenter.org/intelbrief-2026-march-3.
[10] Elizabeth Dwoskin, ‘How Israel Built an “AI Factory” for War, Use in Gaza’, The Washington Post, 29 December 2024.
[11] David Hollingworth, ‘Machine War: How AI and Other Technologies Are Shaping the Fighting in Gaza’, Defence Connect, at: www.defenceconnect.com.au/joint-capabilities/16182-machine-war-how-ai-and-other-technologies-are-shaping-the-fighting-in-gaza.
[12] ‘Reported Impact Snapshot | Gaza Strip (20 August 2025)’, United Nations Office for the Coordination of Humanitarian Affairs (website), at: www.ochaopt.org/content/reported-impact-snapshot-gaza-strip-20-august-2025.
[13] Noa Yachot, ‘“Data Is Control”: What We Learned from a Year Investigating the Israeli Military’s Ties to Big Tech’, The Guardian, 31 December 2025.
[14] Department of Defence, Policy Settings for Responsible Use of Artificial Intelligence in Defence (Canberra: Commonwealth of Australia, 2025).
[15] Netta Goussac and Vincent Boulanin, Responsible Procurement of Military Artificial Intelligence (Stockholm: Stockholm International Peace Research Institute, 2026).
[16] Conn and Bode define the term AIS as ‘artificial intelligence system’.
[17] A Conn and I Bode, ‘Establishing Human Responsibility and Accountability at Early Stages of the Lifecycle for AI-Based Defence Systems’, Ethics and Information Technology 27, no. 51 (2025).
[18] Jessica Dorsey and Marta Bo, ‘AI-Enabled Decision-Support Systems in the Joint Targeting Cycle: Legal Challenges, Risks, and the Human(e) Dimension’, International Law Studies 107 (2026): 184–229.
[19] A recent study by the Boston Consulting Group, published in the Harvard Business Review, outlines implications of ‘AI brain fry’ (AI mental fatigue). The report highlights the capacity limit of human biological brains, noting that the strongest predictor of mental fatigue was the level of oversight required for the AI agent (i.e. greater intimate oversight leads to greater levels of fatigue). See Julie Bedard, Matthew Kropp, Megan Hsu, Olivia T Karaman, Jason Hawes and Gabriella Rosen Kellerman, ‘When Using AI Leads to “Brain Fry”’, Harvard Business Review, 5 March 2026.
[20] Goussac and Boulanin, Responsible Procurement.
[21] Alexandre Verlaine, ‘No-Capability Defence Acquisition: A Literature Review on Policy and Practice of Military Acquisition Within NATO and the EU’, International Journal of Procurement Management 17, no. 4 (2023): 443–487.
[22] Desmond Ball, ‘Arms and Affluence: Military Acquisitions in the Asia-Pacific Region’, International Security 18, no. 3 (1993): 78–112.
[23] W Perry, ‘US Military Acquisition Policy’, Comparative Strategy 13, no. 1 (1994): 19–24.
[24] Cecilia Haskins, Kevin Forsberg and Michael Krueger, Systems Engineering Handbook: A Guide for System Life Cycle Processes and Activities, version 3.1 (International Council on Systems Engineering, 2007).
[25] I Faulconbridge and MJ Ryan, Systems Engineering Practice (Argos Press, 2014).
[26] BS Blanchard and WJ Fabrycky, Instructor’s Guide to Problem Solutions for Systems Engineering and Analysis, fifth edition (Pearson, 2011).
[27] JR Nelson, P Konoske Dey, MR Fiorello, JR Gebman, GK Smith and A Sweetland, A Weapon-System Life-Cycle Overview: The A-7D Experience, R-1452-PR (RAND, 1974).
[28] Z Assaad and A Hepworth, A Systems Engineering Lifecycle Approach to Responsible AI, (The Hague Centre for Strategic Studies, 2025).
[29] S Allik et al., A Framework for Human Decision Making Through the Lifecycle of Autonomous and Intelligent Systems in Defense Applications (New York: IEEE SA Research Group on Issues of Autonomy and AI in Defense Systems, 2024).
[30] Faulconbridge and Ryan, Systems Engineering Practice.
[31] RP Smith, Defence Acquisition and Procurement: How (Not) to Buy Weapons (Cambridge University Press, 2022).
[32] Australian Defence Force Warfare Training Centre, Operations Series ADDP 3.14 Targeting (Australian Defence Doctrine Publications, Defence Publishing Service, 2009)
[33] Mitch Ferry, ‘F3EA: A Targeting Paradigm for Contemporary Warfare’, Australian Army Journal 10, no. 1 (2013).
[34] Australian Defence Force Warfare Training Centre, Operations Series ADDP 3.14.
[35] U.S. Department of Defense, DOD Dictionary of Military and Associated Terms 232, Joint Publication 1-02 (Department of Defense, 2017).
[36] Allied Joint Doctrine for Joint Targeting, AJP-3.9 (NATO, 2008).
[37] Australian Defence Force Warfare Centre, Operations Series ADDP 3.14.
[38] Dorsey and Bo, ‘AI-Enabled Decision-Support Systems in the Joint Targeting Cycle’, pp. 184–229.
[39] Conn and Bode, ‘Establishing Human Responsibility and Accountability at Early Stages of the Lifecycle for AI-Based Defence Systems’; Allik et al., A Framework for Human Decision Making Through the Lifecycle of Autonomous and Intelligent Systems in Defense Applications.
[40] MT Edwards, Measuring the Technical Integrity of a Complex Engineered System (University of South Australia, 2014).
[41] Heidy Khlaaf and Sarah Myers West, ‘Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds’, arXiv Computers and Society, arxiv:2504.15088 (2025).
[42] Michael Edwards, ‘7.3.1 The Impact of Technical Regulation on the Technical Integrity of Complex Engineered Systems’, INCOSE International Symposium 19, no. 1 (2014): 1139–1153.
[43] Z Assaad, ‘A Framework for Safely Scaling Multi-Agent Teams’, Australian Army Journal 21, no. 2 (2025): 132–158.
[44] Dorsey and Bo, ‘AI-Enabled Decision-Support Systems in the Joint Targeting Cycle’, p.187.
[45] For instance, see Allik et al., A Framework for Human Decision Making Through the Lifecycle of Autonomous and Intelligent Systems in Defense Applications, p. 18; Conn and Bode, ‘Establishing Human Responsibility and Accountability at Early Stages of the Lifecycle for AI-Based Defence Systems’; Dorsey and Bo, ‘AI-Enabled Decision-Support Systems in the Joint Targeting Cycle’, p. 226;
Lewis, Michael. 2013. “Human Interaction with Multiple Remote Robots.” Reviews of Human Factors and Ergonomics 9 (1): 131–74. https://doi.org/10.1177/1557234X13506688.
[46] Khlaaf and Myers West, ‘Safety Co-Option and Compromised National Security’.
[47] We acknowledge that the information environment surrounding this conflict is heavily contaminated with AI-generated claims (fabrications); hence this paper attributes contested claims to identified sources rather than as assertions.
[48] Yuval Abraham, ‘“Lavender”: The AI Machine Directing Israel’s Bombing Spree in Gaza’, +972 Magazine, 3 April 2024.
[49] Lauren Gould, Linde Arentze and Marijn Hoijtink, ‘Gaza War: Artificial Intelligence Is Changing the Speed of Targeting and Scale of Civilian Harm in Unprecedented Ways’, The Conversation, 24 April 2024.
[50] Elizabeth Dwoskin, ‘How Israel Built an “AI Factory” for War, Use in Gaza’, The Washington Post, 29 December 2024.
[51] Harry Davies, Bethan McKernan and Dan Sabbagh, ‘“The Gospel”: How Israel Uses AI to Select Bombing Targets in Gaza’, The Guardian, 1 December 2023.
[52] Z Assaad and E Williams, ‘Technology and Tactics: The Intersection of Safety, AI, and the Resort to Force’, Cambridge Forum on AI: Law and Governance 1, no. 49 (2025).
[53] Neta Crawford, ‘Gaza: Civilian Death Toll Outpaces Other Modern Wars’, The Conversation, 27 August 2025.
[54] ‘Reported Impact Snapshot: Gaza Strip (20 August 2025)’, United Nations Office for the Coordination of Humanitarian Affairs (website), at: www.ochaopt.org/content/reported-impact-snapshot-gaza-strip-20-august-2025.
[55] Ozer Khalid, ‘Gaza: Ground Zero for AI Warfare’, Criterion Quarterly 19, no. 4 (2025).
[56] Yachot, ‘“Data Is Control”’.
[57] ‘Israel and the Occupied Palestinian Territory’, Amnesty International (website), at: www.amnesty.org/en/location/middle-east-and-north-africa/middle-east/israel-and-the-occupied-palestinian-territory/report-israel-and-the-occupied-palestinian-territory.
[58] Kevin T Baker, ‘AI Got the Blame for the Iran School Bombing. The Truth Is Far More Worrying’, The Guardian, 26 March 2026.
[59] ‘USA/Iran: Those Responsible for Deadly and Unlawful US Strike on School That Killed Over 100 Children Must Be Held Accountable’, Amnesty International (website), 16 March 2026, at: www.amnesty.org/en/latest/news/2026/03/usa-iran-those-responsible-for-deadly-and-unlawful-us-strike-on-school-that-killed-over-100-children-must-be-held-accountable.
[60] Baker, ‘AI Got the Blame for the Iran School Bombing’.
[61] Jon Harper, ‘Centcom Commander Touts Use of AI in Fight Against Iran During Operation Epic Fury’, DefenseScoop, at: https://defensescoop.com/2026/03/11/us-military-using-ai-against-iran-operation-epic-fury-adm-cooper.
[62] Baker, ‘AI Got the Blame for the Iran School Bombing’.
[63] Katrina Manson, ‘An Analyst’s Missed Remark Surfaced in Deadly Iran School Strike Probe’, Bloomberg, 26 June 2026.
[64] J Floch, C Carrez, P Cieslak, M Roj, RT Sanders and MM Shiaa, ‘A Comprehensive Engineering Framework for Guaranteeing Component Compatibility’, Journal of Systems and Software 83, no. 10 (2010): 1759–1779.
[65] Mikela Chatzimichailidou, Tim Whitcher and Nikola Suzic, ‘Complementarity and Compatibility of Systems Integration and Building Information Management’, IEEE Systems Journal 18, no. 2 (2024): 1198–1207.
[66] Baker, ‘AI Got the Blame for the Iran School Bombing’.
[67] Tejeda, ‘AI Integration in Operation Epic Fury and Cascading Effects’.
[68] Manson, ‘US Military Relying on AI as Tool to Speed Iran Operations’.
[69] V Bergengruen, ‘How Tech Giants Turned Ukraine Into an AI War Lab’, Time, 9 February 2024.
[70] Ibid.
[71] Ibid.
[72] Y Mysyshyn, Advanced Technologies in the War in Ukraine: Risks for Democracy and Human Rights (Washington DC: German Marshall Fund of the United States, 2024).
[73] Ibid.
[74] Ronald F Boisvert, Ronald Cools and Bo Einarsson, ‘Assessment of Accuracy and Reliability’, in Bo Einarsson (ed.), Accuracy and Reliability in Scientific Computing (Society for Industrial and Applied Mathematics, 2005).
[75] Assaad and Williams, ‘Technology and Tactics’.
[76] GG Whitchurch and LL Constantine, ‘Systems Theory’, in PG Boss, WJ Doherty, R LaRossa, WR Schumm and SK Steinmetz (eds), Sourcebook of Family Theories and Methods: A Contextual Approach (New York: Plenum Press, 1993).
[77] D Saidulu and R Sasikala, ‘Machine Learning and Statistical Approaches for Big Data: Issues, Challenges and Research Directions’, International Journal of Applied Engineering Research 12, no. 21 (2017): 11691–11699.
[78] AD Larson and CC Hayes, ‘An Assessment of Weasel: A Decision Support System to Assist in Military Planning’, Proceedings of the Human Factors and Ergonomics Society 49, no. 3 (2005): 287–291.
[79] Harper, ‘Centcom Commander Touts Use of AI in Fight against Iran during Operation Epic Fury’.
[80] Assaad and Williams, ‘Technology and Tactics’.
[81] C Nugent and P Cunningham, ‘A Case-Based Explanation System for Black-Box Systems’, Artificial Intelligence Review 24 (2005): 163–178.
[82] Zena Assaad and Christine Boshuijzen-van Burken, ‘Ethics and Safety of Human-Machine Teaming’, TAS ’23: Proceedings of the First International Symposium on Trustworthy Autonomous Systems (Association for Computing Machinery, 2023), pp. 1–8.
[83] M Wischnewski, N Krämer and E Müller, ‘Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-of-the-Art and Future Directions’, Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Association for Computing Machinery, 2023), pp. 1–16.
[84] Alexander M Aroyo et al., Overtrusting Robots: Setting a Research Agenda to Mitigate Overtrust in Automation’, Journal of Behavioural Robotics 12, no. 1 (2021): 423–436.
[85] Daniel Ullrich, Andreas Butz and Sarah Diefenbach, ‘The Development of Overtrust: An Empirical Simulation and Psychological Analysis in the Context of Human–Robot Interaction’, Frontiers in Robotics and AI 8 (2021).
[86] Dorsey and Bo, ‘AI-Enabled Decision-Support Systems in the Joint Targeting Cycle’, p. 187.
[87] Dan Ariely and Dan Zakay, ‘A Timely Account of the Role of Duration in Decision Making’, Acta Psychologica 108 (2001): 187–207.
[88] Zena Assaad, ‘Artificial Urgency: Reflecting on AI Hype at the 2026 REAIM Summit’, Just Security, at: www.justsecurity.org/132504/ai-hype-2026-reaim-summit.
[89] Michael Biesecker, Sam Mednick and Garance Burke, ‘As Israel Uses US-made AI Models in War, Concerns Arise About Tech’s Role in Who Lives and Who Dies’, AP News, 18 February 2025.
[90] This statement was made by Husham Ahmed, Counsellor at the Permanent Mission of Pakistan to the UN in Geneva, during the informal exchanges on AI in the military domain convened by the UN Office for Disarmament Affairs (UNODA) from 15 to 17 June 2026 in Geneva. Formal permission was granted for attributing this statement to Husham Ahmed.
[91] Yachot, ‘“Data Is Control”, The Guardian, 31 December 2025.
[92] K Klonowska and TK Woodcock, ‘Rhetoric and Regulation: The (Limits of) Human/AI Comparison in Legal Debates on Military AI’, in B Boutin, TK Woodcock and S Soltanzadeh (eds), Legal, Ethical, and Technical Dilemmas in Military Artificial Intelligence (The Hague: MC Asser Press, 2026).
[93] Hyunwoo Kim, Harin Yu and Hanau Yi, ‘The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows’, arXiv, arXiv:2604.14807, at: https://doi.org/10.48550/arXiv.2604.14807.