Back to Blog
AI Governance

Study: Soldiers Trust AI Targeting Less Than Humans — Until You Explain It

A 2,015-participant experiment with a replica Israeli military AI targeting system found algorithmic aversion, not automation bias — but adding explainability features erased that skepticism entirely.

PyramidLedger Research4 min read
Share

Key Takeaways

  • Researchers built a high-fidelity replica of a real Israeli military AI decision-support system and tested it on 2,015 active-duty and veteran personnel across two studies.
  • Participants approved 63% of AI-attributed strike recommendations versus 68% of identical recommendations attributed to human analysts (p < 0.001) — algorithmic aversion, not automation bias.
  • The skepticism gap widened to 7.3 percentage points in high-collateral-damage scenarios but vanished in low-risk ones.
  • Adding three data-point explanations to the AI's output eliminated the aversion gap entirely (67% vs 65%, not statistically significant) — a result that cuts both ways for anyone relying on 'human in the loop' as a safeguard.

Most policy discussion about AI in high-stakes decision-making assumes the risk runs one direction: give operators a confident-sounding algorithm and they will defer to it, oversight becomes theatre, and "human in the loop" turns into rubber-stamping. A new peer-reviewed study, flagged by Bruce Schneier, tests that assumption directly — and finds the opposite, at least under one set of conditions.

What the study actually did

Researchers Ryan Shandler, Michael L. Gross and Yahli Shereshevsky, publishing in the *Journal of Conflict Resolution*, reconstructed the interface and functionality of a real Israeli Defense Forces AI targeting decision-support system, drawing on open-source materials, FOIA requests, and military interviews. They ran two experiments totaling 2,015 active-duty and veteran Israeli soldiers, presenting 12 targeting scenarios — target images, attack methods, success ratings, collateral-damage estimates, confidence scores — with each scenario randomly attributed to either an AI system or a human analyst.

The dependent measures were strike-approval rate, decision confidence, and response time — a clean design for isolating how the *source label* alone shifts human judgment, independent of the underlying recommendation.

Aversion, not bias — but conditional

Participants approved 63% of AI-attributed recommendations versus 68% of identical recommendations attributed to a human analyst (p < 0.001), with lower stated confidence (3.91 vs 4.00 on a 5-point scale). This is algorithmic aversion: people trusted the machine less than an equivalent human source, not more. The effect wasn't uniform — in high-collateral-damage scenarios the approval gap widened to 7.3 percentage points, while low-risk scenarios showed no significant difference. In other words, skepticism of AI scaled with the stakes, which is roughly the behavior a governance framework would want.

The part that should worry oversight designers

The second study added a simple explainability layer — three data points showing what the AI's recommendation was based on. That single change was enough to eliminate the aversion gap entirely (67% vs 65%, no longer statistically significant). Explanation didn't just build appropriate trust; it fully closed the gap between how people treated machine and human recommendations, including in the high-stakes scenarios where healthy skepticism had been strongest.

That is a genuinely double-edged result. Explainable AI is a stated goal of frameworks like ISO/IEC 42001 and the EU AI Act's human-oversight provisions, and this data shows explanation features do change behavior — operators engaged more deliberately with the AI's reasoning rather than reflexively discounting it. But the same data shows explanation can also flatten the very skepticism that acted as a brake in the riskiest scenarios. A three-data-point summary is a thin basis on which to fully re-equalize trust in a collateral-damage decision, and the study doesn't tell us whether that trust was better *calibrated* afterward or simply less resistant.

Why this matters beyond the battlefield

The same dynamic applies to any organization deploying AI decision-support in consequential settings — SOC alert triage, fraud holds, medical or safety-critical recommendations. "Human in the loop" is frequently treated as a checkbox control, but this study is rare empirical evidence that the loop can function as intended — and that interface design, specifically how much explanation is surfaced and how, is what determines whether it does. Auditing an AI system's human-oversight control shouldn't stop at confirming a person can override the model; it should test whether the interface pushes that person toward genuine judgment or toward automatic concurrence.

Frequently Asked Questions

Did the study find automation bias in military AI use?

No. It found the opposite — algorithmic aversion. Soldiers approved AI-attributed strike recommendations at a lower rate (63%) than identical recommendations attributed to a human analyst (68%), with the gap growing in high-collateral-damage scenarios.

Does adding explainability to AI systems make oversight better or worse?

The study shows it changes behavior significantly, closing the trust gap between AI and human sources — but it doesn't establish whether the resulting trust was better calibrated or simply less resistant, which is the open question for anyone designing human-oversight controls around explainable AI.

How was the study conducted?

Researchers built a high-fidelity replica of a real Israeli military AI targeting decision-support system and tested it on 2,015 active-duty and veteran Israeli soldiers across two experiments, randomly attributing identical recommendations to either AI or human analysts.

Sources

  1. 1AI for Military SupportSchneier on Security
  2. 2Black Box Warfare: Human Judgment and Military Decision-Making in the Age of AIJournal of Conflict Resolution
Share

Read next