Executive Guide
Testing and Evaluation of Agentic AI Systems In Military Command and Control
- Author
- Aziz Shuaib Ausi
- Published
- 28 August 2026
- Reading time
- 1 min
- Publication type
- Executive Guide
- Availability
- Open access
Executive Summary
Research identifies significant challenges in the testing and evaluation (T&E) of agentic AI systems for military command and control (C2). Current T&E methodologies, which rely on assumptions about system specifiability, stability, composability, and supervisability, are undermined by the inherent properties of agentic AI. This weakness in the assurance case for these systems raises concerns about the feasibility of fulfilling public commitments to rigorous testing and human oversight.
Research identifies significant challenges in the testing and evaluation (T&E) of agentic AI systems for military command and control (C2). Current T&E methodologies, which rely on assumptions about system specifiability, stability, composability, and supervisability, are undermined by the inherent properties of agentic AI. This weakness in the assurance case for these systems raises concerns about the feasibility of fulfilling public commitments to rigorous testing and human oversight.
Why it matters
The adoption of agentic AI in critical domains like military command and control is a significant technological and operational shift. Challenges in effectively testing and evaluating these systems could lead to unforeseen risks, operational failures, and undermine public trust in AI-driven decision-making, impacting national security and strategic stability.
Key insights
- Agentic AI systems are being acquired for military C2, with public commitments to rigorous testing and human oversight.
- The ability to meet these commitments depends on a robust assurance case, comprising claims, evidence, and argument.
- A review of 240 documented T&E practices identified eight critical assumptions made by established methods regarding their test article.
- These assumptions fall into four clusters: system specifiability, stability, composability, and supervisability.
- The inherent properties of agentic AI systems weaken all eight identified assumptions, complicating their rigorous testing.
- The integrity of the assurance case for agentic AI in C2 is compromised by these weakened assumptions.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2608.20597
Related publications
Previous
Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI
Next
Who Delegates to AI? Evidence from 53,000 Agent Configurations
Embedding inter- and transdisciplinary sustainability skills and knowledge development in higher education: perspectives from an innovative new degree
Executive Guide
Critical thinking as a predictor of task functionality and artificial intelligence use among university students. A PLS-SEM approach
Executive Guide
Cognitive emotion regulation as a statistical mediator of the association between autistic traits and academic performance in university students
Executive Guide
AI self-efficacy as a predictor of satisfaction with studies: the mediating role of research motivation among Peruvian University students
Executive Guide
Generative AI and linguistic creativity in digitally multilingual higher education
Executive Guide
Digital teaching and learning strategies for enhancing self-directed learning in remote ODeL environments: evidence from Zimbabwe Open University
Executive Guide
Download & citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). Testing and Evaluation of Agentic AI Systems In Military Command and Control. Executive Guide. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXG-2026-00652
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXG-2026-00652
- Version
- v1.0 · r0
- Issued
- 28 August 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.