Effective Intelligence, Surveillance, and Reconnaissance (ISR) in contested theaters depends on the ability to extract and transmit decision-relevant information from sensors under severely constrained communications. Modern multimodal sensor payloads, including electro-optical (EO), infrared (IR), event-camera, low-light, and platform-telemetry streams, produce data volumes orders of magnitude beyond what tactical communications links can carry, particularly when those links are degraded, intermittent, jammed, emission-constrained, or forced to single-digit kilobits per second by electronic warfare. The operational consequence is binary: an ISR platform either delivers mission-relevant evidence to the operator under these conditions, or it provides limited operational value.
Small unmanned aircraft systems (sUAS) embody this challenge in its most acute form. Representative Group 1 platforms such as RQ-11B Raven and RQ-20 Puma-class systems carry EO and IR video for dismounted operators with single-digit-watt payload compute budgets. Representative Group 2 platforms such as ScanEagle-class systems provide broader reconnaissance with greater range, endurance, and modular sensors, but remain size, weight, and power (SWaP) constrained. Across both classes, full-motion EO and IR video routinely exceeds available tactical link capacity, particularly when multiple unmanned systems operate simultaneously.
Three concurrent technical failure modes define this problem space and rule out incremental approaches:
• Mission intent is vague and operator-defined. Operators do not request a single class label. They request answers to compositional, context-dependent tasks such as identifying indicators of hostile staging along a ridgeline, judging whether a gathering is consistent with prior pattern-of-life, or noticing whether access patterns to a structure have changed in the last twenty minutes. There is no fixed object dictionary, and what is mission-relevant in one engagement is clutter in the next.
• Regions of interest are not necessarily objects. Mission-relevant regions are frequently not detectable as discrete objects in any single frame. Relevant evidence may be a vehicle stopped where vehicles do not normally stop, a crowd dispersing unusually fast, a heat signature appearing at an unusual time, a person loitering near an asset, or a pattern of vegetation disturbance accumulated over many frames. Object detectors and large vision-language models are weak in this regime because the evidence is distributed across space and time rather than localized in a single frame.
• The platform cannot host heavy artificial intelligence. Group 1 payload compute is typically constrained to single-digit watts; Group 2 platforms allow modest growth but remain size, weight, and power constrained. Large vision-language models and conventional deep-learning trackers are not deployable at this power envelope and are additionally fragile under domain shift, adversarial obscuration, and data-poor conditions characteristic of low-density and low-prior threats.
This topic seeks a new class of mission-aware semantic communication systems that can approach two orders of magnitude data reduction in mission-suitable ISR scenarios while preserving mission-relevant context and semantic content. The required capability is not video compression, object detection, or target tracking; it is an onboard, real-time, low-power information-processing layer that understands operator intent, reasons over spatial and temporal context, captures behavioral evidence distributed across frames, and emits only the compact semantic packets needed for receiver-side reconstruction. Semantic packets shall combine high-fidelity regions of interest, tracklets, scene-change descriptors, temporal evidence, low-rate background context, confidence, uncertainty, and traceable explanations sufficient for the operator to act on the evidence at substantially lower bandwidth.
Phase II prototypes shall demonstrate the following capabilities:
• Mission-intent grounding. Translate vague natural-language or structured operator objectives into compact, executable mission cards that identify relevant cues, contextual relationships, temporal patterns, negative evidence, and uncertainty thresholds. Support operator updates over a constrained link without mission-specific retraining during flight.
• Open-world spatiotemporal reasoning with adaptive ROI coding. Identify regions, events, and behaviors that are mission-relevant even when object categories are unknown, low-resolution, occluded, or camouflaged, by reasoning over multiple frames, tracks, scene context, motion fields, and temporal change rather than per-frame detections. Allocate bits dynamically among high-fidelity ROI patches, event summaries, background context, descriptors, and metadata such that the transmitted payload is reconstructable into operator-usable products at the receiver.
• Multimodal semantic fusion. Extend reasoning beyond single-modality EO video. By end of Phase II, demonstrate EO video plus at least one additional modality or metadata source: long-wave infrared, low-light video, event-camera streams, platform telemetry, gimbal or inertial state, acoustic cueing, or RF or link-state metadata.
• Traceability and operator trust. Each semantic transmission decision shall include a compact evidence chain comprising mission clause, cue type, temporal support, spatial support, confidence, and uncertainty. The receiver shall enable operator audit of why information was transmitted, suppressed, or summarized. This is a hard requirement intended to distinguish the deliverable from black-box neural compression.
• Low-SWaP embedded execution. Operate on edge hardware compatible with small UAS payload constraints. Threshold = 5 W incremental power during interim testing; final objective = 2 W average incremental power excluding camera and radio. Implementations may use quantized neural networks, vector symbolic or hyperdimensional accelerators, FPGA, low-power NPU, microcontroller co-processing, neuromorphic devices, or custom hardware-software co-design. Representative embedded targets include the NVIDIA Jetson Orin Nano in low-power profiles, Hailo-8L, Coral Edge TPU, AMD Kintex UltraScale+ FPGA fabrics, and ARM Cortex-M-class microcontrollers for symbolic primitive execution.
• Robustness under contested operation. Maintain performance under motion blur, compression artifacts, clutter, low contrast, sensor noise, partial occlusion, camera motion, lighting change, intermittent links, packet loss, and adversarially confusing backgrounds. Robustness shall be characterized quantitatively, not asserted.
• Cyber posture and supply chain integrity. Deliver a software bill of materials, secure update path, documented interfaces, and cybersecurity assessment suitable for transition review. The deployable system shall not introduce avoidable vulnerabilities into tactical networks or mission systems.
Approaches of interest may combine lightweight neural perception with mathematically structured spatiotemporal reasoning, including vector symbolic architectures, hyperdimensional computing, temporal graphs, probabilistic scene memory, causal or relational models, neurosymbolic methods, or other compact online-reasoning approaches. Such methods are relevant because they can support compositional representation, efficient similarity search, noise tolerance, online learning, and traceable reasoning under tight power budgets. The Government does not prescribe a specific algorithmic stack; proposers may employ any combination of techniques meeting the performance, traceability, robustness, and SWaP requirements.
Approaches not of interest include solutions relying solely on conventional video compression without mission-aware reasoning; solutions relying solely on object detection or tracking with a fixed taxonomy; solutions relying solely on large vision-language models that cannot execute onboard; approaches requiring cloud processing, continuous high-bandwidth reachback, or extensive per-mission retraining during flight; and approaches requiring replacement of existing UAS flight controllers, autopilots, radios, or ground-control stations. Weapon-release, lethal engagement, and named-person identification capabilities are explicitly excluded from this topic.
The deliverable system shall not autonomously make lethal engagement decisions, shall not identify named persons, and shall not serve as a weapon-release authority. The intended use is bandwidth-efficient ISR sensing, operator cueing, and semantic preservation under constrained communications