Tuesday, August 4, 2026

 The drone video sensing and analytics software market is attractive but structurally competitive, with moderate-to-high rivalry, non trivial entry barriers, and strong pressure to differentiate through vertical focus, data network effects, and AI quality-of-service guarantees. [fortunebusinessinsights]

Below is a Porter’s Five Forces industry analysis tailored to the companies I listed (AuterionOS, Scale AI, Rhoda.AI, AirSentinel.AI, GeneralAgents.AI, FlyPix.AI, NineTen Drones, SkyWays Drones, Agrositech, SkyFoundry, Aurora Flight Sciences, Archer Aviation, Sentinel AI, Cirium, Cireon, etc.), plus adjacent players, with a lens on launching a drone video sensing software firm.

1. Competitive rivalry (current players)

Rivalry is high: there is a crowded field spanning pure-play drone analytics, broader video analytics, and aviation intelligence platforms.[marketresearchfuture] 

Key clusters of competitors:

• Geospatial / drone-native analytics

o FlyPix.AI focuses on geospatial image analysis and change detection with SaaS and data-as-a-service, directly targeting multi-resolution aerial imagery use cases.[flypix]

o Agrositech and similar agriculture-focused firms deliver crop health and field metrics, often with drone-centric workflows.[financialmodelslab] 

o NineTen Drones and Cireon run mission-based inspection and survey reconstitution services, packaging analytics with flight operations. 

• Aviation intelligence and UAS platforms

o Aurora Flight Sciences (Boeing) offers advanced autonomy, ISR (intelligence, surveillance, reconnaissance) solutions, and small UAS platforms, embedding analytics into mission systems.[aurora] 

o Cirium provides aviation data analytics and fleet intelligence, including UAV navigation and services (UTM providers, low-altitude navigation), with rich spatial–temporal analytics.[cirium] 

o Archer Aviation is building eVTOL operations with fleet telemetry and data monetization, emphasizing cloud-scale telemetry observability. 

• Drone software and data platforms

o Rhoda.AI provides mission management and automated pipeline orchestration for enterprise drone programs as SaaS, with strong emphasis on multi-resolution analytics and agentic retrieval. 

o SkyFoundry comes from IoT/building analytics and is integrating drone data as an additional telemetry channel. 

o Auterion and AuterionOS (noted in other sources) deliver open-source PX4-based drone OS and cloud services; AuterionOS is effectively an ecosystem platform that can integrate analytics partners.[360iresearch]

• Security and airspace sensing

o AirSentinel.AI and Sentinel AI focus on real-time UAV threat detection, perimeter security, and intrusion detection, bundling SaaS and hardware (sensors, cameras, RF equipment).[researchdive] 

o GeneralAgents.AI targets agentic task orchestration with QoS and token metering, plugging AI workflows into drone operations and broader enterprise pipelines. 

• Horizontal video analytics / AI infra players

o Scale AI, and similar AI data platforms, offer labeling, data curation, and foundation model infrastructure that can be adapted to drone video streams but are not drone-exclusive.[prnewswire]

o Broader video analytics vendors in the global market (CAGR ~20.4%, reaching around USD 14.9B by 2026) increase rivalry by offering generic video AI applied to drone feeds.[prnewswire]

2. Threat of new entrants

For a drone video sensing software firm, the threat of new entrants is moderate to low: it is easy to start a demo product, but hard to reach enterprise-grade scale with trust, certifications, and data network effects.[studocu]

Entry barriers:

• Technical and infrastructure barriers

o Robust pipelines for high-volume video ingestion, multi-resolution analytics, and spatial–temporal modeling require significant cloud infra, GPU resources, and MLOps maturity.[marketresearchfuture] 

o Enterprise-grade QoS (SLAs for latency, uptime, and cost predictability) demands sophisticated rate limiting, admission control, and telemetry observability, like what GeneralAgents.AI and similar QoS-centric frameworks emphasize. 

• Data and domain knowledge

o Effective models depend on large, domain-specific datasets (e.g., agriculture, utilities, aviation, security). Established players are accumulating proprietary datasets and annotations via repeated missions and partnerships.[cirium] 

o Domain-specific regulatory and safety knowledge (e.g., BVLOS regulations, airspace classes, airport operations) are non-trivial to acquire and institutionalize.[cirium]

• Regulatory and certification constraints

o Operating analytics in regulated airspace or critical infrastructure (airports, utilities, defense applications) often requires certifications, security clearances, or vendor approvals, which increase time-to-entry.[aurora]

• Capital requirements and go-to-market

o While the software stack can be bootstrapped, selling into enterprise and public-sector customers typically demands significant sales, integration, and support teams, plus demonstration missions.[scoop.market]

3. Bargaining power of suppliers

Supplier power is moderate, with some concentration around key sensor, hardware, and cloud/A I infrastructure vendors.[researchdive]

Key supplier categories:

• Hardware and sensor suppliers

o Drone airframe and autopilot vendors (e.g., UAS platforms from Aurora Flight Sciences, Auterion-based systems) influence data formats, APIs, and capabilities.[cirium] 

o Camera and sensor manufacturers (RGB, multispectral, LiDAR, thermal) can exert power when specialized sensors are required and alternatives are limited.[grandviewresearch]

• Cloud and AI infrastructure

o Major cloud providers (Azure, GCP, AWS) and GPU vendors are critical suppliers of compute, storage, and acceleration; pricing and quota policies directly affect our margin stack.[marketresearchfuture]

o Labeling and data platforms (e.g., Scale AI) provide annotation, synthetic data, and model training services that can become bottlenecks or sources of vendor lock-in.[prnewswire]

• UTM and aviation data

o Aviation intelligence suppliers like Cirium providing UTM and navigation data for low-altitude airspace become essential for safe route planning and analytics integration.[cirium] 

4. Bargaining power of buyers

Buyer power is moderate to high, especially among large enterprises and public-sector organizations that can switch between analytics vendors.[studocu]

Buyer segments:

• Enterprise and public-sector drone programs

o Utilities, oil & gas, construction, and infrastructure operators use drones for inspection and have multiple service and software vendors pitching similar capabilities.[grandviewresearch] 

o Aviation and airport operators procure analytics and UTM solutions (e.g., Cirium’s services) and have high expectations regarding reliability and integration with existing systems.[cirium] 

• Agriculture and logistics customers

o Precision agriculture customers often treat drones and analytics as part of broader agronomy solutions, which dilutes vendor differentiation and increases price sensitivity.[financialmodelslab] 

o Logistics providers picking last-mile or warehouse analytics platforms can compare multiple offerings (SkyWays Drones-type solutions, horizontal video analytics) and negotiate based on ROI and contract terms.[marketresearchfuture] 

• Security and surveillance customers

o For perimeter and urban security, buyers can choose among drone-based systems, fixed cameras with video analytics, and broader security platforms, increasing their bargaining power.[researchdive] 

5. Threat of substitutes

Threat of substitutes is moderate, varying by vertical.[studocu]

Substitute categories:

• Non-drone sensing

o Fixed CCTV, ground robots, satellites, and manned aircraft can provide overlapping sensing capabilities for some use cases, especially in security, traffic monitoring, and wide-area surveillance.[prnewswire]

o IoT sensors in buildings and infrastructure (e.g., SkyFoundry’s traditional IoT stack) can partially substitute drone-based inspection.[]

• Manual inspection and legacy workflows

o For some infrastructure and agriculture use cases, manual inspection, ground surveys, or legacy aerial imagery remain entrenched, especially where drone regulations are strict.[scoop.market]

• Horizontal video analytics platforms

o Generic video analytics platforms (non-drone-specific) can process drone feeds and deliver many of the same detection and tracking outputs, particularly in simple surveillance or counting tasks.[prnewswire]


Monday, August 3, 2026

 Moving object classification across a subset of scenes from an aerial drone video.


This article explains the use of Fast-Fourier Transform and Short-Time Fourier Transform for object tracking. 


While standard optical object detection relies heavily on spatial image pixel convolutions, FFT and STFT are critical to detect moving objects. FFT (Fast Fourier Transform) converts a discrete signal from the time domain to the frequency domain where the frequency shift aka Doppler effect or time delay aka frequency beat reveals an objects distance and velocity. STFT (Short-Time Fourier Transform): Applies the FFT to localized, overlapping time windows. This captures how frequency changes over time, producing a spectrogram. It is essential for detecting moving objects, classifying micro-Doppler signatures (e.g., distinguishing a pedestrian from a cyclist). The squared magnitude of the STFT yields the Spectrogram, matrix data that object detection models (like 2D CNNs or Transformers) ingest to localize and classify targets.


To extract frequency-domain features from video pixel tracking, we must perform a 3D-to-1D reduction. A 2D CNN or Transformer cannot directly process a raw spatial image with an FFT/STFT along the temporal axis without creating a spatial-temporal bottleneck. Instead, we extract the 1D spatial trajectory vectors (the X and Y coordinates over time) or the temporal pixel intensity shifts of a tracking bounding box, and apply the STFT to those 1D trajectories. This translates the object's physical acceleration, micro-movements, and brief erratic motion into a 2D Time-Frequency Spectrogram that a standard image-based neural network can classify.


While deep learning frameworks like Swin Transformer 3D or TimeSformer handle video natively, engineering frequency features helps classify fast-moving vs. slow-moving targets (e.g., separating a speeding vehicle from a pedestrian) when data is sparse. 


Below is a complete, working pipeline that takes a sequence of aerial drone frames, tracks an object's spatial coordinates across the scenes, computes the STFT on its movement dynamics, and formats the output into a tensor ready for a 2D CNN or Vision Transformer (ViT):

#! /usr/bin/python

import numpy as np

import cv2

import matplotlib.pyplot as plt

from scipy.signal import stft


# ==========================================

# 1. SIMULATE AERIAL DRONE SCENE DATA

# ==========================================

def generate_mock_drone_video(num_frames=120, height=512, width=512):

    """

    Simulates a continuous aerial video snippet from 100m.

    An object (e.g., a cyclist) moves across the frame with micro-vibrations.

    """

    frames = []

    # Base drone scene background (textured noise)

    background = np.random.randint(100, 130, (height, width), dtype=np.uint8)

    

    # Simulate a target moving diagonally across the frame over time

    for t in range(num_frames):

        frame = background.copy()

        

        # Base trajectory + micro-oscillations (pedaling frequency/road bumps)

        center_x = int(50 + 3.2 * t + 2 * np.sin(0.8 * t))

        center_y = int(80 + 2.5 * t + 1.5 * np.cos(0.8 * t))

        

        # Draw the target if it is within bounds

        if 0 < center_x < width and 0 < center_y < height:

            # Simulated target boundary box footprint

            cv2.circle(frame, (center_x, center_y), radius=6, color=255, thickness=-1)

            

        frames.append(frame)

    return np.array(frames)


# ==========================================

# 2. EXTRACT PIXEL TRACKING TRAJECTORIES

# ==========================================

def track_object_centroid(video_frames):

    """

    Simulates an upstream Object Tracker (e.g., ByteTrack / Kalman Filter).

    Returns a 1D array of spatial positions over time.

    """

    trajectory_x = []

    trajectory_y = []

    

    # Simple centroid extraction loop via thresholding for demo purposes

    for frame in video_frames:

        _, thresh = cv2.threshold(frame, 240, 255, cv2.THRESH_BINARY)

        moments = cv2.moments(thresh)

        

        if moments["m00"] != 0:

            cx = moments["m10"] / moments["m00"]

            cy = moments["m01"] / moments["m00"]

        else:

            # Handle brief occlusion/disappearance by holding last known position

            cx = trajectory_x[-1] if trajectory_x else 0

            cy = trajectory_y[-1] if trajectory_y else 0

            

        trajectory_x.append(cx)

        trajectory_y.append(cy)

        

    return np.array(trajectory_x), np.array(trajectory_y)


# ==========================================

# 3. COMPUTE STFT TENSOR FOR DEEP LEARNING

# ==========================================

def generate_stft_features(trajectory_x, trajectory_y, fps=30):

    """

    Converts 1D motion tracking coordinates into a 2D Time-Frequency Spectrogram map.

    """

    # Convert absolute coordinates to velocity vectors (pixel displacement delta)

    vel_x = np.diff(trajectory_x, prepend=trajectory_x[0])

    vel_y = np.diff(trajectory_y, prepend=trajectory_y[0])

    

    # Compute Magnitude of the velocity vector

    velocity_magnitude = np.sqrt(vel_x**2 + vel_y**2)

    

    # Apply STFT to the velocity sequence

    # Short segment length (nperseg) is vital because targets appear briefly

    nperseg = min(32, len(velocity_magnitude)) 

    frequencies, times, Zxx = stft(velocity_magnitude, fs=fps, nperseg=nperseg, noverlap=nperseg-4)

    

    # Extract Power Spectral Density (Magnitude Squared)

    spectrogram = np.abs(Zxx)**2

    

    # Normalize to 0-255 range for standard 2D Image CNN/Transformer consumption

    log_spectrogram = 10 * np.log10(spectrogram + 1e-10)

    norm_spectrogram = cv2.normalize(log_spectrogram, None, 0, 255, cv2.NORM_MINMAX)

    

    return norm_spectrogram.astype(np.uint8), frequencies, times


# ==========================================

# 4. EXECUTION PIPELINE

# ==========================================

# Step A: Load video sequence (Simulated 30 FPS drone clip)

video_data = generate_mock_drone_video(num_frames=150, height=512, width=512)


# Step B: Get object tracking data across frames

x_coords, y_coords = track_object_centroid(video_data)


# Step C: Generate the 2D STFT target signature

stft_tensor, freqs, timeline = generate_stft_features(x_coords, y_coords, fps=30)


# Step D: Resize to square dimensions for standard networks (e.g., 224x224 for ViT/ResNet)

network_input = cv2.resize(stft_tensor, (224, 224), interpolation=cv2.INTER_CUBIC)


print(f"Processed Tracking Coordinates Shape: {x_coords.shape}")

print(f"Generated STFT Tensor Spectrogram Shape: {stft_tensor.shape}")

print(f"Final 2D CNN/Transformer Input Shape: {network_input.shape} (Ready for network integration)")


# ==========================================

# VISUALIZATION

# ==========================================

plt.figure(figsize=(10, 4))

plt.subplot(1, 2, 1)

plt.plot(x_coords, y_coords, '-o', color='teal', markersize=3)

plt.title("Spatial Pixel Path (Aerial View)")

plt.xlabel("X Coordinate")

plt.ylabel("Y Coordinate")

plt.grid(True)


plt.subplot(1, 2, 2)

plt.imshow(network_input, aspect='auto', cmap='magma', origin='lower')

plt.title("Resized Motion STFT Signature (224x224)")

plt.xlabel("Temporal Windows")

plt.ylabel("Frequency Components")

plt.colorbar(label='Normalized Energy')

plt.tight_layout()

plt.show()


Sample output:

Processed Tracking Coordinates Shape: (150,)

Generated STFT Tensor Spectrogram Shape: (17, 39)

Final 2D CNN/Transformer Input Shape: (224, 224) (Ready for network integration)

 


Conclusion:

The STFT translates continuous time-domain raw sensor signals into structured 2D spatial-frequency representations (spectrograms), enabling standard computer vision models and CFAR filters to accurately segment, classify, and track objects based on their range and Doppler velocity signatures.



Sunday, August 2, 2026

Fourier mathematics is a foundational technology underlying nearly every stage of modern drone image detection and analysis. Rather than viewing the Fourier transform as an outdated signal-processing tool displaced by deep learning, the survey shows that it remains central to both classical and state-of-the-art drone vision systems. Its enduring importance stems from two key advantages: computational efficiency and mathematical invariance. By transforming image operations from the spatial domain into the frequency domain, the Fast Fourier Transform (FFT) reduces the computational cost of many tasks from quadratic or higher complexity to near-linear logarithmic complexity, making real-time processing feasible on power- and weight-constrained UAV platforms. At the same time, Fourier methods naturally provide robustness to common aerial-imaging challenges such as changes in position, altitude, orientation, scale, illumination, and vibration. 

Image registration and orthomosaic generation are two essential functions in aerial imaging. Fourier-based phase correlation exploits the shift theorem to estimate the translational offset between overlapping images quickly and accurately, while remaining resistant to brightness differences. Extensions such as the Fourier-Mellin transform further enable rotation- and scale-invariant registration by converting such transformations into simple shifts in a log-polar frequency representation. These methods complement feature-based approaches such as SIFT and Structure-from-Motion pipelines, offering faster alternatives for many alignment problems while often serving as useful preprocessing stages for more computationally intensive photogrammetric workflows. 

Object detection, tracking, and recognition are other applications. Correlation filter trackers such as MOSSE, CSK, and KCF rely on FFT-based computations to transform expensive spatial searches into efficient frequency-domain multiplications. This allows trackers to operate at high frame rates while consuming relatively little computational power, making them suitable for onboard deployment. Fourier descriptors and Generic Fourier Descriptors provide compact shape representations that remain invariant to translation, rotation, and scale, allowing systems to distinguish drones from birds and other cluttered objects. The survey also highlights the use of micro-Doppler analysis and short-time Fourier transforms to identify the distinctive signatures produced by spinning propellers in radar or optical sensing data. Elliptic Fourier descriptors further extend these ideas by enabling compact contour representations suitable for robust object tracking across video frames. 

A detailed case study demonstrates the practical use of Fourier descriptors for aerial object tracking. Using the example of tracking a vehicle across drone imagery, the study illustrates how object contours can be converted into complex signals, transformed into Fourier coefficients, and compared across frames. Successful tracking depends not merely on applying a transform but on proper normalization procedures. Translation, scale, rotation, and starting-point invariance must be handled carefully to preserve meaningful shape information. The case study serves as a reminder that the theoretical advantages of Fourier descriptors are only realized when classical mathematical principles are implemented correctly. 

Beyond detection and tracking, there is a broader infrastructure that supports drone imaging systems. Frequency-domain methods are widely used for image deblurring, enhancement, and restoration, particularly in compensating for motion blur caused by vibration and rapid aircraft maneuvers. Wiener filtering, homomorphic filtering, and related spectral techniques improve image quality by separating signal from noise and correcting uneven illumination. Fourier and Gabor texture analysis support land-cover classification, crop-health monitoring, and precision agriculture by capturing spatial patterns that are often more informative than raw spectral measurements. In data transmission, discrete cosine transforms power modern image and video compression systems, enabling efficient communication over bandwidth-limited drone links. Frequency-based pansharpening techniques fuse high-resolution spatial information with lower-resolution multispectral or thermal imagery, while FFT-based vibration analysis enables predictive maintenance by identifying rotor imbalance, blade damage, and other mechanical faults from sensor data. 

There is a growing integration of Fourier mathematics into deep learning. Architectures such as Fourier Neural Operators, Fast Fourier Convolution networks, FNet, and Global Filter Networks embed spectral operations directly into neural-network layers rather than treating Fourier transforms solely as preprocessing tools. These architectures leverage frequency-domain representations to expand receptive fields, improve computational efficiency, and enable learning across varying spatial resolutions. In UAV applications, frequency-aware neural networks have proven especially valuable for detecting small objects, enhancing domain robustness across d conditions, and scaling vision models to high-resolution aerial imagery. The common thread is that frequency-domain operations often replace more expensive spatial computations while preserving or improving performance. 

This analysis extends beyond individual algorithms to consider system architecture and economics. Using the Drone Video Sensing Analytics (DVSA) framework as an example, Fourier methods are not merely useful techniques but key enablers of cloud-native drone analytics. Efficient frequency-domain operations support frame selection, change detection, image alignment, object representation, and platform diagnostics at scales that would otherwise be prohibitively expensive. By reducing computational overhead and enabling compact representations of information, Fourier mathematics makes large-scale, catalog-driven, cloud-based drone analytics economically viable.

Finally, there are several open challenges. Researchers must balance the mathematically guaranteed invariances of classical Fourier methods with the adaptability of learned representations. Decisions about which computations belong onboard versus in the cloud remain an active systems-engineering problem. Real-world drone imagery continues to challenge classical methods through motion blur, rolling-shutter effects, illumination variation, and scale changes. Emerging directions include specialized FFT hardware, physics-informed neural operators, frequency-aware transformers, and future integration of spectral features into agentic analytics platforms.  Fourier mathematics remains a core structural component of drone image analysis. From image registration and tracking to enhancement, compression, diagnostics, and deep learning, the same principles of spectral representation, efficiency, and invariance continue to shape both the technical capabilities and economic feasibility of modern UAV analytics. [1](https://1drv.ms/w/c/d609fb70e39b65c8/IQBzyW8MyFSJSKMiUtOG2iVOAXHI0o2iAoOXL9q7V2VmDI0?e=3dUcbS)

Saturday, August 1, 2026

 Sample program to query an aerial drone image:

# filename: vlm_scene_query.py

"""

Download an aerial image from an Azure SAS URL, extract XFIF/GPS metadata,

and query a vision-language model (e.g., Qwen2.5VL-7B VLM) to answer a

scene-level question such as estimating area based on parking spot counts.


Usage:

    export MODEL_ID="your-qwen-model-id-or-hf-repo"

    export HF_API_TOKEN="..." # if required by the model host

    python vlm_scene_query.py --sas-url "<SAS_URL>" --question "Estimate area in square meters"


Notes:

- This script uses the Hugging Face transformers pipeline as a generic interface.

  Qwen VLM may require a provider-specific SDK or a different pipeline name.

  Replace the model loading section with the provider-specific code if needed.

- The script extracts GPS EXIF if present and returns it with the model response.

"""


import os

import sys

import argparse

import tempfile

import json

import math

import logging

from typing import Optional, Dict, Any, Tuple


import requests

from PIL import Image

import exifread


# Optional: transformers pipeline for vision-language models

try:

    from transformers import pipeline, AutoTokenizer, AutoModelForSeq2SeqLM

    HF_AVAILABLE = True

except Exception:

    HF_AVAILABLE = False


logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

logger = logging.getLogger("vlm_scene_query")



def download_image_from_sas(sas_url: str, dest_path: str, timeout: int = 30) -> None:

    """Download an image from an Azure SAS URL to dest_path."""

    logger.info("Downloading image from SAS URL")

    resp = requests.get(sas_url, stream=True, timeout=timeout)

    resp.raise_for_status()

    with open(dest_path, "wb") as f:

        for chunk in resp.iter_content(chunk_size=8192):

            if chunk:

                f.write(chunk)

    logger.info("Downloaded image to %s", dest_path)



def extract_gps_from_exif(image_path: str) -> Dict[str, Any]:

    """Extract GPS EXIF data (if present) using exifread and return a dict."""

    logger.info("Extracting EXIF metadata")

    with open(image_path, "rb") as f:

        tags = exifread.process_file(f, details=False)

    gps = {}

    def _get(tag):

        return tags.get(tag)

    # Common EXIF GPS tags

    lat_ref = _get("GPS GPSLatitudeRef")

    lat = _get("GPS GPSLatitude")

    lon_ref = _get("GPS GPSLongitudeRef")

    lon = _get("GPS GPSLongitude")

    alt = _get("GPS GPSAltitude")

    if lat and lon and lat_ref and lon_ref:

        def _to_deg(value):

            # value is like [Rational(37,1), Rational(46,1), Rational(0,1)]

            try:

                parts = [float(x.num) / float(x.den) for x in value.values]

                deg = parts[0] + parts[1] / 60.0 + parts[2] / 3600.0

                return deg

            except Exception:

                return None

        lat_deg = _to_deg(lat)

        lon_deg = _to_deg(lon)

        if lat_deg is not None and lon_deg is not None:

            if str(lat_ref).upper().startswith("S"):

                lat_deg = -lat_deg

            if str(lon_ref).upper().startswith("W"):

                lon_deg = -lon_deg

            gps["latitude"] = lat_deg

            gps["longitude"] = lon_deg

    if alt:

        try:

            gps["altitude"] = float(alt.values[0].num) / float(alt.values[0].den)

        except Exception:

            pass

    # XFIF or other tags may be present; include raw tags for inspection

    gps["raw_tags"] = {k: str(v) for k, v in tags.items() if k.startswith("GPS")}

    return gps



def default_area_estimate_from_parking_count(parking_count: int,

                                             spot_length_m: float = 4.5,

                                             spot_width_m: float = 1.8,

                                             spacing_factor: float = 1.2) -> Tuple[float, float]:

    """

    Estimate area in square meters and square feet given a count of parking spots.

    Default sedan footprint: 4.5m x 1.8m = 8.1 m^2. spacing_factor accounts for drive lanes and spacing.

    Returns (area_m2, area_ft2).

    """

    single_spot_area = spot_length_m * spot_width_m * spacing_factor

    total_m2 = parking_count * single_spot_area

    total_ft2 = total_m2 * 10.7639

    return total_m2, total_ft2



def build_prompt_for_vlm(question: str, guidance: Optional[str] = None) -> str:

    """

    Build a clear prompt for the vision-language model. Guidance can include

    assumptions to make (e.g., sedan footprint).

    """

    base = (

        "You are given an aerial image. Answer the user's question precisely. "

        "If you need to make reasonable assumptions, state them explicitly. "

        "Return a JSON object with keys: 'answer_text', 'parking_spot_count' (int or null), "

        "'assumptions' (list of strings), and 'computed' (object with numeric fields). "

    )

    if guidance:

        base += guidance + " "

    base += "User question: " + question

    return base




def query_vlm_with_image(model_id: str, image_path: str, prompt: str, hf_token: Optional[str] = None) -> Dict[str, Any]:

    import torch

    from PIL import Image

    from transformers import AutoProcessor, AutoModelForCausalLM


    if hf_token:

        os.environ["HUGGINGFACEHUB_API_TOKEN"] = hf_token


    device = "cuda" if torch.cuda.is_available() else "cpu"


    processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)

    model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).to(device)


    image = Image.open(image_path).convert("RGB")


    # Qwen expects both text + image in the processor

    inputs = processor(text=prompt, images=image, return_tensors="pt").to(device)


    generated_ids = model.generate(

        **inputs,

        max_new_tokens=512

    )


    output_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]


    try:

        return json.loads(output_text)

    except Exception:

        return {"answer_text": output_text, "raw": output_text}




def torch_cuda_available() -> bool:

    try:

        import torch

        return torch.cuda.is_available()

    except Exception:

        return False



def parse_parking_count_from_vlm_response(vlm_resp: Dict[str, Any]) -> Optional[int]:

    """

    Extract parking_spot_count if present in the VLM response dict.

    """

    try:

        count = vlm_resp.get("parking_spot_count")

        if count is None:

            # Try to parse from answer_text heuristically

            text = vlm_resp.get("answer_text", "")

            # naive heuristic: find first integer in text

            import re

            m = re.search(r"\b(\d{1,4})\b", text)

            if m:

                return int(m.group(1))

            return None

        return int(count)

    except Exception:

        return None



def main():

    parser = argparse.ArgumentParser(description="Query a VLM about an aerial image from an Azure SAS URL")

    parser.add_argument("--sas-url", required=True, help="Azure SAS URL to the JPEG image")

    parser.add_argument("--question", required=True, help="Natural language question to ask the VLM")

    parser.add_argument("--model-id", default=os.environ.get("MODEL_ID", "Qwen/Qwen-2.5V-L-7B"), help="VLM model id or repo")

    parser.add_argument("--hf-token", default=os.environ.get("HF_API_TOKEN"), help="Hugging Face API token if required")

    parser.add_argument("--assume-spot-length-m", type=float, default=4.5, help="Assumed parking spot length in meters")

    parser.add_argument("--assume-spot-width-m", type=float, default=1.8, help="Assumed parking spot width in meters")

    parser.add_argument("--spacing-factor", type=float, default=1.2, help="Factor to account for drive lanes and spacing")

    parser.add_argument("--no-vlm", action="store_true", help="Skip VLM and use deterministic heuristic only")

    args = parser.parse_args()


    with tempfile.TemporaryDirectory() as tmpdir:

        img_path = os.path.join(tmpdir, "scene.jpg")

        try:

            download_image_from_sas(args.sas_url, img_path)

        except Exception as e:

            logger.error("Failed to download image: %s", e)

            sys.exit(1)


        gps = extract_gps_from_exif(img_path)

        logger.info("Extracted GPS metadata: %s", gps)


        # Build prompt

        guidance = (

            f"Assume a typical U.S. sedan footprint of {args.assume_spot_length_m}m x {args.assume_spot_width_m}m "

            f"and a spacing factor of {args.spacing_factor} to account for drive lanes. "

            "Count visible parking spots if possible and compute total area in square meters and square feet."

        )

        prompt = build_prompt_for_vlm(args.question, guidance=guidance)


        vlm_response = None

        parking_count = None

        if not args.no_vlm:

            try:

                vlm_response = query_vlm_with_image(args.model_id, img_path, prompt, hf_token=args.hf_token)

                logger.info("VLM response received")

                parking_count = parse_parking_count_from_vlm_response(vlm_response)

            except Exception as e:

                logger.warning("VLM query failed or not available: %s", e)

                vlm_response = {"error": str(e)}

                parking_count = None


        # If VLM didn't provide a parking count, fall back to asking user assumption or using a heuristic

        if parking_count is None:

            # Heuristic fallback: try to detect cars using a very small, dependency-free heuristic is not reliable.

            # Instead, we will ask the model's textual output for a number if available; otherwise, default to 10 spots.

            if vlm_response and isinstance(vlm_response, dict):

                parking_count = parse_parking_count_from_vlm_response(vlm_response)

            if parking_count is None:

                logger.info("No parking count from VLM; using fallback default of 10 spots for estimation")

                parking_count = 10 # conservative default; in production, prefer human-in-the-loop


        area_m2, area_ft2 = default_area_estimate_from_parking_count(

            parking_count,

            spot_length_m=args.assume_spot_length_m,

            spot_width_m=args.assume_spot_width_m,

            spacing_factor=args.spacing_factor

        )


        # Build final structured response

        response = {

            "run_id": f"run-{os.urandom(6).hex()}",

            "agent_id": "qwen-vlm-estimator",

            "start_time": None,

            "end_time": None,

            "model_version": args.model_id,

            "gps": gps,

            "question": args.question,

            "vlm_raw_response": vlm_response,

            "parking_spot_count_used": parking_count,

            "assumptions": [

                f"sedan footprint {args.assume_spot_length_m}m x {args.assume_spot_width_m}m",

                f"spacing factor {args.spacing_factor}"

            ],

            "computed": {

                "area_m2": round(area_m2, 2),

                "area_ft2": round(area_ft2, 2),

                "spot_area_m2": round(args.assume_spot_length_m * args.assume_spot_width_m * args.spacing_factor, 2)

            },

            "answer_text": (

                f"Estimated total area ≈ {round(area_m2,2)} m² ({round(area_ft2,2)} ft²) "

                f"based on {parking_count} parking spots and assumed sedan footprint "

                f"{args.assume_spot_length_m}m x {args.assume_spot_width_m}m with spacing factor {args.spacing_factor}."

            )

        }


        # Print JSON response

        print(json.dumps(response, indent=2))



if __name__ == "__main__":

    main()



Results:

{

  "run_id": "run-186a8c43b7e7",

  "agent_id": "qwen-vlm-estimator",

  "start_time": null,

  "end_time": null,

  "model_version": "Qwen/Qwen2.5-VL-7B-Instruct",

  "gps": {

    "raw_tags": {}

  },

  "question": "Estimate area of unoccupied spots in square meters",

  "parking_spot_count_used": 10,

  "assumptions": [

    "sedan footprint 4.5m x 1.8m",

    "spacing factor 1.2"

  ],

  "computed": {

    "area_m2": 97.2,

    "area_ft2": 1046.25,

    "spot_area_m2": 9.72

  },

  "answer_text": "Estimated total area \u2248 97.2 m\u00b2 (1046.25 ft\u00b2) based on 10 parking spots and assumed sedan footprint 4.5m x 1.8m with spacing factor 1.2."

}


Friday, July 31, 2026

 Agentic judges for drone image analytics

Andrew Ng’s agentic workflow pattern—reflection, tool use, planning, and multi-agent collaboration—applies to drone-vision benchmarking. Reflection lets an agent critique and revise detections, captions, SQL answers, or mission reports. Tool use grounds reasoning in retrieval, code execution, geospatial operators, detector APIs, and database queries. Planning decomposes a workload into explicit steps before execution and supports replanning when a probe fails. Multi-agent orchestration assigns specialized roles: one agent checks geospatial consistency, another evaluates temporal coherence, another audits semantic alignment, and an arbiter aggregates evidence into a score.

Memory has been a design constraint. Loops let agents think; graphs let agents remember. A reflection loop can improve a single answer, but without persistent state the agent forgets why it inspected a frame, which detector disagreed, or which workload constraint failed. A graph turns those transient observations into reusable memory: frames, objects, captions, detections, SQL results, tool calls, critiques, and final judgments become linked evidence rather than buried transcript text. In ezbenchmark, this converts an agentic judge from a one-pass evaluator into a stateful audit system.

A practical build path is incremental. First, add one critique call after every generated answer or score; this is the highest-return change because it catches unsupported claims, missing visual evidence, and weak workload alignment before output is finalized. Second, expose the judge to tools: vector retrieval over frame evidence, SQL over the scene catalog, detector re-runs, spatial predicates, and temporal-neighbor comparisons. Third, require the judge to emit a structured plan before execution, such as JSON steps with expected evidence, tool calls, and failure conditions. Fourth, split evaluation across specialized agents and connect them through a shared graph store. The result is not merely a stronger prompt; it is an architecture in which weak models can outperform stronger single-pass models because the workflow supplies iteration, grounding, decomposition, and durable memory.

Agents are usually dedicated to perception, reasoning, and control in different ways. Sapkota et al. introduce the term “Agentic UAVs” to describe systems that integrate perception, cognition, control, and communication into layered, goal-driven agents that operate with contextual reasoning and memory, rather than fixed scripts or reactive control loops [1]. In their framework, aerial image understanding is only one layer in a broader cognitive stack: perception agents extract structure from imagery and other sensors; cognitive agents plan and replan missions; control agents execute trajectories; and communication agents coordinate with humans and other UAVs. This layered view is useful when we start thinking about agentic frameworks as “judges” for benchmarking: the judging capability can itself be an agent, sitting in the cognition layer, consuming outputs from perception agents and workload metadata rather than raw pixels alone [1].

Vision–language–driven agents are a distinct subclass. Sapkota et al. explicitly highlight vision–language models and multimodal sensing as key enabling technologies for Agentic UAVs, noting that they allow agents to parse complex scenes, follow natural-language instructions, and ground symbolic goals in visual context [1]. These agents differ from traditional planners in that they can reason over image and text jointly, which makes them natural candidates for roles like “mission explainer,” “anomaly triager,” or, in our case, “benchmark judge” for aerial analytics workloads. Instead of judging purely from numeric metrics, a vision–language agent can look at a drone scene, read a workload description, inspect candidate outputs, and form a qualitative judgment about which pipeline better captures the intended analytic semantics [1].


Thursday, July 30, 2026

 Drone Fleet motion propagation

Autonomous drone fleet does not have a centralized controller. Motion of the forward section of the fleet must be followed by the backward section of the fleet via propagation of the direction and speed information. This is like linear heat propagation and is best described by Fourier Transforms. Study of the propagation of fleet movement information is necessary to enable the fleet to behave as coordinated as a whole unit. The movement information is mere payload as it can be enriched with many forms of sensor data that can help the fleet to perform actions such as avoiding obstacles, optimizing flight paths and speed, and achieving formation objectives or cost functions. While these actions can be computed locally or in co-ordination by one or more members of the fleet, the propagation of messages is independent of the computation and involves phenomena like heat propagation. This article discusses the wave propagation by Fourier Transforms assuming peer processing of sensor data can be coordinated with consensus algorithms.

A Fast Fourier Transform converts wave form data in the time domain into the frequency domain. It achieves this by breaking down the original time-based waveform into a series of sinusoidal terms, each with a unique magnitude, frequency, and phase. This process converts a waveform in the time domain into a series of sinusoidal functions which when added together reconstruct the original waveform. Plotting the amplitude of each sinusoidal term versus its frequency creates a power spectrum, which is the response of the original waveform in the frequency domain.

The Fourier transform is a generalization of the Fourier series and allows taking any function as a total of simple sinusoids. A functions’ Fourier transform is a complex-valued function denoting the constituent complex sinusoid’s that contain the original function. For every frequency, the magnitude of the complex value denotes the constituent complex sinusoid’s amplitude with that frequency, and the complex value’s argument denotes the phase offset of the complex sinusoid. If a frequency does not exist, the transform possesses a value of zero for that frequency. Functions that are localized in the domain of time have Fourier transforms that extend out across the domain of frequency and vice versa. The Fourier transform of a Gaussian function is always another Gaussian function. The solutions to heat equations are the Gaussian functions.

The Fourier transform can be generalized to functions of various variables on Euclidean space, forwarding a function of three-dimensional position space to a three-dimensional momentum function (or a space and time function to a 4-momentum function). This helps spatial Fourier Transform to be solutions of waves as functions of either momentum or position or both.

When Fourier transforms are applicable, it means the “earth response” now is the same as the “earth response” later. Switching our point of view from time to space, the applicability of the Fourier transformation means that the “impulse response” here is the same as the “impulse response” there. An impulse is a column vector full of zeros with somewhere a one. An impulse response is a column from the matrix q = Bp The collection of impulse responses in q=Bp defines the convolution operation.

The difference between Fourier Transform and Fourier Series is that the Fourier Transform is applicable for non-periodic signals, while the Fourier Series is applicable to periodic signals. The properties of Fourier transform are duality, linear transform, modulation and Parseval’s theorem. Duality implies that if h(t) possesses a Fourier transform H(f), then the Fourier transform related to H(t) is H(-f). Linear Transform implies that if g(t) and h(t) are two Fourier transforms also represented by G(f) and H(f) respectively, then the linear combination of h and g also has a Fourier transform. The modulation property implies that functions are modulated by other functions if they are multiplied in time. Parseval’s theorem states that the Fourier transform is unitary and the sum of the squares of the H(f) equals that of h(t)

While the dominant interest in the application of wave propagation transforms to fleet movement is the duration it takes for the entire fleet to respond, the drone traffic can be taken as a Gaussian distribution and one that can even be treated as an invariant through different formations. This makes it easier to predict the fleet movements.


Sample FFT application:

import numpy as nm

import scipy

import scipy.fftpack

import pylab


def lowpass_cosine( y, tau, f_3db, width, padd_data=True):

    # padd_data = True means we are going to symmetric copies of the data to the start and stop

    # to reduce/eliminate the discontinuities at the start and stop of a dataset due to filtering

    #

    # False means we're going to have transients at the start and stop of the data


    # kill the last data point if y has an odd length

    if nm.mod(len(y),2):

        y = y[0:-1]


    # add the weird padd

    # so, make a backwards copy of the data, then the data, then another backwards copy of the data

    if padd_data:

        y = nm.append( nm.append(nm.flipud(y),y) , nm.flipud(y) )


    # take the FFT

    ffty=scipy.fftpack.fft(y)

    ffty=scipy.fftpack.fftshift(ffty)


    # make the companion frequency array

    delta = 1.0/(len(y)*tau)

    nyquist = 1.0/(2.0*tau)

    freq = nm.arange(-nyquist,nyquist,delta)

    # turn this into a positive frequency array

    pos_freq = freq[(len(ffty)/2):]


    # make the transfer function for the first half of the data

    i_f_3db = min( nm.where(pos_freq >= f_3db)[0] )

    f_min = f_3db - (width/2.0)

    i_f_min = min( nm.where(pos_freq >= f_min)[0] )

    f_max = f_3db + (width/2);

    i_f_max = min( nm.where(pos_freq >= f_max)[0] )


    transfer_function = nm.zeros(len(y)/2)

    transfer_function[0:i_f_min] = 1

    transfer_function[i_f_min:i_f_max] = (1 + nm.sin(-nm.pi * ((freq[i_f_min:i_f_max] - freq[i_f_3db])/width)))/2.0

    transfer_function[i_f_max:(len(freq)/2)] = 0


    # symmetrize this to be [0 0 0 ... .8 .9 1 1 1 1 1 1 1 1 .9 .8 ... 0 0 0] to match the FFT

    transfer_function = nm.append(nm.flipud(transfer_function),transfer_function)


    # plot up the transfer function

    # since "freq" is only the positive frequencies, select out

    pylab.figure(1)

    pylab.clf()

    pylab.plot(freq,transfer_function)

    pylab.xlabel('Frequency [Hz]')

    pylab.ylabel('Filter Transfer Function')

    pylab.xlim([-10.0,10.0])

    pylab.ylim([-0.05,1.05])


    # apply the filter, undo the fft shift, and invert the fft

    filtered=nm.real(scipy.fftpack.ifft(scipy.fftpack.ifftshift(ffty*transfer_function)))


    # remove the padd, if we applied it

    if padd_data:

        filtered = filtered[(len(y)/3):(2*(len(y)/3))]


    # return the filtered data

    return filtered



# do an example of lowpass filtering

# first make some fake data

# a sine wave fluctuating once every pi seconds

# samples 1000 times per second

fakedata = nm.sin(nm.arange(0,11,0.001)) + nm.random.randn(len(nm.arange(0,11,0.001)))/4.0


# run the filter

# lowpass at 5 Hz, with a 1 Hz width of its roll-off

filtered = lowpass_cosine(fakedata,0.001,5.0,1.0,padd_data=True)


# plot the noisy data, with the filtered data on top

pylab.figure(2)

pylab.clf()

pylab.plot(nm.arange(0,11,0.001),fakedata,label='Noisy Data')

pylab.plot(nm.arange(0,11,0.001),filtered,label='Lowpass Filtered Data')

pylab.xlabel('Time [s]')

pylab.ylabel('Voltage')

pylab.legend()


pylab.ion()

pylab.show()



Wednesday, July 29, 2026

 Agentic judges for drone image analytics

Andrew Ng’s agentic workflow pattern—reflection, tool use, planning, and multi-agent collaboration—applies to drone-vision benchmarking. Reflection lets an agent critique and revise detections, captions, SQL answers, or mission reports. Tool use grounds reasoning in retrieval, code execution, geospatial operators, detector APIs, and database queries. Planning decomposes a workload into explicit steps before execution and supports replanning when a probe fails. Multi-agent orchestration assigns specialized roles: one agent checks geospatial consistency, another evaluates temporal coherence, another audits semantic alignment, and an arbiter aggregates evidence into a score.

Memory has been a design constraint. Loops let agents think; graphs let agents remember. A reflection loop can improve a single answer, but without persistent state the agent forgets why it inspected a frame, which detector disagreed, or which workload constraint failed. A graph turns those transient observations into reusable memory: frames, objects, captions, detections, SQL results, tool calls, critiques, and final judgments become linked evidence rather than buried transcript text. In ezbenchmark, this converts an agentic judge from a one-pass evaluator into a stateful audit system.

A practical build path is incremental. First, add one critique call after every generated answer or score; this is the highest-return change because it catches unsupported claims, missing visual evidence, and weak workload alignment before output is finalized. Second, expose the judge to tools: vector retrieval over frame evidence, SQL over the scene catalog, detector re-runs, spatial predicates, and temporal-neighbor comparisons. Third, require the judge to emit a structured plan before execution, such as JSON steps with expected evidence, tool calls, and failure conditions. Fourth, split evaluation across specialized agents and connect them through a shared graph store. The result is not merely a stronger prompt; it is an architecture in which weak models can outperform stronger single-pass models because the workflow supplies iteration, grounding, decomposition, and durable memory.

Agents are usually dedicated to perception, reasoning, and control in different ways. Sapkota et al. introduce the term “Agentic UAVs” to describe systems that integrate perception, cognition, control, and communication into layered, goal-driven agents that operate with contextual reasoning and memory, rather than fixed scripts or reactive control loops [1]. In their framework, aerial image understanding is only one layer in a broader cognitive stack: perception agents extract structure from imagery and other sensors; cognitive agents plan and replan missions; control agents execute trajectories; and communication agents coordinate with humans and other UAVs. This layered view is useful when we start thinking about agentic frameworks as “judges” for benchmarking: the judging capability can itself be an agent, sitting in the cognition layer, consuming outputs from perception agents and workload metadata rather than raw pixels alone [1].

Vision–language–driven agents are a distinct subclass. Sapkota et al. explicitly highlight vision–language models and multimodal sensing as key enabling technologies for Agentic UAVs, noting that they allow agents to parse complex scenes, follow natural-language instructions, and ground symbolic goals in visual context [1]. These agents differ from traditional planners in that they can reason over image and text jointly, which makes them natural candidates for roles like “mission explainer,” “anomaly triager,” or, in our case, “benchmark judge” for aerial analytics workloads. Instead of judging purely from numeric metrics, a vision–language agent can look at a drone scene, read a workload description, inspect candidate outputs, and form a qualitative judgment about which pipeline better captures the intended analytic semantics [1].

#codingexercise: CodingExercise-07-29-2026.docx