Algorithm interview questions behind modular drones: AI implementation seen through HOVERAir VERSA
I noticed an interesting detail about the HOVERAir VERSA drone from Zero Tech. What attracted me most in its marketing wasn't how high it can fly or how clear the footage is, but that "modular structure"—you can detach the body to use it as a 3-axis gimbal camera, and attach the propellers to turn it back into a drone. Honestly, my first reaction was: what does this design mean for algorithm engineers?
Lately, I've been grinding LeetCode to the point of obsession; whenever I see a new product, I want to break it down into a state machine. The VERSA's "dual-mode" functionality is essentially just a state switch: mode = 0 is handheld gimbal, mode = 1 is flight mode. But when it comes to engineering implementation, sensor fusion, control logic, and AI algorithms all have to follow the state. This reminds me of a question I faced during an interview at a robotics company. They asked, "How would you design a perception system that switches between handheld and autonomous movement?" I stumbled through my answer then. Now, looking at this product, I feel like I've had an epiphany.
Modular Design: State Machine Thinking for Algorithm Deployment
From the product documentation, the VERSA's "Little Black Box" propellers are detachable. This means the computing unit (likely SoC and NPU) needs to support two working modes simultaneously: visual tracking when used as a handheld gimbal, and SLAM obstacle avoidance during flight. These two modes have completely different requirements for compute power, power consumption, and latency.
Let me guess their design approach (purely personal speculation):
// Pseudocode: Modular state management
enum ModMode { HANDHELD, FLIGHT };
void switchMode(ModMode newMode) {
if (newMode == HANDHELD) {
// Disable optical flow sensors, enable touchscreen interaction
// Load lightweight AI composition model (e.g., 2M params)
// Lower IMU sampling rate to save power
} else if (newMode == FLIGHT) {
// Enable all sensors (IMU + Vision + Barometer)
// Load full VIO + Obstacle Avoidance model (possibly 10M+ params)
// Increase NPU clock frequency to ensure real-time performance
}
// Switch sensor drivers, recalibrate coordinate systems
recalibrateSensors();
}
Would this be asked in an interview? I think so. Many companies like to ask, "How do you design an inference pipeline with dynamic multimodal switching?" The key points are:
- Model hot-loading vs. cold start: If both modes share the same DLA hardware, how do you quickly switch weights?
- Memory reuse: Intermediate feature maps occupied in handheld mode might be released in flight mode, but fragmentation must be avoided.
- State consistency: During the switch, the output of the control loop cannot change abruptly, otherwise the gimbal will shake.
When solving these types of problems on LeetCode, I'm used to the "state machine + cache" approach, but actual products have far more non-functional constraints than algorithm questions.
AI Composition: From "Grinding Problems" to "Grinding Faces"
The official site says it supports AI composition features. I'm somewhat interested in this—after all, I've solved 300 LeetCode problems, but I still feel insecure facing open-ended questions like "how to make a machine automatically take good-looking photos."
So-called AI composition is essentially a joint problem of object detection + aesthetic scoring + path planning. For example: if a drone wants to film a person running, it needs to continuously track the subject while adjusting the lens angle to place the person at the golden ratio point of the frame. Behind this is a typical "tracking + planning" problem.
I was recently asked in an interview, "How to use reinforcement learning for automatic drone tracking," and I only answered with the DQN framework. When the interviewer pressed, "How do you design the reward function?", I got stuck. Now thinking about it, VERSA's AI composition might use a more engineering-oriented solution:
# Simplified composition reward function
def reward_function(bbox, frame_width, frame_height):
# Subject position: hope it's close to the golden ratio point
cx, cy = bbox_center(bbox)
ideal_x = frame_width * 0.382
ideal_y = frame_height * 0.618
position_penalty = (cx - ideal_x)**2 + (cy - ideal_y)**2
# Subject size: occupying 1/3 to 1/2 of the frame is better
area_ratio = bbox_area(bbox) / (frame_width * frame_height)
size_penalty = abs(area_ratio - 0.3) # Target 30%
# Smoothness: avoid screen shaking
vel_penalty = current_velocity_norm() # Gimbal motion speed
return - (position_penalty + size_penalty + vel_penalty)
Although this reward function is simple, actual deployment also requires considering inference latency (must complete within 30ms), lighting changes, occlusion, etc. I'm curious how much compute power they use for this; it might also be an NPU bottleneck issue.
3D Worlds Technology: The Practical Version of SLAM Questions
The news mentions "3D Worlds" technology,
Original link: https://www.ithome.com/0/983/809.htm
Physix Frontier