Aiden splits voice interaction from device-task execution: a foreground Realtime Agent keeps the conversation responsive while a Backend Agent handles visually-grounded device work, coordinated asynchronously through a task queue rather than a shared loop. Grounded in real full-duplex speech research (Moshi's parallel-stream model, dual-tower dialogue modeling), addressing the real limitations of both Push-to-Talk (hard turn boundaries) and cascaded VAD→STT→LLM→TTS pipelines (component separation, information loss, special-cased interruption handling).

The piece covers the task-completion vs. task-success distinction, the 500ms notification aggregation window, State/Notice runtime message design, the human-handoff tool pair, and why device tasks run strictly serially by design. Explicitly framed as dev-board stage with a genuine call for community testing feedback.

Full article: aidenai.io/blog/when-voice-meets-the-physical-world-inside-aidens-full-duplex-agent-architecture Firmware: github.com/AidenAI-IO/aiden-firmware