Overview
Large language models (LLMs) are increasingly being deployed as coding agents tasked with editing files, running builds and tests, inspecting execution results, and iteratively repairing software. A specific and demanding application area for these agents is embedded firmware development, where correctness relies on closed-loop behavior under various constraints, including sensing, timing, and safety. This contrasts with static source quality alone. Existing evaluations of embedded agents have often been limited, frequently emphasizing one-shot synthesis or offline correctness.
Research Context
The development and deployment of LLMs as coding agents for embedded systems present unique challenges. Embedded firmware mandates adherence to closed-loop behavior, incorporating sensing, timing, and safety constraints. Traditional evaluation methodologies for these agents have typically focused on single-pass code generation or assessing code correctness without considering dynamic, iterative feedback loops inherent to embedded development.
Approach
A new benchmark was developed specifically for the closed-loop evaluation of embedded coding agents. Each task within this benchmark provides a plain-text engineering description, a constrained workspace, and a visible build-and-runtime surface. The agent's objective is to translate specified requirements into an implementation, generate self-verification steps, and then iterate through these steps until the required device behavior is achieved. The suite includes five embedded-control tasks and incorporates four distinct feedback scenarios:
- One-shot generation
- Realistic self-verification
- CI-style red/green feedback
- Oracle-style detailed feedback
The implementation of this benchmark targets simulated ESP32 firmware, chosen for its reproducibility. The study evaluated seven configurations from the GPT-family and Qwen-family across the five tasks and four feedback scenarios. To ensure robustness, three repetitions were conducted per condition, resulting in a total of 420 runs for the evaluation.
Findings
The evaluation of the LLM agents yielded several specific findings regarding their performance on embedded software development tasks:
- Among the evaluated configurations, gpt-5.4 demonstrated the highest pass rate. However, gpt-5.4 did not saturate the benchmark, indicating room for further performance improvement or task complexity beyond its current capabilities.
- For local models, qwen3.5-27B was identified as the strongest performer among those observed.
- Smaller local models exhibited a significant degradation in performance, reflected in both a sharp decrease in pass rate and reduced search efficiency compared to their larger counterparts.
Why This Matters
The results of this evaluation suggest the emergence of capable local embedded coding agents. The development of a benchmark for closed-loop evaluation addresses a gap in assessing LLM agent performance in the demanding context of embedded firmware, where iterative development and real-time constraints are critical. This structured evaluation framework provides a basis for understanding the current capabilities and limitations of LLMs in generating and debugging embedded software, particularly highlighting performance differences between various model sizes and families.