Closed-Loop Evaluation of LLM Agents for Embedded Software Development

arXiv CS · · 2 min read · Engineering & Technology

Read research and analysis on Closed-Loop Evaluation of LLM Agents for Embedded Software Development published by ICANEWS, a global research journal for emerging researchers.

Key Takeaways

  • gpt-5.4 had the highest pass rate among evaluated configurations but did not saturate the benchmark.
  • qwen3.5-27B was the strongest observed local model.
  • Smaller local models degraded sharply in pass rate and search efficiency.

Why This Matters

The findings suggest that capable local embedded coding agents are emerging, providing a benchmark for assessing their performance in complex, iterative embedded firmware development. This addresses the need for evaluating LLMs beyond one-shot synthesis or offline correctness in a domain critical for sensing, timing, and safety constraints.

Overview

Large language models (LLMs) are increasingly being deployed as coding agents tasked with editing files, running builds and tests, inspecting execution results, and iteratively repairing software. A specific and demanding application area for these agents is embedded firmware development, where correctness relies on closed-loop behavior under various constraints, including sensing, timing, and safety. This contrasts with static source quality alone. Existing evaluations of embedded agents have often been limited, frequently emphasizing one-shot synthesis or offline correctness.

Research Context

The development and deployment of LLMs as coding agents for embedded systems present unique challenges. Embedded firmware mandates adherence to closed-loop behavior, incorporating sensing, timing, and safety constraints. Traditional evaluation methodologies for these agents have typically focused on single-pass code generation or assessing code correctness without considering dynamic, iterative feedback loops inherent to embedded development.

Approach

A new benchmark was developed specifically for the closed-loop evaluation of embedded coding agents. Each task within this benchmark provides a plain-text engineering description, a constrained workspace, and a visible build-and-runtime surface. The agent's objective is to translate specified requirements into an implementation, generate self-verification steps, and then iterate through these steps until the required device behavior is achieved. The suite includes five embedded-control tasks and incorporates four distinct feedback scenarios:

  • One-shot generation
  • Realistic self-verification
  • CI-style red/green feedback
  • Oracle-style detailed feedback

The implementation of this benchmark targets simulated ESP32 firmware, chosen for its reproducibility. The study evaluated seven configurations from the GPT-family and Qwen-family across the five tasks and four feedback scenarios. To ensure robustness, three repetitions were conducted per condition, resulting in a total of 420 runs for the evaluation.

Findings

The evaluation of the LLM agents yielded several specific findings regarding their performance on embedded software development tasks:

  • Among the evaluated configurations, gpt-5.4 demonstrated the highest pass rate. However, gpt-5.4 did not saturate the benchmark, indicating room for further performance improvement or task complexity beyond its current capabilities.
  • For local models, qwen3.5-27B was identified as the strongest performer among those observed.
  • Smaller local models exhibited a significant degradation in performance, reflected in both a sharp decrease in pass rate and reduced search efficiency compared to their larger counterparts.

Why This Matters

The results of this evaluation suggest the emergence of capable local embedded coding agents. The development of a benchmark for closed-loop evaluation addresses a gap in assessing LLM agent performance in the demanding context of embedded firmware, where iterative development and real-time constraints are critical. This structured evaluation framework provides a basis for understanding the current capabilities and limitations of LLMs in generating and debugging embedded software, particularly highlighting performance differences between various model sizes and families.

Research Information

Institution
arXiv CS
Original Study
View Publication
Source
arXiv CS

About ICANEWS

ICANEWS is a global research journal for emerging researchers, publishing student and emerging researcher work across all fields.