AI RACE— The AI Race
Research

AI Coding Agents Can Turn Photos into 3D Scenes but Struggle to Judge Their Own Work

Researchers from the University of Maryland and AWS introduced LEGO-Anything, revealing that while coding agents can iteratively build 3D Blender scenes from photos, their self-assessment of geometric accuracy remains near chance level.

10/03/2026, 15:31
AI viết code có thể biến ảnh chụp thành cảnh 3D nhưng lại bất lực trong việc tự đánh giá sản phẩm của mình

Researchers from the University of Maryland and AWS have unveiled LEGO-Anything, a framework that uses AI coding agents to convert a single two-dimensional photograph into a fully editable 3D scene. By generating and running code step by step in Blender, the system produces transparent, queryable 3D environments.

However, a newly developed benchmark accompanying the project reveals a fundamental bottleneck: while agents can write functional 3D code, they struggle to evaluate whether their reconstructions are geometrically accurate, often sabotaging their own progress.

The Image-to-Code Approach and LEGO-Bench

Traditional generative 3D pipelines often output rigid meshes or neural representations directly. In contrast, LEGO-Anything adopts an "Image-to-Code" paradigm. A coding agent is given a single reference image and iteratively writes Python scripts for Blender, the open-source 3D software. The agent runs the code, inspects the rendered output, and revises the script until the synthesized scene resembles the input. Because the final output is an executable program, users can directly modify objects, adjust camera angles, or inspect geometric layouts.

To evaluate this workflow, the research team created LEGO-Bench. Because real-world photos lack precise 3D ground truth and basic synthetic environments look artificial, LEGO-Bench uses 208 images rendered from 104 professionally designed indoor and outdoor simulator scenes, featuring 443 registered assets. This setup allows researchers to hide exact geometry, depth, and object parameters to use as an automated scoring key.

The benchmark scores agent-generated scenes across three metrics:

  • Validity: Whether the code successfully executes and delivers a usable 3D scene artifact.
  • Reconstruction: How precisely the visible 3D geometry matches the ground truth.
  • Appearance: A pixel-by-pixel visual comparison between the re-rendered scene and the original reference photo.

Strengths, Regressions, and Chance-Level Judgments

Across experiments testing six GPT configurations, all models reliably produced functioning Blender scenes. However, structural precision varied widely across models and scene types.

The top-performing model, GPT-6 Astra, achieved an accuracy score of 53.4 percent on indoor environments and 39.6 percent on outdoor environments. Weaker configurations trailed significantly, scoring around 15 percent. Across the board, complex scenes and outdoor spaces proved significantly more difficult than interior layouts.

Increasing the models' reasoning budgets led to notable performance jumps. On a dedicated office subset, GPT-6 Astra’s score climbed from 32.3 percent to 61.8 percent when given more computation time to think.

Despite these gains, the researchers identified severe flaws during iterative editing. Agents frequently made flawed initial attempts or executed edits that dismantled earlier successes. In one documented instance, GPT-6 Astra degraded its own scene late in the generation pipeline, causing its accuracy score to collapse from 33.9 percent to 4.4 percent.

The root problem lies in spatial evaluation: when prompted to select which of two versions more closely matched the reference image, the models' geometric judgments dropped near or below chance level. Essentially, the AI agents could not determine whether an update improved or ruined the 3D layout.

Fixing Evaluation with LEGO-Plugin

To address the agents' flawed visual self-assessment, the researchers developed LEGO-Plugin, an add-on that requires no fine-tuning. The plugin anchors the initial scene setup directly to the reference image, replaces the agent's subjective visual judgment with deterministic measurements, and prevents regressive edits from overwriting accurate geometry.

LEGO-Plugin improved the outputs across all six tested configurations. Weaker agents saw the most dramatic improvements, with gains reaching up to 62.7 percent. For the top-tier model, the plugin added roughly two percentage points to its already high baseline.

Downstream Applications and the 3D Ecosystem

The team also evaluated whether the code-based 3D reconstructions could assist downstream computer vision tasks without task-specific training. Because the scenes exist as programmatic environments, data such as depth maps, segmentation masks, and bounding boxes can be extracted automatically.

The reconstructed scenes delivered working results, though they still lag behind specialized vision architectures:

  • Object Detection: Achieved approximately half the performance of the dedicated model DINO.
  • Segmentation and Depth Estimation: Showed wider performance gaps when compared against specialized tools such as SAM 3 and Depth Anything 3.

The researchers note that while code-driven 3D generation is a promising direction, a substantial gap remains between generating a runnable script and achieving a geometrically faithful reconstruction.

The findings arrive as the broader AI sector rapidly pivots toward spatial understanding. AI researcher Yoav Artzi noted that GPT-6 Astra represents a significant step forward in spatial reasoning, likely aided by training on extensive 3D datasets such as Blender scripts. Meanwhile, commercial game engines are moving in parallel; Unity recently launched official integrations for coding assistants such as Claude Code and OpenAI Codex. Other research groups are exploring non-code alternatives, including World Labs' Atlas world model and Google DeepMind's GenCeption, which leverages video models to handle depth estimation and segmentation.

◗ Sources

The Decoder10/03

Related stories