AI agents build 3D scenes from photos but have no idea if they got it right
First reported by The Decoder ·
AI can now generate editable 3D scenes from single photos, but their geometric accuracy is not yet reliable enough for precise applications.
Researchers from the University of Maryland and AWS have developed LEGO-Anything, a system where AI coding agents can construct editable 3D scenes from a single photograph. The process, termed "Image-to-Code," involves the agent iteratively writing and refining Blender code. This approach generates explicit representations of objects, geometry, and scene layout, allowing the output to be inspected and modified like any other program. To evaluate performance, the team created LEGO-Bench, a benchmark with 208 images from simulated scenes designed to provide precise ground truth without sacrificing realism. The benchmark assesses scene validity, geometric accuracy, and visual similarity. While the AI agents consistently produced usable scenes, their geometric accuracy remained a significant challenge, with top models like GPT-6 Astra achieving only about 53 percent accuracy on indoor scenes and less on outdoor ones. The system's self-assessment capabilities were found to be unreliable, often failing to correctly identify improvements in geometric accuracy.
The key limitation identified is the AI's inability to accurately self-assess its own geometric reconstructions. This suggests that future progress in 3D scene generation will likely depend on integrating external, objective measurement systems rather than relying on the AI's internal judgment. The LEGO-Plugin, an extension that uses concrete measurements, demonstrated significant improvements by anchoring the scene and preventing regressive edits, indicating a path forward for more robust AI-driven 3D content creation.
This development highlights a critical gap between generating functional 3D representations and achieving faithful, accurate reconstructions, impacting fields that require precise spatial data. While current AI-generated scenes show promise for basic vision tasks, their performance in areas like object detection, segmentation, and depth estimation lags behind specialized models. This signals that while AI can automate initial scene creation, human oversight or more advanced, task-specific AI will be necessary for applications demanding high fidelity and accuracy.
AI-written summary. May contain errors.