Skip to content
Apixo
Blog
news· 4 min read· via MIT News AI

InstructMesh: MIT and Google Researchers Use GPT-4 to Fix AI-Generated 3D Models

MIT, Google, and Northeastern University researchers have developed InstructMesh, an AI tool that lets users repair flawed 3D prints using natural language.

InstructMesh: MIT and Google Researchers Use GPT-4 to Fix AI-Generated 3D Models

Generative artificial intelligence has made significant strides in creating 3D models from simple text or image prompts, but translating these digital designs into functional, real-world objects remains a challenge. Often, AI models capture the visual aesthetics of an object without understanding its physical mechanics. This disconnect frequently results in impractical fabrications, such as a 3D-printed coffee mug that cannot hold liquid. To address this, researchers from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), Google, and Northeastern University have introduced "InstructMesh," a novel design interface that enables users to easily edit and repair AI-generated 3D models before physical fabrication.

Bridging the Gap Between Design and Functionality

Traditional 3D modeling tools require extensive technical expertise, making it difficult for novices to correct errors in AI-generated files. InstructMesh solves this by allowing users to interact with 3D blueprints using a combination of natural language prompts and precise physical controls. By highlighting specific portions of a 3D model, users can describe the adjustments they want to make, and the software translates these instructions into geometric modifications.

The system achieves this capability by pairing Microsoft’s TRELLIS, a model designed to generate 3D structures from textual and visual inputs, with OpenAI’s GPT-4. While TRELLIS provides the foundational 3D generation trained on vast datasets of physical shapes, GPT-4 supplies the reasoning and language capabilities needed to interpret user edits. According to Faraz Faruqi, lead author of the research paper and a recent CSAIL affiliate, the goal was to unite the spatial capabilities of 3D generators with the reasoning power of large language models in an interactive environment. By manipulating the latent space of the generative model, InstructMesh allows users to evaluate and approve geometric changes based on simple descriptions of the issues.

Proving Usability with Real-World Testing

To evaluate the practical utility of InstructMesh, the research team tested the system on popular 3D models from Thingiverse, an online platform for printable designs. When the researchers recreated these models using the TRELLIS generator alone, they discovered that approximately 80 percent of the resulting files contained structural flaws that would prevent successful printing or usage.

The team then asked novice users to identify and correct these structural issues using InstructMesh. Despite having no prior experience in 3D modeling, the participants successfully identified and repaired the flaws about 90 percent of the time, as verified by an expert evaluator. Users reported that the system made it easy to express creative ideas, utilizing sliders to execute precise adjustments like extruding or enlarging specific components.

To demonstrate the tool's versatility, researchers fabricated several highly creative items. These included a mug wrapped in a dragon design where the tail serves as the handle, a shell-shaped whistle, butterfly-wing glasses, and a functional octopus-shaped drink dispenser that distributes liquid into multiple cups simultaneously. They also designed a denim-textured knee brace to match a patient's clothing, and a "bristle bot"—a colorful shrimp-shaped robot containing an internal motor that slides across flat surfaces when powered on.

What it means for developers

For software engineers and AI developers, InstructMesh highlights a shifting paradigm where multimodal AI systems move beyond purely virtual outputs and into physical engineering. The integration of LLMs with specialized spatial models demonstrates how developer teams can leverage latent space manipulation to make complex, domain-specific tools accessible to the general public.

Building applications of this scale requires access to a variety of state-of-the-art foundation models. Developers looking to build or experiment with similar multimodal pipelines can access top AI models cheaply through a single API key at https://apixoai.online, which provides streamlined access to engines like GPT-4, Claude, and Gemini. By lowering the barrier to entry for model orchestration, developers can focus on refining user interfaces and spatial reasoning algorithms rather than managing multiple API subscriptions.

Future Directions and AR Integration

The research team, which includes senior author Stefanie Mueller of MIT CSAIL, alongside co-authors from Google and Northeastern University, will present their findings at the ACM Symposium on User Interface Software and Technology.

Looking ahead, Faruqi, who now works at Google, envisions expanding InstructMesh into augmented reality (AR) environments. In an AR setup, users could prompt the system using physical context from their immediate surroundings to quickly generate and print customized items, such as a phone case designed to match a physical wallet. Future iterations of InstructMesh may also incorporate physics simulations to predict how materials will behave under stress, ensuring a design will not break during everyday use, and integrate the newer TRELLIS.2 model to enable even finer geometric adjustments.


Source: New tool lets users repair AI-generated 3D models, then fabricate them just the way they want — MIT News AI. Written by the Apixo team from that report.

#ai-news#ai#3d-printing#instructmesh#mit#gpt-4
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading