A Blindfold, Not a Ceiling: Why Today's Models Underperform in Engineering
[
]
A Blindfold, Not a Ceiling
Think of the best engineers that you know. Give them one billion dollars and tell them "design a plane." Come back in 10 years.
Whatever they've built will probably be really cool. It also probably won't be what you were looking for.
Maybe they built a B-2, but you were just looking for a Cessna. Maybe they built something that cruises at Mach 3, and you needed something that can loiter over a field for six hours. Maybe they built an elegant two-seater, and you needed to move four hundred people across an ocean.
You would not conclude from this experiment that these are incompetent engineers. You would conclude that you gave them an underspecified task.
We do not extend the same courtesy to large language models. This example is hyperbolic in scale, but it is an accurate description of how we treat models today. Instead of concluding that the problem lies in how we manage and distribute information, we attribute poor model performance to limitations in the models themselves.
Nobody onboards an engineer this way
Today's models can be much better than most engineers expect, but very few people have watched one work with anything close to the amount of information a human engineer would have. Think about what happens when a new engineer joins your company. You onboard them to your internal processes. You teach them the standards in your industry. They learn to use new tools. You put them with a team of more experienced people so they can learn how to reason through the problems specific to your product.
On the other hand, models are usually just handed a few sentences, a screenshot, and maybe a CAD file. This may make for an impressive demo, but in a real production environment these models wouldn't stand a chance. They can't walk across the room and ask the engineer about the part they designed last April. They don't know the implicit requirements that nobody wrote down in your spreadsheet. They don't hear your call with a supplier or read the email thread after it. They don't see the PowerPoint full of design feedback. They don't have the context or the tools that an engineer would.
What we read as a ceiling is really a blindfold, and models never perform at the level they could. Today's engineering organizations were not built for an agent to operate in. The tools we use were never built to gather context in one place a model can reach.
The harness is the variable
Ultimately, the performance of a model is, to some extent, a function of the task you give it and the context and tools it has to do that task.
The advantage of LLMs over other types of models is their ability to process large amounts of unstructured context and generalize across a wide variety of tools and tasks. In both our own internal benchmarking and our work with model partners, we've seen what these models are capable of. But in order for them to work, they need to be harnessed properly.
One place for the context
Tandem is built from the ground up to assemble context in one place. That includes your documents, design reviews, and CAD, as well as context that previously went unrecorded, like what an engineer was thinking when they made a particular CAD choice.
Gathering all this in one place is a useful service on its own. It keeps requirements moving at the speed of design and puts every past decision within reach. But beyond that, it is what removes the blindfold. It is the difference between handing a model a screenshot and actually onboarding it.
-Faz
Keep Reading
[
AI / Agents
]
