An agent flags a clash between a mechanical run and a structural beam, or drafts a code compliance note for an egress path, and the note is wrong. Not catastrophically, not in a way anyone would catch on a quick look, just wrong enough that it ships into a set and someone downstream builds from it. The practice that hired the agent has a question to answer that none of the vendor demos ever raised: whose mistake is that.
The Stamp Doesn't Move
The architect of record carries the liability for a stamped set regardless of what produced any given line in it. A licensing board does not have a category for "the tool was wrong," and an E&O claim does not get weaker because a human reviewed the output of software rather than the output of a junior drafter. This is not a new problem invented by AI agents. It is the same rule that already governed CAD blocks, spec templates, and interns, applied to a tool that happens to write in full sentences and sound confident while doing it.
What is new is how easy that confidence makes it to skip the review step the rule assumes is happening. A junior staffer's redline gets read with some skepticism built in. An agent's output reads like it already checked itself, and a tired reviewer at the end of a long week is the person most likely to wave it through on that impression alone. The stamp still moves with the human who signed, but the actual risk moves with however seriously that human treated the review.
What the Insurer Actually Wants to See
An E&O carrier asking about an AI agent after a claim is not asking whether the agent was any good. It is asking whether the practice can produce a record: what the agent flagged, what a named person did with that flag, and when. A firm that can pull up a log showing an agent surfaced a clash and a project architect dismissed it with a documented reason is in a defensible position even if that dismissal turns out to be wrong. A firm that can only say "the agent checks things like that" is not.
Most practices running an agent today do not have this log, because nobody asked for it until the first claim made it obvious it should exist. Building it after the fact is possible but unconvincing. Building it as part of the rollout, as a plain requirement that every agent flag and every human response gets a timestamp and a name, costs almost nothing and changes the entire conversation with an insurer later.
The Question Nobody Has Answered in the Contract
Read the liability clause in most AI agent vendor agreements and it caps the vendor's exposure at the subscription fee paid, sometimes for the whole contract term, regardless of what the error cost downstream. That is a standard software liability clause, not a targeted decision about this tool, and it means the vendor relationship offers close to no protection if an agent's output causes real damage. The practice absorbs that risk whether or not anyone read the clause closely at signing.
This is not an argument against using an agent. It is an argument for treating the vendor's terms as the floor, not the plan. The actual protection has to come from inside the practice: a written internal protocol naming who reviews what an agent produces, what counts as sign-off, and what triggers a second look before anything goes into a stamped set. A mid-sized practice running this on three concurrent projects put that protocol in writing after a near miss on a fire-rating note, not before, which is the ordinary way firms learn this lesson.
None of this argues for slowing an agent down or trusting it less than it has earned. It argues for writing down, before the first real mistake, who is supposed to catch it and how anyone would know they did.