An architecture practice bringing in an AI agent usually starts by asking what the tool can take off someone's plate. That's the right first question, but it skips a second one that matters just as much: what the agent is structurally unable to do, no matter how well it's trained or how much data it sees.
Knowing that boundary ahead of time saves a practice from two separate failure modes. One is handing off something that should never leave a licensed architect's desk. The other is holding back a task the agent could have safely absorbed months ago, out of a vague worry that it belongs in the first category too.
Stamped Decisions Stay With a Licensed Architect
Anything that ends with a stamp is off the table by definition. An agent can draft a note, flag a code section, or assemble a set of options, but the decision that a design meets code and life-safety requirements has to be made, and owned, by a licensed professional. That isn't a limitation a better model is likely to lift. It's a structural feature of how the profession assigns liability, and it means the agent's output in these areas is always a draft for review, never a final answer on its own.
A mid-sized practice we talked with put the boundary bluntly: the agent can shorten the path to a decision, but it can't be the thing that made the decision. Hold that line and the tool feels safe to expand into new areas. Blur it, and every use of the agent starts a conversation about who's accountable if it turns out wrong.
Client Conversations an Agent Shouldn't Run
A client meeting carries information that never makes it into a written brief: hesitation before answering a budget question, a raised eyebrow at a proposed material, an offhand comment about a board member who will need convincing later. An agent summarizing that meeting from a transcript catches the words. It misses the room. For anything where the read on a person matters as much as the content of what they said, the agent belongs in the loop as a note-taker, not as the one interpreting what happened.
Design Judgment on Genuinely New Problems
Agents built on pattern recognition are strongest where architecture is most repetitive: code checks, RFI language, submittal logs, schedule tracking. They get weaker fast once a project stops resembling anything in its training data. An unusual easement, a client program that doesn't map to a standard building type, a structural condition nobody on the team has solved before, these are the moments a firm pays a senior architect for, and the moments where an agent's confident-sounding answer is least trustworthy.
The output can still be useful as a starting list of precedents to rule out. It just can't be the judgment itself.
Messy Inputs From the Field
An agent works well on structured, legible input: a spreadsheet, a form, a clean PDF export. Construction sites don't produce structured input. A superintendent's handwritten redline, a phone photo of an as-built condition, a verbal change relayed through three people before it reaches the office, these arrive incomplete, and an agent asked to reconcile them against drawings will either guess or stall. The fix in most firms isn't a smarter agent. It's a person who cleans up the field data first.
Deadline Pressure Doesn't Move the Line
The boundary gets tested hardest right before a deadline, when someone reaches for the agent to close a gap it was never built to close. A submittal review compressed from a week to a day is exactly when a firm is tempted to let the agent make a call it should only be assisting with. The tasks that belong to a person don't become safer to hand off just because the schedule got tighter. If anything, a compressed schedule is when a wrong judgment call costs the most.
Where the Boundary Moves, and Where It Doesn't
Some of these limits are temporary. Agents that struggle with messy field photos today may handle them fine in eighteen months, and firms that hesitated a year ago on RFI tracking are already watching that limitation turn into a routine task. Other limits aren't going anywhere. A stamp will keep requiring a license behind it regardless of what the model can do, and a client's trust in being heard by a person isn't a capability gap that better training closes.
The practical move is to sort a practice's own limitation list into those two piles before deciding what to automate next.
A practice that maps this boundary once, honestly, gets more out of an agent over time than one that keeps discovering it by accident, usually after a client conversation went sideways or a redline got misread. The tool doesn't need to do everything to be worth the investment. It needs to do the right half.