Somewhere between hearing about AI agents and actually running one, an architecture practice hits a decision that gets almost no attention. Vendors talk about readiness. Consultants talk about phases and timelines. Almost nobody spends real time on the actual choice of what the agent touches first, and that choice quietly decides whether the whole effort earns its keep or gets shelved by spring.
The Decision Nobody Names
Most practices treat the first task as an afterthought, something picked in the last five minutes of a planning meeting because someone mentioned it in passing. RFI logging comes up. Submittal tracking comes up. Whatever got mentioned last tends to win, not because it's the best candidate but because it's the freshest one in the room.
That's backwards. The first task an agent handles becomes the evidence everyone in the studio uses to decide whether this whole approach is worth trusting. Pick well and the second task gets approved without a fight. Pick poorly and every future proposal starts from a deficit, even if the original failure had nothing to do with the technology itself.
Why Volume Beats Complexity
The instinct is to hand an agent the biggest, thorniest problem in the office. That instinct is almost always wrong. Complex tasks have too many edge cases, too many judgment calls, and too many people who feel ownership over how they're done. An agent that stumbles on a complex task in month one looks like it doesn't work, even when the underlying capability is sound.
High-volume, low-judgment tasks make better first candidates. Think of anything done the same way dozens of times a week: distributing incoming submittals to the right reviewer, chasing consultants for overdue responses, formatting meeting notes into a standard template. None of these are glamorous. All of them generate enough repetitions in the first two weeks to produce a real track record, good or bad, fast.
The Test That Actually Predicts Success
Before committing to a first task, ask one question: can someone describe how to do this correctly in under two minutes, without saying "it depends" more than once? If the answer is yes, the task has enough structure for an agent to handle reliably. If the answer is no, the task needs a human's judgment more than it needs speed, and it should wait.
This test filters out most of the tasks that sound impressive in a pitch meeting. Coordinating design intent across disciplines fails the test. Sorting incoming files by project number passes easily. The unglamorous option is usually the right one, at least for the first thirty days.
Where Practices Get This Wrong
The most common mistake is choosing a task for visibility rather than fit. A firm running three concurrent projects once picked its most public-facing coordination task as the pilot, reasoning that a visible win would build momentum. The task had too many exceptions and the rollout stalled within a month, not because the agent was incapable but because the task was never a fair test.
The second mistake is choosing a task nobody actually minds doing. If the current process barely bothers anyone, replacing it won't generate the kind of relief that gets people talking. The best first tasks sit at the intersection of high repetition and mild irritation, work that's tedious enough to want gone but structured enough to hand off cleanly.
A smaller, less visible task that works reliably will do more for adoption than a big, visible one that half works. Reliability compounds. A shaky first impression rarely does.
Once the first task is chosen and running, resist the urge to add a second one immediately. Give it two or three weeks of quiet operation first. The goal isn't to prove the concept in one dramatic swing, it's to build a small, boring pile of evidence that the studio can point to later, when the next decision, and the one after that, needs to get made.