A drawing review pitch usually leans on one image: an agent opens a set of construction documents and works through them the way a project architect would, catching the same clashes, flagging the same missing dimensions, skipping the same things a tired reviewer skips at six on a Friday. That image sells software well. It is also not what happens underneath the interface, and knowing exactly where the picture breaks down is what tells a firm which parts of a drawing review it can actually hand off, and which parts still need a person looking at the sheet.

The Myth: A Second Set of Trained Eyes

The pitch works because it borrows a mental model everyone already has. A junior architect walks a set sheet by sheet, comparing plan to section to schedule, and flags what does not line up. Sell the agent as a faster version of that junior architect and the value proposition writes itself. It also quietly implies the agent is looking at the drawing the way a person looks at it: taking in the whole sheet, understanding what a hatch pattern means next to a wall type tag, catching that a dimension string does not add up because the eye caught it mid glance.

The Mechanics: Extraction, Not Perception

What is actually happening is closer to indexing than looking. The agent pulls text layers out of a PDF, reads embedded vector data where the file has it, runs optical character recognition on anything that is a raster image, and pulls whatever metadata the authoring software left behind: sheet names, revision tags, layer names if the export kept them. It builds a structured record of what the file contains, then reasons over that record. It does not perceive a drawing as a drawing. It perceives a pile of text strings and coordinate data with a picture attached that it mostly never looks at directly.

That distinction matters most on anything that lives only in the graphic itself. A hand annotation. A revision cloud with no accompanying note. A symbol defined once in a legend on sheet A0.1 and never repeated anywhere else in the set. A human reviewer carries that legend around in their head without thinking about it. The agent has to be told to go find it, every single time, on every sheet.

Where the Gap Shows Up on a Real Set

The gap is not evenly spread across a drawing set. Typed schedules, keynote lists, and door and window tables extract cleanly and get checked reliably, often faster and more consistently than a person checks them, because the data was already structured before the agent ever touched the file. Redline scans from a consultant, handwritten field notes photographed into a submittal, and drawings exported as flattened images with no text layer left in them are a different story. The agent either fails loudly and returns nothing useful, which is at least honest, or fails quietly and reports a clean review that only looks clean because it never actually saw half the sheet.

That second failure mode is the expensive one. A loud failure gets caught in the first week. A quiet one shows up during construction.

A Coordination Check, Worked Through

One mid-sized practice piloted an agent against a few hundred sheets of a renovation set, checking door hardware schedules against plan notes. It caught inconsistencies between the schedule and the notes in under an hour, a task that had taken a staff architect most of a day the month before. The same pilot missed a hardware conflict that existed only as a handwritten note on a scanned consultant redline, because that note lived inside a raster image with no text layer, and nobody had told the agent to run character recognition on that particular sheet. Nobody had thought to, because nobody had asked what format that sheet actually was.

The Question to Ask Before the Pilot Starts

The fix here is not more trust in the tool. It is a better question asked earlier. Before handing any drawing review task to an agent, find out what format every sheet in the set actually is: native vector data, a flattened export, or a scan. Ask which sheets carry information only in the graphic, with no text anywhere backing it up. That one question, asked before a pilot starts rather than after it produces a confidently wrong answer, does more to set real expectations than any claim about how closely the tool reads like a person standing at a light table.