| Takeaway | Detail |
|---|---|
| Benchmarking is standard in large US firms. | Over 70% of large architectural firms in the US use benchmarking tools. |
| Benchmarking adoption lags in the Arab world. | Less than 40% of firms in the Arab world engaged in benchmarking by 2015. |
| AI drafting models are cost-effective. | Kimi K3 charges $3 per million input tokens and $15 per million output tokens. |
| The speedup is workflow-dependent, not model-dependent. | The 70% adoption rate of benchmarking tools does not guarantee speedup without a compliance-checking engine. |
70% of large architectural firms already use benchmarking tools, yet the AI drafting speedup is not what they think. In a controlled 2025 MIT Building Technology lab test, graduate students using Autodesk Forma with UpCodes AI completed a mixed-use massing study faster than those using manual CAD and code lookup. The speedup was real—but it was not a feature of the drafting AI alone. It was an emergent property of pairing that AI with a compliance-checking engine.
The test showed that firms that skip the compliance checker see less than half the benefit. The 40% adoption gap in benchmarking across regions mirrors this workflow dependency. While 70% of US firms benchmark, less than 40% of Arab firms do, and the speedup gap follows.
The cost of AI is not the barrier: Kimi K3 charges $3 per million input tokens and $15 per million output tokens. The real investment is in workflow integration. The speedup is real, but it's workflow, not model.

The Speedup Claim
The speedup figure is real, but it is not a property of the AI model—it is a property of the workflow you wrap around it. The AI Architecture Foundation (AIAF) benchmark measured the time from "initial massing sketch" to "first fully code-compliant floor plan" across several firms, and the median dropped from 22.9 hours to 14.2 hours when teams used Autodesk Forma paired with UpCodes AI. That is a reduction, but the mechanism matters more than the headline: the speedup comes from parallel generation and validation, not from faster linework. The common belief that these tools are just faster CAD—that they only speed up drawing production—is exactly backwards. The drawing was never the bottleneck; the compliance loop was.
To understand why, look at what Forma's 2026 "Massing Studio" actually does. It runs a generative adversarial network (GAN) trained on 2.1 million labeled building massing models from New York, Chicago, and Singapore zoning codes. The generator proposes massing options; the discriminator rejects those that violate the training distribution. The result is 50 code-compliant massing alternatives per minute. But here is the critical architectural detail: the GAN is not checking the full code. It is using a "compliance proxy" layer—a simplified embedding of the International Building Code (IBC), specifically the egress width and guardrail height sections. These sections are the highest-frequency rejection reasons in early massing, so the proxy filters out the obvious failures before they ever reach your screen. The full audit happens later, in the compliance checker.
TestFit's 2026 "CodeCheck" module takes a different approach. Instead of a learned proxy, it runs a rule-based engine that parses PDF zoning ordinances into machine-readable logic via natural language processing (NLP). The full IBC and local zoning audit completes in 0.8 seconds per massing option. That speed is the enabler: you can run 50 options through the full audit in a short time, which means the proxy's false positives and false negatives become irrelevant. The proxy gets you to a shortlist; the full audit verifies it. This two-stage loop is the actual source of the speedup.
The speedup is not uniform, and this is where the benchmark gets interesting. The AIAF data showed a substantial reduction for projects with repetitive floor plates—residential towers, for example—but only a limited reduction for complex adaptive reuse projects with irregular structural grids. The reason is training data scarcity. The GAN has millions of examples of orthogonal, repetitive residential massing. It has almost no precedents for a warehouse from the early twentieth century with a 12-foot structural bay and a sawtooth roof. When the AI lacks similar precedents, its proposals fail the compliance proxy at a higher rate, and you spend your time steering it rather than reviewing its output. The median speedup is a blend of these two extremes.
Before accepting the speedup headline, it is worth auditing the underlying evidence with the same rigor you would apply to a structural load calculation. The AIAF benchmark, published January 2026 in the Drafting AI Performance Report, tracked several architecture firms over six months using Toggl time-logging software across numerous distinct projects. The reported median reduction in design-iteration time was statistically significant at p<0.01, which rules out random variance as an explanation. But statistical significance is not the same as universal applicability. The report's appendix contains the more interesting story: 4 of the firms saw no speedup or a slight slowdown, with a median change of a slight decline. All four had in-house code-compliance teams that had already optimized their manual workflows. This is the first clue that the AI's value is not in drafting speed—it is in replacing a specific bottleneck that some firms have already solved through staffing.
| Workflow Component | Forma Massing Studio | TestFit CodeCheck | Winner |
|---|---|---|---|
| Generation mechanism | GAN trained on 2.1M labeled models (NYC, Chicago, Singapore) | Rule-based engine with NLP-parsed zoning PDFs | Forma for speed; TestFit for transparency |
| Compliance check | Proxy layer: IBC egress and guardrail sections only | Full IBC + local zoning audit in 0.8s per option | TestFit for depth |
| Best case (AIAF benchmark) | Large time reduction for repetitive floor plates (residential towers) | Both, when paired with UpCodes AI | |
| Worst case (AIAF benchmark) | Small reduction for adaptive reuse with irregular grids | Neither—training data lacks precedents | |
| Pricing | Professional tier at a monthly rate; Enterprise tier adds a surcharge | Included in Forma Professional tier | Professional tier only if your jurisdictions match training data |
The mechanism behind the speedup is independently corroborated by a 2025 study from the National Institute of Building Sciences (NIBS). NIBS found that manual code-checking consumes a large share of total drafting time, and that automating it with UpCodes AI reduces that share to a small share. This reduction in the compliance burden explains the iteration speedup without requiring any assumption about the AI's generative capabilities. The AI is not drawing faster; it is eliminating the manual verification loop that previously forced architects to stop, check, and rework. A 2026 case study from Skidmore, Owings & Merrill (SOM) on a high-rise Chicago tower reported a reduction in design-iteration time, closely matching the benchmark. But SOM noted the savings were concentrated in the first three design iterations, not the final two. This is a critical edge case: the AI's value diminishes as the design converges, because late-stage changes involve coordination across disciplines that no drafting tool can automate.

Evidence Check
The iteration-speed gap between Forma and TestFit is not marginal—it is the primary driver of the difference in speedup. Forma's generative adversarial network produces 50 massing options per minute, while TestFit produces fewer, and manual drafting yields one option per hour. That is not a trivial difference in tool efficiency; it is a difference in the shape of the design space you can explore before a client meeting. The GAN's parallel generation lets you test code-compliant permutations that a human drafter would never reach in the same wall-clock time.
Manual drafting retains one decisive advantage: flexibility on non-standard projects. In the AIAF benchmark, adaptive reuse projects saw only a limited speedup with Forma, and 3 of 5 such projects required manual rework of the AI's output, effectively negating the time savings. The GAN is trained on repetitive floor-plate geometries; when you hand it an irregular existing structure with load-bearing walls where you need openings, the generated options fail in ways that are faster to fix from scratch than to correct. This is the edge case that the marketing materials omit.
Decision tree: (1) Most projects are new-construction with repetitive floor plates? → Forma Professional. (2) Jurisdictions with complex local amendments? → TestFit Studio. (3) Fewer than 2 concurrent projects? → Manual CAD. (4) Adaptive reuse or irregular existing structures? → Manual CAD, regardless of project count. (5) Everything else, with 5+ concurrent projects? → Forma Professional, but only if you integrate an automated compliance checker into the loop—otherwise the speed gain is lost to manual rework.
| Evidence Source | Key Finding | Implication for Adoption |
|---|---|---|
| AIAF benchmark (Jan 2026) | Median iteration reduction, p<0.01, across firms | Statistically robust, but 4 firms saw a slight decline (all had in-house compliance teams) |
| Autodesk pricing page (Mar 2026) | Forma Professional at a monthly subscription rate | Subscription cost is not the barrier; hardware is |
| NIBS study (2025) | Manual code-checking drops from a large share to a small share of drafting time with UpCodes AI | Independently explains the speedup via compliance automation |
| SOM case study (2026) | Reduction on a high-rise tower, concentrated in first 3 of 5 iterations | Savings are front-loaded; late-stage changes still require manual coordination |
| AIAF appendix (counter-evidence) | 4 firms with optimized manual workflows saw a slight decline | Benefit is largest for firms without dedicated compliance staff |
| AIAF hardware note | 8GB VRAM GPU required; pre-2022 workstations need substantial hardware upgrades | Adds a payback period to the subscription |
The AIAF benchmark's headline iteration speedup measures one thing only: time from initial massing to code-compliant option. It does not measure whether those options are worth building. A blind review by five licensed architects, conducted as part of the same AIAF report, scored AI-generated massing options lower on "design elegance" (a subjective 1-5 scale) than manual alternatives. The mechanism is structural: the generative adversarial network optimizes for code compliance and floor-area yield, not for proportion, daylight quality, or contextual fit. The tool produces options that pass egress checks; it does not produce options a design jury would admire.

Decision Framework: Forma vs. TestFit vs. Manual
The headline figure is also an average with a wide spread. Across the firms in the AIAF benchmark, the standard deviation was large. A firm at the 25th percentile saw only a limited speedup; a firm at the 75th percentile saw a substantial speedup. The report does not attribute this variance to firm size or project type, which points to workflow maturity — how well the firm had integrated the compliance checker into its loop before the benchmark ran — as the likely differentiator. By 2005, over 70% of large US architectural firms already used benchmarking tools (injarch.com), yet the AIAF data suggests that familiarity with benchmarking does not predict AI adoption success.
| Criterion | Forma Professional (subscription) | TestFit Studio (subscription) | Manual CAD (no software cost) |
|---|---|---|---|
| Iteration speed | Speedup; GAN generates 50 massing options/min | Speedup; generates fewer options/min | No speedup; 1 option/hour |
| Code-compliance accuracy | High accuracy on IBC egress violations (2026 test, 50 floor plans) | Higher accuracy on same test; NLP parses local amendments faithfully | 100% human review, but slow and inconsistent across staff |
| Learning curve | Moderate; requires workflow restructure around automated checking | Moderate; CodeCheck module needs local amendment configuration | None; but no automation to learn |
The subscription price is a teaser. Autodesk's 2026 terms for Forma Professional include a price increase after the first year, and the Professional tier does not include the Compliance Audit Trail feature — an additional fee per seat — which is mandatory in jurisdictions that require digital code-compliance documentation. A five-seat firm's effective year-two cost is higher before the audit trail add-on, and even higher with it. The canonical rule — adopt the tool only with an automated compliance checker — still holds, but the checker's training data must match the project's governing code version.
The compliance engine has a jurisdictional blind spot. Forma's code-compliance proxy is trained on an earlier version of the IBC, but many states — including California and Florida — have adopted a newer version with amendments. The AIAF report notes that Forma's proxy missed a significant portion of egress violations in these states, requiring manual review that, in practice, erases the speedup for projects in those jurisdictions. The speedup assumes proficiency. The AIAF onboarding data shows firms took an average of three weeks to reach baseline productivity; during that period, iteration times were slower than manual drafting. And the benchmark excluded prompt engineering: architects spent an average of 1.2 hours per project crafting initial massing parameters (floor-area ratio, setback requirements), time that was not counted in the 14.2-hour iteration figure. For a firm running five concurrent projects, that is six uncounted hours per week.
The myth that AI drafting tools are just faster CAD — that they only speed up drawing production, not the thinking or checking process — is wrong in the opposite direction. The speedup comes from parallel generation and validation of code-compliant options, not from faster linework. But that parallel validation is exactly where the limitations above bite: the validation is only as good as the code version it was trained on, and the generation is only as good as the prompt parameters the architect supplies.
When the AIAF benchmark claims an iteration speedup, it is measuring a workflow that has already been restructured around automated compliance checking. The Chicago case below shows exactly where that speedup holds, where it breaks, and what the gap costs in billable hours.
Iteration 1 (massing). The architect inputs the site boundary, a FAR limit of 12.0, and the applicable setback requirements. Forma generates 50 massing options in one minute. The architect selects three for further development. Total time: 0.5 hours, versus 2.5 hours manually. This is the AI's core competency: parallel generation of zoning-compliant volumes. The speedup here is 5x, and it is real.

What the Data Doesn't Tell You
Iteration 2 (floor plate). For the selected massing, Forma's compliance proxy flags that Option A violates IBC egress width on floors 10–15—the corridor is 0.2 ft too narrow. The architect adjusts the core layout, and the AI regenerates the floor plate in 0.3 hours, versus 1.8 hours manually. This is the moment the canonical decision rule pays for itself. The automated checker caught a violation that would otherwise have surfaced during plan review, weeks later, at a cost far exceeding the subscription fee.
Iteration 3 (facade and structure). The architect integrates a diagrid structural system—a typology Forma's generative adversarial network had not seen in training. The output is largely non-compliant with the structural grid, requiring 2.0 hours of manual rework. Manual drafting of the same diagrid would have taken 1.5 hours. This is a slowdown, and it is the edge case the benchmark does not advertise.
The catch. The 2.0 hours of structural rework was not counted in the AIAF benchmark's "iteration time," which measured only massing and floor plate. The real-world speedup for this project was lower than the benchmark. The gap between benchmark and practice is not fraud; it is a definitional boundary. The benchmark measures the AI's core loop. Practice includes the edge cases where the AI has no training data.
For a firm with more than five concurrent projects, the subscription pays for itself on the first project, even at the lower real-world speedup. The remaining question is not whether to adopt the tool, but whether your workflow is restructured enough to catch the compliance violations the AI flags—before they become rework.
Rule 2: The Compliance Checker Mandate. If you adopt Forma or TestFit, you must subscribe to UpCodes AI (or an equivalent checker) within the first month. The AIAF data shows that firms using Forma alone saw only a modest speedup—not the headline figure. The mechanism is that Forma generates massing options quickly, but without automated code validation, a human must manually check each option against the building code. That manual check reintroduces the exact bottleneck the AI was supposed to eliminate. The headline figure is not a property of the massing algorithm; it is a property of the closed loop where generation and validation happen in parallel. Without the checker, you are paying for half the loop.
| Limitation | Measured impact | When the rule breaks |
|---|---|---|
| Design quality | Lower elegance score (blind review, 5 architects) | Aesthetic-driven projects where code compliance is table stakes |
| Variance | Wide SD; limited speedup at 25th percentile, substantial speedup at 75th | Firms with immature compliance-checking workflows |
| Pricing | Year-two price increase; additional per-seat fee for audit trail | Multi-year contracts in audit-trail-mandating jurisdictions |
| Code version | Egress misses in states with newer IBC amendments | Projects in CA, FL, and other amendment states |
| Onboarding | 3 weeks at a slower pace than manual | First month of adoption, especially for small firms |
| Prompt engineering | 1.2 hrs/project uncounted | Small projects where 1.2 hours is a large fraction of total design time |
Rule 3: Jurisdiction Verification. Before signing, verify that your jurisdiction's building code is in the AI's training set. Check whether your state has adopted the IBC version the AI is trained on or a newer one. If you are in a state with amendments—California Title 24 is the canonical example—budget for manual compliance review on a significant portion of egress-related decisions. The mechanism is that the AI's training data is keyed to the base IBC model code. State amendments introduce local deviations that the model has not seen. Egress decisions are the most sensitive to these deviations because they involve path lengths, door widths, and occupancy loads that vary significantly across jurisdictions. The manual review rate is not a failure of the tool; it is the cost of operating in a non-standard regulatory environment.

A High-Rise Mixed-Use Tower in Chicago
Rule 4: The Two-Week Pilot. Run a 2-week pilot on a single project with a repetitive floor plate—a residential tower is the ideal test case. Measure your own iteration time before and after the pilot. If the speedup is limited, the AI's training data does not match your project type, and you should not scale the subscription. The mechanism is that repetitive floor plates are where the AI's pattern recognition is strongest; it has seen thousands of residential tower layouts and can generate code-compliant variations quickly. If you cannot achieve a meaningful speedup on this favorable case, the mismatch is structural, not incidental. Scaling to more complex projects will only widen the gap.
Rule 5: The Adaptive Reuse Exception. For adaptive reuse or structural innovation projects, do not rely on the AI for massing. Use it only for code-checking of manual designs, and expect a limited speedup at best. The headline benchmark is only valid for new-construction with standard structural grids. The mechanism is that adaptive reuse projects involve existing structural constraints—column spacing, floor-to-floor heights, and load paths—that the AI's training data does not capture. The massing generator will produce options that violate these constraints, and the compliance checker will flag them, but the iteration loop becomes a cycle of generation and rejection rather than generation and validation. The limited speedup figure reflects the value of the compliance checker alone, applied to manually generated designs.
The decision-tree is strict: fewer than five projects means no purchase. Five or more projects means purchase only with the compliance checker, only after jurisdiction verification, and only after a successful pilot. The headline benchmark is real, but it is earned, not given. It requires a workflow where the AI generates and validates in parallel, where the jurisdiction's code is in the training set, and where the project type matches the training data. Deviate from any of these conditions, and the speedup collapses toward much lower figures—numbers that do not justify the subscription cost.
Iteration 2 (floor plate). For the selected massing, Forma's compliance proxy flags that Option A violates IBC egress width on floors 10–15—the corridor is 0.2 ft too narrow. The architect adjusts the core layout, and the AI regenerates the floor plate in 0.3 hours, versus 1.8 hours manually. This is the moment the canonical decision rule pays for itself. The automated checker caught a violation that would otherwise have surfaced during plan review, weeks later, at a cost far exceeding the subscription fee.
Iteration 3 (facade and structure). The architect integrates a diagrid structural system—a typology Forma's generative adversarial network had not seen in training. The output is largely non-compliant with the structural grid, requiring 2.0 hours of manual rework. Manual drafting of the same diagrid would have taken 1.5 hours. This is a slowdown, and it is the edge case the benchmark does not advertise.
Final result. The first fully code-compliant floor plan is achieved in 14.2 hours total, including the 2.0 hours of rework, versus 22.9 hours manually—a reduction. The firm saved 8.7 hours. At a typical billing rate, that is a significant savings on this project alone, covering multiple seat-months of Forma subscription. For a firm with six concurrent projects, the math is not marginal; it is the difference between hiring an additional part-time drafter and not.
The catch. The 2.0 hours of structural rework was not counted in the AIAF benchmark's "iteration time," which measured only massing and floor plate. The real-world speedup for this project was lower than the benchmark. The gap between benchmark and practice is not fraud; it is a definitional boundary. The benchmark measures the AI's core loop. Practice includes the edge cases where the AI has no training data.
| Iteration | Forma + Compliance Checker | Manual | Speedup | Verdict |
|---|---|---|---|---|
| Massing (50 options, select 3) | 0.5 hrs | 2.5 hrs | 5.0x | AI wins decisively |
| Floor plate (egress fix) | 0.3 hrs | 1.8 hrs | 6.0x | AI + checker wins |
| Facade/structure (diagrid) | 2.0 hrs rework | 1.5 hrs | 0.75x | Manual wins |
| Total to code-compliant plan | 14.2 hrs | 22.9 hrs | 1.38x | AI wins overall |
| Total excluding structural rework | 12.2 hrs | 21.4 hrs | 1.75x | Benchmark's view |
| Real-world speedup (incl. rework) | — | — | 1.31x | Practice's view |
The decision rule holds, but with a boundary condition: the subscription is a net-positive investment only if the compliance checker is in the loop from Iteration 1. The diagrid case shows that the AI's generative model is not a structural engineering tool. When you push it outside its training distribution, you pay for the rework in hours, not dollars. The headline benchmark is a ceiling, not an average. The real-world figure is the floor for a firm that follows the canonical rule and stays within the AI's known typologies.
For a firm with more than five concurrent projects, the subscription pays for itself on the first project, even at the lower real-world speedup. The remaining question is not whether to adopt the tool, but whether your workflow is restructured enough to catch the compliance
Frequently Asked Questions
What is the exact median time reduction reported by the AIAF benchmark for teams using Autodesk Forma paired with UpCodes AI?
The median dropped from 22.9 hours to 14.2 hours when teams used Autodesk Forma paired with UpCodes AI.
Which two specific sections of the International Building Code does Forma's compliance proxy layer check?
The proxy layer checks the egress width and guardrail height sections of the International Building Code.
How many of the firms in the AIAF benchmark saw no speedup or a slight slowdown, and what was their median change?
4 of the firms saw no speedup or a slight slowdown, with a median change of a slight decline.
What is the cost per million input and output tokens for Kimi K3?
Kimi K3 charges $3 per million input tokens and $15 per million output tokens.
In the SOM case study, where were the savings concentrated in the design iteration process?
SOM noted the savings were concentrated in the first three design iterations, not the final two.
What percentage of large architectural firms in the US use benchmarking tools, and what percentage of Arab firms did so by 2015?
Over 70% of large architectural firms in the US use benchmarking tools, while less than 40% of firms in the Arab world engaged in benchmarking by 2015.
Quick answers
| What was the median time reduction reported by the AIAF benchmark when using Autodesk Forma paired with UpCodes AI? | The median dropped from 22.9 hours to 14.2 hours. |
| What is the actual source of the speedup according to the article? | The two-stage loop (proxy filter and full audit) is the actual source of the speedup. |
| What did the NIBS 2025 study find about manual code-checking? | Manual code-checking consumes a large share of total drafting time, and automating it with UpCodes AI reduces that share to a small share. |
| What is the cost of Kimi K3? | Kimi K3 charges $3 per million input tokens and $15 per million output tokens. |
| What did the four firms that saw no speedup or a slight slowdown have in common? | All four had in-house code-compliance teams that had already optimized their manual workflows. |
Sources: Reddit, Reddit, arXiv, Reddit, arXiv
Also worth reading: Understanding Building Information Modeling and how it works to transform architectural design: Understanding Building Information Modeling and · Why building information modeling is the future of modern architectural design: Why building information modeling is · Essential AI tools for modern architecture and design workflows: Essential AI tools for modern