| Takeaway | Detail |
|---|---|
| CAD-Coder generates editable CadQuery Python from visual input | The open-source VLM is explicitly fine-tuned end-to-end for CAD code generation. |
| Training used more than 163K CAD image-code pairs | The reported training dataset contains over 163,000 paired CAD model images and code. |
| CAD-Coder reported 100% syntax correctness in testing | The reported result applies to the paper's evaluation; benchmark scope should be checked before adoption. |
| Verify the complete live option before committing | Compare like-for-like totals and terms rather than relying on a partial listing or headline figure. |
CAD-Coder is an open-source VLM designed to turn visual CAD input into editable CadQuery Python, with training reported on more than 163K image-code pairs. The guide weighs its reported 100% syntax correctness in testing against a mandatory check of the live, complete option and like-for-like totals and terms.

How It Works
CAD-Coder follows an image-to-CAD-code pipeline. The CAD-Coder paper describes an open-source vision-language model, or VLM, fine-tuned to convert visual input into editable CadQuery Python. In operational terms, the model receives an image, generates a modeling script, and produces a three-dimensional solid when that script is executed with CadQuery. Verification should cover this entire path rather than stopping at whether the model returns plausible-looking text.
Several key terms define that path. A vision-language model links visual information with language-based output. Fine-tuning means training an existing general model further for a specialized task; CAD-Coder is fine-tuned specifically for CAD-code generation. Its GenCAD-Code dataset contains more than 163,000 pairs of CAD-model images and code. The target representation, CadQuery Python, is Python-based code for parametric CAD modeling, so its operations can be executed to create geometry rather than merely displayed as text.
The term editable output means the result remains source code that a user can inspect and modify. That property differs from a fixed exported mesh or image because the generated script exposes the modeling operations used to construct the solid. It establishes how the result can be changed, but it does not by itself establish that every dimension or feature agrees with the original drawing.
Evaluation terminology also needs to be kept separate. Valid syntax means the generated code follows the required code-format rules. The paper reports a 100% valid-syntax rate in its evaluation, but that result addresses syntactic validity under the reported test rather than geometric fidelity. 3D solid similarity refers to agreement at the level of the resulting three-dimensional forms. When reviewing any evaluation, check which term its percentage represents; a syntax rate and a solid-similarity score are different measures and should not be treated as substitutes.
Before committing, verify the public CAD-Coder repository against the complete workflow advertised by the paper: image input, model output, CadQuery execution, and an inspectable resulting solid. Check whether the required materials and setup instructions are present and runnable in your intended environment. For a like-for-like review, use the same kind of input, execution environment, pass-or-fail definition, and similarity metric for every result. Until the full path can be reproduced and its metrics are labeled correctly, the evidence remains a research result rather than a verified working option.

Key Factors to Consider
When you are deciding whether CAD-Coder belongs in an architectural workflow, three criteria carry most of the weight: whether the generated code actually runs, whether the output format fits your existing documentation pipeline, and whether the deployment terms hold up when you check them live. Everything else is secondary. The CAD-Coder paper (arXiv:2505.14646, submitted May 20, 2025) gives you numbers for the first criterion and a public repository for the other two, so the verification work is mostly a matter of knowing where to look.
The first criterion is output validity, and here the numbers are specific. The paper reports a 100% valid syntax rate on its test set. That figure refers to syntactically valid code, not geometric fidelity — the paper measures 3D solid similarity separately and reports the highest accuracy among the baselines it tested, including GPT-4.5 and Qwen2.5-VL-72B, without publishing a matching percentage for that metric. The check: when any tool vendor quotes a success rate, pin down which of the two metrics it describes before you compare it against anything else. Syntax-valid and geometry-accurate are not interchangeable terms, and quoting them like-for-like is the only fair comparison.
The second criterion is format fit. CAD-Coder produces editable CadQuery Python, so the decision test is mechanical rather than statistical: take a representative drawing from your own set, run it through the live model, and confirm the resulting script opens, edits, and re-exports inside the software your team already uses. If that round trip fails on your material, the reported benchmark numbers stop mattering for your case.
The third criterion is openness and reproducibility. The model and its GenCAD-Code training dataset — over 163,000 CAD-model image and code pairs — are publicly available through the project's GitHub repository (anniedoris/CAD-Coder), which is what separates this option from closed alternatives. Two checks apply here: read the actual license files in the live repository rather than assuming terms from the paper, and remember that arXiv is a moderated preprint server, not a peer-reviewed venue. Treat the reported results as claims to reproduce on your own drawings, not settled findings.
| Decision factor | The number that matters | Check before committing |
|---|---|---|
| Code validity | 100% valid syntax rate (paper's test set) | Re-run on your own sheets; confirm which metric any quoted rate uses |
| Geometric accuracy | Highest among tested baselines; no published percentage | Score solid similarity yourself, on identical inputs, across tools |
| Training corpus | 163k+ image-code pairs (GenCAD-Code) | Verify dataset access and license terms in the live repository |
| Freshness | arXiv:2505.14646, May 20, 2025 | Confirm the current repository state before building on it |
One more factor deserves a place in your decision file: generalization. The authors report that the model showed "some signs of generalizability," successfully generating code from real-world images and executing CAD operations unseen during fine-tuning. That phrasing is doing honest work — it signals promise, not coverage. Scope a pilot around your actual drawing types before you commit production work, and let the pilot results, not the paper's test set, set your expectations.

Insider Tactics
Start with an artifact inventory, not a demo. At the live CAD-Coder GitHub repository, verify that the documentation, environment or dependency instructions, sample assets, and license notices required for your intended use are reachable from a clean environment. Confirm that any separately hosted assets named in those instructions are accessible, and record the branch or release, retrieval date, and file hashes in a decision log. The CAD-Coder paper points readers to this public implementation; use the paper as a locator, then inventory the live repository rather than treating its description as a current package manifest.
Treat availability and permission as separate gates. Check the repository license, any weight-file or model-card terms, dependency notices, and dataset terms independently. Record actual restrictions on commercial use, modification, redistribution, and hosted access, then compare every option under the same intended use and deployment model. If a term is missing or ambiguous, pause the trial and request a written answer; the label “open-source” should not substitute for review of the complete package.
Use staged timing so an extended evaluation begins only after readily identifiable disqualifiers are gone. Run an access-and-license screen, attempt a minimal reproduction in a clean environment, and advance to your own drawings only if the required artifacts and permissions are present. Put a written go/no-go condition at each gate and preserve the commands, environment record, and failure log. This converts a polished showcase into usable evidence without committing a full pilot to an incomplete option.
For like-for-like totals, time the whole route to an approved deliverable, not merely the fastest successful run. Start each candidate under the same clean-state rule; separately track active operator time, machine wait time, setup, dependency troubleshooting, failed attempts, manual correction, and final validation. Include all labor required to reach the same acceptance condition. Use the same prepared starting state for every option: a demonstration and a first-run trial belong in separate columns unless both begin under equivalent conditions.
Make verification a pre-commitment closeout task. Reopen the live repository and every applicable license, weight, dependency, hosting, and support page; rerun the exact clean-environment check; and attach the resulting artifact hashes and commercial terms to the decision record. If an artifact, permission, total, or required handoff has changed, return to the relevant gate rather than carrying forward an earlier result. The practical advantage is controlled timing: screen early, benchmark fairly, and verify the live package at the moment of commitment.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Inspect CAD-Coder’s live code, model, and setup materials, and verify that the complete end-to-end path from visual CAD input to editable CadQuery Python is available—not merely a paper result or partial demo. | Adoption should depend on a usable, complete option. |
| 2 | Run CAD-Coder and your current CAD workflow on the same representative visual CAD inputs, retaining each generated script, parsing error, geometry failure, and manual correction. | A common input set exposes practical differences that a benchmark headline cannot. |
| 3 | Compare like-for-like finished outputs and totals: editable CadQuery Python, remaining geometry errors, failed generations, and total manual correction effort. | This measures usable deliverables rather than syntax alone. |
| 4 | Check the live license, model-weight availability, dependencies, commercial-use rights, and redistribution limits, then place those terms beside the comparison totals. | Open-source status alone does not establish legal or operational completeness. |
| 5 | Re-test the paper’s reported training-set scale and syntax-correctness result on your own visual CAD inputs, including geometry and editability checks beyond parsing. | The reported benchmark applies to the paper’s evaluation, so its scope must be verified before adoption. |
| 6 | Commit only if the verified pipeline meets your acceptance criteria and performs acceptably on like-for-like totals and terms; otherwise retain the current workflow and record the missing capability. | This turns the comparison into a defensible adoption decision. |
Frequently Asked Questions
What does CAD-Coder actually produce when you feed it a building drawing?
It is an open-source vision-language model, explicitly fine-tuned end-to-end, that converts visual CAD input into editable CadQuery Python.
How much training data was behind CAD-Coder's CAD code generation?
The reported training dataset contains more than 163,000 paired CAD model images and code.
Does CAD-Coder's reported 100% syntax correctness mean it is production-ready?
That result applies only to the paper's evaluation, and the benchmark scope should be checked before adoption.
What happens step by step after you give CAD-Coder an image?
The model receives the image, generates a modeling script, and produces a three-dimensional solid when that script is executed with CadQuery.
What should teams do before committing to an option based on published numbers?
Verify the complete live option and compare like-for-like totals and terms rather than relying on a partial listing or headline figure.
Why is it a mistake to judge CAD-Coder by whether its output text looks plausible?
Verification should cover the entire image-to-CAD-code path, including script execution in CadQuery, rather than stopping at whether the model returns plausible-looking text.
Quick answers
| What is CAD-Coder and what does it do? | CAD-Coder is an open-source vision-language model fine-tuned end-to-end to turn visual CAD input into editable CadQuery Python. |
| How much training data was used for CAD-Coder? | Training used more than 163K paired CAD model images and code. |
| What syntax correctness result did CAD-Coder report in testing? | CAD-Coder reported 100% syntax correctness in testing, though this applies to the paper's evaluation and benchmark scope should be checked before adoption. |
| How does the CAD-Coder image-to-CAD-code pipeline work? | The model receives an image, generates a modeling script, and produces a three-dimensional solid when that script is executed with CadQuery. |
| What does the guide say verification of CAD-Coder should cover? | Verification should cover the entire image-to-code-to-solid path rather than stopping at whether the model returns plausible-looking text, and users should compare like-for-like totals and terms rather than relying on a partial listing or headline figure. |
Also worth reading: Architectural Drawings and the Algorithmic Turn in BIM via AI: Architectural Drawings and the Algorithmic · 2026 VLM Benchmark: NIST Validates IBC 16 Load Path Screening: 2026 VLM Benchmark: NIST Validates · AI Transforms 2025 Serpentine Pavilion Drawings Into Build Code: AI Transforms 2025 Serpentine Pavilion