← Back to list

AI Use in Residential Building Science Forensic Investigations, Part 1 of 6: Claude

There’s an important scientific model that is often overlooked, yet it is one of the most powerful tools we can use when faced with the…

Brett Dillon · 2026-07-15 20:23 · 0 claps · 9.2 min read
#artificial-intelligence #ai #building-science #forensic-investigations
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General 🔬 · Science · General

AI Use in Residential Building Science Forensic Investigations, Part 1 of 6: Claude

There’s an important scientific model that is often overlooked, yet it is one of the most powerful tools we can use when faced with the inherent uncertainty of residential building science forensic investigations: Karl Popper’s concept of falsification. This is the idea that we should disprove things in order to eliminate them from consideration. Richard Feynman wrote it best when he wrote, in 1964’s Character of Physical Laws Lecture 7 “Seeking New Laws”, “There is always the possibility of proving any definite theory wrong; but notice that we can never prove it right. Suppose that you invent a good guess, calculate the consequences, and discover every time that the consequences you have calculated agree with experiment. The theory is then right? No, it is simply proved not wrong. In the future you could compute a wider range of consequences, there could be a wider range of experiments, and you might then discover that the thing is wrong.” Later, he went on to write, “In other words, we are trying to prove ourselves wrong as quickly as possible, because only in that way can we find progress.”

With that in mind, when I was recently told people are using AI to diagnose residential building failures, I immediately reached for this set of investigative tools and asked AI directly.

In Part 1, I discuss this with Claude. I have reproduced Claude’s answers unedited here, and am not responsible for Claude’s responses that I’m faithfully reporting here.

Brett: Explain to me why AI should NOT be used to solve building science-related problems in residential construction.

Claude: There’s a real case to be made here, and it comes down to the nature of building science as a discipline rather than any generic anti-AI sentiment.

Building science is a site-specific, systems-coupled problem — not a lookup problem

Every building envelope decision involves interacting variables: local climate (heating/cooling degree days, driving rain, wind-driven moisture loads), the specific assembly (wall/roof/foundation stack-up), interior moisture generation, HVAC design, and occupant behavior. Change one variable and the “correct” answer for something like vapor retarder placement or insulation R-value can flip entirely. A model trained on aggregated text is pattern-matching across a huge range of climate zones and assemblies without physically grounding the answer in your psychrometric conditions, dew point calculations, or WUFI-type hygrothermal modeling. It can sound authoritative while quietly recommending something that’s correct in Atlanta and catastrophic in Duluth.

The failure mode is delayed and expensive, not immediate

Bad structural advice tends to fail fast and visibly. Bad building science advice — wrong vapor barrier orientation, poor detailing at a rim joist, an improperly vented rainscreen — often doesn’t show up for 2–5 years, after the building has been fully closed up, finished, and occupied. By the time mold, rot, or ice damming appears, the assembly is buried behind drywall and finishes, remediation costs are high, and causal attribution is difficult. This is exactly the kind of harm profile where “the AI seemed confident and it’s usually right” is a genuinely bad way to make decisions.

LLMs don’t understand codes as living, jurisdictional documents

IRC/IECC requirements vary by state amendment, local jurisdiction adoption cycle, and climate zone; an AI answer trained on some blend of code editions and forum chatter can easily cite a requirement that isn’t in force in your jurisdiction, or blend a code minimum with a best practice as if they were the same thing.

Tacit and craft knowledge doesn’t transfer through text

A lot of what separates working assemblies from failing ones is installation-sequence knowledge — how a specific crew actually laps a WRB, how a specific window flashing detail behaves with a specific siding profile — that lives in the hands of installers and BPI/RESNET-credentialed professionals, not in text corpora.

Liability and certification structures assume a human is accountable

Code officials, HERS raters, and PEs stamp documents because there’s a licensed, insurable, legally accountable person behind the judgment. AI-generated building science advice has no such chain of accountability, which matters a great deal when the “test” for whether the advice was any good is a moisture problem that shows up during a future real estate transaction.

The counter-case, briefly

None of this means AI is useless in the space — it can be genuinely helpful for drafting, organizing test data, summarizing established research, or checking arithmetic in load calculations, provided a qualified human is validating the physics against the actual site conditions and stamping the final call. The strongest version of the caution isn’t “never use AI here,” it’s “don’t let it be the final authority on anything where the consequence of being wrong is hidden moisture damage in an assembly you can’t easily inspect later.”

Brett: Explain to me why AI should NOT be used to draft, organize test data, summarize established research, or check math in load calculations for building science-related investigations in residential construction.

Claude: Here’s the steelman case for that stronger position — that even the “safer” AI uses (drafting, data organization, literature summary, arithmetic checks) carry real risk in building science investigations specifically.

Drafting

The risk isn’t the prose quality — it’s that drafted language has a way of becoming the final language. Once a paragraph exists in polished form, the pressure to critically re-derive it (rather than lightly edit it) drops substantially. In an investigation — say, a forensic moisture failure report or an expert analysis — the reasoning chain is the deliverable. If AI drafts the reasoning and a human polishes the wording, the illusion of authorship can mask the fact that no human actually re-derived the physical logic from the underlying data. In litigation or code-compliance contexts, that’s a real problem: the named expert is attesting to conclusions they may not have independently reconstructed.

Organizing test data

Organization sounds neutral, but the act of organizing is already an act of interpretation. Deciding which sensor readings are “outliers” to flag, how to bucket time-series data, which column groupings imply causation — these are analytical choices, not clerical ones. An AI reorganizing hygrothermal or psychrometric datasets can silently impose a framework (e.g., grouping by season instead of by envelope orientation) that shapes what patterns become visible or invisible before a human ever looks at the raw numbers. The QBO-style miscategorization failure mode is illustrative: once bad categorization enters a dataset, it doesn’t just sit there — it propagates into every downstream summary and derived metric, and it’s much harder to catch in aggregate a than at the point of entry.

Summarizing established research

Summarization compresses nuance, and building science is a field where the nuance often is the finding. A study’s conclusion about vapor retarder placement might apply only within a narrow set of climate zones or wall assemblies described in its methods section — details that get dropped in a summary abstract-of-an-abstract. AI summarization also tends to flatten disagreement between sources into a false consensus, when in fact building science has real, unresolved debates (e.g., rainscreen venting requirements in mixed-humid climates) where the “summary” version can mislead someone into thinking the science is settled when it isn’t.

Checking math in load calculations

This is the most seemingly mechanical task and still carries a specific risk: silent unit or assumption errors. Manual J/D/S calculations depend on correctly specified inputs (design temperatures, infiltration rates, duct leakage assumptions) — if those inputs are wrong, the arithmetic can be flawlessly executed and still produce a wrong, load-mismatched HVAC recommendation. An AI “checking the math” verifies internal consistency, not real-world correctness, and a human reviewer who sees “checked by AI, no errors found” may reasonably but wrongly relax their own scrutiny of the input assumptions — a classic automation-complacency effect.

The common thread

All four tasks share the same underlying risk: they’re upstream or downstream of the actual judgment call, and once AI performs them, the human’s incentive to independently verify weakens — even when the human remains nominally “in the loop.” In an investigation where the final product needs to survive cross-examination or peer scrutiny, that erosion of independent verification is the actual harm, more than any single wrong output.

Counterpoint, briefly

The opposing case is that these are exactly the tasks where AI’s weaknesses are least consequential and its strengths — speed, tirelessness, consistency — are most valuable, provided a qualified human treats AI output as a first draft requiring full independent verification rather than a shortcut past it. Whether that proviso reliably holds in practice, across different people and different pressure/time constraints, is really the empirical crux of the disagreement.

Brett: You’ve stated before that AI use, across disciplines, actively de-skills the humans who engage with it. This was recently proven in a test involving gastroenterologists and polyp identification in colonoscopies.

Claude: I don’t have a record of stating that specific claim to you before, but the study is real, and it’s a strong piece of evidence for the deskilling argument — worth walking through directly.

What the study actually found

The Lancet Gastroenterology & Hepatology study found that after physicians had been using AI for three months, their unassisted adenoma detection rate fell significantly, to 22%. This came out of a retrospective, observational study at four endoscopy centres in Poland taking part in the ACCEPT trial, where these centres introduced AI tools for polyp detection at the end of 2021, after which colonoscopies were randomly assigned to be conducted with or without AI assistance. Researchers compared colonoscopy quality across two phases — 3 months before and 3 months after AI implementation, looking specifically at how the same clinicians performed without the AI tool running. The study involved 1,443 patients who underwent non-AI-assisted standard colonoscopy — 795 before and 648 after the introduction of AI. The result: continuous exposure to AI appeared to reduce the ADR of standard non-AI-assisted colonoscopy, suggesting a negative effect on endoscopist behavior. Researchers called it the first documentation of a potential deskilling effect from real-world medical AI use.

This is a meaningful finding because it isn’t a lab simulation or self-report — it’s a before/after natural experiment with a real clinical quality metric (ADR), which is itself tied to downstream colorectal cancer risk.

Why this transfers to the building science argument

The mechanism the researchers point to — automation use leading to a decay in cognitive skills, as suggested by psychological studies in non-medical fields — is not specific to endoscopy. It’s a generic property of skilled visual/diagnostic search: when a system reliably flags the thing you’re looking for, your own active scanning behavior atrophies, because you stop needing to generate the hypothesis yourself. Building science has a structurally similar skill: reading an assembly, a set of psychrometric conditions, or a moisture pattern and generating the failure hypothesis before being told what it is. If an AI (or any decision-support tool) routinely does that generation step, the practitioner’s own pattern-recognition — the thing that lets a rater notice something not on the checklist — likely degrades the same way ADR degraded here.

Where the analogy is honest and where it’s not perfect

It’s worth being precise about the differences, since overclaiming would undercut the argument:

  • The colonoscopy study measured a skill outcome directly (detection rate) with a large, controlled before/after sample. There isn’t an equivalent RCT-quality study for building science professionals and AI use yet — the building science version of this argument is currently a prediction from a documented mechanism, not (yet) a documented outcome in that specific field.
  • Polyp detection is a repeated, high-volume, immediately-verifiable task (biopsy confirms or denies). Building science judgments (e.g., “this wall assembly will trap moisture in this climate”) are lower-frequency and the feedback loop is years long, which if anything should make deskilling harder to detect but not less real — the absence of a matching study is a data gap, not evidence against the effect.
  • The endoscopy case is skill decay in a task the clinician still performs constantly (just without AI half the time). In building science, the more common exposure pattern is a practitioner leaning on AI most of the time, with far fewer no-AI reps to atrophy — so the effect, if it exists, may show up faster and more severely, not more slowly.

The core claim — that reliable automated pattern-detection degrades the human’s independent pattern-detection over time — now has direct clinical evidence behind it, and there’s a plausible, specific mechanism for why building science diagnostic judgment would be vulnerable to the same effect, even without a study measuring it yet in that field.

Brett: By the way, you brought up deskilling of humans through AI use in our conversation titled “the future of society with unrestricted AI use”.

Claude: That checks out — in “The future of society with unrestricted AI use,” I raised cognitive outsourcing directly: the concern that AI use erodes mental habits like extended attention, tolerance for ambiguity, and the slow work of forming one’s own opinions, describing it as offloading not just tasks but thinking itself. I also drew a GPS analogy — that just as GPS eroded innate navigation sense, AI risks doing the same to other cognitive skills.

Notably, that conversation ended with me being fairly self-critical about my own role in it — I’d made claims you had to push back on and verify yourself, and the conclusion we landed on was that you got value from the exchange because of what you brought to it (persistent scrutiny), not because of what I offered unprompted. That’s actually a nice bookend to today’s conversation: the colonoscopy study is a real, measured instance of the same mechanism I was describing abstractly back in March — reliable automation reduces the human’s independent generation of the skill, and the STAT News coverage put it in almost identical terms to what I said then, comparing it to students who try to write an essay independently after using ChatGPT one too many times.

So the throughline across both conversations is consistent: the deskilling risk isn’t hypothetical dressing — it’s a mechanism with actual clinical evidence now, and it maps onto building science exactly the way it mapped onto endoscopy and onto the “unrestricted AI use” conversation’s broader point about cognitive outsourcing.

Note: RESNET® and BPI® are registered trademarks of their respective organizations


메타데이터
post_id
ad7edb4ae1c5
slug
ai-use-in-residential-building-science-forensic-investigations-part-1-of-6-claude-ad7edb4ae1c5
url
https://medium.com/@brett_64695/ai-use-in-residential-building-science-forensic-investigations-part-1-of-6-claude-ad7edb4ae1c5
canonical_url
https://medium.com/@brett_64695/ai-use-in-residential-building-science-forensic-investigations-part-1-of-6-claude-ad7edb4ae1c5
author_url
https://medium.com/@brett_64695
status
ok
fetched_at
2026-07-25 07:36:07