SDK Design for Domain Experts Who Are Not Software Engineers
Domain experts gain an SDK that maps to their mental models, not to software architecture.

An SDK built for a data center engineer succeeds or fails on one question: does it speak the engineer's language, or does it force the engineer to speak the software's? Most tools on the market today fail that test, and the cost appears as hours lost to rework rather than as a missing feature.
Why domain experts hit a wall with tools built for developers
Picture a rack move. The engineer already knows what it touches: the power draw on a circuit shifts, the cooling load in the zone changes, cable pathway lengths stretch or shrink, and the as-built record goes stale the moment the move happens. None of that is a mystery to the person making the call. What's missing is a tool that acts on that knowledge without the engineer having to retype it across separate systems.
That gap is not a skills problem. An MEP engineer, a commissioning specialist, a critical-facility designer knows the domain cold. What they lack is a representation, since their mental model of a rack move, a power budget per zone, or a cooling topology change has no first-class expression in a general-purpose SDK. The fragmentation runs deeper than one bad tool. BIM platforms demand scripting skill to do anything beyond the defaults. DCIM, EPMS, and BMS each keep their own data contracts, refusing to talk to one another out of the box. At every boundary between them, someone retypes the same design intent by hand, a structural problem that repeats on every project, every change order, and every handoff. It's structural, and it repeats on every project, on every change order, on every handoff.
What "mental model" means for data center engineers
Before fixing the tooling, it helps to be precise about what the tooling needs to hold. A data center engineer doesn't reason in nodes, edges, or generic entities. The vocabulary is racks, power paths, cooling zones, cable routes, capacity budgets, and the rules that connect one to the next. Ask an engineer to place a rack: the placement is valid only if it clears the power budget, not merely because it fits on the floor plan. It has to clear the power budget for that zone, sit inside the cooling envelope, and leave a workable cable pathway, all at once, not patched together after the fact one constraint at a time.
Design intent also runs in layers. Change the density on a rack, or the aisle configuration, or the power topology, and consequences ripple downward through the whole facility model. The engineer already carries that dependency graph in their head. They expect the tools to carry it too.
Work doesn't stop at the model, either. A request for information is a question tied to a specific design object, and that question has consequences for procurement, for commissioning, for the as-built record that gets handed to operations. The BMS watches and controls the mechanical and environmental plant, the EPMS watches the electrical path, and DCIM manages IT assets and capacity. An engineer has to hold all three of those pictures in their head at once, even though the platforms behind them were never built to talk to each other.
Why general-purpose SDK design principles fail for domain experts
The standard playbook for good SDK design isn't wrong. Idiomatic code, type-safe models, CI/CD integration, support across multiple languages: all of that reduces friction for a developer who already thinks in terms of APIs. None of it closes the gap for an engineer who doesn't. Look at what "production-ready" actually means in SDK circles: authentication flows, retry logic, pagination, streaming. Real engineering concerns, all of them, and every one built for a developer's problem, not a commissioning specialist's.
Even the parts of SDK design meant to lower the barrier to entry, like interactive "try it out" docs and code samples that auto-populate, assume the user's goal is understanding an API endpoint. A design engineer's goal is never that. The goal is expressing a constraint or pushing a change through a facility model, and no amount of polished documentation changes what the tool is fundamentally asking the user to do.
The deeper issue is one of abstraction level. A general-purpose SDK exposes what the system is capable of. A domain-expert SDK has to expose what the expert is actually trying to accomplish, not merely what the system can do, and those are different things at different levels of specificity. Purdue's GUIDE research makes a version of this same point in interface design: when a system forces experts to work through indirect representations, like prompts or exemplars, instead of their own native way of interacting with the work, design intent gets harder to enforce. GUIDE was built to let designers steer and inspect outputs directly, instead of going through an intermediary layer that doesn't speak their language. The fix for data center engineers is a different foundation. It's a different foundation.
The first principle: abstractions should map to domain objects, not system objects
A domain-expert SDK earns its keep when the nouns in its API match the nouns the expert already uses at the coordination table, reflecting the domain rather than the internal representation of whatever system sits underneath. For a data center engineer, that means the SDK hands over objects called Rack, PowerPath, CoolingZone, CableRoute, and CapacityBudget. Not Layer. Not Node. Not Edge or Entity.
Naming alone doesn't finish the job. The object's behavior has to line up with domain logic too: a Rack object needs to know which zone it belongs to, what power budget it's working within, and which cables depend on it, because that's how the engineer already reasons about a rack. Workshop findings from the ESDA community on design automation back this up directly. Abstractions that model high-level behavior, instead of exposing the low-level machinery underneath, are what let non-specialist users actually drive automation tools with confidence.
The BIM world is the clearest cautionary tale available. A BIM platform's internal object model, geometry, parametric relationships, family instances, is technically sophisticated. None of it maps to how an MEP engineer thinks about a power topology change, and that mismatch is why scripting a BIM platform requires a developer standing between the engineer and the software. There's a useful test for any new abstraction an SDK introduces: would the domain expert use this exact name and description in a coordination meeting? If the answer is no, the abstraction belongs somewhere lower in the stack, away from the engineer.
The second principle: governed rules, not open scripting, for decisions where precision is non-negotiable
Getting the nouns right solves half the problem. One mode is for expressing flexible intent. The other is for decisions where correctness is fixed and non-negotiable, and an SDK for domain experts needs a separate mechanism for each.
Rack power caps, aisle clearances, fire-suppression zones, cooling envelope limits: these are compliance-relevant, and for constraints like these the SDK has to run governed, deterministic rules. Feed the same approved inputs through the system twice, get the same output twice, and if the rule itself needs to change, that change happens through deliberate human action, not silently. The logic behind this holds up in any safety-critical setting. When a domain expert changes a rule, someone made that call on purpose. Handing that decision to an AI agent instead means the organization is outsourcing its reasoning to a system whose training may have nothing to do with the regulatory environment or the risk profile of that specific facility.
Companies building in this space are running into hallucinations and non-deterministic behavior exactly where they can least afford it. The SDK design that serves domain experts makes the choice between deterministic and non-deterministic modes explicit, and puts that choice in the expert's hands rather than the model's. None of this argues against AI showing up in the SDK at all. AI earns its place wherever the instruction space is genuinely open-ended: surfacing candidate configurations, drafting RFI language, summarizing the impact of a change. But whatever executes against a live facility model has to trace back to a governed rule. A probabilistic guess doesn't get to make that call.
In practice, that means rule authoring itself becomes a first-class operation inside the SDK rather than a hidden setting. Named, versioned, reviewable, so the expert can read the logic, modify it, and approve it without writing a line of code.
The third principle: a change in one place must cascade automatically to every place the domain requires
The rack move again. It's worth returning to because it exposes the sharpest failure in current tooling. The failure isn't that an engineer can't express the change. The SDK simply doesn't know what that change implies elsewhere, so the engineer ends up manually chasing down every consequence across power, cooling, cable, schedules, and documentation.
The fix is a dependency graph, configured once against the domain object model, that the system then walks through automatically every time something changes. The expert acts once. The system propagates from there. That propagation has to be visible as well as correct. An engineer needs to see what changed, why it changed, and which rule triggered it, rather than a model that quietly looks different than it did an hour ago. Traceability is a core requirement of the SDK itself, not a logging feature bolted on afterward.
The same pattern occurs in document workflows. An RFI that inherits context from the design object it's questioning, and whose response automatically updates procurement and commissioning records, is the propagation principle applied outside the physical model. The RFI behaves as a dependency in the system rather than as a log entry sitting off to the side. The Greenergy and Siemens deployment offers a real-world version of what happens without this layer: engineers had to map protocols by hand across BACnet, Modbus, and SNMP, and every one of those mappings needed a developer to actually implement it. No layer existed between the engineer and the system to carry the propagation the engineer already understood.
The fourth principle: structured output designed for the handoff boundary, not just for design-time use
Every principle so far lives inside the design phase. None of it survives contact with operations unless the SDK's output is built for that handoff specifically. An SDK that builds rich, well-structured domain objects during design and then exports something loose or unstructured at handoff has just wasted every bit of automation that came before it.
The handoff failure in data center delivery has a specific shape. Asset data generated during design almost never arrives in a form that DCIM, EPMS, or BMS platforms can take in without someone re-keying it by hand. The intelligence the engineer built into the design has to be manually rebuilt by whoever's running operations. An SDK that skips structured handoff output just hands the integration burden to engineers who never signed up to be integrators.
The design implication is straightforward: DCIM, EPMS, and BMS targets need to be first-class output formats inside the SDK, not an export step tacked on at the end⟢. An engineer should be able to declare, the moment an asset is created, that it's headed to DCIM with a specific set of attributes, rather than dealing with that mapping later as a separate chore. And this is a traceability question as much as a formatting one. Structured handoff means every operational parameter is traceable to the design decision and the rule that produced it, so operations teams can audit, update, and validate the data without re-entering it.
What breaks when these principles are violated: the fragmentation tax in real delivery
None of this is theoretical. Skip the first principle and abstractions stop matching domain objects, so experts spend their working hours translating between systems instead of designing. That's how a design change turns into a manual rework cycle instead of a clean, governed propagation event. Skipping the second principle blurs the line between deterministic and AI-driven action, and domain experts stop trusting the outputs. They fall back to manually checking everything by hand, which erases whatever time the automation was supposed to save and adds a review burden on top.
Skipping the third principle turns small changes into coordination failures. Design freeze already lands months earlier than most teams expect, and a change that lands after freeze sets off a procurement cascade that can add real weeks to the critical path. Skipping the fourth principle means operations teams end up manually re-entering data that was already sitting, fully formed, inside the design model. The Uptime Institute's 2025 Annual Outage Analysis found that power was the leading cause of impactful outages for a majority of operators surveyed, and the pattern behind that finding is usually a siloed signal chain: the team watching cooling has no visibility into what's happening on the electrical side.
None of this is a soft inefficiency to shrug off. AI workloads are pushing power density high enough that every design decision now carries real cost. Commissioned equipment sitting idle because of a documentation gap runs up a concrete monthly carrying cost, and that cost scales with the size of the facility.
How ArchiLabs Studio applies these principles in a connected data center workflow
ArchiLabs Studio starts from the same diagnosis: data center design's core problem is fragmentation, the same information rebuilt by hand across every tool in the chain. Its automation layer is built around the four principles laid out above as the operating structure.
The platform's abstractions map to what a critical-facility engineer actually reasons about: layouts, power paths, cable routes, equipment schedules, RFIs. Not generic BIM constructs, not database tables dressed up with friendlier names. On the question of precision, deterministic rules handle power budget enforcement, cooling envelope validation, and cable routing compliance. AI has a role too, speeding up rule authoring and surfacing the impact of a proposed change, but the rule that actually executes against a live facility model stays explicit, reviewable, and owned by the engineering team, not the model.
Moving a rack in Studio cascades automatically to power, cooling, cable, schedules, and sheets, because the dependency graph is configured at the domain level, not patched manually across disconnected tools. And at the handoff boundary, Studio's output is built to reach DCIM, EPMS, and BMS in a form those systems can actually take in without manual re-entry, with every asset attribute traceable back to the design decision that produced it. That's what it looks like when an SDK is built around the engineer's mental model from the start, instead of asking the engineer to adapt to the software's.
Sources
- Workshops on Extreme Scale Design Automation (ESDA) Challenges and Opportunities for 2025 and Beyond
- GUIDE: Designer-in-the-loop Authoring of Conformant Generative User Interfaces
- Mitigating hallucinations and omissions in LLMs for invertible problems: An application to hardware logic design automation
- The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
- Making Software Meaningful
- SDKs: Principles and Best Practices - by Eyal Lantzman
- Making Abstraction Concrete: A Design Space and Interaction Model of Abstraction in Interactive Systems
- ConceptModeller: a Problem-Oriented Visual SDK for Globally Distributed Enterprise Systems


