Event-Driven Architecture for Cascading Design Changes Across Partner Systems
Loose coupling between siloed tools cuts manual coordination overhead.

A data center project in 2026 runs power, cooling, cabling, fiber, structural work, and IT infrastructure through large teams that rarely share a common workflow. Changing one of those systems forces every other system to renegotiate its own design around it. That is the mechanical fact driving cost overruns and schedule slips across critical-facility construction right now, and it is getting worse, not better.
The accelerant is demand. Goldman Sachs Research projects data center power needs will climb substantially by 2027 and land well above 2023 levels by 2030, pushed by AI training and inference workloads. That growth curve means design assumptions do not just get set once at the start of a project and hold.
Picture what one of those shifts actually does. An IT team revises rack density, and that single number forces simultaneous revisions to HVAC sizing, electrical panel schedules, cable tray routing, fiber pathways, and submittal packages. None of those revisions happen automatically. Each discipline has to be told, by a person, in a meeting or an email chain, that something upstream moved.
The cascade is not confined to the drawing board either. One hyperscale operator planning a multi-hundred-megawatt AI training campus in the American Midwest closed on land and signed power purchase agreements, then discovered only one viable fiber route into the property, with no commercially acceptable alternative. A siting decision that should have accounted for fiber access as a first-order constraint treated it as a detail to solve later, and later turned out to be too late.
Inside the building, the same failure runs in the other direction. AI workloads need an ultra-low-latency GPU fabric, and the bandwidth those fabrics require has grown substantially across successive GPU generations. That means ceiling heights and conduit routes, decisions poured into concrete and steel, are increasingly dictated by network architecture rather than the other way around. Power drives cooling. Cooling drives structural. Networking drives building geometry. Every discipline sits downstream of every other discipline, and the coordination load that creates has outgrown what manual, tool-by-tool notification was ever built to handle.
Siloed contracts and tool boundaries turn coordination into manual labor
The cascade problem is built into how data center projects are contracted, since it is not a communication habit that better meetings would fix. The mechanical contractor delivers the building management system. The electrical contractor delivers the electrical power management system. The client's IT team buys the DCIM platform separately. Nobody is paid to own the connections, so the connections do not get built as part of the plan. They get improvised after the fact.
That produces real, physical redundancy in how data moves. UPS status is required simultaneously by the BMS, the EPMS, and the DCIM platform. If the UPS head end cannot speak SNMP natively, that one device now depends on one gateway serving three separate consumers, a gateway nobody budgeted for and nobody's contract explicitly assigned.
BIM was supposed to be the fix for this kind of fragmentation. In practice it often adds to it. Developers commonly hand over a construction BIM model that the operator cannot use for facility management without substantial rework, because BIM was built as a geometry and coordination tool, not a live-linkable data source. It shows where a pipe runs. It does not carry the structured asset data a DCIM platform needs to track that pipe as an asset over the life of the facility. Getting an operational BIM model, one the operator can actually use, requires defining model format, asset-tagging structure, and CMMS/DCIM data requirements at project kickoff, with acceptance criteria written down before design starts. That almost never happens. Those requirements get negotiated at handover instead, when the model is already built and reworking it costs real money and real time.
The BMS integration failures that follow from this are predictable enough to list in advance. Alarm thresholds get set to comfort ranges instead of thermal protection limits. No escalation path gets built, so an alarm logs correctly and nobody on call ever sees it. Each team assumes another team is watching. None of them are wrong to assume that: the contract never told any single party they owned the watch.
The outage data reflects that consequence. The Uptime Institute's 2025 Analysis, built on its 2024 survey, found power was the leading cause of impactful outages for a majority of operators surveyed, with cooling a distant second. Power systems sit at the center of the most contractually fragmented integration path in the entire facility, the one running between EPMS, BMS, and DCIM with no owner assigned to the seams, so the outage pattern reflects that fragmentation rather than physics. Fragmentation is a structural condition, produced by how scopes get written and how tools get bought. A fix for it has to change the architecture connecting those tools.
Event-driven architecture addresses this specific problem
Event-driven architecture, EDA, is a design paradigm where system components communicate through events rather than direct calls. A producer emits a signal that something meaningful just changed. A broker captures that signal and distributes it. Consumers subscribe to the kinds of events they care about and act on their own schedule, asynchronously, without the producer ever knowing who is listening or waiting around for a reply.
The property that matters for project delivery is loose coupling. Producers and consumers run independently. A new consumer can subscribe to an existing event stream without anyone touching the producer's code or process. A failure in one consumer can be isolated and retried without dragging every other consumer down with it. That is the exact inverse of how design coordination runs today, where a rack-density change requires a person to manually track down and notify mechanical, electrical, cabling, and DCIM teams one at a time.
The contrast with polling, the default alternative, is a matter of wasted effort. Event-Driven Microservices Architecture structures systems around producing and consuming events, which gives real-time responsiveness while cutting resource use compared to systems that continuously check for changes that usually have not happened. A system polling every few minutes for a design revision burns compute on thousands of checks that turn up nothing, and still misses the one check that mattered by however many minutes stand between polling cycles.
Applied to design workflows, a rack-density revision is published as a change event to a message broker and consumed simultaneously by mechanical, electrical, structured cabling, and DCIM configuration workflows, without manual re-notification to each team. At a system level, this flips how state changes propagate. Instead of one service directly invoking another's behavior, services emit facts about what already happened, and the rest of the system reacts on its own. Control flow shifts from a sequence someone has to manage by hand to something distributed across independent reactions to the same fact.
None of this replaces the BMS, the EPMS, the DCIM platform, cable management software, or the documentation system a project already runs on. EDA is a coordination layer sitting between those tools, connecting them without forcing a custom, tightly coupled integration to be built and maintained for every possible pair of systems. That distinction is the whole argument for a project-delivery audience: the tools stay the same, the wiring between them changes.
The EDA patterns that matter most for propagating design changes across disciplines
Not every EDA pattern fits this problem equally well. Picking the wrong one produces hidden coupling and fragile schemas, an implementation that calls itself event-driven but behaves like the tightly coupled system it was meant to replace the moment it comes under real pressure.
Publish/Subscribe is the foundational pattern here. Events get published to a broker, and multiple subscribers react to the same event independently, with no direct dependency between them. The rack-density change event, published once, gets consumed by mechanical, electrical, and cabling workflows simultaneously. Broker infrastructure for this kind of workload is well established: Apache Kafka for high-throughput, replayable event logs, AWS EventBridge for cloud-native managed routing, and Google Cloud Pub/Sub for low-latency delivery.
Event-Carried State Transfer matters when a downstream consumer needs the data immediately and cannot afford to depend on an external lookup. In the rack-density case, the event carries the revised kW/rack figure along with its full context: rack ID, zone, new load, previous load, timestamp, and the discipline that submitted it. Every consuming system can act on that event directly, without querying the originating model to get the rest of the picture.
Event Sourcing supplies the audit trail. Rather than storing only the current state of a design, the system stores the full log of every event, and current state gets reconstructed by replaying that log. For critical-facility delivery, commissioning agents and authorities having jurisdiction, who need traceable decisions, get every design change, every RFI answer, and every equipment parameter revision recorded in the log as an immutable event.
The architectural choice of choreography versus orchestration determines who governs the cascade. Under choreography, each service reacts to events and emits new ones with no central controller, fully decoupled but harder to visualize and debug. Under orchestration, a central service directs the others, easier to monitor but introducing a coupling point and a potential bottleneck. For coordination across separate firms on separate contracts, choreography usually governs the peer-to-peer notifications, since no single firm should be directing another firm's internal workflow. Orchestration fits inside one firm's own systems, where a single team already has visibility and control.
IBM's patented approach to function-defined output streams, granted January 2025, shows where this is headed. Output functions combine events from multiple streams into a new derived stream, for instance deriving a new capacity constraint event from simultaneous changes to power load and cooling capacity, with neither original producer aware of the other. That is the pattern working exactly as intended: two disciplines generate facts independently, and a third layer derives the consequence neither one could see on its own.
The failure to avoid is bolting brokers and queues onto an unchanged architecture. Events get published, but nobody owns them. Consumers multiply, but the contracts between producer and consumer stay unstable. The failure chains that should have been designed out appear once the system is under load.
How EDA applies to the RFI and submittal cascade
A design change in the current workflow does not stop at updating a model. It manually triggers an RFI process, a resubmittal cycle, and a fabrication hold, and each of those compounds the delay of the one before it. A single RFI that leads to a resubmittal costs roughly five weeks end to end: about two weeks for the RFI process itself, one week to prepare the revised submittal, and two weeks for the second review. Multiplying that by the typical RFI volume on a critical-facility project makes the schedule exposure enormous.
This is where the cost of fragmentation stops being abstract. Nearly every piece of power and cooling equipment on a data center project sits behind submittal clearance before fabrication can even be released. RFI and submittal latency is a direct multiplier on the project schedule, not administrative overhead that runs quietly in the background.
The Suffolk and MIT white paper, published September 2026, names document management fragmentation as a primary delivery risk and points to AI-assisted RFI and submittal review as a way to speed the process up. Without that architecture, an AI reviewer is reading from the same fragmented, siloed sources causing the delay in the first place.
An RFI event stream built on EDA supplies that architecture. When a design change event publishes, a downstream consumer can cross-reference specifications, drawings, and prior RFIs automatically, identify which submittals the change affects, draft notifications to the relevant disciplines, and track the contractual response clock, all without a project manager manually figuring out who needs to know.
The structural pattern already runs at scale outside construction. Amazon and Shopify's order fulfillment systems use a single triggering event, an order placed, to cascade independent downstream processes: inventory, shipping, customer notifications, each running without blocking the others. A design change event cascading into RFI review, submittal tracking, and discipline notification is the same structure applied to a different industry.
The accountability question that follows is obvious: if an automated cascade triggers a resubmittal, who answers for that decision on a project being reviewed by commissioning agents and AHJs. Event Sourcing answers it directly. Every step in the cascade logs the triggering event, the consuming service, and the action taken, producing a complete, timestamped record for review.
AI inference versus deterministic rules in this architecture
EDA is the plumbing that moves events between systems. What happens when a consumer receives an event is a separate decision, and in critical-facility delivery, the line between AI inference and deterministic rules is a safety and compliance question.
Deterministic rules have to govern anywhere precision and auditability are non-negotiable: applying standard power redundancy rules to a layout change, enforcing rack- and outlet-level power caps, sending certificate-expiry reminders, flagging a design event that violates a thermal protection threshold. These run on codified, versioned policy, and every outcome can be validated ahead of time rather than discovered after the fact.
Under an L4 autonomy model, the rules are deterministic and versioned, and a human only steps in for exceptions. The system executes on its own, but a named engineer stands behind the rule that governs it, which is what satisfies an AHJ or commissioning agent asking who is accountable. Modern data center automation platforms already coordinate EPMS, DCIM, and intelligent PDU data to enforce power caps and shed nonessential load during a constraint, running deterministic rule execution over an event stream.
AI-augmented processing belongs where language, pattern matching, and judgment across unstructured documents speed up work that a person would otherwise do by hand, line by line. An RFI agent reading incoming RFIs, searching specifications and drawings, drafting a cited response, routing it to the correct discipline, tracking the contractual response clock, and flagging RFIs that point to superseded documents fits that description exactly, running at L2, where a human approves the action before it goes out.
That boundary, deterministic rules governing safety and compliance, AI assisting judgment and language work under human approval, is what makes an event-driven architecture usable on a facility where a commissioning agent and an AHJ both have to sign off on how a design change actually got approved.
Sources
- Best architectural patterns for event-driven systems
- Event Driven Architecture Done Right: How to Scale Systems with Quality in 2025 - Growin
- 10 Real-World Event Driven Architecture Examples Transforming Industries in 2025 - Streamkap
- Function defined event streams from multiple event streams and events
- Event-Driven Architecture Patterns for High-Throughput Systems in 2026 - Landskill
- Event-Driven Microservices Architecture for Data Center ...


