ARCHITECTURE
Power, thermal, and the siting problem
Putting accelerator-class compute inside a working facility is a building problem before it is a software problem. Power, heat, dust, noise, physical access, and service are the constraints that decide whether onsite compute can be installed at all.
· 8 min read
Onsite is a software word for a building problem. Once the argument is settled and the compute belongs in the facility, somebody has to find a place to put it, get power to it, reject its heat into a space that already has plenty, keep it alive in air carrying whatever the process throws off, and service it years later without any of a datacenter's conveniences.
None of that is incidental to the architecture. All of it decides whether the architecture can be installed. And it lands on us rather than on the customer because the facility is the unit of deployment — a platform whose answer to power, thermal, and siting is that the site should provide a suitable environment has not been designed for facilities. It has been designed for datacenters and offered to facilities.
Several of these constraints are unfinished work. It is worth walking them in the order a site survey encounters them.
The electrical question comes first
Accelerator-class compute is a meaningful electrical load in a building whose distribution was laid out for machinery, lighting, and offices. The load is not exotic in absolute terms — a plant has much larger ones — but it is a new continuous load in a location chosen for proximity to the fleet rather than proximity to available capacity. The first questions are unglamorous and decisive. Is there a circuit with headroom where the plane needs to sit. Where is the panel, and what else is on it. What does the run cost, and who has to stop work while it is pulled.
Then there is the quality of the supply. A facility feed is not a datacenter feed. Motors start, welders strike, transfers happen between sources, and a sag that machinery shrugs off is a different kind of event for a computer. So the design has to answer a blunt question: does the plane ride through a brief interruption, or does it shut down cleanly.
Both answers cost something. Riding through means stored energy onsite, which means more equipment, more rejected heat, more footprint, and a consumable with a service life that somebody has to track. Shutting down means the fleet loses its coordination plane, its shared world model, and its scheduler at once, which is a facility-wide event rather than a server restart.
Shutting down cleanly is a phrase that conceals most of the work. For a plane the fleet depends on, clean means every machine has somewhere to land: claims released or expired safely, work in progress brought to a defined posture, machines falling back to local autonomy rather than stopping where they stand and blocking an aisle. That behavior has to be designed and exercised deliberately, because the alternative is discovering what it actually does during a brownout. It also means the plane needs early notice of a power event, which makes the electrical interface an input to the software design rather than a detail for the installer.
A sealed box against physics
Sealed, filtered enclosures are the point of the platform. They are how a computer survives an environment with conductive dust, mist, and moisture, and they are part of the physical-integrity argument as well. They also fight thermodynamics head-on, because a sealed box's only route for heat is through its walls.
Sustained load is what makes it hard. A bursty workload can hide behind thermal mass. A plane running continuous inference for a working fleet across a full shift gets no such relief, and it may be doing it on a floor that is already hot, or in an outdoor cabinet in direct sun in the middle of summer.
From there every direction is a trade. Density fights surface area, and surface area fights a footprint the building does not have. Moving more air faster helps the heat and hurts the acoustics, and it shortens filter service intervals in exactly the dirty environments where sealing mattered most. Filters that last longer restrict flow. Every increment of sustained internal temperature is bought from the service life of the components inside. Push density and you have purchased either a shorter life or a louder machine, and the site usually has an opinion about which.
This is genuinely unsolved work rather than a packaging exercise with a known answer, and it is named as an open problem on the architecture page. We would rather say that than publish an enclosure specification we have not earned.
The room is usually a corner
The environmental list is long and none of it is hypothetical. Dust, abrasive in one industry and conductive in another. Coolant mist that reaches everything downwind of a machining center. Vibration carried through the structure rather than the air. Washdown in food and pharmaceutical spaces, where the cleaning regime is aggressive by design and everything in the room is expected to tolerate it. Salt air near water. Temperature and humidity swings across a day and a season that a datacenter never sees. Each is ordinary somewhere, and a platform meant to serve more than one kind of facility has to survive all of them without a different enclosure for each.
Space is the underrated constraint. The room available is usually not a room. It is a corner behind a column, a mezzanine with a load limit, a gap beside a panel that has to stay accessible by code, an outdoor cabinet, or a footprint that must leave a walkway clear. Depth, service clearance, and floor loading all bind, and so does the path from the loading dock to the final position. These are discovered at survey rather than assumed at design, which argues for a physical envelope carrying more tolerance than a rack-shaped product would need.
People work here
Acoustics deserve more weight than they usually get, because the plane goes where the fleet is, and the fleet works where people do. A plant floor already has a noise budget and people wearing protection against it, so adding to that budget is not free even in the loudest environment. A hospital corridor, a laboratory, or a back-of-house area in a public building is far less forgiving, and equipment that would be unremarkable in a datacenter aisle is simply not acceptable there.
Noise and thermal are the same trade seen from two sides, since the cheapest way to get heat out of a box is to move more air through it faster. Every decision that eases the thermal problem tightens the acoustic one, and the acoustic limit is set by the room rather than by the designer.
There is a broader version of the point. The plane is equipment in a workplace. It has to be safe to stand beside, safe to work around, and unremarkable enough that nobody improvises with it — no intake blocked because material got stacked against it, no access door propped open because a filter was awkward to reach.
Inside the boundary is a range, not a place
The security premise of the facility plane is that it sits inside the customer's boundary. That premise is sound, and it is also less specific than it sounds. Inside the boundary covers a locked room with badge access and camera coverage at one end, and at the other a fenced yard, a shared mechanical space that maintenance contractors pass through daily, or a cabinet at a remote site nobody visits for weeks.
Since the enclosure's surroundings vary by site and we do not control them, the room cannot carry the security argument. What has to carry it is what the platform can guarantee everywhere it goes: tamper evidence on the enclosure itself, so that physical access leaves a mark that cannot be undone; a sealed, signed, read-only image with nothing to install into and no shell to reach; and attestation, so that a node which was handled has to prove what it is before it is permitted to serve a fleet again.
That is the same reasoning the platform applies to machines that move, applied to a machine that stays still in a place we cannot describe in advance.
Service without a shell
Sealed-node discipline changes what field service is. There is no shell, no field-installed package, no configuration applied on site by an engineer with a laptop. A node with a problem is not diagnosed and repaired in place in the usual sense. It is replaced, and the replacement earns its way back into the fleet through the same provisioning and attestation path as the original.
The consequences reach past the enclosure. Spares become part of the deployment plan rather than an afterthought, because time to restore service is dominated by logistics rather than diagnosis. Provisioning has to work in the hands of a facility technician rather than a platform engineer, which puts a hard usability requirement on a process normally allowed to be fiddly. Failure granularity becomes a design decision, since whatever cannot be replaced independently defines the size of the thing someone keeps on a shelf. And observability carries more weight than it otherwise would, because anything a technician cannot discover by logging in has to have been reported before the node went quiet.
Every one of these — the circuit, the heat, the corner, the noise, the lock, the spare — is a design input rather than a condition to be written into a requirements document and handed to the customer. That follows directly from treating the facility as the unit of deployment: if the building is where the system lives, the building's constraints are the system's requirements. We have working answers for some of them. The thermal envelope under sustained inference is not among them yet, and of everything on this list it is the constraint most likely to determine what the hardware finally looks like.