The story so far
Twenty-five years, one thread. I decide where compute physically lives, starting from the constraint rather than the diagram. It began in hazardous industrial zones, ran through liquid-cooled GPU farms and the migration of whole regulated enterprises to the cloud, and it converges now on AI factories and edge AI. The constant is range: I can carry an investment case into a steering committee and then go and fix the fabric myself. Here is the arc, in order.
2009 - 2011
Hazardous-area edge
The first data hall I designed sat a few hundred metres from a refinery flare stack, inside a zone where a stray spark is a legal event. Heavy industry does not care about your reference architecture. So I learned to start from the constraint: area classification drawings, ATEX and IECEx zones, and compute engineered to survive an environment that would cook an ordinary server room.
The answer was ruggedised, containerised, modular data centres dropped onto the pad, edge compute years before the industry had that word for it. The edge was a wellhead, a pumping station, a processing train. Control loops needed answers in milliseconds even when the wider network was having a bad day, so OT and IT shared physical infrastructure without sharing failure modes: deterministic low-latency networking on the plant side, a hardened path back to corporate for what genuinely had to leave. The sites were tied together on a proper optical backbone, dark fibre and DWDM and MPLS, with N+1 and 2N redundancy and Tier III to IV targets, all watched for PUE and data sovereignty. None of it looked like AI. All of it taught me to decide where compute lives from the constraint, not the diagram.
Read the deep dive →2012 - 2016
Liquid-cooled GPU farms
Past a certain rack density, thermal design stops being a line item and becomes the whole design. For four years I designed and delivered crypto mine farms end to end: dense GPU clusters, the cooling, the power, and the software that fed the machines work.
A mining rack is a preview of an AI rack. Past a certain density, air just moves hot air into more hot air, so I went to liquid both ways: direct-to-chip cold plates with quick-disconnect manifolds, and full immersion cooling for the densest builds, boards sitting in dielectric fluid with the heat pulled straight into the tank. You cannot design the cooling without designing the power, so I laid out the distribution from the supply in, PDUs and circuits and the very real problem of inrush current. The other half was the pool architecture: miners over the stratum protocol, job distribution and scheduling made fast and fair and low-latency, because a lagging pool means stale work and lost revenue every second. The rigs are long gone. The craft, thermal and power and fabric at density, is exactly what an AI factory needs now.
Read the deep dive →2016 - 2021
Cloud landing zones
A migration goes wrong quietly. It fails weeks later, when a forgotten dependency or a compliance gap surfaces in production and someone senior asks why. My work moved from building the physical layer to moving whole enterprises off it: hybrid and public cloud landing-zone design, and the mass migration of multi-stack platforms out of self-hosted data centres, for the financial sector, where the regulator is a stakeholder in every decision.
You build the landing zone first, on the Cloud Adoption Framework and Well-Architected, hub-and-spoke, the whole estate as Terraform rather than console clicks nobody could reproduce. Then you work the 6 Rs honestly, rehost, replatform, refactor, repurchase, retain, retire, because pretending every application should move the same way is how programs blow their budgets. I did this as principal and technical lead across several multi-million-dollar cloud-adoption programs, coordinating many teams through scrum-of-scrums. In a regulated shop the governance is the product: guardrails as code, a zero-trust posture, and compliance built in from the start against APRA CPS 234 and ISO 27001, with FinOps alongside so the move did not just relocate the waste. This is the chapter where the engineer learned to run a program end to end.
Read the deep dive →2021 - now
AI factories & edge AI
Someone asked me recently what changed when AI arrived. Honestly, less than you would think. The models are new, the economics are new, the physics is the one I have designed around since a refinery pad in 2009. Today I am the enterprise storage subject-matter expert for AI and HPC at a Tier 1 Australian telco, and senior solution designer across on-prem, hybrid cloud and edge.
Storage is where AI clusters actually stall, and training and inference stress it in opposite directions: training wants sustained sequential bandwidth to keep every accelerator fed through checkpointing, inference wants low-latency random reads under a tight tail. So the design authority I hold is over the parts people forget: a tiered data path with NVMe flash close to the GPUs, parallel high-throughput file systems for the hot set, object storage for the lake, and the fabric to match, InfiniBand or RoCEv2 for the east-west collective traffic, with NVMe-oF and GPUDirect-style paths so data moves storage-to-GPU without bouncing through host memory. NVIDIA reference architectures, DGX and HGX and GB200-class, give you the node; making a supercluster out of nodes is a storage and fabric problem. I stay vendor-neutral on the silicon too and track the AMD stack, ROCm and the Instinct line, because an AI-factory architect who can only design for one vendor is a procurement risk.
The other half is governance over the physical envelope, space and power and cooling and redundancy, the refinery and mine-farm discipline renamed, running the air-to-liquid transition on liquid-cooled AI compute where the stakes are a national platform. And the part that keeps the work honest is that none of it matters if the value does not land. I led a review of more than a thousand applications and returned north of twenty million dollars a year, ran multi-million-dollar programs on their economics rather than dogma, and kept FinOps in the design so a move did not just relocate the waste. That is the range the job actually asks for: defend a multi-year investment case to a steering committee in the morning, then sit with the storage and fabric blueprint in the afternoon, and translate faithfully in both directions. The pattern underneath all of it is placement. Centralised training in the factory where the power and cooling justify the density, distributed inference at the edge where the latency budget is tight, sovereign AI where the data cannot leave the ground it sits on.
Read the deep dive →The day job is one half. The other half is the lab and the builds at night: the tools I make because I want to use them, on a DGX and a Mac Studio cluster I actually own and run. That thread runs the whole length of this story too. See the lab, the posts and selected work, or read the full bio.