We built an AI factory for HVAC control

Over the summer I worked with Koja on an AI Champion project: an AI factory for HVAC control optimization. Koja gave me access to three air handling units at a real production site, with live telemetry that had been streaming since the start of 2024. Three zones, each with a completely different occupancy rhythm and its own idea of what “comfortable” means.

That access is the reason the project was worth doing. A benchmark dataset would have been easier and taught me almost nothing.

What do I mean by an AI factory? Not a single model, and not a chatbot pointed at a database. It is a pipeline that takes an objective stated in plain language — improve comfort in this zone without spending more fan energy — and carries it through a fixed sequence of stages: pin the history and its metadata, audit and clean the data, learn a model of each zone, learn a control policy against that model, and evaluate the result against what the building actually did. Specialised AI agents handle the interpretation and planning at each stage. Deterministic programs handle every calculation, and can stop the run outright. Each experiment writes to its own folder, so a result is always traceable to the exact data and settings that produced it.

The goal is not one clever answer for one building. It is a repeatable process that turns raw telemetry into a control recommendation somebody could actually defend.

It also produced the finding I did not expect. The hardest part was not the models. It was deciding what has to be true before a model’s output deserves to be believed — and building a system willing to stop when it isn’t.

In this blog I want to describe what that looked like in practice, why I now think the stopping is the interesting part, and where this work goes next.

Why buildings are worth the trouble

The case for optimizing HVAC is easy to underrate, because ventilation is invisible until it fails.

The US Environmental Protection Agency estimates that we spend around 90% of our time indoors (https://www.epa.gov/indoor-air-quality-iaq/improving-your-indoor-environment), and that concentrations of some indoor pollutants run two to five times higher than typical outdoor levels. Air handling units are what stand between people and that number, in the buildings where they spend nearly all of their lives.

The trend is toward more of this, not less. The UN’s World Urbanization Prospects 2025 (https://www.un.org/en/desa/WUP-2025) finds that 45% of the world’s population now lives in cities, with a further 36% in towns, and projects that two-thirds of all population growth between now and 2050 will happen in cities. More people in more buildings means more air to condition, and more energy spent conditioning it.

The systems are better than they used to be. That’s the interesting part.

There’s a version of this story where HVAC controls are dumb and AI arrives to save them. That version is out of date.

Plenty of systems now run demand-controlled ventilation, heat recovery, weather compensation, and model-predictive control. Koja’s units are well instrumented and thoughtfully engineered. The industry has been actively pursuing optimization for years.

But two things remain true even in good systems. Control logic is commissioned once, against assumptions about how a space will be used, and those assumptions drift — a space commissioned for one pattern of use is rarely used that way five years later. And the operating rules are fixed, so when they are wrong they are wrong quietly: energy is wasted, or comfort is sacrificed, and nothing visibly breaks.

So the opportunity isn’t to replace competent engineering. It is to let a system learn from what the building actually did, and propose operating decisions grounded in measured history rather than in a commissioning-day assumption.

Which brings the question back to data.

The quiet problem underneath

An optimization result is only as trustworthy as the data underneath it.

Real buildings do not hand you clean data. Sensors drift. Signals disappear for hours. The same tag name means different things at different sites, and the meaning often lives in an engineer’s head rather than in the metadata.

None of this announces itself. That is the problem. A pipeline that continues on bad data does not fail loudly — it produces a confident number that happens to be wrong. And a confident wrong number is worse than no number, because someone will act on it.

A workflow built as a graph, not a conversation

My answer was to stop treating the pipeline as something an AI agent talks its way through, and start treating it as an explicit graph.

Every stage is a node. The shared state is inspectable. The edges — not the model — decide what happens next.

Agents propose; deterministic gates dispose. The amber path is a bounded replan; the red path is a halt that only a human can clear.

Agents still do a great deal. They interpret the request, read the site, infer what the sensor labels probably mean, plan the run, adapt the code. That is genuinely useful work, and work that would otherwise take an engineer days.

What agents do not do is grade their own output.

Between the stages sit deterministic checks. Scripts calculate every number. If the incoming data is inconsistent, if a signal falls outside its physical range, if the train/test split leaks, the run stops at a terminal halt and a human is asked. It does not proceed with a caveat attached to the report.

What a gate is actually for

It is tempting to describe these checks as safety features bolted on at the end. They are not. They are what makes the rest of the output meaningful.

A few working principles for where a gate belongs:

1. Where a failure would be silent. Loud failures do not need gates; they announce themselves. Gates belong wherever a problem would pass downstream unnoticed.
2. Where the model has an incentive to be optimistic. Anywhere an agent could report success without anything checking, something else should be checking.
3. Where the fix requires a person. Missing data does not improve on retry, and a site-specific meaning nobody has confirmed cannot be resolved by trying again. These halts go to a human and stay there.
4. Where a stage depends on an earlier stage’s quality. Hyperparameter search is only as meaningful as the data, targets, and features underneath it, so tuning sits behind the same gates as everything else.

The effect is not that the system runs less often. It is that when it does run to completion, the result is tied to a specific dataset, specific settings, and a record of every check it passed.

Learning to control a building without touching it

The modelling has two halves, and they depend on each other.

The surrogate is trained on measured history, then becomes the environment the control policy practises inside.

First, a surrogate: a learned model of each zone that reads recent history and predicts what happens next. Surrogates are comparatively cheap to build — you don’t need insulation values, equipment performance curves, or occupancy counts to get started. That is their appeal and also their limit. They know what the building did, not why.

Second, a control policy that learns from the surrogate’s predictions rather than from the live building. It practises inside the learned model, choosing setpoints and being scored on comfort, air quality, and fan use. The action space is mixed, discrete and continuous together, which adds real complexity. Early training oscillates visibly between options before it settles — which turned out to be a useful diagnostic in its own right.

The factory is designed to compare candidate architectures and return the best surrogate for a zone, and I’m extending the code to bring attention-based models into that candidate set.

Everything here is offline. The policy has never controlled the real building, and the results are simulated results. Saying so plainly isn’t a disclaimer; it is the difference between a claim I can defend and one I cannot.

Not every result improved, either. One objective moved in the wrong direction, which the evaluation surfaced rather than averaged away — and that finding is now shaping the next iteration. A pipeline that only reports its wins isn’t reporting.

Future scope

Two directions interest me most.

Richer context for the surrogate. Adding LiDAR-based room mapping or IFC building models would give it the spatial and structural information it currently lacks. A surrogate learned from telemetry, combined with a structural model of the space, starts to look like a credible route toward a genuine digital twin — rather than the term as it is usually marketed.

Transfer. One building, three zones, one vendor’s telemetry. The models are specific to that site and would need retraining anywhere else. But the governed structure — the graph, the gates, the human checkpoints, the run records — is meant to carry to the next building without rebuilding the reasoning from scratch. A kind of inductive step: from the specific case toward the general one. That’s the claim I’d most like to be wrong about in an interesting way.

Agents propose, the graph disposes

If there’s one sentence I’d keep from this project, it’s that one.

The value of an agentic system isn’t that it can do everything. It’s that it can do a great deal while remaining answerable to something that is not itself.

In GPT-Lab, we’re continuing to work on how governed pipelines like this transfer between sites and sectors. If you’re working on something similar — in buildings, or anywhere else where the data is messier than the demo suggests — don’t hesitate to contact us!

About the author

Rajratan Wankhade

Project Researcher

Scroll to Top