Quick Answer: AI workloads are not a single infrastructure problem. Training and inference have different power requirements, different latency thresholds, and different facility fits. For most Canadian enterprises, colocation sits at the centre of the answer. The specific facility you choose matters as much as the category. Power density per rack, cooling architecture, carrier connectivity, and data sovereignty posture are the four criteria that separate a genuinely capable data centre from one that simply claims to be AI-ready.
The expectation people have from AI has changed a lot in the past few years. What used to be an added benefit and an extra ‘nice-to-have’ has turned into a non-negotiable requirement for consumers. For businesses, the pressure to train better models, serve inference results faster, and do both in a way that satisfies a compliance officer is turning data centre selection into one of the most consequential infrastructure decisions an organisation can make.
Most guidance on the topic stays at the surface. It stops at "choose colocation for AI" without helping buyers evaluate whether a specific facility can actually deliver. What organisations need is a way to match their AI workload type to the right category of infrastructure, and to know exactly which facility-level criteria to evaluate before signing a contract.
AI workloads are often discussed as though they are simply a more demanding version of conventional enterprise compute. They are not. The difference is structural, and it begins with the impact AI has on a facility's power and thermal systems, and why most data centres were never designed to handle it.
Standard enterprise racks are built around 5 to 10 kW of power per cabinet. That figure has been the industry norm for years, and most colocation facilities were designed to accommodate it.
The Uptime Institute's 2025 Global Data Centre Survey found that average server rack power densities are rising, with greater adoption in the 10 to 30 kW range. Facilities capable of exceeding 30 kW comfortably remain relatively uncommon across the broader market.
AI infrastructure operates well above that ceiling. A single NVIDIA H100 GPU draws approximately 700 watts. A training cluster running 256 of those GPUs requires over 180 kW of sustained capacity. NVIDIA's GB200 NVL72 chassis requires approximately 120 kW per rack configuration alone.
The IEA projects that electricity consumption from accelerated servers, driven primarily by AI adoption, will grow at roughly 30 per cent annually through 2030, far outpacing conventional server growth at 9 per cent.
On a smaller scale, the difference between the hardware being used a decade ago and now is substantially different. In our article on RFPs for AI colocation, we covered the maximum power draws of various GPUs and the results were extremely interesting. Compared to NVIDIA’s H100, the Tesla V100 SXM2 from 2017 has a max TDP of 300W. That’s less than half the power needed for a modern server-grade GPU.
When we think about the fact that most of the data centres you’re considering right now were probably made 5-10 years ago, you can see why power has become a whole different game entirely because of the AI boom.
What we call "AI" is really two separate workload categories with very different infrastructure requirements.
Training is how a model learns. It involves processing enormous datasets through parallel compute operations, often continuously for days or weeks. Training clusters prioritise raw compute power, high internal bandwidth between GPU nodes, and sustained high-power delivery. Latency to end users is irrelevant at this stage, which means training workloads can go where power is most available and cost-effective.
AI inference is how a model performs. Every chatbot response, product recommendation, fraud detection signal, and image classification result is an inference operation. Inference infrastructure has almost nothing in common with training from a facility standpoint.
It prioritises low latency, geographic distribution, and reliable carrier connectivity rather than compute density. A single inference request requires far less compute than a training run, but the latency constraints are strict and often non-negotiable for user-facing applications.
These two workloads pull in opposite directions, and assuming one facility can serve both equally well is a mismatch that creates performance or compliance problems after deployment.
With training and inference understood as distinct problems, the next question is which category of infrastructure addresses each one. There are four main types of data centre relevant to enterprise AI deployments, and they are not interchangeable. Most sophisticated AI architectures draw on more than one.
Hyperscale data centres are large-scale facilities built to support very high compute volumes at cost-efficient power and space rates. They are designed primarily for cloud providers and software-as-a-service companies building and running large language models at massive scale.
Because latency to end users is not a concern during training, hyperscale facilities are often located in remote areas where energy costs and land availability favour large-footprint builds.
For Canadian enterprises considering hyperscale as an AI training option, the most significant factor is not technical. It is structural. We’ll talk more about this in a bit more detail later on.
Colocation is the infrastructure model where an organisation houses its own hardware inside a professionally managed third-party facility and pays for space, power, and connectivity.
The organisation owns and controls the compute; the facility operator manages the physical environment. This is the most relevant infrastructure category for the majority of enterprise AI deployments, and the hyperscale vs. colocation comparison is the decision most enterprise buyers eventually face.
What makes colocation particularly well-suited to AI is the combination of hardware control and managed infrastructure. Modern colocation facilities designed for high-density workloads support both training and inference, with carrier-neutral connectivity ecosystems that are especially valuable for inference performance. For organisations running GPU workloads on public cloud and watching the monthly bills climb, migrating GPU clusters to colocation typically delivers significant cost reduction once deployment volume reaches a sustained level.
On-premises data centres give organisations maximum control over hardware, physical security, and regulatory posture. When weighing cloud vs. on-premise data centres for AI specifically, the constraint is almost always power and scale.
An existing on-premises environment is rarely provisioned for the densities AI workloads require, and retrofitting it to support high-density GPU clusters typically involves substantial capital expenditure and multi-year timelines.
For most organisations, on-premises infrastructure plays a supporting role in the AI stack, typically for small inference workloads that require tight physical security or operate in air-gapped environments.
Edge data centres are small, geographically distributed facilities positioned close to end users or data sources. They serve real-time inference applications where millisecond latency matters: autonomous systems, industrial IoT, and regional content delivery.
Their limited compute and power footprint rules them out for training entirely. For most Canadian enterprises, edge plays a supporting role in a broader inference architecture, not the foundation of one.
|
Data Centre Type |
Best Suited For |
Typical Power Density |
Latency Priority |
Canadian Sovereignty Note |
|
Hyperscale |
Large-scale model training |
Very high |
Not a factor |
Often U.S.-parented; CLOUD Act applies |
|
Colocation |
Training and inference |
Medium to high |
Adjustable by deployment |
Varies; verify ownership structure |
|
On-Premises |
Small inference workloads |
Low to medium |
Operator-controlled |
Full control; rarely practical at AI scale |
|
Edge |
Real-time inference |
Low |
Critical |
Full control; limited compute capacity |
"AI-ready" has become one of the most overused phrases in data centre marketing. Most facilities using it are referring to standard high-density colocation capabilities at best.
The criteria below are what actually separate a facility built to support AI workloads from one that will become a bottleneck after you have already committed.
The first number to ask for is the committed power per cabinet in kilowatts. This is the sustained, guaranteed capacity delivered to your rack under normal operating conditions, not a theoretical facility maximum or a shared pool figure.
As a reference point:
Most facilities advertising AI support are actually provisioned for 10 to 20 kW. That is sufficient for lighter inference workloads but insufficient for most training environments. If the answer from a provider is framed around what they could accommodate rather than what they can guarantee, the infrastructure has not been purpose-built for AI demand.
Air cooling is the standard thermal management approach across most of the industry, and it works well up to approximately 15 to 20 kW per rack. Above that threshold, hot spots develop even with hot aisle/cold aisle containment, and the airflow required to compensate becomes inefficient and difficult to scale.
Facilities capable of supporting AI at meaningful density have adopted liquid-assisted cooling in some form. Direct-to-chip liquid cooling, rear-door heat exchangers, and immersion cooling each handle different density thresholds and carry different operational trade-offs.
The key question is not which cooling method a facility uses but whether it can reliably dissipate your hardware's thermal load without requiring a bespoke engineering project after you move in.
For inference workloads, interconnection quality matters as much as the compute environment. Consumer-facing AI applications generally target sub-50 millisecond response times. Reaching that threshold consistently requires diverse carrier options with physically isolated interconnect rooms, not a single-path network where one provider's performance or availability affects everything downstream.
Carrier-neutral facilities let you route traffic across multiple networks, reducing single-carrier dependency as both a latency risk and a resilience risk simultaneously.
Network latency is consistently underestimated in AI infrastructure evaluations. Most buyers prioritise compute and cooling specifications and treat connectivity as an afterthought. For inference workloads specifically, the network fabric of your facility deserves more scrutiny than it usually receives.
SOC 2, ISO 27001, PCI DSS, and HIPAA certifications are evidence of operational discipline. They demonstrate that a facility has been independently audited against a defined security and availability standard, which matters when AI workloads touch regulated data.
For federally regulated financial institutions specifically, OSFI Guideline B-13 places explicit expectations on how technology and third-party infrastructure risk must be managed, and AI workloads touching financial data fall squarely within scope.
Certifications, however, do not answer the sovereignty question independently. A facility holding ISO 27001 certification and operating under a U.S.-headquartered parent remains subject to U.S. legal jurisdiction for data disclosure purposes. For organisations that need data sovereignty rather than just data residency, the ownership and legal structure of the provider is a separate evaluation dimension from its operational certifications.
|
Evaluation Criterion |
Standard Facility |
AI-Ready Facility |
|
Power density per rack |
5-10 kW |
15-50+ kW, pre-provisioned |
|
Cooling above 20 kW |
Air-cooled only |
Liquid-assisted; pre-engineered |
|
Carrier connectivity |
1-3 carriers |
5+ carriers, physically isolated rooms |
|
Certifications |
Basic |
SOC 1/2, ISO 27001, PCI DSS, HIPAA |
|
Data sovereignty posture |
Often unaddressed |
Verified ownership and jurisdiction |
|
Capacity availability |
Often constrained |
Confirm committed deployment lead time |
Location is usually framed as a geography problem: which city minimises latency to your users, or which market offers the best power costs. For Canadian enterprises deploying AI, location carries a legal dimension that most infrastructure discussions skip over entirely. For regulated industries, it may be the most important factor of all.
The U.S. Clarifying Lawful Overseas Use of Data Act, known as the CLOUD Act, requires U.S.-based technology providers to produce data stored anywhere in the world when served with a valid legal order from U.S. law enforcement. This applies to any provider incorporated in the United States, headquartered there, or with a U.S. entity in the corporate ownership chain, regardless of where the physical servers sit.
For AI specifically, this is not a hypothetical concern. Training data is the raw material of your model. If it includes patient records, financial information, or government data, placing it on U.S.-parented infrastructure, even geographically within Canada, creates legal exposure that a data residency clause cannot fix.
The CLOUD Act exposure for Canadian organisations is a function of the provider's corporate structure, not the address on the building. According to CIRA's 2025 Cybersecurity Survey, 69 per cent of Canadian organisations now cite data sovereignty as their primary infrastructure consideration, up from 60 per cent in 2024. An additional 82 per cent say a vendor's country of origin has become more important in the past year alone.
Training workloads are latency-tolerant. Inference is not. Serving AI-powered applications to users across Canada from a single city creates latency degradation for anyone outside the host market. The further a user is from the inference facility, the longer the round trip.
At the millisecond thresholds AI applications demand, that distance translates directly into measurable performance loss.
A provider with facilities across multiple Canadian markets gives organisations the ability to distribute inference workloads geographically without routing traffic across the U.S. border. This matters for both performance and compliance. Routing Canadian inference traffic through U.S. network infrastructure, even briefly, reintroduces the same jurisdictional questions as hosting the data there. Multi-city Canadian coverage removes that exposure at the network layer, not just at the storage layer.
The gap between a facility that genuinely supports AI workloads and one that says it does can be significant. The RFP process for AI-ready colocation needs more rigour than a standard data centre evaluation, because the performance difference between a capable facility and a well-marketed one only becomes apparent after you are already deployed. The right questions, asked before you sign, surface that gap early.
These are the questions that reveal whether an AI-readiness claim is backed by real infrastructure or marketing language:
Available capacity has become a strategic variable in its own right. Toronto is currently operating at approximately 2 per cent vacancy. Most new supply across North American markets has been pre-leased before construction is complete. A facility with strong specifications and a 12 to 18-month deployment lead time is a materially different product from one that can move your workload into production infrastructure within weeks.
This matters most for organisations under compliance timelines, board mandates to move AI workloads off public cloud, or product launch schedules that require inference capacity at a specific date.
Ask every provider not just whether they can support your workload but when. Get a committed deployment lead time in writing before evaluating the rest of the proposal.
Choosing the wrong data centre for AI is not a vendor preference mistake. It is a structural risk that becomes expensive to resolve after hardware has been racked, contracts have been signed, and workloads are running in production. Qu Data Centres addresses the full infrastructure equation that most Canadian organisations cannot find in a single provider.
Qu's nine carrier-neutral facilities across five Canadian markets are built for high-density compute and backed by independent certifications including SOC 2, ISO 27001, PCI DSS, and HIPAA, giving compliance teams the documentation they need before procurement sign-off.
The company is Canadian-incorporated with no U.S. parent in its ownership chain, making the sovereignty posture structural rather than contractual. Four facilities hold Uptime Institute Tier III certification, representing one of the highest concentrations of that standard in Canada. With 17 MW of capacity available today across the national footprint, Qu is positioned to deploy your AI workloads now, not after a multi-year build pipeline. Review all available data centre locations and capacity to find the right facility for your workload.
Book a facility tour to see the infrastructure firsthand and speak with a specialist about your AI deployment requirements.
Training infrastructure is built for sustained, high-power parallel compute over extended periods, and it tolerates latency, so it can be located where power is available and cost-effective. Inference infrastructure is built for speed and proximity to users, serving model results in milliseconds across geographically distributed locations. The two workload types have different power requirements, different cooling demands, and different network priorities.
Standard AI deployments require approximately 15 to 25 kW per rack. High-density configurations need 30 to 50 kW, and extreme-density liquid-cooled systems can exceed 100 kW. Most standard colocation facilities are provisioned for 5 to 10 kW per rack, which is insufficient for most production AI training environments and limits the scope of inference deployments as well.
Colocation typically delivers better cost efficiency than public cloud GPU instances once workload volumes become sustained, generally from 12 months of usage onward. Cloud is faster to start and appropriate for irregular or exploratory training runs. Many organisations use both: cloud for variable workloads and initial testing, colocation for production inference and sustained training where predictable costs matter.
Providers incorporated in the United States or with a U.S. parent in their ownership chain are subject to the CLOUD Act, which requires them to produce data under valid U.S. legal orders regardless of where servers are physically located. For Canadian enterprises handling data governed by PHIPA, OSFI B-13, or PIPEDA, this creates compliance exposure that data residency agreements alone cannot resolve.
Ask for the committed power per cabinet in kilowatts, confirm whether cooling is pre-provisioned for racks above 20 kW, verify which certifications are active at the specific facility you are evaluating, and ask about the ownership structure and legal jurisdiction of the provider. Vague answers to any of these questions suggest the AI-readiness claim has not been validated at the infrastructure level.
Colocation pricing for AI workloads varies based on power density, total capacity, and contract term. Higher per-rack power draws command a premium because they require purpose-built electrical and cooling infrastructure that standard facilities cannot support without significant investment. A detailed breakdown of what drives colocation pricing is covered in the colocation cost guide.