Category Archives: AI

Nine and a half weeks with AI

What I learnt using and working with AI

Exit from Oracle due to AI

Towards the tailend of last year – September 2025 onwards – was when I started using AI (ChatGPT) for my work. It was the internal ChatGPT approved by Oracle, so now and again I would ask it a questions related to my work.

After a couple of questions, I asked it this question: Will AI take over and put me out of a job? Knowing it could not really lie and I was curious to see what it said with the follow up question: What shall I do now to prepare for when AI makes my role redundant…

The answer then as it will be the same but a little less specific if I asked these questions now – that is, it told me it was hard to say – it all depends on what my job is now, and it listed the jobs/roles that would be most affected by AI. As for the “what should I do in preparation – for when AI takes my job” – it told be to become an AI evangelist!

September 2025 was bad at Oracle (and it has been ever since) – there were threats of mass RIF (reduction in force) and the threats did materialise into lots of fellow Oracle employees in India and the US getting laid off. The UK was spared, but the impending doom of RIF was demoralising and people did not, for one moment, thought they were safe as it was all over. Personally, for me I was not hopefully – in fact more the opposite as I had been laid off from Cisco Meraki previous to getting this role at Oracle, so I needed to do something about it now.

I wasn’t going to leave it to chance to be laid off twice, so I starting looking for new roles enabling me to leave before the next round of redundancies. I applied for roles in the SRE and Observability area especially as I liked working as a Observability SRE with Cisco Meraki before being “reduced” prematurely! By the way, Oracle started the RIF process to raise capital expenditure (CapEx) for their AI expansion – the Abilene DC in Texas. They made redundancies where they can to reap the most amount of CapEx – there is no other reason why an individual is made redundant apart from raising as much money from their departure as possible…

To my surprise, two such roles appeared – one at Graphcore and one at Nscale. Both of these companies have a close but differing relationship with AI. I interviewed with both using standard and usual preparation techniques with no help from AI. I was offered a role by Nscale but was rejected by Graphcore. I accepted the role at Nscale and once the contract was signed and handed in notice at Oracle – serving a 1 month notice period before joining Nscale in mid-December of 2025.

Use of AI at Start-ups like Nscale

It is of no surprise that the use of AI has been adopted by small companies and start-ups who need to move fast, launch products, and perform support, sales and marketing tasks fast. AI allows you to do this. One moto at Nscale is to move fast and that good is good enough – don’t let perfection slow you down or halt your progress. With the use and help of AI – in all areas of a start-up company – it will allow you to do this with minimal resources and little time.

I found, on joining Nscale, all employees had access to ChatGPT and could ask for access to Claude, and were encouraged to use AI for all aspects of our work. The company were also using modern applications with AI built-in or enabled so were also encourage to use that AI. For example, traditional applications such as Jira and Confluence were replaced by Linear and Notion. These had AI native functions and behaviour which would speed-up or automate your work. They also integrate with each other and other applications such as Slack and Gmail to enable you to combine AI queries across multiple sources of information.

I joined as an Observability Platform Engineer expecting, as part of my roles, to be creating dashboards and alerts with my skills and experience. But no, AI replaced all of this such that any engineer (without o11y skills or Grafana experiences) could ask AI to simply “create me a dashboard to show the latency of X, Y and Z and associate an alert when X, Y, and Z crosses a threshold of A, B or C” – for example. AI would be able to have a very good attempt at doing this very fast (minutes instead of hours). It seems like I was already out of a job before I even started…

Not surprisingly, as the weeks rolled on at Nscale, the use of AI was very apparent, with each Engineering Weekly meeting having a demo that sang the praises of how AI helped with creating a useful or needed feature or solution in a short amount of time. Later, as I used AI to create applications, I recognised or realised that ALL if not most of the Nscale applications, interfaces, and features were created using AI. The UI to their console is a big give away:

No human would write their code on a a few lines!

All Nscale employees were “faking it until they made it” and using AI to help them do so fast. I found that apps like Notion enables you to use AI to find info, detail and documentation really readily and will also summarise and condense information for digesting in a short amount of time – no more tl;dr – get AI to summarise and read the pertinent snippets…

Addictive Nature of AI

First lession learnt after using AI for work (and also for anything else) is that it is addictive – once you’ve used it and found it helpful – it is hard to not use it and go back to how you use to do things albeit a slower and more laborious.

It’s like using a calculator – if you use it for everything, you lose the skill of doing mental arithmetic and come to depend of it to do the all the calculations. If you imagine AI as a super super magical calculator that can help you solve and perform all your work tasks – it will become addictive and the more you use, the more skills you will lose and the more dependent on it you become…

How does an employee get appraised if AI is doing all the work? It comes down to your manager, and unfortunately, after 2.5 months at Nscale, I was transferred to a new manager who didn’t like the look of me and extended my probation period – setting me up to fail so that he could easily dismiss me during this probation period without causing an issue with the company. So after 4.5 months with Nscale, I was let go.

AI Addiction becomes a habit

Once I knew how to use AI and take advantage of its features and limitations – it was hard to not use it. There are so many areas where it could lighten the burden of tedious work and speed up tasks 10 to 100 times. The first thing to use AI for after been made unemployed is to find a new role and/or job.

 role = what you actually do and how you do
 job = formal title and place in the company
 ATS = applicant tracking system

I started out seeking a new job by setting up a spreadsheet to keep track of my job applications – I knew the job market was going to be a lot tougher than the previous two period of unemployment, so I named this sheet appropriately!

Start off on the right foot by using the “Framing” method – or more specifically strategic naming or linguistic framing. Or strategic framing through naming.

It means choosing a project, programme, policy, or initiative name that shapes how people interpret it before they examine the details. A well-chosen name can influence support, behaviour, priorities, funding, and perceptions of success. Expecting and knowing how tough the job market is I turned job hunting into a “mission” and this certainly influenced my way of working – kicking in my resourcefulness, resourceful thinking and greater innovation

The Resourceful Use of AI

The first and best reason to use AI is: if AI is the cause of your redundancy, then why not use it to attain a new job/role? If AI is really taking over the world (or decimating roles in IT) then your should be able to use it to your advantage to find a suitable role and in turn attain a job with a company.

In a tough job market you will have to apply for many roles, suffer a lot of rejections, ghosting, and no replies to job applications, and this is before you are invited to an interview with a human (companies are using AI interviewers now to the amount of applications!)

Use AI to Improve your CV/Resume

This is the first thing you should do but not the only thing you should do! I wrote a post on this here: https://www.blusas.co.uk/find-a-new-role-or-job-after-redundancy/ but to summarise here with a focus on AI:

AI can help most if you use it to reduce admin, improve targeting, and sharpen your pitch rather than to mass-apply for roles. The biggest wins are tailoring CVs to each role, drafting outreach messages, organizing applications, and preparing for interviews faster.

The use of AI to improve your CV/Resume for ATS is necessary nowadays just to get a talent advisor or recruiter to initiate a contact with you. Without this contact you will just fall wayside and not be seen by anyone – human or AI.

Use AI to Generate Interview Questions

In this latest search for jobs, once I have lined up an interview, I gave AI the JD of that particular role and my CV and simply asked it to give me 10-30 interview questions between basic and advanced level that the interviewer was most likely to ask me.

Although this technique is not foolproof, it is as good as IF the interviewer (new to the task and having lots of candidates to interview or screen) has also used AI to formulate the list of interview questions!

Use AI for Interview Practise

AI could also be instructed to simulate an interview sessions where you could give realtime replies for it to ask you further questions on your answers and give you feedback at the end. I never did this so I can’t really comment but this would be super useful to get interview experience. If you have worked for one company over a long period of time, then not only the job market has changed, but you are also out of practise and rusty at doing well in an interview…

Personally, I use the first few interviews as practice – I would never rely of the first three interview to result in a job offer – I woud also want them to be tough so that they reveal my weaknesses so I can improve. If their is an option to transcribe the interview session, then do that with the intension of handing that to AI to analyse and give suggestions on the improvements.

Use AI in Your End-to-End Job Hunting

Here’s a LinkedIn post from a user on how they use AI in their “job hunting workflow” – end-to-end process: I stopped applying to jobs. Instead, I built a system using Anthropic Claude that applies for me.

I never got to this stage, but potential I could have been desperate enough to resort to this if my period of unemployment continued for a while and I was getting no successes at attaining interviews and job offers (at any salary where I could get into “temporary” employment while carrying on with the job search for a more suitable role).

Use AI to Learn and Improve

I think this is the best use of AI – to learn and improve your skills, knowledge and experience during times of unemployment and working your “Mission for a New Role” strategic framing project!

[more]

AIOps – Artificial Intelligence for IT Operations

AIOps stands for Artificial Intelligence for IT Operations. It is a methodology for using machine learning, statistical analysis, automation, and now LLM-based reasoning to improve how infrastructure and application operations teams detect, investigate, explain, and resolve problems.

At its core, AIOps is trying to solve a very practical problem:

Modern systems produce more operational data than humans can manually inspect, correlate, and act on quickly enough.

That includes metrics, logs, traces, events, alerts, tickets, deployments, topology changes, CI/CD activity, cloud audit logs, Kubernetes events, OpenStack state, Slurm queues, GPU telemetry, network flows, and user-impact signals.


1. The Problem AIOps Is Trying to Solve

Modern IT operations has become too complex for purely manual troubleshooting.

A typical platform may include:

Users

Load balancers

Ingress / API gateways

Kubernetes services

Microservices

Databases / queues / object storage

Cloud / OpenStack / VMware / bare metal

Networks / firewalls / DNS / storage / GPUs

Every layer emits telemetry. The problem is not lack of data. The problem is too much disconnected data.


Problem 1: Alert Fatigue

Operations teams often receive hundreds or thousands of alerts.

Many are:

  • duplicates
  • symptoms rather than root causes
  • low priority
  • transient
  • missing context
  • caused by the same underlying event

Example:

Disk latency high
API latency high
Pod restart count high
Database connection errors
Frontend 500s
SLO burn rate alert
User complaints

A human has to determine whether these are six separate incidents or one cascading failure.

AIOps tries to group these signals into one meaningful incident.


Problem 2: Data Silos

Metrics are in one place.

Logs are in another.

Traces are somewhere else.

Tickets are in Jira or ServiceNow.

Deployments are in GitLab or GitHub.

Infrastructure state is in OpenStack, Kubernetes, Slurm, Ceph, AWS, Azure, or VMware.

The engineer has to jump between tools:

Grafana → Loki → Tempo → Prometheus → Kubernetes → OpenStack → SSH → Jira → GitLab

That is slow, error-prone, and dependent on tribal knowledge.

AIOps tries to connect these sources and reason across them.


Problem 3: Manual Root Cause Analysis

Traditional troubleshooting is often manual correlation.

An engineer asks:

What changed?
What broke?
Who deployed?
Which node is affected?
Is this network, storage, compute, DNS, auth, GPU, database, or app?
Has this happened before?
What fixed it last time?

That investigation may take 30 minutes, 2 hours, or several days.

AIOps attempts to reduce that investigation time by automatically correlating evidence.


Problem 4: Too Much Complexity

Modern platforms are dynamic.

Examples:

  • containers are rescheduled
  • pods are ephemeral
  • cloud instances appear and disappear
  • autoscaling changes capacity
  • CI/CD continuously deploys changes
  • service dependencies shift
  • storage volumes move
  • GPU nodes are drained, allocated, or isolated
  • network paths change
  • certificates expire
  • DNS records update

Humans are not good at mentally tracking all of that in real time.

AIOps tries to build a continuously updated operational view of the environment.


Problem 5: Reactive Operations

Traditional operations is often reactive:

Something breaks → alert fires → engineer investigates → fix applied

AIOps aims to make operations more proactive:

Early warning → anomaly detected → likely cause identified → risk predicted → action recommended

For example:

  • predict disk exhaustion
  • detect memory leak patterns
  • identify increasing error budgets burn
  • spot noisy neighbours
  • detect degraded GPU nodes
  • find abnormal API latency before customers complain
  • flag risky deployments

2. What AIOps Is Trying to Solve

AIOps is trying to make IT operations:

  • faster
  • more accurate
  • less noisy
  • more automated
  • more predictive
  • less dependent on individual experts
  • better aligned with business and service impact

The goal is not simply “AI for dashboards.” The goal is to improve operational outcomes.


3. The Main Capabilities of AIOps

A mature AIOps approach usually includes several capabilities.


1. Data Collection

AIOps needs telemetry from across the environment.

Typical sources include:

Metrics       → Prometheus, Mimir, VictoriaMetrics, CloudWatch
Logs → Loki, Elasticsearch, OpenSearch, Splunk
Traces → Tempo, Jaeger, OpenTelemetry
Events → Kubernetes events, OpenStack events, systemd, audit logs
Tickets → Jira, ServiceNow, GitLab issues
Deployments → GitLab CI/CD, GitHub Actions, ArgoCD
Infrastructure→ OpenStack, Kubernetes, Ceph, Slurm, VMware, AWS, Azure
Network → flow logs, DNS logs, firewall logs, load balancer logs

Without good data, AIOps is weak. The first requirement is solid observability.


2. Noise Reduction

AIOps should reduce alert noise by grouping related alerts.

For example, instead of showing:

NodeDown
PodCrashLooping
HTTP5xxHigh
LatencyHigh
DatabaseConnectionFailure
SLOBurnRateHigh

it should produce something closer to:

Incident: Database node failure causing API errors

Affected services:
- checkout-api
- payment-api
- frontend

Likely root cause:
- PostgreSQL primary unavailable

Evidence:
- node db-03 stopped responding at 10:42
- API connection errors started at 10:43
- customer-facing 500s increased at 10:44

This is one of the biggest practical wins of AIOps.


3. Anomaly Detection

AIOps can learn normal behaviour and detect deviations.

Examples:

CPU usage normally peaks at 70%, now 95%
API latency usually 120 ms, now 900 ms
GPU memory errors normally zero, now increasing
Login failures normally 20/hour, now 5,000/hour
Network packet drops normally rare, now concentrated on one host

This is useful when static thresholds are poor.

A static alert might say:

CPU > 90%

But anomaly detection can say:

This service normally uses 15% CPU at this time of day.
It is now using 65%, which is abnormal for this workload.

That is more context-aware.


4. Event Correlation

AIOps correlates events across systems.

Example:

10:01 - GitLab deployment completed
10:04 - Kubernetes pods restarted
10:05 - latency increased
10:06 - error rate increased
10:07 - SLO burn alert fired

The likely cause is not “latency high.” The likely cause is the deployment.

AIOps should connect those facts.


5. Root Cause Analysis

AIOps tries to identify the underlying cause, not just the symptoms.

Example:

Symptom:
Users cannot access the application.

Possible causes:
- DNS failure
- certificate expiry
- ingress failure
- pod crash
- database outage
- network ACL issue
- storage outage
- failed deployment

AIOps task:
Rank the most likely causes using evidence.

A good AIOps system does not just say:

Application is down.

It says:

The application is down because the ingress controller cannot reach the backend pods.
The backend pods are healthy, but the service selector was changed in the latest deployment.

That is operationally useful.


6. Recommendation

AIOps should recommend next actions.

Example:

Recommended action:
Rollback deployment checkout-api:v2.4.1 to v2.4.0.

Reason:
Errors started within 3 minutes of the deployment.
No infrastructure errors were detected.
Previous version had normal latency and error rate.

The recommendation should include evidence, not just a guess.


7. Automation and Remediation

At higher maturity, AIOps can automate approved actions.

Examples:

Restart a failed service
Scale a deployment
Rollback a release
Drain a bad Kubernetes node
Evacuate an OpenStack compute node
Restart a failed exporter
Open a Jira ticket
Page the correct team
Run a known Ansible playbook

But automation should be controlled carefully.

The safest path is usually:

Detect → Explain → Recommend → Human approval → Execute → Verify

Only mature, low-risk, well-tested actions should be fully automatic.


4. AIOps Compared With Traditional Observability

Traditional observability answers:

What is happening?

AIOps tries to answer:

What is happening?
Why is it happening?
What changed?
What is the blast radius?
What should we do next?
Can we fix it automatically?

Observability provides the evidence.

AIOps provides correlation, reasoning, prioritisation, and action.

They are not competitors. AIOps depends on observability.


5. AIOps in Your MCP Example

In the previous MCP example, the user asks:

Why did gpu-test-01 fail to start?

A traditional engineer might manually check:

openstack server show gpu-test-01
openstack console log show gpu-test-01
openstack port list
openstack hypervisor list
docker logs nova_scheduler
docker logs nova_compute
journalctl on compute nodes
sinfo
nvidia-smi
kubectl get pods
Grafana dashboards
Loki logs
Prometheus GPU metrics

An MCP-enabled AIOps agent could do much of this automatically.

It could query:

OpenStack MCP   → VM state, scheduler errors, Neutron ports
Nova MCP → compute scheduling failure
Neutron MCP → network binding or DHCP issue
Slurm MCP → GPU node allocation or drain state
Prometheus MCP → CPU, RAM, disk, GPU health
Loki MCP → Nova, libvirt, Neutron logs
Kubernetes MCP → GPU Operator / NVIDIA plugin status
Ceph MCP → storage availability
Ansible MCP → known remediation playbooks

Then return something useful:

gpu-test-01 failed because Nova could not schedule the requested PCI device.
The requested alias nvidia-gpu-audio is not defined in nova.conf.
The VM requested a GPU-related PCI alias that the scheduler cannot match.

Evidence:
- Nova API returned PCI alias nvidia-gpu-audio is not defined
- No matching pci_alias exists on the compute configuration
- Hypervisor is otherwise healthy
- Neutron port exists
- Image and flavor are valid

Recommended fix:
Add or remove the correct PCI alias definition, reconfigure Nova, restart nova-scheduler and nova-compute, then retry the server create command.

That is AIOps because the system has moved beyond raw monitoring and into assisted diagnosis.


6. The AIOps Methodology

AIOps is not just a product. It is a way of operating.

A practical methodology looks like this:

1. Instrument everything
2. Centralise telemetry
3. Normalise and enrich the data
4. Correlate events across systems
5. Detect anomalies
6. Identify service impact
7. Recommend actions
8. Automate safe remediation
9. Verify outcomes
10. Learn from incidents

The goal is continuous operational learning.

Every incident should improve the system.


7. The Maturity Model

AIOps adoption usually happens in stages.


Level 1 — Better Visibility

You collect metrics, logs, traces, and events.

Typical tools:

Prometheus
Grafana
Loki
Tempo
OpenTelemetry
Elasticsearch
Alertmanager

At this level, humans still do most of the reasoning.


Level 2 — Alert Correlation

You start grouping alerts into incidents.

Example:

20 alerts → 1 incident

This reduces noise and improves response time.


Level 3 — Assisted Investigation

The system helps engineers investigate.

It can answer:

What changed?
What services are affected?
Are there similar previous incidents?
Which logs matter?
Which deployment caused this?

This is where LLMs and MCP become very useful.


Level 4 — Recommendation

The system recommends fixes.

Example:

Rollback service X
Restart exporter Y
Scale deployment Z
Drain node A
Check Ceph OSD B
Renew certificate C

Humans still approve the action.


Level 5 — Automated Remediation

The system performs low-risk actions automatically.

Example:

Restart crashed exporter
Re-run failed health check
Scale stateless service
Create incident ticket
Attach logs and traces
Notify owning team

High-risk actions still require approval.


8. Who Should Adopt AIOps?

AIOps is most valuable for teams running complex, distributed, high-volume, or business-critical systems.


1. SRE Teams

SRE teams are one of the best fits.

They already care about:

SLIs
SLOs
error budgets
incident response
toil reduction
automation
reliability engineering

AIOps helps SREs reduce repetitive investigation and focus on higher-value engineering.


2. Platform Engineering Teams

Platform teams should adopt AIOps when they operate shared platforms such as:

Kubernetes
OpenStack
VMware
Ceph
GitLab
ArgoCD
CI/CD platforms
internal developer platforms

AIOps helps them understand platform-wide impact and detect shared infrastructure failures.


3. Observability Teams

Observability engineers are central to AIOps.

They provide the telemetry foundation:

metrics
logs
traces
events
dashboards
alerts
instrumentation standards
OpenTelemetry pipelines

Without observability engineering, AIOps becomes guesswork.


4. NOC Teams

Network Operations Centres can use AIOps to reduce noise and improve triage.

Common use cases:

deduplicating network alerts
identifying link degradation
correlating firewall, DNS, BGP, and load balancer events
detecting regional outages
routing incidents to the right team

5. Cloud Infrastructure Teams

Teams running cloud platforms benefit heavily.

Examples:

OpenStack private cloud
AWS landing zones
Azure platforms
GCP platforms
hybrid cloud
multi-cloud

AIOps can correlate compute, network, storage, identity, quota, and deployment issues.


6. HPC and GPU Platform Teams

This is especially relevant for AI infrastructure.

A GPU/HPC platform has many failure domains:

GPU health
PCI passthrough
NVIDIA drivers
CUDA versions
Slurm queues
Kubernetes GPU Operator
RDMA / RoCE
InfiniBand
Ceph / Lustre / WEKA / DDN
container runtimes
job scheduling
tenant quotas
power and thermal limits

AIOps can help detect degraded GPU nodes, scheduling bottlenecks, failed jobs, and noisy tenants.


7. DevOps Teams

DevOps teams can use AIOps to connect deployment activity to runtime impact.

Example:

Deployment happened → error rate increased → SLO burn increased → rollback recommended

This is especially valuable in CI/CD-heavy environments.


8. Enterprises With 24/7 Services

Any organisation with critical always-on services should consider AIOps.

Examples:

financial services
telecoms
cloud providers
SaaS companies
healthcare platforms
universities
government services
e-commerce
AI infrastructure providers

The more expensive downtime is, the more valuable AIOps becomes.


9. Who Does Not Need Heavy AIOps Yet?

AIOps may be overkill for very small or simple environments.

For example:

one small website
one database
low traffic
few alerts
manual operations are still manageable
no 24/7 support requirement

These teams should first focus on:

basic monitoring
good backups
clear alerts
simple runbooks
patching
logging
uptime checks

AIOps should not be used to compensate for weak fundamentals.


10. What AIOps Requires Before It Works Well

AIOps needs a strong foundation.


Good Telemetry

You need reliable metrics, logs, traces, and events.

Bad data produces bad recommendations.


Good Service Ownership

The system must know:

who owns the service
who is on call
what the service depends on
what its SLO is
where the runbook is

Without ownership metadata, routing and remediation are weak.


Good Topology

AIOps needs to understand relationships.

Example:

frontend depends on checkout-api
checkout-api depends on postgres
postgres runs on node db-03
db-03 uses ceph-volume-17
ceph-volume-17 depends on osd-4
osd-4 runs on storage-node-2

Topology allows the system to understand blast radius.


Good Change Data

Most incidents are caused by change.

AIOps should ingest:

deployments
config changes
Terraform changes
Ansible runs
package upgrades
Kubernetes rollouts
OpenStack reconfigures
firewall changes
DNS changes
certificate renewals

Without change data, root cause analysis is incomplete.


Good Runbooks

AIOps automation depends on safe, tested actions.

Examples:

restart service
rollback deployment
clear failed job
rotate certificate
drain node
restart exporter
scale deployment
fail over service

If the runbooks are poor, automation becomes dangerous.


11. Risks and Anti-Patterns

AIOps can fail if implemented badly.


Risk 1: Treating AIOps as Magic

AIOps is not magic.

It cannot fix poor monitoring, poor architecture, missing logs, or unclear ownership.


Risk 2: Automating Too Soon

Do not let AI perform destructive actions before trust is established.

Dangerous actions include:

delete data
restart databases
modify firewall rules
change identity policies
drain production clusters
detach storage
scale expensive GPU workloads

Start with read-only analysis, then human-approved remediation.


Risk 3: Poor Explainability

AIOps must explain why it thinks something is wrong.

Bad:

Root cause: database.

Good:

Root cause is likely PostgreSQL primary saturation.
Evidence:
- connections reached max at 10:42
- API errors began at 10:43
- no deployment occurred
- CPU and disk IO increased on db-01
- similar incident occurred last month

Operations teams need evidence, not vague AI output.


Risk 4: No Human Governance

AIOps should respect operational controls:

approval workflows
audit logs
change windows
RBAC
break-glass access
compliance boundaries
incident commander authority

The AI should support operations, not bypass them.


12. What Success Looks Like

A successful AIOps implementation should improve measurable outcomes.

You should track:

MTTA  - mean time to acknowledge
MTTR - mean time to resolve
MTTD - mean time to detect
alert volume
false positive rate
incident recurrence
toil hours
escalation rate
SLO compliance
change failure rate
automation success rate

The goal is not “we added AI.”

The goal is:

fewer noisy alerts
faster diagnosis
better root cause analysis
lower toil
higher service reliability
safer automation
more consistent operations

13. Practical Adoption Path

A good adoption path would be:

Phase 1: Centralise telemetry
Phase 2: Improve alert quality
Phase 3: Add service ownership and topology
Phase 4: Correlate events and changes
Phase 5: Introduce AI-assisted investigation
Phase 6: Add recommendation workflows
Phase 7: Automate low-risk remediation
Phase 8: Continuously review incidents and improve models/runbooks

For your kind of environment, the most natural starting point would be:

Prometheus/Mimir + Loki + Tempo + OpenTelemetry

Service and infrastructure inventory

Alert correlation

LLM/MCP assistant for investigation

Human-approved Ansible remediation

Closed-loop AIOps

Bottom Line

AIOps is a methodology for making operations smarter by combining:

observability
event correlation
machine learning
LLM reasoning
automation
service ownership
incident management
runbooks
governance

It is trying to solve the operational overload caused by modern distributed systems.

The teams that should adopt it first are:

SRE teams
platform engineering teams
observability teams
cloud infrastructure teams
NOC teams
DevOps teams
HPC/GPU platform teams
enterprises running critical 24/7 services

The best way to think about it is:

Observability tells you what happened.
AIOps helps explain why it happened, what it affects, and what to do next.

Observability Advances for Effective AIOps

Observability is arguably the most important technical component of AIOps.

AIOps is only as good as the operational data it can reason over. The AI layer does not magically understand your systems; it needs evidence. That evidence comes mainly from observability.

You can think of AIOps like this:

AIOps = AI reasoning + Observability data + Automation + ITSM/process + Governance

Or more practically:

Observability provides the evidence.
AI performs correlation and reasoning.
Automation executes safe actions.
ITSM/process manages incidents and ownership.
Governance keeps it controlled and auditable.

Why observability is central

Observability gives the AIOps system the raw material it needs:

Metrics  → What is slow, saturated, failing, or abnormal?
Logs → What actually happened inside the system?
Traces → Where did the request slow down or fail?
Events → What changed in the platform?
Alerts → What conditions crossed operational thresholds?
Topology → What depends on what?

Without this, AI has no reliable basis for diagnosis.

For example, if an AI agent is asked:

Why did gpu-test-01 fail to start?

It needs observability and operational signals from:

OpenStack state
Nova scheduler logs
Neutron events
Libvirt errors
Prometheus metrics
Loki logs
Slurm node state
GPU telemetry
Kubernetes events
Ceph health
Recent Ansible or config changes

The AI then correlates those signals into a root-cause explanation.

The hierarchy of AIOps components

I would rank the major components like this:

1. AI / ML reasoning layer

This includes:

anomaly detection
event correlation
root cause analysis
prediction
recommendation
LLM-based investigation

This is the “intelligence” part.

2. Observability and telemetry

This is the evidence layer:

metrics
logs
traces
events
alerts
service health
infrastructure state
change data

This is probably the most important foundation.

3. Topology and context

AIOps needs to know relationships:

service A depends on service B
pod runs on node X
node X uses storage volume Y
volume Y depends on Ceph OSD Z
tenant workload uses GPU node N

Without topology, the AI may see symptoms but struggle to understand blast radius.

4. Automation and remediation

This turns insight into action:

restart service
rollback deployment
scale workload
drain node
run Ansible playbook
open incident ticket
notify owning team

This is where AIOps becomes operationally powerful.

5. ITSM / incident workflow

This includes:

Jira
ServiceNow
PagerDuty
Opsgenie
Slack
Teams
runbooks
postmortems
ownership metadata

This connects AIOps to how humans actually operate.

6. Governance and security

Especially important if the system can take action:

RBAC
approval gates
audit logs
change windows
human approval
secrets control
tool permissions

So the answer is: yes, but with nuance

Yes — observability is the major foundation after AI.

But in a mature AIOps system, observability alone is not enough. You also need:

good topology
good ownership data
good change tracking
good runbooks
safe automation
clear governance

The best way to phrase it is:

AI is the reasoning engine of AIOps.
Observability is the evidence layer.
Automation is the action layer.
ITSM and governance are the control layer.

For your homelab and OpenStack/GPU/Slurm/Kubernetes direction, the key AIOps foundation would be:

Prometheus/Mimir  → metrics
Loki → logs
Tempo → traces
OpenTelemetry → instrumentation
Alertmanager → alert routing
Grafana → visualisation
MCP → AI tool access
Ansible → controlled remediation
GitLab → change/deployment context

Describe the advances in Observability to make AIOps effectiveThought for 34s

Yes. The big advances in observability that make AIOps genuinely effective are not just “more dashboards.” They are about turning raw telemetry into structured, correlated, machine-readable operational evidence that AI can reason over.

The shift is:

Old monitoring:
Thresholds, dashboards, isolated alerts, manual diagnosis.

Modern observability for AIOps:
Correlated metrics, logs, traces, profiles, events, topology, ownership,
change history, and automation context.

AIOps needs observability to answer five operational questions:

What is happening?
Where is it happening?
Why is it happening?
What changed?
What should we do next?

1. Standardised Telemetry: OpenTelemetry

One of the biggest advances is OpenTelemetry.

Historically, every observability vendor or tool had its own agent, SDK, format, and metadata conventions. That made AIOps hard because the AI had to reason across inconsistent data.

OpenTelemetry helps by giving teams a vendor-neutral way to instrument, generate, collect, and export telemetry such as traces, metrics, and logs. Its Collector provides a common way to receive, process, and export telemetry, reducing the need to run many different agents.

For AIOps, this matters because AI performs better when telemetry has consistent structure.

Example:

Bad:
"error happened on server"

Better:
service.name=checkout-api
deployment.environment=production
k8s.namespace.name=payments
host.name=worker-03
http.response.status_code=500
trace_id=abc123

That structure allows an AI system to correlate across services, clusters, nodes, requests, and deployments.


2. Semantic Conventions

Raw telemetry is not enough. The metadata needs consistent meaning.

OpenTelemetry Semantic Conventions define common names and attributes for operations and data across traces, metrics, logs, profiles, and resources.

This is crucial for AIOps because AI needs to compare like with like.

Without semantic conventions, one team may emit:

app = checkout

another may emit:

service = checkout-api

and another:

component = payments-checkout

An AIOps system then has to guess whether these are the same thing.

With standard conventions, the data becomes more machine-readable:

service.name = checkout-api
service.namespace = payments
deployment.environment = production
k8s.cluster.name = prod-eu-1

That makes correlation, incident grouping, ownership mapping, and root-cause analysis much stronger.


3. Multi-Signal Observability

Traditional monitoring was heavily metrics-focused.

Modern observability combines multiple signals:

Metrics  → What is happening numerically?
Logs → What discrete events occurred?
Traces → How did a request move through the system?
Profiles → Which code consumed CPU, memory, or wall time?
Events → What changed in the platform?

Kubernetes documentation still describes observability around metrics, logs, and traces as the main pillars for understanding cluster state, performance, and health. OpenTelemetry also describes observability signals as system outputs that describe application and platform activity.

For AIOps, this is fundamental.

A metric may say:

API latency is high.

A trace may say:

The latency is in the database query span.

A log may say:

Connection pool exhausted.

A deployment event may say:

New version deployed 5 minutes before the issue.

A profile may say:

CPU is being consumed by JSON serialisation in one function.

The AI can then produce a much better diagnosis than any single signal could provide.


4. Distributed Tracing and Context Propagation

Distributed tracing is one of the most important advances for AIOps.

In a monolith, a request might fail inside one process. In a microservices or cloud-native system, a single user request may cross:

Frontend
API gateway
Auth service
Checkout service
Payment service
Inventory service
Database
Message queue
External SaaS API

A trace connects those hops into one request journey.

For AIOps, tracing gives causal structure. It helps answer:

Where did the request slow down?
Which service returned the error?
Was the failure upstream or downstream?
Which tenant, region, node, or deployment was involved?

This makes root-cause analysis much more precise.

Without tracing, AIOps sees a pile of logs and metrics.

With tracing, it sees a connected execution path.


5. Exemplars: Linking Metrics to Traces

Another important advance is the ability to connect aggregate metrics to specific trace examples.

For example, a dashboard may show:

p99 latency = 2.4 seconds

But the engineer or AI needs to know:

Which actual request was slow?
What did that request do?
Which span caused the delay?

OpenTelemetry metrics support exemplars containing trace and span association fields, and Prometheus/OpenMetrics interoperability includes exemplar conversion rules.

For AIOps, exemplars are powerful because they bridge:

Metric anomaly → actual trace → logs from same request → root cause

That reduces guesswork.


6. Native Histograms and Better Latency Data

AIOps needs good latency distribution data, not just averages.

Average latency hides problems.

Example:

Average latency: 120 ms

That sounds fine, but the distribution may be:

95% of requests: 80 ms
4% of requests: 400 ms
1% of requests: 8 seconds

The 1% tail may be where real user pain exists.

Prometheus native histograms improve how latency and distribution data can be represented, and Prometheus native histograms with standard schemas can map to OpenTelemetry exponential histograms.

For AIOps, this improves:

anomaly detection
SLO burn analysis
performance regression detection
capacity planning
tail-latency investigation

AI needs distribution-aware telemetry to avoid drawing conclusions from misleading averages.


7. Telemetry Pipelines and Data Processing

Another major advance is the rise of programmable telemetry pipelines.

The OpenTelemetry Collector can receive, process, and export telemetry, and its processors can transform, filter, and enrich telemetry as it flows through a pipeline. Grafana Alloy also provides pipelines for telemetry signals such as Prometheus and OpenTelemetry, with support for logs, metrics, traces, and profiles.

This is vital for AIOps because raw telemetry is often messy.

You need to:

drop noisy fields
redact secrets
normalise labels
add environment metadata
add ownership information
route critical data differently
sample high-volume traces
preserve error traces
enrich logs with Kubernetes metadata
convert vendor-specific formats

For AIOps, the telemetry pipeline becomes the data preparation layer.

Bad pipeline:

AI receives noisy, inconsistent, high-volume telemetry.

Good pipeline:

AI receives enriched, normalised, relevant operational evidence.

That is the difference between useful AIOps and expensive confusion.


8. Continuous Profiling

Continuous profiling is another big step forward.

Metrics tell you that CPU is high.

Profiles tell you which code path is consuming CPU.

OpenTelemetry describes profiles as answering which code is responsible for consuming resources, complementing logs, metrics, and traces. The OpenTelemetry Profiles specification describes profiles as an emerging fourth observability signal alongside logs, metrics, and traces. Grafana Pyroscope describes continuous profiling as a systematic method for collecting and analysing performance data from production systems.

For AIOps, profiling helps move from:

The service is slow.

to:

The service is slow because 63% of CPU time is spent in JSON serialisation
inside checkout-api after the latest release.

That is much closer to actionable root cause.


9. eBPF-Based Observability

eBPF has significantly improved infrastructure and network observability.

Cilium describes itself as an eBPF-based solution for networking, observability, and security, providing visibility into workload connectivity. Hubble, built on Cilium, uses eBPF to provide dynamic visibility with detailed insight where needed.

For AIOps, eBPF is valuable because it can observe behaviour at the kernel and network layer without requiring every application to be perfectly instrumented.

It can help answer:

Which pod connected to which service?
Where are packets being dropped?
Is DNS failing?
Is the issue L3, L4, or L7?
Is network policy blocking traffic?
Is the service reachable?
Which process opened this connection?

This is especially important for Kubernetes, OpenStack, service mesh, GPU clusters, and distributed storage platforms.

For your type of environment, eBPF observability is particularly relevant because many failures happen below the application layer:

Neutron networking
Kubernetes CNI
DNS
load balancing
firewalling
pod-to-pod connectivity
GPU node networking
Ceph traffic
Slurm controller-to-worker communication

10. Topology-Aware Observability

AIOps cannot do strong root-cause analysis if it does not understand relationships.

It needs topology.

Example:

frontend
depends on checkout-api
depends on postgres
runs on k8s-worker-03
uses ceph-volume-17
backed by osd-4
runs on storage-node-02

This lets AIOps understand blast radius.

Without topology, AI sees isolated symptoms:

frontend errors
checkout latency
database timeout
Ceph OSD warning
node disk latency

With topology, it can infer:

Ceph OSD degradation on storage-node-02 is affecting postgres,
which is causing checkout-api latency and frontend errors.

This is one of the areas where observability has had to evolve from charts into graph-based operational context.


11. Change-Aware Observability

Most incidents are caused by change.

AIOps becomes much more effective when observability includes change events:

deployments
config changes
Terraform applies
Ansible runs
Kubernetes rollouts
OpenStack reconfigures
package upgrades
certificate renewals
DNS changes
firewall changes
feature flags
autoscaling events

AIOps needs to answer:

What changed before the incident?
Who changed it?
Was it automated?
Which services were affected?
Has this change caused problems before?

This is where GitLab, GitHub Actions, ArgoCD, Terraform, Ansible, Kubernetes events, OpenStack events, and audit logs become part of observability.

For example:

10:01 GitLab deployed checkout-api v2.4.1
10:03 pods restarted
10:04 p99 latency increased
10:05 HTTP 500s increased
10:06 SLO burn alert fired

The likely root cause is not “high latency.”

The likely root cause is the deployment.


12. SLO-Based Observability

Another advance is the move from infrastructure-centric alerts to service-centric SLOs.

Old alerting:

CPU > 90%
Disk > 80%
Pod restarted
Node memory high

Better alerting:

Checkout API availability below SLO
Payment latency budget burning too fast
Login error rate above user-impact threshold

For AIOps, SLOs provide priority.

Not every anomaly matters equally.

A CPU spike on a batch node may be fine.

A small increase in payment failures may be urgent.

SLO-based observability helps AIOps rank incidents by user impact rather than raw technical noise.


13. High-Cardinality and Dimensional Telemetry

Modern systems need dimensional analysis.

You need to slice by:

service
namespace
cluster
region
tenant
customer
version
endpoint
pod
node
GPU model
availability zone
database shard
queue
deployment

Prometheus uses a dimensional data model where time series are identified by a metric name and key-value labels, and PromQL allows teams to query, correlate, and transform time-series data.

For AIOps, dimensions are essential.

Instead of:

API latency is high.

you want:

API latency is high only for:
service=checkout-api
version=v2.4.1
region=eu-west
tenant=customer-a
endpoint=/payment/confirm

That turns a vague incident into a narrowed investigation.


14. Better Log Structure

AIOps performs much better with structured logs.

Bad log:

Something went wrong while processing request.

Better log:

{
"level": "error",
"service.name": "checkout-api",
"trace_id": "abc123",
"user_impact": true,
"order_id": "redacted",
"error.type": "DatabaseConnectionTimeout",
"db.system": "postgresql",
"k8s.pod.name": "checkout-api-7c9fd",
"deployment.version": "v2.4.1"
}

Structured logs let the AI search, group, and correlate events reliably.

For AIOps, this is a major difference.

Unstructured logs require interpretation.

Structured logs provide evidence.


15. Observability for Automation

AIOps is not only about diagnosis. It also needs verification.

Before remediation:

Is the service unhealthy?
What is the likely root cause?
Is the proposed action safe?

After remediation:

Did the error rate fall?
Did latency recover?
Did the pod restart cleanly?
Did the SLO burn rate stabilise?
Did the same alert return?

Observability provides the feedback loop for automation.

Without observability, automation is blind.

A mature AIOps loop looks like this:

Detect

Correlate

Diagnose

Recommend

Approve

Execute

Verify

Learn

The “verify” and “learn” stages depend heavily on observability.


16. AI-Readable Operational Context

The latest practical advance is making observability data usable by AI agents.

That means exposing operational systems through APIs, query layers, or protocols such as MCP-style tool access.

The AI needs controlled access to:

metrics queries
log search
trace lookup
profile analysis
Kubernetes state
OpenStack state
Slurm queue state
Ceph health
GitLab deployments
Ansible runbooks
incident history
service ownership

This turns observability from something humans look at into something AI can query and reason over.

For example:

User asks:
"Why did gpu-test-01 fail to start?"

AI queries:
OpenStack state
Nova logs
Neutron events
Prometheus GPU metrics
Slurm state
Kubernetes GPU operator status
Ceph health
recent config changes

AI replies:
"Nova failed to schedule the VM because the requested PCI alias is not defined.
Neutron and storage are healthy. The failure is isolated to Nova PCI configuration."

That is observability becoming operational intelligence.


How These Advances Make AIOps Effective

The relationship is simple:

Observability advanceWhat it gives AIOps
OpenTelemetryStandard telemetry collection
Semantic conventionsConsistent metadata
MetricsQuantitative system health
LogsEvent-level explanation
TracesRequest-level causality
ProfilesCode-level resource attribution
ExemplarsLink from metric anomaly to trace
eBPFKernel/network visibility
TopologyDependency and blast-radius context
Change events“What changed?” analysis
SLOsBusiness/user-impact priority
Telemetry pipelinesClean, enriched, governed data
Structured logsMachine-readable evidence
Automation feedbackSafe remediation verification

For Your OpenStack / Kubernetes / Slurm / GPU Homelab

For your environment, the observability stack that would make AIOps effective should include:

Metrics:
Prometheus / Mimir

Logs:
Loki

Traces:
Tempo

Profiles:
Pyroscope

Collection and pipelines:
OpenTelemetry Collector or Grafana Alloy

Dashboards and exploration:
Grafana

Alerting:
Alertmanager

Kubernetes network observability:
Cilium / Hubble / eBPF

Change context:
GitLab CI/CD, Ansible logs, OpenStack reconfigure events

Infrastructure state:
OpenStack, Nova, Neutron, Glance, Cinder, Ceph

GPU/HPC state:
Slurm, NVIDIA DCGM exporter, nvidia-smi, GPU Operator if using Kubernetes

Automation:
Ansible playbooks with human approval

AI access:
MCP-style tool interface to query metrics, logs, traces, infrastructure, and runbooks

The important design principle is:

Do not just collect telemetry.
Make telemetry correlated, structured, enriched, searchable, and safe for AI to use.

That is what turns observability into a real AIOps foundation.

AI Observability

AI is not killing observability as a discipline, but it is fundamentally changing how observability is done. The traditional model of “collect everything, store everything, and let humans investigate later” is becoming increasingly impractical in AI-driven infrastructures.

1. Telemetry volume is exploding

Modern systems produce far more telemetry than they did five years ago.

An AI factory may contain:

  • Tens of thousands of GPUs
  • Hundreds of thousands of CPU cores
  • High-speed fabrics (RoCE, InfiniBand)
  • Kubernetes
  • Distributed storage (Ceph, Lustre, GPFS)
  • AI inference services
  • LLM gateways

Each component exports metrics, logs, traces and events.

For example:

2020

100 servers

100 million metrics/day

2026

20,000 GPUs
30,000 CPUs
5,000 switches



Several trillion data points/day

Humans cannot meaningfully explore that volume.


2. Dashboards don’t scale

Traditional observability assumes people sit looking at Grafana dashboards.

Reality:

  • nobody watches 400 dashboards
  • nobody remembers 2,000 PromQL queries
  • nobody notices slow drift

Instead people increasingly ask:

“Why did training become slower?”

AI investigates.

Not humans.


3. Alert fatigue becomes impossible

Large organisations often generate

  • 50,000 alerts/day
  • 100,000 log anomalies/day

Historically:

Prometheus
↓
Alertmanager
↓
PagerDuty
↓
Human

Future:

Prometheus
↓
AI correlation
↓
Root cause
↓
Human receives one explanation

Instead of:

127 alerts

Engineer receives

GPU node gpu-128 experienced ECC errors causing NCCL retries which slowed training by 18%.


4. Humans don’t query telemetry anymore

Traditional workflow

Grafana
↓
Zoom
↓
PromQL
↓
Logs
↓
Tempo
↓
Find issue

Future

"Why are customer requests slower?"
↓
AI
↓
queries everything
↓
returns explanation

Natural language replaces much of manual exploration.


5. AI is becoming the first investigator

Large enterprises increasingly build systems like:

Telemetry
↓
LLM
↓
Reasoning
↓
Correlation
↓
Recommendation

Instead of asking engineers to join the dots.


6. Sampling changes everything

Historically:

Store every log.

Now:

AI decides

Keep
Discard
Summarise
Compress

Observability becomes intelligent instead of passive.


7. Root cause becomes graph reasoning

Today’s tools often correlate:

metric
+
trace
+
log

Future systems correlate:

  • topology
  • Kubernetes
  • network
  • storage
  • deployments
  • Git commits
  • feature flags
  • incidents
  • Slack discussions
  • runbooks

into one knowledge graph.

AI reasons across all of it.


8. AI reduces the need for experts

Today an SRE may spend years learning:

  • PromQL
  • LogQL
  • TraceQL
  • Elasticsearch
  • Loki
  • Tempo
  • Mimir
  • Kubernetes
  • networking

Future engineers may simply ask:

Why did latency increase?

The AI handles the underlying queries.

Expertise shifts from query syntax to validating conclusions and understanding system design.


9. Observability vendors are changing

Every major platform is investing heavily in AI assistants.

Examples include:

  • Grafana Labs
  • Datadog
  • Dynatrace
  • New Relic
  • Elastic
  • Splunk

They’re moving from:

  • dashboards

to:

  • AI copilots
  • automatic investigations
  • causal analysis
  • anomaly explanations
  • remediation suggestions

10. Cost is becoming the biggest problem

Storage costs are growing faster than engineering budgets.

For many enterprises:

Telemetry generated

100 TB/day

Engineers never inspect 99% of it.

AI can:

  • summarize repetitive logs
  • identify low-value telemetry
  • adapt sampling dynamically
  • retain only diagnostically useful data

This can significantly reduce storage and processing costs while preserving investigative value.


What is actually dying?

It is not observability itself, but the manual workflows around it.

Declining:

  • Manual dashboard creation
  • Hand-written alert rules for every scenario
  • Endless log searching
  • Human correlation across tools
  • Memorizing complex query languages

Growing:

  • AI-assisted investigations
  • Automated root-cause analysis
  • Predictive anomaly detection
  • Natural-language querying
  • Knowledge-graph reasoning
  • Automated remediation

What this means for SREs

For someone with your background in observability platforms, this shift is an opportunity rather than a threat. The value moves away from being the person who can write the most intricate PromQL query and toward designing the telemetry architecture, ensuring data quality, integrating AI agents with monitoring systems, and deciding when AI-generated conclusions can be trusted.

The emerging role is less “dashboard builder” and more Observability Platform Engineer or AI Operations (AIOps) Engineer—someone who builds systems that enable AI to investigate production issues safely and accurately. That combination of platform engineering, distributed systems, and AI integration is likely to be in increasing demand over the next decade.

What is AI Observability?

AI-era observability is moving from human-driven inspection to machine-assisted reasoning over telemetry, topology, history and operational knowledge.

The key shift is this:

Old observability:

Metrics + logs + traces

Dashboards and alerts

Human investigates

Human decides

Human fixes


AI-era observability:

Metrics + logs + traces + topology + deployments + runbooks + incidents

AI correlation and reasoning layer

Probable cause, blast radius, next action

Human approval or automated remediation

Below is a detailed breakdown of the six areas.


1. AI-assisted investigations

What it means

AI-assisted investigation is where an AI system acts like a junior SRE investigator sitting beside you.

It does not necessarily fix the issue automatically. Its main job is to reduce the time spent asking basic investigative questions.

Instead of you manually jumping between:

Grafana → Prometheus/Mimir → Loki → Tempo → Kubernetes → Git → Slack → Runbooks

you ask something like:

Why did checkout latency increase after 14:05?

The AI then queries multiple systems and returns a structured investigation.


What it does

A good AI investigation assistant can:

  • Detect the relevant service, namespace, cluster or tenant.
  • Pull related metrics.
  • Search logs around the incident window.
  • Inspect traces for slow spans.
  • Compare current behaviour against baseline behaviour.
  • Check recent deployments.
  • Check Kubernetes events.
  • Check node, pod, container and network health.
  • Retrieve relevant runbooks.
  • Summarise likely causes.
  • Recommend next diagnostic steps.

Example

You ask:

Why is the inference API slower?

The AI investigates:

1. Latency increased at 10:17.
2. p95 rose from 420 ms to 1.8 s.
3. Error rate did not increase.
4. GPU utilisation remained high.
5. Queue depth increased.
6. New model version was deployed at 10:12.
7. Logs show repeated batching timeout warnings.
8. Traces show delay before GPU execution, not during execution.

Result:

Likely issue:
The model service is queueing requests before GPU execution.

Probable cause:
The new batching configuration increased max_batch_wait_ms from 20 ms to 250 ms.

Recommended action:
Rollback batching config or reduce batch wait threshold.

That is much faster than manually checking ten dashboards.


What data it needs

AI-assisted investigation works best when it has access to:

Metrics:
- RED metrics: rate, errors, duration
- USE metrics: utilisation, saturation, errors
- Kubernetes pod/node metrics
- GPU metrics
- Network metrics
- Storage metrics

Logs:
- Application logs
- Kubernetes events
- System logs
- Ingress/controller logs
- Deployment logs

Traces:
- Request path
- Slow spans
- Upstream/downstream dependencies
- Database/storage/API calls

Context:
- Deployment history
- Git commits
- Feature flags
- Config changes
- Runbooks
- Incident history
- Service ownership

Without context, AI just summarises telemetry. With context, it can investigate.


SRE value

For SREs, this means less time doing mechanical investigation and more time validating the diagnosis.

The future SRE skill is not just:

Can I write PromQL?

It becomes:

Can I design telemetry so AI can reason correctly?
Can I validate the AI's conclusion?
Can I prevent unsafe remediation?
Can I encode good operational knowledge into the platform?

2. Automated root-cause analysis

What it means

Automated root-cause analysis, or automated RCA, is the process of identifying the most likely initiating cause of a production issue without relying entirely on manual human correlation.

It tries to answer:

What actually started the incident?

Not merely:

What symptoms are currently visible?

This distinction matters.


Symptom versus root cause

Example incident:

Customer latency is high.
API pods are slow.
Database queries are slow.
Storage latency is high.
Ceph OSDs are rebalancing.
One storage node has a failing disk.

The symptoms are:

High API latency
Slow database responses
Increased request duration
More timeout warnings

The probable root cause is:

A failing disk caused Ceph recovery/rebalancing,
which increased storage latency,
which slowed the database,
which slowed the API.

Automated RCA attempts to build that causal chain.


How automated RCA works

There are several techniques.

1. Temporal correlation

The system checks what changed first.

10:01 disk errors begin
10:03 Ceph recovery starts
10:05 storage latency rises
10:07 database latency rises
10:09 API latency rises
10:10 customer alerts fire

The earliest credible abnormal event is often close to the root cause.


2. Topology-aware analysis

The system understands dependencies.

frontend

checkout-api

postgres

ceph/rbd

osd-node-07

If osd-node-07 is unhealthy, and all dependent systems are degraded, the RCA engine can infer blast radius.


3. Change correlation

The system checks recent changes:

Deployments
Config changes
Feature flags
Kernel updates
Node drains
Network changes
Storage migrations
Certificate rotations
DNS changes
Autoscaling events

Many incidents are change-induced. A useful RCA system always asks:

What changed recently?

4. Statistical anomaly ranking

The system ranks abnormal signals.

For example:

Signal                         Abnormality score
GPU ECC errors 0.98
NCCL retry count 0.94
Training step duration 0.91
CPU usage 0.22
Memory usage 0.18

The AI focuses on the strongest abnormal signals.


5. Causal graph reasoning

This is more advanced.

Instead of treating metrics as isolated time series, the system builds a causal model:

Bad disk
→ Ceph recovery
→ Storage latency
→ Database latency
→ API latency
→ Customer impact

This is much closer to how an experienced SRE thinks.


Example automated RCA output

Incident:
Checkout latency p95 increased from 300 ms to 2.4 s.

Likely root cause:
PostgreSQL read latency increased due to degraded Ceph RBD volume performance.

Evidence:
- API latency increased at 13:42.
- PostgreSQL read latency increased at 13:39.
- Ceph pool latency increased at 13:36.
- OSD 12 reported slow ops and disk errors at 13:34.
- No relevant application deployment occurred in the previous hour.

Blast radius:
- checkout-api
- payment-api
- order-history-api

Recommended action:
- Mark OSD 12 out if disk errors continue.
- Move affected workload if possible.
- Check Ceph recovery/backfill limits.
- Consider temporarily scaling API timeout thresholds.

What makes automated RCA hard

Automated RCA is difficult because distributed systems are messy.

Common problems:

Correlation is not causation.
Multiple things can break at once.
Telemetry may be missing.
Logs may be noisy.
Clocks may not be perfectly synchronised.
Service dependency maps may be stale.
The root cause may be outside the monitored system.

This is why good automated RCA usually gives:

Probable cause
Confidence level
Supporting evidence
Contradicting evidence
Recommended next checks

It should not pretend to be certain when it is not.


3. Predictive anomaly detection

What it means

Predictive anomaly detection tries to detect abnormal behaviour before it becomes a major incident.

Traditional alerting says:

Alert when disk usage > 90%.

Predictive alerting says:

Disk usage is growing at a rate that will hit 90% in 11 hours.

That is a major shift.


Traditional threshold alerting

Example:

Alert: DiskAlmostFull
Condition: disk_used_percent > 90

This is simple and useful, but it misses context.

A disk at 85% may be fine if it grows slowly.

A disk at 60% may be dangerous if it is growing rapidly.


Predictive anomaly detection

Predictive systems look at behaviour over time:

Normal pattern:
- CPU rises during business hours
- drops overnight
- spikes during batch processing

Abnormal pattern:
- CPU rises at midnight
- no scheduled job exists
- memory grows continuously
- request rate is normal

The system detects that the pattern is unusual, even if no hard threshold has been crossed.


Types of predictive anomalies

1. Trend-based prediction

Useful for capacity planning.

Disk usage will reach 90% in 3 days.
Mimir object storage will exceed budget in 12 days.
Kafka partition disk will fill in 9 hours.
Ceph pool will hit near-full ratio this weekend.

2. Seasonality-aware anomaly detection

Useful for normal daily/weekly cycles.

Example:

CPU at 80% at 10:00 Monday may be normal.
CPU at 80% at 03:00 Sunday may be abnormal.

The system learns expected patterns.


3. Multivariate anomaly detection

Looks at several signals together.

For example:

Request rate: normal
Error rate: normal
Latency: high
CPU: normal
Database latency: high
Network retransmits: high

Individually, some metrics may not trigger alerts. Together, they reveal an abnormal condition.


4. Behavioural drift detection

Useful in AI and ML platforms.

Example:

Training jobs are completing successfully,
but average step time has increased by 12% over two weeks.

No incident has occurred yet, but performance is drifting.


5. Saturation prediction

Very useful for SRE.

GPU memory saturation likely within 40 minutes.
Kubernetes node memory pressure likely in 2 hours.
Ceph recovery will saturate backend network.
Kafka consumer lag will exceed SLO in 25 minutes.

Example

A predictive anomaly detector observes:

Mimir ingest rate: stable
Object storage write latency: slowly increasing
Compactor duration: increasing
Query latency: increasing
Store-gateway cache hit rate: decreasing

It predicts:

Within 6 hours, users will experience slow dashboard loads.

The remediation might be:

Scale store-gateways.
Check object storage latency.
Increase cache.
Review compactor backlog.

Why it matters

Predictive anomaly detection changes operations from:

React after customer impact

to:

Intervene before customer impact

That is the core SRE value.


4. Natural-language querying

What it means

Natural-language querying allows engineers to ask operational questions in plain English instead of writing PromQL, LogQL, TraceQL, SQL or Elasticsearch queries manually.

Example:

Show me p95 latency for checkout-api over the last 6 hours,
split by Kubernetes namespace.

The AI converts that into the right query.


Traditional workflow

You need to know the query language:

histogram_quantile(
0.95,
sum by (le, namespace) (
rate(http_request_duration_seconds_bucket{
service="checkout-api"
}[5m])
)
)

With natural-language querying:

What is checkout-api p95 latency by namespace for the last 6 hours?

The AI generates and executes the query.


Where this is useful

Natural-language querying is useful across:

Metrics:
- Prometheus
- Mimir
- Thanos
- VictoriaMetrics

Logs:
- Loki
- Elasticsearch/OpenSearch
- ClickHouse

Traces:
- Tempo
- Jaeger
- OpenTelemetry backends

Databases:
- PostgreSQL
- BigQuery
- Snowflake
- ClickHouse

Cloud APIs:
- Kubernetes
- AWS
- Azure
- GCP
- OpenStack

Example questions

Which services had the largest increase in error rate in the last hour?

Show me pods that restarted after the latest deployment.

Find logs for payment-api where timeout errors increased.

Which traces spent the most time waiting on PostgreSQL?

Which Kubernetes nodes have high network retransmits?

Show me Ceph OSDs with rising latency and degraded placement groups.

Which GPU nodes show ECC errors or thermal throttling?

The important part: semantic mapping

Natural-language querying is not just text-to-query.

It needs to understand your telemetry naming.

For example, you may ask:

Show API latency.

But your metrics may be called:

http_request_duration_seconds_bucket
nginx_ingress_controller_request_duration_seconds_bucket
istio_request_duration_milliseconds_bucket
app_http_server_duration_bucket

The AI needs a semantic layer that maps human concepts to real telemetry.


Good natural-language querying needs

Metric catalogue
Label documentation
Service ownership map
Namespace conventions
Dashboard metadata
Runbook links
Known-good query examples
SLO definitions
Deployment metadata

Without that, the AI may generate syntactically valid but operationally useless queries.


Risk: hallucinated queries

Natural-language querying can be dangerous if it invents metric names.

Bad output:

rate(checkout_latency_seconds[5m])

But that metric may not exist.

Better behaviour:

I could not find a metric named checkout_latency_seconds.
I found http_request_duration_seconds_bucket with service="checkout-api".
Using that instead.

The AI should verify queries against the actual telemetry backend.


SRE impact

SREs will still need to understand PromQL, LogQL and traces, but less time will be spent manually composing queries.

The valuable skill becomes designing the semantic layer:

Good metric names
Useful labels
Consistent service metadata
Accurate ownership data
Clear runbooks
Well-documented SLOs

5. Knowledge-graph reasoning

What it means

Knowledge-graph reasoning connects operational facts into a graph so AI can reason over relationships.

Traditional observability stores data like this:

Metric:
checkout-api p95 latency = 2.1s

Log:
timeout connecting to postgres

Trace:
checkout-api → postgres took 1.8s

Kubernetes:
postgres pod moved to node-12

Infrastructure:
node-12 has disk pressure

A knowledge graph connects those facts:

checkout-api
depends_on → postgres
runs_in → namespace prod
owned_by → payments-team

postgres
runs_on → node-12
uses → ceph-rbd-volume-44

node-12
has_condition → disk_pressure

ceph-rbd-volume-44
backed_by → ceph-pool-prod

Now the AI can reason across relationships.


Why graphs matter

Most incidents are not isolated.

They involve chains:

Application → runtime → Kubernetes → node → network → storage → hardware

Dashboards show symptoms. Graphs show relationships.


Example graph

Customer impact

checkout-api latency

postgres query latency

RBD volume latency

Ceph OSD slow ops

failing NVMe device

A graph-based system can move up and down this chain.


What goes into the graph

A strong observability knowledge graph includes:

Services
APIs
Databases
Queues
Kubernetes namespaces
Pods
Nodes
Clusters
Storage volumes
Ceph pools
Network devices
Load balancers
Ingress controllers
Deployments
Git commits
Feature flags
SLOs
Alerts
Incidents
Runbooks
Owners
Escalation paths
Cloud resources
OpenStack projects
GPU nodes
Training jobs

How the graph is built

Data sources may include:

Kubernetes API
Prometheus/Mimir labels
OpenTelemetry resource attributes
Service mesh telemetry
CMDB
Terraform state
GitOps repositories
CI/CD systems
Incident management tools
Cloud APIs
OpenStack APIs
Ceph APIs
Network controllers
Runbooks and docs

The graph is continuously updated.


Example reasoning

Question:

Why are training jobs slower on rack 3?

The graph helps the AI discover:

Training-job-982
runs_on → gpu-node-31, gpu-node-32, gpu-node-33
located_in → rack-3
uses_network → leaf-switch-3a
uses_storage → lustre-client
depends_on → metadata-server-2

Telemetry shows:

leaf-switch-3a has rising packet drops
NCCL retries increased
GPU utilisation has sawtooth pattern
training step time increased

AI conclusion:

The training slowdown is likely caused by network instability on rack 3,
not by GPU compute saturation.

That is knowledge-graph reasoning.


Why this matters for AI data centres

AI/HPC environments are dependency-heavy.

A single training workload may depend on:

GPU health
GPU memory
NVLink/NVSwitch
PCIe
RoCE/InfiniBand
Leaf-spine network
Storage bandwidth
Metadata servers
Container runtime
Kubernetes scheduler
Slurm scheduler
Image registry
Secrets
DNS
Authentication
Object storage

A flat dashboard cannot represent that well. A graph can.


6. Automated remediation

What it means

Automated remediation is when the system not only detects and diagnoses an issue, but also takes corrective action.

This is the most powerful and most dangerous part of AI-era observability.

It moves from:

Observe → Alert → Human fixes

to:

Observe → Diagnose → Decide → Act → Verify

Simple automated remediation

Low-risk examples:

Restart a failed pod.
Scale a deployment from 3 to 5 replicas.
Clear a stuck job.
Rotate a saturated log file.
Drain a bad Kubernetes node.
Open an incident ticket.
Create a Slack/PagerDuty summary.
Rollback a known-bad deployment.
Increase queue consumers.

Advanced automated remediation

Higher-risk examples:

Move workloads away from degraded storage.
Change Ceph recovery/backfill settings.
Disable a feature flag.
Rebalance Kafka partitions.
Quarantine a GPU node.
Remove a bad node from a load balancer.
Apply a network policy change.
Trigger disaster recovery failover.
Patch a vulnerable service.

These require stronger guardrails.


The remediation loop

A safe remediation system should work like this:

1. Detect
Something abnormal happened.

2. Diagnose
Determine probable cause and confidence.

3. Propose
Generate a remediation plan.

4. Check policy
Is this action allowed?
Is the blast radius acceptable?
Is approval required?

5. Act
Execute the change.

6. Verify
Did the metric improve?
Did errors reduce?
Did customer impact stop?

7. Roll back
If not improved, revert or escalate.

8. Learn
Record the incident and outcome.

Example

Issue:

checkout-api error rate increased after deployment.

AI investigation:

New version deployed at 09:03.
Errors began at 09:05.
Only pods running version v2.7.4 are affected.
Previous version v2.7.3 had no errors.

Remediation proposal:

Rollback checkout-api from v2.7.4 to v2.7.3.

Policy check:

Allowed because:
- service has rollback automation
- error rate exceeds SLO threshold
- last known-good version exists
- no database migration detected

Action:

kubectl rollout undo deployment/checkout-api

Verification:

Error rate returned to baseline after 4 minutes.
p95 latency returned to normal.
Incident summary created.

Guardrails are essential

Automated remediation must not be a reckless agent with production write access.

Good guardrails include:

Read-only by default
Approval required for high-risk actions
Change windows
Blast-radius limits
Dry-run mode
Policy-as-code
RBAC
Audit logs
Rollback plans
Rate limits
Canary execution
Human confirmation for destructive actions

For example:

Allowed automatically:
- restart one unhealthy pod
- scale a stateless service within limits
- create an incident ticket

Requires approval:
- drain production node
- rollback payment service
- modify firewall/network policy
- change Ceph recovery settings
- fail over database

How these six areas fit together

They are not separate ideas. They form a pipeline.

Natural-language querying

Lets humans ask better questions

AI-assisted investigations

Gathers evidence automatically

Knowledge-graph reasoning

Understands relationships and dependencies

Automated root-cause analysis

Identifies probable initiating cause

Predictive anomaly detection

Finds issues before they become incidents

Automated remediation

Fixes or mitigates the issue

A mature AI observability platform combines all six.


Practical architecture for an AI observability platform

A realistic architecture could look like this:

Telemetry sources
├─ Prometheus / Mimir metrics
├─ Loki logs
├─ Tempo traces
├─ Kubernetes events
├─ Ceph / storage metrics
├─ GPU metrics
├─ Network telemetry
├─ CI/CD events
├─ Git commits
└─ Incident history



Data normalization layer
├─ OpenTelemetry attributes
├─ Service naming standards
├─ Environment labels
├─ Owner/team labels
└─ SLO metadata



Context layer
├─ Runbooks
├─ Architecture docs
├─ Past incidents
├─ Known failure modes
├─ Deployment history
└─ Dependency maps



AI reasoning layer
├─ LLM
├─ RAG over runbooks/docs
├─ Query generation
├─ Anomaly detection
├─ Causal graph reasoning
└─ RCA ranking



Action layer
├─ Human-readable incident summary
├─ Suggested next steps
├─ Ticket creation
├─ Slack/PagerDuty update
├─ Safe automation
└─ Approved remediation

What you would build first as an SRE

I would not start with fully automated remediation. That is too risky.

The sensible maturity path is:

Stage 1: AI-assisted read-only investigation

Build a tool that can answer:

What changed?
What alerts fired?
What services are affected?
What logs are unusual?
What traces are slow?
What runbook applies?

No write actions.


Stage 2: Natural-language query assistant

Allow engineers to ask:

Show me p95 latency by service.
Find logs for this incident window.
Show me failed pods after the deployment.
Compare today’s error rate with yesterday.

The assistant should show the generated query so the engineer can verify it.


Stage 3: Incident summariser

Generate structured summaries:

Incident:
Impact:
Start time:
Affected services:
Probable cause:
Evidence:
Actions taken:
Current status:
Recommended next steps:

This alone saves huge operational time.


Stage 4: RCA recommendation engine

Add correlation with:

Deployments
Kubernetes events
Node health
Storage health
Network telemetry
Recent config changes

Output probable root cause with confidence.


Stage 5: Predictive alerting

Start with safer predictions:

Disk will fill.
Object storage usage will exceed budget.
Kafka lag will breach SLO.
Ceph pool will hit near-full.
Certificate will expire.
GPU nodes are showing increasing ECC errors.

Stage 6: Human-approved remediation

The AI proposes actions, but humans approve.

Example:

Recommended action:
Drain node gpu-17 and reschedule workloads.

Reason:
GPU ECC errors increased and training retries are affecting jobs.

Approval required:
Yes.

Stage 7: Limited automatic remediation

Only allow automation for narrow, reversible, low-risk actions.

Restart crashed pod
Scale stateless deployment
Reopen failed consumer
Create incident ticket
Disable noisy alert temporarily with expiry

Main risks

AI observability can go wrong if the system has poor telemetry or too much authority.

1. Bad telemetry in, bad reasoning out

If labels are inconsistent, traces are incomplete, or logs are unstructured, AI conclusions will be weak.


2. Hallucinated root cause

The AI may sound confident while being wrong.

Always require:

Evidence
Confidence
Alternative theories
Query links
Raw data references

3. Unsafe remediation

A bad automated action can make an incident worse.

Example:

AI sees high memory.
AI restarts all pods.
All pods restart at once.
Outage gets worse.

That is why blast-radius control matters.


4. Hidden cost explosion

AI investigation can generate expensive backend queries.

A poorly controlled AI assistant may run huge queries across logs, traces and metrics.

You need:

Query limits
Timeouts
Caching
Sampling
Tenant controls
Cost visibility

5. Security and access control

The AI should not see or do everything.

It needs RBAC:

Read-only access for most users
Sensitive log masking
No secret exposure
Audit trail
Approval for write actions
Tenant isolation

The big picture

These six capabilities are the future of observability:

CapabilityMain purposeHuman role
AI-assisted investigationsSpeed up incident analysisValidate findings
Automated RCAIdentify probable causeJudge evidence
Predictive anomaly detectionPrevent incidents earlierTune models and thresholds
Natural-language queryingMake telemetry easier to accessVerify generated queries
Knowledge-graph reasoningUnderstand system relationshipsMaintain accurate topology
Automated remediationFix or mitigate issuesDefine guardrails and approve risk

The core change is this:

Observability is no longer just about collecting telemetry.

It is becoming a reasoning system over telemetry.

For SREs, the opportunity is to become the person who builds and governs that reasoning system: telemetry quality, context, automation safety, incident workflows, and trust boundaries.

Commercial AI Observability

Commercial companies are building AI into observability in two directions:

  1. AI for observability — using AI to investigate, correlate, explain, predict and remediate production issues.
  2. Observability for AI — monitoring LLMs, agents, RAG pipelines, vector databases, model quality, hallucinations, token cost, latency, drift and safety.

So the product shift is not just “add a chatbot to dashboards.” The bigger move is toward an AI operations layer that sits above metrics, logs, traces, events, topology and runbooks.

Telemetry + topology + deployments + logs + traces + incidents + runbooks

AI reasoning layer

Explain issue → find cause → predict risk → recommend/execute action

1. Datadog

Datadog is building AI into its platform around Bits AI, Watchdog, and LLM/Agent Observability.

Datadog’s Watchdog is its AI engine for automated alerts, insights and root-cause analysis across Datadog telemetry. It continuously monitors infrastructure and surfaces important signals to help teams detect, troubleshoot and resolve issues.

Datadog’s Bits AI SRE is positioned as an always-on AI SRE agent that helps handle troubleshooting and alerts, with Datadog describing it as able to pinpoint root causes faster by using Datadog’s incident and telemetry context.

Datadog is also pushing Bits AI Agents and Agent Builder, where the platform can build custom AI agents that investigate issues, make decisions and take action using Datadog and third-party data, with prebuilt actions across cloud, security, CI/CD and collaboration tooling.

For the second direction, Datadog has Agent Observability / LLM Observability, aimed at tracing, evaluating and improving LLM-powered applications and AI agents. Datadog says each LLM application request can be represented as a trace, allowing teams to investigate root cause, operational performance, quality, privacy and safety.

In plain SRE terms, Datadog is building:

Datadog AI direction:

Watchdog
→ automatic anomaly detection
→ automated insights
→ RCA suggestions

Bits AI
→ natural-language investigation
→ AI SRE assistant
→ incident summarisation
→ workflow automation

Bits AI Agents
→ custom agentic workflows
→ investigation agents
→ remediation/documentation agents

LLM / Agent Observability
→ traces for LLM calls
→ prompt/response monitoring
→ quality, privacy, safety checks
→ AI-agent debugging

Datadog is also doing deeper model work: its Toto time-series foundation model is specifically designed for observability time-series forecasting and was trained partly on Datadog observability data.


2. Dynatrace

Dynatrace has probably been the most explicit about putting causal AI at the centre of observability.

Its AI engine is Davis AI / Dynatrace Intelligence. Dynatrace describes its AI approach as combining predictive AI, causal AI and generative AI over unified observability and security data to automate workflows.

Dynatrace’s key differentiator is that it does not want the AI to merely correlate metrics. It wants the platform to understand causality: what caused what, what depends on what, and what failure actually triggered the incident. Dynatrace describes causal AI as using causal and deterministic techniques to determine underlying causes and effects rather than just relying on correlation.

Dynatrace also presents Dynatrace Intelligence as combining deterministic insights with agentic action for prevention, remediation and optimisation at scale.

For AI workloads, Dynatrace has AI and LLM Observability for monitoring, optimising and securing generative AI apps, LLMs and agentic workflows, with emphasis on performance, explainability and compliance.

In SRE terms, Dynatrace is building:

Dynatrace AI direction:

Davis AI / Dynatrace Intelligence
→ anomaly detection
→ causal root-cause analysis
→ topology-aware problem detection
→ predictive risk detection
→ generative explanations
→ workflow automation

Causal AI
→ dependency-aware analysis
→ fault-tree-style reasoning
→ root cause, not just symptom correlation

AI and LLM Observability
→ GenAI app monitoring
→ LLM and agentic workflow visibility
→ explainability
→ compliance-oriented monitoring

The important point: Dynatrace is trying to make observability less like “search through telemetry” and more like automated dependency-aware diagnosis.


3. Splunk

Splunk is building AI into observability through Splunk AI Assistant in Observability Cloud, broader AI Observability, and AI/agent monitoring.

Splunk’s AI Assistant in Observability Cloud uses observability data from metrics, traces, logs and alerts through a chat interface inside Splunk Observability Cloud.

Splunk says the AI Assistant can analyze data across APM, Infrastructure Monitoring, Database Monitoring, RUM and log analytics to help with root-cause analysis.

Splunk is also building “observability for AI” capabilities. Its Splunk Observability for AI is described as full-fidelity monitoring and troubleshooting across AI applications and the AI infrastructure components used to build them.

Splunk’s AI Agent Monitoring aims to correlate degraded AI agent/model performance and track operational metrics such as latency and errors alongside quality/security metrics such as hallucinations, bias, drift, accuracy, cost and token usage.

Splunk’s AI Observability positioning is broader: observe and optimise performance, quality, cost and security across agents, LLMs, vector databases and infrastructure.

In SRE terms, Splunk is building:

Splunk AI direction:

AI Assistant in Observability Cloud
→ natural-language investigations
→ logs + metrics + traces + alerts analysis
→ RCA assistance
→ incident summarisation

AI Observability
→ AI application monitoring
→ AI infrastructure monitoring
→ agent performance tracking
→ LLM quality and safety monitoring

AI Agent Monitoring
→ latency and errors
→ hallucination tracking
→ bias/drift/accuracy
→ token and cost visibility
→ model and agent reliability

Splunk’s direction is very aligned with its historical strength: search, correlation and operational analytics, now wrapped in AI-assisted investigation and AI workload monitoring.


4. New Relic

New Relic is building AI into its platform through New Relic AI, AI-powered observability features, and AI Monitoring / LLM observability.

New Relic says New Relic AI can help instrument systems, generate system health reports and identify alert coverage gaps for full-stack observability.

New Relic has also positioned its platform as AI-powered observability that correlates telemetry across the stack to isolate root cause and reduce operational toil.

For LLM applications, New Relic AI monitoring captures telemetry from AI-powered apps through APM agents and collects data from external LLMs and vector stores.

New Relic’s AI monitoring focuses on troubleshooting, comparing and optimising LLM prompts and responses for performance, cost and quality issues such as hallucination, bias and toxicity.

It also supports LLM observability through OpenLIT integration, which automatically generates traces and metrics for LLM and VectorDB performance and cost analysis.

In SRE terms, New Relic is building:

New Relic AI direction:

New Relic AI
→ AI assistant for DevOps
→ system health reports
→ alert coverage analysis
→ instrumentation help

AI-powered observability
→ telemetry correlation
→ root-cause isolation
→ faster troubleshooting

AI Monitoring / LLM Observability
→ prompt/response analysis
→ LLM latency and error tracking
→ cost analysis
→ hallucination, bias and toxicity signals
→ VectorDB visibility

New Relic’s direction is about making its “all-in-one observability” platform more assistant-driven and making AI workloads first-class observable systems.


What they are all converging on

All four vendors are converging on the same broad architecture:

1. Collect telemetry
metrics, logs, traces, events, profiles, topology

2. Normalize context
services, owners, deployments, dependencies, SLOs, runbooks

3. Apply AI
anomaly detection, query generation, summarisation, RCA, prediction

4. Explain
what happened, why it happened, what changed, what is affected

5. Act
create ticket, page team, suggest fix, trigger workflow, remediate safely

6. Observe AI itself
LLM calls, prompts, responses, token cost, model quality, hallucinations,
safety, drift, vector DBs, RAG pipelines, agent workflows

The big product categories are:

AI capabilityWhat vendors are building
AI assistantChat interface over observability data
AI SRE agentInvestigates incidents and proposes actions
Automated RCAFinds likely root cause using telemetry and topology
Predictive anomaly detectionSpots problems before thresholds are breached
Natural-language queryingConverts plain English into PromQL, LogQL, SQL, trace/log queries
Incident summarisationExplains impact, timeline, evidence and next steps
Runbook automationRecommends or triggers operational workflows
AI workload monitoringMonitors LLMs, agents, prompts, responses, cost and quality
Governance/safetyTracks hallucination, toxicity, bias, privacy and compliance risks
Cost optimisationReduces telemetry waste and tracks LLM/token spend

The strategic reason they are doing this

The observability market is under pressure from three directions.

First, telemetry volumes are exploding. Kubernetes, microservices, edge, GPU clusters, AI workloads and distributed storage produce far more telemetry than humans can manually inspect.

Second, SRE teams are overloaded. Vendors are trying to sell “lower MTTR” and “less operational toil” by making the platform do more triage and correlation automatically.

Third, AI applications create new observability requirements. Traditional APM can tell you latency and error rate, but AI systems also need visibility into prompts, responses, hallucinations, drift, token usage, model quality, RAG retrieval quality, vector database behaviour and agent decisions.

So vendors are not just adding AI because it is fashionable. They are defending and expanding their core observability business.

What this means for an SRE / Observability Platform Engineer

The skill shift is significant.

Old value:

Build dashboards.
Write alert rules.
Know PromQL and LogQL.
Search logs manually.
Correlate incidents by experience.

New value:

Design telemetry that AI can reason over.
Standardise labels and service metadata.
Maintain accurate topology and ownership maps.
Connect observability to deployment and incident data.
Create safe remediation workflows.
Validate AI-generated RCA.
Control cost, access and blast radius.

The winners will not simply be the engineers who know the most dashboards. The winners will be the engineers who can build a trusted operational intelligence layer over metrics, logs, traces, topology and automation.

AI Strategies of New Observability Products

Coralogix is releasing the most explicit “AI observability product suite.” Cribl is positioning itself as the telemetry data layer for AI-era observability. Tsuga is newer and appears to be building an AI-native, bring-your-own-cloud observability architecture rather than simply adding an AI assistant to an old SaaS model.

Quick comparison

CompanyAI directionProduct maturity from public material
CoralogixAI Center, AI guardrails, AI evaluations, AI-SPM, Olly AI observability agentVery explicit productised AI offering
CriblCribl AI, Copilot, AI-guided Search Investigations, telemetry for humans and agentsStrong AI-assisted telemetry/data-management direction
TsugaBYOC observability for the AI era, agent-native observability, MCP/CLI for customer-owned agentsNewer; more architectural and agent-native positioning

1. Coralogix: AI observability as a full product suite

Coralogix is clearly releasing AI-focused products. Its main AI platform is AI Center, which Coralogix describes as a complete platform for AI-powered applications combining observability, guardrails, evaluations, and AI Security Posture Management in one place. It monitors LLM interactions for health, performance, cost, latency, errors, security and quality issues.

The key Coralogix AI products are:

Coralogix AI Center
├─ AI Observability
├─ AI Guardrails
├─ AI Evaluations
├─ AI Security Posture Management
├─ AI Application Discovery
└─ AI Explorer / Application Drilldown

What Coralogix is targeting

Coralogix is not just monitoring servers. It is monitoring AI application behaviour:

Prompt

LLM call

Response

Evaluation

Guardrail decision

Security / quality / cost signal

Its AI Center monitoring gives an organisation-level view of LLM usage and lets teams drill from a trend down to a specific application and even a specific prompt/response interaction.

It also supports OpenTelemetry GenAI semantic conventions, so teams can send GenAI spans into Coralogix AI Center without needing a Coralogix-specific SDK.

Olly: Coralogix’s AI observability agent

Coralogix also has Olly, which it describes as an AI-native observability agent. Olly lets users ask natural-language questions and get answers across logs, metrics, traces and alerts.

In practice, this is the “AI SRE assistant” layer:

Human asks:
“Why is payment latency rising?”

Olly checks:
├─ logs
├─ metrics
├─ traces
├─ alerts
├─ correlations
└─ possible root causes

Then returns:
├─ explanation
├─ evidence
├─ affected services
└─ recommended next steps

Coralogix also positions Olly as more than a simple assistant: it says Olly uses specialised agents for log analysis, trace exploration, metrics interpretation, security research, code debugging, correlation analysis and hypothesis generation.

My read on Coralogix

Coralogix is trying to own AI production reliability:

Monitor AI apps
Evaluate AI outputs
Detect prompt injection / PII / toxicity
Track token cost
Find bad model behaviour
Use AI to investigate normal production incidents

So yes: Coralogix is strongly AI-focused.

2. Cribl: AI platform for telemetry, not classic dashboard observability

Cribl’s AI angle is different. Cribl is not primarily trying to be another Datadog-style full-stack UI. It is positioning itself as the AI Platform for Telemetry: the collection, routing, shaping, searching and governance layer for machine data used by humans and AI agents. Cribl’s homepage describes the platform as giving enterprises choice and control for telemetry, and says it helps manage and analyse telemetry for both humans and agents.

The AI-focused Cribl areas are:

Cribl AI
├─ Copilot
├─ Copilot Editor
├─ AI-guided Search Investigations
├─ Natural-language queries
├─ AI-assisted pipeline creation
├─ AI telemetry parsing
└─ AI-ready telemetry routing

Cribl AI and Copilot

Cribl says its AI capabilities help teams create and modify pipelines, queries and configurations using natural language. It also says Cribl Copilot provides troubleshooting guidance, answers product/configuration questions and helps teams resolve issues faster.

This matters because a lot of observability toil is not just dashboards. It is:

Parse this log format.
Map this schema.
Route this data.
Drop this noisy field.
Mask this sensitive value.
Send this stream to the SIEM.
Send this other stream to cheaper storage.

Cribl’s AI is aimed at reducing that data-engineering toil.

Copilot Editor

Cribl’s Copilot Editor uses AI to help with schema mapping, translating logs across systems and building telemetry pipelines that clean, filter and route events.

That is important because AI-era observability needs clean, standardised telemetry. A reasoning agent is only useful if the data has usable structure.

Raw logs

AI-assisted parsing

Schema mapping

Enrichment / masking / routing

Search / SIEM / observability backend / AI agent

AI-guided Cribl Search Investigations

Cribl Search has an Investigations feature in preview. The docs describe it as a guided workspace where users explore incidents and telemetry using natural-language prompts. It helps analyse telemetry, identify patterns and document findings without manually building every query.

That means Cribl is moving into the AI-assisted investigation workflow:

Alert or question

Natural-language investigation

Generated queries

Pattern discovery

Findings captured in one workspace

Cribl’s AI observability thesis

Cribl’s recent AI observability messaging is that AI observability is a telemetry problem, not just a dashboard problem. It argues that LLM apps generate prompts, completions, tool calls, retrieval steps, token counts, model choices, policy events and infrastructure signals, and that those need to be collected and shaped for different teams and tools.

My read on Cribl

Cribl is not saying:

“We are the AI RCA dashboard.”

It is saying:

“We are the telemetry control plane that makes AI investigations possible.”

That is strategically clever. AI agents need cheap, governed, high-quality access to large telemetry volumes. Cribl wants to be the pipe, filter, schema and search layer underneath that.

3. Tsuga: AI-native observability architecture, still early

Tsuga is the newest and least mature publicly compared with Coralogix and Cribl, but it is very clearly positioning itself around the AI-era observability problem.

Tsuga describes itself as a bring-your-own-cloud observability platform for logs, metrics, traces and APM, deployed inside the customer’s AWS account using infrastructure-as-code. It says customers get the control of self-hosted infrastructure without the operational burden of running it.

Its newer positioning is explicitly AI-era focused. Tsuga announced a $35 million Series A on June 23, 2026, saying it is building “observability for the AI era” inside the customer’s cloud so the customer’s data and AI do not leave their control.

Tsuga’s AI claim

Tsuga’s argument is architectural:

Traditional observability:
telemetry leaves your cloud
vendor stores it
cost rises with volume
AI agents require broad access to vendor-hosted data

Tsuga model:
observability runs inside your cloud
telemetry stays inside your perimeter
AI runs on your own data
agents can use complete telemetry without exporting sensitive context

Tsuga says its AI tools run on the customer’s data inside the customer’s perimeter. It also says automated root-cause analysis runs on complete, unsampled data, and that its MCP server and CLI let engineering teams build their own agents on that foundation inside their own security boundary.

That MCP point is important. It suggests Tsuga is not only building an observability UI; it is exposing observability context to AI agents.

Agent-native observability

Tsuga has a specific Agent-Native Observability page. It says Tsuga is built so AI agents can use observability data effectively, affordably and inside the customer environment. It highlights agent-first APIs, MCPs, CLIs and query interfaces designed to return relevant context rather than raw data dumps.

That is a very modern product angle.

AI agent asks:
“What changed before this incident?”

Tsuga should return:
├─ relevant metrics
├─ relevant logs
├─ deployment context
├─ service ownership
├─ topology
└─ probable causal evidence

Not:
└─ 10GB of raw logs

What is less clear with Tsuga

Publicly, Tsuga looks less like:

Named AI assistant with lots of screenshots and feature modules

and more like:

AI-native observability architecture:
BYOC
complete telemetry
agent APIs
MCP
automated RCA
customer-owned AI boundary

So my assessment is: yes, Tsuga is AI-focused, but the public product story is currently more architectural and agent-native than feature-by-feature like Coralogix.

The strategic differences

Coralogix: “Observe and govern AI applications”

Coralogix is focused on production AI application reliability:

LLM monitoring
AI guardrails
Evaluations
AI security posture
Prompt/response visibility
Olly AI investigation agent

Best fit:

Teams deploying LLM apps and agents who need monitoring, safety, cost tracking and AI-assisted troubleshooting.

Cribl: “Prepare and control telemetry for AI”

Cribl is focused on the telemetry substrate:

Collect once
Shape data
Mask sensitive fields
Route anywhere
Search cheaply
Let humans and agents investigate
Use AI to build pipelines and queries

Best fit:

Large enterprises drowning in telemetry volume, SIEM costs, log routing complexity and multi-tool data sprawl.

Tsuga: “Run AI-era observability inside your own cloud”

Tsuga is focused on sovereign, cost-controlled, agent-native observability:

BYOC deployment
Telemetry stays in your cloud
AI and agents run inside your boundary
Automated RCA on unsampled data
MCP/CLI for custom SRE agents

Best fit:

Regulated, European, AI-native or high-scale companies that do not want telemetry, prompts, incident history and operational context exported to a third-party SaaS cloud.

The bigger market pattern

These newer players are attacking the incumbents from three angles:

1. Cost
AI generates more telemetry.
Per-GB SaaS observability becomes painful.

2. Data control
AI telemetry includes prompts, responses, business context and security-sensitive data.
Customers do not always want that in a vendor cloud.

3. Agent-readiness
Future observability is not just dashboards for humans.
AI agents need APIs, context retrieval, governed telemetry access and automated RCA.

So the new wave is less about “AI as a dashboard chatbot” and more about building the data foundation for AI-driven operations.

The sharpest summary is:

Coralogix = AI observability product suite
Cribl = AI-ready telemetry control plane
Tsuga = AI-native sovereign observability architecture

For an SRE/observability platform engineer, these companies are worth watching because they indicate where the next jobs and platform designs are going: telemetry engineering, AI-readable context, agent-safe access, automated RCA, guardrails and cost-controlled observability architectures.

Opensource AI Observability

AI adoption in open-source observability is happening, but it is different from what Datadog, Dynatrace, Splunk and New Relic are doing.

The commercial vendors are embedding AI directly into their SaaS platforms. The open-source ecosystem is mostly building the standards, collectors, SDKs, self-hostable platforms and agent interfaces that allow AI observability to work without vendor lock-in.

The big shift is this:

Old open-source observability:

Prometheus / Loki / Tempo / Grafana / OpenTelemetry

Collect, store, query, dashboard, alert


AI-era open-source observability:

OpenTelemetry + collectors + traces + logs + metrics + AI metadata

LLM / agent / RAG / GPU / vector DB visibility

AI assistants, AI SRE agents, natural-language querying, RCA

1. Grafana: open observability stack + AI features around it

Grafana Labs is moving in two directions.

First, it is keeping the open observability stack relevant for AI-era workloads: Grafana, Loki, Mimir, Tempo, Pyroscope and Alloy remain the core telemetry stack.

Second, it is adding AI-powered layers on top, especially in Grafana Cloud.

Grafana’s AI Observability product is built on OpenTelemetry and is aimed at teams running LLM agents in production. It monitors agent activity, traces conversations, tracks costs and evaluates quality. Grafana documents SDK support for Go, Python, TypeScript, Java and .NET, plus integrations with frameworks such as LangChain, LangGraph, OpenAI Agents and Vercel AI SDK.

Grafana also has Grafana Assistant, an AI-powered observability agent. It lets users ask questions like “Show me CPU usage” or “Create a dashboard for my database,” and it works across metrics, logs, traces, profiles and databases. Grafana says it can run investigations, manage dashboards, build/refine queries and help users navigate Grafana resources.

The important nuance: Grafana Assistant is not the same thing as open-source Grafana itself. It is primarily a Grafana Cloud AI capability, though Grafana documents a self-managed Assistant app that connects to a Grafana Cloud stack with reduced functionality.

Grafana’s most open-source-relevant AI move is probably Grafana Alloy. Alloy is Grafana Labs’ open-source OpenTelemetry Collector distribution with built-in Prometheus pipelines and support for metrics, logs, traces and profiles. It gives Grafana a standard collector layer for AI-era telemetry pipelines.

So Grafana’s strategy is:

Grafana AI strategy:

Open-source base:
Grafana
Loki
Mimir
Tempo
Pyroscope
Alloy

AI observability:
LLM / agent traces
cost tracking
quality evaluation
AI workload dashboards

AI assistant:
natural-language querying
dashboard creation
investigation assistance
query generation

Strategic direction:
keep the OSS stack open,
but place high-value AI workflows in Grafana Cloud.

2. OpenTelemetry: the standard layer for AI observability

OpenTelemetry is not a company; it is a CNCF open-source project. Its role is different from Grafana’s.

OpenTelemetry is becoming the standard telemetry schema and instrumentation layer for AI systems.

OpenTelemetry describes itself as an open-source observability framework for cloud-native software, providing APIs, libraries, agents and collector services for capturing telemetry. It also emphasises vendor-neutral instrumentation, meaning you instrument once and export to different backends.

For AI, the key development is OpenTelemetry semantic conventions for generative AI. OpenTelemetry has been extending its conventions so GenAI telemetry can capture model parameters, response metadata, token usage, traces, metrics and events for model interactions.

That matters because LLM systems need new telemetry fields that normal web apps did not need:

Traditional app telemetry:
service.name
http.status_code
duration
error
route
database call

AI app telemetry:
model name
prompt
completion
token count
tool call
retrieval step
vector DB query
embedding model
cost
temperature
hallucination score
safety evaluation

OpenTelemetry is not trying to become an AI assistant. Its value is that it gives the ecosystem a common language for AI telemetry.

So OpenTelemetry’s strategy is:

OpenTelemetry AI strategy:

Standardise:
spans
metrics
logs/events
attributes
semantic conventions

Support:
LLM calls
model interactions
prompts/responses
token usage
latency
errors
provider metadata

Enable:
Grafana
SigNoz
Langfuse
OpenLIT
Elastic
New Relic
Datadog
custom platforms

Strategic direction:
become the neutral telemetry contract for AI applications.

3. OpenLIT: open-source LLM observability on OpenTelemetry

OpenLIT is a good example of the new generation of open-source AI observability projects.

It describes itself as an open-source LLM observability and AI engineering platform built on OpenTelemetry. Its positioning is self-hosted, privacy-first and vendor-neutral.

This is important because many companies do not want prompts, responses, user inputs, sensitive data or AI-agent traces going straight into a third-party SaaS.

OpenLIT’s direction is:

OpenLIT strategy:

Monitor:
LLM calls
latency
token usage
cost
model behaviour
vector DBs
GPU usage

Deploy:
self-hosted
OpenTelemetry-native
privacy-first

Best fit:
teams building AI apps who want open-source AI observability
without committing to a commercial platform first.

4. Langfuse: open-source LLM tracing and evaluation

Langfuse is another major open-source AI observability project.

It focuses on LLM application tracing: capturing prompts, model responses, token usage, latency, tool calls and retrieval steps. Langfuse also provides AI-engineering features such as LLM-as-judge evaluation, prompt management, experiments and datasets, and it can be self-hosted.

Langfuse is less like “Grafana for all infrastructure” and more like “observability and evaluation for LLM applications.”

Its strategy is:

Langfuse strategy:

Trace:
prompt
response
tool call
RAG step
latency
token usage
cost

Evaluate:
quality
scoring
experiments
prompt versions
datasets

Best fit:
AI product teams who need to debug and improve LLM apps,
not just monitor infrastructure.

5. SigNoz: open-source observability with AI-agent access

SigNoz is moving from being an open-source Datadog/New Relic alternative into a more AI-aware observability platform.

SigNoz describes itself as an open-source observability tool powered by OpenTelemetry, covering logs, metrics, traces, dashboards, alerts and LLM/AI observability. It also advertises an MCP server for bringing telemetry into coding agents and an AI teammate called Noz for incident investigation, alert tuning and dashboard building.

This is significant because it shows a broader open-source pattern: observability platforms are not just adding AI dashboards; they are exposing telemetry to AI agents.

SigNoz direction:

OpenTelemetry-native observability
+
LLM/AI observability
+
MCP access for coding agents
+
AI teammate for investigations and dashboards

That is where open-source observability is going: not just dashboards for humans, but context APIs for agents.

6. HolmesGPT: open-source AI SRE agent

HolmesGPT is another important example because it is not primarily about observing LLM apps. It is about using AI to investigate production incidents.

HolmesGPT describes itself as an open-source AI agent for investigating production incidents and finding root causes across Kubernetes, VMs, cloud providers, databases and SaaS platforms. It is listed as a CNCF sandbox project.

That puts it closer to the Datadog Bits AI / Dynatrace Davis AI direction, but in open-source form.

HolmesGPT strategy:

Input:
alerts
Kubernetes state
metrics
logs
cloud context
runbooks

AI task:
investigate incident
gather evidence
find probable root cause
explain next action

Best fit:
platform teams wanting an open-source AI SRE layer
over existing observability tools.

The overall open-source adoption pattern

Open-source observability is adopting AI in four layers.

1. AI telemetry standards

This is where OpenTelemetry is most important.

Goal:
make AI applications observable in a standard way

Examples:
GenAI semantic conventions
token usage attributes
model request spans
prompt/response events
tool-call spans

This is foundational. Without standard AI telemetry, every vendor and OSS project invents incompatible schemas.

2. AI workload observability

This is where Grafana AI Observability, OpenLIT, Langfuse and SigNoz fit.

Goal:
monitor LLM apps, agents and RAG pipelines

Signals:
latency
token cost
prompt/response quality
hallucination risk
model errors
vector DB retrieval
tool calls
agent steps

3. AI-assisted operations

This is where Grafana Assistant, HolmesGPT, SigNoz Noz and similar tools fit.

Goal:
help humans investigate production systems faster

Capabilities:
natural-language querying
alert explanation
dashboard generation
root-cause hints
log summarisation
incident summaries

4. Agent-native observability

This is the newest layer.

Goal:
let AI agents consume observability data safely

Interfaces:
MCP servers
CLI tools
API access
context retrieval
guarded query execution
evidence-based RCA

This matters because future AI coding agents and SRE agents will need access to production telemetry to debug issues. The observability stack must become queryable by both humans and machines.

The key difference from commercial observability

Commercial vendors are building polished AI experiences inside their own SaaS platforms.

Open-source observability is building the portable foundations:

LayerOpen-source approach
InstrumentationOpenTelemetry SDKs and semantic conventions
CollectionOpenTelemetry Collector, Grafana Alloy
Storage/queryGrafana LGTM, SigNoz, ClickHouse-based stacks
AI app tracingOpenLIT, Langfuse, OTel GenAI conventions
AI SREHolmesGPT, MCP-enabled tools
Agent accessMCP, APIs, CLI workflows

The strategic difference is:

Commercial vendors:
"Use our platform and our AI will help you."

Open-source ecosystem:
"Instrument once, own your data, expose telemetry to any backend or AI agent."

What this means for SREs and observability engineers

The valuable skill is moving from only operating dashboards to building an AI-readable telemetry platform.

That means:

You need:
consistent OpenTelemetry attributes
clean service names
good resource metadata
deployment markers
trace/log/metric correlation
AI workload spans
token/cost metrics
evaluation signals
MCP or API access for agents
guardrails around sensitive telemetry

For a homelab or professional platform, the modern open-source direction would be:

Applications / AI agents

OpenTelemetry SDKs + GenAI semantic conventions

Grafana Alloy or OpenTelemetry Collector

Mimir / Loki / Tempo / ClickHouse / SigNoz / Langfuse / OpenLIT

Grafana dashboards + AI assistant / HolmesGPT / MCP-enabled agents

The sharp summary:

Grafana is making the open observability stack AI-aware.

OpenTelemetry is becoming the standard language for AI telemetry.

OpenLIT, Langfuse and SigNoz are making LLM apps observable.

HolmesGPT-style tools are turning open telemetry into AI-assisted SRE investigations.

So, yes: open-source observability is adopting AI quickly, but the centre of gravity is different. The open-source world is less about one vendor-owned AI brain and more about open telemetry, self-hostable AI observability, and agent-ready operations.

AI DC Buildouts, Changing Jobs & Roles of the 4th Industrial Revolution

The AI infrastructure race is being led by a relatively small number of corporations, but together they represent well over US$1 trillion of planned investment over the remainder of this decade. Many figures below are approximate because companies often announce campuses or regions rather than exact building counts, and projects evolve rapidly.

CorporationOperational data centres (approx.)AI data centres planned / under constructionMain locations
Amazon Web Services100+ availability zones across 36+ regionsDozens of new AI campuses through 2028 (including Project Rainier)USA (Virginia, Pennsylvania, Georgia, Mississippi, Oregon), Europe, UK, Germany, India, Japan, Australia
Microsoft300+ data centres globallyTens of new AI campuses; ~$80B AI infrastructure investmentUSA, Sweden, Finland, UK, Germany, Australia, Japan, Texas, Wisconsin
Google40+ cloud regions and many hyperscale campusesMultiple new AI mega-campusesOhio, Nebraska, Oklahoma, Texas, Iowa, Europe, Asia
Meta20+ hyperscale campusesNumerous AI campuses under expansionLouisiana, Ohio, Iowa, Texas, Alabama, with additional capacity from Crusoe
Oracle80+ cloud regionsMulti-gigawatt AI campuses via Stargate plus Oracle Cloud expansionTexas, New Mexico, Ohio, Michigan and other US states
OpenAIOperates via partners rather than owning a global DC fleetStargate aims for roughly 20 major AI campusesTexas, New Mexico, Ohio, Wisconsin, Michigan and additional US sites
SoftBankNo major hyperscale cloud estateCo-investor in StargateUnited States (multiple campuses)
CoreWeave~30+ AI data centresContinuing rapid expansionUSA, UK, Norway, Spain and additional European sites
xAI1 flagship AI supercluster (Colossus) plus expansionsExpanding toward one million GPUsMemphis, Tennessee and additional US locations
CrusoeSeveral AI campuses under operationMultiple campuses for OpenAI, Meta and MicrosoftTexas, Oklahoma and other US states
NscaleEarly-stage AI infrastructureUK and European sovereign AI facilities plannedUnited Kingdom, Norway and Europe (build-out still in early stages)

Where the biggest build-out is happening

The current hotspots are:

  • Texas – by far the largest concentration, with Stargate, Oracle, Microsoft, Google and xAI all investing heavily.
  • Ohio – Google, Meta and Oracle are all expanding there.
  • Louisiana – Meta’s enormous AI campus.
  • Virginia – still the world’s largest concentration of conventional cloud data centres.
  • Pennsylvania, Georgia and Oklahoma – major AWS and Google investments.
  • Wisconsin, Michigan and New Mexico – emerging AI infrastructure hubs.

The scale is unprecedented

The six largest AI infrastructure builders (Amazon, Microsoft, Google, Meta, Oracle and the Stargate consortium) have collectively committed around US$690–700 billion in AI-related capital expenditure, with 74 new AI-focused projects breaking ground in the US during 2026 alone. Longer-term projections suggest total AI infrastructure investment could exceed US$5 trillion globally by 2030.

One notable trend is that these companies are no longer building isolated data centres. They are constructing AI campuses consisting of anywhere from 8 to more than 20 individual data-centre buildings, all linked by ultra-high-speed networking so they function as a single giant AI supercomputer. A single campus can consume 500 MW to over 1 GW of power, equivalent to the electricity demand of a medium-sized city.

The largest AI campuses consume enormous quantities of resources. Some impacts are already measurable, while others remain uncertain and depend on how utilities allocate costs. It’s important to distinguish local effects (which can be substantial) from national effects (which are often much smaller).

ResourceHow AI campuses use itImpact on consumers
ElectricityHundreds of MW to several GW continuouslyHigher utility investment, possible higher electricity bills in constrained regions, increased need for new power stations
WaterCooling systems can consume millions of gallons per day, although newer designs increasingly use closed-loop or air coolingCompetition for water in drought-prone areas; pressure on municipal supplies
LandCampuses often occupy hundreds to thousands of acresIndustrial land values rise; reduced land available for other development
Construction materialsSteel, concrete, copper, fibre-optic cableHigher demand can contribute to material price increases, though AI is only one of several drivers
Electrical equipmentTransformers, switchgear, substationsLonger lead times for utilities and industrial customers
GPUs and serversHundreds of thousands of accelerators per campusSemiconductor manufacturing capacity diverted toward AI, increasing demand for advanced chips
Skilled labourElectrical engineers, construction workers, data-centre techniciansWage competition and labour shortages in some regions
Natural gasSome campuses are building dedicated gas-fired generationIncreased demand for gas infrastructure and fuel in certain markets

Electricity prices

Electricity is the area where households are most likely to notice an effect.

Large AI campuses require utilities to invest in:

  • New transmission lines
  • New substations
  • Additional generation
  • Grid upgrades

Who pays depends on regulation.

In some regions, regulators are trying to ensure that AI companies pay most of these costs. In others, some infrastructure costs are spread across all customers, which can increase household bills.

For example:

RegionReported effect
PJM (eastern U.S.)Wholesale electricity prices rose sharply as demand from AI data centres increased, prompting calls for tech companies to fund more of the required infrastructure.
ArizonaUtilities warn that electricity infrastructure may need to roughly double within a few years because of AI growth.
VirginiaData centres already account for a very large share of electricity demand in some parts of the state.

It’s also worth noting that recent academic work found that, historically (2015–2024), data centres slightly reduced average U.S. electricity prices by helping spread fixed grid costs over more customers. The authors caution that this may not hold if future supply constraints become severe.

Water

Water is highly location-dependent.

Older evaporative cooling systems can use several million gallons of water per day. Newer AI facilities increasingly employ:

  • Closed-loop liquid cooling
  • Direct-to-chip liquid cooling
  • Air cooling where practical

These approaches can significantly reduce freshwater consumption, but water remains a concern in arid regions.

Housing

AI campuses can affect local housing markets by:

  • Bringing thousands of construction workers
  • Creating highly paid engineering jobs
  • Increasing demand for rental accommodation

The effect is usually local rather than national.

Employment

Benefits include:

  • Construction employment
  • Electrical contracting
  • Operations and maintenance jobs
  • Security
  • Network engineering
  • Mechanical engineering

However, once operational, AI campuses employ far fewer people than factories of similar size.

Have prices increased?

Evidence is mixed:

ItemObserved trend
ElectricitySome U.S. regions have seen higher wholesale prices and concerns about retail bills where AI demand is concentrated.
WaterMostly local impacts in water-stressed regions rather than broad consumer price rises.
HousingLocal increases around major developments are common, though driven by multiple factors.
Construction materialsIncreased demand contributes to pressure, but AI is only one of many drivers.
Consumer goodsThere is currently little evidence that AI data centres have directly increased the prices of everyday retail goods.

Overall, the greatest measurable impact today is on electricity infrastructure. The International Energy Agency projects that global data-centre electricity consumption will more than double to about 945 TWh by 2030, driven largely by AI. Whether households ultimately pay more depends on regulatory decisions about who funds the new power plants, transmission lines and substations needed to support these AI campuses.

Changing Jobs and Roles

The AI infrastructure boom is creating the largest shift in infrastructure engineering since the rise of public cloud around 2006–2015. Traditional cloud providers needed engineers to build reliable, scalable services for virtual machines, storage and networking. AI Factories require all of that plus expertise in GPUs, ultra-high-speed networking, power engineering, liquid cooling and AI software platforms.

Evolution of Infrastructure Engineering

EraPrimary GoalMain InfrastructureTypical Employer
Enterprise IT (1990–2010)Business applicationsServers, SAN, LANBanks, government, enterprises
Cloud (2006–2024)Multi-tenant cloud servicesHyperscale datacentersAWS, Azure, Google Cloud
AI Factory (2024–2035+)Massive AI computationGPU supercomputers, AI campusesOpenAI, Meta, xAI, Oracle, CoreWeave, Nscale, AWS

Traditional Cloud Provider Jobs

Cloud providers traditionally organised engineering into around a dozen major disciplines.

DisciplineTypical Roles
Datacenter FacilitiesFacilities Engineer, Mechanical Engineer, Electrical Engineer
ComputeServer Engineer, Linux Engineer, Virtualisation Engineer
StorageStorage Engineer, Ceph Engineer, SAN Engineer
NetworkingNetwork Engineer, Network Architect
Cloud PlatformKubernetes Engineer, OpenStack Engineer, VMware Engineer
ReliabilitySite Reliability Engineer (SRE), DevOps Engineer
SecuritySecurity Engineer, IAM Engineer
ObservabilityMonitoring Engineer, Logging Engineer
AutomationAnsible Engineer, Terraform Engineer
SoftwareBackend Engineer, Platform Engineer
OperationsNOC Engineer, Incident Manager
CapacityCapacity Planner, Performance Engineer

A large hyperscale datacenter typically employs 100–300 permanent staff, with many more contractors during construction.


AI Factory Engineering

AI Factories introduce entirely new engineering domains.

New DisciplineExample Roles
GPU InfrastructureGPU Systems Engineer, GPU Cluster Engineer
AI NetworkingInfiniBand Engineer, RoCE Engineer, Ethernet Fabric Engineer
AI StorageHigh-performance Storage Engineer, Parallel Filesystem Engineer
AI CoolingLiquid Cooling Engineer, Thermal Systems Engineer
AI SchedulingSlurm Engineer, Kubernetes AI Platform Engineer
AI RuntimeCUDA Engineer, Distributed Training Engineer
AI OptimisationML Infrastructure Engineer
AI Datacenter PowerHigh-voltage Power Engineer
AI Chip EngineeringAccelerator Integration Engineer
AI OperationsAI Infrastructure SRE

Engineering Stack

Traditional cloud:

Applications
Containers
Virtual Machines
Hypervisor
Servers
Storage
Networking
Power

AI Factory:

AI Models
Distributed Training
Kubernetes / Slurm
CUDA / ROCm
100,000+ GPUs
InfiniBand / RoCE
Parallel Storage
Liquid Cooling
Gigawatt Power

Traditional Cloud Skills

  • Linux
  • VMware
  • Kubernetes
  • OpenStack
  • AWS
  • Azure
  • Terraform
  • Ansible
  • Prometheus
  • Grafana
  • Python
  • Go
  • Storage
  • Networking

New AI Factory Skills

Additional skills now becoming highly valuable include:

  • NVIDIA GPU architecture
  • AMD Instinct
  • CUDA
  • NCCL
  • GPUDirect RDMA
  • InfiniBand
  • RoCE v2
  • Slurm
  • Ray
  • Kubeflow
  • MLFlow
  • Triton Inference Server
  • Parallel file systems (Lustre, IBM Storage Scale/GPFS, BeeGFS)
  • High-performance Ethernet (400/800 GbE)
  • Direct-to-chip liquid cooling
  • Rack-scale power engineering

Jobs Growing Fastest

RoleGrowth Outlook
GPU Infrastructure EngineerExtremely High
AI Platform EngineerExtremely High
HPC Systems EngineerExtremely High
Kubernetes Platform EngineerVery High
Storage EngineerVery High
Site Reliability EngineerVery High
Network Fabric EngineerExtremely High
Power Systems EngineerExtremely High
Mechanical Cooling EngineerExtremely High
AI Operations EngineerExtremely High

Approximate Current Workforce (2025–2026)

The exact numbers are difficult to measure because many roles overlap, but industry estimates suggest:

ProfessionEstimated Global Workforce
Cloud Engineers2–3 million
DevOps Engineers1.5–2 million
Site Reliability Engineers400,000–700,000
Kubernetes Engineers500,000–900,000
Datacenter Engineers300,000–500,000
Storage Engineers200,000–350,000
HPC Engineers80,000–150,000
GPU Infrastructure Specialists20,000–40,000
AI Infrastructure Engineers50,000–100,000

Estimated Workforce Needed by 2030

As AI campuses proliferate worldwide, demand is expected to increase significantly.

ProfessionEstimated Demand by 2030
AI Infrastructure Engineers300,000–500,000
GPU Cluster Engineers150,000–250,000
HPC Engineers250,000–400,000
SREs (AI/Cloud)800,000–1.2 million
Kubernetes Platform Engineers1–1.5 million
Network Fabric Engineers300,000–500,000
Storage Engineers500,000+
Power Engineers400,000–700,000
Cooling Engineers250,000–500,000

These are indicative estimates derived from announced AI infrastructure expansion plans and broader industry workforce analyses rather than official forecasts.


Where the Talent Is Coming From

Most AI Factory engineers are not newly trained graduates. Companies are recruiting experienced professionals from adjacent disciplines:

Previous RoleTransition To
Cloud EngineerAI Platform Engineer
Kubernetes EngineerAI Infrastructure Engineer
SREAI Operations Engineer
HPC EngineerGPU Cluster Engineer
Linux EngineerGPU Systems Engineer
Network EngineerInfiniBand/RoCE Fabric Engineer
Storage EngineerAI Storage Architect
OpenStack EngineerAI Cloud Platform Engineer
Ceph EngineerHigh-performance Storage Engineer
DevOps EngineerML Platform Engineer

Why This Matters

The next decade is likely to see a shift similar to the transition from enterprise IT to cloud computing. During the 2010s, the most sought-after roles were Cloud Engineers, DevOps Engineers and SREs. Through the late 2020s and into the 2030s, many of the highest-demand infrastructure roles are expected to centre on AI Factories: designing, building and operating gigawatt-scale GPU campuses, high-performance storage systems, ultra-low-latency networks and AI platforms.

For someone with expertise in Linux, Kubernetes, observability, automation, storage and cloud infrastructure, the progression into AI infrastructure engineering is relatively direct. Adding knowledge of GPU platforms, HPC networking (InfiniBand/RoCE), parallel storage (such as Lustre or GPFS), Slurm, CUDA and liquid-cooled datacenter design positions engineers for many of the roles expected to see the strongest demand over the coming decade.

Part of the 4th Industrial Revolution

Yes — this is plausibly the tail-end phase of the Forth Industrial Revolution, but with one caveat: we do not yet know whether AGI/ASI will arrive, or when. What is clear is that capital, land, power, water, chips, networks and engineering labour are being redirected toward AI factories.

The simplest framing:

Industrial phaseCore machineMain resourceMain labour shift
1stSteam engineCoalFarm → factory
2ndElectrified production lineOil, steel, electricityCraft → mass production
3rdComputerSilicon, softwareClerical → digital
4thCloud + automationData, networks, platformsIT → cloud/SRE/DevOps
5thAI factoryCompute, power, GPUs, dataHuman labour → AI-augmented/AI-directed labour

The AI factory is the new “mill.” Instead of spinning cotton or stamping cars, it converts electricity + chips + data + models into intelligence services: code, design, analysis, customer support, robotics control, synthetic media, drug discovery and eventually autonomous decision systems.

The resource pull is already visible. The IEA projects global data-centre electricity consumption could roughly double to about 945 TWh by 2030, growing far faster than general electricity demand. That is why hyperscalers, AI labs and neoclouds are racing to secure power, grid connections, GPUs, cooling, land and engineering staff.

On jobs, the likely pattern is not “all jobs disappear.” It is task compression: fewer people needed for routine cognitive work, more people needed for infrastructure, supervision, security, robotics, energy, regulation and high-complexity design. Goldman Sachs has estimated that AI could expose the equivalent of 300 million full-time jobs globally to automation, while the World Economic Forum projects by 2030 about 170 million roles created and 92 million displaced, for a net gain of 78 million under its surveyed-employer scenario.

Likely traditional jobs under pressure:

AreaJobs most exposed
Admin/officeData entry, scheduling, basic document processing
Customer serviceTier-1 support, call-centre scripts, helpdesk triage
SoftwareBoilerplate coding, simple QA, basic web/app work
Finance/legalDocument review, reconciliation, compliance paperwork
Media/marketingGeneric copywriting, SEO text, simple design production
EducationBasic tutoring, marking, lesson-content generation
Transport/logisticsDispatch, route planning, warehouse coordination
RetailCheckout, product support, inventory admin

New and expanded jobs:

Future areaRoles likely to grow
AI infrastructureGPU cluster engineer, AI SRE, AI platform engineer
Power/gridSubstation engineer, energy systems engineer, microgrid operator
Cooling/facilitiesLiquid-cooling engineer, thermal engineer, datacenter mechanic
NetworkingInfiniBand/RoCE engineer, optical network engineer
Storage/dataParallel storage engineer, data governance engineer
AI safety/securityModel auditor, AI red-team engineer, AI incident responder
RoboticsRobot fleet supervisor, autonomy technician, human-robot workflow designer
RegulationAI compliance officer, algorithmic accountability auditor
Human-AI workAgent orchestrator, prompt/workflow architect, AI operations manager
Synthetic worldsSimulation designer, digital twin engineer, synthetic-data engineer

If AGI arrives, the shift accelerates. If ASI arrives, the shift becomes civilisational: the scarce resources may become energy, compute rights, physical materials, robotics capacity, trusted governance and human legitimacy, rather than ordinary labour.

So yes: the AI build-out looks like the physical foundation of a Fifth Industrial Revolution — not just software, but a new industrial base built around manufactured intelligence.

Climate change and broader sociological factors are arguably the largest long-term uncertainties for the Fifth Industrial Revolution. Unlike technical bottlenecks, they can alter not just the pace of AI adoption but also where, how, and for whom AI infrastructure is built.

I don’t think climate change will stop the AI revolution, but it could fundamentally reshape it. History suggests industrial revolutions adapt to resource constraints rather than ending because of them.

Climate change

1. Energy transition

Today’s AI factories consume enormous amounts of electricity.

If climate policies tighten globally, AI companies may no longer be able to rely on inexpensive fossil-fuel generation.

This is already pushing investment towards:

  • Nuclear power
  • Small Modular Reactors (SMRs)
  • Geothermal
  • Offshore wind
  • Utility-scale solar
  • Long-duration batteries
  • Grid-scale storage

By the 2040s, a successful AI company may be judged as much by its carbon intensity per AI token as by its model quality.


2. Water shortages

Many AI campuses currently use water-intensive cooling.

Increasing droughts could force AI factories to relocate.

Future AI campuses are likely to favour:

  • Scotland
  • Norway
  • Sweden
  • Finland
  • Iceland
  • Canada
  • Pacific Northwest
  • Patagonia

Cool climates reduce cooling costs while providing more reliable water supplies.


3. Sea-level rise

Many current datacentres sit near coasts because they benefit from:

  • Fibre landing stations
  • Major cities
  • Existing infrastructure

Over decades, flood risks may encourage more inland development.


4. Extreme weather

Increasingly frequent:

  • Heatwaves
  • Wildfires
  • Hurricanes
  • Flooding

all increase operational risks.

Future campuses may need:

  • Greater redundancy
  • Fire-resistant designs
  • Multiple grid connections
  • Larger battery systems
  • Independent power generation

Resource nationalism

Countries increasingly recognise compute as a strategic asset.

Competition may intensify over:

  • Lithium
  • Copper
  • Rare earth elements
  • Uranium
  • Semiconductor-grade silicon
  • Freshwater
  • Electricity

The next century may see competition over compute capacity much as the twentieth century saw competition over oil.


Demographics

Many developed nations face ageing populations.

This may actually accelerate AI adoption.

Examples include:

  • Japan
  • South Korea
  • Germany
  • Italy

If fewer working-age people are available, automation becomes economically attractive.


Education

Universities are already adapting.

Future curricula may emphasise:

  • AI engineering
  • Robotics
  • HPC
  • Power engineering
  • Semiconductor engineering
  • AI governance

Routine programming skills alone may become less valuable than systems integration, critical thinking and domain expertise.


Public trust

AI adoption depends heavily on social acceptance.

Concerns include:

  • Surveillance
  • Privacy
  • Bias
  • Deepfakes
  • Autonomous weapons
  • Job displacement

Public backlash could lead to stricter regulation or slower deployment in some sectors.


Wealth inequality

One of the most significant risks is that AI could concentrate wealth among those who own:

  • AI models
  • Compute infrastructure
  • Semiconductor intellectual property
  • Energy assets
  • Data

If productivity gains are not widely shared, inequality could increase.

Possible policy responses include:

  • Expanded education and retraining
  • Wage insurance
  • Stronger competition policy
  • Tax reforms
  • New social safety nets

Different countries are likely to pursue different approaches.


Employment transition

Industrial revolutions historically eliminate some jobs while creating others.

The challenge is timing.

If AI removes work faster than new roles appear, societies may experience:

  • Higher unemployment
  • Political instability
  • Reduced consumer spending
  • Pressure for labour-market reforms

Managing this transition is likely to be one of the defining policy challenges of the coming decades.


Geopolitics

Compute is becoming a strategic resource.

This may encourage blocs centred around:

  • North America
  • Europe
  • China
  • India
  • Middle East

Each could develop increasingly independent AI ecosystems, supply chains and regulations.


Alternative futures

ScenarioAI build-outSociety
Green AI RevolutionAI powered largely by low-carbon energy; highly efficient hardwareAI helps accelerate decarbonisation and scientific progress
AI Arms RaceNational security drives rapid expansion despite environmental costsFragmented AI ecosystems and geopolitical competition
AI BubbleInfrastructure investment slows after poor returnsAI remains important but grows more gradually
Climate Adaptation AIAI prioritises climate modelling, energy optimisation and resilient infrastructureAI becomes a key tool for adapting to climate change
Post-Scarcity Transition (speculative)Abundant clean energy and highly capable AI dramatically reduce production costsWork shifts towards creativity, care, governance and exploration

The “AI Factory Economy”

A useful way to think about the long term is that AI factories may become a new class of critical infrastructure, similar to:

  • Power stations
  • Railways
  • Ports
  • Telecommunications
  • The Internet

The economy could evolve around interconnected systems:

Clean Energy


AI Factories


Robotics + Software + Scientific Discovery


Higher Productivity


Lower Cost of Goods and Services


More Resources Available for Society

That is an optimistic pathway. A less favourable outcome is also possible if productivity gains are unevenly distributed, infrastructure cannot keep pace, or environmental constraints become more severe.

The most important sociological question

The defining issue may not be whether AI becomes powerful enough—it almost certainly will continue to improve significantly. The larger question is who benefits from the productivity gains.

Previous industrial revolutions eventually raised average living standards, but they also brought decades of disruption, labour conflict and institutional change. The Fifth Industrial Revolution, if it unfolds as many expect, is likely to follow a similar pattern: technological progress may be rapid, but the economic and social institutions needed to distribute its benefits will evolve more slowly.

In other words, the success of the Fifth Industrial Revolution may depend less on building bigger AI factories and more on how societies adapt their education systems, labour markets, energy infrastructure and governance to make effective use of the capabilities those AI factories create.

What is a Neocloud? CoreWeave, Crusoe, Nscale and Oracle vs Radiant

“Neocloud” (sometimes written neo cloud) is a term for a new generation of cloud providers that specialize in AI computing rather than offering the full range of traditional cloud services. They focus heavily on providing high-performance GPUs for AI training and inference.

How neoclouds differ from traditional cloud providers

Traditional cloud (AWS, Azure, Google Cloud)Neocloud
Broad range of services (databases, storage, networking, analytics, etc.)Primarily focused on AI and GPU computing
Designed for many types of workloadsOptimized specifically for AI/ML workloads
Large hyperscale platformsOften smaller, AI-focused companies
GPU capacity can be limited or expensiveAim to provide faster access to GPUs and lower costs

Why neoclouds became popular

The explosion of generative AI created huge demand for GPUs such as NVIDIA H100 and Blackwell chips. Many organizations struggled to obtain enough AI compute from traditional cloud providers, creating an opportunity for specialized GPU cloud companies.

Examples of neocloud providers

Some well-known neocloud companies include:

  • CoreWeave
  • Lambda
  • Crusoe
  • Nebius
  • Together AI

These companies provide GPU-as-a-Service (GPUaaS) and AI-focused infrastructure.

Simple analogy

Think of traditional cloud providers as a large supermarket that sells everything, while a neocloud is a specialty store focused almost entirely on AI computing power. It may offer fewer services overall, but it is optimized for AI workloads and often provides better access to GPUs.

Nscale a European Neocloud?

Today, a more representative list of major neoclouds would include:

CompanyRegionNotes
NscaleUK / EuropeFull-stack AI infrastructure, sovereign AI cloud, GPU cloud, data centre developer.
CoreWeaveUSOften regarded as the archetypal neocloud.
NebiusEuropeAI cloud and GPU infrastructure provider.
LambdaUSGPU cloud focused on AI training and inference.
CrusoeUSAI data centres and GPU cloud infrastructure.
Together AIUSAI platform plus infrastructure.

Nscale’s positioning is actually slightly different from some of the others because it is trying to be vertically integrated:

  • Building or owning AI data centres.
  • Procuring GPU fleets at massive scale.
  • Operating AI cloud services.
  • Offering sovereign AI infrastructure for governments and enterprises.
  • Running full-stack AI platforms rather than just renting GPUs.

Some analysts now classify Nscale as an AI hyperscaler rather than merely a neocloud because of the scale it is targeting. ABI Research ranked Nscale as the overall leader among 14 neocloud providers in its 2026 assessment.

What’s interesting is that the neocloud landscape appears to be splitting into three tiers:

  1. GPU rental companies – essentially GPU-as-a-Service.
  2. AI cloud platforms – GPUs plus AI tooling.
  3. AI hyperscalers – own data centres, networking, power, GPUs, and cloud platform.

Nscale is deliberately pursuing category 3. The company describes itself as a vertically integrated AI cloud and has announced very large-scale deployments in Europe and the US.

If you compare Nscale, CoreWeave, and Crusoe specifically, I’d place them like this:

AreaNscaleCoreWeaveCrusoe
Sovereign European AIStrongestLimitedLimited
GPU CloudStrongVery StrongStrong
Data Centre OwnershipExtensive strategyGrowingExtensive
AI Hyperscaler AmbitionVery HighHighHigh
European PresenceStrongestModerateModerate
Microsoft PartnershipsSignificantSignificantSignificant

From a European perspective, Nscale is probably the closest thing Europe currently has to a home-grown AI hyperscaler.

No. If we’re talking about Europe specifically, I would actually argue the opposite:

CoreWeave is currently ahead in deployed AI infrastructure, while Nscale is ahead in announced future European capacity.

Those are very different things.

CoreWeave’s position in Europe

CoreWeave already has:

  • European headquarters in London.
  • Two operational UK data centres.
  • Expansion into Norway, Sweden, and Spain.
  • Billions already committed and deployed into European infrastructure.
  • A mature GPU cloud platform that is already serving customers globally.

By 2025, CoreWeave had announced European expansion into Norway, Sweden, and Spain alongside its existing UK footprint.

More importantly, CoreWeave entered Europe after already becoming a large-scale AI cloud provider in the US. They brought:

  • Operational expertise
  • Existing customers
  • Existing software platform
  • Existing GPU fleet

That is a major advantage.

Where Nscale is stronger

Nscale’s strength is the future build pipeline.

Publicly announced projects include:

  • Stargate Norway
  • Sines (Portugal)
  • UK AI campus developments
  • Iceland expansion plans

Some of these projects are absolutely enormous on paper. The Norway Stargate project alone targets 100,000 NVIDIA GPUs.

Portugal is also positioned as one of Nscale’s flagship European hubs, with 12,600+ Blackwell GPUs initially and much larger Rubin deployments planned later.

The key distinction

If you compare today’s operational reality:

MetricCoreWeaveNscale
Operational GPU cloudAheadBehind
Existing customer workloadsAheadBehind
Software/cloud platform maturityAheadBehind
European operational experienceAheadBehind
Publicly visible deployed GPU capacityAheadBehind

If you compare future announced European capacity:

MetricCoreWeaveNscale
Norway buildoutLargeVery large
PortugalLimited public presenceMajor flagship site
Sovereign AI initiativesSomeStrong focus
OpenAI-linked projectsLimitedSignificant
Future European MW pipelineLargePotentially larger

A useful analogy

Today, CoreWeave is closer to:

“We already run a large AI cloud and are expanding into Europe.”

Nscale is closer to:

“We are building some of Europe’s largest AI campuses and will become a major AI cloud.”

Those are different stages of maturity.

The question investors are asking

The debate isn’t really:

“Can Nscale catch CoreWeave?”

The debate is:

“Can Nscale turn announced power, land, and GPU commitments into revenue-producing clusters before demand or financing conditions change?”

CoreWeave has already demonstrated it can operate large GPU fleets and monetize them. Nscale is in the process of proving that at the same scale.

One interesting point: some recent reporting has questioned the extent to which both companies’ European investment announcements translate into immediately operational facilities, noting that some “new data centre” claims are actually deployments into existing colocation facilities rather than brand-new campuses. That criticism has been directed at both Nscale and CoreWeave.

So as of mid-2026:

  • Operationally: CoreWeave is ahead in Europe.
  • Announced future European capacity: Nscale may have the larger headline pipeline.
  • Execution risk: Nscale has more to prove because a larger proportion of its European footprint is still future-dated.

Is Nscale’s IPO still on target for late 2026?

As of June 2026, there is no publicly filed prospectus, no announced exchange, and no confirmed IPO date for Nscale.

The strongest public indication that an IPO is still being pursued comes from industry reports stating that Nscale was planning a fall/late-2026 IPO and was pursuing additional US data-centre acquisitions ahead of that listing.

However, there are several reasons to be cautious about assuming it is “on target”:

Reasons it could still happen in late 2026

  • The AI infrastructure sector remains one of the hottest areas in public markets.
  • Investors have rewarded AI infrastructure companies such as CoreWeave since its public debut.
  • Nscale has announced very large infrastructure commitments involving Microsoft and multiple multi-hundred-megawatt campuses, which is the type of growth story public investors currently like.

Reasons it could slip into 2027

The challenge is that public-market investors increasingly want proof of:

  • Revenue growth
  • Actual GPU deployments
  • Utilization rates
  • Long-term customer contracts
  • Cash-flow visibility

rather than just power agreements and future construction plans.

Unlike CoreWeave, which entered public markets after operating large GPU fleets for years, much of Nscale’s most ambitious capacity remains future-dated. That creates execution risk that investors will scrutinize heavily.

What I would watch for

If Nscale is genuinely targeting a late-2026 IPO, I would expect to see during the next few months:

  1. Appointment of lead underwriters (Goldman Sachs, Morgan Stanley, JPMorgan, etc.).
  2. Public filing activity or confidential filing reports.
  3. More detailed revenue disclosures.
  4. Announcements of operational GPU deployments, not just planned deployments.
  5. Additional long-term customer agreements.

My assessment

If I had to assign probabilities today:

OutcomeProbability
IPO in Q4 2026~40%
IPO slips into H1 2027~45%
IPO delayed beyond 2027~15%

That’s not based on any insider information—just on where Nscale appears to be in its infrastructure rollout compared with where most AI infrastructure companies are when they ring the bell.

The most important signal is not the IPO date itself. It’s whether Nscale can demonstrate that its Norway, Texas, Portugal, and future UK capacity are being converted into revenue-generating GPU clusters with high utilization. If that evidence emerges during 2026, a late-2026 IPO becomes much more plausible.

CoreWeave

CoreWeave is an AI cloud provider that specializes in delivering large-scale GPU infrastructure for AI training, inference, HPC, rendering, and scientific computing.

The company started life as a GPU-focused cloud provider and has evolved into one of the largest independent AI infrastructure companies in the world.

Unlike AWS, Azure, and Google Cloud, which offer AI as part of a broader cloud portfolio, CoreWeave is almost entirely focused on GPU-accelerated workloads.

CategoryDetails
Founded2017
HeadquartersRoseland, New Jersey, USA
FocusAI Cloud Infrastructure
Primary BusinessGPU-as-a-Service
Main CustomersOpenAI, Microsoft, NVIDIA ecosystem, AI startups
Major HardwareNVIDIA H100, H200, GB200, Blackwell
CompetitorsAWS, Azure, Google Cloud, Crusoe, Lambda, Nscale

How CoreWeave Started

The company originally operated in cryptocurrency mining.

Management realized early that:

  • GPUs used for mining
  • GPUs used for AI training
  • GPUs used for rendering

all required similar infrastructure.

When the AI boom began following the success of ChatGPT, CoreWeave pivoted aggressively into AI compute.

This turned out to be one of the best-timed pivots in the technology industry.


CoreWeave’s Business Model

Think of CoreWeave as:

NVIDIA

CoreWeave

AI Companies

Instead of:

NVIDIA

Microsoft Azure
AWS
Google Cloud

AI Companies

CoreWeave sits between NVIDIA and AI customers.


What Services Does CoreWeave Offer?

1. AI Training Clusters

Used for:

  • Large Language Models (LLMs)
  • Foundation Models
  • Multimodal Models
  • Scientific AI

Examples:

  • GPT-style models
  • Image generation models
  • Robotics models

Typical infrastructure:

  • Thousands of GPUs
  • InfiniBand networking
  • Petabytes of storage

2. AI Inference

After a model is trained:

Training

Model

Inference

Inference is what happens when:

  • You ask ChatGPT a question
  • Generate an image
  • Run a chatbot

CoreWeave provides infrastructure for this at scale.


3. HPC

High Performance Computing workloads:

  • Weather modelling
  • Genomics
  • Drug discovery
  • CFD
  • Physics simulations

This is an area where CoreWeave competes with traditional HPC centres.


4. GPU Cloud

Instead of buying:

  • H100s
  • H200s
  • Blackwell systems

Customers rent them by:

  • Hour
  • Day
  • Month

Why NVIDIA Likes CoreWeave

NVIDIA has invested in CoreWeave because CoreWeave helps NVIDIA:

  • Deploy GPUs faster
  • Reach AI startups
  • Increase GPU utilization
  • Expand GPU cloud capacity

NVIDIA has been both a supplier and investor.


CoreWeave Infrastructure

Typical CoreWeave clusters contain:

NVIDIA GPUs

InfiniBand

GPU Nodes

High-speed Storage

Kubernetes

Customer Workloads

Technologies typically include:

  • NVIDIA DGX
  • HGX
  • InfiniBand
  • RoCE
  • Kubernetes
  • Slurm
  • Object Storage

How Big is CoreWeave?

By 2026, CoreWeave is operating or building infrastructure measured in:

  • Hundreds of thousands of GPUs
  • Multiple gigawatts of power
  • Dozens of AI data centres

This puts them among the largest AI-focused cloud providers globally.


Why Microsoft Matters

One of CoreWeave’s biggest customers has been Microsoft.

Microsoft has used CoreWeave capacity to supplement Azure AI infrastructure when Azure could not provision GPUs quickly enough.

This relationship helped accelerate CoreWeave’s growth enormously.


CoreWeave vs Nscale

AreaCoreWeaveNscale
Founded20172024
StageMature AI cloudEmerging AI hyperscaler
GPUs Deployed TodayVery LargeMore Limited
RevenueMuch HigherEarlier Growth
Operational ExperienceExtensiveBuilding
US PresenceMajorGrowing
Europe PresenceGrowingLarge Future Pipeline
Data CentresOperating TodayMany Future Builds
AI Cloud PlatformMatureDeveloping

What Would Interest an SRE?

For someone coming from:

  • Kubernetes
  • Observability
  • OpenTelemetry
  • Prometheus
  • Mimir
  • Loki
  • Tempo
  • HPC

CoreWeave is fascinating because it combines:

Infrastructure Scale

Thousands of servers per cluster.

AI Networking

  • InfiniBand
  • RoCE
  • GPUDirect RDMA

Storage

  • High-throughput parallel storage
  • Object storage
  • Checkpointing

Reliability

When a training run consumes:

10,000 GPUs
×
7 days

a single infrastructure failure can cost millions of dollars.

This creates unique SRE challenges around:

  • Cluster reliability
  • GPU scheduling
  • Capacity management
  • Fleet automation
  • Telemetry at hyperscale
  • AI workload observability

Why CoreWeave is Important

CoreWeave is one of the first companies to prove that a specialist AI cloud provider can compete with traditional hyperscalers.

The company effectively created a new category:

Traditional Cloud
AWS
Azure
GCP

vs

AI Cloud
CoreWeave
Crusoe
Lambda
Nscale

That category is now one of the fastest-growing areas of infrastructure technology and is driving much of the current AI infrastructure build-out worldwide.

CoreWeave’s stock has had one of the most volatile post-IPO journeys in the AI infrastructure sector.

Share Price Since IPO

CoreWeave completed its Nasdaq IPO in March 2025 under the ticker CRWV. The IPO was downsized before launch, raising about $1.5 billion rather than the larger amount initially targeted.

The broad trajectory has been:

PeriodApproximate Story
Mar 2025 IPOWeak initial reception and downsized offering
Apr–Jun 2025Strong AI enthusiasm drove shares sharply higher
Jun 2025Reached all-time highs around $187/share
H2 2025Significant correction as investors focused on debt, losses, and data-centre execution
Early 2026Recovery driven by AI demand, Anthropic, Meta, OpenAI and enterprise growth
Jun 2026Trading around $107/share

Recent trading puts the company at a market capitalization of roughly $56 billion.


The Good News Financially

Revenue Growth Is Extraordinary

CoreWeave is one of the fastest-growing infrastructure companies in the market.

Examples include:

  • Revenue more than doubled year-over-year in multiple recent quarters.
  • Enterprise adoption is expanding beyond AI labs into financial services and large enterprises.
  • Revenue backlog reached approximately $99.4 billion as of Q1 2026.

That backlog is enormous and provides strong visibility into future revenue.


Major Customers

CoreWeave has secured relationships with:

These are arguably the most important AI infrastructure customers on the planet.


Scale Advantage

Reuters recently noted that CoreWeave has:

  • More than 1 GW already deployed
  • More than 3.5 GW contracted for future deployment

This places it among the largest dedicated AI infrastructure operators globally.


The Risks

Massive Debt Load

This is the biggest concern.

CoreWeave financed much of its growth through:

  • Asset-backed debt
  • Infrastructure loans
  • GPU-backed financing
  • Convertible notes

Multiple analysts and investors have pointed to the company’s very large debt burden as its primary financial risk.

The business model requires spending billions before revenue arrives.


Still Losing Money

Despite explosive revenue growth, CoreWeave remains unprofitable on a net-income basis.

Investors are essentially betting that:

Revenue Growth
>
Interest Costs + Depreciation + Expansion Costs

over the long term.

Recent earnings showed revenue beating expectations while margins and profitability remained under pressure.


Customer Concentration

Historically, a large portion of revenue has come from a relatively small number of customers.

If:

  • OpenAI
  • Microsoft
  • Meta
  • Anthropic

decide to build more capacity themselves, future growth could be affected.

This is one reason investors closely watch customer mix and backlog growth.


Why Investors Still Like It

The bullish thesis is straightforward:

  1. AI demand continues growing.
  2. GPU supply remains constrained.
  3. Training and inference workloads keep increasing.
  4. CoreWeave owns and operates the infrastructure needed to satisfy that demand.

In that scenario, today’s debt becomes manageable because revenue grows faster than financing costs.


Compared with Nscale

If I compare the two today:

AreaCoreWeaveNscale
Public CompanyYesNot yet
Market Cap~$56BPrivate
RevenueMulti-billionMuch smaller
Operational GPU CapacityVery largeLimited publicly visible
Revenue Backlog~$99BNot publicly disclosed at same level
DebtVery highMuch lower today
Execution RiskModerateHigh
Infrastructure MaturityEstablishedEmerging

CoreWeave’s biggest challenge is financial leverage.

Nscale’s biggest challenge is execution.

CoreWeave has already proven it can build and operate AI infrastructure at scale. The question investors are asking is whether it can generate enough cash flow to justify the enormous capital expenditure and debt required to stay ahead in the AI compute race.

Crusoe

Crusoe is arguably the third major AI infrastructure challenger behind CoreWeave and the large hyperscalers, and alongside Nscale and Radiant in the race to build AI factories.

What makes Crusoe unique is that it evolved from an energy company into an AI infrastructure company.

Its progression has been roughly:

Flared Gas Capture

Power Generation

Bitcoin Mining

GPU Infrastructure

AI Cloud

AI Factories

Today the company describes itself as an “AI Factory Company” rather than a traditional cloud provider.


Current Position

Valuation

Crusoe raised:

  • $600M Series D (2024)
  • $1.375B Series E (2025)

at a valuation exceeding $10 billion.

There are also industry reports suggesting private-market discussions at significantly higher valuations during 2026, though these are not official company figures.


Funding Strength

Crusoe has now raised approximately:

  • $3.8B+ equity funding
  • Additional billions in project finance and credit facilities

including a $750M Brookfield-backed credit facility.

Compared with many startups, Crusoe has become exceptionally well capitalized.


The Abilene AI Campus

The company’s flagship project is:

Abilene, Texas

This has become one of the largest AI infrastructure projects in the world.

Public reports describe:

  • 1.2 GW campus
  • Up to ~400,000 NVIDIA GB200-class GPUs planned
  • $15B+ joint venture funding
  • Major Oracle/OpenAI involvement
  • Multiple operational buildings already online

This campus is one of the key foundations of the Stargate ecosystem.


Relationship With OpenAI, Oracle & Microsoft

Crusoe sits at the center of a fascinating triangle:

OpenAI

Oracle

Crusoe

Microsoft

Recent developments have been mixed:

Positive

Oracle states:

  • Abilene remains on schedule
  • Two buildings are operational
  • Additional Stargate capacity remains under development

Complicated

Several planned expansions have changed tenants or scope.

Reports indicate:

  • OpenAI and Oracle stepped back from some expansion plans.
  • Microsoft subsequently agreed to lease part of the adjacent capacity.
  • Meta has reportedly evaluated some available capacity.

This isn’t necessarily bad news—it may actually demonstrate that demand is broad enough that multiple hyperscalers are competing for capacity.


Revenue Performance

Industry estimates suggest:

YearRevenue
2024~$276M
2025~$998M
2026Potentially >$2B

These are not audited public-company figures but are widely cited estimates reflecting the company’s rapid growth trajectory.

If accurate, Crusoe would be among the fastest-growing infrastructure companies globally.


Why Investors Like Crusoe

1. Speed

Crusoe has developed a reputation for building AI infrastructure extremely quickly.

Some investors explicitly cite build speed as a competitive advantage versus traditional data-center developers.


2. Vertical Integration

Unlike many competitors, Crusoe controls:

Power

Generation

Infrastructure

Data Centres

GPU Cloud

This resembles Radiant’s strategy and increasingly resembles Nscale’s.


3. AI Factory Focus

The company is moving beyond:

GPU Rental

toward:

Complete AI Factories

which is where the largest contracts are emerging.


Current Challenges

1. Customer Concentration

Much of Crusoe’s growth is tied to:

  • OpenAI
  • Oracle
  • Microsoft

This creates concentration risk.

If one customer changes strategy, large projects can be affected.


2. Capital Intensity

Like CoreWeave, Crusoe requires enormous capital expenditures.

Building:

  • Multi-GW campuses
  • Power infrastructure
  • GPU fleets

requires tens of billions of dollars.


3. Project Volatility

Recent examples include:

  • Wyoming project pause
  • Changing Stargate scope
  • Customer reallocations between OpenAI, Oracle, Microsoft and others

This demonstrates that even the hottest AI infrastructure projects are not immune to execution risk.


How Crusoe Compares

CategoryCoreWeaveCrusoeNscaleRadiant
Public CompanyYesNoNoNo
Valuation~$56B market cap$10B+ privatePrivatePrivate
AI Cloud PlatformMatureGrowing rapidlyEmergingOri platform
Operational AI InfrastructureVery largeLargeSmaller todayEarly
AI Factory FocusStrongVery strongVery strongVery strong
Energy IntegrationModerateStrongStrongExceptional
IPO CandidateAlready publicLikely future IPOPotential IPOLong-term possibility

What I Think of Crusoe

Among the “new hyperscalers”:

  1. CoreWeave is currently the operational leader.
  2. Crusoe is probably the most advanced private AI infrastructure company.
  3. Nscale has one of the largest future pipelines.
  4. Radiant may have the strongest long-term capital structure because of Brookfield.

Crusoe’s biggest strength is that it has already proven it can deliver and operate very large AI campuses while still retaining startup-level speed. Its biggest challenge is moving from a few gigantic flagship projects into a diversified, repeatable AI infrastructure business that is less dependent on any single customer or project.

CoreWeave vs Crusoe vs Nscale

These are arguably the three most important “Neoclouds” today.

All three are trying to become the AI-era equivalent of hyperscalers, but they are taking very different paths.

Executive Summary

CompanyCoreWeaveCrusoeNscale
Founded201720182024
StatusPublic companyLarge private companyLarge private company
Core IdentityAI cloud providerAI factory builderAI infrastructure hyperscaler
Geographic StrengthUSUSEurope
Operational MaturityHighestHighEmerging
AI Cloud PlatformMost matureGrowingDeveloping
Energy OwnershipLimitedStrongStrong
Future Capacity PipelineLargeVery LargeEnormous
Biggest RiskDebtCustomer concentrationExecution
Biggest StrengthOperational excellenceInfrastructure deliveryPower + future capacity

CoreWeave is currently winning on execution. Crusoe is winning on AI factory construction. Nscale is winning on future infrastructure ambition.


1. CoreWeave

What CoreWeave Is

CoreWeave is fundamentally an AI-native cloud provider.

Think:

AWS for GPUs

except purpose-built for:

  • AI training
  • AI inference
  • LLMs
  • HPC

Its cloud platform is already mature and heavily used by large AI companies. CoreWeave operates dozens of data centres, hundreds of thousands of GPUs, and has become one of NVIDIA’s most important cloud partners.

Strengths

  • Most mature software platform
  • Largest operational fleet
  • Strong OpenAI, Microsoft, Meta, Anthropic relationships
  • Fastest revenue growth
  • Proven ability to monetize GPUs

CoreWeave reported more than $5B revenue and a backlog approaching $67B-$88B depending on reporting period.

Weaknesses

  • Huge debt load
  • Heavy capex requirements
  • Customer concentration
  • Public market scrutiny

2. Crusoe

What Crusoe Is

Crusoe is best described as:

Energy Company
+
AI Factory Builder
+
GPU Cloud

It started by monetizing stranded energy and evolved into building some of the largest AI campuses in the world.

The Abilene campus in Texas has become one of the flagship AI infrastructure projects globally and is tied to Oracle and OpenAI’s broader Stargate ecosystem.

Strengths

  • Extremely fast construction capability
  • Strong energy expertise
  • Large-scale AI factory delivery
  • Deep OpenAI/Oracle ecosystem integration

Weaknesses

  • Smaller cloud platform than CoreWeave
  • Less diversified customer base
  • Still heavily tied to a few mega-projects

What Crusoe Wants To Become

Crusoe appears to be evolving toward:

AI Factory Company

rather than simply a GPU cloud.


3. Nscale

What Nscale Is

Nscale is pursuing the most ambitious infrastructure vision.

Their strategy is:

Power

Land

Data Centres

GPUs

Cloud Platform

They are effectively trying to build a European AI hyperscaler from scratch.

Strengths

  • Massive future pipeline
  • Strong sovereign AI positioning
  • European leadership position
  • Large power commitments
  • Strong Microsoft/OpenAI/NVIDIA relationships

Weaknesses

  • Much of capacity remains future-dated
  • Less operational experience
  • Less mature cloud platform
  • Execution risk

Public reporting has highlighted that several headline projects remain in buildout or planning phases rather than being fully operational today.


The Strategic Difference

CoreWeave

Started with:

GPUs

Then added:

Cloud
→ Data Centres
→ Power

Crusoe

Started with:

Energy

Then added:

Data Centres
→ GPUs
→ AI Factories

Nscale

Started with:

Power + Infrastructure

Then added:

GPUs
→ Cloud
→ Sovereign AI

Which Company Is Furthest Ahead Today?

Operational AI Cloud

Winner:

🥇 CoreWeave

Reason:

  • Largest operational fleet
  • Most mature software platform
  • Largest customer base

AI Factory Construction

Winner:

🥇 Crusoe

Reason:

  • Abilene
  • Stargate involvement
  • Proven delivery capability

Future Capacity Pipeline

Winner:

🥇 Nscale

Reason:

  • Norway
  • Portugal
  • Texas
  • UK projects
  • Sovereign AI initiatives

Which Is Closest To Becoming a New Hyperscaler?

Today

CoreWeave

|
Crusoe
|
Nscale

By 2030 (Potential)

CoreWeave
Crusoe
Nscale

All three could be major AI infrastructure providers, but they will likely specialize differently:

CompanyLikely Long-Term Identity
CoreWeaveAI Cloud Hyperscaler
CrusoeAI Factory & Energy Infrastructure Leader
NscaleSovereign AI & European AI Hyperscaler

From an SRE / Cloud Infrastructure Perspective

If you wanted to work on the most technically mature environment today:

CoreWeave

If you wanted to build some of the world’s largest AI campuses:

Crusoe

If you wanted to help create a new AI hyperscaler from the ground up:

Nscale

That is the clearest distinction between the three companies as of mid-2026.

Who is Radiant?

Radiant/Ori is one of the more interesting challengers because they are not trying to copy CoreWeave or Nscale exactly.

Instead, they are attempting to combine:

  • Brookfield’s enormous infrastructure and energy assets
  • Ori’s AI cloud software platform
  • NVIDIA’s AI factory ecosystem
  • Sovereign AI demand from governments and large enterprises

into a vertically integrated AI infrastructure company.

What is Ori?

Before the merger, Ori Industries was a UK AI cloud company founded in 2019.

Ori built:

  • Distributed GPU cloud infrastructure
  • AI model training platforms
  • AI deployment services
  • Multi-location AI compute services

The company operated AI infrastructure across more than 20 global locations and developed software to orchestrate AI workloads across GPU infrastructure.

Think of Ori as:

What CoreWeave built:
GPU Cloud Platform

What Ori built:
Distributed AI Infrastructure Platform

Ori’s technology is arguably the key intellectual property in the merger.


What is Radiant?

Radiant is Brookfield’s AI infrastructure company.

Brookfield is one of the world’s largest infrastructure investors with hundreds of billions under management spanning:

  • Power generation
  • Transmission
  • Renewable energy
  • Real estate
  • Data centres
  • Infrastructure projects

Radiant was created to become Brookfield’s AI compute platform.


Why Brookfield Matters

This is where Radiant becomes potentially disruptive.

Most AI clouds have a structure like:

Raise Venture Capital

Buy GPUs

Rent Datacentre Space

Sell Compute

CoreWeave largely grew this way.

Nscale is evolving toward:

Power

Datacentres

GPUs

Cloud Platform

Radiant starts with:

Brookfield Capital
+
Brookfield Power
+
Brookfield Land
+
Brookfield Datacentres
+
Ori Software

That means they potentially have access to cheaper capital than most AI startups.


Their Stated Strategy

Radiant has publicly described itself as a vertically integrated AI infrastructure platform.

Target customers include:

  • Sovereign governments
  • Hyperscalers
  • Tier-1 telecom operators
  • Large enterprises

Rather than simply renting GPUs to startups.

Their focus appears to be:

AI Factories

Large installations of:

  • NVIDIA GPUs
  • AI networking
  • AI storage
  • AI orchestration software

built for nations and large corporations.


The NVIDIA Connection

Radiant is built around NVIDIA’s AI factory vision.

Public statements indicate:

  • NVIDIA contributed capital to Brookfield’s AI fund.
  • NVIDIA will supply GPUs.
  • Radiant will deploy NVIDIA DSX AI factories.

This places them squarely in the same ecosystem as:

  • CoreWeave
  • Crusoe
  • Lambda
  • Nscale

but with a heavier focus on sovereign infrastructure.


How They Intend to Join the Hyperscaler Club

The strategy appears to be:

Phase 1: Acquire Software

Acquire Ori.

Result:

GPU Cloud Software
AI Orchestration
AI Platform Expertise

✓ Completed.


Phase 2: Leverage Brookfield Infrastructure

Use Brookfield’s:

  • powered land
  • data centres
  • energy assets

instead of building everything from scratch.

This is a major advantage versus startups.


Phase 3: Build Sovereign AI Factories

Target:

  • governments
  • national AI initiatives
  • regulated industries

This aligns well with Europe’s push toward sovereign AI and AI factories.


Phase 4: Scale Like a Utility

This is probably the most important difference.

Several executives have stated they want AI infrastructure financed like:

Power Stations
Utilities
Rail Networks
Airports

rather than venture-backed cloud startups.

That could significantly lower financing costs compared with many GPU cloud providers.


How Do They Compare?

CompanyCoreWeaveNscaleRadiant
Founded201720242026
PublicYesNoNo
Core StrengthOperating GPU cloudsBuilding AI campusesInfrastructure + software
Main BackerPublic marketsInvestors/NVIDIABrookfield
FocusAI cloudAI hyperscalerAI utility model
Sovereign AIModerateStrongVery Strong
Capital AccessGoodGoodPotentially Exceptional
Operational GPU Scale TodayHighestLowerVery Early

What Could Make Radiant Dangerous?

If you look at this as an SRE or infrastructure engineer, the biggest threat to competitors is not technology.

It is cost of capital.

CoreWeave’s biggest weakness is debt.

Nscale’s biggest challenge is execution.

Radiant’s pitch is:

“We already own the power, land, infrastructure financing, and data-centre expertise. We just needed the AI cloud software.”

That is precisely what the Ori acquisition gives them.

If Brookfield genuinely deploys the AI Infrastructure Fund at the scale discussed publicly (up to $10B fund commitments and potentially much larger through co-investment structures), Radiant could become one of the few companies capable of competing with CoreWeave, Nscale, Crusoe, and the hyperscalers in the sovereign AI factory market.

For someone with a background in Kubernetes, OpenStack, HPC, AI infrastructure, observability, Ceph, Slurm, and GPU platforms, Radiant is arguably one of the most interesting companies to watch over the next 2–3 years because they are trying to build the “AI utility company” rather than just another GPU cloud.

Is Radiant Ramping Up Recruitment?

If I were advising Radiant’s leadership after the Brookfield + Ori merger, I would not primarily hire more software developers or more data-centre staff initially.

The biggest challenge is integrating:

Energy Infrastructure
+
Data Centres
+
GPU Factories
+
Cloud Platform
+
Sovereign AI

into a single operating model.

That requires a very specific set of engineers.


Tier 1 — Recruit Immediately

These are the highest-priority hires.

1. Principal AI Infrastructure Architects

Need 5–10 globally.

Background:

  • CoreWeave
  • Microsoft Azure
  • AWS
  • Google
  • Oracle Cloud
  • NVIDIA
  • Crusoe
  • Nscale

Skills:

  • AI factories
  • Multi-GW campuses
  • GPU fabrics
  • Infrastructure strategy

These people define the architecture.

Without them everyone builds different solutions.


2. Staff/Principal GPU Platform Engineers

Need 20–50.

Skills:

  • Kubernetes
  • GPU Operator
  • Slurm
  • CUDA
  • MIG
  • NCCL
  • DGX/HGX

Responsibilities:

GPU lifecycle
GPU scheduling
GPU utilization
GPU observability

These are the people that actually make expensive GPUs productive.


3. Staff Network Engineers

Need 20–40.

The AI industry is becoming:

Network Limited
rather than
GPU Limited

Experience:

  • InfiniBand
  • RoCE
  • EVPN/VXLAN
  • Arista
  • NVIDIA Spectrum
  • Mellanox

Sources:

  • Meta
  • Microsoft
  • NVIDIA
  • Oracle OCI
  • Azure

4. Site Reliability Engineers

Need 30–60.

Not generic web SREs.

Need:

  • Kubernetes
  • Linux
  • GPU clusters
  • Storage
  • Automation

Focus:

Reliability
Capacity
Performance
Automation

5. Observability Platform Engineers

Need 10–20.

This is where many AI companies are currently weak.

Technology:

  • OpenTelemetry
  • Prometheus
  • Mimir
  • Loki
  • Tempo
  • ClickHouse
  • Kafka

Mission:

Observe
Everything

including:

  • GPUs
  • Power
  • Cooling
  • Storage
  • Training jobs
  • Networks

This is one of the areas where someone with your background would be valuable.


Tier 2 — Build During Year One

6. OpenStack Engineers

Many sovereign customers still want:

Private Cloud

rather than:

Public GPU Cloud

Need:

  • Nova
  • Neutron
  • Cinder
  • Ironic

Especially for government customers.


7. Storage Engineers

Need 15–30.

Experience:

  • Ceph
  • Lustre
  • BeeGFS
  • Weka
  • VAST

AI clusters consume storage at enormous scale.


8. Infrastructure Software Engineers

Need 20–50.

Build:

  • Fleet management
  • Provisioning
  • Capacity systems
  • Internal developer platforms

Languages:

  • Go
  • Python
  • Rust

9. Platform Security Engineers

Need 10–20.

Focus:

  • Supply chain security
  • GPU isolation
  • Sovereign compliance
  • Zero trust

Tier 3 — The Secret Weapon

These are the hires that separate a cloud provider from an AI hyperscaler.

10. HPC Engineers

Need 20–40.

Backgrounds:

  • National labs
  • Universities
  • Supercomputing centres

Skills:

  • Slurm
  • MPI
  • InfiniBand
  • Parallel filesystems

These people understand:

10,000 GPU training jobs

better than most cloud engineers.


11. Power Systems Engineers

This is where Brookfield can dominate.

Need:

  • Utility engineers
  • Grid engineers
  • High-voltage engineers

Most AI companies have very few.

Brookfield already has many.

Radiant should integrate them directly.


12. Cooling Engineers

Future AI factories may be:

100MW+
500MW+
1GW+

Cooling becomes strategic.

Need expertise in:

  • Liquid cooling
  • Direct-to-chip
  • Immersion

The Leadership Layer

Radiant’s biggest risk is organizational silos.

Avoid:

Brookfield Team
|
|
Ori Team

Instead build:

AI Infrastructure
|
+-- Energy
+-- Datacentres
+-- GPU Platform
+-- SRE
+-- Observability
+-- Security

If I Had £100M Hiring Budget

I’d prioritize:

RoleApprox Headcount
GPU Platform Engineers40
SREs40
Network Engineers30
Infrastructure Software Engineers30
Storage Engineers20
Observability Engineers15
HPC Engineers20
Security Engineers15
AI Infrastructure Architects10
Power/Cooling Specialists20

Total: ~240 specialist engineers.


The Three Most Valuable Hires

If Radiant could only hire three categories tomorrow:

  1. Principal GPU Platform Engineers
  2. Principal AI Networking Engineers
  3. Principal Observability/SRE Engineers

Those three groups determine whether a 100,000-GPU AI factory operates at:

95% utilization

or

60% utilization

The difference is potentially hundreds of millions of dollars per year in infrastructure efficiency. For a company trying to become an AI utility, those engineering disciplines are arguably more important than almost any other technical hiring category.

Oracle’s Journey

Phase 1: Database Company (1977-2010)

For decades Oracle was essentially:

Databases
+
Enterprise Software

Revenue came from:

  • Oracle Database
  • Enterprise applications
  • Middleware
  • Support contracts

Oracle dominated enterprise IT but missed the early public cloud wave.


Phase 2: Late Cloud Entrant (2010-2020)

AWS, Azure and Google Cloud were already well established.

Oracle’s first cloud attempts struggled because they largely tried to:

Move Oracle Products

Into Oracle Cloud

rather than building a cloud-native platform.

OCI v1 wasn’t competitive.


Phase 3: OCI Rebuild (2018-2024)

This is where Oracle changed direction.

Under Clay Magouyrk’s leadership, OCI was essentially rebuilt from scratch.

Key design decisions:

Bare Metal First

Unlike AWS:

Physical Server

Hypervisor

VM

OCI emphasized:

Physical Server

Customer

This became attractive for:

  • HPC
  • AI
  • Databases

RDMA Networking

Oracle invested heavily in:

  • RoCE
  • RDMA
  • HPC fabrics

Years before AI made these mainstream.

This is one reason OCI became attractive for GPU clusters.


Autonomous Infrastructure

OCI automated large parts of:

  • provisioning
  • patching
  • operations

allowing Oracle to run cloud regions with fewer people.


Phase 4: AI Pivot (2023-Present)

ChatGPT changed everything.

Oracle suddenly found that:

Their Strengths Were AI Strengths

They already had:

✓ Bare metal

✓ HPC networking

✓ RDMA

✓ Large data centres

✓ Enterprise customers

These are exactly what AI workloads need.


The OpenAI Relationship

This is where Oracle became a serious AI player.

Oracle started providing infrastructure for:

  • OpenAI
  • Microsoft
  • Stargate

through extremely large GPU deployments.

Oracle is now one of the biggest buyers of NVIDIA GPUs in the world.


Oracle’s AI Infrastructure Today

Oracle is building:

GB200 Clusters

Blackwell Clusters

RoCE Fabrics

AI Superclusters

At a scale that rivals many neoclouds.

Some deployments involve:

10,000+
50,000+
100,000+ GPUs

depending on project.


Why Oracle Is Different From CoreWeave

CoreWeave started with:

GPUs

Cloud

Oracle started with:

Cloud

GPUs

This gives Oracle advantages.


Existing Customers

Oracle already has:

  • banks
  • governments
  • telecoms
  • healthcare

These customers are now buying AI services.

CoreWeave must acquire those customers.

Oracle already has them.


Existing Revenue

Oracle generates tens of billions annually.

This means they can fund AI expansion from operating cash flow.

CoreWeave relies more heavily on:

  • debt
  • equity
  • project financing

Existing Global Footprint

OCI already operates dozens of regions.

Nscale and Crusoe are still building much of theirs.


Is Oracle Becoming a Hyperscaler?

Oracle already is one.

OCI is generally considered the fourth major hyperscaler after:

  1. AWS
  2. Azure
  3. Google
  4. Oracle

The question is really:

Is Oracle becoming an AI hyperscaler?

The answer is:

Yes.


Is Oracle Becoming a Neocloud?

Not really.

Neoclouds are generally:

AI First

Examples:

  • CoreWeave
  • Crusoe
  • Nscale
  • Radiant

Oracle is:

Cloud First

AI Enhanced

A different origin story.


What Oracle Is Morphing Into

I would describe Oracle as:

Traditional Hyperscaler
+
AI Factory Operator
+
GPU Supercluster Provider

In fact Oracle increasingly resembles:

AWS + CoreWeave

combined.

AWS scale.

CoreWeave-style GPU infrastructure.


Why This Matters for the AI Race

The biggest threat to CoreWeave, Nscale and Crusoe may not be each other.

It may be Oracle.

Because Oracle has:

✓ Existing cloud

✓ Existing customers

✓ Existing revenue

✓ Existing data centres

✓ Existing support organisation

✓ Existing enterprise sales force

✓ Massive GPU procurement

The neoclouds must build these capabilities.

Oracle already has them.


The Next 5 Years

If current trends continue:

CompanyLikely Position 2030
AWSLargest general cloud
AzureLargest enterprise AI cloud
GoogleAI + data platform leader
OracleAI infrastructure hyperscaler
CoreWeaveLargest independent AI cloud
CrusoeAI factory leader
NscaleSovereign AI hyperscaler
RadiantAI utility platform

My view is that Oracle is not becoming a neocloud.

Instead, Oracle is doing something arguably more powerful:

It is transforming from a traditional hyperscaler into an AI hyperscaler while retaining all the advantages of an established cloud provider.

That combination of existing scale, enterprise relationships, and AI infrastructure investment is why Oracle has suddenly become one of the most important players in the AI infrastructure market.

Is Oracle the opposite of Radiant and vice versa?

Not exactly, but they are surprisingly close to being mirror images of each other.

If you look at their origins:

OracleRadiant
Started with softwareStarted with infrastructure
Database companyInfrastructure company
Built cloud platformAcquired cloud platform (Ori)
Added AI laterAdded AI from day one
Enterprise customers firstSovereign AI first
Compute-centricPower-centric
Cloud → AIInfrastructure → AI

A useful way to think about it is:

Oracle
-------
Software

Database

Cloud

AI Infrastructure

Radiant
--------
Infrastructure

Power

Data Centres

AI Infrastructure

So they are converging on a similar destination from opposite directions.


Oracle’s DNA

Oracle fundamentally thinks like a software company.

Its worldview is:

Application

Database

Cloud Platform

Infrastructure

Its biggest assets are:

  • Enterprise customers
  • Databases
  • SaaS products
  • Sales organisation
  • OCI platform

AI is an extension of those assets.

Oracle asks:

“How do we deliver AI to our existing customers?”


Radiant’s DNA

Radiant fundamentally thinks like an infrastructure company.

Its worldview is:

Power

Land

Data Centre

GPU Factory

AI Services

Its biggest assets are:

  • Brookfield capital
  • Brookfield power
  • Brookfield real estate
  • Brookfield infrastructure expertise
  • Ori’s AI platform

Radiant asks:

“How do we build the infrastructure that powers AI?”


The Biggest Difference

Oracle’s bottleneck is usually:

Customer Demand

They already have:

  • Data centres
  • Customers
  • Revenue

They need more GPUs and power.


Radiant’s bottleneck is usually:

Software & Customer Acquisition

They already have:

  • Capital
  • Infrastructure expertise
  • Energy

They need:

  • AI cloud adoption
  • Enterprise relationships
  • Platform scale

What They Are Trying To Become

Oracle is evolving toward:

AI Hyperscaler

Radiant is evolving toward:

AI Utility

Those are related but different.

Oracle Vision

Oracle Cloud
+
AI Superclusters
+
Enterprise AI

Think:

“AWS/Azure with massive AI capability.”


Radiant Vision

Power
+
Data Centres
+
AI Factories
+
Long-term Infrastructure Contracts

Think:

“National Grid meets CoreWeave.”


Why Radiant Could Learn From Oracle

Radiant lacks:

  • Enterprise software experience
  • Large-scale customer operations
  • Decades of cloud platform evolution

Oracle has all of that.


Why Oracle Could Learn From Radiant

Radiant understands:

  • Power economics
  • Infrastructure financing
  • Long-duration capital
  • Utility-scale thinking

areas where Oracle historically has less expertise.


If They Met In The Middle

The interesting thing is that both companies are converging toward something like:

Power

Data Centre

GPU Factory

Cloud Platform

Enterprise AI

The difference is where they started.

LayerOracle StrengthRadiant Strength
PowerModerateExceptional
Data CentresStrongExceptional
GPUsStrongEmerging
Cloud PlatformExceptionalGood (via Ori)
Enterprise SalesExceptionalDeveloping
Sovereign AIModerateStrong
Long-Term Infrastructure FinanceModerateExceptional

The More Interesting Comparison

I actually think the closest opposite of Radiant is not Oracle.

It’s CoreWeave.

CoreWeave

Started with:

GPUs

Cloud

Data Centres

Power

Radiant

Started with:

Power

Data Centres

Cloud

GPUs

Those are almost exact inverses.

Oracle sits somewhere else entirely because it arrived carrying:

Databases
+
Enterprise Software
+
Cloud Platform

which neither CoreWeave nor Radiant possessed.

So my assessment would be:

  • CoreWeave and Radiant are the closest opposites.
  • Oracle and Radiant are converging from opposite ends of the technology stack.
  • By 2030, Oracle and Radiant may end up looking surprisingly similar externally, even though one began as a software giant and the other as an infrastructure and energy giant.

Reflective Journeys: Oracle vs Radiant

Yes, in many cases Oracle employees affected by AI-related restructuring could be strong candidates for Radiant, but it depends heavily on which part of Oracle they came from.

The interesting thing is that Oracle and Radiant are moving toward the same destination from opposite directions:

Oracle
Database

Cloud

AI Infrastructure

Radiant
Power

Infrastructure

AI Infrastructure

That creates a surprising amount of skill overlap.

Oracle Employees Radiant Should Recruit Aggressively

OCI Engineers

These are probably the highest-value hires.

Experience:

  • OCI regions
  • Cloud operations
  • Bare metal
  • Networking
  • Cloud automation

Radiant needs people who know how to operate cloud infrastructure at scale.

These engineers bring exactly that.


AI Infrastructure Engineers

Oracle has been building:

  • GPU superclusters
  • RDMA fabrics
  • RoCE networks
  • AI training environments

Those skills are directly transferable to:

  • Radiant AI factories
  • GPU clouds
  • Sovereign AI deployments

OCI SREs

Particularly valuable:

  • Capacity planning
  • Reliability engineering
  • Infrastructure automation
  • Fleet management

Radiant will need these people immediately as AI factories scale.


Data Centre Engineers

Oracle has been building data centres globally.

Skills:

  • Capacity planning
  • Facility operations
  • Power
  • Cooling
  • Commissioning

These map extremely well to Radiant’s infrastructure-first strategy.


Network Engineers

Potentially the most valuable category.

Particularly if they have:

  • RoCE
  • RDMA
  • EVPN/VXLAN
  • High-performance networking

AI infrastructure is increasingly network-limited rather than GPU-limited.


Observability Engineers

This is a category many AI infrastructure companies underestimate.

Skills:

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Logging platforms
  • Distributed tracing

Radiant will eventually need to observe:

Power
Cooling
Networks
Storage
GPUs
Training Jobs
Cloud Platform

at enormous scale.


Oracle Employees Radiant May Need Less Of

Traditional ERP / Applications Teams

Experience in:

  • E-Business Suite
  • HR systems
  • Legacy applications

is less directly relevant.

Radiant is building infrastructure rather than enterprise applications.


Traditional Database Administration

Still useful, but lower priority.

Radiant’s biggest bottlenecks are more likely:

  • GPUs
  • Networking
  • Data centres
  • Cloud platforms

than Oracle Database administration.


Would It Be Good For The Employees?

Potentially yes.

Oracle is becoming:

Large AI Hyperscaler

Radiant is becoming:

AI Infrastructure Startup
with Brookfield backing

Some engineers prefer:

Oracle

  • Stability
  • Massive scale
  • Mature processes
  • Existing customer base

Radiant

  • Building from scratch
  • More influence
  • Faster decision making
  • Potentially larger individual impact

If I Were Radiant’s CTO

The first Oracle hires I would target would be:

  1. OCI Principal SREs
  2. OCI Network Architects
  3. OCI GPU Platform Engineers
  4. OCI Capacity Engineers
  5. OCI Observability Platform Engineers
  6. OCI Data Centre Build Engineers

These people have already operated infrastructure at scales that Radiant wants to achieve.


Looking at Your Background

Based on the areas you’ve worked deeply in—observability, OpenTelemetry, Prometheus/Mimir/Loki/Tempo, Kubernetes, HPC, storage, automation, cloud platforms, and AI infrastructure—the type of role that would likely be most valuable to a company like Radiant is not a generic SRE.

It would be something closer to:

  • Principal Observability Engineer
  • AI Infrastructure Observability Architect
  • Staff SRE (AI Platforms)
  • Platform Engineering Lead
  • AI Factory Telemetry Architect

because one of the hardest problems these emerging AI infrastructure companies will face is creating observability across the entire stack:

Power

Data Centre

Network Fabric

GPU Cluster

Kubernetes / Slurm

AI Workloads

Very few engineers have practical experience spanning that many layers.

One caveat: public reporting has discussed Oracle workforce reductions in various parts of the business, but I have not seen reliable evidence supporting a single confirmed figure of “30,000 layoffs” across Oracle as a whole. When evaluating career moves, it’s better to focus on the strategic trend—Oracle investing heavily in AI infrastructure and cloud—rather than any specific layoff number unless confirmed by Oracle itself.

Where is all the Money?

The short answer is:

The money is real, but most of it is not sitting in a bank account waiting to be spent.

What you’re seeing is a combination of:

  1. Cash flow
  2. Debt financing
  3. Equity financing
  4. Project finance
  5. Infrastructure finance
  6. Customer pre-commitments
  7. Stock market valuations

The AI infrastructure boom is probably the largest capital deployment into technology infrastructure since the construction of the Internet and mobile networks.


Where Does The Money Actually Come From?

Imagine a company announces:

$20 Billion AI Campus

Many people picture:

Bank Account

$20 Billion

That’s almost never what happens.

Instead:

Equity
+
Debt
+
Customer Contracts
+
Infrastructure Loans
+
Future Revenue

fund the project.


CoreWeave

CoreWeave is the easiest example.

They need:

  • GPUs
  • Data centres
  • Power
  • Networking

worth billions.

They fund this through:

Equity

Investors buy shares.

Debt

Banks lend money.

GPU-backed loans

This is fascinating.

CoreWeave can buy:

100,000 H100s

and lenders treat those GPUs almost like collateral.

Similar to:

Mortgage
↔ House

Loan
↔ GPU Fleet

Oracle

Oracle is different.

Oracle generates tens of billions in annual revenue.

Their funding comes primarily from:

Operating Cash Flow

Database Revenue
SaaS Revenue
Support Contracts
OCI Revenue

This is actual cash arriving every quarter.

Oracle can invest from profits.

Corporate Debt

Oracle also issues bonds.

For example:

Oracle Bond

Investors buy it

Oracle receives cash

This is normal corporate finance.


AWS

AWS funding is even simpler.

Amazon generates huge cash flows.

When AWS builds a data centre:

Retail Business
+
AWS Revenue
+
Debt Markets

fund it.

The money is real.


Microsoft

Microsoft is currently spending at extraordinary levels.

Funding comes from:

Windows

Office 365

Azure

LinkedIn

GitHub

Copilot

All producing cash.

Microsoft can spend tens of billions annually because they generate enormous free cash flow.


Nscale

Nscale is much more interesting.

Nscale doesn’t have Oracle’s cash flow.

Instead funding comes from:

Equity Investors

Strategic Investors

Infrastructure Finance

Project Finance

Future Customer Contracts

Think:

Power Agreement
+
Land
+
Customer Demand

Banks lend money

Crusoe

Crusoe is heavily project-finance oriented.

Example:

OpenAI

Needs Capacity

Oracle

Needs Capacity

Crusoe

Build Campus

The campus may be funded by:

  • Equity
  • Infrastructure loans
  • Project financing
  • Long-term customer commitments

Similar to how airports and power stations are financed.


Radiant

Radiant may have the strongest financing model.

Why?

Because Brookfield already finances:

  • Power stations
  • Airports
  • Ports
  • Railways
  • Data centres

worth hundreds of billions.

Brookfield understands:

Build Asset

Generate Revenue

Repay Debt

better than almost anyone.

Radiant can potentially tap into infrastructure capital that many neoclouds cannot.


Is The Money “Real”?

Yes.

But there are three different meanings.


Real Cash

Example:

Microsoft earns:

$100

Customer pays.

Microsoft receives:

$100 cash

Real money.


Debt

Example:

Bank lends:

$10 Billion

to build AI infrastructure.

Also real money.

But must be repaid.


Market Valuation

This is where people get confused.

Example:

CoreWeave market cap:

$56 Billion

That does NOT mean:

$56 Billion cash

exists.

It means:

Share Price × Shares Outstanding

equals $56B.

Much of that value exists “on paper.”


The Hidden Fuel: Pension Funds

Most people don’t realize who ultimately finances much of this.

The money often comes from:

  • Pension funds
  • Sovereign wealth funds
  • Insurance companies
  • Infrastructure funds

For example:

Teacher Pension

Infrastructure Fund

Brookfield

AI Data Centre

The chain can be surprisingly long.


Why Everyone Is Comfortable Lending

The reason banks are willing to lend is simple:

They believe AI demand will continue growing.

Their assumption is:

GPU Demand
>
Debt Cost

If true:

  • Loans get repaid.
  • Investors make money.
  • Infrastructure grows.

If false:

  • Some companies will fail.
  • Some campuses will be underutilized.
  • Some lenders will take losses.

The Biggest Risk

The entire AI infrastructure sector is making a giant bet:

Future AI Demand

If AI demand keeps growing:

  • Oracle wins.
  • Microsoft wins.
  • CoreWeave wins.
  • Crusoe wins.
  • Nscale wins.
  • Radiant wins.

If demand slows dramatically:

The most leveraged companies suffer first.

That is why investors currently view:

CompanyFinancial Risk
MicrosoftLow
OracleLow
AWSLow
GoogleLow
Radiant/BrookfieldModerate
CrusoeModerate
NscaleModerate-High
CoreWeaveHigh

The hyperscalers are largely spending from enormous existing cash flows. The neoclouds are spending mostly against future growth, future contracts, and infrastructure financing. The money is real, but much more of the neocloud funding stack depends on future demand continuing to justify today’s investments.

The 2026 AI Funding Diagram

Key changes since the Bloomberg/Morgan Stanley “AI Money Machine” chart from late 2025

OpenAI

  • Valuation increased from roughly $500B to over $730B–850B after its record funding rounds in 2026.
  • New major funding sources:
    • Amazon
    • SoftBank
    • Nvidia
  • OpenAI is now much less dependent on Microsoft than the original chart suggests.

Amazon (missing from the original)

Amazon is arguably the biggest omission now:

  • Invested approximately $50B in OpenAI.
  • Expanded AWS compute commitments.
  • OpenAI agreed to use AWS infrastructure and Trainium capacity.

SoftBank (missing from the original)

  • Became one of OpenAI’s largest financial backers.
  • Major participant in Stargate-style infrastructure funding.

CoreWeave

  • OpenAI relationship expanded to about $22.4B in AI infrastructure contracts.
  • Now one of OpenAI’s largest compute suppliers.
  • Public company rather than private neocloud startup.

Nvidia

  • Still sits at the center.
  • Added direct investment into OpenAI.
  • Continues investing in CoreWeave while simultaneously selling GPUs to it and buying capacity from it.

Nscale and Nebius

  • The original chart correctly anticipated their importance.
  • They now fit into a larger category of “GPU-native neoclouds” alongside CoreWeave.
  • Their role is increasingly as infrastructure providers for AI model companies rather than model developers themselves.

If Bloomberg redrew it today

The centre would probably look like:

                    Microsoft
|
|
Amazon ----\
\
SoftBank ----> OpenAI <---- Nvidia
/ | \
/ | \
/ | \
Oracle CoreWeave AMD
| |
| |
GPU Clouds (Nebius, Nscale)
|
Mistral / xAI / Figure / Cursor

The biggest differences

2025 Chart2026 Reality
Microsoft dominates OpenAI fundingAmazon + SoftBank now rival Microsoft
CoreWeave is peripheralCoreWeave is a central infrastructure supplier
Amazon absentAmazon is one of the largest players
SoftBank absentSoftBank is one of the largest financiers
OpenAI ≈ $500BOpenAI > $730B valuation
Nscale/Nebius nicheNscale/Nebius increasingly recognized as AI infrastructure providers

For your interests in AI infrastructure, neoclouds, and hyperscalers, a more useful 2026 version would actually be an “AI Infrastructure Ecosystem Map” showing:

  • Nvidia
  • AMD
  • OpenAI
  • Microsoft
  • Amazon
  • Oracle
  • SoftBank
  • CoreWeave
  • Nebius
  • Nscale
  • Crusoe
  • xAI
  • Mistral
  • Figure AI
  • Stargate

with arrows for:

  • Capital investment
  • GPU purchases
  • Cloud contracts
  • Equity stakes
  • AI model consumption

That would better reflect where the industry sits today than the original Bloomberg graphic.

An AI Job Revolution?

I asked this question at the start of the year: Is 2026 going to be the worst year for IT layoffs?

Now, it is mid year, so let’s look at IT/tech layoffs through mid-2026 compared to the past 10 years:

2026 Tech Layoffs (through June 2026)

Current total: ~156,000–172,000 tech workers laid off

  • As of June 9, 2026: 156,058 jobs cut across 50 tech companies
  • Some trackers show 172,130+ jobs cut for 2026
  • First 5 months (Jan-May): 128,940 tech workers laid off
  • March 2026 was the worst single month: 49,452 layoffs
  • Major employers: Oracle (30,000), Amazon (16,000–30,000), Meta, Microsoft, Dell

Comparison with Previous 10 Years

YearTech LayoffsKey Driver
2026 (through June)~156,000–172,000AI restructuring, over-hiring correction 
2025~105,000–244,851AI-led efficiency, economic uncertainty 
2023~263,000Peak: Overhiring reversal, ad market collapse 
2024~152,000AI restructuring, cost discipline 
2022~165,000Rate hikes, post-ZIRP correction 
2021~10,000Near-zero; record hiring year 
2020~80,000COVID shock 
2019~60,000Strategic pivots (ride-sharing, large tech) 
2018~60,000Strategic pivots 
2017~20,000–50,000Restructuring (Intel, Yahoo, Oracle) 
2016~20,000–50,000Restructuring 

Key Findings

2026 is NOT the worst year — it’s still behind the 2023 peak:

  • 2023 remains the worst year on record with ~263,000 tech jobs cut
  • 2026 is on track to potentially reach ~250,000–300,000 if the pace continues (currently at ~156K through June)
  • 2026 has already surpassed 2024 (~152K) and 2022 (~165K) in just 5–6 months

The trajectory is concerning: With 128,940 layoffs in just 5 months and March being the worst month in over a year, 2026 could rival 2023 if the second half remains severe. However, as of mid-2026, it’s still approximately 40–45% below the 2023 peak.

The primary driver has shifted from 2022–23’s funding correction to AI-driven substitution — roles being cut may not return even when revenue grows.

Are we talking about a Revolution?

So 2023 was financial, 2026 is AI restructuring, so is this a revolution in the IT job market?

Yes, this represents a genuine revolution in the IT job market, but it’s more complex than a simple “AI replacing jobs” narrative. Here’s what’s actually happening:

Two Different Layoff Waves

Aspect2022-2023 (Financial)2026 (AI-Led)
DriverOverhiring + rate hikes + funding correctionAI efficiency + cost-cutting + restructuring 
Cited reason” restructuring”AI now the leading reason (May 2026) 
AI’s shareMinimal49,135 cuts (13-25% of total) 

Why This Is a Revolution

1. Jobs Aren’t Just Being Cut — They’re Being Redefined

  • Routine, automation-prone roles fell 13% after ChatGPT’s debut
  • Analytical/technical/creative jobs grew 20% — AI is augmenting these roles
  • Not uniform elimination: AI reshapes more jobs than it replaces

2. The “AI Washing” Problem

Some companies exaggerate AI’s role to justify cuts that are actually about:

  • Cost-cutting and slowing demand
  • Correcting over-hiring from previous years
  • Cloud market leveling off (AWS example)

Reality: It’s both real displacement AND convenient scapegoating

3. Skill Requirements Are Shrinking for Some Roles

  • 7% fewer skills needed in automation-prone occupations
  • New skills emerging: prompt writing, AI tools, human-AI collaboration
  • Workers need non-automatable skills: judgment, interpersonal communication

4. The Funding Narrative Changed

“Whether or not specific jobs are being replaced by AI, the funding for those positions certainly is.”
— Andy Challenger, workplace specialist

CEOs face pressure to demonstrate AI investment value, making job cuts politically convenient.

The Bottom Line

Yes, this is revolutionary, but in a nuanced way:

Real transformation: Routine tasks are genuinely being automated
New demand: AI-augmented roles require different skills
⚠️Mixed motives: Some companies use AI as a blanket excuse
⚠️ Uncertain long-term: Short-term data shows reshaping, not mass elimination

For someone in your field (SRE, Kubernetes, observability), the outlook is different from entry-level roles:

  • Your work involves judgment, complex systems, and human coordination — harder to automate
  • But you’ll need AI literacy (prompting, AI tools in workflows)
  • The risk is not elimination but reduced team sizes with AI doing more of the routine monitoring/debugging

The tech industry’s guarantee of job stability is gone — this is the fundamental shift.

What are the Safe and Stable Jobs in IT?

Based on the latest data, here are the safe and stable IT jobs going into late 2026 and beyond, ranked by their resistance to AI disruption and growth trajectory:

Top Safe & Stable IT Roles (UK 2026–2035)

1. Cybersecurity Engineer 🛡️

Why safe:

  • Cyber threats continue to increase
  • AI creates new attack surfaces
  • Regulatory requirements keep growing

UK Salary:

  • Mid-Level: £55,000–£75,000
  • Senior: £75,000–£100,000
  • Principal/Lead: £100,000–£140,000

Growth Outlook:
One of the most resilient technology careers for the next decade.


2. Cloud Architect ☁️

Why safe:

  • Strategic infrastructure design
  • Multi-cloud and hybrid-cloud complexity
  • Requires business and technical judgement

UK Salary:

  • Senior Cloud Architect: £90,000–£130,000
  • Principal Architect: £130,000–£170,000

Growth Outlook:
Still seeing strong demand as enterprises modernise infrastructure.


3. AI / Machine Learning Engineer 🤖

Why safe:

  • Building and operating AI systems
  • Demand exceeds supply
  • Critical for AI adoption

UK Salary:

  • Mid-Level: £70,000–£100,000
  • Senior: £100,000–£140,000
  • Staff/Principal: £140,000–£220,000+

Growth Outlook:
Among the strongest growth areas through 2035.


4. Senior Cloud Engineer ☁️

Why safe:

  • Designs and operates large cloud platforms
  • Increasing focus on automation and reliability
  • Deep infrastructure expertise remains difficult to automate

UK Salary:

  • £75,000–£110,000

Typical Skills:

  • AWS/Azure/GCP
  • Kubernetes
  • Terraform
  • Observability
  • Security

Growth Outlook:
Strong demand across SaaS, fintech, AI and hyperscale companies.


5. Staff Cloud Engineer ☁️🚀

Why safe:

  • Technical leadership role
  • Cross-team architectural influence
  • Requires experience, judgement and organisational impact

UK Salary:

  • £100,000–£160,000
  • Elite AI/Hyperscaler firms: £160,000–£220,000+

Typical Employers:

  • Nscale
  • CoreWeave
  • Google
  • Microsoft

Growth Outlook:
One of the safest senior technical career paths available.


6. Site Reliability Engineer (SRE) ⚙️

Why safe:

  • Reliability remains business-critical
  • AI infrastructure requires even more operational excellence
  • Combines software, operations, cloud and observability

UK Salary:

  • Mid-Level: £65,000–£85,000
  • Senior SRE: £85,000–£120,000
  • Staff SRE: £120,000–£180,000+

Growth Outlook:
Particularly strong in AI, fintech and hyperscale environments.


7. Data Engineer 📊

Why safe:

  • Data pipelines underpin AI systems
  • Data governance requirements increasing
  • Real-time analytics demand growing

UK Salary:

  • £60,000–£90,000
  • Senior: £90,000–£130,000

Growth Outlook:
Consistently one of the most in-demand engineering roles.


8. Technical Project Manager 📋

Why safe:

  • Human coordination remains difficult to automate
  • AI increases project complexity

UK Salary:

  • £60,000–£90,000
  • Senior: £90,000–£130,000

9. Technical Product Manager 🧭

Why safe:

  • Strategy, prioritisation and stakeholder alignment
  • Strong human and business focus

UK Salary:

  • £70,000–£110,000
  • Senior: £110,000–£150,000

Key Upgrades to Stay Safe:

  1. AI literacy — prompt writing, AI tools in workflows
  2. Security focus — automation security is the 2026 priority
  3. Observability + AI — AI-driven analytics becoming standard

Bottom Line: What Makes Jobs “Safe”

FactorSafe JobsRisky Jobs
Task typeJudgment, coordination, creativityRoutine, repetitive, automation-prone
SkillsStrategic design, security, AI literacyStructured, predictable tasks
Human elementCross-functional managementSolo execution
Growth17-20%+ through 2030Declining or flat

The safest path: Combine your existing SRE/Cloud expertise with AI tools + security focus — this is the sweet spot for 2026-2035.

Find a new role or job after redundancy

The job market is still pretty tight in April 2026, and IT is tougher than the headline unemployment numbers suggest because employers are hiring more selectively, keeping vacancy growth subdued, and raising the bar for experience. UK labour-market reporting says hiring is close to stabilising, but conditions remain challenging, and technology roles are resilient relative to the wider market rather than broadly easy to land [1][2].

What is happening

  • Employers are still cautious after a long slowdown in vacancies and hiring confidence, with UK reports describing the market as close to bottoming out rather than clearly recovering [1][2].
  • Competition is high because more candidates are chasing fewer openings, especially in entry-level and mid-level roles [3][2].
  • In IT, companies are still investing, but they are being very selective about which roles they open and often prefer people with niche, immediately useful skills [4][1].

Why IT feels harder

  • AI is reshaping hiring, and some employers are explicitly reducing junior or commodity-type roles because automation can handle part of that work [4][2].
  • Layoffs in tech have added experienced candidates back into the market, which makes competition worse for everyone else [5][6].
  • Many postings now expect broader skill sets than before, so “good enough” candidates often get filtered out quickly [4][3].

Where demand still exists

  • Cyber security, data, AI, cloud, and other specialist infrastructure roles are still among the strongest areas in the UK IT market [4][1].
  • Engineering and technology hiring is described as relatively more resilient than the wider labour market, even though demand is still weak compared with boom periods [1].
  • Employers are still looking for people who can deliver immediately, particularly in roles tied to automation, digital transformation, and AI enablement [4][1].

Practical read

If you are already in IT, the market is difficult but not dead: experienced people with scarce skills are still getting opportunities, while generic support, junior dev, and broad “all-rounder” roles are the hardest to place [4][2]. For job seekers, the main challenge in 2026 is not absolute lack of jobs, but a mismatch between what many employers want and what most applicants can show on paper [3][2].

A simple way to think about it: 2026 is not a “no jobs” market, it is a “harder to get shortlisted” market, especially in IT [1][3].

How long to find IT job after layoff 2026?

In 2026, a realistic IT job-search timeline after a layoff is often 4 to 6 months, with some people landing in 6 to 8 weeks and others taking much longer depending on seniority, specialization, location, and how targeted their search is [1][2]. Broader job-market data also suggests the average job search after a layoff is around five to six months, which lines up with the tech-specific estimates [2].

Typical timeline

  • Fast outcomes: about 6 to 8 weeks if your skills are in demand, you have strong referrals, and you apply very selectively [1][3].
  • Common outcome: about 4 to 6 months for many IT professionals in 2026 [1][2].
  • Slower cases: 8 months or more if you are aiming for remote roles, a narrow niche, or senior positions with very few openings [4][5].

Why it takes longer

  • ATS filtering and high application volumes mean many strong candidates never reach a recruiter [1].
  • Tech layoffs have increased the supply of experienced applicants, so competition is tighter than in a normal year [6][7].
  • Employers are hiring more cautiously and often want people who can contribute immediately with minimal ramp-up [8][9].

What affects your speed

  • Seniority matters: mid-level specialists often move faster than generalists because they can show clear value [3][1].
  • Location and flexibility matter: being open to onsite or hybrid roles can shorten the search compared with insisting on fully remote work [4].
  • Targeting matters: referred candidates and focused applications usually outperform broad mass-applying [1].

Practical expectation

For someone with solid IT experience, a good planning assumption in 2026 is three to six months, with a faster result possible if you have in-demand cloud, security, observability, or platform skills and a strong network [8][9][1]. If your profile is broader or your target is very specific, plan for the search to last longer and budget accordingly [2][5]

What strategies cut IT job search to under 3 months after layoff

To get under 3 months, the winning pattern is: target fewer roles, use warm introductions, tailor aggressively, and move fast in the first 2 weeks [1][2]. The fastest recoveries are not from spraying applications everywhere; they come from building a shortlist of target companies, speaking to people inside them, and getting referred before the role is crowded [1][6].

What works best

  • Build a target list of about 10 to 15 companies and focus on them hard rather than applying broadly [1].
  • Reach out to 3 people at each target company, ideally future teammates or adjacent peers rather than only recruiters [1].
  • Ask for short, specific conversations, then follow up every 2 to 3 weeks with something useful or relevant [1].
  • Keep your CV tightly matched to each role so it is easy to read and directly aligned with the posting [3][7].

Speed levers

  • Apply in the first 24 to 72 hours after a role is posted, when fewer candidates have piled in [1].
  • Use on-site or hybrid options if you can tolerate them, because sticking to fully remote roles can slow the search [1].
  • Stay in your current lane unless a pivot is truly justified; searches are faster when you sell proven experience rather than a brand-new direction [1].
  • Add contract, interim, or freelance work as a bridge if the market is slow; that keeps income coming and preserves momentum [2][9].

First 30 days

  • Days 1 to 3: fix CV, LinkedIn, references, and a target list [2].
  • Days 4 to 10: start outreach and referrals before spending heavy time on applications [1][6].
  • Days 10 to 30: run parallel tracks of networking, direct applications, and interview prep so you are not waiting on any one channel [2][3].
  • Keep a tracker so you can see which companies and contacts actually produce interviews [2].

For IT roles

The best odds of getting under 3 months are in higher-demand areas like cloud, security, data, platform engineering, and AI-adjacent infrastructure work [11][12]. Generalist support or commoditised roles usually take longer, so narrowing your pitch to scarce, business-critical skills matters more than ever [11][13]. In IT, referrals and a very specific value proposition often beat raw application volume [1][6].

Simple rule

If you want a sub-3-month outcome, think in terms of 10 target firms, 30 meaningful contacts, 2 tailored applications per day, and interview prep from day one [1][2][3]. That combination is much more likely to produce momentum than waiting for job boards to do the work [1][7].

How to use AI to help job searching

AI can help most if you use it to reduce admin, improve targeting, and sharpen your pitch rather than to mass-apply for roles. The biggest wins are tailoring CVs to each role, drafting outreach messages, organizing applications, and preparing for interviews faster [1][2][5].

Best uses

  • Tailor your CV to a job description by extracting keywords and matching your experience to the role [1][4][6].
  • Draft cover letters and recruiter messages quickly, then edit them so they sound like you [1][5].
  • Build a shortlist of target companies and roles from your skills, location, and preferences [1][7].
  • Track applications, follow-ups, interview dates, and contacts in one place [3][5].
  • Prepare for interviews with role-specific questions, mock answers, and STAR-story prompts [2][10].

A good workflow

  1. Paste the job description into AI and ask for the top skills, likely screening keywords, and gaps in your CV [1][4].
  2. Ask it to rewrite your summary and bullets around measurable outcomes, not responsibilities [2][10].
  3. Generate a tailored outreach note for a hiring manager, recruiter, or employee referral contact [1][2].
  4. Use AI to turn your notes into a cleaner application tracker and follow-up plan [3][5].
  5. Before interviews, ask for likely technical and behavioral questions based on the role and company [2][10].

What works especially well in IT

For IT roles, AI is most useful when you use it to map your experience to specific stacks and outcomes, such as cloud migration, observability, DevOps, security, data engineering, or platform reliability [1][4]. It can help you turn broad experience into stronger role-specific language, which matters a lot when recruiters are filtering for exact keywords [1][4]. It is also helpful for finding adjacent roles you may not have considered, especially if you want to pivot within infrastructure or operations [7][10].

What not to do

  • Do not send AI-written applications without editing them for accuracy and voice [1][6].
  • Do not rely on AI scores alone; they can miss context or overrate generic keyword stuffing [8][5].
  • Do not use it to invent experience, certifications, or achievements [10].
  • Do not mass-apply just because AI makes it easy; the best results still come from targeted roles and real networking [11][2].

Simple prompt pattern

A strong prompt is: “Here is my CV and this job description. Identify missing keywords, rewrite my summary for this role, suggest 5 stronger bullets, and draft a short recruiter message.” That gives you a focused output instead of a generic blob [1][4][10].

What is AI, AGI and ASI?

We know what AI is but what is AGI and ASI?

AI refers to machines performing tasks that typically require human intelligence. AGI is AI with general, human‑level cognitive abilities across a wide range of tasks. ASI is a hypothetical AI that surpasses human intelligence in virtually all domains.

Overview

  • AI or ANI (Artificial Narrow Intelligence): specialized AI that excels at specific tasks (e.g., image recognition, playing chess, language translation). It’s the most common form in use today.
  • AGI (Artificial General Intelligence): systems with broad, human‑level capabilities—understanding, learning, reasoning, and applying knowledge across many domains.
  • ASI (Artificial Superintelligence): a level of intelligence that greatly exceeds human capabilities in all areas.

Scope of tasks

  • ANI: narrow scope, task-specific
  • AGI: wide, adaptable reasoning across domains
  • ASI: superior performance in everything, including creativity and problem‑solving
  • Learning and adaptability:
    • ANI: learns within fixed parameters and datasets
    • AGI: can learn from diverse experiences and transfer knowledge
    • ASI: continuously self‑improves beyond human constraints
  • Current state and timelines
    • ANI is ubiquitous today, powering search, assistants, recommendations, and more.
    • AGI remains aspirational; estimates vary widely among researchers, with no consensus on when or even if it will be achieved.
    • ASI is speculative science fiction at present; most experts agree it would require breakthroughs beyond AGI.
  • Potential implications
    • Economic and labor impacts: automation of complex tasks could shift job roles and demand new skills.
    • Safety and governance: AGI/ASI would raise significant ethical, safety, and governance questions, including alignment with human values.
    • Research and science: AGI could accelerate discovery across fields, from medicine to physics.
  • Common misconceptions
    • AGI does not imply immediate, conscious machines with emotions; it implies broad cognitive capabilities.
    • ASI does not mean instant, uncontrollable intelligent beings; it depends on many speculative breakthroughs and safety frameworks.

Summary

  • ANI: specialized AI for specific tasks
  • AGI: human‑level general intelligence across tasks
  • ASI: intelligence far surpassing human capabilities

How do AGI and ASI differ in capabilities

AGI is defined as AI that can match human-level intelligence across many domains, while ASI is a hypothetical future AI that would far surpass the best human minds in virtually all areas of cognition. Both are more capable than today’s narrow AI, but ASI adds superhuman scale, speed, and depth along with the ability to improve itself far beyond human limits.softbank+3

Core capability difference

  • AGI: Human-level performance on most intellectual tasks, including learning, reasoning, planning, and adapting across domains, similar to a broadly educated person.wikipedia+1
  • ASI: Superhuman performance in essentially every intellectual task, including science, strategy, creativity, and long-term planning, not just faster computation.moontechnolabs+2

In short, AGI aims to do what humans can do; ASI aims to do far more than humans can do, in both breadth and depth.netguru+1

Learning and self‑improvement

  • AGI: Can learn from diverse data and experiences, transfer knowledge between domains, and adapt to new tasks, but its self‑improvement is still constrained by design and human oversight.softbank+2
  • ASI: Typically defined as recursively self‑improving—able to redesign its own algorithms, generate its own training data, and continually increase its capabilities without direct human guidance.creolestudios+2

This recursive self‑improvement is a key reason ASI is often linked to the “intelligence explosion” or technological singularity.netguru+1

Scope and problem‑solving

  • AGI: Expected to handle any task a human knowledge worker could, from scientific research to teaching, software engineering, or policy analysis, with strong but roughly human‑comparable judgment.kanerika+2
  • ASI: Would solve problems beyond human comprehension, discover patterns humans cannot see, and generate new scientific theories or technologies at a pace and sophistication that humans could not match.moontechnolabs+2

Where AGI might collaborate with humans as a peer, ASI would function more like an extremely powerful, alien-level expert.

Table: AGI vs ASI capabilities

AspectAGI (General)ASI (Superintelligence)
Intelligence levelRoughly human-level across many domains.wikipedia+2Vastly beyond the best human experts in all domains.softbank+2
Learning styleLearns and generalizes across tasks like a human.softbank+2Creates its own learning strategies and data; recursively self‑improving.softbank+1
Problem‑solvingSolves most tasks humans can, with high flexibility.moontechnolabs+1Cracks problems humans cannot even fully specify or understand.softbank+2
Speed and scaleSimilar cognitive range to humans, but faster on some tasks.moontechnolabsOperates at machine speed on global, massive‑scale data streams.moontechnolabs+1
Novel discoveryCan assist and sometimes lead human‑level research.moontechnolabs+1Drives radical breakthroughs in science, technology, and strategy autonomously.moontechnolabs+1
Self‑modificationLimited or controlled self‑improvement.softbank+1Open‑ended self‑modification and optimization of its own code and goals (in many visions).softbank+2

Status today

  • AGI: Still not achieved; current systems are closer to advanced narrow or “weak” forms of generality, and AGI remains an aspirational research target.viso+2
  • ASI: Entirely speculative at present, dependent on first reaching robust AGI and then enabling powerful recursive self‑improvement.kanerika+2

If you want to go deeper, the next useful dimension is risk and governance differences: why AGI raises big alignment questions, and why ASI raises civilization‑level control questions.

What are real world uses of AGI today

There are no widely accepted, deployed AGI systems in the strict “human‑level general intelligence” sense today; what exists are advanced narrow or “proto‑AGI” models being used in many real‑world applications. Most concrete use cases branded as “AGI” are actually powerful generative or multimodal AI applied across multiple tasks rather than true general intelligence.sidetool+3

Important clarification

  • No consensus that AGI exists yet: Surveys and 2025 overviews still describe AGI as a future milestone that requires breakthroughs in general problem‑solving and knowledge transfer.kanerika+1
  • Marketing vs reality: Many articles and vendors use “AGI” aspirationally for advanced models, but technical write‑ups usually treat them as steps toward AGI, not fully general minds.techrevolt+2

So “real‑world uses of AGI” today are better described as uses of advanced AI that show some generality but remain below true AGI.

Advanced AI uses often framed as AGI‑like

  • Autonomous task agents in business: End‑to‑end agents can plan, call tools, and complete tasks such as drafting contracts, generating code, and running simple workflows with minimal oversight.cloud.google+1
  • Cross‑domain copilots: Enterprise copilots (e.g., Google’s Gemini‑based Workspace assistants) summarize mail, generate documents, analyze sheets, and answer questions over internal knowledge, acting as a general knowledge worker assistant within one organization.cloud.google

These systems show broader versatility than classic narrow AI but still lack robust, human‑level general understanding and autonomy.scientificamerican+1

Examples often cited as “AGI use cases”

  • Healthcare decision support: Systems like IBM’s Watson Health analyze patient records and large medical literature corpora to suggest treatment options and support oncology decisions, sometimes cited as early AGI‑style reasoning in a specialized domain.ibm+1
  • Industrial optimization and predictive maintenance: Large AI systems monitor sensor data, detect anomalies, and optimize complex production lines or logistics networks, adjusting schedules and predicting failures in near real time.ibm+1

These fit “powerful domain‑general within a vertical” more than “fully general intelligence.”

Multimodal “proto‑AGI” applications

  • Customer service and sales: Multimodal models that understand text, voice, and sometimes images are used for richer customer support, sentiment‑aware responses, and context‑aware recommendations.weforum+2
  • Knowledge and research assistance: Large models are used for legal analysis, financial modeling aid, and code generation, giving human experts faster access to cross‑domain information and draft reasoning.sidetool+1

Some commentators argue these systems are “chasing and occasionally catching” AGI capabilities in narrow slices, but they are not yet considered true AGI by most researchers.techrevolt+2

How to interpret “AGI in the real world” today

  • In research and journalism, “AGI use cases” usually mean: “Here’s what AGI would do in healthcare, manufacturing, logistics, etc.” rather than deployments of a confirmed AGI.kanerika+1
  • In industry and marketing, the term often labels very capable general‑purpose assistants or agents built on frontier generative models that are still narrow compared with human general intelligence.sidetool+2

If you want, the next step can be mapping specific “AGI‑like” patterns (tool‑using agents, long‑horizon planners, cross‑modal perception) to concrete systems you’re likely to encounter in enterprise or developer workflows.

What is the progress to ASI – how near are we?

Artificial superintelligence (ASI) does not exist today, and expert views on when or whether it will emerge vary enormously, from “possibly this century” to “not with current approaches, if ever.” The only honest summary is that progress is rapid toward stronger AI systems, but ASI remains speculative and timeline estimates are highly uncertain.theaireport+2

Where we are now

  • Current frontier models show impressive gains in coding, reasoning benchmarks, and multimodal tasks, and some can already write and improve their own software in constrained settings.scientificamerican+1
  • However, these systems still fail in robust general reasoning, long‑horizon planning, and reliable self‑improvement—capabilities most definitions of AGI or ASI would require.hai.stanford+1

So the field is in a phase of “advanced but brittle” systems, not in sight of true superintelligence.

Expert timelines and disagreement

  • Recent expert surveys put the median 50% probability for AGI sometime between roughly 2040 and 2061, though a vocal minority predicts much earlier dates.forbes+1
  • Some commentators then assume a relatively fast transition from AGI to ASI, with speculative ranges from a few years to several decades; one analysis frames a “useful estimate” of about 10 years from AGI to ASI.theaireport+1

At the same time, many experts (around three‑quarters in one survey) think simply scaling today’s techniques is not enough for AGI, let alone ASI.theaireport

Indicators of movement toward ASI

  • Rapid capability scaling: Benchmarks show steep year‑on‑year jumps in areas like multimodal understanding and software engineering, plus large drops in compute cost for a given capability level.hai.stanford+1
  • Early self‑improvement: Leading models can already optimize their own code or assist in model design, raising the question of whether this could ever “snowball” into open‑ended self‑improvement.scientificamerican+1

These trends are why some researchers argue we might be seeing the first steps toward systems that could one day underpin superintelligence, even if that point is still far away.scientificamerican

Cautious views and roadmaps

  • Researchers like Yann LeCun emphasize that even ambitious roadmaps aiming at “advanced machine intelligence” still expect only animal‑level world understanding (e.g., “cat or rat” level) prototypes within a few years, not human‑ or super‑human minds.eetimes
  • Groups branding themselves around “ASI” publish roadmaps that explicitly admit human‑level AGI has not yet been reached and present ASI as a longer‑term goal after major research advances.superintelligence

These more conservative views highlight that achieving robust world models, reasoning, and planning is still an unsolved problem.

How near are we, realistically?

  • No ASI today: There are no systems that fit the common definition of intelligence vastly beyond the best human experts across all domains.theaireport+1
  • AGI first, ASI later (if at all): Most structured analyses treat ASI as contingent on first achieving AGI, with a further uncertain period—possibly years to decades—before any superintelligent phase.forbes+1

Taken together, a reasonable reading of the evidence is: ASI is not imminent, serious people disagree on whether it is this‑century or ever, and the main near‑term focus is on managing rapidly strengthening but still fallible AI systems rather than on a concrete ASI arrival date.hai.stanford+2

AI Circular Financing

We’ve all seen the above image when Oracle started making waves about their AI contract to provide OpenAI compute power at their Abilene DC with tonnes of NVIDIA GPUs – well there’s now news about circular financing fraud involving the big AI players…

We look at what this article is saying and assess whether a postive circular financing fraud case will be the pin that bursts the AI bubble: https://substack.com/home/post/p-179453867

The Algorithm That Detected a $610 Billion Fraud: How Machine Intelligence Exposed the AI Industry’s Circular Financing Scheme

The article describes how advanced trading algorithms and machine intelligence detected an unprecedented $610 billion circular financing fraud involving major AI industry players, including Nvidia, Microsoft, OpenAI, Oracle, and xAI. This scheme involved companies repeatedly recycling the same money through chip purchases, cloud credits, and investments, allowing them to recognize fictitious revenue and inflate financial numbers without actual cash being exchanged.binance+3

What Happened

On November 20, 2025, AI-driven trading systems analyzed Nvidia’s earnings and discovered massive inconsistencies: billions in accounts receivable remained unpaid, and customers were taking longer to pay, resulting in aging bills and swelling inventories. The investigation revealed funds circulated between firms—such as Nvidia investing in xAI, which would use pushed loans to buy Nvidia chips, while Microsoft and Oracle funneled cloud credits through OpenAI—getting booked as revenue at multiple companies, but with no real payment behind the transactions.instagram+2

Impact on Financial Markets

Once the fraud was flagged, Nvidia’s stock, which had surged by $130 billion after earnings, quickly crashed as algorithms triggered selloffs faster than human analysts could respond. Bitcoin, used by many AI startups as loan collateral, also plunged as market confidence evaporated, with experts predicting further Bitcoin sell-offs as defaults increase.stockmarketanalysis04.substack+2

Industry Reactions and Fallout

Executives—including those from Airbnb and OpenAI—acknowledged the reliance on artificial “vibe revenue,” admitting the industry’s expected future profits were vastly overstated, and many projects would never deliver the anticipated returns. Famous investors quickly exited positions, and credit agencies are expected to downgrade the sector as payment cycles lengthen and the circular financing unwinds.news.futunn+2

Warnings and Criticism

Analysis of the reports cautioned that while some claims align with public financial data, other details appear speculative or exaggerated, casting doubt on the scale and specifics of the fraud presented in the article. However, the exposure of circular accounting practices has sparked renewed scrutiny of AI industry finances and the systemic risks they pose to both tech and broader financial markets.x+1

What would be the market implications if the fraud is confirmed?

If the $610 billion AI industry fraud is confirmed, the market implications would likely be severe and multifaceted:

  • Stock prices of the companies involved would likely experience sharp declines, similar to historic fraud cases where stock prices dropped significantly upon fraud discovery and investigation announcements. For example, firms have seen cumulative abnormal returns fall by 29% at fraud discovery and further 8% after regulatory investigation announcements, particularly when fraud involves revenue recognition or asset overstatement.nacva+1
  • Investor trust across the AI and related tech sectors would erode drastically, causing not only plummeting stock demand for the implicated companies but also collateral damage to wider market confidence. This loss of trust can depress sales, revenue, and overall financial performance beyond the direct fraud perpetrators.financemagnates
  • Increased regulatory scrutiny and enforcement actions would follow, including potential SEC investigations, fines, and legal consequences for perpetrators, shaking market stability and inviting tighter oversight on AI industry financial reporting.fraud+1
  • The revelation of such a large-scale circular financing scheme would raise concerns over information opacity and systemic risks in the AI sector and adjacent markets. This could raise the risk of future stock price crashes and long-term market volatility due to shaken investor confidence and greater caution toward AI-related investments.sciencedirect
  • Credit downgrades and withdrawal of investment capital across affected companies and startups would likely accelerate, hampering innovation financing and causing a sector-wide funding crunch.binance

Together, these effects imply a potential market shock comparable to major corporate fraud scandals, with profound short-to-medium term disruptions in AI industry valuations and investor behavior, and increased regulatory and due diligence focus on AI sector financial transparency.nacva+2

​Which peer firms and sectors would face contagion risk?

If the $610 billion AI industry fraud is confirmed, the peer firms and sectors facing contagion risk would include:

  • Other major AI and technology firms interconnected via investments, chip sales, and cloud services, such as Microsoft, Oracle, OpenAI, and Nvidia’s competitors. Because the fraud involved circular financing among these players, their financial health and market valuations would be at risk of contagion through direct exposures and reputational damage.
  • The broader technology sector, including semiconductor manufacturers, cloud infrastructure providers, and AI startups, which depend heavily on investment flows and customer confidence linked to the implicated firms. Funding shortages and tighter credit conditions would ripple across these linked companies.
  • Financial institutions such as banks, investment funds, hedge funds, and insurers holding equity, debt, or derivatives tied to AI and tech firms. Market shocks from the fraud revelation could trigger liquidity strains, asset fire sales, and valuation drops across these financial intermediaries.
  • Sectors with strong dependence on AI-driven growth or technology investments, including software service companies, cloud computing, and innovative tech hardware industries, which would face heightened volatility and funding risks from cascading investor withdrawals.

In summary, the contagion risk would span a network of interlinked AI/tech firms, their financial backers (banks, funds, insurers), and adjacent sectors relying on AI-driven growth and innovation capital. These risks would be amplified by dense financial linkages, cross-holdings, and overlapping investor exposure, potentially escalating systemic risks and market instability across these interwoven sectors.ecb.europa+2

How likely is this a fraud?

The likelihood of this $610 billion AI industry fraud being real appears mixed, with some evidence supporting concerns but also significant skepticism about the full scale and details.

  • Support for fraud possibility: Automated algorithms detected unusual circular financing flows and discrepancies in accounts receivable, payments, and inventory levels among key players like Nvidia, OpenAI, Microsoft, and Oracle. Executives and industry insiders have acknowledged artificial revenue recognition practices, “vibe revenue,” and unsustainable burn rates, fueling suspicions of inflated financial reporting and reliance on self-referential capital cycling.reddit+2
  • Skepticism: Several data points in the original fraud claim articles do not fully align with official financial statements or audited reports. For example, some growth rates and cash conversion claims appear unverified or exaggerated, and robust revenue projections remain in place for major firms like Nvidia. Critiques suggest parts of the narrative may be speculative or amplified to support short-selling or competitive narratives.news.futunn+1
  • Market actions such as large insider sales, stock volatility, and algorithmic trading responses reflect real concerns but do not alone prove deliberate fraud.

Overall, while there is credible reason for regulatory scrutiny and investor caution, it is not yet conclusively proven that a $610 billion-scale fraud has occurred as portrayed. Financial markets and regulators will need to conduct thorough audits and investigations to confirm or refute the full extent of circular financing and unverifiable revenues in the AI sector.reddit+1

​What additional evidence would confirm or refute the fraud claim?

To confirm or refute the $610 billion AI industry fraud claim, the following additional evidence would be critical:

  • Detailed audit reports and forensic accounting analyses that reveal whether reported revenues and assets align with actual cash flows, eliminating circular financing and fictitious transactions. Clear documentation of payment trails and verification of accounts receivable would be vital.
  • Access to internal communications, contracts, emails, and loan agreements showing intent, knowledge, or concealment of fraudulent practices among the companies involved. These would demonstrate deliberate misrepresentation rather than innocent errors.
  • Independent verification of inventories, customer payments, and cloud credit usage that purport to back financial figures, to detect any discrepancies or inflated valuations.
  • Testimonies or whistleblower accounts from insiders who can provide direct evidence of deception, revenue manipulation, or artificially inflated accounts.
  • Examination of algorithmic data and trading patterns that initially flagged the fraud, corroborated by hard financial evidence to remove speculation.
  • Legal scrutiny and regulatory investigations with power to compel document production and testimony, leading to formal findings of fraud based on a balance of probabilities.

These types of evidence collectively would build a persuasive case by proving the key elements of fraud: false statements of fact, knowledge of falsity, intent to deceive, reliance on the falsehood by others, and resulting financial damage.lincolnandrowe+3

​Due to the big players involved, how like is this to be recognised as fraud?

Given the involvement of major public companies such as Nvidia, Microsoft, Oracle, and OpenAI, it is relatively less likely that the alleged $610 billion circular financing scheme would go unrecognized or uninvestigated by regulators, auditors, and independent analysts. These firms operate under strict financial reporting standards, are subject to continuous market scrutiny, and have audits performed by reputable firms. Nvidia, for example, is audited by PwC and rated highly by credit agencies, and recently reported strong revenue guidance despite allegations.news.futunn+1

However, the complexity and scale of circular financing and vendor financing deals—spread across multiple firms and private entities like OpenAI—and the extraordinary sums involved create challenges in transparency. Given the high stakes, any material misstatements or abuses would likely draw regulatory attention from bodies like the SEC, and investigations would follow, as evidenced by past fraud cases in tech and finance.peterwildeford.substack

Market reactions such as rapid stock sell-offs, insider selling, and investor withdrawals suggest heightened suspicion and the potential for increased scrutiny. But actual recognition of fraud depends on whether clear evidence emerges proving intent and knowing misrepresentation, beyond complex but possibly legitimate financial engineering.substack+1

Therefore, while these large players’ reputations, audit requirements, and regulator oversight make outright fraud recognition more probable than in less regulated sectors, proving the full scope involves thorough forensic audits and investigations. The high-profile nature ensures any verified fraud will be highly publicized and swiftly acted upon.tomtunguz+3

​The public seems to be able to see this circular financing already, but there has been no obvious denies or evidence until now?

The public’s ability to observe the circular financing is likely due to the following:

  • Many of the so-called circular financing deals in AI are part of legitimate strategic partnerships, vendor financing, and ecosystem-building efforts. Companies like Nvidia, OpenAI, Microsoft, AMD, Broadcom, and others are deeply intertwined financially and operationally, investing heavily in each other through chip sales, equity stakes, and cloud infrastructure deals. This creates a “circular economy” that keeps capital flowing within a compact network, which is visible in market disclosures and company announcements but can appear as circular funding to outside observers.forbes+2
  • These circular arrangements often reflect business models designed to ensure long-term demand for AI hardware and cloud services while fostering innovation through close alliances. While they may inflate revenue temporarily on paper by recognizing sales within the network, such practices are also common vendor financing strategies used in many industries historically and are not necessarily fraudulent if transparently disclosed and compliant with accounting standards.reddit+1
  • Until now, there has been no obvious denial or strong rebuttal partly because companies benefit from these arrangements to drive AI ecosystem growth and financial performance, and the accounting treatments remain within legal frameworks. The lack of clear external evidence against them also means regulators and auditors have not formally challenged these practices as fraudulent yet, despite increasing scrutiny and skepticism among investors and algorithms detecting anomalies.bbc+1
  • The recent algorithmic detection of potential fraud has heightened attention and raised the stakes for transparency, but distinguishing aggressive vendor financing from deliberate fraud requires detailed forensic audit evidence and regulatory investigation.

Thus, the public is seeing signals of circular financing because it is foundational to how AI ecosystem finance currently operates, but formal fraud recognition demands deeper proof beyond visible financial flows circulating within interconnected big players.ubs+3

A(I) Risky Move?

It is no secret that I am leaving Oracle to join Nscale at the end of this month. My main aim was to avoid the Oracle RIF/layoffs and while putting the worse case scenario in place, I managed to get a role at Nscale.

Both Oracle and Nscale are hedging their bets on AI and with the current worry that the AI bubble is going to burst, with the boss of Google saying: ‘no company is going to be immune’ [BBC], I asked AI these questions…

How Oracle would suffer in an AI crash?

Oracle would experience significant financial and operational setbacks if an AI crash occurs, mainly because its recent strategy has focused heavily on AI cloud infrastructure and hyperscaler partnerships. Here’s how Oracle and Oracle Cloud Infrastructure (OCI) would suffer:

Revenue Slowdown and Underutilized Data Centers

  • Oracle expects that AI-driven cloud infrastructure could make up the majority of its revenue by 2030, with multi-billion dollar GPU investments and flagship deals supporting OpenAI, Meta, and other hyperscale model trainers.webpronews+1
  • In an AI crash, client demand for large-scale training and inference workloads would sharply decline. That means Oracle’s new data centers and GPU fleets (including hundreds of thousands of Nvidia GPUs) could become underutilized, forcing write-downs and lowering ROI.fortune

Margin Compression and Volatility

  • Oracle’s current margins on AI cloud offerings are modest (around 14% in recent quarters), and plans for improvement depend on scale and full utilization of infrastructure.bloomberg+1
  • Reduced client activity would keep margins low or negative, intensifying investor scrutiny and driving downward pressure on stock price—recent volatility already shows concern over fulfilling ambitious AI cloud predictions.cnbc+1
  • Oracle would also struggle to justify continued massive investments to shareholders if growth fails to materialize.

Exposure to Contract and Customer Risk

  • Oracle is especially exposed to a few large anchor clients (e.g. OpenAI, Microsoft), who may defer or downsize contracts if the economics of large-model training weaken.webpronews+1
  • Cancellations or renegotiations on these hyperscale deals would lead to revenue instability and possible reputational harm in the cloud sector.

Delayed Application Growth and Broader Impact

  • While Oracle’s core application business (ERP, NetSuite, traditional cloud services) offers some revenue buffer, these divisions would also be affected by a broader tech and enterprise spending slowdown typical in an AI-induced market downturn.constellationr+1
  • Expansion plans and R&D investment would likely be scaled back, slowing Oracle’s ability to compete with more insulated rivals.

Summary Table: Oracle’s Pain Points in an AI Crash

Effect AreaOracle Cloud & Business Impact
Infrastructure ROIUnderutilized data centers, GPU fleets
Margins & ProfitMargin compression, stock volatility
Contract RiskDeferred/canceled hyperscale projects
App/Enterprise SpendRevenue dampening, slower growth

In summary, Oracle would suffer via underperforming cloud assets, margin pressure, dependency on a handful of high-risk AI clients, and contraction of enterprise demand. While traditional non-AI applications offer some shelter, the scale of Oracle’s recent AI cloud bets makes it highly exposed to a major market correction.reuters+6

How Nscale would be affected by an AI crash?

Nscale would be severely affected by an AI crash due to its intense reliance on hyperscaler contracts, rapid capital deployment, and debt-financed AI infrastructure builds. Its business model is specifically geared toward serving large-scale AI model training and compute workloads, making it one of the highest-risk startups in a downturn scenario.

Revenue Freeze and Stranded AI Data Centers

  • If AI demand collapses, hyperscalers like Microsoft and OpenAI—Nscale’s anchor clients—would scale back their contracts, delay GPU purchases, or walk away from expansion phases, freezing Nscale’s main cashflow pipeline.computing
  • Data centers built for intensive model training could be left mostly idle, forcing Nscale to absorb heavy operational losses while continuing to service debt and lease obligations.

Investor and Debt Pressure

  • Nscale’s recent $1.1 billion Series B and $433 million SAFE financing involve a mix of industrial and venture capital (Aker, Nvidia, Dell, Blue Owl), and depend on optimistic growth projections and future contract fulfillment.cnbc+3
  • In a crash, these investors would likely mark down their stakes significantly, making further fundraising difficult. If cash flow dries up, Nscale could face covenant breaches, mass layoffs, and possibly insolvency.

Asset Write-Downs and Market Value Declines

  • Large inventories of Nvidia GPUs and purpose-built data centers for AI workloads could lose much of their resale value if a flood of surplus infrastructure hits the market simultaneously.computing
  • Hardware and campuses not suited for general-purpose workloads would be especially vulnerable—mirroring the fate of unused “dark fiber” and crypto-mining assets after previous speculative bubbles burst.

Strategic Pivot and Survival Pathways

  • With deep exposure to hyperscaler deals, Nscale’s best shot at survival would be pivoting toward government-backed sovereign compute projects, or repositioning under-utilized capacity for scientific, academic, or inference-driven workloads.computing
  • However, the overall scale-back in AI investment would make even these pivots challenging and may require restructuring, asset sales, or merger activity.

Summary Table: Nscale’s Crash Exposure

Pain PointCrash Impact
AI contract relianceRevenue freeze, phased-out deals, layoffs
Infrastructure assetsWrite-downs, idle data centers, GPU oversupply
Investor relationsMarkdowns, covenant risk, fundraising difficulty
Pivot opportunitiesLimited to sovereign or non-AI workloads

In summary, Nscale would be among the hardest-hit AI infrastructure startups during a crash, facing immediate cash flow problems, asset devaluation, and existential threats unless it rapidly adapts or receives protective strategic support.uktech+4

As an employee of Oracle or Nscale – who would be worst affected in an AI crash?

As an employee, Nscale staff would likely be much worse affected by an AI crash than Oracle employees. This is due to the relative scale, business diversity, and risk tolerance of the two companies:

Nscale Employees: High Risk

  • Startup Vulnerability: Nscale is a high-growth, venture-backed startup focused almost entirely on the AI cloud hyperscaler market. If an AI crash occurs, Nscale would face abrupt revenue freezes, contract withdrawals, and rapid asset devaluation, with far less financial cushion to weather losses than a large incumbent like Oracle.
  • Layoffs and Insolvency Risk: In a crash, Nscale may be forced to lay off a substantial portion of its workforce quickly, pause projects, or even enter administration if funding evaporates or debts cannot be serviced. Employees could face sudden job loss, delayed salaries, and loss of equity value, with fewer internal opportunities for reassignment.
  • Equity and Options: Most compensation upside for Nscale staff is likely tied to startup equity. A market correction would deeply devalue or wipe out employee options/shareholdings, and prospects for another lucrative exit (e.g., IPO or acquisition) would diminish drastically.

Oracle Employees: More Stability, Less Immediate Risk

  • Diversified Business: Oracle is a global enterprise with diversified revenue streams across cloud, applications, hardware, licensing, and services. Even with heavy AI investment, its overall risk is spread out, and core business lines would remain resilient.
  • Less Severe Layoffs: Large tech companies like Oracle do enact layoffs in downturns, but they tend to be more measured and spread across divisions. Employees have more options to transfer internally or leverage severance and career support programs.
  • Compensation Impact: Public company employees with RSUs or stock options would feel a drop in value, but these tend to recover after corrections, and overall job security remains higher due to Oracle’s operational scale and financial reserves.

Summary Table

FactorNscale Employee ImpactOracle Employee Impact
Layoff riskExtremely high, abruptModerate, measured
Job securityLowHigh
Equity/option valueLikely wiped outLower, but more resilient
Internal mobilityMinimalSignificant
Corporate backup/supportMinimal/noneSubstantial

In summary:
Nscale employees are far more exposed to job loss, equity wipeout, and uncertainty during an AI crash, while Oracle employees—even in affected AI-focused roles—would generally be shielded by the company’s size, global reach, and business diversity.

What is the benefits of joining Nscale for my career over staying at Oracle

Joining Nscale instead of staying at Oracle offers a profoundly different set of career benefits and risks. Your choice should be shaped by your appetite for growth, impact, compensation profile, and risk tolerance.

Career Acceleration and Impact

  • Rapid Responsibility & Visibility: At Nscale, a high-growth AI infrastructure startup, you’re likely to hold broader responsibilities, have a direct line to executives, and see tangible impact from your contributions—versus a more specialized, hierarchical environment at Oracle.linkedin+1
  • Cutting-Edge Technology Exposure: Nscale’s business is centered around the latest in AI hyperscale, GPU clusters, and sovereign compute, putting you at the heart of Europe’s most ambitious AI and cloud projects. This hands-on work accelerates technical learning and keeps you current with elite tooling.linkedin
  • Networking Opportunity: Being at a much talked-about scaleup, you’ll interact with top-tier industry partners (Microsoft, Nvidia, blue-chip VCs), as well as the fast-moving European AI ecosystem, which can boost professional reputation and future opportunities.cnbc+1

Compensation and Equity

  • Equity Upside: As with most startups, joining Nscale comes with employee equity—potentially life-changing financial upside if the company’s rapid growth leads to a successful IPO or acquisition. At Oracle, equity is usually granted as RSUs whose upside, while substantial, is tied to a mature company’s slower growth trajectory.
  • Compensation Mix Changes: Startup salary may be lower or matched to Oracle’s, but total comp potential is usually tied more to company performance and exit events.

Innovation, Learning and Influence

  • Innovation Environment: Nscale’s pace and hunger for innovation generally surpass what’s possible at a large, diversified incumbent like Oracle. Progression can be much faster for high performers at a startup, and you have more say in technical and strategic direction.
  • Learning Curve: The breadth of challenges—from scaling infra to winning deals and iterating products for hyperscaler clients—means you’ll develop skills quickly that are valuable for future founding roles, CTO posts, or senior technical leadership.

Risks, Trade-Offs and Stability

  • Job and Financial Risk: As shown in your AI crash scenario queries, Nscale is exposed to sector volatility. Layoff, equity wipeout, or strategic pivots are real risks in a downturn, while Oracle offers much greater job security and stable, long-term career pathways.
  • Work-Life Balance: Startups typically expect longer hours and intense focus, often trading stability for the chance at rapid advancement and financial gain.

Summary Table: Nscale vs Oracle Career Benefits

Career FactorNscale (Startup)Oracle (Enterprise)
Career accelerationVery high—fast progression, impactSlower, more structured
Tech exposureLatest/bleeding-edge AI/GPU workSolid, broad enterprise stack
Compensation upsideHigh equity, riskyStable salary, predictable RSUs
Job securityLow—volatile, high layoff riskHigh, robust to market downturns
InfluenceLarge, direct access to leadershipSmaller, part of larger organization
Work-life balanceOften demandingGenerally more moderate
NetworkingTop-tier VCs, hyperscalers, foundersGlobal enterprise contacts

In summary: Nscale offers a “rocket ship” growth and learning environment for someone hungry for ownership, innovation, and possible equity windfall—but this comes with higher risk and volatility. Oracle provides greater long-term job security, measured career progression, and reliable compensation, but less potential upside and less direct influence. The right choice depends on your risk appetite, ambition, and career goals.uktech+2