Exploring how AI reshapes the way humans think, learn, and create.

Learning With AI

Will AI Replace IT Operations Jobs? Where Human Judgment Still Matters

AI is absorbing standardized work in IT operations and predictive maintenance. See what it can't replace, plus a three-question test for your own tasks.

Intelligenr Research Team·
AI in IT operationsAIOps skills roadmapAI in predictive maintenancepredictive maintenance vs. maintenance technicians

Editorial Note: This article examines where AI is changing operations work, and where human judgment, accountability, and physical-world constraints still shape the boundary. The examples are illustrative composites; claims about products, regulations, and survey findings are based on publicly available sources.


AI in IT Operations and Predictive Maintenance: Two Rooms Where This Is Already Happening

At 2 a.m., an on-call engineer at a regional bank in North America gets woken by his phone. An alert: latency on a payment path has climbed. He spends twenty minutes confirming it isn't the same known nightly jitter he's seen before, restarts an instance per the runbook, latency comes back down, he goes back to sleep. In the morning the record lands in a spreadsheet and nobody looks at it again.

That same year, in a plant in the American Midwest, a maintenance technician presses the back of his hand against an air compressor housing during a walkaround, decides it feels hotter than last week, and writes a note in the work order.

Both of those moments depend on some of the same things: a person being present, a person having experience, and a person being willing to act when the system cannot resolve the situation on its own. Over the past year and change, that category of work has been getting rewritten — not deleted, rewritten.

This article starts from those two rooms. A bank and a manufacturer, different industries, different stacks, different regulators — yet the work AI has taken over looks nearly identical in both, and so does the work it hasn't. That resemblance is more worth explaining than any opinion about whether AI replaces operations.

AI in IT Operations: How a Regional Bank's On-Call Rotation Actually Changed

The bank has roughly 3,000 employees. Its architecture is mixed: core accounting still runs on legacy hosts, most of the business sits in containers, and there are around 400 services. On-call is 24/7, six people in rotation, each pulling roughly one full week a month.

Before the change, a typical night produced two to three hundred alerts. Fewer than a fifth of them required a human to act; the rest got suppressed by correlation rules and silence windows — never cleanly, though, and a few always slipped through needing manual confirmation. Most of the engineer's night went into determining whether a given alert was the same known jitter again. During the day, L1 and L2 tickets were dominated by repetitive remediation: restarts, scaling, cache clears, running a documented procedure end to end. The two most senior people on the team had their time shredded by recurring firefighting, leaving very little uninterrupted stretch for architecture work. That's a common condition in operations teams, not something specific to this bank.

The first change was alert correlation. Unsexy, immediate payoff: noise dropped, and people could finally see the actual events. Next came investigation agents — a category that reached commercial availability in late 2025. Datadog shipped Bits AI SRE as its first generally available AI agent at the end of that year, after a limited release earlier in the year. When an alert fires, the agent reads the runbooks, pulls telemetry, walks the service topology, separates signal from noise, verifies its own conclusions, and pushes an evidence-backed summary into the collaboration tool before the on-call engineer has even logged in.

The change in day-to-day experience is concrete. What's on the desk at handover is no longer a bare alert; it's a document saying "I think this is X, here are the metrics and the change that support it, start by checking Y." The engineer's job shifts from investigating from zero to reviewing a hypothesis that already has a shape.

The parts that didn't work out matter more than the parts that did.

Autonomous remediation never got switched on. Technically it was feasible; the blocker sat elsewhere. Core accounting paths, databases, cross-AZ traffic shifts all require a named owner and rollback authorization under the bank's change management. The accountability structure for "agent acts autonomously, humans audit afterward" simply hadn't been written by most organizations. One team tried enabling it on low-risk services, took one bad call, and pulled the permission back. That's governance catching up, and in many organizations, governance moves much more slowly than product development.

On cross-system coupled failures, the agent produced conclusions that were wrong in direction and plausible in detail. Those failures are hard precisely because there's no clean historical sample: a network blip, stacked on a deploy, stacked on a connection pool setting, with the chain crossing four or five systems. Hallucination is more dangerous in operations than in most contexts, because invented config keys, parameter names, and metric names all look real, and someone who doesn't know the system well won't catch it on the spot. The cost lands on human trust and human hours, not on compute.

One more thing gets consistently underestimated: an agent's conclusion quality is capped by the quality of the data it can see. This bank's CMDB — the configuration database recording every host, service, and dependency — had a meaningful share of stale entries and an incomplete dependency graph. The same model wired to clean data and wired to dirty data performs like two different products.

A year later, headcount was unchanged. What changed was the content of the rotation: L1 repetitive remediation dropped sharply, and "review the agent's conclusion plus handle what the agent can't" became the main body of on-call work.

There's a market-side signal pointing the same direction. PagerDuty disclosed seat compression among large enterprise customers on a recent earnings call, attributing part of it to AI tooling changing seat demand, and moved toward usage-based rather than per-seat pricing. That's a vendor's own disclosure and shouldn't be read as an industry average — but a pricing unit migrating from heads to usage tells you something about where head demand is going.

Predictive Maintenance in Manufacturing: What AI Takes Over and What It Can't

The manufacturer runs six plants. Critical assets include CNC spindles, air compressors, injection molding hydraulic systems, and conveyor drives. The prior model was time-based maintenance plus reactive repair: three reliability engineers at the corporate level, with maintenance crews at each plant.

The instrumentation details determine whether a program like this succeeds, and they're also what makes it credible to anyone who's done the work. Vibration is sampled at high frequency — with sampling rates in the tens of kilohertz used in some spindle and rotating-equipment applications — while evaluation criteria may reference standards such as ISO 20816 for machine-vibration measurement and severity. Bearing defects, imbalance, and misalignment can produce characteristic spectral signatures, but diagnosing those conditions requires methods and domain knowledge beyond the ISO 20816 severity criteria alone. Temperature is sampled at a much lower frequency, often one point every minute or five, good for trends and less useful for fast transients. Motor current can be analyzed using MCSA-type methods, where broken rotor bars and eccentricity may produce identifiable sidebands. The three signal types have different sampling requirements, diagnostic criteria, and false-alarm behavior. Treating them as one undifferentiated "anomaly detection" layer can obscure those differences and create avoidable failure modes.

What got absorbed is clear: meter reading and inspection log entry, trend interpretation, first-pass fault diagnosis, spare parts demand forecasting, work order dispatch and scheduling. Condition-based maintenance — replace on evidence rather than on the calendar — took over part of the scheduled work. That recovers the cost of over-maintenance, and it also moves risk from missed inspection to misclassification.

The physical world adds constraints the data center doesn't have, and that's the key difference from the bank.

Safety constraints are hard constraints. Lockout/tagout — de-energizing and physically locking the energy source before service so nobody can re-energize it — confined space entry, energy isolation: these procedures exist precisely because humans can be wrong. They don't get relaxed as model accuracy improves. If anything, a model recommending a shutdown adds a human confirmation step.

Shutdown decisions are economic. The cost of a four-hour line stoppage, the delivery window on an order, spare parts lead time, downstream customer impact — none of those variables live in a sensor. A model can say "failure probability for this asset is rising over the next 72 hours." Whether to stop now is a judgment with commercial consequences attached.

Non-routine field work still runs on human hands and eyes. The failure mode the model has never seen, the sensor caked in oil, the piping rerouted temporarily during a modification — unstructured anomalies in a physical environment are still handled by a technician.

Technicians didn't disappear. The "read the gauge plus go by feel" portion compressed, and "verify the model's call, execute the physical work, hold the safety line" became the core of the job. One observation worth noting: the technicians who diverged fastest from their peers were the ones who started recording "the model said A, I judged B, it turned out to be C." That record is the raw material for every improvement that follows.

Intent Layer vs. Execution Layer: Where AI Actually Replaces Operations Work

Put the two rooms side by side and the absorbed work sits in one layer.

The execution layer is still rules and code. Ansible batch changes, Kubernetes self-healing restarts, CI/CD deploy and rollback, PLC interlocks, protective relay trip logic — none of that got replaced by AI, because none of it ever needed understanding, only determinism. AI's role at that layer is to invoke those mechanisms.

What compressed is the intent layer. A meaningful share of an operations engineer's value used to be knowing which command to type and which gauge to read. That's now covered by describing intent in natural language and having an agent call the tools. The substance of the change is a shift in where judgment can sit: from "human judges, then executes" toward "machine investigates, human reviews," and, in some low-risk cases, toward machine-initiated or machine-executed actions under predefined guardrails.

That reframes the useful question. When people argue about AI replacing operations, the two things actually in motion are who holds the judgment and who bears the consequences of a wrong judgment. The first is migrating fast. The second is migrating slowly — and the answer to the second largely sets the boundary on the first.

Where AI Is Less Likely to Fully Replace Operations Work: Six Structural Reasons

Each of these six comes with a mechanism rather than an adjective.

1. Authorization and accountability. High-risk changes often require a defined human approval or oversight process, and that requirement is institutional rather than purely technical. The EU AI Act, as adjusted by the 2026 AI Omnibus, applies the rules for certain high-risk AI systems in areas including critical infrastructure from December 2, 2027, while high-risk AI systems embedded in regulated products have an extended timeline to August 2, 2028. The requirements include human oversight, logging, documentation, and traceability. The legal responsibilities fall on providers and deployers rather than on the model as an independent legal actor. This boundary doesn't dissolve as models improve; it moves as accountability frameworks evolve.

2. Causal reasoning about unfamiliar failures. Models are strong where a similar pattern exists in history. Hard failures often have no clean sample and require hypothesis-and-elimination, a process that generates new information as it runs. Combined with the hallucination profile described above, a plausible-but-wrong conclusion costs more than an obviously wrong one, because it has to be disproven first.

3. Architecture and trade-offs. How many nines an SLO should promise, what tier of disaster recovery to build, how much availability two million dollars buys — these have no optimum, only a solution matched to the business's loss tolerance. A model can produce options, parameter comparisons, and historical references. The value judgment stays with a person.

4. Non-routine physical work. Racking and de-racking, cabling, power cutover, on-site anomaly handling. Inspection robots and manipulators are advancing, but unstructured physical environments still need human hands in the near term, and they're governed by safety procedures on top of that.

5. Coordination and negotiation. Cross-team mobilization during an incident, negotiating a degraded-mode plan with the business, assigning fault with a vendor, finding a change window. These distribute interests among parties; they're negotiation, not computation.

6. Writing the rules and holding the fallback. Who decides whether alert correlation dropped a critical event? Who turns a senior engineer's tacit knowledge into runbooks and case libraries an agent can consume? Who sets the agent's guardrails, writes its rollback conditions, defines what counts as success? Who judges its accuracy and miss rate in the retrospective?

That sixth item is the real bottleneck in both rooms, and it's directly a knowledge problem: the ceiling on an AI deployment is largely set by how much tacit knowledge an organization has converted into structured assets. Whether a team's runbooks are three-year-old prose notes or structured documents with trigger conditions, hypothesis lists, verification steps, and rollback conditions determines the performance gap between two companies running the same agent. The people doing that conversion sit upstream of the AI.

Will AI Replace My Job? A Three-Question Test for Operations Tasks

For any task on your plate, three questions:

  1. Is there a standardized procedure? Anything with stable, explicit steps gets automated first.
  2. Is there enough data and a clear pass/fail criterion? If yes, it's learnable.
  3. How bad is an error, and who owns it? High-consequence work with a named owner keeps a human confirmation step for a long time.

Three yeses means the task is a strong candidate for automation or AI-assisted execution. Any single no is a signal to examine what human involvement, verification, or oversight is still required.

The test is worth running rather than just reading. Take last week's tickets and tasks and label each one against the three questions, then compute the share that scores yes-yes-yes. The value isn't the number itself; it's that a vague anxiety ("am I going to be replaced") becomes a measurable ratio you can track quarter over quarter, and trend carries more information than absolute value.

Three ways people get this wrong:

Documented process and actual process frequently differ. People label against the SOP document while the real execution is full of exceptions and verbal agreements. The question asks whether execution has stable steps, not whether a document exists. If a procedure only runs with a senior person supplying ten unwritten rules out loud, it scores no on question one.

Having data isn't the same as having usable labels. If the root-cause field in incident records was filled in casually by whoever was on shift and never verified, that data is worth little for training or evaluation. Question two should really ask who labeled it and whether the label was validated.

Consequence gets systematically underestimated. An operation that "has never caused a problem" reads as low-stakes, but it may simply be a small sample. Judge consequence by worst-case blast radius, not by historical frequency.

One addition: don't leave out the work that was never ticketed — the throwaway script, the quick health check you ran for someone, the alert you looked at as a favor. Those are often the most repetitive and most exposed to automation, and they're the easiest to miss in any time report.

What Would Break This Conclusion About AI and Operations Jobs

An honest version states its own failure conditions. Three of them, each mapping to one of the six reasons above.

Inference costs drop another order of magnitude and long context makes coupled cross-system failures exhaustively searchable. That hits reason two directly. If "a human running hypothesis-and-elimination is cheaper" stops being true, that moat gets noticeably shallower.

Accountability frameworks change. That hits reason one. If insurance, industry standards, or regulators permit agents to carry responsibility for a defined class of change — say, low-risk, automatically rollbackable, fully audited — the boundary moves inward. That hasn't happened in most jurisdictions yet, but it's a policy variable rather than a technical one, so it can move faster than expected.

Embodied field operation matures. That hits reason four. As robots become capable in unstructured environments, the cost of physical non-routine work falls and that moat shallows too.

The reverse risk is more immediate than "the AI isn't good enough." DORA's 2025 survey of nearly 5,000 technical practitioners found roughly 90% use AI at work, with a median of about two hours a day, but only 4% report very high trust and 20% somewhat high — about a quarter combined. The report's framing is that AI acts as an amplifier: it magnifies the existing strengths of effective teams and the existing process and technical-debt problems of ineffective ones. That gap shows "using it" and "trusting it" are separate states, and the dangerous combination runs the other way: trusting it without verifying it. In operations, that combination costs one bad autonomous change.

One self-correction: this article may be conservative about the speed at which L1 and L2 work gets absorbed. Vendor product cadence — general availability, agent-ification, usage-based pricing — moves faster than most organizations' institutional adjustments, and the gap between the two tends to exist as "tooling deployed, responsibility undefined." Real risk during that window is higher than a paper assessment suggests.

AIOps Skills Roadmap: Rebuilding Operations Skills in 2026

This section describes what's happening in teams where it's already working, rather than prescribing general upskilling.

The change in knowledge form is the main thread. Tacit experience ("this service jitters at night, don't panic") becomes structured assets: trigger conditions, hypothesis lists, verification steps, rollback conditions, historical cases and their outcomes. Whoever does that conversion holds the ceiling on agent performance.

There's a low-cost move available immediately: take the last three real incidents, rewrite them in a format an agent can consume, then have the agent run the procedure and see where it goes off the rails. That exercise produces two things. First, it exposes the parts of the documentation you never actually thought through — people writing runbooks habitually skip the why and record only the how, and the why is the only part that still works when the situation deviates. Second, it gives you an empirical record of where the agent's capability boundary sits, which is more reliable than any vendor material.

The core new skill is verification, not operation. Concretely: reading an agent's reasoning chain and locating its evidence gaps; designing a counterexample that falsifies a plausible conclusion; distinguishing a correct conclusion from a correct conclusion reached by faulty reasoning, since the latter fails again in a different form next time.

Observability and the data foundation become the new source of leverage. Without data quality, an agent can only invent. Conversely, the person who can define what to collect, how to label it, and how to validate those labels gains standing.

The moves differ sharply by tenure.

0–3 years. The most exposed position is pure button operation, which is exactly the layer compressing fastest. A reliable path to accumulating judgment is claiming incident retrospectives and runbook authorship: that work forces you to answer why, which is one of the few fast routes to judgment. Pair it with replaying one real incident into a reproducible drill in a test environment — worth more than ten documents.

3–8 years. Shift from handling failures to designing boundaries: which changes may execute autonomously, how guardrails and rollback conditions get written, how to measure agent accuracy and miss rate, how to degrade when the agent misjudges. This work doesn't have a settled job title yet, and it's among the scarcest capabilities of the next few years.

8+ years / management. Make human-in-the-loop a process rather than a slide: an approval matrix for which change classes an agent may execute autonomously, audit logging of who approved what and when, retrospective metrics on agent-involved incidents (conclusion accuracy, miss rate), and a named owner for the knowledge assets. Without an owner, that work stalls within a quarter.

From Execution to Judgment: What Operations Roles Become

In both rooms, what got taken was hand work. What remains is head work and shoulder work: judgment, and carrying the consequences of it.

That boundary will move, probably faster than most organizations' institutional adjustments. Its direction is predictable: standardized process with data and clear criteria goes first; work requiring accountability, cross-party coordination, or handling of genuinely novel situations goes later.

For an individual, the actionable question is narrow: what share of last week's hours fell on the side that goes first, and is that share rising or falling.

One question to leave with readers: in your team, is there a defined process and a named owner for "the agent's judgment was wrong"? If so, which document is it in?


Read More of Intelligenr


References

  1. Datadog, 2025 Bits AI SRE Agent to Resolve Incidents Faster https://www.datadoghq.com/about/latest-news/press-releases/datadog-launches-bits-ai-sre-agent-to-resolve-incidents-faster/

  2. Dadfarnia, M., Sharp, M., & Herrmann, J., 2025 Comprehensive evaluations of condition monitoring-based technologies in industrial maintenance: A systematic review https://www.nist.gov/publications/comprehensive-evaluations-condition-monitoring-based-technologies-industrial

  3. DORA, 2025 State of AI-assisted Software Development https://dora.dev/research/2025/dora-report/


Author Note: The goal here is not to predict which operations jobs will disappear, but to examine which parts of the work are most exposed to automation — and which forms of judgment remain harder to transfer. The boundary will continue to move as AI capabilities, organizational practices, and accountability frameworks evolve.

This article was published in Learning With AI.

← Back to Learning With AI