Exploring how AI reshapes the way humans think, learn, and create.

Learning With AI

Generative AI in Security Operations: Cases, Risks, Skills

Generative AI in security operations: what agentic triage really fixes, where context sets the ceiling, and how analysts build judgment without repetition.

Intelligenr Research Team·
generative ai security operationsagentic socai log analysisalert triage automation

This article combines published security research and industry reports with fictional composite scenarios. The scenarios illustrate operational risks and are not accounts of verified incidents at specific organizations. Vendor-reported performance figures are identified as such and should not be treated as independently validated results.

Generative AI in Security Operations: Two Cases, One Bottleneck


Why Security Operations Is a Demanding Test for Generative AI

Security operations has three properties that make it a demanding environment for generative AI: a continuous supply of machine-readable data, feedback that can often be used to evaluate decisions, and time-based metrics that make operational changes measurable. Logs, alerts, tickets and change records arrive every day. Some incorrect decisions surface within hours or days, while others remain undetected until an incident or retrospective review. MTTD, MTTR and false-positive rates give teams measurable indicators, although none alone establishes that an AI deployment has improved security outcomes. Together, these conditions make security operations a useful environment for testing generative AI under production constraints rather than only in controlled demonstrations.

The results have not been uniform. SANS Institute's AI Survey Insights 2026 (published 13 July 2026; 536 practitioners and 57 executives) reports that AI use in security rose from 50% to 78% year over year, with red-team usage climbing from 33% to 61%. Over the same period, only 27% of respondents described their deployments as mature and in production, and 63% reported material AI deficiencies in detection and response, up from 45% in 2025. The adoption curve and the maturity curve are clearly moving at different speeds, and that gap is where most of the interesting work sits.

What follows runs along two tracks. The first starts on the defensive side, in the logs, passes through a failed pilot, and arrives at the question of who owns a conclusion. The second starts on the offensive side, via an incident disclosed publicly in autumn 2026, and arrives at the fact that AI systems have become an attack surface in their own right. The two tracks converge on the same place: the constraint has moved from detection speed to context quality and responsibility boundaries.

Case One: A Bank, 900 Log Sources, and a Pilot That Looked Too Good

Bank A is a mid-sized European retail bank with roughly 6,000 employees and three acquisitions over the past decade. Its security operations center runs across three SIEM platforms. More than 900 log sources feed in, a substantial share of them inherited from acquired entities. Time synchronization has been a chronic problem — NTP drift on some hosts reaches the seconds, so stitching a cross-device timeline requires manual correction. Tier-1 monitoring is outsourced to an MSSP; the in-house team handles escalations and forensic work. Daily alert volume runs into the tens of thousands, far beyond what humans can triage, so a large fraction closes as "unhandled." That is common in the industry and rarely appears in a monthly report.

In the first half of 2026, Bank A stood up an agentic triage pilot. The scope was deliberately narrow: two log sources (endpoint EDR and authentication logs), read-only analysis, output limited to a verdict plus cited evidence, and no response actions of any kind. The first two weeks looked strong on the dashboard. On the vendor's own figures, alert noise reduction exceeded 95% and average per-alert triage time fell from tens of minutes to single digits (vendor-reported, not independently verified; the numbers came from the vendor's self-reported results at comparable-scale customers, and Bank A had not established its own baseline measurement). Shift pressure visibly eased, and the project's weekly status began to include language about expanding scope.

In week three, a real incident was judged benign and closed.

The underlying event was unremarkable in shape. An internal user exported a bulk set of customer contact records through an internal reporting system outside business hours. That pattern genuinely existed in Bank A's historical baseline — quarter-end reconciliation produced similar query shapes. This time it was not reconciliation. The exported data appeared on an external channel four days later. Reviewing the record afterward, the agent's reasoning chain was complete, it cited real log entries, it produced a confidence score, and on the surface there was nothing to fault in the form of the output.

The first line of the post-mortem did not mention the model.

What the Post-Mortem Actually Found: Context Sets the Ceiling

The review located four gaps, all of them outside the model.

Change windows were not in the context. Eleven days before the event, an internal system had undergone a permission-model change. The record existed in the change management system. It had never been connected to the triage pipeline. The agent saw an anomalous privilege elevation; it could not see that the elevation had been approved days earlier.

The asset inventory was missing three classes of internal systems. The reporting system in question was not tagged as a sensitive data source, because it had never been reclassified after the acquisition. Without that tag, downstream severity scoring automatically dropped a level.

The behavioral baseline covered a single quarter. Quarter-end reconciliation query shapes had been learned as normal. The system did not distinguish bulk export inside a reconciliation window from bulk export outside one, because business rhythm on the time axis had never been modeled explicitly.

There was no evaluation set. This was the most consequential gap. The "improvement" observed in the first two weeks could not be separated from a degradation in judgment quality. A rising noise-reduction rate is equally consistent with genuine capability gains and with boundary cases being suppressed along with the noise. Bank A had no labeled corpus of historical incidents — including the ones that had been closed unhandled — and therefore could not answer the question that mattered: would the missed event have been caught two weeks earlier?

The four gaps share a characteristic. They are data-contract problems rather than algorithm problems. Log normalization, field mapping, time synchronization, asset and business tagging, change-event streaming — these determine the ceiling on what an AI output can be, and they usually sit outside the budget line of an AI project.

That leads to a practical rule: without a regression evaluation set, any claim of improvement lacks falsifiability. After the review, Bank A assembled the previous twelve months of real incidents (false positives and misses included) into a fixed test set, and re-ran it on every model swap, prompt change, or rule modification, tracking recall, precision, and misaction rate. The exercise adds no new capability. It replaces "it feels better" with "the numbers moved."

The table below organizes the current capability boundary by workflow stage, along with the failure mode typical to each.

Table 1: Capability boundaries of generative AI across security operations stages

Stage Tasks AI can assist with under defined conditions Decisions requiring human control or independent validation Typical failure mode
Log ingestion and parsing Format detection, field mapping, semantic extraction from unstructured text Collection scope, data-contract maintenance, retention and compliance Missing upstream fields silently dropped; downstream analysis built on partial data
Alert noise reduction and prioritization Clustering, deduplication, context correlation, priority suggestions Alert closure policy and business-criticality weighting Boundary cases suppressed along with noise; misses remain invisible in aggregate metrics
Investigation and attribution Evidence collection, timeline reconstruction, ATT&CK technique mapping Causal determination, blast radius and business impact Real logs cited, but the wrong chain inferred; formally complete, substantively wrong
Response and remediation Drafting response plans, preparing low-risk actions, writing post-incident drafts Authorization for state-changing actions, rollback decisions and external communication Actions executed on stale context or without a viable rollback path
Offense and assessment Reconnaissance orchestration, bulk validation and report generation within an approved scope Attack-path selection, authorization and scope boundaries Assets outside the authorized scope pulled into testing
Securing AI systems themselves Prompt-injection detection, content auditing and anomalous-behavior monitoring Permission-model design, trust-relationship design and audit policy Tool output treated as trusted input; permissions inherited along the call chain

The defensive constraint turned out to be context. The offensive change arrived faster, and needed less context to take effect.

Case Two: Agentic Tools and the Perimeter Nobody Scoped

In an analysis published on 7 October 2026, CrowdStrike described a campaign against South Korean financial organizations that was active from late September to early October. The activity involved ARTEX, a recently released open-source agentic penetration-testing tool developed in China, alongside several large language models. Reported targets included a loan-progress inquiry service used by financial brokers and an employee mobile-work support system, rather than the institutions' core accounting platforms. One particularly revealing finding was the discovery of exposed directories containing Claude Code session histories, ARTEX configuration files, and Claude memory files. These records provided a view into the attacker's tooling and operational workflow.

CrowdStrike's assessment drew a specific distinction. In incidents of this type, AI compresses the time spent on reconnaissance, result analysis, script authoring, tool invocation, and record-keeping, which shortens attack cycles and lowers the skill floor. It does not generate novel exploits from scratch. That distinction matters because it determines where defenders should put resources.

Company B, a payments fintech with roughly 1,200 employees, ran an internal red-team exercise in the third quarter of 2026. The scope followed convention: core transaction systems, API gateways, cloud infrastructure. When the team extended scope to the internal AI toolchain, the density of findings rose sharply:

  • A customer-service agent had 14 MCP tools mounted, three of them third-party, with no provenance or behavioral validation performed before mounting.
  • The agent ran on a long-lived service credential whose privilege envelope exceeded what any single support scenario required, and privileges propagated along the tool-call chain without per-call downscoping.
  • The vectorized knowledge base accepted writes from externally accessible documents, creating a context-poisoning path: an injected block of instructional text continued to shape agent behavior across later sessions.
  • A failed call to one tool triggered retry and fallback logic in the upstream agent, and the fallback path bypassed an approval node that the primary path enforced — a textbook cascading failure.

Independent research reinforces the pattern. In a July 2025 study, Trend Micro identified 492 MCP servers without client-side authentication or transport encryption, exposing how agent tool integrations can become gateways to sensitive data. A subsequent update reported 1,467 exposed MCP servers, indicating that the problem had widened. OWASP's Top 10 for Agentic Applications 2026 identifies risks including goal hijacking, tool misuse, identity and permission abuse, agent supply-chain compromise, memory and context poisoning, cascading failures, and rogue agents. Separately, the OWASP GenAI Security Project's LLM Top 10 2026 (published 4 August 2026) moved Excessive Agency from rank 6 to rank 3 and Unbounded Consumption from rank 10 to rank 6. Its ranking combines 75% community voting with 25% incident data drawn from a corpus of 7,714 incidents.

Shipped products have supplied their own evidence. EchoLeak (CVE-2025-32711, CVSS 9.3), disclosed in June 2025, allowed a crafted email to trigger data exfiltration through Microsoft 365 Copilot; Microsoft stated the server-side issue was remediated and that no customer impact was recorded.

Both cases point at the same shift: tempo changed, and the two sides did not change tempo in step.

The Attacker's Economics Moved Before the Defender's Did

Verizon's 2026 Data Breach Investigations Report provides several indicators of expanding AI-related exposure. Shadow AI activity in its data-loss-prevention datasets increased fourfold year over year. Source code was the most commonly observed data type submitted to external AI services, accounting for 28% of the reported data types, and 45% of employees regularly used AI on company devices, up from 15% the previous year. These findings point to growing exposure inside organizations even without a formal procurement decision. The further implication for individual practitioners is that assessing agent supply chains, tool poisoning, memory and context poisoning, and cascading failures may become a valuable specialization. Company B's customer-service agent illustrates how such systems can remain outside established security ownership until an assessment brings them into scope.

Read together, those figures suggest that AI bought attackers speed and coverage rather than new principles. IBM and Ponemon's Cost of a Data Breach Report 2026 (29 July 2026) sketches the same picture from the cost side: the global average breach cost reached $4.99M, up 12% year over year; AI-driven attacks accounted for a quarter of malicious breaches, up 56%, averaging about $6M each; among organizations hit by such events, the most frequently cited contributing factors were weak peripheral systems and misconfigured cloud environments, rather than direct compromise of core platforms.

The operational implication for defenders is a scope change. Broker portals, employee mobile-office systems, internal support services, and externally exposed APIs need to sit inside routine attack-surface assessment instead of being an appendix to an annual penetration test. The corresponding controls are fairly concrete: enforce multi-factor authentication on those systems, rate-limit automated requests, build dedicated alerting for bulk queries and bulk exports, and network-separate sensitive data sources from low-trust entry points.

Company B's second move after the exercise was to bring AI systems into its asset inventory and change management. Models, agents, MCP tools, vector stores, and prompt templates had previously appeared in no register, which placed them outside change control, rollback procedures, and audit scope. The technical difficulty of this step is low. Its importance is that it gives every subsequent control something to attach to.

Authorization boundaries for response actions are the other item that needs to be written down explicitly. Table 2 shows a three-tier division that is common in financial and critical-infrastructure environments.

Table 2: Action authorization tiers for AI-assisted response

Tier Criteria Example actions Audit requirement
Autonomous within defined permissions Read-only, no state change, and no material side effects Querying logs and asset data, retrieving threat intelligence, aggregating context, drafting a verdict Full logging; conclusions must carry evidence citations
Human confirmation required Changes system state, with a bounded blast radius and a viable rollback path Isolating a single host, blocking a single IP, disabling a single account, pushing a temporary policy Record approver, timestamp, and stated basis
Not permitted for autonomous execution Irreversible actions, changes to core transactions or customer data, or external disclosure without explicit authorization Deleting data, altering core transaction configuration, bulk account operations, external notification Follow the established change and incident-management process; record any required authorization

The tiering itself is not the point. The point is that it must live in the agent's permission configuration and execution engine, where it can be traced in an audit log, rather than in a policy document and informal understanding.

Once the technical boundaries are clear, the remaining problem sits with people, and it is harder to solve than the technical one.

The Judgment Problem: How Do You Build Skill When the Repetition Is Gone?

Tier-1 monitoring in a SOC has historically served two functions at once: processing alerts, and building judgment. High-volume repetitive triage is how newcomers develop intuition — enough false positives seen is what produces a feel for what counts as normal in a given environment. When agents absorb that work, the first thing removed is training volume, ahead of the role itself.

The risk of deskilling deserves attention as AI absorbs repetitive SOC work. Analysts whose daily tasks increasingly consist of reviewing agent output may accumulate experience in verification without getting the same volume of hands-on forensic practice. Yet the ability to evaluate an agent's conclusions still depends on understanding how incidents unfold, how evidence connects, and where an investigation can go wrong. Without deliberate opportunities to perform that work directly, the feedback loop can weaken the very expertise that effective oversight requires.

Industry discussion points to both changing responsibilities and the continued need for foundational skills. SANS's The Augmented Analyst (2026) describes a growing need for analysts who can supervise AI-assisted workflows and apply generalist judgment across investigation tasks. Infosecurity Magazine likewise highlights the importance of making AI-supported investigations inspectable and retaining human oversight. These discussions suggest that entry-level work and training may need to be redesigned, rather than assuming that repetitive tasks can disappear without affecting how analysts develop expertise.

Held together, those observations support a few concrete readings.

Skill formation has to be redesigned rather than assumed. After Bank A's review, the team created a responsibility that had not existed before: senior analysts annotate and attribute every incorrect agent verdict, and reproducible patterns get promoted into detection rules or context-enrichment items. This turns human correction from passive backstopping into active production of judgment standards. It is also the work newcomers can most easily contribute to, and the work that accumulates context fastest.

An auditable reasoning chain is the vehicle for transferring skill. If a conclusion arrives without original log IDs, timestamps, and the rule or detection logic that fired, a senior analyst cannot quickly locate the faulty link, and a newcomer has nothing to learn from. Requiring "no conclusion without evidence" improves accuracy and teachability at the same time.

Organizational knowledge needs to be externalized. Business baselines, incident history, cross-team trust relationships, known traps — these have long lived in oral form. In the agent era their value rises, because they are precisely what a model cannot acquire on its own, and the risk of them living only in individual heads rises with it. Writing them into a maintainable context layer (asset tags, business rhythm, change-event streams, a historical false-positive library) serves as both a dependency for AI projects and a hedge against staff turnover.

One more data point from the SANS 2026 survey is relevant: 73% of responding practitioners said AI had already changed their team's training requirements, up from 51% in 2025. Changes in training requirements can surface before changes in job titles.

Measuring It Honestly: Metrics That Survive a Vendor Demo

Demo-environment metrics and production metrics are frequently not the same set of numbers. A few recurring traps:

MTTD and MTTR comparability. Denominators and start/stop definitions vary widely between organizations, so cross-organization comparison has limited value, and even longitudinal comparison only holds if the measurement definition stays fixed. Writing the definitions into a measurement document and holding them constant across an AI rollout is the safer path.

Noise reduction in isolation misleads. Rising noise reduction and rising misses can occur simultaneously, and the first is visible in a monthly report while the second is not. Miss rate and misaction rate therefore need to be presented alongside noise reduction, and miss accounting must include alerts closed as "unhandled" — Bank A's event escaped through exactly that category.

Adoption rate and override rate of AI recommendations. An adoption rate approaching 100% usually signals that review has become ceremonial rather than that the model has become accurate. The override rate and the distribution of override reasons carry more information about the system's real state.

Drift monitoring. Model version updates, log schema changes, and shifts in business rhythm all degrade judgment quality over time under an unchanged configuration. The value of a regression evaluation set is that it makes such degradation detectable before an incident reveals it.

How to read vendor figures. Vendor-supplied noise-reduction, automation-rate, and time-compression numbers typically come from the vendor's own customer environments with unpublished measurement methodology. Treating them as directional references is reasonable; treating them as acceptance criteria is not. One of Bank A's lessons was writing vendor figures into project status reports without first building an internal baseline.

Where the Opportunity Sits: Value Migration in the AI-Augmented SOC

The previous two sections dealt with systems and measurement. Moving the lens to individuals, the public data suggests several directions in which professional value may be shifting, and these can appear before changes in job titles. In the SANS 2026 survey, 73% of responding practitioners said AI had already changed their team's training requirements, up from 51% in 2025. This is evidence of changing skill demands, although it does not by itself establish a universal sequence in which training changes always precede role restructuring.

First migration: from consuming alerts to producing the standard of judgment. The Case One review points to a structural fact — the context layer (asset and sensitivity tags, business rhythm, change-event streams, historical false-positive library) sets the ceiling on AI output, and that layer is maintained by people, updated by people, and owned by people for accuracy. In most organizations this work has no clear owner. It sits at the seam between operations, detection engineering, and the business. Whoever takes it on and keeps it maintainable holds the ceiling on judgment quality rather than merely the execution of judgment. The observable signal is plain: check whether misattribution of false positives and misses has a fixed place to be recorded. Where it does not, there is an opening.

Second migration: defensive scope expanding from core systems to AI toolchains and peripheral connected systems. Verizon's DBIR 2026 carries a set of numbers worth noting: shadow AI grew fourfold year over year in DLP datasets, the most commonly exfiltrated data type was source code at 28%, and 45% of employees routinely used AI on company devices, up from 15% the prior year. Attack surface is growing inside organizations without any procurement decision. At the same time, people with hands-on experience assessing agent supply chains, tool poisoning, memory and context poisoning, and cascading failure are scarce. Newly opened domains usually have few incumbents, which makes them both a risk exposure and the lowest-cost entry point for an individual. Company B's customer-service agent toolchain, uncovered by the red team, sat outside anyone's remit before the exercise.

Third migration: from executing actions to designing boundaries and validating conclusions. The three-tier authorization in Table 2, auditable reasoning chains, and rollback-capable execution paths are design work rather than execution work; their output is a constraint, not a ticket. A distribution in IBM's 2026 report reflects the same phase: only 18% of organizations use agents for vulnerability remediation, while more than half use them for detection and containment. Automation currently concentrates on the read side, with the write side still gated by humans. That distribution will move, but its direction and rough pace are predictable, and understanding the write-side authorization model in advance costs less than learning it after the shift.

Collapsing these three migrations, the dividing line repeated across industry discussion lands on three checkable self-assessments. First, what share of daily judgment consists of replicable fixed patterns — that share will be automated eventually, and people who hand it over deliberately while retaining acceptance responsibility generally end up in a better position than those who wait. Second, whether one holds context others cannot obtain — business baselines, incident history, cross-system causality — which is exactly what a model cannot acquire on its own and exactly what was missing in Bank A's misjudgment. Third, whether there is demonstrable work to show: a labeled evaluation set, a set of detection rules grown out of false positives, a post-mortem that measurably reduced misaction rate. All three are verifiable by a third party, which makes them more informative about a person's actual position than a title.

One counter-observation on deskilling is worth recording. AI-assisted review can also accelerate learning when analysts can inspect the evidence behind an agent's conclusions and understand how an investigation was conducted. The risk arises when reviewing and approving output replaces, rather than supports, hands-on forensic practice. Generalist judgment holds up only if analysts can independently verify each layer of output; otherwise, specialization gives way to uncertainty across more stages. At the individual level, occasionally reconstructing an incident from raw logs and writing a detection rule by hand can help preserve foundational skills. At the organizational level, review exercises and rotation mechanisms can make that practice part of the workflow.

Bank A's review did not point primarily to model capability; it pointed to a missing asset inventory and missing change context. Company B's weakest link was not necessarily its core transaction platform; it was an agent toolchain that had never entered a formal register. Machines can already process many detection and triage tasks faster than people, but speed alone does not establish judgment quality. Two questions remain central to accountable security operations: what counts as normal in a given environment, and who is responsible for a conclusion. The gap between 27% of SANS survey respondents describing their AI deployments as mature in production and the broader adoption figures illustrates the challenge of turning deployment into operationally reliable practice.


Read More of Intelligenr


References

  1. CrowdStrike (2026) Unknown Threat Actor Uses AI-Driven ARTEX to Target South Korean Finance https://www.crowdstrike.com/en-us/blog/unknown-threat-actor-uses-artex-to-target-south-korean-finance/

  2. Verizon (2026) 2026 Data Breach Investigations Report https://www.verizon.com/business/resources/T1e0/reports/2026-dbir-data-breach-investigations-report.pdf

  3. IBM and the Ponemon Institute (2026) Cost of a Data Breach Report 2026 https://www.ibm.com/reports/data-breach

  4. OWASP GenAI Security Project (2026) OWASP Top 10 for Agentic Applications 2026 https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/

  5. SANS Institute (2026) AI Survey Insights 2026 https://www.sans.org/white-papers/2026-sans-ai-survey-insights

  6. SANS Institute (2026) The Augmented Analyst: How AI is Reshaping Security Operations in 2026 https://www.sans.org/blog/augmented-analyst-how-ai-reshaping-security-operations-2026

  7. Trend Micro (2026) Update on Exposed MCP Servers: The Threat Widens to the Cloud https://www.trendmicro.com/vinfo/de/security/news/vulnerabilities-and-exploits/update-on-exposed-mcp-servers-the-threat-widens-to-the-cloud

  8. OWASP GenAI Security Project (2026) OWASP GenAI LLM Top 10 2026 https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/

  9. Microsoft Security Response Center (2025) CVE-2025-32711 — Microsoft 365 Copilot Information Disclosure Vulnerability https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711


Author Note: Intelligenr examines how AI changes human judgment, expertise, and decision-making. This analysis considers not only what AI can automate in security operations, but also what human oversight and sustained analyst development still require.

This article was published in Learning With AI.

← Back to Learning With AI