AppSec: The Next Generation
AI has created a gold rush in software security, and much of the industry is using it to rebuild the same tools we already had.
We now have agentic SAST, agentic DAST, autonomous pentesting, AI remediation, AI-generated threat models, AI-assisted triage, and increasingly creative acronyms for products that still produce roughly the same artifacts and hand them to roughly the same people. The models are better and the interfaces are more conversational and some of the results are great but the underlying vision is often unchanged.
It feels as if we are taking the existing application security market and just replacing pieces of its machinery with agents rather than taking a holistic view of what bits of our workstreams survive, which should go away, and where we should be leaning in. Everything in software development changed over night. Now it’s our turn!
In my last article, I used vulnerability chaining to make a narrower version of this argument. An intelligent security system should not report several related weaknesses independently when their combination produces a more serious attacker outcome. It should validate the individual conditions, understand how they interact, account for controls and deployment context, and present the actual risk. The point was not that every scanner needs a vulnerability-chaining feature. It was that AI gives us the ability to perform work that older security tools could not, and we should stop using those tools as the specification for what comes next.
This article takes the larger swing. AI is changing how software is conceived, designed, written, tested, deployed, and operated. At the same time, it gives machines a new ability to interpret and synthesize the messy details that application security has always depended on. That combination should cause us to reconsider the discipline itself: what AppSec is for, how it operates, what practitioners do, which categories still matter, and what we can finally build that was not practical before.
We are not bug janitors
Application security has become heavily associated with vulnerability backlogs. We scan applications, validate findings, argue about severity, create tickets, negotiate remediation, retest fixes, and report on whether the pile is growing or shrinking. Vulnerability management consumes so much of the function that people have begun to confuse it with the purpose of the function, even though the vulnerability backlog was never the reason AppSec existed.
It was the tax imposed on AppSec by tools that could identify suspicious conditions but could not do much of the work surrounding them. They could not reliably determine whether a finding was real, understand the application well enough to establish consequence, account for compensating controls, find meaningful variants, decide which team owned the problem, work through remediation tradeoffs, or preserve what they learned for the next decision. Humans supplied the missing cognition and context, and as scanning scaled, the cleanup work scaled with it.
Eventually, the burden created by the tools became mistaken for the identity of the discipline.
An AppSec practitioner, is an adversarially minded software practitioner experienced in defensive controls and processes. They understand enough about code, architecture, identities, data flows, deployment, and operations to see how risk forms as systems are designed and built. They learn how an application is intended to work, identify the assumptions holding that design together, and find ways around those assumptions. They may not be the best application engineers in the company, but they need to be software-literate. If you work in software security and cannot understand how software is made, you are working at a permanent disadvantage.
That expertise has always been applied far beyond bugs. AppSec practitioners review early designs, revisit them when requirements change, build threat models, create secure standards, teach developers, advise on modern defenses, review consequential code changes, support launches, manage external researchers, help investigate incidents, assess acquired companies, and build guardrails that prevent entire categories of mistakes. Consultation is often the highest-fidelity work we do because it lets us apply years of experience breaking software to decisions that may affect hundreds of systems.
Many of us arrived in AppSec because finding vulnerabilities was fun. After enough years, finding another IDOR or proving that one more application can be compromised stops feeling like the highest use of that experience. The greater opportunity is to change how software is designed and delivered so the same weakness does not keep appearing.
At GitHub, for example, substantial engineering effort went into deploying and refining the Content Security Policy of their prime assets as part of a broader strategy against Cross-Site Scripting (XSS). The objective was not to become more efficient at filing XSS bugs forever. It was to make XSS an anomaly and make successful exploitation far more difficult when a bug did appear. That is categorical risk reduction. It is a different and more valuable outcome than closing a larger number of XSS tickets.
The OWASP Software Assurance Maturity Model (SAMM) reflects this broader reality. SAMM spans governance, design, implementation, verification, and operations. It includes strategy, education, threat assessment, security requirements, architecture, secure build and deployment, defect management, testing, incident management, and environmental management. AppSec has never been only a scanner program, even when scanners consumed an absurd amount of its time.
We are not bug janitors. Our job is to understand why the software factory produces risk and change the system so it produces less risk with the same, or hopefully less, effort.
Two constraints broke at the same time
The current moment matters because two constraints are changing together. The first is how software is delivered. AI is no longer confined to autocomplete or isolated code generation. Developers increasingly delegate design, implementation, debugging, testing, documentation, codebase exploration, and other work to agents. This transition is uneven, and claims of completely autonomous development remain far ahead of what most organizations can operate responsibly. Still, the direction is clear: software professionals are becoming accountable for more work than they personally perform.
Anthropic's own research offers a useful snapshot of this middle ground. Its engineers and researchers reported using Claude in roughly 60 percent of their work while saying they could fully delegate only a small fraction of their tasks. They described active supervision, validation, greater output, broader capability, concerns about skill atrophy, and a changing relationship with the craft of software engineering. That is a much more credible near-term model than either extreme: AI does not disappear as an elaborate guessing engine, and humans do not disappear because software suddenly builds itself perfectly.
The second constraint is what machines can understand about software. For the first time, we have broadly available systems that can interpret and synthesize information across source code, design documents, infrastructure, deployment configuration, security controls, issue history, organizational standards, and runtime evidence. They are imperfect, probabilistic, and capable of producing persuasive nonsense. But they can attempt classes of work that were previously too contextual, judgment-heavy, or expensive to automate.
What these systems understand did not materialize from the models themselves. Their apparent expertise is downstream of decades of human curiosity and experimentation. Humans discovered the vulnerabilities, developed the exploitation techniques, reverse-engineered the frameworks, documented the architectural failures, built the tools, proposed the defenses, and published enough of that work for a model to learn its patterns.
That matters because software security does not stand still. New frameworks introduce unfamiliar assumptions. New architectures create unexpected trust boundaries. Attackers discover ways to combine legitimate features into outcomes their designers never considered. Defenders learn that yesterday’s recommended control has a bypass, an operational weakness, or an interaction nobody anticipated. A model trained on yesterday’s knowledge inherits yesterday’s blind spots. Retrieval can provide newer information, but only if someone first did the work required to discover, validate, and document it.
AI may help generate hypotheses, explore variations, and accelerate research. It may even contribute to the discovery of techniques that no person has previously documented. But it cannot be allowed to establish its own output as truth. Someone still has to determine whether the technique works, understand why it works, distinguish novelty from nonsense, and turn the result into knowledge that other people and systems can trust.
If human experts gradually stop conducting original research, deeply investigating strange behavior, challenging accepted guidance, and publishing what they learn, the knowledge supply begins to stagnate. Without this human intervention the world continues changing while models increasingly reasons from stale human knowledge and derivative machine-generated material. Its answers may remain convincing even as their connection to current reality weakens.
This is another reason the rise of agentic security does not eliminate the need for skilled practitioners. We should delegate work that machines can perform effectively, but we cannot abandon the human frontier that makes those systems useful. Future AppSec practitioners will not only supervise agents. They will continue to experiment, discover, validate, teach, and contribute the new knowledge on which the next generation of security systems depends.
These changes create the conditions for the next generation of AppSec.
I am not predicting a distant world where humans express an intention and an autonomous machine handles everything else. I do not know whether that future is ten years away, fifty years away, or a mirage. Some organizations will push much further toward autonomy, and the lessons from that frontier will help shape the transition for everyone else.
The world I am concerned with is already emerging. Developers, QA engineers, platform teams, SREs, and security practitioners remain accountable for their respective responsibilities, but they increasingly manage agentic workforces that perform more of the execution.
The factory-floor analogy is useful here. Industrialization did not eliminate expertise from manufacturing. It changed where expertise was applied. People stopped fabricating every screw and began designing production systems, operating specialized machinery, establishing tolerances, inspecting output, diagnosing failures, and improving the factory. Output increased, but systemic mistakes could also contaminate thousands of products.
AppSec cannot stand outside this new software factory and periodically inspect what rolls off the line. We have never had the luxury of being separate from development. Our capabilities must operate inside the same flows, using the same evolving context, while preserving the independence needed to challenge assumptions and verify outcomes.
Tacking security at the end was an early experiment that failed spectacularly. Integrating controls directly into the developer workflow is now essential, yet it cannot become a bottleneck that delays shipping. Our fundamental goal in application security must be designing guardrails so that the easiest way to deliver code is naturally the safest one.
The old categories become techniques
The security market is already showing signs that its familiar categories are collapsing. Static Application Security Testing (SAST) is becoming broader agentic software analysis. Once a system can reason across first-party code, dependencies, configuration, infrastructure, data flow, architecture, and intended behavior, the separation between code analysis and component analysis becomes less meaningful. A dependency is part of how an application behaves. Its importance depends on whether it is loaded, reachable, attacker-controllable, exposed under actual deployment conditions, and capable of producing a meaningful outcome.
That absorbs part of Software Composition Analysis (SCA), but not all of it. Supply-chain security also requires inventory, provenance, end-of-life tracking, licensing where relevant, and continuous intelligence about newly vulnerable or malicious packages. Those capabilities persist, but they become evidence streams feeding a larger security system rather than a separate queue of component findings.
Dynamic Application Security Testing (DAST) is dividing into at least two distinct functions. Autonomous pentesting and red teaming explore deployed systems adversarially, looking for attack paths through runtime behavior. Targeted runtime verification begins with a hypothesis produced elsewhere and interacts with the application to establish whether a suspected condition can actually be exercised. Autonomous red teaming (slight misnomer if you ask me) automates the offensive-security work performed by AppSec where bugs are discovered, attacked, and validated at runtime.
Interactive Application Security Testing (IAST) may have less of a future as a standalone category. Few teams are asking for more operationally difficult instrumentation sold under another acronym. But its underlying capability, observing execution and data flow from inside the application, may become valuable telemetry for an agentic verification system.
The techniques do not become useless; their boundaries become less important. SAST, DAST, SCA, and IAST remain valuable ways of collecting evidence. What no longer makes sense is treating them as separate workflows and sources of truths. Those categories describe how and in which way evidence is collected, not the security outcome AppSec is responsible for producing. AI-Native AppSec should be able to draw from any of them as needed to help produce an accurate picture of risk. An intelligent security system should begin with what it needs to establish, then select the evidence and technique appropriate to the question. Sometimes source and configuration will be enough. Sometimes the system will need runtime observation. Sometimes it will need an autonomous adversarial test. Sometimes it will need a human who understands a business decision the machine cannot own.
Rebuilding each acronym as an autonomous product may improve individual tools. It will not produce the operating model required for agentic software development.
The problem is trustworthy risk state
The most dangerous and challenging problem AppSec faces is maintaining a trustworthy understanding of risk across thousands of services changing faster than humans can continuously reconstruct.
Risk state is not a score painted red, yellow, or green on an executive dashboard. It is a defensible understanding of a service assembled from many kinds of evidence: its purpose, architecture, trust boundaries, exposed attack surface, data sensitivity, identities, controls, technology stack, dependencies, ownership, maintenance activity, known vulnerabilities, deployment conditions, production behavior, external reports, incident history, and the freshness and confidence of every conclusion.
Today, much of that understanding lives in fragments. Design intent sits in a product document. Architectural decisions are buried in meetings and chat. Threat models live in diagrams that began aging before the meeting ended. Scanner findings live in several products. Runtime evidence belongs to another team. External reports enter through a separate system. Remediation appears in a pull request. The incident timeline lands in a postmortem. An experienced AppSec practitioner becomes the human memory and reasoning layer connecting all of it. Especially important in an era where it is becoming increasingly rare amongst stakeholders to have a shared or mutual agreement on the state of “risk”.
That approach cannot survive machine-scale software change. Threat modeling shows what a redesigned capability could look like. The old version is usually a workshop attached to a moment in the lifecycle. AppSec reconstructs the architecture, asks questions, records threats, assigns follow-up work, and produces an artifact that gradually diverges from reality. In the best of scenarios, the lazy AI version generates the same document with an LLM. At its worst, it creates a deceptively plausible yet highly inaccurate version of that same document.
The useful version continuously maintains the security implications of the service as it evolves. Imagine a team introducing a background worker that processes customer-supplied archives. The system recognizes the new data flow, parsing boundary, storage interaction, execution context, identity, and network access. It relates those changes to approved patterns, organizational knowledge, prior incidents, and the existing threat model. It checks code and deployment configuration for evidence of isolation and control enforcement. It distinguishes verified controls from claimed controls, proposes adversarial tests for unresolved assumptions, and escalates design tradeoffs that require accountable human judgment.
More importantly, it keeps watching. If implementation deviates from the original product requirements, the change does not update only a design document. It may alter the threat model, invalidate hardening recommendations, change which controls require verification, and affect the service's risk state. Those dependent conclusions must move together.
This is not automated threat-model generation. It is continuous alignment among intent, design, implementation, deployment, and observed behavior.
That service-level capability should operate alongside a peer program-level system. The service system lives inside delivery workflows and maintains security as individual applications change. The program system works across the organization: identifying recurring conditions, aging technology stacks, unsupported frameworks, ownership gaps, weak platform capabilities, incident patterns, and categories of risk that should be eliminated centrally.
Evidence and capabilities flow in both directions. Service-level work reveals systemic weakness. Program-level work turns those patterns into safer frameworks, libraries, standards, skills, controls, and paved roads. Those capabilities return to the delivery environment, where service-level verification determines whether they work under real conditions.
AI-native AppSec operates within individual software lifecycles while learning across all of them. It protects each service, then uses what it learns to change how the organization produces software.
AppSec gets to become security engineering again
This creates an important convergence between Application Security and Product Security. Companies use those names inconsistently, and I am not proposing a universal org chart. I am describing a convergence of capabilities.
AppSec has often concentrated on adversarial analysis inside applications and services: understanding how they fail, challenging assumptions, finding attack paths, and determining whether controls work. Product Security, in the distinction I am drawing, has tended to operate at a more systemic engineering layer: building authentication capabilities, secure libraries, supported frameworks, customer-facing security features, guardrails, and paved roads that eliminate or constrain risk across many services.
Agentic systems can give AppSec practitioners more capacity to turn adversarial knowledge into systemic engineering. If ten services repeatedly implement authorization incorrectly, the goal should not be to find and close authorization bugs faster. The security function should recognize the organizational pattern, determine why the factory keeps producing it, help design and build the safer primitive, deploy it through the development environment, and verify that the control works.
The durable output is not a finding or even a patch. It is a changed organizational condition: the unsafe pattern has become harder to introduce, existing instances have been addressed, and evidence shows that the control continues to work.
This evolution raises the bar for AppSec practitioners. They will need to be builders. They will need enough engineering depth to understand the software systems they govern and enough AI systems knowledge to understand the security capabilities acting on their behalf. They will need to design workflows, choose where probabilistic judgment is appropriate, define evidence requirements, establish permissions, build evaluations, investigate failures, and improve the system over time.
The uncomfortable reality is that the field does not uniformly have those capabilities today. Every profession has a broad middle. Many AppSec practitioners entered organizations where success meant configuring purchased tools, triaging output, meeting ticket SLAs, and following established playbooks. They are now being asked to govern systems they may not understand while models, frameworks, threats, and development practices change underneath them.
The old guard has a problem too, and I include myself in it. Experience teaches us why prior approaches failed, but it can also anchor us to the categories we spent our careers improving. Meanwhile, the new guard may understand agents while lacking enough AppSec scar tissue to recognize the failure modes they are recreating.
The next generation needs both. Experienced practitioners must contribute the history without demanding that every historical category survive. New builders must bring imagination without assuming software security began with LLMs. We need to return to first principles, but with an education. Our core secure engineering principles are foundational and still apply even in an AI-centric world.
AI can generate findings without reducing risk
Software development is already experiencing its own reckoning. AI can produce enormous amounts of code, but producing code is not the same as building software. Building software requires understanding intent, making architectural tradeoffs, integrating with existing systems, testing assumptions, preserving maintainability, operating the result, resilience, and caring whether it continues to solve the right problem after deployment. The same distinction applies to security: AI systems can generate code without building software, just as AI security systems can generate findings without reducing risk.
More findings, patches, tickets, threat-model documents, and automated reviews can create the appearance of security productivity while weakening ownership and judgment. If the system cannot establish which observations matter, understand how conditions interact, recognize when its assumptions become stale, and connect its work to a change in risk, then we have automated a potentially time and resource wasting activity rather than built a security capability.
This is why model capability is not system capability. A strong model can read code, invoke tools, produce plausible or convincing explanations, and perform well in a demonstration. A production security system also needs durable state, evidence provenance, permission boundaries, failure containment, evaluation, observability, escalation, cost control, and a way to learn from what happened after the original decision.
A harness is not an orchestration layer, and an agent loop is not an operating model.
Humans govern; agents operate within bounds
If agentic development increases the volume and velocity of software work, AppSec cannot respond by inserting a human approval into every security decision. We cannot spend the day interrupting the factory and call it governance. Human-in-the-loop designs that require constant routine approval will either become bottlenecks or devolve into ceremonial clicking (which is already becoming the default). Humans need to be on the loop, establishing the conditions under which autonomous work is permitted and intervening when those conditions no longer hold.
Practitioners should establish what a specialized agent is allowed to do, which evidence it must collect, what error rate is acceptable for that task, what permissions it receives, which conditions are considered normal, and what must trigger escalation. The agent can then operate autonomously inside that evaluated envelope.
The escalation policy cannot be universal because there is no universal AppSec agent. A threat-modeling agent, external-report triage agent, remediation agent, supply-chain monitoring agent, production-testing agent, and guardrail-building agent have different consequences and authority. Still, several general conditions hold. The system should escalate when operating conditions move outside validated bounds, when an action requires elevated permissions or additional tools, when a decision requires accountable risk ownership, or when evidence is contested and confidence remains below the threshold required for the action.
Consensus among agents can be useful, but agreement is not automatically independent confirmation. If the agents share a model, retrieval source, context, or flawed assumption, consensus can be correlated error. Retrospective evaluations using captured traces, multiple models, and eventual ground truth can establish where disagreement predicts failure and where agreement has historically been reliable. The orchestration layer must know the difference between corroboration and repetition.
Oversight also does not happen only in real time. Some of the most valuable correction happens after the original decision, when later evidence establishes what was true. Production traces can be compared with ground truth, failures can be clustered, biases can be identified, and the operating envelope can be expanded, narrowed, or revoked. Do not only evaluate outputs. Evaluate the conditions that produced them.
Trust in autonomous security systems has to be continuously earned and continuously verified.
That trust is multidimensional. Accuracy matters, but it is not the only measure. We must consider false positives, false negatives, and the consequences of incorrect action. We must also consider the full economic cost of models, orchestration, observability, evaluation, maintenance, and human correction. We should measure whether the system removes meaningful toil or creates a new class of persuasive garbage that practitioners must babysit.
Quality of life belongs in this calculation. A capability can be financially neutral and still provide real value if it removes repetitive triage, late-night dependency hunts, administrative follow-up, and other work that drains practitioners without using their expertise. For a low-consequence and recoverable task, a small reduction in accuracy may be acceptable when the system dramatically improves efficiency, remains inexpensive, and gives people better work. As consequence rises, the error envelope tightens and the evidence requirement increases.
Every probabilistic operation is another opportunity to get something wrong and poison downstream work. The answer is not to remove probabilism from a system whose value depends on interpretation and synthesis. The answer is to introduce it deliberately, bound its authority, preserve its evidence, and use deterministic mechanisms wherever judgment is unnecessary.

Figure 1: The bounded-autonomy control loop
Humans define operating bounds, authority, evidence requirements, and escalation conditions. Specialized agents operate within those validated conditions. Traces, outcomes, and anomalies feed continuous evaluation. Human correction updates the system and can expand, narrow, or revoke autonomy. Decisions that are contested, consequential, or outside the envelope return to accountable people.
Move carefully without thinking small
The strongest argument for preserving the existing AppSec model is not that it works particularly well. It is that its tools are bounded, understandable, auditable, and already integrated into organizations. Rules, dependency inventories, dynamic tests, and human design reviews may be incomplete, but teams know how to operate them. AI systems are probabilistic, expensive to observe, vulnerable to bad… everything (tooling misuse, outages, bad context, cost overruns, etc.), and capable of reaching confident conclusions that are wrong.
Most organizations also lack the solid foundation of clean service inventory, ownership data, architecture context, production telemetry, and engineering maturity required for the integrated system I have described. Replacing everything at once would be irresponsible, and it is not what I am recommending.
The destination can be ambitious while the migration remains controlled. Start with a bounded human activity. Capture how experienced practitioners perform it, including the evidence they use, the judgments they make, and the conditions that cause escalation. Introduce the agentic capability in shadow mode. Compare it with historical traces and ground truth. Find the dark corners, including missing context, unusual architectures, correlated errors, cost spikes, permission problems, and failures that contaminate downstream work.
Correct the system. Evaluate it again. Grant limited authority where performance remains inside the acceptable envelope. Observe production outcomes, not only benchmark accuracy. Expand to the next activity when trust has been earned.
Autonomy should not be a product setting labeled on or off. It should be conditional permission: this capability may take this action, with these tools, under these conditions, because the complete system has demonstrated acceptable performance and the consequences are understood.
Some familiar tools will remain during this transition, and some will remain indefinitely because deterministic analysis is the right technique for the job. The problem is not incremental implementation. The problem is allowing an incremental migration to define an unimaginative destination.
Move incrementally in implementation, but do not think incrementally about the destination.

Figure 2: Before and after the rebirth
The rebirth of AppSec
The security industry has a rare opportunity to rethink what it builds. Unfortunately, established categories are easy to explain, benchmark, fund, purchase, and sell. Agentic SAST fits into a known budget. Autonomous pentesting fits into a known engagement. AI remediation fits into a known ticket workflow. A system that continuously maintains security intent, reasons about service-level risk, learns across the organization, helps build systemic controls, and governs its own uncertainty is harder to package into a familiar box.
That packaging difficulty is not evidence that the familiar box is the future. The old AppSec categories were shaped by the limitations of the technology available when they were created. Some represent durable techniques. Others represent boundaries we no longer need. We should decide what survives based on whether it helps us reduce risk in the environment now forming, not because an analyst taxonomy or procurement process expects it.
AppSec practitioners also have to change. We cannot demand more capable systems while remaining passive operators of tools we do not understand. We will be responsible for security work performed on our behalf, including work at a scale no human team could execute directly. That requires more engineering depth, stronger evaluation discipline, better judgment about autonomy, and a willingness to build.
This is not the death of AppSec; it is a chance to recover its purpose. The discipline was never supposed to be defined by the backlog. It was supposed to apply software knowledge, adversarial thinking, and security engineering to burn down risk. Limited tools buried practitioners in the labor surrounding vulnerabilities. AI can remove part of that constraint, but only if we resist using it to make the same machinery run faster.
Stop asking how AI can perform yesterday's AppSec tasks. Start asking what security capabilities are required when software is increasingly produced by agentic systems, changes beyond human review capacity, and still leaves humans accountable for the result.
Do not rebuild what already existed with moar AI. Build what AppSec was trying to become all along.
Sources and further reading
- OWASP SAMM: The Model
- Anthropic: How AI is transforming work at Anthropic
- Anthropic: Scaling Agentic Coding Across Your Organization
- GitHub's CSP journey
CTA:

