You open your Software Composition Analysis (SCA) dashboard: hundreds of open findings, dozens marked critical, and a backlog that grows faster than anyone can clear it. Your project pull requests are full of autogenerated library version bumps which need review or aren’t passing the build at all. This was quite often the reality of basic or poorly configured SCA tools in software development projects. Now, with AI-driven vulnerability discovery, it has become dramatically worse. This post explains why CVSS-based triage can’t keep up and how to approach the triage process with layered, context-aware analysis, enhanced with AI tools aiding the decision of what to ignore.
TL;DR
CVEs are projected for 2026, while the number of actually exploited vulnerabilities has stayed roughly flat.
|
The flood
Most SCA findings come from third-party dependencies, often transitive ones that nobody on the team chose directly.
In 2025, 48,185 CVEs were published, 20.6% more than in 2024. 2026 is running far ahead. Disclosures in the first half of the year were nearly 50% above the same period of 2025, and FIRST’s vulnerability forecasting team now expects around 66,000 CVEs for the full year.

The main reason is that finding bugs got significantly cheaper.
- Mozilla’s CVE Numbering Authority reported a 164% jump in Q1 2026 disclosures over its prior-year baseline, which FIRST attributes directly to AI-assisted tooling run against the Firefox engine.
- In Anthropic’s Project Glasswing, the Mythos Preview model flagged more than 23,000 issues across over 1,000 open-source projects, about 6,200 of them estimated as high or critical. Of 1,752 high/critical findings reviewed by Anthropic and six independent security firms, more than 90% were confirmed as real.
- Sometimes volume has nothing to do with quality. Daniel Stenberg, lead maintainer of the `curl` library, tested the model on `curl` and found that most of its reports weren’t real vulnerabilities.
AI speeds up another issue: mean time to exploit. Mandiant’s M-Trends 2026 estimates it at minus seven days.
However, volume is not risk in itself. While disclosures are exploding, most of them often aren’t exploitable or even relevant to systems depending on them. That is the core problem of SCA now. Detection is solved, but what about prioritization? Engineers have to verify and rank findings at a volume the industry has never had to handle before. To an extent, this issue was present before the AI finding surge. Now it’s simply much more painful, and rule-based auto-triaging needs to evolve as well.
Manual triage is toil
Ask an engineer whether a CVE in some transitive dependency matters for their service. An honest answer requires knowing:
- Is the vulnerable code used at all? Declaring a dependency doesn’t mean calling the vulnerable function. Maybe it’s used only in tests?
- Can an attacker reach it? Is there a path from an entry point, such as an HTTP endpoint, a message consumer or a file upload, to the vulnerable code?
- Does attacker-controlled data get there? Many vulnerabilities only matter when untrusted input reaches one specific API.
- Does the vulnerability require a particular OS, infrastructure or application configuration?
- Is there a working exploit? A public proof of concept, or active exploitation in the wild? How hard is the attack? Does it need authentication, or a chain of other bugs?
- What does the service do? An internet-facing payments API and an internal batch job carry very different risks.
- Would updating the dependency introduce other risks? There could be breaking changes or performance considerations, often quite expensive to resolve.
Sometimes the answers can be given instantly, but in most cases it takes quite a while to assess each aspect. Relying on base scoring is far from enough. Yes, the standards were designed with context in mind. CVSS v4 has threat and environmental metrics for exploit maturity and deployment context. There are good complementary signals too: the CISA Known Exploited Vulnerabilities (KEV) catalog, EPSS exploit-prediction scores, and decision frameworks such as Stakeholder-Specific Vulnerability Categorization (SSVC) and Vulnerability Exploitability eXchange (VEX). Even if your SCA tool enriches the findings with these properties, they only cover a part of the equation. They help answer if a vulnerability is dangerous somewhere, not necessarily in your code, so we can treat them as the first layer of false-positive/false-criticality filtering.
This still leaves a situation where teams are overloaded with dashboards where everything is high/critical, nothing is prioritized, and engineers learn to ignore it.
Layered triage
Modern SCA tools filter findings in layers. Each one removes a different kind of false positive and passes the rest on. Cheap, deterministic filters go first, then the expensive, probabilistic one goes last and only sees what’s left. But filtering is not the whole story. Context analysis can also raise issue severity.

A worked example: SnakeYAML, CVE-2022-1471
SnakeYAML is widely used in the Java ecosystem, often as a transitive dependency: Spring Boot's core starter pulls it in, so many services include it without anyone choosing it. In versions before 2.0, its default `Constructor` doesn’t restrict which types can be instantiated during deserialization, so parsing attacker-supplied YAML can lead to remote code execution. The recommended mitigation is `SafeConstructor` for untrusted content. Public exploits exist, and the finding shows up in a lot of JVM repositories.
Consider two services in the same organization, both on SnakeYAML 1.x.
Service A imports configuration uploaded by users
Service B reads its own bundled settings
A basic SCA tool would report both as the same high-severity finding. Let’s see how an advanced layered analysis would rate it:
- Reachability. Is the vulnerable code executed on production? The constructor is instantiated and `load()` is called, so for both services the answer is yes.
- Exploit intelligence. Public proof-of-concept exploits exist, so priority goes up for both.
- Deployment context. Service A exposes an endpoint that accepts arbitrary user uploads. Service B has no such endpoints. This fact pushes the score of Service A up and Service B down.
- Code logic analysis. In A, the HTTP request body goes straight into `yaml.load()`. In case B, the input is a safe resource packaged into the application.
The verdict for Service A would then be “fix now”, while for Service B it can be reduced to “upgrade can be performed within normal maintenance cadence”. Telling A from B requires reading the code the way a reviewer would, which is where AI assistance becomes very helpful.
Why not just upgrade everything?
As I already mentioned, upgrades aren’t free. SnakeYAML 2.0 makes the safe constructor the default, which breaks applications that rely on deserializing their own types. That’s fine to schedule, but not something you do blindly across a hundred services in a week. A common effect of introducing tools like Dependabot, which automatically bump dependencies, is a flood of pull requests, many of which may be breaking the build and annoying developers, who eventually start ignoring them. This doesn’t mean that tools like Dependabot are useless - they indeed help to deal with keeping the dependency hygiene. Automated update PRs and fewer unnecessary dependencies shrink the backlog before triage even starts. Still, what we are left with grows and quickly may become unmanageable.
Another notable mentions are infrastructure (DB/MQ/etc.) client libraries, which may introduce performance issues after upgrading if not treated with proper config changes. Not every system can and should eagerly use all the latest dependencies as soon as they are released.
What we see in practice
We are partners of Aikido Security and we roll out vulnerability scan tooling for clients in different sectors with different levels of regulations, as well as on our own code. In our experience, a few patterns repeat:
- Most raw findings in a typical JVM or TypeScript service turn out not to be exploitable in context: the package is declared but unused, the vulnerable function is never called, or untrusted data never reaches it.
- Teams that triage by a simplistic scoring scheme eventually start neglecting the backlog and treating the dashboard as background noise.
- After layered triage, the backlog turns into a shorter list engineers actually act on. The biggest change is cultural: people start reading alerts again.
- A combination of complex static analysis and AI exploitability analysis brings the false positives to even better levels, significantly unburdening the engineers. It’s still auditable - every decision on lowered scoring is auto-documented. Example from the Aikido CVE Exploitability Analysis:

Questions to ask your SCA vendor
If you are responsible for a secure software development life cycle and you are about to evaluate vendors on their capabilities regarding SCA, here are some questions you should consider asking:
When the tool isn’t sure
determine reachability?
public exploits) override a “not reachable” verdict? | Who makes the call
its own? Can we require human approval per action type or per repository?
boxed so findings resurface? |
Good answer: Unknown is shown as unknown, never silently as unreachable, and exploitation evidence always raises priority. | Good answer: Per-action automation levels, snooze as well as ignore, and an AI that can raise severity, not only lower it. |
Your stack
reflection, dependency injection and dynamic dispatch in our languages?
container images, or only direct dependencies? | Evidence for auditors
reasoning and timestamp? |
Good answer: A specific answer for your stack, demonstrated on one of your own repositories during the trial. | Good answer: A complete, reviewable decision log. No verdict without a rationale a human can check. |
The AI’s own attack surface
where is it processed, and is it ever used for training?
code. How is it protected against instructions planted in that code? | Data qualityWhich vulnerability scores and sources do you use beyond NVD, and how do you handle CVEs that were never enriched? |
Good answer: Inference-only processing, clear data residency, and a concrete answer on prompt injection rather than a shrug. | Good answer: Several independent sources. Enriching with EPSS. |
Related compliance requirements
|
What about malware?
The term SCA covers more than “known vulnerabilities”: it also concerns malicious packages, license issues and end-of-life components. Malware in particular is growing fast. AI coding agents install dependencies with less human oversight, and attackers register package names that LLMs tend to hallucinate (“slopsquatting”).
A malicious package usually has no CVE and no severity score. It often executes at install time, before any of your code touches it, so reachability is irrelevant. Such dependencies need immediate blocking, on a different level, usually at the registry and install layer: package firewalls, install-time scanning, and quarantine windows for freshly published versions. This is a subject that requires its own treatment. I will cover it in another blog post, together with free solutions like Safe Chain or more sophisticated like Aikido Device Protection.
End-of-life components, on the other hand, belong to this story. An EOL library gets no fixes, so a reachable CVE in it can’t be solved by upgrading. Knowing which findings are actually exploitable matters then even more.
Summary
CVEs were always a pain point in the secure software development life cycle. In the age of AI they have become extremely difficult to manage manually. Countering circumstances caused by AI with tools built around AI seems to be the only reasonable direction, and soon will become non-negotiable.
A step further - AI-driven SDLC
At VirtusLab, we built Visdom - a full toolkit for building software using AI. As software delivery experts, we know exactly how to put together the entire process efficiently, orchestrating LLMs, building context and properly governing and tracing AI usage. Visdom Security is the component responsible for scanning vulnerabilities in all the SDLC stages, powered by Aikido. Reach out and learn how your organization can deploy Visdom and turn the development process into a modern, secure, controllable, robust AI-driven one.

