Every cloud security team we talk to has the same quiet problem. The tools work. Security Hub, Config, GuardDuty, Inspector, and Access Analyzer all do their jobs, and between them they surface more signal than a small team could hope to have five years ago. That is exactly why the team is drowning. The scanners never stop, the findings never stop, and the number of people available to read them has not changed at all.
The industry answer for the last two years has been to add intelligence at the point of triage. Put a model on the alert, let it summarize, let it rank, let it suggest a fix. That helps. We build in that direction ourselves, and our AI Analyst does deterministic scoring with model narration today. But there is a ceiling to a single agent reading one finding after another, and we have started to hit it. So we are building something different alongside it. We call it the audit swarm, and this is why.
The noise is a structural problem, not a tuning problem
It is tempting to treat alert fatigue as a configuration issue. Tighten the rules, suppress the noisy checks, snooze the categories nobody acts on. Teams do this constantly, and it buys a few weeks. Then a new service ships, a new account is onboarded, a new conformance pack turns on, and the backlog refills.
The reason tuning never wins is that the work is genuinely large. A modern AWS estate has many accounts, several environments per account, and half a dozen distinct security domains inside each one: public exposure, identity and privilege, encryption, network posture, logging and audit coverage, and credential hygiene. Reading a finding well means holding the context of its domain and its environment at the same time. A public bucket in a sandbox is a shrug. The same bucket in a regulated production account is a Friday night. The judgment lives in that intersection, and the intersection is exactly what a long flat queue destroys.
So the problem is not that any one finding is hard. It is that there are thousands of them, each needs a small amount of the right context, and there is one queue and one reader.
One agent is still one reader
The obvious move is to put a model on the queue. We did, and we stand behind it. Our AI Analyst scores each finding deterministically, clusters findings into themes, assigns a disposition of act now, schedule, or monitor, and lets a model narrate the reasoning in plain language. Crucially the model never sets the priority. The math does. The model explains.
That design solves the trust problem. It does not solve the throughput problem. A single agent working a queue is faster than a person working the same queue, but it is still one worker pulling one item at a time. When you point it at an organization with dozens of accounts, it either serializes through everything, which is slow, or you widen the batch and lose the per-domain context that made the triage good in the first place. You end up back at the same tradeoff, just at a higher speed.
This is the same wall the rest of the industry hit in 2026. Amazon shipped a GuardDuty investigation agent that can assess a finding, an account, or an organization and return a risk level with a MITRE ATT&CK mapping in minutes. It is genuinely useful, and it is still fundamentally a single investigator you point at one scope at a time. The frontier is no longer whether a model can triage. It clearly can. The frontier is coordination.
What a swarm actually is
A swarm is not a bigger agent. It is many small agents, each deliberately narrow, running at the same time, coordinated by something that decides who runs, when, and with which model.
Our model for cloud auditing is one specialist per domain, per environment. The identity auditor for the production account is a different worker from the identity auditor for staging, and both are different from the encryption auditor for either one. Each agent carries only the rules and the context for its own square of the grid, which is what keeps its judgment sharp. Instead of one reader trying to hold the whole estate in its head, you get a fleet of readers who each hold exactly one well-defined piece.
Then the important part happens. The swarm does not hand you a hundred separate reports. It synthesizes. Every agent’s output flows back through the same deterministic scoring, gets deduplicated across environments, and collapses into a single ranked queue. The operator sees one prioritized view of the entire estate, not a stack of per-account PDFs to reconcile by hand.
The operator controls concurrency and model, on purpose
The two dials that matter most in a swarm are how many agents run at once and how capable each agent needs to be. We put both in the operator’s hands rather than hiding them.
Concurrency is a cost and blast-radius control. A team running an interactive triage on a quiet afternoon might fan out wide and get a whole-estate read in a minute. A team running a scheduled overnight sweep might cap concurrency low to stay gentle on rate limits and keep spend predictable. The point is that the same swarm can be tuned to the moment instead of being one fixed speed.
Model choice is where the economics get interesting, and it maps cleanly onto where the industry has landed. Not every square of the grid deserves the same horsepower. A fast, inexpensive model is the right tool for high-volume, well-structured findings, and it is where most of the grid should run. A stronger workhorse model handles the domains where the reasoning is denser. The most capable models are reserved for the hardest synthesis, the cross-domain correlation where a modest finding in one place plus a modest finding in another add up to a real attack path. Letting the operator route models per run means the swarm spends compute where judgment is scarce and saves it where the answer is obvious.
The guardrail scales with the swarm
The faster and more autonomous a system gets, the more its safety has to be built into the structure rather than bolted on afterward. This is where a lot of agentic security tooling worries us. An agent that both decides a resource is dangerous and holds the credentials to change it is a very efficient way to take down a production workload on a confident wrong answer.
Our contract does not change when one agent becomes a hundred. The math computes, the model narrates, the human decides. Every agent in the swarm scores its findings deterministically, so priority is reproducible and never invented by a model under pressure to sound sure. Any currency figure a model emits is audited back to a computed fact before it is ever shown, and a figure that does not reconcile is flagged rather than presented as truth. When a model hallucinates a finding or a resource that the deterministic layer never produced, that output is dropped on the way back, not merged into your queue.
And the whole system stays read-only. CloudDefender connects through a role that can look but not touch. When a finding deserves a fix, the swarm produces a ready-to-deploy artifact, a CloudFormation template, a Terraform module, a service control policy, or a CLI command, that you launch in your own account after you have read it. The swarm can cover an entire multi-account estate in parallel and still not have the ability to change a single resource. That is deliberate. Coverage should scale. Authority should not.
Why now
Three things converged in 2025 and 2026 that make this the right moment to build a swarm rather than a faster single agent.
The first is that agentic patterns finally became reliable enough to coordinate. Running many narrow agents and merging their work is no longer a research demo. It is a pattern with real operational discipline around it, and cloud providers are shipping first-party versions of the individual pieces.
The second is model routing. A year ago, using a different model for different work meant gluing together different vendors. Now a single family spans a fast tier, a workhorse tier, and a frontier tier, and choosing between them per task is a normal engineering decision. A swarm that routes models per agent was awkward to build before and is straightforward now.
The third is that the noise got worse, not better. Every improvement in cloud security scanning adds signal, and signal without triage capacity is just a larger backlog. The teams who feel this most acutely are the lean ones, the security team of three covering forty accounts, and they are precisely the teams that cannot hire their way out. A swarm is leverage for exactly those teams.
What is live today, and what we are building
We want to be precise about this, because the difference matters.
Live today is the foundation the swarm stands on. The AI Analyst performs deterministic scoring, ranking, clustering, and disposition over your AWS findings, with model narration on top and the faithfulness guardrail underneath. Per-finding remediation guidance and ready-to-deploy fix artifacts are live. The read-only connection model is live. All of it honors the math computes, model narrates, human decides contract.
What we are building, and what is not in the current release of CloudDefender.io, is the swarm itself: the parallel fan-out of one specialist per domain per environment, the operator controls for concurrency and model routing, the multi-account and eventually multi-cloud scope, and the synthesis of every agent’s work into one ranked queue. You can see the direction taking shape as a prototype inside the product, and the screenshots in this article are from that prototype. It runs on sample data while we build the real fan-out behind it, and it is labeled a prototype in the product so no one mistakes the preview for a shipped capability. To be unambiguous: you cannot run an audit swarm against your own account today. When it ships, we will say so plainly.
The shape of the bet
The last decade of cloud security was about seeing more. That fight is largely won. The scanners see almost everything, and the result is that seeing more is no longer the constraint. Acting on what you see is.
A swarm is our bet on how lean teams close that gap without pretending a model can be handed the keys. Cover the whole estate at once by running many narrow specialists in parallel. Route the right model to the right work so compute lands where judgment is scarce. Fold everything into one ranked queue so a person sees priority, not volume. Keep the math in charge of what matters most, keep the connection read-only, and keep the final decision human.
That is why we are building an audit swarm. Not to replace the analyst, but to give the analyst a fleet.
Sources and reporting notes
- AWS: introducing the Amazon GuardDuty investigation agent, July 20, 2026. Consulted for the finding, account, and organization scopes, the risk-level and MITRE ATT&CK output, and the minutes-scale assessment referenced above.
- Claims about CloudDefender describe capabilities that are live in the product as of publication: deterministic scoring and ranking in the AI Analyst, per-finding remediation guidance, ready-to-deploy fix artifacts, and the read-only connection model.
- The audit swarm itself is an in-product prototype running on sample data and is labeled as such inside the product. This article describes it as work in progress and the direction we are building toward, not a shipped capability.
CloudDefender is a read-only AWS security platform. It surfaces misconfigurations and findings, triages and ranks them with deterministic scoring and audited AI narration, and hands back ready-to-deploy remediation artifacts you launch in your own account. It never gets a write role to your resources and never applies changes itself.