Every security team feeling the weight of alert volume is being sold the same cure this year. Put a large language model on the queue, let it read the findings, and let it tell you what matters. The demos are genuinely impressive. The model summarizes a wall of alerts into a paragraph, ranks them, and explains its reasoning in fluent prose. For a team drowning in signal, it is intoxicating.
Then the auditor arrives, or the incident review starts, and asks the question that deflates the whole thing. How do you know the model did not make this up? Why was this finding ranked critical and that one ignored? Can you reproduce that decision, or did the ranking depend on the mood of a stochastic model on a Tuesday? For a regulated business, an unanswerable version of that question is not a rough edge. It is a reason not to deploy the tool at all.
This is the real barrier to AI in the security operations center, and it is not capability. Models are more than capable of triaging cloud findings. The barrier is trust, specifically the kind of trust you can defend to an auditor, a regulator, or a board. The good news is that this is an architecture problem, and it has a clean solution. You can have the fluency of a model and the defensibility of a deterministic system at the same time, if you are disciplined about which one is allowed to do what.
The mistake is letting the model decide
Most AI triage tools make one architectural choice that quietly forfeits the audit trail: they let the model set the priority. The finding goes in, the model reads it, and the model outputs the rank. That single decision is what makes the whole thing indefensible.
A model asked to rank is doing something fundamentally different from a model asked to explain. Ranking is a decision, and when a model makes it, the decision inherits every property you do not want in a security control. It is not reproducible, because the same input can yield a different order on a different run. It is not inspectable, because there is no rule you can point to, only a distribution the model sampled from. And it is not trustworthy under adversarial conditions, because a model under pressure to sound confident will produce a confident ranking whether or not it has grounds for one.
You cannot audit a vibe. When the reviewer asks why finding A outranked finding B, “the model felt it was more important” is not an answer that survives a compliance conversation. And it should not, because it means you have no idea whether the ranking reflects real risk or a fluent hallucination.
Math computes, the model narrates, the human decides
The fix is a division of labor we hold to strictly. Every decision that has to be defensible is made by deterministic code. The model is allowed to do exactly one thing: put that decision into words.
In practice this means the backend, not the model, computes each finding’s priority. A score is built from components you can write down and reproduce: a base weight from the finding’s severity, additional weight for specific risk signals matched against the finding’s own text, such as public exposure, excessive privilege, or a logging gap, and a modifier for the finding’s status. From those scores the code produces the rank, groups findings into themes, and assigns each one a disposition of act now, schedule, or monitor. All of it is pure arithmetic over the real finding data. Run it twice on the same input and you get the same answer, every time.
Only then does the model enter, and only to narrate. It writes the overall summary, the per-theme detail, and the plain-language reason a given finding sits where it does. It makes the output readable. It does not make the output true. The priority was already set by the math before the model wrote a single word, so no amount of eloquence can change what matters. The model explains the decision. It never gets to make it.
This is not a limitation grudgingly accepted. It is the point. The reproducibility that makes a security control auditable comes from the deterministic layer, and the clarity that makes it usable comes from the model, and because the two jobs are cleanly separated you get both instead of trading one for the other.
Verify what the model says before anyone sees it
Splitting the work is necessary but not sufficient, because the model still produces text, and text can contain claims. A narration can assert a dollar figure, name a resource, or reference a finding. Any of those can be wrong, and a plausible wrong claim is more dangerous than an obvious one because it slips past a busy reader. So the last stage before a human sees anything is verification.
Two checks matter most. The first is numeric faithfulness. Every currency figure the model emits is audited back against the computed facts before it is shown. If the model writes a number that reconciles with a real calculated value, within a small tolerance, it passes. If it writes a figure that matches nothing the deterministic layer produced, that figure is flagged rather than presented as truth. A model cannot smuggle a fabricated number into a report simply by stating it confidently, because a number that does not trace to a fact does not get to wear the costume of a fact.
The second check handles invented structure. When the model references a finding or a resource that the deterministic layer never produced, that reference is dropped on the way back rather than merged into your queue. The narration is stitched onto the deterministic skeleton by exact identifier, so a hallucinated finding has nothing to attach to and falls out. Where the model failed to say something useful, the output falls back to the deterministic rationale, so the result is always complete. What you read is the real, computed set of findings, described by the model only where the description checks out.
The effect is that the model’s creativity, the thing that makes it fluent and also the thing that makes it hallucinate, is contained. It can help. It cannot invent. And the report a human acts on is demo-safe and audit-safe by construction, not by hoping the model behaved.
What this buys the people who have to sign off
For the analyst, this is the difference between a tool they can lean on and a tool they have to double-check. Because the priority is deterministic, they can trust the order of the queue without re-deriving it. Because the prose was verified, they can read the explanation without wondering whether a number was invented. The AI removes the reading burden without adding a fact-checking burden, which is the whole reason to adopt it.
For the security leader, this is the answer to the governance question. When the board or an auditor asks how you use AI in security operations, you can describe a control where the model never makes a decision, every figure is checked against computed truth, and anything the model invents is discarded before a human sees it. That is a defensible posture. It is the difference between “we let an AI triage our alerts and we trust it” and “our triage is deterministic and reproducible, and the AI only writes the explanations, which we verify.” One of those statements ends the meeting. The other starts a much longer one.
And for the business, it removes the false choice. You do not have to pick between the fluency that makes AI useful and the rigor that makes it safe to deploy in a regulated environment. The architecture gives you both, because it assigns each to the layer that can actually be trusted with it.
The principle generalizes
None of this is specific to security findings. Any place you want to use a language model on decisions that have to be defended, the same three-part discipline applies. Compute the decision deterministically, so it is reproducible and inspectable. Let the model narrate, so the output is clear. Verify every checkable claim against the computed facts before anyone acts on it, so the model’s fluency can never become the source of a false statement.
The teams that will get real, durable value from AI in operations are not the ones that hand the model the most authority. They are the ones that give it the least authority compatible with usefulness, and build the verification to make even that safe. In security especially, the model should be the most articulate member of the team and the one with the least power to decide anything on its own. That is not a lack of ambition. It is what makes the ambition survive contact with an auditor.
Sources and reporting notes
- This article describes CloudDefender’s AI Analyst as of publication: deterministic scoring, ranking, clustering, and disposition computed in code over the customer’s own findings, with a language model generating only the narrative layer, and a faithfulness check that audits currency figures against computed facts and drops references the deterministic layer did not produce. The “math computes, model narrates, human decides” contract is the product’s stated design principle.
- The specific scoring weights and tolerances referenced are illustrative of the approach and may change as the product evolves; the invariant is that priority is computed deterministically and the model does not set it.
- No specific competing product is named or characterized. The pattern of “let the model rank the queue” is described as a general industry approach, not an account of any particular tool.
CloudDefender’s AI Analyst computes finding priority deterministically, uses a language model only to explain the result in plain language, and audits every figure the model emits against a computed fact before you see it. The math sets what matters, the model makes it readable, and you make the call.