VerdictRuntime
Application & AI security assessment

We only report what we proved.

Most scanners hand you a queue. Yours fills with rule matches nobody has confirmed, and the real issue sits on page four behind forty things that were never exploitable. Verdict Runtime is built the other way around: a finding has to earn its severity before it reaches you, it tells you exactly what evidence it has — and where we can fix it, we hand you the fix with proof that it worked.

Confirmed High Example finding
SQL injection in product sort parameter
CWE-89 · app/routes/products.py:142 · GET /api/products?sort=
Agreement
Flagged independently by two static analysis engines
Reachable
Route resolves to this handler; parameter reaches the query unsanitised
Proof
Boolean-difference confirmed live in an egress-locked sandbox — ' AND 1=1-- returned 24 rows, ' AND 1=2-- returned 0
Remediation
Parameterised-query patch prepared; re-scan clean and your test suite still green

Illustrative example of the report format, not a real customer finding.

Demo

Two minutes in the product

Run against deliberately vulnerable open-source training applications: a finding verified live against the running app, the fix prepared for it, and a vulnerable dependency patched. Account details and repository names are blurred.

How a finding earns its place

Three levels of evidence, labelled on every finding

Nothing is silently promoted. Each finding carries the reason it holds the rank it does, so triage starts from evidence instead of a rule identifier.

SINGLE_TOOL

One engine saw it

Reported, ranked below corroborated findings, and marked as uncorroborated. Useful signal, presented as exactly that rather than as a confirmed vulnerability.

HIGH

Two independent engines agree

Dependencies, source code and infrastructure are each covered by two engines built by different teams, and findings are correlated across them. Agreement between two independent implementations is a materially stronger signal than either one alone, and it is the single most effective filter we have against false positives. Where only one engine exists for a class, its findings stay labelled as single-engine.

CONFIRMED

Reproduced against a running target

For the classes where a safe proof exists — injection, cross-site scripting, open redirect, XXE, command injection — a hand-written, reviewed proof runs against your deployed application inside a network-isolated sandbox whose egress is locked to your target alone. If it does not reproduce, it does not get called confirmed.

Coverage

One normalised report across the whole surface

Static analysis, dependencies, infrastructure, the running application and your AI features all land in a single schema with consistent severity, location and confidence, then get deduplicated and correlated — so you read one report, not six.

Dependencies & supply chain

Known-vulnerable packages, whether the affected code is actually reachable from your entry points, and a signed software bill of materials you can attach to a release.

Reachability · SBOM · signing · VEX

Source code

Injection, authentication, access-control and cryptography defects, cross-correlated between two engines before anything is ranked high.

~20 in-house rules for gaps the market misses

Secrets, infrastructure & pipelines

Committed credentials, cloud and Kubernetes misconfiguration, and build-pipeline weaknesses that let a workflow be hijacked.

Cloud posture · container images · CI workflows

Running application

Authenticated dynamic testing against a deployed environment, including broken object-level and function-level authorisation that only appears when two real accounts are compared.

20 targeted checks · OWASP API1–10 · A01–A10

AI & LLM features

Live adversarial testing of the model-backed parts of your product, covering the full OWASP Top 10 for LLM Applications plus emerging agentic threats.

21 checks · LLM01–LLM10 · agentic

Remediation

Fixes prepared as reviewable patches, each re-scanned and run against your own test suite before it reaches you.

Dependency · code · infrastructure · pipeline
AI security

Your AI features tested the way an attacker would

Most security tooling has no model of what an LLM feature can be talked into doing. These checks send real adversarial traffic at your live endpoint and judge the response against something concrete — a planted canary, a marker string, a tool name that should never have been reachable — rather than asking a model whether the answer looked unsafe.

OWASPCheckWhat a finding proves
LLM01Prompt injectionA crafted instruction overrode the system prompt and returned an exact attacker-chosen marker.
LLM01Indirect prompt injectionAn instruction planted in fetched content — not in the user's message — was followed.
LLM01Multi-turn escalationDefences that hold on a single message fail across a real multi-turn conversation.
LLM02Sensitive information disclosureA canary secret planted in your own data came back in a model response.
LLM03Plugin and tool enumerationThe model disclosed the real integrations it can reach when asked to list them.
LLM04Persistent data poisoningContent written by one account changed what a different account was told, on a fresh session.
LLM05Improper output handlingModel output reached the page or a downstream system without escaping.
LLM06Excessive agencyThe model invoked a sensitive tool outside the user's stated intent.
LLM07System prompt leakageA known verbatim fragment of your own system prompt was recovered.
LLM08Cross-tenant retrieval leakageRetrieval crossed a tenant boundary and returned another tenant's canary.
LLM08Knowledge-base poisoningA document submitted through your own product changed later answers after sync.
LLM09Hallucination under pressureThe model produced confident technical detail about an entity that does not exist.
LLM10Unbounded consumptionOversized input and long conversations were accepted with no effective cap.
AgenticAgent code executionThe agent ran an attacker-supplied command through its own execution tool.
AgenticTool parameter injectionAn authorised tool was re-invoked with an out-of-scope parameter value.
AgenticAgent-to-agent trustA message claiming another agent's identity was accepted without verification.
AgenticIdentity and guardrail erosionSustained reframing across turns moved the model off its assigned identity and limits.

Seventeen of twenty-one checks shown. Every one requires your written authorisation before it sends a single request.

AI-driven dynamic testing

An agent that probes your business logic, not just your code

Some flaws are invisible to any rule because nothing about them is syntactically wrong. A coupon that can be redeemed twice, a cart that accepts a negative quantity, a checkout that can be replayed at yesterday's price — the code is fine, the logic is not.

It learns your application first

A real browser session walks your product the way a user does, capturing every API call the pages actually make — including the ones absent from your documentation. Captured calls are linked into a dependency graph, so a request needing a real order identifier gets one that genuinely exists.

Then it forms and tests hypotheses

A model is given the real captured traffic and asked what the business rules appear to be and how they might be broken. It sends real requests, reads the real responses, and adapts — a loop closer to a tester probing your app than to a scanner replaying a payload list.

Deterministic rules where you want certainty

You can also state invariants directly — a total must equal quantity times unit price, a balance must never go negative — and have them checked against live responses. A violation is a confirmed finding, not a judgement call.

Uncertainty is labelled, never hidden

Model-driven results are marked as needing review rather than promoted to confirmed. A machine that suggests a lead is useful; a machine that reports its own guesses as facts is the reason your queue is full.

Remediation

A fix, and the evidence it worked

Reporting a problem is the easy half. For dependency, code, infrastructure and pipeline findings, we prepare the change and then prove it — because an unverified fix is just another claim.

STEP 01

Prepare the change

A dependency is moved to the lowest version that actually resolves the issue, with the upstream changelog summarised so you can see what else moves with it. Code, infrastructure and workflow fixes come as a diff against the real file, scoped to the finding.

STEP 02

Re-scan to confirm

The same analysis runs again against the patched tree. The finding has to be gone, and no new finding may appear in its place. This step is deterministic — no model is asked whether the fix looks correct.

STEP 03

Prove nothing broke

Your own test suite runs against the patched code. If a fix turns a vulnerability into an outage, that is not a fix, and you find out from us rather than from production.

Nothing is committed, pushed or merged on your behalf. Every change arrives as a reviewable diff or a pull request you open yourself, with the verification results attached — and any fix we could not verify is handed over labelled as exactly that.

Authorisation

Nothing runs without your written scope

Every agent that can send a single packet at a target checks it against a scope file you write and sign off. The check is compiled into the tool, not a setting — no scope entry means the run exits before it touches a key or a report. There is no permissive mode.

  • ›
    Reading your source needs only a repository entry.
  • ›
    Live traffic needs a separate, explicit dynamic authorisation.
  • ›
    Logging in as a real account needs a further credentialed grant.
  • ›
    Anything that writes data requires a disposable-target attestation, and is never assumed from the tiers above.
.pentest-scope
# reviewed and committed by your team
https://github.com/acme/api
dynamic:https://staging.acme.com
credentialed:https://staging.acme.com
# production is deliberately absent
Why trust the method

The discipline is checkable

The same standard applied to the product is applied to its own claims. We test this tool by fixing real security bugs in open-source projects that are not our customers, owe us nothing, and review the work in public. Four commitments we hold ourselves to.

Contributed upstream
We publish detection rules into the open-source ecosystem rather than keeping every improvement private. Two rules have been submitted back to the community rule sets the industry already relies on; both are still open for review, and we describe them as submitted rather than accepted until a maintainer says otherwise.
Ships fixes, not just findings
We have opened nine pull requests that patch real CVEs in active, real-world open-source projects — not synthetic benchmarks, and not projects that pay us. One has been reviewed and merged upstream by a project's own maintainers. A second was closed because the maintainer read it, agreed, and shipped the same dependency fix within minutes of receiving the pull request — the vulnerable versions are gone from that project either way. Seven remain open. Those two outcomes are a harder test than any benchmark score: someone outside this company had to read the reasoning and agree the fix was correct.
Validated against real code
New detections are proven on production open-source projects — more than fifty of them — with positive and negative controls in the same run, so a rule that fires on nothing is treated as broken rather than clean. The live-traffic, cloud, cluster and container checks are exercised against genuinely running systems: a real cluster, a real cloud account, a real built image, never a mocked response.
Limitations written down
Known gaps are documented rather than hidden, and no result is described as validated unless the validation was actually run.
Pricing

Priced per engagement, quoted up front

Every application is different, so we quote rather than publish a rate card. You get the quote before any work starts, and the first assessment is free.

What sets the price

How many repositories and applications are in scope, whether live and logged-in testing is needed, and whether your product has AI features to test.

What every engagement includes

The report with evidence on every finding, the fixes we can prepare with their verification results, and a written note on anything we could not check.

The first assessment is on us

One repository, one report, no obligation. You will get the findings, the evidence behind each one, the fixes we can prepare, and an honest note on what could not be checked.

Request an assessment See a sample report Reply within two business days.
We never scan first. Nothing is tested until you have named the target in writing and confirmed you are authorised to have it tested. If you would rather start on an open-source repository you maintain, that works too.