GenNXT · AI & emerging technology

AI adds attack paths a scanner can’t see.

An LLM application can be talked into leaking data or taking actions it should not. We test the prompt, the model, the tools it calls and what it returns, then help you govern AI safely.

  • OWASP Top 10 for LLMs

    AI applications tested against the OWASP categories for LLM risk, from prompt injection to excessive agency.

  • Mapped to MITRE ATLAS

    Attack techniques against AI systems described in the language your threat team already uses.

  • ISO 42001 and NIST AI RMF

    Governance built on the recognised frameworks for managing AI risk.

  • EU AI Act readiness

    Know which obligations apply to your AI systems and what evidence you will need.

What changes with AI

Three assumptions AI breaks

Security teams already know how to test inputs, dependencies and users. AI changes what each of those means.

Input is a form field

The prompt is an input

Anything the model reads, including documents and web pages it retrieves, can carry instructions. It has to be tested like any untrusted input.

Dependencies are libraries

The model is a dependency

Models, datasets and plugins come from third parties. They belong in your supply chain review, with their provenance known.

Users are people

The agent is a user

An agent that can call tools acts on your systems. Give it least privilege, require approval for risky actions and log what it does.

Services

Secure what you are building next

Test and govern AI, design architecture that limits damage, and keep validating as things change.

  1. Area 01

    Test AI

    Attack your AI applications before someone else does.

  2. Area 02

    Govern AI

    Decide who may use AI, for what, and with which data.

  3. Area 03

    Modern architecture

    Designs that assume a breach and limit what it can reach.

  4. Area 04

    Continuous validation

    Keep testing, because your attack surface keeps changing.

Questions

Frequently asked questions

What is LLM security testing?
It is a penetration test aimed at an application built on a large language model. We test how the model, its prompts, the data it retrieves and the tools it can call respond to hostile input, using the OWASP Top 10 for LLM Applications as the baseline.
Why can a normal web application test not find these issues?
A web scanner looks for known patterns such as SQL injection. An LLM application can be manipulated in plain language, through a document it reads or a tool it is allowed to call. Finding that takes testers who understand how the model and its integrations behave.
What is prompt injection?
Prompt injection is input that makes the model ignore its instructions and follow the attacker instead. It can be typed directly, or hidden in content the model retrieves, such as a web page, email or uploaded file.
What is excessive agency?
It is when an AI agent has more permissions, tools or autonomy than its task needs. If it is manipulated, it can then take harmful actions, such as sending data or changing records, without a person approving them.
Which frameworks do you use for AI governance?
We build AI governance on ISO/IEC 42001 and the NIST AI Risk Management Framework, and map obligations under the EU AI Act where it applies to you. MITRE ATLAS is used to describe attack techniques against AI systems.
Does the EU AI Act apply to companies outside the EU?
It can. It applies to providers placing AI systems on the EU market and to deployers whose AI output is used in the EU, wherever they are based. An assessment tells you which of your systems are in scope and at which risk level.
We only use a third-party AI service. Do we still need testing?
Yes. The provider secures its model, but your prompts, the data you connect, the tools you expose and how you handle the output are your responsibility. Most AI application issues sit in that integration layer.
What is continuous security validation?
It is testing that runs regularly rather than once a year, using breach and attack simulation and attack surface monitoring. It shows whether your controls still work as systems, cloud resources and threats change.

Launching an AI feature? Test it first.

Tell us what your AI application can read and what it can do. We will show you how an attacker would try to misuse it.