Many security teams want AI to run a pentest, but they cannot send source code, prompts, or findings to a cloud model. Policy or regulation keeps that work on-prem, meaning the software and the data stay inside the building, which rules out tools that call an outside API for every test.
Local inference is the other half of that job: the AI actually runs on hardware the buyer controls, not on a vendor’s servers.
In this guide, we compared 7 on-prem and self-hosted AI pentest options for organisations that need that constraint, including air-gapped setups with no public internet and white-box tests that can see source, not only the live app.
Quick Overview
|
Tool |
Best for |
Local inference |
What it pentests
|
|
Aikido |
Organisations that need local AI inference and cannot send source code or findings to cloud models |
Yes: models run on a 4U GPU server in the data center |
White-box application pentest from source, docs, and runtime, with working exploits and merge-ready pull requests |
|
Zero Hunt |
Teams that want a racked red-team appliance with its own models and no cloud callbacks |
Yes: ZeroHunt Apex models run on the appliance GPU |
Generative network pentest, traffic analysis, and compliance evidence |
|
RidgeBot |
Teams that want automated black-box pentest as on-prem software |
On-prem software on bare metal or VMs (cloud also available); not a dedicated local-inference GPU appliance |
Agentless black-box internal, external, and lateral-movement tests with exploit proof |
|
Pentera Core |
Teams validating exploitable paths on the internal production network |
On-prem network testing with an AI copilot; not the same as all inference staying off cloud models |
Internal network pentest: vulns, misconfigs, identities, credentials, ransomware emulation |
|
Horizon3 NodeZero |
Teams that want autonomous exploit chains with a runner on their network |
Runner can sit on-prem; control plane and results live in Horizon3’s cloud |
Internal tests from a Docker host or OVA; external tests from Horizon3 cloud |
|
Strix |
Teams that will self-host open-source agents and bring their own model |
Possible with Ollama or vLLM; most local models struggle with agentic pentest |
Autonomous app pentest with proven exploits and merge-ready pull requests |
|
PentAGI |
Teams building a self-hosted agent stack and wiring a local model themselves |
Optional via Ollama; default path is cloud LLM APIs |
Autonomous pentest agent with Kali tools, browser, editor, and internet search |
Key Criteria to Consider
We ranked the list on the following things that matter when a pentest has to stay inside the building.
- The first is local inference, which just means the AI that plans and runs the test sits on hardware the team actually controls, instead of calling a model that lives in someone else’s cloud.
- The second is the data path: source code, prompts, and findings should not leave the network, because that is why some teams cannot use a typical AI pentest product in the first place.
- The third criterion worth your attention is whether the tool is really pentesting, meaning it proves an exploit, not only draws a map of paths that might work.
- Finally, check whether it can run without the public internet, the air-gap case, where there are no outbound calls and updates come on a USB or through a tightly limited channel.
#1. Aikido
Aikido Machine is the on-prem pentest setup we put first. It is a 4U GPU server in the data center, and the pentest models run on that hardware, so source, docs, runtime, findings, and telemetry stay on the network instead of going out to a cloud model. That local view of the code is what white-box testing actually means here: the tool can see how the app is built, not only what shows up on a live endpoint.
Additionally, it can work without internet access. Updates come through one allowed domain, on an encrypted USB when the site is fully air-gapped, or from a separate machine used only for those pulls. An Aikido engineer comes on site to plug it in, join it to the network, and start the first pentest.
Tests can run on an ongoing basis across apps and releases. A harness handles login, sessions, a proxy, and scope, so the run stays inside the targets the team set. Findings come with a working exploit, a pull request that is ready to merge, and a retest on the same scope. AI pentesting and AI Code Analysis both run on the machine, which is the setup we used for teams that have to keep this work on their own hardware, including banks, defence, and healthcare.
Pros
- 4U GPU server in the buyer’s data center; models and inference stay on that hardware
- Source, docs, runtime, findings, and telemetry stay inside the network
- Air-gap with USB updates, or one whitelisted domain for updates only
- White-box access to source, docs, and runtime
- Continuous pentests of every application on every release
- Working exploits, ready-to-merge pull requests, and retest on the same scope
Why Choose Aikido
We ranked Aikido first because this job needs the AI to run on hardware in the data center, and Aikido Machine is that server, not a cloud pentest with a runner dropped on the network.
The models run on GPUs the buyer controls, so source and findings stay on the network instead of going to a cloud model. It still has to prove the hole with a working exploit, and it can run without the public internet when the site is air-gapped.
A 4U box still needs space, power, and network rules, and the team still reviews the pull requests it opens. We would pick it for organisations that need that local setup and cannot send source or findings to a cloud model.
#2. Zero Hunt
Zero Hunt is an on-premise AI security appliance the buyer racks, with its own ZeroHunt Apex models running on the appliance GPU. There are no cloud callbacks, no telemetry, and no external LLM APIs, and air-gapping is supported.

The same box covers generative pentest, traffic analysis, and compliance evidence, so the day-to-day job is a network and red-team appliance rather than a white-box application pentest from the buyer’s source with merge-ready pull requests.
Pros
- Rack appliance with GPU inference on site
- Own ZeroHunt Apex models; no external LLM APIs
- No cloud callbacks and no telemetry
- Air-gap support
- Generative pentest plus traffic analysis and compliance evidence
#3. RidgeBot
Another AI-powered automated pentest software that finds issues, exploits them with proof, and writes the results up is RidgeBot. On-prem, it is a software package on bare metal or VMs, and a cloud deployment is available too, with agentless black-box coverage for internal tests, external tests, and lateral movement. RidgeBrain, the dual AI engine, is part of that package. On-prem keeps findings local as a software install, but it is not a dedicated local-inference GPU appliance for white-box source pentest.

Pros
- Automated pentest that exploits with proof
- On-prem as software on bare metal or VMs (cloud also available)
- Agentless black-box: internal, external, and lateral movement
- Dual AI engine / RidgeBrain
- Documents validated risks for the security team
#4. Pentera Core
Pentera Core uses AI-driven pentest on internal production networks, chaining vulnerabilities, misconfigurations, identities, and credentials into full attack paths. An agentic AI copilot helps manage those tests in real time.

Pros
- AI-driven pentest of internal production networks
- Chains vulns, misconfigs, identities, and credentials
- Agentic AI copilot to manage tests
- Black box, grey box, ransomware emulation, AD password, OWASP Top 10, CISA KEV
- Re-run paths to confirm fixes
#5. Horizon3 NodeZero
Then, the next tool is NodeZero; it chains weaknesses, shows proof, gives remediation guidance, and offers Quick Verify after a fix.

Internal tests start from a Docker host or an OVA on the network, while external tests run from Horizon3’s cloud. The buyer keeps an account in Horizon3, and dedicated short-lived resources spin up in an isolated VPC for the run. The runner can run on-prem, but the control plane and results live in Horizon3’s cloud, so code and findings are not fully local-inference air-gapped.
Pros
- Autonomous exploit chains with proof
- Remediation guidance and Quick Verify
- Internal tests from a Docker host or OVA
- External tests from Horizon3 cloud
- Short-lived resources in an isolated VPC
#6. Strix
Strix is an open-source agent path the buyer self-hosts. It proves exploits and can open merge-ready pull requests, and it runs on the buyer’s hardware with a model they bring, including Ollama or vLLM. Most local models struggle with agentic pentesting, especially under 70B parameters, and cloud models remain the stronger fit when privacy is not the absolute priority. It is not a turnkey 4U appliance with models included.

Pros
- Open-source autonomous agents
- Proven exploits and merge-ready pull requests
- Self-host on buyer hardware
- Air-gapped possible with a local model and a local target
#7. PentAGI
PentAGI lands as a Docker stack: a self-hosted autonomous AI pentest agent with Kali tools, a browser, and an editor. It supports more than twelve LLM providers, including OpenAI, Anthropic, Gemini, Bedrock, and Ollama for local inference.

The default path is still a cloud LLM API unless the buyer wires Ollama, and internet search is part of the design. That is not a white-box application pentest appliance with merge-ready pull requests.
Pros
- Self-hosted autonomous pentest agent
- Docker with Kali tools, browser, and editor
- Twelve-plus LLM providers, including Ollama
- Local inference available when Ollama is wired
- Internet search is part of the workflow
Final Choice
The constraint that matters on this list is local inference: the AI has to run on hardware the organisation controls, and source code plus findings cannot go to a cloud model.
When that is the rule, including an air gap, the product that is actually built as that server is Aikido Machine. The others still have a place, as a network pentest on-prem, as a runner with a cloud control plane, or as a do-it-yourself local LLM, but they do not replace a box whose models stay on the local network.

