AI Pen-Testing ConferenceNov. 3, 12-3pm ETSave your seat →

The AI Pen-Testing
Toolbox

AI tools and resources for pen-testers, red teamers, and bug bounty hunters.

November 3, 2026 · 12-3pm ET

AI Pen-Testing Conference

A free, three-hour online event on AI-driven pen-testing and red teaming, with the people doing the work.

Save your seat →

What the Toolbox is

The AI Pen-Testing Toolbox is the ultimate guide for pen-testers, red teams, and bug bounty hunters looking to get an edge with AI. The toolbox lists some of the best agents and automated scanners for conducting instant AI-driven pen-testing. It also identifies leading testing platforms and offensive security firms that specialize in pen-testing for auditing purposes. It also highlights other helpful resources, such as CVE lists, frameworks, bug bounty communities, certifications, and free open-source tools to aid pen-testing in the AI age.

  • 8 categories — from autonomous pen-testing agents to bug bounty platforms, frameworks, certifications, and open-source tools.
  • For practitioners — pen-testers, red teams, and bounty hunters looking to get an edge with AI.
  • Independent & non-sponsored — nothing here is ranked, scored, or paid for.

Every category, tool by tool

Pen-Testing AI Agents

Next-gen agentic pen-testing tools that autonomously plan, run, and validate pen-tests, turning weeks-long engagements into continuous, on-demand testing.

XBOW

An autonomous offensive security platform that continuously runs agents against your applications, chaining flaws into validated exploits. Topped the HackerOne leaderboard, and combines deterministic logic and AI creativity to find flaws.

Horizon3.aiNodeZero

AI pen-testing that applies AI to autonomously hack a system, discover vulnerabilities, and determine if they are exploitable, with remediation recommendations. Its NodeZero platform is used by the NSA and four of the Fortune 10.

Novee

Offers a proprietary offensive AI model optimized for hacking that continuously discovers vulnerabilities, determines exploitability, and provides tailored fixes.

Simbian

An agent for conducting context-aware penetration testing and analysis with remediation guidance. Fixed, transparent per-test pricing.

Terra Security

Continuous agentic pen-testing with humans in the loop. Covers web apps, networks, and AI systems, with testing aligned to code changes.

RedVeil

Autonomous pen-testing that involves agents in multi-step attack chains. Launches on demand and delivers audit-ready reports within hours.

Penligent

An agentic AI hacker that orchestrates 200+ security tools to find, verify, and exploit vulnerabilities, with human-in-the-loop control and compliance-ready reports.

CodeAnt AI

Runs black-box, white-box, and gray-box pen-tests with agents that reason across code, infrastructure, and runtime to prove what's exploitable and fix it.

MindFort

YC-backed autonomous agents that continuously pen-test live apps and APIs, attach a proof-of-concept exploit to every finding, and open pull requests with patches.

Sableby Vulnetic

Sable, by Vulnetic, can run penetration tests on web apps, APIs, internal networks, and Active Directory. Runs real exploits and backs each finding with reproducible evidence.

Xintby Theori

Offensive security tester from Theori that provides both source code and runtime penetration testing. Confirms exploitable findings and offers remediation advice.

RunSybilSybil

Sybil reasons like an attacker to continuously test code, APIs, cloud, and infrastructure, chaining exploitable paths and giving security feedback on every pull request.

Pen-Testing Tools & Services

Penetration testing as a service (PTaaS) providers and AI-assisted platforms or scanners that pair automation with human testers and existing pen-test workflows.

BreachLock

A full-stack PTaaS platform that pairs agentic AI with certified human testers to expose gaps in applications, APIs, cloud, networks, and IoT.

Escape

Runs agentic workflows that conduct penetration testing, discover exploits, offer next steps, and retain memory. Covers API- and GraphQL-based testing in addition to standard application vulnerabilities.

Hadrian

Agentic offensive security platform. Atlas continuously maps and tests external exposure, while Nova runs on-demand agentic pen-tests that return verified, evidence-backed findings.

Burp Suiteby PortSwigger

The pen-tester's staple web testing toolkit. Burp AI adds an agentic partner in Repeater that suggests attack angles, validates exploits, and drafts reports.

Cobalt

PTaaS provider that pairs 500+ vetted pen-testers with a new autonomous pen-test offering. Also runs dedicated AI and LLM pen-tests.

StealthNet AI

Hybrid pen-testing that combines autonomous AI agents with senior US-based testers, delivering audit-ready reports in 48 hours. Also tests AI and LLM systems.

Astra

Combines pen-testing with API security and cloud vulnerability scanning. Offers features like human verification for pen-testing, AI-powered threat modeling, and SOC 2, ISO 27001, HIPAA, and PCI reports.

Synack

PTaaS that pairs Sara, its autonomous red agent, with 1,500+ vetted researchers. FedRAMP Moderate authorized, with AI and LLM pen-testing.

Dreadnode

Infrastructure for running your own offensive AI agents, with ready-made capabilities for AI red teaming, web app pen-testing, and network operations.

Traditional App Testing Platforms

Popular traditional application testing platforms, such as DAST, SAST, and vulnerability scanning tools, that have added AI capabilities or can be used to pen-test AI-based applications, agents, or LLMs for vulnerabilities.

Aikido

Popular application security platform that provides an autonomous penetration testing tool, which simulates attacks with hundreds of AI agents. Can quickly produce SOC 2 and ISO 27001-ready reports.

Semgrep

Popular SAST platform that blends deterministic static analysis with AI to detect logic flaws like broken authorization, triage findings, and generate fixes.

Invicti

DAST-first AppSec platform with proof-based scanning for apps and APIs. Now adds agentic pen-testing and AI-driven DAST alongside SAST, SCA, and ASPM.

Nessusby Tenable

Tenable's long-standing vulnerability scanner for finding flaws, missing patches, and misconfigurations. A baseline complement to AI-driven pen-testing, not an AI tool itself.

Corgea

AI-native SAST that catches business logic and authorization flaws and generates fixes. Recently added autonomous AI pen-testing with auditor-ready reports.

Offensive Pen-Testing Firms

Firms, consultancies, or key individuals-for-hire that can perform modern AI-aware pen-testing. Often used for third-party compliance checks or by resellers.

Oneleet

Provides human pen-testing, code scanning, and a virtual CISO as part of its compliance platform. Can perform tests to quickly support SOC 2 and other compliance frameworks.

Maltek Solutions

Syracuse-based offensive security firm led by practicing operators. Offers AI-augmented, human-validated web app assessments alongside network pen-testing and vCISO services.

MSP Pentesting

A white-labeled solution for MSPs and resellers that offers penetration testing services. Can conduct AI pen-testing with human-validated findings across networks, web apps, and cloud.

Bishop Fox

Veteran offensive security firm whose AI/LLM security assessments pressure-test guardrails, prompt injection, and model behavior. Also runs AI-focused red team engagements.

NetSPI

Enterprise pen-testing firm with dedicated AI/ML services, including LLM application pen-testing, jailbreak benchmarking, and continuous AI penetration testing.

AI Bug Bounty Communities

The bug bounty platforms that aggregate bug bounties from companies. Bounty hunters and pen-testers should be following these, and companies should consider listing their programs on these sites.

Bugcrowd

Platform for posting bug bounty programs and crowdsourcing detection. Also offers PTaaS, AI pen-tests, and AI bias assessments.

HackerOne

Leading bug bounty platform. Also pairs AI agents with human experts for agentic pen-testing, and offers AI red teaming for models and apps.

huntr

A bug bounty platform for AI/ML, backed by Palo Alto Networks. Rewards researchers for finding flaws in AI frameworks, models, and guardrails.

0DIN

Mozilla-backed GenAI bug bounty program. Pays researchers for verified jailbreaks, prompt injections, and other exploits against frontier models and AI agents.

Intigriti

European bug bounty and PTaaS platform with 150,000+ vetted researchers, including AI specialists. A strong fit for EU data sovereignty.

AI-Related CVE Lists & Frameworks

The vulnerability databases and risk frameworks covering AI-related exploits that red teams should be following to stay informed on threats.

AVIDAI Vulnerability Database

The AI Vulnerability Database, an open-source knowledge base of failure modes across AI models, tools, and agents, with reproducible evidence.

Vulnerable MCP Project

A community database of Model Context Protocol vulnerabilities, from tool poisoning to remote code execution in popular MCP servers.

MITRE ATLAS

MITRE's knowledge base of adversary tactics and techniques against AI systems, modeled on ATT&CK and grounded in real-world case studies.

OWASP Top 10 for LLMs and Generative AI

The standard list of risks for LLM apps, from prompt injection to excessive agency. The 2025 edition is current.

OWASP Top 10 for Agentic Applications

Released in December 2025, this list covers the most critical security risks facing autonomous and agentic AI systems.

OWASP API Security Top 10

The go-to list of API risks, led by broken object-level authorization. Agents hit APIs directly, so it still applies.

OWASP MCP Top 10

A beta OWASP project ranking MCP risks such as token mismanagement, tool poisoning, and shadow MCP servers.

Certifications and Trainings

Hands-on certifications, courses, and communities for building and proving AI pen-testing and red-teaming skills.

AI Security University

Free courses, including LLM Security Fundamentals and MCP Security Fundamentals, that teach the OWASP Top 10 lists through real-world attacks.

Wraith Certified AI PentesterWCAP

A hands-on, 10-scenario exam covering prompt injection, tool abuse, data exfiltration, and guardrail bypass. Free during its launch phase.

OffSec AI Red TeamerOSAI+

OffSec's AI-300 course and 24-hour proctored exam on attacking LLMs, multi-agent systems, RAG pipelines, and AI infrastructure.

HTB Certified Offensive AI ExpertHTB COAE

Hack The Box certification built on its AI Red Teamer path, developed with Google. Ends in a seven-day practical assessment and report.

SANS SEC536 and GIAC GAIPT

A three-day SANS course on attacking LLMs, RAG pipelines, agents, and MCP servers. Pairs with the new GIAC AI Penetration Tester certification.

Practical AI Pentest AssociatePAPA · TCM Security

TCM Security's practical exam: two days to pen-test an agentic AI application, then two days to write a professional report.

NHIMGNon-Human Identity Management Group

The Non-Human Identity Management Group, an independent authority on NHI and agentic AI security. Offers CPD-accredited training on securing machine identities and AI agents.

Red Team Village

Nonprofit offensive security community behind the DEF CON village. Runs workshops, CTFs, and live online training for red teamers.

Open-Source AI Pen-Testing Tools

Helpful agents, open-source projects, local LLMs, and other tools for free, self-hosted pen-testing.

Strix

An open-source AI penetration testing tool to find and fix your app's vulnerabilities. The agent can run locally and works with any model of your choice.

Hex

A terminal-native agent powered by DeepSeek for hackers and red teamers that runs security tools through natural language and chains them into reusable attack workflows.

Deadend CLIby Stray Labs

Built by Stray Labs, this is another self-hosted, free open-source tool for AI-driven penetration testing. Scores about 81% on XBOW's validation benchmark.

Transilience

Open-source skills, agents, and commands for AI-powered penetration testing and security research with Claude Code.

Shannonby Keygraph

Keygraph's white-box AI pen-tester for web apps and APIs. Reads your source code, then reports only the vulnerabilities it can exploit.

PentAGI

A fully autonomous multi-agent system for complex pen-testing tasks. Self-hosted with Docker and works with multiple LLM providers.

PentestGPT

The agentic pen-testing framework published at USENIX Security 2024. Now runs autonomous pipelines on Claude Code and Codex, or local models through Ollama.

HexStrike AI

An MCP server that lets agents like Claude or Copilot autonomously run 150+ security tools for pen-testing and bug bounty workflows.

Cybersecurity AICAI · Alias Robotics

Alias Robotics' framework for building your own offensive and defensive agents, with support for 300+ models. Free for research use.

garakby NVIDIA

NVIDIA's LLM vulnerability scanner. Like nmap for models, it probes for prompt injection, jailbreaks, data leakage, and more.

PyRITby Microsoft

Microsoft's open-source red-teaming framework for proactively finding risks in generative AI systems.

KimiOpen-weight LLM

Moonshot AI's open-weight model family. Deadend CLI ran its XBOW benchmark on Kimi K2.5.

DeepSeekOpen-weight LLM

Open-weight models that power Hex and plug into most open-source pen-testing agents.

QwenOpen-weight LLM

Alibaba Cloud's open-weight model family, with sizes small enough to run on local hardware.

November 3, 2026 · 12-3pm ET

The AI Pen-Testing
Conference

Organized by AI Security University — a free, three-hour online event on how AI is changing pen-testing and red teaming: the agents, the techniques, and what still needs a human.

3 hours · online
Free to attend
Practitioner speakers

Register for the conference

Save your seat — we'll send the joining link and agenda.

Save your seat →

Missing a tool?

If you build or use an AI pen-testing tool that belongs in the Toolbox, we'd like to hear from you.

Suggest a tool →
AI Pen-Testing Conference
Nov. 3 · 12-3pm ET
Register →