The AI Pen-Testing
Toolbox
AI tools and resources for pen-testers, red teamers, and bug bounty hunters.
AI Pen-Testing Conference
A free, three-hour online event on AI-driven pen-testing and red teaming, with the people doing the work.
Save your seat →What the Toolbox is
The AI Pen-Testing Toolbox is the ultimate guide for pen-testers, red teams, and bug bounty hunters looking to get an edge with AI. The toolbox lists some of the best agents and automated scanners for conducting instant AI-driven pen-testing. It also identifies leading testing platforms and offensive security firms that specialize in pen-testing for auditing purposes. It also highlights other helpful resources, such as CVE lists, frameworks, bug bounty communities, certifications, and free open-source tools to aid pen-testing in the AI age.
- 8 categories — from autonomous pen-testing agents to bug bounty platforms, frameworks, certifications, and open-source tools.
- For practitioners — pen-testers, red teams, and bounty hunters looking to get an edge with AI.
- Independent & non-sponsored — nothing here is ranked, scored, or paid for.
The Toolbox
AI pen-testing in one map — 8 categories, 65 entries. Click a category to jump to the details.
SemgrepInvictiNessusCorgeaEvery category, tool by tool
Pen-Testing AI Agents
Next-gen agentic pen-testing tools that autonomously plan, run, and validate pen-tests, turning weeks-long engagements into continuous, on-demand testing.
An autonomous offensive security platform that continuously runs agents against your applications, chaining flaws into validated exploits. Topped the HackerOne leaderboard, and combines deterministic logic and AI creativity to find flaws.
AI pen-testing that applies AI to autonomously hack a system, discover vulnerabilities, and determine if they are exploitable, with remediation recommendations. Its NodeZero platform is used by the NSA and four of the Fortune 10.
Offers a proprietary offensive AI model optimized for hacking that continuously discovers vulnerabilities, determines exploitability, and provides tailored fixes.
An agent for conducting context-aware penetration testing and analysis with remediation guidance. Fixed, transparent per-test pricing.
Continuous agentic pen-testing with humans in the loop. Covers web apps, networks, and AI systems, with testing aligned to code changes.
Autonomous pen-testing that involves agents in multi-step attack chains. Launches on demand and delivers audit-ready reports within hours.
An agentic AI hacker that orchestrates 200+ security tools to find, verify, and exploit vulnerabilities, with human-in-the-loop control and compliance-ready reports.
Runs black-box, white-box, and gray-box pen-tests with agents that reason across code, infrastructure, and runtime to prove what's exploitable and fix it.
YC-backed autonomous agents that continuously pen-test live apps and APIs, attach a proof-of-concept exploit to every finding, and open pull requests with patches.
Sable, by Vulnetic, can run penetration tests on web apps, APIs, internal networks, and Active Directory. Runs real exploits and backs each finding with reproducible evidence.
Offensive security tester from Theori that provides both source code and runtime penetration testing. Confirms exploitable findings and offers remediation advice.
Pen-Testing Tools & Services
Penetration testing as a service (PTaaS) providers and AI-assisted platforms or scanners that pair automation with human testers and existing pen-test workflows.
A full-stack PTaaS platform that pairs agentic AI with certified human testers to expose gaps in applications, APIs, cloud, networks, and IoT.
Runs agentic workflows that conduct penetration testing, discover exploits, offer next steps, and retain memory. Covers API- and GraphQL-based testing in addition to standard application vulnerabilities.
Agentic offensive security platform. Atlas continuously maps and tests external exposure, while Nova runs on-demand agentic pen-tests that return verified, evidence-backed findings.
The pen-tester's staple web testing toolkit. Burp AI adds an agentic partner in Repeater that suggests attack angles, validates exploits, and drafts reports.
PTaaS provider that pairs 500+ vetted pen-testers with a new autonomous pen-test offering. Also runs dedicated AI and LLM pen-tests.
Hybrid pen-testing that combines autonomous AI agents with senior US-based testers, delivering audit-ready reports in 48 hours. Also tests AI and LLM systems.
Combines pen-testing with API security and cloud vulnerability scanning. Offers features like human verification for pen-testing, AI-powered threat modeling, and SOC 2, ISO 27001, HIPAA, and PCI reports.
PTaaS that pairs Sara, its autonomous red agent, with 1,500+ vetted researchers. FedRAMP Moderate authorized, with AI and LLM pen-testing.
Traditional App Testing Platforms
Popular traditional application testing platforms, such as DAST, SAST, and vulnerability scanning tools, that have added AI capabilities or can be used to pen-test AI-based applications, agents, or LLMs for vulnerabilities.

Popular application security platform that provides an autonomous penetration testing tool, which simulates attacks with hundreds of AI agents. Can quickly produce SOC 2 and ISO 27001-ready reports.
Popular SAST platform that blends deterministic static analysis with AI to detect logic flaws like broken authorization, triage findings, and generate fixes.
DAST-first AppSec platform with proof-based scanning for apps and APIs. Now adds agentic pen-testing and AI-driven DAST alongside SAST, SCA, and ASPM.
Tenable's long-standing vulnerability scanner for finding flaws, missing patches, and misconfigurations. A baseline complement to AI-driven pen-testing, not an AI tool itself.
Offensive Pen-Testing Firms
Firms, consultancies, or key individuals-for-hire that can perform modern AI-aware pen-testing. Often used for third-party compliance checks or by resellers.
Provides human pen-testing, code scanning, and a virtual CISO as part of its compliance platform. Can perform tests to quickly support SOC 2 and other compliance frameworks.
Syracuse-based offensive security firm led by practicing operators. Offers AI-augmented, human-validated web app assessments alongside network pen-testing and vCISO services.
A white-labeled solution for MSPs and resellers that offers penetration testing services. Can conduct AI pen-testing with human-validated findings across networks, web apps, and cloud.
Veteran offensive security firm whose AI/LLM security assessments pressure-test guardrails, prompt injection, and model behavior. Also runs AI-focused red team engagements.
AI Bug Bounty Communities
The bug bounty platforms that aggregate bug bounties from companies. Bounty hunters and pen-testers should be following these, and companies should consider listing their programs on these sites.
Platform for posting bug bounty programs and crowdsourcing detection. Also offers PTaaS, AI pen-tests, and AI bias assessments.
Leading bug bounty platform. Also pairs AI agents with human experts for agentic pen-testing, and offers AI red teaming for models and apps.
A bug bounty platform for AI/ML, backed by Palo Alto Networks. Rewards researchers for finding flaws in AI frameworks, models, and guardrails.
Mozilla-backed GenAI bug bounty program. Pays researchers for verified jailbreaks, prompt injections, and other exploits against frontier models and AI agents.
AI-Related CVE Lists & Frameworks
The vulnerability databases and risk frameworks covering AI-related exploits that red teams should be following to stay informed on threats.
The AI Vulnerability Database, an open-source knowledge base of failure modes across AI models, tools, and agents, with reproducible evidence.
A community database of Model Context Protocol vulnerabilities, from tool poisoning to remote code execution in popular MCP servers.
MITRE's knowledge base of adversary tactics and techniques against AI systems, modeled on ATT&CK and grounded in real-world case studies.
The standard list of risks for LLM apps, from prompt injection to excessive agency. The 2025 edition is current.
Released in December 2025, this list covers the most critical security risks facing autonomous and agentic AI systems.
The go-to list of API risks, led by broken object-level authorization. Agents hit APIs directly, so it still applies.
A beta OWASP project ranking MCP risks such as token mismanagement, tool poisoning, and shadow MCP servers.
Certifications and Trainings
Hands-on certifications, courses, and communities for building and proving AI pen-testing and red-teaming skills.
Free courses, including LLM Security Fundamentals and MCP Security Fundamentals, that teach the OWASP Top 10 lists through real-world attacks.
A hands-on, 10-scenario exam covering prompt injection, tool abuse, data exfiltration, and guardrail bypass. Free during its launch phase.
OffSec's AI-300 course and 24-hour proctored exam on attacking LLMs, multi-agent systems, RAG pipelines, and AI infrastructure.
Hack The Box certification built on its AI Red Teamer path, developed with Google. Ends in a seven-day practical assessment and report.
A three-day SANS course on attacking LLMs, RAG pipelines, agents, and MCP servers. Pairs with the new GIAC AI Penetration Tester certification.
TCM Security's practical exam: two days to pen-test an agentic AI application, then two days to write a professional report.
The Non-Human Identity Management Group, an independent authority on NHI and agentic AI security. Offers CPD-accredited training on securing machine identities and AI agents.
Open-Source AI Pen-Testing Tools
Helpful agents, open-source projects, local LLMs, and other tools for free, self-hosted pen-testing.
An open-source AI penetration testing tool to find and fix your app's vulnerabilities. The agent can run locally and works with any model of your choice.
A terminal-native agent powered by DeepSeek for hackers and red teamers that runs security tools through natural language and chains them into reusable attack workflows.
Built by Stray Labs, this is another self-hosted, free open-source tool for AI-driven penetration testing. Scores about 81% on XBOW's validation benchmark.
Open-source skills, agents, and commands for AI-powered penetration testing and security research with Claude Code.
Keygraph's white-box AI pen-tester for web apps and APIs. Reads your source code, then reports only the vulnerabilities it can exploit.
A fully autonomous multi-agent system for complex pen-testing tasks. Self-hosted with Docker and works with multiple LLM providers.
The agentic pen-testing framework published at USENIX Security 2024. Now runs autonomous pipelines on Claude Code and Codex, or local models through Ollama.
An MCP server that lets agents like Claude or Copilot autonomously run 150+ security tools for pen-testing and bug bounty workflows.
Alias Robotics' framework for building your own offensive and defensive agents, with support for 300+ models. Free for research use.
NVIDIA's LLM vulnerability scanner. Like nmap for models, it probes for prompt injection, jailbreaks, data leakage, and more.
Microsoft's open-source red-teaming framework for proactively finding risks in generative AI systems.
Moonshot AI's open-weight model family. Deadend CLI ran its XBOW benchmark on Kimi K2.5.
Open-weight models that power Hex and plug into most open-source pen-testing agents.
Alibaba Cloud's open-weight model family, with sizes small enough to run on local hardware.
The AI Pen-Testing
Conference
Organized by AI Security University — a free, three-hour online event on how AI is changing pen-testing and red teaming: the agents, the techniques, and what still needs a human.
Missing a tool?
If you build or use an AI pen-testing tool that belongs in the Toolbox, we'd like to hear from you.
