What counts instead of a pen test?
- SOC 2 never required a pen test, and the follow-up question has a clean answer. Five questions decide whether any security evaluation counts, whoever or whatever performed it.
- The five: who performed the evaluation, do they operate the systems they checked, what did they look at, is there a dated report with findings, and did you respond in writing.
- You may already hold a qualifying evaluation without buying anything: a customer's security review of you, a cloud partner's assessment, a certification an app store required, or last year's examination.
- AI agents now run security testing cheaply. The output counts when a named party outside your company directs the work, reviews it, and signs. Run the tool yourself and it is a self-check, however smart the tool.
Five questions decide what counts
Once a founder learns that SOC 2 never required a pen test, the next question arrives in the same breath: then what do we need instead?
The answer is smaller and more useful than a product name. A SOC 2 examination, for anyone new to this, measures your controls, meaning the specific things you do to keep your system safe, against the published criteria of the AICPA, the body that writes US audit standards. One of those criteria asks you to evaluate whether your own controls are working, partly through the eyes of someone who does not run them. Any instrument that delivers that outside look can serve. When an evaluation lands in my examination file, these are the questions that decide its weight.
1. Who performed it, and are they outside your company? A named organization or person has to stand behind the result. A report nobody signs is a document nobody is accountable for, and accountability is what makes a claim usable by the customer who reads it.
2. Do they run the systems they checked? Nobody can independently check their own work. This is why independence survives every wave of technology: it describes the relationship between you and whoever checks you, and a better tool does not change that relationship.
3. What did they look at? The criterion asks about your control system: access, changes, monitoring, recovery. Testing that only probes your attack surface answers a narrower question. It is valuable, and it is narrower.
4. Is there a dated report with a stated scope and real findings? A clean certificate with no method and no findings is a logo. Findings are what make an evaluation inspectable later.
5. Did you respond in writing? What you fixed, what you scheduled, what you decided to accept and why. Detection without a recorded response is a smoke detector with no one home.
You may already hold one
Before buying any evaluation, check what already exists. The most common qualifying artifacts I see cost nothing, because someone else already paid for them.
An enterprise customer's security team reviewed you before signing the contract. Their questionnaire, their findings, and your replies add up to an independent evaluation sitting in your inbox. Confirm you are allowed to share the artifact before handing it over. A cloud provider's partner may run a security review of your workload. These are often free, and they produce a findings report. Your app may have gone through a security certification an app store or platform required, where an outside lab checks it. That counts in proportion to how deep the lab went, so a questionnaire the lab only reviewed carries less weight than a hands-on lab test. And from your second audit cycle onward, last year's examination is itself an independent evaluation that was performed on you.
Every one of those artifacts answers the five questions without a new purchase. The gap I see most often is question five: the evaluation happened, and nobody wrote down what the company did about the findings. Close that loop and the artifact you already hold starts carrying weight.
AI testing counts when someone outside your company stands behind it
The newest instrument deserves its own look, because the price is collapsing. AI agents now probe systems, chain findings, and write reports that used to take a specialist team a week. Whether that output counts has nothing to do with the model, and everything to do with the five questions. Run them against the two ways AI testing shows up.
You run the agent yourself. The first two questions already fail. The tool got smarter, and the evaluator is still you, because independence attaches to whoever directs the work. A founder running an AI security agent against their own stack has produced a self-check. That is genuinely useful for finding problems, and it is still not an independent evaluation. The same goes for the export of any scanner or monitoring dashboard, however green the tiles.
A third party runs the agent and signs the report. Now it can count, on exactly the same basis as any human-performed assessment. The AI is the instrument; the firm is the evaluator. Two conditions keep that honest, and they are worth checking before you pay anyone.
First, the signer directs the work or independently verifies it. If you ran the tool and handed the output to someone for a signature, the evidence was generated under your control, and you could have scoped the runs away from anything embarrassing. An evaluator has to obtain their own evidence, which is the same rule that applies to auditors.
Second, the signer documents what they themselves did. Which findings they reviewed, what they validated, how they judged severity. A signature on machine output the signer never read is a signature for rent, and as AI makes the technical work nearly free, that is the failure mode this market will produce at scale. Ask any vendor one question: what does your reviewer do before signing? A firm doing the work has a specific answer.
The pen test keeps its place
None of this makes outside testing a bad purchase. A penetration test remains a legitimate instrument, it answers the five questions cleanly when a qualified independent firm performs it, and sometimes your customers' contracts require one regardless of what the criteria say, so read what you signed first. The point is simple: the requirement was always the five questions, and the market around them just got much cheaper to shop.
Frequently asked questions
What counts as an independent evaluation for SOC 2 if I don't buy anything?
Can I use an AI penetration test for SOC 2?
Does running a security scanner on my own systems count as an independent evaluation?
What should I ask an AI security testing vendor before buying?
Is a signature on an AI-generated security report enough?
Keep reading
Do you need an evidence collection tool?
Nothing in the standards requires one. Evidence is records your systems already produce, and the auditor should be the one collecting them.
How many controls does SOC 2 require?
None, as a number. The rulebook holds 61 criteria and no control list at all. What actually decides how many you end up writing.
How much evidence does a SOC 2 audit need?
About twenty sources, and one export often answers several criteria at once. What auditors ask for, and what does not count.
How many samples does an auditor actually test?
No standard sets a number. How often the control runs does. The table firms work from, and why which items get picked matters more.
Sources
- The Trust Services Criteria require the entity to perform ongoing and/or separate evaluations to ascertain whether the components of internal control are present and functioning, and state that use of the criteria does not require an assessment of whether each point of focus is addressed.
- In an examination under the attestation standards, the practitioner must obtain sufficient appropriate evidence; inquiry alone is not sufficient.
- NIST SP 800-115 describes penetration testing and vulnerability scanning as distinct security assessment techniques.