SOC 2 sample sizes scale with how often a control operates. An annual control is tested once. A quarterly control is usually tested with two or three items, a monthly control with two to five, a weekly control with five to eight, and a daily or high-volume control with roughly twenty to forty items. The AICPA does not mandate specific sample sizes for SOC examinations, so the exact numbers come from each firm's documented methodology, but the scaling pattern is consistent across every reputable firm.

Knowing those numbers is useful. Knowing what happens before sampling is what determines whether you pass. An auditor cannot draw a sample until you produce a complete population, and more Type 2 exceptions originate in populations that could not be produced completely than in controls that genuinely failed.

How does a SOC 2 auditor decide what to test?

A SOC 2 Type 2 examination under AT-C section 205 requires the auditor to obtain sufficient appropriate evidence that controls operated effectively throughout the period, not on the day you were asked. For any control that operates more than once, testing every instance is usually impractical, so the auditor samples.

The sequence is always the same four steps:

  1. Identify the control and its frequency. Annual, quarterly, monthly, weekly, daily, or event-driven.
  2. Define and obtain the population. Every instance the control should have operated on during the period.
  3. Establish that the population is complete. Independently, not by taking your word for it.
  4. Select and test the sample. Then evaluate whether the result supports a conclusion about the whole population.

Step three is where audits break. Steps one, two, and four are mechanical. Step three requires you to be able to prove that the list you handed over is the whole list, and most organizations have never been asked to do that.

SOC 2 sample size by control frequency

The table below reflects how sample sizes are commonly set for tests of operating effectiveness in SOC engagements, drawing on the AICPA Audit Guide's recommendations for small populations and infrequently operating controls, together with the sampling principles in AU-C section 530.

Treat it as a planning tool, not a rule. Your auditor's methodology governs, and higher assessed risk, prior-period exceptions, or a control that relies on manual judgment will push sizes up.

Control frequencyPopulation / 12 monthsTypical sample sizeWhat the auditor asks for
Annual11The single instance, tested in full. No sampling.
Semi-annual22Both instances.
Quarterly42 to 3Selected quarters, plus the population to confirm all four occurred.
Monthly122 to 5Selected months, plus evidence the other months happened.
Semi-monthly243 to 5Selected instances from across the period.
Weekly525 to 8Selected weeks spread across the period, not clustered.
Dailyapprox. 25015 to 25Selected days, usually random rather than judgmental.
Multiple times per day250+25 to 40Random selection from a system-generated population.
Event-driven (onboarding, termination, change, incident, vendor onboarding)VariesCommonly 5 to 40The full population first, then a selection from it.
Automated, configuration-based1 configuration1Reperformance or observation, sometimes at two points in the period.

Four things this table implies that people miss:

Annual controls carry disproportionate risk. A population of one means a sample of one, which means a single failure is a 100 percent exception rate. Your annual penetration test, annual risk assessment, annual policy review, annual disaster recovery test, and annual management review are the highest-risk items in your entire program precisely because there is no room to absorb a miss. If your annual DR test slipped by a quarter and landed outside the period, that is an exception with no mitigation available.

Automated controls are not free. An auditor may test a configuration once, but they will also want evidence the configuration held for the whole period. A screenshot taken during fieldwork evidences the day of fieldwork. Change logs on the configuration, or configuration snapshots at two points, evidence the period.

Sample selection is spread deliberately. Auditors do not take the first five weeks or the most recent five. Selections are spread across the period, and often include the first and last month, because control performance degrades in predictable places: immediately after implementation, over holidays, and during team turnover.

Higher risk means larger samples. If a control failed in the prior period, if it depends entirely on one person remembering, or if the population is heavily manual, expect sizes at the top of the range or above it.

What is SOC 2 population completeness, and why does it fail?

Before an auditor can sample, they must be satisfied that the population is complete. AT-C 205 does not let the auditor accept a list on assertion. If the population might be missing items, a clean sample from it proves nothing, because the failures could be sitting in the part you did not hand over.

Concretely: you say there were 23 terminations in the period. The auditor tests 8 and every one had access revoked within 24 hours. Excellent. Unless there were actually 26 terminations, and the 3 missing from your list are the contractors whose access is still live.

The auditor asks how you know it was 23. That question is where readiness programs succeed or fail.

The five populations that are hardest to prove complete

1. Terminations and leavers. The population must come from a source independent of the system whose provisioning you are testing. If you produce the termination list from your identity provider, you have proven only that the identity provider knows about the people it knows about. Pull it from the HR system of record, reconcile the count, and expect the auditor to look for contractors, interns, and anyone offboarded outside the normal process. The gap between the HR list and the identity provider list is the most reliable finding in any SOC 2 audit.

2. Production changes. The population is every change deployed to production during the period, and the auditor needs it from a source that cannot omit changes made outside the normal path. A ticket list proves the changes that got tickets. Version control history or deployment pipeline logs prove what actually shipped. Reconcile the two, because the delta is the population of unticketed changes, and that delta is a change management exception waiting to be found.

3. Security incidents. Every organization's incident population is understated, because the definition of incident is applied inconsistently in the moment. If the auditor finds evidence of an event in Slack, a status page, a monitoring alert, or a customer email that does not appear in your incident register, the register is incomplete, and incident response cannot be tested against an incomplete register.

4. New vendors and subservice organizations. The population is every vendor onboarded in the period who touches your data or systems. Procurement records, accounts payable, and SaaS discovery tooling rarely agree. Where they do not agree, the vendor risk assessment control cannot be evidenced as operating across the population.

5. Access requests and privilege grants. Requests approved through the ticketing system are easy. Privileges granted directly in a console by an engineer with admin rights are the population you cannot produce, and they are precisely the ones that matter for the logical access criteria.

How to prove completeness before fieldwork

For every population you will need, document three things in advance and keep them with the evidence:

  1. The source of record. Which system is authoritative, and why it cannot omit items. HR for people. Version control or CI logs for changes. Not the tool being tested.
  2. The extraction method. The query, filter, or report used, with the date range, saved so it can be rerun and reproduced identically.
  3. The reconciliation. A second independent source compared against the first, with the difference explained. HR headcount to identity provider accounts. Deployment logs to change tickets. Monitoring alerts to the incident register.

A population that arrives with its source, its extraction method, and a reconciliation is accepted. A population that arrives as a spreadsheet with no provenance generates a request for more evidence, then a scope expansion, then delay, and sometimes an exception on the grounds that completeness could not be established.

This is also the single clearest limit of compliance automation. A platform can produce a list quickly. It cannot establish that the list is complete, because it can only report on the systems it is connected to and the events those systems chose to log. Completeness is a human verification step, and it is the step that gets skipped when speed is the selling point.

What counts as evidence for a sampled item?

For each selected item, the auditor is testing whether the control operated as described. Evidence needs four attributes, and evidence missing any one of them is commonly rejected.

The three most common evidence rejections:

Screenshots with no timestampA screenshot of a configuration proves the configuration existed when the screenshot was taken. If it was taken during fieldwork, it evidences fieldwork, not the period. Capture system-generated exports with dates, or capture at intervals through the period.

Reviews with no exceptions and no evidence of examination. A quarterly access review that removes nothing, every quarter, with no annotation, reads as a signature rather than a review. Retain the list reviewed, the annotations, and the removal tickets that resulted.

Approvals after the fact. A change approved two weeks after deployment is an exception even if the approver would have approved it. The control as described says approval precedes deployment. Sequence is part of the control.

What happens when a sampled item fails?

An exception is a sampled item where the control did not operate as described. Understanding how one exception propagates is worth more than any amount of pre-audit optimism.

The auditor evaluates it, not just records it. One late access revocation out of eight sampled from a population of twenty-three is not automatically a qualified opinion. The auditor considers the nature and cause of the deviation, whether it is isolated or systemic, and whether it affects the conclusion about the population.

Sample sizes may expand. If an exception suggests the control is unreliable, the auditor may extend the sample to assess whether the failure is isolated. Extended testing frequently finds more, because the first exception is rarely the only one.

Compensating controls may be considered. If a secondary control would have caught the failure, and it operated, and it can be evidenced, the auditor can consider it. Emphasis on evidenced.

The disclosure decision follows. Exceptions appear in Section 4 with the test, the result, and usually a management response. A documented exception with a stated root cause and remediation reads very differently from an exception with no response. If the exception is material enough, the opinion is qualified.

None of this is fatal. A qualified opinion is recoverable. What is not recoverable inside the period is the evidence you did not collect while it was happening, which is why sampling is a readiness problem rather than a fieldwork problem.

The practical takeaway: work backward from the sample

Most organizations prepare for a SOC 2 by implementing controls and hoping the evidence exists. The inversion that works is to start from the sample the auditor will draw.

For every control in your matrix, before the period opens, write down four things:

  1. Frequency, which sets the sample size from the table above.
  2. The population source of record, which cannot be the system the control operates in.
  3. The reconciliation that will prove that population is complete.
  4. The artifact that will evidence one sampled instance, containing date, actor, subject, and outcome.

Then, one month into the period, produce a population and a sample for three controls and test them yourself. Not at the end. One month in, while there is still eleven months of runway to fix what the exercise reveals.

That exercise is the cheapest thing in a SOC 2 program and the closest thing to a guarantee available. Almost every Type 2 exception is discoverable in month one by someone who tries to produce a population and cannot.

Primary source: the AICPA's Audit Sampling Audit Guide, together with AU-C section 530.

Frequently asked questions

How many samples will a SOC 2 auditor take?

It depends on how often the control operates. An annual control is tested once, a quarterly control usually with two or three items, monthly with two to five, weekly with five to eight, and daily or high-volume controls with roughly twenty to forty. Higher assessed risk or a prior-period exception pushes those numbers up. Ask your auditor for their sampling methodology at planning, because firms differ and you are entitled to know before the period opens.

Does the AICPA specify required SOC 2 sample sizes?

No. Neither the trust services criteria in TSP section 100 nor AT-C section 205 prescribes sample sizes. The AICPA Audit Guide on audit sampling provides recommendations, particularly useful for small populations and infrequently operating controls, and AU-C section 530 sets out the sampling principles. Actual sizes come from each firm's documented and peer-reviewed methodology, applied with judgment based on assessed risk.

What is population completeness in a SOC 2 audit?

Population completeness is the auditor's satisfaction that the list of control instances you provided contains every instance from the period. A sample can only support a conclusion about a population that is known to be complete. In practice this means providing the population from an independent source of record, documenting how it was extracted, and reconciling it against a second source. Incomplete populations are a more common cause of Type 2 exceptions than controls that actually failed.

Why does my auditor want the termination list from HR and not from our identity provider?

Because the identity provider is the system whose provisioning process is being tested. If a person was never fully offboarded, they may not appear in the identity provider's termination data at all, which is exactly the failure the test is designed to detect. HR is independent of that process and is the source of record for who left and when.

What happens if one sampled item fails?

The auditor evaluates the deviation rather than automatically qualifying the opinion. They consider its nature and cause, whether it is isolated or systemic, and whether a compensating control operated and can be evidenced. Sample sizes are often extended. The exception, if it stands, is disclosed in Section 4 with a management response. A qualified opinion follows only where the matter is material to the conclusion.

Can compliance automation platforms produce the populations my auditor needs?

Partly. Platforms are good at producing lists from the systems they are connected to, and that saves real time. They cannot establish completeness, because they only see what they are integrated with and what those systems logged. Reconciling a population against an independent source of record is a human verification step, and it is the step most commonly missing when a report is produced quickly.

How far in advance should I prepare populations?

Before the observation period opens, not before fieldwork. By the time fieldwork begins, the period is closed and any population you cannot produce is a population you cannot retroactively create. Document your population sources and reconciliation approach during scoping, then test three of them one month into the period.

Related reading: how to gather evidence for SOC 2 Type 2, what happens if you fail a SOC 2 audit, and what compliance automation will not do for your SOC 2.