How many samples? A practical guide to SOX sample sizes by control frequency

"How many should I test?" is one of the first questions a control owner asks once a control is scoped for the year, and it is also one of the most inconsistently answered. Ask five practitioners and you will often get five different numbers, because sample size in SOX testing is not a single formula. It is a judgment that starts with how often a control operates, adjusts for population size (sometimes), and then gets filtered through your external auditor's expectations and your own internal audit methodology.

This guide walks through the logic that most testing methodologies use to size a sample, gives you an illustrative reference table you can use as a sanity check, and explains where that logic breaks down if you apply it mechanically instead of thinking about what the control actually does.

Why frequency drives sample size

Sample sizes in control testing are built around a simple idea: you are trying to get comfortable that a control operated effectively across the full population of times it ran during the period, without testing every single instance. The fewer times a control runs, the larger a fraction of the population you need to look at to draw a reasonable conclusion, and often the more consequential each individual instance is. The more times a control runs, the more a well-chosen sample can stand in for the whole population.

That is why testing methodologies anchor sample size to control frequency first, not to some flat number applied across the board. A control that runs once a year (an annual budget approval, for instance) typically has its single instance tested. A control that runs multiple times a day (an automated system edit check, a daily reconciliation posting) can be sampled from a much larger population, so a modest sample can still be representative.

Frequency is a reasonable starting point precisely because it is a proxy for population size and for how much can go wrong between tests. A weekly review has 52 opportunities a year for something to slip through; an annual approval has one. The sample needs to reflect that difference.

A reference table of illustrative sample sizes

The table below reflects sample sizes commonly seen in SOX testing programs, organized by how often the control operates during the period being tested.

Control frequencyIllustrative sample size
Annual1
Quarterly2
Monthly2 to 5
Weekly5 to 15
Daily20 to 40
Multiple times per day25 to 60

Warning

This table is illustrative only. It is not a professional standard, and it is not a substitute for the sample sizes your external audit firm requires or the minimums set by your own internal audit or SOX methodology. Sample size guidance varies between firms, between engagement teams, and over time. Always confirm the actual minimums that apply to your program before finalizing a test plan, and treat any published table, including this one, as a starting point for discussion rather than a rule to cite.

In practice, most programs land close to these ranges because they trace back to similar underlying logic about population coverage and risk. But "close to" is not "identical to," and the difference matters when a reviewer or an external auditor pushes back on a sample that looks light.

Start your 14-day free trial

Bring one control and its evidence. See the AI test it in minutes.

Start free trial

Letting the company set its own numbers

Because sample size minimums are firm-specific and methodology-specific, a single hardcoded table is the wrong long-term answer for a testing tool. In SOXLayer, a company defines its own sample-size table by control frequency in settings: annual, quarterly, monthly, weekly, daily, and multiple times per day each get a minimum the company sets, based on whatever its external auditor or internal policy requires. Those numbers, not a generic default, are what SOXLayer uses when it builds test plans, so the sample sizes that show up in the app are always the company's own.

Company sample-size table by control frequency
A company's own sample-size table by control frequency in SOXLayer.

This matters beyond convenience. When a company changes external audit firms, or when a firm updates its methodology, the sample-size minimums can shift. Having that table live in one place, applied consistently to every test plan generated for every control, means the change happens once instead of getting manually re-applied (or missed) across dozens of individual control workpapers.

How population size interacts with frequency, or doesn't

Frequency tells you how many times a control ran. Population size tells you how many transactions or events sat underneath those runs. The two interact differently depending on what kind of testing you are doing.

Attributes-based testing

Most SOX control testing is attributes-based: you are testing whether a control operated (was the reconciliation reviewed, was the access request approved, did the system edit check fire) rather than sampling dollar amounts for a substantive conclusion. For attributes testing, sample size is typically driven by frequency, not by the underlying population size. A monthly reconciliation over a population of 200 line items and a monthly reconciliation over a population of 20,000 line items are both, at the level of the control, one review event per month. The sample size for testing whether that monthly review happened is the same in both cases, because you are not sampling line items, you are sampling instances of the control operating. This is the point that trips people up most often: a large underlying transaction population does not, by itself, justify a larger sample when the control itself only ran a handful of times.

When population does matter

Population size becomes directly relevant when the "instances" you are testing are the individual transactions rather than periodic reviews. If a control operates on every transaction (an automated three-way match, a system-enforced approval limit) and you are testing a sample of transactions to confirm the control applied consistently, then the transaction population, and its makeup, does drive sampling decisions: how many distinct locations, systems, or transaction types feed into it, whether the population is homogeneous, and how it is stratified if not. In those cases, frequency still sets a floor, but the transaction-level population shapes how the sample is drawn from within it.

The practical rule of thumb: ask what one "test unit" represents for this specific control. If a test unit is "the control ran this period," frequency governs. If a test unit is "this transaction was processed," population matters.

Rounding conventions

Sample sizes rarely come out as clean whole numbers when derived from a percentage of population or a statistical formula, so most methodologies apply a rounding convention. The common approach is to round up, never down, when a calculated sample size falls between two whole numbers. This is deliberate: rounding down to save one test undermines the logic of choosing a sample size in the first place, since the sample already represents a bare minimum for reasonable assurance, not a comfortable cushion.

A related convention worth stating explicitly in your methodology: minimums are floors, not targets. If your table says "2 to 5" for monthly controls, that range usually reflects risk tiering (lower risk toward the low end, higher risk or a control with a history of exceptions toward the high end), not a number to be negotiated downward whenever testing capacity is tight.

Common mistakes

  • Using the wrong frequency bucket. The most common error is applying an annual or quarterly sample size to a control that actually operates monthly or weekly once you look closely at how it is performed. A "quarterly access review" that is actually run every month because the control owner finds it easier to review smaller batches needs a monthly sample size, not a quarterly one, even if the control narrative still says "quarterly."
  • Testing the documented frequency instead of the actual frequency. Control narratives and RCMs go stale. Always confirm how often the control is genuinely being performed this period, not how it was described when the control was first documented.
  • Inflating sample size because the population is large, for an attributes-based control. As covered above, a big transaction population under a periodic review does not automatically call for a bigger sample of that review.
  • Applying one flat sample size across all frequencies. Using "3 samples for everything" regardless of whether the control ran once a year or fifty times a year defeats the purpose of frequency-based sizing and will draw scrutiny from an external auditor.
  • Rounding down under time pressure. Shaving a sample from 3 to 2 to save time on a busy testing cycle looks reasonable in the moment and looks bad in review.
  • Treating a default table as policy. Any illustrative table, including the one in this guide, is a starting point. Skipping the step of confirming actual firm or internal minimums is a recurring finding in testing programs that get audited by a second party.

Note

If your organization works with an external audit firm, ask them directly for their expected sample sizes by control frequency and any conditions under which they expect larger samples (control deficiencies in a prior period, first-year testing, higher inherent risk). Document that guidance and use it as your company's baseline rather than a generic table pulled from an article.

Key takeaways

  • Sample size in SOX testing is primarily driven by how often a control operates, from one test for an annual control up to dozens for controls that run multiple times a day.
  • The table in this guide (1 annual, 2 quarterly, 2 to 5 monthly, 5 to 15 weekly, 20 to 40 daily, 25 to 60 multiple times per day) is illustrative only, not a professional standard or a substitute for your external auditor's or internal policy's minimums.
  • For attributes-based testing of periodic controls, population size usually does not increase the sample; frequency does. Population matters more when you are sampling individual transactions rather than periodic review instances.
  • Sample sizes should be rounded up, not down, and treated as floors rather than negotiable targets.
  • A frequent and costly mistake is testing a control against the frequency written in an old control narrative instead of how it is actually being performed today.
  • SOXLayer lets each company define its own sample-size table by control frequency in settings, so test plans always use the company's own minimums rather than a generic default.
SOXLayer Team
Product

Start your 14-day free trial.

Bring one control and its evidence. See the AI test it in minutes.