SOX control testing 101: from RCM to signed-off test

Every SOX program eventually comes down to the same question from an auditor or an audit committee: can you show that this control operated, and can you show it for every instance you claim it did. Everything between scoping and a signed off conclusion exists to answer that question with evidence instead of assurance. Teams that treat testing as a checklist to clear tend to end up with rework, PBC churn, and last minute scrambles before the external auditor's walkthrough. Teams that treat it as a lifecycle, each stage feeding the next, tend to move faster and get fewer questions back.

This post walks through that lifecycle end to end: scoping, building the risk and control matrix, walkthroughs, the split between design and operating effectiveness, sampling, evidence collection, review, and conclusion. It assumes you already know what SOX is and why it exists. The goal here is a practical refresher on how a control actually moves from "in scope" to "tested and signed off," and where programs typically lose time or quality along the way.

Scoping: deciding what gets tested

Scoping sets the boundary for the entire testing cycle. It usually starts with a top down, risk based assessment: material account balances and disclosures, the processes that feed them, and the locations or business units that matter based on quantitative thresholds (often tied to a percentage of revenue, assets, or a materiality calculation) and qualitative factors like fraud risk, recent restatements, or system changes.

Scoping decisions typically cover:

  • Significant accounts and disclosures, and the relevant assertions (existence, completeness, valuation, rights and obligations, presentation and disclosure) tied to each.
  • In scope locations and business units, based on materiality and risk, not just size.
  • Significant processes and the IT systems and interfaces that support them, including any IT general controls (ITGCs) that the process controls depend on.
  • Changes since last year: new systems, new processes, M&A activity, reorganizations, or prior year deficiencies that push additional controls into scope.

Under scoped programs get flagged by auditors and end up doing rework mid cycle. Over scoped programs waste tester hours on controls that carry little risk. Most mature programs revisit scoping at least annually and treat it as a living document that gets updated when the business changes, not a one time exercise from the year SOX first applied.

The risk and control matrix (RCM)

The RCM is the backbone of the testing cycle. It is the document that ties a financial statement risk to the specific control that addresses it, and from there to the test that will be performed. A weak RCM produces weak testing no matter how careful the tester is, because the tester is only as precise as the control description they were handed.

A well built RCM row usually includes:

FieldWhat it captures
Risk statementThe specific way a misstatement could occur (for example, revenue recognized before delivery criteria are met)
Control descriptionWho performs the control, what they do, how often, and what evidence it leaves behind
Control typePreventive or detective, manual, automated, or IT dependent manual
FrequencyDaily, weekly, monthly, quarterly, annual, or as needed (event driven)
Assertion(s) addressedExistence, completeness, accuracy, valuation, cutoff, rights and obligations, presentation
Key or non keyWhether the control is relied on to prevent or detect a material misstatement on its own
Control ownerThe individual or role accountable for performing the control

A control description that says "manager reviews reconciliation" is not testable. A description that says "the accounting manager reviews the bank reconciliation monthly, agrees the book balance to the bank statement, investigates reconciling items over a defined threshold, and evidences review by signing and dating the reconciliation" gives the tester something concrete to check against. Vague control descriptions are one of the most common root causes of testing delays, because the tester has to go back to the process owner mid test to figure out what the control was actually supposed to do.

Note

Frequency drives sample size more than almost any other RCM field. Getting frequency wrong (for example, treating a control performed as needed as if it were weekly) throws off the sample calculation downstream and often is not caught until review.

Start your 14-day free trial

Bring one control and its evidence. See the AI test it in minutes.

Start free trial

Walkthroughs: confirming the control as designed

A walkthrough is a single, detailed trace of one transaction or instance through the control, done with the process owner, usually early in the cycle before formal testing begins. The point is not to test operating effectiveness yet. It is to confirm that the control as described in the RCM is actually the control as performed in practice, and to identify the specific evidence that will support testing later.

During a walkthrough, the tester typically:

  • Inquires with the person who performs the control about what they do, step by step.
  • Observes the control being performed, or reviews the system configuration if it is automated.
  • Inspects the evidence that a completed instance would leave behind.
  • Confirms the control addresses the risk and assertion it is mapped to in the RCM, and flags gaps if it does not.

Walkthroughs are also where design gaps get caught early, before a tester spends hours pulling a sample against a control that was never going to hold up. A control that has no documented threshold for what counts as a reconciling item worth investigating, or that relies on an email approval with no record retained, usually surfaces here. Catching that in a walkthrough in March is a design conversation. Catching it in November during testing is a deficiency.

Test of design vs. test of operating effectiveness

These two are often confused, but they answer different questions and usually happen at different points in the cycle.

Test of design (TOD)

TOD asks: if this control operated exactly as described, would it actually prevent or detect a material misstatement. It does not require pulling a sample of instances. It is a one time evaluation of whether the control, as designed, is capable of addressing the risk. A control can be operating exactly as intended every single time and still fail TOD, if the design itself has a gap, such as a review threshold set too high to catch a realistic error, or a segregation of duties conflict built into the process.

Test of operating effectiveness (TOE)

TOE asks a different question: did this control actually operate, consistently, throughout the period under review. This is where sampling comes in. TOE only makes sense once TOD has passed. There is no point pulling twenty five samples to check consistency of a control that would not have caught the risk even if performed perfectly every time.

In practice, most programs perform TOD during or right after the walkthrough, and TOE later in the cycle, often split between an interim testing window and a roll forward window closer to year end that covers the remaining months.

Sampling

Sample size is typically driven by control frequency, since frequency is a reasonable proxy for how many total instances of the control occurred during the period. Programs generally work from a documented sampling methodology with company defined minimums, adjusted upward for higher risk controls (for example, ones with a history of exceptions, or ones judged to be more susceptible to error or override).

Control frequencyTypical minimum sample size
Annual1
Quarterly2
Monthly2 to 3
Weekly5 to 10
Daily or multiple times per day15 to 25 (often 25, or 40 to 60 for fully automated population testing)

These are common starting points, not a universal standard. Every program should document its own methodology and be able to explain why it chose it, since this is one of the first things an external auditor will probe. Sampling should also be genuinely representative of the full period, covering different weeks or months rather than clustering around a single convenient stretch, and the selection itself should be evidenced, not just the results.

Tip

Interim testing usually covers a window ending a month or two before year end, with a smaller roll forward sample covering the remaining months. Planning the roll forward sample size at the same time as the interim plan, rather than scrambling for it in December, keeps the year end crunch manageable.

Evidence collection

Evidence is what turns a tester's judgment into something a reviewer, and eventually an external auditor, can independently verify. For each sample selected, the tester needs to obtain evidence that the control was performed, for that specific instance, in the way the RCM describes it.

This is also the stage where testing effort tends to balloon, because evidence rarely arrives clean. A single sample might mean a signed reconciliation PDF, a screenshot of a system approval queue, an exported spreadsheet of exception items, and an email confirming a threshold was investigated, all of which need to be read, matched against the test attributes, and cited. This is the part of the job that tools like SOXLayer now handle the first pass on: reading the PDFs, images, and spreadsheets submitted as evidence and proposing a Pass, Fail, Insufficient, or N/A per attribute, with a confidence score and a citation back to the exact quote and page. Low confidence results get escalated to a stronger model rather than guessed at, and a human still has to accept, edit, or override every proposed result with a reason. The status never changes without a person doing it.

Regardless of how the first pass gets done, a complete test file needs to hold up on its own, without anyone re-explaining it later. A useful checklist for what that file should contain:

  • The control description and reference back to the RCM row it tests.
  • The population the sample was drawn from, and how the sample was selected.
  • Each sample item, individually identified (date, transaction ID, or equivalent).
  • The evidence obtained for each sample item, attached or clearly referenced.
  • The specific test attributes evaluated for each sample (what exactly was checked, not just "looks fine").
  • A pass or fail conclusion per attribute, per sample, with a citation to where in the evidence that conclusion comes from.
  • Any exceptions noted, with enough detail to support a deficiency evaluation later.
  • Tester name and date performed.
  • Reviewer name, date, and review notes or sign off.

Anything missing from that list is usually the first thing a reviewer, or an external auditor sampling the workpapers, asks about.

Review and conclusion

Review is a distinct step, not a formality tacked onto the end of testing. A reviewer should be checking that the sample was appropriate, the evidence actually supports the conclusion reached, the citations hold up against the source documents, and any exceptions were evaluated and, where needed, escalated. This is also where an append only, hash chained audit trail earns its keep: reviewers and auditors alike can see exactly who accepted, edited, or overrode each result and when, without anyone needing to reconstruct the history from memory or a chat thread.

Once individual controls are tested and reviewed, the program level conclusion aggregates everything: were there control deficiencies, and if so, were they design deficiencies, operating deficiencies, or both. Each deficiency gets evaluated for severity, typically starting with whether it is a control deficiency, a significant deficiency, or a material weakness, based on the magnitude of potential misstatement and the likelihood it would occur. Compensating controls, if any genuinely exist and are themselves tested, can offset an individual gap, but they need to be identified and evaluated on their own merits, not assumed.

The final conclusion, and the evidence trail behind it, is what management ultimately relies on for its assessment of internal control over financial reporting, and what the external auditor will re-perform or test on a sample basis during their own audit. A clean trail from RCM to walkthrough to test to review to conclusion is what makes that re-performance fast instead of painful.

Key takeaways

  • Scoping sets the boundary for the whole cycle and should be revisited whenever the business changes, not treated as a one time exercise.
  • A testable RCM row spells out who does the control, what they check, how often, and what evidence it leaves behind. Vague descriptions cause rework later.
  • Walkthroughs catch design gaps early, before hours get spent testing a control that was never going to hold up.
  • Test of design and test of operating effectiveness answer different questions. TOE sampling only matters once TOD has passed.
  • Sample sizes should follow a documented, frequency based methodology, and selections should genuinely represent the full period.
  • A complete test file stands on its own: population, sample, evidence, attribute level conclusions, citations, exceptions, and sign off.
SOXLayer Team
Product

Start your 14-day free trial.

Bring one control and its evidence. See the AI test it in minutes.