Key Takeaways
- Every control needs two kinds of proof. Design evidence shows the control exists and is built correctly. Operating evidence shows it actually ran, consistently, across the whole audit period.
- A screenshot is almost never operating evidence. Auditors test recurrence and period coverage. One capture proves one moment.
- Conflating the two is the most common compliance data-quality bug we see. It makes a program look 95% complete when the operating half is largely empty.
- The fix is structural, not clerical: tag evidence by type at capture time, pick an authoritative source per capability, and let status be computed from what's actually there.
The moment it goes wrong
Here's a scene we've watched play out more times than we can count.
A company walks into fieldwork feeling good. Every control has a green check. Policies are written, approved, and version-controlled. Architecture diagrams are current. Configuration screenshots are neatly filed against each control. The dashboard says the program is ready.
Then the auditor picks one control — say, quarterly access reviews — and asks a question that sounds almost casual:
"Show me the four reviews you performed during the audit period, with the reviewer, the date, and what changed as a result."
And the room goes quiet. Because what's in the folder is the access review policy and a screenshot of the review UI, taken last Tuesday, while preparing for this meeting.
The control exists. It's well designed. Nobody can prove it operated.
Two questions, not one
Auditors are asking two separate questions about every control, and they grade them separately.
Design: Is this control built appropriately to address the risk?
This is point-in-time. A policy document, an approved procedure, a system architecture diagram, a configuration standard, a screenshot of an enforced setting. Upload it once and the design question is answered — until the design changes.
Operating effectiveness: Did this control actually work, as designed, throughout the review window?
This is period-based, and that word does all the work. Twelve months of logs. Four quarterly review exports with real reviewer names on them. Recurring training completion records. Ticket histories. Evidence that recurs at the cadence the control claims.
A control with strong design evidence and no operating evidence is not "mostly done." In a Type II examination, it's an exception waiting to be written up.
Why teams get this wrong
Not because they're lazy. Because the failure mode is invisible.
The tooling encourages it. Most GRC tools have one "upload evidence" button. Whatever you attach counts as evidence for that control, full stop. Nothing in the interface distinguishes a policy PDF from twelve months of export files, so the completion percentage treats them identically — and the number goes up either way.
Design evidence is easy and operating evidence is a habit. You can write a policy in an afternoon. You cannot retroactively perform four quarterly reviews. Teams naturally finish the tractable half first, then mistake that progress for readiness.
Screenshots feel like proof. They're the most common artifact in every evidence library and the weakest one. A screenshot shows a state, not a history. It answers "is MFA on right now?" and never touches "was MFA enforced for all 340 employees for the full twelve months?"
Automated collectors get mislabeled. An integration that pulls your IAM configuration nightly is producing operating evidence — a recurring record of enforcement. A one-time snapshot captured during setup is design evidence. Same connector, same API, completely different audit meaning. If your system doesn't record which one it was, and when it was captured, an auditor can't rely on either.
What good operating evidence actually looks like
Three properties, and all three are non-negotiable:
Recurrence at the stated cadence. If the control says quarterly, four captures. If monthly, twelve. Gaps are exceptions — and a gap that lines up with a personnel change or a holiday period is exactly where sampling lands.
Period coverage. Evidence has to span the review window, not cluster in the last three weeks before fieldwork. A dense burst of captures right before the audit is a tell, and experienced auditors read it immediately.
Provenance. When was this captured, and from where? An export with no capture timestamp and no source system is an assertion, not evidence. Every automated record should carry both.
There's a fourth property worth naming: population completeness. For controls with a population that must do something — complete training, acknowledge a policy, attest to access — operating evidence means the whole population, with exceptions explicitly accepted and justified. "Most employees completed training" is not a passing result. It's a finding with a number attached.
The structural fix
You can solve this with discipline and a shared drive. It just doesn't survive contact with a second audit cycle, staff turnover, or scale. The durable version is structural.
Tag evidence type at capture, not at review time. Every evidence record should carry an explicit design/operating/both classification the moment it lands. Retroactive classification is guesswork — and it's guesswork performed by whoever is under the most deadline pressure.
Make completion status honest. A control shouldn't advance to operating when only design evidence exists. If the underlying data can't support the status, the status is wrong, and a status that's wrong in your favor is worse than no status at all. The healthiest programs have a category between "nothing here" and "done": design satisfied, operating incomplete. That's an accurate description of most controls, most of the time.
Pick an authoritative source per capability. If three systems could theoretically prove your MFA enforcement, you need to decide which one does — for your organization, for this control. Without that decision, you get noise: cards failing for platforms you don't actually use for that function, and real gaps buried underneath them.
Compute the verdict; never store it. Overdue, complete, failing — these are functions of the evidence and the current date. Stored verdicts go stale silently and are always stale in the optimistic direction. Derive them at read time and they're correct by construction.
Build recurrence into the collector, not the calendar. If a control needs monthly evidence, the job that captures it should run monthly and keep every capture. Recurrence enforced by a human reminder is recurrence that stops working the first busy quarter.
A useful self-test
Pick five controls at random. For each one, try to answer, using only what's in your evidence library right now:
- What proves this control is designed correctly?
- What proves it ran in month three of the audit period?
- Who or what captured that, and when?
- If there's a population, who in it didn't comply — and where is that documented?
If question one is easy and questions two through four require a Slack thread, you've found the gap. It's almost certainly not limited to those five controls.
The takeaway
The distinction between design and operating evidence isn't audit trivia. It's the difference between a program that looks ready and one that is ready — and the two are indistinguishable on most dashboards right up until fieldwork.
We built Huduku around this split deliberately: evidence is classified by type at capture, each capability has a declared system of record, operating status is computed from what's actually present, and population-based controls get a real pass/fail check with per-person exceptions and documented risk acceptance.
None of that is glamorous. But it's the difference between answering "show me month three" in ten seconds and answering it in ten days.