Proctoring Evidence: How Enterprises Should Evaluate Suspicious Activity
Modern proctoring systems can detect unusual behaviour at a scale no human invigilator ever could. A single AI-proctored exam session can generate dozens of data points: eye movement, tab switches, background audio, second-screen signals.

What is in this guide
It captures all of it second by second, across thousands of candidates at once.
But detection isn't the same as determination. A tab switch, a second-screen signal, or a candidate glancing away from the screen may be relevant evidence. That's not the same as proof of misconduct.
This isn't just a technical nuance. Proctoring flags are widely understood as alerts that suspicious behaviour occurred - not confirmed findings of misconduct. Legal reviewers who handle exam-integrity disputes describe them the same way: flags are review triggers, not definitive evidence of cheating.
For an enterprise, this distinction isn't academic. Think of a bank running a compliance certification. A hospital validating clinical competency. A university administering a high-stakes final. A recruiter screening thousands of candidates.
For all of them, it's the difference between a defensible decision and a legal, reputational, or fairness problem. This is where online proctoring stops being a surveillance question. It becomes an evidence-management question.
This guide covers:
What proctoring evidence actually is
Why a flag should never be treated as a verdict
What a mature evaluation workflow looks like
What enterprises should demand from any vendor before trusting a system with high-stakes decisions
What Is Proctoring Evidence?
Proctoring evidence is the complete set of data an AI proctoring system captures during an exam session.
Not just the moment something looks unusual. Everything around it that gives that moment context.
Good evidence answers three questions at once:
- What happened? The observable event, a tab switch, a second voice, a copy-paste action.
- When did it happen? Timestamped and relative to the question being answered.
- What else was happening at the same time? Webcam feed, screen activity, session pattern.
A raw AI flag isn't evidence on its own. It's a pointer to a moment in a much larger session timeline.
Evidence is what you get when that pointer is backed by verifiable, time-stamped artifacts a human can actually review. Well-built systems don't just flag - they produce audit-ready records with clear notes, so investigations move fast and enforcement holds up.
This is a foundational idea behind how Tunnel Quiz approaches exam security. The system's job is to preserve evidence faithfully. Not to render a verdict.
Why a Flag Isn't a Verdict
Every AI proctoring system, no matter how advanced, works probabilistically. It assigns a confidence level to a behaviour pattern based on training data.
It doesn't "know" a candidate cheated, any more than a smoke detector "knows" there's a fire. It knows conditions matched a pattern.
That matters, because false positives aren't rare edge cases. They're a structural part of how detection models work.
Industry data puts false-positive rates at roughly 5-15%, depending on system sensitivity. Accuracy also varies significantly by behaviour type - facial recognition is far more reliable than gaze or audio analysis.
A slow internet connection. A glare on glasses. A sibling walking past a webcam. A screen-reader tool used for accessibility.
All of these can resemble "suspicious" patterns to a model that has never seen that candidate's normal environment.
Treating a flag as an automatic verdict punishes normal human variation as if it were intent to cheat, nervousness, disability accommodations, poor lighting, a noisy household. That's why every reputable framework treats an AI flag as the start of a review process, not the end of one.
The Difference Between a Signal, a Risk Level, and a Decision
Enterprises evaluating cheating detection technology need a shared vocabulary for what's actually happening inside the system.
There are five distinct stages between "something happened" and "a decision was made." Most proctoring failures, and most unfair outcomes, come from collapsing these into one step.
Signal → Risk → Evidence → Human Review → Decision
Signal
- A raw event is detected: a tab switch, a face leaving frame, a second voice on the microphone.Risk
- The system weighs that signal against context, frequency, timing, other signals, and assigns a risk level. Not a guilt level.Evidence
- The signal, its risk score, and the surrounding session data are packaged into a reviewable record.Human Review
- A trained reviewer examines the evidence in context, not the signal in isolation.Decision
- The institution, not the software, makes the final call: clear, retest, escalate, or investigate further.
A system that skips straight from "Signal" to "Decision" isn't proctoring. It's an unaccountable automated judgment machine.
This five-stage separation is the backbone of how enterprise-grade proctoring should be designed and audited.
What Evidence Should an Enterprise Proctoring System Capture?
To make the "Human Review" step actually possible, the underlying system needs a complete, structured evidence set. Not just an alert.
A defensible enterprise proctoring stack should log:
Identity verification
- ID-document match, face match at login, periodic re-verification through the sessionTimestamp
- precise time-alignment of every event to the exam clock and the specific question being attemptedTab switches
- count, duration, and timing of any navigation away from the exam windowFull-screen exits
- whether and how often the candidate left full-screen/lockdown modeSecond-screen events
- detection of additional connected displays or mirrored outputCopy/paste activity
- clipboard actions during the session, correlated to question contentWebcam captures
- periodic or event-triggered image/video snapshots of the candidateScreen captures
- periodic or event-triggered snapshots of on-screen activityAudio/activity signals
- background voices, unusual noise patterns, or prolonged silence where activity was expectedExam answers
- the actual response data, so reviewers can correlate suspicious timing with content changesSession timeline
- one chronological record tying every signal above into a single reviewable narrative
This is the exact philosophy behind Tunnel Quiz's approach to online proctoring. Capture everything relevant. Structure it clearly. Hand institutions a complete record, not a black-box accusation.
Why Context Matters
Context is what separates noise from a genuine risk pattern.
- One tab switch ≠ cheating.
A candidate might tab away to check instructions, silence a notification, or recover from a stray click. On its own, that tells you almost nothing.
- But: repeated tab switches + second-screen detection + suspicious timing = a higher-risk session worth reviewing.
When multiple independent signals converge - frequency, device pattern, timing relative to difficult questions - the odds that something's actually worth investigating rise sharply.
This is the core design principle behind how Tunnel Quiz weighs signals. No single event triggers an automatic penalty.
Instead, the platform correlates signals over the session timeline and surfaces a risk level. That tells reviewers where to look first, not what to conclude.
This mirrors how experienced proctoring teams already operate. Every proctored exam generates an audit trail, so there's always evidence to review. The goal of the technology is to make that review fast and accurate, not to replace it.
Why Automated Decisions Create Problems
Letting software issue final verdicts, instead of surfacing evidence for a human to weigh, creates risk on several fronts.
False positives: Detection accuracy is uneven across behaviour types. Audio analysis in particular runs at only 70-80% accuracy. Fully automated penalties will regularly punish innocent behaviour.
Candidate fairness: A student who was falsely flagged deserves the chance to explain. Legal reviewers note that candidates should be able to respond with facts, screenshots, timestamps, and original correspondence before any penalty is finalized.
Accessibility: Assistive technology, screen readers, extra monitors used for accommodation, atypical eye or head movement linked to disabilities - all of these can resemble "risk" to a naive model. Systems that don't account for this end up discriminating against the exact candidates who most need protection.
Human review: Remove a trained reviewer from the loop, and you remove the one safeguard that catches context an algorithm can't: a legitimate work tool, a documented accommodation, a technical glitch.
Institutional accountability: If a bank, university, or certification body can't show why a candidate was disqualified, only that "the AI flagged them", they have no defensible position if the decision is challenged or litigated. Best-practice frameworks recommend hybrid models: human reviewers validate every AI flag before penalties are applied. That step alone cuts false-positive impact by roughly 85%.
None of this means the technology doesn't work. Real-world deployments show that introducing proctoring measurably changes exam outcomes - average scores drop, which researchers read as evidence that unproctored exams had been letting cheating slide through.
The risk isn't in the detection. It's in what happens after detection.
What Banks, Universities, Hospitals and Certification Bodies Should Ask Vendors
Before adopting any proctoring platform for high-stakes use, get clear answers to these:
- Does the system separate risk scoring from final decisions? If the vendor can't describe a human-review step distinct from the AI flag, the tool isn't enterprise-ready.
- What evidence is retained, for how long, and in what format?You need session timelines, not just alert summaries, and retention that matches your audit and appeal windows.
- What's the false-positive rate, by behaviour type?A vendor who can't share this by category (video, audio, network, behavioural) hasn't measured it properly.
- How are accessibility accommodations handled in the detection model?Ask specifically about screen readers, extra displays for accommodation, and atypical movement.
- Who can access candidate video and biometric data, and under what retention policy?Critical for healthcare and banking clients bound by strict data-privacy rules.
- Can a non-technical reviewer understand the evidence package without vendor support?If review requires a data scientist, it won't scale to real institutional workflows.
- Does the platform support your specific exam type?Professional certification, academic assessment, or recruitment screening - each needs configurable rules, not one-size-fits-all thresholds.
- What does the appeal process look like end-to-end?From flag to candidate notification to final institutional decision.
The questions are the same across sectors. The stakes aren't.
A false disqualification in a banking compliance exam can affect someone's licensing status. A missed detection in a hospital credentialing exam can affect patient safety. A flawed recruitment screen can expose an employer to discrimination claims.
What Good Proctoring Looks Like
A well-designed enterprise proctoring workflow shares a few consistent traits, regardless of vendor:
- Every flag traces back to a specific, timestamped piece of evidence - never a bare "risk score" with no supporting detail
- Risk levels are shown to reviewers as a starting point - never delivered to candidates as an accusation
- Human reviewers see the full session timeline, not an isolated clip stripped of context
- Candidates have a defined path to explain a flagged moment before any penalty is applied
- Evidence retention and access controls meet the regulatory bar for the industry involved: education, finance, healthcare, employment
- Reporting is built for institutional decision-makers, not just IT administrators
This is the model Tunnel Quiz was built around: automated detection, paired with clear, structured evidence and a human-review layer. Institutions get speed and scale, without giving up defensibility.
The goal isn't to accuse candidates.
It's to give institutions enough evidence to make defensible decisions.
Questions, answered
No. A flag means a system detected a pattern statistically associated with risk, not that misconduct occurred. Reputable institutions treat every flag as a trigger for human review, not an automatic finding.
Document the moment in detail as soon as possible - what happened, any technical issues, the exact time. Then request the underlying evidence (video, timeline, screen capture) instead of accepting the flag summary alone. Most appeal processes are built around exactly this kind of evidence exchange.
It varies a lot by signal type. Identity and face matching tend to be the most reliable. Behavioural and audio-based detection are meaningfully less precise. That's exactly why a layered evidence approach, not one automated score, is the safer design for high-stakes decisions.
Yes, regularly. Institutions that retain incomplete evidence - an alert with no supporting video or timeline - are in a weak position if a decision is disputed. Comprehensive, well-organized evidence protects both the candidate's right to a fair process and the institution's ability to defend its decision.
No, A banking certification exam, a hospital credentialing assessment, a university final, and a recruitment screening test each carry different risk profiles and regulatory requirements. Rules and evidence-retention policies should be configured per use case, not applied as a single default.