AI detection

How accurate is AI proctoring?

How accurate is AI proctoring, really? A clear look at what accuracy means, what affects it, and why human review still decides the outcome.

6 min read 11 September 2026
Share
Illustration of a target with several dots grouped near the centre and two dots clearly outside the rings, the stray dots outlined rather than filled

In short

AI proctoring accuracy isn't one number - it's a mix of reliable event detection, less reliable intent interpretation, and a review process that determines whether flags translate into fair outcomes.

AI proctoring is reasonably accurate at detecting specific, well-defined events, a second screen, a tab switch, a mismatched ID photo - but far less reliable at determining whether a flagged event actually means someone cheated. Those are two different accuracy questions, and vendors don't always separate them clearly when they cite a headline number. This piece breaks down what accuracy actually means in this context, what affects it, and why human review still has to be part of the answer. For the broader picture of how these systems work, see our complete guide to AI proctoring software.

What "Accuracy" Actually Means in AI Proctoring

"Is AI proctoring accurate?" sounds like one question, but it's really three, and conflating them is where most vendor marketing goes wrong.

A two-by-two grid of four equal squares, one square filled pale mint and one outlined in a heavier line, showing four possible outcomes without naming them
  • Identity verification accuracy

    - did the system correctly confirm the test-taker is who they claim to be? This is usually the most reliable of the three, especially with a simple photo-ID match or OTP-based check.
  • Event detection accuracy

    - did the system correctly log that a tab switch, second screen, or copy-paste action happened? This is close to binary and generally reliable, since it's detecting a technical event rather than interpreting intent.
  • Cheating determination accuracy

    - did a flagged event actually indicate cheating? This is the weakest link by far, because a tab switch or a glance away from the camera can have entirely innocent explanations.

Most public accuracy claims describe the first or second category while implying the third. That's the gap worth watching for. For more on what these systems technically flag and why, see our guide on how AI proctoring detects cheating.

What Affects AI Proctoring Accuracy

Even within the categories above, accuracy isn't fixed - it shifts based on conditions that have nothing to do with the candidate's honesty.

  • Camera and lighting quality

    - Poor lighting or a low-resolution webcam makes any face-based check less reliable, through no fault of the candidate.
  • Internet connection stability

    - A dropped connection can look identical to a candidate leaving the test window, depending on how the system logs it.
  • Device and browser setup

    - Some second-screen or tab-detection methods behave differently across operating systems and browsers, which means detection consistency isn't always uniform.
  • Test design itself

    - A test that requires candidates to reference external material will naturally trigger more tab-switch flags than one that doesn't - the flag rate says as much about the test as the candidate.

None of these factors are about whether someone is cheating. They're about whether the environment gives the system a clean signal to work with, which is a meaningfully different thing.

Where AI Proctoring Gets It Wrong

Every automated flagging system produces false positives, sessions flagged as suspicious that turn out to have a normal explanation. A notification popup that steals focus, a candidate reaching for water and briefly leaving frame, a shared office Wi-Fi that momentarily disconnects - all of these can generate a flag that looks identical, at a glance, to something that matters.

This isn't a flaw unique to any one vendor. It's a structural property of flagging systems that err toward catching more rather than missing real issues, since missing genuine cheating is the costlier mistake for most organizations. The trade-off is a higher rate of flags that need a human to sort out. We cover this specific failure mode, and how to handle it fairly, in AI proctoring false positives.

Why Human Review Still Matters

The honest conclusion from all of this is that AI proctoring accuracy is really a two-stage question: how good is the system at flagging, and how good is the process at reviewing what it flags. A highly accurate detection system paired with no human review process still produces unfair outcomes, because a flag isn't proof.

A flag icon on the left joined by a dotted line to a magnifying glass on the right, and from the magnifying glass two short branches ending in a tick and a cross

This is why platforms built around evidence rather than automated verdicts tend to hold up better under scrutiny. TunnelQuiz, for example, logs tab/focus events, second-screen detection, and periodic photo checks, then packages the session into a shareable, view-only report - it's built to give a reviewer what they need to make the call, not to render one on its own. That distinction matters more for actual fairness than any single accuracy figure would.

How to Evaluate Accuracy Claims From Vendors

If you're comparing proctoring platforms and one of them cites a specific accuracy percentage, a few questions are worth asking before taking it at face value:

  • Accuracy at what, specifically? Identity match, event detection, or cheating determination are different claims - ask which one the number describes.
  • Tested under what conditions? A number from a controlled lab environment with good lighting won't hold up the same way across thousands of real candidates on varied hardware.
  • What happens to a flagged session? If the answer is "automatically rejected," that's a red flag regardless of the accuracy number attached to it.
  • Is the number independently verifiable? A percentage with no methodology behind it is a marketing claim, not a measurement.
  • What's the false positive rate, not just the true positive rate? A system that catches 99% of real cheating but flags 20% of honest candidates isn't necessarily the better trade-off for your use case.

A vendor that answers these clearly is generally more trustworthy than one that leads with a single impressive-sounding number.

The Bottom Line

AI proctoring accuracy isn't one number - it's a mix of reliable event detection, less reliable intent interpretation, and a review process that determines whether flags translate into fair outcomes. Evaluate all three before trusting a vendor's headline claim. TunnelQuiz focuses on the part it can do reliably, logging tab/focus events, second-screen detection, and identity checks, and hands the judgment call to a human with a clear, shareable report.

Frequently asked questions

Is AI proctoring accurate?

It's generally reliable for detecting specific technical events like tab switches or a second screen, and reasonably accurate for identity verification. It's considerably less reliable at determining whether a flagged event actually means someone cheated - that judgment usually still needs a human reviewer.

Can AI proctoring be wrong?

Yes, regularly. False positives - flags triggered by innocent behavior like a notification pop-up or a brief camera absence - are a normal part of automated flagging, not a rare malfunction. This is exactly why review processes matter as much as detection technology.

Does AI proctoring replace human judgment?

It shouldn't, and most well-designed systems aren't built to. AI proctoring is best understood as a way to surface evidence and narrow attention, with a human still making the final call on anything flagged as suspicious.

What's the difference between AI proctoring accuracy and reliability?

Accuracy refers to how correctly the system detects and classifies specific events. Reliability includes accuracy but also consistency across different devices, connections, and conditions - a system can be accurate in ideal conditions but unreliable in real, varied ones.

Why do proctoring vendors cite different accuracy numbers?

Because they're often measuring different things - identity match rate, event detection rate, or overall "integrity score", under different testing conditions. Without knowing the methodology, comparing two vendors' accuracy claims directly isn't meaningful.

Should low accuracy concerns stop me from using AI proctoring?

Not necessarily. The more useful question is whether your process includes human review of flagged sessions and a fair way for candidates to explain a false flag. A system with imperfect detection paired with good review can still be fair; a "highly accurate" system with no review process is riskier than it sounds.

Run exams you can actually stand behind.

Human review on every flag, transparent room-scan and lockdown policies, and a pilot-first rollout.

  • Free plan, 50 credits a month
  • Works in any browser
  • Proctoring on every attempt
  • Scored the moment they submit