How Goodfit measures candidate integrity

How to read a candidate's integrity report: what the verdict means, what every row in the checks grid measures, and what each value it shows you is counting.

Reading a candidate's integrity report

Every AI video interview produces an integrity report. This page is about reading one. For what the system actually runs during an interview, see Inside Goodfit's interview proctoring.

A flag means something happened that is worth a look. The report shows you the evidence behind it.

Opening it

Open the candidate, then find the assessment card for the interview. The integrity chip sits on that card. Click it and the full report opens.

Phone and WhatsApp interviews carry no chip at all.

The verdict, at the top

The first line of the report is the verdict. There are four:

VerdictWhat it means
High riskAt least one serious signal. Someone else on camera, a voice that does not match the video, a phone in shot, developer tools opened, or strong evidence of AI-written answers.
Needs reviewSomething worth a look, short of the above. Repeated tab switches, sustained attention away from the screen, earbuds, notes in hand, a partial voice match.
Looks clearEverything finished, nothing flagged.
No flags yetNothing flagged so far, but a check is still running. It can still become High risk when the remaining analysis lands.

Video and voice analysis take a few minutes after the interview ends and run on their own. If you see No flags yet, come back shortly.

Under the verdict is a summary of what was found, and where identity was checked, the reference image next to the interview itself.

What each section shows you

Headline findings. The chips directly under the verdict are the specific concerns, most serious first. Click one and the report jumps to the evidence behind it, or to that exact moment in the recording.

Session timeline. The whole interview as one bar, with every flagged moment marked in place. This is the fastest way to see whether concerns are clustered in one stretch, which usually has an ordinary explanation, or spread across the whole call, which usually does not.

What we checked. The full grid of checks, each with its own result. Every row is listed in detail further down. The states are:

  • Passed — ran, found nothing
  • Flagged — ran, found something
  • Analyzing — still running, no result yet
  • Doesn't apply — not relevant to this assessment type

When every applicable check has run and come back clean, the report says so plainly at the top of this section.

AI-written answers. Present when the transcript showed strong signs of reading a prepared or generated answer. Each flagged answer is quoted with the reason it was flagged. Click one to hear the candidate actually say it.

Voice ↔ lip match. Whether the voice on the recording belongs to the person on camera, and whether more than one voice was present.

Visual evidence. Still frames pulled from the recording at the moments that were flagged. Click any frame to jump to that point in the video and see it in motion. A single frame can mislead, the surrounding seconds rarely do.

Event log. Everything recorded during the interview, grouped by Browser, Camera, Microphone and Tab, with timestamps. Repeated events of the same kind are collapsed into one line with a count.

Every row in "What we checked"

Each row shows a result and the measurement behind it.

RowWhenWhat it measuresValues you will see
IdentityAfterThe face on camera against the selfie taken before the interviewmatched selfie · mismatch · 12s · insufficient
Extra personAfterAnyone other than the candidate appearing in framenone · 1 person max · 18s · max 2 people
AI writingAfterAnswers read off a script or a screen, judged from the words and from the timingno flags · 0 of 14 answers flagged · 2 answers · 78%
Voice matchAfterWhether the voice belongs to the person on camera93% in sync · 41% out of sync · partial match · mismatch
Reading notesAfterWhere the eyes go while the candidate is speakingnone detected · 36s off-screen gaze
Tab focusLiveThe candidate leaving the interview tabno switches · 7 switches
BrowserLiveDeveloper tools, leaving full screen, AI helper extensionsno flags · AI extension used · AI extension installed · DevTools ×2, 3 fullscreen exits
Phone / spoofAfterA phone in shot, or a photo or screen held up to the cameranone detected · 22s

Live rows are recorded as they happen and are final the moment the interview ends. After rows are worked out from the finished recording and take a few minutes to arrive, which is why they read analyzing… if you open the report immediately.

Identity is the one row that spans both: the reference selfie is captured before the interview starts, and the comparison against it runs afterwards.

Browser-based assessments with no camera show a different set: Browser, Clipboard, Tab focus, Fullscreen, DevTools and Shortcuts. All six are live rows, each counting attempts the same way.

Seconds are totals, not stretches. 18s · max 2 people means a second person was visible for eighteen seconds in total across the interview, and at the busiest moment there were two people in frame. Those eighteen seconds may be one continuous stretch or six separate appearances. The session timeline shows which.

Counts like 7 switches or DevTools ×2 are raw event totals, each with a timestamp in the event log.

Percentages on Voice match are the share of the interview where the lip movement and the audio agreed, or did not. 93% in sync is a pass. 41% out of sync is the same measurement reported from the other side, because the row is flagged.

Reading notes starts from the camera, one frame at a time: which way the eyes are pointing, the angle of the head, whether the mouth is moving, whether a hand is at the face, and whether paper, a pen or a phone is in it.

The judgement comes from putting the frames back in order. Blinks are discarded. Looking away counts only while the candidate is speaking, which separates someone thinking from someone reading, and it is measured as a share of the interview so the result holds whatever the length. When that share is high enough, the frames are read a second time as a sequence, which is what tells ordinary glances around a room from eyes returning to the same fixed point again and again.

The percentage on AI writing is the single highest-scoring answer in the interview, on a scale of 0 to 100. 2 answers · 78% means two answers crossed the flagging threshold, and the strongest evidence on any one of them scored 78. A single answer at 90 will show 90 even if every other answer was clean.

That score comes from two independent engines, and the row turns red only when both agree. The first reads the words, looking for the marks of a text being spoken aloud rather than composed: the same phrase delivered twice verbatim, a sentence restarted three times on the same clause, one misread word in an otherwise fluent passage, lists too clean to be spoken, and answers that keep going after the question has been fully answered.

The second reads only the clock. A long silence followed by a long fluent answer. A speaking rate no one sustains unrehearsed. Across the whole interview: response times that barely vary, almost no filler words, and nobody ever talking over the interviewer. It is measuring how even the conversation is, because real ones are not.

Fluency on its own is never evidence. Excellent, accented or second-language English, domain jargon and practised interview technique leave this score where it is. Only the timing engine agreeing puts the row in the red.

On this page