A professor described giving an oral exam in an online graduate course last month to find out whether his students were reading the assigned book or reading AI summaries of it. He asked them what a “latifundio” was. By his count the word appears twenty eight times in the first book on the syllabus, and he had sent the class a short glossary in week one that defined it. Most of them had nothing. One student asked if he meant “a lot of fun.”
The story went around because the punchline is good. The part worth sitting with is how fast it worked. A written assignment on that book would probably have come back clean, competent, and difficult to challenge. A few seconds of conversation settled a question that a stack of essays could not.
That gap explains the renewed interest. Written work has become a weaker signal of what a student understands, and faculty are reaching for the oldest assessment format there is to recover it. This guide covers what oral exams give you, the objections that keep most professors from running them, how to build a rubric and score a hinted answer, the design choices that separate a useful oral exam from a stressful one, and the places where the format still fails.
What an oral exam is
An oral exam is an assessment where an examiner asks questions out loud and the student answers in speech, usually one to one, usually with the examiner free to follow up based on what the student says. The format is old enough that it predates written examinations entirely, and it survives in the Oxbridge tutorial, the European viva voce, doctoral defenses, and board certification in medicine.
Three related things get confused with it, and the distinctions matter when you design one.
A class presentation is prepared and delivered. The student controls the content, the sequencing, and the rehearsal, which is exactly what makes it weak evidence of understanding once AI can draft and script it.
A thesis or dissertation defense is narrow and deep. The questions are specific to work the candidate spent months on, and the examiner is testing ownership of that particular contribution.
An oral exam in the sense this guide uses covers the material of a course. The student does not know which questions are coming, cannot control the sequence, and has to respond in real time. That loss of control is the entire source of the signal.
A useful hybrid sits between the last two. The student submits a paper, problem set, or project, and then answers questions about the specific choices they made in it. Cornell’s engineering faculty call this an oral defense, and it is probably the easiest first version to run because the artifact already exists in your course.
Can you tell if a student used AI?
Almost never from the document. This matters because it is the first thing most people try, and the time spent trying is time not spent on assessment design.
Two peer-reviewed evaluations published in the International Journal for Educational Integrity in 2026 tested this directly, and the results were not close. In one study, Turnitin classified one hundred percent of the fully AI-generated papers in the sample as false negatives. GPTZero missed seventy percent. Copyleaks missed seventy five. One newer tool, Pangram, did perform well, so the honest summary is that detector quality varies enormously and the tools most institutions have already deployed performed worst on that sample.
The failure runs in both directions. A second study compared detector behaviour on authentic writing by students working in a second language against professionally authored text, and found neither of the two commercial detectors it tested reached the reliability needed for high-stakes decisions. Both struggled with hybrid documents where a human and a model had both contributed, which describes most real student work now. False positives concentrated on predictable groups, which means the students most likely to be wrongly accused are the ones with the least room to absorb an accusation.
Every one of these papers reaches the same conclusion about consequential use, which is that a detector score should never be the sole basis for an integrity finding. So the question has to change. Instead of asking what produced this document, ask the student something the document cannot answer for them.
What a spoken answer shows that a document does not
The useful signal in an oral exam lives in the shape of the thinking around an answer, more than in the answer itself.
A commenter on that same thread put it better than I can: a model cannot reproduce the specific stumbling pattern of someone who has actually engaged with a text, and it cannot fake the tonal difference between a student who read something and forgot it and a student who never read it at all. You hear the difference immediately, and you cannot un-hear it.
Faculty who have adopted the format describe it in similar terms. In AP’s reporting on the shift, Chris Schaffer, who teaches biomedical engineering at Cornell, introduced twenty minute Socratic sessions he calls oral defenses after students submit written problem sets. His summary is that you will not be able to AI your way through an oral exam, which is stronger than the position I would defend, and I come back to that later. At Penn, Emily Hammer pairs oral exams with written papers in her seminars and is careful about why. She forbids AI on writing assignments and tells her students plainly that she cannot enforce it, and she frames the oral component around what students are losing, which she describes as skills, cognitive capacity, and creativity. Olivia Piserchia, a Cornell junior on the receiving end, described the mechanism from the student side. It is a lot harder, she said, to look someone in the eyes and say out loud that you do not know something.
This part of the argument is well covered elsewhere, so I will leave it there. The harder question is why, given how well the format works, so few courses actually use it.
The objections are good, and most of them are right
Three years ago a sessional instructor in anthropology asked a room full of faculty whether anyone had actually done this, or whether the enthusiasm for oral exams was talk. The replies are the most useful thing I have read on the subject, because they are a catalogue of real reasons written by people with no incentive to be diplomatic.
Class size comes up more than anything else. One reply does the arithmetic out loud and stops there: imagine scheduling sixty oral exams at ten minutes each for online students. Another says simply that if he only had twenty students he would do it. The constraint usually wins.
Discipline fit is a genuine limit. A finance instructor said he was not going to listen to someone walk through a balance sheet orally or calculate the beta of a stock out loud, and he is right about that. Oral exams suit conceptual reasoning, tradeoffs, and defending a position. They are poor at computation and at anything whose answer is a number.
And the cheapest objection is the strongest one. The same instructor asked why, if the goal is preventing AI use, you would not simply run in-class handwritten work, which costs less and everyone already understands it. If that is your only goal, he is correct and you should do that instead. Handwritten in-class assessment defeats AI more completely than a remote oral exam does, at a fraction of the scheduling cost. The case for oral assessment rests on the other thing it produces, which is evidence of how a student reasons when the conditions change under them. A closed-book written exam tells you what a student can reproduce. It does not tell you whether they will notice that the premise you just handed them is wrong. If you do not need that second signal, the bluebook is the better tool and this guide is not for you.
The written record objection is the sharpest one. The same instructor wanted a record of what was done and how it was graded, because he expects grade disputes and does not want to defend a score from memory. Anyone who has sat on an appeals committee will recognise this immediately.
External moderation makes it harder in some systems. A lecturer in Ireland pointed out that external examiners from another university review all assessment there, which means an oral exam is only workable for the extern if it is recorded, and recording changes what the room is.
Scheduling breaks first in asynchronous courses. Students in an async course have no shared time, faculty are not available around the clock, and sixty sessions spread across the days people are actually free turns a week of assessment into three.
Accessibility concerns came from someone speaking about themselves. A commenter with auditory processing difficulties explained that they would struggle to hold a complex multi-part question in memory long enough to answer it, and that having the prompt in writing in front of them would let them refer back to the main points. That is a specific and cheaply fixable problem.
And there is the career risk nobody writes about. The instructor who started the thread worried, before anything else, that she would bomb her course evaluations. When an assessment format threatens your teaching reviews, “it produces better evidence” is not a sufficient argument.
Three years on, here is my honest accounting of where each of those stands. Two of them have good answers now. Three have partial ones. Two do not have answers at all, and I would not pretend otherwise to anyone considering this.
| Objection | Status | What addresses it |
|---|---|---|
| Examiner hours in a large class | Partly addressed | Self-scheduled AI-conducted sessions, shorter formats, or dropping the written grading the oral replaces. For human-led exams the hours are unchanged |
| Computation and numerical work | Not addressed | A real limit of the format. Keep anything whose answer is a number in written work |
| No written record for grade disputes | Addressed | A transcript tied to rubric criteria, with the specific evidence quoted against each criterion |
| External examiner moderation | Not addressed | Moderation wants recordings, and recording changes the conversation. A genuine tradeoff with no clean resolution |
| Scheduling in asynchronous courses | Addressed | Students book a slot whenever suits them, with no shared meeting time required |
| Accessibility and processing time | Partly addressed | Question visible on screen, practice attempts, opening with straightforward questions. None of this removes the need for formal accommodations |
| Risk to course evaluations | Partly addressed | Low weighting on the first run, an early low-stakes conversation before the graded one, format and rubric published in advance |
Why most professors still do not run oral exams
Read those objections together and a pattern shows up. Very few of them dispute that oral exams reveal understanding. Almost all of them are about affordability.
The arithmetic is unforgiving. Schaffer’s Cornell course has seventy students and he runs twenty minute sessions several times a semester, which is roughly twenty three hours of examiner time per round. He pays for it in the only currency available. He and his teaching assistants no longer grade the written problem sets at all. They grade only the oral defenses. Other Cornell courses have made the same trade in different shapes, including a religious studies course that replaced its final exam with thirty minute conversations, and an engineering course that compressed the format down to four minute interviews to fit a class of a hundred and eighty.
One reply in that 2023 thread is the most honest cost accounting I have seen. An instructor replaced the written proposal for a final paper with individual meetings of twenty to thirty minutes, weighted at ten percent, graded against a simple checklist to keep the tone conversational. She put the cost at ten to fifteen hours. What she did not expect was the offset. She got noticeably fewer student questions about the final project afterwards, students told her the feedback was better than anything they got on a written pitch, and she thought she might come close to breaking even on time across the term. She also described, almost in passing, the detection benefit that actually works: if a student who could not articulate their thinking in the meeting submits an excellent paper a month later, she has a specific reason to ask questions.
That is the honest picture. The cost is structural. An oral exam consumes examiner attention in real time, one student at a time, and the only ways to pay for it are to add examiners, shorten the sessions, or stop doing something else. Schaffer did the third. Splitting sessions across teaching assistants helps with throughput and introduces its own problem, which is that four examiners will not grade the same performance the same way unless the rubric is unusually specific about what counts.
What changes when an AI conducts the conversation
This is where the current wave of tooling enters, and it is worth being precise about what it changes, because the marketing around it usually is not.
An AI examiner does not make assessment cheat-proof. What it changes is how many conversations a course can hold, and when they can happen. Panos Ipeirotis at NYU Stern built an AI-conducted oral final for a course on AI product management and describes it as fighting fire with fire. Students log in from home at whatever time fits their schedule, which dissolves the asynchronous scheduling problem that stopped the anthropology instructor cold. He now wants to pair an oral component with every written assignment, on the grounds that he no longer trusts written work to be the product of thinking.
He and Konstantinos Rizakos also published what happened, including the parts that went badly, which makes it far more useful than a case study. Thirty six personalized oral exams cost fifteen dollars in total, about forty two cents per student. Seventy percent of students agreed the format tested genuine understanding. Eighty three percent found it more stressful than a written exam. Separately, eighty three percent had never taken an oral exam of any kind before. The study does not report how much those two groups overlap, and matching percentages do not establish that they do, so treat the obvious explanation as a hypothesis rather than a finding. It is a hypothesis worth acting on, because the fix is cheap either way. The agent also bundled several questions into one turn, failed to randomize reliably, and used a cloned voice that some students experienced as aggressive.
The student reaction reported alongside it is worth reading closely, because it is a design brief. One student found the voice surprisingly human but the conversation choppy, with odd pauses, and said it asked multiple questions at once. Her sharpest complaint was that it felt awkward to be talking to what was essentially a blank screen.
Those are all fixable, and the rest of this guide is mostly about how.
Separator questions, the design move that makes an oral exam hard to fake
Most published advice about oral exams stops at “ask probing follow-up questions,” which is true and nearly useless as instruction, because it does not tell you what to probe or how to tell a good probe from a bad one. There is a narrower technique underneath it that does most of the work.
Build questions whose correct answer is disagreement.
The examiner states something plausible and wrong, in a confident voice, with a reason attached. The student has to notice and push back. In the exams I have seen work best, these are marked explicitly in the rubric as the questions that separate a strong student from an adequate one, and they are placed deliberately rather than scattered.
The reason this is resistant to AI preparation is behavioural. A student who has been working from summaries has been trained by that material to agree. Summaries present settled conclusions in a confident register, and the reader’s job is absorption. When an authority figure then asserts something fluent and false, the student with no independent model of the material has no basis to object, and the social pressure runs entirely toward agreement. A student who genuinely understands the concept experiences the claim as wrong before they have finished processing it, and you can hear that happen.
Three examples, in different disciplines, none of them taken from a real exam:
| Field | What the examiner asserts | What a real rebuttal contains | What the wrong answer sounds like |
|---|---|---|---|
| Statistics | ”Your p-value came out at 0.04, so there is a four percent chance the null hypothesis is true.” | The p-value gives the probability of data this extreme assuming the null holds, which is a claim about the data given the hypothesis. Reversing it to a claim about the hypothesis given the data requires a prior. | Agreement, often helpfully restated as being ninety six percent confident the effect is real. |
| Accounting | ”The company posted its best profit in three years, so it can fund the expansion out of cash on hand.” | Accrual profit is not cash. Recognised revenue is not collected revenue. Record profit alongside growing receivables or inventory can sit with negative free cash flow. Asks to see the cash flow statement. | Agreement that the company is performing strongly, sometimes with a comment about reinvesting the profit. |
| Anthropology | ”Cultural relativism means an anthropologist can never criticize a practice in another society.” | Separates the methodological stance, understand a practice in its own context before evaluating it, from the moral claim that no cross-cultural judgment is legitimate. Accepting the first does not commit you to the second. | Agreement, usually with an added remark that judging would be ethnocentric. |
The accounting row is worth pausing on. The finance instructor was right that computation does not survive the transfer to speech. The profit and cash confusion shows what does: a conceptual error that a well-executed spreadsheet will happily carry all the way to a correct-looking answer, and that collapses the moment someone asks about it out loud. Computation and concepts come apart here, and the oral format only ever had a claim on the second.
Notice that all three wrong answers are polite, fluent, and confident. None of them sound like a student who does not know the material, which is exactly the problem with reading fluency as understanding in written work.
Writing one is a four step job. Start from the misconception you see every single year in the written work, because you already know what it is. State it as a confident claim in the examiner’s voice with a plausible reason attached, since a bare false statement reads as a trick while a reasoned one reads as a position. Decide in advance what a real rebuttal has to contain, described as an idea rather than a phrase, so you are not rewarding students who happen to say the magic word. Then write down what the wrong answer sounds like. That last step matters more than it looks, because after a term of running the question you will know whether a low pass rate means the cohort has a gap or your question is bad, and you cannot tell those apart from scores alone.
Three guardrails, and I would treat all of them as non-negotiable.
The separators have to be written by faculty in advance. A model improvising false assertions during a graded exam is a different and much worse product. Sooner or later it will challenge a student who was correct, or endorse one who was not, on a point nobody reviewed, and you will find out from the complaint.
The examiner has to concede when the student is right. If it keeps pressing after a correct rebuttal, the exam turns into a trap, students will correctly describe it as unfair, and you will deserve the complaint. The point of the question is to find out whether the student holds a correct position under pressure, and once you know, the pressure has served its purpose.
Cap how many appear in one exam. Two or three in a twenty minute conversation reads as rigour. Six reads as hostility, and the students who suffer most from a relentlessly adversarial examiner are the anxious ones rather than the unprepared ones.
What goes in an oral exam rubric
A written exam rubric can lean on the artifact, because the evidence sits there on the page and you can reread it. An oral rubric has to work in real time, while you are also conducting the conversation, and it has to produce the same result when a colleague uses it on a different student. That pushes the design in a few specific directions.
Score concepts, never keywords. A criterion like “mentions informed consent” is easy to game and it fails the student who handled consent correctly in her own words. Write the criterion as the idea you want to see: did the student check that the other party understood what they were agreeing to, and did they respect the refusal when it came? A student can satisfy that in three different vocabularies, and a student who has memorised the phrase cannot satisfy it at all.
Define what each level sounds like. For every criterion, write a sentence describing strong evidence, partial evidence, and insufficient evidence, in the register a student would actually use. This is the part people skip and it is the part that makes two examiners agree. “Exceptional: identifies the constraint before being asked, and names a second-order consequence” is usable at speed. “Excellent understanding” is not.
State the weight on delivery out loud. Decide how much of the grade is communication and put the number in the rubric where students can see it. One rubric I looked at recently sets it at fifteen percent and says so. The alternative is not a rubric without delivery in it, because fluency influences every human judgment whether or not you wrote it down. The alternative is a rubric where delivery counts an unknown amount and nobody can audit it.
Require the evidence, not just the score. Every criterion should be recorded with a moment from the conversation attached. That is the difference between “she scored 6 on reasoning” and “she scored 6 because she identified the bottleneck only after a prompt and could not extend it to the second case.” The first invites an appeal you cannot win. The second usually ends the conversation.
Keep it short. Three to five criteria is workable while you are listening and asking follow-ups. Nine is a form you will fill in from memory afterwards, which defeats the purpose.
One structural note on weighting. If the exam covers several topics in sequence and a student runs out of time before the last one, decide in advance whether the unreached topic scores zero or gets dropped from the denominator. Both are defensible. Only one of them is what your students expect, so tell them which.
Scoring, and what to do about hints
Buried in that 2023 thread is a question a faculty member raised and nobody answered. If a student cannot get there on their own and you nudge them, how should that affect the grade?
It is the right question, and inconsistency on it is probably the largest source of unfairness in oral assessment. Two examiners with the same rubric will grade the same performance differently if one of them counts a hinted answer as correct and the other does not.
The cleanest solution I have seen scores every sub-question into one of three buckets instead of on a continuous scale.
| What happened | Marks | What it tells you |
|---|---|---|
| Student gets there unprompted | Full | They hold the concept independently and can retrieve it under mild pressure |
| Student gets there after a nudge | Roughly half | The knowledge is present but not available without a cue, which is a real and gradeable difference |
| Student does not get there | Zero | Recorded with the hint that was given, so the gap is documented rather than inferred |
That is the whole model. Three buckets, applied per sub-question, averaged with whatever weights the rubric assigns.
The virtue of this is that depth of understanding gets encoded in the scoring model rather than left to the examiner’s impression at the end of a long day. It also makes hinting safe. The examiner is free to help a stuck student, which is humane and produces a better conversation, without that help quietly inflating the result. And it gives you something specific to tell students in advance, which does more for exam anxiety than reassurance does.
One related rule is worth copying. If your question bank is large enough that not every student gets asked every follow-up, pro-rate the score across the questions that were actually asked. Students should not lose marks for coverage gaps that belong to the examiner.
The operational details that decide whether this works
The difference between an oral exam students describe as fair and one they describe as a nightmare is mostly logistics. None of the following is intellectually interesting and all of it matters.
Run a separate technical check, days before the exam. This is the single highest-return thing on the list. A short session, three or four minutes, that confirms the student can hear the examiner, confirms the examiner can hear the student, confirms any visual material renders on their screen, and answers questions about the format. It should refuse to discuss exam content, and it should produce a simple pass or fail record so you know before exam day who has a broken microphone. Remember that eighty three percent of the NYU cohort had never taken an oral exam of any kind. A large share of what gets recorded as exam anxiety is really unfamiliarity with the room, and this removes it cheaply.
Put hard constraints in the workflow rather than the prompt. The NYU study found that instructing the model to ask one question at a time and to randomize did not reliably produce either behaviour. That result should be taken at face value. Time limits, wrap-up behaviour, and turn structure need enforcement from something outside the prompt, such as a timer that forces the wrap-up at a fixed point regardless of where the conversation has got to. Treat anything written in the prompt as a preference the model will usually honour and occasionally will not.
Handle silence deliberately. A student thinking for twenty seconds is not a student who has disconnected, and an examiner that fills every pause is training students to answer before they have thought. Set a genuine threshold, and decide explicitly what happens when it is crossed.
Write for speech. The choppy delivery with odd pauses that NYU’s students complained about is a text-to-speech artifact with a text-side fix. Short sentences, because a full stop becomes a natural pause. One idea per sentence. No markdown, no bullet symbols, no URLs, numbers spelled out. Anything that only makes sense on a page will sound wrong out loud.
Give the student something to look at. The blank screen complaint is real and it is easy to fix. A face, whether an avatar or a person, changes the register of the conversation. Showing the question and any supporting exhibit on screen alongside the audio is even more useful, because it directly solves the accessibility problem the commenter with auditory processing difficulties described. The question stays visible, and they can refer back to it instead of holding a multi-part prompt in working memory.
Randomize from a bank. Students in the same course talk to each other, and the 2023 thread worried specifically about coordination over Discord and group chats. Multiple questions per topic, drawn at random, mostly handles it. Publishing the full question list in advance is also a legitimate strategy: one instructor’s colleague did exactly that for an intermediate maths course and reported that students still did not all ace it, because knowing the question and being able to prove the result are different things.
Let students self-schedule. This is the piece that makes asynchronous courses possible at all, and it is the clearest practical gain from an AI examiner.
I should say where this list comes from. I work on Tough Tongue AI, and every item above exists because an earlier version of something we shipped was worse. The technical check exists because students arrived at a graded exam with microphones that had never been tested. The silence rule exists because an early agent talked over people who were thinking. The timer exists because asking the model to watch the clock did not work. One of the exams these lessons came from was an operations final for a Kellogg executive programme.
Where oral exams still break
I want to be direct about this, because the enthusiastic version of this argument is starting to overclaim and faculty will notice.
An oral exam raises the cost of faking understanding a very long way. It does not make faking impossible, and the gap is widening.
Remote sessions are the weak point. A student can run a language model in voice mode on a second device and read the output, and one commenter noted this within minutes of the format being proposed. Smart glasses can now display text to the wearer or speak into their ear, and at least one instructor in that thread has already concluded she has to police eyewear. Another described running what he thought was a healthy discussion-based seminar for a full term, only to learn afterwards that several students had been reading AI-generated answers aloud in real time. He instituted a technology ban and said it helped enormously.
An in-person oral exam is dramatically harder to defeat than a remote one. If integrity is the primary goal and you have the option, do it in person. Many programmes do not have that option, including online degrees and doctoral programmes whose candidates have already left campus, and pretending otherwise does not help them.
The external examiner conflict is also unresolved, and it is worth separating two things that get collapsed together. A transcript is a text record of what was said. A recording is the audio or video itself. You can keep the first and discard the second, which covers the grade dispute problem without retaining a student’s voice or face. The exams I described earlier are configured that way, with transcription on and recording off.
External moderation is the case where that is not enough. An examiner from another institution assessing whether standards were applied consistently may reasonably want the recording, because tone, hesitation, and how hard the examiner pushed are part of what they are judging. Turning recording back on to satisfy them changes the event, since a filmed conversation is not the same conversation. I do not have a clean answer here. There is a choice between privacy and moderation depth, and it should be made deliberately and disclosed to students either way.
Finally, an oral exam does not do the job of a writing assignment. A composition instructor said this plainly in response to the enthusiasm, and she was right to. Some faculty are responsible for teaching students to write, and that obligation does not go away because grading writing has become harder. Oral assessment belongs alongside written work, as a complement to it.
Fairness, anxiety, and access
Oral exams surface understanding that written work misses. They also introduce barriers around speech, hearing, processing time, language, anxiety, and access to a quiet room, and those barriers are not distributed randomly.
The core discipline is to keep the thing you are measuring separate from the channel you are measuring it through. If the course assesses economic reasoning, an accent or a slow speaking pace should not quietly become part of the score. The bounded delivery weight in the rubric is what holds that line, and the specific number matters less than the fact that students can see it.
Carolyn Aslan, who runs Cornell’s oral exam training, recommends clarifying the format ahead of time and opening with straightforward questions, which costs nothing. She also makes a point that gets lost in the anxiety conversation: getting a quiet student one to one is sometimes the breakthrough, and you finally hear from someone who never speaks in class.
Practice attempts matter more than any single accommodation. So does publishing the rubric with the format. Beyond that, offer the alternatives your institution has approved, whether that is extra time, a text mode, a human-led session, or a different way to demonstrate the same outcome, and disclose recording, retention, access, and the appeal path before the activity rather than after a complaint. The UCL toolkit on inclusive oral assessment and the University of Delaware guide, which includes sample questions and rubrics across disciplines, are both worth reading before a graded pilot.
It is also worth remembering that students are not uniformly opposed to this. A first-year student in that thread described her chemistry professor’s oral midterm and final in detail, including the scheduling form and the randomized flash cards, and her verdict was that it made her learn across the whole term instead of cramming for three nights. Her observation about her classmates was that the format pulled up the people who would otherwise have coasted.
A first oral exam you could run this semester
Keep the first one small enough that failure is cheap.
- Pick one assignment that already exists. Do not design a new one.
- Add a five minute conversation about the work the student submitted.
- Publish the format, the rubric, and the three-bucket scoring rule when you announce it.
- Write two separator questions from misconceptions you saw last year.
- Give everyone one ungraded practice run, and a technical check if it is remote.
- Weight it low enough that a bad first attempt is survivable.
- Keep the final grading decision yours.
That is enough to find out whether the conversation tells you something the assignment did not. If it does, you will know what to expand. If it does not, you have lost a week rather than a term.
The main idea
Oral exams work because speech makes thinking visible in a way a finished document no longer can. That has always been true, and AI has made it matter more by degrading the alternative.
What held the format back was mostly hours, scheduling, and record-keeping, and those are the constraints that have moved. A conversation that can happen at a time the student chooses, that leaves a transcript tied to rubric criteria, and that can run more than once in a term is a different proposition from a format that only works in classes of twenty.
The judgment stays with you. Deciding what is worth assessing, writing the rubric, resolving an accommodation, and owning a contested grade are all still yours, and none of them get easier because the conversation scaled.
If you want the mechanics in more depth, the guide to conducting an AI oral exam covers question design, grounding an agent in course material, and Canvas workflows, and the platform comparison covers what is on the market. For the operational side specifically, our playbook chapters on oral exam operations and concept-based grading go further than this post does.
One term of running one yourself will tell you more about whether this belongs in your teaching than any amount of further reading.