Stanford's "Responsible Assessment in the AI Era" Already Exists. It’s Called Debate.
Stanford’s Accelerator for Learning and ETS just published a blueprint for continuous, transparent, socially situated evaluation of durable human skills. Debate has been running that exact system,
On January 29, the Stanford Accelerator for Learning and ETS convened about a hundred researchers, technologists, and education leaders for a gathering called Responsible Assessment in the AI Era. The resulting white paper opens with Candace Thille’s framing question — “How do we shape the current moment?” — and spends twenty pages sketching what assessment should become now that AI can produce polished final products on demand.
The word “debate” does not appear anywhere in the document.
It doesn’t need to. The paper describes a debate tournament on nearly every page.
Here is the authors’ own definition of responsible assessment. It should be “grounded in individuals’ sociocultural contexts.” It should generate “valid, trustworthy, and context-specific inferences.” And it should rest on the “accumulation of multiple sources of evidence through socially situated performance over time.”
Read that last phrase again. Accumulated evidence. Multiple sources. Socially situated performance. Over time.
That is not a description of any test currently administered at scale in American education. It is a description of a debate season: five to seven live performances every tournament weekend, each evaluated by an independent trained judge who writes reasons for the decision, aggregated across months into a longitudinal public record. The assessment field is trying to design, from first principles, a system debate has been operating — and iterating on, and arguing about, and improving — since before ETS existed.
I want to walk through the report carefully, because the mapping is not loose or metaphorical. It is close to exact.
The mismatch they name
The report’s diagnosis starts with Millie García, Chancellor of the California State University: traditional assessments “may tell us if a student arrived at a correct answer,” but “they often tell us little about how [students] got there.”
Debate is the inversion of this problem. The “how” is the entire assessed object. Reasoning is performed aloud, in sequence, under hostile questioning. The flow is a contemporaneous process trace. Cross-examination is a live audit of warrants. It is structurally impossible to present a conclusion in a round without exposing the path to it — because if you won’t walk the path, your opponent will walk it for you, unfavorably.
Cindy Mazow of Stanford’s Graduate School of Business names the deeper tension: learning requires room to take risks, but in assessment mode students “don’t want to make a mistake because they know that there’s a consequence to that assessment, whether it’s high-stakes or low-stakes.” The paper treats this as nearly inescapable. Debate’s architecture blunts it. Any single round is one data point in a long season. Every loss arrives with reasons attached. The next round starts in two hours. Debaters lose constantly — the activity is built on iterative public failure, and the culture treats the ballot as feedback rather than verdict. Trying a new aff, a new strategy, an unfamiliar argument is rewarded over a season precisely because no single event is terminal. That is Mazow’s “room to make mistakes” achieved inside a consequential environment, not by sanitizing the environment.
Then Kadriye Ercikan of ETS: “Full self includes family history, home context, and societal context,” and it follows students into assessment. Her observation that “many current systems still emphasize end products over the pathways that produce them” points at one of debate’s most distinctive properties, which is that it standardizes the format, not the content. Speech times and order are uniform. What students argue, from what literature, in what voice, is theirs. The diversity of debate leagues, performance debate, and the kritik are institutionalized proof that lived experience can be the substance of an assessed performance rather than construct-irrelevant noise.
Even the report’s Vygotsky material is implemented structurally in debate. Learners can do “hard and gritty things” when tasks are right-sized to their current abilities, the paper notes, invoking the Zone of Proximal Development. Debate operationalized that decades ago: novice, JV, and varsity divisions; lab groupings at camp; and power-matched pairings, where, after the early rounds, the tab software pairs you against opponents at your record. The tournament literally adapts difficulty to demonstrated ability, round by round.
The assessment AI can’t take for you
The report’s core validity crisis is the one every teacher in America now knows: learners can “generate high quality products without fully engaging in the learning processes those products are meant to represent,” and so assessment “risks measuring technological proficiency rather than human skill or understanding.”
A debate round is synchronous, extemporaneous, and adversarial. No one can outsource the 2NR. No one can outsource the final focus. AI can and should transform preparation — research, blocks, redos, practice reps — but the assessed moment demands internalized understanding, because an opponent is probing it live. This makes debate the rare assessment where AI assistance raises the floor of preparation without contaminating the measurement. In fact the relationship runs the other way: the better your AI-assisted prep, the more cleanly the round measures what you actually understood and can deploy under pressure.
That is exactly the complementarity OpenAI’s Sara Caldwell describes when she argues that people who combine their own competencies with AI will outperform those who rely on AI alone. Her image of the future worker “orchestrating a set of AI agents” already describes what a well-coached debater’s prep stack looks like in 2026 — with the round assessing the judgment layer the agents can’t supply.
And look at the durable-skills inventory the report lands on: “critical thinking, creativity, curiosity, collaboration, agency, and adaptability,” plus “emotional intelligence, relational intelligence, ethical judgment, and challenging assumptions.” That is close to a construct map of the ballot. Every item on the list has an observable, scored debate behavior: refutation, case innovation, self-directed research, partnerships, student-owned strategy, judge adaptation, framework and values debate, and — the whole point of the activity — challenging assumptions. This is the fourth-literacy argument I have been making for two years, written in the assessment field’s own vocabulary: as AI absorbs routine cognition, the capacities that remain distinctively human are the ones debate both trains and, crucially for this report, observes.
Their toolbox is a tournament, disassembled
The report’s most striking section is its inventory of tools for measuring human-centered skills. Victor Lee concedes that for AI literacy “we don’t really have a clear articulation for what that is,” and the paper admits adaptability “lacks a standard definition with operational indicators.” Fine. But look at the eight tools the convening surfaced, and notice what they assemble into:
Cognitive interviews. Cross-examination is a cognitive interview, conducted by a maximally motivated peer. The think-aloud is debate’s native genre.
Live demonstrations. ETS’s Lydia Liu calls these “the most authentic way to collect evidence” about what learners know and can do, then flags the problem: wide application “will require operational efficiency for feasibility.” Feasibility is debate’s solved problem. A weekend tournament runs hundreds of students through five to seven live demonstrations each, with trained volunteer raters and mature tab software, at trivial marginal cost. A century of logistics R&D, free for the field to borrow.
Role-playing. Congressional Debate is sustained legislative role-play. Policy debate role-plays the policymaker. World Schools assigns positions regardless of belief. Patrick Kyllonen’s AI role-play partner for measuring social skills is an automation of a thing debate does with humans.
Scenario-based assessments. Every resolution is a scenario. Disadvantages and counterplans are structured consequence reasoning about non-traditional, real-world situations.
Portfolio-based assessments. A debate career is a portfolio: case files, redos, ballots, records across four years. Growth demonstrated through “a curated collection of work over time” is simply what the activity produces as exhaust.
Conversation-based assessments. Diego Zapata-Rivera’s evidence-centered-design dialogue agents are an engineered version of what the round and CX already do. Debate is the original conversation-based assessment. His line — “With AI, we could create expert agents to generate questions or create opportunities even if we are not there” — describes AI sparring partners and practice judges, the most obvious near-term AI application in debate coaching, and one I use with my own students weekly.
Virtual robotics tasks for persistence and resilience signals. Debate makes persistence visible without instrumenting anything: rebuilding after an 0-3 start, prepping out a bad matchup, the multi-year novice-to-varsity arc.
Game-based assessment. Debate is a game-based assessment, and the competitive frame elicits the maximal authentic effort that test designers struggle to motivate.
Eight tools. Eight components of a tournament, described separately by people trying to invent each one.
Continuous assessment without the surveillance
Here is where debate resolves a tension the report raises but cannot dissolve.
The paper is excited about stealth assessment — systems that “continuously extract signals from a learner’s everyday learning environment” — and about continuous assessment, which Kyllonen suggests may reduce test anxiety and assessment fatigue. Then come the costs. Privacy, per Meta’s Vikas Wadhwani. Learner awareness as a contaminant. And Ben Domingue’s warning, the most quotable caution in the document: “The more data we’re able to collect passively, the harder those questions are going to become.”
Debate offers continuous, longitudinal, multi-source assessment that is entirely overt. Nothing is collected passively. Students know exactly when they are being evaluated, by whom, against what published criteria — and performing under known evaluation is itself part of the construct rather than a distortion of it. A season’s record is longitudinal data with the consent problem solved by design.
Google’s Julia Wilkowski cautions, in the report’s paraphrase, that more evidence does not automatically produce better insight. Also solved, socially: in debate, every data point arrives pre-paired with a trained human’s interpretive rationale. The RFD is the insight, shipped with the data.
Explainability was solved socially, decades before it became a research program
The trust section of the report belongs to Emma Brunskill: “The reason we often want transparency is really an issue of trust.” Her microwave example — users don’t need to understand every mechanism, but they need faith in a broader system of safeguards — is a good frame. Now apply it to debate.
Published judge paradigms are pre-disclosed evaluation criteria. The oral RFD is explainability delivered to your face, in minutes, not in a technical appendix. The post-round is contestation — the assessed party interrogating the rater about the rationale, immediately, in person. What other assessment in American education lets you ask the scorer to justify the score to your face?
And debate goes one level deeper than anything in the report’s co-design section. Victor Lee says the field’s strongest tools are “to open participation and to be inquisitive,” and that “there’s amazing brilliance that exists in so many communities that may not always appear in assessments.” Beautiful — and debate is the only assessment ecosystem I know of where the assessed openly contest and revise the assessment criteria inside the assessed performance. Theory debates are students arguing about what the norms of their own evaluation should be, and winning those arguments changes community practice. Students vote on topics. Norms evolve through argument. That is co-design at a depth the assessment field only gestures toward.
Temple Lovelace defines success as learners becoming “self-regulated enough to chart their own learning pathways.” That is a job description for a varsity debater: self-directed research agendas, chosen literatures, owned strategic identities. Nobody assigns a senior their 1AC.
What debate has to concede
To claim this ground credibly, we have to concede what we haven’t done — and I want to be precise here, because the concessions cut toward partnership rather than against the argument.
Speaker points are noisy. Judge variability is real, and human judging has its own biases; debate’s answer is transparency and contestability, not the elimination of variance. Access is uneven and the activity reaches a small fraction of American students. Competitive intensity can narrow learning when coaching cultures let it. And debate has never systematically documented its own validity evidence in the assessment field’s terms — no construct maps, no inter-rater reliability studies at scale, no published technical documentation of the kind ETS produces for automated scoring.
But notice: that last gap is precisely the report’s research agenda. The paper calls on researchers to “advance ecological validity,” to design “assessments embedded in authentic tasks and environments,” to study “performance across contexts rather than isolated conditions.” Debate is a natural laboratory for exactly this work — authentic tasks, real stakes, decades of transcripts and outcomes, and human rationales sitting there as ground truth for AI-scoring research. Debate doesn’t need to become a test. The assessment field needs to study the thing that already works, and debate needs the psychometric partnership to make its century of authentic performance data legible.
The ask
Run down the report’s action checklist for education systems: “pilot continuous, embedded assessment models”; adopt portfolio approaches; “reduce reliance on high-stakes, one-time testing”; build educator capacity; engage learners and families.
A school that invests in a debate program does all five simultaneously, this year, with no new technology. Continuous embedded assessment: the season. Portfolios: files, ballots, records. Many low-stakes events instead of one high-stakes one: the tournament calendar. Educator capacity: judge training is rater training. Engaging families: public rounds that parents literally judge.
The infrastructure exists — Tabroom, the NSDA, judge pools, topic committees, a national calendar — and the hard part, a trained human rater community with shared norms, is the part debate already has. Expanding access on the urban debate league model costs orders of magnitude less than building new assessment platforms. Stanford’s own ROAR project shows the pilot pattern the report endorses: partner with districts, run the authentic assessment, publish what you learn. The adolescent version of the MAGIC project’s constructs — curiosity, creative thinking, problem-solving — already runs at national scale every weekend from September to June. It just needs the research partnership, and the sustained funding Sandip Sinharay rightly warns these initiatives die without.
The report ends on its best line: AI “cannot be held accountable for the consequences of assessment decisions. People can.” Accountability is debate’s grammar — this is what I mean when I talk about a pedagogy of standing. Claims defended by a person. Decisions rendered, with reasons, by a person. Each answerable to the other, out loud, in a room.
Thille’s opening question — “How do we shape the current moment?” — is itself a debate question. And the activity’s answer is the one this excellent paper circles for twenty pages without naming: put students in situations where they must know, reason, respond, and answer for it, in front of people who give reasons back.
We don’t have to invent the future of assessment. We have to recognize it, fund it, study it, and open the doors wider.
Dan Schwartz told the convening: “We need instruction that produces adaptive learners and assessment that can tell.” I know where there’s some. If you work in assessment research and want to talk about debate as a research site — ballots, transcripts, longitudinal records, a rater community a century deep — my inbox is open.


