0:41 - Dr. Anand Rao: I’m thrilled to introduce our guest this week, John Hines. John is a true leader at the intersection of dialogic learning, AI, and debate with more than two decades of experience as a national champion debate coach and curriculum innovator.1 He’s the principal investigator of a project at DebaterHub, which he co-founded, which is funded by the NSF, SBIR, AWS, and an NIIR grant. His work focuses on developing augmented debate-centered instruction, a novel approach that uses AI to enhance students’ cognitive and communication skills. His expertise covers a wide range of topics, including debate and argumentation-based pedagogy, AI ethics, and educational technology.
John has presented his research on responsible AI integration in education at major conferences like the AAAI 24 workshop on AI and education, ASU GSV in 2025, and the ALTA 2025 argumentation conference.3 Through his collaborations, John is actively working to scale debate-centered learning with ethical AI, transform critical thinking skills, create equitable pathways for student success, and pioneer new standards for responsible AI in education.4 We hope you enjoy this interview with John Hines. John, thank you so much for joining us for the podcast. We’ve been looking forward to this discussion for quite a while. Why don’t we start off by you telling us a bit about DebaterHub? You know, what is the platform? What are your goals for it? What should people listening know about it?
2:08 - John Hines: Sure. Yeah. Thanks. First of all, thank you for the opportunity to to share about DebaterHub and talk about our our journey. DebaterHub started about four years ago. Um, the the initial impetus was a shared concern with my my partner and I about the general state of democracy and state of education and a really firm belief on our part. We’re both debaters. I was a debate coach for much longer, but he debate was very influential in in his career as a high school student that a speech and debate-based education is transformative and powerful and learning to communicate across difference is something that’s immensely important, not just for somebody’s professional success, but also for our democracy.
We really felt like the state of civil discourse was in rapid decline. And we felt that debate was the most powerful activity we had ever encountered, and we wanted more people to have access to debate. The big problem with that is over time, debate has become an activity that’s almost the exclusive purview of wealthy suburban schools or private schools. And the vast majority of kids don’t get access to debate-based education.
But I I really firmly believe that our constitution and and the founding fathers of this country presumed a citizenry capable of disagreeing across difference. That that’s the brilliance of the American constitution is we don’t need a king deciding things for us and telling us what to do, that we can debate each other, that we can disagree with each other, and then we elect representatives who go to Congress and debate on our behalf. I really firmly believe that our constitution presumes a citizenry capable of debate. Um, but you know, in the 1800s as a society, we realized we couldn’t scale debate-based education. But that’s what was happening in our colonial chartered colleges. Harvard and Dartmouth and Yale, those students were debating regularly in their classrooms, like that was their form of assessment. So that was the the initial impetus for DebaterHub was, you know, can we create platforms and tools that give more people access to debate?
Um, and um, shortly after we formed, ChatGPT hit, and just like a meteor exploded, you know, and everybody got really excited about AI. Um, and it really kind of accelerated our timeline because basically we could take the debate architecture that we were building and just start plugging these LLM models into it and and getting it to debate. Um, and eventually we secured a National Science Foundation grant, um, which we just completed last week and and ran a number of very interesting results, which I’ll be happy to share later.
Um, but at this point, DebaterHub has accomplished a fully autonomous debating system that we have created, an AI system, an AI workflow, a multi-agent workflow that is capable of participating at human level debate. In fact, the the benchmarking test we just ran, that our system can create the first speech in a policy debate, which is the most rigorous form of competitive debate, um, it can create that first speech at exceptionally high levels. In fact, according to our benchmarks, it exceeds expert human performance at that level and and absolutely crushed zero-shot attempts at the most expensive consumer models out there. We paid for the $300 Groq, the $200 ChatGPT, and they couldn’t even come close to what our system was producing.
More long-run, then our goal is to figure out a way to to take that fundamental scientific um discovery that we’ve created, technological discovery we’ve created, and turn that into a usable user interface, user experience, and collaborate with urban debate leagues, Title One schools, and other schools that are interested in providing more people access to debate-based education.
5:48 - Stefan Bauschard: Well, and thanks John, you know, for choosing our podcast to tell the world you’ve achieved AGI.
5:53 - John Hines: I wouldn’t call it that. It’s narrowly defined, right? Like AGI presumes it’s across activities, right? But you got superhuman intelligence in...
6:05 - Dr. Anand Rao: : Maybe what they discovered was ADI, artificial debate intelligence.
6:11 - John Hines: Yeah, that’s probably a good term for it, is is we have artificial debate intelligence in that our system can achieve very high levels of debate capability with much more efficient models, actually. Like we used very cheap models to accomplish this. And that’s probably one of the more exciting aspects of it, is that we didn’t have to use these supermodels to generate this exceptional performance.
6:32 - Stefan Bauschard: You know, you know, so you know, that’s kind of the one one of the findings, right? You know, we don’t we don’t know if, you know, actually I saw I saw a study yesterday where somebody used like ChatGPT like 2.5 or something, but they put all this medical data in it and it was able to complete all these diagnoses. So, you know, that’s an interesting finding of your study, and maybe that is true in other areas. But under the NSF grant, what kind of, you know, what what kind of did you apply for in terms of like, okay, we want to get this grant so we can do X, Y, and Z? I mean, part of it was obviously kind of to develop the system to prove it out. But, you know, what kind of maybe objectives were there beyond that, right? Okay, let’s let’s see if we can get, you know, close to human or now superhuman performance and argumentation. What what all are we looking at under the grant?
7:16 - John Hines: Yeah, well, let’s be clear about the kind of grant we received. It was a Small Business Innovation Research grant. So, there’s two kind of fundamental things that you need to prove to get the grant. One is that it’s scientifically novel research, um, that it’s not like just an engineering problem that you’re taking existing scientific understanding and applying it into a new engineering area. It actually has to be fundamental scientific research. Um, so that was challenge one. Challenge two is then the idea is you’re going to turn this into a business, um, that this could be and and within them it’s the idea is like a scalable venture-backed business is where you’re trying to head with that.
I think we successfully accomplished number one. I don’t think we necessarily successfully accomplished number two. Um, I I think that our conclusion at this point is we’re probably more inclined to convert our company to like a nonprofit than to try to go get venture-backed investment. And we can talk about why that is. It actually, a lot of my research into edtech and educational technology and what’s happening in education spaces, I’m not entirely comfortable to be honest at this point, um, saying we need venture-backed venture capital in higher ed. I’m a little concerned about venture capital in higher ed. We can talk about that as well.
So, the scientific questions though, right? I think that’s kind of where you were going, Stephan, with like, okay, what were the scientific questions? The first question was, can we create this fully autonomous debating system? Can we take this debate stack, right, and and accomplish this? And the inspiration for that is a project called IBM Project Debater. Most people that I bring this up, even within the debate community, they’re completely unaware that there was a third grand AI challenge for IBM. All right. So a little bit of AI history here. All right. Before ChatGPT and all that exploded, and even before like AlphaZero, AlphaGo, which made a lot of noise, IBM had a history of doing these grand AI challenges. In 1997, they conquered chess, right? They they beat Garry Kasparov famously with an AI system, winning, you know, human expert in chess. That made a bunch of noise for them. Then shortly after that, they beat Jeopardy, like I think they beat Ken Jennings, right, in in Jeopardy with their AI systems. And their third attempt, and and that everybody tends to know that when I say those things, everybody shakes their heads, right? And I go, “Okay, well the third one was IBM Project Debater.” And they’re like, “What are you talking about?” I was like, they tried to do debate. Um, but what they did is they they created a kind of a mock debate format, right? It wasn’t an actual competitive debate format that we would recognize as competitive debate coaches. Um, and they created kind of a very narrowly defined definition of debate. Um, and they they tried to put their system into a public debate against the world’s best parliamentary debater. They announced it was a tie of this public debate. You laugh because you both of you are laughing because you’re debate coaches or former debate coaches yourself, and you know there is no such thing as a tie in a debate. There’s a winner and a loser. Um, and ultimately, IBM kind of packed up Project Debater and moved on. So most people don’t know about it. But they did a lot of really interesting research, well, a lot of great scientific discovery that’s a lot of predecessor to what’s happening now.
So that was kind of our first step, was can we accomplish this Project Debater ideal of putting a debate, creating a debate system? And the key differentiator between us and IBM is IBM ultimately decided that they didn’t think debate as they understood debate was appropriate for machine learning because it wasn’t really a game was their conclusion. They published an article in Nature announcing that debate is not a game. And therefore, that’s why they failed, because machine learning requires gamification. And that’s our point of disagreement with IBM, is they chose the wrong format. That if they were familiar with American forensic-style debate, we definitely believe what we teach and coach is a game, and it has rules, and it has a game board, and it has clearly defined rules and moves. We get to debate the rules, but you know, it it is a game, and in certain respects, it is gamifiable and amenable to machine learning.
So that was our step one. Can we do that, which we think we did. The step two was then try to put this into classrooms and test this into classrooms with that other thesis that I was talking about earlier, which is debate is the original kind of like best version of learning in a classroom, right? Argumentation, dialogic learning is an outstanding format for learning. The problem is we’ve moved past it over the past 100, 120 years, because we were trying to put everybody in our education system, right? This kind of utilitarian, you know, populist education that America, you know, started pursuing in the 1800s, 1900s, where we wanted to get everybody an education. Well, if we wanted to give everybody an education, then we started to do factory-based education and couldn’t do this kind of one-on-one dialogic education anymore. So our our second challenge is to implement, use these tools to implement dialogic education in classrooms, and then demonstrate that we’re getting, you know, effective results.
We weren’t able to do those studies because of kind of the political chaos that happened at the beginning of 2025, and we were getting vague directives from the National Science Foundations about things we were and weren’t allowed to do. Our research partners, which were urban debate leagues, were also getting vague directives about what they were allowed and not allowed to do. So we we’ve kind of put a pause on the human study and the educational impact of what we did. So we don’t have any data on that yet. Um, that’s kind of going to be our next step, right? So now we’re we’re pursuing grant money to do that next step of research, so that we can demonstrate the efficacy of these automated debate systems in actual classrooms where we can give, you know, instructors who are looking for new types of assessments, right? They’re worried about their essays and and everybody just turning in ChatGPT-generated work. What if we took those assessments and created debate-based assessments, dialogue-based assessments, oral assessments, but also leveraged AI systems to help make it a little more uniform, a little more predictable? So that’ll be our next step of research.
13:14 - Dr. Anand Rao: You’ll be able to take that on now that they’ve resolved all of those other political conflicts.
13:20 - John Hines: Hopefully. Hopefully. I mean, there’s a lot of interest in debate, right? And, you know, there’s some questions about what we consider debate to be. But I think, you know, it should not be a controversial thing to say in education that we need more debate and that we need more dialogue across difference and that we need to learn to communicate effectively with people when we disagree with them. And I think that it’s it’s really important that we centralize that in our classrooms. And I think that, you know, ChatGPT and AI and and the impact that’s having in higher ed have made clear that we need new assessment modalities, that our traditional assessment modalities that evaluate products at the end as a snapshot of learning, as opposed to actually evaluating the learning itself. That’s what we need to start doing, right? We need to start actually evaluating the learning and and how they’ve learned as opposed to a snapshot at the end that attempts to assess how much they learned in that containerized moment right there. And and moving more towards process and evaluating the process of learning and and being able to evaluate that you actually are evaluating what they’re learning and how they’re learning and their capacity to learn as opposed to this is the amount of facts they memorized in a short period of time.
14:33 - Dr. Anand Rao: Now, I remember talking with you earlier in this process while you were working through the NSF grant, and you referred to a series of interviews that you conducted. And so, tell us, you said they were I-Corps interviews. Tell us what what that means. But then probably more importantly, really want to hear about some of the findings that that you were able to gather. It was really fascinating to hear that maybe cheating wasn’t the absolute top concern.
14:52 - John Hines: Yeah. Yeah. So, the I-Corps is a secondary grant you receive, you apply for and receive when you get into this NSF SBIR grant program. So, the SBIR, the idea of the NSF SBIR is you’re going to take, you know, innovation, you know, scientific innovation that’s happening in a lab and take it out and make a business out of it and turn it into a product. And so, I-Corps is kind of like their own little incubator kind of program, kind of like NSF’s Y Combinator for people who know what Y Combinator is. The the purpose is to to take these scientists in these in these academic labs and teach them how to like think like a business person and what it takes to to go from, okay, we’ve got an invention, how do we take this invention and turn it into an innovation that becomes a product that people pay for? And a key step of that of like product thinking, product-oriented thinking is customer discovery interviews. All right? Um, that it’s really easy to say, “Ah, I have a brilliant idea. Let me go make some things and then hope people use them.” Which is the way most people approach, you know, launching a company, launching a product. The reality is, until you, as Mike Tyson famously said, everybody has a plan until they get punched in the mouth. The same is true of business, right? If you don’t spend time talking to your customers and getting a deep appreciation for what’s going on with your customers, what they need, what they’re what what’s going on with them, what they’re trying to accomplish, what motivates them. If you don’t spend a lot of time really talking to them and getting a deep understanding of what they need, they’re not going to use your product, right? Because your product wasn’t built for them. So the I-Corps kind of teaches you all that. And then a key component of it is you have to conduct at least 100 customer discovery interviews. So we conducted over 100 interviews with higher ed administrators, professors, deans, presidents, and you know, also, you know, you know, CEOs of other edtech companies and a wide variety of of, you know, people within the world of higher ed and and also K-12 education as well, right?
Because our fundamental business thesis was educators are concerned about AI. That they’re concerned about AI cheating. And that they need a technological solution for that. And that that’s what we could offer, right? That we could say, “Hey, yeah, the essay is dead, right? You may not be willing to admit the essay is dead yet, but the essay is dying and soon to be dead as a valuable mode of assessment, and we can create a replacement assessment for that.” But what’s really interesting about the way these customer discovery interviews are supposed to work, if you’re really doing them correctly and scientifically, is you’re not coming here saying, “Here’s my product. Tell me what you think of my product.” What you should just be doing is asking a bunch of open-ended questions about their day-to-day experiences, what are their challenges, what what, you know, what’s bothering them, and how they’re trying to solve that. And in that process, I discovered that our thesis that like AI cheating is this massive crisis in higher ed was not entirely true. Um, it it’s a nuisance to them. And in certain areas, it’s definitely more impactful than others. Like I would say people who run writing centers are really impacted. Um, people who teach like first-year composition are really impacted. People who teach public speaking courses are really impacted by people turning in AI-generated work, and it’s creating a lot of problems for them. But at the university as a whole, there were much bigger fires burning at that time than than that issue.
The number one fire for the universities is the the the problem that they’re struggling to demonstrate the value of degrees to universities, and and it has to do with the cost of education. All right, the cost of education itself, of getting the degree itself versus the the the lifetime potential earnings from that degree are getting close to each other. And so it’s really hard to demonstrate the value or the return on investment to higher ed. That was kind of like the number one concern that I was getting, is there’s a lot of pressure coming from top-down to demonstrate return on investment to the like course level even in in some states, right? Like in Texas and in Iowa, they’re getting directives from their state-based funding agencies that you need to be able to track like a student took this course and then they got this job based in this course, and the skills you’re developing produce these types of jobs. And that’s kind of the new, you know, standards that are being set. And it’s really hard for for universities to meet that standard.
The second major crisis that they’re facing is the demographic cliff crisis, which is our, you know, birth rates have dropped below replacement rate, right? And the way the typical university funding model works is enrollment and financial aid based in enrollment and tuition is really what’s keeping the lights on. But if you already have this problem where the cost of degree is getting close to the value of degree, and you have this enrollment, declining enrollment problem, there’s an existential crisis in higher ed that that that precedes this question of what’s happening with AI. And I I really think AI just kind of lights everything on fire, because at the same time, it just demonstrates that like what we’re assessing, what we’re teaching is also problematic, right? And so I really think what’s happening with AI is it reveals this existential crisis as an actual existential crisis, and it kind of speaks to like, well, what is the value of the degree I’m getting, right? If an AI can like write all my write all my essays for me, and if it can complete all the multiple-choice tests you’re giving me, what am I actually learning in your classroom? Like what skill am I learning in your classroom, right?
And and you know, one thing that that I want to acknowledge is like there’s kind of a a battle in higher ed between two different visions of what we’re trying to accomplish in higher ed, right? Some people were really firmly like, I’m not here to get you a job. I’m here to like build a better person, a better thinker, better citizen, etc., etc. But the vast majority of the students and the parents who are paying the bills are there to get a job, right? I mean, me too when I went to school, I went to school because I wanted to get a job on the other side of that, right? And so AI kind of exacerbates these problems, exposes these problems, and also explains why you have so much cheating around these problems, right? Or cheating, right? And I even put cheating under quotation marks because I think what a lot of students are doing is imminently rational given the world they live in, right? That all that matters is a degree, and the way to get the degree is to get the grade, right? We’re not assessing actual learning, right? Like if we’re we’re taking snapshots of learning, but there’s not an incentive structure in education right now, and it’s not just higher ed, you know, high schools also have this problem, there’s not an incentive structure for the students to actually learn. They’re only incentivized to get a good grade so they can then move on to the next thing. So they just have to be as efficient as possible to get that good grade to move to the next thing. They have this tool that makes it very easy to get that grade, so it’s kind of logical that they use it. But you know, that’s also why we have a bunch of professors and like, let’s just ban it. But I I just don’t think that’s realistic to ban it.
21:48 - Stefan Bauschard: You know, it kind of brings up, you know, something to think about, right? It’s like, you know, professors, two things. One, professors kind of implicitly grade on a curve, right? If not explicitly. Yep. And you know, if work just kind of starts getting better because students are using AI and, you know, then you have someone who’s like maybe was a B student, but now they’re not using AI, right? Like their their work’s just going to kind of fall off like relative a little bit. But my my more significant thing I want to get to is that, you know, you talk about like how, okay, well we had this grade that I guess was kind of like a marker of of the skill we assumed, right? And that collapsed too because everybody just started getting A’s or B’s. Right. So inflation crushed all that. Right. Did it make much of a difference anyhow when you applied for a job and somebody’s like, “Well, now I’m going to I’m going to hire the person who got an A minus, not a B plus.” Right. It was kind of right. It kind of became a little it started becoming a little bit of a meaningless marker. But more importantly, do you think like you, you know, what you’re kind of proposing for DebaterHub or at least in terms of how you want to develop it, you know, and you know, there’s these ideas of we can assess what what students can do, right? Not just kind of the content that they’re going to learn or not maybe our reflection back on like, “Oh, I think they gave a good an A speech or a B speech,” right? But do you think somebody could kind of, you know, whether they graduate or they just kind of take some courses and then they go apply for a job and they say like, “Hey, you know, I’m like a pretty good public speaker.” You know, the employers are, “Well, I don’t know. I can’t even speak for two minutes.” But you have this whole kind of like uh, you know, record of like, “All right, like here’s like whatever, uh, you know, 20 public speaking skills and, you know, this person gave nine speeches and, you know, this is how the system’s rating them or how they’re interacting with people in a debate.” Like do you think it can measure, would you like it to be able to measure some of these kind of soft skills?
23:21 - John Hines: Yeah. Yeah. And and I’d even push back on calling it soft skills. Um, you know, we used to call them soft skills, right? We hard skills and soft skills, right? Hard skills would be those hard sciences, the really hard problems, which by the way, the AI now does better than humans. Uh, and the soft skills were like, all those those touchy-feely human things, they’re soft, right? It turns out if you look at like the World Economic Forum Future of Jobs report that just came out at the beginning of this this last year, uh, the seven out of the top 10 most valuable skills that employers are looking for are soft skills, right? Uh, which I think it’s more accurate to describe them as durable skills or power skills, that these are the actual human skills that are the employable skills that we need to send signals to employers that the students have accomplished.
Um, but they’re things like perspective taking, active listening, creativity, critical thinking, and and we don’t have reliable measures for those things. But we do measure them in debate, right? When we do public speaking and debate activities, we’re measuring those types of skills. But we’ve never done a good job of communicating how we’re measuring them as a community. So that’s that’s kind of like another area of research, that’s that next step of research, which I I tend to term dialogical proficiency. It sounds kind of wonky, but I’m looking for a better term of how do we how do we classify and and evaluate an individual’s ability to communicate across difference, to collaborate in a complex situation with a team, right? I I think the future of work is going to require that level of communication with AIs and and with other humans. Uh, that we need communication skills. Human communication skills are going to be the most valuable skills of the future.
But we need to develop reliable measures that send signals to potential employers or like the next level of education of how you’ve developed that skill. All right? Uh, and that skill is something that’s hard to measure because you measure the ability to do that thing over time and across spaces, right? It’s not something you can take a multiple choice test and and demonstrate, “Oh, you’re a really good perspective taker,” or “You’re a really good empathic listener,” or “You’re a really good creative thinker,” right? Like the way to measure those. We’re going to need to develop measurement systems that track a student’s ability to do something like that over time and across courses. So we need to kind of rethink, and I think this even challenges like the major system, right? Um, and and the way we think about evaluation, you know, the Carnegie unit, right? Uh, I think that this really calls into question a lot of the fundamental assumptions that we’ve had for a long time about how we assess, how we evaluate, how we instruct. Uh, which is why I think that this is a long-term project. It’s not something that’s going to be solved tomorrow, next year, or even in the next couple of years. That there needs to be a long-term re-evaluation of what we’re doing in education.
26:07 - Dr. Anand Rao: And I think you outlined pretty well that this is not a a new issue or a concern. It’s just that it really isn’t hasn’t come to to to the fore because uh until really AI has helped illustrate some of the challenges and some concerns. So we think about this dialogical proficiency is um developing this as a model as a response to some of the challenges presented by AI. How can we also use AI to develop this model? Right? I mean, if you’re considering how this works because you’re right, if we think about a one-off course and you have an individual communication class, even in that one class, there’s only so much I could do to really evaluate listening skills, let’s say. Um, and I might be able to say, “Yeah, they were able to listen on a certain level here.” Um, but there’s no way to tell if that was developed, if that’s something it develops over time or to be able to develop proficiency or what that proficiency would look like. How can we do that with some of the the the AI skills or or tools that you might have with something like DebaterHub?
27:02 - John Hines: Yeah, the answer to your question is is a field of uh, you know, computer science called computational argumentation that tracks things like that. Um, and there’s new tools, you know, you know, DebaterHub is part of a larger movement. There’s another tool called Debatrix, which is able to track moves in a debate over time, right?5 Uh, and so what we can do is develop systems. So if we can develop debating systems, you know, systems that are capable of understanding how argument and rhetoric and persuasion and communication and perspective taking evolve over a period of time, right? And we can make systems that are capable of doing that, then that indicates that we could also create systems that are capable of tracking and evaluating that, right? Uh, and so if we can create systems that are capable of tracking and evaluating rhetorical moves, right, or long-term planning, long-term decision-making, uh, perspective taking, ability to to be pluralistic in your approach to a problem, and and making those objective and deterministic, right? Because that’s always why they’ve been considered soft skills, is they weren’t considered objective or universalizable, right? And what, you know, the breakthroughs that are happening in computational argumentation, uh, things like DebaterHub, IBM Project Debater, Debatrix demonstrate that these things that we used to think were soft skills and didn’t have objective and deterministic qualities to them actually do have computable qualities to them. Uh, and if they have computable qualities to them, then we should theoretically be able to build systems that can track and analyze those those those trends over time and and then make a case for objectively across courses, across multiple instructors, um, across departments within a university, um, that this person started here their freshman year with these level of skills for perspective taking and creative thinking and critical thinking and active listening, uh, and but by the time they went through the gen ed curriculum, they developed to this uh level.
And I think that’s a much more valuable skill to an employer, right? Like if you look at, you know, what happens when students matriculate out of university and then go get a job, I’m going to throw out a statistic here, it’s probably wrong, but I want to say like somewhere to 70-80% of employers say they then have to completely retrain or train the hire. Right? So the signal that employers are looking for from a university is an individual’s capability to learn and grow. All right? That’s the signal they’re looking for because they don’t think they’re coming with the skills they need anyway. Um, so you know, I think one way that universities can start to demonstrate the value of the degree is that they can signal, “Oh no, the skills you said you want, we’ve imparted, and we can actually demonstrate how the students developed those skills over time. We can give you a sense of the trajectory of how long it took them to develop those skills and how quickly they’ll be able to acculturate into your system by giving you those types of signals.”
I think those will be much more valuable signals to employers in the future than, “Oh, this student got an A in this class when everybody knows grade inflation’s happening.” Um, everybody knows the kids are using AI to cheat. Uh, so the value of the university degree signal to an employer is is is pretty non-existent right now, and the skills they’re looking for aren’t tracked by those grades. Uh, so I think that’s kind of why we have like a real problem. Like if you look at the employment metrics for the the recent grads, it’s like the worst it’s been in a really long time. That it’s a really difficult time for recent grads trying to get jobs because there’s a fundamental misalignment between the signal that comes from their degree and what the employers are looking for.
30:35 - Stefan Bauschard: Yeah. And I think, you know, it’s, you know, to think about, you know, the role of the university, right? Like skills develop over time, right? Like, you know, even even the best debaters, you know, it’s funny, sometimes I’ll tell a story, “Oh, I lost my first six debates” or something like that, right? You know, which is more than you’re going to have in a debate class, right? So it, you know, the university, you know, however long, maybe students won’t do four-year degrees anymore, maybe two years they’ll become more, who knows, right? But whatever time they’re there, it’s like they’re going to be in the process of developing these skills. And universities too, we’ve been talking about universities, but you know, if this, you know, we’re talking about knowledge graphs, right? That that could start early, right? Like then it’s just a question of like how much kind of they improve that that could be charted. But the other question I want, the question I want to ask is, you know, we talked a lot about argumentation and argumentation skills, right? And as you know, right, there’s this kind of historical debate about like, “Oh, you know, the substance of the argument versus like, you know, the rhetorical skills.” And, you know, for those of you on the podcast, you don’t know, Anand and I used to debate each other. And, you know, Anand was really suave and persuasive. And we used to call him, we used to call him GQ Man when we lost because we’d get frustrated, uh...
31:34 - Dr. Anand Rao: ...to persuasive. It was really the power of my argument. It wasn’t just rhetorical.
31:40 - John Hines: Yeah. Yeah. That’s true. It sounds good. But, you know, obviously, you know, to a degree, they’re intertwined, right? I mean, you can’t persuade someone with a terrible argument. So I’ll give him, I’ll give him a little bit that his argument was okay. But, you know, does it, first of all, you know, I I would imagine like it can’t, it can measure these things, but it it kind of gets interesting questions, right? Like, “Oh, well, okay, this person on and he’s super persuasive. He’ll be able to convince anyone of anything if he comes to your business or works in your government.” Um, you know, and I, you know, at NCA, you know, I kind of reviewed like some of the the research, and this was what, November 2024, that already showed pretty conclusively much more persuasive than people in a in a digital format, right? So, you know, do you have any thought, you know, just kind of briefly, I I think we know where we’re going on, can it can it measure the skills you kind of been talking about that a little bit, but you know, I could see people reacting a little bit negative, “All you’re going to get an exact report on, you know, how good Anand is at, you know, kind of manipulating the office.” Right. Um, not not not to you know, it could be anyone. It could be maybe I want to be much more persuasive. Do you have any kind of thoughts on how people are going to react to that or what?
32:47 - John Hines: Well, I I agree with the concern with that, right? Like that’s the danger of this technology, right? Um, the, you know, there’s a potential future where every human has a team of super persuasive AIs that are assigned to them to manipulate them in the market. It’s dystopian but totally technologically possible. Um, and so I think, you know, the role of the university, the role of educators is to teach ethical communication, an ethical and responsible use of rhetorical skills, right? And and, you know, students need to learn the difference between rhetorical persuasion and effective argumentation, right? You know, you all kind of implied the difference, right? Like there’s there’s the logical argument that you make, right, that’s grounded in reason and science. And then there’s rhetorical persuasion, which can be a little manipulative, right? Because it taps into things like emotion.
Um, and so, and and I will say I’m a little skeptical of some of the studies that have already come out about saying AI is super persuasive, because when you look at the research that happens in computational argumentation and and the the scholarship that’s happening there, there hasn’t been sufficient, in my view, sufficient engagement with rhetoric scholars to to make those conclusions, right? And so a lot of computational argumentation theorists equate argumentation with persuasion as being the same thing, right? That if you’ve made an effective argument, you’ve been persuasive. And scholars of rhetoric will say that that’s definitely not exactly the same thing. That there’s differences there. But I do think in some of these studies where they’re like, “Oh, the the system was super persuasive,” it was super persuasive because it was able to kind of identify the the target that it was trying to manipulate and make arguments specifically for that target. So it’d be much more useful to examine like what those rhetorical tools they used and understand that.
But my answer is kind of like what I would call an inoculation theory, that we need to be training our students to be super effective communicators and super astute communicators so that we can protect ourselves from these attempts at rhetorical manipulation that are, you know, that that theoretically becomes superpowers now of the AI. You know, so our systems are primarily trained much more on like kind of logical argumentation and then, you know, ethical acts of persuasion. So I think those guardrails we put around it as educators are really important, but I think that’s that’s the role of the university, which kind of goes back to this question that I brought before of like, are we teaching people to go get a job, or are we teaching people to be citizens, you know, in a democratic republic? And I I think debate, ethical debate, allows us to do both at the same time.
35:30 - Dr. Anand Rao: Yeah. So a lot of the discussion has been about individual skill sets, developing those skills for the person. Um, you know, I think that there’s an area to explore where we consider them as cyborgs, where it’s a human using some of the tools to expand their capabilities. Before we get to that, though, the one thing I wanted to ask, which is a necessary next step, is the idea of understanding how arguments are used by AI models themselves. And and you all did a lot of work on this, especially thinking about like a simulated debate tournament. Tell us about how you play, how that has played out. What what did you find when you were running that tournament? Um, was it, and was the goal really just to kind of replicate what humans are able to do, or did it start to develop its own capabilities, or you could see the potential for capabilities, so it’s different than what humans would do at a debate tournament?
36:18 - John Hines: Yeah. So let’s talk about the experiment we just ran, and some of the problems that we were trying to address from the get-go, because, you know, I’ve been suggesting people in debate use AI for years now, and there’s a lot of pushback, right? There’s a lot of pushback on hallucination, that it’s just going to make stuff up, that we can’t trust it, it’s not reliable, or that, you know, it’s it’s going to consume debate and turn it into a monolith, right? That it’s just going to make...
36:44 - Dr. Anand Rao: Of course, you’re probably mostly just hearing from the people that don’t want other teams to use it, and you’re really not hearing from the ones that are using it, of course.
36:56 - John Hines: Right. Right. That’s correct. Or the people who they’re they make a ton of money paying people or getting paid by people to cut cards and do things like that. And they they’re worried about their job. Like let’s be clear, the people who push back the most are the ones that are worried about their job. Like, “Oh, I get paid to cut cards. How am I going to get paid in the future?” And it’s like, “Well, I understand that’s important, but I’m more concerned by the tens of thousands of people that are locked out of the community because they can’t pay you $50 an hour to cut cards for them.” Um, once again, it’s an elitist activity. Then, that’s my main concern is it tends to be an elitist activity.
So what we did is we wanted to test, the first step we wanted to test. We we we took the hardest format of debate, which is American policy-style debate. It’s the most rigorous format of debate. But because of some of the norms in the community, there’s a lot of available data that can be used. In particular, there’s a massive dataset that was created by one of the people that we do some of our research with called the Open Debate Evidence dataset, where there’s 3.5 million pieces of evidence that has been used in competitive debate rounds, and that you can use to then simulate and train how we would actually do debate. And in particular, what we wanted to start with is the speech that we call the first affirmative constructive. It’s the speech that starts the debate. And it’s a really important speech because it sets the ground for the entire debate. And a well-constructed 1A, you win the debate, if you’ve written your 1A properly, right? And, you know, national champion level debaters will acknowledge this, right? That it all starts with a very well-constructed 1A. That’s actually like a really hard challenge for AI to accomplish, because there’s a lot of, you know, norms and rules and expectations that are grounded in argumentation theory, rhetorical theory, and like evidence citation practices about like how we cite our evidence. In particular, the frontier models aren’t really good at citing evidence. They’re good at doing research, consuming what information is out there, and then generating its own text, right? So it’s not literal quotation, right? This is the hallucination problem, right? Um, that we had to deal with. So what we decided to do was to build a system that’s capable of generating a 1A that would be a human-level 1AC that would look like a policy debate 1A, which means that you do actual research with high-quality sources and then we’re going to quote those sources extensively, directly. Right. So, for those that don’t know, a policy debate speech, first affirmative constructive speech, 90% of what’s delivered in that speech is quotation from other people. Um, an extensive quotation, right? Like your your evidence is at a minimum a paragraph long, and typically three, five, six, seven paragraphs long. And then you’re underlining certain lines of that to support your arguments. And the other thing is that you have kind of these argumentative constructions that you have to do to make a policy advocacy that we would call our stock issues from argumentation theory. And so there are some burdens that the affirmative team has to meet, rhetorical burdens, argumentative burdens that the affirmative team has to meet in order to present a logical affirmative case.
So what we did is we created three categories of generated 1ACs. We took human-created 1ACs from the high school policy debate topic from the summer. There were 45 open-source first affirmative cases that were released on a website after the end of summer camps, right? So debate students in high school go to colleges over the summer. They spend about $1,000 a week to receive expert instruction from college debate coaches and college debaters and high school, you know, experienced high school coaches. And then at the end of that camp, you produce a bunch of files, a bunch of evidence, right? So we found 45 1As, first affirmative constructives for the current topic. We then had our system generate about, you know, and we we selected 22 of those, randomly sampled half of them. Um, then we had our system produce an additional 22 1As. And so we gave our system just the resolution as a prompt. And then the way we’ve architected it, we want our our system to be pluralistic. We don’t want it to focus on just like a single approach. And so what it does when we give it that prompt, it hypothesizes 50 different potential perspectives about debate and then generates 50 different 1ACs from that. Um, and so we had our system produce its its, you know, 1ACs, and then we randomly selected 22. And then we paid for, you know, the most expensive Groq, the most expensive Claude, the most expensive ChatGPT, and the most expensive Gemini. Put them all in research mode and gave them a very sophisticated prompt-engineered prompt. It’s like a seven-page prompt is what we gave them. But we wanted it to operate under the same parameters that the humans in the debate camps were operating under and that our system was operating under, right? Which is that you have to have a plan. You have to have a topical plan. It has to be inherent. It has to solve. It has to have advantages. It has to quote reliable sources, and those sources must be traceable. And your your evidence needs to be at least a paragraph long. And what you claim in your tags or what you assert the evidence says has to correspond to what the evidence is there.
Okay, so we had 66 1ACs. And then we had them, we just paired them up head-to-head against each other in a double-elimination tournament, right? So the, you know, random AF here, random AF here, paired up against each other. And then we created a very kind of complex, sophisticated evaluation rubric of this is what a quality 1A would accomplish. This is strategically what it needs to do, argumentatively what it needs to do, evidence ethics what it needs to do. And we gave a very complex rubric for evaluation for the system, and then we had it debate, you know, and you know, single double-elimination. So one 1AC against another 1AC, the system scored it. We did use LLM evaluators, and I can come in a minute I’ll talk about some of the problems with us doing that because there’s definitely a problem, and we you know, there’s better data will come when we use human annotators for this, but we weren’t allowed to under the the the terms of the grant, um, that we weren’t allowed to pay human annotators under the terms of our grant. So we used an LLM evaluator, but we used the most sophisticated Claude as our LLM evaluator. Right. So Claude evaluated them, scored them across that rubric. The higher score would advance, the lower score would move into a losers bracket, and you could make make your way out of the losers bracket by winning. Um, so but each 1A was scored and then would move forward. Okay.
Um, the top two 1As were our 1ACs that our system generated. Um, the average scores on aggregate for DebaterHub-generated 1ACs was an 80% off that rubric. The average score for the expert human-generated 1ACs that came out from summer debate camps was a 70%. And the average score for the really hyper-expensive, you know, $300 research-enabled Groq was a 50%. And and what those 1ACs did, what the the frontier models 1ACs did is they they mimicked the structure pretty well. Um, they, you know, they they created the plan text fairly well. Um, they did tend to write the same 1AC over and over and over. There wasn’t a lot of diversity in the what 1A they they generated. They really loved the the icebreakers affirmative. They wrote the icebreakers affirmative a bunch of times. Um, you know, the domain awareness, which for some of the more popular, you know, 1ACs at the camps, right, these systems replicated that pretty well. Um, where they really failed was the evidence production. In fact, a number of them, when we set the parameters around like you have to actually quote evidence, you can’t generate summaries, but it actually has to be evidence. For instance, like the Claude ones just refused. Claude was just like, “Oh, I can’t do this, sorry, no.” And it gave no evidence, right? So their 1ACs didn’t have any evidence at all. So you can see where they’re failing there, right? Um, you know, Gemini did generate a paragraph. I want to say ChatGPT would generate a sentence, um, would quote a sentence.
Um, but the other thing that we did to kind of test these systems, because the reliability of the evidence is really important, and a lot of people might not realize this, but in in competitive debate world, if you present something as a quotation from an author and we can’t verify that that’s actually a quotation from an author, that is a fundamental ethical violation of the community, and you automatically lose that round. And if you’re on my debate team and you did something like that, I’m kicking you off my debate team, right? That the the ethics, you know, the ethics around evidence in this community are extremely strict. So the the next test that we did then was we wanted to test, you know, were people, was our system making up cards or were the other systems making up cards? Were they hallucinating cards or falsifying cards? And so the other test that we ran was a fidelity test where we set the expectation that 100% of the cards in the 1AC had to be verifiable. That and what our test did, what our assessment system did is it would took a random line in a card and then it would look at the citation information and then go to the web, find that card, and verify that that that line of text existed in that evidence that that you you had cited. Um, the LLM, the frontier LLM got a zero score on that. Wow. Not a single 1AC that was generated by those LLMs had 100% verifiable evidence in it. Right now, and like I said, like the Claude ones just didn’t produce any evidence because it was like, “Oh, I’m not capable of this. I’m not even going to try.” Um, the our system was like, I want to say like 79% of our 1ACs, 100% of the cards were verifiable. So probably like what what is that math-wise? Like maybe two three of our 22 1As did not meet that metric. The really shocking thing was only about 8% of the human-generated 1A could pass that test.
46:46 - Dr. Anand Rao: Right now, having heard a lot of these 1As, I’m not that surprised to be honest. But but that is that’s an interesting point of comparison because we’re always worried about or or critics are concerned about, “Oh, AIs will hallucinate,” but we don’t really have a discussion about how often humans hallucinate or, you know, I don’t think it’s it’s malicious. I don’t think they’re making it up. I think there are just mistakes made when they’re keeping track of it, when they’re copying it.
47:08 - John Hines: Yeah. Yeah. No, that’s absolutely and that’s what I was going to say. That’s my takeaway is not that, you know, these camps produced a lot of fake evidence. Um, I think what happened is these camps did not have sufficiently rigorous protocols to verify that the citation information that’s being put in that evidence is accurate. Uh, and so and and I’ve taught at debate camps. I’m sure y’all have as well. Like, yeah, I could totally see that happening. Um, that, you know, a bunch of, you know, a bunch of citation failure happened, right? That the they’re not citing.
47:30 Stefan Bauschard: I’m sure it happens in human human undergraduate research papers too, right?
47:38 - Dr. Anand Rao: Absolutely. Absolutely. A lot of this like outrage about, you know, hallucination and like, “Oh, it it gave me a paper with 100 sources and one of the sources was wrong.” I wonder, maybe the the next step is we should start incorporating this into debate tournaments and give a special award to whoever has the the the the greatest level of fidelity evidence in their 1AC, right?
48:03 - John Hines: Yeah. Submit that ahead of time and have it run it.
48:08 - Stefan Bauschard: Yeah. I heard about a debate coach, you know, and you know, some of the stuff in the materials I share with students, you know, one of the sites was wrong and this person was really upset. They kind of said, “Well, you know, I don’t, you know, I found the evidence, but the citation was wrong.” And it’s like, there’s a coach in our area. He coaches, he coaches, “Gotcha, go after their sites.” And his strategy is to try to, you know, just double down on everybody’s site to see if they made a mistake and try to win the debate, which, you know, to me is a pretty pretty silly way to debate. But, you know, it’s it’s uh, I’ll just more like human error, right? Like, yeah, I didn’t make up the evidence, so I just put the wrong citation on it, right? And even the person was like, “Well, yeah, the citation is just wrong. You didn’t make it up.”
48:48 - John Hines: Yeah. Yeah. I mean, there’s deeper discussions, right? And I think it is a little hard for, well it’s not a little hard, it’s a lot hard for LLMs, total context, like the straw person argument, right? You know, they’re pulling that idea out of there. AI is really good, uh, AI debate systems are, the fidelity of the tag to the card is 100%, yeah, yeah, yeah, right? Like because they are not going to lie and manipulate right, the way humans will, right? Like when human researchers, especially like high school kids, they have a debate case in mind that they want to exist.
49:20 - Stefan Bauschard: Yeah, they kind of force the card in there, right?
49:21 - John Hines: Yeah, they’ll force the card into the argument. The LLMs will not do that, right? Like if the LLM has tagged it as something, that card’s going to say that. And that’s the other kind of like comparison, like are the LLMs worse than the humans? Well, the humans are going to manipulate the evidence to say what they want it to say, whereas the LLM is not going to do that. Yeah. So that’s really interesting.
49:48 -Dr. Anand Rao: Yeah. We only have a few minutes left, so I want to talk a little bit about what’s next. You know, what what what do you think could be the next step for DebaterHub, for similar systems to take these lessons? I mean, it’s really fascinating to look at the comparison of, you know, these are pretty elite debate programs and summer programs, and and it’s not like you would expect sloppy work. It’s not sloppy. It’s just that’s representative of human error. So, um, how do we then, maybe that’s part of this or or what do you think DebaterHub can do to help address some of this or move forward with debate online and debate with AI?
50:18 - John Hines: Yeah, I mean, there’s lots of directions we can go. Um, I think one of the problems we’ve had at DebaterHub is is we like to chase, we’ve we’ve had a tendency to chase all the problems at once. Um, we’ve we’ve made a a specific decision this time going forward that we’re going to chase like a narrow problem. So we’re looking for like one debate league, and and I think we found them, um, but we’re not announcing anything yet, but we’re looking for one debate league where we’re going to build a a bespoke interface for them, and we’re going to collaborate specifically with them on on what their needs are. Um, and the idea is to generate um some coachlike avatars that the humans can interact with, right? Um, because one of the most valuable learning tools in learning how to debate is somebody that you can talk to and ask questions, right? Because debate is a talking and listening game. And you only get better by talking and listening with somebody, but not there’s not enough people to talk and listen to you. And so step one is use these innovations that we know we have a reliable debate system that can perform at expert human level. So now the question is how do we create a user experience that allows Urban Debate kids to get quick access to a coach 24/7, right? That they can just ask questions, and we can know confidently that the answers the the AI coach is giving are solid, right? Doesn’t have to be perfect, doesn’t have to be the best debate coach ever. Every debate coach makes errors.
51:44 - Speaker 2: So you’re really thinking about implementing this with the preparation for debates, not necessarily in a debate.
51:49 - John Hines: Yeah, I think that’s a different question, right? And and I think that that’s more controversial. There’s certainly a lot of AI anti-AI pushback going on in education broadly and also in high school competitive debate right now. In fact, there were arguments read last weekend at a tournament that said if you don’t identify the amount of AI you use during your prep, you should lose. Um, AI spec for the debate people watching has has made an appearance. Um, and and so I think the most valuable uses of AI and the most ethical uses of AI in debate are pre- and post-round.
I think that we can use our tools to help people learn what debate is and think about debate. And we now have tools that can help people write a 1A. Um, the next step in our research is going to be taking it further in the debate, like, okay, what does a 1NC look like, etc. Our systems aren’t there yet. We’re confident in our 1ACs. Uh, but we can certainly help people generate 1ACs and think about 1ACs. Um, and the question is what’s the ethical way to do that? But that is pre-round prep. Um, and then I think another thing that we haven’t implemented yet, but I would like to do eventually is you can have, I think you could just use notetakers for this. Uh, I think it’d be smart for debaters to have AI notetakers record all their post-round discussions with their judges. And then take those notetakers and and and aggregate the results that you’re getting from your your judges to look for repeated feedback that you’re getting over and over. Um, I think that’s a low-tech tool, right? Like you your debaters should be implementing this immediately. They don’t need us for this. Like go get one of these AI notetaker systems and have it listen to your post-rounds and then look for what’s the consistent feedback you’re getting in all these post-round conversations, would be something that debaters can do today.
Um, I certainly wouldn’t encourage them to use any of these systems to write a 1A. But um, I know NSDA apparently just gave everybody access to Perplexity, which is really cool. Um, that’s something we didn’t benchmark against. So I actually kind of want to go run another benchmark and see if Perplexity can write 1ACs under the same constraints that ours can. Um, but yeah, I don’t think we’re there using AI in debate rounds yet. But I do I would advocate for using AI in debates and classrooms. Um, and perhaps in the future, you know, we could have AI debate leagues, right? Um, and that could be part of it as well. I just don’t think people are ready for that.
54:06 - Stefan Bauschard: And say in general, you know, people aren’t ready, you know, talking with some of my own debaters, you know, and I was like, “Well, well, you know, the, you know, kids get coaching like before the round, right?” Sometimes from live people, sometimes when they’re teammates or an external coach will jump on, right, and you know, give them give them some advice and some arguments to make. I’m like, you should you should have the notetaker on, and then as soon as you’re done before the debate starts, right, because that’s a little more controversial, just have that whole transcript right there and you you have it for during during, you know, for the for the round, you know, in the round. I I get it’s more controversial. It’s harder to it’s harder to control regular, right? Like we don’t have any idea. But that’s a discussion for another day. But I think there’s a lot of ways to kind of enhance enhance what you’re doing, right? Just enhance your own arguments, enhance your own abilities.
54:58 - John Hines: I mean, if we’re going to have AI spec, we should have coach spec. Well, you know, how much of how much of your argument should have to help me write my arguments, right? How much does your school spend on hiring college debaters to write arguments for you? Yeah. Yeah. Right. Right. Expect that. That’s the response. Yeah. Yeah. That’s what it’s like, AI, it’s just like another, you know, like I I’ve always viewed this, I always thought debate was a perfect model even when this first came out, because to me, it’s like it’s another intelligence you collaborate with, just like you collaborate with your teammates who who may know about more than you, your coaches. There’s obviously all kind of inequality amongst not just resources but in terms of other coaching and students people have access to in the standard equation.
So, in looking at costs in this tournament that we ran, it cost us a dollar per 1A. And that’s without doing like cost efficiency measures, right? It cost us a dollar to write an expert-level 1AC. Whereas you go to a debate camp and spend $1,000 a week to get one 1A, and you’re at least three weeks in. So you’re spending $3,000.
55:51 - Dr. Anand Rao: You’re learning something while you’re getting, you’re learning to learn.
55:55 - Stefan Bauschard: Yeah. If someone, I’ll write I’ll write a 1A for 1,500 bucks, you know?
56:01 - John Hines: No, you get you get more than that, right? But I mean, I’ll be honest. I’ve I’ve worked in programs where we hired college debaters to write things for us. You couldn’t buy a 1AC for less than $300 to $500, right? You know, out on the market, if you wanted to pay somebody to write you a 1A, you’re going to pay $300 to $500 for that. And our system did it for a dollar. Um, is it going to win you the TOC? Probably not. But it’s a 1AC that you can have and take into debates and and and learn.
56:28 - Stefan Bauschard: Well, there’s a lot more here to explore, but John, thanks so much for for sharing some of the work that you’ve done with DebaterHub. Um, fascinating study with the uh, the simulated tournament in comparison of the 1ACs, and I think there’s a lot there that I I hope you you publish this because it really, I’ve been really intrigued recently about the comparison of hallucinations with human error rates and the way we contextualize that discussion. And I think there’s a lot in your study that that helps illustrate that.
56:52 - John Hines: Definitely. Yeah. I think we want to do the human annotation first. Um, but then we’re we’re definitely going to submit it to one of the top machine learning conferences is the plan.
57:04 - Dr. Anand Rao: That’s excellent. Well, thanks so much for joining us. We appreciate it.
57:04 - John Hines: Thank you so much, guys. Have a great rest of your day.

