THANK YOU FOR SUBSCRIBING
Be first to read the latest tech news, Industry Leader's Insights, and CIO interviews of medium and large enterprises exclusively from Education Technology Insights
THANK YOU FOR SUBSCRIBING
A featured contribution from Leadership Perspectives: a curated forum reserved for leaders nominated by our subscribers and vetted by the Education Technology Insights Advisory Board.

Julio Zelaya, Academic Director – “Implementing Artificial Intelligence (AI) in the Classroom”


Julio Zelaya is Academic Director of the Implementing AI in the Classroom program at the University of Pennsylvania's Graduate School of Education, where he is also a doctoral candidate in education, and Senior Principal in the Innovation and AI practice at The RBL Group. He is the author of five bestselling books, including one on artificial intelligence and organizational learning, and speaks internationally in English and Spanish. He works with universities and companies across the Americas on how leaders and institutions actually learn.
Education was built around the scarcity of good answers. That scarcity is ending, and our gradebooks have not noticed.
A student handed me a strategy recommendation last term. Clean, structured, confident, every heading in the right place.
I asked her one question. What would have to be true for this to be wrong?
She read her own page as if seeing it for the first time. She had produced the document. She had made no decision inside it.
I work in two rooms. One is academic, at Penn, where we study how people learn. The other is operational, where leaders decide by Friday and live with it. Both are full of machine-generated answers. Only one still grades the answer instead of the decision.
So the question I cannot put down: why did a student capable of producing a good answer have nothing to say about it?
The Answer Got Cheap
Stanford's AI Index tracked the cost of running a model at GPT-3.5 level performance. Twenty dollars per million tokens in November 2022. Seven cents by October 2024. A 280-fold drop in under two years.
By this year, the collapse had reached the classroom. Stanford's 2026 index reports that four in five U.S. high school and college students use AI for schoolwork, while only 6 percent of teachers say their school's policy is clear.
Access is the cheapest thing in the building now.
The Crutch and the Ladder
Researchers at the University of Pennsylvania ran a randomized trial with nearly a thousand Turkish high school students learning mathematics. One group practiced with a plain chatbot. One used a tutor built with guardrails that withheld answers and pushed for reasoning. One had neither.
With the chatbot in hand, students improved 48 percent on practice problems. Then it was taken away for the exam, and they scored 17 percent below the students who never had it.
“Education was built around the scarcity of good answers. That scarcity is ending, and our gradebooks have not noticed.”
The tutored group improved 127 percent during practice and showed no measurable penalty on the exam.
Same technology, opposite outcome. The variable was design.
Polished Is Not the Same as Right
A field experiment with 758 consultants at Boston Consulting Group, published this year in Organization Science, found quality gains above 30 percent on tasks within the model's competence. On one task placed deliberately outside it, AI users were 19 percentage points less likely to be correct. Their wrong answers were better written.
That is the trap our graduates walk into. Failure now arrives well formatted.
Here is what surprised me in the research. The frameworks are already ahead of us. UNESCO's student framework leads with a human-centered mindset. The AI literacy framework published this year by the OECD and the European Union asks learners to evaluate whether AI outputs should be accepted, revised or rejected, and to decide whether to use AI at all. The vocabulary of judgment is endorsed and written down. The gradebook has not caught up.
The Confianza Test
Confianza is the Spanish word for trust, the earned kind, the trust you extend to work you can vouch for. Four questions decide whether you have earned it.
• Stakes: What does it cost if this is wrong and nobody catches it?
• Standing: Do I know enough to catch the error at all?
• Source: Can I trace the claim to something real?
• Signature: Am I willing to put my name on it?
Standing is the one thing people skip, and it carries the paradox. The easier it becomes to borrow intelligence, the more of your own you need to evaluate what you borrowed. A novice with a powerful model is just a novice who can be wrong faster.
What to Change on Monday
First, rewrite one assignment so the deliverable is the decision, not the document. Students submit the machine's output, their revision, and a paragraph defending what they rejected.
• Second, put a wrong answer on the table on purpose. Give everyone the same persuasive, flawed output and grade those who find the flaw.
• Third, make standing explicit. Before students use AI, have them write what they would need to know to evaluate the result.
• Fourth, add a signature line. Name attached, use disclosed, ownership stated.
For centuries, good answers were hard to reach, so we built institutions around reaching them. That era is closing quietly, one price cut at a time. What stays scarce is the person who can look at a good-sounding answer and decide whether to stand behind it.
Our students will graduate surrounded by more intelligence than any generation in history. What they will need from us is the moments we deliberately design, where they have to choose, defend and sign.
Designing those moments well: eso sí es ganar.
I agree We use cookies on this website to enhance your user experience. By clicking any link on this page you are giving your consent for us to set cookies. More info

However, if you would like to share the information in this article, you may use the link below:
www.educationtechnologyinsights.com/leadership-perspective/julio-zelaya-nid-3952.html