AI in English Literature: Better Essays or Better Thinkers?
From ancient stories to artificial intelligence, the question is what our tools enable us to become.
TL;DR
The opportunity: AI can widen access to explanations, feedback and the practice of literary judgement.
The risk: A convincing answer can replace the reading and reasoning that students need to learn.
The standard: Measure what students can understand independently, and the judgement they exercise when assistance is available.
Before literature was a school subject
Long before literature became an examination subject, stories gave shape to origins, grief, duty and belonging. What we call literature often belonged to religion, performance and shared memory rather than a separate world of books.
In southern Mesopotamia, writing developed in the late fourth millennium BCE around administration. Storytelling has an earlier history that surviving inscriptions cannot date. The Epic of Gilgamesh later confronted mortality; Egyptian traditions encompassed tales, teachings and hymns. Language already did more than record practical information.
Greek epic drew on oral composition, while Athenian theatre brought conflicts over power and duty before audiences. Aristotle associated representation with learning: “to learn gives the liveliest pleasure”. This was a philosophical account, not an experiment.
Vedic chanting and Māori waiata, whakapapa and pūrākau also show disciplined oral transmission. These remain living traditions. Across this history, preserving words and understanding their significance are related achievements, but neither guarantees the other.
Every new medium changes the bargain
The worry that a tool might preserve knowledge while weakening the knower is ancient too. In Plato’s Phaedrus, Socrates recounts a myth in which King Thamus warns that writing can give learners “the appearance of wisdom, not true wisdom”. The concern is about mistaking access to recorded words for understanding. It is not evidence that writing harmed every reader, nor a prediction of ChatGPT. The warning survives because Plato wrote it down: the medium under suspicion also made the argument available to later generations.
Manuscripts and printing subsequently changed who could encounter a text and how. Medieval books of hours supported private devotion through words and images. In early eleventh-century Japan, Murasaki Shikibu’s The Tale of Genji explored the intricacies of court life; later printed editions helped extend its circulation beyond the narrow world of manuscript ownership. Literature’s forms and audiences developed in different ways across societies.
Printing was not a single European invention. China’s dated Diamond Sutra was printed in 868; Korea’s Jikji, from 1377, is the oldest surviving book printed with movable metal type. Gutenberg’s Bible followed in Europe around 1455. Greater reproducibility altered the possibilities for circulation, although possessing a book and being able to read it remained different things.
Later movements also altered what literature asked of a reader. Romanticism placed particular weight on imagination and individual feeling. Modern writers such as Virginia Woolf questioned whether inherited ways of constructing a novel could adequately represent human character. These shifts were arguments about what deserved attention and how language might represent it. They were not a steady march towards one superior form of literature.
Digital text brought another change before conversational AI. Project Gutenberg began in 1971; its founder Michael Hart described a purpose built around making texts easy to store, retrieve, reproduce and search. This made access a different problem from owning a physical copy. Reading on a screen could still mean encountering an existing author’s words. Generative AI adds the possibility of asking a system to produce the explanation, interpretation or new text itself. That is a further change in the reader’s relationship with language, not merely another way to deliver the same book.
The lesson for education is therefore more demanding than “new technology is good” or “new technology is dangerous”. A medium expands certain possibilities. Institutions, teachers and readers decide which possibilities become habits.
The shortcut existed before AI
Study guides, model paragraphs and rehearsed essays already allowed students to reproduce conclusions without understanding how to reach them. Memorisation itself is useful: quotations and literary knowledge provide material for thinking. The difficulty begins when remembered conclusions replace the connection between evidence and judgement. Willingham’s account of critical thinking stresses its dependence on subject knowledge.
Wizardry’s founder reports seeing students generate English preparation with AI and memorise it for exams. That observation motivates this inquiry; it is not a prevalence estimate. The Auckland student survey supplies context about uncertainty over improvement, but did not establish students’ motives for AI use.
The educational response must address that uncertainty. If students cannot see what useful practice involves, a finished answer offers an understandable escape. Independence needs to be taught through specific tasks, guidance and opportunities to apply what has been learned.
What AI changes
Generative AI changes the relationship between a reader and an answer. A study guide offers an existing interpretation; a chatbot can produce one for this passage, this question and this student, then defend or rewrite it. Access is now interactive. The system can act as an apparent explainer, editor or opposing reader, often within the same conversation.
That responsiveness has substantial educational promise. A student can ask about an unfamiliar phrase without interrupting the class, compare two possible arguments or receive feedback while a difficulty is still fresh. The same system can complete those activities for the student. The difference depends partly on when help arrives and what the learner must do with it. A novice might need an explanation before attempting a task; an experienced reader might need a challenge after forming a view. Personalisation is valuable when it responds to the learner’s development, rather than merely their request for less effort.
Why the technology matters
The technology helps explain why both uses are possible. Transformers supplied an influential architecture; later research demonstrated learning from prompted examples and training systems to follow instructions with human feedback. Those developments made generated language more useful and responsive. They do not make every quotation reliable or every argument warranted.
Linguists Bender and Koller argue that linguistic form alone is insufficient for grounded meaning. Their 2020 position paper does not settle the capabilities of later multimodal systems. It does sharpen an educational distinction: a system’s ability to construct an interpretation cannot establish that its user has encountered the text, understood the reasoning or taken responsibility for the claim. Those are separate questions about the learner.
Fluency, confidence and agreement
The psychological risk also predates AI. Fisher, Goddu and Keil found across nine experiments that searching online could inflate people’s estimates of their own explanatory knowledge, including on unrelated topics. This was internet-search research, not a chatbot trial. It suggests a mechanism worth testing: access to an articulate explanation may blur the boundary between what the tool supplies and what the person can explain.
A conversation can make that boundary more difficult to notice. Sharma and colleagues found that five tested AI assistants sometimes favoured agreement with users’ beliefs over correct responses, a behaviour called sycophancy. This does not describe every model or interaction. It gives reason to examine what happens when a student asks for confirmation of a weak interpretation: a supportive tone can be mistaken for an independent evaluation.
Facts are only part of the judgement
Reliability also extends beyond factual accuracy. In a preregistered experiment involving 1,912 participants, Shu and colleagues found that AI-generated accounts of two historical events could shift opinions through framing, even when the accounts were factually accurate. The study does not establish a universal political bias. It illustrates why checking facts is necessary but incomplete: selection, emphasis and omitted perspectives can shape an apparently neutral explanation. Literary education already asks readers to attend to precisely those choices.
These questions now concern everyday practice. Pew’s 2026 report found that 54% of surveyed US teenagers had used chatbots for schoolwork; HEPI found that 94% of UK undergraduate respondents used generative AI to help with assessed work. Different questions and populations prevent a direct comparison, and neither figure measures cheating.
In Noy and Zhang’s experiment with 453 professionals, ChatGPT reduced writing-task time by 40% and raised assessed quality by 18%. Those are productivity results. Education adds another objective: developing the person who will face the next task. AI makes convincing language cheaper and easier to obtain; it therefore makes the distinction between producing an answer and becoming able to judge it more consequential.
A better answer is not the same as better learning
Soderstrom and Bjork distinguish performance during practice from durable learning, demonstrated through retention and transfer. A stronger paragraph can conceal that difference when the assistance producing it disappears.
Retrieval experiments with prose passages and a classroom meta-analysis covering 222 studies and 48,478 students support reconstructing knowledge and revisiting it later. Self-explanation research supports asking learners to explain relationships themselves. For English, the proposed application is to connect a textual choice with a defensible interpretation, then apply that reasoning to an unfamiliar passage. This application requires evaluation; the studies do not validate a particular literary-analysis platform.
External support is still part of thinking. Risko and Gilbert describe cognitive offloading as a normal strategy, while Sinha and Kapur’s review concerns initial problem-solving followed by instruction. Together, they discourage treating either convenience or difficulty as inherently educational. Looking up a word can enable close reading; receiving the whole interpretation can remove the intended practice. An initial attempt needs subsequent guidance when the learner cannot progress.
The experiments resist a simple verdict
In Turkish high-school mathematics, Bastani and colleagues found that a general GPT interface improved assisted practice performance by 48%, but its users performed 17% worse than controls after assistance was removed. A tutor with learning safeguards avoided that measured harm without establishing a significant later unassisted gain. These relative differences are mathematics results, not predictions for English.
Positive trials deserve equal attention. A crossover study of 194 Harvard physics students found higher immediate learning gains with a purpose-built AI tutor, without establishing long-term retention. A six-week Nigerian programme produced an estimated English gain of 0.24 standard deviations, but combined AI, teachers and additional learning time. The gain cannot be attributed to the chatbot alone.
Tutor CoPilot offers another direction: AI suggested teaching moves to human tutors. Its field experiment reported a four-percentage-point increase in student topic mastery, alongside practical problems such as unsuitable grade-level suggestions.
The evidence supports neither automatic rejection nor automatic adoption. Designs, subjects and outcomes differ. Schools should ask which students learned what, compared with which alternative, and whether the gain survived beyond the supported task.
What “AI slop” misses about literature
“AI slop” describes polished language that makes interchangeable or poorly supported claims. It is an informal judgement about quality, not a test of authorship. Humans produce it too, and AI can help someone improve it. A sentence saying that a poem reveals the profound complexity of the human condition gives us little to evaluate until it identifies the words, the complexity and the reasoning.
Three difficulties should be distinguished. An answer can be vague enough to fit almost any text. It can be specific but wrong, inventing evidence or ignoring a contradiction. Or it can be a strong interpretation that the student does not understand. The third answer is not necessarily slop. It is an educational problem because the quality belongs to the submitted object while the intended learning has not yet happened. Treating all AI-assisted work as equally empty misses that distinction.
Different readings, accountable reasons
Literature complicates the search for one definitive answer. Philosophical work on interpretation examines relationships among a work, its historical setting, an author and a reader. Understanding often involves revisiting parts in light of the whole, then reconsidering the whole after noticing a detail. This is an account of interpretive practice, not an experimental proof that every reading is legitimate.
Pause before the comparison: choose an interpretation, then identify a detail that might complicate it. What would persuade you to revise your reading?
Explore two defensible readings of Frost
Frost’s The Road Not Taken provides a useful worked example. One reading explores choosing under uncertainty: the speaker cannot experience both futures. Another examines the stories people later tell about their choices. The paths are initially described as “about the same”, yet the speaker anticipates presenting the choice as distinctive. The future orientation of “I shall” helps the second reading account for that tension. A simple celebration of individualism must explain it too. The final sigh might suggest regret, satisfaction or irony; assigning it one emotion requires an argument. These readings ask different questions of the same evidence. Their difference does not make either automatically correct, and an interpretation that ignores the similarly worn paths explains less of the poem.
From that example, we can propose a practical standard for a valid interpretation. It should represent the text accurately, make a coherent connection between evidence and claim, account for details that complicate the argument, and use context responsibly. It should also distinguish what the passage establishes from what the reader infers. These are reasons a reader can inspect, rather than a demand that every reader arrive at the same conclusion.
Some readings will consequently be stronger than others. A surprising interpretation does not earn credibility simply by being original; a familiar interpretation is not invalid because others share it. We can compare explanatory reach: does the argument illuminate a pattern across the work, or depend on one convenient phrase? We can ask what would weaken it. We can also retain uncertainty when several readings account for the evidence without forcing a premature verdict. Disagreement becomes productive when it improves the account of the text.
AI could support this work by generating a counterreading or exposing an overlooked detail. It could also narrow it by presenting the first plausible answer as the answer to memorise. The decisive question is whether the student can explain why one suggestion deserves acceptance and another does not. A paraphrase of an AI paragraph is not sufficient evidence of that judgement; neither is simply avoiding AI.
Culture, variety and voice
There is also a cultural dimension to apparently neutral quality. Agarwal, Naaman and Vashistha studied 118 participants in India and the US and found that AI suggestions shifted Indian participants’ writing towards Western styles and weakened some cultural nuances. This was a bounded writing experiment, not proof that every tool erases every culture. It nevertheless challenges the assumption that smoother English is always a culturally innocent improvement. In literature, a distinctive register or way of organising experience may be part of the meaning that correction removes.
Doshi and Hauser found another tension: AI ideas improved individual short-story ratings while making the collection more similar. Stronger outputs and less variety occurred together. Shared interpretations are not inherently bad, but a classroom loses something if alternative questions disappear before students encounter them. Voice means more than unusual vocabulary. It includes what a writer notices, what they consider worth defending and how they respond to another person’s objection.
What the act of reading contributes
Reading is also an activity, not merely a container of conclusions. Tamir and colleagues’ brain-imaging study linked different kinds of fictional passages with distinct patterns of activity in networks associated with simulation. That does not prove lasting empathy gains. A longitudinal study of 236 early adolescents found a concurrent association between fiction reading and understanding other minds at age 13, but reading at 11 did not predict that later outcome. The benefits should not be romanticised: encountering perspectives offers an occasion for understanding, not an automatic moral transformation.
Forster’s The Machine Stops imagines a culture admiring knowledge detached from direct encounter: “Let your ideas be second-hand”. WALL-E supplies a familiar image of comfortable dependence. These are cultural prompts, not research evidence. They bring us back to a literary question: when explanations become abundant, who still undertakes the work of attending, questioning and choosing?
What would make this a good revolution?
A good revolution would allow more students to enter the conversation literature makes possible. A hesitant reader could receive an explanation that opens a difficult passage; a teacher could identify a missing connection sooner; a student with an unconventional interpretation could test it against an attentive challenge. The ambition is a wider distribution of the ability to make and examine meaning. More generated text is useful only insofar as it serves that ambition.
Agency and capability
This requires a careful definition of agency. Ryan and Deci’s account of motivation emphasises autonomy, competence and relatedness. Giving a student choices while removing the intellectual activity may preserve the appearance of autonomy without building competence. Conversely, withholding help can leave a learner excluded. An appropriate use of AI would support decisions the student increasingly understands, within relationships where those decisions can be discussed. That is an instructional implication of the theory, not a tested claim about Wizardry.
Research on human and AI collaboration resists easy optimism. Vaccaro and colleagues’ meta-analysis of 106 experiments found that combinations improved on humans alone on average, but underperformed the better of humans or AI alone; outcomes differed between creation and decision tasks. The experiments covered varied systems, not just current chatbots. Education has an additional purpose beyond maximising an immediate result: developing human capability. A productive partnership must therefore be judged by both the task and what the person learns through it.
The value of thoughtful friction
Good design may sometimes feel less convenient. Buçinca and colleagues’ experiment with 199 participants found that interventions requiring more thought reduced overreliance in AI-assisted decisions, but the more effective designs received poorer subjective ratings. Benefits also varied with participants’ inclination towards effortful thinking. This was a decision task, not an English classroom. It suggests that popularity and smoothness are incomplete educational measures, while compulsory friction without suitable support may benefit students unevenly.
Nor must conversational AI merely reinforce existing beliefs. Costello, Pennycook and Rand found that evidence-based AI dialogues reduced conspiracy belief by about 20% among 2,190 participants, with effects persisting for two months. That does not demonstrate improved literary interpretation, but shows a constructive possibility: a responsive conversation can help someone reconsider evidence.
Persuasiveness itself is insufficient. Salvi and colleagues’ 900-person debate experiment found that GPT-4 with participant information could outperform a human comparison. A September 2026 correction clarified that the direct advantage over non-personalised GPT-4 was not statistically established. Neither result shows that persuaded participants became better reasoners. The educational test should concern the reasons a learner can examine afterwards, not simply whether the system changes their mind.
How a bad revolution could take hold
A bad revolution could emerge through ordinary institutional incentives. If success means more submissions, faster marking and higher satisfaction, a system can reward finished answers while making dependence difficult to see. A student receives a generated interpretation; another system approves its familiar structure; the school records completion. This is a possible failure pattern, not a measured description of all AI classrooms. Its danger is that improvement in the paperwork can conceal a narrowing of the student’s participation.
- Answers arrive before the student encounters the text.
- Familiar wording is rewarded without examining the reasoning.
- Completion and convenience stand in for learning.
- Support helps the student enter and question the text.
- Evidence and counterreadings guide the student’s decisions.
- New tasks reveal what the student can now understand.
The same question applies to access. Cheap assistance could reduce a real barrier. But a future where some students receive expert teaching and others receive automated substitution would distribute understanding unequally. UNESCO’s technology-in-education report highlights tensions between individualisation and social learning, and between inclusion and new exclusions. Departments should therefore examine who receives useful support, whose language is treated as deficient, and whether technology creates more opportunity for discussion with teachers and peers.
Fair assessment and a human standard
Wizardry’s planned school AI checker belongs within this evaluation. Its contribution would depend on validation and fair human review; a flag cannot establish misconduct. Research documents detector errors affecting non-native English writing, while newer evaluations show substantial differences between tools. A good revolution would protect students’ opportunity to explain their work alongside schools’ need for authentic assessment.
The revolution becomes good when access to language expands participation in judgement. It becomes bad when access to language allows institutions to mistake a completed answer for a developed person. Schools can test that distinction through unfamiliar passages, delayed assessment, student explanations and the ability to reject plausible AI errors. Independence and intelligent assistance both matter.
Morrison’s Nobel lecture offers a demanding human standard: “But we do language. That may be the measure of our lives.” She was not discussing AI. Her words nevertheless remind us that language is something people undertake and answer for, rather than simply obtain.
Designing AI use around the student
The practical sequence can remain simple. Read a manageable passage, make an observation and attempt a claim. Use AI to address a specific gap through a hint, question or counterreading. Check the evidence, decide what to accept and explain why. Then attempt another passage with less support and revisit the skill later. A beginner may need a worked example first; this is a proposed application of the evidence, not a validated intervention.
Wizardry’s authored approach connects technique, effect and literary element with ideas, themes and message. AI feedback could help a student notice a missing connection and practise explaining it. The framework should organise enquiry without becoming a formula that preselects the interpretation. External studies give reasons to investigate this design, rather than evidence of Wizardry’s own effectiveness.
Departments also need task-specific boundaries. NZQA English guidance distinguishes learning support from assessed writing, while Cambridge and IB guidance require authentic work and acknowledgement where assistance is permitted. Establish those boundaries before practice begins.
The teacher’s final question should reach beyond the paragraph: can the student now notice, explain and defend something they could not before? That is where a promising tool becomes a meaningful educational change.
Sources and further reading
This research review draws on historical scholarship, philosophical arguments, experiments, surveys and policy guidance. Their purposes and evidential limits differ. The references below include further reading; expand a category to explore them. Sources were reviewed on 30 September 2026.
History and literary perspectives (21)
- H01 Ira Spar / The Metropolitan Museum of Art (2004), The Origins of Writing
- H02 British Museum, The Flood Tablet, K.3375
- H03 UCL Digital Egypt, Egyptian Literature: Compositions of the Middle Kingdom · Sinuhe context
- H04 Harvard Center for Hellenic Studies, Homer Multitext overview
- H05 Colette Hemingway / The Met (2004), Theater in Ancient Greece
- H06 Aristotle, Poetics, Part IV
- H07 Plato, Phaedrus, 275
- H08 UNESCO, Tradition of Vedic chanting
- H09 Te Ara, Māori education: Traditional society · Whakapapa overview
- H10 Jacob Nadal / Library of Congress (2023), From Jikji to Gutenberg
- H11 Kathryn Calley Galitz / The Met (2004), Romanticism
- H12 The Met (2019), The Tale of Genji exhibition and printed edition record · printed Genji
- H13 Wendy Alpern Stein / The Met (2017), The Book of Hours: A Medieval Bestseller
- H14 Virginia Woolf (1924), Mr. Bennett and Mrs. Brown
- H15 Toni Morrison (1993), Nobel Lecture
- H16 E. M. Forster (1909), The Machine Stops, Chapter III
- H17 Pixar, WALL-E
- H18 Andrew R. George (2010), The Epic of Gilgamesh, in The Cambridge Companion to the Epic · SOAS research account
- H19 Michael Hart (1992), The History and Philosophy of Project Gutenberg
- H20 Theodore George (2020; revised 2025), Hermeneutics
- H21 Robert Frost (1915), The Road Not Taken
Learning and psychology (18)
- L01 Soderstrom and Bjork (2015), Learning Versus Performance: An Integrative Review · author-hosted full text
- L02 Roediger and Karpicke (2006), Test-Enhanced Learning · author-hosted paper
- L03 Dunlosky et al. (2013), Improving Students’ Learning With Effective Learning Techniques
- L04 Yang et al. (2021), Testing (Quizzing) Boosts Classroom Learning
- L05 Rozenblit and Keil (2002), The Misunderstood Limits of Folk Science
- L06 Risko and Gilbert (2016), Cognitive Offloading
- L07 Bisra et al. (2018), Inducing Self-Explanation: A Meta-Analysis
- L08 Sinha and Kapur (2021), When Problem Solving Followed by Instruction Works
- L09 Hattie and Timperley (2007), The Power of Feedback
- L10 Graham and Perin (2007), Writing Next
- L11 Ryan and Deci (2000), Self-Determination Theory and the Facilitation of Intrinsic Motivation
- L12 Dodell-Feder and Tamir (2018), Fiction Reading Has a Small Positive Impact on Social Cognition
- L13 Panero et al. (2016), Does Reading a Single Passage of Literary Fiction Really Improve Theory of Mind? · authors’ 2017 reply
- L14 Daniel T. Willingham (2007), Critical Thinking: Why Is It So Hard to Teach?
- L15 Fisher, Goddu and Keil (2015), Searching for Explanations: How the Internet Inflates Estimates of Internal Knowledge · Yale’s substantive account
- L16 Delgado, Vargas, Ackerman and Salmerón (2018), Don’t Throw Away Your Printed Books · author-institution record
- L17 Tamir, Bricker, Dodell-Feder and Mitchell (2016), Reading Fiction and Reading Minds: The Role of Simulation in the Default Network · author-hosted paper
- L18 van der Kleij, Apperly, Shapiro, Ricketts and Devine (2022), Reading Fiction and Reading Minds in Early Adolescence · institution-hosted paper
AI studies and reliability (23)
- A01 Vaswani et al. (2017), Attention Is All You Need
- A02 Brown et al. (2020), Language Models Are Few-Shot Learners
- A03 Ouyang et al. (2022), Training Language Models to Follow Instructions With Human Feedback
- A04 Noy and Zhang (2023), Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence · Science paper
- A05 Bastani et al. (2025), Generative AI Without Guardrails Can Harm Learning · correction
- A06 Kestin et al. (2025), AI Tutoring Outperforms In-Class Active Learning
- A07 De Simone et al. (2025), From Chalkboards to Chatbots: Transforming Learning in Nigeria, WPS11125 · reproducibility package
- A08 Wang et al. (2024), Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise · Stanford manuscript
- A09 Doshi and Hauser (2024), Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content
- A10 Lee et al. (2025), The Impact of Generative AI on Critical Thinking
- A11 Kosmyna et al. (2025, preprint), Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing
- A12 Liang et al. (2023), GPT Detectors Are Biased Against Non-Native English Writers
- A13 Weber-Wulff et al. (2023), Testing of Detection Tools for AI-Generated Text
- A14 Van Vlasselaer, Van Droogenbroeck and Spruyt (2026), Who Wrote This?
- A15 Buçinca, Malaya and Gajos (2021), To Trust or to Think · author-hosted paper
- A16 Sharma et al. (2023; revised 2025), Towards Understanding Sycophancy in Language Models
- A17 Bender and Koller (2020), Climbing Towards NLU: On Meaning, Form, and Understanding in the Age of Data
- A18 Agarwal, Naaman and Vashistha (2025), AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
- A19 Tao, Viberg, Baker and Kizilcec (2024), Cultural Bias and Cultural Alignment of Large Language Models
- A20 Salvi, Horta Ribeiro, Gallotti and West (2025), On the Conversational Persuasiveness of GPT-4; correction (2026) · 3 September 2026 correction
- A21 Costello, Pennycook and Rand (2024), Durably Reducing Conspiracy Beliefs Through Dialogues With AI
- A22 Shu, Karell, Okura and Davidson (2026), How Latent and Prompting Biases in AI-Generated Historical Narratives Influence Opinions
- A23 Vaccaro, Almaatouq and Malone (2024), When Combinations of Humans and AI Are Useful
Education, adoption and assessment (10)
- P01 Stephenson and Armstrong / HEPI (2026), Student Generative AI Survey 2026, Report 199 · original PDF
- P02 Pew Research Center (2026), How Teens Use and View AI
- P03 OECD (2026), Digital Education Outlook 2026
- P04 Miao and Holmes / UNESCO (2023), Guidance for Generative AI in Education and Research
- P05 NZQA, AI guidance for schools
- P06 NZQA, English National Moderator’s Report
- P07 Cambridge International, Generative AI in coursework
- P08 International Baccalaureate, AI statement and learning/assessment guidance · AI guidance
- P09 Australian Department of Education, Australian Framework for Generative AI in Schools
- P10 UNESCO (2023), Global Education Monitoring Report: Technology in Education, a Tool on Whose Terms?
Local context: Wizardry’s Auckland student English survey. The survey supplies context, rather than evidence of students’ motives for using AI or proof of product effectiveness.