AI can now draft PhD-level work, but learning is still the missing proof

AI can now draft PhD-level work, but learning is still the missing proof

University writing has crossed a line that many assessment systems were not built to see. The change is not that students can ask a chatbot for paragraphs. That was already visible when ChatGPT first entered classrooms. The deeper shift is that an AI system can now search, plan, read sources, compare claims, build an argument, cite passages and revise a report through several research steps before a student has opened a library database. OpenAI describes deep research in ChatGPT as an agentic capability for multi-step internet research on complex tasks, with updates that added broader access, higher query limits and later connections to external apps through the Model Context Protocol.

The rupture is not the essay, but the evidence trail

That matters because the university assignment has long served two jobs at once. It was a piece of work to be graded, and it was also a trail of evidence about the student’s thinking. When a student submitted a literature review, seminar essay, case study or policy memo, the lecturer inferred that the student had searched, selected, read, compared, planned, drafted and revised. Generative AI breaks that inference. Agentic research breaks it more completely, because it does not merely write fluently. It simulates the surrounding academic labour that used to leave traces in notes, bibliographies, false starts and drafts.

“PhD-level” is the phrase now doing the rounds because the newest research tools produce work that feels closer to the tone, structure and breadth of advanced scholarship than the early chatbot essays did. The phrase is tempting, but it needs care. A generated report may resemble a doctoral literature review in surface form: dense references, sober caveats, disciplinary vocabulary, a neat research gap and a confident argument. It is not the same as doctoral work. A PhD is not a long essay with sources. It is a trained act of original judgment under methodological, ethical and supervisory discipline.

The trouble for universities is that assessment often rewards the same surface forms that AI now handles well. A polished report with citations, formal prose and a balanced argument may receive a high mark even if the student has not built the knowledge that the mark is meant to certify. The OECD’s 2026 Digital Education Outlook puts the problem in sharp terms: using generative AI to perform tasks may raise performance without producing real learning gains, and unguided AI use risks cognitive offloading rather than durable understanding.

That is the central news value of this moment. AI has not only given students a shortcut. It has exposed a weak assumption inside higher education: the finished assignment was never a perfect measure of learning, but for decades it was treated as good enough. Deep research tools make “good enough” much harder to defend.

Deep research changed the unit of academic work

The first wave of AI-assisted cheating was easy to picture. A student typed a prompt, received an essay, edited it lightly and submitted it. Universities responded with warnings, detection tools, updated misconduct policies and classroom conversations about acceptable use. That response was imperfect, but it matched the threat as most people understood it: the machine was a writer.

Deep research changes the object. The machine is no longer only a writer. It behaves more like a junior research assistant with a browser, a citation habit and a plan. OpenAI’s system card says deep research is powered by a version of o3 optimized for web browsing and data analysis; it can search, interpret and analyze text, images and PDFs, pivot during a research task, read user-uploaded files and write or execute Python for data work. Google’s Gemini Deep Research was framed in similar terms when it launched: it creates a multi-step plan, browses and repeats searches, then generates a report with links to original sources and the option to export to Google Docs.

That workflow attacks the old assignment at its strongest point. A good university paper used to require many separate acts: formulating a question, finding credible sources, avoiding weak sources, extracting useful claims, spotting contradictions, planning the structure and building a written answer. AI systems now bundle those acts into one interface. The unit of work has moved from “write this essay” to “conduct this research task.”

The difference is visible in the kind of output a student can produce. Earlier AI writing often had a shallow centre: clean prose, vague claims, few specific references, weak awareness of current debates. Agentic research systems improve the surrounding scaffolding. They can retrieve a new policy document, compare it with a journal article, find a recent news report, identify a regulatory timeline, and place those materials into a coherent answer. The prose may still contain errors, but the submission looks more academically plausible because the tool has supplied context.

The benchmark results are relevant, not because benchmarks map neatly onto university learning, but because they show the direction of capability. OpenAI reported that the model behind deep research scored 26.6% on Humanity’s Last Exam, a benchmark of more than 3,000 expert-level questions across more than 100 subjects, and performed strongly on GAIA, which tests real-world questions requiring reasoning, web browsing, multimodal work and tool use. These are not essay marks. They are signs that systems are being trained to handle the messy, multi-step tasks that academic assessment often assumes only a student can perform.

The phrase “PhD-level” enters here because doctoral students often spend much of their early work doing exactly this kind of synthesis: mapping a field, reading across sources, finding a gap, locating a debate, and drafting an argument with references. An AI report can now imitate enough of that labour to unsettle markers. Yet imitation is not mastery. The machine can assemble a research-like product without inhabiting the research problem. It has no disciplinary apprenticeship, no obligation to defend a method, no long-term memory of failed approaches unless engineered into the workflow, and no personal accountability for the claim.

Universities therefore face a new kind of assignment problem. The old question was, “Did the student write this?” The better question is now, “Which parts of the intellectual process did the student genuinely perform, and how do we know?”

The old outsourcing model no longer explains the risk

Academic integrity policies were built around human misconduct. A student copied text, bought an essay, reused work, fabricated data, colluded with a classmate or used unauthorized help. Those acts still exist. AI has not replaced contract cheating; it has made a cheaper and more private version of it available inside ordinary study tools.

Yet the outsourcing model is too narrow. Students are not only asking AI to do the whole assignment. They use it to explain concepts, brainstorm questions, summarize papers, rewrite paragraphs, translate ideas, produce outlines, generate references, debug code, clean data and rehearse arguments. A 2025 Anthropic analysis of university student use of Claude found that students used it most often to create or improve educational content, including practice questions, essay editing and summaries, while a large share used it for technical explanations, coding, algorithms and mathematics.

That pattern matters because not every use is misconduct. A student who asks for a definition of heteroskedasticity, a critique of a paragraph, or a practice quiz is doing something closer to tutoring than cheating. A student who asks for a full literature review with citations and submits it as personal work is doing something else. The policy challenge sits between those poles. AI use is not a single behaviour. It is a chain of micro-decisions across the life of an assignment.

The old outsourcing frame also misses the emotional and practical reasons students use AI. Some are under pressure, working jobs, caring for family, writing in a second language or trying to catch up after weak preparation. Some use AI because everyone else seems to be using it, and a ban feels like a rule that only honest students will obey. Some use it because lecturers assign tasks that feel artificial: write 2,500 words on a broad topic with predictable arguments and little personal connection to the course.

Research on student attitudes supports this mixed picture. A large student survey published in the International Journal for Educational Integrity found that only 7% of respondents had not heard of generative AI, over half had used or considered using it for academic purposes, and students were much more supportive of tools that assist writing than of ChatGPT writing a whole essay. The same study reported that students with lower confidence in academic writing were more likely to use generative AI and that many wanted university-wide policy clarity.

This is why moral panic alone fails. It treats the student as the only problem. The deeper problem is a mismatch between available tools and assessment designs. If an assignment can be completed convincingly by a tool that performs research, writing and citation in one workflow, then the assignment may no longer test what the course claims to teach. AI has turned weak assessment design into a visible institutional risk.

That does not absolve students. It changes the burden on universities. Rules still matter. Honesty still matters. Disclosure still matters. But higher education cannot police its way out of a technical shift that has entered ordinary writing, search and productivity software. The assignment itself has to carry more of the proof.

A capable student can now work like a small research team

The strongest AI users are not always the weakest students. A weak student may paste a vague prompt and receive a generic essay. A strong student may use deep research as a force multiplier: ask sharper questions, compare source lists, test counterarguments, request methodological critiques, generate alternative structures and check whether the literature supports a claim. In that setting, AI does not replace academic work. It expands the speed and scale of the student’s first pass.

This is the uncomfortable part for universities. The same tool that enables low-effort cheating also rewards high-skill supervision. A student who understands a field can push the model harder, catch errors, ask for better sources and use AI output as raw material rather than a finished submission. A student who does not understand the field may accept plausible nonsense. The gap between the two students may grow, not shrink.

Agentic research makes the gap wider because prompting becomes a form of research management. A good user specifies the scope, defines source quality, asks for disagreements, separates evidence from speculation, checks dates, requests direct citations, and uses follow-up prompts to narrow the argument. A poor user asks for “an essay about climate policy” and trusts the answer. The tool is powerful in both cases, but the academic value differs sharply.

This creates a new skill layer inside university writing. Students now need to know how to interrogate AI output. They need to ask whether a cited source exists, whether a claim follows from the source, whether the model has confused a concept, whether the argument has ignored a major school of thought, and whether the prose conceals a weak method. AI literacy is becoming part of disciplinary literacy, not a separate digital skill.

UNESCO’s student AI competency framework points in that direction. It frames AI education around a human-centred mindset, ethics, AI techniques and applications, and AI system design, with progression from understanding to applying to creating. Its teacher framework also argues that the old teacher-student relationship is becoming a teacher-AI-student relationship, and it sets out competencies across ethics, pedagogy, AI foundations and professional learning.

The research-team analogy also has limits. A real research team has accountability. It has named contributors, ethical approvals, methods, data provenance and peer review. AI outputs often lack that chain. A deep research report may cite sources, but it does not carry responsibility. A student may turn that report into a submission, but the intellectual ownership is blurred unless the course requires disclosure and process evidence.

That is why the best university response is not to pretend students will stop using AI. They will not. The better response is to teach students to use it under constraints that preserve learning. A strong assignment might allow AI for source discovery but require annotated source verification, forbid AI-generated prose in the submitted draft, require a viva or oral defence, and ask the student to submit a reflection on which AI outputs were rejected. The student’s judgment becomes the assessable object.

PhD-level does not mean PhD-worthy

The claim that deep research can produce work at a PhD level needs a sharper definition. If “PhD-level” means a fluent report with many sources, a refined structure, a research gap, disciplinary language and a sober tone, then the claim is often credible. If it means original scholarship worthy of a doctorate, the claim collapses.

Doctoral work is not measured only by writing quality. It depends on original contribution, methodological competence, sustained engagement with a field, ethical research conduct, resilience under critique and the ability to defend choices before experts. AI can assist parts of this work. It can summarize literature, identify patterns, rewrite sections and suggest objections. It cannot replace the lived process of producing new knowledge under scrutiny.

This distinction matters because universities may misread the risk. The danger is not that AI has become a real PhD candidate. The danger is that many undergraduate and master’s assignments reward outputs that resemble academic maturity without requiring enough proof of the underlying cognition. If a course asks for a generic “critical essay” on a well-covered topic, AI can generate a passable or even strong-looking draft. If a course asks for a tightly evidenced analysis tied to seminar debate, local data, oral defence, staged drafts and personal methodological choices, AI has less room to hide the student.

Benchmarks also need humility. Humanity’s Last Exam and GAIA measure advanced reasoning and tool use, not doctoral authorship. OpenAI’s reported 26.6% score on Humanity’s Last Exam shows real capability against expert-level questions, but it also shows many failures. The system card itself treats deep research as powerful but still bounded, with safety evaluations covering cyber, biological and other risk areas rather than declaring the system academically equivalent to a researcher.

The real academic threat sits in the middle. AI output does not have to be dissertation-worthy to damage assessment. It only has to be good enough to receive credit in a course where the marker cannot see the process. Many university assignments are graded under time pressure, with large class sizes and limited feedback capacity. A polished, well-cited paper may move through that system smoothly.

There is also a danger for students who believe the “PhD-level” label too literally. A generated assignment may satisfy surface expectations while leaving the student unable to explain core ideas in class, use concepts in a new problem, or defend the sources chosen. The performance looks advanced, but the student’s transferable knowledge may remain thin. That gap becomes visible later: in exams, projects, placements, lab work, interviews or professional practice.

A university degree is a trust product. It tells employers, professional bodies and society that a person can do certain kinds of thinking and work. If AI allows students to submit advanced-looking assignments without gaining that capacity, the loss is not confined to one module. It weakens confidence in the credential itself.

The assignment has stopped being a proxy for the mind

For generations, the written assignment worked because it was costly to fake at scale. A student had to spend time in libraries, take notes, read enough material to sound credible, and produce a coherent answer. Contract cheating existed, but it required money, risk and coordination. Plagiarism existed, but it left textual evidence. AI lowers the friction so far that the finished paper can no longer stand alone as proof of learning.

The proxy relationship was always imperfect. Some students wrote beautifully without deep understanding. Others understood the material but struggled with academic prose. Some received heavy family or peer support. Some learned the rules of the essay game better than the subject itself. AI has not created every unfairness; it has made them harder to ignore.

A useful way to frame the change is this: the assignment is no longer enough evidence by itself. It may still be part of the evidence. It may still matter. But the marker needs more than the final text when the stakes are high. Draft history, source notes, short oral checks, in-class writing, project logs, lab notebooks, data files, peer discussion and reflective commentary now carry more weight.

This is not only about catching cheating. It is about protecting learning. The OECD notes that performing a task with generative AI does not automatically lead to learning, and it warns of metacognitive laziness and cognitive offloading when students rely on AI without pedagogical guidance. It also states that pedagogically grounded AI systems can support learning when they question, nudge and guide rather than simply deliver answers.

That difference should guide assessment reform. A tool that gives the answer too early can suppress struggle. A tool that asks the student to explain, compare, justify and revise can support learning. The same model may play both roles depending on design. This is why OpenAI’s Study Mode and Anthropic’s Learning Mode are telling product signals. OpenAI frames Study Mode as step-by-step guidance rather than quick answers, built around active participation, cognitive load management, metacognition and feedback. Anthropic’s Claude for Education similarly describes a learning mode that guides reasoning rather than simply providing answers.

Universities should read these features defensively. The major AI labs know that education cannot accept answer machines without conditions. They are moving toward tutor-like interfaces because the legitimacy of AI in education depends on visible learning, not only output quality. The finished assignment has become a weaker proxy; the interaction around the assignment has become stronger evidence.

That does not mean every essay must become a surveillance exercise. A university that records every keystroke may damage trust and privacy. The better path is proportionality. Low-stakes formative work can permit broad AI experimentation. High-stakes summative work should require human verification of process. The design principle is simple: the more a mark certifies individual competence, the more evidence of individual cognition the assessment should demand.

Search, synthesis and citation have collapsed into one interface

Academic writing used to be partly protected by fragmentation. Search happened in one place, reading in another, note-taking somewhere else, writing in a document editor, referencing in a citation manager, and feedback through a tutor or peer. Each stage required student effort. Each stage created chances for learning.

Now those stages are converging. A student can ask a deep research tool for an overview, approve a plan, receive a source-backed report, request an argument map, ask for gaps in the literature, turn the report into an outline, convert the outline into paragraphs, and ask for a citation style. Google’s Deep Research launch described the workflow as plan, browse, repeat searches, synthesize and export, supported by Gemini’s context window and Google’s search infrastructure. OpenAI describes deep research as able to browse, interpret files, use Python and cite sources for complex research tasks.

The collapse of stages is useful for professionals. Analysts, consultants, lawyers, marketers, journalists and scientists all want faster ways to map a field. The same capability inside a university course is more delicate because the process is the point. A student is meant to learn by moving through those stages, not only by receiving their combined output.

Citation is especially fragile. Many students and lecturers once treated citations as evidence that reading had occurred. AI weakens that signal. A report may cite a genuine source without the student ever reading it. It may quote accurately but select selectively. It may list credible sources while missing the most relevant debate. It may paraphrase a paper in a way that sounds accurate yet changes the meaning. A citation now proves that a text contains a reference. It does not prove that the student understands the source.

This forces a change in how source work is assessed. Instead of rewarding the presence of citations, lecturers need to ask for evidence of use. An annotated bibliography can ask students to state why each source matters, what claim it supports, what limitation it has and how it changed the argument. A literature matrix can require comparison across methods, samples, findings and theoretical assumptions. A short oral check can ask the student to explain the strongest source without notes.

AI also changes what counts as a good research question. Broad prompts are now poor assessment prompts. “Discuss the impact of AI on education” invites a generic generated answer. A stronger prompt asks for a position within a defined course debate, using a named framework, a local institutional policy, two required readings and one contested source found by the student. That design makes AI use less decisive because the student must connect materials that are specific to the learning environment.

There is no perfect defence. A skilled student can still use AI inside those constraints. But the work becomes harder to outsource cleanly. The student must make choices, defend them and connect them to course-specific knowledge. Assessment should make the student’s selection and judgment visible, not merely the final prose.

Student use is no longer marginal

The scale of student AI use has moved beyond experimentation. HEPI’s 2025 Student Generative AI Survey, based on 1,041 UK students, found that AI use for assessments rose from 53% in 2024 to 88% in 2025, while non-use dropped from 47% to 12%. Common uses included explaining concepts, summarizing articles and suggesting research ideas. Stanford’s 2026 AI Index education chapter reported that four out of five U.S. high school and college students use AI for schoolwork, with common uses including research, essay editing and brainstorming.

Those figures do not mean that most students are submitting AI-written work dishonestly. They do mean that AI is now part of ordinary study behaviour. The classroom conversation has changed from “Will students use it?” to “Which uses are allowed, which must be disclosed, and which uses undermine the learning outcome?” A policy that treats AI as rare or deviant is already out of date.

Student uptake also differs by discipline and task. Coding, mathematics, business writing, law, social science essays and literature reviews invite different kinds of AI use. A code assignment may involve debugging and explanation. A humanities essay may involve argument structure and prose revision. A lab report may involve data analysis and methods explanation. A design project may involve ideation and critique. Universities need discipline-specific rules because AI touches disciplines unevenly.

The adoption curve also changes student expectations. Students now encounter AI inside search engines, office software, note apps, learning management systems and phones. They may not experience AI use as a special act requiring disclosure. To them, it may feel like spellcheck, grammar support, translation or search. Universities that do not define boundaries clearly leave students guessing.

The student survey literature suggests that students want clarity, even when they use AI. The Springer study on student perspectives found that 41.1% wanted university-wide policy, while attitudes varied sharply between assistive writing tools and whole-essay generation. This creates an opening for universities. Students may accept rules if those rules are specific, fair and tied to learning rather than framed as panic.

The risk grows when staff rules differ wildly across modules. One lecturer allows AI for brainstorming; another bans it; a third says nothing; a fourth encourages it but requires disclosure; a fifth uses a detector. The student receives a hidden curriculum: AI is everywhere, but institutional expectations are fragmented. That fragmentation invites both accidental misconduct and deliberate boundary-pushing.

A serious university policy must define baseline categories. It should state what is allowed without disclosure, what is allowed with disclosure, what requires prior permission and what is forbidden. It should then let disciplines add stricter rules where learning outcomes require them. The policy must be clear enough for a first-year student to follow at midnight while writing an assignment.

Policies are moving faster, but not evenly enough

Higher education is no longer ignoring generative AI. Many institutions have built guidance pages, model syllabus statements, staff workshops, student advice and assessment reform projects. A 2025 study of policies and guidelines at the top 50 U.S. universities found that 94% had faculty-facing AI guidelines, often stressing course-specific policies, syllabus statements, assignment redesign, detection issues and ethical use. It also found fewer student and researcher guidelines, which points to uneven maturity across campus roles.

The move toward course-level policy makes sense. A blanket rule cannot handle every learning outcome. A translation exercise, legal memo, coding task, reflective essay and statistics project require different AI boundaries. Still, course-level policy without institutional structure creates confusion. A student taking five modules should not have to decode five incompatible moral philosophies of AI use.

Regulators and sector bodies are pushing assessment reform rather than simple prohibition. Australia’s TEQSA has published resources on assessment reform in a time of AI, focusing on learning assurance, academic integrity and responsible use. Its broader knowledge hub gathers institutional practices and case studies on academic integrity and generative AI in higher education. In the UK, Jisc has built resources and maturity tools for AI adoption across tertiary education.

These initiatives signal a shift from emergency response to governance. The first stage of the university AI response was often improvised: statements from provosts, warnings in syllabi, temporary bans, detector pilots, staff webinars. The second stage is harder. It requires revising assessment calendars, workload models, quality assurance, staff development, procurement, accessibility policy, student support, appeals procedures and data governance.

Policy maturity is not the same as having a PDF. A strong policy changes what happens in classes. Students receive examples of permitted and forbidden use. Lecturers receive templates that are not vague. Departments review assignments for AI exposure. Appeals processes stop relying on detector scores alone. Academic integrity boards understand how false positives occur. Libraries teach source verification. Writing centres teach AI-assisted revision without ghostwriting.

There is also a tension between flexibility and fairness. If each lecturer sets their own rules, assignments can be aligned to learning outcomes. If rules vary too much, students face uncertainty and unequal risk. The best model is layered: institutional baseline, programme-level interpretation, assignment-level permission statement. Every assignment should say plainly whether AI may be used for brainstorming, research, outlining, drafting, editing, translation, coding, data analysis and citation management.

The deep research era raises the stakes because vague permission is dangerous. “You may use AI as a study aid” is not enough when the same system can produce a source-backed report. Students need to know whether a deep research output counts as a source, a tutor, a draft, an unauthorized collaborator or all of those at different stages. A policy that does not name agentic research will be read around, not followed.

Detection became the wrong centre of gravity

AI detection appealed to universities because it promised a technical answer to a technical disruption. Feed in a text, receive a probability, investigate the suspicious cases. The appeal is obvious, especially for large courses. The problem is that detection cannot carry the weight placed on it.

Turnitin itself warns that its AI writing report is not a misconduct decision. Its guidance says low-percentage scores are less reliable, and its false-positive discussion stresses that instructors must judge cases rather than treat detection as proof. Stanford researchers also found that AI detectors can be biased against non-native English writers, reporting that more than half of TOEFL essays by non-native English students were classified as AI-generated by detectors, with nearly all flagged by at least one of seven tools.

That finding is devastating for universities with international cohorts. A detector that penalizes predictable prose or lower lexical complexity may punish students who are already writing under linguistic pressure. It may also be gamed by adding awkward variation, paraphrasing or human editing. Detection is weakest exactly where the institution needs fairness most: contested cases with high consequences.

The deeper flaw is conceptual. Detection focuses on text origin. Assessment needs evidence of learning. A student may use AI heavily but still learn if the task requires critique, verification and defence. Another student may write without AI and learn little through rote paraphrase. Text origin matters for honesty, but it is not the whole educational question.

Deep research tools make detection even less useful because the final prose may be rewritten by the student. A student can use AI for research, planning and source selection, then draft manually. A detector may show nothing, but the intellectual process was still partly outsourced. Another student can write the argument and use AI for grammar polishing, then receive a suspicious score. The detector sees surface patterns, not process.

Universities should therefore treat detection as a weak signal, not a foundation. It may prompt a conversation. It should not replace evidence. Better evidence includes draft history, source annotations, in-class tasks, oral examination, version records, data files, coding logs and the student’s ability to explain choices. This is slower than detection. It is also more aligned with education.

There is a cost issue here. Process-based assessment takes staff time. Oral defences do not scale easily. Draft feedback requires workload. Large first-year courses need practical designs. The answer is not to do everything everywhere. It is to triage. High-stakes assignments, capstones, professional accreditation tasks and research-heavy modules need stronger evidence. Low-stakes practice tasks can use AI more openly. Detection should sit at the edge of an integrity system, not at its centre.

The equity problem cuts both ways

AI in university writing creates a fairness problem with two opposing faces. On one side, paid AI tools may give wealthier students better research, writing and feedback support. On the other side, bans may harm students who use AI for language support, disability accommodation, study confidence or basic access to explanation. A fair policy has to hold both truths at once.

The paid-tool issue is growing. OpenAI’s deep research access expanded over time, with different monthly limits across Free, Plus, Team, Enterprise, Edu and Pro accounts. Google tied early Deep Research access to Gemini Advanced. If better tools sit behind subscriptions, students with money can receive more advanced support. Universities may then face a strange version of the old tutoring gap: private AI assistance at scale.

The access issue does not only concern money. Students differ in technical confidence. Some know how to prompt, verify and iterate. Others do not. Some know which sources to request. Others ask for broad summaries and accept weak answers. Some study in disciplines where AI tools are already woven into professional practice. Others receive little guidance. Unequal AI literacy may become a new academic advantage.

The ban issue is just as real. The Springer student-perspectives study warned that outright bans may disadvantage students who need writing support, including disabled students, neurodiverse students and students writing in a non-native language, while paid versions raise inequality concerns. Stanford’s detector-bias findings deepen that concern because non-native English writers may face a higher risk of false suspicion.

The most credible equity position is not “ban AI” or “let everyone use it freely.” It is structured access with clear boundaries. Universities can provide institutionally approved tools, teach verification, require disclosure for substantive use, protect accommodations, and avoid punitive reliance on detectors. They can also design assignments that reduce the advantage of expensive tools by tying work to local teaching, oral defence, in-class steps and course-specific materials.

Equity also requires attention to language. A policy that says “AI-generated text is forbidden” may sound clear until a student uses AI to translate their own ideas into academic English. Is that generated text, editing, language support or unauthorized authorship? The answer should depend on the learning outcome. In a writing course, language production may be central. In a science course, clarity and accuracy may matter more than unaided prose. The rule needs to say so.

Fairness now means transparency about which parts of the task must be human. Students should not have to guess whether grammar support is allowed, whether AI summaries count as reading, whether a chatbot may explain feedback, or whether a deep research report can be used as a starting point. Ambiguity punishes the cautious and rewards the bold.

The safest assignments now make the process visible

A finished paper used to hide enough of the process that markers learned to read between the lines. They looked for argument quality, source use, voice, structure and disciplinary judgment. AI weakens those signals. A safe assignment now needs to surface the process directly.

Process evidence can be simple. A student submits a research log with search terms, databases used, source selection criteria and rejected sources. A draft includes comments explaining changes between versions. An annotated bibliography names the claim each source supports. A short reflection explains where AI was used and which suggestions were rejected. A seminar task asks students to defend one source orally. The goal is not bureaucracy. The goal is to make thinking harder to counterfeit.

This approach also supports learning. Students often struggle because they do not know what good research looks like before the final draft. Requiring interim steps gives lecturers a chance to correct weak source selection, shallow reading or unfocused questions earlier. It makes the hidden curriculum visible. If AI is allowed, the process record can show whether the student treated AI as a shortcut or as a tool to test ideas.

There is a design risk. Process evidence can become performative. A student may ask AI to generate a fake research log or reflection. The answer is to connect process tasks to live course activity. Ask students to refer to a seminar debate, use a source discussed in class, bring notes to a workshop, explain a change made after feedback, or answer a brief oral question. The more the process is embedded in teaching, the harder it is to fabricate cleanly.

Process evidence should also be proportionate. A 500-word weekly reflection does not need the same verification as a final-year dissertation. Excessive documentation can become busywork and overload staff. The most useful process tasks are those that also teach the discipline: source comparison, method justification, evidence ranking, concept application, peer critique and revision rationale.

The deep research era makes one process question central: did the student merely accept a generated pathway, or did they make defensible choices? A student can be asked to submit the AI prompt and output, then identify three claims they accepted, three they rejected and one source they verified independently. That turns AI use into assessable judgment. It also discourages blind submission because the student knows they must account for the tool’s work.

The safest assignment is not AI-proof. It is learning-rich enough that AI alone cannot complete the educational task. That distinction matters. No mass university system can make every assignment impossible to game. It can make gaming less aligned with success and make genuine engagement more visible.

Writing still matters because judgment lives in revision

Some observers treat AI writing as evidence that writing itself is becoming less central. That is the wrong lesson. Writing matters more when machines can produce fluent prose because the human task moves toward judgment: deciding what should be said, what should be cut, what evidence is strong, what caveat is honest and what structure serves the argument.

A student who lets AI produce a first draft may still learn if they revise deeply. But many will not revise deeply unless the course teaches and rewards it. The danger is that AI creates a polished draft so early that students skip the hard middle stage where understanding is formed. They may not wrestle with the source, notice contradictions, discover that their thesis is too broad, or learn why a paragraph fails.

Research on AI-assisted essay writing raises this concern. A 2024 study in Computers in Human Behavior assigned university students to use ChatGPT or Google for argument-writing tasks and found that ChatGPT users experienced lower cognitive load but produced weaker reasoning and argumentation quality, even though perspective diversity did not differ. The finding does not prove that AI always harms learning. It does show a plausible trade-off: lower effort may come at the cost of deeper reasoning.

A 2025 MIT-linked preprint popularly discussed as “Your Brain on ChatGPT” reported weaker neural connectivity and lower ownership among LLM users in an essay-writing experiment, though the work has also drawn methodological critiques and should be treated as provisional rather than settled science. The broader lesson is cautious: educators should not assume that a better-looking draft means stronger learning.

Revision is the place to rebuild the link. Ask students to submit a paragraph before and after revision with a note explaining the change. Ask them to identify the weakest claim in their own draft. Ask them to show how feedback altered the argument. Ask them to compare an AI-generated paragraph with their own and explain which is more accurate. Revision turns writing from output into evidence of thought.

This also changes feedback. Lecturers may need to comment less on surface grammar and more on conceptual control. If AI can clean a sentence, the teacher’s scarce time should be spent on argument, evidence, method and interpretation. Writing centres may shift from “make this more polished” to “make this more yours, more accurate and more defensible.”

Student writers also need permission to sound human. AI prose often pushes toward smooth neutrality. Academic writing at its best is not neutral mush; it is precise, situated and accountable. A student who has read deeply should be able to make a claim with texture. Universities should teach students to avoid AI-flattened prose not because style is sacred, but because flat prose often conceals flat thinking. The future of student writing is not less writing. It is more explicit ownership of choices.

Literature reviews are the first casualty

The literature review is the assignment format most exposed to deep research tools. It asks for exactly the skills these systems are being trained to imitate: search widely, group sources, summarize findings, identify gaps and present a coherent map of a field. A student can now request a review on a topic, ask for themes, receive source links and turn the material into a structured narrative.

This does not make literature reviews obsolete. It makes generic literature reviews weak assessment. A broad prompt such as “review the literature on social media and mental health” is now a gift to AI. A stronger task might require students to compare three theoretical traditions, evaluate two methods, trace a debate across five assigned readings and find one recent empirical study that challenges the course narrative. The review must test judgment, not only coverage.

The quality problem is subtle. Deep research tools may retrieve plausible sources and organize them neatly, but they can still miss field-defining works, overweight accessible sources, summarize abstracts rather than arguments, blur methods, or invent a false consensus. Students who know the field can catch this. Students who do not may submit a clean but shallow map.

Lecturers can redesign literature review tasks around source accountability. Students might submit a source matrix listing research question, sample, method, finding, limitation and relevance to their thesis. They might rank sources by evidentiary strength. They might identify one paper that changed their view and explain why. They might include a short “non-use” note for two sources they rejected. These tasks make AI assistance possible but not sufficient.

Deep research also changes postgraduate training. Doctoral students will use these tools; banning them would be artificial in many fields. Supervisors should teach students to audit AI-generated literature maps: check database coverage, verify source existence, inspect citation chains, identify missing journals, compare with expert bibliographies and test whether the claimed “gap” is real. A generated gap is especially dangerous because it may be a gap in the model’s retrieval, not a gap in knowledge.

The phrase “AI can write a PhD-level review” should alarm supervisors less for its literal truth than for what it reveals about weak review practices. If a review is only a thematic summary, AI will do it well enough to disrupt assessment. If it is a disciplined act of positioning a research problem within a field, with method-aware critique and evidence of intellectual risk, the student still has work to do.

Universities should stop treating the literature review as a static genre. It needs oral defence, process logs, field maps, source matrices and supervisor questioning. It should become less about producing a polished chapter and more about proving that the researcher can make reliable choices in a crowded field.

Method sections expose the difference between knowledge and performance

AI can explain a method. It can compare qualitative interviews with surveys. It can describe regression assumptions, thematic analysis, randomized trials or archival research. It can produce a tidy methods section for a proposal. That does not mean the student can do the method.

Methods are where surface fluency often breaks. A student who submits an AI-assisted research proposal may describe sampling, validity, reliability, ethics and analysis with confidence. Ask them why the sample size is defensible, how the recruitment strategy biases the findings, what happens if the data are missing, or how the coding framework will be tested, and the gap may appear quickly. Methods require operational understanding, not only correct terminology.

Deep research tools raise the quality of methods prose, which can fool assessment if the task is purely written. The stronger assessment asks students to perform part of the method: code a short excerpt, interpret a regression output, build an interview guide, justify inclusion criteria, clean a dataset, write an ethics risk note, or respond to a methodological objection. The student must show that the method is usable in their hands.

The same issue appears in data analysis. OpenAI’s system card says deep research can write and execute Python for data analysis. For professional analysts, that is useful. For students, it creates a grading problem. If the learning outcome is “interpret data using appropriate methods,” AI-assisted coding may be allowed with disclosure. If the outcome is “learn to write statistical code,” AI-generated scripts may undercut the task unless students must explain and modify them.

The answer depends on disciplinary purpose. A public policy course may reasonably allow AI for data cleaning if the mark rests on interpretation and policy reasoning. A statistics course may forbid it for the coding portion. A computer science course may allow AI suggestions but require students to annotate each function and pass an oral code review. The same AI action can be acceptable or unacceptable depending on the skill being certified.

Methods also reveal why “PhD-level” output is not the same as doctoral competence. A doctoral student must live with the consequences of methods choices. If recruitment fails, data are messy or assumptions break, the student must adapt. AI can propose a method, but it does not face the fieldwork, ethics board, failed pilot, hostile archive or contradictory dataset. The doctoral craft sits in those frictions.

University assignments should therefore put friction back into methods assessment. Let students use AI to propose options, then require them to choose and defend one under constraints. Give them flawed data. Give them a rejected ethics application. Give them a method that does not match the question and ask them to repair it. The goal is not to make life harder for its own sake. It is to certify that students can think when the template fails.

The new stack behind one assignment

The modern university assignment increasingly sits on a hidden stack of tools. A student may use a browser with AI summaries, a chatbot with deep research, a PDF reader with automatic extraction, a citation manager, a grammar tool, a translation system, a note app, a code assistant and a document editor with embedded AI. The submission arrives as one file, but the work behind it may have passed through many systems.

Assignment work before and after agentic research tools

Assignment stageOlder student workflowAI-assisted workflowAcademic risk
Topic framingStudent narrows question after readingTool proposes research questions and gapsStudent may adopt a question they cannot defend
Source discoveryDatabase search, library guides, citation chasingDeep research gathers and sorts sourcesCoverage may look strong while missing core literature
ReadingStudent reads papers and takes notesTool summarizes PDFs and extracts claimsSummary may replace reading rather than support it
PlanningStudent builds outline from notesTool turns findings into structureArgument may reflect model logic, not student judgment
DraftingStudent writes and revisesTool drafts sections or full paperAuthorship and learning evidence blur
ReferencingStudent checks citation style and source useTool formats and links referencesCitation presence may be mistaken for understanding

This table does not imply that every AI-assisted step is misconduct. It shows why assignment rules need to name stages. A policy that only mentions “writing with AI” misses the research, planning and source-selection work that now shapes the final submission.

The stack matters because students may not see all these acts as “AI use.” Search summaries feel like search. Grammar rewriting feels like editing. A PDF summary feels like accessibility. A citation suggestion feels like reference management. A deep research report feels like a starting point. From the institution’s perspective, however, each stage may affect the student’s authorship and learning.

The practical challenge is disclosure. A demand to disclose every AI interaction may be impossible and intrusive. A demand to disclose only whole-text generation is too narrow. The better approach is threshold-based disclosure. Students disclose AI use when it contributes substantively to content, argument, source selection, data analysis or wording beyond routine spelling and grammar checks. Routine tools can be exempted if the course allows them.

Yet “routine” will keep changing. A grammar checker that once corrected commas may now rewrite tone, restructure paragraphs and suggest arguments. A search engine may now answer in synthesized prose. A document editor may generate a draft from notes. Universities need policies based on academic function, not tool brand. The question is not whether the student used ChatGPT, Gemini, Claude or Copilot. The question is whether a system performed a task the student was supposed to learn or be assessed on.

The stack also creates privacy and data concerns. Students may upload lecture notes, unpublished data, copyrighted readings, patient-like case material or personal reflections into third-party systems. They may not understand retention settings or model-training policies. Universities cannot tell students to “use AI responsibly” while leaving them to choose tools with different data practices. Institutional procurement and approved-tool lists are now part of academic integrity.

Deep research sharpens the issue because it can combine user files with web research. That is powerful for a dissertation draft, a lab report or a policy memo. It is also risky if students upload sensitive material. The educational question and the governance question now meet inside the same prompt box.

Assessment needs friction by design

Friction has a bad reputation in technology because friction often means inconvenience. In learning, friction is not always a defect. Struggle, delay, uncertainty and revision help students build knowledge. The deep research era forces universities to distinguish harmful friction from educational friction.

Harmful friction is administrative clutter: unclear policies, duplicate disclosure forms, inconsistent lecturer expectations, inaccessible tools and punitive suspicion. Educational friction is different. It asks the student to explain, choose, test, defend and revise. The goal is not to slow students down. The goal is to keep the learning work inside the task.

A frictionless assignment is now easy to outsource. Ask for a broad essay, mark only the final text, give no oral check and require no process evidence; AI will fit neatly. Add purposeful friction: a seminar-based claim, a source matrix, a class dataset, a short defence, a revision note and a disclosure statement. AI remains usable, but the student must work through it.

This is especially relevant for deep research outputs. These tools are built to reduce the time cost of exploration. In professional settings, that is a gain. In education, exploration is often the learning. A first-year student needs to learn why some sources are poor, why search terms matter, why abstracts can mislead, why method sections are hard, and why a neat answer may conceal uncertainty. If AI removes those discoveries, the student loses more than time.

Friction by design does not require every course to become an exam. It can be built into assignments through staged submission. Week one: research question and rationale. Week two: source matrix. Week three: argument plan. Week four: draft paragraph and peer review. Final submission: essay plus revision note and AI disclosure. This structure gives the lecturer more evidence and gives students less incentive to outsource at the end.

For large classes, friction can be sampled. Not every student needs an oral defence every week. Departments can use random short vivas for final papers, require group poster defences, run in-class applied tasks linked to essays, or ask students to answer one personalized question after submission. The knowledge that the student may have to explain the work changes behaviour.

Assessment should make the cheapest path to a good mark pass through learning. At present, many assignments make the cheapest path pass through polished production. AI exploits that. Redesign should redirect effort toward judgement, application and defence.

Responsible use cannot be reduced to citation

Many AI policies tell students to cite AI use. Disclosure is needed, but citation alone is a weak fix. A student can cite a chatbot and still submit work they do not understand. Another student may use AI as a tutor and produce genuinely learned work. The citation tells the marker that a tool was involved. It does not settle the academic question.

Responsible use has at least four layers. The first is permission: was AI allowed for this task? The second is transparency: did the student disclose substantive use? The third is verification: did the student check claims, sources and outputs? The fourth is ownership: can the student defend the final work as their own intellectual product? Citation addresses only one layer.

The ownership layer is the hardest. If AI generated a literature review plan and the student followed it, who owns the argument? If AI proposed the research gap and the student wrote the prose, who owns the contribution? If AI drafted the methods section and the student edited it, what is being assessed? These questions cannot be answered by a footnote. They require course-level boundaries.

Universities should avoid two false comforts. The first is treating AI as just another source. It is not. A source is a stable object that can be checked. A generative system is a process that produces different outputs and may synthesize from many materials, some visible and some not. The second false comfort is treating disclosure as moral cleansing. A disclosed violation may still be a violation if the tool performed the assessed work.

Copyright and authorship debates add another layer. The U.S. Copyright Office has been examining copyright issues raised by AI, including the copyrightability of outputs created with generative AI and registration guidance for works containing AI-generated material. Academic authorship is not identical to copyright, but the policy logic overlaps: human contribution and control matter.

Responsible use also requires source discipline. If a deep research report cites a source, the student should open the source, verify the claim, check the context and decide whether it belongs in the argument. A useful assignment can require a “verified source note” for each AI-suggested source: claim, page or section, relevance, limitation and student judgment. That is more work than pasting a disclosure line, but it teaches the skill universities now need.

The future academic integrity statement should read less like a confession and more like a methods note. It should say what tool was used, for which stage, under which permission, how outputs were checked and what the student changed. That kind of statement is harder to fake convincingly and more useful to the marker.

Learning modes reveal the industry’s defensive pivot

AI companies have understood that education is a legitimacy test. If their tools are seen mainly as cheating engines, universities will resist procurement and regulators will pay closer attention. The product response is visible: study modes, learning modes, campus offerings and educator controls.

OpenAI’s Study Mode, introduced in July 2025, is presented as a way to give step-by-step guidance rather than quick answers, with features such as interactive prompts, scaffolded responses, knowledge checks and personalization. OpenAI also acknowledges limitations, saying the mode is powered by custom system instructions and may behave inconsistently while the company studies learning outcomes. Anthropic’s Claude for Education takes a similar public stance, with a learning mode intended to guide reasoning rather than provide direct answers.

This pivot matters because it shows where the battle line is moving. The first consumer AI tools competed on answer quality. Education tools now need to compete on learning quality, institutional trust and governance. A chatbot that refuses to write a student’s answer but asks probing questions may be less attractive to a student in a hurry, but more acceptable to a university.

There is a commercial logic here. OpenAI launched ChatGPT Edu for universities that want to deploy AI across students and campus communities. Anthropic announced campus partnerships and education programs around Claude. Google has integrated Deep Research into Gemini subscriptions and Workspace-related workflows. The campus is becoming a market for managed AI, not only a place where students bring consumer tools.

Learning mode is also a policy compromise. It lets universities say they are not banning AI while still discouraging answer outsourcing. It gives students a supported tool rather than leaving them to private subscriptions. It gives vendors a route into institutional procurement. It gives lecturers a way to distinguish tutoring from ghostwriting.

The open question is evidence. Do learning modes actually improve learning, or do they only wrap answer systems in educational language? OpenAI’s own note that Study Mode can be inconsistent and that learning outcomes are still being studied is a useful warning. Universities should ask vendors for evidence, not only features: retention, transfer, student understanding, equity effects, accessibility, privacy and error rates.

Learning modes should also be inspected by discipline. A Socratic prompt may work in one field and frustrate another. A coding tutor, language tutor, law tutor and statistics tutor need different guardrails. A generic “teach me” mode is not enough for high-stakes assessment. The more universities integrate AI, the more they must test whether the tool supports the intended learning outcome.

There is a wider cultural issue. Students may learn to treat AI as the default interlocutor before peers, tutors or libraries. That may be useful at times, but universities are social institutions. Learning includes debate, feedback, disagreement and intellectual identity. AI tutoring should not replace the human situations where students learn to defend and revise their ideas in public.

Universities are now platform customers

Higher education used to buy learning management systems, library databases, plagiarism tools, office software and research infrastructure. AI adds a new category: campus-scale cognitive infrastructure. The decision to adopt, block or ignore an AI platform now shapes assessment, student support, staff workload, privacy, equity and institutional reputation.

This creates procurement questions that academic departments cannot answer alone. What data are stored? Are student prompts used for training? Can the institution control retention? Are copyrighted readings protected? Can the tool access internal files? Are audit logs available? Are accessibility needs met? Can students opt out? What happens when a vendor changes model behaviour, pricing or access limits?

Deep research features heighten these questions because they blend web search, user files, reasoning and report generation. A university that allows students to upload lecture materials or research data into a tool should know where those materials go. A university that recommends a tool should know whether it is equally available to all students. A university that embeds a tool in assessment should know how failures are handled.

The market pressure is strong. AI spending has become a major technology trend. Stanford’s 2025 AI Index reported that U.S. private AI investment reached $109.1 billion in 2024 and that private investment in generative AI reached $33.9 billion, while business use of AI rose sharply. Education is not outside that economy. It is a target market, a data environment, a training ground and a legitimacy arena.

The risk is that universities become dependent before they become competent buyers. A platform may appear to solve student support, writing feedback or staff workload, but it may also create lock-in, uneven access, hidden costs and new forms of surveillance. Procurement should be tied to pedagogy and governance, not only innovation branding.

Platform dependence also affects academic freedom. If a university’s approved AI tool filters topics, refuses certain tasks, ranks sources through opaque systems or changes behaviour without notice, teaching may be shaped by vendor policy. This is not a reason to reject every tool. It is a reason to keep human control, transparent documentation and alternative routes for students.

The library has a major role. Librarians understand source quality, database access, copyright, metadata and information literacy. The deep research era makes those skills more central, not less. A campus AI strategy that excludes libraries will likely overfocus on tools and underfocus on evidence.

Writing centres also need a seat at the table. They know how students actually draft, panic, revise, misunderstand prompts and respond to feedback. AI governance designed only by senior management and IT will miss the lived reality of assignments. The university platform decision is now an academic decision, not a software purchase.

Copyright, privacy and data governance enter the classroom

The student assignment used to raise copyright issues mostly through plagiarism and reuse. AI adds new questions. Students may paste copyrighted articles into tools. They may ask a model to summarize a paywalled chapter. They may submit AI-generated material whose copyright status is uncertain. They may upload unpublished research notes, interviews or institutional data. The assignment becomes a data governance event.

The U.S. Copyright Office’s AI work shows how unsettled the broader field remains. Its AI initiative examines the copyrightability of works containing AI-generated material, the use of copyrighted works in training and registration guidance for AI-assisted works. Universities do not need to resolve every copyright debate to write classroom rules. They do need to tell students what materials may be uploaded and what cannot.

Privacy is more immediate. A student may paste identifiable interview excerpts into a chatbot for coding. A trainee teacher may upload pupil work. A health student may summarize a case. A business student may upload company data from a placement. Even when the student intends no harm, the tool may not be approved for that data. AI misconduct is not only about cheating. It can also be careless disclosure.

This is why approved tools matter. An institutionally managed AI service may offer stronger data controls than a consumer account. But approval is not magic. The university still needs rules for sensitive data, research ethics, copyright-protected readings and assessment submissions. Students should receive examples, not abstract warnings.

Data governance also affects staff. Lecturers may be tempted to upload student essays into AI tools to generate feedback. That raises consent, privacy and intellectual property questions. A university cannot forbid students from uploading material while quietly allowing staff to upload student work into unapproved systems. The rule should apply across roles.

Deep research workflows bring another governance risk: source contamination and overcollection. A tool may browse many sites, retrieve uncertain material and generate a report. Students may assume the system has used sources lawfully and appropriately. But academic use still requires source checking. A generated report is not a clean substitute for library access, scholarly databases or course readings.

The privacy conversation should not be framed as a reason to avoid all AI. It should be framed as part of professional training. Students entering law, medicine, education, public administration, journalism, engineering and business will need to understand when AI use is unsafe because data are sensitive. Universities should teach AI data judgment as a professional competence.

The EU AI Act makes education a compliance question

AI in education is no longer only a teaching issue. It is becoming a regulatory issue. The European Commission describes the EU AI Act as a risk-based framework for developers and deployers, with different obligations depending on risk level. The Act bans certain practices, including emotion recognition in workplaces and education institutions, and introduces transparency obligations for generative AI and rules for general-purpose AI models.

The timeline matters for universities in Europe and for vendors serving them. The Act entered into force on August 1, 2024; prohibited practices and AI literacy obligations began applying on February 2, 2025; general-purpose AI rules began applying in August 2025; many transparency obligations for generative AI are set for August 2026, while high-risk education-related provisions follow later under the Commission’s timeline.

Education appears in the Act because AI systems may affect access, assessment, progression and opportunity. A tool used for tutoring is not the same as a tool used to decide admission, grade exams or flag misconduct. Universities must therefore separate low-risk support from higher-stakes decision systems. An AI detector used in a disciplinary process is not just a classroom gadget; it may become part of an institutional decision chain.

The AI literacy obligation is especially relevant. Universities cannot treat AI literacy as optional enrichment when students and staff are using these tools across learning and assessment. Staff need to understand tool limits, bias, privacy, copyright, academic integrity and assessment design. Students need the same, framed around their disciplines and future professions.

UNESCO has made a similar policy argument at global level. Its guidance for generative AI in education and research warns that the release of public AI tools has outpaced many national regulatory frameworks and has left institutions unprepared to validate tools in many settings. The EU AI Act gives European institutions a legal frame, but the educational work still has to be done locally.

Compliance should not be reduced to paperwork. A university may update policies and still fail students if teaching practice does not change. The practical questions are concrete: Is AI literacy part of induction? Do assessment boards know how to handle AI disputes? Are staff trained not to rely blindly on detectors? Are procurement decisions documented? Do students know which tools are approved? Are high-stakes AI systems reviewed?

The deep research era means universities cannot wait for perfect legal clarity. Assessment redesign, staff training and student guidance are needed now. Regulation will shape the floor. Academic standards must set the ceiling.

The business impact reaches far beyond academic integrity

The AI assignment problem is also a labour-market problem. Employers use degrees as signals of writing, analysis, research and problem-solving ability. If graduates can submit strong work without gaining those abilities, employers will adjust. They may rely more on interviews, work samples, probation tasks, technical tests or institution-specific reputation. The market value of a degree depends on trust that the graduate can perform without hidden assistance.

At the same time, workplaces themselves are adopting AI. Employers may not want graduates who avoid AI. They may want graduates who use it well: verify sources, protect data, improve drafts, analyze information, document decisions and know when not to use it. That creates a double requirement for universities. They must stop AI from hollowing out learning while also preparing students for AI-rich work.

The old academic integrity frame is too narrow for this. A business school graduate who can use AI to produce a market analysis but cannot judge the sources is a risk. A law student who can draft with AI but cannot check authority is a risk. A health student who uploads sensitive case information into a public tool is a risk. An engineer who accepts generated code without testing is a risk. The issue is professional reliability, not only cheating.

This changes curriculum design. AI use should be taught inside disciplinary tasks rather than as a generic digital skills workshop. Business students should audit AI-generated market claims. Law students should verify legal citations and jurisdiction. Science students should check methods and data interpretation. Humanities students should critique AI-generated readings of texts. Computing students should inspect generated code for security and maintainability.

Employers may also change what they ask from universities. They may seek clearer evidence of assessed human competence, especially in professional programmes. Accreditation bodies may require AI-aware assessment policies. Placement providers may demand data governance training. Graduate recruiters may ask candidates to explain AI use in portfolios.

Universities that respond well can strengthen their reputation. They can say their graduates are not AI-free; they are AI-literate and accountable. They can show assessment designs that test independent judgment, collaborative work, tool use and oral defence. They can teach students to document AI use as professionals document methods, assumptions and limitations.

The risk for weaker institutions is credibility drift. If students, employers and staff believe assignments are easily gamed, the degree loses signal value. That loss may not happen suddenly. It may accumulate through stories: graduates who cannot write under pressure, interns who cannot explain their own reports, employers who add extra tests, students who see honest work as naive. Trust erodes through repeated small doubts.

Faculty workload is becoming the silent bottleneck

AI assignment reform sounds sensible until it reaches staff workload. Process evidence, oral checks, staged drafts, AI literacy teaching and policy explanation all take time. Many lecturers already teach large classes, manage research expectations, handle administration and support students with complex needs. Telling them to redesign assessment without time, tools or recognition is not a plan.

The faculty workload issue explains why detection became attractive. It promised scale. Process-based assessment does not scale as neatly. A ten-minute oral defence for 300 students is not practical without structural change. A detailed source matrix for every paper may overwhelm marking. Department leaders need to treat AI assessment reform as workload reform, not only pedagogy.

There are practical mitigations. Oral checks can be short and sampled. Draft review can use peer workshops before staff grading. Rubrics can focus on process evidence without adding long narrative feedback. AI disclosure statements can be standardized. Libraries and writing centres can teach source verification at scale. Programmes can identify a smaller number of high-stakes assignments that need heavier verification rather than redesigning every task equally.

The staff development burden is real. Many lecturers are not AI experts, and they should not have to become prompt engineers overnight. They need usable examples by discipline: allowed-use statements, redesigned prompts, sample rubrics, process evidence templates, appeal procedures and detector guidance. Sector bodies such as TEQSA and Jisc are useful because they gather practices and frameworks, but local adaptation still requires time.

A second workload problem is emotional. AI suspicion can poison teaching relationships. Staff may feel forced to police students. Students may feel presumed guilty. Academic integrity meetings can become adversarial and stressful. A well-designed policy reduces this by making expectations clear and using process evidence before accusations. It also gives staff a route for concern that does not depend on gut feeling or detector scores.

The workload issue also cuts across disciplines. Writing-heavy courses face one pressure; coding courses face another; lab courses another. Central policy should not dump the same solution everywhere. Departments need local assessment audits: Which assignments are most vulnerable? Which learning outcomes are most at risk? Which tasks already include human verification? Which courses need support first?

Institutions should also count AI reform work in workload models. Redesigning assessments, attending training, updating rubrics and handling AI disclosure are academic labour. If universities treat this as unpaid invisible work, implementation will be patchy and resentment will grow. The deep research era cannot be managed by individual lecturer heroics. It requires institutional capacity.

The strongest answer is not a ban

A total ban sounds clean until it meets reality. AI is already embedded in search, writing tools, productivity software and study habits. Students can use it privately. Staff may use it for lesson planning, feedback drafts or research. Employers expect AI familiarity. A ban may be justified for particular assessments, especially where unaided performance is the learning outcome. As a general institutional stance, it is brittle.

A free-for-all is no better. If students may use AI without disclosure for all stages, the degree stops distinguishing between tool-assisted production and student competence. The university becomes an output-certification service rather than a learning institution. Students who want to learn may feel foolish when others outsource. Staff lose confidence in marks.

The stronger answer is permission with boundaries. AI should be allowed where it supports learning and forbidden where it replaces the assessed skill. This sounds simple, but it requires each assignment to name the skill. Is the course assessing source discovery, source evaluation, argument writing, coding fluency, data interpretation, language production, professional judgment, creativity, collaboration or oral defence? The AI rule follows from that answer.

Some tasks should be AI-open. A course on digital marketing might ask students to use AI tools to draft campaign variants, then critique bias, evidence and brand fit. A policy course might ask students to compare an AI-generated briefing with official documents. A research methods class might ask students to audit a generated literature map. In these tasks, AI use is part of the learning.

Some tasks should be AI-restricted. A language exam, closed-book problem set, foundational coding exercise or diagnostic writing sample may need unaided work. The restriction should be explained. Students accept rules more readily when they understand the skill being protected. “No AI because we say so” is weaker than “No AI because this task certifies your ability to construct an argument unaided in timed conditions.”

Some tasks should be AI-disclosed. Many assignments will sit here: students may use AI for brainstorming, source discovery or editing, but they must disclose substantive use and verify outputs. The mark can reward critical use rather than secret use.

The ban debate is a distraction when the real design question is alignment. Match the AI rule to the learning outcome, then create enough evidence to support the mark. That is harder than a ban. It is also more honest.

A usable policy has to name the allowed move

Students do not need abstract statements about integrity. They need to know what they may do on Tuesday night when the assignment is due on Wednesday. A usable AI policy names actions. It distinguishes brainstorming from drafting, summarizing from reading, editing from rewriting, translation from authorship, and source discovery from source evaluation.

A cleaner assignment policy model

Permission levelStudent can doStudent must showStronger assessment check
No AI useComplete the task unaided except approved accessibility toolsStatement of unaided work if requiredIn-class or oral verification
AI for study onlyAsk for explanations, quizzes or feedback before writingNo AI text or source list in submissionShort concept check
AI for research supportUse AI to find sources or map debatesVerified source notes and disclosureAnnotated bibliography or source defence
AI for drafting with limitsUse AI for outline or paragraph feedbackPrompt/output record, revision note, disclosureDraft comparison and oral question
AI-integrated taskUse AI as part of the assignmentCritical evaluation of tool outputRubric rewards verification and judgment

The table is deliberately practical. It treats AI policy as an assignment design tool, not a moral slogan. Students should see this kind of statement on each assignment brief, with examples from the discipline.

The phrase “allowed move” matters. Academic writing is a game with rules, but many rules were tacit. AI makes tacit rules unfair. A student may think asking for an outline is harmless. A lecturer may see it as outsourcing structure. A student may think summarizing a paper is like reading an abstract. A lecturer may see it as skipping reading. Both may be acting in good faith until the policy is specific.

Assignment briefs should therefore include short AI-use blocks. For example: “You may use AI to brainstorm possible angles and to check grammar. You may not use AI to generate paragraphs, choose sources or write the final argument. If you use AI for brainstorming, include a short disclosure naming the tool and how it changed your plan.” That is clearer than a page of principles.

The policy should also tell students what good disclosure looks like. Bad disclosure: “I used ChatGPT.” Useful disclosure: “I used Claude on March 3 to generate possible counterarguments after I had drafted my outline. I rejected two suggestions because they did not match the course readings. I used Grammarly for sentence-level editing. No AI-generated paragraphs are included.” The second statement gives the marker evidence of control.

A usable policy reduces accidental misconduct and makes deliberate misconduct easier to identify. When the allowed moves are clear, a student who secretly submits a generated deep research report has less room to claim confusion. At the same time, a student who uses AI properly is not punished for ordinary study support.

Policy should also avoid tool-specific fragility. Naming tools can be helpful in examples, but the rule should survive product changes. A student may use ChatGPT, Gemini, Claude, Copilot, Perplexity, a local model or an AI feature inside a document editor. The policy should define functions: generate, summarize, translate, analyze, plan, code, edit, cite, search and verify.

The final piece is enforcement. A policy without assessment design is a sign on an unlocked door. Students must know that they may be asked to explain sources, defend their argument, show drafts or discuss AI use. Not every student needs heavy checking. The possibility of checking, aligned with a clear policy, changes the risk calculation.

The future assignment will ask for evidence of thinking

The university assignment is not dead. It is being forced to become more honest. A strong future assignment will not ask only for a polished product. It will ask for evidence of thinking: choices, revisions, source judgments, failed attempts, ethical constraints, data decisions and oral defence.

This is a return to what good teaching already valued. The best supervisors never judged a thesis only by the final bound copy. They asked how the student arrived there. They challenged interpretations. They probed sources. They watched the argument mature. AI pushes undergraduate and master’s teaching toward that same evidence-rich model, though with practical limits.

The future assignment may include more live components. Not all assessment will move to exams, but timed in-class writing, oral explanation and applied problem-solving will gain weight. A student who submits a report may need to answer a five-minute question about the strongest source. A group project may include individual defence. A dissertation may include a methods viva. The student’s voice will matter again, not as style, but as accountable explanation.

Portfolios may also grow. A portfolio can show drafts, feedback, revisions, source notes, data work, AI disclosures and reflections. It gives a richer picture than one essay. It also aligns with professional practice, where people document decisions and iterations. The challenge is marking load, so portfolios need clear rubrics and selective evidence.

Authentic assessment will become more attractive, but it needs care. Asking students to produce workplace-like tasks does not automatically solve AI misuse, because workplaces also use AI. The task must include professional accountability: verify the briefing, check legal or ethical constraints, document assumptions, defend recommendations and explain tool use. Authenticity without accountability becomes another polished output.

Course-specificity will matter. Assignments tied to local class discussions, live datasets, institutional contexts, field visits, labs, studios or community partners are harder to outsource. AI can still assist, but the student must connect output to experiences and materials not fully available online. This benefits learning as well as integrity.

The future assignment will also teach students to challenge AI. Instead of banning generated output entirely, a course may ask students to produce a critique of an AI literature review, identify missing sources, find hallucinated claims, test bias, and rewrite the argument with verified evidence. That makes AI the object of analysis rather than a hidden assistant.

The strongest assignments will not pretend AI does not exist. They will require students to prove they can think with it, against it and without it. Those are three different competencies, and universities need all of them.

The university degree is being re-priced by trust

The deep research era puts a price on trust. A degree is valuable because people believe it certifies more than attendance. It certifies ability: to read, reason, write, solve, interpret, apply, create, critique and learn. If AI makes academic production easier to fake, trust becomes more expensive to maintain.

Universities can spend that trust badly. They can over-rely on detectors, punish false positives, issue vague bans, ignore equity, let staff improvise and keep marking final essays as if nothing changed. That path may preserve routines for a short time, but it will not preserve confidence.

They can also spend it well. They can redesign vulnerable assignments, teach AI literacy, give students approved tools, require process evidence where marks are high, protect privacy, train staff, clarify policy and engage employers and professional bodies. That path costs time and money. It also makes the degree more defensible.

The issue is not whether AI can produce a PhD-level-looking paper. It can often produce work that resembles advanced academic prose closely enough to disturb old marking habits. The real issue is whether universities can define and assess the human capacities that still matter. Original judgment, ethical responsibility, methodological control, source criticism, oral defence and domain understanding are harder to automate than fluent paragraphs.

AI will keep improving. Deep research tools will read more, plan better, cite more cleanly, use private databases under institutional agreements and integrate with writing environments. The arms race between detection and generation will not restore the old essay. The better path is to stop treating the final essay as the whole proof.

A university that adapts will not ask, “Did AI touch this?” as its only question. It will ask: Was AI allowed? Was it disclosed? Did the student verify it? Did the student make the core judgments? Can the student explain the work? Does the assessment still measure the learning outcome? Those questions are less dramatic than accusations of cheating. They are more useful.

The deep research era does not end university writing. It ends the comfortable fiction that polished writing alone proves learning. The next phase of higher education will be judged by whether it can rebuild that proof in public, fair and intellectually serious ways.

Questions universities, students and employers now ask about AI-written assignments

Can AI deep research really prepare a university paper at PhD level?

AI deep research tools can produce work that resembles advanced academic writing, especially literature reviews, policy briefings and source-backed reports. That does not make the work PhD-worthy. A real doctorate requires original contribution, methodological control, defence before experts and sustained accountability.

What changed from early ChatGPT essays to deep research tools?

Early chatbot essays mainly generated prose. Deep research tools can plan a research task, browse sources, analyze files, compare claims and produce a source-linked report. The risk moved from ghostwriting alone to outsourcing parts of the research process.

Does using AI for brainstorming count as cheating?

It depends on the assignment rule. Brainstorming may be allowed in one course and banned in another if idea generation is part of the assessed skill. Students should follow the assignment-specific AI statement and disclose substantive use when required.

Are AI detectors reliable enough for academic misconduct cases?

They should not be treated as proof by themselves. Turnitin warns that its AI report is not a misconduct decision, and research has found detector bias against non-native English writers. Detectors may support a conversation, but fair cases need wider evidence.

What is the biggest risk for universities?

The biggest risk is not only cheating. It is that assignments may award credit for polished outputs without proving that students learned the underlying skills. That weakens trust in grades and degrees.

Should universities ban AI completely?

A total ban may be justified for specific tasks where unaided performance is being tested. As a general policy, it is hard to enforce and may harm students who use AI for legitimate support. A better model defines allowed, disclosed and forbidden uses by assignment.

What should students disclose when they use AI?

Students should disclose AI use that affects content, argument, structure, source selection, data analysis or wording beyond routine spelling and grammar. A useful disclosure names the tool, the stage of work, the purpose and how the student checked or changed the output.

Is AI use fair if some students can pay for better tools?

Paid access creates an equity problem. Universities can reduce it by providing approved tools, designing tasks tied to course-specific evidence, and teaching AI literacy to all students rather than leaving tool skill to private advantage.

Can AI help students learn rather than cheat?

Yes, when it works like a tutor: asking questions, giving feedback, testing understanding and prompting revision. It becomes harmful when it replaces the thinking, reading or writing that the assignment is meant to develop.

Why are literature reviews especially vulnerable?

Literature reviews ask for searching, summarizing, grouping sources and identifying gaps. Those are exactly the tasks deep research systems are built to perform. Literature review assignments now need source matrices, oral defence and evidence of student judgment.

What should lecturers assess instead of only the final essay?

They should assess process evidence: research logs, annotated bibliographies, draft changes, source verification, revision notes, oral explanation, data files and the student’s ability to defend choices.

Does citing AI solve the academic integrity problem?

No. Citation or disclosure is only one layer. The student must also have permission to use AI, verify its claims, avoid prohibited uses and retain ownership of the final intellectual work.

How does AI affect international students?

AI can support language development and comprehension, but international students may also face unfair suspicion from detectors that misclassify non-native English writing. Clear policy and fair evidence standards are needed.

What role should oral exams or vivas play?

Short oral checks are useful for high-stakes assignments because they reveal whether the student understands sources, methods and argument choices. They do not need to replace written work, but they add evidence.

Are AI-generated citations trustworthy?

Not automatically. Students should open every cited source, verify the claim, check context and confirm that the source supports the argument. A citation in an AI report is only a starting point.

What does the EU AI Act mean for universities?

For European institutions, the EU AI Act makes AI governance, literacy and risk management part of compliance. Universities need to distinguish low-risk learning support from systems used in high-stakes decisions such as grading or misconduct processes.

How should universities handle privacy when students use AI?

They should tell students which tools are approved, what data may be uploaded, and which materials are forbidden, such as personal data, confidential placement data, unpublished research or sensitive case information.

Will AI make essays obsolete?

No. Essays still test argument, evidence and disciplinary writing when designed well. Generic essays are weaker now. Strong essays will include process evidence, source defence and course-specific reasoning.

What should employers expect from graduates in the AI era?

Employers should expect graduates to use AI responsibly, verify outputs, protect data, explain their reasoning and work without AI when needed. Universities should certify those abilities through assessment, not assume them.

What is the best single change a university can make now?

Every assignment brief should include a clear AI-use statement. It should say what is allowed, what must be disclosed, what is forbidden and how the student may be asked to prove their understanding.

Author:
Jan Bielik
CEO & Founder of Webiano Digital & Marketing Agency

AI can now draft PhD-level work, but learning is still the missing proof
AI can now draft PhD-level work, but learning is still the missing proof

This article is an original analysis supported by the sources cited below

Introducing deep research
OpenAI’s launch and update page for deep research in ChatGPT, used for product capabilities, benchmark context and access changes.

Deep research system card
OpenAI’s technical and safety document for deep research, used for details on browsing, file analysis, Python use, benchmark evaluation and risk classification.

Google Gemini Deep Research
Google’s launch article for Gemini Deep Research, used for the plan-browse-report workflow and export features.

Deep Research updates in Gemini
Google Workspace Updates post on Deep Research improvements, used for current Gemini education and workspace research capabilities.

Introducing ChatGPT Edu
OpenAI’s announcement of ChatGPT Edu, used for institutional deployment context in higher education.

ChatGPT study mode
OpenAI’s article on Study Mode, used for tutoring-oriented AI design, scaffolding features and stated limitations.

Introducing Claude for Education
Anthropic’s education announcement, used for learning mode, campus partnerships and institutional AI adoption.

Anthropic Education Report
Anthropic’s analysis of university student Claude use, used for evidence on educational content creation, essay support and technical assignment support.

Student generative AI survey 2025
HEPI’s 2025 student survey, used for data on the rapid growth of generative AI use in university assessments.

Guidance for generative AI in education and research
UNESCO’s global guidance, used for policy context, regulatory lag and institutional readiness.

AI competency framework for students
UNESCO’s student AI competency framework, used for AI literacy, ethics and student capability development.

AI competency framework for teachers
UNESCO’s teacher AI competency framework, used for the teacher-AI-student relationship and professional learning needs.

OECD Digital Education Outlook 2026
OECD’s 2026 education report, used for analysis of generative AI, cognitive offloading, teacher attitudes and learning risks.

Artificial intelligence and education and skills
OECD topic page, used for broader policy framing on AI, education, assessment and human-AI skill development.

2025 AI Index report
Stanford HAI’s 2025 AI Index, used for market and adoption context around generative AI investment and business use.

Education in the 2026 AI Index report
Stanford HAI’s education chapter, used for evidence on student AI use for schoolwork, research, editing and brainstorming.

AI detectors biased against non-native English writers
Stanford HAI article on detector bias, used for evidence on false positives and risks for non-native English writers.

Student perspectives on generative AI in higher education
Peer-reviewed study in the International Journal for Educational Integrity, used for student attitudes, confidence, policy demand and equity concerns.

Higher education institutional guidelines and policies on generative AI
Peer-reviewed study of AI policies at top U.S. universities, used for evidence on faculty guidelines, course-level policy and student guidance gaps.

Using the AI Writing Report
Turnitin guidance page, used for the limits of AI writing scores and reliability warnings.

Understanding false positives within our AI writing detection capabilities
Turnitin’s explanation of false positives, used for the distinction between detection signals and misconduct judgments.

Enacting assessment reform in a time of artificial intelligence
TEQSA resource on assessment reform, used for regulatory and sector guidance on AI-related learning assurance.

Gen AI, academic integrity and assessment reform
TEQSA knowledge hub page, used for institutional practice context around generative AI, academic integrity and assessment.

Artificial intelligence
Jisc’s AI resource page, used for UK tertiary education guidance, maturity tools and responsible AI adoption.

AI Act regulatory framework
European Commission page on the EU AI Act, used for risk-based regulation, education-related prohibitions, transparency duties and implementation timeline.

Copyright and artificial intelligence
U.S. Copyright Office AI page, used for copyrightability, registration and policy questions around AI-generated and AI-assisted works.

Is it OK for scientists to use AI to write papers
Nature news feature, used for broader research-writing context, disclosure debates and concerns about AI traces in scholarly papers.

Cognitive ease at a cost
Computers in Human Behavior study, used for evidence on ChatGPT use, lower cognitive load and weaker argumentation quality in student writing tasks.

Your brain on ChatGPT
Preprint on LLM-assisted essay writing and cognitive engagement, used cautiously as emerging evidence on ownership and cognitive effort.

How do students use ChatGPT as a writing support
Preprint study of student ChatGPT interactions during essay writing, used for context on usage patterns and writing support.

Artificial intelligence usage and perceptions in an elite college
Preprint on AI adoption at a selective college, used for context on discipline, demographic and policy effects in student AI use.

Citing this article? Brief excerpts are welcome. Please credit Webiano.digital, name the author where stated, and include a link to https://webiano.digital and to this original article. Full or substantial republication requires prior written permission. Read our Copyright and Content Use Policy.

This article was prepared with the assistance of artificial intelligence tools. The content underwent expert human review, and Webiano Digital & Marketing Agency assumes editorial responsibility for its final version and publication.